Enrichment of abnormally methylated DNA
The method enhances the sensitivity and efficiency of detecting abnormal methylation patterns in cfDNA by degrading uninformative sequences and partitioning for cancer-specific markers, addressing inefficiencies in existing liquid biopsy methods.
Patent Information
- Application Number
- JP2024575741
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-06-30
- Filing Date
- 2023-06-29
- Publication Date
- 2025-07-17
AI Technical Summary
Existing methods for analyzing cell-free DNA (cfDNA) in liquid biopsies are inefficient, labor-intensive, and lack sensitivity for detecting abnormally modified DNA associated with cancer, particularly due to the low concentration and heterogeneity of cfDNA.
A method involving the degradation of uninformative DNA sequences using modification-independent sequence-specific nucleases, followed by partitioning and detection of abnormally modified DNA sequences, such as those with cytosine methylation, to enrich for cancer-specific markers.
This approach enhances the sensitivity and efficiency of detecting abnormal methylation patterns in cfDNA, providing a faster and more cost-effective method for early cancer detection without requiring prior knowledge of modified sequences.
Smart Images

Figure 2025522763000001_ABST
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims the benefit of U.S. Provisional Application No. 63 / 367,462, filed Jun. 30, 2022, which is incorporated herein by reference for all purposes.
[0002] Field of the Invention The present disclosure provides methods and compositions related to the analysis using sequence - specific cleavage of abnormally modified DNA in a sample, e.g., abnormally methylated cell - free DNA (cfDNA). In some embodiments, the sample comprises cfDNA from a subject having or suspected of having cancer and / or the sample comprises DNA from cancer cells. In some embodiments, the abnormally modified cfDNA in the sample is enriched by cleaving DNA sequences containing modifications commonly found in cfDNA from healthy subjects, and the remaining DNA containing the modifications is detected.
Background Art
[0003] Introduction and Summary Cancer is the cause of millions of deaths worldwide each year. Early detection of cancer can lead to improved outcomes since early - stage cancer tends to be more responsive to treatment.
[0004] Inappropriately controlled cell growth is a characteristic of cancer and generally results from the accumulation of genetic and epigenetic changes, such as copy number variations (CNVs), single nucleotide variations (SNVs), gene fusions, insertions and / or deletions (indels), cytosine modifications (e.g., 5-methylcytosine, 5-hydroxymethylcytosine, and other more oxidized forms), and epigenetic diversity including the association of DNA with chromatin proteins and transcription factors. Thus, cancer can be manifested by non-sequence modifications such as methylation. Examples of methylation changes in cancer include the local gain of DNA methylation in the CpG islands of the TSS of genes involved in normal growth control, DNA repair, cell cycle regulation, and / or cell differentiation. Hypermethylation can be associated with abnormal loss of transcriptional ability of the genes involved and occurs at least as frequently as point mutations and deletions as a cause of altered gene expression. Furthermore, without being bound by any particular theory, cells within or surrounding a cancer or neoplasm can excrete more DNA than cells of the same tissue type in a healthy subject. DNA derived from such cells can be epigenetically different from the DNA excreted in a healthy subject. Thus, during carcinogenesis, the distribution of epigenetically modified (e.g., methylated) DNA, such as cell-free DNA (cfDNA), in a particular DNA sample may change. Thus, sufficient sensitive epigenetic (e.g., DNA methylation) profiling can be used to detect abnormal methylation of DNA in a sample.
[0005] Biopsy represents a conventional approach for detecting or diagnosing cancer, which involves extracting cells or tissues from a potential cancer site and analyzing the associated phenotypic and / or genotypic features. Biopsy has the drawback of being invasive.
[0006] Cancer detection based on the analysis of body fluids such as blood (“liquid biopsy”) is an interesting alternative based on the observation that DNA is released into the body fluid from cancer cells. Liquid biopsy is non-invasive (sometimes only blood collection is required). However, considering the low concentration and heterogeneity of cell-free DNA, it has been difficult to develop an accurate and sensitive method for analyzing liquid biopsy materials that provides detailed information on nucleobase modifications. The contribution of DNA from cells within or surrounding a cancer or neoplasm to a sample may be relatively small compared to the contribution from other cells, and the DNA contributed by other cells may be uninformative regarding the cancer state. Isolating and processing the fraction of cell-free DNA that is useful for further analysis in liquid biopsy procedures is an important part of these methods. Thus, improved methods and compositions for analyzing cell-free DNA are still needed, for example, in liquid biopsy. SUMMARY OF THE INVENTION MEANS FOR SOLVING THE PROBLEMS
[0007] The methods according to the present disclosure can include, for example, the degradation of DNA sequences that are modified and / or uninformative regarding the cancer state in a healthy subject. The degradation can replace more labor-intensive steps in existing DNA profiling methods. Thus, the methods herein can provide a method that is faster, more efficient, more cost-effective, and / or more sensitive for detecting abnormally modified DNA. Further, no prior knowledge or identification of abnormally modified sequences is required to practice the methods herein.
[0008] In some embodiments, DNA methylation involves the addition of a methyl group to a cytosine residue in a CpG dinucleotide (cytosine-phosphate-guanine dinucleotide, i.e., cytosine followed by guanine in the 5’→3’ direction of a nucleic acid sequence). In some embodiments, DNA methylation involves the addition of a methyl group to an adenine residue, for example, in N6-methyladenine. In some embodiments, DNA methylation is 5-methylation (modification of the fifth carbon of the six-carbon ring of cytosine). In some embodiments, 5-methylation involves the addition of a methyl group to the 5C position of a cytosine residue to produce 5-methylcytosine (m5c or 5-mC or 5mC). In some embodiments, methylation includes derivatives of m5c. Examples of derivatives of m5c include, but are not limited to, 5-hydroxymethylcytosine (5-hmC or 5hmC), 5-formylcytosine (5-fC), and 5-carboxylcytosine (5-caC). In some embodiments, DNA methylation is 3C-methylation (modification of the third carbon of the six-carbon ring of a cytosine residue). In some embodiments, 3C-methylation involves the addition of a methyl group to the 3C position of a cytosine residue to produce 3-methylcytosine (3mC). Methylation can occur at sites other than CpG, for example, methylation can occur at CpA, CpT, or CpC sites. DNA methylation can alter the activity of the methylated DNA region. For example, when DNA within a promoter region is methylated, transcription of the gene can be repressed. DNA methylation is extremely important for normal development, and abnormalities in methylation can interfere with epigenetic regulation. Interference with epigenetic regulation, such as repression, can lead to diseases such as cancer. Promoter methylation in DNA can be indicative of cancer.
[0009] The present disclosure is directed to meeting the need for improved analysis of cfDNA and / or providing other benefits. Accordingly, the following exemplary embodiments are provided.
[0010] Embodiment 1 is as follows: A method for analyzing DNA in a sample, comprising: a) distributing the sample into a plurality of secondary samples by contacting the DNA with an agent that recognizes a modification associated with the DNA, wherein the plurality of secondary samples includes a first secondary sample and a second secondary sample, and the first secondary sample includes a higher proportion of DNA with an associated modification than the second secondary sample; b) sequence-specifically degrading a plurality of DNA sequences in the first secondary sample that contain a modification and are commonly found in cell-free DNA (cfDNA) derived from a healthy subject, comprising contacting the first secondary sample with a modification-independent sequence-specific nuclease to thereby produce a processed sample; c) detecting the presence or absence of one or more DNA sequences with an associated modification in the processed sample and comprising a method.
[0011] Embodiment 2 is as follows: The method according to Embodiment 1, wherein the distributing step is performed before the degrading step.
[0012] Embodiment 3 is as follows: A method for analyzing cfDNA in a sample, comprising: a) contacting the sample or a secondary sample thereof with an MSRE to thereby degrade DNA containing an unmethylated recognition site of the MSRE; b) sequence-specifically degrading a plurality of DNA sequences having a methylated sequence commonly found in cfDNA derived from a healthy subject and a plurality of sequences lacking a CpG motif, comprising contacting the sample with a modification-independent sequence-specific nuclease to thereby produce a processed sample; c) detecting the presence or absence of one or more DNA sequences in the processed sample and comprising a method.
[0013] Embodiment 4 is as follows: The method according to the immediately preceding embodiment, wherein the methylation includes cytosine methylation.
[0014] Embodiment 5 is as follows: A step of distributing a sample into a plurality of secondary samples by contacting DNA with an agent that recognizes a modification associated with the DNA, wherein the plurality of secondary samples includes a first secondary sample and a second secondary sample, and the first secondary sample includes DNA with an associated modification at a higher rate than the second secondary sample, the method according to embodiment 3 or embodiment 4 further comprising the step.
[0015] Embodiment 6 is as follows: The method according to any one of the embodiments, wherein the distributing step includes distributing based on the methylation level of the DNA.
[0016] Embodiment 7 is as follows: The method according to any one of the embodiments, wherein the modification is methylation.
[0017] Embodiment 8 is as follows: The method according to the immediately preceding embodiment, wherein the methylation includes cytosine methylation.
[0018] Embodiment 9 is as follows: The method according to any one of embodiments 1, 2, or 5 to 8, wherein the distributing step includes distributing based on the hydroxymethylation level of the DNA.
[0019] Embodiment 10 is as follows: The method according to the immediately preceding embodiment, wherein hydroxymethyl is labeled before the distributing step, and optionally, the label includes biotin, glucosyl, or sulfonyl.
[0020] Embodiment 11 is as follows: The method according to any one of embodiments 1 to 3 or 7 to 10, wherein the modification is hydroxymethylation.
[0021] Embodiment 12 is as follows: The method according to any one of the embodiments, wherein the agent that recognizes the modification associated with DNA is methyl, hydroxymethyl, or a labeled hydroxymethyl-binding reagent.
[0022] Embodiment 13 is as follows: The method according to the immediately preceding embodiment, wherein the methyl, hydroxymethyl, or labeled hydroxymethyl-binding reagent is an antibody.
[0023] Embodiment 14 is as follows: The method according to embodiment 12 or 13, wherein the methyl-binding reagent specifically recognizes 5-methylcytosine, 5-hydroxymethylcytosine, biotinylated 5-hydroxymethylcytosine, glucosylated 5-hydroxymethylcytosine, or sulfonylated 5-hydroxymethylcytosine.
[0024] Embodiment 15 is as follows: The method according to any one of embodiments 12 to 14, wherein the methyl-binding reagent is immobilized on a solid support.
[0025] Embodiment 16 is as follows: The method according to any one of the embodiments, wherein the step of partitioning comprises immunoprecipitation of methylated, hydroxymethylated, or labeled hydroxymethylated DNA.
[0026] Embodiment 17 is as follows: The method according to any one of embodiments 1, 2, or 5 to 16, wherein the step of partitioning comprises partitioning based on binding to a protein, and optionally, the protein is a methylated protein, an acetylated protein, a non-methylated protein, a non-acetylated protein, and / or, optionally, the protein is a histone.
[0027] Embodiment 18 is as follows: The method according to the immediately preceding embodiment, wherein the step of partitioning comprises contacting the collected cfDNA with a binding reagent that is specific for a protein and immobilized on a solid support.
[0028] Embodiment 19 is as follows: The method according to any one of the embodiments, wherein the step of digesting comprises contacting a second secondary sample with a modification-independent sequence-specific nuclease, thereby producing a second processed sample.
[0029] Embodiment 20 is as follows: The method according to any one of Embodiments 1, 2, or 5 to 19, comprising contacting one or more of the plurality of secondary samples with a methylation-sensitive restriction enzyme (MSRE), thereby digesting DNA containing an unmethylated recognition site of the MSRE.
[0030] Embodiment 21 is as follows: The method according to the immediately preceding embodiment, wherein the step of contacting one or more of the plurality of secondary samples with MSRE is performed after the step of dispensing.
[0031] Embodiment 22 is as follows: The method according to Embodiment 20 or 21, wherein the step of contacting one or more of the plurality of secondary samples with MSRE is performed before the step of digesting.
[0032] Embodiment 23 is as follows: The method according to Embodiment 20 or 21, wherein the step of contacting one or more of the plurality of secondary samples with MSRE is performed after the step of digesting.
[0033] Embodiment 24 is as follows: The method according to Embodiment 20 or 21, wherein the step of contacting one or more of the plurality of secondary samples with MSRE is performed simultaneously with the step of digesting.
[0034] Embodiment 25 is as follows: The method according to any one of Embodiments 21 to 24, wherein the first secondary sample is contacted with MSRE.
[0035] Embodiment 26 is as follows: The method according to any one of Embodiments 20 to 25, wherein the step of contacting the sample or the secondary sample with the MSRE is performed before the step of sequence-specific cleavage.
[0036] Embodiment 27 is as follows: The method according to any one of the embodiments, wherein the modification-independent sequence-specific nuclease is a CRISPR nuclease.
[0037] Embodiment 28 is as follows: The method according to the immediately preceding embodiment, wherein the CRISPR nuclease is a Cas12a, Cas12b, or CasX nuclease.
[0038] Embodiment 29 is as follows: The method according to Embodiment 27, wherein the CRISPR nuclease is a Cas9 nuclease.
[0039] Embodiment 30 is as follows: The method according to the immediately preceding embodiment, wherein the Cas9 nuclease is a multi-turnover Cas9 nuclease.
[0040] Embodiment 31 is as follows: The method according to Embodiment 29, wherein the Cas9 nuclease is a Streptococcus pyogenes Cas9 nuclease or a variant thereof.
[0041] Embodiment 32 is as follows: The method according to Embodiment 29, wherein the Cas9 nuclease is a Staphylococcus aureus Cas9 nuclease or a variant thereof.
[0042] Embodiment 33 is as follows: The method according to any one of Embodiments 29 to 32, wherein the Cas9 nuclease is a high-fidelity variant.
[0043] Embodiment 34 is as follows: The method according to any one of the embodiments, wherein the step of specifically degrading the array comprises contacting the DNA with a plurality of guide RNAs.
[0044] Embodiment 35 is as follows: The method according to the immediately preceding embodiment, wherein at least one guide RNA comprises one or more modifications.
[0045] Embodiment 36 is as follows: The method according to the immediately preceding embodiment, wherein the one or more modifications comprise phosphorothioate internucleoside linkages, 2'-substitutions, or UNA, LNA, cEt, or ENA nucleotide sugars.
[0046] Embodiment 37 is as follows: The method according to the immediately preceding embodiment, wherein the 2'-substitution is 2'-fluoro, 2'-hydro, 2'-O-methoxyethyl, or 2'-O-alkyl.
[0047] Embodiment 38 is as follows: The method according to any one of Embodiments 34 to 37, wherein at least one guide RNA is a sgRNA.
[0048] Embodiment 39 is as follows: The method according to any one of Embodiments 34 to 38, wherein at least one guide RNA specifically binds to DNA comprising a CpG motif that is methylated in cfDNA derived from healthy tissue or a healthy subject.
[0049] Embodiment 40 is as follows: The method according to any one of Embodiments 34 to 39, wherein at least one guide RNA specifically binds to a DNA sequence lacking CpG dinucleotides.
[0050] Embodiment 41 is as follows: The method according to any one of Embodiments 34 to 39, wherein each guide RNA of the plurality of guide RNAs comprises a modification and is configured to specifically bind to a DNA sequence commonly found in cell-free DNA (cfDNA) derived from a healthy subject.
[0051] Embodiment 42 is as follows: The method according to any one of Embodiments 34 to 40, wherein each guide RNA of the plurality of guide RNAs comprises a modification and is configured to specifically bind to a DNA sequence commonly found in cell-free DNA (cfDNA) derived from a healthy subject or a DNA sequence lacking CpG dinucleotides.
[0052] Embodiment 43 is as follows: The method according to any one of Embodiments 1 to 26, 34 to 37, or 39 to 42, wherein the modification-independent sequence-specific nuclease is an Argonaute nuclease.
[0053] Embodiment 44 is as follows: The method according to any one of Embodiments 1 to 26, 34 to 37, or 39 to 42, wherein the modification-independent sequence-specific nuclease is a zinc finger nuclease.
[0054] Embodiment 45 is as follows: The method according to any one of Embodiments 1 to 26, 34 to 37, or 39 to 42, wherein the modification-independent sequence-specific nuclease is a TALEN.
[0055] Embodiment 46 is as follows: The method according to any one of the embodiments, wherein the step of detecting comprises sequencing.
[0056] Embodiment 47 is as follows: The method according to the immediately preceding embodiment, wherein the step of detecting comprises sequencing a plurality of target regions within at least one set of target regions.
[0057] Embodiment 48 is as follows: The method according to any one of the embodiments, further comprising enriching one or more of a plurality of target regions within at least one set of target regions.
[0058] Embodiment 49 is as follows: The method according to the immediately preceding embodiment, wherein the enriching step comprises contacting the DNA with a target-specific probe specific for one or more of a plurality of target regions within at least one set of target regions.
[0059] Embodiment 50 is as follows: The method according to any one of embodiments 47 to 49, wherein at least one set of target regions comprises target regions that are not generally found in methylated form in cfDNA from healthy subjects or not generally found in methylated form in healthy tissues.
[0060] Embodiment 51 is as follows: The method according to embodiments 47 to 50, wherein at least one set of target regions comprises target regions that are generally found in methylated form in a tissue and do not substantially contribute to cfDNA in healthy subjects.
[0061] Embodiment 52 is as follows: The method according to any one of embodiments 47 to 51, wherein at least one set of target regions comprises target regions that are generally found in methylated form in cancerous tissue.
[0062] Embodiment 53 is as follows: The method according to any one of embodiments 47 to 52, wherein at least one set of target regions comprises a set of sequence-variable target regions and a set of epigenetic target regions.
[0063] Embodiment 54 is as follows: The method according to any one of embodiments 47 to 53, wherein at least one set of target regions comprises a set of hypermethylated variable target regions.
[0064] Embodiment 55 is as follows: The method according to the immediately preceding embodiment, wherein the hypermethylated variable target region set includes regions where the degree of methylation in at least one tissue type is higher than the degree of methylation in cfDNA from a healthy subject.
[0065] Embodiment 56 is as follows: The method according to any one of Embodiments 47 to 55, wherein at least one target region set includes a hypomethylated variable target region set.
[0066] Embodiment 57 is as follows: The method according to the immediately preceding embodiment, wherein the hypomethylated variable target region set includes regions where the degree of methylation in at least one tissue type is lower than the degree of methylation in cfDNA from a healthy subject.
[0067] Embodiment 58 is as follows: The method according to any one of Embodiments 47 to 57, wherein at least one target region set includes a methylation control target region set.
[0068] Embodiment 59 is as follows: The method according to any one of Embodiments 47 to 57, wherein at least one target region set includes a fragmentation variable target region set.
[0069] Embodiment 60 is as follows: The method according to the immediately preceding embodiment, wherein the fragmentation variable target region set includes a transcription start site region.
[0070] Embodiment 61 is as follows: The method according to Embodiment 59 or 60, wherein the fragmentation variable target region set includes a CTCF binding region.
[0071] Embodiment 62 is as follows: The method according to any one of embodiments 53 to 61, wherein the array-variable target region set comprises at least one sequence not commonly found in cfDNA from a healthy subject.
[0072] Embodiment 63 is as follows: The method according to any one of embodiments 46 to 62, wherein sequencing comprises sequencing a gene or a part thereof selected from Table 1, Table 2, Table 3, Table 4, and / or Table 5.
[0073] Embodiment 64 is as follows: The method according to any one of embodiments 46 to 63, wherein sequencing comprises sequencing all of the DNA sequences in the processed sample.
[0074] Embodiment 65 is as follows: The method according to any one of the embodiments, wherein 20 to 250,000 sequences are resolved.
[0075] Embodiment 66 is as follows: The method according to any one of the embodiments, wherein 50 to 100,000 sequences are resolved.
[0076] Embodiment 67 is as follows: The method according to the immediately preceding embodiment, wherein 100 to 10,000 sequences are resolved.
[0077] Embodiment 68 is as follows: The method according to any one of the embodiments, wherein the sequences to be resolved comprise repetitive elements, and optionally, the repetitive elements comprise SINE, LINE, and / or Alu elements.
[0078] Embodiment 69 is as follows: The method according to any one of embodiments 1 to 45, wherein the detecting step comprises performing qPCR.
[0079] Embodiment 70 is as follows: The method according to any one of the embodiments, wherein DNA is collected from a test subject.
[0080] Embodiment 71 is as follows: The method according to any one of the embodiments, wherein DNA comprises cfDNA obtained from a test subject.
[0081] Embodiment 72 is as follows: The method according to embodiment 70 or 71, wherein DNA comprises DNA obtained from a tissue sample of a test subject.
[0082] Embodiment 73 is as follows: The method according to the immediately preceding embodiment, wherein the tissue sample is a biopsy material, a fine needle aspirate, or a formalin-fixed paraffin-embedded tissue sample.
[0083] Embodiment 74 is as follows: The method according to any one of the embodiments, further comprising ligating a barcode-containing adapter to the DNA, optionally before or simultaneously with the amplification of the DNA.
[0084] Embodiment 75 is as follows: The method according to the immediately preceding embodiment, wherein the plurality of sequences specifically cleaved comprises sequences comprising a barcode-containing adapter dimer junction.
[0085] Embodiment 76 is as follows: The method according to the immediately preceding embodiment, comprising contacting the DNA with a plurality of guide RNAs configured to specifically bind to each of the possible adapter dimer junctions.
[0086] Embodiment 77 is as follows: The method according to any one of the embodiments, wherein the DNA is amplified before the detecting step.
[0087] Embodiment 78 is as follows: The step of array-specific decomposition is after the step of ligating a barcode-containing adapter to DNA and (a) before the step of partitioning the sample into a plurality of secondary samples by contacting the DNA with an agent that recognizes a modification associated with the DNA, (b) after the step of partitioning the sample into a plurality of secondary samples by contacting the DNA with an agent that recognizes a modification associated with the DNA, (c) before the step of contacting the sample or its secondary sample with an MSRE, (d) after the step of contacting the sample or its secondary sample with an MSRE, (e) before the step of amplifying the DNA before the step of detecting, (f) after the step of amplifying the DNA before the step of detecting, or (g) any one or two of (a) and any one or two of (c)-(f), any one or two of (b) and any one or two of (c)-(f), any one or two of (c) and any one or two of (a), (b), (e), and (f), any one or two of (d) and any one or two of (a), (b), (e), and (f), any one or two of (e) and any one or two of (a)-(d); or any one or two of (f) and any one or two of (a)-(d), in any combination, the method according to any one of the embodiments.
[0088] Embodiment 79 is as follows: The method according to any one of the embodiments, comprising the step of partitioning the sample, wherein the DNA molecules from the first secondary sample and the DNA molecules from the second secondary sample are differentially tagged.
[0089] Embodiment 80 is as follows: The method according to the immediately preceding embodiment, wherein the DNA molecules from the first secondary sample and the DNA molecules from the second secondary sample are sequenced within the same sequencing cell.
[0090] Embodiment 81 is as follows: The method according to embodiment 79 or 80, comprising the step of differentially tagging and pooling a first secondary sample and a second secondary sample.
[0091] Embodiment 82 is as follows: The method according to any one of embodiments 79-81, wherein the DNA of the first secondary sample and the DNA of the second secondary sample are differentially tagged, and after the differential tagging, a part of the DNA derived from the second secondary sample is added to the first secondary sample or at least a part thereof, thereby forming a pool.
[0092] Embodiment 83 is as follows: The method according to the immediately preceding embodiment, wherein the pool comprises about 45% or less, about 40% or less, about 35% or less, about 30% or less, about 25% or less, about 20% or less, about 15% or less, about 10% or less, or about 5% or less of the DNA of the second secondary sample.
[0093] Embodiment 84 is as follows: The method according to the immediately preceding embodiment, wherein the pool comprises about 70-90%, about 75-85%, or about 80% of the DNA of the second secondary sample.
[0094] Embodiment 85 is as follows: The method according to any one of embodiments 82-84, wherein the pool comprises substantially all of the DNA of the first secondary sample.
[0095] Embodiment 86 is as follows: The method according to any one of the embodiments, comprising the step of distributing a sample into a plurality of secondary samples, wherein the plurality of secondary samples includes a third secondary sample that contains DNA having cytosine modification at a higher proportion than the second secondary sample but at a lower proportion than the first secondary sample.
[0096] Embodiment 87 is as follows: The method according to the immediately preceding embodiment, further comprising the step of differentially tagging a third secondary sample.
[0097] Embodiment 88 is as follows: Before the step of sequence-specifically degrading a plurality of DNA sequences in a first secondary sample that contains modifications and is commonly found in cell-free DNA, the first secondary sample is enriched for one or more sets of target regions, the method according to any one of the embodiments.
[0098] Embodiment 89 is as follows: Before the step of sequence-specifically degrading a plurality of DNA sequences that contain modifications and are commonly found in cell-free DNA, a plurality of first secondary samples are pooled, and optionally, the plurality of first secondary samples are from different subjects and / or are differentially tagged with sample tags, the method according to any one of the embodiments.
[0099] Embodiment 90 is as follows: The method according to any one of the embodiments, further comprising the step of determining the likelihood that the subject has cancer.
[0100] Embodiment 91 is as follows: By sequencing, a plurality of sequencing reads are generated, the method comprising mapping the plurality of sequence reads to one or more reference sequences to generate mapped sequence reads, and processing the mapped sequence reads corresponding to the set of sequence variable target regions and the set of epigenetic target regions to determine the likelihood that the subject has cancer, the method according to any one of embodiments 45 to 90.
[0101] Embodiment 92 is as follows: The method according to any one of embodiments 70 to 91, wherein the test subject has been previously diagnosed with cancer and has received one or more previous cancer treatments.
[0102] Embodiment 93 is as follows: The method according to the immediately preceding embodiment, wherein cfDNA is obtained and detected at one or more preselected time points after one or more previous cancer treatments, and the detecting step comprises sequencing a DNA sequence, thereby providing a set of sequence information.
[0103] Embodiment 94 is as follows: The method according to the immediately preceding embodiment, further comprising detecting the presence or absence of DNA originating from or derived from tumor cells using the set of sequence information at a preselected time point.
[0104] Embodiment 95 is as follows: The method according to the immediately preceding embodiment, further comprising determining a cancer recurrence score indicative of the presence or absence of DNA originating from or derived from tumor cells for a test subject, and optionally further comprising determining a cancer recurrence status based on the cancer recurrence score, wherein when the cancer recurrence score is determined to be at or above a predetermined threshold, the cancer recurrence status of the test subject is determined to be at risk of cancer recurrence, or when the cancer recurrence score is below a predetermined threshold, the cancer recurrence status of the test subject is determined to be at low risk of cancer recurrence.
[0105] Embodiment 96 is as follows: The method according to the immediately preceding embodiment, further comprising comparing the cancer recurrence score of the test subject with a predetermined cancer recurrence threshold, wherein if the test subject has a cancer recurrence score above the cancer recurrence threshold, the test subject is classified as a candidate for subsequent cancer treatment, or if the cancer recurrence score is below the cancer recurrence threshold, the test subject is classified as not a candidate for subsequent cancer treatment. BRIEF DESCRIPTION OF THE DRAWINGS
[0106]
Fig. 1-1
Fig. 1-2
[0107]
Fig. 2
[0108]
Fig. 3
DETAILED DESCRIPTION OF THE INVENTION
[0109] DETAILED DESCRIPTION OF CERTAIN EMBODIMENTS Reference will now be made in detail to certain embodiments of the invention. It will be understood that the invention is described in connection with such embodiments, but the invention is not intended to be limited to those embodiments. On the contrary, the invention is intended to cover all alternatives, modifications, and equivalents that may be included within the scope of the invention as defined by the appended claims.
[0110] Prior to elaborating on the present teachings, it should be understood that the present disclosure is not limited to specific compositions or process steps and can therefore vary. As used in this specification and the appended claims, it should be noted that the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to "a nucleic acid" includes a plurality of nucleic acids, reference to "a cell" includes a plurality of cells, and so on.
[0111] A numerical range includes the numbers defining the range. It is understood that measured values and measurable values are approximate, taking into account significant figures and errors associated with the measurement. Also, the use of "comprise", "comprises", "comprising", "contain", "contains", "containing", "include", "includes", and "including" is not intended to be limiting. It should be understood that both the foregoing general description and the detailed description are exemplary and for illustrative purposes only, and do not limit the present teachings.
[0112] Unless otherwise specified in the above specification, embodiments in this specification listing various components as "comprising" also contemplate them as "consisting of" or "consisting essentially of" the listed components, and embodiments in this specification listing various components as "consisting of" also contemplate them as "comprising" or "consisting essentially of" the listed components, and embodiments in this specification listing various components as "consisting essentially of" also contemplate them as "consisting of" or "comprising" the listed components (this interchangeability does not apply to the use of these terms in the claims).
[0113] The section headings used in this specification are for the purpose of organization and should not be construed as limiting the disclosed subject matter in any way. Where any document or other material incorporated by reference conflicts with any explicit content of this specification, including definitions, this specification shall prevail. I. Definitions
[0114] “Cell-free DNA,” “cfDNA molecule,” or simply “cfDNA” that exists naturally in extracellular form in a subject (e.g., in blood, serum, plasma, or other body fluids such as lymph, cerebrospinal fluid, urine, or sputum) contains DNA molecules. CfDNA previously existed within large, complex biological organisms, e.g., mammalian cells (s) but has been released from the cells (s) into the fluid found in the organism and can be obtained from a sample of the fluid without performing an in vitro cell lysis step. CfDNA molecules can exist as DNA fragments.
[0115] As used herein, a “CpG motif” is a continuous DNA sequence that includes at least one CpG dinucleotide. In some embodiments, a CpG motif includes one CpG dinucleotide. In some embodiments, a CpG motif includes two CpG dinucleotides. In some embodiments, a CpG motif includes a sequence that is specifically recognized by a sequence-specific nuclease. In some embodiments, a CpG motif includes a recognition sequence for a methylation-sensitive nuclease or a methylation-dependent nuclease. The terms “CpG dinucleotide” and “CpG site” are used interchangeably herein.
[0116] As used herein, “prevalent” in the context of a DNA sequence (e.g., cfDNA) in a sample means present at a detectable level. In some embodiments, a sequence that is prevalent in a sample is detectable in a majority or all samples of a particular sample type. For example, a methylated sequence that is prevalent in cfDNA from a healthy subject is a methylated sequence that is detectable by standard techniques in the art in samples from healthy subjects. In some such embodiments, therefore, a methylated sequence that is prevalent in cfDNA from a healthy subject is not associated with a disease or the likelihood or risk of developing a disease.
[0117] As used herein, when the modification is a covalent modification of DNA or is a covalent modification of a protein (e.g., histone) bound to DNA, the modification is "associated with" the DNA.
[0118] As used herein, "partitioning" of a nucleic acid such as a DNA molecule means separating, fractionating, or sorting a sample or population of the nucleic acid into a plurality of secondary samples or subpopulations of the nucleic acid, based on one or more modifications or characteristics that differ in proportion in each of the plurality of secondary samples or subpopulations. Partitioning can include physically partitioning nucleic acid molecules based on the presence or absence of one or more methylated nucleobases. A sample or population can be partitioned into one or more partitioned secondary samples or subpopulations based on characteristics indicative of a genetic or epigenetic change or condition.
[0119] As used herein, when a fraction of nucleotides having a modification or other characteristic is greater in a first sample or population than in a second population, the modification or other characteristic is present at a "higher proportion" in the first sample or population of the nucleic acid than in the second sample or population. For example, if one tenth of the nucleotides in a first sample are mC and one twentieth of the nucleotides in a second sample are mC, the first sample contains the 5-methylated cytosine modification at a higher proportion than the second sample.
[0120] As used herein, "substantially without alteration of base pairing specificity" for a given nucleobase means that the majority of molecules containing that nucleobase that can be sequenced do not have an alteration in the base pairing specificity of that nucleobase as compared to its base pairing specificity in the originally isolated sample. In some embodiments, 75%, 90%, 95%, or 99% of the molecules containing that nucleobase that can be sequenced do not have an alteration in base pairing specificity as compared to its base pairing specificity in the originally isolated sample. As used herein, "altered base pairing specificity" for a given nucleobase means that the majority of molecules containing that nucleobase that can be sequenced have a base pairing specificity in that nucleobase as compared to its base pairing specificity in the originally isolated sample.
[0121] As used herein, "base pairing specificity" refers to the standard DNA base (A, C, G, or T) to which a given base most preferentially pairs. For example, unmodified cytosine and 5-methylcytosine have the same base pairing specificity (i.e., specificity for G), while uracil and cytosine have different base pairing specificities in that uracil has a base pairing specificity for A whereas cytosine has a base pairing specificity for G. The ability of uracil to form wobble pairs with G is not important as uracil most preferentially pairs with A among the four standard DNA bases.
[0122] A "converted nucleobase" is a nucleobase having an altered base pairing specificity in which the original base pairing specificity of the nucleobase has been changed by a procedure. For example, by a particular procedure, non-methylated or unmodified cytosine is converted to dihydrouracil, or more generally, at least one modified or unmodified form of cytosine undergoes deamination to yield uracil (which is considered a modified nucleobase in the context of DNA) or a further modified form of uracil. As used herein, a "converted sample" is a sample containing DNA that includes at least one converted nucleobase.
[0123] As used herein, a "combination" that includes a plurality of members refers to either a single composition that includes the members, or a set of compositions that are in proximity, e.g., in separate containers, or in compartments within a larger container such as a multi-well plate, test tube rack, refrigerator, freezer, incubator, water bath, ice bucket, machine, or other storage format. Thus, one combination, a plurality of combinations, or combinations thereof refer to any permutation and combination of the terms listed before the term "combination". For example, "A, B, C, or combinations thereof" is intended to include at least one of A, B, C, AB, AC, BC, or ABC, and, if order is important in a particular context, also BA, CA, CB, ACB, CBA, BCA, BAC, or CAB. Continuing with this example, combinations that contain repeats of one or more items or terms, e.g., BB, AAA, AAB, BBC, AAABCCCC, CBBAAA, CABABB, etc. are clearly included. In general, and unless it is clear from the context otherwise, those skilled in the art will understand that there is no limit to the number of items or terms in any combination.
[0124] "Specifically binds" means that, in the context of an oligonucleotide such as a guide RNA and a nucleic acid containing a sequence that is partially or completely complementary to the oligonucleotide, under appropriate hybridization conditions, the oligonucleotide hybridizes to the complementary sequence to form a stable oligonucleotide:complementary sequence hybrid, while at the same time minimizing the formation of stable oligonucleotide:non-complementary sequence hybrids. Thus, the oligonucleotide hybridizes to the complementary sequence to a sufficiently greater extent than to non-complementary sequences. Appropriate hybridization conditions are well known in the art, can be predicted based on sequence composition, or can be determined by using routine testing methods (see, e.g., Sambrook et al., Molecular Cloning, A Laboratory Manual, 2nd ed. (Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1989) §§ 1.90-1.91, 7.37-7.57, 9.47-9.51 and 11.47-11.57, particularly §§ 9.50-9.51, 11.12-11.13, 11.45-11.47 and 11.55-11.57, which are incorporated herein by reference).
[0125] "Target region set" refers to a plurality of genomic loci that include regions sharing at least one common feature. In some embodiments, the target region set is identified by at least one common feature. For example, a hypermethylated variable target region set includes regions of hypermethylated DNA.
[0126] "Sequence variable target region set" refers to a target region set that includes target regions that can exhibit sequence changes, such as nucleotide substitutions (i.e., single nucleotide polymorphisms), insertions, deletions, or gene fusions or translocations, in abnormal cells such as nascent cells (e.g., tumor cells and cancer cells) compared to normal cells.
[0127] An "epigenetic target region set" refers to a set of target regions that can exhibit sequence-independent changes in abnormal cells such as newborn cells (e.g., tumor cells and cancer cells) compared to normal cells, or in cfDNA derived from a subject with cancer compared to cfDNA derived from a healthy subject. Examples of sequence-independent changes include, but are not limited to, changes in methylation (increase or decrease), changes in nucleosome distribution, changes in cfDNA fragmentation patterns, changes in the binding of CCCTC-binding factor ("CTCF"), changes in transcription start sites, and changes in regulatory protein-binding regions. Thus, the epigenetic target region set includes, but is not limited to, a hypermethylation variable target region set, a hypomethylation variable target region set, and a fragmentation variable target region set, such as CTCF-binding sites and transcription start sites. For the purposes of this disclosure, loci that are prone to local amplifications and / or gene fusions associated with neoplasia, tumors, or cancer can also be included in the epigenetic target region set, because, for example, the detection of local amplifications and / or gene fusions does not depend on the base call accuracy at one or a few individual positions, and thus they can be detected by relatively shallow depth sequencing, in that the detection of copy number changes by sequencing, or the detection of map fusion sequences located at more than one locus in the reference genome, tends to be similar to the detection of the exemplary epigenetic changes discussed above compared to the detection of nucleotide substitutions, insertions, or deletions.
[0128] If it is derived from tumor cells, the nucleic acid is "produced by the tumor" or "circulating tumor DNA" ("ctDNA"). Tumor cells are newborn cells that originate from a tumor and can either remain within the tumor or separate from the tumor (e.g., in the case of metastatic cancer cells and circulating tumor cells).
[0129] In the context of nucleic acid molecules, the term "methylation" refers to the addition of a methyl group to a nucleobase within a nucleic acid molecule. In some embodiments, methylation refers to the addition of a methyl group to cytosine at a CpG site (cytosine-phosphate-guanine site, i.e., cytosine followed by guanine in the 5’→3’ direction of a nucleic acid sequence). In some embodiments, DNA methylation refers to, for example, the addition of a methyl group to adenine in N 6 -methyladenine. In some embodiments, DNA methylation is 5-methylation (modification of the 5th carbon of the 6-carbon ring of cytosine). In some embodiments, 5-methylation refers to the addition of a methyl group to the 5C position of cytosine, which generates 5-methylcytosine (5mC). In some embodiments, methylation includes derivatives of 5mC. Examples of derivatives of 5mC include, but are not limited to, 5-hydroxymethylcytosine (5-hmC), 5-formylcytosine (5-fC), and 5-carboxylcytosine (5-caC). In some embodiments, DNA methylation is 3C methylation (modification of the 3rd carbon of the 6-carbon ring of cytosine). In some embodiments, 3C methylation includes the addition of a methyl group to the 3C position of cytosine, which generates 3-methylcytosine (3mC). Methylation can also occur at non-CpG sites. For example, methylation can occur at CpA, CpT, or CpC sites. DNA methylation can change the activity of the methylated DNA region. For example, when the DNA within a promoter region is methylated, gene transcription can be inhibited. DNA methylation is extremely important for normal development, and abnormal methylation may interfere with epigenetic regulation. Interference with epigenetic regulation, such as inhibition, may cause diseases such as cancer. Promoter methylation in DNA can indicate cancer.
[0130] The term "hypermethylation" refers to an increase in the level or degree of methylation of nucleic acid molecules (plural possible) within a population of nucleic acid molecules (e.g., a sample) compared to other nucleic acid molecules. In some embodiments, hypermethylated DNA may include DNA molecules that contain at least 1 methylated residue, at least 2 methylated residues, at least 3 methylated residues, at least 5 methylated residues, or at least 10 methylated residues.
[0131] The term "hypomethylation" refers to a reduction in the level or degree of methylation of nucleic acid molecules (plural possible) within a population of nucleic acid molecules (e.g., a sample) compared to other nucleic acid molecules. In some embodiments, hypomethylated DNA includes unmethylated DNA molecules. In some embodiments, hypomethylated DNA may include DNA molecules that contain 0 methylated residues, up to 1 methylated residue, up to 2 methylated residues, up to 3 methylated residues, up to 4 methylated residues, or up to 5 methylated residues.
[0132] The term "epigenetic state" refers to a particular level or extent of sequence-independent variability that can exist in a DNA sequence. In some embodiments, the epigenetic state of a DNA sequence refers to the degree or level of methylation of the sequence, nucleosome distribution, cfDNA fragmentation pattern, binding of CCCTC-binding factor ("CTCF"), transcription start site, or regulatory protein-binding region. Thus, the epigenetic state includes, but is not limited to, hypermethylation, hypomethylation, and the presence or absence of CTCF-binding sites or transcription start sites. The epigenetic state of a sequence can be a "reference epigenetic state" that can be used in comparison to the epigenetic state of a corresponding sequence within other DNA molecules. Examples of reference epigenetic states are states that are obtained from healthy subjects and are commonly found in samples that are not cancer-related.
[0133] As used herein, "methylation state" refers to the presence or absence of a methyl group at a DNA nucleobase (e.g., cytosine) at a specific genomic position within a nucleic acid, the degree of methylation of the nucleic acid (e.g., high, low, medium, or unmethylated), or the number of methylated nucleotides in a particular nucleic acid molecule. A nucleic acid that is "in a methylated form" means that the nucleic acid contains a sequence that includes a methylated DNA nucleobase, such as a methylated cytosine in a CpG dinucleotide.
[0134] As used herein, "consensus sequence" refers to a sequence derived from a redundant sequence of a parental molecule that is intended to represent the sequence of the original parental molecule. A consensus sequence includes base identity at a single position. In some embodiments, a consensus sequence can represent a single nucleotide base at a specific genomic position. In some embodiments, a consensus sequence can represent a run of nucleotide bases at multiple genomic positions. A consensus sequence can be generated by other methods such as voting (where the nucleotide that occupies the majority among the sequences, e.g., the nucleotide most commonly observed at a given base position, is the consensus nucleotide for each) or comparison to a reference genome. A consensus sequence can be generated, for example, by tagging the original parental molecule with a unique or non-unique molecular tag, thereby enabling tracking of progeny sequences (e.g., after amplification) by tracking of the tag and / or use of sequence read internal information. Examples of tagging or barcoding, and the use of tags or barcodes, are provided herein and in, for example, U.S. Patent Publications Nos. 2015 / 0368708, 2015 / 0299812, 2016 / 0040229, and 2016 / 0046986, each of which is incorporated herein by reference in its entirety.
[0135] As used herein, "sequence-specific nuclease" means a nuclease that cleaves only nucleic acid sequences that contain a specific sequence or consensus sequence. In some embodiments, the sequence-specific nuclease cleaves only nucleic acid sequences that contain a specific sequence or consensus sequence that is at least 15 nucleotides, at least 16 nucleotides, or at least 17 nucleotides in length. In some embodiments, the sequence-specific nuclease is "modification-independent" and cleaves a specific sequence regardless of the presence of modifications such as methylation. In some embodiments, the sequence-specific nuclease is "modification-dependent" and cleaves sequences that contain a specific sequence or consensus sequence and that further contain (or do not contain) one or more nucleic acid base modifications, such as methylation. In some embodiments, the modification-independent sequence-specific nuclease binds to a guide RNA that hybridizes to a sequence that is cleaved by the nuclease or in the vicinity thereof. Examples of sequence-specific nucleases include, but are not limited to, CRISPR (e.g., Cas nucleases such as Cas9), Argonaute, TALEN, and zinc finger nucleases. A "variant" sequence-specific nuclease contains at least one modification to at least one amino acid compared to the sequence-specific nuclease from which it is derived, has at least 80% sequence identity (e.g., at least 85%, 90%, 95%, 98%, or 99% sequence identity) to the sequence-specific nuclease from which it is derived, and retains at least some or enhanced nuclease function. The modification may be a substitution, deletion, or insertion of a natural or non-natural amino acid.
[0136] As used herein, "sequence-specific degradation" or "sequence-specifically degrading" means degradation that is dependent on the sequence of the site to be degraded and that is independent of modifications. Degrading includes any form of nucleic acid degradation, such as nucleotide strand cleavage or nucleic acid cleavage.
[0137] As used herein, "methylation-sensitive nuclease" refers to a nuclease that preferentially cuts unmethylated DNA compared to methylated DNA. For example, a methylation-sensitive nuclease can cut a recognition sequence such as a restriction site or its vicinity in a manner that depends on the absence of methylation of at least one of the nucleobases within the recognition sequence, for example, cytosine methylation. In some embodiments, the nuclease activity of a methylation-sensitive nuclease is at least 10-fold, 20-fold, 50-fold, or 100-fold higher at an unmethylated recognition site compared to a methylated control in a standard nuclease assay. Methylation-sensitive nucleases include methylation-sensitive restriction enzymes.
[0138] As used herein, "methylation-sensitive restriction enzyme" or "MSRE" refers to a methylation-sensitive nuclease that is a restriction enzyme. An MSRE is sensitive to the methylation state of DNA (e.g., cytosine methylation), i.e., the rate at which the enzyme cleaves DNA is altered by the presence or absence of a methyl group at a nucleotide base within its recognition sequence. In some embodiments, a methylation-sensitive restriction enzyme does not cleave DNA when a specific nucleotide base within the recognition sequence is methylated. For example, HpaII is a methylation-sensitive restriction enzyme that has a recognition sequence of "CCGG" and does not cleave DNA when the second cytosine within the recognition sequence is methylated.
[0139] As used herein, "methylation-dependent nuclease" refers to a nuclease that preferentially cuts methylated DNA compared to unmethylated DNA. For example, a methylation-dependent nuclease can cut a recognition sequence such as a restriction site or its vicinity in a manner that depends on the methylation of at least one of the nucleobases within the recognition sequence, for example, cytosine methylation. In some embodiments, the nuclease activity of a methylation-dependent nuclease is at least 10-fold, 20-fold, 50-fold, or 100-fold higher at a methylated recognition site compared to an unmethylated control in a standard nuclease assay. Methylation-dependent nucleases include methylation-dependent restriction enzymes.
[0140] As used herein, the term "methylation-dependent restriction enzyme" or "MDRE" refers to a methylation-dependent nuclease that is a restriction enzyme. MDREs are dependent on DNA methylation (e.g., cytosine methylation), i.e., the presence or absence of a methyl group in a nucleotide base alters the rate at which the enzyme cleaves DNA. In some embodiments, a methylation-dependent restriction enzyme does not cleave DNA when a specific nucleotide base in the recognition sequence is unmethylated. For example, MspJI is a methylation-dependent restriction enzyme that has the recognition sequence "mCNNR(N9)" and does not cleave DNA when there is no methylated cytosine (mC) within the recognition sequence.
[0141] The term "agent that recognizes a modified nucleobase in DNA" refers to a molecule or reagent that binds to or detects one or more modified nucleobases in DNA. A "modified nucleobase" is a nucleobase that includes a difference in chemical structure from an unmodified nucleobase. In the case of DNA, unmodified nucleobases are adenine, cytosine, guanine, or thymine. In some embodiments, the modified nucleobase is a modified cytosine. In some embodiments, the modified nucleobase is a methylated nucleobase. In some embodiments, the modified cytosine is methylcytosine, e.g., 5-methylcytosine. In such embodiments, the cytosine modification is a methyl group. Agents that recognize methylcytosine in DNA include, but are not limited to, "methyl-binding reagents", which as used herein refer to reagents that bind to methylcytosine. Methyl-binding reagents include, but are not limited to, methyl-binding domains (MBDs), methyl-binding proteins (MBPs), and antibodies specific for methylcytosine. In some embodiments, such antibodies bind to 5-methylcytosine in DNA. In some such embodiments, the DNA may be single-stranded or double-stranded.
[0142] "Or" is used in an inclusive sense, i.e., it is equivalent to "and / or" unless the context requires a different interpretation. II. Exemplary Methods A. Step of sequence-specifically degrading a plurality of nucleic acid sequences
[0143] The methods disclosed herein include the step of sequence - specifically degrading a plurality of nucleic acid sequences including modifications. In some embodiments, the step of sequence - specific degradation includes contacting the nucleic acids in a sample or a sub - sample thereof with a sequence - specific nuclease. The modified sequences to be degraded are those commonly found in nucleic acid sequences obtained from healthy subjects. In some embodiments, the nucleic acid sequence is a cfDNA sequence. In some embodiments, the modification is methylation. In some embodiments, the methylation comprises or consists of cytosine methylation. In some embodiments, the methylated sequences to be degraded include CpG sites that are generally methylated in cfDNA obtained from healthy subjects. In some embodiments, sequences containing CpG sites that are generally not methylated are not degraded by the sequence - specific nuclease. In some embodiments, the methods herein include additional elements or steps, such as those disclosed herein, to deplete unmodified sequences and sequences containing methylations or modifications other than the modification of interest that are commonly found in cfDNA obtained from healthy subjects. A combination of the step of depleting sequences modified in cfDNA from healthy subjects and the step of depleting sequences unmodified or containing other modifications in cfDNA from healthy subjects can result in a processed sample enriched in DNA molecules that are generally found only in non - healthy subjects. In some embodiments, the sequences remaining after the sequence - specific degradation step without being cleaved are detected. Thus, some of the methods herein provide for the detection of the presence or absence of abnormally modified nucleic acid sequences by specifically degrading normally modified nucleic acid sequences commonly found in samples obtained from healthy subjects. In some embodiments, the processed sample is enriched for modified sequences associated with cancer. Generally, the methods described herein can facilitate more efficient capacity use in downstream analysis steps (e.g., detection of target sequences, such as by sequencing).This is achieved by degrading arrays that have no information value, such as arrays that have modifications such as methylation in all, essentially all, or a majority of the samples, such that such arrays with no information value do not participate in and / or consume capacity during downstream analysis steps. In some embodiments, arrays with no information value include repetitive DNA elements (e.g., SINE, LINE, and / or Alu elements) and / or DNA that is methylated in the erythroid lineage in healthy cells.
[0144] In some embodiments, the step of sequence-specific degradation includes cleavage of both strands of double-stranded DNA that results in the formation of double-stranded ends. In some embodiments, the double-stranded ends are blunt ends. In some embodiments, the double-stranded ends include overhangs of one or more nucleosides.
[0145] In some embodiments, the step of sequence-specific cleavage is performed after the step of partitioning DNA based on methylation status. For example, see FIG. 1. In some such embodiments, the method includes the step of contacting the DNA with a methyl-binding domain (MBD) specific for methylcytosine. In some embodiments, the step of sequence-specific cleavage is performed on hypermethylated-partitioned DNA, but not on hypomethylated-partitioned DNA. In some embodiments, the step of sequence-specific cleavage is performed before or after an adapter containing a barcode is ligated to the DNA. In certain embodiments, the step of sequence-specific cleavage is performed after an adapter containing a barcode is ligated to the DNA. In some embodiments, the step of sequence-specific cleavage is after the step of ligating a barcode-containing adapter to the DNA and (a) before the step of partitioning the sample into a plurality of secondary samples by contacting the DNA with an agent that recognizes a modification associated with the DNA, (b) after the step of partitioning the sample into a plurality of secondary samples by contacting the DNA with an agent that recognizes a modification associated with the DNA, (c) before the step of contacting the sample or its secondary sample with an MSRE, (d) after the step of contacting the sample or its secondary sample with an MSRE, (e) before the step of amplifying the DNA before the step of detecting, (f) after the step of amplifying the DNA before the step of detecting, (before the step of enriching one or more of the plurality of target regions within at least one set of target regions, after the step of enriching one or more of the plurality of target regions within at least one set of target regions, (i) before two or more samples (or two or more secondary samples) are pooled (e.g., within the same flow cell in a next-generation sequencing reaction) before the step of sequencing the sample, (j) after two or more samples (or two or more secondary samples) are pooled (e.g., within the same flow cell in a next-generation sequencing reaction) before the step of sequencing the sample, or (k) (a) and any one, two, three, or four of (c)-(j), (b) and(c) to (j), any one, two, three, or four of them, (c), and any one, two, three, or four of (a), (b), and (e) to (j), (d), and any one, two, three, or four of (a), (b), and (e) to (j), (e), and any one, two, three, or four of (a) to (d) and (g) to (j), (f), and any one, two, three, or four of (a) to (d) and (g) to (j), (g), and any one, two, three, or four of (a) to (f) and (i) or (j), (h), and any one, two, three, or four of (a) to (f) and (i) to (j), (i), and any one, two, three, or four of (a) to (h); or (j) and any one, two, three, or four of (a) to (h) are performed in any combination. In some embodiments, the DNA is then amplified, sequenced, and analyzed.,
[0146] In some embodiments, the step of sequence-specific cleavage is performed after the steps of partitioning the DNA based on methylation status, ligating an adapter containing a barcode to the DNA, and cleaving the hypomethylated partition DNA with an MSRE. See, for example, FIG. 3. In some such embodiments, the DNA remaining after the step of sequence-specific cleavage is amplified by PCR. In some embodiments, the adapter containing a barcode is ligated to the hypomethylated partition DNA, and then the DNA is amplified by PCR and enriched using a DNA or RNA probe for the set of target regions. In some such embodiments, the hypomethylated partition DNA is not sequence-specifically cleaved. In some embodiments, the enriched DNA of the hypermethylated partition is combined with the DNA remaining within the hypermethylated partition prior to sequencing and analysis., 1. Sequence-specific nuclease
[0147] In some embodiments, the step of sequence-specifically degrading is performed by contacting a nucleic acid such as DNA with a sequence-specific nuclease. In some embodiments, the sequence-specific nuclease is a modification-independent sequence-specific nuclease. Examples of modification-independent sequence-specific nucleases include, but are not limited to, CRISPR nucleases, TALENs, zinc fingers, and Argonaute nucleases.
[0148] In some embodiments, the modification-independent sequence-specific nuclease is a CRISPR nuclease. Exemplary CRISPR nucleases include Cas9, such as Streptococcus pyogenes Cas9 nuclease or variants thereof, Staphylococcus aureus Cas9 or variants thereof, Cas12, such as Cas12a or Cas12b nuclease or variants thereof, and type II and type V Cas nucleases including CasX nuclease or variants thereof. In some embodiments, the Cas nuclease is a multi-turnover Cas nuclease or a high-fidelity variant. Some exemplary CRISPR nucleases are further described, for example, in Yourik et al. Staphylococcus aureus Cas9 is a multiple-turnover enzyme. RNA 25:35-44 (2019) and Kleinstiver et al. High-fidelity CRISPR-Cas9 variants with undetectable genome-wide off-targets. Nature 529: 490-495 (2016), which are hereby incorporated by reference in their entirety. 2. Guide RNA and Sequences Specifically Degraded
[0149] In some embodiments, the step of sequence-specifically degrading comprises contacting the DNA with one or more guide RNAs that guide the DNA to a specific sequence (s) to be degraded by a nuclease. Suitable guide RNA sequences are selected based on the specific sequence to be degraded and the nuclease used. For example, when using a CRISPR nuclease, one or more suitable guide RNAs recognized by the CRISPR nuclease are used. In some embodiments, a single guide RNA ("sgRNA") or other type of fusion or truncated CRISPR guide RNA is used with a suitable CRISPR nuclease. Exemplary guide RNAs are described, for example, in Fu et al. Improving CRISPR-Cas nuclease specificity using truncated guide RNAs. Nat. Biotechnol. 32:279-284 (2014), which is hereby incorporated by reference in its entirety.
[0150] In some embodiments, the one or more guide RNAs comprise one or more modifications. In some such embodiments, the guide RNA comprises a modified internucleoside linkage. In some embodiments, the modified internucleoside linkage is a phosphorothioate internucleoside linkage. In some embodiments, the guide RNA comprises a modified sugar. In some embodiments, the modified sugar comprises a 2'-substitution. In some embodiments, the 2'-substitution is 2'-fluoro, 2'-O-methoxyethyl, 2'-O-alkyl substitution, or 2'-hydroxy substitution. In the context of a guide RNA, a sugar comprising a 2'-hydroxy substitution, such as a 2'-deoxyribosyl sugar, is a modified sugar. In some embodiments, the 2'-O-alkyl substitution is 2'-O-methyl substitution. In some embodiments, the modified sugar is a bicyclic sugar. In some embodiments, the bicyclic sugar is an LNA, cEt, or ENA sugar. In some embodiments, the modified sugar is a linear sugar. In some embodiments, the linear sugar is a UNA sugar.
[0151] A guide RNA configured to bind to a particular sequence includes a portion that specifically binds to the particular sequence. In some embodiments, a portion of each guide RNA specifically binds to a target region of DNA or a portion of the target region. In some embodiments, a portion of each guide RNA specifically binds to a series of DNA containing modifications, and the modified version of the sequence is generally found in cfDNA obtained from healthy subjects.
[0152] In some embodiments, the step of array-specific degradation comprises contacting the DNA with a plurality of guide RNAs. In some embodiments, each of the unique guide RNAs, if present, specifically binds to a different member of a plurality of specific sequences in the sample or a secondary sample thereof. In some embodiments, the plurality of guide RNAs includes a guide RNA that specifically binds to a modified sequence in cfDNA obtained from a healthy subject. In some embodiments, the plurality of guide RNAs includes a guide RNA that specifically binds to a sequence containing a CpG motif that is methylated in cfDNA obtained from a healthy subject. In some embodiments, the plurality of sequences to be degraded are from 1 million to 100 million sequences. In some embodiments, the plurality of sequences to be degraded are 20 to 250,000, 20 to 100,000, 20 to 10,000, 20 to 1,000, or 20 to 100 sequences. In some embodiments, the plurality of sequences to be degraded are 50 to 100,000, 50 to 10,000, 50 to 1,000, or 50 to 100 sequences. In some embodiments, the plurality of sequences to be degraded are 100 to 100,000, 100 to 10,000, or 100 to 1,000 sequences. In some embodiments, each of the plurality of sequences to be degraded is a sequence present at a different gene or locus. In some embodiments, the sequences to be degraded include or are selected from repetitive elements (e.g., LINE, such as LINE1; Alu elements; and / or SINE). In some embodiments, the sequences to be degraded include sequences known to be modified in cfDNA obtained from a healthy subject. In some embodiments, the sequences to be degraded include sequences determined using existing methods for detecting modification sites in the DNA sequence that are modified in cfDNA obtained from a healthy subject.
[0153] In some embodiments, the methods herein include an element or step that depletes unmodified or unmethylated sequences commonly found in cfDNA obtained from a healthy subject. In some such embodiments, the step of sequence-specific cleavage includes cleaving sequences lacking motifs that can be modified or methylated. In some such embodiments, the sequences to be cleaved include sequences lacking CpG motifs. In some embodiments, the step of sequence-specific cleavage includes contacting the DNA with a plurality of guide RNAs, where the plurality of guide RNAs includes guide RNAs that specifically bind to sequences that are not methylated in cfDNA obtained from a healthy subject. In some embodiments, each of the plurality of guide RNAs specifically binds to a sequence that is commonly found in a form that includes a modification in cfDNA from a healthy subject or a sequence that is commonly found in a form that lacks a modification in cfDNA from a healthy subject. B. Degradation by methylation-sensitive nucleases of DNA
[0154] In some embodiments, the methods herein include contacting the DNA with a methylation-sensitive nuclease, thereby including the step of degrading DNA that includes unmethylated sequences or sequences having a low level of methylation. In some such embodiments, the methylation-sensitive nuclease is a methylation-sensitive restriction enzyme (MSRE), and thus degrades DNA that includes the unmethylated recognition site of the MSRE. Thus, methylation-sensitive nucleases can be used in the methods herein that include one or more steps that deplete unmodified or unmethylated sequences commonly found in cfDNA from a healthy subject.
[0155] In some embodiments, the step of contacting the DNA with a methylation-sensitive nuclease is performed before the step of sequence-specifically cleaving the DNA with a sequence-specific nuclease. In some embodiments, the step of contacting the DNA with a methylation-sensitive nuclease is performed after the step of sequence-specifically cleaving the DNA with a sequence-specific nuclease. In some embodiments, the step of contacting the DNA with a methylation-sensitive nuclease is performed simultaneously with the step of sequence-specifically cleaving the DNA with a sequence-specific nuclease. C. Distribution of the sample into a plurality of secondary samples
[0156] The methods disclosed herein include analyzing DNA in a sample. In such methods, different forms of DNA (e.g., hypermethylated DNA and hypomethylated DNA) can be physically distributed into a plurality of secondary samples based on one or more characteristics of the DNA. This approach can be used, for example, to determine whether a particular sequence is hypermethylated or hypomethylated.
[0157] In some embodiments, the distributing step includes contacting the DNA with an agent that recognizes a modification associated with (e.g., within) the DNA. In some embodiments, the agent that recognizes the modification is an antibody. In some embodiments, the agent is immobilized on a solid support. In some embodiments, the distributing step includes immunoprecipitation using an antibody agent, such as an antibody immobilized on a solid support.
[0158] In some embodiments, the modification is methylation, and in some such embodiments, the distributing step comprises distributing based on the methylation level. In some such embodiments, the agent is a methyl-binding reagent. In some embodiments, the methyl-binding reagent specifically recognizes 5-methylcytosine. In some such embodiments, the agent is a hydroxymethyl-binding reagent. In some embodiments, the methyl-binding reagent specifically recognizes 5-hydroxymethylcytosine, biotinylated 5-hydroxymethylcytosine, glucosylated 5-hydroxymethylcytosine, or sulfonylated 5-hydroxymethylcytosine. In some embodiments, the distributing step comprises distributing based on binding to a protein, which comprises contacting a sample containing DNA with a binding reagent specific for the protein. In some such embodiments, the binding reagent specifically binds to a methylated protein, an acetylated protein, for example, a methylated or acetylated histone. In some embodiments, the binding reagent specifically binds to a non-methylated or non-acetylated protein epitope.
[0159] In some embodiments, the modification is hydroxymethylation, and in some such embodiments, the distributing step comprises distributing based on the hydroxymethylation level. In some such embodiments, the agent is a hydroxymethyl-binding reagent, such as an antibody. In some embodiments, the hydroxymethyl-binding reagent (e.g., antibody) specifically recognizes 5-hydroxymethylcytosine (5-hmC). In some embodiments, a modification such as hydroxymethylation is labeled (e.g., biotinylated, glucosylated, or sulfonated) prior to contacting with an agent that recognizes the labeled form of the modification. For example, 5-hmC can be enzymatically glucosylated and then distributed based on binding to J-binding protein 1. Exemplary methods for labeling and / or distributing 5-hmC are provided, for example, in Song et al., Nat. Biotech. 29:68-72 (2010); Ko et al., Nature 468:839-843 (2010); and Robertson et al., Nucleic Acids Res. 39:e55 (2011).
[0160] If immunoprecipitation is used and an antibody that recognizes single-stranded DNA is involved, the DNA can be converted to a double-stranded form by complementary strand synthesis prior to the step of specifically degrading the sequence. For such synthesis, an adapter can be used as the primer-binding site, or random priming can be used.
[0161] In some embodiments, a sample containing DNA is distributed into a plurality of secondary samples. In some embodiments, the plurality of distributed secondary samples includes two secondary samples, a first secondary sample and a second secondary sample. In some embodiments, the plurality of distributed secondary samples includes three secondary samples, a first secondary sample, a second secondary sample, and a third secondary sample. In some embodiments, the method includes a distributing step that is performed prior to the sequence-specific cleavage step. Some such embodiments include contacting the first secondary sample with a sequence-specific nuclease. Some such embodiments include contacting the first secondary sample with a sequence-specific nuclease and contacting the second secondary sample with a sequence-specific nuclease. In some embodiments that include a third distributed secondary sample, the third secondary sample includes DNA associated with modifications at a higher rate than in the second secondary sample and at a lower rate than in the first secondary sample. In some embodiments, the method includes a distributing step that is performed prior to the step of sequence-specifically cleaving the DNA and prior to the step of contacting the DNA with a methylation-sensitive nuclease.
[0162] By partitioning nucleic acid molecules in a sample, for example, enriching rare nucleic acid molecules that are more commonly seen in one partition of the sample, rare signals can be increased. For example, genetic diversity that exists in highly methylated DNA but is less (or not at all) present in lowly methylated DNA can be more easily detected by partitioning the sample into highly methylated nucleic acid molecules and lowly methylated nucleic acid molecules. By analyzing multiple partitions of a sample, multi-dimensional analysis of single molecules can be performed, and thus higher sensitivity can be achieved. Partitioning can include physically partitioning or sub-sampling nucleic acid molecules based on the presence or absence of one or more methylated nucleobases. The sample can be partitioned or sub-sampled based on differential gene expression or characteristics indicative of a disease state. The sample can be partitioned during the analysis of nucleic acids, such as cell-free DNA (cfDNA), non-cfDNA, tumor DNA, circulating tumor DNA (ctDNA), and cell-free nucleic acids (cfNA), based on characteristics or combinations thereof that result in a difference in signal between a normal state and a diseased state.
[0163] In some embodiments, highly methylated and / or lowly methylated variable target regions are analyzed to determine whether they exhibit differential methylation characteristic of tumor cells, or cell types that do not normally contribute to the DNA sample being analyzed (e.g., cfDNA), and / or certain immune cell types.
[0164] In some embodiments, each partition is differentially tagged. The tagged partitions can then be pooled together for collective sample preparation and / or sequencing. The partition-tag-pool steps can be performed more than once, with each round of the partitioning step being performed based on different characteristics (examples provided herein) and tagging using differential tags that distinguish it from other partitioning and partitioning means. In other cases, differentially tagged partitions are sequenced separately.
[0165] In some embodiments, the differentially tagged and pooled sequence reads obtained from the DNA are analyzed in silico. Tags are used to sort the reads from different pools. Analyses for detecting genetic variants can be performed at the level of each pool, as well as at the level of the entire nucleic acid population. For example, the analysis can include in silico analysis to determine genetic variants in the nucleic acids within each pool, such as CNVs, SNVs, indels, fusions. In some cases, the in silico analysis can include determining chromatin structure. For example, the coverage of the sequence reads can be used to determine nucleosome positioning of chromatin. Higher coverage can correlate with higher nucleosome occupancy within a genomic region, while lower coverage can correlate with lower nucleosome occupancy or nucleosome depleted regions (NDRs).
[0166] Examples of features that can be used for the step of partitioning include the length of the array, methylation level, array mismatches, immunoprecipitation, and / or proteins that bind to DNA. The resulting partitioning may include one or more of the following nucleic acid forms: single-stranded DNA (ssDNA), double-stranded DNA (dsDNA), shorter DNA fragments, and longer DNA fragments. In some embodiments, a partitioning step based on cytosine modification (e.g., cytosine methylation) or methylation is generally performed and, optionally, combined with at least one additional partitioning step that may be based on any of the aforementioned DNA features or forms. In some embodiments, a heterogeneous population of nucleic acids is partitioned into nucleic acids having one or more epigenetic modifications and nucleic acids having no epigenetic modifications. Examples of epigenetic modifications include the presence or absence of methylation; the level of methylation; the type of methylation (e.g., is it 5-methylcytosine or another type of methylation, such as adenine methylation and / or cytosine hydroxymethylation); and the association and level of association with one or more proteins such as histones. Alternatively or in addition, a heterogeneous population of nucleic acids can be partitioned into nucleic acid molecules with nucleosomes and nucleic acid molecules lacking nucleosomes. Alternatively or in addition, a heterogeneous population of nucleic acids can be partitioned into single-stranded DNA (ssDNA) and double-stranded DNA (dsDNA). Alternatively or in addition, a heterogeneous population of nucleic acids can be partitioned based on the length of the nucleic acid (e.g., molecules up to 160 bp and molecules longer than 160 bp).
[0167] Agents used to partition a population of nucleic acids in a sample can be affinity agents, such as antibodies with desired specificities, natural binding partners or variants thereof (Bock et al., Nat Biotech 28: 1106-1114 (2010); Song et al., Nat Biotech 29: 68-72 (2011)), or artificial peptides selected, for example, by phage display to have specificity for a given target. In some embodiments, the agent used in the partitioning step is an agent that recognizes a modified nucleobase. In some embodiments, the modified nucleobase recognized by the agent is a modified cytosine such as methylcytosine (e.g., 5-methylcytosine). In some embodiments, the modified nucleobase recognized by the agent is the product of a procedure that affects the first nucleobase in the DNA of the sample in a different form than the second nucleobase in the DNA. In some embodiments, the modified nucleobase may be a "converted nucleobase", i.e., one whose base pairing specificity has been changed by a procedure. For example, by a particular procedure, non-methylated or unmodified cytosine is converted to dihydrouracil, or more generally, at least one modified or unmodified form of cytosine undergoes deamination, thereby producing uracil (considered a modified nucleobase in the context of DNA) or a further modified form of uracil. Examples of partitioning agents include antibodies, such as antibodies that recognize modified nucleobases that can be modified cytosines such as methylcytosine (e.g., 5-methylcytosine). In some embodiments, the partitioning agent is an antibody that recognizes a modified cytosine other than 5-methylcytosine, such as 5-carboxylcytosine (5caC). Exemplary partitioning agents include methyl-binding domains (MBDs) and methyl-binding proteins (MBPs) described herein, including proteins such as MeCP2, MBD2, and antibodies that preferentially bind 5-methylcytosine. When using an antibody to immunoprecipitate methylated DNA, the methylated DNA can be recovered in single-stranded form. In such embodiments, a second strand can be synthesized.Next, the hypermethylated (and, optionally, intermediate-methylated) secondary sample can be contacted with a methylation-sensitive nuclease that does not cleave hemimethylated DNA, such as HpaII, BstUI, or Hin6I. Alternatively, or in addition, the hypomethylated (and, optionally, intermediate-methylated) secondary sample can then be contacted with a methylation-dependent nuclease that cleaves hemimethylated DNA.
[0168] Additional non-limiting examples of partitioning agents or binding reagents are histone-binding proteins that can separate nucleic acids bound to histones from free or unbound nucleic acids. Examples of histone-binding proteins that can be used in the methods disclosed herein include RBBP4, RbAp48, and SANT domain peptides.
[0169] In some embodiments, the partitioning step can include both binary partitioning and partitioning based on degree / level of modification. For example, methylated fragments can also be partitioned by methylated DNA immunoprecipitation (MeDIP), and all methylated fragments can be separated and partitioned from unmethylated fragments using a methyl-binding domain protein (e.g., MethylMinder Methylated DNA Enrichment Kit (ThermoFisher Scientific)). Subsequently, additional partitioning steps can involve eluting fragments with different levels of methylation by adjusting the salt concentration in a solution containing the methyl-binding domain and the bound fragments. As the salt concentration increases, fragments with higher methylation levels elute.
[0170] In some cases, the final distribution is enriched in nucleic acids with different degrees of modification (overrepresentative or underrepresentative of the modification). Overrepresentation and underrepresentation can be defined by the number of modifications contributed by the nucleic acid compared to the median number of modifications per strand in the population. For example, if the median number of 5-methylcytosine residues in the nucleic acids of a sample is 2, nucleic acids containing more than 2 5-methylcytosine residues are overrepresented with respect to this modification, and nucleic acids with 1 or zero 5-methylcytosine residues are underrepresented. The effect of affinity separation is that nucleic acids with overrepresented modifications are enriched in the binding phase and nucleic acids with underrepresented modifications are enriched in the non-binding phase (i.e., in solution). After eluting the nucleic acids in the binding phase, they can then be processed.
[0171] When using MeDIP or the MethylMiner® Methylated DNA Enrichment Kit (ThermoFisher Scientific), sequential elution can be used to partition various levels of methylation. For example, hypomethylated fractions (unmethylated) can be separated from methylated fractions by contacting a nucleic acid population with MBD from the kit that is attached to magnetic beads. The beads are used to separate methylated nucleic acids from unmethylated nucleic acids. One or more elution steps are then performed sequentially to elute nucleic acids with different levels of methylation. For example, a first set of methylated nucleic acids can be eluted at a salt concentration of 160 mM or higher, such as at least 150 mM, at least 200 mM, 300 mM, 400 mM, 500 mM, 600 mM, 700 mM, 800 mM, 900 mM, 1000 mM, or 2000 mM. After eluting such methylated nucleic acids, magnetic separation is used again to separate nucleic acids with a higher level of methylation from nucleic acids with a lower level of methylation. The elution and magnetic separation steps are repeated to generate various fractions such as hypomethylated fractions (enriched in nucleic acids without methylation), methylated fractions (enriched in nucleic acids with low levels of methylation), and hypermethylated fractions (enriched in nucleic acids with high levels of methylation).
[0172] In some methods, nucleic acids bound to the agent used for partitioning based on affinity separation are subjected to a washing step. In the washing step, nucleic acids that weakly bind to the affinity agent are washed away. Such nucleic acids may be enriched in nucleic acids that are modified to an extent close to the average or median (i.e., intermediate between the nucleic acids that remained bound to the solid phase and the nucleic acids that were not bound to the solid phase upon initial contact of the sample with the agent).
[0173] By affinity separation, at least two, and sometimes three or more, distributions of nucleic acids with different degrees of modification are brought about. The distributions are still separate, but at least one distribution, and usually two or three (or more) distributions of nucleic acids, are linked to nucleic acid tags that are usually brought about as components of the adapter, and nucleic acids within different distributions receive different tags that distinguish members of one distribution from members of another distribution. Tags linked to nucleic acid molecules within the same distribution may be the same or different from each other. However, if they are different from each other, the tags may share a portion of their code such that it is identified that the molecules to which they are attached are from a particular distribution.
[0174] For further details regarding the partitioning of nucleic acid samples based on features such as methylation, see WO2018 / 119452, which is incorporated herein by reference.
[0175] In some embodiments, the partitioning step is performed after contacting the DNA with a methylation-sensitive restriction enzyme (MSRE) and / or a methylation-dependent restriction enzyme (MDRE). After treatment of the DNA with MSRE or MDRE, the DNA is partitioned based on size to generate a highly methylated sub-sample (the longest DNA molecules after MSRE treatment and the shortest DNA fragments after MDRE treatment), a moderately methylated sub-sample (DNA molecules of intermediate length after MSRE or MDRE treatment), and a lowly methylated sub-sample (the shortest DNA molecules after MSRE treatment and the longest DNA fragments after MDRE treatment).
[0176] In some embodiments, the step of partitioning is performed by contacting the nucleic acid with the methyl-binding domain ("MBD") of a methyl-binding protein ("MBP"). In some such embodiments, the nucleic acid is contacted with the entire MBP. In some embodiments, the MBD binds to 5-methylcytosine (5mC), the MBP includes the MBD, and is herein referred to interchangeably as a methyl-binding protein or a methyl-binding domain protein. In some embodiments, the MBD is attached via a biotin linker to paramagnetic beads, such as Dynabeads® M-280 streptavidin. Partitioning into fractions of different degrees of methylation can be performed by eluting the fractions by increasing the NaCl concentration.
[0177] In some embodiments, the bound DNA is eluted by contacting the antibody or MBD with a protease such as proteinase K. This can be done instead of or in addition to the elution step using NaCl discussed herein.
[0178] Examples of agents that recognize modified nucleobases contemplated herein include, but are not limited to, the following:
[0179] (a) MeCP2 and MBD2 are proteins that preferentially bind 5-methyl-cytosine over unmodified cytosine.
[0180] (b) RPL26, PRP8, and DNA mismatch repair protein MHS6 preferentially bind 5-hydroxymethyl-cytosine over unmodified cytosine.
[0181] (c) FOXK1, FOXK2, FOXP1, FOXP4, and FOXI3 preferably bind to 5-formyl-cytosine over unmodified cytosine (Iurlaro et al., Genome Biol. 14: R119 (2013)).
[0182] (d) An antibody specific for one or more methylated or modified nucleobases or their conversion products, e.g., 5mC, 5caC, or DHU.
[0183] Generally, elution is according to the number of modifications per molecule, e.g., the number of methylation sites, and molecules with more methylation elute at higher salt concentrations. To elute DNA into separate populations based on the degree of methylation, a series of elution buffers with gradually increasing NaCl concentrations can be used. The salt concentration can range from about 100 nM to about 2500 mM NaCl. In one embodiment, the process results in three fractions. The molecule is contacted with a solution of a first salt concentration that contains an agent that recognizes the modified nucleobase and a molecule that can bind to a capture moiety such as streptavidin. At the first salt concentration, one population of molecules binds to the agent and one population remains unbound. The unbound population can be separated as the "low methylation" population. For example, the first fraction enriched in the low-methylated form of DNA is the one that remains unbound at a low salt concentration, e.g., 100 mM or 160 mM. The second fraction enriched in moderately methylated DNA is eluted using an intermediate salt concentration, e.g., a concentration between 100 mM and 2000 mM. This is also separated from the sample. The third fraction enriched in the high-methylated form of DNA is eluted using a high salt concentration, e.g., at least about 2000 mM.
[0184] In some embodiments, a monoclonal antibody raised against 5-methylcytidine (5mC) is used to purify methylated DNA. The DNA is denatured, e.g., at 95 °C, to obtain single-stranded DNA fragments. The DNA bound to the antibody is immunoprecipitated using protein G bound to standard or magnetic beads and washing after incubation with the anti-5mC antibody. The DNA can then be eluted. The fractionation can include the DNA that did not precipitate and one or more fractions eluted from the beads. In some embodiments, the DNA fractionation is desalted and concentrated for use in the enzymatic steps of library preparation.
[0185] In some embodiments, the method includes preparing a first pool that includes at least a portion of the DNA with low methylation distribution. In some embodiments, the method includes preparing a second pool that includes at least a portion of the DNA with high methylation distribution. In some embodiments, the first pool further includes a portion of the DNA with high methylation distribution. In some embodiments, the second pool further includes a portion of the DNA with low methylation distribution. In some embodiments, the first pool includes the majority of the DNA with low methylation distribution and, optionally, a minority of the DNA with high methylation distribution. In some embodiments, the second pool includes the majority of the DNA with high methylation distribution and a minority of the DNA with low methylation distribution. In some embodiments involving medium methylation distribution, the second pool includes at least a portion of the DNA with medium methylation distribution, e.g., the majority of the DNA with medium methylation distribution. In some embodiments, the first pool includes the majority of the DNA with low methylation distribution and the second pool includes the majority of the DNA with high methylation distribution and the majority of the DNA with medium methylation distribution.
[0186] In some embodiments, the method comprises the step of capturing at least a first set of target regions from a first pool, where, for example, the first pool is as described in any of the embodiments herein. In some embodiments, the first set comprises an array-variable target region. In some embodiments, the first set comprises a hypomethylated variable target region and / or a fragmented variable target region. In some embodiments, the first set comprises an array-variable target region and a fragmented variable target region. In some embodiments, the first set comprises an array-variable target region, a hypomethylated variable target region, and a fragmented variable target region. The step of amplifying the DNA in the first pool can be performed prior to this capture step. In some embodiments, the step of capturing a first set of target regions from the first pool comprises contacting the DNA of the first pool with a first set of target-specific probes. In some embodiments, the first set of target-specific probes comprises target-binding probes specific to an array-variable target region. In some embodiments, the first set of target-specific probes comprises target-binding probes specific to an array-variable target region, a hypomethylated variable target region, and / or a fragmented variable target region.
[0187] In some embodiments, the method comprises the step of capturing a second set or a plurality of sets of target regions from a second pool, wherein, for example, the first pool is as described in any of the embodiments herein. In some embodiments, the plurality of second regions includes epigenetic target regions, such as hypermethylation variable target regions and / or fragmentation variable target regions. In some embodiments, the plurality of second regions includes sequence variable target regions as well as epigenetic target regions, such as hypermethylation variable target regions and / or fragmentation variable target regions. The step of amplifying the DNA in the second pool can be performed prior to this capture step. In some embodiments, the step of capturing a plurality of sets of second target regions from the second pool includes contacting the DNA of the first pool with a second set of target-specific probes, wherein the second set of target-specific probes includes target-binding probes specific for sequence variable target regions, and the target-binding probes are specific for epigenetic target regions. In some embodiments, the first set of target regions and the second set of target regions are not the same. For example, the first set of target regions may include one or more target regions that are not present in the second set of target regions. Alternatively, or in addition, the second set of target regions may include one or more target regions that are not present in the first set of target regions. In some embodiments, at least one hypermethylation variable target region is captured from the second pool but not from the first pool. In some embodiments, a plurality of hypermethylation variable target regions are captured from the second pool but not from the first pool. In some embodiments, the first set of target regions includes sequence variable target regions, and / or the second set of target regions includes epigenetic target regions. In some embodiments, the first set of target regions includes sequence variable target regions and fragmentation variable target regions, and the second set of target regions includes epigenetic target regions, such as hypermethylation variable target regions and fragmentation variable target regions.In some embodiments, the first set of target regions includes variable array target regions, fragmented variable target regions, and hypomethylated variable target regions, and the second set of target regions includes epigenetic target regions, such as hypermethylated variable target regions and fragmented variable target regions.
[0188] In some embodiments, the first pool includes the majority of DNA with hypomethylated distribution and a portion (e.g., about half) of DNA with hypermethylated distribution, and the second pool includes a portion (e.g., about half) of DNA with hypermethylated distribution. In some such embodiments, the first set of target regions includes variable array target regions and / or the second set of target regions includes epigenetic target regions. The variable array target regions and / or epigenetic target regions may be as described in any of the embodiments described elsewhere herein.
[0189] Methylation profiling can involve determining methylation patterns across different regions of the genome. For example, molecules can be distributed based on the degree of methylation (e.g., the relative number of methylated nucleobases per molecule), sequenced, and then the sequences of the molecules within different distributions can be mapped to a reference genome. This can show regions of the genome that are more highly methylated or less highly methylated compared to other regions. Thus, genomic regions can have different degrees of methylation as opposed to individual molecules. D. Adapter Ligation or Addition; Tagging
[0190] In some embodiments, the disclosed method includes the step of adding an adapter to the DNA. In some embodiments, the adapter is added prior to the step of sequence-specifically cleaving the DNA sequence. In some embodiments, the adapter is added to the DNA after the step of partitioning and prior to the step of sequence-specifically cleaving. In some embodiments, the adapter is added to the DNA prior to the step of partitioning and prior to the step of sequence-specifically cleaving. In some embodiments, the adapter can be added to the DNA before or after the amplification step, for example, by providing the adapter to the 5' portion of the primer (which, when using PCR, can be referred to as library preparation-PCR or LP-PCR), in parallel with the amplification procedure. In some embodiments, the adapter is added by other means. In some such methods, the first adapter is added to the nucleic acid by ligation to its 3' end, which may include ligation to single-stranded DNA. The adapter can be used, for example, as a priming site for second-strand synthesis using a universal primer and DNA polymerase. The second adapter can then be ligated to at least the 3' end of the second strand of the molecule, which is double-stranded at that point. In some embodiments, the first adapter includes an affinity tag such as biotin, and the nucleic acid ligated to the first adapter is bound to a solid support (e.g., beads) that includes a binding partner for the affinity tag, such as streptavidin. For further considerations regarding related procedures, see Gansauge et al., Nature Protocols 8: 737-748 (2013). Commercially available kits for sequencing library preparations that are compatible with single-stranded nucleic acids, such as the Accel-NGS® Methyl-Seq DNA Library Kit from Swift Biosciences, are available. In some embodiments, the nucleic acid is amplified after adapter ligation. In some embodiments, end repair of the DNA is performed prior to the addition of the adapter.
[0191] In some embodiments, after the adapter is ligated, the nucleic acid is subjected to amplification. For amplification, for example, universal primers that recognize primer-binding sites within the adapter can be used.
[0192] In some embodiments, DNA is ligated at both ends to a Y-shaped adapter that contains primer-binding sites and tags. In some such embodiments, the DNA is amplified.
[0193] In some embodiments, unwanted adapter dimers may form. In some such embodiments, the methods herein include the step of sequence-specifically cleaving sequences that contain barcode-containing adapter dimer junctions. In some such embodiments, the plurality of guide RNAs includes a plurality of guide RNAs configured to specifically bind to each of the possible adapter dimer junctions. In such embodiments, the plurality of guide RNAs includes at least one guide RNA that contains a sequence that is complementary to each one of the sequences of the possible adapter dimer junctions.
[0194] Tagging of DNA molecules is the procedure of attaching or associating a tag to a DNA molecule. Such a tag can be a molecule such as a nucleic acid that contains information indicative of the identity of the molecule to which the tag is associated. The tag can enable identification of the molecule from which a sequence read originated. For example, the molecule can have a sample tag (distinguishing molecules in one sample from molecules in a different sample) or a molecule tag / molecular barcode / barcode (distinguishing different molecules from one another, in both unique tagging scenarios and non-unique tagging scenarios). For methods involving a dispensing step, a dispensing tag (distinguishing molecules in one dispense from molecules in a different dispense) can be included. In some embodiments, an adapter added to a DNA molecule contains a tag. In some such embodiments, the tag contains one barcode or a combination of barcodes. As used herein, the term “barcode” refers, depending on the context, to a nucleic acid molecule having a specific nucleotide sequence, or to the nucleotide sequence itself. A barcode can have, for example, between 10 and 100 nucleotides. A set of barcodes can have degenerate sequences, or sequences having a particular Hamming distance desirable for a particular purpose. Thus, for example, a molecular barcode can consist of one barcode, or a combination of two barcodes each bound to a different end of the molecule. In addition to, or instead of, this, different sets of molecular barcodes or molecular tags can be used for different dispensings and / or samples, and thus the barcodes function as molecular tags through their individual sequences, and also function to identify the corresponding dispensing and / or sample based on the set of which they are members. Tags containing barcodes can be incorporated into an adapter or joined in other ways. Among other methods, tags can be incorporated by ligation, overlap extension PCR.
[0195] Tagging strategies can be divided into unique tagging and non-unique tagging strategies. In unique tagging, all or substantially all of the molecules in a sample have different tags, and thus, reads can be assigned to the original molecules based on tag information alone. Tags used in such methods are sometimes referred to as "unique tags." In non-unique tagging, different molecules in the same sample can have the same tag, and thus, additional information is used in addition to tag information to assign sequence reads to the original molecules. Such information can include start and stop coordinates, coordinates to which the molecule is mapped, start coordinate alone or stop coordinate alone, etc. Tags used in such methods are sometimes referred to as "non-unique tags." Thus, it is not necessary to uniquely tag every molecule in a sample. It is sufficient to uniquely tag the molecules that fall within an identifiable class within the sample. Thus, molecules within different identifiable families can have the same tag without loss of information regarding the identity of the tagged molecules.
[0196] In some embodiments, the adapter includes a sufficient number of different tags such that the number of combinations of tags results in a low probability, e.g., 95%, 99%, or 99.9%, that two nucleic acids having the same start and stop points receive the same combination of tags. The adapter can include the same or different primer binding sites, whether having the same tag or different tags. In some embodiments, the adapter includes the same primer binding site.
[0197] In certain embodiments with non-unique tagging, the number of different tags used can be sufficient such that there is a very high likelihood (e.g., at least 99%, at least 99.9%, at least 99.99% or at least 99.999%) that all the molecules of a particular group have different tags. For example, in some embodiments where barcode linkages are randomly included at both ends of the molecule, the combination of barcodes constitutes the tag together. This number in turn depends on the number of molecules that enter the call. For example, a class can be all the molecules that map to the same start - stop position on the reference genome. A class can be all the molecules that map across a particular locus, e.g., a particular base or a particular region (e.g., up to 100 bases or a gene or an exon of a gene). In certain embodiments, the number of different tags used to uniquely identify the number of molecules, z, within a class is 2 * z, 3 * z, 4 * z, 5 * z, 6 * z, 7 * z, 8 * z, 9 * z, 10 * z, 11 * z, 12 * z, 13 * z, 14 * z, 15 * z, 16 * z, 17 * z, 18 * z, 19 * z, 20 * z or 100 * either z (e.g., lower limit) ~ 100,000 * z, 10,000 * z, 1000 * z or 100 * can be anywhere between either z (e.g., upper limit) and 100,000.
[0198] For example, in a sample of about 5 ng to 30 ng of cell-free DNA, approximately 3000 molecules are mapped to a specific nucleotide coordinate, and it is expected that between about 3 and 10 molecules with any starting coordinate share the same termination coordinate. Thus, to uniquely tag all such molecules, between about 50 and about 50,000 different tags (e.g., barcode combinations between about 6 and 220) may be sufficient. To uniquely tag all 3000 molecules mapped across the nucleotide coordinates, between about 1 million and about 20 million different tags are required.
[0199] Generally, the assignment of unique or non-unique tag barcodes in a reaction follows the methods and systems described in U.S. Patent Application Nos. 20010053519, 20030152490, 20110160078, as well as U.S. Patent Nos. 6,582,908, 7,537,898, and 9,598,731. Tags can be ligated to the sample nucleic acid either randomly or non-randomly.
[0200] In some embodiments, the tagged nucleic acids are sequenced after being loaded into a microwell plate. The microwell plate may have 96, 384, or 1536 microwells. In some cases, the tagged nucleic acids are introduced at an expected ratio relative to the microwells of unique tags. For example, the unique tags can be loaded such that more than about 1, more than about 2, more than about 3, more than about 4, more than about 5, more than about 6, more than about 7, more than about 8, more than about 9, more than about 10, more than about 20, more than about 50, more than about 100, more than about 500, more than about 1000, more than about 5000, more than about 10000, more than about 50,000, more than about 100,000, more than about 500,000, more than about 1,000,000, more than about 10,000,000, more than about 50,000,000, or more than about 1,000,000,000 unique tags are loaded per genomic sample. In some cases, the unique tags can be loaded such that less than about 2, less than about 3, less than about 4, less than about 5, less than about 6, less than about 7, less than about 8, less than about 9, less than about 10, less than about 20, less than about 50, less than about 100, less than about 500, less than about 1000, less than about 5000, less than about 10000, less than about 50,000, less than about 100,000, less than about 500,000, less than about 1,000,000, less than about 10,000,000, less than about 50,000,000, or less than about 1,000,000,000 unique tags are loaded per genomic sample.In some cases, the average number of unique tags loaded per sample genome is less than about 1, less than about 2, less than about 3, less than about 4, less than about 5, less than about 6, less than about 7, less than about 8, less than about 9, less than about 10, less than about 20, less than about 50, less than about 100, less than about 500, less than about 1000, less than about 5000, less than about 10000, less than about 50,000, less than about 100,000, less than about 500,000, less than about 1,000,000, less than about 10,000,000, less than about 50,000,000 or less than about 1,000,000,000 per genome sample, or more than about 1, more than about 2, more than about 3, more than about 4, more than about 5, more than about 6, more than about 7, more than about 8, more than about 9, more than about 10, more than about 20, more than about 50, more than about 100, more than about 500, more than about 1000, more than about 5000, more than about 10000, more than about 50,000, more than about 100,000, more than about 500,000, more than about 1,000,000, more than about 10,000,000, more than about 50,000,000 or more than about 1,000,000,000 unique tags.
[0201] In a preferred format, 20 - 50 different tags (e.g., barcodes) ligated to both ends of the target nucleic acid are used. For example, 35 different tags (e.g., barcodes) ligated to both ends of the target molecule generate a 35×35 permutation, which is equal to 1225 for 35 tags. Such a number of tags is sufficient so that different molecules having the same starting and ending points receive different tag combinations with a high probability (e.g., at least 94%, 99.5%, 99.99%, 99.999%). Other barcode combinations include any number between 10 - 500, such as about 15×15, about 35×35, about 75×75, about 100×100, about 250×250, about 500×500.
[0202] In some cases, the unique tag may be an oligonucleotide of a predetermined or random or semi-random sequence. In other cases, multiple barcodes can be used, and thus the barcodes do not necessarily have to be unique to each other within the multiple. In this example, the barcodes can be ligated to individual molecules, and thus a unique sequence that can be individually tracked is generated by the combination of the barcode and the sequence to which it can be ligated. As described herein, the detection of a combination of non-unique barcodes and sequence data at the beginning (start) and end (stop) portions of the sequence read may enable the assignment of a unique identity to a particular molecule. The base pair length or number of individual sequence reads can also be used to assign a unique identity to such molecules. As described herein, a unique identity can be assigned to a fragment from a single strand of nucleic acid, thereby enabling subsequent identification of the fragment from the parental strand.
[0203] In some embodiments, two or more populations, samples, subsamples, or aliquots are differentially tagged. To associate a tag (or tags) with a particular population or aliquot, the tags can be used to label individual DNA populations. In some embodiments, a single tag can be used to label a particular population or aliquot. In some embodiments, multiple different tags can be used to label a particular population or aliquot. In embodiments using multiple different tags to label a particular aliquot, the set of tags used to label one aliquot can be readily distinguishable from the set of tags used to label other aliquots. In some embodiments, the tags may have additional functionality, for example, the tags can be used to index the sample source, or can be used as unique molecular identifiers (e.g., as described in Kinde et al., Proc Nat'l Acad Sci USA 108: 9530-9535 (2011), Kou et al., PLoS ONE, 11: e0146638 (2016), to improve the quality of sequencing data by differentiating sequencing errors from mutations), or can be used as non-unique molecular identifiers, for example, as described in U.S. Patent No. 9,598,731. Similarly, in some embodiments, the tags may have additional functionality, for example, the tags can be used to index the sample source, or can be used as non-unique molecular identifiers (which can be used to improve the quality of sequencing data by differentiating sequencing errors from mutations).
[0204] In some embodiments, partitioning tagging involves tagging the molecules within each partition with a partitioning tag. After re - combining the partitions (e.g., to reduce the number of sequencing runs required and avoid unnecessary costs) and sequencing the molecules, the partitioning tag is used to identify the original partition. In another embodiment, different partitions are tagged with different sets of molecular tags, for example, pairs of barcodes. Thus, each molecular barcode is useful not only for indicating the original partition but also for distinguishing the molecules within the partition. For example, a first set of 35 barcodes can be used to tag the molecules within a first partition, while a second set of 35 barcodes can be used to tag the molecules within a second partition.
[0205] In some embodiments, after tagging, the molecules can be pooled for sequencing in a single run. In some embodiments, a sample tag is added to the molecules, for example, in a step after the addition and pooling of other tags. The sample tag can facilitate pooling the materials generated from multiple samples for sequencing in a single sequencing run.
[0206] In some embodiments, a partitioning tag can be associated with a sample as well as a partition. As a simple example, a first tag can indicate the first partition of a first sample, a second tag can indicate the second partition of the first sample, a third tag can indicate the first partition of a second sample, and a fourth tag can indicate the second partition of the second sample.
[0207] Tags can be attached to molecules based on one or more characteristics, but the final tagged molecules in the library may no longer have those characteristics. For example, single-stranded DNA molecules can be dispensed and / or tagged, but the final tagged molecules in the library may be double-stranded. Similarly, DNA can be subjected to dispensing based on different levels of methylation, but in the final library, the tagged molecules derived from these molecules may be unmethylated. Thus, tags attached to molecules in the library generally indicate the characteristics of the "parent molecule" from which the final tagged molecule is derived, and do not necessarily indicate the characteristics of the tagged molecule itself.
[0208] As an example, use barcodes 1, 2, 3, 4, etc. to tag and label the molecules in the first dispensing, use barcodes A, B, C, D, etc. to tag and label the molecules in the second dispensing, and use barcodes a, b, c, d, etc. to tag and label the molecules in the third dispensing. The differentially tagged dispensings can be pooled prior to sequencing. The differentially tagged dispensings can also be sequenced separately or together, for example, in the same flow cell of an Illumina sequencer.
[0209] After sequencing, the reads can be analyzed at the level of each dispensing as well as at the level of the pooled DNA. Tags are used to sort the reads from different dispensings. The analysis can include in silico analysis to determine genetic diversity and epigenetic diversity (one or more of methylation, chromatin structure, etc.) using sequence information, length of genomic coordinates, coverage, and / or copy number. E. Amplification
[0210] In some embodiments, the DNA is amplified. For example, the DNA sandwiched by adapters added to the DNA as described herein can be amplified by PCR or other amplification methods. The amplification methods used herein can include any suitable methods such as those known to those skilled in the art. In some embodiments, the amplification is primed by a primer that binds to a primer-binding site within an adapter adjacent to the DNA molecule to be amplified. The amplification method may involve cycles of denaturation, annealing, and extension obtained from thermocycling, such as polymerase chain reaction (PCR), and may be isothermal, such as in linear amplification methods, transcription-mediated amplification, recombinase polymerase amplification (RPA), helix-dependent amplification (HDA), loop-mediated isothermal amplification (LAMP) (Notomi et al., Nuc. Acids Res., 28, e63, 2000), rolling circle amplification (RCA) (Blanco et al., J. Biol. Chem., 264, 8935-8940, 1989), or hyperbranched rolling circle amplification (Lizard et al., Nat. Genetics, 19, 225-232, 1998). Other amplification methods include ligase chain reaction, strand displacement amplification, nucleic acid sequence-based amplification, and self-sustained sequence-based replication.
[0211] In some embodiments, the step of detecting the presence or absence of one or more DNA sequences includes embodiments that involve amplification, such as qPCR or digital PCR. Some such embodiments that involve targeted detection of DNA sequences using qPCR or digital PCR do not include standard DNA library preparation steps such as adapter ligation or tagging. In some embodiments, the qPCR may be library-wide qPCR (i.e., not targeted PCR). In such embodiments, standard DNA library preparation steps such as adapter ligation or tagging are performed before the step of contacting the DNA with a methylation-sensitive nuclease (e.g., the step of contacting the DNA with an MSRE) and / or before the step of sequence-specific degradation (e.g., using depletion by CRISPR-Cas9).
[0212] In some embodiments, dsDNA ligation using a T-tailed adapter and a C-tailed adapter that results in amplification of at least 50%, 60%, 70%, or 80% of the double-stranded nucleic acid can be performed before ligation to the adapter. F. Detecting; Sequencing
[0213] In some embodiments, detecting the presence or absence of a DNA sequence involves sequencing. Generally, a sample nucleic acid, including nucleic acids flanked by adapters, can be subjected to sequencing, with or without prior amplification. Examples of sequencing methods include, for example, Sanger sequencing, high-throughput sequencing, pyrosequencing, sequencing by synthesis, single molecule sequencing, nanopore sequencing, semiconductor sequencing, sequencing by ligation, sequencing by hybridization, Digital Gene Expression (Helicos), next-generation sequencing (NGS), Single Molecule Sequencing by Synthesis (SMSS) (Helicos), enzymatic methyl sequencing (EM-Seq), Tet-assisted pyridine borane sequencing (TAPS), massively parallel sequencing, Clonal Single Molecule Array (Solexa), shotgun sequencing, Ion Torrent, Oxford Nanopore, Roche Genia, Maxam-Gilbert sequencing, primer walking, and sequencing using the PacBio, SOLiD, Ion Torrent, or Nanopore platform. The sequencing reaction can be carried out in various sample processing units that may include means for substantially simultaneously processing multiple lanes, multiple channels, multiple wells, or other multiple sets of samples. The sample processing unit may also include multiple sample chambers that allow for the simultaneous processing of multiple runs.
[0214] In some embodiments, the DNA is sequenced in a modification-sensitive manner (i.e., detecting and / or distinguishing between unmodified and modified nucleobases). For example, as a long-read sequencing (also referred to herein as third-generation sequencing) method, longer sequencing reads can be generated as compared to short-read sequencing methods that generally result in reads up to about 600 bases in length, for example, reads exceeding 10 kilobases. In the case of long reads, de novo assembly, identification of transcript isoforms, and detection and / or mapping of structural variants can be improved as compared to short reads. Further, long-read sequencing of native DNA or RNA molecules reduces amplification bias and preserves base modifications such as methylation status. Long-read sequencing techniques useful herein include, but are not limited to, Pacific Biosciences (PacBio) single molecule real-time (SMRT) sequencing, Oxford Nanopore Technologies (ONT) nanopore sequencing, and synthetic long-read sequencing approaches, such as any suitable long-read sequencing method including ligation reads, proximity ligation strategies, and optical mapping. Synthetic long-read approaches involve assembling short reads from the same DNA molecule to generate synthetic long reads and can be used in conjunction with "true" long-read sequencing techniques such as SMRT and nanopore sequencing methods.
[0215] Single molecule real-time (SMRT) sequencing facilitates the direct detection of, for example, 5-methylcytosine and 5-hydroxymethylcytosine as well as unmodified cytosine. See, for example, Schatz., Nature Methods. 14 (4): 347-348 (2017); Weirather JL, et al., F1000Research, 6: 100, 2017; and US9,150,918. Next-generation sequencing methods detect increased signals from a clonal population of amplified DNA fragments, whereas in SMRT sequencing, a single DNA molecule is captured and base modifications are maintained during sequencing. Since the signal-to-noise ratio from a single DNA molecule is not high, the error rate of raw data generated by PacBio SMRT sequencing is approximately 13-15%. To increase accuracy, circular DNA templates are used in this platform by ligating hairpin adapters to both ends of the target double-stranded DNA. Since the polymerase traverses and replicates the circular molecule repeatedly, the DNA template is sequenced multiple times to generate continuous long reads (CLRs). The CLRs can be split into multiple reads ("subreads") by removing the adapter sequences, and circular consensus sequence ("CCS") reads are generated with higher accuracy from the multiple subreads. The average length of the CLRs exceeds 10 kb and can reach up to 60 kb, and the length depends on the lifetime of the polymerase. Thus, the length and accuracy of the CCS reads depend on the fragment size. PacBio sequencing is utilized in genomic research (e.g., de novo assembly, detection of structural variants, and haplotype determination) and transcriptome research (e.g., gene isoform reconstruction and discovery of novel genes / isoforms).
[0216] ONT is a nanopore-based single-molecule sequencing technology (Weirather JL, et al., F1000Research, 6: 100, 2017). ONT directly sequences native single-stranded DNA (ssDNA) molecules by measuring the characteristic current changes as bases pass through the nanopore with a molecular motor protein. In ONT, a hairpin library structure similar to the PacBio circular DNA template, where the DNA template and its complement are joined by a hairpin adapter, is used. Thus, the DNA template, followed by the hairpin, and finally the complement pass through the nanopore. By removing the adapter, the raw reads can be split into two "1D" reads ("template" and "complement"). The consensus sequence of the two "1D" reads is a more accurate "2D" read.
[0217] As 5-letter and 6-letter sequencing methods, in addition to 5mC and 5hmC, A, C, T, and G are sequenced to obtain 5-letter (either A, C, T, G, and either 5mC or 5hmC) or 6-letter (A, C, T, G, 5mC, and 5hmC) digital readout information in a single workflow, including whole-genome sequencing methods. The processing of DNA samples is entirely enzymatic, avoiding DNA degradation and genomic coverage bias in bisulfite treatment. In an exemplary 5-letter sequencing method developed by Cambridge Epigenetix, sample DNA is first fragmented by sonication and then ligated to short synthetic DNA hairpin adapters at both ends (Fuellgrabe, et al. 2022, bioRxiv doi: https: / / doi.org / 10.1101 / 2022.07.08.499285). The construct is then split to separate the sense and antisense sample strands. For each original sample strand, a complementary copy strand is synthesized by DNA polymerase extension at the 3' end, generating a hairpin construct in which the original sample DNA strand and its complementary strand lacking epigenetic modifications are connected via a synthetic loop. Next, sequencing adapters are ligated to the ends. Modified cytosines are enzymatically protected. Then, unprotected C is deaminated to uracil, which is subsequently read as thymine. The deaminated construct is no longer fully complementary, and the duplex stability is substantially reduced, so the hairpin can be easily opened and amplified by PCR. The construct can be sequenced in a paired-end format where read 1 (primed by P1) is the original strand and read 2 (primed by P2) is the copy strand. The read data is pairwise aligned such that read 1 is aligned to its complementary read 2. Homologous residues from both reads are computationally resolved to yield a single genetic or epigenetic letter. The formation of five different homologous base pairs that are allowed is the result of imperfect fidelity at several stages, including sample preparation and amplification, or incorrect base calls during sequencing.Since these errors occur independently for the same type of bases on each strand, substitutions result in unacceptable pairs. Unacceptable pairs are masked (marked as N) within the elucidated reads, the reads themselves are retained, resulting in minimal information loss and high accuracy at the read level. The elucidated reads are aligned against the reference genome. Counts of genetic variants and methylation are generated by read counts at the base level.
[0218] 5hmC has been shown to be valuable as a marker of biological states and diseases, including early cancer detection from cell-free DNA. In adapting 5-letter sequencing to 6-letter sequencing, the ambiguity between 5mC and 5hmC is resolved without compromising genetic base calling within the same sample fragment. The first three steps of the workflow that generate synthetic copy strands ligated to adapter sample fragments are identical to the 5-letter sequencing described above. Methylation at 5mC over CpG units is enzymatically copied to C on the copy strand, while 5hmC is enzymatically protected from such copying. Thus, unmodified C, 5mC, and 5hmC in each of the original CpG units are distinguished by unique two-base combinations. Next, unmodified cytosine is deaminated to uracil, which is then read as thymine. The DNA is subjected to PCR amplification and sequencing as previously described. Reads are pairwise aligned and decoded using the two-base code. Since the three CpG units are distinct sequencing contexts of the two-base code, each of unmodified C, 5mC, and 5hmC can be decoded.
[0219] In some embodiments, sequencing includes targeted sequencing in which one or more target genomic regions are sequenced. In some such embodiments, the target genomic region includes a region presenting one or more genes selected from Tables 1, 2, 3, 4, and / or 5. In some such embodiments, DNA sequences that do not include the target region are not sequenced. Some embodiments include non-targeted sequencing, for example, all genomic regions of DNA in a processed sample or a secondary sample are sequenced, or the genomic regions to be sequenced are randomly selected. In other embodiments, the step of detecting the presence or absence of DNA sequences includes sequencing DNA that is not enriched for the target genomic region (non-targeted sequencing), for example, detectable sequences are obtained in a substantially unbiased manner.
[0220] In some embodiments, the sequencing step is performed on a library comprising one or more captured target region sets, which may include any one or more of the target region sets described herein. In some embodiments, the sequencing step is performed on a library comprising a secondary sample that has not undergone a capture / enrichment step (e.g., a whole genome secondary sample). In some embodiments, the target regions may be captured from a first secondary sample and a second secondary sample and then sequenced. In other embodiments, the target regions are captured from a first secondary sample and, for example, after contacting the DNA of the first sample with a methylation-sensitive nuclease and after attaching one or more tags to the DNA of the first and / or second secondary samples, may be combined with the second secondary sample after processing. In yet other embodiments, the target regions are captured from a second secondary sample and, for example, after contacting the DNA of the first sample with a methylation-sensitive nuclease and after attaching one or more tags to the DNA of the first and / or second secondary samples, may be combined with the first secondary sample after processing. In other embodiments, both the first secondary sample and the second secondary sample are processed (e.g., tagged with one or more tags) and may be combined without undergoing capture / enrichment.
[0221] The sequencing reaction can be performed on one or more forms of nucleic acid, at least one of which is known to contain a marker for cancer or other disease. The sequencing reaction can also be performed on any nucleic acid fragment present in the sample. In some embodiments, the genomic sequencing coverage can be less than 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, 99.9% or less than 100%. In some embodiments, the sequencing reaction can result in a sequencing coverage of at least 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, or 80% of the genome. The sequencing coverage can be performed on at least 5, 10, 20, 70, 100, 200 or 500 different genes, or up to 5000, 2500, 1000, 500 or 100 different genes.
[0222] Multiplex sequencing can be used to perform simultaneous sequencing reactions. In some cases, the cell-free nucleic acid can be sequenced in at least 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, 100,000 sequencing reactions. In other cases, the cell-free nucleic acid can be sequenced in less than 1000, less than 2000, less than 3000, less than 4000, less than 5000, less than 6000, less than 7000, less than 8000, less than 9000, less than 10000, less than 50000, less than 100,000 sequencing reactions. The sequencing reactions can be performed sequentially or simultaneously. Subsequent data analysis can be performed on all or part of the sequencing reactions. In some cases, the data analysis can be performed on at least 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, 100,000 sequencing reactions. In other cases, the data analysis can be performed on less than 1000, less than 2000, less than 3000, less than 4000, less than 5000, less than 6000, less than 7000, less than 8000, less than 9000, less than 10000, less than 50000, less than 100,000 sequencing reactions. Exemplary read depths are 1000 to 50000 reads per locus (base).
[0223] In some embodiments, sequences that were not cleaved during the cleavage step are sequenced. In some embodiments, less than 50%, less than 40%, less than 30%, less than 20%, less than 10%, less than 5%, less than 4%, less than 3%, less than 2%, or less than 1% of the sequences cleaved during the cleavage step are sequenced.
[0224] In some embodiments, the sequencing step includes generating a plurality of sequencing reads and mapping the plurality of sequencing reads to one or more reference sequences (e.g., one or more human reference sequences) to generate mapped array reads. In some embodiments, at least a portion of the DNA, RNA, or cDNA generated from the RNA of at least the first and second secondary samples is sequenced within the same sequencing cell. G. Target Region Set
[0225] In some embodiments, certain genomic regions of interest are detected and / or enriched. The genomic regions of interest can include one or more target region sets. In some embodiments, the target region set includes diversities that are not commonly found in cfDNA from healthy subjects or in DNA obtained from healthy tissue regions. In some embodiments, the target region set includes diversities that are present in healthy cells but not normally present in sample types such as blood samples. In some embodiments, the diversities are present in abnormal cells (e.g., hyperplastic cells, metaplastic cells, or neoplastic cells). Exemplary target region sets include sequence variable target region sets and epigenetic target region sets. In some embodiments, the first secondary sample is enriched for one or more target region sets prior to the step of sequence-specifically digesting a plurality of DNA sequences in the first secondary sample that contain modifications and are commonly found in cell-free DNA. This can be useful in that, as a result of the enrichment procedure, if off-target sequences are present in the enriched DNA, the subsequent digestion step can then reduce the amount of off-target sequences or eliminate off-target sequences to preserve or maximize the efficiency of downstream steps such as sequencing.
[0226] In some embodiments, a first set of target regions is detected that includes at least epigenetically targetable regions. In some embodiments, the epigenetically targetable regions detected in the first secondary sample include hypermethylation variable target regions. In some embodiments, the hypermethylation variable target regions are CpG-containing regions that are unmethylated or have low methylation (e.g., methylation below the average compared to the entire cfDNA) in cfDNA from healthy subjects. In some embodiments, the hypermethylation variable target regions exhibit type-specific hypermethylation in healthy cfDNA from one or more associated cell or tissue types. Without being bound by any particular theory, the presence of cancer cells can increase the excretion of DNA into the bloodstream (e.g., from the cancer and / or surrounding tissues). Thus, the distribution of the origin tissues of cfDNA can change during carcinogenesis. Thus, an increase in the level of hypermethylation variable target regions in the first secondary sample can be an indicator of the presence of cancer (or recurrence depending on the subject's medical history).
[0227] In some embodiments, the methods herein include detecting a second captured set of target regions from a sample or a second secondary sample that includes at least epigenetically targetable regions. In some embodiments, the second set of epigenetically targetable regions includes hypomethylation variable target regions. In some embodiments, the hypomethylation variable target regions are CpG-containing regions that are methylated or have high methylation (e.g., methylation above the average compared to the entire cfDNA) in cfDNA from healthy subjects. Without being bound by any particular theory, cancer cells can excrete more DNA into the bloodstream than healthy cells of the same tissue type. Thus, the distribution of the origin tissues of cfDNA can change during carcinogenesis. Thus, an increase in the level of hypomethylation variable target regions in the second secondary sample can be an indicator of the presence of cancer (or recurrence depending on the subject's medical history).
[0228] Furthermore, the set of target regions can include DNA corresponding to a set of sequence variable target regions. 1. Set of Epigenetically Targetable Regions
[0229] In some embodiments, the set of target regions is or includes an epigenetic target region set. The epigenetic target region set can include one or more types of target regions that can potentially distinguish DNA from neoplastic (e.g., tumor or cancer) cells and from healthy cells, such as non-neoplastic circulating cells. Exemplary types of such regions are discussed in detail herein. The epigenetic target region set can also include, for example, one or more control regions as described herein.
[0230] In some embodiments, the epigenetic target region set has a footprint of at least 100 kb, such as at least 200 kb, at least 300 kb, or at least 400 kb. In some embodiments, the epigenetic target region set has a footprint in the range of 100 - 1000 kb, such as 100 - 200 kb, 200 - 300 kb, 300 - 400 kb, 400 - 500 kb, 500 - 600 kb, 600 - 700 kb, 700 - 800 kb, 800 - 900 kb, and 900 - 1,000 kb. a. Hypermethylated and hypomethylated variable target regions
[0231] In some embodiments, the epigenetic target region set includes hypermethylated variable target regions. In some embodiments, the hypermethylated variable target regions are differentially or exclusively hypermethylated in one or more associated cell or tissue types. Such hypermethylated variable target regions can be hypermethylated in other cell or tissue types as well, but not to the extent observed in one or more associated cell or tissue types. In some embodiments, the hypermethylated variable target regions exhibit even higher methylation in cfDNA derived from diseased cells of one or more associated cell or tissue types.
[0232] A comprehensive review of methylation variable target regions in colorectal cancer is provided by Lam et al., Biochim Biophys Acta. 1866: 106-20 (2016). These include VIM, SEPT9, ITGA4, OSM4, GATA4, and NDRG4. An exemplary set of hypermethylated variable target regions based on colorectal cancer (CRC) studies is provided in Table 1. Many of these genes may have relevance not only to colorectal cancer but also to other cancers. For example, TP53 is widely recognized as a very important tumor suppressor, and inactivation based on hypermethylation of this gene may be a common tumorigenic mechanism. Table 1. Exemplary hypermethylated target regions based on CRC studies.
Table 1-1
Table 1-2
[0233] In some embodiments, the genomic regions targeted for sequencing include a plurality of loci listed in Table 1, e.g., at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% of the loci listed in Table 1. In some embodiments, the genomic regions are captured using probes. For example, for each locus included as a target region, there may be one or more probes having hybridization sites that bind between the transcription start site and the stop codon of the gene (the last stop codon of alternatively spliced genes) or within the promoter region of the gene. In some embodiments, one or more probes bind within 300 bp, e.g., within 200 bp or 100 bp, from the transcription start site of the genes in Table 1.
[0234] Methylation variable target regions in various types of lung cancer have been discussed in detail, for example, in Ooki et al., Clin. Cancer Res. 23: 7141-52 (2017); Belinksy, Annu. Rev. Physiol. 77: 453-74 (2015); Hulbert et al., Clin. Cancer Res. 23: 1998-2005 (2017); Shi et al., BMC Genomics 18: 901 (2017); Schneider et al., BMC Cancer. 11: 102 (2011); Lissa et al., Transl Lung Cancer Res 5 (5): 492-504 (2016); Skvortsova et al., Br. J. Cancer. 94 (10): 1492-1495 (2006); Kim et al., Cancer Res. 61: 3419-3424 (2001); Furonaka et al., Pathology International 55: 303-309 (2005); Gomes et al., Rev. Port. Pneumol. 20: 20-30 (2014); Kim et al., Oncogene. 20: 1765-70 (2001); Hopkins-Donaldson et al., Cell Death Differ. 10: 356-64 (2003); Kikuchi et al., Clin. Cancer Res. 11: 2954-61 (2005); Heller et al., Oncogene 25: 959-968 (2006); Licchesi et al., Carcinogenesis. 29: 895-904 (2008); Guo et al., Clin. Cancer Res. 10: 7917-24 (2004); Palmisano et al., Cancer Res. 63: 4620-4625 (2003); and Toyooka et al., Cancer Res. 61: 4556-4560, (2001).
[0235] Table 2 provides an exemplary set of hypermethylated variable target regions based on lung cancer research. Many of these genes may have relevance not only to lung cancer but also to other cancers. For example, Casp8 (caspase 8) is an important enzyme in programmed cell death, and inactivation based on hypermethylation of this gene may be a common tumorigenic mechanism not limited to lung cancer. Furthermore, some genes appear in both Table 1 and Table 2, thereby indicating generality. Table 2. Exemplary Hypermethylated Target Regions Based on Lung Cancer Research [Table 2]
[0236] Any of the above-described embodiments regarding the target regions identified in Table 2 can be combined with any of the above-described embodiments regarding the target regions identified in Table 1. In some embodiments, the genomic regions targeted for sequencing include a plurality of loci listed in Table 1 or Table 2, for example, at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% of the loci listed in Table 1 or Table 2.
[0237] Additional hypermethylated target regions can be obtained from, for example, the Cancer Genome Atlas described in CancerLocator, a probabilistic method constructed using hypermethylated target regions from breast, colon, kidney, liver, and lung, as described in Kang et al., Genome Biology 18: 53 (2017). In some embodiments, the hypermethylated target regions can be specific to one or more types of cancer. Thus, in some embodiments, the hypermethylated target regions include one, two, three, four, or five subsets of hypermethylated target regions that collectively exhibit hypermethylation in one, two, three, four, or five of breast cancer, colon cancer, kidney cancer, liver cancer, and lung cancer.
[0238] In some embodiments, the set of epigenetic target regions includes hypomethylated variable target regions. In some embodiments, the hypomethylated variable target regions are hypomethylated exclusively in one or more related cell or tissue types. Such hypomethylated variable target regions may be hypomethylated in other cell or tissue types as well, but not to the extent observed in one or more related cell or tissue types.
[0239] In some embodiments, when capturing different epigenetic target regions, the epigenetic target regions include hypermethylated and / or hypomethylated variable target regions.
[0240] Additional exemplary hypermethylated variable target regions and hypomethylated variable target regions useful for distinguishing between various cell types have been identified, for example, as described in Scott, C.A., Duryea, J.D., MacKay, H. et al., "Identification of cell type-specific methylation signals in bulk whole genome bisulfite sequencing data," Genome Biol 21, 156 (2020)(doi.org / 10.1186 / s13059-020-02065-5), by analyzing DNA obtained from various cell types by whole genome bisulfite sequencing. Whole genome bisulfite sequencing data is available from the Blueprint consortium, which is available online at dcc.blueprint-epigenome.eu. b. CTCF binding regions
[0241] In some embodiments, the epigenetic target region set includes CTCF-binding regions. CTCF is a DNA-binding protein that contributes to chromatin organization and often co-localizes with cohesin. Disruptions of CTCF-binding sites have been reported in various different cancers. See, for example, Katainen et al., Nature Genetics, doi: 10.1038 / ng.3335, published online on June 8, 2015; Guo et al., Nat. Commun. 9: 1520 (2018). CTCF binding results in a recognizable pattern of cfDNA that can be detected, for example, by sequencing using fragment length analysis. Thus, disruption of CTCF binding results in a diversity of cfDNA fragmentation patterns. Thus, CTCF-binding sites are one type of fragmentation variable target region.
[0242] There are many known CTCF-binding sites. See, for example, the CTCFBSDB (CTCF Binding Site Database) available at insulatordb.uthsc.edu / on the Internet, each of which is incorporated by reference; Cuddapah et al., Genome Res. 19: 24-32 (2009); Martin et al., Nat. Struct. Mol. Biol. 18: 708-14 (2011); Rhee et al., Cell. 147: 1408-19 (2011). Exemplary CTCF-binding sites are nucleotides 56014955-56016161 on chromosome 8 and nucleotides 95359169-95360473 on chromosome 13.
[0243] In some embodiments, the CTCF-binding regions include at least 10, 20, 50, 100, 200, or 500 CTCF-binding regions, or 10-20, 20-50, 50-100, 100-200, 200-500, or 500-1000 CTCF-binding regions, such as, for example, the CTCF-binding regions in one or more of the above or CTCFBSDB or the papers by Cuddapah et al., Martin et al., or Rhee et al. cited above. In some embodiments, at least a portion of the CTCF sites can be methylated or unmethylated, where the methylation state is associated with whether the cell is a cancer cell. In some embodiments, the set of epigenetic target regions includes regions at least 100 bp, at least 200 bp, at least 300 bp, at least 400 bp, at least 500 bp, at least 750 bp, at least 1000 bp upstream and downstream of the CTCF-binding sites. c. Transcription start site
[0244] In some embodiments, the set of epigenetic target regions includes variable transcription start sites. The transcription start sites can exhibit perturbations in nascent cells. For example, the nucleosome organization at various transcription start sites in healthy cells of the hematopoietic lineage that substantially contribute to cfDNA in a healthy individual can be different from the nucleosome organization at transcription start sites in nascent cells. Thereby, as generally discussed in Snyder et al., Cell 164: 57-68 (2016); WO2018 / 009723; and US20170211143A1, different cfDNA patterns that can be detected by sequencing are brought about. In another example, the transcription start sites may not necessarily be epigenetically different in cancerous tissue compared to DNA from healthy tissue of the same type, but may be epigenetically different (e.g., with respect to nucleosome organization) compared to cfDNA that is typical in healthy subjects. Perturbations of the transcription start sites also result in diversity in the fragmentation pattern of cfDNA. Thus, transcription start sites are also one type of fragmentation variable target region.
[0245] The human transcriptional start site is available from the DBTSS (Database of Transcriptional Start Sites in Humans) described in Yamashita et al., Nucleic Acids Res. 34 (Database issue): D86 - D89 (2006), which is available on the Internet at dbtss.hgc.jp and is incorporated herein by reference. In some embodiments, the transcriptional start site can be at least 10, 20, 50, 100, 200, or 500 transcriptional start sites, or ranges of 10 - 20, 20 - 50, 50 - 100, 100 - 200, 200 - 500, or 500 - 1000 transcriptional start sites, such as those listed in the DBTSS. In some embodiments, at least a portion of the transcriptional start site can be methylated or unmethylated, where the methylation state is associated with whether the cell is a cancer cell. In some embodiments, the set of epigenetic target regions includes regions at least 100 bp, at least 200 bp, at least 300 bp, at least 400 bp, at least 500 bp, at least 750 bp, at least 1000 bp upstream and downstream of the transcriptional start site. d. Local amplification
[0246] Local amplification is a somatic mutation but can be detected in a manner similar to techniques for detecting certain epigenetic changes such as changes in methylation based on read frequency by sequencing. Thus, regions that can exhibit local amplification in cancer can be included in the set of epigenetic target regions and can include one or more of AR, BRAF, CCND1, CCND2, CCNE1, CDK4, CDK6, EGFR, ERBB2, FGFR1, FGFR2, KIT, KRAS, MET, MYC, PDGFRA, PIK3CA, and RAF1. e. Methylation control region
[0247] To facilitate data verification, it may be useful to include a control region. In some embodiments, the set of epigenetic target regions includes control regions that are expected to be methylated or unmethylated in essentially all samples regardless of whether the DNA is from cancer cells or normal cells. In some embodiments, the set of epigenetic target regions includes control hypomethylated regions that are expected to be hypomethylated in essentially all samples. In some embodiments, the set of epigenetic target regions includes control hypermethylated regions that are expected to be hypermethylated in essentially all samples. 2. Set of sequence-variable target regions
[0248] In some embodiments, the set of target regions is or includes a set of variable target regions. The set of variable target regions can include one or more types of target regions that may be capable of discriminating DNA from neoplastic (e.g., tumor or cancer) cells and from healthy cells, such as non-neoplastic circulating cells. Exemplary types of such regions are discussed in detail herein. The set of variable target regions can also include, for example, one or more control regions as described herein. In some embodiments, the set of variable target regions includes a plurality of regions known to undergo somatic mutations in cancer. In some aspects, the set of variable target regions targets a plurality of different genes or genomic regions (a “panel”) selected such that a determined percentage of subjects with cancer exhibit genetic variants or tumor markers in one or more different genes or genomic regions within the panel. The panel can be selected such that the region to be sequenced is limited to a fixed number of base pairs. The panel can be selected such that a desired amount of DNA is sequenced. The panel can further be selected such that a desired sequence read depth is achieved. The panel can be selected such that a desired sequence read depth or sequence read coverage is achieved for a given amount of base pairs to be sequenced. The panel can be selected such that a theoretical sensitivity, theoretical specificity, and / or theoretical accuracy for detecting one or more genetic variants in a sample is achieved.
[0249] Examples of lists of genomic locations of interest can be found, for example, in Tables 3 and 4 of this specification. In some embodiments, the set of array variable target regions comprises at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, or 70 portions of the genes in Table 3. In some embodiments, the set of array variable target regions comprises at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, or 70 of the SNVs in Table 3. In some embodiments, the set of array variable target regions comprises at least 1, at least 2, at least 3, at least 4, at least 5, or 6 portions of the fusions in Table 3. In some embodiments, the set of array variable target regions comprises at least 1, at least 2, or at least 3 portions of the indels in Table 3. In some embodiments, the set of array variable target regions comprises at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, or 73 portions of the genes in Table 4. In some embodiments, the set of array variable target regions comprises at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, or 73 of the SNVs in Table 4. In some embodiments, the set of array variable target regions comprises at least 1, at least 2, at least 3, at least 4, at least 5, or 6 portions of the fusions in Table 4.In some embodiments, the set of array-variable target regions includes at least a portion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, or 18 of the indels in Table 4. Each of these genomic locations of interest can be identified as a backbone region or a hotspot region for a given panel. Table 5 shows an example list of the genomic locations of the target hotspot. In some embodiments, the set of array-variable target regions includes at least a portion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, or at least 20 of the genes in Table 5. Each hotspot genomic region is listed along with the associated gene of the given genomic region of interest, the chromosome on which it is located, the start and stop positions of the genome representing the locus of the gene, the length of the locus of the gene in base pairs, the exons encompassed by the gene, as well as highly important features (e.g., type of mutation). Table 3
Table 3
Table 4
Table 5-1
Table 5-2
Table 5-3
[0250] Examples of a list of target regions of interest can also be found, for example, in Table 4 of WO2020 / 160414. As an additional example, a locus disclosed in Gale et al., PLoS one 13: e0194630 (2018), which is incorporated herein by reference and describes a group of 35 cancer-related gene targets: AKT1, ALK, BRAF, CCND1, CDK2A, CTNNB1, EGFR, ERBB2, ESR1, FGFR1, FGFR2, FGFR3, FOXL2, GATA3, GNA11, GNAQ, GNAS, HRAS, IDH1, IDH2, KIT, KRAS, MED12, MET, MYC, NFE2L2, NRAS, PDGFRA, PIK3CA, PPP2R1A, PTEN, RET, STK11, TP53, and U2AF1 is provided. In some embodiments, the set of sequence-variable target regions includes target regions from at least 10, 20, 30, or 35 cancer-related genes, such as the cancer-related genes listed herein and in WO2020 / 160414.
[0251] In some embodiments, the set of sequence-variable target regions has a footprint of at least 50 kbp, for example, at least 100 kbp, at least 200 kbp, at least 300 kbp, or at least 400 kbp. In some embodiments, the set of sequence-variable target regions has a footprint in the range of 100 - 2000 kbp, for example, 100 - 200 kbp, 200 - 300 kbp, 300 - 400 kbp, 400 - 500 kbp, 500 - 600 kbp, 600 - 700 kbp, 700 - 800 kbp, 800 - 900 kbp, 900 - 1,000 kbp, 1 - 1.5 Mbp or 1.5 - 2 Mbp. In some embodiments, the set of sequence-variable target regions has a footprint of at least 2 Mbp. H. Subject
[0252] In some embodiments, DNA (e.g., cfDNA) is obtained from a subject (e.g., a test subject) having cancer or a pre-cancerous condition. In some embodiments, the subject has stage I cancer, stage II cancer, stage III cancer, or stage IV cancer. In some embodiments, the DNA derived from the subject is obtained from and / or derived from a sample obtained from the subject. In some embodiments, the DNA is obtained from a subject suspected of having cancer or a pre-cancerous condition. In some embodiments, the DNA is obtained from a subject having a tumor. In some embodiments, the DNA is obtained from a subject suspected of having a tumor. In some embodiments, the DNA is obtained from a subject having a neoplasm. In some embodiments, the DNA is obtained from a subject suspected of having a neoplasm. In some embodiments, the DNA is obtained from a subject in a state of remission from a tumor, cancer, or neoplasm (e.g., after chemotherapy, surgical resection, radiation, or a combination thereof). In any of the above embodiments, the pre-cancerous condition, cancer, tumor, or neoplasm or suspected pre-cancerous condition, cancer, tumor, or neoplasm can be of the bladder, head and neck, lung, colon, rectum, kidney, breast, prostate, skin, or liver. In some embodiments, the pre-cancerous condition, cancer, tumor, or neoplasm or suspected pre-cancerous condition, cancer, tumor, or neoplasm is of the lung. In some embodiments, the pre-cancerous condition, cancer, tumor, or neoplasm or suspected pre-cancerous condition, cancer, tumor, or neoplasm is of the colon or rectum. In some embodiments, the pre-cancerous condition, cancer, tumor, or neoplasm or suspected pre-cancerous condition, cancer, tumor, or neoplasm is of the breast. In some embodiments, the pre-cancerous condition, cancer, tumor, or neoplasm or suspected pre-cancerous condition, cancer, tumor, or neoplasm is of the prostate. In any of the above embodiments, the subject can be a human subject. In any of the above embodiments, the subject can be a test subject. I. Sample
[0253] The sample may be any biological sample isolated from a subject. The sample may be a bodily sample. The sample may be a body tissue or a body fluid, such as a known or suspected solid tumor, whole blood, platelets, serum, plasma, feces, red blood cells, white blood cells (white blood cell or leucocyte), endothelial cells, tissue biopsy material, cerebrospinal fluid, synovial fluid, lymph, ascites, interstitial fluid or extracellular fluid, fluid in the intercellular space, gingival crevicular fluid, bone marrow, pleural effusion, pleura fluid, cerebrospinal fluid, saliva, mucus, sputum, semen, sweat, and urine. The sample is preferably a body fluid, particularly blood and its fractions, cerebrospinal fluid, pleura fluid, saliva, sputum, or urine. The sample may be in the form originally isolated from the subject, or may be subjected to further processing to remove or add components such as cells or to enrich one component relative to another component.
[0254] In some embodiments, the population of nucleic acids is obtained from a serum, plasma, or blood sample from a subject suspected of having a neoplasm, tumor, pre-cancerous condition, or cancer, or previously diagnosed with a neoplasm, tumor, pre-cancerous condition, or cancer. The population includes nucleic acids having various levels of sequence diversity, epigenetic diversity, post-translational modification (PTM) of chromatin, and / or post-replicative or post-transcriptional modification. Post-replicative modifications include, in particular, modification of cytosine at the 5-position of the nucleobase, such as 5-methylcytosine, 5-hydroxymethylcytosine, 5-formylcytosine, and 5-carboxylcytosine.
[0255] A sample can be isolated or obtained from a subject and transported to a location for sample analysis. The sample can be stored and shipped at a desired temperature, such as room temperature, 4°C, -20°C, and / or -80°C. The sample can be isolated or obtained from the subject at the location for sample analysis. The subject can be a human, mammal, animal, companion animal, assistance animal, or pet animal. The subject can have cancer, a pre-cancerous condition, an infection, a transplant rejection reaction, or other disease or disorder related to changes in the immune system. The subject may not have cancer or detectable cancer symptoms. The subject may be treated with one or more cancer treatments, such as chemotherapy, antibodies, vaccines, or biologies. The subject may be in a state of remission. The subject may or may not be diagnosed as being susceptible to cancer or any cancer-related genetic mutation / disorder.
[0256] A reference molecule or control molecule can be added or spiked into the sample as a control or normalization standard. For example, a specific amount of modified DNA derived from a species other than the species of the subject from which the sample was obtained, or a synthetic nucleic acid containing a specific modification, can be added to the sample. In some embodiments, the reference molecule or control molecule is distinguishable from the molecules originally present in the sample. In some embodiments, the detected DNA sequences are normalized against the reference molecule or control molecule.
[0257] In some embodiments, the sample includes plasma. The volume of plasma obtained can depend on the desired read depth of the region to be sequenced. Exemplary volumes are 0.4 - 40 ml, 5 - 20 ml, 10 - 20 ml. For example, the volume can be 0.5 mL, 1 mL, 5 mL, 10 mL, 20 mL, 30 mL, or 40 mL. The volume of plasma collected for the sample can be 5 - 20 mL.
[0258] In some embodiments, a plurality of first secondary samples (e.g., derived from different subjects and / or tagged to be distinguishable by sample tags) are pooled before sequence-specific cleavage of a plurality of DNA sequences commonly found in cell-free DNA that contain modifications. This approach can reduce costs in that, for example, the reagents that may be required per secondary sample to be processed are reduced. J. Analysis
[0259] This method can be used to diagnose or classify a condition in a subject. In some embodiments, the condition is cancer or a pre-cancerous condition. In some embodiments, the response to treatment of a condition that is characterized (e.g., staging cancer or determining cancer heterogeneity), is monitored, or the prognostic risk of developing a condition or the subsequent course of a condition is determined. The present disclosure may also be useful in determining the effectiveness of a particular treatment option. In a successful treatment option, the amount of DNA sequences associated with cancer detected in a subject's blood may decrease because there may be fewer cancer cells shedding DNA. In other instances, this may not occur. In another example, a particular treatment option can be associated with the genetic profile of cancer over time. This correlation may be useful in treatment selection.
[0260] Furthermore, if cancer is observed to be in a remission state after treatment, this method can be used to monitor for residual disease or recurrence of the disease.
[0261] Cancer types and numbers that can be detected include blood cancer, brain cancer, lung cancer, skin cancer, nasal cancer, throat cancer, liver cancer, bone cancer, lymphoma, pancreatic cancer, skin cancer, intestinal cancer, rectal cancer, colon cancer, prostate cancer, thyroid cancer, bladder cancer, head and neck cancer, kidney cancer, oral cancer, stomach cancer, solid tumors, heterogeneous tumors, homogeneous tumors, and the like. The cancer type and / or stage can be detected from genetic diversity including mutations, rare mutations, indels, copy number polymorphisms, base conversions, translocations, recombinations, inversions, deletions, aneuploidy, partial aneuploidy, ploidy, chromosomal instability, chromosomal structural changes, gene fusions, chromosomal fusions, gene shortening, gene amplification, gene duplication, chromosomal damage, DNA damage, abnormal changes in nucleic acid chemical modifications, abnormal changes in epigenetic patterns, and abnormal changes in nucleic acid 5-methylcytosine.
[0262] In some embodiments, the methods described herein include detecting the presence or absence of nucleic acids, such as DNA, produced by a tumor (or neoplastic cells, or cancer cells) or by pre-cancerous state cells.
[0263] The information and data generated by the methods disclosed herein can also be used to characterize specific forms of cancer. Cancer is often heterogeneous both in terms of composition and staging. The methods disclosed herein can enable the characterization of specific subtypes of cancer that may be important in the diagnosis or treatment of specific subtypes. This information can also provide clues regarding the prognosis of a specific type of cancer to the subject or practitioner, and enable either the subject or practitioner to adapt treatment options to the progression of the disease. Some cancers can progress to a higher malignancy and become genetically unstable. Other cancers can remain benign, inactive, or in a quiescent state. The systems and methods of the present disclosure can be useful in determining disease progression.
[0264] Furthermore, the methods of the present disclosure can be used to characterize the heterogeneity of a condition in a subject. Such methods can include, for example, generating an aggregated profile of extracellular nucleic acids derived from the subject, the aggregated profile including a plurality of data obtained from various nucleic acid assays. In some embodiments, the aggregated profile includes epigenetic and mutational assays. In some embodiments, the aggregated profile includes a summary of information from different cells in a heterogeneous disease. This summary can include identity and levels of structural diversity, copy number polymorphisms, epigenetic diversity, or other mutational assays.
[0265] The methods can be used to diagnose, prognose, monitor, or observe a pre-cancerous condition, cancer, or other disease. In some embodiments, the methods herein do not involve diagnosing, prognosing, or monitoring a fetus and thus are not directed to non-invasive prenatal testing. In other embodiments, these methods can be used in a subject during pregnancy to diagnose, prognose, monitor, or observe cancer or other disease in a prenatal subject in which DNA and other polynucleotides may be co-circulating with maternal molecules.
[0266] Exemplary methods for analyzing DNA include the following steps: 1. Preparing an extracted DNA sample (e.g., extracting plasma DNA from a human sample). 2. Distributing the sample into a plurality of sub-samples based on the methylation state of the DNA. 3. Applying different tags and adapter sequences that enable NGS to each distribution. 4. Sequence-specifically degrading the DNA of at least one sub-sample using a modification-independent sequence-specific nuclease and a plurality of guide RNAs. 5. Amplifying the DNA using adapter-specific DNA primer sequences. 6. Sequencing the DNA on an NGS instrument. 7. A step of performing bioinformatics analysis of NGS data, the step comprising identifying unique molecules using tags and deconvoluting a sample into differentially distributed molecules.
[0267] In some embodiments, detecting the presence, level, or absence of a DNA sequence facilitates the diagnosis of a disease or the identification of an appropriate treatment. In some embodiments, the presence or a change in the level of one or more sequences indicates the presence of a disease or disorder in a subject, such as cancer or a pre-cancerous condition, or another disorder that causes a change in nucleic acids as compared to a healthy subject. III. Additional Features of Some of the Disclosed Methods A. A procedure that affects the first nucleobase of DNA in a different manner than the second nucleobase of DNA
[0268] In some embodiments, the methods disclosed herein include subjecting DNA to a procedure that affects the first nucleobase in DNA in a different manner than the second nucleobase in DNA, wherein the first nucleobase is a modified or unmodified nucleobase, the second nucleobase is a modified or unmodified nucleobase different from the first nucleobase, and the first and second nucleobases have the same base pairing specificity. In some embodiments, the procedure chemically converts the first or second nucleobase such that the base pairing specificity of the converted nucleobase is altered. In some embodiments, after the step of sequence-specifically degrading DNA, the DNA is subjected to a procedure that affects the first nucleobase in DNA in a different manner than the second nucleobase in DNA. Alternatively, the DNA is subjected to the procedure prior to sequence-specific degradation, where sequence-specific degradation includes the use of a plurality of guide RNAs, including a guide RNA configured to specifically bind to DNA containing the converted nucleobase.
[0269] In some embodiments, when the first nucleobase is a modified or unmodified adenine, the second nucleobase is a modified or unmodified adenine; when the first nucleobase is a modified or unmodified cytosine, the second nucleobase is a modified or unmodified cytosine; when the first nucleobase is a modified or unmodified guanine, the second nucleobase is a modified or unmodified guanine; and when the first nucleobase is a modified or unmodified thymine, the second nucleobase is a modified or unmodified thymine (for the purposes of this step, modified and unmodified uracils are included within modified thymine).
[0270] In some embodiments, the first nucleobase is a modified or unmodified cytosine, and thus the second nucleobase is a modified or unmodified cytosine. For example, the first nucleobase may include unmodified cytosine (C), and the second nucleobase may include one or more of 5-methylcytosine (mC) and 5-hydroxymethylcytosine (hmC). Alternatively, the second nucleobase may include C, and the first nucleobase may include one or more of mC and hmC. Other combinations are possible, such as when one of the first and second nucleobases includes mC and the other includes hmC.
[0271] In some embodiments, procedures that affect a first nucleobase in DNA in a different way than a second nucleobase in the DNA include bisulfite conversion. Performing bisulfite conversion can facilitate the identification of positions containing mC or hmC using sequence reads. Treatment with bisulfite converts unmodified cytosine and certain modified cytosine nucleotides (e.g., 5-formylcytosine (fC) or 5-carboxylcytosine (caC)) to uracil, while other modified cytosines (e.g., 5-methylcytosine, 5-hydroxymethylcytosine) are not converted. Thus, when using bisulfite conversion, the first nucleobase includes one or more of unmodified cytosine, 5-formylcytosine, 5-carboxylcytosine, or other cytosine forms affected by bisulfite, and the second nucleobase can include one or more of mC and hmC, e.g., mC, and optionally hmC. Sequencing bisulfite-treated DNA identifies positions read as cytosine as being mC positions or hmC positions. On the other hand, positions read as T are identified as T or C in a form susceptible to bisulfite, e.g., unmodified cytosine, 5-formylcytosine, or 5-carboxylcytosine. Thus, for example, performing bisulfite conversion on a DNA sample described herein facilitates the identification of positions containing mC or hmC using sequence reads obtained from an exemplary sample. For an exemplary description of bisulfite conversion, see, for example, Moss et al., Nat Commun. 2018;9: 5068.
[0272] In some embodiments, procedures that affect the first nucleobase in DNA in a different manner than the second nucleobase in DNA include oxidative bisulfite (Ox - BS) conversion. In this procedure, first, hmC is converted to fC, which is susceptible to bisulfite, and then bisulfite conversion is performed. Thus, when using oxidative bisulfite conversion, the first nucleobase includes one or more of unmodified cytosine, fC, caC, hmC, or other cytosine forms susceptible to bisulfite, and the second nucleobase includes mC. By sequencing the Ox - BS - converted DNA, positions read as cytosine are identified as mC positions. On the other hand, positions read as T are identified as T, hmC, or C in a form susceptible to bisulfite, such as unmodified cytosine, fC, or hmC. Thus, for example, by performing Ox - BS conversion on a DNA sample described herein, it becomes easy to identify positions containing mC using sequence reads obtained from the sample. For an exemplary description of oxidative bisulfite conversion, see, for example, Booth et al., Science 2012; 336: 934 - 937.
[0273] In some embodiments, a procedure that affects a first nucleobase in DNA in a different way than a second nucleobase in the DNA involves Tet-assisted bisulfite (TAB) conversion. In TAB conversion, hmC is protected from conversion, and mC is oxidized prior to bisulfite treatment, such that positions originally occupied by mC are converted to U, while positions originally occupied by hmC remain as cytosine in a protected form. For example, as described in Yu et al., Cell 2012; 149: 1368-80, β-glucosyltransferase can be used to protect hmC (forming 5-glucosylhydroxymethylcytosine (ghmC)), and then a TET protein such as mTet1 can be used to convert mC to caC, and then bisulfite treatment can be used to convert C and caC to U, while ghmC remains unaffected. Thus, when using TAB conversion, the first nucleobase includes one or more of unmodified cytosine, fC, caC, mC, or other cytosine forms that are affected by bisulfite, and the second nucleobase includes hmC. By sequencing the TAB-converted DNA, positions that are read as cytosine are identified as hmC positions. On the other hand, positions that are read as T are identified as T, mC, or a form of C that is susceptible to bisulfite effects, such as unmodified cytosine, fC, or caC. Thus, for example, by performing TAB conversion on a DNA sample described herein, it becomes easy to identify positions containing hmC using sequence reads obtained from the sample.
[0274] In some embodiments, the procedure for differentially affecting a first nucleobase in DNA from a second nucleobase in the DNA involves Tet-assisted conversion with a substituted borane reducing agent, which, optionally, is 2-picolyl borane, borane pyridine, tert-butylamine borane, or ammonia borane. In Tet-assisted pic-borane conversion using a substituted borane reducing agent conversion, TET protein is used to convert mC and hmC to caC, with unmodified C unaffected. caC, and fC if present, are then converted to dihydrouracil (DHU) by treatment with 2-picolyl borane (pic-borane) or another substituted borane reducing agent, such as borane pyridine, tert-butylamine borane, or ammonia borane, again with unmodified C unaffected. See, for example, Liu et al., Nature Biotechnology 2019; 37:424-429 (e.g., Supplementary Figure 1 and Supplementary Note 7). DHU is read as T in sequencing. Thus, when using this type of conversion, the first nucleobase includes one or more of mC, fC, caC, or hmC, and the second nucleobase includes unmodified cytosine. Sequencing the converted DNA identifies positions read as cytosine as unmodified C positions, while positions read as T are identified as T, mC, fC, caC, or hmC. Thus, performing TAP conversion on a DNA sample as described herein, for example, facilitates identifying positions containing unmodified C using sequence reads obtained from the sample. This procedure encompasses Tet-assisted pyridine borane sequencing (TAPS), described in more detail above in Liu et al. 2019.
[0275] In some embodiments, protection of hmC (e.g., using βGT) can be combined with Tet-assisted conversion using a substituted borane reducing agent. Performing TAPSβ conversion can facilitate distinguishing, using sequence reads, positions containing unmodified C or hmC on the one hand from positions containing mC. hmC can be protected by glucosylation using βGT as described above to form ghmC. Then, upon treatment with a TET protein such as mTet1, mC is converted to caC, but neither C nor ghmC is converted. Then, caC is converted to DHU by treatment with pic-borane or another substituted borane reducing agent, e.g., borane pyridine, tert-butylamine borane, or ammonia borane, and again neither unmodified C nor ghmC is affected. Thus, when using Tet-assisted conversion with a substituted borane reducing agent, the first nucleobase contains mC and the second nucleobase contains unmodified cytosine or hmC, e.g., unmodified cytosine and, optionally, one or more of hmC, fC, and / or caC. By sequencing the converted DNA, positions read as cytosine are identified as being either hmC or unmodified C positions. Positions read as T, on the other hand, are identified as being T, fC, caC, or mC. Thus, for example, performing TAPSβ conversion on a DNA sample as described herein facilitates distinguishing, using sequence reads obtained from the sample, positions containing unmodified C or hmC on the one hand from positions containing mC. For an exemplary illustration of this type of conversion, see, e.g., Liu et al., Nature Biotechnology 2019; 37: 424-429.
[0276] In some embodiments, procedures that differentially affect a first nucleobase in DNA relative to a second nucleobase in DNA include chemical-assisted conversion using a substituted borane reducing agent, which optionally is 2-picolyl borane, borane pyridine, tert-butylamine borane, or ammonia borane. In chemical-assisted conversion using a substituted borane reducing agent, an oxidizing agent such as potassium perruthenate (KRuO4, also suitable for use in ox-BS conversion) is used to specifically oxidize hmC to fC. Treatment with pic-borane or another substituted borane reducing agent, such as borane pyridine, tert-butylamine borane, or ammonia borane, converts fC and caC to DHU, but does not affect mC or unmodified C. Thus, when using this type of conversion, the first nucleobase includes one or more of hmC, fC, and caC, and the second nucleobase includes one or more of unmodified cytosine or mC, e.g., unmodified cytosine and optionally mC. By sequencing the converted DNA, positions read as cytosine are identified as being either mC or unmodified C positions. Positions read as T, on the other hand, are identified as being T, fC, caC, or hmC. Thus, for example, by performing this type of conversion on a DNA sample described herein, sequence reads obtained from the sample can be used to readily distinguish positions containing unmodified C or, on the one hand, mC from positions containing hmC. For an exemplary description of this type of conversion, see, e.g., Liu et al., Nature Biotechnology 2019; 37: 424-429.
[0277] In some embodiments, procedures that differentially affect a first nucleobase in DNA from a second nucleobase in DNA include APOBEC-coupled epigenetic (ACE) conversion. In ACE conversion, DNA deaminase enzymes of the AID / APOBEC family, such as APOBEC3A (A3A), are used such that unmodified cytosine and mC are deaminated, while hmC, fC, or caC are not deaminated. Thus, when using ACE conversion, the first nucleobase includes unmodified C and / or mC (e.g., unmodified C and optionally, mC), and the second nucleobase includes hmC. By sequencing the ACE-converted DNA, positions that are read as cytosine are identified as hmC, fC, or caC positions. In contrast, positions that are read as T are identified as T, unmodified C, or mC. Thus, performing ACE conversion on a DNA sample described herein facilitates distinguishing positions containing hmC from positions containing mC or unmodified C using sequence reads obtained from the sample. For an exemplary description of ACE conversion, see, e.g., Schutsky et al., Nature Biotechnology 2018; 36: 1083-1090.
[0278] In some embodiments, a procedure that affects a first nucleobase in DNA in a different manner than a second nucleobase in the DNA includes, for example, enzymatic conversion of the first nucleobase, as in the case of EM-Seq. See, for example, Vaisvila R, et al. (2019) EM-seq: Detection of DNA methylation at single base resolution from picograms of DNA. bioRxiv; DOI: 10.1101 / 2019.12.20.884692, available at www.biorxiv.org / content / 10.1101 / 2019.12.20.884692v1. For example, 5mC and 5hmC can be converted to substrates that cannot be deaminated by a deaminase (e.g., APOBEC3A) using TET2 and T4-βGT, and then a deaminase (e.g., APOBEC3A) can be used to deaminate unmodified cytosine and convert it to uracil.
[0279] In some embodiments, the procedure for affecting the first nucleobase in the DNA of the first secondary sample in a form different from the second nucleobase of the DNA converts the modified nucleoside. In some embodiments, the conversion procedure for converting the modified nucleoside includes enzymatic conversion such as DM-seq described in, for example, WO2023 / 288222A1. In DM-seq, unmodified cytosine in the DNA is enzymatically protected from the subsequent deamination step in which 5mC of 5mCpG is converted to T. The enzymatically protected unmodified (e.g., unmethylated) cytosine is not converted and is read as "C" during sequencing. Cytosine read as thymine in the DNA (in the context of CpG) is identified as methylated cytosine. Thus, when using this type of conversion, the first nucleobase includes unmodified (e.g., unmethylated) cytosine and the second nucleobase includes modified (e.g., methylated) cytosine. By sequencing the converted DNA, the positions read as cytosine are identified as unmodified C positions. On the other hand, the positions read as T are identified as either T or 5mC. Thus, by performing DM-seq conversion, it becomes easy to identify the positions containing 5mC using the resulting sequence reads.
[0280] In some embodiments, a procedure for affecting a first nucleobase in DNA in a manner different from a second nucleobase in the DNA includes separating DNA that originally contains the first nucleobase from DNA that does not originally contain the first nucleobase. In some such embodiments, the first nucleobase is hmC. Using a labeling procedure that includes a biotinylated position that originally contained the first nucleobase, DNA that originally contains the first nucleobase can be separated from other DNA. In some embodiments, the first nucleobase is first derivatized using an azide-containing moiety, such as a glucosyl azide-containing moiety. The azide-containing moiety can then function as a reagent for attaching biotin, for example by Huisgen cycloaddition chemistry. A biotin binder, such as avidin, neutravidin (deglycosylated avidin having an isoelectric point of about 6.3), or streptavidin, can then be used to separate the DNA that was biotinylated at that point and that originally contained the first nucleobase from DNA that does not originally contain the first nucleobase. An example of a procedure for separating DNA that originally contains the first nucleobase from DNA that does not originally contain the first nucleobase is hmC-seal, which labels hmC to form β-6-azido-glucosyl-5-hydroxymethylcytosine, then attaches a biotin moiety by Huisgen cycloaddition, and then uses a biotin binder to separate the biotinylated DNA from other DNA. For an exemplary description of hmC-seal, see, for example, Han et al., Mol. Cell 2016; 63: 711-719. This technique is useful for identifying fragments containing one or more hmC nucleobases.
[0281] In some embodiments, after such separation, the method further includes differentially tagging each of the DNA that originally contains the first nucleobase and the DNA that originally does not contain the first nucleobase. The method may further include pooling the DNA that originally contains the first nucleobase and the DNA that originally does not contain the first nucleobase after the differentially tagging step. The DNA that originally contains the first nucleobase and the DNA that originally does not contain the first nucleobase can then be used for downstream analysis. For example, while sequencing the pooled DNA that originally contains the first nucleobase and the DNA that originally does not contain the first nucleobase within the same sequencing cell (e.g., after subjecting to further processing such as those described herein), the ability to elucidate whether a given read is from a molecule of DNA that originally contains the first nucleobase or from a molecule of DNA that originally does not contain the first nucleobase using the differential tags can be retained.
[0282] In some embodiments, the first nucleobase is a modified or unmodified adenine and the second nucleobase is a modified or unmodified adenine. In some embodiments, the modified adenine is N 6 -methyladenine (mA). In some embodiments, the modified adenine is N 6 -methyladenine (mA), N 6 -hydroxymethyladenine (hmA), or N 6 -formyladenine (fA), or one or more of them.
[0283] Techniques involving partitioning based on methylation status or methylated DNA immunoprecipitation (MeDIP) can be used to separate DNA containing modified bases, such as mC, mA, caC (which can be generated, for example, by oxidation of mC or hmC using Tet2 prior to enzymatic conversion of unmodified C to U using a deaminase such as APOBEC3A), or dihydrouracil, from other DNA. See, for example, Kumar et al., Frontiers Genet. 2018; 9: 640; Greer et al., Cell 2015; 161: 868-878. Antibodies specific for mA are described in Sun et al., Bioessays 2015; 37: 1155-62. Antibodies are commercially available against various modified nucleobases, including thymine / uracil forms, such as mC, caC, and dihydrouracil, or halogenated forms, such as 5-bromouracil. Various modified bases can also be detected based on changes in their base pairing specificities. For example, hypoxanthine is a modified form of adenine that can result from deamination and is read as a G in sequencing. See, for example, U.S. Patent 8,486,630; Brown, Genomes, 2 nd Ed., John Wiley & Sons, Inc., New York, N.Y., 2002, chapter 14, "Mutation, Repair, and Recombination." See also. Capture of B.DNA; Capture moiety
[0284] In some embodiments, the methods herein include the step of capturing or enriching nucleic acid molecules containing sequences that are present within a set of target regions for subsequent analysis. Such enrichment or capture can be performed on any of the samples or subsamples described herein using any suitable technique known in the art. In some embodiments, the step of capturing includes contacting the DNA with a probe specific for such target regions. In some embodiments, the probe includes an oligonucleotide and a capture moiety, such as biotin or other examples described below. The probe can have a sequence selected to be arrayed across a group of regions, such as genes.
[0285] Methods that include DNA capture using a probe that includes a capture moiety, such as a target-specific probe labeled with biotin, can also include a second moiety or binding partner that binds to the capture moiety, such as streptavidin. In some embodiments, the capture moiety and the binding partner can have higher and lower capture yields, respectively, for sets of different probes, such as those used to capture sets of sequence-variable target regions and epigenetic target regions, as discussed elsewhere herein. Methods that include a capture moiety are further described, for example, in U.S. Patent No. 9,850,523, issued December 26, 2017, which is incorporated herein by reference.
[0286] Examples of capture moieties include, without limitation, biotin, avidin, streptavidin, nucleic acids containing a specific nucleotide sequence, haptens recognized by antibodies, and magnetically attractable particles. In some embodiments, the capture moiety bound to the subject is captured by a binding partner bound to an isolatable moiety, such as a magnetically attractable particle or a large particle that can be sedimented by centrifugation. The capture moiety can be any type of molecule that enables affinity separation from a nucleic acid lacking the capture moiety of the nucleic acid having the capture moiety. Exemplary capture moieties include biotin, which enables affinity separation by binding to streptavidin linked or linkable to a solid phase, or an oligonucleotide, which enables affinity separation by binding to a complementary oligonucleotide linked or linkable to a solid phase.
[0287] In some embodiments, non-specifically bound DNA that does not contain the target region is washed away from the captured DNA. In some embodiments, the DNA is then dissociated from the probe and eluted from the solid support using a buffer containing a salt wash or another DNA denaturing agent. In some embodiments, the probe is also eluted from the solid support, for example, by disrupting the biotin-streptavidin interaction. In some embodiments, the captured DNA is amplified after elution from the solid support. In some such embodiments, DNA containing an adapter is amplified using PCR primers that anneal to the adapter. In some embodiments, the captured DNA is amplified while bound to the solid support. In some such embodiments, the amplification involves the use of PCR primers that anneal to sequences within the adapter and PCR primers that anneal to sequences within the probe that anneal to the target region of the DNA.
[0288] In some embodiments, the target region is captured from an aliquot, portion, or secondary sample of a sample (e.g., a sample that has undergone adapter ligation and amplification), while the step of partitioning the DNA can be performed on separate aliquots, portions, or secondary samples of the sample. The step of enriching or capturing DNA comprising the target region can include contacting the DNA with a first or second set of target-specific probes. Such target-specific probes can have any of the features described herein, including, but not limited to, the embodiments described in this specification and in the section of this specification regarding probes. The capturing step can be performed on one or more secondary samples prepared during the methods disclosed herein. In some embodiments, the DNA is captured from a first secondary sample or a second secondary sample. In some embodiments, the secondary samples are differentially tagged (e.g., as described herein) and then pooled prior to undergoing capture. Exemplary methods for capturing DNA comprising epigenetic target regions and / or sequence-variable target regions can be found, for example, in WO2020 / 160414, which is hereby incorporated by reference herein.
[0289] The capturing step(s) can be performed using conditions suitable for specific nucleic acid hybridization, which generally depend to some extent on the features of the probe, such as length, base composition, etc. Suitable conditions are well known to those of skill in the art in view of general knowledge in the art regarding nucleic acid hybridization.
[0290] In some embodiments, the methods described herein include capturing a plurality of sets of target regions of cfDNA obtained from a subject. The target regions may include differences depending on whether their origin is tumor, healthy cells, or a particular cell type. The capturing step results in a captured set of cfDNA molecules. In some embodiments, cfDNA molecules corresponding to sets of variable target regions are captured with a higher capture yield in the captured set of cfDNA molecules than cfDNA molecules corresponding to sets of epigenetic target regions. In some embodiments, the methods described herein include contacting cfDNA obtained from a subject with a set of target-specific probes, the set of target-specific probes being configured to capture cfDNA corresponding to sets of variable target regions with a higher capture yield than cfDNA corresponding to sets of epigenetic target regions.
[0291] To analyze variable target regions with sufficient confidence or accuracy, a deeper sequencing depth may be required than may be necessary to analyze epigenetic target regions, so it can be beneficial to capture cfDNA corresponding to sets of variable target regions with a higher capture yield than cfDNA corresponding to sets of epigenetic target regions. The amount of data required to determine fragmentation patterns (e.g., to test for disruption of transcription start sites or CTCF binding sites) or fragment abundance (e.g., in high and low methylation profiles) is generally less than the amount of data required to determine the presence or absence of cancer-related sequence variations. Capturing sets of target regions at different yields can facilitate sequencing the target regions to different sequencing depths in the same sequencing run (e.g., using pooled mixtures and / or within the same sequencing cell). Copy number polymorphisms such as local amplifications are somatic mutations, but can be detected in a manner similar to techniques for detecting certain epigenetic changes such as changes in methylation by sequencing based on read frequency.
[0292] In some embodiments, the captured DNA is amplified. In various embodiments, the method further includes, consistent with the discussion herein, sequencing the captured DNA to different degrees of sequencing depth for, e.g., an epigenetic target region set and a variable target region set. In some embodiments, RNA probes are used. In some embodiments, DNA probes are used. In some embodiments, single-stranded probes are used. In some embodiments, double-stranded probes are used. In some embodiments, single-stranded RNA probes are used. In some embodiments, double-stranded DNA probes are used.
[0293] In some embodiments, the capturing step is performed using probes for a variable target region set and probes for an epigenetic target region set simultaneously in the same vessel, e.g., the probes for the variable target region set and the epigenetic target region set and the capture probes are in the same composition. This approach results in a relatively streamlined workflow.
[0294] In some embodiments, a group of target-specific probes is used in a method that includes the step of capturing the DNA described herein. In some embodiments, a group of target-specific probes includes target-binding probes specific to one or more target region sets. In some embodiments, the capture yield of the target-binding probes specific to the variable target region set is higher (e.g., at least two times higher) than the capture yield of the target-binding probes specific to the epigenetic target region set. In some embodiments, a group of target-specific probes is configured to have a capture yield for the variable target region set that is higher (e.g., at least two times higher) than the capture yield for the epigenetic target region set. C. Computer System
[0295] The methods of the present disclosure can be implemented using or leveraging a computer system. For example, such methods can include subjecting DNA or a secondary sample in a sample to a procedure that affects a first nucleobase in the DNA in a manner different from a second nucleobase in the DNA, where the first nucleobase is a modified or unmodified nucleobase, the second nucleobase is a modified or unmodified nucleobase different from the first nucleobase, and the first nucleobase and the second nucleobase have the same base pairing specificity; sequence-specifically cleaving a target DNA sequence; and sequencing the remaining DNA in a manner that distinguishes the first nucleobase and the second nucleobase in the DNA.
[0296] FIG. 2 shows a computer system 201 programmed or otherwise configured to implement the methods of the present disclosure. The computer system 201 can regulate various aspects of sample preparation, sequencing, and / or analysis. In some examples, the computer system 201 is configured to perform sample analysis, including sample preparation and nucleic acid sequencing, according to any of the methods disclosed herein, for example.
[0297] The computer system 201 may include a central processing unit (CPU, also referred to herein as "processor" and "computer processor") 205, which may be a single-core or multi-core processor, or multiple processors for parallel processing. The computer system 201 also includes a memory or memory location 210 (e.g., random access memory, read-only memory, flash memory), an electronic storage device 215 (e.g., hard disk), a communication interface 220 (e.g., network adapter) for communicating with one or more other systems, and peripheral devices 225, such as a cache, other memory, data storage, and / or an electronic display adapter. The memory 210, storage device 215, interface 220, and peripheral devices 225 communicate with the CPU 205 via a communication network or a bus, such as a motherboard (solid lines). The storage device 215 may be a data storage device (or data repository) for storing data. The computer system 201 may be operatively coupled to a computer network 230 using the communication interface 220. The computer network 230 may be the Internet, an intranet and / or extranet of the Internet, or an intranet and / or extranet that communicates with the Internet. The computer network 230 may be, in some cases, a telecommunications and / or data network. The computer network 230 may include one or more computer servers that enable distributed computing, such as cloud computing. The computer network 230 may, in some cases, implement a peer-to-peer network that enables devices coupled to the computer system 201 to act as clients or servers using the computer system 201.
[0298] The CPU 205 can execute a series of machine-readable instructions that can be embodied in a program or software. The instructions may be stored in a memory location such as the memory 210. Examples of operations performed by the CPU 205 can include fetch, decode, execute, and write-back.
[0299] The storage device 215 can store files, such as drivers, libraries, and saved programs. The storage device 215 can store programs and recorded sessions generated by the user, as well as outputs associated with the programs. The storage device 215 can store user data, such as user preferences and user programs. In some cases, the computer system 201 can include one or more additional data storage devices located external to the computer system 201, for example, on a remote server that communicates with the computer system 201 through an intranet or the Internet. Data can be transferred by moving it from one location to another, for example, using a communication network or physical data transfer (e.g., using a hard drive, thumb drive, or other data storage mechanism).
[0300] The computer system 201 can communicate with one or more remote computer systems through the network 230. As an embodiment, the computer system 201 can communicate with a remote computer system of a user (e.g., an operator). Examples of remote computer systems can include personal computers (e.g., laptop PCs), slate or tablet PCs (e.g., Apple® iPad®, Samsung® Galaxy Tab), telephones, smartphones (e.g., Apple® iPhone®, Android®-compatible devices, Blackberry®), or personal digital assistants. The user can access the computer system 201 via the network 230.
[0301] The methods described herein can be implemented by code executable by a machine (e.g., a computer processor) stored in an electronic memory location of a computer system 201, such as in a memory 210 or an electronic storage device 215. The machine-executable or machine-readable code may be provided in the form of software. In use, the code can be executed by a processor 205. In some cases, the code can be retrieved from the storage device 215 and stored in the memory 210 so that it can be readily accessed by the processor 205. In some situations, the electronic storage device 215 may be excluded and the machine-executable instructions may be stored in the memory 210.
[0302] In one aspect, the present disclosure provides a non-transitory computer-readable medium comprising computer-executable instructions that, when executed by at least one electronic processor, perform at least part of a method comprising: collecting cfDNA from a subject; distributing the DNA into a plurality of secondary samples and / or contacting the sample with a methylation-sensitive nuclease based on the level of modification associated with (e.g., in) the DNA; sequence-specifically digesting the DNA using a modification-dependent sequence-specific nuclease; sequencing the remaining cfDNA; obtaining a plurality of sequence reads generated by a nucleic acid sequencer from sequencing the cfDNA molecules; mapping the plurality of sequence reads to one or more reference sequences to generate mapped sequence reads; and processing the mapped sequence reads to determine the likelihood that the subject has cancer.
[0303] The code can be precompiled and configured for use on a machine having a processor adapted to execute the code, or can be compiled during run-time. The code can be provided in a programming language that can be selected to enable the code to be executed in a precompiled mode or a just-in-time compilation mode.
[0304] Aspects of the systems and methods provided herein, such as computer system 201, can be embodied in programming. Various aspects of the technology can generally be considered a "product" or "manufacture" in the form of machine (or processor) executable code and / or associated data that is implemented or embodied as a certain type of machine-readable medium. The machine executable code can be stored in an electronic storage device such as a memory (e.g., read only memory, random access memory, flash memory) or a hard disk. A "storage" type medium can include any or all of tangible memories of a computer, processor, etc., or associated modules thereof, such as various semiconductor memories, tape drives, disk drives, etc., that can provide non-transitory storage at any point in time for software programming.
[0305] All or part of the software may sometimes communicate through the Internet or various other telecommunications networks. Such communication can enable, for example, the loading of software from one computer or processor to another, such as from a management server or host computer to an application server's computer platform. Thus, as another type of medium on which software elements can be carried, there are those used through wired and optical line networks and various air links beyond the physical interface between local devices, including light waves, radio waves, and electromagnetic waves. Physical elements that carry such waves, such as wired or wireless links, optical links, etc., can also be regarded as media on which software is carried. As used herein, unless restricted to non-transitory tangible "memory" media, terms such as computer or machine "readable media" refer to any medium involved in bringing instructions for execution to a processor.
[0306] Thus, machine-readable media such as computer-executable code can take many forms including, but not limited to, tangible storage media, carrier wave media, or physical transmission media. Non-volatile storage media includes, for example, any computer(s) such as optical or magnetic disks, such as those that can be used to implement a database shown in the drawings, or any of the storage devices in any computer. Volatile storage media includes dynamic memory such as the main memory of such a computer platform. Tangible transmission media includes coaxial cables; copper wires and fiber optics including lines that include buses within computer systems. Carrier wave transmission media may take the form of electrical or electromagnetic signals, or acoustic or light waves such as those generated during high-frequency (RF) and infrared (IR) data communications. Thus, common forms of computer-readable media include, for example, floppy (registered trademark) disks, flexible disks, hard disks, magnetic tapes, any other magnetic media, CD-ROM, DVD or DVD-ROM, any other optical media, punch cards, paper tapes, any other physical storage media having patterns of holes, RAM, ROM, PROM and EPROM, FLASH (registered trademark)-EPROM, any other memory chip or cartridge, carrier wave-transmitted data or instructions, cables or links that transport such carrier waves, or any other media that a computer can read programming code and / or data from. Many of these forms of computer-readable media may be involved in carrying one or more sequences of one or more instructions to a processor for execution.
[0307] Computer system 201 can include or communicate with an electronic display that includes a user interface (UI) for providing, for example, one or more results of a sample analysis. Examples of UIs include, without limitation, graphical user interfaces (GUIs) and web-based user interfaces.
[0308] Additional details regarding computer systems and networks, databases, and computer program products can be found, for example, in Peterson, Computer Networks: A Systems Approach, Morgan Kaufmann, 5th Ed. (2011), Kurose, Computer Networking: A Top-Down Approach, Pearson, 7 th Ed. (2016), Elmasri, Fundamentals of Database Systems, Addison Wesley, 6th Ed. (2010), Coronel, Database Systems: Design, Implementation, & Management, Cengage Learning, 11 th Ed. (2014), Tucker, Programming Languages, McGraw-Hill Science / Engineering / Math, 2nd Ed. (2006), and Rhoton, Cloud Computing Architected: Solution Design Handbook, Recursive Press (2011). D. Applications 1. Cancer and Other Diseases
[0309] The present method can be used to diagnose the presence of a condition in a subject, such as the presence of cancer or a pre-cancerous condition, to characterize the condition (e.g., to determine the stage of cancer or the heterogeneity of cancer), to monitor the response of the subject to having received treatment for the condition (e.g., the response to a chemotherapeutic agent or an immunotherapeutic agent), to assess the prognosis of the subject (e.g., to predict the survival outcome in a subject having cancer), to determine the risk that a subject will develop the condition, to predict the subsequent course of the condition in the subject, to determine the metastasis or recurrence of cancer (or the risk of cancer metastasis or recurrence) in the subject, and / or to monitor the health of the subject as part of a preventive health monitoring program (e.g., to determine whether further diagnostic screening is needed for the subject and / or to determine when further diagnostic screening is needed). The present disclosure can also be useful in determining the effectiveness of a particular treatment option. In the case of a successful treatment option, more cancer cells can be killed and DNA can be excreted, so that the amount of copy number polymorphisms, rare mutations, and / or cancer-related epigenetic signatures (e.g., hypermethylated regions or hypomethylated regions) detected in the blood of the subject (e.g., cfDNA isolated from a blood sample from the subject (e.g., a whole blood sample, a leukocyte-depleted sample, a buffy coat sample, or a PBMC sample)) may increase, or an increase or decrease in the quantity of a particular immune cell type in the blood may result if the treatment is successful, and no change will result if the treatment is unsuccessful. In other examples, this may not occur. In another example, a particular treatment option can be associated with the genetic profile of cancer over time. This correlation can be useful in the selection of treatment for a subject.
[0310] In some embodiments, the method is used to screen for cancer or in a method for screening for cancer. For example, the sample may be a sample from a subject who has not previously been diagnosed with cancer. In some embodiments, the subject may or may not have cancer. In some embodiments, the subject may or may not have early-stage cancer. In some embodiments, the subject has one or more risk factors for cancer, such as tobacco use (e.g., smoking), being overweight or obese, having a high body mass index (BMI), being elderly, having malnutrition, having a high alcohol consumption, or having a family history of cancer.
[0311] In some embodiments, the subject has used tobacco for, for example, at least 1 year, 5 years, 10 years, or 15 years. In some embodiments, the subject has a high BMI, for example, a BMI of 25 or higher, 26 or higher, 27 or higher, 28 or higher, 29 or higher, or 30 or higher. In some embodiments, the subject is at least 40 years old, 45 years old, 50 years old, 55 years old, 60 years old, 65 years old, 70 years old, 75 years old, or 80 years old. In some embodiments, the subject is undernourished, for example, having a high consumption of one or more of red and / or processed meat, trans fat, saturated fat, and refined sugar, and / or having a low consumption of fruits and vegetables, complex carbohydrates, and / or unsaturated fat. High or low consumption can be defined, for example, as exceeding or falling below the recommendations in the Dietary Guidelines for Americans 2020 - 2025 available at www.dietaryguidelines.gov / sites / default / files / 2021-03 / Dietary_Guidelines_for_Americans-2020-2025.pdf. In some embodiments, the subject has a high alcohol consumption, for example, an average of at least 3, 4, or 5 drinks per day (where one drink is about 1 ounce or 30 mL of 80-proof hard liquor or equivalent). In some embodiments, the subject has a family history of cancer, for example, at least 1, 2, or 3 blood relatives have previously been diagnosed with cancer. In some embodiments, the blood relatives are at least third-degree relatives (e.g., great-grandparents, great-aunts or great-uncles, cousins), at least second-degree relatives (e.g., grandparents, aunts or uncles, or siblings with a common parent), or first-degree relatives (e.g., parents or siblings with a common parent).
[0312] Generally, the disease under consideration is a certain type of cancer. Non-limiting examples of such cancers include biliary tract cancer, bladder cancer, transitional cell carcinoma, urothelial carcinoma, brain cancer, glioma, astrocytoma, breast cancer, metaplastic carcinoma, cervical cancer, cervical squamous cell carcinoma, rectal cancer, colorectal cancer, colon cancer, hereditary non-polyposis colorectal cancer, colorectal adenocarcinoma, gastrointestinal stromal tumor (GIST), endometrial cancer, endometrial stromal sarcoma, esophageal cancer, esophageal squamous cell carcinoma, esophageal adenocarcinoma, intraocular melanoma, choroidal melanoma, gallbladder cancer, gallbladder adenocarcinoma, renal cell carcinoma, clear cell renal cell carcinoma, transitional cell carcinoma, urothelial carcinoma, Wilms tumor, leukemia, acute lymphoblastic leukemia (ALL), acute myeloid leukemia (AML), chronic lymphocytic leukemia (CLL), chronic myeloid leukemia (CML), chronic myelomonocytic leukemia (CMML), liver cancer, hepatoma, hepatocellular carcinoma, cholangiocarcinoma, hepatoblastoma, lung cancer, non-small cell lung cancer (NSCLC), mesothelioma, B-cell lymphoma, non-Hodgkin lymphoma, diffuse large B-cell lymphoma, mantle cell lymphoma, T-cell lymphoma, non-Hodgkin lymphoma, precursor T-lymphoblastic lymphoma / leukemia, peripheral T-cell lymphoma, multiple myeloma, nasopharyngeal carcinoma (NPC), neuroblastoma, oropharyngeal cancer, oral squamous cell carcinoma, osteosarcoma, ovarian cancer, pancreatic cancer, pancreatic ductal adenocarcinoma, pseudopapillary neoplasm, acinar cell carcinoma, prostate cancer, prostate adenocarcinoma, skin cancer, melanoma, malignant melanoma, cutaneous melanoma, small intestine cancer, gastric cancer, gastrointestinal stromal tumor (GIST), uterine cancer, or uterine sarcoma. The type and / or stage of cancer can be detected from genetic diversity including mutations, rare mutations, indels, copy number polymorphisms, base conversions, translocations, inversions, deletions, aneuploidy, segmental aneuploidy, ploidy, chromosomal instability, chromosomal structural changes, gene fusions, chromosomal fusions, gene shortening, gene amplification, gene duplication, chromosomal damage, DNA damage, abnormal changes in nucleic acid chemical modification, abnormal changes in epigenetic patterns, and abnormal changes in nucleic acid 5-methylcytosine.
[0313] The methods of this specification can also be used to characterize certain forms of cancer. Cancers are often heterogeneous with respect to both composition and staging. Characterization of a particular subtype of cancer can be important in the diagnosis or treatment of that particular subtype. This information can also provide clues regarding the prognosis of a particular type of cancer to a subject or practitioner, enabling either the subject or practitioner to adapt treatment options in accordance with the progression of the disease. Some cancers can progress to a higher malignancy and become genetically unstable. Other cancers can remain benign, inactive, or in a quiescent state. The systems and methods of the present disclosure can be useful in determining disease progression.
[0314] Furthermore, the methods of the present disclosure can be used to characterize the heterogeneity of an abnormal condition in a subject. Such methods can include, for example, generating a profile of cfDNA derived from the subject. In some embodiments, the abnormal condition is cancer. In some embodiments, the abnormal condition can result in a heterogeneous genomic population. In the example of cancer, some tumors have been found to contain tumor cells at different stages of cancer. In other examples, the heterogeneity can include multiple lesions of a disease. Again, in the example of cancer, multiple tumor lesions can be present, perhaps one or more of the lesions being the result of metastases that have spread from the primary site. The origin tissue(s) can be useful in identifying the organ(s) affected by cancer, including the primary cancer and / or metastatic tumors.
[0315] In some embodiments, the methods described herein include detecting the presence or absence of nucleic acids originating from or derived from tumor cells at a preselected time point after a previous cancer treatment of a subject previously diagnosed with cancer. The method can further include determining, for the subject, a cancer recurrence score indicative of the presence or level of DNA originating from or derived from tumor cells.
[0316] When determining a cancer recurrence score, the score can be further used to determine the cancer recurrence status. The cancer recurrence status can be, for example, that there is a risk of cancer recurrence if the cancer recurrence score exceeds a predetermined threshold. The cancer recurrence status can be, for example, that the risk of cancer recurrence is low or lower if the cancer recurrence score exceeds a predetermined threshold. In certain embodiments, a cancer recurrence score equal to the predetermined threshold can result in a cancer recurrence status of either having a risk of cancer recurrence or having a low or lower risk of cancer recurrence.
[0317] In some embodiments, the cancer recurrence score is compared to a predetermined cancer recurrence threshold, and the subject is classified as a candidate for subsequent cancer treatment if the cancer recurrence score exceeds the cancer recurrence threshold or not a candidate for treatment if the cancer recurrence score is below the cancer recurrence threshold. In certain embodiments, a cancer recurrence score equal to the cancer recurrence threshold can result in a classification as either a candidate for subsequent cancer treatment or not a candidate for treatment.
[0318] In some embodiments, the methods herein do not involve diagnosing, prognosticating, or monitoring a fetus, and thus are not directed to non-invasive prenatal testing. In other embodiments, these methods can be used in a subject during pregnancy to diagnose, prognosticate, monitor, or observe cancer or other diseases in a pre-born subject in which DNA and other polynucleotides may be co-circulating with maternal molecules. Optionally, non-limiting examples of other gene-based diseases, disorders, or conditions that can be evaluated using the methods and systems disclosed herein include achondroplasia, alpha-1 antitrypsin deficiency, antiphospholipid syndrome, autism, autosomal dominant polycystic kidney disease, Charcot-Marie-Tooth (CMT), meowing cat, Crohn's disease, cystic fibrosis, Dercum's disease, Down syndrome, Duane syndrome, Duchenne muscular dystrophy, factor V Leiden thrombophilia, familial hypercholesterolemia, familial Mediterranean fever, fragile X syndrome, Gaucher disease, hemochromatosis, hemophilia, holoprosencephaly, Huntington's disease, Klinefelter syndrome, Marfan syndrome, myotonic dystrophy, neurofibromatosis, Noonan syndrome, osteogenesis imperfecta, Parkinson's disease, phenylketonuria, Poland anomaly, porphyria, progeria, retinitis pigmentosa, severe combined immunodeficiency (SCID), sickle cell disease, spinal muscular atrophy, Tay-Sachs, thalassemia, trimethylaminuria, Turner syndrome, velocardiofacial syndrome, WAGR syndrome, Wilson's disease, and the like.
[0319] This method can also be used to quantify the levels of different cell types, including rare immune cell types such as activated lymphocytes and bone marrow cells at specific differentiation stages. Such quantification can be based on the number of molecules corresponding to a given cell type in the sample. In some embodiments, the quantity of each of a plurality of cell types, such as immune cell types, is determined based on the sequencing and analysis (e.g., determination of epigenetic and / or genomic signatures) of DNA (e.g., cfDNA) isolated from at least one sample from the subject (e.g., a whole blood sample, a buffy coat sample, a leukapheresis sample, or a PBMC sample). The plurality of immune cell types include, but are not limited to, macrophages (including M1 macrophages and M2 macrophages), activated B cells (including regulatory B cells, memory B cells, and plasma cells); T cell subsets, such as central memory T cells, naive-like T cells, and activated T cells (including cytotoxic T cells, regulatory T cells (Tregs), CD4 effector memory T cells, CD4 central memory T cells, CD8 effector memory T cells, and CD8 central memory T cells); immature myeloid cells (including myeloid-derived suppressor cells (MDSCs), low-density neutrophils, immature neutrophils, and immature granulocytes); and natural killer (NK) cells. As disclosed herein, differences in the levels and / or presence of specific genetic and / or epigenetic signatures in DNA isolated from a blood sample from a subject can be used to quantify cell types such as immune cell types within the sample.
[0320] The array information obtained in the present method may include nucleic acid sequence reads generated by a nucleic acid sequencer. In some embodiments, the nucleic acid sequencer performs pyrosequencing, single molecule sequencing, nanopore sequencing, semiconductor sequencing, sequencing by synthesis, five-letter sequencing, six-letter sequencing, sequencing by ligation, or sequencing by hybridization on the nucleic acid to generate sequencing reads. In some embodiments, the method further includes the step of grouping the sequence reads into families of sequence reads, each family including sequence reads generated from nucleic acids in the sample. In some embodiments, the method includes the step of determining the likelihood that the subject from whom the sample was obtained has cancer, a pre-cancerous condition, an infectious disease, graft rejection, or another disease or disorder associated with a change in the proportion of immune cell types. As discussed herein, by comparing the identity and / or the quantity / proportion of immune cells between two or more samples collected from a subject at two different time points, one or more aspects of the condition in the subject can be monitored over time, such as the subject's response to treatment, the severity of the condition in the subject (e.g., cancer stage), recurrence of the condition (e.g., cancer), and / or the risk that the subject will develop the condition (e.g., cancer).
[0321] The methods discussed herein may further include any suitable feature(s) described elsewhere herein, including sections related to methods for determining the risk of cancer recurrence in a subject and / or classifying a subject as a candidate for subsequent cancer treatment. 2. Method for determining the risk of cancer recurrence in a test subject and / or classifying a test subject as a candidate for subsequent cancer treatment
[0322] In some embodiments, the methods provided herein are methods for determining the risk of cancer recurrence in a subject. In some embodiments, the methods provided herein are methods for classifying a subject as a candidate for subsequent cancer treatment.
[0323] Any such method may include the step of collecting nucleic acids (e.g., DNA or RNA originating from or derived from tumor cells) at one or more preselected time points after one or more previous cancer treatments for the subject from a subject diagnosed with cancer. The subject may be any of the subjects described herein. The DNA may include cfDNA. The DNA may be DNA from a blood sample (e.g., a whole blood sample), such as cfDNA. The DNA may include DNA obtained from a tissue sample.
[0324] Any such method may include the step of contacting the sample or a secondary sample thereof with a modification-independent sequence-specific nuclease according to any of the embodiments described herein. Any such method may include the step of sequencing a DNA molecule, thereby generating a set of sequence information. Any such method may include the step of using the set of sequence information to detect the presence or absence of DNA originating from or derived from tumor cells at a preselected time point. Detection of the presence or absence of DNA originating from or derived from tumor cells, such as cfDNA, can be performed according to any of the embodiments described elsewhere herein.
[0325] In any such method, the previous cancer treatment may include surgery, administration of a therapeutic composition, and / or chemotherapy.
[0326] A method for determining the risk of cancer recurrence in a subject may include determining, for the subject, a cancer recurrence score indicative of the presence, absence, or amount of a target region originating from or derived from the target genomic region and tumor cells. The cancer recurrence score can be further used to determine the cancer recurrence state. The cancer recurrence state can be, for example, that there is a risk of cancer recurrence if the cancer recurrence score exceeds a predetermined threshold. The cancer recurrence state can be, for example, that the risk of cancer recurrence is low or lower if the cancer recurrence score exceeds a predetermined threshold. In certain embodiments, a cancer recurrence score equal to the predetermined threshold can result in a cancer recurrence state of either a risk of cancer recurrence or a low or lower risk of cancer recurrence.
[0327] A method for classifying a subject as a candidate for subsequent cancer treatment may include comparing the subject's cancer recurrence score to a predetermined cancer recurrence threshold, whereby the subject is classified as a candidate for subsequent cancer treatment if the cancer recurrence score exceeds the cancer recurrence threshold, or not a candidate for treatment if the cancer recurrence score is below the cancer recurrence threshold. In certain embodiments, a cancer recurrence score equal to the cancer recurrence threshold can result in a classification of either a candidate for subsequent cancer treatment or not a candidate for treatment. In some embodiments, the subsequent cancer treatment includes administration of chemotherapy or a therapeutic composition.
[0328] Any of such methods may include determining, for the subject, a disease-free survival (DFS) period based on the cancer recurrence score. For example, the DFS period can be 1 year, 2 years, 3 years, 4 years, 5 years, or 10 years.
[0329] In some embodiments, the set of sequence information includes a set of target region sequences, and the step of determining the cancer recurrence score can include determining at least a first subscore indicative of the level of a specific cell type, SNV, insertion / deletion, CNV, and / or fusion present within the sequence-variable target region sequences.
[0330] In some embodiments, some mutations within the array-variable target region selected from one, two, three, four, or five are sufficient to result in a cancer recurrence score in which the first subscore is classified as positive for cancer recurrence. In some embodiments, the number of mutations is selected from one, two, or three.
[0331] In some embodiments, the set of array information includes an epigenetic target region array, and the step of determining the cancer recurrence score includes determining a second subscore indicative of the amount of molecules (obtained from the epigenetic target region array) that represent an epigenetic state different from the DNA found in a sample from a corresponding healthy subject (e.g., DNA from a blood sample (e.g., whole blood sample), e.g., cfDNA, and / or DNA found in a tissue sample from a healthy subject, where the tissue sample is of the same tissue type as that obtained from the subject). These abnormal molecules (i.e., molecules having an epigenetic state different from the DNA found in a sample from a corresponding healthy subject) may coincide with epigenetic changes associated with cancer, such as methylation of hypermethylated variable target regions and / or disruption of fragmentation of fragmented variable target regions, where "disruption" means different from the DNA found in a sample from a corresponding healthy subject.
[0332] In some embodiments, a proportion of molecules corresponding to a hypermethylated variable target region set and / or a fragmented variable target region set that exhibits hypermethylation in the hypermethylated variable target region set and / or abnormal fragmentation in the fragmented variable target region set being greater than or equal to a value in the range of 0.001% to 10% is sufficient for the second subscore to be classified as positive for cancer recurrence. The range may be 0.001% to 1%, 0.005% to 1%, 0.01% to 5%, 0.01% to 2%, or 0.01% to 1%.
[0333] In some embodiments, any such method may include determining a fraction of tumor DNA that exhibits one or more characteristics indicative of originating from tumor cells from a fraction of molecules within a set of sequence information. This can be done, for example, for molecules corresponding to some or all of the target regions, including hypermethylated variable target regions, one or both of hypomethylated variable target regions, and fragmented variable target regions (hypermethylation of hypermethylated variable target regions and / or abnormal fragmentation of fragmented variable target regions can be considered indicative of originating from tumor cells). This can be done for molecules corresponding to sequence variable target regions, such as molecules containing changes consistent with cancer, such as SNVs, indels, CNVs, and / or fusions. The fraction of tumor DNA can be determined based on a combination of molecules corresponding to epigenetic target regions and molecules corresponding to sequence variable target regions.
[0334] Determination of the cancer recurrence score may be based at least in part on the fraction of tumor DNA, where a fraction of tumor DNA exceeding a threshold in the range of 10 -11 ~1 or 10 -10 ~1 is sufficient for the cancer recurrence score to be classified as positive for cancer recurrence. In some embodiments, a fraction of tumor DNA in the range of 10 -10 ~10 -9 、10 -9 ~10 -8 、10 -8 ~10 -7 、10 -7 ~10 -6 、10 -6 ~10 -5 、10 -5 ~10 -4 、10 -4 ~10 -3 、10 -3 ~10 -2 、or 10 -2 ~10 -1 exceeding or equal to the threshold in the range is sufficient for the cancer recurrence score to be classified as positive for cancer recurrence. In some embodiments, the fraction of tumor DNA is at least 10 -7Exceeding the threshold is sufficient for the cancer recurrence score to be classified as positive for cancer recurrence. Determining that the fraction of tumor DNA exceeds a threshold, such as the threshold corresponding to any of the foregoing embodiments, can be performed based on the cumulative probability. For example, if the cumulative probability that the tumor fraction exceeds a threshold within any of the foregoing ranges exceeds a probability threshold of at least 0.5, 0.75, 0.9, 0.95, 0.98, 0.99, 0.995, or 0.999, the sample was considered positive. In some embodiments, the probability threshold is at least 0.95, such as 0.99.
[0335] In some embodiments, the set of sequence information includes a sequence variable target region sequence and an epigenetic target region sequence, and the step of determining the cancer recurrence score includes a first subscore indicating the amount of SNVs, insertions / deletions, CNVs, and / or fusions present in the sequence variable target region sequence, and a second subscore indicating the amount of abnormal molecules in the epigenetic target region sequence, and combining the first subscore and the second subscore to yield the cancer recurrence score. When combining the subscores, a threshold can be independently applied to each subscore within the sequence variable target region and to abnormal molecules (i.e., molecules having an epigenetic state different from the DNA found in a sample from a corresponding healthy subject; e.g., a tumor) exceeding a predetermined fraction in the epigenetic target region, or they can be combined by training a machine learning classifier to determine the state based on a plurality of positive and negative training samples.
[0336] In some embodiments, the combined score value being in the range of -4 to 2 or -3 to 1 is sufficient for the cancer recurrence score to be classified as positive for cancer recurrence.
[0337] In any embodiment where the cancer recurrence score is classified as positive for cancer recurrence, the subject's cancer recurrence status may be at risk of cancer recurrence and / or the subject may be classified as a candidate for subsequent cancer treatment. In some embodiments, the cancer is any one of the cancer types described elsewhere herein. 3. Method for Monitoring Cancer in a Subject over Time; Sample Collection at Two or More Time Points
[0338] In some embodiments, the method can be used to monitor one or more aspects of a condition in a subject over time, such as the subject's response to treatment for the condition (e.g., response to a chemotherapeutic or immunotherapeutic agent), the severity of the condition in the subject (e.g., cancer stage), recurrence of the condition (e.g., cancer), and / or the risk that the subject will develop the condition (e.g., cancer), and / or to monitor the health of the subject as part of a preventive health monitoring program (e.g., to determine whether further diagnostic screening is needed for the subject and / or when further diagnostic screening is needed). In some embodiments, the monitoring includes the analysis of at least two samples collected from the subject at at least two different time points as described herein.
[0339] The methods according to the present disclosure can be useful, for example, in predicting a subject's response to a particular treatment option over a period of time. As described elsewhere herein, in a successful treatment option, for example, more cancer may die and DNA may be excreted if the treatment is successful, so that the amount of cancer-related DNA sequences detected in the subject's blood can increase. In such an example, a particular treatment option can be associated with the genetic profile of the cancer over time. This correlation can be useful in treatment selection.
[0340] As disclosed herein, in some embodiments, the quantity of each of a plurality of cell types, such as immune cell types, is determined based on sequencing and analysis (e.g., determination of epigenetic and / or genomic signatures) of DNA isolated from at least one sample comprising cells from a subject (e.g., a tissue sample or a blood sample, such as a whole blood sample, a buffy coat sample, a leukapheresis sample, or a PBMC sample). In some embodiments, differences in the levels and / or presence of specific genetic and / or epigenetic signatures in DNA isolated from a blood sample from a subject can be used to quantify cell types, such as immune cell types, within the sample. Thus, comparison of the disclosed genetic and / or epigenetic signatures in DNA isolated from blood samples collected from a subject at two or more time points can be used to monitor changes in the quantity of cell types in a subject under different conditions (e.g., before and after treatment), or over time (e.g., as part of a preventive health monitoring program).
[0341] The disclosed method may include the step of evaluating (e.g., quantifying) and / or interpreting the cell types (e.g., immune cell types) present in one or more samples (e.g., tissue samples or blood samples, such as whole blood samples, buffy coat samples, leukapheresis samples, or PBMC samples) collected from a subject at one or more time points by comparing them to a selected baseline value or reference standard (or a set of selected baseline values or reference standards). The baseline value or reference standard may be the quantity of cell types (e.g., the average quantity or range of quantities of cell types present in at least two samples) measured in one or more samples collected from a subject at one or more time points, e.g., before treatment, before diagnosis of a condition (e.g., cancer), or as part of a preventive health monitoring program. The baseline value or reference standard may be the quantity of cell types (e.g., the average quantity or range of quantities of cell types present in at least two samples) measured in one or more samples collected from one or more subjects without the condition (e.g., healthy subjects without cancer), one or more subjects who responded favorably to a treatment, or one or more subjects who have not received treatment at one or more time points. In certain embodiments, the baseline value or reference standard utilized is a standard or profile derived from a single reference subject. In other embodiments, the baseline value or reference standard utilized is a standard or profile derived from average data from multiple reference subjects. In various embodiments, the reference standard may be a single value, average value, mean, numerical mean or range of numerical means, numerical pattern, or graphical pattern generated from quantity data of cell types derived from a single reference subject or multiple reference subjects. The selection of a particular baseline value or reference standard, or the selection of one or more reference subjects, depends on the use of the methods described herein, e.g., by a research scientist or clinician (e.g., a physician).
[0342] In some embodiments, one or more samples (e.g., tissue samples or blood samples, e.g., whole blood samples, buffy coat samples, leukapheresis samples, or PBMC samples) can be collected from a subject at two or more time points to assess changes in cell types (e.g., changes in the quantity of cell types) between two or more time points. In some embodiments, the sample collected at the first time point is a tissue sample or a blood sample, and the sample collected at a subsequent time point (e.g., the second time point) is a blood sample. In some embodiments, the sample collected at the first time point is a tissue sample, and the sample collected at a subsequent time point (e.g., the second time point) is a blood sample. By monitoring cell types in samples collected from a subject at two or more time points and identifying differences between cell types, this method can be used to determine, for example, the presence or absence of a condition (e.g., cancer), the response of the subject to a treatment, one or more characteristics of a condition (e.g., the stage of cancer) in the subject, the recurrence of a condition (e.g., cancer), and / or the risk that the subject will develop a condition (e.g., cancer). Thus, in some embodiments, provided is a method of comparing the quantity of cell types present in at least one sample (e.g., at least one tissue sample and / or at least one blood sample, e.g., whole blood sample, buffy coat sample, leukapheresis sample, or PBMC sample) collected from a subject at one or more time points (e.g., before receiving treatment) with the quantity of cell types present in at least one sample collected from the subject at one or more different time points (e.g., after receiving treatment). The disclosed method can enable patient-specific monitoring, and thus, for example, changes (e.g., the presence or absence of a condition, the response to a treatment, prognosis, etc.) that are meaningful with respect to the subject but can still fall within the normal range of the general healthy population can be indicated by differences in the quantity of cell types between samples collected from the subject at different time points.
[0343] As disclosed herein, methods are provided for monitoring one or more aspects of a condition in a subject over time, such as, but not limited to, the subject's response to having received treatment for the condition (e.g., response to a chemotherapeutic or immunotherapeutic agent). In certain embodiments, one or more samples are collected from the subject at at least 1 to 10, at least 1 to 5, at least 2 to 5, or at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, or at least 20 time points before the subject receives treatment. In certain embodiments, one or more samples are collected from the subject at at least 1 to 10, at least 1 to 5, at least 2 to 5, or at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, or at least 20 time points after the subject receives treatment. Continued sample collection from the subject during and / or after treatment can monitor the subject's response to treatment.
[0344] In some embodiments, samples are not collected from the subject before diagnosis of the condition (e.g., cancer) or before treatment is received. In such embodiments, the response of the subject to treatment, or the progression or stage of the condition (e.g., cancer) in the subject is monitored over time, and cell types are compared between samples obtained at at least 2 to 10, at least 2 to 5, at least 3 to 6, or at least 2, such as at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, or at least 20 time points after the subject is diagnosed and / or after the subject receives treatment. Continued sample collection from the subject during and / or after treatment can monitor the subject's response to treatment.
[0345] In some embodiments of the disclosed method, one or more samples (e.g., one or more tissue, whole blood, buffy coat, leukapheresis, or PBMC samples) are collected from a subject at least once per year, e.g., about 1 to 12 times or about 2 to 6 times per year, e.g., about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 times per year. In other embodiments, one or more samples are collected from a subject less than once per year, e.g., once every about 13 months, once every about 14 months, once every about 15 months, once every about 16 months, once every about 17 months, once every about 18 months, once every about 19 months, once every about 20 months, once every about 21 months, once every about 22 months, once every about 23 months, or once every about 24 months. In some embodiments, one or more samples are collected from a subject once every about 1 to 5 years or once every about 1 to 2 years, e.g., once every about 1 year, once every about 1.5 years, once every about 2 years, once every about 2.5 years, once every about 3 years, once every about 3.5 years, once every about 4 years, once every about 4.5 years, or once every about 5 years.
[0346] In other embodiments of the disclosed method, one or more samples (e.g., one or more tissue samples or blood samples, e.g., or one or more buffy coat samples, whole blood samples, leukoreduced samples, or PBMC samples) are collected from a subject at least once per week, e.g., on 1 to 4 days of the week, on 1 to 2 days, or on 1 day, 2 days, 3 days, 4 days, 5 days, 6 days, or 7 days. In certain embodiments, one or more samples are collected from a subject at least once per month, e.g., 1 to 15 times per month, 1 to 10 times, 2 to 5 times, or 1 time, 2 times, 3 times, 4 times, 5 times, 6 times, 7 times, 8 times, 9 times, 10 times, 11 times, 12 times, 13 times, 14 times, or 15 times. In other embodiments, one or more samples are collected from a subject monthly, every 2 months, every 3 months, every 4 months, every 5 months, every 6 months, every 7 months, every 8 months, every 9 months, every 10 months, every 11 months, or every 12 months. In some embodiments, one or more samples are collected from a subject at least once per day, e.g., 1 time, 2 times, 3 times, 4 times, 5 times, or 6 times per day. The selection of the one or more sample collection times (e.g., the frequency of sample collection) or the number of samples collected at each time depends on the use to be made of the method described herein, e.g., by a research scientist or a clinician (e.g., a physician). 4. Treatment and Related Administration
[0347] In certain embodiments, the methods disclosed herein relate to identifying and administering to a subject a treatment, e.g., a customized treatment. In some embodiments, determination of the levels of specific nucleic acids facilitates the selection of an appropriate treatment. In some embodiments, the subject has a given disease, disorder, or condition, e.g., any of the conditions described elsewhere herein.
[0348] In some embodiments, the treatment is customized based on whether the nucleic acid variant is of somatic origin or germline origin. In some embodiments, essentially any cancer treatment (e.g., surgery, radiation therapy, chemotherapy, immunotherapy, etc.) can be included as part of these methods. In certain embodiments, the treatment administered to the subject comprises at least one chemotherapeutic agent. In some embodiments, chemotherapeutic agents include alkylating agents (e.g., but not limited to chlorambucil, cyclophosphamide, cisplatin, and carboplatin), nitrosoureas (e.g., but not limited to carmustine and lomustine), antimetabolites (e.g., but not limited to fluorouracil, methotrexate, and fludarabine), plant alkaloids and natural products (e.g., but not limited to vincristine, paclitaxel, and topotecan), antitumor antibiotics (e.g., but not limited to bleomycin, doxorubicin, and mitoxantrone), hormonal agents (e.g., but not limited to prednisone, dexamethasone, tamoxifen, and leuprolide), and biological response modifiers (e.g., but not limited to herceptin and avastin, arcitumomab and rituxan). In some embodiments, the chemotherapy administered to the subject can include FOLFOX or FOLFIRI. In certain embodiments, a treatment comprising at least one PARP inhibitor can be administered to the subject. In certain embodiments, examples of PARP inhibitors can include, inter alia, olaparib, talazoparib, rucaparib, niraparib (trade name ZEJULA). Generally, the treatment comprises at least one immunotherapy (or immunotherapeutic agent). Immunotherapy generally refers to a method of enhancing the immune response against a given cancer type. In certain embodiments, immunotherapy refers to a method of enhancing the T cell response against a tumor or cancer.
[0349] In some embodiments, the immunotherapy or immunotherapeutic agent targets immune checkpoint molecules. Certain tumors can evade the immune system by exploiting the immune checkpoint pathway. Thus, targeting immune checkpoints has emerged as an effective approach to block the ability of tumors to evade the immune system and activate anti-tumor immunity against certain cancers. Pardoll, Nature Reviews Cancer, 2012, 12: 252-264.
[0350] In certain embodiments, the immune checkpoint molecule is an inhibitory molecule that reduces signals involved in the T cell response to an antigen. For example, CTLA4 is expressed on T cells and plays a role in the downregulation of T cell activation by binding to CD80 (also known as B7.1) or CD86 (also known as B7.2) on antigen-presenting cells. PD-1 is another inhibitory checkpoint molecule expressed on T cells. PD-1 restricts the activity of T cells in peripheral tissues during an inflammatory response. Additionally, the ligands of PD-1 (PD-L1 or PD-L2) are generally upregulated on the surface of many different tumors, resulting in the downregulation of the anti-tumor immune response in the tumor microenvironment. In certain embodiments, the inhibitory immune checkpoint molecule is CTLA4 or PD-1. In other embodiments, the inhibitory immune checkpoint molecule is a ligand of PD-1, such as PD-L1 or PD-L2. In other embodiments, the inhibitory immune checkpoint molecule is a ligand of CTLA4, such as CD80 or CD86. In other embodiments, the inhibitory immune checkpoint molecule is lymphocyte activation gene 3 (LAG3), killer cell immunoglobulin-like receptor (KIR), T cell membrane protein 3 (TIM3), galectin 9 (GAL9), or adenosine A2a receptor (A2aR).
[0351] Antagonists that target these immune checkpoint molecules can be used to enhance antigen-specific T cell responses against certain cancers. Thus, in certain embodiments, an immunotherapy or immunotherapeutic agent is an antagonist of an inhibitory immune checkpoint molecule. In certain embodiments, the inhibitory immune checkpoint molecule is PD-1. In certain embodiments, the inhibitory immune checkpoint molecule is PD-L1. In certain embodiments, the antagonist of the inhibitory immune checkpoint molecule is an antibody (e.g., a monoclonal antibody). In certain embodiments, the antibody or monoclonal antibody is an anti-CTLA4, anti-PD-1, anti-PD-L1, or anti-PD-L2 antibody. In certain embodiments, the antibody is a monoclonal anti-PD-1 antibody. In some embodiments, the antibody is a monoclonal anti-PD-L1 antibody. In certain embodiments, the monoclonal antibody is a combination of an anti-CTLA4 antibody and an anti-PD-1 antibody, an anti-CTLA4 antibody and an anti-PD-L1 antibody, or an anti-PD-L1 antibody and an anti-PD-1 antibody. In certain embodiments, the anti-PD-1 antibody is one or more of pembrolizumab (Keytruda®) or nivolumab (Opdivo®). In certain embodiments, the anti-CTLA4 antibody is ipilimumab (Yervoy®). In certain embodiments, the anti-PD-L1 antibody is one or more of atezolizumab (Tecentriq®), avelumab (Bavencio®), or durvalumab (Imfinzi®).
[0352] In certain embodiments, the immunotherapy or immunotherapeutic agent is an antagonist (e.g., an antibody) to CD80, CD86, LAG3, KIR, TIM3, GAL9, or A2aR. In other embodiments, the antagonist is a soluble version of an inhibitory immune checkpoint molecule, e.g., a soluble fusion protein comprising the extracellular domain of an inhibitory immune checkpoint molecule and the Fc domain of an antibody. In certain embodiments, the soluble fusion protein comprises the extracellular domain of CTLA4, PD-1, PD-L1, or PD-L2. In some embodiments, the soluble fusion protein comprises the extracellular domain of CD80, CD86, LAG3, KIR, TIM3, GAL9, or A2aR. In one embodiment, the soluble fusion protein comprises the extracellular domain of PD-L2 or LAG3.
[0353] In certain embodiments, the immune checkpoint molecule is a co-stimulatory molecule that amplifies signals involved in the T cell response to an antigen. For example, CD28 is a co-stimulatory receptor expressed on T cells. When a T cell binds to an antigen through its T cell receptor, CD28 binds to CD80 (also known as B7.1) or CD86 (also known as B7.2) on the antigen-presenting cell, amplifying the T cell receptor signaling and promoting T cell activation. Since CD28 binds to the same ligands (CD80 and CD86) as CTLA4, CTLA4 can cancel or regulate the co-stimulatory signaling mediated by CD28. In certain embodiments, the immune checkpoint molecule is a co-stimulatory molecule selected from CD28, inducible T cell co-stimulator (ICOS), CD137, OX40, or CD27. In other embodiments, the immune checkpoint molecule is a ligand of a co-stimulatory molecule, e.g., including CD80, CD86, B7RP1, B7-H3, B7-H4, CD137L, OX40L, or CD70.
[0354] Using agonists that target these co-stimulatory checkpoint molecules can enhance antigen-specific T cell responses against certain cancers. Thus, in certain embodiments, the immunotherapy or immunotherapeutic agent is an agonist of a co-stimulatory checkpoint molecule. In certain embodiments, the agonist of the co-stimulatory checkpoint molecule is an agonist antibody, preferably a monoclonal antibody. In certain embodiments, the agonist antibody or monoclonal antibody is an anti-CD28 antibody. In other embodiments, the agonist antibody or monoclonal antibody is an anti-ICOS, anti-CD137, anti-OX40, or anti-CD27 antibody. In other embodiments, the agonist antibody or monoclonal antibody is an anti-CD80 antibody, anti-CD86 antibody, anti-B7RP1 antibody, anti-B7-H3 antibody, anti-B7-H4 antibody, anti-CD137L antibody, anti-OX40L antibody, or anti-CD70 antibody.
[0355] In certain embodiments, the somatic or germline origin status of nucleic acid variants from a sample from a subject can be compared to a database of comparative results from a reference population to identify a customized or targeted treatment for that subject. Generally, the reference population includes patients having the same cancer or disease type as the subject and / or patients who are receiving or have received the same treatment as the subject. Customized or targeted treatment(s) can be identified if the nucleic acid variant and the comparative result meet certain classification criteria (e.g., substantial or approximate match).
[0356] In certain embodiments, the customized treatments described herein are generally administered parenterally (e.g., intravenously or subcutaneously). Pharmaceutical compositions containing immunotherapeutic agents are generally administered intravenously. Certain therapeutic agents are administered orally. However, customized treatments (e.g., immunotherapeutic agents, etc.) may also be administered by any method known in the art, e.g., buccally, sublingually, rectally, vaginally, intraurethrally, topically, intraocularly, intranasally, and / or intratympanically, and the administration may include tablets, capsules, granules, aqueous suspensions, gels, sprays, suppositories, plasters, ointments, etc.
[0357] Treatment options for treating diseases, disorders, or conditions based on specific genes other than cancer are generally well-known to those of skill in the art and will be apparent in view of the specific disease, disorder, or condition under consideration. IV. Kits
[0358] Kits are also provided that include the compositions described herein. The kits may be useful for practicing the methods described herein. In some embodiments, the kit includes a first reagent for sequence-specific cleavage of a DNA sequence. In some such embodiments, the reagent for sequence-specific cleavage includes a modification-independent sequence-specific nuclease and a plurality of guide RNAs. In some embodiments, the nuclease is a CRISPR nuclease and the one or more guide RNAs are sgRNAs and / or include one or more modifications. In some embodiments, the kit further includes a second reagent for dispensing each sample described herein into a plurality of secondary samples, e.g., any of the dispensing reagents described elsewhere herein. In some embodiments, the reagent for dispensing includes an agent that recognizes a modification in DNA, e.g., a modified nucleobase, e.g., a methylated nucleobase. In some embodiments, the agent that recognizes a modified nucleobase in DNA is an antibody. In some embodiments, the modification is a modified cytosine, e.g., methylcytosine or hydroxymethylcytosine (e.g., labeled hydroxymethylcytosine). In some embodiments, the agent that recognizes a modification in DNA is an antibody specific for methylcytosine in DNA. In some embodiments, the dispensing reagent includes a solid support. In some embodiments, the kit includes a methylation-sensitive nuclease. In some embodiments, the kit includes the first reagent, the second reagent, and, optionally, a methylation-sensitive nuclease and / or other additional elements discussed elsewhere herein. In some embodiments, the kit includes instructions for practicing the methods described herein.
[0359] The kit may include at least 4, 5, 6, 7, or 8 different adapters having distinct molecular barcodes and / or identical sample barcodes. The adapter may not be a sequencing adapter. For example, the adapter does not include an array for a flow cell for sequencing nor an array that enables the formation of a hairpin loop. Different variations and combinations of molecular barcodes and sample barcodes are described throughout and are applicable to the kit. Further, in some cases, the adapter is not a sequencing adapter. Additionally, the adapter provided in the kit may also include a sequencing adapter. The sequencing adapter may include an array that hybridizes to one or more sequencing primers. The sequencing adapter may further include an array that hybridizes to a solid support, such as a flow cell array. For example, the sequencing adapter may be a flow cell adapter. The sequencing adapter may be attached to one or both ends of a polynucleotide fragment. In some cases, the kit may include at least 8 different adapters having distinct molecular barcodes and identical sample barcodes. The adapter may not be a sequencing adapter. The kit may further include a sequencing adapter having a first array that selectively hybridizes to the adapter and a second array that selectively hybridizes to the flow cell array. In another example, the sequencing adapter may be hairpin-shaped. For example, the hairpin-shaped adapter may include a complementary double-stranded portion and a loop portion, where the double-stranded portion may be bound (e.g., ligated) to a double-stranded polynucleotide. A hairpin-shaped sequencing adapter may be attached to both ends of a polynucleotide fragment to generate a circular molecule that can be sequenced multiple times.The sequencing adapter may be from end to end, up to 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, or more bases. The sequencing adapter may from end to end, include 20 - 30 bases, 20 - 40 bases, 30 - 50 bases, 30 - 60 bases, 40 - 60 bases, 40 - 70 bases, 50 - 60 bases, 50 - 70 bases. In certain examples, the sequencing adapter may from end to end, include 20 - 30 bases. In another example, the sequencing adapter may from end to end, include 50 - 60 bases. The sequencing adapter may include one or more barcodes. For example, the sequencing adapter may include a sample barcode. The sample barcode may include a predetermined sequence. The sample barcode can be used to identify the source of the polynucleotide. The sample barcode may be at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, or more (or any length described throughout) nucleic acid bases, e.g., at least 8 bases. The barcode may, as described above, be a continuous sequence or a non - continuous sequence.
[0360] The adapter may be blunt - ended and Y - shaped and may be less than or equal to 40 nucleic acid bases in length. Other variations can be found throughout and are applicable to the kit.
[0361] The kit may further include a plurality of oligonucleotide probes that selectively hybridize to at least 5, 6, 7, 8, 9, 10, 20, 30, 40, or all of the genes selected from the group consisting of ALK, APC, BRAF, CDKN2A, EGFR, ERBB2, FBXW7, KRAS, MYC, NOTCH1, NRAS, PIK3CA, PTEN, RBI, TP53, MET, AR, ABL1, AKT1, ATM, CDH1, CSF1R, CTNNB1, ERBB4, EZH2, FGFR1, FGFR2, FGFR3, FLT3, GNA11, GNAQ, GNAS, HNF1A, HRAS, IDH1, IDH2, JAK2, JAK3, KDR, KIT, MLH1, MPL, NPM1, PDGFRA, PROC, PTPN11, RET, SMAD4, SMARCB1, SMO, SRC, STK11, VHL, TERT, CCND1, CDK4, CDKN2B, RAF1, BRCA1, CCND2, CDK6, NF1, TP53, ARID1A, BRCA2, CCNE1, ESR1, RIT1, GATA3, MAP2K1, RHEB, ROS1, ARAF, MAP2K2, NFE2L2, RHOA, and NTRK1. The number of genes to which the oligonucleotide probes can selectively hybridize may vary. For example, the number of genes may include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, or 54. The kit may include a container containing the plurality of oligonucleotide probes and instructions for performing any of the methods described herein.
[0362] All patents, patent applications, websites, other publications or documents, accession numbers, etc. cited in this specification are incorporated by reference in their entirety for all purposes to the same extent as if each individual item was specifically and individually indicated to be incorporated by reference. Where different versions of a sequence are associated with an accession number at different times, it means the version associated with the accession number on the effective filing date of this application. The effective filing date means either the actual filing date or, if applicable, the earlier of the filing dates of the priority applications to which the accession number is referred. Similarly, unless otherwise specified, where different versions of a publication, website, etc. are published at different times, it means the version published most recently on the effective filing date of this application.
Example
[0363] (Example 1) Analysis of cfDNA for detecting the presence or absence of tumors in a subject The workflow described in this example is illustrated in Figure 3. A set of patient samples is analyzed by a blood-based NGS assay at Guardant Health (Redwood City, CA, USA) to detect the presence or absence of cancer. cfDNA is extracted from the plasma of these patients. The cfDNA of the patient samples is then combined with magnetic beads conjugated to MBD protein and an appropriate buffer and incubated overnight. During this incubation, methylated cfDNA binds to the MBD protein. Unmethylated or less methylated DNA is washed off the beads using buffers containing increasing concentrations of salt. Finally, a buffer with a high salt concentration is used to wash off the heavily methylated DNA from the MBD protein. The combination of the unbound DNA and these washes results in at least three fractions of gradually methylated cfDNA (low methylation fraction, residual methylation fraction, and high methylation fraction). The cfDNA molecules within the fractions are washed to remove salt and concentrated in preparation for the enzymatic steps of library preparation.
[0364] After concentrating cfDNA within an aliquot, a first adapter is added by ligating it to the cfDNA at its 3’ end. The adapter is used as a priming site for second strand synthesis using a universal primer and DNA polymerase. Next, a second adapter is ligated to the 3’ end of the second strand, which has then become a double-stranded molecule at that point. These adapters contain non-unique molecular barcodes, and for each aliquot, an adapter having a non-unique molecular barcode that is distinguishable from the barcodes of the adapters used within other aliquots is ligated.
[0365] After ligation, the DNA is washed, concentrated, and then enriched. Once concentrated, the DNA of the hypermethylated aliquot is combined with a methylation-sensitive restriction enzyme (MSRE) for digestion of non-specifically aliquoted DNA (e.g., unmethylated DNA within the hypermethylated aliquot). Next, the digested DNA is combined with Cas9 nuclease and a plurality of sgRNAs specific to sequences containing methylation modifications commonly found in DNA from healthy subjects, such that sequences of hypermethylated aliquots that are expected to be present in a sample obtained from a healthy subject are degraded. The remaining hypermethylated DNA, in which sequences not commonly found in methylated form in healthy subjects are enriched at that point, is amplified by PCR.
[0366] Optionally, the hypomethylated DNA is amplified by PCR, then washed and concentrated. Once concentrated, the amplified DNA is combined with a biotinylated RNA probe of a set of target regions, including a salt buffer and a probe for a set of sequence variable target region probes and / or a set of epigenetic target region probes, to enrich for the target region of interest. Probes for the set of epigenetic target regions include oligonucleotides targeting hypomethylated variable target regions (e.g., sequences that are hypomethylated in cancer and / or not normally hypomethylated in healthy cfDNA), CTCF binding target regions, transcription start site target regions, local amplification target regions, and / or methylation control regions. The mixture is incubated overnight. The biotinylated RNA probe and any hybridized DNA are captured by streptavidin magnetic beads and separated from the DNA that was not captured by a series of salt-based washes, thereby enriching the sample. The enriched hypomethylated DNA can then be combined with the amplified DNA from hypermethylated DNA enriched for sequences not commonly found in methylated form in healthy subjects.
[0367] Sequencing is performed using a next-generation sequencer (e.g., Illumina NovaSeq sequencer). Subsequently, the sequence reads generated by the sequencer are analyzed using bioinformatics tools / algorithms. Molecular barcodes are used for identifying unique molecules and for deconvolution of differentially distributed samples. Analyze the sequences of unique molecules to call SNVs, insertions, deletions, and fusions, etc. genomic alterations with sufficient support to distinguish actual tumor variants from technical errors (e.g., PCR errors, sequencing errors). Independently analyze the sequences of molecules from high methylation profiling to detect methylated cfDNA molecules in regions shown to be differentially methylated in cancer compared to normal cells and / or regions shown to be differentially methylated in cell types that do not substantially contribute to cfDNA in healthy individuals. Finally, combine the results of both analyses to yield a final tumor presence / absence call. (Example 2) Analysis of cfDNA without the step of distribution
[0368] Cell-free DNA is extracted from a set of patient samples and adapters are added to the cfDNA basically as described in Example 1. After ligation, the DNA is washed, concentrated, and then enriched. Once concentrated, the DNA is combined with MSRE to degrade unmethylated DNA. After MSRE treatment, the DNA is then combined with Cas9 nuclease and multiple sgRNAs. Some of the multiple sgRNAs are specific for sequences containing methyl modifications commonly found in DNA from healthy subjects. The other multiple sgRNAs are specific for sequences that do not contain CpG dinucleotides. Thus, sequences expected to be highly methylated in samples obtained from healthy subjects and sequences that do not contain certain motifs that can be methylated are degraded. Amplify the remaining DNA by PCR, which is enriched at that time for sequences not commonly found in methylated form in healthy subjects.
[0369] After enrichment of sequences not commonly found in methylated form in healthy subjects, an aliquot of the enriched sample is sequenced using a next-generation sequencer (e.g., an Illumina NovaSeq sequencer). The sequence reads generated by the sequencer are then analyzed using bioinformatics tools / algorithms. Molecular barcodes are used to identify unique molecules. Tumor presence / absence calls are made based at least in part on the presence and / or amount of methylated cfDNA molecules in regions shown to be differentially methylated in cancer compared to normal cells and / or regions shown to be differentially methylated in cell types that do not substantially contribute to cfDNA in healthy individuals. (Example 3) Analysis of single-site methylation in cfDNA
[0370] As described in Example 1 or Example 2, a set of patient samples is analyzed at Guardant Health (Redwood City, CA, USA) by a blood-based NGS assay with an additional step of subjecting the DNA of the sample or distributed secondary samples to a procedure that modifies the first nucleobase in a form different from the second nucleobase to detect the presence or absence of cancer. For example, prior to Cas9 treatment, the DNA is treated with bisulfite to convert unmethylated cytosine to thymine. To disrupt sequences that were originally unmethylated, multiple sgRNAs are included with sequences specific to sequences containing the converted thymine. In another example, the DNA is treated with bisulfite after Cas9 treatment to identify positions of abnormal cytosine methylation or positions of cytosine not methylated in sequences commonly found in cfDNA from healthy subjects. (Example 4) Additional methods for analyzing cfDNA
[0371] As described in Examples 1, 2, or 3, a set of patient samples was analyzed at Guardant Health (Redwood City, CA, USA) using a blood-based NGS assay with the additional feature that the sgRNAs included sequences specific for unique junctions and a portion of molecular barcodes generated by adapter dimer ligation events, to detect the presence or absence of cancer.
Claims
1. A method for analyzing DNA in a sample, comprising: a) distributing the sample into a plurality of secondary samples by contacting the DNA with an agent that recognizes a modification associated with the DNA, wherein the plurality of secondary samples includes a first secondary sample and a second secondary sample, and the first secondary sample includes a higher proportion of DNA with the modification associated therewith than the second secondary sample; b) sequence-specifically degrading a plurality of DNA sequences in the first secondary sample that contain the modification and are commonly found in cell-free DNA (cfDNA) derived from a healthy subject, comprising contacting the first secondary sample with a modification-independent sequence-specific nuclease to thereby produce a processed sample; c) detecting the presence or absence of one or more DNA sequences with the modification associated therewith in the processed sample. A method comprising the above steps.
2. The method according to claim 1, wherein the distributing step is performed before the degrading step.
3. A method for analyzing cfDNA in a sample, comprising: a) contacting the sample or a secondary sample thereof with an MSRE to thereby degrade DNA containing an unmethylated recognition site of the MSRE; b) sequence-specifically degrading a plurality of DNA sequences having a methylated sequence commonly found in cfDNA derived from a healthy subject and a plurality of sequences lacking a CpG motif, comprising contacting the sample with a modification-independent sequence-specific nuclease to thereby produce a processed sample; c) detecting the presence or absence of one or more DNA sequences in the processed sample. A method comprising the above steps.
4. The method according to the immediately preceding claim, wherein the methylation includes cytosine methylation.
5. The method according to claim 3 or claim 4, further comprising distributing the sample into a plurality of secondary samples by contacting the DNA with an agent that recognizes a modification associated with the DNA, wherein the plurality of secondary samples includes a first secondary sample and a second secondary sample, and the first secondary sample includes a higher proportion of DNA with the modification associated therewith than the second secondary sample.
6. The method according to any one of the preceding claims, wherein the distributing step includes distributing based on the methylation level of the DNA.
7. The method according to any one of the preceding claims, wherein the modification is methylation.
8. The method according to the immediately preceding claim, wherein the methylation comprises cytosine methylation.
9. The method according to any one of claims 1, 2, or 5 to 8, wherein the distributing step comprises distributing based on the hydroxymethylation level of the DNA.
10. The method according to the immediately preceding claim, wherein hydroxymethyl is labeled prior to the distributing step, and optionally, the label comprises biotin, glucosyl, or sulfonyl.
11. The method according to any one of claims 1 to 3 or 7 to 10, wherein the modification is hydroxymethylation.
12. The method according to any one of the preceding claims, wherein the agent that recognizes the modification associated with the DNA is a methyl, hydroxymethyl, or labeled hydroxymethyl-binding reagent.
13. The method according to the immediately preceding claim, wherein the methyl, hydroxymethyl, or labeled hydroxymethyl-binding reagent is an antibody.
14. The method according to claim 12 or 13, wherein the methyl-binding reagent specifically recognizes 5-methylcytosine, 5-hydroxymethylcytosine, biotinylated 5-hydroxymethylcytosine, glucosylated 5-hydroxymethylcytosine, or sulfonylated 5-hydroxymethylcytosine.
15. The method according to any one of claims 12 to 14, wherein the methyl-binding reagent is immobilized on a solid support.
16. The method according to any one of the preceding claims, wherein the distributing step comprises immunoprecipitation of methylated, hydroxymethylated, or labeled hydroxymethylated DNA.
17. The method according to any one of claims 1, 2, or 5 to 16, wherein the distributing step comprises distributing based on binding to a protein, and optionally, the protein is a methylated protein, acetylated protein, non-methylated protein, non-acetylated protein, and / or optionally, the protein is a histone.
18. The method according to the immediately preceding claim, wherein the distributing step comprises contacting the collected cfDNA with a binding reagent that is specific for the protein and immobilized on a solid support.
19. The method according to any one of the preceding claims, wherein the step of digesting comprises contacting the second secondary sample with a modification-independent sequence-specific nuclease, thereby producing a second processed sample.
20. The method according to any one of claims 1, 2, or 5 to 19, comprising the step of contacting one or more of the plurality of secondary samples with a methylation-sensitive restriction enzyme (MSRE), thereby digesting DNA containing an unmethylated recognition site of the MSRE.
21. The method according to the immediately preceding claim, wherein the step of contacting one or more of the plurality of secondary samples with the MSRE is performed after the step of distributing.
22. The method according to claim 20 or 21, wherein the step of contacting one or more of the plurality of secondary samples with the MSRE is performed before the step of digesting.
23. The method according to claim 20 or 21, wherein the step of contacting one or more of the plurality of secondary samples with the MSRE is performed after the step of digesting.
24. The method according to claim 20 or 21, wherein the step of contacting one or more of the plurality of secondary samples with the MSRE is performed simultaneously with the step of digesting.
25. The method according to any one of claims 21 to 24, wherein the first secondary sample is contacted with the MSRE.
26. The method according to any one of claims 20 to 25, wherein the step of contacting a sample or a secondary sample with the MSRE is performed before the step of sequence-specific digestion.
27. The method according to any one of the preceding claims, wherein the modification-independent sequence-specific nuclease is a CRISPR nuclease.
28. The method according to the immediately preceding claim, wherein the CRISPR nuclease is a Cas12a, Cas12b, or CasX nuclease.
29. The method according to claim 27, wherein the CRISPR nuclease is a Cas9 nuclease.
30. The method according to the immediately preceding claim, wherein the Cas9 nuclease is a multi-turnover Cas9 nuclease.
31. The method according to claim 29, wherein the Cas9 nuclease is a Streptococcus pyogenes Cas9 nuclease or a variant thereof.
32. The method according to claim 29, wherein the Cas9 nuclease is a Staphylococcus aureus Cas9 nuclease or a variant thereof.
33. The method according to any one of claims 29 to 32, wherein the Cas9 nuclease is a high-fidelity variant.
34. The method according to any one of the preceding claims, wherein the step of specifically cleaving the sequence comprises contacting the DNA with a plurality of guide RNAs.
35. The method according to the immediately preceding claim, wherein at least one guide RNA comprises one or more modifications.
36. The method according to the immediately preceding claim, wherein the one or more modifications comprise phosphorothioate internucleotide linkages, 2'-substitutions, or UNA, LNA, cEt, or ENA nucleotide sugars.
37. The method according to the immediately preceding claim, wherein the 2'-substitution is 2'-fluoro, 2'-hydro, 2'-O-methoxyethyl, or 2'-O-alkyl.
38. The method according to any one of claims 34 to 37, wherein at least one guide RNA is a sgRNA.
39. The method according to any one of claims 34 to 38, wherein at least one guide RNA specifically binds to DNA containing a CpG motif that is methylated in cfDNA derived from healthy tissue or a healthy subject.
40. The method according to any one of claims 34 to 39, wherein at least one guide RNA specifically binds to a DNA sequence lacking CpG dinucleotides.
41. The method according to any one of claims 34 to 39, wherein each guide RNA of the plurality of guide RNAs comprises the modification and is configured to specifically bind to a DNA sequence commonly found in cell-free DNA (cfDNA) derived from a healthy subject.
42. The method according to any one of claims 34 to 40, wherein each guide RNA of the plurality of guide RNAs comprises the modification and is configured to specifically bind to a DNA sequence commonly found in cell-free DNA (cfDNA) derived from a healthy subject or a DNA sequence lacking CpG dinucleotides.
43. The method according to any one of claims 1 to 26, 34 to 37, or 39 to 42, wherein the modification-independent sequence-specific nuclease is an Argonaute nuclease.
44. The method according to any one of claims 1 to 26, 34 to 37, or 39 to 42, wherein the modification-independent sequence-specific nuclease is a zinc finger nuclease.
45. The method according to any one of claims 1 to 26, 34 to 37, or 39 to 42, wherein the modification-independent sequence-specific nuclease is a TALEN.
46. The method according to any one of the preceding claims, wherein the detecting step comprises sequencing.
47. The method according to the preceding claim, wherein the detecting step comprises sequencing a plurality of target regions within at least one set of target regions.
48. The method according to any one of the preceding claims, further comprising enriching one or more of the plurality of target regions within at least one set of target regions.
49. The method according to the preceding claim, wherein the enriching step comprises contacting the DNA with a target-specific probe specific for the one or more of the plurality of target regions within at least one set of target regions.
50. The method according to any one of claims 47 to 49, wherein the at least one set of target regions comprises target regions that are not commonly found in methylated form in cfDNA from a healthy subject or not commonly found in methylated form in healthy tissue.
51. The method according to claims 47 to 50, wherein the at least one set of target regions comprises target regions that are commonly found in methylated form in a tissue and that do not substantially contribute to cfDNA in a healthy subject.
52. The method according to any one of claims 47 to 51, wherein the at least one set of target regions comprises target regions that are commonly found in methylated form in a cancerous tissue.
53. The method according to any one of claims 47 to 52, wherein the at least one set of target regions comprises a set of sequence-variable target regions and a set of epigenetic target regions.
54. The method according to any one of claims 47 to 53, wherein the at least one set of target regions comprises a set of hypermethylated variable target regions.
55. The method according to the preceding claim, wherein the set of hypermethylated variable target regions comprises regions in which the degree of methylation in at least one tissue type is higher than the degree of methylation in cfDNA from a healthy subject.
56. The method according to any one of claims 47 to 55, wherein the at least one set of target regions comprises a set of hypomethylated variable target regions.
57. The method according to the immediately preceding claim, wherein the set of hypomethylated variable target regions comprises regions where the degree of methylation in at least one tissue type is lower than the degree of methylation in cfDNA from a healthy subject.
58. The method according to any one of claims 47 to 57, wherein the at least one set of target regions comprises a set of methylated control target regions.
59. The method according to any one of claims 47 to 57, wherein the at least one set of target regions comprises a set of fragmented variable target regions.
60. The method according to the immediately preceding claim, wherein the set of fragmented variable target regions comprises a transcription start site region.
61. The method according to claim 59 or 60, wherein the set of fragmented variable target regions comprises a CTCF binding region.
62. The method according to any one of claims 53 to 61, wherein the set of sequence variable target regions comprises at least one sequence not commonly found in cfDNA from a healthy subject.
63. The method according to any one of claims 46 to 62, wherein the sequencing comprises sequencing a gene or a part thereof selected from Table 1, Table 2, Table 3, Table 4, and / or Table 5.
64. The method according to any one of claims 46 to 63, wherein the sequencing comprises sequencing all of the DNA sequences in the processed sample.
65. The method according to any one of the preceding claims, wherein 20 to 250,000 sequences are resolved.
66. The method according to any one of the preceding claims, wherein 50 to 100,000 sequences are resolved.
67. The method according to the immediately preceding claim, wherein 100 to 10,000 sequences are resolved.
68. The method according to any one of the preceding claims, wherein the sequences to be resolved contain repetitive elements, and optionally, the repetitive elements contain SINE, LINE, and / or Alu elements.
69. The method according to any one of claims 1 to 45, wherein the detecting step comprises performing qPCR.
70. The method according to any one of the preceding claims, wherein the DNA is collected from a test subject.
71. The method according to any one of the preceding claims, wherein the DNA comprises cfDNA obtained from a test subject.
72. The method according to claim 70 or 71, wherein the DNA comprises DNA obtained from a tissue sample of the test subject.
73. The method according to the immediately preceding claim, wherein the tissue sample is a biopsy material, a fine needle aspirate, or a formalin-fixed paraffin-embedded tissue sample.
74. The method according to any one of the preceding claims, further comprising a step of ligating a barcode-containing adapter to the DNA, optionally before or simultaneously with the amplification of the DNA.
75. The method according to the immediately preceding claim, wherein the plurality of sequences specifically cleaved comprises sequences comprising a barcode-containing adapter dimer junction.
76. The method according to the immediately preceding claim, comprising contacting the DNA with a plurality of guide RNAs configured to specifically bind to each of the possible adapter dimer junctions.
77. The method according to any one of the preceding claims, wherein the DNA is amplified before the step of detecting.
78. The step of specifically cleaving the sequence is after the step of ligating a barcode-containing adapter to the DNA and (a) before the step of partitioning the sample into a plurality of secondary samples by contacting the DNA with an agent that recognizes a modification associated with the DNA, (b) after the step of partitioning the sample into a plurality of secondary samples by contacting the DNA with an agent that recognizes a modification associated with the DNA, (c) before the step of contacting the sample or its secondary sample with MSRE, (d) after the step of contacting the sample or its secondary sample with MSRE, (e) before the step of amplifying the DNA before the step of detecting, (f) after the step of amplifying the DNA before the step of detecting, or The method according to any one of the preceding claims, which is performed by any combination of: (g) one or two of (a) and any one or two of (c) to (f); (b) and any one or two of (c) to (f); (c) and any one or two of (a), (b), (e), and (f); (d) and any one or two of (a), (b), (e), and (f); (e) and any one or two of (a) to (d); or (f) and any one or two of (a) to (d).
79. The method according to any one of the preceding claims, comprising the step of dispensing the sample, wherein the DNA molecules from the first secondary sample and the DNA molecules from the second secondary sample are differentially tagged.
80. The method according to the preceding claim, wherein the DNA molecules from the first secondary sample and the DNA molecules from the second secondary sample are sequenced in the same sequencing cell.
81. The method according to claim 79 or 80, comprising the step of differentially tagging and pooling the first secondary sample and the second secondary sample.
82. The method according to any one of claims 79 to 81, wherein the DNA of the first secondary sample and the DNA of the second secondary sample are differentially tagged, and after differential tagging, a part of the DNA from the second secondary sample is added to the first secondary sample or at least a part thereof, thereby forming a pool.
83. The method according to the preceding claim, wherein the pool comprises about 45% or less, about 40% or less, about 35% or less, about 30% or less, about 25% or less, about 20% or less, about 15% or less, about 10% or less, or about 5% or less of the DNA of the second secondary sample.
84. The method according to the preceding claim, wherein the pool comprises about 70 - 90%, about 75 - 85%, or about 80% of the DNA of the second secondary sample.
85. The method according to any one of claims 82 to 84, wherein the pool comprises substantially all of the DNA of the first secondary sample.
86. A step of distributing the sample into a plurality of secondary samples, wherein the plurality of secondary samples include a third secondary sample having a higher proportion of DNA having cytosine modification than the second secondary sample but a lower proportion than the first secondary sample, the method according to any one of the preceding claims.
87. The method according to the immediately preceding claim, further comprising the step of differentially tagging the third secondary sample.
88. Before the step of sequence-specifically degrading the plurality of DNA sequences in the first secondary sample that contain the modification and are generally found in cell-free DNA, the first secondary sample is enriched for one or more target region sets, the method according to any one of the preceding claims.
89. Before the step of sequence-specifically degrading the plurality of DNA sequences that contain the modification and are generally found in cell-free DNA, a plurality of first secondary samples are pooled, and optionally, the plurality of first secondary samples are from different subjects and / or are differentially tagged with sample tags, the method according to any one of the preceding claims.
90. The method according to any one of the preceding claims, further comprising the step of determining the likelihood that the subject has cancer.
91. By the sequencing, a plurality of sequencing reads are generated, and the method further comprises the steps of mapping the plurality of sequence reads to one or more reference sequences to generate mapped sequence reads, and processing the mapped sequence reads corresponding to the set of sequence variable target regions and the set of epigenetic target regions to determine the likelihood that the subject has cancer, the method according to any one of claims 45 to 90.
92. The method according to any one of claims 70 to 91, wherein the test subject has been previously diagnosed with cancer and has received one or more previous cancer treatments.
93. The cfDNA is obtained at one or more preselected time points after the one or more previous cancer treatments, and the detecting step comprises sequencing the DNA sequence, thereby providing a set of sequence information, the method according to the immediately preceding claim.
94. The method according to the immediately preceding claim, further comprising the step of detecting the presence or absence of DNA originating from or derived from tumor cells using the set of array information at a preselected point in time. **Claim 95** The method according to the immediately preceding claim, further comprising the step of determining, for the subject, a cancer recurrence score indicative of the presence or absence of the DNA originating from or derived from the tumor cells, and optionally further comprising the step of determining a cancer recurrence status based on the cancer recurrence score, wherein if the cancer recurrence score is determined to be at or above a predetermined threshold, the cancer recurrence status of the subject is determined to be at risk of cancer recurrence, or if the cancer recurrence score is below the predetermined threshold, the cancer recurrence status of the subject is determined to be at low risk of cancer recurrence. **Claim 96** The method according to the immediately preceding claim, further comprising the step of comparing the cancer recurrence score of the subject with a predetermined cancer recurrence threshold, wherein if the subject has a cancer recurrence score that exceeds the cancer recurrence threshold, the subject is classified as a candidate for subsequent cancer treatment, or if the cancer recurrence score is below the cancer recurrence threshold, the subject is classified as not being a candidate for subsequent cancer treatment.