Methods for Analysing Cytosine Methylation and Hydroxymethylation

JP2025507763A5Pending Publication Date: 2026-03-10GUARDANT HEALTH INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-03-01
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

It is difficult to develop efficient and sensitive methods for analyzing methylation and hydroxymethylation in cellular free DNA, especially in low concentration and heterogeneous DNA samples.

Method used

A library preparation workflow is adopted to generate DNA libraries that do not reduce sequencing efficiency by synthesizing the first and second complementary strands, glycosylated 5-hydroxymethylated cytosines, methylated cytosines, deaminated unmethylated cytosines, and serialization treatment methods, and labeling them through a Y-type adapter or bubble adapter for subsequent sequencing analysis.

Benefits of technology

This method can effectively retain the encoding information and pairing analysis capabilities in the DNA library, and improve the analysis accuracy and sensitivity of DNA methylation and hydroxymethylation, especially in the detection capabilities in low-concentration and heterogeneous DNA samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The present disclosure aims to meet the need for improved analysis of DNA, e.g., cell-free DNA, and / or provide other benefits. In some embodiments, the present disclosure provides a library preparation workflow that does not reduce assay efficiency and generates libraries that retain the encoded information and pairing resolution of the "original" and "copy" strands. The present disclosure provides methods and systems for reducing the signal-to-noise ratio of methylation splitting assays.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Patent Application No. 63 / 268,750, filed March 1, 2022, which is hereby incorporated by reference herein for all purposes.

[0002] FIELD OF THEINVENTION The present disclosure provides compositions and methods for analyzing DNA, for example cell-free DNA, for cytosine methylation and hydroxymethylation.In some embodiments, cell-free DNA is from a subject who has or is suspected of having cancer, and / or cell-free DNA includes DNA from cancer cells.In some embodiments, DNA is glycosylated. [Background technology]

[0003] Introduction and Overview Cancer causes millions of deaths per year worldwide. Early detection of cancer can result in improved outcomes, as early stage cancers tend to be more susceptible to treatment.

[0004] Inappropriately controlled cell growth is a hallmark of cancer, which generally results from the accumulation of genetic and epigenetic alterations, including copy number variations (CNVs), single nucleotide variations (SNVs), gene fusions, insertions and / or deletions (indels), cytosine modifications (e.g., 5-methylcytosine, 5-hydroxymethylcytosine, and other more oxidized forms), and the association of DNA with chromatin proteins and transcription factors.

[0005] Biopsy represents a traditional approach to detecting or diagnosing cancer in which cells or tissue are extracted from a possible cancer site and analyzed for relevant phenotypic and / or genotypic traits. Biopsies have the disadvantage of being invasive.

[0006] Cancer detection based on the analysis of bodily fluids such as blood ("liquid biopsy") is an interesting alternative based on the observation that DNA from cancer cells is released into bodily fluids. Liquid biopsy is non-invasive (sometimes only requiring blood sampling). Current methods of cancer diagnostic assays of cell-free nucleic acids (e.g., cell-free DNA or cell-free RNA) may focus on the detection of tumor-associated somatic variants, including single nucleotide variants (SNVs), copy number variations (CNVs), fusions and indels (i.e., insertions or deletions), all of which are mainstream targets for liquid biopsies. There is growing evidence that non-sequence modifications in cell-free DNA, such as methylation status and fragmentomic signals, can provide information about the source of cell-free DNA and disease levels. Furthermore, different types of modifications, such as 5-methylation and 5-hydroxymethylation, may have various implications regarding the presence or absence of disease. Detailed knowledge of non-sequence modifications of cell-free DNA (e.g., when combined with somatic mutation calling) can improve the assessment of tumor status. However, it has been difficult to develop accurate and sensitive methods for analyzing DNA from sources, such as liquid biopsies, that take into account the low concentration and heterogeneity of cell-free DNA and provide detailed information on distinct nucleobase modifications, such as methylation and hydroxymethylation.

[0007] Recently, a complex single-site methylation workflow has been developed that encodes information about a copied DNA strand that is physically linked to the original strand via a hairpin DNA structure in the library preparation process (WO2022023753A1, US11078529B2). Both the original strand and the copied strand are sequenced together in a single NGS read. However, the hairpin nature of DNA library molecules is prone to self-hybridization, which poses challenges to various assay efficiencies, e.g., amplification, hybridization-based target enrichment, and sequencing itself. Thus, there is a need for improved methods and compositions for analyzing DNA samples, such as cell-free DNA, for example, in liquid biopsies. [Prior art documents] [Patent documents]

[0008] [Patent Document 1] International Publication No. 2022023753 [Patent Document 2] U.S. Pat. No. 1,107,8529 Summary of the Invention

[0009] The present disclosure aims to meet the need for improved analysis of DNA, e.g., cell-free DNA, and / or provide other benefits. In some embodiments, the present disclosure provides a library preparation workflow that does not reduce assay efficiency and produces libraries that retain the encoded information and pairing resolution of the "original" and "copy" strands. The following exemplary embodiments are provided:

[0010] Embodiment 1 is as follows. 1. A method for analyzing a DNA molecule in a sample, the DNA molecule comprising a first and a second strand and an asymmetric adaptor, the method comprising: a) synthesizing a first complementary strand that is complementary to the first strand and a second complementary strand that is complementary to the second strand; b) glycosylating 5-hydroxymethylated cytosines in at least one of the first or second strands before or after synthesizing said first and second complementary strands; c) methylating cytosines in at least one of the first complementary strand or the second complementary strand, said methylation converting hemi-methylated CpGs to fully methylated CpGs; d) deaminating unmodified cytosines in at least one of the first or second strands, thereby producing a treated DNA molecule; and e) sequencing at least a portion of the treated DNA molecules. Includes; Optionally, the asymmetric adaptor is a Y-type adaptor or a bubble adaptor.

[0011] The second embodiment is as follows. 1. A method for analyzing DNA molecules in a sample, the DNA molecules comprising first and second strands and asymmetric adaptors, at least one asymmetric adaptor comprising a deamination-sensitive cytosine, the method comprising: a) synthesizing a first complementary strand that is complementary to the first strand and a second complementary strand that is complementary to the second strand; b) glycosylating 5-hydroxymethylated cytosines in at least one of the first or second strands before or after synthesizing said first and second complementary strands; c) methylating cytosines in at least one of the first complementary strand or the second complementary strand, said methylation converting hemi-methylated CpGs to fully methylated CpGs; d) deaminating unmodified cytosines in at least one of the first or second strands, thereby producing a treated DNA molecule; and e) sequencing at least a portion of the treated DNA molecules. Includes; Optionally, the asymmetric adaptor is a Y-type adaptor or a bubble adaptor.

[0012] The third embodiment is as follows. 3. The method of embodiment 1 or embodiment 2, wherein each asymmetric adapter comprises at least one deamination-sensitive cytosine, and / or said deamination-sensitive cytosine is an unmethylated cytosine.

[0013] The fourth embodiment is as follows. 2. The method of any one of the preceding embodiments, wherein each asymmetric adaptor comprises one deamination-sensitive cytosine and at least one deamination-resistant cytosine, and optionally, the deamination-resistant cytosine is a 5-methylcytosine, and / or each cytosine other than the one deamination-sensitive cytosine in each asymmetric adaptor is a deamination-resistant cytosine.

[0014] Embodiment 5 is as follows. 5. The method of any one of embodiments 2 to 4, wherein the nucleotide immediately 3' to the deamination susceptible cytosine comprises a nucleobase other than guanine, and optionally the non-guanine nucleobase is adenine, cytosine, thymine or uracil.

[0015] Embodiment 6 is as follows. 2. The method of any one of the preceding embodiments, wherein the step of deaminating unmodified cytosines comprises bisulfite conversion.

[0016] Embodiment 7 is as follows. 1. A method for analyzing a DNA molecule in a sample, the DNA molecule comprising a first and a second strand and an asymmetric adaptor, the method comprising: a) oxidizing 5-hydroxymethylated cytosines in at least one first or second strand to 5-formylcytosines; b) synthesizing a first complementary strand that is complementary to the first strand and a second complementary strand that is complementary to the second strand; c) methylating cytosines in at least one of the first complementary strand or the second complementary strand, said methylation converting hemi-methylated CpGs to fully methylated CpGs; d) converting the modified cytosine in at least one of the first or second strands to a thymine or a base that is read as a thymine, thereby producing a processed DNA molecule; and e) sequencing at least a portion of the treated DNA molecules. Includes; Optionally, the asymmetric adaptor is a Y-type adaptor or a bubble adaptor.

[0017] Embodiment 8 is as follows. 1. A method for analyzing DNA molecules in a sample, the DNA molecules comprising first and second strands and asymmetric adaptors, at least one asymmetric adaptor comprising an unmodified cytosine, the method comprising: a) oxidizing 5-hydroxymethylated cytosines in at least one first or second strand to 5-formylcytosines; b) synthesizing a first complementary strand that is complementary to the first strand and a second complementary strand that is complementary to the second strand; c) methylating cytosines in at least one of the first complementary strand or the second complementary strand, said methylation converting hemi-methylated CpGs to fully methylated CpGs; d) converting at least one modified cytosine in the first or second strand to a thymine or a base that is read as a thymine; and e) sequencing at least a portion of the treated DNA molecules. Includes; Optionally, the asymmetric adaptor is a Y-type adaptor or a bubble adaptor.

[0018] Embodiment 9 is as follows. 13. The method of any one of the preceding embodiments, wherein steps a)-e) are performed in the order a) through e).

[0019] Embodiment 10 is as follows. 10. The method of any one of embodiments 8 or 9, wherein converting the modified cytosine in at least one of the first or second strands to a thymine or a base that is read as a thymine comprises oxidizing a hydroxymethylcytosine.

[0020] Embodiment 11 is as follows. The method of the immediately preceding embodiment, wherein said hydroxymethylcytosine is oxidized to formylcytosine.

[0021] Embodiment 12 is as follows. Oxidizing the hydroxymethylcytosine to formylcytosine comprises contacting the hydroxymethylcytosine with a ruthenium salt, and optionally the ruthenium salt is KRuO 4 3. The method of claim 2, wherein

[0022] Embodiment 13 is as follows. 13. The method of any one of embodiments 7 to 12, wherein the modified cytosine is converted to thymine, uracil, or dihydrouracil.

[0023] Embodiment 14 is as follows. 14. The method according to any one of embodiments 7 to 13, comprising, as part of the step of converting said modified cytosine in at least one first or second strand to a thymine or a base that is read as a thymine, converting formylcytosine and / or methylcytosine to a carboxylcytosine.

[0024] Embodiment 15 is as follows. 5. The method of claim 1, wherein the step of converting the formylcytosine and / or the methylcytosine to a carboxyl cytosine comprises contacting the formylcytosine and / or the methylcytosine with a TET enzyme, optionally wherein the TET enzyme is TET1, TET2 or TET3.

[0025] Embodiment 16 is as follows. 16. The method of any one of embodiments 14 to 15, comprising reducing the carboxyl cytosine as part of converting the modified cytosine in at least one of the first or second strands to a thymine or a base that is read as a thymine.

[0026] Embodiment 17 is as follows. The method of the immediately preceding embodiment, wherein said carboxyl cytosine is reduced to dihydrouracil.

[0027] Embodiment 18 is as follows. 18. The method of any one of embodiments 16 to 17, wherein the step of reducing the carboxyl cytosine comprises contacting the carboxyl cytosine with a reducing agent, optionally wherein the reducing agent is a borane or borohydride reducing agent.

[0028] Embodiment 19 is as follows. The borane or borohydride reducing agent may be pyridine borane, 2-picoline borane, borane, tert-butylamine borane, ammonia borane, sodium cyanoborohydride (NaBHCN), lithium borohydride (LiBH 4 ), sodium borohydride, ethylenediamine borane, dimethylamine borane, sodium triacetoxyborohydride, morpholine borane, 4-methylmorpholine borane, trimethylamine borane, dicyclohexylamine borane, or a salt thereof.

[0029] Embodiment 19.1 includes a step of oxidizing 5-hydroxymethylated cytosines to 5-formylcytosines (e.g., oxidizing the hydroxymethylcytosines in the first and second strands with KRuO 4 20. The method of any one of embodiments 7 to 19, wherein the DNA is rendered single-stranded prior to contacting the DNA with the

[0030] Embodiment 19.2 is the method of embodiment 19.1, wherein the DNA is made single-stranded using heat denaturation.

[0031] Embodiment 20 is as follows. 2. The method of any one of the preceding embodiments, wherein the asymmetric adapter comprises a molecular barcode.

[0032] Embodiment 21 is as follows. 2. The method of any one of the preceding embodiments, comprising preparing the DNA molecule by attaching the Y-shaped adaptor to a precursor DNA molecule, and optionally, said attaching comprises ligating.

[0033] Embodiment 22 is as follows. The method of the immediately preceding embodiment, wherein said precursor DNA molecule is cell-free DNA.

[0034] Embodiment 23 is as follows. 23. The method of any one of embodiments 21 to 22, wherein the precursor DNA molecule is obtained from a sample.

[0035] Embodiment 24 is as follows. The method of the immediately preceding embodiment, wherein the sample is from a mammal.

[0036] Embodiment 25 is as follows. 25. The method of any one of embodiments 23 to 24, wherein the sample is a blood sample.

[0037] Embodiment 26 is as follows. 23. The method of any one of the preceding embodiments, comprising dividing the sample into at least the first sub-sample and a second sub-sample prior to attaching the Y-shaped adaptors to DNA molecules of at least a first sub-sample.

[0038] Embodiment 27 is as follows. 4. The method of the previous embodiment, wherein said dividing step is based on epigenetic modifications.

[0039] Embodiment 28 is as follows. The method of the immediately preceding embodiment, wherein said epigenetic modification is a cytosine modification.

[0040] Embodiment 29 is as follows. The method of the immediately preceding embodiment, wherein said cytosine modification is a 5-methylation.

[0041] Embodiment 30 is as follows. 30. The method of any one of embodiments 27 to 29, wherein said first sub-sample comprises DNA molecules enriched for said epigenetic modifications.

[0042] Embodiment 31 is as follows. 31. The method of any one of embodiments 26 to 30, wherein the dividing step comprises contacting the DNA molecule with a methyl-binding reagent immobilized on a solid support.

[0043] Embodiment 32 is as follows. 32. The method of any one of embodiments 26 to 31, comprising a step of distinctly tagging said first sub-sample and said second sub-sample.

[0044] Embodiment 33 is as follows. The method of the immediately preceding embodiment, wherein DNA from said first sub-sample and said second sub-sample is pooled.

[0045] Embodiment 34 is as follows. 34. The method of any one of embodiments 32 to 33, wherein DNA from the first sub-sample and the set of target regions or the second sub-sample is sequenced in the same sequencing cell.

[0046] Embodiment 35 is as follows. 35. The method of any one of embodiments 26 to 34, wherein the DNA of the first sub-sample and the DNA of the second sub-sample are differentially tagged; after differential tagging, a portion of the DNA from the second sub-sample or processed sub-sample is added to the first sub-sample or an additional processed sub-sample or at least a portion thereof, thereby forming a pool; and sequence variable target regions and epigenetic target regions are captured from the pool.

[0047] Embodiment 36 is as follows. The method of the immediately preceding embodiment, wherein said pool comprises less than or equal to about 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10% or 5% of said DNA of said second sub-sample.

[0048] Embodiment 37 is as follows. The method of the immediately preceding embodiment, wherein the pool comprises about 70-90%, about 75-85%, or about 80% of the DNA of the second sub-sample.

[0049] Embodiment 38 is as follows. 38. The method of any one of embodiments 35 to 37, wherein the pool comprises substantially all of the DNA of the first sub-sample.

[0050] Embodiment 39 is as follows. 39. The method of any one of embodiments 35 to 38, wherein said pool comprises substantially all of said DNA of said first sub-sample or of said processed first sub-sample.

[0051] Embodiment 40 is as follows. 40. The method of any one of embodiments 35 to 39, wherein a first set of target regions is captured from at least a portion of said first sub-sample or processed first sub-sample after formation of said pool.

[0052] Embodiment 41 is as follows. 41. The method of any one of embodiments 37 to 40, wherein the plurality of sub-samples comprises a third sub-sample comprising a higher proportion of DNA with said epigenetic modifications than said second sub-sample but a lower proportion than said first sub-sample.

[0053] Embodiment 42 is as follows. The method of the immediately preceding embodiment, further comprising the step of differentially tagging said third sub-sample.

[0054] Embodiment 43 is as follows. 2. The method of claim 1, wherein the DNA from the first sub-sample, the DNA from the third sample, and the set of target regions are pooled, and optionally the DNA from the first, second and third sub-samples is sequenced in the same sequencing cell.

[0055] Embodiment 44 is as follows. 2. The method of any one of the preceding embodiments, wherein the DNA molecule is amplified.

[0056] Embodiment 45 is as follows. 13. The method of any one of the preceding embodiments, wherein the DNA molecule is amplified after the deaminating step.

[0057] Embodiment 46 is as follows. 2. The method of any one of the preceding embodiments, wherein the DNA molecules are amplified prior to the sequencing step.

[0058] Embodiment 47 is as follows. 13. The method of any one of the preceding embodiments, wherein the sequencing step is next generation sequencing.

[0059] Embodiment 48 is as follows. 2. The method of any one of the preceding embodiments, further comprising a step of capturing a set of target regions of the processed DNA molecules prior to the sequencing step, wherein the captured processed DNA molecules are sequenced or amplified and sequenced.

[0060] Embodiment 49 is as follows. The method of the immediately preceding embodiment, wherein said set of target regions comprises epigenetic target regions.

[0061] Embodiment 50 is as follows. 50. The method of any one of embodiments 48 to 49, wherein the set of target regions comprises a set of hypermethylated variable target regions.

[0062] Embodiment 51 is as follows. The method of the immediately preceding embodiment, wherein the set of hypermethylated variable target regions comprises regions that have a higher degree of methylation in at least one type of tissue than the degree of methylation in cell-free DNA from a healthy subject.

[0063] Embodiment 52 is as follows. 52. The method of any one of embodiments 48 to 51, wherein the set of target regions comprises a set of hypomethylated variable target regions.

[0064] Embodiment 53 is as follows. The method of the immediately preceding embodiment, wherein the set of hypomethylated variable target regions comprises regions that have a lower degree of methylation in at least one type of tissue than the degree of methylation in cell-free DNA from a healthy subject.

[0065] Embodiment 54 is as follows. 54. The method of any one of embodiments 48 to 53, wherein the set of target regions comprises a set of methylation control target regions.

[0066] Embodiment 55 is as follows. 55. The method of any one of embodiments 48 to 54, wherein the set of target regions comprises a set of fragmented variable target regions.

[0067] Embodiment 56 is as follows. The method of the immediately preceding embodiment, wherein said set of fragmented variable target regions comprises a transcription start site region.

[0068] Embodiment 57 is as follows. 57. The method of any one of embodiments 55 to 56, wherein the set of fragmented variable target regions comprises a CTCF binding region.

[0069] Embodiment 58 is as follows. 41. The method of any one of embodiments 31 to 40, wherein the set of target regions comprises a set of hydroxymethylated variable target regions.

[0070] Embodiment 59 is as follows. The method of the immediately preceding embodiment, wherein the set of hydroxymethylated variable target regions comprises regions that have a higher degree of hydroxymethylation in at least one type of tissue than the degree of hydroxymethylation in cell-free DNA from a healthy subject.

[0071] Embodiment 60 is as follows. 60. The method of any one of embodiments 58 to 59, wherein DNA molecules corresponding to the set of hydroxymethylated variable target regions are captured with a higher capture yield than DNA molecules corresponding to at least one other set of target regions.

[0072] Embodiment 61 is as follows. 61. The method of any one of embodiments 48 to 60, wherein the set of target regions comprises sequence variable target regions.

[0073] Embodiment 62 is as follows. The method of the immediately preceding embodiment, wherein DNA molecules corresponding to said set of sequence variable target regions are captured with a higher capture yield than DNA molecules corresponding to said set of epigenetic target regions.

[0074] Embodiment 63 is as follows. 2. The method of any one of the preceding embodiments, wherein the DNA molecule comprises insert DNA from a subject, and the method further comprises determining a likelihood that the subject has cancer.

[0075] Embodiment 64 is as follows. The method of the immediately preceding embodiment, wherein the sequencing step generates a plurality of sequencing reads; the method further comprises mapping the plurality of sequence reads to one or more reference sequences to generate mapped sequence reads, and processing the mapped sequence reads corresponding to the set of sequence variable target regions and the mapped sequence reads corresponding to the set of epigenetic target regions to determine a likelihood that the subject has cancer.

[0076] Embodiment 65 is as follows. The method of any one of embodiments 1 to 63, wherein the DNA molecules comprise insert DNA derived from a subject, the subject having previously been diagnosed with cancer and having undergone one or more previous cancer treatments, and optionally, the cfDNA is obtained at one or more preselected time points after the one or more previous cancer treatments, and a captured set of cfDNA molecules is sequenced, thereby generating a set of sequence information.

[0077] Embodiment 66 is as follows. The method of the immediately preceding embodiment, further comprising using said set of sequence information to detect the presence or absence of DNA originating from or derived from a tumor cell at a preselected time point.

[0078] Embodiment 67 is as follows. The method of the immediately preceding embodiment, further comprising determining for the subject a cancer recurrence score indicative of the presence or absence of the DNA originating or derived from the tumor cells, and optionally further comprising determining a cancer recurrence status based on the cancer recurrence score, wherein the cancer recurrence status of the subject is determined to be at risk of cancer recurrence if the cancer recurrence score is determined to be at or above a predetermined threshold, or the cancer recurrence status of the subject is determined to be at a lower risk of cancer recurrence if the cancer recurrence score is below the predetermined threshold.

[0079] Embodiment 68 is as follows. The method of the immediately preceding embodiment, further comprising comparing the subject's cancer recurrence score to a predetermined cancer recurrence threshold, wherein the subject is classified as a candidate for subsequent cancer treatment if the cancer recurrence score is above the cancer recurrence threshold, or is not classified as a candidate for subsequent cancer treatment if the cancer recurrence score is below the cancer recurrence threshold.

[0080] Embodiment 69 is as follows. 2. The method of any one of the preceding embodiments, wherein the step of glucosylating the 5-hydroxymethylated cytosine comprises contacting the 5-hydroxymethylated cytosine with a β-glucosyltransferase.

[0081] Embodiment 70 is as follows. 2. The method of any one of the preceding embodiments, wherein the step of glucosylating the 5-hydroxymethylated cytosine produces 5-glucosylhydroxymethylcytosine.

[0082] Embodiment 71 is as follows. 2. The method of any one of the preceding embodiments, wherein methylating the cytosine in at least one of the first complementary strand or the second complementary strand comprises contacting the cytosine with a DNA methyltransferase.

[0083] Embodiment 72 is as follows. The method of the immediately preceding embodiment, wherein said DNA methyltransferase is DNMT1 or DNMT5.

[0084] Embodiment 73 is as follows. 2. The method of any one of the preceding embodiments, further comprising identifying methylated positions in the DNA molecule.

[0085] Embodiment 74 is as follows. 2. The method of any one of the preceding embodiments, further comprising identifying hydroxymethylated positions in the DNA molecule.

[0086] Embodiment 75 is as follows. 2. The method of any one of the preceding embodiments, comprising identifying methylated positions in said DNA molecules and hydroxymethylated positions in said DNA molecules.

[0087] Embodiment 76 is as follows. 2. The method of any one of the preceding embodiments, comprising the steps of: (a) identifying methylated positions in said DNA molecule and hydroxymethylated positions in said DNA molecule, and (b) identifying the gene sequence of said DNA molecule.

[0088] Embodiment 77 is as follows. The method of any one of the preceding embodiments, comprising identifying at least one position in the DNA molecule that contained a hydroxymethylated cytosine; at least one position in the DNA molecule that contained a methylated cytosine; at least one position in the DNA molecule that contained a cytosine that was not methylated or hydroxymethylated; at least one position in the DNA molecule that contained an adenine; at least one position in the DNA molecule that contained a guanine; and at least one position in the DNA molecule that contained a thymine.

[0089] Embodiment 78 is as follows. The method of the immediately preceding embodiment, wherein said cytosine that is not methylated or hydroxymethylated is an unmodified cytosine.

[0090] Embodiment 79 is as follows. 2. The method of any one of the preceding embodiments, wherein the step of synthesizing a first complementary strand complementary to the first strand and a second complementary strand complementary to the second strand comprises extending a primer using unmethylated dNTPs.

[0091] Embodiment 80 is as follows. The method of the immediately preceding embodiment, wherein said dNTPs consist of unmethylated dNTPs.

[0092] Embodiment 81 is as follows. 2. The method of any one of the preceding embodiments, wherein synthesizing the first complementary strand complementary to the first strand and the second complementary strand complementary to the second strand converts at least one methylated CpG to a hemi-methylated CpG.

[0093] Embodiment 82 is as follows. 2. The method of any one of the preceding embodiments, wherein synthesizing the first complementary strand complementary to the first strand and the second complementary strand complementary to the second strand converts at least one hydroxymethylated CpG to a hemi-hydroxymethylated CpG.

[0094] Embodiment 83 is as follows. 2. The method of any one of the preceding embodiments, wherein the 5-hydroxymethylated cytosine contained in the hemihydroxymethylated CpG is glycosylated in at least one of the first or second strands.

[0095] Embodiment 84 is as follows. 2. The method of any one of the preceding embodiments, wherein the DNA molecule is free in solution during one or more of the synthesizing, glucosylating, methylating and deaminating steps, and optionally the DNA molecule is free in solution during two, three or four of the synthesizing, glucosylating, methylating and deaminating steps, and further optionally the DNA molecule is free in solution during each of the synthesizing, glucosylating, methylating and deaminating steps.

[0096] Embodiment 85 is as follows. a) reagents for synthesizing a first complementary strand that is complementary to the first strand and a second complementary strand that is complementary to the second strand; b) a reagent for glycosylating 5-hydroxymethylated cytosines in at least one of the first or second strands before or after the step of synthesizing said first and second complementary strands; c) a reagent for methylating cytosines in at least one of the first complementary strand or the second complementary strand, said methylation converting a hemi-methylated CpG to a fully methylated CpG; d) a reagent for deaminating unmodified cytosines in at least one of the first or second strands; e) a plurality of oligonucleotide probes; f) primers for synthesizing a first complementary strand complementary to the first strand and a second complementary strand complementary to the second strand; g) primers for the sequencing step; and h) Library adapters with distinct molecular barcodes A kit comprising one or more of:

[0097] Embodiment 85.1 is a kit according to embodiment 85, comprising at least 2, 3, 4, 5, 6, 7 or each of a) through h).

[0098] Embodiment 85.2 is a kit according to embodiment 85.1, comprising at least a) to d); at least a) to d) and f); at least a) to d) and h); or at least a) to d), f) and h).

[0099] Embodiment 85.3 is a kit according to embodiment 85.2, comprising at least a) through g); at least b) through h); at least a) and c) through h); at least a) through b) and d) through h); at least a) through c) and e) through h); at least a) through d) and e) through h); at least a) through e) and f) through h); or at least a) through f) and h).

[0100] Embodiment 85.4 is the kit of any one of embodiments 85 to 85.3, further comprising an amplification primer, optionally wherein said amplification primer comprises a sample barcode.

[0101] Embodiment 86 is as follows. The kit according to any one of embodiments 85 to 85.4, wherein the reagents for synthesizing a first complementary strand complementary to the first strand and a second complementary strand complementary to the second strand comprise a polymerase and / or dNTPs, and optionally, the dNTPs are unmethylated.

[0102] Embodiment 87 is as follows. The kit according to any one of embodiments 85 to 85.4 or 86, wherein the reagent for glucosylating 5-hydroxymethylated cytosines in at least one first or second strand before or after the step of synthesizing the first and second complementary strands is a glucosyltransferase and / or uridine diphosphate glucose, and optionally the glucosyltransferase is a β-glucosyltransferase.

[0103] Embodiment 88 is as follows. 88. The kit of any one of embodiments 85 to 87, wherein the reagents for methylating cytosines in at least one of the first complementary strand or the second complementary strand comprise one or more of a DNA methyltransferase and a methyl donor, optionally wherein the DNA methyltransferase is DNMT1 or DNMT5, and optionally wherein the methyl donor is S-adenosylmethionine.

[0104] Embodiment 89 is as follows. 89. The kit of any one of embodiments 85 to 88, wherein the reagents for deaminating unmodified cytosines in at least one first or second strand comprise one or more of sodium bisulfite or APOBEC3A.

[0105] Embodiment 90 is as follows. A kit described in any one of embodiments 85 to 89, further comprising reagents for the step of converting 5mC and 5hmC into a substrate that cannot be deaminated by a deaminase, and optionally, wherein the reagents for the step of converting 5mC and 5hmC into the substrate that cannot be deaminated by a deaminase comprise a TET enzyme or T4-βGT.

[0106] Embodiment 91 is a) a reagent for the step of oxidizing 5hmC to formylcytosine, optionally comprising KRuO 4 A reagent; b) a TET enzyme; and / or c) borane or borohydride reducing agents, including pyridine borane, 2-picoline borane, borane, tert-butylamine borane, ammonia borane, sodium borohydride, ethylenediamine borane, dimethylamine borane, sodium triacetoxyborohydride, morpholine borane, 4-methylmorpholine borane, trimethylamine borane, dicyclohexylamine borane, or salts thereof. 91. The kit of any one of embodiments 85 to 90, further comprising one or more of:

[0107] Embodiment 92 is as follows. a) the plurality of oligonucleotide probes are selected from the group consisting of ALK, APC, BRAF, CDKN2A, EGFR, ERBB2, FBXW7, KRAS, MYC, NOTCH1, NRAS, PIK3CA, PTEN, RBI, TP53, MET, AR, ABLl, AKTl, ATM, CDHl, CSFIR, CTNNBl, ERBB4, EZH2, FGFRl, FGFR2, FGFR3, FLT3, GNA11, GNAQ, GNAS, HNF1A, HRAS, IDH1, IDH2, JAK2, JAK3, KDR, KIT, MLH1, MPL, NPM1, PDG selectively hybridizes to at least 5, 6, 7, 8, 9, 10, 20, 30, 40 or all of the genes selected from FRA, PROC, PTPN11, RET, SMAD4, SMARCB1, SMO, SRC, STK11, VHL, TERT, CCND1, CDK4, CDKN2B, RAF1, BRCA1, CCND2, CDK6, NF1, TP53, ARID1A, BRCA2, CCNE1, ESR1, RIT1, GATA3, MAP2K1, RHEB, ROS1, ARAF, MAP2K2, NFE2L2, RHOA, and NTRKl; b) the library adaptors do not contain flow cell sequences or sequences that allow for the formation of hairpin loops for sequencing; c) the library adapters are blunt ended and Y-shaped; and / or d) the library adaptors are less than or equal to 40 nucleobases in length; 92. The kit according to any one of embodiments 85 to 91.

[0108] Embodiment 93 is as follows. The kit according to any one of embodiments 85 to 92, further comprising instructions for carrying out the method according to any one of embodiments 1 to 84. I. Brief description of the drawings [Brief description of the drawings]

[0109] [Figure 1A-1] FIG. 1A shows an exemplary workflow according to certain embodiments of the present disclosure. "Wa-orig" and "Cr-orig" are used to indicate the first and second strands of a cfDNA molecule. "Wa-copy" and "Cr-copy" respectively indicate the first and second complementary strands generated by extension. This figure illustrates a method according to the present disclosure in which a Y-shaped adapter containing a molecular barcode is ligated to a cfDNA molecule, followed by synthesis of the first and second complementary strands, glucosylating 5-hydroxymethylcytosine (5hmC) to form 5-glucosylhydroxymethylcytosine (5ghmC), methylating cytosines in hemimethylated CpGs to convert them to fully methylated CpGs (e.g., using DNMT1), deaminating unmethylated cytosines (e.g., by treatment with bisulfite), amplifying by uracil-resistant DNA polymerase, and sequencing. Using the rules shown below the figure, including for identifying strand status as original or copy, the original sequences of the first and second (Wa and Cr) strands can be derived, including whether the cytosines were unmethylated, methylated or hydroxymethylated. [Figure 1A-2] Same as above. [Figure 1A-3] Same as above. [Figure 1A-4] Same as above.

[0110] [Figure 1B-1]FIG. 1B shows an exemplary workflow according to certain embodiments of the present disclosure involving a Y-shaped adapter containing an unmethylated cytosine that can be used to function as a reporter base, i.e., to distinguish strand status as original or copy. Such use of a reporter base makes the method independent of the need for an unmethylated cytosine in the inserted DNA molecule. The steps of ligating a Y-shaped adapter containing a molecular barcode to a cfDNA molecule, followed by synthesizing the first and second complementary strands, glucosylating 5-hydroxymethylcytosine (5hmC) to form 5-glucosylhydroxymethylcytosine (5ghmC), methylating cytosines in hemimethylated CpGs to convert them to fully methylated CpGs (e.g., using DNMT1), deaminating unmethylated cytosines (e.g., by treatment with bisulfite), amplifying with uracil-resistant DNA polymerase, and sequencing are performed essentially similarly to the method shown in FIG. 1A. The rules for base identification are provided below the figure. [Figure 1B-2] Same as above. [Figure 1B-3] Same as above. [Figure 1B-4] Same as above.

[0111] [Figure 1C-1] 1C shows an exemplary workflow according to certain embodiments of the present disclosure that corresponds to the workflow of FIG. 1A with the additional presence of a reporter base in the Y-shaped adapter. As shown, cytosines other than the reporter base in the adapter must be methylated to avoid being deaminated by bisulfite treatment. [Figure 1C-2] Same as above. [Figure 1C-3] Same as above. [Figure 1C-4] Same as above.

[0112] [Figure 1D-1]FIG. 1D shows an exemplary workflow according to certain embodiments of the present disclosure in which unmethylated cytosines remain intact, but methylated cytosines and hydroxymethylcytosines are converted to bases read as thymine, comprising the steps of: oxidizing 5-hydroxymethylated cytosines to 5-formylcytosines (e.g., by contacting hydroxymethylcytosines in a first strand and a second strand with KRuO4); synthesizing a first complementary strand complementary to the first strand and a second complementary strand complementary to the second strand; methylating cytosines in at least one of the first complementary strand or the second complementary strand, where the methylation converts a hemi-methylated CpG to a fully methylated CpG; converting modified cytosines in at least one of the first or second strands to thymine or a base read as thymine, thereby producing a processed DNA molecule; and sequencing at least a portion of the processed DNA molecule. In some embodiments, the DNA is rendered single-stranded (e.g., by thermal denaturation) prior to the step of oxidizing 5-hydroxymethylated cytosines to 5-formylcytosines (e.g., by contacting hydroxymethylcytosines in the first and second strands with KRuO4). In some embodiments, the step of converting modified cytosines in at least one of the first or second strands to thymine or a base that is read as thymine includes converting formylcytosine and / or methylcytosine to carboxyl cytosine (e.g., by contacting formylcytosine and / or methylcytosine with TET enzyme) and reducing carboxyl cytosine to dihydrouracil (e.g., by contacting carboxyl cytosine with borane or borohydride reducing agent). As shown, in such a method, the asymmetric adapter can have a modified cytosine, e.g., a methylated cytosine, that functions as a reporter base. [Figure 1D-2] Same as above. [Figure 1D-3] Same as above. [Figure 1D-4] Same as above.

[0113] [Diagram 2] FIG. 2 is a schematic diagram of one example of a system suitable for use in some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0114] II. Detailed Description of Certain Embodiments Reference will now be made in detail to certain specific embodiments of the invention. While the invention will be described in conjunction with such embodiments, it will be understood that they are not intended to limit the invention to those embodiments. On the contrary, the invention is intended to cover all alternatives, modifications, and equivalents which may be included within the scope of the invention as defined by the appended claims.

[0115] Before describing the teachings of the present invention in detail, it should be understood that the present disclosure is not limited to specific compositions or process steps, which may vary as such. It should be noted that, as used in this specification and the appended claims, the singular forms "a", "an" and "the" include plural references unless the context clearly indicates otherwise. Thus, for example, reference to "a nucleic acid" includes a plurality of nucleic acids, reference to "a cell" includes a plurality of cells, and so forth.

[0116] Numerical ranges are inclusive of the numbers defining the range. Measured and measurable values ​​are understood to be approximations taking into account significant digits and errors associated with measurement. Also, the use of "comprise," "comprises," "comprising," "contain," "contains," "containing," "include," "includes," and "including" are not intended to be limiting. It is to be understood that both the foregoing general and detailed description are exemplary and explanatory only and are not restrictive of the teachings.

[0117] Unless specifically noted in the specification above, embodiments in the specification that recite "comprising" various components can also be considered to "consist of" or "consist essentially of" the recited components; embodiments in the specification that recite "consisting of" various components can also be considered to "comprise" or "consist essentially of" the recited components; and embodiments in the specification that recite "consisting essentially of" various components can also be considered to "consist of" or "comprise" the recited components (this interchangeability does not apply to the use of these terms in the claims).

[0118] The section headings used herein are for organizational purposes only and should not be construed as limiting the disclosed subject matter in any way. In the event that any document or other material incorporated by reference conflicts with any express content of this specification, including definitions, the present specification will control. A.Definition

[0119] "Cell-free DNA", "cfDNA molecule" or simply "cfDNA" includes DNA molecules that naturally exist in a subject in an extracellular form (e.g., in blood, serum, plasma, or other bodily fluids, such as lymph, cerebrospinal fluid, urine, or sputum). cfDNA originally existed in a cell or cells in a large complex organism, such as a mammal, but has been released from the cell into the fluid found in the organism, and can be obtained from a sample of the fluid without the need to perform an in vitro cell lysis step.

[0120] As used herein, "cellular nucleic acids" refers to nucleic acids that are located within the cell or cells from which they originate, at least at the time the sample is obtained or collected from a subject, even if those nucleic acids are subsequently removed (e.g., via cell lysis) as part of a given analytical process.

[0121] As used herein, a modification or other feature is present in a "higher percentage" in a first sample or population of nucleic acids than in a second sample or population if the percentage of nucleotides having the modification or other feature is higher in the first sample or population than in the second population. For example, if in a first sample, 1 / 10 of the nucleotides are mC, and in a second sample, 1 / 20 of the nucleotides are mC, then the first sample contains a higher percentage of 5-methylated cytosine modifications than the second sample.

[0122] As used herein, "without substantially altering the base-pairing specificity" of a given nucleobase means that the majority of molecules that contain that nucleobase that can be sequenced have no alteration of the base-pairing specificity of the second nucleobase compared to its base-pairing specificity when it was in the sample from which it was originally isolated. In some embodiments, 75%, 90%, 95% or 99% of molecules that contain that nucleobase that can be sequenced have no alteration of the base-pairing specificity of the second nucleobase compared to its base-pairing specificity when it was in the sample from which it was originally isolated.

[0123] As used herein, "base pairing specificity" refers to the standard DNA base (A, C, G or T) with which a given base most preferentially pairs. Thus, for example, unmodified cytosine and 5-methylcytosine have the same base pairing specificity (i.e., specificity for G), but uracil and cytosine have different base pairing specificities, since uracil has base pairing specificity for A, while cytosine has base pairing specificity for G. The ability of uracil to form a wobble pair with G is not important, since uracil still most preferentially pairs with A among the four standard DNA bases.

[0124] As used herein, "modified cytosine" refers to a cytosine in which at least one position of the cytosine is substituted with a chemical moiety, e.g., methyl or hydroxymethyl, which is different from the substituent at that position in an unmodified cytosine. For the avoidance of doubt, "modified cytosine" does not include unmodified cytosine.

[0125] As used herein, a "combination" containing multiple members refers to either a single composition containing those members, or a set of adjacent compositions, for example, in separate containers or compartments within a larger container, such as a multi-well plate, tube rack, refrigerator, freezer, incubator, water bath, ice bucket, machine, or other form of storage.

[0126] The "capture yield" of a collection of probes for a given target set refers to the amount of nucleic acid corresponding to the target set that the collection of probes captures under typical conditions (e.g., the amount relative to another target set, or the absolute amount). Exemplary typical capture conditions are incubation of sample nucleic acid and probes at 65°C for 10-18 hours in a small reaction volume (about 20 μL) containing a stringent hybridization buffer. Capture yields can be expressed in absolute terms, or for multiple collections of probes, in relative terms. When capture yields for multiple sets of target regions are compared, they are normalized with respect to the footprint size of the target region set (e.g., on a per kilobase basis). Thus, for example, if the footprint sizes of the first and second target regions are 50 kb and 500 kb, respectively (giving a normalization factor of 0.1), when the mass concentration per volume of the captured DNA corresponding to the first set of target regions is higher than 0.1 times the mass concentration per volume of the captured DNA corresponding to the second set of target regions, the DNA corresponding to the first set of target regions will be captured with a higher yield than the DNA corresponding to the second set of target regions. As a further example, using the same footprint size, if the captured DNA corresponding to the first set of target regions has a mass concentration per volume that is 0.2 times the mass concentration per volume of the captured DNA corresponding to the second set of target regions, the DNA corresponding to the first set of target regions will be captured with a capture yield that is 2 times higher than the DNA corresponding to the second set of target regions.

[0127] "Capturing" one or more target nucleic acids refers to preferentially isolating or separating one or more target nucleic acids from non-target nucleic acids.

[0128] A "captured set" of nucleic acids refers to the nucleic acids that have undergone capture.

[0129] A "target region set" or "set of target regions" refers to multiple genomic loci that are targeted for capture and / or targeted by a set of probes (e.g., via sequence complementarity).

[0130] "Corresponding to a set of target regions" means that a nucleic acid, e.g., cfDNA, originates from a locus in the set of target regions or specifically binds to one or more probes for the set of target regions.

[0131] As used herein, a "differentially methylated region" (DMR) or a "differentially hydroxymethylated region" (DhMR) refers to a region of DNA that has a detectably different degree of methylation or hydroxymethylation, respectively, in at least one cell or tissue type compared to the degree of methylation or hydroxymethylation in the same region of DNA from at least one other cell or tissue type; or that has a detectably different degree of methylation or hydroxymethylation in at least one cell or tissue type obtained from a subject with a disease or disorder compared to the degree of methylation or hydroxymethylation in the same region of DNA in the same cell or tissue type obtained from a healthy subject. In some embodiments, a DMR has a detectably higher degree of methylation (e.g., a hypermethylated region) in at least one cell or tissue type compared to the degree of methylation in the same region of DNA from at least one other cell or tissue type or from the same cell or tissue type from a healthy subject. In some embodiments, the DMR has a detectably lower degree of methylation in at least one cell or tissue type compared to the degree of methylation in the same region of DNA from at least one other cell or tissue type or from the same cell or tissue type from a healthy subject (e.g., a hypomethylated region). Similarly, in some embodiments, the DhMR has a detectably higher degree of hydroxymethylation in at least one cell or tissue type compared to the degree of hydroxymethylation in the same region of DNA from at least one other cell or tissue type or from the same cell or tissue type from a healthy subject (e.g., a hyperhydroxymethylated region). In some embodiments, the DhMR has a detectably lower degree of hydroxymethylation in at least one cell or tissue type compared to the degree of hydroxymethylation in the same region of DNA from at least one other cell or tissue type or from the same cell or tissue type from a healthy subject (e.g., a hypohydroxymethylated region).

[0132] "Specifically binds" in the context of a probe or other oligonucleotide and a target sequence means that, under appropriate hybridization conditions, the oligonucleotide or probe hybridizes to its target sequence or a copy thereof to form a stable probe:target hybrid, while at the same time minimizing the formation of stable probe:non-target hybrids. Thus, the probe hybridizes to the target sequence or a copy thereof to a substantially greater extent than to a non-target sequence, allowing capture or detection of the target sequence. Suitable hybridization conditions are well known in the art and can be predicted based on sequence composition or determined using routine testing methods (see, e.g., §§ 1.90-1.91, 7.37-7.57, 9.47-9.51, and 11.47-11.57 of Sambrook et al., Molecular Cloning, A Laboratory Manual, 2nd ed. (Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1989), which is hereby incorporated by reference, in particular §§ 9.50-9.51, 11.12-11.13, 11.45-11.47, and 11.55-11.57).

[0133] A "sequence variable target region set" refers to a set of target regions that may exhibit changes in sequence, such as nucleotide substitutions (i.e., single nucleotide variations), insertions, deletions, or gene fusions or rearrangements, in neoplastic cells (e.g., tumor cells and cancer cells).

[0134] "Epigenetic target region set" refers to a set of target regions that may show sequence-independent changes in neoplastic cells (e.g., tumor cells and cancer cells) or that may show sequence-independent changes in cfDNA from subjects with cancer compared with cfDNA from healthy subjects.Examples of sequence-independent changes include, but are not limited to, changes in methylation (increase or decrease), nucleosome distribution, CTCF binding, transcription start site and regulatory protein binding region. For purposes of the present invention, loci that are subject to neoplasia-, tumor- or cancer-associated focal amplifications and / or gene fusions may also be included in the set of epigenetic target regions, as detection of changes in copy number by sequencing or by fused sequences that map to more than one locus in a reference genome tends to be more similar to detection of the exemplary epigenetic changes discussed above than detection of nucleotide substitutions, insertions or deletions, in that, for example, focal amplifications and / or gene fusions can be detected at a relatively shallow depth of sequencing, since their detection does not depend on the accuracy of base calls at one or a few individual positions.

[0135] A nucleic acid is "produced by a tumor" or ctDNA or circulating tumor DNA if it originates from a tumor cell. Tumor cells are neoplastic cells that originate from a tumor, regardless of whether they remain in the tumor or are separated from the tumor (e.g., as is the case for metastatic cancer cells and circulating tumor cells).

[0136] The term "methylation" or "DNA methylation" refers to the addition of a methyl group to a nucleotide base in a nucleic acid molecule. In some embodiments, methylation refers to the addition of a methyl group to a cytosine at a CpG site (a cytosine-phosphate-guanine site (i.e., a cytosine followed by a guanine in the 5'→3' direction of a nucleic acid sequence)). In some embodiments, DNA methylation is 5-methylation (modification of the fifth carbon of the 6-carbon ring of cytosine). In some embodiments, 5-methylation refers to the addition of a methyl group to the 5C position of cytosine, generating 5-methylcytosine (5mC). Methylation can also occur at non-CpG sites, e.g., methylation can occur at CpA, CpT, or CpC sites. DNA methylation can change the activity of a methylated DNA region. For example, if DNA in a promoter region is methylated, transcription of a gene can be repressed. DNA methylation is important for normal development, and abnormalities in methylation can disrupt epigenetic regulation. Disruptions, e.g., repression, in epigenetic regulation can cause disease, e.g., cancer. Promoter methylation in DNA can be indicative of cancer.

[0137] The term "hydroxymethylation" or "DNA hydroxymethylation" refers to the addition of a hydroxymethyl group to a nucleotide base in a nucleic acid molecule. In some embodiments, hydroxymethylation refers to the addition of a hydroxymethyl group to a cytosine at a CpG site (a cytosine-phosphate-guanine site (i.e., a cytosine followed by a guanine in the 5'→3' direction of a nucleic acid sequence)). In some embodiments, DNA hydroxymethylation is 5-hydroxymethylation (modification of the fifth carbon of the 6-carbon ring of cytosine). In some embodiments, 5-hydroxymethylation refers to the addition of a hydroxymethyl group to the 5C position of cytosine, generating 5-hydroxymethylcytosine (5hmC). Hydroxymethylation can also occur at non-CpG sites, e.g., hydroxymethylation can occur at CpA, CpT, or CpC sites. DNA hydroxymethylation can alter the activity of the hydroxymethylated DNA region. Aberrant hydroxymethylation in DNA is associated with and can be indicative of various cancers.

[0138] The term "hypermethylated" refers to an increased level or degree of methylation of a nucleic acid molecule compared to other nucleic acid molecules within a population (e.g., a sample) of nucleic acid molecules. In some embodiments, hypermethylated DNA can include DNA molecules that contain at least one methylated residue, at least two methylated residues, at least three methylated residues, at least five methylated residues, or at least ten methylated residues.

[0139] The term "hypomethylation" refers to a level or degree of methylation of a nucleic acid molecule that is reduced compared to other nucleic acid molecules in a population (e.g., a sample) of nucleic acid molecules. In some embodiments, hypomethylated DNA includes DNA molecules that are unmethylated. In some embodiments, hypomethylated DNA can include DNA molecules that contain zero methylated residues, at most one methylated residue, at most two methylated residues, at most three methylated residues, at most four methylated residues, or at most five methylated residues.

[0140] As used herein, "methylation status" can refer to the presence or absence of a methyl group on a DNA base (e.g., cytosine) at a particular genomic position in a nucleic acid molecule. Methylation status can also refer to the degree of methylation in a nucleic acid sequence (e.g., a highly methylated, low methylated, intermediately methylated, or unmethylated nucleic acid molecule). Methylation status can also refer to the number of methylated nucleotides in a particular nucleic acid molecule.

[0141] As used herein, "mutation" refers to a variation from a known reference sequence, including mutations such as single nucleotide variants (SNVs) and insertions or deletions (indels).Mutation can be germline mutations or somatic mutations.In some embodiments, the reference sequence for comparison purposes is the wild-type genomic sequence of the species of the subject that provides the test sample, typically the human genome.

[0142] As used herein, the terms "neoplasm" and "tumor" are used interchangeably. They refer to the abnormal growth of cells in a subject. A neoplasm or tumor can be benign, potentially malignant, or malignant. A malignant tumor is also called a cancer or cancerous tumor.

[0143] As used herein, "next-generation sequencing" or "NGS" refers to a sequencing technology that has increased throughput compared to traditional Sanger-based and capillary electrophoresis-based approaches, e.g., has the ability to generate hundreds of thousands of relatively small sequence reads at a time. Some examples of next-generation sequencing techniques include, but are not limited to, sequencing by synthesis, sequencing by ligation, and sequencing by hybridization. In some embodiments, next-generation sequencing involves the use of an instrument that can sequence a single molecule. Examples of commercially available instruments for performing next-generation sequencing include, but are not limited to, NextSeq, HiSeq, NovaSeq, MiSeq, Ion PGM, and Ion GeneStudio S5.

[0144] As used herein, "nucleic acid tag" refers to short nucleic acids (e.g., less than about 500, 100, 50 or 10 nucleotides in length) of different types or different processes that are used to distinguish nucleic acids from different samples (e.g., to indicate a sample index), to distinguish nucleic acids from different compartments (e.g., to indicate a split tag), or to distinguish different nucleic acid molecules in the same sample (e.g., to indicate a molecular barcode). Nucleic acid tags include predetermined, fixed, non-random, random or semi-random oligonucleotide sequences. Such nucleic acid tags can be used to label different nucleic acid molecules or different nucleic acid samples or sub-samples. Nucleic acid tags can be single-stranded, double-stranded, or at least partially double-stranded. Nucleic acid tags can be the same length or vary in length, as appropriate. Nucleic acid tags can also include double-stranded molecules with one or more blunt ends, can include 5' or 3' single-stranded regions (e.g., overhangs), and / or can include one or more other single-stranded regions at other locations within a given molecule. Nucleic acid tags can be attached to one or both ends of other nucleic acids (e.g., sample nucleic acids to be amplified and / or sequenced). Nucleic acid tags can be decoded to reveal information such as the origin, form or processing of a given nucleic acid sample. For example, nucleic acid tags can also be used to enable pooling and / or parallel processing of multiple samples containing nucleic acids carrying different molecular barcodes and / or sample indexes, where the nucleic acids are subsequently deconvoluted by detecting (e.g., reading) the nucleic acid tags. Nucleic acid tags can also be referred to as identifiers (e.g., molecular identifiers, sample identifiers). Additionally or alternatively, nucleic acid tags can be used as molecular identifiers (e.g., to distinguish between different molecules or amplicons of different parent molecules in the same sample or subsamples). This includes, for example, uniquely tagging different nucleic acid molecules in a given sample or non-uniquely tagging such molecules.For non-unique tagging applications, a limited number of tags (i.e., molecular barcodes) may be used to tag each nucleic acid molecule such that different molecules may be identified based on their intrinsic sequence information (e.g., start and / or stop positions where they are mapped to a selected reference genome, subsequences at one or both ends of the sequence, and / or length of the sequence) in combination with at least one molecular barcode. Typically, a sufficient number of different molecular barcodes are used such that there is a low probability (e.g., less than about 10%, less than about 5%, less than about 1%, or less than about 0.1%) that any two molecules may have the same intrinsic sequence information (e.g., start and / or stop positions, subsequences at one or both ends of the sequence, and / or length) and also have the same molecular barcode.

[0145] As used herein, "non-immobilized" or "free in solution" DNA refers to DNA that is not covalently or non-covalently attached to a solid support, such as a bead. Such DNA can be free in solution during any step (e.g., all steps) of the disclosed methods.

[0146] As used herein, "partitioning" refers to physically separating or fractionating a mixture of nucleic acid molecules in a sample based on the characteristics of the nucleic acid molecules. Partitioning can be a physical partitioning of the molecules. Partitioning can include separating nucleic acid molecules into groups or sets based on the level of epigenetic features (e.g., methylation). For example, nucleic acid molecules can be partitioned based on the level of methylation of the nucleic acid molecules. In some embodiments, methods and systems used for partitioning can be found in PCT Patent Application No. PCT / US2017 / 068329, which is hereby incorporated by reference in its entirety.

[0147] As used herein, a "partitioned set" or "partition" refers to a set of nucleic acid molecules divided into sets or groups based on the differential binding affinity of the nucleic acid molecules or proteins associated with the nucleic acid molecules to a binder. A partitioned set may also be referred to as a subsample. A binder preferentially binds to nucleic acid molecules that contain nucleotides with epigenetic modifications. For example, if the epigenetic modification is methylation, the binder may be a methyl-binding domain (MBD) protein. In some embodiments, a partitioned set may include nucleic acid molecules that belong to a particular level or degree of epigenetic trait (e.g., methylation). For example, the nucleic acid molecules may be divided into three sets - one set (first subsample, hyperpartition, overpartitioned set or overmethylated partitioned set) for highly methylated nucleic acid molecules, a second set (second subsample, hypopartition, hypopartitioned set or hypomethylated partitioned set) for low methylated nucleic acid molecules, and a third set (third subsample, intermediate partitioned set, intermediate methylated partitioned set, remaining partitioned set or remaining partition) for intermediate methylated nucleic acid molecules. In another example, the nucleic acid molecules may be divided based on the number of methylated nucleotides - one partitioned set may have nucleic acid molecules with nine methylated nucleotides and another partitioned set may have unmethylated nucleic acid molecules (zero methylated nucleotides).

[0148] As used herein, "polynucleotide", "nucleic acid", "nucleic acid molecule" or "oligonucleotide" refers to a linear polymer of nucleosides (including deoxyribonucleosides, ribonucleosides, or analogs thereof) connected by internucleoside linkages. Typically, a polynucleotide contains at least three nucleosides. Oligonucleotides often range in size from a few monomeric units, e.g., 3-4, to hundreds of monomeric units. Whenever a polynucleotide is represented by a sequence of letters, e.g., "ATGCCTG", the nucleotides are in 5'→3' order from left to right, and in the case of DNA, "A" refers to deoxyadenosine, "C" refers to deoxycytidine, "G" refers to deoxyguanosine, and "T" refers to deoxythymidine, unless otherwise specified. The letters A, C, G, and T may be used to refer to the bases themselves, nucleosides, or nucleotides that contain the bases.

[0149] As used herein, "processing" refers to a set of steps used to generate a library of nucleic acids suitable for sequencing. The set of steps may include, but is not limited to, splitting, end repair, adding sequencing adaptors, tagging, and / or PCR amplification of the nucleic acids.

[0150] As used herein, "quantitative measure" refers to an absolute or relative measure. A quantitative measure can be, without limitation, a number, a statistical measurement (e.g., frequency, mean, median, standard deviation, or quantile), or a degree or relative amount (e.g., high, medium, and low). A quantitative measure can be a ratio of two quantitative measures. A quantitative measure can be a linear combination of quantitative measures. A quantitative measure can be a standardized measure.

[0151] As used herein, "reference sequence" refers to a known sequence that is used for comparison with experimentally determined sequences.For example, the known sequence can be the whole genome, chromosome, or any segment thereof.The reference sequence can be aligned with a single continuous sequence of genome or chromosome or chromosome arm, or can include non-contiguous segments that align with different regions of genome or chromosome.Examples of reference sequences include, for example, human genome, for example, hg19 and hg38.

[0152] As used herein, a "sample" means anything capable of being analyzed by the methods and / or systems disclosed herein.

[0153] As used herein, "sequencing" refers to any of several techniques used to determine the sequence (e.g., the identity and order of monomeric units) of a biomolecule, e.g., a nucleic acid, e.g., DNA or RNA. Examples of sequencing methods include, but are not limited to, targeted sequencing, single molecule real-time sequencing, exon or exome sequencing, intron sequencing, electron microscopy-based sequencing, panel sequencing, transistor-mediated sequencing, direct sequencing, random shotgun sequencing, Sanger dideoxy termination sequencing, whole genome sequencing, sequencing by hybridization, pyrosequencing, duplex sequencing, cycle sequencing, single base extension sequencing, solid-phase sequencing, high-throughput sequencing, massively parallel signature sequencing, emulsion PCR, co-amplification-PCR at lower denaturation temperatures (COLD-PCR), multiplex PCR, reversible dye terminator sequencing, paired-end sequencing, short-term sequencing, exonuclease sequencing, sequencing by ligation, short read sequencing, single molecule sequencing, sequencing by synthesis, real-time sequencing, reverse-terminator sequencing, nanopore sequencing, 454 sequencing, Solexa Genome These include Genetic Analyzer sequencing, SOLiD™ sequencing, MS-PET sequencing, and combinations thereof. In some embodiments, sequencing can be performed by a genetic analyzer, such as the genetic analyzer commercially available from Illumina, Inc., Pacific Biosciences, Inc., or Applied Biosystems / Thermo Fisher Scientific, among many others.

[0154] As used herein, "sequence information" in the context of a nucleic acid polymer means the order and identity of the monomeric units (eg, nucleotides) in that polymer.

[0155] As used herein, a "sequence variable target region set" refers to a set of target regions that can exhibit changes in sequence, such as nucleotide substitutions, insertions, deletions, or gene fusions or rearrangements, in neoplastic cells (e.g., tumor cells and cancer cells).

[0156] As used herein, the terms "somatic mutation" or "somatic variation" are used interchangeably. They refer to mutations in the genome that occur after conception. Somatic mutations can occur in any cell of the body except germ cells, and therefore are not passed on to offspring.

[0157] As used herein, "specifically bind" in the context of a probe or other oligonucleotide and a target sequence means that under suitable hybridization conditions, the oligonucleotide or probe hybridizes to its target sequence or a copy thereof to form a stable probe:target hybrid, while at the same time minimizing the formation of stable probe:non-target hybrids. Thus, the probe hybridizes to the target sequence or a copy thereof to a substantially higher degree than to a non-target sequence, allowing the capture or detection of the target sequence. Suitable hybridization conditions are well known in the art and can be predicted based on sequence composition or determined by using routine testing methods (see, e.g., §§ 1.90-1.91, 7.37-7.57, 9.47-9.51 and 11.47-11.57 of Sambrook et al., Molecular Cloning, A Laboratory Manual, 2nd ed. (Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1989), which is incorporated herein by reference, in particular §§ 9.50-9.51, 11.12-11.13, 11.45-11.47 and 11.55-11.57).

[0158] As used herein, "subject" refers to an animal, such as a mammalian species (e.g., human) or avian (e.g., avian) species, or other organism, such as a plant. More specifically, the subject may be a vertebrate, such as a mammal, such as a mouse, a primate, a monkey, or a human. Animals include farm animals (e.g., production cattle, dairy cows, poultry, horses, pigs, etc.), sport animals, and companion animals (e.g., pets or support animals). The subject may be a healthy individual, an individual having or suspected of having a disease or predisposition to disease, or an individual in need of treatment or suspected of needing treatment. The terms "individual" or "patient" are intended to be interchangeable with "subject." For example, the subject may be an individual diagnosed with cancer, an individual about to undergo cancer treatment, and / or an individual who has undergone at least one cancer treatment. The subject may be in remission from cancer. As another example, the subject may be an individual diagnosed with an autoimmune disease. As another example, the subject may be a female individual who is pregnant or planning to become pregnant, who may have been diagnosed with a disease, e.g., cancer, an autoimmune disease, or who is suspected of having such a disease.

[0159] As used herein, "target region set" or "set of target regions" or "target region" or "target region of interest" or "region of interest" or "genomic region of interest" refers to multiple genomic loci or multiple genomic regions that are targeted for capture and / or targeted by a set of probes (e.g., via sequence complementarity).

[0160] As used herein, "tumor fraction" refers to the proportion of cfDNA molecules that originate from tumor cells for a given sample or sample-region pair.

[0161] As used herein, an "asymmetric adapter" is a double-stranded adapter in which the two strands are not perfectly complementary or are otherwise distinguishable such that synthesis of a complementary sequence on one strand of the adapter results in a sequence that is distinguishable from the sequence on the other strand of the adapter. Examples of asymmetric adapters are Y-shaped adapters and bubble adapters.

[0162] As used herein, a "Y-shaped adapter" refers to an adapter that includes two DNA strands, including a complementary portion and a non-complementary portion, where the non-complementary portion forms a single-stranded arm. The adapter can be attached to a sample or insert DNA molecule, for example, by ligation, such that the complementary (double-stranded) portion of the adapter is proximal to the sample or insert DNA molecule. Prior to attachment, the double-stranded portion of the Y-shaped adapter can be blunt-ended or have an overhang of, for example, 1-3 nucleotides. The single-stranded arms may or may not be of the same length.

[0163] As used herein, a "bubble adapter" refers to an adapter that includes two DNA strands that include a non-complementary portion flanked by complementary portions such that the adapter has a single-stranded region located between the double-stranded regions. The adapter can be attached to a sample or insert DNA molecule, e.g., by ligation, such that one of the complementary (double-stranded) portions of the adapter is proximal to the sample or insert DNA molecule. Prior to attachment, the double-stranded portion of the Y-shaped adapter that is attached to the insert or sample molecule can have a blunt end or an overhang of, e.g., 1-3 nucleotides. The single-stranded portions of the two strands may or may not be of identical length.

[0164] The terms "or combinations thereof" and "or combinations thereof" as used herein refer to any and all permutations and combinations of the listed terms preceding the term. For example, "A, B, C, or combinations thereof" is intended to include at least one of the following: A, B, C, AB, AC, BC, or ABC, and also BA, CA, CB, ACB, CBA, BCA, BAC, or CAB, if the order is important in the particular context. Extending this example, combinations containing one or more repeats of an item or term are expressly included, such as BB, AAA, AAB, BBC, AAABCCCC, CBBAAA, CABABB, etc. One of skill in the art will understand that there is typically no limitation on the number of items or terms in any combination, unless otherwise clear from the context.

[0165] "Or" is used in its inclusive sense, ie, equivalent to "and / or," unless the context requires otherwise. B. Exemplary Methods 1. Overview

[0166] The formation and progression of cancer can result from both genetic modification and epigenetic features of deoxyribonucleic acid (DNA).The present disclosure provides a method and system for analyzing DNA, such as cell-free DNA (cfDNA).The present disclosure provides a method and system for reducing the signal-to-noise ratio of methylation splitting assay.

[0167] Without wishing to be bound by any particular theory, cells in or around cancer or neoplasms may shed more DNA than cells of the same tissue type in healthy subjects. Thus, the distribution of tissues from which a particular DNA sample, e.g., cfDNA, originates, may change during carcinogenesis. Thus, for example, an increase in the level of hypermethylated variable target regions that show lower methylation in healthy cfDNA than in at least one other tissue type may be indicative of the presence of cancer (or recurrence, depending on the subject's medical history). Similarly, an increase in the level of hypomethylated variable target regions in a sample may be indicative of the presence of cancer (or recurrence, depending on the subject's medical history). DNA methylation and / or hydroxymethylation profiling may be used to detect aberrant methylation or hydroxymethylation, respectively, in the DNA of a sample. The DNA may correspond to certain genomic regions ("differentially methylated regions" or "DMRs" or "differentially hydroxymethylated regions" or "DhMRs") that are normally hyper- or hypomethylated and / or hyper- or hypohydroxymethylated in a given sample type (e.g., cfDNA from the bloodstream) but may show abnormal degrees of methylation and / or hydroxymethylation that correlate with neoplasms or cancer, for example, due to an abnormally increased contribution of tissue to the sample type (e.g., due to increased DNA shedding in or around neoplasms or cancers) and / or from the degree of genomic methylation and / or hydroxymethylation that is altered during development or perturbed by disease, e.g., cancer or any cancer-related disease. Cytosine methylation and hydroxymethylation are two different forms of modifications that can provide separate information, but existing techniques may not easily facilitate discriminative detection of both of these modifications in the same workflow.

[0168] Thus, methods are provided herein in which asymmetric adaptors, e.g., Y-shaped adaptors, complementary strand synthesis, glycosylation, methylation of hemimethylated CpG, and deamination are combined to facilitate the identification of methylation and hydroxymethylation of cytosine in the same workflow. In any of the methods described herein, DNA may be free in solution during any step (e.g., all steps) of the method, i.e., not immobilized (e.g., not covalently or non-covalently attached to a solid support, e.g., beads). In some embodiments, the DNA molecule is free in solution during one or more of the synthesis, glycosylation, methylation, and deamination steps described herein, and optionally the DNA molecule is free in solution during two, three, or four of the synthesis, glycosylation, methylation, and deamination steps, and further optionally the DNA molecule is free in solution during each of the synthesis, glycosylation, methylation, and deamination steps. Asymmetric adapters may contain strand reporter bases (e.g., deamination-sensitive cytosines, e.g., unmethylated cytosines) that can be used to determine whether a strand sequence originated from the original strand or the synthesized complementary strand. Methylation and hydroxymethylation of cytosines can be identified in the same molecule with single-base resolution. The approaches herein can also avoid the complications associated with looped or hairpin adapters. Additionally, these methods may further include either or both of glucosylation of hydroxymethylcytosines and methylation of hemimethylated CpGs.Methylation of hemimethylated CpGs may prevent deamination of CpGs in the synthesized complementary strand at the methylated positions in the original molecule, and glucosylation of hydroxymethylcytosine may reduce or eliminate methylation opposite the hydroxymethylcytosine positions (because some methyltransferases, e.g., DNMT1, may have some activity to methylate the unmethylated CpG opposite the hydroxymethylCpG, which is inhibited or blocked by glucosylation). Thus, as a result of glucosylation, the unmethylated CpG opposite the glucosylated hydroxymethylCpG is not recognized as a substrate for methylation by methyltransferases, e.g., DNMT1, and therefore remains unmethylated throughout the remaining steps of the disclosed method, so that the site containing the glucosylated hydroxymethylcytosine remains hemihydroxymethylated. These features may simplify data analysis and / or reduce error rates (e.g., avoiding misidentification of hydroxymethylcytosine and methylcytosine positions). In such embodiments, a six letter (A, C, T, G, 5mC and 5hmC) digital readout can be provided in a single workflow. In embodiments of the disclosed methods that do not include a step of glycosylating hydroxymethylcytosines, a five letter (A, C, T, G, and either 5mC or 5hmC) digital readout can be provided in a single workflow, rather than a six letter (A, C, T, G, 5mC and 5hmC) digital readout.

[0169] In some embodiments, the methods include synthesizing a first complementary strand complementary to a first strand of a DNA molecule and a second complementary strand complementary to a second strand of the DNA molecule, the DNA molecule comprising an asymmetric adaptor, e.g., a Y-shaped adaptor. The asymmetric adaptor is generally not a looped adaptor or a hairpin adaptor. As shown in FIG. 1A, synthesizing the complementary strand produces a double-stranded molecule in which the newly synthesized strand comprises a sequence complementary to the adaptor sequence of the original strand. Thus, the first complementary strand can be distinguished from the second strand based on the difference in its adaptor sequence, and similarly for the second complementary strand and the first strand. In some embodiments, synthesizing the first complementary strand complementary to the first strand of a DNA molecule and the second complementary strand complementary to the second strand of the DNA molecule comprises extending a primer using unmethylated deoxyribonucleotide triphosphates (dNTPs). In some embodiments, the dNTPs consist of unmethylated dNTPs.

[0170] In some embodiments, as shown in Figure 1C and Figure 1B, the adaptor comprises a deamination sensitive nucleotide (e.g., cytosine), for example, an unmethylated cytosine. In some embodiments of the method comprising a step in which unmodified cytosine is deaminated, one or more of the primers used in the step of synthesizing a first complementary strand complementary to the first strand of the DNA molecule and a second complementary strand complementary to the second strand of the DNA molecule also comprise one or more unmethylated cytosines. The unmethylated cytosines are deaminated in the deamination step and can also function to identify which strand the sequence corresponds to and whether it corresponds to the original insert sequence or the sequence of the complementary strand created during the step of synthesizing. In particular, as shown in FIG. 1C, the deamination step facilitates the identification of reads corresponding to the original first and second strands, separate from the first and second complementary ("copy") strands, in that if the read corresponds to the original strand, the reporter bases proximal and distal to the read start are T and G, respectively, whereas if the read corresponds to the copy strand, the reporter bases proximal and distal to the read start are C and A, respectively. The use of reporter bases makes this method independent of the presence of unmethylated cytosines in the sequence of the original sample molecule. Otherwise, if all cytosines were methylated, it would be difficult to identify the original and copy strands. In some embodiments, the deamination-sensitive cytosines are in the strand of an asymmetric adaptor that undergoes ligation to the 5' end of the sample or insert DNA molecule. In some embodiments, the nucleotide immediately 3' to the deamination-sensitive cytosine contains a nucleobase other than guanine, e.g., adenine, cytosine, thymine, or uracil (including modified forms thereof, e.g., methylated cytosine). Thus, the deamination-sensitive cytosine is not part of a CpG and is not recognized as a substrate by a methyltransferase, e.g., DNMT1 or DNMT5.Thus, the inclusion of a deamination-sensitive cytosine, e.g., an unmethylated cytosine, in the sequenced portion of the asymmetric adapter (e.g., adjacent to the molecular barcode) serves as a reporter of the "original" vs. "copy" strand status of the library molecule. As outlined in Figure 1B, due to the molecular biology logic of the workflow, the "strand reporter" cytosine is read as a thymine (T) base in read 1 of the "original" strand library molecule (that has undergone deamination) and as a cytosine (C) base in read 1 of the "copy" strand library molecule.

[0171] In other embodiments of the method comprising the step of deaminating unmodified cytosines, the asymmetric adaptor comprises one or more methylated cytosines. In such embodiments, one or more of the primers used in the step of synthesizing the first complementary strand complementary to the first strand of the DNA molecule and the second complementary strand complementary to the second strand of the DNA molecule also comprise one or more methylated cytosines. When the step of synthesizing the first complementary strand complementary to the first strand and the second complementary strand complementary to the second strand is carried out using dNTPs containing unmethylated cytosines, the unmethylated CpG cytosines in the first and second complementary strands (distal adaptor end copies) are methylated during the subsequent methylation step (when the hemimethylated CpG is converted to fully methylated CpG), while the unmethylated non-CpG cytosines are deaminated during the subsequent deamination step. As a result, the sequences of the distal adaptor ends in the first and second complementary strands may differ from those of the original strands after conversion.

[0172] The step of synthesizing the first complementary strand complementary to the first strand and the second complementary strand complementary to the second strand is performed before the deamination step, so that the adapter sequences after deamination of the original and copy strands differ at the positions of the cytosines that are differentially methylated between the first strand and the first complementary strand or between the second strand and the second complementary strand. In such an embodiment, two sets of primers can be utilized in the amplification step, for example during the optional library amplification step, the optional enrichment PCR step, or the sequencing step. In such an embodiment, the first set of primers can be complementary to the original non-deaminated sequence, and the second set of primers can be complementary to the deaminated sequence. In some embodiments, the amplification step can be used to add standard sequencing primers to the DNA, so that custom primers are not required in the sequencing step.

[0173] Optionally, these methods may include a step of protecting, e.g., glucosylating, hydroxymethylated cytosine (hmC), e.g., using a glucosyltransferase, e.g., β-glucosyltransferase (β-GT). In some embodiments, a glucose donor, e.g., uridine diphosphate glucose, is also provided. The glucosylating step may be performed before or after the step of synthesizing the complementary strand. Glucosylation may protect hmC from undesired modification in subsequent steps. Sites in DNA that contain hemihydroxymethylated glucosylated hmC remain hemihydroxymethylated throughout the remaining steps of the disclosed method.

[0174] In some embodiments, these methods include, for example, using a DNA methyltransferase, such as DNMT1 or DNMT5, methylating cytosines in at least one of the first or second complementary strands, which converts hemimethylated CpGs to fully methylated CpGs. As shown in FIG. 1A, this step affects hemimethylated CpGs (i.e., CpGs in which the cytosine on one strand is methylated but the cytosine on the other strand is not methylated), but does not affect glucosylated hmC. If the mode of methylating hemimethylated CpGs does not recognize hemihydroxymethylated CpGs, glucosylating or otherwise protecting hmC is unnecessary.

[0175] In some embodiments, the methods include deaminating unmodified cytosines in at least one of the first or second strands, thereby producing a treated DNA molecule. In some embodiments, deamination includes contacting the DNA with sodium bisulfite or an APOBEC enzyme (e.g., APOBEC3A). As shown in FIG. 1A, this deamination step functions to convert unmodified cytosines to bases that are read as T during sequencing (e.g., uracil). In some embodiments, unmodified cytosines in at least one of the first or second strands are deaminated using sodium bisulfite or a deaminase, e.g., a cytidine deaminase, e.g., an APOBEC enzyme (or a fragment thereof), e.g., APOBEC3A. In particular, as also shown in FIG. 1A, CpGs that are fully methylated (e.g., as a result of a methylation step disclosed herein (e.g., using a methyltransferase (e.g., DNMT1)), whereby methylation converts hemi-methylated CpGs to fully methylated CpGs) are unaffected by the deamination step; CpGs with unmodified cytosines undergo deamination on both strands; hemi-hydroxymethylated (or hemi-glucosylhydroxymethylated) CpGs undergo deamination on only one strand. This differential deamination can facilitate the determination of whether individual positions in the original sample molecule are methylated, hydroxymethylated, or unmethylated.

[0176] In some embodiments, these methods include sequencing at least a portion of the processed DNA molecule. For example, by obtaining sequences for each of the first strand, the first complementary strand, the second strand and the second complementary strand, the deamination pattern for each CpG reveals the methylation or hydroxymethylation status of the original molecule without the need to divide the sample into subsamples, process the subsamples separately and analyze them for methylation and hydroxymethylation. This approach also avoids the complications associated with hairpin adaptors or ligated molecules.

[0177] In some embodiments, the adapters comprise molecular barcodes, which, as discussed in detail elsewhere herein, can be used to identify sequence reads that originate from the same original sample molecule.

[0178] DNA methylation (including hydroxymethylation) can change the activity of methylated DNA regions. For example, if DNA in a promoter region is methylated, transcription of a gene can be suppressed. DNA methylation is important for normal development, and abnormalities in methylation can disrupt epigenetic regulation. Disruptions in epigenetic regulation, such as suppression, can cause disease, such as cancer. Promoter methylation in DNA can indicate cancer. Methylation profiling can include determining methylation patterns across different regions of a genome. For example, the sequences of molecules in different distributions can be mapped to a reference genome. This can indicate regions of the genome that are more highly methylated or less highly methylated compared to other regions. In this way, genomic regions can differ in the degree of methylation as opposed to individual molecules.

[0179] In some embodiments, combining signals obtained from methylation profiling with signals obtained from somatic variations (e.g., SNVs, indels, CNVs and gene fusions) facilitates the detection of cancer.

[0180] In some embodiments, the nucleic acid molecules in a sample may be fractionated or divided based on the methylation status of the nucleic acid molecules prior to the adaptor ligation step. Partitioning the nucleic acid molecules in a sample may increase rare signals. For example, genetic variations that are present in hypermethylated DNA but less abundant (or absent) in hypomethylated DNA may be more easily detected by partitioning the sample into hypermethylated and hypomethylated nucleic acid molecules. By analyzing multiple fractions of a sample, multidimensional analysis of a single molecule may be performed, thus achieving greater sensitivity. Partitioning may include physically partitioning the nucleic acid molecules into subsets or groups based on the presence or absence of one or more methylated nucleobases. The sample may be fractionated or distributed into one or more partitioned sets based on features indicative of differential gene expression or disease state. Samples may be fractionated based on features or combinations thereof that provide a difference in signal between normal and diseased states during analysis of nucleic acids, e.g., cell-free DNA ("cfDNA"), non-cfDNA, tumor DNA, circulating tumor DNA ("ctDNA"), and cell-free nucleic acid ("cfNA"). In some embodiments, the workflow described in Figures 1A-1D may be performed on one or more partitioned sets. In an embodiment, the workflow described in Figures 1A-1D may be performed on a hypermethylated partitioned set.

[0181] The partitioning procedure may result in incomplete sorting of DNA molecules between sub-samples. For example, a minority of the molecules in the second sub-sample may be highly modified (e.g., hypermethylated), and / or a minority of the molecules in the first sub-sample may be unmodified or mostly unmodified (e.g., unmethylated or mostly unmethylated). The highly modified molecules in the second sub-sample and the unmodified or mostly unmodified molecules in the first sub-sample are considered to be non-specifically partitioned. The methods described herein may include a step that can reduce technical noise from non-specifically partitioned DNA, for example, by decomposing it. Thus, the methods described herein may provide improved sensitivity and / or streamlined analysis.

[0182] In some embodiments, the method may further comprise detecting the presence or absence of cancer in a subject, for example, based on the methylation status at one or more loci of the nucleic acid molecules in at least one divided set. In some embodiments, the method further comprises determining a level of DNA from tumor cells in the polynucleotide sample.

[0183] In certain embodiments, the method comprises the following steps, performed in the following order: a) synthesizing a first complementary strand complementary to the first strand and a second complementary strand complementary to the second strand (e.g., using a polymerase, e.g., the Klenow fragment); b) glucosylating 5-hydroxymethylated cytosines in at least one of the first or second strands (e.g., using βGT) before or after synthesizing the first and second complementary strands; c) methylating cytosines in at least one of the first or second complementary strands, where the methylation converts a hemimethylated CpG to a fully methylated CpG (e.g., using a methylase). transferase, e.g., using DNMT1 or DNMT5); d) deaminating unmodified cytosines in at least one of the first or second strands, thereby producing a treated DNA molecule (e.g., using bisulfite, e.g., sodium bisulfite, or an APOBEC enzyme, e.g., APOBEC3A); and e) sequencing at least a portion of the treated DNA molecule; wherein the DNA molecule comprises a first and a second strand and an asymmetric adaptor, optionally wherein at least one asymmetric adaptor comprises a deamination-sensitive cytosine, and optionally wherein the asymmetric adaptor is a Y-shaped adaptor or a bubble adaptor. In such an embodiment, step a is performed first, step b is performed second (after step a), step c is performed third (after steps a and b), step d is performed fourth (after steps a-c), and step e is performed fifth (after steps a-d).

[0184] In another particular embodiment, the method comprises the following steps, performed in the following order: a) oxidizing 5-hydroxymethylated cytosines in at least one first or second strand to 5-formylcytosines (e.g., using a ruthenium salt, e.g., KRuO 4b) synthesizing a first complementary strand complementary to the first strand and a second complementary strand complementary to the second strand (e.g., using a polymerase, e.g., the Klenow fragment); c) methylating cytosines in at least one of the first or second complementary strands, where the methylation converts a semi-methylated CpG to a fully methylated CpG (e.g., using a methyltransferase, e.g., DNMT1 or DNMT5); d) synthesizing modified cytosines in at least one of the first or second strands. converting cytosines to thymines or bases that are read as thymines, thereby producing treated DNA molecules (e.g., using a reducing agent, e.g., picoline borane or pyridine borane); and e) sequencing at least a portion of the treated DNA molecules; wherein the DNA molecules comprise first and second strands and asymmetric adaptors, and optionally, at least one asymmetric adaptor comprises a deamination-susceptible cytosine, and optionally, the asymmetric adaptor is a Y-shaped adaptor or a bubble adaptor. In such embodiments, step a is performed first, step b is performed second (after step a), step c is performed third (after steps a and b), step d is performed fourth (after steps a-c), and step e is performed fifth (after steps a-d). In some embodiments, the DNA molecules are single-stranded prior to the oxidizing step. 2. Splitting the sample into multiple subsamples; sample characteristics

[0185] In certain embodiments described herein, a population of different forms of nucleic acids (e.g., hypermethylated and hypomethylated DNA in a sample, e.g., cfDNA) may be physically divided based on one or more characteristics of the nucleic acids prior to further analysis, e.g., one or more of adapter ligation, complementary strand synthesis, glycosylation, methylation, deamination, sequencing, etc., or all of the following. In certain embodiments, the dividing step is performed prior to adapter ligation, e.g., prior to adapter ligation in the methods disclosed in Figures 1A-1D. The adapters ligated after the dividing step may include a division tag that facilitates identification of the compartment into which the molecule was sorted, e.g., after the compartments are pooled and processed together in a subsequent step. This approach may be used, for example, to determine whether a particular sequence is hypermethylated or hypomethylated. Additionally, partitioning a heterogeneous nucleic acid population may increase rare signals, e.g., by enriching for rare nucleic acid molecules that are more prevalent in one fraction (or partition) of the population. For example, genetic variations that are present in hypermethylated DNA but less abundant (or absent) in hypomethylated DNA can be more easily detected by partitioning the sample into hypermethylated and hypomethylated nucleic acid molecules. By analyzing multiple fractions of a sample, multidimensional analysis of a single locus or nucleic acid species of the genome can be performed, and thus greater sensitivity can be achieved.

[0186] When a splitting step is used, the subsequent steps of the method (e.g., synthesizing, glucosylating, methylating and deaminating; or oxidizing, synthesizing, methylating and converting) can be performed on one, more than one, or each compartment. For example, the subsequent steps of the method (e.g., synthesizing, glucosylating, methylating and deaminating; or oxidizing, synthesizing, methylating and converting) can be performed on the hypermethylated compartment. In another example, the subsequent steps of the method (e.g., synthesizing, glucosylating, methylating and deaminating; or oxidizing, synthesizing, methylating and converting) can be performed on the hypomethylated compartment. In another example, the subsequent steps of the method (e.g., synthesizing, glucosylating, methylating and deaminating; or oxidizing, synthesizing, methylating and converting) can be performed on the hypomethylated compartment and the hypermethylated compartment. These compartments may be combined at a stage in the workflow if the remaining steps performed for each compartment (e.g., including at least a sequencing step, or including (i) synthesizing, glucosylating, methylating and deaminating steps, or (ii) oxidizing, synthesizing, methylating and converting steps followed by a sequencing step in either case (i) or (ii)) are the same.

[0187] In some examples, a heterogeneous nucleic acid sample is divided into two or more partitions (e.g., at least 3, 4, 5, 6 or 7 partitions). A partition of a sample is also referred to herein as a sub-sample. In some embodiments, each partition is differentially tagged. The tagged partitions can then be pooled together for collective sample preparation and / or sequencing. The partitioning-tagging-pooling steps can be performed more than once, and each round of partitioning is performed based on different characteristics (examples provided herein) and tagged using differential tags and partitioning means that are distinguished from other partitions.

[0188] Examples of features that can be used for partitioning include sequence length, methylation level, nucleosome binding, sequence mismatch, immunoprecipitation, and / or proteins that bind to DNA. The resulting partitioning can include one or more of the following nucleic acid forms: single-stranded DNA (ssDNA), double-stranded DNA (dsDNA), shorter DNA fragments, and longer DNA fragments. In some embodiments, a partitioning step based on cytosine modification (e.g., cytosine methylation) or methylation is generally performed, and is optionally combined with at least one additional partitioning step that can be based on any of the above-mentioned characteristics or forms of DNA. In some embodiments, a heterogeneous population of nucleic acids is partitioned into nucleic acids that have one or more epigenetic modifications and nucleic acids that do not have one or more epigenetic modifications. Examples of epigenetic modifications include the presence or absence of methylation; the level of methylation; the type of methylation (e.g., 5-methylcytosine vs. other types of methylation, e.g., adenine methylation and / or cytosine hydroxymethylation); and the association and level of association with one or more proteins, e.g., histones. Alternatively or additionally, the heterogeneous population of nucleic acids can be divided into nucleosome-associated nucleic acid molecules and nucleosome-free nucleic acid molecules.Alternatively or additionally, the heterogeneous population of nucleic acids can be divided into single-stranded DNA (ssDNA) and double-stranded DNA (dsDNA).Alternatively or additionally, the heterogeneous population of nucleic acids can be divided based on nucleic acid length (e.g., molecules up to 160bp and molecules having a length longer than 160bp).

[0189] In some instances, each compartment (representing a different nucleic acid form) is differentially labeled and the compartments are pooled together prior to the sequencing step, in other instances, the different forms are sequenced separately.

[0190] In some embodiments, the population of different nucleic acids is divided into two or more different compartments. Each compartment represents a different nucleic acid morphology, with the first compartment (also called sub-sample) containing a higher proportion of DNA with cytosine modifications than the second sub-sample. Each compartment is tagged separately. The tagged nucleic acids are pooled together before the sequencing step. Sequence reads are obtained and analyzed, including to distinguish in silico the first nucleic acid base from the second nucleic acid base in the DNA of the first sub-sample. The tags are used to sort the reads from the different distributions. Analysis to detect genetic variants can be performed at the level of each distribution and at the level of the entire nucleic acid population. For example, the analysis can include in silico analysis to determine genetic variants, such as CNVs, SNVs, indels, fusions, in the nucleic acids in each distribution. In some examples, the in silico analysis can include determining chromatin structure. For example, the coverage of sequence reads can be used to determine the positioning of nucleosomes in chromatin. Higher coverage may correlate with higher nucleosome occupancy in a genomic region, while lower coverage may correlate with lower nucleosome occupancy or nucleosome depleted regions (NDRs).

[0191] The sample may contain nucleic acids that vary in modifications, including post-replication modifications to nucleotides, and in their binding, usually non-covalently, to one or more proteins.

[0192] In an embodiment, the nucleic acid population is obtained from serum, plasma or blood samples from subjects suspected of having neoplasm, tumor or cancer, or subjects previously diagnosed with neoplasm, tumor or cancer.The nucleic acid population comprises nucleic acids with varying levels of methylation.Methylation can result from any one or more post-replicative or post-transcriptional modifications.Post-replicative modifications include, in particular, modifications of the cytosine of the nucleotide at the 5th position of the nucleobase, such as 5-methylcytosine, 5-hydroxymethylcytosine, 5-formylcytosine and 5-carboxylcytosine.

[0193] Affinity agents can be antibodies with the desired specificity, natural binding partners or variants thereof (Bock et al., Nat Biotech 28: 1106-1114 (2010); Song et al., NatBiotech 29: 68-72 (2011)), or artificial peptides selected, e.g., by phage display, to have specificity for a given target.

[0194] Examples of capture moieties contemplated herein include proteins, such as MeCP2, MBDs, such as MBD2, and methyl-binding domains (MBDs) and methyl-binding proteins (MBPs) described herein, including antibodies that preferentially bind to 5-methylcytosine. When antibodies are used to immunoprecipitate methylated DNA, the methylated DNA can be recovered in single-stranded form. In such an embodiment, a second strand can be synthesized. The hypermethylated (and optionally intermediately methylated) subsample can then be contacted with a methylation-sensitive nuclease, such as HpaII, BstUI, or Hin6i, that does not cleave hemimethylated DNA. Alternatively or additionally, the hypomethylated (and optionally intermediately methylated) subsample can then be contacted with a methylation-dependent nuclease that cleaves hemimethylated DNA.

[0195] Similarly, the division of different forms of nucleic acid can be performed using histone-binding proteins that can separate histone-bound nucleic acid from free or unbound nucleic acid. Examples of histone-binding proteins that can be used in the methods disclosed herein include RBBP4, RbAp48 and SANT domain peptides.

[0196] For some affinity agents and modifications, binding to the agent may occur in an essentially all-or-none manner, depending on whether the nucleic acid carries the modification, but separation may be a matter of degree.In such an example, nucleic acids with over-represented modifications bind to the agent to a higher degree than nucleic acids with under-represented modifications.Alternatively, nucleic acids with modifications may bind in an all-or-none manner.However, different levels of modifications may then be sequentially eluted from the binding agent.

[0197] For example, in some embodiments, partitioning can be binary or based on the degree / level of modification. For example, all methylated fragments can be partitioned from unmethylated fragments using a methyl-binding domain protein (e.g., MethylMinder Methylated DNA Enrichment Kit (ThermoFisher Scientific)). Subsequently, additional partitioning can include eluting fragments with different levels of methylation by adjusting the salt concentration in the solution containing the methyl-binding domain and bound fragments. As the salt concentration increases, fragments with higher methylation levels are eluted.

[0198] In some cases, the final compartments show nucleic acids with different degrees of modification (over- or under-representation of the modification). Over- and under-representation can be defined by the number of modifications made by the nucleic acid compared to the median number of modifications per strand in the population. For example, if the median number of 5-methylcytosine residues in the nucleic acids in the sample is 2, then nucleic acids containing more than two 5-methylcytosine residues will be over-represented with this modification, and nucleic acids with one or zero 5-methylcytosine residues will be under-represented. The effect of affinity separation is to enrich nucleic acids with over-represented modifications in the bound phase and under-represented modifications in the unbound phase (i.e., in solution). Nucleic acids in the bound phase can be eluted before subsequent processing.

[0199] When using MethylMiner Methylated DNA Enrichment Kit (ThermoFisher Scientific), different levels of methylation can be partitioned using sequential elution. For example, low methylated partition (no methylation) can be separated from methylated partition by contacting the nucleic acid population with MBD from this kit bound to magnetic beads. The beads are used to separate methylated nucleic acids from unmethylated nucleic acids. Subsequently, one or more elution steps are performed sequentially to elute nucleic acids with different levels of methylation. For example, the first set of methylated nucleic acids can be eluted with a salt concentration of 160 mM or higher, for example, at least 150 mM, at least 200 mM, 300 mM, 400 mM, 500 mM, 600 mM, 700 mM, 800 mM, 900 mM, 1000 mM or 2000 mM. After such methylated nucleic acids are eluted, magnetic separation is once again used to separate the more highly methylated nucleic acids from those with lower levels of methylation. The elution and magnetic separation steps can be repeated multiple times to generate various compartments, e.g., hypomethylated compartments (indicating no methylation), methylated compartments (indicating low levels of methylation), and hypermethylated compartments (indicating high levels of methylation).

[0200] In some methods, the nucleic acids bound to the agent used for affinity separation are subjected to a washing step, which washes away the nucleic acids that are weakly bound to the affinity agent. Such nucleic acids may be enriched for nucleic acids with modifications to an extent close to the average or median (i.e., halfway between the nucleic acids that remain bound to the solid phase and those that are not bound to the solid phase upon initial contact of the sample with the agent).

[0201] Affinity separation produces at least two, and sometimes three or more, distributions of nucleic acids with different degrees of modification. These distributions are still separate, but the nucleic acids of at least one distribution, usually two or three (or more) distributions, are linked to nucleic acid tags, usually provided as components of an adapter, and the nucleic acids in different distributions receive different tags that distinguish the members of one distribution from another distribution. The tags linked to the nucleic acid molecules of the same distribution can be the same or different from each other. However, when different from each other, these tags can have part of their code in common to identify the molecules to which they are bound as being from a particular distribution.

[0202] For further details regarding partitioning nucleic acid samples based on features such as methylation, see WO2018 / 119452, hereby incorporated by reference.

[0203] In some embodiments, the nucleic acid molecules may be partitioned into different partitions based on the nucleic acid molecules that are bound to a specific protein or fragment thereof and the nucleic acid molecules that are not bound to the specific protein or fragment thereof.

[0204] Nucleic acid molecules can be partitioned based on DNA-protein binding. Protein-DNA complexes can be partitioned based on the specific properties of proteins. Examples of such properties include various epitopes, modifications (e.g., histone methylation or acetylation) or enzymatic activity. Examples of proteins that can bind to DNA and serve as the basis for fractionation can include, but are not limited to, protein A and protein G. Any suitable method can be used to partition nucleic acid molecules based on protein-bound regions. Examples of methods used to partition nucleic acid molecules based on protein-bound regions include, but are not limited to, SDS-PAGE, chromatin-immunoprecipitation (ChIP), heparin chromatography and asymmetrical field flow fractionation (AF4).

[0205] In some embodiments, the division of the nucleic acid is carried out by contacting the nucleic acid with the methylation binding domain ("MBD") of the methylation binding protein ("MBP"). The MBD binds to 5-methylcytosine (5mC). The MBD is coupled to paramagnetic beads, e.g., Dynabeads® M-280 streptavidin, via a biotin linker. The division into fractions with different degrees of methylation can be carried out by eluting the fractions by increasing the NaCl concentration.

[0206] Examples of MBPs contemplated herein include, but are not limited to, the following:

[0207] (a) MeCP2 and MBD2 are proteins that preferentially bind 5-methyl-cytosine over unmodified cytosine.

[0208] (b) RPL26, PRP8 and the DNA mismatch repair protein MHS6 bind preferentially to 5-hydroxymethyl-cytosine over unmodified cytosine.

[0209] (c) FOXK1, FOXK2, FOXP1, FOXP4 and FOXI3 preferably bind 5-formyl-cytosine over unmodified cytosine (Iurlaro et al., Genome Biol. 14: R119 (2013)).

[0210] (d) An antibody specific for one or more methylated nucleotide bases (e.g., MeDIP).

[0211] In some embodiments, the dividing step comprises immunoprecipitation of methylated DNA. For example, dividing by immunoprecipitation of methylated DNA can be used in a method in which the target region set is captured before the dividing step is performed.

[0212] In general, elution is a function of the number of methylated sites per molecule, with molecules with more methylation eluting under increasing salt concentrations. A series of elution buffers with increasing NaCl concentrations can be used to elute DNA into separate populations based on the degree of methylation. Salt concentrations can range from about 100 nM to about 2500 mM NaCl. In one embodiment, the process produces a three-partition. The molecules are contacted with a solution containing a molecule containing a methyl-binding domain at a first salt concentration, which can be bound to a capture moiety, e.g., streptavidin. At the first salt concentration, one population of molecules binds to the MBD and one population remains unbound. The unbound population can be separated as a "hypomethylated" population. For example, the first compartment showing a hypomethylated form of DNA is the compartment that remains unbound at a low salt concentration, e.g., 100 mM or 160 mM. The second compartment, which represents intermediate methylated DNA, is eluted using an intermediate salt concentration, for example, between 100 mM and 2000 mM. This is also separated from the sample. The third compartment, which represents hypermethylated forms of DNA, is eluted using a high salt concentration, for example, at least about 2000 mM.

[0213] In certain embodiments, sample DNA (e.g., between 5 ng and 200 ng) is mixed with methyl-binding domain (MBD) buffer and magnetic beads conjugated with MBD protein and incubated overnight. Methylated DNA (hypermethylated DNA) binds to the MBD protein on the magnetic beads during this incubation. Unmethylated DNA (hypomethylated DNA) or less methylated DNA (intermediately methylated) is washed away from the beads with buffers containing increasing concentrations of salt. For example, one, two or more fractions containing unmethylated, hypomethylated, and / or intermediately methylated DNA may be obtained from such washing. Finally, a high salt buffer is used to elute the strongly methylated DNA (hypermethylated DNA) from the MBD protein. In some embodiments, these washes result in three compartments of DNA with increasing levels of methylation (hypomethylated compartment, intermediately methylated fraction, and hypermethylated compartment).

[0214] In some embodiments, the three compartments of DNA are desalted and concentrated in preparation for the enzymatic step of library preparation. In some embodiments (e.g., after concentrating the DNA in the compartments), the divided DNA is made ligatable, for example, by extending the terminal overhangs of the DNA molecules, adding adenosine residues to the 3' ends of the fragments, and phosphorylating the 5' ends of each DNA fragment. DNA ligase and adapters are added to ligate each divided DNA molecule with an adapter on each end. These adapters contain split tags (e.g., non-random, non-unique barcodes) that are distinguishable from the split tags in the adapters used in other compartments. Two, three, or more compartments are then pooled together and amplified (e.g., by PCR using adapter-specific primers, etc.).

[0215] After PCR, the amplified DNA can be purified and concentrated before enrichment. The amplified DNA is contacted with a group of probes (which can be, for example, biotinylated RNA probes) described herein that target a specific region of interest. The mixture is incubated, for example, overnight, in a salt buffer. The probes are captured (for example, using streptavidin magnetic beads) and separated from the amplified DNA that is not captured, for example, by a series of salt washes, thereby enriching the sample. After enrichment, the enriched sample is amplified by PCR. In some embodiments, the PCR primer contains a sample tag, thereby incorporating the sample tag into the DNA molecule. In some embodiments, DNA from different samples is pooled together and then multiplex sequenced, for example, using an Illumina NovaSeq sequencer. 3. Tag; barcode; amplify

[0216] In some embodiments, a tag is present in the DNA molecule that may be or may include a barcode. The tag may be included as part of the asymmetric adapter described elsewhere herein and / or may be incorporated in the ligation step or in a separate amplification step. The tag may facilitate the identification of the origin of the nucleic acid. For example, a barcode (a type of tag) may be used to allow the source (e.g., subject) from which the DNA originated to be identified after pooling multiple samples for parallel sequencing. This may be performed in parallel with the amplification procedure, for example, by providing a barcode in the 5' portion of the primer, as described above. In some embodiments, the adapter and the tag / barcode are provided by the same primer or primer set. For example, the barcode may be located 3' of the adapter and 5' of the target hybridizing portion of the primer. Alternatively, the barcode may be added by other approaches, for example, ligation, optionally together with the adapter in the same ligation substrate.

[0217] In some embodiments, the DNA is amplified. In some embodiments, the amplification is performed before the capturing step. In some embodiments, the amplification is performed after the capturing step.

[0218] A tag or index can be a molecule, such as a nucleic acid, that contains information that indicates the characteristics of the molecule with which the tag is associated. The tag can allow for differentiation of the molecule from which the sequence read originates. For example, a molecule can carry a sample tag or sample index (which distinguishes a molecule in one sample from a molecule in a different sample), a split tag (which distinguishes a molecule in one compartment from a molecule in a different compartment), or a molecular tag / molecular barcode / barcode (which distinguishes different molecules from each other (in both unique and non-unique tagging scenarios)). In certain embodiments, a tag can include one or a combination of barcodes. As used herein, the term "barcode" refers to a nucleic acid molecule with a specific nucleotide sequence, or the nucleotide sequence itself, depending on the context. A barcode can have, for example, between 10 and 100 nucleotides. A collection of barcodes can have degenerate sequences, or sequences with a certain Hamming distance, if desired for a particular purpose. Thus, for example, a molecular barcode can be composed of one barcode or a combination of two barcodes, each attached to a different end of the molecule. Additionally or alternatively, different sets of molecular barcodes, molecular tags or molecular indexes can be used for different distributions and / or samples, such that the barcodes function as molecular tags through their individual sequences and also function to identify the distributions and / or samples they correspond to based on the set they are members of. The tags, including barcodes, can be incorporated into or otherwise connected to the adapter. The tags can be incorporated by ligation, overlap extension PCR, among other methods.

[0219] Tagging strategies can be divided into unique tagging strategies and non-unique tagging strategies. In unique tagging, all or substantially all of the molecules in a sample carry different tags, so that the reads can be assigned to the original molecule based on tag information alone. The tags used in such methods are sometimes called "unique tags". In non-unique tagging, different molecules in the same sample can carry the same tags, so that other information besides tag information is used to assign sequence reads to the original molecule. Such information can include start and stop coordinates, coordinates to which the molecule is mapped, start or stop coordinates alone, etc. The tags used in such methods are sometimes called "non-unique tags". Thus, it is not necessary to uniquely tag all molecules in a sample. This is sufficient to uniquely tag molecules in a sample that fall within an identifiable class. Thus, molecules in different identifiable families can carry the same tag without loss of information about the identity of the tagged molecule.

[0220] In certain embodiments of non-unique tagging, the number of different tags used may be sufficient so that there is a very high probability (e.g., at least 99%, at least 99.9%, at least 99.99%, or at least 99.999%) that all molecules in a particular group carry different tags. It should be noted that when barcodes are used as tags, and when barcodes are attached, for example randomly, to both ends of a molecule, a combination of barcodes together may constitute a tag. This number in terms is a function of the number of molecules that go into the call. For example, a class may be all molecules that map to the same start-stop position on a reference genome. A class may be all molecules that map to a particular locus, for example, over a particular base or a particular region (e.g., up to 100 bases, or a gene, or an exon of a gene). In certain embodiments, the number z of different tags used to uniquely identify some molecules in a class may be 2 * z, 3 * z, 4 * z, 5 * z, 6 * z, 7* z, 8 * z, 9 * z, 10 * z, 11 * z, 12 * z, 13 * z, 14 * z, 15 * z, 16 * z, 17 * z, 18 * z, 19 * z, 20 * z or 100 * Either z (for example, the lower limit) and 100,000 * z, 10,000 * z, 1000 * z or 100 * z can be between any (e.g., upper limit)

[0221] For example, in a sample of about 5 ng to 30 ng of cell-free DNA, roughly 3000 molecules are expected to map to a particular nucleotide coordinate, with between about 3 and 10 molecules having any start coordinate such that they share the same stop coordinate. Thus, about 50 to about 50,000 different tags (e.g., between about 6 and 220 barcode combinations) may be sufficient to uniquely tag all such molecules. To uniquely tag all 3000 molecules mapped across nucleotide coordinates, about one million to about 20 million different tags would be required.

[0222] In general, the assignment of unique or non-unique tag barcodes in the reaction follows the methods and systems described by U.S. Patent Application Nos. 20010053519, 20030152490, 20110160078, and U.S. Patent Nos. 6,582,908, 7,537,898, and 9,598,731. Tags can be randomly or non-randomly linked to the sample nucleic acids.

[0223] In some embodiments, the tagged nucleic acid is sequenced after being loaded into a microwell plate.The microwell plate can have 96, 384 or 1536 microwells.In some cases, they are introduced with a predicted ratio of unique tags to microwells.For example, unique tags can be loaded so that more than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 10000, 50,000, 100,000, 500,000, 1,000,000, 10,000,000, 50,000,000 or 1,000,000,000 unique tags are loaded per genome sample. In some cases, unique tags may be loaded such that less than about 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 10000, 50,000, 100,000, 500,000, 1,000,000, 10,000,000, 50,000,000 or 1,000,000,000 unique tags are loaded per genomic sample. In some cases, the average number of unique tags loaded per sample genome is less than or greater than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 10000, 50,000, 100,000, 500,000, 1,000,000, 10,000,000, 50,000,000, or 1,000,000,000 unique tags per genome sample.

[0224] A preferred format uses 20-50 different tags (e.g., barcodes) ligated to both ends of a target nucleic acid. For example, 35 different tags (e.g., barcodes) ligated to both ends of a target molecule would create 35x35 permutations, which is equal to 1225 for 35 tags. Such a number of tags is sufficient so that different molecules with the same start and stop points have a high probability (e.g., at least 94%, 99.5%, 99.99%, 99.999%) of receiving different combinations of tags. Other barcode combinations include any number between 10 and 500, such as about 15x15, about 35x35, about 75x75, about 100x100, about 250x250, about 500x500.

[0225] In some cases, the unique tag can be a predetermined or random or semi-random sequence oligonucleotide. In other cases, multiple barcodes can be used, where the barcodes do not necessarily need to be unique to each other in the plurality. In this example, the barcode can be ligated to each molecule such that the combination of the barcode and the sequence to which it can be ligated creates a unique sequence that can be tracked individually. As described herein, detection of a non-unique barcode in combination with sequence data at the beginning (start) and end (stop) positions of the sequence read can allow for the assignment of a unique identity to a particular molecule. The length or number of base pairs of each sequence read can also be used to assign a unique identity to such a molecule. As described herein, fragments from a single strand of a nucleic acid that are assigned a unique identity can thereby allow subsequent identification of fragments from the parent strand.

[0226] Tags can be used to label individual polynucleotide population distributions to correlate the tag(s) with a specific distribution. Alternatively, tags can be used in embodiments of the invention that do not use a distribution step. In some embodiments, a single tag can be used to label a specific distribution. In some embodiments, multiple different tags can be used to label a specific distribution. In embodiments that use multiple different tags to label a specific distribution, the set of tags used to label one distribution can be easily differentiated from the set of tags used to label another distribution. In some embodiments, the tag may have additional functions, e.g., the tag may be used to index a sample source, or may be used as a unique molecular identifier (which may be used to improve the quality of sequencing data by differentiating sequencing errors from mutations, e.g., as in Kinde et al., Proc Nat'l Acad Sci USA 108: 9530-9535 (2011); Kou et al., PLoS ONE,11: e0146638 (2016)), or may be used as a non-unique molecular identifier, e.g., as described in U.S. Pat. No. 9,598,731. Similarly, in some embodiments, the tag may have additional functions, e.g., the tag may be used to index a sample source, or may be used as a non-unique molecular identifier (which may be used to improve the quality of sequencing data by differentiating sequencing errors from mutations).

[0227] In one embodiment, tagging the distributions includes tagging the molecules in each distribution with a distribution tag. After recombining the distributions (e.g., to reduce the number of required sequencing runs and avoid unnecessary costs) and sequencing the molecules, the distribution tag identifies the source distribution. In another embodiment, different distributions are tagged with different sets of molecular tags, e.g., composed of pairs of barcodes. In this way, each molecular barcode not only indicates the source distribution, but is also useful for identifying the molecules within the distribution. For example, a first set of 35 barcodes can be used to tag the molecules in the first distribution, and a second set of 35 barcodes can be used to tag the molecules in the second distribution. a. Tagging of parcels

[0228] In some embodiments in which a dividing step is performed, two or more partitions, for example, each partition is differentially tagged (eg, with a different partition tag).

[0229] In some embodiments, after partitioning and tagging with partition tags, the molecules can be pooled for sequencing in a single run. In some embodiments, sample tags are added to the molecules, for example, in a step subsequent to the addition of the partition tag and pooling. Sample tags can facilitate pooling of material generated from multiple samples for sequencing in a single sequencing run.

[0230] Alternatively, in some embodiments, the distribution tag can be correlated to sample and distribution.As a simple example, the first tag can indicate the first distribution of the first sample; the second tag can indicate the second distribution of the first sample; the third tag can indicate the first distribution of the second sample; and the fourth tag can indicate the second distribution of the second sample.

[0231] Tags may be attached to molecules that have already been distributed based on one or more characteristics, but the final tagged molecules in the library may no longer have that characteristic. For example, single-stranded DNA molecules may be distributed and tagged, but the final tagged molecules in the library are likely to be double-stranded. Similarly, DNA may be subjected to distribution based on different levels of methylation, but in the final library, the tagged molecules derived from these molecules are likely to be unmethylated. Thus, the tags attached to molecules in the library typically represent the characteristics of the "parent molecule" from which the final tagged molecules are derived, but not necessarily the characteristics of the tagged molecules themselves.

[0232] As an example, barcodes 1, 2, 3, 4, etc. are used to tag and label molecules in a first distribution; barcodes A, B, C, D, etc. are used to tag and label molecules in a second distribution; barcodes a, b, c, d, etc. are used to tag and label molecules in a third distribution. Differentially tagged distributions can be pooled before sequencing. Differentially tagged distributions can be sequenced separately or can be sequenced together in parallel, for example, in the same flow cell of an Illumina sequencer.

[0233] After sequencing, analysis of reads to detect genetic variants can be performed at the level of each distribution and at the level of the entire nucleic acid population. Tags are used to sort the reads from different distributions. Analysis can include in silico analysis to determine genetic and epigenetic variations (one or more of methylation, chromatin structure, etc.) using sequence information, genome coordinate length, coverage and / or copy number. In some embodiments, higher coverage can be correlated with higher nucleosome occupancy in genome regions, while lower coverage can be correlated with lower nucleosome occupancy or nucleosome depleted regions (NDRs).

[0234] In some embodiments, the adaptors are added to the nucleic acids after partitioning the nucleic acids, while in other embodiments the adaptors may be added to the nucleic acids before partitioning the nucleic acids. In some such methods, a population of nucleic acids carrying different degrees of modification (e.g., 0, 1, 2, 3, 4, 5 or more methyl groups per nucleic acid molecule) is contacted with adaptors prior to fractionation of the population, depending on the degree of modification. b. Alternative methods of analysis of modified nucleic acids after the partitioning step

[0235] The adaptor is attached to either one or both ends of the nucleic acid molecules in the population. Preferably, the adaptor contains a sufficient number of different tags to generate a low probability of the number of tag combinations, for example, 95, 99 or 99.9% of two nucleic acids with the same start and stop points receive the same combination of tags. The adaptors may contain the same or different primer binding sites, whether carrying the same or different tags, but preferably the adaptors contain the same primer binding sites. After the adaptor is attached, the nucleic acid is contacted with an agent (e.g., such an agent as previously described) that preferentially binds to nucleic acids carrying modifications. The nucleic acid is distributed into at least two subsamples that differ in the degree to which the nucleic acid carries the modification from binding to the agent. For example, if the agent has an affinity for nucleic acids carrying modifications, nucleic acids that are over-represented with the modification (compared to the median representation in the population) will preferentially bind to the agent, while nucleic acids that are under-represented with the modification will not bind to the agent or will be more easily eluted from the agent. After the partitioning step, the first and / or second sub-samples are subjected to the method steps described elsewhere herein. Sequence data from the different partitions can then be compared.

[0236] The present disclosure provides further methods for analyzing a population of nucleic acids, where at least a portion of the nucleic acids comprises one or more modified cytosine residues, e.g., 5-methylcytosine, and any of the other modifications previously described. In these methods, after the dividing step, the nucleic acid subsample is contacted with an adaptor that comprises one or more cytosine residues modified at the 5C position, e.g., 5-methylcytosine. In some embodiments, all or all but one of the cytosine residues in such adaptors are also modified, or all such cytosines in the primer binding region of the adaptor are modified. The adaptors are attached to both ends of the nucleic acid molecules in the population. Preferably, the adaptors contain a sufficient number of different tags to produce a low probability that the number of tag combinations is low, e.g., 95, 99, or 99.9% of two nucleic acids with the same start and stop points receive the same combination of tags. The primer binding sites in such adaptors can be the same or different, but are preferably the same. After the adaptor binding, the nucleic acid is amplified from primers that bind to the primer binding sites of the adaptors. The amplified nucleic acid is divided into a first and a second aliquot. The first aliquot is assayed for sequence data with or without further processing. Thus, the sequence data for the molecules in the first aliquot is determined independently of the initial methylation state of the nucleic acid molecule. The nucleic acid molecules in the second aliquot are subjected to a procedure that affects the first nucleic acid base in the DNA differently from the second nucleic acid base in the DNA, the first nucleic acid base comprises a modified cytosine at position 5, and the second nucleic acid base comprises an unmodified cytosine. This procedure can be bisulfite treatment, or another procedure that converts the unmodified cytosine to uracil. The nucleic acid subjected to this procedure is then amplified using a primer against the original primer binding site of the adaptor linked to the nucleic acid. Only the nucleic acid molecules originally ligated to the adapter are now amplifiable (separate from their amplification products) because these nucleic acids retain cytosines in the primer binding sites of the adapters, but the amplification products have lost methylation of these cytosine residues and undergone conversion to uracil in the bisulfite treatment.Thus, only the original molecules in the population that are at least partially methylated undergo amplification. After amplification, these nucleic acids are subjected to sequence analysis. Comparison of the sequences determined from the first and second aliquots can indicate, among other things, which cytosines in the nucleic acid population have been subjected to methylation.

[0237] Such an analysis can be carried out using the following exemplary procedure: After partitioning, the methylated DNA is ligated at both ends to a Y-shaped adapter that contains a primer binding site and a tag. The cytosine in the adapter is modified at the 5-position (e.g., 5-methylated). The modification of the adapter serves to protect the primer binding site in a subsequent conversion step (e.g., bisulfite treatment, TAP conversion, or any other conversion that does not affect the modified cytosine but affects the unmodified cytosine). After the binding of the adapter, the DNA molecule is amplified. The amplification product is divided into two aliquots for sequencing with or without conversion. The aliquot that has not been subjected to conversion can be subjected to sequence analysis with or without further treatment. The other aliquot is subjected to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA, the first nucleobase containing a modified cytosine at the 5-position and the second nucleobase containing an unmodified cytosine. This procedure can be bisulfite treatment or another procedure that converts unmodified cytosines to uracils. When contacted with a primer specific to the original primer binding site, only the primer binding site protected by the modification of cytosine can support amplification. Thus, only the original molecule, not the copy from the first amplification, is subjected to further amplification. The further amplified molecule is then subjected to sequence analysis. The sequences from the two aliquots can then be compared. As with the separation scheme discussed above, the nucleic acid tag in the adapter is not used to distinguish between methylated and unmethylated DNA, but is used to distinguish nucleic acid molecules within the same distribution. 4. Deamination

[0238] In some embodiments, the step of deaminating unmodified cytosines in at least one first or second strand comprises bisulfite conversion. Treatment with bisulfite converts unmodified cytosines and certain modified cytosine nucleotides (e.g., 5-formylcytosine (fC) or 5-carboxylcytosine (caC)) to uracil, but does not convert other modified cytosines (e.g., 5-methylcytosine, 5-hydroxymethylcystosine). Sequencing of the bisulfite-treated DNA identifies positions that are read as cytosines as mC or hmC positions. Whether mC or hmC was present can be determined using additional sequence data from the other strand, as discussed above. Meanwhile, positions that are read as T are identified as T or bisulfite-sensitive forms of C, such as unmodified cytosine, 5-formylcytosine or 5-carboxylcytosine. Thus, performing bisulfite conversion as described herein facilitates identifying positions containing mC or hmC. For an exemplary description of bisulfite conversion, see, e.g., Moss et al., Nat Commun. 2018; 9: 5068.

[0239] In some embodiments, at least one unmodified cytosine in at least one first or second strand is deaminated using a deaminase, such as a cytidine deaminase, such as an APOBEC enzyme (or a fragment thereof), such as APOBEC3A, APOBEC2, APOBEC3B, APOBEC3C, APOBEC3E, APOBEC3F, APOBEC3G, APOBEC3H, APOBEC4. In some embodiments, the deaminase is APOBEC3A. For an exemplary description of the deamination procedure using APOBEC3A, see Fullgrabe, et al. (2023) Simultaneous sequencing of genetic and epigenetic bases in DNA. Nature Biotechnology, doi: 10.1038 / s41587-022-01652-0, available at www.nature.com / articles / s41587-022-01652-0.

[0240] In some embodiments, deamination comprises enzymatic conversion of unmethylated cytosines, e.g., as in EM-Seq. See, e.g., Vaisvila R, et al. (2019) EM-seq: Detection of DNA methylation at single base resolution from picograms of DNA. bioRxiv; DOI:10.1101 / 2019.12.20.884692v1, available at: www.biorxiv.org / content / 10.1101 / 2019.12.20.884692v1. For example, a TET enzyme (e.g., TET1, TET2, or TET3) and T4-βGT can be used to convert 5mC and 5hmC into a substrate that cannot be deaminated by a deaminase (e.g., APOBEC3A), which can then be used to deaminate unmodified cytosines, converting them to uracil. In some embodiments, the TET enzyme is TET2. 5. Alternative embodiments involving conversion of methylated and / or hydroxymethylated cytosines

[0241] Methods are also provided herein that use alternative base conversion schemes.For example, unmethylated cytosine can remain intact, while methylated cytosine and hydroxymethylcytosine are converted to base reading as thymine (e.g., uracil, thymine or dihydrouracil).See, for example, Figure 1D.

[0242] Thus, there is provided a method for analyzing DNA molecules in a sample, the DNA molecules comprising a first and a second strand and an asymmetric adapter, the method comprising the steps of: a) oxidizing 5-hydroxymethylated cytosines to 5-formylcytosines (e.g., oxidizing 5-hydroxymethylcytosines in the first and second strands with a ruthenium salt, e.g., KRuO 4 b) synthesizing a first complementary strand complementary to the first strand and a second complementary strand complementary to the second strand; c) methylating cytosines in at least one of the first complementary strand or the second complementary strand, where the methylation converts a hemi-methylated CpG to a fully methylated CpG; d) converting a modified cytosine in at least one of the first or second strand to a thymine or a base that is read as a thymine, thereby producing a processed DNA molecule; and e) sequencing at least a portion of the processed DNA molecule; optionally, the asymmetric adaptor is a Y-shaped adaptor or a bubble adaptor.

[0243] In some embodiments, methylating cytosines in at least one of the first or second complementary strands comprises contacting the cytosines with a methyltransferase, e.g., DNMT1 or DNMT5. In such embodiments, oxidizing 5-hydroxymethylated cytosines to 5-formylcytosines (e.g., oxidizing 5-hydroxymethylcytosines in the first and second strands with KRuO4 (by contacting with) may be optional.

[0244] In some embodiments, at least one asymmetric adaptor comprises a modified cytosine, e.g., a methylated or hydroxymethylated cytosine, which functions as a strand reporter base similar to the deamination-sensitive cytosine used as a reporter in the method discussed in the preceding section. In some embodiments, each asymmetric adaptor comprises a modified cytosine, e.g., a methylated or hydroxymethylated cytosine. In some embodiments, an asymmetric adaptor comprises one modified cytosine, e.g., a methylated or hydroxymethylated cytosine, and all other cytosines in the asymmetric adaptor are unmodified. In some embodiments, each asymmetric adaptor comprises one modified cytosine, e.g., a methylated or hydroxymethylated cytosine, and all other cytosines in the asymmetric adaptor are unmodified. In some embodiments, the modified cytosine is in the strand of the asymmetric adaptor that undergoes ligation to the 5' end of the sample or insert DNA molecule. In some embodiments, the nucleotide immediately 3' to the modified cytosine contains a nucleobase other than guanine, e.g., adenine, cytosine, thymine, or uracil (including modified forms thereof, e.g., methylated cytosine). Thus, the modified cytosine is not part of a CpG and is not recognized as a substrate by a methyltransferase, e.g., DNMT1.

[0245] In some embodiments, the step of converting the modified cytosine in at least one of the first or second strands to thymine or a base that is read as thymine comprises oxidizing hydroxymethylcytosine, e.g., hydroxymethylcytosine is oxidized to formylcytosine. In some embodiments, oxidizing hydroxymethylcytosine to formylcytosine comprises oxidizing hydroxymethylcytosine to formylcytosine with a ruthenium salt, e.g., potassium ruthenate (KRuO 4 ) in contact with the

[0246] In some embodiments, the modified cytosine is converted to thymine, uracil, or dihydrouracil.

[0247] In some embodiments, the method includes converting formylcytosine and / or methylcytosine to carboxylcytosine as part of converting at least one modified cytosine in the first or second strand to thymine or a base that is read as thymine. For example, converting formylcytosine and / or methylcytosine to carboxylcytosine may include contacting formylcytosine and / or methylcytosine with a TET enzyme, such as TET1, TET2 or TET3. In some embodiments, the method includes reducing carboxylcytosine as part of converting at least one modified cytosine in the first or second strand to thymine or a base that is read as thymine, and / or the carboxylcytosine is reduced to dihydrouracil. In some embodiments, reducing carboxylcytosine includes contacting carboxylcytosine with a borane or borohydride reducing agent.

[0248] In some embodiments, the borane or borohydride reducing agent is pyridine borane, 2-picoline borane, borane, tert-butylamine borane, ammonia borane, sodium borohydride, sodium cyanoborohydride (NaBH 3 CN), lithium borohydride (LiBH 4), ethylenediamine borane, dimethylamine borane, sodium triacetoxyborohydride, morpholine borane, 4-methylmorpholine borane, trimethylamine borane, dicyclohexylamine borane, or a salt thereof. In other embodiments, the reducing agent comprises lithium aluminum hydride, sodium amalgam, amalgam, sulfur dioxide, dithionate, thiosulfate, iodide, hydrogen peroxide, hydrazine, diisobutylaluminum hydride, oxalic acid, carbon monoxide, cyanide, ascorbic acid, formic acid, dithiothreitol, beta-mercaptoethanol, or any combination thereof. 6. Enrichment / Capture Step; Captured Set

[0249] In some embodiments, the method disclosed herein comprises capturing one or more sets of target regions of DNA, for example, cfDNA.Capturing can be carried out using any suitable approach known in the art.In some embodiments, capturing one or more sets of target regions of DNA is carried out after deamination.Capturing can be carried out, for example, using a group of target-specific probes as described elsewhere herein.

[0250] In some embodiments, the capture step is performed before the deamination step. In such embodiments, when target-specific probes specific to a set of sequence-variable target regions and / or target-specific probes specific to a set of epigenetic target regions are used, these probes contain sequences complementary to the non-deaminated target regions. In some embodiments, DNA molecules from the sample are not amplified until after capture and deamination.

[0251] In some embodiments, the capture step is performed after the deamination step. In such embodiments, when target-specific probes specific to a set of sequence-variable target regions and / or target-specific probes specific to a set of epigenetic target regions are used, these probes may contain sequences complementary to the deaminated target regions or sequences complementary to the deaminated target regions at non-CpG positions. In some embodiments, the probes include a first subpopulation that contains sequences complementary to the deaminated target regions and a second subpopulation that contains sequences complementary to the deaminated target regions at non-CpG positions only. Sequences complementary to the deaminated target regions generally use A residues (hybridizing to uridine resulting from deamination of cytosine, or thymidine in progeny molecules) instead of G residues.

[0252] In some embodiments, the capture step is performed before the step of converting modified cytosines in at least one first or second strand to thymine or a base that is read as thymine. In such an embodiment, when target-specific probes specific to a set of sequence-variable target regions and / or target-specific probes specific to a set of epigenetic target regions are used, these probes comprise sequences complementary to unconverted target regions. In some embodiments, DNA molecules from a sample are not amplified until after the capture and the step of converting modified cytosines in at least one first or second strand to thymine or a base that is read as thymine.

[0253] In some embodiments, the capture step is performed after the step of converting modified cytosines in at least one first or second strand to thymine or a base that is read as thymine. In such an embodiment, when target-specific probes specific to a set of sequence-variable target regions and / or target-specific probes specific to a set of epigenetic target regions are used, these probes comprise sequences that are complementary to the converted target regions. In some embodiments, the probes comprise a first subpopulation that comprises sequences that are complementary to the unconverted target regions and a second subpopulation that comprises sequences that are complementary to the converted target regions. The sequences that are complementary to the converted target regions generally use, for example, A residues in CpG positions instead of G residues (which hybridize to thymine or a base that is read as thymine resulting from the conversion of the modified cytosine).

[0254] In some embodiments, the capturing step comprises contacting the DNA to be captured with a set of target-specific probes. The set of target-specific probes may have any of the features described herein for a set of target-specific probes, including but not limited to the embodiments shown above and in the section on probes below. The capturing step may be performed on one or more aliquots prepared during the method disclosed herein. In some embodiments, DNA is captured from at least the first aliquot or the second aliquot, for example, at least the first aliquot and the second aliquot. In some embodiments, the aliquots are differentially tagged (e.g., as described herein) and then pooled before undergoing capture.

[0255] The capturing step can be carried out using conditions suitable for specific nucleic acid hybridization, which generally depend to some extent on the characteristics of the probe, such as length, base composition, etc. Those skilled in the art are familiar with suitable conditions, taking into account general knowledge in the art of nucleic acid hybridization. In some embodiments, a complex between the target-specific probe and DNA is formed.

[0256] In some embodiments, the methods described herein include capturing cfDNA obtained from a subject for a plurality of sets of target regions. The target regions include epigenetic target regions that may show differences in methylation levels and / or fragmentation patterns depending on whether they originate from a tumor or a healthy cell. The target regions also include sequence-variable target regions that may show differences in sequence depending on whether they originate from a tumor or a healthy cell. The capturing step generates a captured set of cfDNA molecules, and cfDNA molecules corresponding to the sequence-variable target region set are captured with a higher capture yield in the captured set of cfDNA molecules than cfDNA molecules corresponding to the epigenetic target region set. For additional discussion of the capturing step, capture yield, and related aspects, see WO2020 / 160414, which is hereby incorporated by reference for all purposes.

[0257] In some embodiments, the methods described herein include contacting cfDNA obtained from a subject with a set of target-specific probes configured to capture cfDNA corresponding to a set of sequence-variable target regions with a higher capture yield than cfDNA corresponding to a set of epigenetic target regions.

[0258] It may be beneficial to capture cfDNA corresponding to a set of sequence-variable target regions with a higher capture yield than cfDNA corresponding to a set of epigenetic target regions, because a deeper sequencing depth than that required when analyzing epigenetic target regions may be required to analyze sequence-variable target regions with sufficient confidence or accuracy.The volume of data required to determine fragmentation patterns (e.g., to test for disruption of transcription start sites or CTCF binding sites) or fragment abundances (e.g., in hypermethylated and hypomethylated distributions) is generally smaller than the volume of data required to determine the presence or absence of cancer-related sequence mutations.Capturing a set of target regions with different yields may facilitate sequencing target regions to different depths of sequencing in the same sequencing run (e.g., using pooled mixtures and / or in the same sequencing cell).

[0259] In various embodiments, the methods further include sequencing the captured cfDNA, e.g., to different degrees of sequencing depth for the set of epigenetic target regions and the set of sequence variable target regions, consistent with the discussion herein.

[0260] In some embodiments, the complex of target-specific probe and DNA is separated from the DNA that is not bound to the target-specific probe.For example, when the target-specific probe is covalently or non-covalently bound to solid support, washing or suction step can be used to separate unbound material.Alternatively, when the complex has a chromatographic property that is different from that of unbound material (for example, when the probe contains a ligand that binds to a chromatographic resin), chromatography can be used.

[0261] As discussed in detail elsewhere herein, the set of target-specific probes may include multiple sets, for example, probes for sequence variable target regions and probes for epigenetic target regions. In some such embodiments, the capturing step is performed simultaneously in the same vessel with the probes for sequence variable target regions and the probes for epigenetic target regions, for example, the probes for sequence variable target regions and the probes for epigenetic target regions are in the same composition. This approach provides a relatively streamlined workflow. In some embodiments, the concentration of the probes for the sequence variable target regions is higher than the concentration of the probes for the epigenetic target regions.

[0262] Alternatively, the capturing step is performed with a sequence variable target region probe set in a first container and an epigenetic target region probe set in a second container, or the contacting step is performed with a sequence variable target region probe set at a first time point and in the first container, and with an epigenetic target region probe set at a second time point before or after the first time point. This approach allows for the preparation of separate first and second compositions that contain captured DNA corresponding to the sequence variable target region set and captured DNA corresponding to the epigenetic target region set. These compositions can be processed separately as desired (e.g., to fractionate based on methylation, as described elsewhere herein), and recombined in appropriate proportions to provide material for further processing and analysis, e.g., sequencing.

[0263] In some embodiments, a captured set of DNA (e.g., cfDNA) is provided. In relation to the disclosed method, the captured set of DNA can be provided, for example, by carrying out a capturing step before the sequencing step described herein. For example, the capturing step can be carried out after one or more of the synthesis step, the glycosylation step, the methylation step or the deamination step, or, if applicable, after the division step. The captured set can include DNA corresponding to a sequence variable target region set, an epigenetic target region set, or a combination thereof. In some embodiments, the capturing step is carried out before the conversion step or after the conversion step.

[0264] In some embodiments, the first set of target regions is captured (e.g., from a sample or a first subsample) and includes at least an epigenetic target region. The epigenetic target region captured from the first subsample may include a hypermethylated variable target region. In some embodiments, the hypermethylated variable target region is a CpG-containing region that is unmethylated or has low methylation (e.g., methylation below average compared to bulk cfDNA) in cfDNA from a healthy subject. In some embodiments, the hypermethylated variable target region is a region in healthy cfDNA that shows lower methylation than in at least one other tissue type. Without wishing to be bound by any particular theory, cancer cells may shed more DNA into the bloodstream than healthy cells of the same tissue type. Thus, the distribution of tissues from which cfDNA originates may change during cancer development. Thus, an increase in the level of a hypermethylated variable target region in the first subsample may be an indication of the presence (or recurrence, depending on the subject's medical history) of cancer. Alternatively or additionally, the epigenetic target region may include a hydroxymethylated variable target region. Hydroxymethylated variable target regions are similar to hypermethylated variable target regions, except that the associated modification is hydroxymethylation rather than hypermethylation.

[0265] In some embodiments, the second set of target regions is captured from the second subsample, which includes at least epigenetic target regions. The epigenetic target regions may include hypomethylated variable target regions. In some embodiments, the hypomethylated variable target regions are CpG-containing regions that are methylated or have high methylation (e.g., methylation above average compared to bulk cfDNA) in cfDNA from healthy subjects. In some embodiments, the hypomethylated variable target regions are regions in healthy cfDNA that show higher methylation than in at least one other tissue type. Without wishing to be bound by any particular theory, cancer cells may shed more DNA into the bloodstream than healthy cells of the same tissue type. Thus, the distribution of tissues from which cfDNA originates may change during cancer development. Thus, an increase in the level of hypomethylated variable target regions in the second subsample may be an indication of the presence (or recurrence, depending on the subject's medical history) of cancer.

[0266] In some embodiments, the amount of captured sequence variable target region DNA is greater than the amount of captured epigenetic target region DNA when normalized for differences in size (footprint size) of the targeted regions.

[0267] Alternatively, a first and second captured set may be provided that includes DNA corresponding to a set of sequence variable target regions and DNA corresponding to a set of epigenetic target regions, respectively. The first and second captured sets may be combined to provide a combined captured set.

[0268] In some embodiments, where the captured set includes DNA corresponding to the set of sequence variable target regions and DNA corresponding to the set of epigenetic target regions, including the combined captured sets discussed above, the DNA corresponding to the set of sequence variable target regions is at a higher concentration than the DNA corresponding to the set of epigenetic target regions, e.g., 1.1-fold to 1.2-fold higher, 1.2-fold to 1.4-fold higher, 1.4-fold to 1.6-fold higher, 1.6-fold to 1.8-fold higher, 1.8-fold to 2.0-fold higher, 2.0-fold to 2.2-fold higher, 2.2-fold to 2.4-fold higher. 2.4x to 2.6x higher concentration, 2.6x to 2.8x higher concentration, 2.8x to 3.0x higher concentration, 3.0x to 3.5x higher concentration, 3.5x to 4.0x, 4.0x to 4.5x higher concentration, 4.5x to 5.0x higher concentration, 5.0x to 5.5x higher concentration, 5.5x to 6.0x higher concentration, 6.0x to 6.5x higher concentration, 6.5x to 7.0x higher, 7.0x to 7.5x higher concentration, 7.5x to 8.0x higher concentration, 8.0x to 8.5x higher concentration, 8.5x to 9.0x higher concentration, 9.0x to 9.5x higher concentration, 9.5x to 10.0x higher concentration, 10x to 11x higher concentration, 11x to 12x higher concentration It may be present at 12-13x higher concentration, 13-14x higher concentration, 14-15x higher concentration, 15-16x higher concentration, 16-17x higher concentration, 17-18x higher concentration, 18-19x higher concentration, 19-20x higher concentration, 20-30x higher concentration, 30-40x higher concentration, 40-50x higher concentration, 50-60x higher concentration, 60-70x higher concentration, 70-80x higher concentration, 80-90x higher concentration or 90-100x higher concentration. The degree of difference in concentration is the primary cause of standardization with respect to the footprint size of the target region, as discussed in the definition section. A. Epigenetic target region set

[0269] The epigenetic target region set may include one or more types of target regions that are likely to differentiate DNA from neoplastic (e.g., tumor or cancer) cells from DNA from healthy cells, e.g., non-neoplastic circulating cells. Exemplary types of such regions are discussed in detail herein. The epigenetic target region set may also include one or more control regions, e.g., as described herein.

[0270] In some embodiments, the set of epigenetic target regions has a footprint of at least 100 kbp, e.g., at least 200 kbp, at least 300 kbp, or at least 400 kbp. In some embodiments, the set of epigenetic target regions has a footprint in the range of 100-20 Mbp, e.g., 100-200 kbp, 200-300 kbp, 300-400 kbp, 400-500 kbp, 500-600 kbp, 600-700 kbp, 700-800 kbp, 800-900 kbp, 900-1,000 kbp, 1-1.5 Mbp, 1.5-2 Mbp, 2-3 Mbp, 3-4 Mbp, 4-5 Mbp, 5-6 Mbp, 6-7 Mbp, 7-8 Mbp, 8-9 Mbp, 9-10 Mbp, or 10-20 Mbp. In some embodiments, the set of epigenetic target regions has a footprint of at least 20 Mbp. i. Hypermethylated and / or hydroxymethylated variable target regions

[0271] In some embodiments, the epigenetic target region set includes one or more hypermethylated variable target regions. In general, a hypermethylated variable target region refers to a region where, for example, an increase in the level of observed methylation in a cfDNA sample indicates an increased likelihood that the sample (e.g., of cfDNA) contains DNA produced by a neoplastic cell, e.g., a tumor or cancer cell. For example, hypermethylation of promoters of tumor suppressor genes has been repeatedly observed. See, e.g., Kang et al., Genome Biol. 18:53 (2017) and references cited therein. In another example, as discussed above, a hypermethylated variable target region may include a region that is not necessarily differently methylated in cancerous tissue compared to DNA from a healthy tissue of the same type, but is differentially methylated (e.g., has more methylation) compared to typical cfDNA in a healthy subject. For example, if the presence of a cancer results in increased cell death, e.g., apoptosis, of cells of a tissue type corresponding to the cancer, such cancer may be detected, at least in part, using such hypermethylated variable target regions. Alternatively, or in addition, the epigenetic target region may comprise a hydroxymethylation variable target region, which is similar to a hypermethylation variable target region, except that the associated modification is hydroxymethylation rather than hypermethylation.

[0272] An extensive discussion of methylation variable target regions in colorectal cancer is provided in Lam et al., Biochim Biophys Acta. 1866:106-20 (2016). These include VIM, SEPT9, ITGA4, OSM4, GATA4 and NDRG4. An exemplary set of hypermethylated variable target regions based on colorectal cancer (CRC) studies is provided in Table 1. Many of these genes likely have relevance to cancer beyond colorectal cancer; for example, TP53 is widely recognized as a critical tumor suppressor, and hypermethylation-based inactivation of this gene may be a common oncogenic mechanism.

[0273] Table 1. Exemplary hypermethylation target regions based on CRC studies [Table 1]

[0274] In some embodiments, the hypermethylated variable target region comprises a plurality of loci listed in Table 1, e.g., at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90% or 100% of the loci listed in Table 1. For example, for each locus included as a target region, there may be one or more probes having hybridization sites that bind between the transcription start site and the stop codon of the gene (for alternatively spliced ​​genes, the final stop codon) or in the promoter region of the gene. In some embodiments, the one or more probes bind within 300 bp, e.g., within 200 or 100 bp, of the transcription start site of a gene in Table 1.

[0275] Methylation variable target regions in various types of lung cancer are described, for example, in Ooki et al., Clin. Cancer Res. 23:7141-52 (2017);Belinksy, Annu. Rev. Physiol. 77:453-74 (2015);Hulbert et al., Clin. Cancer Res. 23:1998-2005 (2017);Shi et al., BMC Genomics 18:901 (2017);Schneider et al., BMC Cancer. 11:102 (2011);Lissa et al., Transl Lung Cancer Res 5(5):492-504 (2016);Skvortsova et al., Br. J. Cancer. 94(10):1492-1495 (2006);Kim et al., Cancer Res. 61:3419-3424 (2001);Furonaka et al., Pathology International 55:303-309 (2005);Gomes et al., Rev. Port. Pneumol. 20:20-30 (2014);Kim et al., Oncogene. 20:1765-70 (2001);Hopkins-Donaldson et al., Cell Death Differ. 10:356-64 (2003);Kikuchi et al., Clin. Cancer Res. 11:2954-61 (2005);Heller et al., Oncogene 25:959-968 (2006);Licchesi et al., Carcinogenesis. 29:895-904 (2008);Guo et al., Clin. Cancer Res. 10:7917-24 (2004);Palmisano et al., Cancer Res. 63:4620-4625 (2003); and Toyooka et al., Cancer Res. 61:4556-4560, (2001).

[0276] An exemplary set of hypermethylated variable target regions based on lung cancer studies is provided in Table 2. Many of these genes likely have relevance to cancer beyond lung cancer; for example, Casp8 is a key enzyme in programmed cell death, and hypermethylation-based inactivation of this gene may be a common oncogenic mechanism not limited to lung cancer. Additionally, several genes appear in both Tables 1 and 2, indicating universality.

[0277] Table 2. Exemplary hypermethylated target regions based on lung cancer studies [Table 2-1] [Table 2-2]

[0278] Any of the above-described embodiments relating to target regions identified in Table 2 may be combined with any of the above-described embodiments relating to target regions identified in Table 1. In some embodiments, the hypermethylated variable target regions include a plurality of loci listed in Table 1 or Table 2, e.g., at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90% or 100% of the loci listed in Table 1 or Table 2.

[0279] Additional hypermethylated target regions can be obtained, for example, from the Cancer Genome Atlas. Kang et al., Genome Biology 18:53 (2017) describes the construction of a probabilistic method called CancerLocator, which uses hypermethylated target regions from breast, colon, kidney, liver and lung. In some embodiments, the hypermethylated target regions can be specific to one or more types of cancer. Thus, in some embodiments, the hypermethylated target regions include one, two, three, four or five subsets of hypermethylated target regions that collectively show hypermethylation in one, two, three, four or five of breast, colon, kidney, liver and lung cancers.

[0280] In some embodiments, when different epigenetic target regions are captured from the first and second sub-samples, the epigenetic target region captured from the first sub-sample comprises a hypermethylated variable target region. ii. Hypomethylated variable target regions

[0281] Global hypomethylation is a commonly observed phenomenon in various cancers. See, e.g., Hon et al., Genome Res. 22:246-258 (2012) (breast cancer); Ehrlich, Epigenomics 1:239-259 (2009) (colon, ovarian, prostate, leukemia, hepatocellular and cervical cancers). See, for example, a review article that focuses on the observation of hypomethylation in tumor cells. For example, regions, such as repetitive elements, such as LINE1 elements, Alu elements, centromeric tandem repeats, pericentromeric tandem repeats, and satellite DNA, as well as intergenic regions that are normally methylated in healthy cells, may show reduced methylation in tumor cells. Thus, in some embodiments, the epigenetic target region set includes hypomethylated variable target regions, where a decrease in the level of observed methylation indicates an increased likelihood that the sample (e.g., of cfDNA) contains DNA produced by neoplastic cells, such as tumor or cancer cells. In another example, as discussed above, hypomethylated variable target regions may include regions that are not necessarily differentially methylated in cancerous tissue compared to DNA from healthy tissue of the same type, but are differentially methylated (e.g., less methylated) compared to typical cfDNA in healthy subjects. For example, if the presence of a cancer results in increased cell death, e.g., apoptosis, of cells of the tissue type corresponding to the cancer, then such cancer may be detected, at least in part, using such hypomethylated variable target regions.

[0282] In some embodiments, the hypomethylated variable target region comprises a repetitive element and / or an intergenic region, hi some embodiments, the repetitive element comprises one, two, three, four or five of a LINE1 element, an Alu element, a centromeric tandem repeat, a pericentromeric tandem repeat and / or satellite DNA.

[0283] Exemplary specific genomic regions that exhibit cancer-associated hypomethylation include nucleotides 8403565-8953708 and 151104701-151106035 of human chromosome 1. In some embodiments, the hypomethylated variable target region overlaps with or includes one or both of these regions.

[0284] In some embodiments, when different epigenetic target regions are captured from the first and second sub-samples, the epigenetic target region captured from the second sub-sample comprises a hypomethylated variable target region. In some embodiments, the epigenetic target region captured from the second sub-sample comprises a hypomethylated variable target region and the epigenetic target region captured from the first sub-sample comprises a hypermethylated variable target region. iii.CTCF binding region

[0285] CTCF is a DNA-binding protein that contributes to chromatin organization and often colocalizes with cohesin. Disruption of CTCF binding sites has been reported in a variety of different cancers. See, for example, Katainen et al., Nature Genetics, doi:10.1038 / ng.3335, published online June 8, 2015; Guo et al., Nat. Commun. 9:1520 (2018). CTCF binding produces recognizable patterns in cfDNA that can be detected, for example, by sequencing through fragment length analysis. Details regarding sequencing-based fragment length analysis are provided in Snyder et al., Cell 164:57-68 (2016); WO2018 / 009723; and US20170211143A1, each of which is hereby incorporated by reference herein.

[0286] Thus, disruption of CTCF binding results in variation in the fragmentation pattern of cfDNA. Thus, CTCF binding sites represent one type of fragmentation variable target region.

[0287] There are many known CTCF binding sites.See, for example, CTCFBSDB (CTCF Binding Site Database) available on the Internet at insulatordb.uthsc.edu / ;Cuddapah et al., Genome Res. 19:24-32 (2009);Martin et al., Nat. Struct. Mol. Biol. 18:708-14 (2011);Rhee et al., Cell. 147:1408-19 (2011).Each of these is incorporated by reference.Exemplary CTCF binding sites are located at nucleotides 56014955-56016161 on chromosome 8 and 95359169-95360473 on chromosome 13.

[0288] Thus, in some embodiments, the set of epigenetic target regions comprises CTCF binding regions, in some embodiments, the CTCF binding regions include at least 10, 20, 50, 100, 200, or 500 CTCF binding regions, or 10-20, 20-50, 50-100, 100-200, 200-500, or 500-1000 CTCF binding regions, such as those described in the CTCFBSDB or one or more of the articles by Cuddapah et al., Martin et al., or Rhee et al., cited above or cited above.

[0289] In some embodiments, at least a portion of the CTCF sites can be methylated or unmethylated, and the methylation status correlates with whether the cell is a cancer cell. In some embodiments, the set of epigenetic target regions includes at least 100 bp, at least 200 bp, at least 300 bp, at least 400 bp, at least 500 bp, at least 750 bp, at least 1000 bp upstream and downstream of the CTCF binding site. iv. Transcription start site

[0290] Transcription start sites may also show perturbations in neoplastic cells. For example, hematopoietic lineages contribute substantially to cfDNA in healthy individuals, but the nucleosome organization at various transcription start sites in the healthy cells may differ from the nucleosome organization at the transcription start sites in the neoplastic cells. This results in different cfDNA patterns that can be detected by sequencing, as generally discussed in Snyder et al., Cell 164:57-68 (2016); WO2018 / 009723; and US20170211143A1. In another example, transcription start sites may not necessarily be epigenetically different in cancerous tissue compared to DNA from healthy tissue of the same type, but may be epigenetically different (e.g., in terms of nucleosome organization) compared to typical cfDNA in healthy subjects. For example, if the presence of cancer results in increased cell death, e.g., apoptosis, of cells of the tissue type corresponding to the cancer, such cancer may be detected, at least in part, using such transcription start sites.

[0291] Thus, perturbations in transcription start sites also result in variations in the fragmentation patterns of cfDNA. Thus, transcription start sites also represent a type of fragmentation variable target region.

[0292] Human transcription start sites are available from the Database of Human Transcription Start Sites (DBTSS) described in Yamashita et al., Nucleic Acids Res. 34(Database issue): D86-D89 (2006), available on the Internet at dbtss.hgc.jp, and hereby incorporated by reference.

[0293] Thus, in some embodiments, the set of epigenetic target regions includes a transcription start site. In some embodiments, the transcription start site includes at least 10, 20, 50, 100, 200, or 500 transcription start sites, or 10-20, 20-50, 50-100, 100-200, 200-500, or 500-1000 transcription start sites, such as those listed in the DBTSS. In some embodiments, at least a portion of the transcription start sites can be methylated or unmethylated, and the methylation status correlates with whether the cell is a cancer cell. In some embodiments, the set of epigenetic target regions includes at least 100 bp, at least 200 bp, at least 300 bp, at least 400 bp, at least 500 bp, at least 750 bp, at least 1000 bp upstream and downstream of the transcription start site. v. Localized amplification

[0294] Although focal amplification is somatic mutation, it can be detected by sequencing based on read frequency in a similar manner to the approach of detecting certain epigenetic changes, such as methylation changes.Therefore, the region that may show focal amplification in cancer can be included in epigenetic target region set, and can include one or more of AR, BRAF, CCND1, CCND2, CCNE1, CDK4, CDK6, EGFR, ERBB2, FGFR1, FGFR2, KIT, KRAS, MET, MYC, PDGFRA, PIK3CA and RAF1.For example, in some embodiments, epigenetic target region set includes at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17 or 18 of the above-mentioned targets. vi. Methylation control region

[0295] It may be useful to include a control region to facilitate data validation. In some embodiments, the epigenetic target region set includes a control region that is expected to be methylated or unmethylated in essentially all samples, regardless of whether the DNA is from cancer cells or normal cells. In some embodiments, the epigenetic target region set includes a hypomethylated control region that is expected to be hypomethylated in essentially all samples. In some embodiments, the epigenetic target region set includes a hypermethylated control region that is expected to be hypermethylated in essentially all samples. b. Set of sequence variable target regions

[0296] In some embodiments, the captured set comprises a set of sequence variable target regions. In some embodiments, the set of sequence variable target regions comprises a plurality of regions known to undergo somatic mutation in cancer.

[0297] In some embodiments, the sequence variable target region set targets a plurality of different genes or genomic regions ("panels") selected such that a determined percentage of subjects with cancer exhibits genetic variants or tumor markers in one or more different genes or genomic regions in the panel. The panels can be selected to limit the regions to be sequenced to a fixed number of base pairs. The panels can be selected to sequence a desired amount of DNA, for example, by adjusting the affinity and / or amount of probes, as described elsewhere herein. The panels can further be selected to achieve a desired sequence read depth. The panels can be selected to achieve a desired sequence read depth or sequence read coverage for a certain amount of sequenced base pairs. The panels can be selected to achieve a theoretical sensitivity, theoretical specificity and / or theoretical accuracy for detecting one or more genetic variants in a sample.

[0298] The probes for detecting the panel of regions may include probes for detecting genomic regions of interest (hotspot regions) as well as nucleosome recognition probes (e.g., KRAS codons 12 and 13), and may be designed to optimize capture based on the analysis of cfDNA coverage and fragment size variation influenced by nucleosome binding patterns and GC sequence composition. The regions used herein may also include non-hotspot regions that are optimized based on nucleosome positions and GC models.

[0299] Examples of lists of genomic locations of interest can be found in Tables 3 and 4. In some embodiments, the sequence variable target region set used in the methods of the present disclosure comprises at least a portion of at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, or 70 of the genes in Table 3. In some embodiments, the sequence variable target region set used in the methods of the present disclosure comprises at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, or 70 of the SNVs in Table 3. In some embodiments, the sequence variable target region set used in the methods of the present disclosure comprises at least 1, at least 2, at least 3, at least 4, at least 5, or 6 of the fusions in Table 3. In some embodiments, the sequence variable target region set used in the disclosed method comprises at least a portion of at least 1, at least 2, or 3 of the indels in Table 3. In some embodiments, the sequence variable target region set used in the disclosed method comprises at least a portion of at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, or 73 of the genes in Table 4. In some embodiments, the sequence variable target region set used in the disclosed method comprises at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, or 73 of the SNVs in Table 4. In some embodiments, the sequence variable target region set used in the disclosed method comprises at least 1, at least 2, at least 3, at least 4, at least 5, or 6 of the fusions in Table 4.In some embodiments, the sequence variable target region set used in the method of the present disclosure comprises at least a portion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, or at least 18 of the indels in Table 4. Each of these genomic locations of interest can be identified as a scaffold region or a hotspot region for a given panel. An example of a list of hotspot genomic locations of interest can be found in Table 5. In some embodiments, the sequence variable target region set used in the method of the present disclosure comprises at least a portion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, or at least 20 of the genes in Table 5. Each hotspot genomic region is listed with several characteristics, including the associated gene, the chromosome on which it resides, the genomic start and stop positions representing the locus, the length of the locus in base pairs, the exons covered by the gene, and important features that a given genomic region of interest may seek to capture (e.g., types of mutations). Table 3 [Table 3] Table 4 [Table 4-1] [Table 4-2] Table 5 [Table 5-1] [Table 5-2] [Table 5-3]

[0300] Additionally or alternatively, suitable target region set can be obtained from literature.For example, Gale et al., PLoS One 13: e0194630 (2018), which is hereby incorporated by reference, describes a panel of 35 cancer-related gene targets that can be used as part or all of sequence variable target region set.These 35 targets are AKT1, ALK, BRAF, CCND1, CDK2A, CTNNB1, EGFR, ERBB2, ESR1, FGFR1, FGFR2, FGFR3, FOXL2, GATA3, GNA11, GNAQ, GNAS, HRAS, IDH1, IDH2, KIT, KRAS, MED12, MET, MYC, NFE2L2, NRAS, PDGFRA, PIK3CA, PPP2R1A, PTEN, RET, STK11, TP53 and U2AF1.

[0301] In some embodiments, the set of sequence variable target regions includes target regions from at least 10, 20, 30, or 35 cancer associated genes, such as those listed above.

[0302] In some embodiments, the set of sequence variable target regions has a footprint of at least 50 kbp, e.g., at least 100 kbp, at least 200 kbp, at least 300 kbp, or at least 400 kbp. In some embodiments, the set of sequence variable target regions has a footprint in the range of 100-2000 kbp, e.g., 100-200 kbp, 200-300 kbp, 300-400 kbp, 400-500 kbp, 500-600 kbp, 600-700 kbp, 700-800 kbp, 800-900 kbp, 900-1,000 kbp, 1-1.5 Mbp, or 1.5-2 Mbp. In some embodiments, the set of sequence variable target regions has a footprint of at least 2 Mbp. 7. Sequencing

[0303] In general, the adaptor-flanked sample nucleic acid can be subjected to sequencing with or without prior amplification. Sequencing methods include, for example, Sanger sequencing, high-throughput sequencing, pyrosequencing, sequencing by synthesis, single molecule sequencing, nanopore sequencing, semiconductor sequencing, sequencing by ligation, sequencing by hybridization, digital gene expression (Helicos), next-generation sequencing (NGS), single molecule sequencing by synthesis (SMSS) (Helicos), massively parallel sequencing, clonal single molecule array (Solexa), shotgun sequencing, Ion Torrent, Oxford Nanopore, Roche Genia, Maxim-Gilbert sequencing, primer walking, and sequencing using PacBio, SOLiD, Ion Torrent or Nanopore platforms. Sequencing reactions can be carried out in various sample processing units, which can be multiple lanes, multiple channels, multiple wells, or other means of processing multiple sample sets substantially simultaneously. Sample processing units can also include multiple sample chambers that can process multiple runs simultaneously.

[0304] In some embodiments, the step of sequencing comprises detecting and / or identifying unmodified and modified nucleic acid bases. For example, long-read sequencing (also referred to herein as third generation sequencing) methods include methods that can generate longer sequencing reads, e.g., reads of more than 10 kilobases, compared to short-read sequencing methods that generally generate reads of up to about 600 bases in length. Compared to short reads, long reads can improve de novo assembly, transcript isoform identification, and detection and / or mapping of structural variants. Furthermore, long-read sequencing of native DNA or RNA molecules reduces amplification bias and preserves base modifications, e.g., methylation status. Long-read sequencing techniques useful herein can include any suitable long-read sequencing method, including, without limitation, Pacific Biosciences (PacBio) single molecule real-time (SMRT) sequencing, Oxford Nanopore Technologies (ONT) nanopore sequencing, and synthetic long-read sequencing approaches, e.g., linked reads, proximity ligation strategies, and optical mapping. The synthetic long-read approach involves the assembly of short reads from the same DNA molecule to generate synthetic long reads and can be used in conjunction with "true" long-read sequencing technologies, such as SMRT and nanopore sequencing.

[0305] Single molecule real-time (SMRT) sequencing, for example, facilitates the direct detection of 5-methylcytosine and 5-hydroxymethylcytosine as well as unmodified cytosine (Weirather JL, et al., "Comprehensive comparison of PacificBiosciences and oxford Nanopore Technologies and their applications to transcriptome analysis," F1000Research, 6:100, 2017). Next-generation sequencing detects increased signals from clonal populations of amplified DNA fragments, whereas SMRT sequencing captures single DNA molecules and maintains base modifications during sequencing. Because the signal-to-noise ratio from a single DNA molecule is not high, the error rate of raw data generated by PacBio SMRT sequencing is about 13-15%. To increase accuracy, this platform uses a circular DNA template by ligating hairpin adapters to both ends of the target double-stranded DNA. As the polymerase repeatedly traverses and replicates the circular molecule, the DNA template is sequenced multiple times to generate continuous long reads (CLRs). The CLRs can be split into multiple reads ("subreads") by removing adapter sequences, which generate circular consensus sequence ("CCS") reads with higher accuracy. The average length of a CLR is >10 kb and up to 60 kb, with the length depending on the polymerase lifetime. Thus, the length and accuracy of the CCS reads depend on the fragment size. PacBio sequencing has been utilized for genomic (e.g., de novo assembly, structural variant detection, and haplotyping) and transcriptomic (e.g., gene isoform reconstruction and novel gene / isoform discovery) studies.

[0306] ONT is a nanopore-based single molecule sequencing technology (Weirather JL, et al., F1000Research, 6:100, 2017). ONT directly sequences native single-stranded DNA (ssDNA) molecules by measuring the characteristic current changes as the bases are threaded through the nanopore by molecular motor proteins. ONT uses a hairpin library structure similar to the PacBio circular DNA template: a DNA template, whose complement is bound by a hairpin adapter. Thus, the DNA template passes through the nanopore, then the hairpin, and finally the complement. The raw read can be split into two "1D" reads ("template" and "complement") by removing the adapters. The consensus sequence of the two "1D" reads is the "2D" read, which has higher accuracy.

[0307] Five-letter and six-letter sequencing methods include whole genome sequencing methods capable of sequencing A, C, T and G in addition to 5mC and 5hmC to provide a digital readout of five letters (A, C, T, G, and either 5mC or 5hmC) or six letters (A, C, T, G, 5mC and 5hmC) in a single workflow. DNA sample processing is fully enzymatic, avoiding DNA degradation and the genome coverage bias of bisulfite treatment. In an exemplary five-letter sequencing method developed by Cambridge Epigenetix, the sample DNA is first fragmented via sonication and then ligated to short synthetic DNA hairpin adapters at both ends (Fuellgrabe, et al. 2022, bioRxiv doi: https: / / doi.org / 10.1101 / 2022.07.08.499285). The construct is then split to separate the sense and antisense sample strands. For each original sample strand, a complementary copy strand is synthesized by DNA polymerase extension of the 3' end to generate a hairpin construct with the original sample DNA strand connected to its complementary strand lacking epigenetic modification via a synthesis loop. A sequencing adaptor is then ligated to the end. The modified cytosine is enzymatically protected. The unprotected C is then deaminated to uracil, which is subsequently read as thymine. The deaminated construct is no longer fully complementary and has substantially reduced duplex stability, so the hairpin can be easily opened and amplified by PCR. The construct can be sequenced in a paired-end format, whereby read 1 (P1 primed) is the original strand and read 2 (P2 primed) is the copy strand. The read data is aligned pairwise, so that read 1 is aligned to its complementary read 2. Cognate residues from both reads are computationally determined to generate a single genetic or epigenetic letter.Cognate base pairings that differ from the five allowed are the result of imperfect fidelity at some stage(s), including sample preparation, amplification, or incorrect base calling during sequencing. These errors occur independently for the cognate base on each strand, so substitutions result in disallowed pairs. Disallowed pairs are masked (denoted as N) within the resolved read, and the read itself is retained, resulting in minimal information loss and high accuracy at the read level. The resolved read is aligned to the reference genome. Genetic variants and methylation counts are generated by read counting at the base level.

[0308] 5hmC has been shown to have value as a marker of biological states and diseases, including early cancer detection from cell-free DNA. In adapting 5-letter sequencing to 6-letter sequencing, 5mC is disambiguated from 5hmC without compromising genetic base calling within the same sample fragment. The first three steps of the workflow to generate adaptor-ligated sample fragments with synthetic copy strands are identical to 5-letter sequencing described above. Methylation at 5mC is enzymatically copied to C on the copy strand across CpG units, while 5hmC is enzymatically protected from such copying. Thus, unmodified C, 5mC and 5hmC in each of the original CpG units are identified by unique two-base combinations. Unmodified cytosine is then deaminated to uracil, which is subsequently read as thymine. DNA is subjected to PCR amplification and sequencing as previously described. The reads are pairwise aligned and sequenced using the two-base code, and since the three CpG units are separate sequencing environments of the two-base code, each of the unmodified C, 5mC, and 5hmC can be determined.

[0309] In some embodiments, the step of sequencing comprises targeted sequencing, in which one or more genomic regions of interest are sequenced. In some such embodiments, the genomic regions of interest comprise a region that represents one or more genes selected from Tables 1, 2, 3, 4 and / or 5. In some such embodiments, DNA sequences that do not include a region of interest are not sequenced. Some embodiments comprise non-targeted sequencing, for example, all genomic regions of the DNA in the processed sample or sub-sample are sequenced, or genomic regions are randomly selected for sequencing. In other embodiments, detecting the presence or absence of DNA sequences comprises sequencing DNA that is not enriched for the genomic regions of interest (non-targeted sequencing), for example, detectable sequences are obtained in a substantially unbiased manner.

[0310] In some embodiments, the sequencing step is performed on a library that includes a captured set of target regions, which may include any of the target region sets described herein.In some embodiments, the sequencing step is performed on a library that includes a sub-sample (e.g., a whole genome sub-sample) that has not undergone capture / enrichment.For example, target regions can be captured from a first sub-sample and a second sample, and then sequenced; or target regions can be captured from a first sub-sample, and after processing such as contacting and tagging steps, can be combined with a second sub-sample; or target regions can be captured from a second sub-sample, and after processing such as contacting and tagging steps, can be combined with a first sub-sample; or the first sub-sample and the second sub-sample can be processed together and combined without undergoing capture / enrichment.

[0311] The sequencing reaction may be performed on one or more forms of nucleic acid, at least one of which is known to contain a marker for cancer or other disease. The sequencing reaction may be performed on any nucleic acid fragment present in the sample. In some embodiments, the sequence coverage of the genome may be less than 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, 99.9% or 100%. In some embodiments, the sequence reaction may provide at least 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70% or 80% sequence coverage of the genome. Sequence coverage may be performed for at least 5, 10, 20, 70, 100, 200 or 500 different genes, or at most 5000, 2500, 1000, 500 or 100 different genes.

[0312] Simultaneous sequencing reaction can be carried out using multiplex sequencing.In some cases, cell-free nucleic acid can be sequenced in at least 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, 100,000 sequencing reactions.In other cases, cell-free nucleic acid can be sequenced in less than 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, 100,000 sequencing reactions.Sequencing reaction can be carried out sequentially or simultaneously.Subsequent data analysis can be carried out on all or part of sequencing reaction. In some cases, data analysis may be performed on at least 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, 100,000 sequencing reactions. In other cases, data analysis may be performed on less than 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, 100,000 sequencing reactions. Exemplary read depths are 1000-50000 reads per locus (base). a. Differential depth of sequencing

[0313] In some embodiments, the nucleic acids corresponding to the set of sequence variable target regions are sequenced to a greater depth of sequencing than the nucleic acids corresponding to the set of epigenetic target regions. In some embodiments, the nucleic acids corresponding to the set of hydroxymethylation variable target regions are sequenced to a greater depth of sequencing than the nucleic acids corresponding to at least one other set of target regions. For example, the depth of sequencing for the nucleic acids corresponding to the set of sequence variable and / or hydroxymethylation variable target regions is at least 1.25x, 1.5x, 1.75x, 2x, 2.25x, 2.5x, 2.75x, 3x, 3.5x, 4x, 4.5x, 5x, 6x, 7x, 8x, 9x, 10x, 11x, 12x, 13x, 14x, 16x, 18x, 19x, 20x, 21x, 22x, 23x, 24x, 25x, 26x, 27x, 28x, 29x, 30x, 31x, 32x, 33x, 34x, 35x, 36x, 37x, 38x, 39x, 40x, 41x, 42x, 43x, 44x, 45x, 46x, 47x, 48x, 49x, 50x, 51x, 52x, 53x, 54x, 55x, 56x, 57x, 58x, 59x, 60x, 61x, 62x, 63x, 64x, 65x, 66x, 67x, 68x, 69x, 70x, 71x, 72x, 73x, 74 1x or 15x deeper, or may be 1.25x to 1.5x, 1.5x to 1.75x, 1.75x to 2x, 2x to 2.25x, 2.25x to 2.5x, 2.5x to 2.75x, 2.75x to 3x, 3x to 3.5x, 3.5x to 4x, 4x to 4.5x, 4.5x to 5x, 5x to 5.5x, 5.5x to 6x, 6x to 7x, 7x to 8x, 8x to 9x, 9x to 10x, 10x to 11x, 11x to 12x, 13x to 14x, 14x to 15x, or 15x to 100x deeper. In some embodiments, the depth of sequencing is at least 2x deeper. In some embodiments, the depth of sequencing is at least 5x deeper. In some embodiments, the depth of sequencing is at least 10x deeper. In some embodiments, the depth of sequencing is between 4x and 10x deeper. In some embodiments, the depth of sequencing is between 4x and 100x deeper. Each of these embodiments refers to the extent to which nucleic acids corresponding to a set of sequence variable target regions are sequenced to a deeper sequencing depth than nucleic acids corresponding to a set of epigenetic target regions.

[0314] In some embodiments, the captured cfDNA corresponding to the set of sequence variable target regions and the captured cfDNA corresponding to the set of epigenetic target regions are sequenced in parallel, e.g., in the same sequencing cell (e.g., a flow cell of an Illumina sequencer), and / or in the same composition, which can be a pooled composition resulting from recombining separately captured sets, or a composition obtained by capturing cfDNA corresponding to the set of sequence variable target regions and the captured cfDNA corresponding to the set of epigenetic target regions in the same container.

[0315] In some embodiments, the captured cfDNA corresponding to the set of hydroxymethylated variable target regions and the captured cfDNA corresponding to at least one other set of target regions are sequenced in parallel, e.g., in the same sequencing cell (e.g., a flow cell of an Illumina sequencer), and / or in the same composition, which can be a pooled composition obtained from recombining the separately captured sets or a composition obtained by capturing cfDNA corresponding to the set of hydroxymethylated variable target regions and the captured cfDNA corresponding to at least one other set of target regions in the same vessel. 8.Analysis

[0316] In some embodiments, the methods described herein include identifying the presence of DNA produced by a tumor (or neoplastic or cancer cell).

[0317] The method of the present invention can be used to diagnose the presence of a condition, particularly cancer, in a subject, characterize the condition (e.g., stage cancer or determine the heterogeneity of cancer), monitor the response of the condition to treatment, and provide a prognostic risk of developing the condition or the subsequent course of the condition. The present disclosure can also be useful in determining the effectiveness of a particular treatment option. If treatment is successful, a successful treatment option can increase the amount of copy number variation or rare mutation detected in the blood of the subject, since more cancers may die and shed DNA if treatment is successful. In other examples, this may not occur. In another example, perhaps a particular treatment option can be correlated with the genetic profile of cancer over time. This correlation can be useful in selecting a treatment.

[0318] In some embodiments, the method of the present invention is used to screen for cancer or in a method for screening for cancer. For example, the sample may be a sample from a subject who has not been previously diagnosed with cancer. In some embodiments, the subject may or may not have cancer. In some embodiments, the subject may or may not have early stage cancer. In some embodiments, the subject has one or more risk factors for cancer, such as tobacco use (e.g., smoking), being overweight or obese, having a high body mass index (BMI), being elderly, nutritional deficiencies, high alcohol intake, or a family history of cancer.

[0319] In some embodiments, the subject has been using tobacco for at least 1, 5, 10 or 15 years, for example.In some embodiments, the subject has a high BMI, for example, a BMI of 25 or higher, 26 or higher, 27 or higher, 28 or higher, 29 or higher, or 30 or higher.In some embodiments, the subject is at least 40, 45, 50, 55, 60, 65, 70, 75 or 80 years old.In some embodiments, the subject has a nutritional deficiency, for example, a high intake of one or more of red meat and / or processed meat, trans fat, saturated fat and refined sugar, and / or a low intake of fruits and vegetables, complex carbohydrates and / or unsaturated fat. High intake and low intake can be defined as, for example, above or below the recommendations in the Dietary Guidelines for Americans 2020-2025 available at www.dietaryguidelines.gov / sites / default / files / 2021-03 / Dietary_Guidelines_for_Americans-2020-2025.pdf, respectively. In some embodiments, the subject has a high alcohol intake, e.g., averaging at least 3, 4, or 5 drinks per day (wherein a drink is about 1 ounce or 30 mL of 80 proof hard liquor or equivalent). In some embodiments, the subject has a family history of cancer, e.g., at least 1, 2, or 3 blood relatives have been previously diagnosed with cancer. In some embodiments, a relative is at least a third degree relative (e.g., a great-grandparent, a great aunt or uncle, a cousin), at least a second degree relative (e.g., a grandparent, an aunt or uncle, or a half-sibling), or a first degree relative (e.g., a parent or full sibling).

[0320] Furthermore, if a cancer is observed to be in remission following treatment, the methods of the invention can be used to monitor for residual disease, or recurrence of disease.

[0321] The type and number of cancers that can be detected can include blood cancer, brain cancer, lung cancer, skin cancer, nose cancer, throat cancer, liver cancer, bone cancer, lymphoma, pancreatic cancer, skin cancer, intestinal cancer, rectal cancer, thyroid cancer, bladder cancer, kidney cancer, oral cancer, stomach cancer, solid tumors, heterogeneous tumors, homogeneous tumors, etc. The type and / or stage of cancer can be detected from genetic variations including mutations, rare mutations, indels, copy number variations, transversions, translocations, inversions, deletions, aneuploidy, partial aneuploidy, polyploidy, chromosomal instability, chromosomal structural alterations, gene fusions, chromosomal fusions, gene truncations, gene amplifications, gene duplications, chromosomal damage, DNA damage, abnormal changes in chemical modifications of nucleic acids, abnormal changes in epigenetic patterns, and abnormal changes in 5-methylcytosines of nucleic acids.

[0322] Genetic data can also be used to characterize specific forms of cancer. Cancers are often heterogeneous in both composition and staging. Genetic profile data can allow characterization of specific subtypes of cancer, which can be important in the diagnosis or treatment of that specific subtype. This information can also provide clues to the subject or practitioner regarding the prognosis of a particular type of cancer, allowing either the subject or practitioner to adapt treatment options as the disease progresses. Some cancers can progress to become more aggressive and genetically unstable. Other cancers can remain benign, inactive or dormant. The disclosed system and method can be useful in determining disease progression.

[0323] Furthermore, the method of the present disclosure can be used to characterize the heterogeneity of abnormal conditions in a subject. Such a method can include, for example, generating a genetic profile of extracellular polynucleotides from a subject, the genetic profile including a plurality of data resulting from the analysis of copy number variations and rare mutations. In some embodiments, the abnormal condition is cancer. In some embodiments, the abnormal condition can be one that produces a heterogeneous genomic population. In the example of cancer, it is known that some tumors contain tumor cells that are at different stages of cancer. In other examples, the heterogeneity can include multiple foci of disease. Again, in the example of cancer, there can be multiple tumor foci, perhaps this time one or more foci are the result of metastasis that has spread from the primary site.

[0324] The methods of the present invention can be used to generate or profile a fingerprint or set of data that is the sum of genetic information from different cells in a heterogeneous disease. This set of data can include analysis of copy number variations, epigenetic variations, and mutations, either alone or in combination.

[0325] The method of the present invention can be used to diagnose, prognose, monitor or observe cancer or other diseases.In some embodiments, the method herein does not involve diagnosing, prognosing or monitoring fetus, and therefore does not concern non-invasive prenatal testing.In other embodiments, these methodologies can be used in pregnant subjects to diagnose, prognose, monitor or observe cancer or other diseases in subjects before birth, whose DNA and other polynucleotides may co-circulate with maternal molecules. 9. Target

[0326] In some embodiments, the DNA molecule (e.g., cfDNA molecule) (which may be referred to simply as DNA or cfDNA for brevity) is obtained from a subject having cancer. In some embodiments, the DNA (e.g., cfDNA) is obtained from a subject having cancer. In some embodiments, the DNA (e.g., cfDNA) is obtained from a subject suspected of having cancer. In some embodiments, the DNA (e.g., cfDNA) is obtained from a subject having a tumor. In some embodiments, the DNA (e.g., cfDNA) is obtained from a subject suspected of having a tumor. In some embodiments, the DNA (e.g., cfDNA) is obtained from a subject suspected of having a neoplasm. In some embodiments, the DNA (e.g., cfDNA) is obtained from a subject suspected of having a neoplasm. In some embodiments, the DNA (e.g., cfDNA) is obtained from a subject in remission from a tumor, cancer, or neoplasm (e.g., after chemotherapy, surgical resection, radiation, or a combination thereof). In any of the above-mentioned embodiments, the cancer, tumor, or neoplasm or suspected cancer, tumor, or neoplasm may be of the lung, colon, rectum, kidney, breast, prostate, or liver. In some embodiments, the cancer, tumor or neoplasm or suspected cancer, tumor or neoplasm is of the lung. In some embodiments, the cancer, tumor or neoplasm or suspected cancer, tumor or neoplasm is of the colon or rectum. In some embodiments, the cancer, tumor or neoplasm or suspected cancer, tumor or neoplasm is of the breast. In some embodiments, the cancer, tumor or neoplasm or suspected cancer, tumor or neoplasm is of the prostate. In any of the above embodiments, the subject may be a human subject. In some embodiments, the sample is obtained from a subject with stage I cancer, stage II cancer, stage III cancer, or stage IV cancer.

[0327] In some embodiments, the subject is a human, a mammal, an animal, a companion animal, a service animal, or a pet. The subject may have cancer, precancer, an infectious disease, transplant rejection, or other disease or disorder associated with changes in the immune system. The subject may not have cancer, detectable symptoms of cancer, or detectable symptoms of metastasis. The subject may have been treated with one or more cancer therapies, such as any one or more of chemotherapy, antibodies, vaccines, or biotherapeutics. The subject may be in remission. The subject may or may not be diagnosed as susceptible to cancer or any cancer-related genetic mutation / disorder. C. Additional Features of Certain Disclosed Methods 1. Sample

[0328] The sample can be any biological sample isolated from a subject. The sample can be a body sample. The sample can include body tissues, such as known or suspected solid tumors, whole blood, platelets, serum, plasma, feces, red blood cells, white blood cells or leucocytes, endothelial cells, tissue biopsies, cerebrospinal fluid, synovial fluid, lymphatic fluid, ascites, interstitial or extracellular fluid, fluid in the space between cells, including gingival crevicular fluid, bone marrow, pleural fluid, cerebrospinal fluid, saliva, mucus, sputum, semen, sweat, urine. The sample is preferably a body fluid, in particular blood and its fractions, as well as urine. The sample can be in the form originally isolated from the subject, or can have been subjected to further processing to remove or add components, such as cells, or to enrich one component relative to another. Thus, the preferred body fluid for analysis is plasma or serum, which contains cell-free nucleic acid. The sample can be isolated or obtained from the subject and transported to the location of sample analysis. The sample may be stored and shipped at a desired temperature, e.g., room temperature, 4°C, -20°C, and / or -80°C. The sample may be isolated or obtained from the subject at the site of sample analysis. The subject may be a human, a mammal, an animal, a companion animal, a service animal, or a pet. The subject may have cancer. The subject may not have cancer or detectable symptoms of cancer. The subject may have been treated with one or more cancer therapies, e.g., any one or more of chemotherapy, antibodies, vaccines, or biologies. The subject may be in remission. The subject may or may not have been diagnosed as susceptible to cancer or any cancer-related genetic mutation / disorder.

[0329] The volume of plasma may depend on the desired read depth for the region to be sequenced. Exemplary volumes are 0.4-40 ml, 5-20 ml, 10-20 ml. For example, volumes may be 0.5 mL, 1 mL, 5 mL 10 mL, 20 mL, 30 mL, or 40 mL. The volume of sampled plasma may be 5-20 mL.

[0330] A sample can contain various amounts of nucleic acid, including genome equivalents. For example, a sample of about 30 ng of DNA contains about 10,000 (10 4 ) haploid human genome equivalent, and for cfDNA, approximately 200 billion (2 × 10 11 ) individual polynucleotide molecules. Similarly, a sample of about 100 ng of DNA may contain about 30,000 haploid human genome equivalents, and in the case of cfDNA, about 600 billion individual molecules.

[0331] The sample may include nucleic acids from different sources, e.g., from cells and acellular of the same subject, from cells and acellular of different subjects. The sample may include nucleic acids carrying mutations. For example, the sample may include DNA carrying germline mutations and / or somatic mutations. Germline mutations refer to mutations present in the germline DNA of a subject. Somatic mutations refer to mutations originating from the somatic cells of a subject, e.g., cancer cells. The sample may include DNA carrying cancer-associated mutations (e.g., cancer-associated somatic mutations). The sample may include epigenetic variants (i.e., chemical or protein modifications), where the epigenetic variants are associated with the presence of genetic variants, e.g., cancer-associated mutations. In some embodiments, the sample includes epigenetic variants associated with the presence of a genetic variant, where the sample does not include the genetic variant.

[0332] Exemplary amounts of cell-free nucleic acid in a sample prior to amplification range from about 1 fg to about 1 μg, e.g., 1 pg to 200 ng, 1 ng to 100 ng, 10 ng to 1000 ng. For example, the amount can be up to about 600 ng, up to about 500 ng, up to about 400 ng, up to about 300 ng, up to about 200 ng, up to about 100 ng, up to about 50 ng, or up to about 20 ng of cell-free nucleic acid molecules. The amount can be at least 1 fg, at least 10 fg, at least 100 fg, at least 1 pg, at least 10 pg, at least 100 pg, at least 1 ng, at least 10 ng, at least 100 ng, at least 150 ng, or at least 200 ng of cell-free nucleic acid molecules. The amount can be up to 1 femtogram (fg), 10 fg, 100 fg, 1 picogram (pg), 10 pg, 100 pg, 1 ng, 10 ng, 100 ng, 150 ng or 200 ng of cell-free nucleic acid molecules. The method can include obtaining between 1 femtogram (fg) and 200 ng.

[0333] Cell-free DNA refers to DNA that is not contained within cells at the time of its isolation from a subject. For example, cfDNA can be isolated from a sample as DNA remaining in the sample after removing intact cells without lysing the cells or otherwise extracting intracellular DNA. Cell-free nucleic acids include DNA, RNA, and hybrids thereof, including genomic DNA, mitochondrial DNA, siRNA, miRNA, circular RNA (cRNA), tRNA, rRNA, small nucleolar RNA (snoRNA), Piwi-interacting RNA (piRNA), long non-coding RNA (long ncRNA), or fragments of any of these. Cell-free nucleic acids can be double-stranded, single-stranded, or hybrids thereof. Cell-free nucleic acids can be released into bodily fluids through secretion or cell death processes, such as necrosis and apoptosis of cells. Some cell-free nucleic acids, such as circulating tumor DNA (ctDNA), are released into bodily fluids from cancer cells. Others are released from healthy cells. In some embodiments, cfDNA is cell-free fetal DNA (cffDNA). In some embodiments, the cell-free nucleic acid is produced by tumor cells. In some embodiments, the cell-free nucleic acid is produced by a mixture of tumor cells and non-tumor cells.

[0334] Cell-free nucleic acids have an exemplary size distribution of about 100-500 nucleotides, with molecules of 110 to about 230 nucleotides accounting for about 90% of the molecules, the mode being about 168 nucleotides, and a second minor peak ranging between 240-440 nucleotides.

[0335] Cell-free nucleic acids can be isolated from bodily fluids through a fractionation or partitioning step, where the cell-free nucleic acids found in solution are separated from intact cells and other non-soluble components of the bodily fluid. Partitioning can include techniques such as centrifugation or filtration. Alternatively, cells in the bodily fluid can be lysed and the cell-free and cellular nucleic acids can be processed together. In general, after the addition of buffer and a washing step, the nucleic acid can be precipitated with alcohol. Further cleaning steps such as silica-based columns to remove contaminants or salts can be used. Non-specific bulk carrier nucleic acids for bisulfite sequencing, hybridization and / or ligation, such as C1 DNA, DNA or proteins, can be added throughout the reaction to optimize certain aspects of the procedure, such as yield.

[0336] After such processing, the sample may contain various forms of nucleic acid, including double-stranded DNA, single-stranded DNA and single-stranded RNA. In some embodiments, single-stranded DNA and RNA may be converted to double-stranded form so that they can be included in subsequent processing and analysis steps.

[0337] The double-stranded DNA molecules in the sample and the single-stranded nucleic acid molecules converted to double-stranded DNA molecules can be ligated to adapters at either one or both ends. Typically, the double-stranded molecules are blunt-ended by treatment with a polymerase that has a 5'-3' polymerase and a 3'-5' exonuclease (or proofreading function) in the presence of all four standard nucleotides. Klenow large fragment and T4 polymerase are examples of suitable polymerases. The blunt-ended DNA molecules can be ligated with at least partially double-stranded adapters (e.g., Y-shaped or bell-shaped adapters). Alternatively, complementary nucleotides can be added to the blunt ends of the sample nucleic acid and adapter to facilitate ligation. Both blunt-end ligation and sticky-end ligation are contemplated herein. In blunt-end ligation, both the nucleic acid molecule and the adapter tag have blunt ends. In sticky end ligation, typically the nucleic acid molecule carries an "A" overhang and the adaptor carries a "T" overhang. 2. Amplification

[0338] The sample nucleic acid flanked by adaptors can be amplified by PCR and other amplification methods. Amplification is typically primed by primers that bind to primer binding sites in the adaptors flanking the DNA molecule to be amplified. Amplification methods can include cycles of denaturation, annealing and extension resulting from thermocycling, or can be isothermal in transcription-mediated amplification. Other amplification methods include ligase chain reaction, strand displacement amplification, nucleic acid sequence-based amplification, and self-sustaining sequence-based replication.

[0339] In some embodiments, the methods of the invention perform dsDNA ligation with T-tailed and C-tailed adapters resulting in at least 50, 60, 70 or 80% amplification of double stranded nucleic acid prior to ligation to the adapters. Preferably, the methods of the invention increase the amount or number of amplified molecules by at least 10, 15 or 20% compared to a control method performed with T-tailed adapters alone. 3.Bait set; capture part

[0340] As discussed above, the nucleic acid in the sample may be subjected to a capture step, in which molecules with target sequences are captured for subsequent analysis. Target capture may include the use of a bait set that includes a capture moiety, e.g., an oligonucleotide bait labeled with biotin or other examples noted below. The probe may have a sequence selected to tile across a region, e.g., a panel of genes. In some embodiments, the bait set may have higher and lower capture yields for a set of target regions, e.g., a sequence variable target region set and an epigenetic target region set, respectively, as discussed elsewhere herein. Such a bait set is combined with the sample under conditions that allow hybridization of the target molecule with the bait. The captured molecules are then isolated using a capture moiety, e.g., a bead-based streptavidin biotin capture moiety. Such methods are further described, for example, in U.S. Patent No. 9,850,523, issued December 26, 2017, which is hereby incorporated by reference.

[0341] Capture moieties include, but are not limited to, biotin, avidin, streptavidin, nucleic acids containing specific nucleotide sequences, haptens recognized by antibodies, and magnetically attractable particles. Extraction moieties can be members of binding pairs, such as biotin / streptavidin or hapten / antibody. In some embodiments, the capture moiety bound to the analyte is captured by its binding partner bound to an isolable moiety, such as a magnetically attractable particle or a large particle that can be sedimented via centrifugation. Capture moieties can be any type of molecule that allows affinity separation of nucleic acids bearing the capture moiety from nucleic acids that lack the capture moiety. Exemplary capture moieties are biotin, which allows affinity separation by binding to streptavidin that is or can be linked to a solid phase, or oligonucleotides, which allow affinity separation via binding to complementary oligonucleotides that are or can be linked to a solid phase. D. Collection of target-specific probes

[0342] In some embodiments, a collection of target-specific probes is used in the methods described herein. In some embodiments, the group of target-specific probes comprises target-binding probes specific to a set of sequence-variable target regions. In some embodiments, the group of target-specific probes comprises target-binding probes specific to a set of epigenetic target regions. In some embodiments, the collection of target-specific probes comprises target-binding probes specific to a set of sequence-variable target regions and target-binding probes specific to a set of epigenetic target regions. The target-specific probes may comprise sequences complementary to non-deaminated target regions, sequences complementary to non-deaminated target regions, sequences complementary to converted target regions, or sequences complementary to non-converted target regions, as discussed in detail elsewhere herein, for example, in the above discussion of the enrichment / capture step.

[0343] In some embodiments, the capture yield of a target binding probe specific to a set of sequence variable target regions is higher (e.g., at least 2-fold higher) than the capture yield of a target binding probe specific to a set of epigenetic target regions. In some embodiments, a collection of target-specific probes is configured to have a capture yield specific to a set of sequence variable target regions that is higher (e.g., at least 2-fold higher) than its capture yield specific to a set of epigenetic target regions. When a target-specific probe specific to a set of hydroxymethylation variable target regions is used, in some embodiments, such a target-specific probe has a capture yield higher (e.g., at least 2-fold higher) than the capture yield of a target binding probe specific to at least one other set of target regions. In some embodiments, a collection of target-specific probes is configured to have a capture yield specific to a set of hydroxymethylation variable target regions that is higher (e.g., at least 2-fold higher) than its capture yield specific to at least one other set of target regions.

[0344] In some embodiments, the capture yield of target binding probes specific for the set of sequence variable target regions is at least 1.25x, 1.5x, 1.75x, 2x, 2.25x, 2.5x, 2.75x, 3x, 3.5x, 4x, 4.5x, 5x, 6x, 7x, 8x, 9x, 10x, 11x, 12x, 13x, 14x or 15x higher than the capture yield of target binding probes specific for the set of epigenetic target regions. In some embodiments, the capture yield of target binding probes specific for the set of sequence variable target regions is 1.25x to 1.5x, 1.5x to 1.75x, 1.75x to 2x, 2x to 2.25x, 2.25x to 2.5x, 2.5x to 2.75x, 2.75x to 3x, 3x to 3.5x, 3.5x to 4x, 4x to 4.5x, 4.5x to 5x, 5x to 5.5x, 5.5x to 6x, 6x to 7x, 7x to 8x, 8x to 9x, 9x to 10x, 10x to 11x, 11x to 12x, 13x to 14x, or 14x to 15x higher than the capture yield of target binding probes specific for the set of epigenetic target regions.

[0345] In some embodiments, the collection of target-specific probes is configured to have a capture yield specific for a set of sequence variable target regions that is at least 1.25x, 1.5x, 1.75x, 2x, 2.25x, 2.5x, 2.75x, 3x, 3.5x, 4x, 4.5x, 5x, 6x, 7x, 8x, 9x, 10x, 11x, 12x, 13x, 14x or 15x higher than its capture yield for the set of epigenetic target regions. In some embodiments, the collection of target-specific probes is configured to have a capture yield specific for a set of sequence variable target regions that is 1.25x-1.5x, 1.5x-1.75x, 1.75x-2x, 2x-2.25x, 2.25x-2.5x, 2.5x-2.75x, 2.75x-3x, 3x-3.5x, 3.5x-4x, 4x-4.5x, 4.5x-5x, 5x-5.5x, 5.5x-6x, 6x-7x, 7x-8x, 8x-9x, 9x-10x, 10x-11x, 11x-12x, 13x-14x, or 14x-15x higher than its capture yield specific for the set of epigenetic target regions.

[0346] A collection of probes can be configured to provide higher capture yields for a set of sequence-variable target regions in a variety of ways, including concentration, different lengths and / or chemistries (e.g., to affect affinity), and combinations thereof. Affinity can be modulated by adjusting probe length and / or including nucleotide modifications, as discussed below.

[0347] In some embodiments, the target-specific probes specific to the set of sequence variable target regions are present at a higher concentration than the target-specific probes specific to the set of epigenetic target regions. In some embodiments, the concentration of the target-binding probes specific to the set of sequence variable target regions is at least 1.25 times, 1.5 times, 1.75 times, 2 times, 2.25 times, 2.5 times, 2.75 times, 3 times, 3.5 times, 4 times, 4.5 times, 5 times, 6 times, 7 times, 8 times, 9 times, 10 times, 11 times, 12 times, 13 times, 14 times, or 15 times higher than the concentration of the target-binding probes specific to the set of epigenetic target regions. In some embodiments, the concentration of target binding probes specific to the set of sequence variable target regions is 1.25-1.5x, 1.5-1.75x, 1.75-2x, 2-2.25x, 2.25-2.5x, 2.5-2.75x, 2.75-3x, 3-3.5x, 3.5-4x, 4-4.5x, 4.5-5x, 5-5.5x, 5.5-6x, 6-7x, 7-8x, 8-9x, 9-10x, 10-11x, 11-12x, 13-14x, or 14-15x higher than the concentration of target binding probes specific to the set of epigenetic target regions. In such embodiments, concentration may refer to the average mass concentration per volume of the individual probes in each set.

[0348] In some embodiments, the target-specific probes specific to the sequence variable target region set have a higher affinity for their targets than the target-specific probes specific to the epigenetic target region set. In some embodiments, the target-specific probes specific to the sequence variable target region set have a modification that increases their affinity for their targets. In some embodiments, alternatively or additionally, the target-specific probes specific to the epigenetic target region set have a modification that decreases their affinity for their targets. In some embodiments, the target-specific probes specific to the sequence variable target region set have a longer average length and / or a higher average melting temperature than the target-specific probes specific to the epigenetic target region set. These embodiments can be combined with each other and / or with the differences in concentration discussed above to achieve a desired fold difference in capture yield, such as any of the fold differences or ranges described above.

[0349] In some embodiments, the capture yield of target binding probes specific for a set of hydroxymethylated variable target regions is at least 1.25x, 1.5x, 1.75x, 2x, 2.25x, 2.5x, 2.75x, 3x, 3.5x, 4x, 4.5x, 5x, 6x, 7x, 8x, 9x, 10x, 11x, 12x, 13x, 14x or 15x higher than the capture yield of target binding probes specific for at least one other set of target regions. In some embodiments, the capture yield of target binding probes specific for the set of hydroxymethylated variable target regions is 1.25-fold to 1.5-fold, 1.5-fold to 1.75-fold, 1.75-fold to 2-fold, 2-fold to 2.25-fold, 2.25-fold to 2.5-fold, 2.5-fold to 2.75-fold, 2.75-fold to 3-fold, 3-fold to 3.5-fold, 3.5-fold to 4-fold, 4-fold to 4.5-fold, 4.5-fold to 5-fold, 5-fold to 5.5-fold, 5.5-fold to 6-fold, 6-fold to 7-fold, 7-fold to 8-fold, 8-fold to 9-fold, 9-fold to 10-fold, 10-fold to 11-fold, 11-fold to 12-fold, 13-fold to 14-fold, or 14-fold to 15-fold higher than the capture yield of target binding probes specific for at least one other set of target regions.

[0350] In some embodiments, a panel of target-specific probes is configured to have a capture yield specific for a set of hydroxymethylated variable target regions that is at least 1.25x, 1.5x, 1.75x, 2x, 2.25x, 2.5x, 2.75x, 3x, 3.5x, 4x, 4.5x, 5x, 6x, 7x, 8x, 9x, 10x, 11x, 12x, 13x, 14x or 15x higher than its capture yield for at least one other set of target regions. In some embodiments, a panel of target-specific probes is configured to have a capture yield specific for a set of hydroxymethylated variable target regions that is 1.25x to 1.5x, 1.5x to 1.75x, 1.75x to 2x, 2x to 2.25x, 2.25x to 2.5x, 2.5x to 2.75x, 2.75x to 3x, 3x to 3.5x, 3.5x to 4x, 4x to 4.5x, 4.5x to 5x, 5x to 5.5x, 5.5x to 6x, 6x to 7x, 7x to 8x, 8x to 9x, 9x to 10x, 10x to 11x, 11x to 12x, 13x to 14x, or 14x to 15x higher than its capture yield specific for at least one other set of target regions.

[0351] A panel of probes can be configured to provide higher capture yields for a set of hydroxymethylated variable target regions in a variety of ways, including concentration, different lengths and / or chemistries (e.g., to affect affinity), and combinations thereof. Affinity can be modulated by adjusting probe length and / or including nucleotide modifications, as discussed below.

[0352] In some embodiments, the target specific probes specific for the set of hydroxymethylated variable target regions are present at a higher concentration than the target specific probes specific for at least one other set of target regions. In some embodiments, the concentration of the target binding probes specific for the set of hydroxymethylated variable target regions is at least 1.25x, 1.5x, 1.75x, 2x, 2.25x, 2.5x, 2.75x, 3x, 3.5x, 4x, 4.5x, 5x, 6x, 7x, 8x, 9x, 10x, 11x, 12x, 13x, 14x or 15x higher than the concentration of the target binding probes specific for at least one other set of target regions. In some embodiments, the concentration of target binding probes specific to a set of hydroxymethylated variable target regions is 1.25-1.5x, 1.5-1.75x, 1.75-2x, 2-2.25-2.5-2.5-2.5-2.75-3x, 3-3.5-4x, 4-4.5-5x, 5-5.5-5x, 5.5-6x, 6-7x, 7-8x, 8-9x, 9-10x, 10-11x, 11-12x, 13-14x, or 14-15x higher than the concentration of target binding probes specific to at least one other set of target regions. In such embodiments, concentration may refer to the average mass per volume concentration of the individual probes in each set.

[0353] In some embodiments, the target-specific probes specific to the hydroxymethylated variable target region set have a higher affinity for their targets than the target-specific probes specific to at least one other target region set. In some embodiments, the target-specific probes specific to the hydroxymethylated variable target region set have a modification that increases their affinity for their targets. In some embodiments, alternatively or in addition, the target-specific probes specific to at least one other target region set have a modification that decreases their affinity for their targets. In some embodiments, the target-specific probes specific to the hydroxymethylated variable target region set have a longer average length and / or a higher average melting temperature than the target-specific probes specific to at least one other target region set. These embodiments can be combined with each other and / or with the differences in concentration discussed above to achieve a desired fold difference in capture yield, such as any of the fold differences or ranges described above.

[0354] Affinity can be modulated in any way known to those skilled in the art, including by using different probe chemistries.For example, certain nucleotide modifications, such as cytosine 5-methylation (in the context of certain sequences), modifications that provide heteroatoms at the 2' sugar position, and LNA nucleotides, can increase the stability of double-stranded nucleic acids, indicating that oligonucleotides with such modifications have relatively high affinity to their complementary sequences.See, for example, Severin et al., Nucleic Acids Res. 39: 8740-8751 (2011);Freier et al., Nucleic Acids Res. 25: 4429-4443 (1997);U.S. Patent No. 9,738,894. Also, longer sequence lengths generally provide increased affinity.Other nucleotide modifications, such as the substitution of guanine with the nucleobase hypoxanthine, reduce affinity by reducing the amount of hydrogen bonds between the oligonucleotide and its complementary sequence.

[0355] In some embodiments, the target-specific probe comprises a capture moiety. The capture moiety can be any of the capture moieties described herein, e.g., biotin. In some embodiments, the target-specific probe is linked, e.g., covalently or non-covalently, to a solid support, e.g., via the interaction of the binding pair of the capture moiety. In some embodiments, the solid support is a bead, e.g., a magnetic bead.

[0356] In some embodiments, the target specific probes specific to a set of sequence variable target regions and / or the target specific probes specific to a set of epigenetic target regions are probes that include a bait set as discussed above, e.g., a capture moiety and a region, e.g., a sequence selected to tile across a panel of genes.

[0357] In some embodiments, the target-specific probes are provided in a single composition. The single composition can be a solution (liquid or frozen). Alternatively, the single composition can be a lyophilizate.

[0358] Alternatively, target-specific probes can be provided as multiple compositions, including, for example, a first composition comprising a probe specific to a set of epigenetic target regions and a second composition comprising a probe specific to a set of sequence-variable target regions.These probes can be mixed in a suitable ratio to provide a combined probe composition with any of the above-mentioned fold differences in concentration and / or capture yield.Alternatively, they can be used in separate capture procedures (e.g., with an aliquot of sample, or sequentially with the same sample) to provide a first and second composition comprising captured epigenetic target regions and sequence-variable target regions, respectively. 1. Probes specific to epigenetic target regions

[0359] The probes for the epigenetic target region set may include probes specific to one or more types of target regions that are likely to differentiate DNA from neoplastic (e.g., tumor or cancer) cells from healthy cells, e.g., non-neoplastic circulating cells. Exemplary types of such regions are discussed in detail herein, for example, in the section above on captured sets. The probes for the epigenetic target region set may also include probes for one or more control regions, for example, as described herein.

[0360] In some embodiments, the probes for the set of epigenetic target regions have a footprint of at least 100 kbp, e.g., at least 200 kbp, at least 300 kbp, or at least 400 kbp. In some embodiments, the set of epigenetic target regions have a footprint in the range of 100-20 Mbp, e.g., 100-200 kbp, 200-300 kbp, 300-400 kbp, 400-500 kbp, 500-600 kbp, 600-700 kbp, 700-800 kbp, 800-900 kbp, 900-1,000 kbp, 1-1.5 Mbp, 1.5-2 Mbp, 2-3 Mbp, 3-4 Mbp, 4-5 Mbp, 5-6 Mbp, 6-7 Mbp, 7-8 Mbp, 8-9 Mbp, 9-10 Mbp, or 10-20 Mbp. In some embodiments, the set of epigenetic target regions has a footprint of at least 20 Mbp. A. Hypermethylated variable target region

[0361] In some embodiments, the probes for the epigenetic target region set include probes specific for one or more hypermethylated variable target regions. The hypermethylated variable target region may also be referred to herein as a hypermethylated DMR (differentially methylated region). The hypermethylated variable target region may be any of those listed above. For example, in some embodiments, the probes specific for the hypermethylated variable target region include probes specific for a plurality of loci listed in Table 1, e.g., at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90% or 100% of the loci listed in Table 1. In some embodiments, the probes specific for the hypermethylated variable target region include probes specific for a plurality of loci listed in Table 2, e.g., at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90% or 100% of the loci listed in Table 2. In some embodiments, the probes specific for the hypermethylated variable target region include probes specific for multiple loci listed in Table 1 or Table 2, e.g., at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90% or 100% of the loci listed in Table 1 or Table 2. In some embodiments, for each locus included as a target region, there may be one or more probes with hybridization sites that bind between the transcription start site of the gene and the stop codon (for alternatively spliced ​​genes, the final stop codon). In some embodiments, the one or more probes bind within 300 bp, e.g., within 200 or 100 bp, of the listed position. In some embodiments, the probes have hybridization sites that overlap with the positions listed above. In some embodiments, the probes specific for hypermethylated target regions include probes specific for a subset of 1, 2, 3, 4 or 5 hypermethylated target regions that collectively exhibit hypermethylation in 1, 2, 3, 4 or 5 of breast, colon, kidney, liver and lung cancer. b. Hypomethylated variable target regions

[0362] In some embodiments, the probe for epigenetic target region set comprises one or more hypomethylated variable target region specific probes. Hypomethylated variable target region may also be referred to herein as hypomethylated DMR (differentially methylated region). Hypomethylated variable target region may be any of those shown above. For example, one or more hypomethylated variable target region specific probes may include probes for regions that may show reduced methylation in tumor cells, such as repetitive elements, such as LINE1 elements, Alu elements, centromeric tandem repeats, pericentromeric tandem repeats and satellite DNA, and intergenic regions that are normally methylated in healthy cells.

[0363] In some embodiments, the probes specific for the hypomethylated variable target region include probes specific for repetitive elements and / or intergenic regions, hi some embodiments, the probes specific for repetitive elements include probes specific for one, two, three, four or five of the following: LINE1 elements, Alu elements, centromeric tandem repeats, pericentromeric tandem repeats and / or satellite DNA.

[0364] Exemplary probes specific for genomic regions exhibiting cancer-associated hypomethylation include probes specific for nucleotides 8403565-8953708 and / or 151104701-151106035 of human chromosome 1. In some embodiments, probes specific for hypomethylated variable target regions include probes specific for regions overlapping with or including nucleotides 8403565-8953708 and / or 151104701-151106035 of human chromosome 1. c.CTCF binding region

[0365] In some embodiments, the probes for the set of epigenetic target regions include probes specific for CTCF binding regions. In some embodiments, the probes specific for CTCF binding regions are specific for at least 10, 20, 50, 100, 200, or 500 CTCF binding regions, or 10-20, 20-50, 50-100, 100-200, 200-500, or 500-1000 CTCF binding regions, such as those described in the CTCFBSDB or one or more of the articles by Cuddapah et al., Martin et al., or Rhee et al., cited above or cited above. In some embodiments, the probes for the set of epigenetic target regions include at least 100 bp, at least 200 bp at least 300 bp, at least 400 bp, at least 500 bp, at least 750 bp, or at least 1000 bp upstream and downstream of the CTCF binding site. D transcription start site

[0366] In some embodiments, the probes for the set of epigenetic target regions include probes specific to transcription start sites. In some embodiments, the probes specific to transcription start sites include probes specific to at least 10, 20, 50, 100, 200, or 500 transcription start sites, or 10-20, 20-50, 50-100, 100-200, 200-500, or 500-1000 transcription start sites, such as those listed in the DBTSS. In some embodiments, the probes for the set of epigenetic target regions include probes for sequences at least 100 bp, at least 200 bp, at least 300 bp, at least 400 bp, at least 500 bp, at least 750 bp, or at least 1000 bp upstream and downstream of the transcription start site. e. focal amplification

[0367] As noted above, focal amplifications are somatic mutations, but they can be detected by sequencing based on read frequency in a manner similar to the approach for detecting certain epigenetic changes, such as changes in methylation.Thus, the region that may show focal amplification in cancer can be included in the epigenetic target region set, as discussed above.In some embodiments, the probe specific to the epigenetic target region set includes a probe specific to focal amplification.In some embodiments, the probe specific to focal amplification includes a probe specific to one or more of AR, BRAF, CCND1, CCND2, CCNE1, CDK4, CDK6, EGFR, ERBB2, FGFR1, FGFR2, KIT, KRAS, MET, MYC, PDGFRA, PIK3CA and RAF1. For example, in some embodiments, focal amplification specific probes include probes specific for one or more of at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, or 18 of the above-mentioned targets. f. Control area

[0368] To facilitate data validation, it may be useful to include control regions. In some embodiments, the probes specific to the set of epigenetic target regions include probes specific to methylated control regions that are expected to be methylated in essentially all samples. In some embodiments, the probes specific to the set of epigenetic target regions include probes specific to hypomethylated control regions that are expected to be hypomethylated in essentially all samples. 2. Probes specific to sequence-variable target regions

[0369] The probes for the sequence variable target region set may include probes specific to multiple regions known to undergo somatic mutation in cancer. The probes may be specific to any sequence variable target region set described herein. Exemplary sequence variable target region sets are discussed in detail herein, for example, in the section above regarding captured sets.

[0370] In some embodiments, the sequence variable target region probe set has a footprint of at least 0.5 kb, e.g., at least 1 kb, at least 2 kb, at least 5 kb, at least 10 kb, at least 20 kb, at least 30 kb, or at least 40 kb. In some embodiments, the epigenetic target region probe set has a footprint in the range of 0.5-100 kb, e.g., 0.5-2 kb, 2-10 kb, 10-20 kb, 20-30 kb, 30-40 kb, 40-50 kb, 50-60 kb, 60-70 kb, 70-80 kb, 80-90 kb, and 90-100 kb. In some embodiments, the sequence variable target region probe set has a footprint of at least 50 kbp, e.g., at least 100 kbp, at least 200 kbp, at least 300 kbp, or at least 400 kbp. In some embodiments, the sequence variable target region probe set has a footprint in the range of 100-2000 kbp, e.g., 100-200 kbp, 200-300 kbp, 300-400 kbp, 400-500 kbp, 500-600 kbp, 600-700 kbp, 700-800 kbp, 800-900 kbp, 900-1,000 kbp, 1-1.5 Mbp, or 1.5-2 Mbp. In some embodiments, the sequence variable target region set has a footprint of at least 2 Mbp.

[0371] In some embodiments, the probes specific for the set of sequence variable target regions include probes specific for at least a portion of at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, or 70 of the genes in Table 3. In some embodiments, the probes specific for the set of sequence variable target regions include probes specific for at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, or 70 of the SNVs in Table 3. In some embodiments, the probes specific for the set of sequence variable target regions include probes specific for at least 1, at least 2, at least 3, at least 4, at least 5, or 6 of the fusions in Table 3. In some embodiments, the probes specific for the set of sequence variable target regions include probes specific for at least a portion of at least 1, at least 2, or 3 of the indels in Table 3. In some embodiments, the probes specific for the set of sequence variable target regions include probes specific for at least a portion of at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70 or 73 of the genes in Table 4. In some embodiments, the probes specific for the set of sequence variable target regions include probes specific for at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70 or 73 of the SNVs in Table 4. In some embodiments, the probes specific for the set of sequence variable target regions include probes specific for at least 1, at least 2, at least 3, at least 4, at least 5 or 6 of the fusions in Table 4.In some embodiments, the probes specific for the set of sequence variable target regions include probes specific for at least a portion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, or 18 of the indels in Table 4. In some embodiments, the probes specific for the set of sequence variable target regions include probes specific for at least a portion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, or at least 20 of the genes in Table 5.

[0372] In some embodiments, the probes specific for the set of sequence variable target regions include probes specific for target regions from at least 10, 20, 30, or 35 cancer-associated genes, such as AKT1, ALK, BRAF, CCND1, CDK2A, CTNNB1, EGFR, ERBB2, ESR1, FGFR1, FGFR2, FGFR3, FOXL2, GATA3, GNA11, GNAQ, GNAS, HRAS, IDH1, IDH2, KIT, KRAS, MED12, MET, MYC, NFE2L2, NRAS, PDGFRA, PIK3CA, PPP2R1A, PTEN, RET, STK11, TP53, and U2AF1. E. Computer Systems

[0373] The method of the present disclosure can be implemented using or with the aid of a computer system.For example, such a method can include: synthesizing a first complementary strand that is complementary to the first strand and a second complementary strand that is complementary to the second strand; b) optionally before or after synthesizing the first and second complementary strands, glucosylating the 5-hydroxymethylated cytosine in at least one of the first or second strands; c) optionally methylating the cytosine in at least one of the first or second complementary strands, wherein the methylation converts hemimethylated CpG to fully methylated CpG; d) optionally deaminating the unmodified cytosine in at least one of the first or second strands, thereby producing a processed DNA molecule; e) optionally sequencing at least a portion of the processed DNA molecule; the DNA molecule comprises a first and second strand and an asymmetric adaptor, and optionally the asymmetric adaptor is a Y-shaped adaptor or a bubble adaptor.

[0374] 2 shows a computer system 201 programmed or otherwise configured to carry out the methods of the present disclosure. The computer system 201 can coordinate various aspects of sample preparation, sequencing and / or analysis. In some examples, the computer system 201 is configured to perform sample preparation and sample analysis, including nucleic acid sequencing.

[0375] The computer system 201 includes a central processing unit (CPU, also referred to herein as "processor" and "computer processor") 205, which may be a single-core or multi-core processor, or multiple processors for parallel processing. The computer system 201 also includes memory or memory locations 210 (e.g., random access memory, read-only memory, flash memory), electronic storage 215 (e.g., hard disk), communication interface 220 (e.g., network adapter) for communicating with one or more other systems, and peripheral devices 225, such as cache, other memory, data storage, and / or electronic display adapters. The memory 210, storage 215, interface 220, and peripheral devices 225 are in communication with the CPU 205 via a communication network or bus (solid lines), e.g., a motherboard. The storage 215 may be a data storage device (or data repository) for storing data. The computer system 201 may be operatively coupled to a computer network 230 with the aid of the communication interface 220. The computer network 230 may be the Internet, an Internet and / or an extranet, or an intranet and / or an extranet in communication with the Internet. The computer network 230 may in some cases be a remote communication and / or data network. The computer network 230 may include one or more computer servers, which may enable distributed computing, e.g., cloud computing. The computer network 230 may in some cases, with the aid of the computer system 201, implement a peer-to-peer network, which may enable devices coupled to the computer system 201 to behave as clients or servers.

[0376] CPU 205 may execute sequences of machine-readable instructions, which may be embodied in a program or software. These instructions may be stored in memory locations, such as memory 210. Examples of operations performed by CPU 205 may include fetch, decode, execute, and writeback.

[0377] The storage device 215 can store files, such as drivers, libraries, and saved programs. The storage device 215 can store user generated programs and recorded sessions, as well as program-related output. The storage device 215 can store user data, such as user preferences and user programs. The computer system 201 can include one or more additional data storage devices, in some cases located external to the computer system 201, e.g., on a remote server in communication with the computer system 201 via an intranet or the Internet. Data can be transferred from one location to another, for example, using a communications network or physical data transfer (e.g., using a hard drive, thumb drive, or other data storage mechanism).

[0378] Computer system 201 can communicate with one or more remote computer systems via network 230. For an embodiment, computer system 201 can communicate with a remote computer system of a user (e.g., an operator). Examples of remote computer systems include a personal computer (e.g., a portable PC), a slate or tablet PC (e.g., Apple® iPad®, Samsung® Galaxy Tab), a phone, a smartphone (e.g., Apple® iPhone®, Android® enabled device, Blackberry®), or a personal digital assistant. A user can access computer system 201 via network 230.

[0379] The methods described herein may be performed by machine (e.g., a computer processor) executable code stored on an electronic storage location of the computer system 201, such as, for example, on memory 210 or electronic storage 215. The machine executable or machine readable code may be provided in the form of software. During use, the code may be executed by the processor 205. In some cases, the code may be retrieved from storage 215 and stored on memory 210 for ready access by the processor 205. In some situations, the electronic storage 215 may be omitted and the machine executable instructions are stored on memory 210.

[0380] In an aspect, the disclosure, when executed by at least one electronic processor, comprises the steps of: a) synthesizing a first complementary strand complementary to the first strand and a second complementary strand complementary to the second strand; b) optionally, before or after the steps of synthesizing the first and second complementary strands, glucosylating 5-hydroxymethylated cytosines in at least one of the first or second strands; c) methylating cytosines in at least one of the first or second complementary strands, wherein the methylation converts hemi-methylated CpGs to fully methylated CpGs. a) converting unmodified cytosines in at least one of the first or second strands to methylated CpGs, thereby producing a treated DNA molecule; b) sequencing at least a portion of the treated DNA molecule; and c) sequencing at least a portion of the treated DNA molecule; wherein the DNA molecule comprises a first and a second strand and an asymmetric adaptor, and optionally the asymmetric adaptor is a Y-type adaptor or a bubble adaptor. In an aspect, the disclosure, when executed by at least one electronic processor, comprises the steps of: synthesizing a first complementary strand complementary to the first strand and a second complementary strand complementary to the second strand; b) optionally, before or after the steps of synthesizing the first and second complementary strands, glucosylating 5-hydroxymethylated cytosines in at least one of the first or second strands; c) methylating cytosines in at least one of the first or second complementary strands, wherein methylation converts hemimethylated CpGs to fully methylated CpGs. a) converting unmodified cytosines in at least one of the first or second strands to methylated CpGs, thereby producing a treated DNA molecule; b) sequencing at least a portion of the treated DNA molecule; and c) sequencing at least a portion of the treated DNA molecule; wherein the DNA molecule comprises a first and a second strand and an asymmetric adaptor, and optionally the asymmetric adaptor is a Y-type adaptor or a bubble adaptor.In an aspect, the disclosure, when executed by at least one electronic processor, comprises the steps of: synthesizing a first complementary strand complementary to the first strand and a second complementary strand complementary to the second strand; b) optionally, before or after the steps of synthesizing the first and second complementary strands, glucosylating 5-hydroxymethylated cytosines in at least one of the first or second strands; c) methylating cytosines in at least one of the first or second complementary strands, wherein methylation converts hemimethylated CpGs to fully methylated CpGs. A non-transitory computer readable medium is provided, comprising computer executable instructions for performing at least a portion of the method, comprising: converting unmodified cytosines in at least one of the first or second strands to methylated CpG; d) deaminating unmodified cytosines in at least one of the first or second strands, thereby producing a treated DNA molecule; e) sequencing at least a portion of the treated DNA molecule; the DNA molecule comprises a first and second strand and an asymmetric adaptor, and optionally the asymmetric adaptor is a Y-shaped adaptor or a bubble adaptor. In some embodiments, the method further comprises obtaining a plurality of sequence reads generated by a nucleic acid sequencer from the sequencing; mapping the plurality of sequence reads to one or more reference sequences to produce mapped sequence reads; and processing the mapped sequence reads to determine the likelihood that the subject has cancer.

[0381] The code may be pre-compiled and configured for use with a machine having a processor adapted to execute the code, or may be compiled during run-time. The code may be provided in a programming language that may be selected to enable the code to be executed in a pre-compiled or as-compiled fashion.

[0382] An embodiment of the system and method provided herein, e.g., computer system 201, may be embodied with programming. The technology of various embodiments may be considered as a "product" or "article of manufacture," typically in the form of machine (or processor) executable code and / or associated data carried on or embodied in some type of machine-readable medium. The machine-executable code may be stored on electronic storage, e.g., memory (e.g., read-only memory, random access memory, flash memory) or hard disk. A "storage" type medium may include any or all of the tangible memory of a computer, processor, etc., or its associated modules, e.g., various semiconductor memories, tape drives, disk drives, etc., that may provide non-transitory storage for software programming at any time.

[0383] All or parts of the software may be communicated from time to time via the Internet or various other telecommunications networks. Such communication may, for example, allow loading of the software from one computer or processor into another, for example, from a management server or host computer into the computer platform of an application server. Thus, other types of media that may bear software elements include light waves, radio waves, and electromagnetic waves, such as those used across physical interfaces between local devices, through wired and optical landline networks, and through various air-links. Physical elements that carry such waves, such as wired or wireless links, optical links, and the like, may also be considered media bearing the software. As used herein, unless limited to non-transitory tangible "storage" media, terms such as computer or machine "readable medium" refer to any medium that participates in providing instructions to a processor for execution.

[0384] Thus, machine-readable media, such as computer executable code, may take many forms, including but not limited to tangible storage media, carrier wave media, or physical transmission media. Non-volatile storage media include, for example, optical or magnetic disks, such as any of the storage devices in any computer, such as those used to execute databases, such as those shown in the drawings. Volatile storage media include dynamic memory, such as the main memory of such a computer platform. Tangible transmission media include coaxial cables; copper wires and optical fibers, including the wires that comprise a bus in a computer system. Carrier wave transmission media may take the form of electric or electromagnetic signals, or sound or light waves, such as those generated during radio frequency (RF) and infrared (IR) data communications. Thus, common forms of computer readable media include, for example, a floppy disk, a flexible disk, a hard disk, a magnetic tape, any other magnetic medium, a CD-ROM, a DVD or DVD-ROM, any other optical medium, punch cards, paper tape, any other physical storage medium having a pattern of holes, RAM, ROM, PROM and EPROM, FLASH-EPROM, any other memory chip or cartridge, a carrier wave carrying data or instructions, a cable or link carrying such a carrier wave, or any other medium from which a computer may read programming code and / or data. Many of these forms of computer readable media may be involved in carrying one or more sequences of one or more instructions to a processor for execution.

[0385] The computer system 201 may include, or be in communication with, an electronic display 235 that includes, for example, a user interface (UI) 240 for providing one or more results of a sample analysis. Examples of UIs include, but are not limited to, graphical user interfaces (GUIs) and web-based user interfaces.

[0386] For additional details regarding computer systems and networks, databases, and computer program products, see, for example, Peterson, Computer Networks: A Systems Approach, Morgan Kaufmann, 5th Ed. (2011); Kurose, Computer Networking: A Top-Down Approach, Pearson, 7 th Ed. (2016), Elmasri, Fundamentals of Database Systems, Addison Wesley, 6th Ed. (2010), Coronel, Database Systems: Design, Implementation, & Management, Cengage Learning, 11 th Ed. (2014), Tucker, Programming Languages, McGraw-Hill Science / Engineering / Math, 2nd Ed. (2006), and Rhoton, Cloud Computing Architected: Solution Design Handbook, Recursive Press (2011), each of which is hereby incorporated by reference in its entirety. F. Application 1. Cancer and other diseases

[0387] The method of the present invention can be used to diagnose the presence of a condition, particularly cancer, in a subject, characterize a condition (e.g., stage cancer or determine cancer heterogeneity), monitor the response of a condition to treatment, and provide a prognostic risk of developing a condition or subsequent course of a condition. The present disclosure can also be useful in determining the effectiveness of a particular treatment option. If treatment is successful, a successful treatment option may increase the amount of copy number variation or rare mutation detected in the subject's blood, since more cancers may die and shed DNA if treatment is successful. In other examples, this may not occur. In another example, perhaps a particular treatment option may be correlated with the genetic profile of cancer over time. This correlation may be useful in selecting a treatment. In some embodiments, hypermethylated variable epigenetic target regions are analyzed to determine whether they show the hypermethylation characteristic of tumor cells or cells that do not normally contribute significantly to cfDNA, and / or hypomethylated variable epigenetic target regions are analyzed to determine whether they show the hypomethylation characteristic of tumor cells or cells that do not normally contribute significantly to cfDNA.

[0388] In some embodiments, the methods of the present invention are used to screen for cancer, e.g., metastasis, or in methods for screening for cancer, e.g., in methods for detecting the presence or absence of metastasis. For example, the sample may be from a subject who has or has not been previously diagnosed with cancer. In some embodiments, one or more, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more samples are collected from a subject described herein, e.g., before and / or after the subject is diagnosed with cancer. In some embodiments, the subject may or may not have cancer. In some embodiments, the subject may or may not have early stage cancer. In some embodiments, the subject has one or more risk factors for cancer, e.g., tobacco use (e.g., smoking), being overweight or obese, having a high body mass index (BMI), being elderly, nutritional deficiencies, high alcohol intake, or a family history of cancer.

[0389] In some embodiments, the subject has been using tobacco for at least 1, 5, 10 or 15 years, for example.In some embodiments, the subject has a high BMI, for example, a BMI of 25 or higher, 26 or higher, 27 or higher, 28 or higher, 29 or higher, or 30 or higher.In some embodiments, the subject is at least 40, 45, 50, 55, 60, 65, 70, 75 or 80 years old.In some embodiments, the subject has a nutritional deficiency, for example, a high intake of one or more of red meat and / or processed meat, trans fat, saturated fat and refined sugar, and / or a low intake of fruits and vegetables, complex carbohydrates and / or unsaturated fat. High intake and low intake can be defined as, for example, above or below the recommendations in the Dietary Guidelines for Americans 2020-2025 available at www.dietaryguidelines.gov / sites / default / files / 2021-03 / Dietary_Guidelines_for_Americans-2020-2025.pdf, respectively. In some embodiments, the subject has a high alcohol intake, e.g., averaging at least 3, 4, or 5 drinks per day (wherein a drink is about 1 ounce or 30 mL of 80 proof hard liquor or equivalent). In some embodiments, the subject has a family history of cancer, e.g., at least 1, 2, or 3 blood relatives have been previously diagnosed with cancer. In some embodiments, a relative is at least a third degree relative (e.g., a great-grandparent, a great aunt or uncle, a cousin), at least a second degree relative (e.g., a grandparent, an aunt or uncle, or a half-sibling), or a first degree relative (e.g., a parent or full sibling).

[0390] Furthermore, if a cancer is observed to be in remission following treatment, the methods of the invention can be used to monitor for residual disease or recurrence of the disease.

[0391] In some embodiments, the methods and systems disclosed herein can be used to identify customized or targeted therapies for treating a given disease or condition in a patient based on the classification of a nucleic acid variant as being of somatic or germline origin. Typically, the disease under consideration is a type of cancer. Non-limiting examples of such cancers include biliary tract cancer, bladder cancer, transitional cell carcinoma, urothelial carcinoma, brain cancer, glioma, astrocytoma, breast cancer, metaplastic carcinoma, cervical cancer, cervical squamous cell carcinoma, rectal cancer, colorectal cancer, colon cancer, hereditary nonpolyposis colorectal cancer, colorectal adenocarcinoma, gastrointestinal stromal tumor (GIST), endometrial cancer, endometrial stromal sarcoma, esophageal cancer, esophageal squamous cell carcinoma, esophageal adenocarcinoma, ocular melanoma, uveal melanoma, gallbladder cancer, gallbladder adenocarcinoma, renal cell carcinoma, clear cell renal cell carcinoma, transitional cell carcinoma, urothelial carcinoma, Wilms' tumor, leukemia, acute lymphocytic leukemia (ALL), acute myeloid leukemia (AML), chronic lymphocytic leukemia (CLL), chronic myelogenous leukemia (CML), chronic myelomonocytic leukemia (CMML), liver cancer, hepatoma, hepatoma, hepatocellular carcinoma, cholangiocarcinoma, hepatoblastoma, lung cancer, non-small cell lung cancer (NSCLC), mesothelioma, B cell lymphoma, non-Hodgkin's lymphoma, diffuse large B cell lymphoma, mantle cell lymphoma, T cell lymphoma, non-Hodgkin's lymphoma, precursor T lymphoblastic lymphoma / leukemia, peripheral T cell lymphoma, multiple myeloma, nasopharyngeal carcinoma (NPC), neuroblastoma, oropharyngeal carcinoma, oral squamous cell carcinoma, osteosarcoma, ovarian cancer, pancreatic cancer, pancreatic ductal adenocarcinoma, pseudopapillary neoplasm, acinar cell carcinoma, prostate cancer, prostate adenocarcinoma, skin cancer, melanoma, malignant melanoma, cutaneous melanoma, small intestine cancer, gastric cancer, gastrointestinal stromal tumor (GIST), uterine cancer or uterine sarcoma. The type and / or stage of cancer may be detected from genetic variations including mutations, rare mutations, indels, copy number variations, transversions, translocations, inversions, deletions, aneuploidy, partial aneuploidy, polyploidy, chromosomal instability, chromosomal structural alterations, gene fusions, chromosomal fusions, gene truncations, gene amplifications, gene duplications, chromosomal damage, DNA damage, abnormal changes in chemical modifications of nucleic acids, abnormal changes in epigenetic patterns, and abnormal changes in 5-methylcytosines of nucleic acids.

[0392] Genetic data can also be used to characterize specific forms of cancer. Cancers are often heterogeneous in both composition and staging. Genetic profile data can allow characterization of specific subtypes of cancer, which can be important in the diagnosis or treatment of that specific subtype. This information can also provide clues to the subject or practitioner regarding the prognosis of a particular type of cancer, allowing either the subject or practitioner to adapt treatment options as the disease progresses. Some cancers can progress to become more aggressive and genetically unstable. Other cancers can remain benign, inactive or dormant. The disclosed system and method can be useful in determining disease progression.

[0393] Furthermore, the methods of the present disclosure can be used to characterize the heterogeneity of abnormal conditions in a subject. Such methods can include, for example, generating a genetic profile of extracellular polynucleotides from a subject, the genetic profile including a plurality of data resulting from the analysis of copy number variations and rare mutations. In some embodiments, the abnormal condition is cancer. In some embodiments, the abnormal condition can be one that results in a heterogeneous genomic population. In the example of cancer, it is known that some tumors contain tumor cells that are at different stages of cancer. In other examples, the heterogeneity can include multiple foci of disease, for example, where one or more foci (e.g., one or more tumor foci) are the result of metastasis that has spread from the primary site of cancer. The tissue(s) of origin can be useful to identify the organ affected by the cancer, including the primary cancer and / or metastatic tumors.

[0394] The methods of the invention can also be used to quantify the levels of different cell types, including rare immune cell types, such as activated lymphocytes and myeloid cells, at a particular stage of differentiation. Such quantification can be based on the number of molecules corresponding to a given cell type in the sample. The sequence information obtained in the methods of the invention can include sequence reads of nucleic acids generated by a nucleic acid sequencer. In some embodiments, the nucleic acid sequencer performs pyrosequencing, single molecule sequencing, nanopore sequencing, semiconductor sequencing, sequencing by synthesis, five-letter sequencing, six-letter sequencing, sequencing by ligation, or sequencing by hybridization on the nucleic acid to generate sequencing reads. In some embodiments, the methods further include grouping the sequence reads into families of sequence reads, each family including sequence reads generated from nucleic acids in the sample. In some embodiments, these methods include determining the likelihood that the subject from whom the sample was taken has cancer or precancer, or has metastasis, which is related to changes in the proportion of immune cell types.

[0395] The methods of the present invention can be used to generate or profile a fingerprint or set of data that is the sum of genetic information from different cells in a heterogeneous disease. This set of data can include analysis of copy number variations, epigenetic variations, and mutations, either alone or in combination.

[0396] The method of the present invention can be used to diagnose, prognose, monitor or observe cancer or other diseases.In some embodiments, the method herein does not involve diagnosing, prognosing or monitoring fetus, and therefore does not concern non-invasive prenatal testing.In other embodiments, these methodologies can be used in pregnant subjects to diagnose, prognose, monitor or observe cancer or other diseases in subjects before birth, whose DNA and other polynucleotides may co-circulate with maternal molecules.

[0397] Non-limiting examples of other genetically based diseases, disorders or conditions that may be optionally evaluated using the methods and systems disclosed herein include achondroplasia, alpha 1 antitrypsin deficiency, antiphospholipid syndrome, autism, autosomal dominant polycystic kidney disease, Charcot-Marie-Tooth (CMT), cricketing, Crohn's disease, cystic fibrosis, Dercum's disease, Down's syndrome, Duane's syndrome, Duchenne muscular dystrophy, factor V Leiden thrombophilia, familial hypercholesterolemia, familial Mediterranean fever, fragile X syndrome, Gaucher's disease, hemochromatosis, hemophilia, holoprosencephaly, Huntington's disease, Klinefelter's syndrome, Marfan's syndrome, myotonic dystrophy, neurofibromatosis, Noonan's syndrome, osteogenesis imperfecta, Parkinson's disease, phenylketonuria, Poland anomaly ... These include porphyria, progeria, retinitis pigmentosa, severe combined immunodeficiency (SCID), sickle cell disease, spinal muscular atrophy, Tay-Sachs disease, thalassemia, trimethylaminuria, Turner syndrome, palatocardiofacial syndrome, WAGR syndrome, and Wilson's disease.

[0398] In some embodiments, the method described herein comprises using a set of sequence information obtained as described herein to detect the presence or absence of DNA originating from or derived from tumor cells at a preselected time point after a previous cancer treatment in a subject previously diagnosed with cancer. The method may further comprise determining a cancer recurrence score for the subject, the cancer recurrence score indicating the presence or absence of DNA originating from or derived from tumor cells.

[0399] When the cancer recurrence score is determined, it can be further used to determine the cancer recurrence status. The cancer recurrence status can be, for example, at risk of cancer recurrence when the cancer recurrence score is above a predetermined threshold. The cancer recurrence status can be, for example, at low or lower risk of cancer recurrence when the cancer recurrence score is above a predetermined threshold. In certain embodiments, a cancer recurrence score equal to a predetermined threshold can result in a cancer recurrence status of either at risk of cancer recurrence or at low or lower risk of cancer recurrence.

[0400] In some embodiments, the cancer recurrence score is compared to a predetermined cancer recurrence threshold, and the subject is classified as a candidate for subsequent cancer treatment if the cancer recurrence score is above the cancer recurrence threshold, or not as a candidate for treatment if the cancer recurrence score is below the cancer recurrence threshold. In certain embodiments, a cancer recurrence score equal to the cancer recurrence threshold may result in classification as either a candidate for subsequent cancer treatment or as not a candidate for treatment.

[0401] The methods discussed above may further include any suitability feature(s) set forth elsewhere herein, including in the sections relating to methods for determining the risk of cancer recurrence in a subject and / or methods for classifying a subject as a candidate for subsequent cancer treatment. 2. Methods for determining the risk of cancer recurrence in a subject and / or classifying a subject as a candidate for subsequent cancer treatment

[0402] In some embodiments, the method provided herein is or comprises the method of determining the risk of cancer recurrence in a subject.In some embodiments, the method provided herein is or comprises the method of detecting the presence or absence of metastasis in a subject.In some embodiments, the method provided herein is or comprises the method of classifying a subject as a candidate for subsequent cancer treatment.

[0403] Any of such methods may include collecting a sample (e.g., DNA, e.g., DNA originating from or derived from tumor cells) from a subject diagnosed with cancer at one or more preselected time points after one or more previous cancer treatments for the subject. The subject may be any of the subjects described herein. The sample may include chromatin, cfDNA, or other cellular material. The sample, e.g., a DNA sample, may be a tissue sample.

[0404] Any of such methods may include capturing a plurality of sets of target regions from DNA from a subject, the plurality of sets of target regions including a set of sequence variable target regions and a set of epigenetic target regions, thereby producing a captured set of DNA molecules. The capturing step may be performed according to any of the embodiments described elsewhere herein.

[0405] In any of such methods, the previous cancer treatment may include surgery, administration of a therapeutic composition, and / or chemotherapy.

[0406] Any such method includes a step of sequencing the captured DNA molecules, thereby producing a set of sequence information. The captured DNA molecules of the set of sequence-variable target regions may be sequenced to a greater sequencing depth than the captured DNA molecules of the set of epigenetic target regions.

[0407] Any of such methods may include using the set of sequence information to detect the presence or absence of DNA originating from or derived from a tumor cell at a preselected time point. Detecting the presence or absence of DNA originating from or derived from a tumor cell may be performed according to any of the embodiments thereof described elsewhere herein.

[0408] The method for determining the risk of cancer recurrence in a subject may include determining a cancer recurrence score for the subject, which indicates the presence or absence or amount of DNA originating from or derived from tumor cells, such as genomic regions of interest and target regions.The cancer recurrence score can be further used to determine a cancer recurrence status.The cancer recurrence status can be, for example, at risk of cancer recurrence when the cancer recurrence score is above a predetermined threshold.The cancer recurrence status can be, for example, at low or lower risk of cancer recurrence when the cancer recurrence score is above a predetermined threshold.In certain embodiments, a cancer recurrence score equal to a predetermined threshold may result in a cancer recurrence status of either at risk of cancer recurrence or at low or lower risk of cancer recurrence.

[0409] Methods for detecting the presence or absence of metastasis in a subject may include a step of comparing the presence or level of tissue-specific cellular material to the presence or level of tissue-specific cellular material obtained from the subject at a different time point, to a reference level of tissue-specific cellular material, or to a comparative cellular material. The methods herein may include an additional step for determining whether metastasis is present.

[0410] A method of classifying a subject as a candidate for subsequent cancer treatment may include comparing the subject's cancer recurrence score to a predetermined cancer recurrence threshold, whereby the subject is classified as a candidate for subsequent cancer treatment if the cancer recurrence score is above the cancer recurrence threshold, or is not classified as a candidate for treatment if the cancer recurrence score is below the cancer recurrence threshold. In certain embodiments, a cancer recurrence score equal to the cancer recurrence threshold may result in classification as either a candidate for subsequent cancer treatment or as not a candidate for treatment. In some embodiments, the subsequent cancer treatment includes administration of chemotherapy or a therapeutic composition.

[0411] Any of such methods may include determining a disease-free survival (DFS) period for the subject based on the cancer recurrence score; for example, the DFS period may be 1 year, 2 years, 3 years, 4 years, 5 years or 10 years.

[0412] In some embodiments, a sequence variable target region sequence is obtained and determining the cancer recurrence score may comprise determining at least a first subscore indicative of the amount of SNVs, insertions / deletions, CNVs and / or fusions present in the sequence variable target region sequence.

[0413] In some embodiments, the number of mutations in the sequence variable target region selected from 1, 2, 3, 4 or 5 is sufficient to generate a cancer recurrence score in which the first subscore is classified as positive for cancer recurrence. In some embodiments, the number of mutations is selected from 1, 2 or 3.

[0414] In some embodiments, the epigenetic target region sequence is obtained and determining the cancer recurrence score includes determining a second subscore indicative of the amount of molecules (obtained from the epigenetic target region sequence) that exhibit an epigenetic state different from the DNA found in a corresponding sample from a healthy subject (e.g., cfDNA found in a blood sample from a healthy subject, or DNA found in a tissue sample from a healthy subject if the tissue sample is the same type of tissue as that obtained from the subject). These abnormal molecules (i.e., molecules that have an epigenetic state different from the DNA found in the corresponding sample from the healthy subject) may be consistent with epigenetic changes associated with cancer (such as those with metastasis), such as methylation of a hypermethylated variable target region and / or disrupted fragmentation of a fragmented variable target region, where "disrupted" means different from the DNA found in the corresponding sample from the healthy subject.

[0415] In some embodiments, a percentage of molecules corresponding to the set of hypermethylated variable target regions and / or the set of fragmented variable target regions that exhibit hypermethylation in the set of hypermethylated variable target regions and / or aberrant fragmentation in the set of fragmented variable target regions that is greater than or equal to a value in the range of 0.001% to 10% is sufficient for the subscore to be classified as positive for cancer recurrence. This range may be 0.001% to 1%, 0.005% to 1%, 0.01% to 5%, 0.01% to 2%, or 0.01% to 1%.

[0416] In some embodiments, any of such methods may include determining the fraction of tumor DNA from the fraction of molecules in the set of sequence information that exhibit one or more features that indicate a tumor cell-derived origin. This may be performed for molecules that correspond to some or all of the target regions, including, for example, one or more of hypermethylated variable target regions, hypomethylated variable target regions, and fragmented variable target regions (hypermethylation of a hypermethylated variable target region and / or abnormal fragmentation of a fragmented variable target region may be considered to indicate a tumor cell-derived origin). This may be performed for molecules that correspond to sequence variable target regions, for example, molecules that include alterations that are consistent with cancer, such as SNVs, indels, CNVs, and / or fusions. The fraction of tumor DNA may be determined based on a combination of molecules that correspond to epigenetic target regions and molecules that correspond to sequence variable target regions.

[0417] The determination of the cancer recurrence score may be based at least in part on the fraction of tumor DNA, -11 ~1 or 10 -10 A fraction of tumor DNA greater than a threshold in the range of 1 to 1 is sufficient for the Cancer Recurrence Score to be classified as positive for cancer recurrence. -10 ~10 -9 , 10 -9 ~10 -8 , 10 -8 ~10 -7 , 10 -7 ~10 -6 , 10 -6 ~10 -5 , 10 -5 ~10-4 , 10 -4 ~10 -3 , 10 -3 ~10 -2 or 10 -2 ~10 -1 In some embodiments, a fraction of tumor DNA greater than or equal to a threshold in the range of at least 10 is sufficient for the Cancer Recurrence Score to be classified as positive for cancer recurrence. -7 The fraction of tumor DNA greater than the threshold is sufficient for the cancer recurrence score to be classified as positive for cancer recurrence. The determination that the fraction of tumor DNA is greater than a threshold, for example, a threshold corresponding to any of the above-mentioned embodiments, can be based on cumulative probability. For example, a sample is considered positive if the cumulative probability that the tumor fraction is greater than a threshold in any of the above-mentioned ranges exceeds a probability threshold of at least 0.5, 0.75, 0.9, 0.95, 0.98, 0.99, 0.995 or 0.999. In some embodiments, the probability threshold is at least 0.95, for example 0.99.

[0418] In some embodiments, the set of sequence information includes sequence variable target region sequences and epigenetic target region sequences, and determining the cancer recurrence score includes determining a subscore indicative of the amount of SNVs, insertions / deletions, CNVs and / or fusions present in the sequence variable target region sequences, and a subscore indicative of the amount of abnormal molecules in the epigenetic target region sequences, and combining the subscores to provide a cancer recurrence score. When subscores are combined, they can be combined by applying a threshold value (e.g., for sequence variable target regions, greater than a predetermined number of mutations (e.g., >1), and for epigenetic target regions, greater than a predetermined fraction of abnormal molecules (i.e., molecules with an epigenetic state different from that found in the corresponding sample from a healthy subject; e.g., tumor)) to each subscore independently, or by training a machine learning classifier to determine the status based on multiple positive and negative training samples.

[0419] In some embodiments, a value for the combined score in the range of -4 to 2 or -3 to 1 is sufficient for the Cancer Recurrence Score to be classified as positive for cancer recurrence.

[0420] In any embodiment in which the Cancer Recurrence Score is classified as positive for cancer recurrence, the subject's Cancer Recurrence Status may be at risk of cancer recurrence and / or the subject may be classified as a candidate for subsequent cancer treatment.

[0421] In some embodiments, the cancer is any one of the types of cancer described elsewhere herein, e.g., colorectal cancer. 3. Treatment and Related Administration

[0422] In certain embodiments, the methods disclosed herein relate to customized treatment, for example, identifying customized treatment and administering it to a patient.In some embodiments, the patient or subject has a given disease, disorder or condition, for example, any of the cancer or other conditions described elsewhere herein.Essentially any cancer treatment (e.g., surgical treatment, radiation therapy, chemotherapy, immunotherapy, etc.) can be included as part of these methods.In certain embodiments, the treatment administered to the subject includes at least one chemotherapy drug. In some embodiments, chemotherapy drugs may include alkylating agents (e.g., without limitation, chlorambucil, cyclophosphamide, cisplatin and carboplatin), nitrosoureas (e.g., without limitation, carmustine and lomustine), antimetabolites (e.g., without limitation, Fluorauracil, methotrexate and fludarabine), plant alkaloids and natural products (e.g., without limitation, vincristine, paclitaxel and topotecan), antitumor antibiotics (e.g., without limitation, bleomycin, doxorubicin and mitoxantrone), hormonal agents (e.g., without limitation, prednisone, dexamethasone, tamoxifen and leuprolide) and biological response modifiers (e.g., without limitation, herceptin and avastin, erbitux and rituxan). In some embodiments, chemotherapy administered to a subject may include FOLFOX or FOLFIRI. In certain embodiments, the treatment comprises at least one PARP inhibitor can be administered to the subject.In certain embodiments, the PARP inhibitor can include, among others, OLAPARIB, TALAZOPARIB, RUCAPARIB, NIRAPARIB (brand name ZEJULA).Typically, the treatment comprises at least one immunotherapy (or immunotherapeutic agent).Immunotherapy generally refers to the method of enhancing immune response to a given cancer type.In certain embodiments, immunotherapy refers to the method of enhancing T cell response to tumor or cancer.

[0423] In some embodiments, the immunotherapy or immunotherapeutic agent targets immune checkpoint molecules.Certain tumors can evade the immune system by utilizing immune checkpoint pathways.Therefore, targeting immune checkpoints has emerged as an effective approach to negate tumors' ability to evade the immune system and activate antitumor immunity against certain cancers.Pardoll, Nature Reviews Cancer, 2012, 12:252-264.

[0424] In certain embodiments, the immune checkpoint molecule is an inhibitory molecule that reduces signals involved in T cell responses to antigens. For example, CTLA4 is expressed on T cells and plays a role in downregulating T cell activation by binding to CD80 (also known as B7.1) or CD86 (also known as B7.2) on antigen-presenting cells. PD-1 is another inhibitory checkpoint molecule expressed on T cells. PD-1 limits the activity of T cells in peripheral tissues during inflammatory responses. In addition, the ligand for PD-1 (PD-L1 or PD-L2) is commonly upregulated on the surface of many different tumors, resulting in downregulation of anti-tumor immune responses in the tumor microenvironment. In certain embodiments, the inhibitory immune checkpoint molecule is CTLA4 or PD-1. In other embodiments, the inhibitory immune checkpoint molecule is a ligand for PD-1, such as PD-L1 or PD-L2. In other embodiments, the inhibitory immune checkpoint molecule is a ligand for CTLA4, such as CD80 or CD86. In other embodiments, the inhibitory immune checkpoint molecule is lymphocyte activation gene 3 (LAG3), killer cell immunoglobulin-like receptor (KIR), T cell membrane protein 3 (TIM3), galectin 9 (GAL9) or adenosine A2a receptor (A2aR).

[0425] Antagonists targeting these immune checkpoint molecules can be used to enhance antigen-specific T cell responses to certain cancers. Thus, in certain embodiments, the immunotherapy or immunotherapeutic agent is an inhibitory immune checkpoint molecule antagonist. In certain embodiments, the inhibitory immune checkpoint molecule is PD-1. In certain embodiments, the inhibitory immune checkpoint molecule is PD-L1. In certain embodiments, the inhibitory immune checkpoint molecule antagonist is an antibody (e.g., a monoclonal antibody). In certain embodiments, the antibody or monoclonal antibody is an anti-CTLA4, anti-PD-1, anti-PD-L1 or anti-PD-L2 antibody. In certain embodiments, the antibody is a monoclonal anti-PD-1 antibody. In some embodiments, the antibody is a monoclonal anti-PD-L1 antibody. In certain embodiments, the monoclonal antibody is a combination of an anti-CTLA4 antibody and an anti-PD-1 antibody, an anti-CTLA4 antibody and an anti-PD-L1 antibody, or an anti-PD-L1 antibody and an anti-PD-1 antibody. In certain embodiments, the anti-PD-1 antibody is one or more of pembrolizumab (Keytruda®) or nivolumab (Opdivo®). In certain embodiments, the anti-CTLA4 antibody is ipilimumab (Yervoy®). In certain embodiments, the anti-PD-L1 antibody is one or more of atezolizumab (Tecentriq®), avelumab (Bavencio®), or durvalumab (Imfinzi®).

[0426] In certain embodiments, the immunotherapy or immunotherapeutic agent is an antagonist (e.g., an antibody) against CD80, CD86, LAG3, KIR, TIM3, GAL9, or A2aR. In other embodiments, the antagonist is a soluble version of an inhibitory immune checkpoint molecule, such as a soluble fusion protein comprising the extracellular domain of an inhibitory immune checkpoint molecule and the Fc domain of an antibody. In certain embodiments, the soluble fusion protein comprises the extracellular domain of CTLA4, PD-1, PD-L1, or PD-L2. In some embodiments, the soluble fusion protein comprises the extracellular domain of CD80, CD86, LAG3, KIR, TIM3, GAL9, or A2aR. In one embodiment, the soluble fusion protein comprises the extracellular domain of PD-L2 or LAG3.

[0427] In certain embodiments, the immune checkpoint molecule is a costimulatory molecule that amplifies the signal involved in T cell response to antigen. For example, CD28 is a costimulatory receptor expressed on T cells. When T cells bind to antigen through their T cell receptor, CD28 binds to CD80 (also known as B7.1) or CD86 (also known as B7.2) on antigen-presenting cells to amplify T cell receptor signaling and promote T cell activation. Because CD28 binds to the same ligands (CD80 and CD86) as CTLA4, CTLA4 can counter or regulate the costimulatory signaling mediated by CD28. In certain embodiments, the immune checkpoint molecule is a costimulatory molecule selected from CD28, inducible T cell costimulator (ICOS), CD137, OX40 or CD27. In other embodiments, the immune checkpoint molecule is a ligand for a costimulatory molecule, including, for example, CD80, CD86, B7RP1, B7-H3, B7-H4, CD137L, OX40L, or CD70.

[0428] Agonists that target these costimulatory checkpoint molecules can be used to enhance antigen-specific T cell responses to certain cancers. Thus, in certain embodiments, the immunotherapy or immunotherapeutic agent is an agonist of costimulatory checkpoint molecules. In certain embodiments, the agonist of costimulatory checkpoint molecules is an agonist antibody, preferably a monoclonal antibody. In certain embodiments, the agonist antibody or monoclonal antibody is an anti-CD28 antibody. In other embodiments, the agonist antibody or monoclonal antibody is an anti-ICOS, anti-CD137, anti-OX40 or anti-CD27 antibody. In other embodiments, the agonist antibody or monoclonal antibody is an anti-CD80, anti-CD86, anti-B7RP1, anti-B7-H3, anti-B7-H4, anti-CD137L, anti-OX40L or anti-CD70 antibody.

[0429] In certain embodiments, the status of nucleic acid variants from a sample from a subject, such as somatic or germline origin, can be compared with a database of comparative subject results from a reference population to identify customized or targeted treatment for the subject.Typically, the reference population includes patients with the same cancer or disease type as the subject and / or patients who are undergoing or have undergone the same treatment as the subject.Customized or targeted treatment (or multiple treatments) can be identified when the nucleic acid variant and comparative subject results meet certain classification criteria (e.g., are substantial or close match).

[0430] In certain embodiments, the customized treatment described herein is typically administered parenterally (e.g., intravenously or subcutaneously). Pharmaceutical compositions containing immunotherapeutic agents are typically administered intravenously. Certain therapeutic agents are administered orally. However, the customized treatment (e.g., immunotherapeutic agents, etc.) can also be administered by any method known in the art, for example, buccal, sublingual, rectal, vaginal, intraurethral, ​​topical, intraocular, intranasal, and / or intraauricular, and the administration can include tablets, capsules, granules, aqueous suspensions, gels, sprays, suppositories, salves, ointments, etc.

[0431] In some embodiments, treatment is customized based on the status of the nucleic acid variant, such as somatic or germline origin. In some embodiments, determining the level of a particular cell type, e.g., an immune cell type, including a rare immune cell type, facilitates the selection of an appropriate treatment.

[0432] The method of the present invention can be used to diagnose the presence of a condition in a subject, such as cancer or precancer, characterize a condition (e.g., to determine the cancer stage or cancer heterogeneity), monitor the subject's response to receiving treatment for a condition (e.g., response to a chemotherapeutic or immunotherapeutic agent), evaluate a subject's prognosis (e.g., to predict survival outcome in a subject with cancer), determine a subject's risk of developing a condition, predict the subsequent course of a condition in a subject, determine cancer metastasis or recurrence (or risk of cancer metastasis or recurrence) in a subject, and / or monitor the health of a subject as part of a preventive health monitoring program (e.g., to determine whether and / or when a subject requires further diagnostic screening).The method according to the present disclosure can also be useful in predicting a subject's response to a particular treatment option. If the treatment is successful, more cancer cells may die and shed DNA, or if successful treatment results in an increase or decrease in the amount of a particular immune cell type in the blood, and unsuccessful treatment results in no change, a successful treatment option may increase the amount of copy number variations, rare mutations, and / or cancer-associated epigenetic signatures (e.g., hypermethylated or hypomethylated regions) detected in the subject's blood (e.g., DNA isolated from a buffy coat sample from the subject, or any other sample containing cells, such as a blood sample (e.g., a whole blood sample, a leukapheresis sample, or a PBMC sample)). In other examples, this may not occur. In another example, certain treatment options may be correlated with the genetic profile of the cancer over time. This correlation may be useful in selecting a treatment for the subject. In some embodiments, determining the metastatic site facilitates the selection of an appropriate treatment.

[0433] In some embodiments, the treatment is customized based on the status of the detected nucleic acid variant, such as somatic or germline origin. In some embodiments, essentially any cancer treatment (e.g., surgical treatment, radiation therapy, chemotherapy, etc.) can be included as part of these methods. Typically, the customized treatment includes at least one immunotherapy (or immunotherapeutic agent). Immunotherapy generally refers to a method of enhancing immune response to a given cancer type. In certain embodiments, immunotherapy refers to a method of enhancing T cell response to tumor or cancer.

[0434] In certain embodiments, the status of nucleic acid variants from samples from subjects, that is, somatic or germline origin, can be compared with a database of comparative subject results from a reference population to identify customized or targeted treatment for the subject.Typically, the reference population comprises patients with the same cancer or disease type as the subject and / or patients who are undergoing or have undergone the same treatment as the subject.Customized or targeted treatment(s) can be identified when the nucleic acid variants and comparative subject results meet certain classification criteria (e.g., are substantial or close match).

[0435] The disclosed methods may include evaluating (e.g., quantifying) and / or interpreting at least one cellular material (e.g., at least one cellular material in a sample from a subject) released from a potential metastatic site and / or cell type that contributes to DNA, e.g., cfDNA, in one or more samples collected from the subject at one or more time points, relative to a selected baseline value or reference standard (or a selected set of baseline values ​​or reference standards). The baseline value or reference standard may be the presence or level of at least one cellular material and / or amount of a cell type (e.g., the average amount or range of amounts of a cell type present in at least two samples) measured in one or more samples collected from the subject at one or more time points, e.g., before undergoing treatment, before diagnosis of a condition (e.g., cancer), or as part of a preventative health monitoring program. The baseline value or reference standard can be the presence or level of at least one cellular material and / or amount of a cell type (e.g., the average amount or range of amounts of a cell type present in at least two samples) measured for one or more samples collected at one or more time points from one or more subjects without a condition (e.g., healthy subjects without cancer), one or more subjects who have responded favorably to a treatment, or one or more subjects who have not received a treatment. In certain embodiments, the baseline value or reference standard utilized is a standard or profile derived from a single reference subject. In other embodiments, the baseline value or reference standard utilized is a standard or profile derived from data averaged from multiple reference subjects. The reference standard can, in various embodiments, be a single value, a mean, an average, a numerical mean or range of numerical means, a numerical pattern, or a graphic pattern derived from a single reference subject or generated from cell type amount data derived from multiple reference subjects. The selection of a particular baseline value or reference standard, or the selection of one or more reference subjects, depends on how the methods described herein are used, for example, by a research scientist or clinician (e.g., physician).

[0436] In some embodiments, methods are provided for monitoring a subject's response to a treatment (e.g., chemotherapy or immunotherapy) (e.g., a change in disease state, e.g., the presence or absence of metastasis in a subject, as measured, such as by assessing the presence or level of at least one cellu...

Claims

1. 1. A method for analyzing DNA molecules in a sample, wherein the DNA molecules comprise first and second strands and asymmetric adapters, optionally wherein at least one asymmetric adapter comprises a deamination-sensitive cytosine, the method comprising: a) synthesizing a first complementary strand complementary to the first strand and a second complementary strand complementary to the second strand; b) glucosylating 5-hydroxymethylated cytosines in at least one of the first or second strands before or after synthesizing the first and second complementary strands; c) methylating cytosines in at least one of the first complementary strand or the second complementary strand, wherein said methylation converts semi-methylated CpGs to fully methylated CpGs; d) deaminating unmodified cytosines in at least one of the first or second strands, thereby producing a treated DNA molecule; and e) sequencing at least a portion of said treated DNA molecules Including; Optionally, the asymmetric adaptor is a Y-type adaptor or a bubble adaptor.

2. (i) each asymmetric adapter comprises at least one deamination-sensitive cytosine, and / or the deamination-sensitive cytosine is an unmethylated cytosine; (ii) each asymmetric adaptor comprises one deamination-sensitive cytosine and at least one deamination-resistant cytosine, and optionally, the deamination-resistant cytosine is a 5-methylcytosine, and / or each cytosine other than the one deamination-sensitive cytosine in each asymmetric adaptor is a deamination-resistant cytosine; (iii) the nucleotide immediately 3′ to the deamination-susceptible cytosine comprises a nucleobase other than guanine, and optionally the nucleobase other than guanine is adenine, cytosine, thymine, or uracil; and / or (iv) the step of deaminating unmodified cytosines comprises bisulfite conversion. The method of claim 1.

3. 1. A method for analyzing DNA molecules in a sample, wherein the DNA molecules comprise first and second strands and asymmetric adapters, optionally wherein at least one asymmetric adapter comprises an unmodified cytosine, the method comprising: a) oxidizing 5-hydroxymethylated cytosines in at least one first or second strand to 5-formylcytosines; b) synthesizing a first complementary strand complementary to the first strand and a second complementary strand complementary to the second strand; c) methylating cytosines in at least one of the first complementary strand or the second complementary strand, wherein said methylation converts semi-methylated CpGs to fully methylated CpGs; d) converting the modified cytosine in at least one of the first or second strands to a thymine or a base that is read as a thymine, thereby producing a treated DNA molecule; and e) sequencing at least a portion of said treated DNA molecules Including; Optionally, the asymmetric adaptor is a Y-type adaptor or a bubble adaptor.

4. (i) the step of converting the modified cytosine in at least one of the first or second strands to thymine or a base that is read as thymine comprises oxidizing hydroxymethylcytosine; for example, the hydroxymethylcytosine is oxidized to formylcytosine; and optionally, oxidizing the hydroxymethylcytosine to formylcytosine comprises contacting the hydroxymethylcytosine with a ruthenate, and optionally, the ruthenate is KRuO4; (ii) the modified cytosine is converted to thymine, uracil, or dihydrouracil; and / or (iii) the method comprises, as part of the step of converting the modified cytosine in at least one first or second strand to thymine or a base that is read as thymine, converting formylcytosine and / or methylcytosine to carboxylcytosine; for example, (A) the step of converting the formylcytosine and / or the methylcytosine to a carboxylcytosine comprises contacting the formylcytosine and / or the methylcytosine with a TET enzyme, optionally wherein the TET enzyme is TET1, TET2 or TET3; and / or (B) the method includes, as part of converting the modified cytosine in at least one of the first or second strands to thymine or a base that is read as thymine, reducing the carboxyl cytosine; e.g., the carboxyl cytosine is reduced to dihydrouracil; and / or the step of reducing the carboxyl cytosine comprises contacting the carboxyl cytosine with a reducing agent, optionally wherein the reducing agent is a borane or borohydride reducing agent, further optionally wherein the borane or borohydride reducing agent comprises pyridine borane, 2-picoline borane, borane, tert-butylamine borane, ammonia borane, sodium cyanoborohydride (NaBH3CN), lithium borohydride (LiBH4), sodium borohydride, ethylenediamine borane, dimethylamine borane, sodium triacetoxyborohydride, morpholine borane, 4-methylmorpholine borane, trimethylamine borane, dicyclohexylamine borane, or a salt thereof; The method of claim 3.

5. the asymmetric adapter comprises a molecular barcode; and / or the method comprises preparing the DNA molecule by attaching the Y-shaped adaptor to a precursor DNA molecule, optionally wherein the attaching comprises ligating, e.g., the precursor DNA molecule is cell-free DNA and / or the precursor DNA molecule is obtained from a sample; optionally, the sample is from a mammal and / or the sample is a blood sample. The method according to claim 1 or claim 3.

6. dividing the sample into at least the first sub-sample and a second sub-sample before attaching the Y-shaped adaptors to the DNA molecules of at least a first sub-sample; and optionally (i) the dividing step is based on an epigenetic modification; e.g., the epigenetic modification is a cytosine modification, e.g., 5-methylation, and / or the first sub-sample comprises DNA molecules enriched for the epigenetic modification; (ii) the dividing step comprises contacting the DNA molecules with a methyl-binding reagent immobilized on a solid support; and / or (iii) differentially tagging the first sub-sample and the second sub-sample; optionally, pooling DNA from the first sub-sample and the second sub-sample, and / or sequencing DNA from the first sub-sample and the set of target regions or the second sub-sample in the same sequencing cell. The method according to claim 1 or claim 3.

7. The DNA of the first sub-sample and the DNA of the second sub-sample are differentially tagged; after differential tagging, a portion of the DNA from the second sub-sample or processed sub-sample is added to the first sub-sample or additional processed sub-sample or at least a portion thereof, thereby forming a pool; sequence variable target regions and epigenetic target regions are captured from the pool; e.g., (i) (A) the pool comprises less than or equal to about 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, or 5% of the DNA of the second sub-sample; or the pool comprises about 70-90%, about 75-85%, or about 80% of the DNA of the second sub-sample; and / or (B) the pool comprises substantially all of the DNA of the first sub-sample; or the pool comprises substantially all of the DNA of the first sub-sample or the processed first sub-sample; (ii) a first set of target regions is captured from at least a portion of said first sub-sample or processed first sub-sample after forming said pool; and / or (iii) the plurality of sub-samples includes a third sub-sample comprising a higher proportion of DNA with the epigenetic modification than the second sub-sample but a lower proportion than the first sub-sample; optionally, the method further comprises differentially tagging the third sub-sample, and further optionally, the DNA from the first sub-sample, the DNA from the third sample, and the set of target regions are pooled, and optionally the DNA from the first, second, and third sub-samples is sequenced in the same sequencing cell. The method of claim 6. (A) the DNA molecule is amplified, e.g., the DNA molecule is amplified after the deaminating step and / or before the sequencing step; and / or (B) the sequencing step is next-generation sequencing; The method according to claim 1 or claim 3.

9. The method may further include a step of capturing a set of target regions of the processed DNA molecules prior to the sequencing step, and the captured processed DNA molecules may be sequenced or amplified and sequenced; for example, the set of target regions may be: (i) comprises an epigenetic target region; (ii) comprises a set of hypermethylated variable target regions; e.g., the set of hypermethylated variable target regions comprises regions having a higher degree of methylation in at least one type of tissue than the degree of methylation in cell-free DNA from a healthy subject; (iii) comprises a set of hypomethylated variable target regions; e.g., the set of hypomethylated variable target regions comprises regions that have a lower degree of methylation in at least one type of tissue than the degree of methylation in cell-free DNA from a healthy subject; (iv) includes a set of methylation control target regions; (v) comprising a set of fragmented variable target regions; for example, the set of fragmented variable target regions comprises a transcription start site region and / or a CTCF binding region; (vi) comprises a set of hydroxymethylated variable target regions; e.g., the set of hydroxymethylated variable target regions comprises regions having a higher degree of hydroxymethylation in at least one type of tissue than the degree of hydroxymethylation in cell-free DNA from healthy subjects, and / or DNA molecules corresponding to the set of hydroxymethylated variable target regions are captured with a higher capture yield than DNA molecules corresponding to at least one other set of target regions; and / or (vii) comprise sequence-variable target regions; e.g., DNA molecules corresponding to the set of sequence-variable target regions are captured with a higher capture yield than DNA molecules corresponding to the set of epigenetic target regions; The method according to claim 1 or claim 3.

10. (i) the DNA molecule comprises an inserted DNA derived from a subject, and the method further comprises obtaining mapped sequence reads as an indication of likelihood that the subject has cancer; for example, the sequencing step generates a plurality of sequencing reads, and the method further comprises mapping the plurality of sequence reads to one or more reference sequences to generate mapped sequence reads, and processing the mapped sequence reads corresponding to the set of sequence variable target regions and the mapped sequence reads corresponding to the set of epigenetic target regions to indicate likelihood that the subject has cancer; and / or (ii) sequencing a captured set of cfDNA molecules obtained at one or more preselected time points after the one or more previous cancer treatments, wherein the DNA molecules comprise insert DNA from a subject, the subject having previously been diagnosed with cancer and having undergone one or more previous cancer treatments, whereby a set of sequence information is generated; and optionally, (A) using said set of sequence information to detect the presence or absence of DNA originating from or derived from a tumor cell at a preselected time point; and optionally, (B) determining for the subject a cancer recurrence score indicative of the presence or absence of the DNA originating from or derived from the tumor cells, and optionally determining based on the cancer recurrence score that the cancer recurrence score is indicative of a cancer recurrence status, wherein a cancer recurrence score at or above a predetermined threshold indicates that the cancer recurrence status of the subject is determined to be at risk of cancer recurrence if the cancer recurrence score is determined to be at or above a predetermined threshold, or a cancer recurrence score below the predetermined threshold indicates that the cancer recurrence status of the subject is determined to be at a lower risk of cancer recurrence if the cancer recurrence score is below the predetermined threshold; optionally, (C) further comprising the step of comparing the subject's cancer recurrence score to a predetermined cancer recurrence threshold, wherein a cancer recurrence score above the cancer recurrence threshold indicates that the subject will be classified as a candidate for subsequent cancer treatment if the cancer recurrence score is above the cancer recurrence threshold, or a cancer recurrence score below the cancer recurrence threshold indicates that the subject will not be a candidate for subsequent cancer treatment if the cancer recurrence score is below the cancer recurrence threshold. The method according to claim 1 or claim 3.

11. (i) (A) the step of glucosylating the 5-hydroxymethylated cytosine comprises contacting the 5-hydroxymethylated cytosine with a β-glucosyltransferase and / or produces 5-glucosylhydroxymethylcytosine; and / or (B) methylating the cytosine in at least one of the first complementary strand or the second complementary strand comprises contacting the cytosine with a DNA methyltransferase; e.g., the DNA methyltransferase is DNMT1 or DNMT5; (ii) identifying methylated positions in said DNA molecule and / or identifying hydroxymethylated positions in said DNA molecule; optionally further comprising identifying the gene sequence of said DNA molecule; (iii) identifying at least one position in the DNA molecule that contained a hydroxymethylated cytosine; at least one position in the DNA molecule that contained a methylated cytosine; at least one position in the DNA molecule that contained a cytosine that was neither methylated nor hydroxymethylated; at least one position in the DNA molecule that contained an adenine; at least one position in the DNA molecule that contained a guanine; and at least one position in the DNA molecule that contained a thymine; e.g., the cytosine that was neither methylated nor hydroxymethylated was an unmodified cytosine. (iv) the step of synthesizing a first complementary strand complementary to the first strand and a second complementary strand complementary to the second strand comprises: (A) extending a primer using unmethylated dNTPs; e.g., the dNTPs consist of unmethylated dNTPs; (B) converting at least one methylated CpG to a hemi-methylated CpG; and / or (C) converting at least one hydroxymethylated CpG to a hemi-hydroxymethylated CpG; (v) the 5-hydroxymethylated cytosine contained in the hemihydroxymethylated CpG is glycosylated in at least one of the first or second strands; (vi) the DNA molecule is free in solution during one or more of the synthesizing, glucosylating, methylating and deaminating steps, optionally the DNA molecule is free in solution during two, three or four of the synthesizing, glucosylating, methylating and deaminating steps, and further optionally the DNA molecule is free in solution during each of the synthesizing, glucosylating, methylating and deaminating steps; and / or (vii) steps a) to e) are carried out in the order a) to e); The method according to claim 1 or claim 3.

12. a) reagents for synthesizing a first complementary strand complementary to the first strand and a second complementary strand complementary to the second strand; b) a reagent for glycosylation of 5-hydroxymethylated cytosines in at least one of the first or second strands before or after synthesis of said first and second complementary strands; c) a reagent for methylating cytosines in at least one of the first complementary strand or the second complementary strand, said methylation converting semi-methylated CpGs to fully methylated CpGs; d) a reagent for deaminating unmodified cytosines in at least one of the first or second strands; e) a plurality of oligonucleotide probes; f) primers for synthesizing a first complementary strand complementary to the first strand and a second complementary strand complementary to the second strand; g) a primer for sequencing; and h) Library adapters with distinct molecular barcodes A kit comprising one or more of: (i) the reagents for synthesizing a first complementary strand complementary to the first strand and a second complementary strand complementary to the second strand comprise a polymerase and / or dNTPs, and optionally, the dNTPs are unmethylated; (ii) the reagent for glucosylating 5-hydroxymethylated cytosines in at least one of the first or second strands before or after synthesizing the first and second complementary strands is a glucosyltransferase and / or uridine diphosphate glucose, and optionally the glucosyltransferase is a β-glucosyltransferase; (iii) the reagent for methylating cytosines in at least one of a first complementary strand or a second complementary strand comprises one or more of a DNA methyltransferase and a methyl donor, optionally wherein the DNA methyltransferase is DNMT1 or DNMT5, and optionally wherein the methyl donor is S-adenosylmethionine; and / or (iv) the reagent for deaminating unmodified cytosines in at least one first or second strand comprises one or more of sodium bisulfite or APOBEC3A; The kit of claim 12. (i) a reagent for converting 5mC and 5hmC into substrates that cannot be deaminated by a deaminase, optionally wherein the reagent for converting 5mC and 5hmC into said substrates that cannot be deaminated by a deaminase comprises a TET enzyme or T4-βGT; (j) A) a reagent for oxidizing 5hmC to formylcytosine, optionally wherein the reagent for oxidizing 5hmC to formylcytosine is a ruthenate, optionally wherein the ruthenate is KRuO 4 ; B) TET enzymes, optionally TET1, TET2 and / or TET3; C) borane or borohydride reducing agents, including pyridine borane, 2-picoline borane, borane, tert-butylamine borane, ammonia borane, sodium borohydride, ethylenediamine borane, sodium cyanoborohydride (NaBH3CN), lithium borohydride (LiBH4), dimethylamine borane, sodium triacetoxyborohydride, morpholine borane, 4-methylmorpholine borane, trimethylamine borane, dicyclohexylamine borane, or salts thereof; and / or D) lithium aluminum hydride, sodium amalgam, amalgam, sulfur dioxide, dithionate, thiosulfate, iodide, hydrogen peroxide, hydrazine, diisobutylaluminum hydride, oxalic acid, carbon monoxide, cyanide, ascorbic acid, formic acid, dithiothreitol, beta-mercaptoethanol, or any combination thereof and / or (k) instructions for carrying out the method of any one of claims 1 to 11. further comprising: The kit of claim 12.

15. a) the plurality of oligonucleotide probes are selected from the group consisting of ALK, APC, BRAF, CDKN2A, EGFR, ERBB2, FBXW7, KRAS, MYC, NOTCH1, NRAS, PIK3CA, PTEN, RBI, TP53, MET, AR, ABL1, AKT1, ATM, CDH1, CSFIR, CTNNBl, ERBB4, EZH2, FGFR1, FGFR2, FGFR3, FLT3, GNA11, GNAQ, GNAS, HNF1A, HRAS, IDH1, IDH2, JAK2, JAK3, KDR, KIT, MLH1, MPL, NPM1, PDG selectively hybridizes to at least 5, 6, 7, 8, 9, 10, 20, 30, 40 or all genes selected from FRA, PROC, PTPN11, RET, SMAD4, SMARCB1, SMO, SRC, STK11, VHL, TERT, CCND1, CDK4, CDKN2B, RAF1, BRCA1, CCND2, CDK6, NF1, TP53, ARID1A, BRCA2, CCNE1, ESR1, RIT1, GATA3, MAP2K1, RHEB, ROS1, ARAF, MAP2K2, NFE2L2, RHOA, and NTRK1; b) the library adapters do not contain flow cell sequences or sequences that allow for the formation of hairpin loops for sequencing; c) the library adapters are blunt-ended and Y-shaped; and / or d) the library adapters are less than or equal to 40 nucleobases in length A kit according to any one of claims 12 to 14.

16. Use of the method according to claim 1 or claim 3, (A) to characterize the heterogeneity of an abnormal condition in a subject; e.g., the method comprises generating a genetic profile of extracellular polynucleotides from the subject, the genetic profile comprising a plurality of data resulting from analysis of copy number variations and rare mutations; e.g., the abnormal condition is cancer; and / or (B) To generate a profile, fingerprint, or data set that is a summation of genetic information from different cells in a heterogeneous disease, for example, the data set includes analysis of copy number variation, epigenetic variation, and mutation, alone or in combination; Use of.