Integrated targeted and whole-genome somatic and DNA methylation sequencing workflow
An integrated library preparation method for cell-free DNA analysis in liquid biopsies improves detection of methylation status and fragmentomic signals by enriching specific target regions and sequencing, addressing the challenges of low concentration and heterogeneity in current methods.
Patent Information
- Application Number
- JP2025536229
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-12-22
- Filing Date
- 2023-12-22
- Publication Date
- 2025-12-25
AI Technical Summary
Current methods for analyzing cell-free DNA in liquid biopsies face challenges in accurately detecting non-sequence modifications such as methylation status and fragmentomic signals due to the low concentration and heterogeneity of cell-free DNA, making it difficult to isolate and process a useful fraction for detailed analysis.
An integrated library preparation method that involves dividing DNA into sub-samples, enriching specific target regions, and sequencing them to improve nucleic acid-based analysis, including methods for adjusting epigenetic and sequence variable target regions to enhance cancer screening.
This approach allows for improved detection of both specific changes in target regions and broad signatures in a single workflow, reducing assay costs and enhancing the sensitivity of cancer screening methods.
Smart Images

Figure 2025542261000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 476,923, filed December 22, 2022, which is incorporated herein by reference for all purposes.
[0002] FIELD OF THE INVENTION The present disclosure provides compositions and methods related to analyzing DNA, for example, cell-free DNA. In some embodiments, the DNA is DNA from a subject who has or is suspected of having cancer, and / or the DNA comprises DNA from cancer cells. In some embodiments, the DNA is divided into at least a first sub-sample and a second sub-sample. In certain embodiments, the first sub-sample comprises DNA with a higher proportion of nucleotide modifications (e.g., cytosine modifications) than the second sub-sample. In some embodiments, the DNA of the first and second sub-samples is combined before sequencing. [Background technology]
[0003] Introduction and Abstract Cancer causes millions of deaths worldwide each year. Early detection of cancer can lead to improved outcomes, as early-stage cancers tend to be more susceptible to treatment.
[0004] Improperly controlled cell growth is a hallmark of cancer, typically resulting from the accumulation of genetic and epigenetic alterations, such as copy number variations (CNVs), single nucleotide variations (SNVs), gene fusions, insertions and / or deletions (indels), cytosine modifications (e.g., 5-methylcytosine, 5-hydroxymethylcytosine, and other more oxidized forms), and epigenetic variations involving the association of DNA with chromatin proteins and transcription factors. Thus, cancer can be manifested by non-sequence alterations, such as methylation. Examples of methylation changes in cancer include localized gain of DNA methylation in CpG islands at the TSSs of genes involved in normal growth control, DNA repair, cell cycle regulation, and / or cell differentiation. Hypermethylation can be associated with abnormal loss of transcriptional capacity of the involved genes and occurs at least as frequently as point mutations and deletions as a cause of altered gene expression. Furthermore, without wishing to be bound by any particular theory, cells in or around cancer or neoplasia may shed more DNA than cells of the same tissue type in healthy subjects. The DNA from such cells may be epigenetically different from the shed DNA of healthy subjects. Therefore, the distribution of epigenetically modified (e.g., methylated) DNA in a certain DNA sample, such as cell-free DNA (cfDNA), may change during carcinogenesis. Thus, sufficiently sensitive epigenetic (e.g., DNA methylation) profiling can be used to detect abnormal methylation in the DNA of a sample.
[0005] Biopsy represents a traditional approach to detecting or diagnosing cancer in which cells or tissue are extracted from a potential cancer site and analyzed for relevant phenotypic and / or genotypic traits. Biopsies have the disadvantage of being invasive.
[0006] Cancer detection based on the analysis of bodily fluids, such as blood ("liquid biopsy"), is an interesting alternative based on the observation that DNA from cancer cells is released into bodily fluids. Liquid biopsies are non-invasive (sometimes requiring only a blood draw). However, due to the low concentration and heterogeneity of cell-free DNA, developing accurate and sensitive methods for analyzing liquid biopsy material that provide detailed information about nucleic acid base modifications has been a challenge. The contribution of DNA from cells in or around the cancer or neoplasm to a sample can be relatively small compared to the contribution from other cells, and the DNA contributed from other cells can be useless regarding the cancer state. Isolating and processing a useful fraction of cell-free DNA for further analysis in liquid biopsy procedures is an important part of these methods.
[0007] Furthermore, current methods for cancer diagnostic assays of cell-free nucleic acids (e.g., cell-free DNA or cell-free RNA) may focus on detecting tumor-associated somatic variants, including single nucleotide variations (SNVs), copy number variations (CNVs), fusions, and indels (i.e., insertions or deletions), all of which are mainstream targets for liquid biopsies. There is growing evidence that non-sequence modifications, such as methylation status and fragmentomic signals in cell-free DNA, can provide information about the origin and disease level of cell-free DNA. In addition, different types of modifications, such as 5-methylation and 5-hydroxymethylation, can have different implications regarding the presence or absence of disease. Detailed knowledge of non-sequence modifications in cell-free DNA (e.g., when combined with somatic mutation calling) can improve the assessment of tumor status. Summary of the Invention [Problem to be solved by the invention]
[0008] Thus, there is a continuing need for improved methods and compositions for analyzing DNA, including cell-free DNA, for example, in liquid biopsies. [Means for solving the problem]
[0009] The present disclosure aims to meet the need for improved analysis of DNA, such as cell-free DNA, and / or provide other advantages. In some embodiments, the present disclosure provides an integrated library preparation method for detecting both specific changes in target regions and broad signatures in the same workflow. In some embodiments, the present disclosure provides a method for adjusting epigenetic target regions and / or sequence variable target regions to improve nucleic acid-based analysis, such as cancer screening methods, and / or reduce assay costs.
[0010] The following exemplary embodiments are provided:
[0011] Embodiment 1 is a method for analyzing DNA, comprising: a) dividing the DNA into a plurality of sub-samples, the plurality of sub-samples comprising a first sub-sample and a second sub-sample; b) enriching for one or more sets of target regions of DNA from the first sub-sample, thereby providing enriched DNA of the first sub-sample, wherein the one or more sets of target regions comprise a set of sequence variable target regions. c) combining the enriched DNA of the first sub-sample and the DNA of the second sub-sample, thereby providing a combined sub-sample, wherein the DNA of the second sub-sample is not enriched for one or more sets of target regions of DNA; and d) sequencing the DNA of the combined sub-samples The method includes:
[0012] Embodiment 1.1 is a method for analyzing DNA, comprising: a) dividing the DNA into a plurality of sub-samples, the plurality of sub-samples comprising a first sub-sample and a second sub-sample; b) enriching for one or more sets of target regions of DNA from the first sub-sample, thereby providing enriched DNA of the first sub-sample, wherein the one or more sets of target regions comprise a set of sequence variable target regions; and c) sequencing the enriched DNA of the first sub-sample and the DNA of the second sub-sample. The method includes:
[0013] Embodiment 2 is a method for analyzing DNA, comprising: a) dividing the DNA into a plurality of sub-samples, the plurality of sub-samples comprising a first sub-sample and a second sub-sample; b) enriching for one or more sets of target regions of DNA from the first sub-sample, thereby providing enriched DNA of the first sub-sample, wherein the one or more sets of target regions comprise a set of epigenetic target regions; c) combining the enriched DNA of the first sub-sample and the DNA of the second sub-sample, thereby providing a combined sub-sample, wherein the DNA of the second sub-sample is not enriched for one or more sets of target regions of DNA; and d) sequencing the DNA of the combined sub-samples The method includes:
[0014] Embodiment 2.1 is a method for analyzing DNA, comprising: a) dividing the DNA into a plurality of sub-samples, the plurality of sub-samples comprising a first sub-sample and a second sub-sample; b) enriching for one or more sets of target regions of DNA from the first sub-sample, thereby providing enriched DNA of the first sub-sample, wherein the one or more sets of target regions comprise a set of epigenetic target regions; and c) sequencing the enriched DNA of the first sub-sample and the DNA of the second sub-sample. The method includes:
[0015] Embodiment 3 is the method of embodiment 1 or 1.1, wherein the one or more sets of target regions further comprise one or more sets of epigenetic target regions.
[0016] Embodiment 4 is the method of embodiment 2 or 2.2, wherein the one or more sets of target regions further comprise one or more sets of sequence variable target regions.
[0017] Embodiment 5 is the method of any one of the preceding embodiments, further comprising the step of subjecting the DNA or sub-sample thereof to a procedure that affects a first nucleobase of the DNA differently from a second nucleobase of the DNA, wherein the first nucleobase is a modified or unmodified nucleobase and the second nucleobase is a modified or unmodified nucleobase different from the first nucleobase, and wherein the first nucleobase and the second nucleobase have the same base-pairing specificity; optionally, wherein the step of subjecting the DNA or sub-sample thereof to a procedure that affects the first nucleobase of the DNA differently from the second nucleobase of the DNA is performed before dividing the DNA into multiple sub-samples.
[0018] Embodiment 6 is a method for analyzing DNA, comprising: a) subjecting the DNA to a procedure that affects a first nucleobase of the DNA differently from a second nucleobase of the DNA, wherein the first nucleobase is a modified or unmodified nucleobase, the second nucleobase is a modified or unmodified nucleobase that is different from the first nucleobase, and the first nucleobase and the second nucleobase have the same base-pairing specificity; b) dividing the DNA into a plurality of sub-samples, the plurality of sub-samples comprising a first sub-sample and a second sub-sample; c) enriching for one or more sets of target regions of DNA from the first sub-sample, thereby providing enriched DNA of the first sub-sample, wherein the one or more sets of target regions comprise one or more of a set of sequence variable target regions and a set of epigenetic target regions; d) combining the enriched DNA of the first sub-sample and the DNA of the second sub-sample, thereby providing a combined sub-sample, wherein the DNA of the second sub-sample is not enriched for one or more sets of target regions of DNA; and e) sequencing the DNA of the combined sub-samples The method includes:
[0019] Embodiment 6.1 is a method for analyzing DNA, comprising: a) subjecting the DNA to a procedure that affects a first nucleobase of the DNA differently from a second nucleobase of the DNA, wherein the first nucleobase is a modified or unmodified nucleobase, the second nucleobase is a modified or unmodified nucleobase that is different from the first nucleobase, and the first nucleobase and the second nucleobase have the same base-pairing specificity; b) dividing the DNA into a plurality of sub-samples, the plurality of sub-samples comprising a first sub-sample and a second sub-sample; c) enriching for one or more sets of target regions of DNA from the first sub-sample, thereby providing enriched DNA of the first sub-sample, wherein the one or more sets of target regions comprise one or more of a set of sequence variable target regions and a set of epigenetic target regions; and d) sequencing the enriched DNA of the first sub-sample and the DNA of the second sub-sample. The method includes:
[0020] Embodiment 7 is the method of any one of embodiments 1 to 6.1, wherein dividing the DNA into a plurality of sub-samples comprises distributing the DNA into a plurality of sub-samples such that a first sub-sample contains a higher proportion of DNA with cytosine modifications than a second sub-sample, or wherein at least one of the first sub-sample and the second sub-sample is distributed into a plurality of further sub-samples comprising at least a first further sub-sample and a second further sub-sample such that the first further sub-sample contains a higher proportion of DNA with cytosine modifications than the second further sub-sample.
[0021] Embodiment 8 is a method for distributing the sample, the method comprising the steps of: a) prior to enriching the DNA for one or more sets of epigenetic and / or sequence variable target regions of the DNA; b) after enriching the DNA for one or more sets of epigenetic and / or sequence variable target regions of DNA; c) before subjecting the DNA of the first sub-sample to a procedure that affects a first nucleobase of the DNA differently from a second nucleobase of the DNA; d) after subjecting the DNA of the first sub-sample to a procedure that affects a first nucleobase of the DNA differently from a second nucleobase of the DNA; or e) Any combination of (a) to (d) 8. The method of embodiment 7, wherein the method is carried out by
[0022] Embodiment 9 is the method of any one of embodiments 1 to 8, further comprising contacting at least one sub-sample with at least one restriction enzyme prior to enrichment or sequencing, optionally wherein the contacting occurs before performing a procedure that affects a first nucleobase of the DNA differently from a second nucleobase of the DNA, and / or optionally contacting the first sub-sample with at least one restriction enzyme.
[0023] Embodiment 10 is a method for analyzing DNA, comprising: a) dividing the DNA into a plurality of sub-samples, the plurality of sub-samples comprising a first sub-sample and a second sub-sample; b) subjecting DNA of at least a first sub-sample to a procedure that affects a first nucleobase of the DNA differently from a second nucleobase of the DNA, thereby providing converted DNA of the first sub-sample, wherein the first nucleobase is a modified or unmodified nucleobase, and the second nucleobase is a modified or unmodified nucleobase that is different from the first nucleobase, and wherein the first nucleobase and the second nucleobase have the same base-pairing specificity; c) combining the converted DNA of the first sub-sample and the DNA of the second sub-sample, thereby providing a combined sub-sample; and d) sequencing the DNA of the combined sub-samples The method includes:
[0024] Embodiment 11 is a method for analyzing DNA, comprising: a) dividing the DNA into a plurality of sub-samples, the plurality of sub-samples comprising a first sub-sample and a second sub-sample; b) contacting at least a first sub-sample with at least one restriction enzyme, thereby providing digested DNA of the first sub-sample; c) combining the digested DNA of the first sub-sample and the DNA of the second sub-sample, thereby providing a combined sub-sample; and d) sequencing the DNA of the combined sub-samples The method includes:
[0025] Embodiment 11.1 is a method for analyzing DNA, comprising: a) dividing the DNA into a plurality of sub-samples, the plurality of sub-samples comprising a first sub-sample and a second sub-sample; b) contacting at least a first sub-sample with at least one restriction enzyme, thereby providing digested DNA of the first sub-sample; and c) sequencing the digested DNA of the first sub-sample and the DNA of the second sub-sample. The method includes:
[0026] Embodiment 12 is the method of any one of embodiments 10 to 11.1, further comprising contacting the second sub-sample with at least one restriction enzyme prior to combining, thereby providing digested DNA of the second sub-sample.
[0027] Embodiment 12.1 is the method of any one of embodiments 10 to 11.1, wherein the second sub-sample is not contacted with at least one restriction enzyme.
[0028] Embodiment 13 is the method of any one of embodiments 10 to 11.1, further comprising, prior to combining, subjecting the DNA of the second sub-sample to a procedure that affects a first nucleobase of the DNA differently from a second nucleobase of the DNA, thereby providing converted DNA of the second sub-sample, wherein the first nucleobase is a modified or unmodified nucleobase and the second nucleobase is a modified or unmodified nucleobase that is different from the first nucleobase, and wherein the first nucleobase and the second nucleobase have the same base-pairing specificity.
[0029] Embodiment 13.1 is the method of any one of embodiments 10 to 11.1, wherein the second sub-sample is not subjected to a procedure that affects the first nucleobase of the DNA differently from the second nucleobase of the DNA.
[0030] Embodiment 14 is the method of any one of embodiments 10 to 13.1, wherein prior to combining, the first sub-sample is contacted with at least one methylation-sensitive restriction enzyme, thereby generating hypermethylated DNA of the first sub-sample.
[0031] Embodiment 15 is the method of any one of embodiments 10 to 14, wherein the second sub-sample is contacted with at least one methylation-dependent restriction enzyme, thereby generating hypomethylated DNA of the second sub-sample.
[0032] Embodiment 16 is a method for analyzing DNA, comprising: a) distributing at least a portion of the DNA into a plurality of sub-samples, including a first sub-sample and a second sub-sample, wherein the first sub-sample contains a higher proportion of DNA with cytosine modifications than the second sub-sample; b) subjecting DNA of at least a first sub-sample to a procedure that affects a first nucleobase of the DNA differently from a second nucleobase of the DNA, thereby providing converted DNA of the first sub-sample, wherein the first nucleobase is a modified or unmodified nucleobase, and the second nucleobase is a modified or unmodified nucleobase that is different from the first nucleobase, and wherein the first nucleobase and the second nucleobase have the same base-pairing specificity; c) combining the converted DNA of the first sub-sample and the DNA of the second sub-sample, thereby providing a combined sub-sample; and d) sequencing the DNA of the combined sub-samples The method includes:
[0033] Embodiment 16.1 is a method of analyzing DNA, comprising: a) distributing at least a portion of the DNA into a plurality of sub-samples, including a first sub-sample and a second sub-sample, wherein the first sub-sample contains a higher proportion of DNA with cytosine modifications than the second sub-sample; b) subjecting DNA of at least a first sub-sample to a procedure that affects a first nucleobase of the DNA differently from a second nucleobase of the DNA, thereby providing converted DNA of the first sub-sample, wherein the first nucleobase is a modified or unmodified nucleobase, the second nucleobase is a modified or unmodified nucleobase that is different from the first nucleobase, and the first nucleobase and the second nucleobase have the same base-pairing specificity; and c) sequencing the converted DNA of the first sub-sample and the DNA of the second sub-sample. The method includes:
[0034] Embodiment 17 is a method for analyzing DNA, comprising: a) distributing at least a portion of the DNA into a plurality of sub-samples, including a first sub-sample and a second sub-sample, wherein the first sub-sample contains a higher proportion of DNA with cytosine modifications than the second sub-sample; b) contacting at least a first sub-sample with at least one restriction enzyme, thereby providing digested DNA of the first sub-sample; c) combining the digested DNA of the first sub-sample and the DNA of the second sub-sample, thereby providing a combined sub-sample; and d) sequencing the DNA of the combined sub-samples The method includes:
[0035] Embodiment 17.1 is a method for analyzing DNA, comprising: a) distributing at least a portion of the DNA into a plurality of sub-samples, including a first sub-sample and a second sub-sample, wherein the first sub-sample contains a higher proportion of DNA with cytosine modifications than the second sub-sample; b) contacting at least a first sub-sample with at least one restriction enzyme, thereby providing digested DNA of the first sub-sample; and c) sequencing the digested DNA of the first sub-sample and the DNA of the second sub-sample. The method includes:
[0036] Embodiment 18 is the method of any one of embodiments 16 to 17.1, further comprising contacting the second sub-sample with at least one nuclease prior to combining, thereby providing digested DNA of the second sub-sample; optionally, the nuclease is a methylation-dependent nuclease; and further optionally, the methylation-dependent nuclease is a methylation-dependent restriction enzyme.
[0037] Embodiment 19 is the method of any one of embodiments 16 to 17.1, further comprising, prior to combining, subjecting the DNA of the second sub-sample to a procedure that affects a first nucleobase of the DNA differently from a second nucleobase of the DNA, thereby providing converted DNA of the second sub-sample, wherein the first nucleobase is a modified or unmodified nucleobase and the second nucleobase is a modified or unmodified nucleobase that is different from the first nucleobase, and wherein the first nucleobase and the second nucleobase have the same base-pairing specificity.
[0038] Embodiment 20 is the method of any one of embodiments 16 to 19, wherein prior to combining, the first sub-sample is contacted with at least one methylation-sensitive restriction enzyme, thereby generating hypermethylated DNA of the first sub-sample.
[0039] Embodiment 21 is the method of any one of embodiments 16 to 19, wherein the second sub-sample is contacted with at least one methylation-dependent restriction enzyme, thereby generating hypomethylated DNA of the second sub-sample. a) Embodiment 22 is the method of any one of the preceding embodiments, further comprising amplifying the enriched DNA of the first sub-sample before combining the enriched DNA of the first sub-sample and the DNA of the second sub-sample.
[0040] Embodiment 22.1 is a) amplifying the DNA before dividing the DNA into multiple sub-samples; and / or b) amplifying the DNA prior to enriching for one or more sets of target regions of DNA from the first sub-sample. 3. The method of any one of the preceding embodiments, further comprising:
[0041] Embodiment 23 is the method of embodiment 22 or 22.1, wherein the amplifying step is performed (a) after subjecting the DNA to a procedure that affects a first nucleobase of the DNA differently from a second nucleobase of the DNA; (b) after contacting at least the first and / or second sub-sample with at least one restriction enzyme; (c) after enriching for one or more sets of target regions of the DNA; or (d) after any combination of (a)-(c).
[0042] Embodiment 24 is the method of embodiment 22, 22.1, or 23, wherein the amplifying step comprises one or more of polymerase chain reaction, linear amplification, rolling circle amplification, ligase chain reaction, strand displacement amplification, nucleic acid sequence-based amplification, and sequence-based self-sustained replication.
[0043] Embodiment 25 is the method according to any one of embodiments 22 to 24, wherein the amplification comprises thermocycling amplification.
[0044] Embodiment 26 is the method according to any one of embodiments 22 to 24, wherein the amplification comprises isothermal amplification.
[0045] Embodiment 27 is the method of any one of the preceding embodiments, wherein the DNA comprises a barcode.
[0046] Embodiment 27.1 is the method of any one of the preceding embodiments, wherein the barcodes are molecular barcodes, and wherein the molecular barcodes distinguish between different DNA molecules in the same sample.
[0047] Embodiment 27.2 is the method of embodiment 27.1, wherein the molecular barcode is added to the DNA molecule, and optionally the molecular barcode is added by ligation.
[0048] Embodiment 28 is the method of any one of the preceding embodiments, comprising ligating an adapter comprising a barcode to the DNA prior to sequencing.
[0049] Embodiment 29 is the method according to any one of embodiments 22 to 28, comprising ligating an adapter comprising a barcode to the DNA before amplification.
[0050] Embodiment 30 is the method of any one of embodiments 7-9 or 16-29, wherein the partitioning comprises partitioning based on methylation levels.
[0051] Embodiment 31 is the method of any one of embodiments 7-9 or 16-30, wherein the partitioning step comprises contacting the DNA with an agent that recognizes modified cytosines in DNA, wherein the first sub-sample comprises DNA with a higher proportion of modified cytosines than the second sub-sample.
[0052] Embodiment 32 is the method of embodiment 31, wherein the agent that recognizes modified nucleobases in DNA is a methyl-binding reagent.
[0053] Embodiment 33 is the method of embodiment 32, wherein the methyl-binding reagent is a methyl-binding domain (MBD) protein or antibody.
[0054] Embodiment 34 is the method of embodiment 32 or embodiment 33, wherein the methyl-binding reagent is specific for one or more methylated nucleotide bases, and optionally, the one or more methylated nucleotide bases are 5-methylcytosine.
[0055] Embodiment 35 is the method of any one of embodiments 32 to 34, wherein the methyl-binding reagent is immobilized on a solid support.
[0056] Embodiment 36 is the method of any one of embodiments 7 to 9 or 16 to 35, wherein the partitioning step comprises immunoprecipitation of methylated DNA.
[0057] Embodiment 37 is the method of any one of embodiments 7-9 or 16-36, wherein the partitioning step comprises partitioning based on binding to a protein, optionally wherein the protein is a methylated protein, an acetylated protein, an unmethylated protein, or an unacetylated protein; and / or optionally wherein the protein is a histone.
[0058] Embodiment 38 is the method of embodiment 37, wherein the distributing step comprises contacting the DNA with a binding reagent specific for the protein and immobilized on a solid support.
[0059] Embodiment 39 is the method of any one of embodiments 7-9 or 16-38, wherein a first dispensed sub-sample of the plurality of dispensed sub-samples is differentially tagged from a second dispensed sub-sample of the plurality of dispensed sub-samples.
[0060] Embodiment 40 is the method of any one of embodiments 5-10, 13-16.1, or 19-39, wherein the first nucleobase is an unmodified cytosine and the second nucleobase is a modified cytosine, and optionally the modified cytosine is 5-methylcytosine or 5-hydroxymethylcytosine.
[0061] Embodiment 41 is the method of any one of embodiments 5-10, 13-16.1, or 19-40, wherein the procedure affecting a first nucleobase of the DNA differently from a second nucleobase of the DNA chemically converts the first or second nucleobase such that the base-pairing specificity of the converted nucleobase is altered.
[0062] Embodiment 42 is the method of any one of embodiments 5 to 10, 13 to 16.1, or 19 to 41, wherein the procedure affecting the first nucleobase of the DNA differently from the second nucleobase of the DNA is a methylation-sensitive conversion.
[0063] Embodiment 43 is the method of embodiment 42, wherein the methylation-sensitive conversion is bisulfite conversion, oxidative bisulfite (Ox-BS) conversion, Tet-assisted bisulfite (TAB) conversion, APOBEC-linked epigenetic (ACE) conversion, or enzymatic conversion.
[0064] Embodiment 44 is the method of embodiment 43, wherein the Tet-assisted conversion further comprises a substituted borane reducing agent, optionally wherein the substituted borane reducing agent is 2-picoline borane, borane pyridine, tert-butylamine borane, or ammonia borane.
[0065] Embodiment 45 is the method of any one of embodiments 9, 11-15, or 17-44, wherein at least one restriction enzyme is a methylation-sensitive restriction enzyme (MSRE).
[0066] Embodiment 46 is the method of embodiment 45, wherein the first sub-sample is contacted with an MSRE.
[0067] Embodiment 47 is the method of any one of embodiments 9, 11-15, or 17-44, wherein at least one restriction enzyme is a methylation-dependent restriction enzyme (MDRE).
[0068] Embodiment 48 is the method of embodiment 47, wherein the second sub-sample is contacted with an MDRE.
[0069] Embodiment 49 is the method of any one of embodiments 9, 11-15, or 17-48, comprising contacting at least one sub-sample with at least two restriction enzymes prior to enrichment or sequencing, optionally wherein the contacting occurs before performing a procedure that affects a first nucleobase of the DNA differently from a second nucleobase of the DNA.
[0070] Embodiment 50 is the method of embodiment 49, wherein the at least two restriction enzymes comprise or consist of two or three restriction enzymes.
[0071] Embodiment 51 is the method of any one of embodiments 9, 11-15, or 17-50, wherein the at least one restriction enzyme or at least two restriction enzymes are selected from the group consisting of FspEI, LpnPI, MspJI, SgeI, AatII, AccII, AciI, Aor13HI, Aor15HI, BspT104I, BssHII, BstUI, Cfr10I, ClaI, CpoI, Eco52I, HaeII, HapII, HhaI, Hin6I, HpaII, HpyCH4IV, MluI, MspI, NaeI, NotI, NruI, NsbI, PmaCI, Pspl406I, PvuI, SacII, SalI, SmaI, and SnaBI.
[0072] Embodiment 52 is the method of any one of embodiments 9, 11-15, or 17-51, further comprising the step of attaching one or more adapters to at least one end of at least some of the DNA molecules in the plurality of distributed sets before contacting the at least one sub-sample with the at least one restriction enzyme or at least two restriction enzymes.
[0073] Embodiment 53 is the method of embodiment 52, wherein one or more adapters comprise at least one tag.
[0074] Embodiment 54 is the method of embodiment 53, wherein at least one tag comprises a molecular barcode.
[0075] Embodiment 55 is the method according to any one of embodiments 52 to 54, wherein one or more adapters are resistant to digestion by a methylation-sensitive or methylation-dependent restriction enzyme.
[0076] Embodiment 56 is a method for preparing a methylation-sensitive restriction enzyme-resistant nucleotide sequence comprising: a) one or more methylated nucleotides, optionally wherein the methylated nucleotides comprise 5-methylcytosine and / or 5-hydroxymethylcytosine; b) one or more nucleotide analogs that are resistant to methylation-sensitive restriction enzymes; or c) a nucleotide sequence that is not recognized by a methylation-sensitive restriction enzyme 56. The method of embodiment 55, comprising:
[0077] Embodiment 56.1 is a) each cytosine in the adaptor is methylated, and optionally the methylated cytosines include 5-methylcytosine and / or 5-hydroxymethylcytosine; or b) one or more adapters that are resistant to digestion by methylation-sensitive restriction enzymes consist of a nucleotide sequence that is not recognized by the methylation-sensitive restriction enzyme; 57. The method of embodiment 56.
[0078] Embodiment 57 is the method of any one of embodiments 10 to 56.1, further comprising enriching for one or more sets of target regions of DNA from at least a first sub-sample, thereby providing enriched DNA of the first sub-sample, wherein the one or more sets of target regions comprise one or more of a set of sequence variable target regions and a set of epigenetic target regions.
[0079] Embodiment 58 is the method of any one of embodiments 1 to 9 or 22 to 57, wherein the enriching step comprises contacting the DNA with target-specific probes specific for one or more sets of epigenetic target regions and / or one or more sets of sequence variable target regions.
[0080] Embodiment 59 is the method of any one of embodiments 2 to 9 or 22 to 58, wherein the set of epigenetic target regions comprises a set of hypermethylated variable target regions and / or a set of hypomethylated variable target regions.
[0081] Embodiment 60 is the method of any one of embodiments 2 to 9 or 22 to 59, wherein the set of epigenetic target regions comprises a set of fragmented variable target regions.
[0082] Embodiment 61 is the method of embodiment 60, wherein the set of fragmented variable target regions comprises a transcription start site region.
[0083] Embodiment 62 is the method of embodiment 60 or embodiment 61, wherein the set of fragmented variable target regions comprises a CTCF binding region.
[0084] Embodiment 63 is the method of any one of embodiments 2 to 9 or 22 to 62, wherein the set of epigenetic target regions comprises one or more type-specific epigenetic target regions.
[0085] Embodiment 64 is the method of embodiment 63, wherein the one or more type-specific epigenetic target regions comprise type-specific differentially methylated regions and / or type-specific fragments.
[0086] Embodiment 65 is the method of embodiment 63, wherein the one or more type-specific epigenetic target regions comprise a type-specific hypomethylated region and / or a type-specific hypermethylated region.
[0087] Embodiment 66 is the method of any one of embodiments 63 to 65, wherein the one or more type-specific epigenetic target regions comprise cell type-specific, cell cluster type-specific, tissue type-specific, and / or cancer type-specific epigenetic target regions.
[0088] Embodiment 67 is a method for determining whether the one or more type-specific epigenetic target regions are: hypermethylated in immune cells compared to non-immune cell types present in blood samples; differentially methylated in the colon compared with other tissue types; differentially methylated in breast compared with other tissue types; differentially methylated in liver compared with other tissue types; differentially methylated in kidney compared with other tissue types; differentially methylated in the pancreas compared with other tissue types; differentially methylated in the prostate compared with other tissue types; is differentially methylated in skin compared to other tissue types; or Differentially methylated in bladder compared with other tissue types 67. The method of any one of embodiments 63 to 66, comprising a target region.
[0089] Embodiment 68 is the method of any one of embodiments 65 to 67, wherein the hypermethylated target region is methylated to an extent that is at least 10%, 20%, 30%, or at least 40% greater than the average methylation of the target region in the sample or compared to other cell or tissue types.
[0090] Embodiment 69 is the method of any one of embodiments 63 to 68, wherein the one or more type-specific epigenetic target regions comprise:
[0091] a target region that is hypomethylated in non-immune cell types present in the sample compared to the methylation level of the target region in different cell or tissue types in the sample;
[0092] a fragment that is specific for immune cells relative to non-immune cell types present in the sample; or
[0093] Fragments specific for colon, lung, breast, liver, kidney, pancreas, prostate, skin, or bladder compared to other tissue types.
[0094] Embodiment 70 is the method of any one of embodiments 63 to 69, wherein the level of one or more type-specific epigenetic target regions of cell type or tissue type origin is determined.
[0095] Embodiment 71 is the method of embodiments 63-70, wherein the level of one or more immune cells, non-immune cell types present in the blood sample, and / or one or more type-specific epigenetic target regions originating from colon, lung, breast, liver, kidney, prostate, skin, bladder, or pancreatic cells is determined.
[0096] Embodiment 72 is the method of any one of embodiments 63 to 71, further comprising identifying at least one cell type, cell cluster type, tissue type, and / or cancer type from which the one or more type-specific epigenetic target regions originated.
[0097] Embodiment 73 is the method of any one of embodiments 63 to 72, comprising determining the methylation level of the type-specific epigenetic target region.
[0098] Embodiment 74 is the method of any one of the preceding embodiments, wherein the DNA of the first sub-sample and the DNA of the second sub-sample are differentially tagged.
[0099] Embodiment 75 is the method of any one of the preceding embodiments, wherein the combined sub-sample comprises (a) at least a portion of the converted DNA of the first sub-sample or at least a portion of the enriched DNA of the first sub-sample, and (b) at least a portion of the DNA of the second sub-sample.
[0100] Embodiment 76 is the method of any one of the preceding embodiments, wherein the combined sub-sample further comprises at least a portion of the DNA of a third sub-sample.
[0101] Embodiment 77 is the method of any one of the preceding embodiments, wherein the combined sub-sample comprises less than or equal to about 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, 5%, or 1% of the DNA of the first sub-sample.
[0102] Embodiment 78 is the method of any one of the preceding embodiments, wherein the combined sub-sample comprises about 50-70%, 70-90%, about 75-85%, or about 80% of the DNA of the first sub-sample.
[0103] Embodiment 79 is the method of any one of the preceding embodiments, wherein the combined sub-sample comprises less than or equal to about 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, or 5%, or 1% of the DNA of the second sub-sample.
[0104] Embodiment 80 is the method of any one of the preceding embodiments, wherein the combined sub-sample comprises about 50-70%, 70-90%, about 75-85%, or about 80% of the DNA of the second sub-sample.
[0105] Embodiment 81 is the method of any one of embodiments 76 to 80, wherein the combined sub-sample comprises less than or equal to about 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, or 5%, or 1% of the DNA of the third sub-sample.
[0106] Embodiment 82 is the method of any one of embodiments 76 to 81, wherein the combined sub-sample comprises about 50-70%, 70-90%, about 75-85%, or about 80% of the DNA of the third sub-sample.
[0107] Embodiment 83 is the method of any one of the preceding embodiments, wherein the combined sub-sample comprises substantially all of the DNA of the first sub-sample.
[0108] Embodiment 84 is the method of any one of the preceding embodiments, wherein the combined sub-sample comprises substantially all of the DNA of the second sub-sample.
[0109] Embodiment 85 is the method of any one of embodiments 76 to 84, wherein the combined sub-sample comprises substantially all of the DNA of the third sub-sample.
[0110] Embodiment 86 is the method of any one of the preceding embodiments, wherein sequencing the DNA of the combined sub-samples comprises sequencing the DNA in a manner that distinguishes the first nucleobase from the second nucleobase.
[0111] Embodiment 87 is the method of any one of the preceding embodiments, wherein sequencing the DNA of the combined sub-samples comprises sequencing at least a portion of the DNA of at least the first and second sub-samples in the same sequencing cell.
[0112] Embodiment 88 is the method of any one of the preceding embodiments, wherein sequencing the DNA of the combined sub-samples comprises sequencing the DNA in a manner that is sensitive to the modification.
[0113] Embodiment 89 is the method of embodiment 88, wherein the step of sequencing in a manner sensitive to the modification comprises long-read sequencing.
[0114] Embodiment 90 is the method of embodiment 88 or embodiment 89, wherein the step of sequencing in a manner sensitive to the modification comprises nanopore sequencing.
[0115] Embodiment 91 is the method of embodiment 88, wherein the step of sequencing in a manner sensitive to modification comprises five-letter or six-letter sequencing.
[0116] Embodiment 92 is the method according to any one of the preceding embodiments, wherein the step of sequencing the DNA of the combined sub-samples comprises next generation sequencing.
[0117] Embodiment 93 is the method of any one of the preceding embodiments, wherein sequencing the DNA of the combined sub-samples comprises generating a plurality of sequencing reads, and the method further comprises mapping the plurality of sequence reads to one or more reference sequences to generate mapped sequence reads, and processing the mapped sequence reads corresponding to the set of sequence variable target regions and the set of epigenetic target regions.
[0118] Embodiment 94 is a method according to any one of the preceding embodiments, comprising detecting:
[0119] at least one epigenetic state in the DNA of the first subsample; and
[0120] At least one fragmentation feature and / or at least one somatic variant in the DNA of the second subsample.
[0121] Embodiment 95 is the method of the immediately preceding embodiment, wherein at least one epigenetic state is the methylation state of a region or nucleotide, and optionally the region is a hypermethylated or hypomethylated variable target region, or the nucleotide is in a hypermethylated or hypomethylated variable target region.
[0122] Embodiment 96 is the method of embodiment 95 or 96, wherein the at least one somatic variant comprises a single nucleotide variation (SNV), a copy number variation (CNV), a gene fusion, or an insertion or deletion (indel).
[0123] Embodiment 97 is the method of any one of the preceding embodiments, wherein the DNA is from a blood sample and / or a tissue sample.
[0124] Embodiment 98 is the method of embodiment 97, wherein the blood sample is a whole blood sample, a plasma sample, a buffy coat sample, a leukoreduced sample, or a PBMC sample.
[0125] Embodiment 99 is the method of any one of the preceding embodiments, wherein the DNA is cell-free DNA.
[0126] Embodiment 100 is the method of embodiment 99, wherein the cell-free DNA is in an amount of 1 ng to 500 ng.
[0127] Embodiment 101 is the method of any one of the preceding embodiments, wherein the DNA and / or sample is from a subject.
[0128] Embodiment 102 is the method of embodiment 101, wherein the subject is an animal.
[0129] Embodiment 103 is the method of embodiment 101 or embodiment 102, wherein the subject is a human.
[0130] Embodiment 104 is the method of any one of embodiments 97 to 103, wherein the blood sample is fractionated before enrichment for at least one set of epigenetic target regions of DNA.
[0131] Embodiment 105 is the method of any one of embodiments 101 to 104, wherein the subject has or is at risk of having cancer.
[0132] Embodiment 106 is the method of any one of embodiments 101 to 105, further comprising determining the presence or status of cancer in the subject.
[0133] Embodiment 107 is the method of any one of embodiments 101 to 106, further comprising determining the likelihood that the subject has an infection.
[0134] Embodiment 108 is the method of any one of embodiments 101 to 107, further comprising determining the likelihood that the subject will have transplant rejection.
[0135] Embodiment 109 is the method according to any one of the preceding embodiments, wherein the DNA of the first and second sub-samples is not combined prior to sequencing.
[0136] Embodiment 110 is the method according to any one of the preceding embodiments, wherein the step of sequencing the second sub-sample comprises whole genome sequencing.
[0137] Embodiment 111 is the method of embodiment 110, further comprising the step of determining a hypomethylation score for each sample after whole genome sequencing.
[0138] Embodiment 112 is a method for determining a hypomethylation score for each sample, comprising: a) determining a signal ratio for the second sub-sample for each of a plurality of partially methylated domains (PMDs), wherein for each of the plurality of PMDs, determining the signal ratio for the PMD comprises dividing the number of fully unmethylated molecules identified within the PMD by the total number of molecules identified within the PMD; and b) for each of the plurality of selected PMDs, comparing the signal ratio for the second sub-sample to a predetermined threshold to identify a PMD comprising reduced methylation in the second sub-sample compared to the same PMD in the reference standard; and c) determining a proportion of positive PMDs of the plurality of PMDs for the second sub-sample, thereby determining a sample-by-sample hypomethylation score for the second sub-sample. 112. The method of embodiment 111, comprising:
[0139] Embodiment 113 is the method of embodiment 112, wherein each of the plurality of molecules identified within the PMD comprises at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 CpGs.
[0140] Embodiment 114 is the method of embodiment 112, wherein each of the plurality of molecules identified within the PMD comprises at least 3, 4, 5, or 6 CpGs.
[0141] Embodiment 115 is the method of embodiment 112, wherein each of the plurality of molecules identified in the PMD comprises at least three CpGs.
[0142] Embodiment 116 is the method of any one of embodiments 112 to 115, wherein each of the PMDs is 5 to 1000, 5 to 900, 5 to 800, 5 to 700, 5 to 600, 5 to 500, 5 to 400, 5 to 300, 5 to 200, 5 to 100, 50 to 500, 50 to 400, 50 to 300, 50 to 200, or 50 to 150 kilobases in length.
[0143] Embodiment 117 is the method of any one of embodiments 112 to 115, wherein each of the PMDs is 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000 or 10000 kilobases in length.
[0144] Embodiment 118 is the method of any one of embodiments 112 to 117, wherein the plurality of PMDs comprises 10 to 500, 20 to 500, 50 to 500, 50 to 400, 50 to 300, 50 to 200, 50 to 150, or 75 to 125 PMDs.
[0145] Embodiment 119 is a method described in any one of embodiments 112 to 117, wherein the plurality of PMDs includes 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 125, 150, 175, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, or 1000 PMDs.
[0146] Embodiment 120 is a method according to any one of embodiments 112 to 119, wherein the predetermined threshold comprises a signal ratio of at least 0.5 to 10%, at least 0.5 to 5%, at least 1 to 10%, or at least 1 to 5%.
[0147] Embodiment 121 is a method described in any one of embodiments 112 to 119, wherein the predetermined threshold comprises a signal ratio of at least 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, or at least 10%.
[0148] Embodiment 122 is the method of any one of embodiments 112 to 119, wherein the predetermined threshold comprises a signal ratio of at least 1%.
[0149] Embodiment 123 is the method according to any one of embodiments 113 to 119, wherein each of the plurality of molecules identified in the PMD comprises at least 3 CpGs, the predetermined threshold being at least 1%.
[0150] Embodiment 124 is the method according to any one of embodiments 113 to 119, wherein each of the plurality of molecules identified in the PMD comprises at least 4 CpGs, and the predetermined threshold is at least 1%.
[0151] Embodiment 125 is a method according to any one of embodiments 113 to 119, wherein each of the plurality of molecules identified within the PMD comprises at least 5 CpGs, and the predetermined threshold is at least 1%.
[0152] Embodiment 126 is the method according to any one of embodiments 113 to 119, wherein each of the plurality of molecules identified within the PMD comprises at least 6 CpGs, and the predetermined threshold is at least 1%.
[0153] Embodiment 127 is the method according to any one of embodiments 113 to 119, wherein each of the plurality of molecules identified in the PMD comprises at least one CpG, the predetermined threshold being at least 5%. [Brief explanation of the drawings]
[0154] [Figure 1A] FIG. 1A illustrates an exemplary workflow according to certain embodiments disclosed herein.
[0155] [Figure 1B] FIG. 1B illustrates an exemplary workflow according to certain embodiments disclosed herein.
[0156] [Figure 2] FIG. 2 is a schematic diagram of an example system suitable for use with some embodiments of the present disclosure.
[0157] [Figure 3] Figure 3 is a box plot of sample-specific hypomethylation (PSH) scores for cancer-free, AA, and CRC samples. A score of 1 on the y-axis of the plot indicates that at least 1% of the molecules in each partially methylated domain (PMD) were unmethylated.
[0158] [Figure 4] Figure 4 shows the overlay of the hypomethylation scores per sample for cancer-free, AA, and CRC samples, and either the subject hypermethylation or subject hypomethylation metrics measured for the same samples. A score of 1 on the y-axis of these plots indicates that all 100 PMDs were altered.
[0159] [Figure 5-1] Figure 5 shows receiver operating characteristic (ROC) curves for five models: (a) per-sample hypomethylation score ("PSH"), (b) subject hypermethylation metric ("High"), (c) subject hypomethylation metric ("Low"), (d) combined per-sample hypomethylation score and subject hypermethylation metric ("PSH+High"), and (e) combined per-sample hypomethylation score and subject hypomethylation metric ("PSH+Low") for AA and CRC samples. Area under the curve (AUC) values are shown for each model as the numbers to the right of the model name. [Figure 5-2] Same as above. DETAILED DESCRIPTION OF THE INVENTION
[0160] Detailed Description of Certain Embodiments Reference will now be made in detail to certain specific embodiments of the invention. While the invention will be described in conjunction with such embodiments, it will be understood that they are not intended to limit the invention to those embodiments. On the contrary, the invention is intended to cover all alternatives, modifications, and equivalents, which may be included within the scope of the present invention as defined by the appended claims.
[0161] Before describing the teachings of the present invention in detail, it is understood that the disclosure is not limited to specific compositions or process steps, as these may vary. It should be noted that as used in this specification and the appended claims, the singular forms "a," "an," and "the" include the plural forms unless the context clearly dictates otherwise. Thus, for example, a reference to "a nucleic acid" includes a plurality of nucleic acids, a reference to "a cell" includes a plurality of cells, etc.
[0162] Numerical ranges are inclusive of the numbers defining the range. Measured and measurable values are understood to be approximations, taking into account significant digits and errors associated with measurement. Similarly, the use of "comprise," "comprises," "comprising," "contain," "contains," "containing," "include," "includes," and "including" is not intended to be limiting. It is understood that both the foregoing general and detailed descriptions are exemplary and explanatory only and do not limit the present teachings.
[0163] Unless specifically noted in the specification above, embodiments herein that recite various components as "comprising" are also contemplated as "consisting of" or "consisting essentially of" the recited components, and embodiments herein that recite various components as "consisting essentially of" are also contemplated as "comprising" or "consisting essentially of" the recited components (this interchangeability does not apply to the use of these terms in the claims).
[0164] The section headings used herein are for organizational purposes only and are not to be construed as limiting the disclosed subject matter in any way. In the event that any document or other material incorporated by reference contradicts the express contents of this specification, including definitions, the present specification will control. I. Definition
[0165] "Buffy coat" refers to the portion of a blood (e.g., whole blood) or bone marrow sample that contains all or most of the sample's white blood cells and platelets. A buffy coat fraction of a sample can be prepared from the sample using centrifugation, which separates sample components by density. For example, after centrifugation of a whole blood sample, the buffy coat fraction is located between the plasma and erythrocyte (red blood cell) layers. Buffy coats can contain both mononuclear (e.g., T cells, B cells, NK cells, dendritic cells, and monocytes) and polymorphonuclear (e.g., granulocytes such as neutrophils and eosinophils) white blood cells.
[0166] "Cell-free DNA," "cfDNA molecules," or simply "cfDNA" includes DNA molecules naturally present in a subject in an extracellular form (e.g., in blood, serum, plasma, or other bodily fluids, such as lymph, cerebrospinal fluid, urine, or sputum). cfDNA was previously present in one or more cells of a large, complex biological organism, e.g., a mammal, but has been released from the cell(s) into fluids found in the organism, and can be obtained from a sample of the fluid without the need to perform an in vitro cell lysis step. cfDNA molecules may exist as DNA fragments.
[0167] As used herein, "fragment" refers to a biological component, such as a nucleic acid molecule (e.g., DNA or RNA), that is broken or separated from one or more other pieces.Fragmentation, such as DNA fragmentation, can occur spontaneously (as in the cfDNA fragments that can be obtained from blood samples), or can be intentionally induced, for example, using standard laboratory procedures as described herein.DNA fragmentation can be carried out, for example, to prepare DNA (e.g., genomic DNA and / or DNA isolated from samples containing cells) for sequencing.In some samples, such as cfDNA samples, artificial fragmentation may be unnecessary.
[0168] As used herein, "fragmentation characteristics" refers to any trait related to the endpoints, midpoints, size, presence, absence, and / or quantity of DNA fragments isolated from a subject, e.g., DNA fragments having midpoints or one or both endpoints at a particular genomic location or within a particular range of locations, and / or having a particular value or range of lengths.
[0169] As used herein, a modification or other feature is "present in a higher proportion" in a first sample or population of nucleic acids than in a second sample or population if the proportion of nucleotides having the modification or other feature is higher in the first sample or population than in the second population. For example, if one in ten nucleotides in a first sample are mC and one in twenty nucleotides in a second sample are mC, then the first sample contains a higher proportion of 5-methylated cytosine modifications than the second sample.
[0170] As used herein, "leukapheresis" refers to a procedure in which white blood cells (leukocytes) are isolated from a sample of blood collected from a subject. Leukapheresis can be performed, for example, to obtain cells, such as those described herein, for research, diagnostic, prognostic, or monitoring purposes. Thus, as used herein, a "leukapheresis sample" refers to a sample containing white blood cells collected from a subject using leukapheresis.
[0171] As used herein, " peripheral blood mononuclear cells " or " PBMC " refers to the immune cells that originate from bone marrow and have a single round nucleus, and are found in peripheral circulation. Such cells include, for example, lymphocytes (T cells, B cells, and NK cells) and monocytes, and are isolated from blood samples (for example, from the whole blood sample collected from a subject) using density gradient centrifugation.
[0172] As used herein, "without substantially changing the base-pairing specificity" of a given nucleobase means that the majority of molecules that comprise the nucleobase that can be sequenced have no change in the base-pairing specificity of the given nucleobase compared to its base-pairing specificity when present in the original isolated sample.In some embodiments, 75%, 90%, 95% or 99% of molecules that comprise the nucleobase that can be sequenced have no change in the base-pairing specificity when present in the original isolated sample.As used herein, "change in base-pairing specificity" of a given nucleobase means that the majority of molecules that comprise the nucleobase that can be sequenced have the base-pairing specificity of the nucleobase compared to its base-pairing specificity in the original isolated sample.
[0173] As used herein, "base-pairing specificity" refers to the standard DNA base (A, C, G, or T) with which a given base most preferentially pairs. For example, unmodified cytosine and 5-methylcytosine have the same base-pairing specificity (i.e., specificity for G), but uracil and cytosine have different base-pairing specificities, with uracil having base-pairing specificity for A, whereas cytosine has base-pairing specificity for G. The ability of uracil to form a wobble base pair with G is irrelevant, since uracil nevertheless pairs most preferentially with A among the four standard DNA bases.
[0174] The "capture yield" of a collection of probes for a given target set refers to the amount of nucleic acid corresponding to the target set that the collection of probes captures under typical conditions (e.g., the amount or absolute amount relative to another target set). Exemplary typical capture conditions are incubating sample nucleic acid and probes at 65°C for 10-18 hours in a small reaction volume (approximately 20 μL) containing a stringent hybridization buffer. Capture yields can be expressed in absolute terms, or, in the case of a collection of multiple probes, in relative terms. When comparing capture yields for a set of multiple target regions, they are normalized with respect to the footprint size (e.g., on a per kilobase basis) of the target region set. Thus, for example, if the footprint sizes of the first and second target regions are 50 kb and 500 kb, respectively (using a normalization factor of 0.1), and the mass / volume concentration of the captured DNA corresponding to the first set of target regions is greater than 0.1 times the mass / volume concentration of the captured DNA corresponding to the second set of target regions, the DNA corresponding to the first set of target regions will be captured with a higher yield than the DNA corresponding to the second set of target regions. As a further example, using the same footprint size, if the captured DNA corresponding to the first set of target regions has a mass / volume concentration that is 0.2 times the mass / volume concentration of the captured DNA corresponding to the second set of target regions, the DNA corresponding to the first set of target regions will be captured with a capture yield that is 2 times higher than the DNA corresponding to the second set of target regions.
[0175] "Enriching" or "capturing" one or more target nucleic acids or one or more nucleic acids comprising at least one target region refers to preferentially isolating or separating one or more target nucleic acids or one or more nucleic acids comprising at least one target region from non-target nucleic acids or from nucleic acids that do not contain at least one target region.
[0176] An "enriched set" or "captured set" of nucleic acids, or an "enriched" or "captured" nucleic acid, refers to nucleic acids that have undergone capture.
[0177] As used herein, a "capture moiety" is a molecule that allows for affinity separation of a molecule, such as a nucleic acid, linked to the capture moiety from molecules that lack the capture moiety. Exemplary capture moieties include biotin, which allows for affinity separation by binding to streptavidin that is or can be linked to a solid phase, or oligonucleotides that allow for affinity separation through binding to complementary oligonucleotides that are or can be linked to a solid phase.
[0178] As used herein, a "cell type" is a set of cells that share common characteristics. For example, a cell type may include cells of different origins, differentiation types, activation types, or any combination of different origins, differentiation types, and activation types. In fact, the differentiation state and activation state may overlap and often change together in a given cell, such as an immune cell or cancer cell. For example, activation of an immune cell may induce differentiation of the cell. In some embodiments, cell types may be distinguished based on characteristics such as one or more cell surface markers, gene signatures (e.g., the expression (or expression level) of a particular gene or set of genes), and / or epigenetic signatures such as regions of DNA hypermethylation or hypomethylation.
[0179] As used herein, a "cell cluster" or "cluster" refers to a plurality of related cell types, e.g., immune cell types, tissue-specific cell types, and / or cancer cell types. In some embodiments, the cell types within a cluster have similar DNA methylation profiles, e.g., in multiple hypermethylated and / or hypomethylated variable target regions.
[0180] "Converted nucleobase" is a nucleobase that has a change in base pairing specificity, and the original base pairing specificity of the nucleobase has been changed by a procedure.For example, a certain procedure converts unmethylated or unmodified cytosine into dihydrouracil, or more generally, at least one modified or unmodified form of cytosine undergoes deamination, resulting in uracil (considered a modified nucleobase in the context of DNA) or a further modified form of uracil.As used herein, "converted sample" refers to a sample that contains DNA that contains at least one converted nucleobase.
[0181] As used herein, a "combination" of steps or other elements refers to the performance or presence of two or more steps or elements in a method or product; where appropriate, the elements may be together in a single composition, device, or the like, or may be in close proximity, e.g., in separate containers or compartments within a larger container, such as a multiwell plate, tube rack, refrigerator, freezer, incubator, water bath, ice bucket, machine, or other storage configuration. A combination, combinations, or combinations thereof refers to any and all permutations and combinations of the terms listed before the term "combination." For example, "A, B, C, or a combination thereof" is intended to include at least one of A, B, C, AB, AC, BC, or ABC, and also includes BA, CA, CB, ACB, CBA, BCA, BAC, or CAB if order is important in the particular context. Continuing with this example, combinations containing repeats of one or more items or terms are expressly included, e.g., BB, AAA, AAB, BBC, AAABCCCC, CBBAAA, CABABB, etc. Those skilled in the art will understand that there is typically no limit to the number of items or terms in any combination, unless otherwise clear from the context.
[0182] "Specifically binds" in the context of a primer, probe, or other oligonucleotide and a target sequence (e.g., a nucleic acid comprising a sequence partially or fully complementary to the primer, probe, or other oligonucleotide) means that, under appropriate hybridization conditions, the primer, probe, or other oligonucleotide hybridizes to the target sequence, or a copy thereof, to form a stable hybrid, while minimizing the formation of stable non-target hybrids. Thus, the primer, probe, or other oligonucleotide hybridizes to the target sequence, or a copy thereof, to a sufficiently greater extent than to non-target sequences, ultimately allowing enrichment or detection of the target sequence. Suitable hybridization conditions are well known in the art and can be predicted based on sequence composition or determined by using routine testing methods (see, e.g., §§ 1.90-1.91, 7.37-7.57, 9.47-9.51, and 11.47-11.57 of Sambrook et al., Molecular Cloning, A Laboratory Manual, 2nd ed. (Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1989), which is incorporated herein by reference, especially §§ 9.50-9.51, 11.12-11.13, 11.45-11.47, and 11.55-11.57).
[0183] A "target region" refers to a genomic locus that is targeted for identification and / or capture, e.g., by using a probe (e.g., through sequence complementarity). A "target region set" or "set of target regions" refers to multiple genomic loci that are targeted for identification and / or capture, e.g., by using a set of probes (e.g., through sequence complementarity). A "target region set" may include regions that share at least one common feature. In some embodiments, a target region set is identified by at least one common feature. For example, a hypermethylated variable target region set includes regions of DNA that are hypermethylated.
[0184] A "sequence variable target region" refers to a target region that may exhibit sequence changes, such as nucleotide substitutions (i.e., single-base mutations), insertions, deletions, or gene fusions or rearrangements, in neoplastic cells (e.g., tumor cells and cancer cells) compared to normal cells. A "set of sequence variable target regions" refers to a set of sequence variable target regions. In some embodiments, the sequence variable target regions are target regions that may exhibit changes affecting less than or equal to 50 consecutive nucleotides, e.g., less than or equal to 40, 30, 20, 10, 5, 4, 3, 2, or 1 nucleotide.
[0185] "Epigenetic target region" refers to a target region that may exhibit sequence-independent differences in different cell or tissue types (e.g., different types of immune cells) or abnormal cells, such as neoplastic cells (e.g., tumor cells and cancer cells), compared to normal cells, or in DNA, e.g., DNA from different cell types or from subjects with cancer, compared to DNA from healthy subjects, that may exhibit sequence-independent differences (i.e., differences in methylation, nucleosome distribution, or other epigenetic features, without changes to the nucleotide sequence). Examples of sequence-independent changes include, but are not limited to, changes in methylation (increase or decrease), nucleosome distribution, fragmentation patterns, CCCTC-binding factor ("CTCF") binding, transcription start sites (e.g., with respect to any one or more of the binding of RNA polymerase components, regulatory protein binding, fragmentation characteristics, and nucleosome distribution), and regulatory protein binding regions. Thus, epigenetic target region sets include, but are not limited to, hypermethylated variable target region sets, hypomethylated variable target region sets, and fragmented variable target region sets, such as CTCF binding sites and transcription start sites.For the purpose of the present invention, the loci that are prone to neoplasia, tumor, or cancer-related local amplification and / or gene fusion can also be included in epigenetic target region sets, because the detection of copy number changes by sequencing or fusion sequences that map to more than one locus in a reference genome tends to be more similar to the detection of the exemplary epigenetic changes discussed above than the detection of nucleotide substitutions, insertions, or deletions, for example, in that the detection of local amplification and / or gene fusion can be detected at a relatively low sequencing depth because it does not depend on the accuracy of base calls at one or a few individual positions." Epigenetic target region set " is a set of epigenetic target regions.
[0186] As used herein, a "differentially methylated region" refers to a region of DNA that has a detectably different degree of methylation in at least one cell or tissue type compared to the degree of methylation in the same region of DNA from at least one other cell or tissue type, or a region of DNA that has a detectably different degree of methylation in at least one cell or tissue type obtained from a subject with a disease or disorder compared to the degree of methylation in the same region of DNA in the same cell or tissue type obtained from a healthy subject. In some embodiments, a differentially methylated region has a detectably higher degree of methylation (e.g., a hypermethylated region) in at least one cell or tissue type, e.g., at least one immune cell type, compared to the degree of methylation in the same region of DNA from at least one other cell or tissue type, e.g., another immune cell type, or from the same cell or tissue type from a healthy subject. In some embodiments, a differentially methylated region has a detectably lower degree of methylation (e.g., a hypomethylated region) in at least one cell or tissue type, e.g., at least one immune cell type, compared to the degree of methylation in the same region of DNA from at least one other cell or tissue type, e.g., another immune cell type, or from the same cell or tissue type from a healthy subject.
[0187] A nucleic acid is "produced by a tumor" if it originates from a tumor cell. A tumor cell is a neoplastic cell that originates from a tumor, regardless of whether it remains within the tumor or leaves the tumor (e.g., in the case of metastatic cancer cells and circulating tumor cells). As used herein, a "precancer" or "precancerous condition" refers to an abnormality that has the potential to become cancerous, and the likelihood of becoming cancerous is higher than if the potential abnormality were not present, i.e., normal. Examples of precancer include, but are not limited to, adenoma, hyperplasia, dysplasia, dysplasia, benign neoplasm (benign tumor), premalignant intramucosal carcinoma, and polyp. It should be noted that certain types of intramucosal carcinoma are recognized in the art as cancerous, e.g., stage 0 cancer, as opposed to premalignant.
[0188] The term "methylation" or "DNA methylation" refers to the addition of a methyl group to a nucleic acid base in a nucleic acid molecule. In some embodiments, methylation refers to the addition of a methyl group to cytosine at a CpG site (cytosine-phosphate-guanine site (i.e., cytosine followed by guanine in the 5' to 3' direction of a nucleic acid sequence). In some embodiments, DNA methylation refers to the addition of a methyl group to an adenine, e.g., N6-methyladenine. In some embodiments, DNA methylation is 5-methylation (modification of the 5 carbon of the 6-membered ring of cytosine). In some embodiments, 5-methylation refers to the addition of a methyl group to the 5C position of cytosine to produce 5-methylcytosine (5mC). In some embodiments, methylation includes derivatives of 5mC, including, but not limited to, 5-hydroxymethylcytosine (5-hmC), 5-formylcytosine (5-fC), and 5-caryboxylcytosine (5- In some embodiments, DNA methylation is 3C methylation (modification of the nitrogen at the 3 position of the 6-membered ring of cytosine). In some embodiments, 3C methylation involves adding a methyl group to the 3C position of cytosine to generate 3-methylcytosine (3mC). Methylation can also occur at non-CpG sites, for example, methylation can occur at CpA, CpT, or CpC sites. DNA methylation can alter the activity of methylated DNA regions. For example, if DNA in a promoter region is methylated, gene transcription can be suppressed. DNA methylation is crucial for normal development, and abnormal methylation can disrupt epigenetic regulation. Disruption of epigenetic regulation, for example, suppression, can cause diseases such as cancer. Promoter methylation in DNA can indicate cancer.
[0189] The term "hypermethylation" refers to an increased level or degree of methylation of a nucleic acid molecule compared to other nucleic acid molecules containing the same genetic information within a population (e.g., sample) of nucleic acid molecules. In some embodiments, hypermethylated DNA can include DNA molecules containing at least one methylated residue, at least two methylated residues, at least three methylated residues, at least five methylated residues, or at least ten methylated residues.
[0190] The term "hypomethylated" refers to a decreased level or degree of methylation of a nucleic acid molecule compared to other nucleic acid molecules containing the same genetic information within a population (e.g., sample) of nucleic acid molecules. In some embodiments, hypomethylated DNA can include DNA molecules that contain zero methylated residues, at most one methylated residue, at most two methylated residues, at most three methylated residues, at most four methylated residues, or at most five methylated residues.
[0191] The term "agent that recognizes modified nucleobases in DNA," e.g., "agent that recognizes modified cytosines in DNA," refers to a molecule or reagent that binds to or detects one or more modified nucleobases in DNA, e.g., methylcytosines. A "modified nucleobase" is a nucleobase that comprises a difference in chemical structure from an unmodified nucleobase. In the case of DNA, the unmodified nucleobase is adenine, cytosine, guanine, or thymine. In some embodiments, the modified nucleobase is a modified cytosine. In some embodiments, the modified nucleobase is a methylated nucleobase. In some embodiments, the modified cytosine is a methylcytosine, e.g., 5-methylcytosine. In such embodiments, the cytosine modification is methyl. Agents that recognize methylcytosines in DNA include, but are not limited to, "methyl-binding reagents," which herein refer to reagents that bind to methylcytosines. Methyl-binding reagents include, but are not limited to, methyl-binding domains (MBDs) and methyl-binding proteins (MBPs), as well as antibodies specific for methylcytosines. In some embodiments, such antibody binds to 5-methylcytosine in DNA.In some such embodiments, DNA can be single-stranded or double-stranded.Suitable agents include those that recognize double-stranded DNA, single-stranded DNA, and modified nucleotides in both double-stranded and single-stranded DNA.
[0192] The term "epigenetic state" refers to a certain level or degree of sequence-independent variation that may exist in a DNA sequence. In some embodiments, the epigenetic state of a DNA sequence refers to the degree or level of sequence methylation, nucleosome distribution, cfDNA fragmentation pattern, CCCTC-binding factor ("CTCF") binding, transcription start site, or regulatory protein binding region. Thus, epigenetic states include, but are not limited to, hypermethylation, hypomethylation, and the presence or absence of CTCF binding site or transcription start site. The epigenetic state of a sequence may also be a "reference epigenetic state" that can be used to compare the epigenetic state of corresponding sequences in other DNA molecules. An example of a reference epigenetic state is a state that exists in a sample obtained from a healthy subject and is not associated with cancer.
[0193] As used herein, "methylation state" refers to the presence or absence of a methyl group on a DNA nucleobase (e.g., cytosine) at a particular genomic position in a nucleic acid, the degree of methylation of a nucleic acid (e.g., high, low, intermediate, or unmethylated), or the number of methylated nucleotides in a particular nucleic acid molecule. A "methylated form" of a nucleic acid is meant to include a sequence containing a methylated DNA nucleobase, e.g., a methylated cytosine in a CpG dinucleotide.
[0194] As used herein, "methylation-sensitive nuclease" refers to a nuclease that preferentially cleaves unmethylated DNA compared to methylated DNA. For example, a methylation-sensitive nuclease can cleave at or near a recognition sequence, such as a restriction site, in a manner that depends on the absence of methylation of at least one nucleic acid base in the recognition sequence, such as cytosine. In some embodiments, the nucleolytic activity of a methylation-sensitive nuclease is at least 10, 20, 50, or 100 times higher at an unmethylated recognition site compared to a methylated control in a standard nucleolytic assay. Methylation-sensitive nucleases include methylation-sensitive restriction enzymes.
[0195] As used herein, "methylation-sensitive restriction enzyme" or "MSRE" refers to a methylation-sensitive nuclease that is a restriction enzyme. MSREs are sensitive to the methylation state of DNA (e.g., cytosine methylation), i.e., the presence or absence of a methyl group in a nucleotide base in its recognition sequence alters the rate at which the enzyme cleaves DNA. In some embodiments, a methylation-sensitive restriction enzyme does not cleave DNA if a specific nucleotide base is methylated in the recognition sequence. For example, HpaII is a methylation-sensitive restriction enzyme with the recognition sequence "CCGG," and does not cleave DNA if the second cytosine in the recognition sequence is methylated.
[0196] As used herein, "methylation-dependent nuclease" refers to a nuclease that preferentially cleaves methylated DNA compared to unmethylated DNA. For example, a methylation-dependent nuclease can cleave at or near a recognition sequence, such as a restriction site, in a manner that depends on the methylation of at least one nucleic acid base in the recognition sequence, such as cytosine. In some embodiments, the nucleolytic activity of a methylation-dependent nuclease is at least 10, 20, 50, or 100 times higher at a methylated recognition site compared to an unmethylated control in a standard nucleolytic assay. Methylation-dependent nucleases include methylation-dependent restriction enzymes.
[0197] As used herein, "methylation-dependent restriction enzyme" or "MDRE" refers to a methylation-dependent nuclease that is a restriction enzyme. MDREs depend on DNA methylation (e.g., cytosine methylation), i.e., the presence or absence of a methyl group in a nucleotide base alters the rate at which the enzyme cleaves DNA. In some embodiments, a methylation-dependent restriction enzyme does not cleave DNA if a specific nucleotide base is not methylated in the recognition sequence. For example, MspJI is a methylation-dependent restriction enzyme with the recognition sequence "mCNNR(N9)" and does not cleave DNA if a methylated cytosine (mC) is not present in the recognition sequence.
[0198] As used herein, "digestion efficiency" or "cutting efficiency" refers to the efficiency of restriction enzyme digestion. Digestion efficiency can be calculated based on the number of control molecules observed upon digestion with a restriction enzyme and the number of control molecules observed in the absence of restriction enzyme digestion. MSRE digestion efficiency is calculated as follows: Efficiency = 1 - (negative control molecules) [MSRE] Number of negative control molecules [Mock] The MDRE digestion efficiency can be calculated by: Efficiency = 1 - (number of positive control molecules) [MDRE] Number of positive control molecules [Mock] It can be calculated by the number of
[0199] As used herein, "mutation" refers to a variation from a known reference sequence, including, for example, single nucleotide variations (SNVs) and mutations such as insertions or deletions (indels). Mutations can be germline mutations or somatic mutations. In some embodiments, the reference sequence for comparison purposes is the wild-type genomic sequence of the target species from which the test sample is provided, typically the human genome.
[0200] As used herein, the terms "neoplasm" and "tumor" are used interchangeably. They refer to an abnormal growth of cells in a subject. A neoplasm or tumor can be benign, potentially malignant, or malignant. A malignant tumor is called a cancer or cancerous tumor.
[0201] As used herein, "nucleic acid tag" refers to a short nucleic acid (e.g., less than about 500 nucleotides, about 100 nucleotides, about 50 nucleotides, or about 10 nucleotides in length) that is used to distinguish nucleic acids from different samples (e.g., representing a sample index), different types, different fractions in the same sample (e.g., representing a fraction tag), or different nucleic acid molecules (e.g., representing a molecular barcode), or that has undergone different processing. Nucleic acid tags comprise predetermined, fixed, non-random, random, or semi-random oligonucleotide sequences. Such nucleic acid tags may be used to label different nucleic acid molecules or different nucleic acid samples or sub-samples. Nucleic acid tags can be single-stranded, double-stranded, or at least partially double-stranded. Nucleic acid tags can have the same length or various lengths, as desired. Nucleic acid tags can also include double-stranded molecules with one or more blunt ends, include 5' or 3' single-stranded regions (e.g., overhangs), and / or include one or more other single-stranded regions elsewhere within a given molecule. Nucleic acid tags can be attached to one or both ends of other nucleic acids (e.g., sample nucleic acids to be amplified and / or sequenced). Nucleic acid tags can be decoded to reveal information such as the origin, morphology, or processing of a given nucleic acid. For example, nucleic acid tags can also be used to enable pooling and / or parallel processing of multiple samples containing nucleic acids with different molecular barcodes and / or sample indices, with the nucleic acids subsequently deconvoluted by detecting (e.g., reading) the nucleic acid tags. Nucleic acid tags can also be referred to as identifiers (e.g., molecular identifiers, sample identifiers). Additionally or alternatively, nucleic acid tags can be used as molecular identifiers (e.g., to distinguish between amplicons of different molecules or different parent molecules in the same sample or subsample). This includes, for example, unique tagging of different nucleic acid molecules in a given sample or non-unique tagging of such molecules.In the case of non-unique tagging applications, a limited number of tags (i.e., molecular barcodes) may be used to tag each nucleic acid molecule such that different molecules can be distinguished based on their intrinsic sequence information (e.g., start and / or stop positions where they map to a selected reference genome, subsequences at one or both ends of the sequence, and / or sequence length) in combination with at least one molecular barcode. Typically, a sufficient number of different molecular barcodes are used so that there is a low probability (e.g., less than about 10%, less than about 5%, less than about 1%, or less than about 0.1% chance) that any two molecules will have the same intrinsic sequence information (e.g., start and / or stop positions, subsequences at one or both ends of the sequence, and / or length) and also have the same molecular barcode.
[0202] As used herein, "partitioning" refers to physically separating, sorting, and / or fractionating a mixture of nucleic acid molecules in a sample into multiple subsamples or subpopulations of nucleic acids based on characteristics of the nucleic acid molecules. A sample or population can be partitioned into one or more partitioned subsamples or subpopulations based on characteristics indicative of genetic or epigenetic alterations or disease states. The partitioning step can be a physical partitioning of the molecules. The partitioning step can include separating nucleic acid molecules into groups or sets based on the level of an epigenetic trait (e.g., methylation). For example, nucleic acid molecules can be partitioned based on the level of methylation of the nucleic acid molecules. In other words, the partitioning step can include physically partitioning nucleic acid molecules based on the presence or absence of one or more methylated nucleic acid bases. In some embodiments, methods and systems used for partitioning can be found in PCT Patent Application No. PCT / US2017 / 068329, which is incorporated herein by reference in its entirety.
[0203] As used herein, a "distributed set" or "fraction" refers to a set of nucleic acid molecules distributed into sets or groups based on the different binding affinities of the nucleic acid molecules or proteins associated with the nucleic acid molecules for a binder. A distributed set may also be referred to as a subsample. A binder preferentially binds to nucleic acid molecules containing nucleotides with epigenetic modifications. For example, if the epigenetic modification is methylation, the binder may be a methyl-binding domain (MBD) protein. In some embodiments, a distributed set may include nucleic acid molecules that belong to a particular level or degree of epigenetic trait (e.g., methylation). For example, the nucleic acid molecules may be distributed into three sets: one set for highly methylated nucleic acid molecules (first sub-sample, high fraction, highly distributed set, or highly methylated distributed set), a second set for low methylated nucleic acid molecules (second sub-sample, low fraction, low distributed set, or low methylated distributed set), and a third set for intermediately methylated nucleic acid molecules (third sub-sample, intermediate distributed set, intermediate methylation distributed set, residual fraction, or residual distributed set). In another example, the nucleic acid molecules may be distributed based on the number of methylated nucleotides: one distributed set may have nucleic acid molecules with nine methylated nucleotides, and another distributed set may have unmethylated nucleic acid molecules (zero methylated nucleotides).
[0204] As used herein, "sample" means anything that can be analyzed by the methods and / or systems disclosed herein.
[0205] As used herein, "sequencing" refers to any of several techniques used to determine the sequence (e.g., the identity and order of monomeric units) of a biomolecule, e.g., a nucleic acid such as DNA or RNA. Examples of sequencing methods include, but are not limited to, targeted sequencing, single-molecule real-time sequencing, exon or exome sequencing, intron sequencing, electron microscope-based sequencing, panel sequencing, transistor-mediated sequencing, direct sequencing, random shotgun sequencing, Sanger dideoxy termination sequencing, whole genome sequencing, sequencing by hybridization, pyrosequencing, double-strand sequencing, cycle sequencing, single-base extension sequencing, solid-phase sequencing, and hybridization sequencing. These include throughput sequencing, massively parallel signature sequencing, emulsion PCR, co-amplification-PCR (COLD-PCR) at low denaturation temperature, multiplex PCR, reversible dye terminator sequencing, paired-end sequencing, near-term sequencing, exonuclease sequencing, ligation sequencing, short-read sequencing, single molecule sequencing, sequencing by synthesis, real-time sequencing, reverse terminator sequencing, long-read sequencing, nanopore sequencing, 454 sequencing, Solexa Genome analyzer sequencing, SOLiD sequencing, MS-PET sequencing, and combinations thereof.In some embodiments, sequencing can be carried out by genetic analyzer, such as the genetic analyzer commercially available from Illumina, Inc., Pacific Biosciences, Inc., or Applied Biosystems / Thermo Fisher Scientific, and many others.
[0206] As used herein, "next-generation sequencing" or "NGS" refers to a sequencing technology that has increased throughput compared to traditional Sanger and capillary electrophoresis-based approaches, for example, with the ability to generate hundreds of thousands of relatively small sequence reads at a time.Some examples of next-generation sequencing methods include, but are not limited to, sequencing by synthesis, sequencing by ligation, and sequencing by hybridization.In some embodiments, next-generation sequencing involves the use of an instrument that can sequence single molecules.Examples of commercially available instruments for performing next-generation sequencing include, but are not limited to, NextSeq, HiSeq, NovaSeq, MiSeq, Ion PGM, and Ion GeneStudio S5.
[0207] As used herein, the terms "somatic mutation" and "somatic variation" are used interchangeably. They refer to mutations in the genome that occur after conception. Somatic mutations can occur in any cell of the body except germ cells and are therefore not passed on to offspring.
[0208] As used herein, "subject" refers to an animal, such as a mammalian species (e.g., a human) or an avian (e.g., an avian) species, or other organism, such as a plant. More specifically, the subject may be a vertebrate, e.g., a mammal, such as a mouse, a primate, a monkey, or a human. Animals include livestock (e.g., beef cattle, dairy cattle, poultry, horses, pigs, and the like), sport animals, and companion animals (e.g., pets or support animals). A subject may be a healthy individual, an individual having or suspected of having a disease or predisposition to a disease, or an individual in need of treatment or suspected of needing treatment. The terms "individual" or "patient" are intended interchangeably with "subject." For example, a subject may be an individual who has been diagnosed with cancer, an individual who will undergo cancer treatment, and / or an individual who has undergone at least one cancer treatment. A subject may be in remission from cancer. As another example, a subject may be an individual who has been diagnosed with an autoimmune disease. As another example, the subject may be a female individual who may be diagnosed with or suspected of having a disease, e.g., cancer, an autoimmune disease, who is pregnant or planning to become pregnant.
[0209] "Or" is used in the inclusive sense, ie, equivalent to "and / or" unless the context requires otherwise. II. Exemplary Methods A. Overview
[0210] The formation and progression of cancer can be caused by both genetic alteration and the epigenetic characteristics of deoxyribonucleic acid (DNA).The present disclosure provides a method and system for analyzing DNA, such as cell-free DNA (cfDNA).The present disclosure provides a method for analyzing epigenetic target region and / or sequence variable target region together with whole genome analysis in the same workflow.
[0211] Without wishing to be bound by any particular theory, the cells in or around cancer or neoplasm may shed more DNA than the cells of the same tissue type of healthy subjects.Therefore, the distribution of the tissue of origin of a certain DNA sample, for example, cfDNA, may change during carcinogenesis.Thus, for example, the increased level of hypermethylated variable target region, which shows lower methylation in healthy cfDNA than in at least one other tissue type, can be an indication of the presence of cancer (or recurrence depending on the subject's medical history).Similarly, the increased level of hypomethylated variable target region in sample can be an indication of the presence of cancer (or recurrence depending on the subject's medical history).
[0212] In addition, cancer can be manifested by non-sequence alterations such as methylation. Examples of methylation changes in cancer include localized gain of DNA methylation in CpG islands at the TSSs of genes involved in normal growth control, DNA repair, cell cycle regulation, and / or cell differentiation. This hypermethylation can be associated with abnormal loss of transcriptional capacity of the involved genes, occurring at least as frequently as point mutations and deletions as a cause of altered gene expression.
[0213] In this way, DNA methylation profiling can be used to detect the abnormal methylation in the DNA of sample.For example, because of the abnormal increase in the contribution of tissue to sample type (for example, due to the increased DNA loss in or around neoplasia or cancer), and / or the abnormal increase in the contribution from the degree of genomic methylation that changes during development or is disrupted by disease, for example, cancer or any cancer-related disease, DNA can correspond to certain genomic regions (" differentially methylated regions " or " DMR ") that are usually hypermethylated or hypomethylated in given sample type (for example, cfDNA from bloodstream), but can show the abnormal degree of methylation that correlates with neoplasia or cancer.
[0214] In some embodiments, DNA methylation comprises the addition of a methyl group to a cytosine residue at a CpG site (cytosine-phosphate-guanine site (i.e., a cytosine followed by a guanine in the 5' to 3' direction of a nucleic acid sequence). In some embodiments, DNA methylation comprises the addition of a methyl group to an adenine residue, e.g., N6-methyladenine. In some embodiments, DNA methylation is 5-methylation (modification of the 5 carbon of the 6-membered ring of cytosine). In some embodiments, 5-methylation comprises the addition of a methyl group to the 5C position of a cytosine residue to produce 5-methylcytosine (m5c or 5-mC or 5mC). In some embodiments, methylation comprises derivatives of m5c, including, but not limited to, 5-hydroxymethylcytosine (5-hmC or 5hmC), 5-formylcytosine (5-fC), and 5-carboxyl Examples of DNA methylation include cytosine (5-caC). In some embodiments, DNA methylation is 3C methylation (modification of the nitrogen at the 3 position of the 6-membered ring of a cytosine residue). In some embodiments, 3C methylation involves the addition of a methyl group to the 3C position of a cytosine residue to generate 3-methylcytosine (3mC). Methylation can also occur at non-CpG sites, for example, methylation can occur at CpA, CpT, or CpC sites. DNA methylation can alter the activity of methylated DNA regions. For example, if DNA in a promoter region is methylated, gene transcription can be suppressed. DNA methylation is crucial for normal development, and abnormal methylation can disrupt epigenetic regulation. Disruption of epigenetic regulation, for example, suppression, can cause diseases such as cancer. Promoter methylation in DNA can indicate cancer.
[0215] Methylation profiling can involve determining the methylation pattern across different regions of genome.For example, after dividing molecules based on methylation level (for example, the relative number of methylated nucleotides per molecule) and sequencing, the sequences of molecules in different fractions can be mapped to a reference genome.This can show the regions of genome that are more highly methylated or less highly methylated compared to other regions.In this way, genome regions can differ in their methylation level, as opposed to individual molecules.
[0216] Whole genome sequencing is a useful method for detecting global hypomethylation (for example, whole genome (WG) methylation sequencing), fragmentomics approach based on some length, and copy number variation (CNV).Whole genome sequencing, as discussed in detail elsewhere herein, generally refers to the sequencing of a library that is not enriched for one or more sets of target regions of DNA.For fragmentomics and some methylation analysis, whole genome sequencing can be used to provide epigenetic information, but in some cases, can assay relatively short DNA molecules (for example, relatively short cfDNA molecules), which may be too small to be enriched in standard hybrid capture assay, for example, to be adjusted to nucleosome fragments.
[0217] In view of the above, a method that enables the analysis of both target features in specific regions and broad genome-wide signatures associated with cancer in DNA samples (e.g., cfDNA samples) is important for improving clinical outcomes. However, many commercial and developmental methods target specific cancer alterations that occur at very low frequencies in early-stage cancers and precancers. Assaying such alterations requires ultra-deep sequencing and / or enrichment. As such, these methods may not be suitable for detecting broad genome-wide cancer signatures, such as global hypomethylation, which may occur in (semi-)random locations. The present disclosure describes an integrated method (e.g., workflow) that assays both targeted features (e.g., cancer biomarkers, e.g., DMRs) and broad genome-wide signatures (e.g., global hypomethylation, certain somatic alterations (e.g., certain SNVs, indels, gene fusions, and / or CNVs), and / or certain fragmentomics signatures) in an economical manner while maintaining biomarker sensitivity. The methods disclosed herein enable region-specific adjustment of sequencing depth in epigenetic and gene detection assays to optimize, for example, cancer screening performance and assay cost.
[0218] As described herein, dividing nucleic acid molecules in a sample (e.g., a sample from a subject) into multiple sub-samples can allow for region-specific adjustment of sequencing depth (sequencing coverage) in, for example, epigenetic target regions and / or sequence-variable target regions. The sub-samples of the multiple sub-samples can be processed as needed using one or more of the same or different methods described herein (e.g., subjecting one or more DNAs of the multiple sub-samples to a procedure that affects a first nucleic acid base in the DNA differently from a second nucleic acid base in the DNA, enriching one or more sub-samples for one or more target region sets, contacting one or more sub-samples with one or more methylation-sensitive or methylation-dependent nucleases, and / or dividing one or more of the multiple sub-samples based on, for example, methylation status), and then combined to provide a combined sub-sample before sequencing. The combined sub-samples can be sequenced in the same flow cell, reducing assay costs.
[0219] In some embodiments, sub-samples are not combined before sequencing, and are not sequenced in the same flow cell.In such embodiments, sequencing data (for example, the NGS sequencing data from at least the first and second sub-samples described herein) can be analyzed and combined in situ, for example, as described in Example 7.Generally, the embodiments described herein, including these embodiments, can provide improved sensitivity and specificity, for example, by adjusting the depth of both subject and whole genome sequencing, and enabling the analysis of biomarkers from both subject and whole genome assays to be combined into integrated analysis when detecting cancer, pre-cancer, or other conditions.
[0220] In some embodiments, at least one sub-sample of the combined sub-samples (or at least one sub-sample that is not combined before sequencing) is not enriched for one or more sets of target regions of DNA (which may or may not have been subjected to a procedure that affects a first nucleobase in DNA differently from a second nucleobase in DNA, as described herein). For clarity, a sub-sample that is not enriched for one or more sets of target regions of DNA may have been subjected to treatment with a methylation-sensitive or methylation-dependent nuclease, and / or a step of partitioning for epigenetic features, and / or amplification with a universal primer (e.g., that binds to an adapter sequence), and / or any other operation that does not select for a specific target sequence. DNA from such a sub-sample may be used in a sequencing step to evaluate broad genome-wide signatures (e.g., global hypomethylation, certain somatic variations, and / or certain fragment mix signatures).
[0221] While whole-genome assayed biomarkers can be detected at relatively low genome coverage, targeted sequencing of specific DMRs (or somatic variations) may require relatively high coverage, for example, because such features are present at low frequencies in samples such as blood. For example, a higher sequencing depth may be required to analyze sequence-variable target regions with sufficient confidence or precision than may be required to analyze epigenetic target regions. Splitting DNA into multiple subsamples may enable enrichment of DNA corresponding to a set of sequence-variable target regions with a higher enrichment (capture) yield than cfDNA corresponding to the set of epigenetic target regions. Thus, the disclosed method allows for adjusting the depth of both subject and whole-genome sequencing and combining biomarker analysis from both subject and whole-genome assays into an integrated analysis. As described herein, signals obtained from methylation profiling can be combined with signals obtained from somatic variations (e.g., SNVs, indels, CNVs, and / or gene fusions) to facilitate cancer detection.
[0222] In some embodiments, whole genome sequencing can be used to determine a hypomethylation score for each sample, for example, as described in Example 7. Without being bound by theory, cancer-associated differential methylation has been identified in partially methylated domains (PMDs) in DNA. Such domains can span megabases and are estimated to cover approximately 50% of the genome. PMDs can be located in CpG-poor regions of the genome and can coincide with late-replicating regions and nuclear-lamina-associated domains (e.g., the nuclear periphery). In some embodiments of the methods disclosed herein, at least a second (non-enriched) subsample is analyzed to determine a hypomethylation score for each sample. In some embodiments, the PMDs are bins of 5 to 1000 kb in size, such as 50 to 200 kb in size, such as 50 to 150 kb in size, such as 5 kb, 10 kb, 15 kb, 20 kb, 25 kb, 50 kb, 75 kb, 100 kb, 150 kb, 200 kb, 250 kb, 300 kb, 350 kb, 400 kb, 450 kb, 500 kb, 550 kb, 600 kb, 650 kb, 700 kb, 750 kb, 800 kb, 850 kb, 900 kb, 950 kb, or 1000 kb. Then, for each PMD, all molecules with at least the specified number of CpGs (i.e., unique sequences within a given PMD, as indicated by their attached barcodes and / or their genomic start and stop positions) can be added together. In some embodiments, the specified number of CpGs is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or at least 10 CpGs. In some embodiments, the specified number of CpGs is at least 3, 4, 5, or at least 6. In some embodiments, the specified number of CpGs is at least 3. The ratio of "signal molecules" per PMD can then be calculated by dividing the number of fully unmethylated molecules ("signal molecules") by the total number of molecules (e.g., using samples from subjects without cancer, e.g., to determine a predetermined threshold) (this ratio is referred to herein as the "signal ratio").In some embodiments, the predetermined threshold (specifying the minimum percentage of unmethylated molecules in a given PMD) comprises a signal ratio of at least 0.5-10%, at least 0.5-5%, at least 1-10%, or at least 1-5%. In some embodiments, the predetermined threshold comprises a signal ratio of at least 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, or at least 10%. In certain embodiments, the specified number of CpGs is at least 3 and the predetermined threshold is at least 1%. In certain embodiments, the specified number of CpGs is at least 5 and the predetermined threshold is at least 1%. In certain embodiments, the specified number of CpGs is at least 3 and the predetermined threshold is at least 1%. In certain embodiments, the specified number of CpGs is at least 6 and the predetermined threshold is at least 1%. In certain embodiments, the specified number of CpGs is at least 1 and the predetermined threshold is at least 5%.
[0223] A selected number of the "quietest" PMDs (i.e., PMDs that had the lowest ratio of signal molecules) are retained for further analysis. In some embodiments, 10-1000 of the quietest PMDs are retained, such as 10-500, 20-500, 50-500, 50-400, 50-300, 50-200, 50-150, or 75-125 of the quietest PMDs. In some embodiments, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 125, 150, 175, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, or 1000 of the quietest PMDs are retained. The ratio of signal molecules in each of the selected (e.g., 100) most silent PMDs can then be calculated for samples collected from subjects (e.g., subjects with cancer, subjects suspected of having cancer, or subjects at risk of having cancer) using the above method. The distribution of signal ratios per sample can be evaluated across all samples, and signal ratio limits can be identified to determine "positive" PMDs (e.g., those that best distinguish non-cancer samples from cancer samples). The ratio of positive PMDs among the 100 most silent PMDs can then be calculated for each sample as a "per-sample hypomethylation score" (PSH). B. Subjecting the DNA or a subsample thereof to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA.
[0224] In some embodiments, the methods disclosed herein include subjecting DNA or a subsample thereof to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA, wherein the first nucleobase is a modified or unmodified nucleobase, the second nucleobase is a modified or unmodified nucleobase different from the first nucleobase, and the first nucleobase and the second nucleobase have the same base-pairing specificity. In some embodiments, the procedure chemically converts the first or second nucleobase so that the base-pairing specificity of the converted nucleobase changes. In some embodiments, the DNA is subjected to the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA before library preparation using the DNA, before the first amplification of the DNA, before dividing the DNA into multiple subsamples, or any combination thereof. In certain embodiments, the DNA is subjected to the procedure before or after contacting the DNA with a methylation-sensitive nuclease.
[0225] In some embodiments, when the first nucleobase is modified or unmodified adenine, the second nucleobase is modified or unmodified adenine; when the first nucleobase is modified or unmodified cytosine, the second nucleobase is modified or unmodified cytosine; when the first nucleobase is modified or unmodified guanine, the second nucleobase is modified or unmodified guanine; when the first nucleobase is modified or unmodified thymine, the second nucleobase is modified or unmodified thymine (modified and unmodified uracil are encompassed within modified thymine for the purposes of this step).
[0226] In some embodiments, when the first nucleobase is modified or unmodified cytosine, the second nucleobase is modified or unmodified cytosine.For example, the first nucleobase can comprise unmodified cytosine (C), and the second nucleobase can comprise one or more of 5-methylcytosine (mC) and 5-hydroxymethylcytosine (hmC).Alternatively, the second nucleobase can comprise C, and the first nucleobase can comprise one or more of mC and hmC.Other combinations are also possible, such as when one of the first and second nucleobases comprises mC, and the other comprises hmC.
[0227] In some embodiments, the procedure that affects the first nucleic acid base in DNA differently from the second nucleic acid base in the DNA of the first sub-sample comprises bisulfite conversion.Bisulfite treatment converts unmodified cytosine and certain modified cytosine nucleotides (for example, 5-formylcytosine (fC) or 5-carboxylcytosine (caC)) to uracil, while other modified cytosines (for example, 5-methylcytosine, 5-hydroxymethylcytosine) are not converted.Thus, when using bisulfite conversion, the first nucleic acid base comprises one or more of unmodified cytosine, 5-formylcytosine, 5-carboxylcytosine, or other cytosine types that are affected by bisulfite, and the second nucleic acid base can comprise one or more of mC and hmC, for example, mC and optionally hmC.Sequencing of bisulfite-treated DNA identifies the position that is read as cytosine as mC or hmC position. On the other hand, positions that read as T are identified as T or bisulfite-sensitive forms of C, such as unmodified cytosine, 5-formylcytosine, or 5-carboxylcytosine. Thus, performing bisulfite conversion on a DNA sample as described herein facilitates identifying positions containing mC or hmC using sequence reads obtained from an exemplary sample. For an exemplary description of bisulfite conversion, see, e.g., Moss et al., Nat Commun. 2018; 9: 5068.
[0228] In some embodiments, the procedure that affects a first nucleobase in DNA differently from a second nucleobase in the DNA of a first subsample comprises oxidative bisulfite (Ox-BS) conversion. This procedure first converts hmC to fC, which is bisulfite-sensitive, followed by bisulfite conversion. Thus, when using oxidative bisulfite conversion, the first nucleobase comprises one or more of unmodified cytosine, fC, caC, hmC, or other cytosine forms that are affected by bisulfite, and the second nucleobase comprises mC. Sequencing of the Ox-BS-converted DNA identifies positions that are read as cytosine as mC positions. Meanwhile, positions that are read as T are identified as T, hmC, or bisulfite-sensitive forms of C, such as unmodified cytosine, fC, or hmC. Thus, performing Ox-BS conversion on a DNA sample as described herein facilitates identifying mC-containing positions using sequence reads obtained from the sample. For an exemplary description of oxidative bisulfite conversion, see, e.g., Booth et al., Science 2012; 336: 934-937.
[0229] In some embodiments, the procedure for affecting a first nucleobase in DNA differently from a second nucleobase in the DNA of a first sub-sample comprises Tet-assisted bisulfite (TAB) conversion. In TAB conversion, hmC is protected from conversion and mC is oxidized prior to bisulfite treatment, thereby converting the position originally occupied by mC to U and leaving the position originally occupied by hmC as a protected form of cytosine. For example, as described in Yu et al., Cell 2012; 149: 1368-80, after protecting hmC (forming 5-glucosylhydroxymethylcytosine (ghmC)) using β-glucosyltransferase, a TET protein such as mTet1 can be used to convert mC to caC, and then bisulfite treatment can be used to convert C and caC to U, while leaving ghmC unaffected. Alternatively, after protecting hmC using a carbamoyltransferase enzyme such as 5-hydroxymethylcytosine carbamoyltransferase (by converting hmC to 5-carbamoyloxymethylcytosine (5cmC)) as described in Yang et al., Bio-protocol, 2023; 12(17): e4496, a TET protein such as mTet1 can be used to convert mC to caC, followed by bisulfite treatment to convert C and caC to U, while leaving 5cmC unaffected. Thus, when using TAB conversion, the first nucleobase comprises one or more of unmodified cytosine, fC, caC, mC, or other types of cytosine affected by bisulfite, and the second nucleobase comprises hmC. Sequencing of the TAB-converted DNA identifies positions read as cytosine as hmC positions. On the other hand, positions read as T are identified as T, mC, or a bisulfite-sensitive form of C, such as unmodified cytosine, fC, or caC. Thus, performing TAB conversion on a DNA sample as described herein facilitates identifying positions containing hmC using sequence reads obtained from the sample.
[0230] In some embodiments, the procedure for affecting a first nucleobase in DNA differently from a second nucleobase in the DNA of the first subsample includes Tet-assisted conversion with a substituted borane reducing agent, optionally 2-picoline borane, borane pyridine, tert-butylamine borane, or ammonia borane. Tet-assisted pic-borane conversion with a substituted borane reducing agent converts mC and hmC to caC using a TET protein without affecting unmodified C. Then, caC and, if present, fC are converted to dihydrouracil (DHU) by treatment with 2-picoline borane (pic-borane) or another substituted borane reducing agent, such as borane pyridine, tert-butylamine borane, or ammonia borane, similarly without affecting unmodified C. See, for example, Liu et al., Nature Biotechnology 2019; 37:424-429 (e.g., Supplementary Figure 1 and Supplementary Note 7). DHU is read as T in sequencing. Thus, when using this type of conversion, the first nucleobase contains one or more of mC, fC, caC, or hmC, and the second nucleobase contains an unmodified cytosine. Sequencing the converted DNA identifies positions that read as cytosine as unmodified C positions, while positions that read as T are identified as T, mC, fC, caC, or hmC. Thus, performing TAP conversion on a DNA sample as described herein facilitates identifying positions containing unmodified C using sequence reads obtained from the sample. This procedure encompasses Tet-assisted pyridine borane sequencing (TAPS), as described in further detail in Liu et al. 2019, supra.
[0231] Alternatively, protection of hmC (e.g., using βGT or 5-hydroxymethylcytosine carbamoyltransferase) can be combined with Tet-assisted conversion with a substituted borane reducing agent. hmC can be protected as described above through glucosylation using βGT to form ghmC, or through carbamoylation using 5-hydroxymethylcytosine carbamoyltransferase to form 5cmC. Treatment with a TET protein, e.g., mTet1, then converts mC to caC but not C, ghmC, or 5cmC. caC is then converted to DHU by treatment with pic-borane or another substituted borane reducing agent, e.g., borane pyridine, tert-butylamine borane, or ammonia borane, similarly without affecting ghmC, 5cmC, or unmodified C. Thus, when using Tet-assisted conversion with a substituted borane reducing agent, the first nucleobase comprises mC, and the second nucleobase comprises unmodified cytosine or hmC, e.g., unmodified cytosine, and optionally one or more of hmC, fC, and / or caC. Sequencing of the converted DNA identifies positions that read as cytosine as hmC or unmodified C positions. Meanwhile, positions that read as T are identified as T, fC, caC, or mC. Thus, performing TAPSβ conversion on a DNA sample or the like as described herein facilitates distinguishing positions containing either unmodified C or hmC from positions containing mC using sequence reads from the sample. For an exemplary description of this type of conversion, see, e.g., Liu et al., Nature Biotechnology 2019; 37:424-429.
[0232] In some embodiments, the procedure for affecting a first nucleobase in DNA differently from a second nucleobase in the DNA of the first sub-sample comprises chemical-assisted conversion with a substituted borane reducing agent, optionally 2-picoline borane, borane pyridine, tert-butylamine borane, or ammonia borane. In the chemical-assisted conversion with a substituted borane reducing agent, an oxidizing agent, such as potassium perruthenate (KRuO4) (also suitable for use in the ox-BS conversion), is used to specifically oxidize hmC to fC. Treatment with pic-borane or another substituted borane reducing agent, such as borane pyridine, tert-butylamine borane, or ammonia borane, converts fC and caC to DHU, but does not affect mC or unmodified C. Thus, when using this type of conversion, the first nucleobase comprises one or more of hmC, fC, and caC, and the second nucleobase comprises one or more of unmodified cytosine or mC, for example, unmodified cytosine and, optionally, mC. Sequencing the converted DNA identifies positions that read as cytosine as mC or unmodified C positions, while positions that read as T are identified as T, fC, caC, or hmC. Thus, performing this type of conversion on a DNA sample, such as described herein, facilitates distinguishing positions containing either unmodified C or mC from positions containing hmC using sequence reads from the sample. For an exemplary description of this type of conversion, see, e.g., Liu et al., Nature Biotechnology 2019; 37:424-429. 5-hydroxymethylcytosine carbamoyltransferase is described in Yang et al., Bio-protocol, 2023; 12(17): e4496.
[0233] In some embodiments, the procedure for affecting a first nucleobase in DNA differently from a second nucleobase in the DNA of a first subsample comprises APOBEC-linked epigenetic (ACE) conversion. ACE conversion uses an AID / APOBEC family DNA deaminase enzyme, such as APOBEC3A (A3A), to deaminate unmodified cytosine and mC without deaminating hmC, fC, or caC. Thus, when using ACE conversion, the first nucleobase comprises unmodified C and / or mC (e.g., unmodified C and optionally mC), and the second nucleobase comprises hmC. Sequencing of the ACE-converted DNA identifies positions that read as cytosine as hmC, fC, or caC positions. Meanwhile, positions that read as T are identified as T, unmodified C, or mC. Thus, performing ACE conversion on a DNA sample as described herein facilitates using sequence reads from the sample to distinguish positions containing hmC from positions containing mC or unmodified C. For an exemplary description of ACE conversion, see, e.g., Schutsky et al., Nature Biotechnology 2018; 36: 1083-1090.
[0234] In some embodiments, the procedure affecting a first nucleobase in the DNA differently from a second nucleobase in the DNA of the first sub-sample comprises enzymatic conversion of the first nucleobase, e.g., enzymatic conversion in EM-Seq. See, e.g., Vaisvila R, et al. (2019) EM-seq: Detection of DNA methylation at single base resolution from picograms of DNA. bioRxiv; DOI: 10.1101 / 2019.12.20.884692, available at www.biorxiv.org / content / 10.1101 / 2019.12.20.884692v1. For example, TET2 and T4-βGT or 5-hydroxymethylcytosine carbamoyltransferase (described in Yang et al., Bio-protocol, 2023; 12(17): e4496) can be used to convert 5mC and 5hmC into substrates that cannot be deaminated by a deaminase (e.g., APOBEC3A), which can then be used to deaminate unmodified cytosines, converting them to uracil.
[0235] In some embodiments, the procedure of affecting a first nucleobase in DNA differently from a second nucleobase in the DNA comprises enzymatic conversion of the first nucleobase using a double-stranded DNA deaminase sensitive to non-specific modifications, e.g., enzymatic conversion in SEM-seq. See, e.g., Vaisvila et al. (2023) "Discovery of novel DNA cytosine deaminase activities enables a nondestructive single-enzyme methylation sequencing method for base resolution high-coverage methylome mapping of cell-free and ultra-low input DNA." bioRxiv; DOI: 10.1101 / 2023.06.29.547047, available at https: / / www.biorxiv.org / content / 10.1101 / 2023.06.29.547047v1. SEM-Seq uses a nonspecific modification-sensitive double-stranded DNA deaminase (MsddA) in a non-destructive, single-enzyme 5-methylcytosine sequencing (SEM-seq) method to deaminate unmodified cytosines. Therefore, SEM-seq does not require the denaturing step used in protocols based on TET2 and T4-βGT or 5-hydroxymethylcytosine carbamoyltransferase protection, as well as APOEC3A. Additionally, MsddA does not deaminate 5-formylated cytosine (5fC) or 5-carboxylated cytosine (5caC). In SEM-seq, unmodified cytosines in DNA are deaminated to uracil and read as "T" during sequencing. Modified cytosines (e.g., 5mC) are not converted and are read as "C" during sequencing. Cytosines that are read as thymine are identified in DNA as unmodified (e.g., unmethylated) cytosines or as thymine. Thus, performing SEM-seq conversion makes it easy to identify positions containing 5mC using the resulting sequence reads.In some embodiments, the step of affecting a first nucleobase in the DNA differently from a second nucleobase in the DNA comprises enzymatic conversion of the first nucleobase using MsddA.
[0236] In some embodiments, the procedure for affecting the first nucleobase in DNA so that it differs from the second nucleobase in the DNA of the first subsample comprises separating DNA that originally contains the first nucleobase from DNA that does not originally contain the first nucleobase. In some such embodiments, the first nucleobase is hmC. DNA that originally contains the first nucleobase can be separated from other DNA using a labeling procedure that includes a biotinylation site that originally contains the first nucleobase. In some embodiments, the first nucleobase is first derivatized with an azide-containing moiety, for example, a glucosyl-azide-containing moiety. The azide-containing moiety can then serve as a reagent for binding biotin, for example, through Huisgen cycloaddition chemistry. Next, DNA that originally contains the now biotinylated first nucleobase can be separated from DNA that does not originally contain the first nucleobase using a biotin-binding agent, such as avidin, neutravidin (deglycosylated avidin with an isoelectric point of about 6.3), or streptavidin. An example of a procedure for separating DNA that originally contains the first nucleobase from DNA that does not originally contain the first nucleobase is hmC-sealing, which involves labeling hmC to form β-6-azido-glucosyl-5-hydroxymethylcytosine, then attaching a biotin moiety via Huisgen cycloaddition, and then using a biotin-binding agent to separate the biotinylated DNA from other DNA. For an exemplary description of hmC-sealing, see, for example, Han et al., Mol. Cell 2016; 63: 711-719. This approach is useful for identifying fragments containing one or more hmC nucleobases.
[0237] In some embodiments, after such separation, the method further comprises the step of differentially tagging the DNA that originally comprises the first nucleobase and the DNA that does not originally comprise the first nucleobase.The method can further comprise the step of pooling the DNA that originally comprises the first nucleobase and the DNA that does not originally comprise the first nucleobase after differential tagging.The DNA that originally comprises the first nucleobase and the DNA that does not originally comprise the first nucleobase can then be used in downstream analysis.For example, the pooled DNA that originally comprises the first nucleobase and the DNA that does not originally comprise the first nucleobase can be sequenced in the same sequencing cell (for example, after being subjected to further processing such as the processing described herein), while retaining the ability to use differential tag to determine whether a given read originates from the molecule of DNA that originally comprises the first nucleobase or from the DNA that does not originally comprise the first nucleobase.
[0238] In some embodiments, the first nucleobase is a modified or unmodified adenine and the second nucleobase is a modified or unmodified adenine. In some embodiments, the modified adenine is N 6 In some embodiments, the modified adenine is N-methyladenine (mA). 6 -Methyladenine (mA), N 6 -hydroxymethyladenine (hmA), or N 6 -formyl adenine (fA).
[0239] Techniques involving partitioning based on methylation status or methylated DNA immunoprecipitation (MeDIP) can be used to separate DNA containing modified bases such as mC, mA, caC (e.g., which can be generated by oxidation of mC or hmC by Tet2, e.g., using a deaminase such as APOBEC3A, prior to enzymatic conversion of the unmodified C to U), or dihydrouracil, from other DNA. See, e.g., Kumar et al., Frontiers Genet. 2018;9:640; Greer et al., Cell 2015;161:868-878. An antibody specific for mA is described in Sun et al., Bioessays 2015;37:1155-62. Antibodies against various modified nucleobases, such as mC, caC, and forms of thymine / uracil, including halogenated forms such as dihydrouracil or 5-bromouracil, are commercially available. Various modified bases can also be detected based on changes in their base pairing specificity. For example, hypoxanthine is a modified form of adenine that can result from deamination and is read as G in sequencing. See, e.g., U.S. Patent No. 8,486,630; Brown, Genomes, 2002; nd Ed., John Wiley & Sons, Inc., New York, NY, 2002, chapter 14, "Mutation, Repair, and Recombination."
[0240] In some embodiments, the conversion procedure is an enzymatic conversion procedure that converts the base-pairing specificity of a modified nucleoside (e.g., a DM-seq conversion that involves adding a protecting group (e.g., a carboxymethyl group) to an unmodified cytosine and deaminating 5mC, e.g., using an APOBEC enzyme) or an enzymatic conversion procedure that converts the base-pairing specificity of an unmodified nucleoside (e.g., SEM-seq).
[0241] In some cases, the conversion procedure used in the method of the present disclosure is a procedure that changes the base pairing specificity of modified nucleosides (e.g., methylated cytosine), but does not change the base pairing specificity of corresponding unmodified nucleosides (e.g., cytosine), or does not change the base pairing specificity of any unmodified nucleosides (e.g., cytosine, adenosine, guanosine, and thymidine (or uracil)).The advantages of a method that does not change the base pairing specificity of unmodified nucleosides include reduced loss of sequence complexity, higher sequencing efficiency, and reduced alignment loss.In addition, methods such as DM-seq may in some cases be preferable to methods such as bisulfite sequencing and EM-seq because they are less destructive (particularly important for low-yield samples such as cfDNA), do not require denaturation, and non-conversion errors are theoretically more likely to be random.In methods that require denaturation for conversion, failure to denature DNA molecules results in the non-conversion of all bases in the DNA molecule. Because biological changes in methylation are preferentially coordinated to the localized region of interest, these non-random (localized) conversions may appear as false negatives (unmethylated regions). Random non-conversion methods can maximize the effect on low base percentages within a region, thus, by setting a threshold for the percentage of bases within a methylated / unmethylated region, the specificity of methylation change detection can be maximized (false positives can be reduced). Therefore, in some cases, conversion procedures that do not involve denaturation are preferred.
[0242] In other cases, the conversion procedure used in the disclosed methods is one that alters the base pairing specificity of an unmodified nucleoside (e.g., cytosine) but does not alter the base pairing specificity of the corresponding modified nucleoside (e.g., methylated cytosine).
[0243] Those skilled in the art can select a suitable method according to their needs, including which nucleoside modifications are to be detected and / or identified.
[0244] In some embodiments, the conversion procedure converts modified nucleosides. In some embodiments, the conversion procedure for converting modified nucleosides includes enzymatic conversion, such as DM-seq, as described in WO2023 / 288222A1. In DM-seq, unmodified cytosines in DNA are enzymatically protected from a subsequent deamination step, which converts 5mC in 5mCpG to T. Enzymatically protected unmodified (e.g., unmethylated) cytosines are not converted and are read as "C" during sequencing. Cytosines (in CpG contexts) that are read as thymine are identified as methylated cytosines in DNA.
[0245] Thus, when using this type of conversion, the first nucleobase contains an unmodified (e.g., unmethylated) cytosine, and the second nucleobase contains a modified (e.g., methylated) cytosine. Sequencing of the converted DNA identifies positions that are read as cytosine as unmodified C positions. Meanwhile, positions that are read as T are identified as T or 5mC. In this way, performing DM-seq conversion makes it easy to use the resulting sequence reads to identify positions containing 5mC.
[0246] Exemplary cytosine deaminase for use herein includes APOBEC enzyme, for example, APOBEC3A.Generally, AID / APOBEC family DNA deaminase enzyme, for example, APOBEC3A (A3A), is used to deaminate (unprotected) unmodified cytosine and 5mC.For exemplary explanation of APOBEC conversion, see, for example, Schutsky et al., Nature Biotechnology 2018; 36: 1083-1090.
[0247] Enzymatic protection of unmodified cytosines in DNA involves adding a protecting group to the unmodified cytosine. Such protecting groups can include alkyl groups, alkyne groups, carboxyl groups, carboxyalkyl groups, amino groups, hydroxymethyl groups, glucosyl groups, glucosylhydroxymethyl groups, isopropyl groups, or dyes. For example, DNA can be treated with a methyltransferase, such as a CpG-specific methyltransferase, which adds a protecting group to the unmodified cytosine. The term methyltransferase is used broadly herein to refer to an enzyme that can transfer methyl or a substituted methyl (e.g., carboxymethyl) to a substrate (e.g., cytosine in a nucleic acid). In some embodiments, DNA is contacted with a CpG-specific DNA methyltransferase (MTase), such as a CpG-specific carboxymethyltransferase (CxMTase), and a substituted methyl donor, such as a carboxymethyl donor (e.g., carboxymethyl-S-adenosyl-L-methionine). See, for example, WO2021 / 236778A2. In certain embodiments, CxMTase can facilitate the addition of a protective carboxymethyl group to unmethylated cytosine. In some embodiments, the unmethylated cytosine is unmodified cytosine. The carboxymethyl group can prevent deamination of cytosine during a deamination step (e.g., a deamination step using an APOBEC enzyme such as A3A). Substituted methyl or carboxymethyl donors useful in the disclosed methods include, but are not limited to, S-adenosyl-L-methionine (SAM) analogs, and optionally, the SAM analog is carboxy-S-adenosyl-L-methionine (CxSAM). SAM analogs are described, for example, in WO2022 / 197593A1. The MTase may be, for example, CpG methyltransferase from Spiroplasma sp. strain MQ1 (M.SssI), DNA-methyltransferase 1 (DNMT1), DNA-methyltransferase 3 alpha (DNMT3A), DNA-methyltransferase 3 beta (DNMT3B), or DNA adenine methyltransferase (Dam).The CxMTase may be a CpG methyltransferase from Mycoplasma penetrans (M.MpeI). In certain embodiments, the methyltransferase enzyme is a variant of M.MpeI having SEQ ID NO:1 or SEQ ID NO:2, or a sequence at least 90%, at least 92%, at least 94%, at least 96%, at least 97%, at least 98%, or at least 99% identical thereto, and optionally, the amino acid corresponding to position 374 is R or K.
[0248] In one embodiment, the methyltransferase enzyme is a variant of M.MpeI having an N374R or N374K substitution. The methyltransferase of SEQ ID NO:1 or SEQ ID NO:2 can further comprise one or more amino acid substitutions selected from: a) substitution of one or both of residues T300 and E305 with S, A, G, Q, D, or N; b) substitution of one or more of residues A323, N306, and Y299 with a positively charged amino acid selected from K, R, or H; and / or c) substitution of S323 with A, G, K, R, or H, which may enhance the activity of the enzyme.
[0249] Optionally, the conversion procedure further includes enzymatic protection of 5hmC in DNA, such as by glucosylation of 5hmC (e.g., using βGT) or by carbamoylation of 5hmC (e.g., using 5-hydroxymethylcytosine carbamoyltransferase) before deamination of unprotected modified cytosines. In this method, 5hmC can be protected from conversion through glucosylation using, for example, β-glucosyltransferase (βGT) to form (5-glucosylhydroxymethylcytosine) 5ghmC, or through carbamoylation using 5-hydroxymethylcytosine carbamoyltransferase to form 5cmC. This is described, for example, in Yu et al., Cell 2012; 149: 1368-80, and Yang et al., Bio-protocol, 2023; 12(17): e4496. Glucosylation or carbamoylation of 5hmC can reduce or eliminate deamination of 5hmC by deaminases such as APOBEC3A. Treatment with MTase or CxMTase then adds a protecting group to unmodified (unmethylated) cytosines in DNA. 5mC (but not the protected unmodified cytosine, not 5ghmC or 5cmC) is then deaminated (in the case of 5mC, converted to T) by treatment with a deaminase, e.g., an APOBEC enzyme (e.g., APOBEC3A). Sequencing of the converted DNA identifies positions that read as cytosine as 5hmC or unmodified C positions. Meanwhile, positions that read as T are identified as T or 5mC. Thus, performing DM-seq conversion of 5hmC glycosylation on a sample as described herein facilitates using the resulting sequence reads to distinguish positions containing either unmodified C or 5hmC from positions containing 5mC.
[0250] Also provided herein are methods that use alternative base conversion schemes, for example, unmethylated cytosine can remain intact, while methylated and hydroxymethyl cytosine are converted to a base that is read as thymine (e.g., uracil, thymine, or dihydrouracil).
[0251] In some embodiments, methylating cytosines in at least one of the first or second complementary strands comprises contacting the cytosines with a methyltransferase, such as DNMT1 or DNMT5. In such embodiments, oxidizing 5-hydroxymethylated cytosines to 5-formylcytosines (e.g., by contacting the 5-hydroxymethylcytosines in the first and second strands with KRuO4) can be optional.
[0252] In some embodiments, converting the modified cytosine in at least one of the first or second strands to thymine or a base that is read as thymine comprises oxidizing the hydroxymethylcytosine, e.g., oxidizing the hydroxymethylcytosine to formylcytosine. In some embodiments, oxidizing the hydroxymethylcytosine to formylcytosine comprises contacting the hydroxymethylcytosine with a ruthenate, e.g., potassium ruthenate (KRuO).
[0253] In some embodiments, the modified cytosine is converted to thymine, uracil, or dihydrouracil. In any such embodiment, the amplification method may comprise a uracil- and / or dihydrouracil-resistant amplification method, such as PCR using a uracil- and / or dihydrouracil-resistant DNA polymerase.
[0254] In some embodiments, the method includes converting formylcytosine and / or methylcytosine to carboxyl cytosine as part of converting at least one modified cytosine in the first or second strand to thymine or a base that is read as thymine. For example, converting formylcytosine and / or methylcytosine to carboxyl cytosine may include contacting the formylcytosine and / or methylcytosine with a TET enzyme, such as TET1, TET2, or TET3. In some embodiments, the method includes reducing the carboxyl cytosine as part of converting at least one modified cytosine in the first or second strand to thymine or a base that is read as thymine, and / or the carboxyl cytosine is reduced to dihydrouracil. In some embodiments, reducing the carboxyl cytosine includes contacting the carboxyl cytosine with a borane reducing agent or a borohydride reducing agent.
[0255] In some embodiments, the borane or borohydride reducing agent comprises pyridine borane, 2-picoline borane, borane, tert-butylamine borane, ammonia borane, sodium borohydride, sodium cyanoborohydride (NaBHCN), lithium borohydride (LiBH), ethylenediamine borane, dimethylamine borane, sodium triacetoxyborohydride, morpholine borane, 4-methylmorpholine borane, trimethylamine borane, dicyclohexylamine borane, or a salt thereof. In other embodiments, the reducing agent comprises lithium aluminum hydride, sodium amalgam, amalgam, sulfur dioxide, dithionate, thiosulfate, iodide, hydrogen peroxide, hydrazine, diisobutylaluminum hydride, oxalic acid, carbon monoxide, cyanide, ascorbic acid, formic acid, dithiothreitol, beta-mercaptoethanol, or any combination thereof.
[0256] Various TET enzymes can be used in the disclosed methods, as desired. In some embodiments, one or more TET enzymes include TETv. TETv is described in U.S. Patent No. 10,260,088, and its sequence is SEQ ID NO: 1 therein (SEQ ID NO: 3 herein). In some embodiments, one or more TET enzymes include TETcd. TETcd is described in U.S. Patent No. 10,260,088, and its sequence is SEQ ID NO: 3 therein (SEQ ID NO: 4 herein). In some embodiments, one or more TET enzymes include TET1. In some embodiments, one or more TET enzymes include TET2. TET2 can be expressed and used as a fragment comprising residues 1129-1480 of TET2 linked to residues 1844-1936 of TET2 by a linker (SEQ ID NO: 5 herein), e.g., as described in U.S. Patent No. 10,961,525. In some embodiments, one or more TET enzymes include TET1 and TET2. In some embodiments, the one or more TET enzymes comprise a V1900 TET mutant, such as a V1900A, V1900C, V1900G, V1900I, or V1900P TET mutant. In some embodiments, the one or more TET enzymes comprise a V1900 TET2 mutant, such as a V1900A, V1900C, V1900G, V1900I, or V1900P TET2 mutant. Examples of the V1900A, V1900C, V1900G, V1900I, and V1900P TET2 mutants are provided as SEQ ID NOs: 6-10. In some embodiments, the V1900 TET mutant has at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity to SEQ ID NOs: 6, 7, 8, 9, or 10. Position 1900 of the wild-type TET2 sequence corresponds to position 438 in each of SEQ ID NOs: 5-10. Because 5-caC is not a substrate for enzymatic deamination by APOBEC enzymes, such as APOBEC3A, it may be beneficial to use TET enzymes that maximize the formation of 5-carboxylcytosine (5-caC) relative to less oxidized modified cytosines, particularly 5-formylcytosine.Thus, maximizing the formation of 5-caC reduces the risk of false calls in which a base is identified as unmethylated because it underwent deamination even if it was methylated (or hydroxymethylated) in the original sample. Thus, in some embodiments, the TET enzyme contains a mutation that increases the formation of 5-caC. Exemplary mutations are described above. A "mutation that increases the formation of 5-caC" means that a TET enzyme with the mutation produces more 5-caC than a TET enzyme lacking the mutation, all else being equal. 5-caC production can be measured, for example, as described in Liu et al., Nat Chem Biol 13:181-187 (2017) (see the Online Methods section, TET reactions in vitro subsection, "driving" conditions). Any of the variants and / or mutants described in Liu et al. (2017) can be used in the disclosed methods, as desired. C. Adapter ligation or addition; tagging
[0257] In some embodiments, the disclosed methods include adding an adapter to DNA. In some embodiments, the adapter is added to the DNA before or after dividing the DNA into multiple sub-samples, e.g., before dividing, and / or before or after subjecting the DNA to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA, e.g., after subjecting the DNA to such a procedure. In some embodiments, the adapter may be added to the DNA before or after the amplification step, simultaneously with the amplification procedure, e.g., by providing the adapter at the 5' portion of the primer (when PCR is used, this may be referred to as library prep-PCR or LP-PCR). In some embodiments, the adapter is added by other approaches. In some such methods, a first adapter is added to the nucleic acid by ligation to its 3' end, which may include ligation to single-stranded DNA. The adapter can be used as an initiation site for double-stranded synthesis, e.g., using a universal primer and DNA polymerase. In other such methods, a first adapter is added to the nucleic acid by ligation to its 5' end, which may include ligation to single-stranded DNA. A second adaptor can then be ligated to at least the 3' end of the second strand of the now double-stranded molecule. In some embodiments, the first adaptor includes an affinity tag, such as biotin, and the nucleic acid ligated to the first adaptor is attached to a solid support (e.g., beads) that can include a binding partner for the affinity tag, such as streptavidin. For further discussion of related procedures, see Gansauge et al., Nature Protocols 8:737-748 (2013). Commercially available kits for sequencing library preparation compatible with single-stranded nucleic acids are available, such as the Accel-NGS® Methyl-Seq DNA Library Kit from Swift Biosciences. In some embodiments, after adaptor ligation, the nucleic acid is amplified. In some embodiments, DNA end repair is performed before the addition of the adaptor.
[0258] In some embodiments, single-stranded DNA library preparation is performed using a one-step phosphorylation / ligation reaction combination, as described, for example, in Troll et al., BMC Genomics, 20:1023 (2019), available at https: / / doi.org / 10.1186 / s12864-019-6355-0. This method, called Single-Reaction Single-Stranded LibrarY ("SRSLY"), can be performed without end polishing. SRSLY can be useful for converting short fragmented DNA molecules, such as cfDNA fragments, into a sequencing library while retaining their native length and ends. The SRSLY method can generate sequencing libraries (e.g., Illumina sequencing libraries) from fragmented or degraded template (input) DNA. In certain embodiments, the template DNA is first heat-denatured and then immediately subjected to a cold shock to render the template DNA molecules single-stranded. The DNA can be maintained as single-stranded throughout the ligation reaction by the inclusion of a thermostable single-strand binding protein (SSB). The template DNA, which is now single-stranded and may be coated with SSB, is then subjected to a phosphorylation / ligation duplex reaction using directional dsDNA NGS adapters containing single-stranded overhangs. Both forward and reverse sequencing adapters may share a similar structure, but differ in that the ends are unblocked to facilitate proper ligation. Both sequencing adapters may contain a dsDNA portion and single-stranded splint overhangs of random nucleotides that occur at the 3-prime end of the lower strand of the forward adapter and the 5-prime end of the lower strand of the reverse adapter. In this way, the forward adapter (e.g., (P5) Illumina adapter) can be delivered to the 5-prime end of the template molecule, and the reverse adapter (e.g., (P7) Illumina adapter) can be delivered to the 3-prime end of the template molecule. In this way, the native polarity of the input DNA molecule can be preserved.
[0259] During the dual phosphorylation / ligation reaction, T4 polynucleotide kinase (PNK) can be used to phosphorylate the 5-prime end and dephosphorylate the 3-prime end to prepare template DNA ends for ligation. T4 PNK works on both ssDNA and dsDNA molecules and has no activity against the phosphorylation state of proteins. Simultaneously, random nucleotides of the splint adapter can be annealed to the single-stranded template molecule. This creates a short, localized dsDNA molecule, allowing ligation of the template to the adapter by a ligase such as T4 DNA ligase, which has high ligation efficiency for dsDNA templates but low efficiency for ssDNA. After the single phosphorylation / ligation reaction is complete, the library DNA can be purified and directly subjected to standard NGS indexing PCR, compatible with both conventional single- and dual-index primers.
[0260] In some embodiments, after binding of the adapters, the nucleic acids are subjected to amplification, which can, for example, use universal primers that recognize the primer binding sites in the adapters.
[0261] In some embodiments, the DNA is ligated at both ends to Y-shaped adapters that contain primer binding sites and tags. In some such embodiments, the DNA is amplified.
[0262] Tagging a DNA molecule is the process of attaching or associating a tag to a DNA molecule. Such a tag can be a molecule, such as a nucleic acid, that contains information that characterizes the molecule to which the tag is attached. The tag can enable distinguishing the molecule from which a sequence read originated. For example, a molecule can have a sample tag (which distinguishes molecules in one sample from molecules in a different sample) or a molecular tag / molecular barcode / barcode (which distinguishes different molecules from each other in both unique and non-unique tagging scenarios). For methods that include a partitioning step, a fraction tag (which distinguishes molecules in one fraction from molecules in a different fraction) may also be included. In some embodiments, the adapter added to the DNA molecule comprises a tag. In some such embodiments, the tag comprises a barcode or a combination of barcodes. As used herein, the term "barcode" refers to a nucleic acid molecule having a specific nucleotide sequence or the nucleotide sequence itself, depending on the context. A barcode can have, for example, between 10 and 100 nucleotides. A collection of barcodes may have degenerate sequences, or may have sequences with a certain Hamming distance if desired for a particular purpose. Thus, for example, a molecular barcode may be composed of one barcode or a combination of two barcodes, each attached to a different end of a molecule. Additionally or alternatively, different sets of molecular barcodes or molecular tags may be used for different fractions and / or samples, so that the barcodes serve as molecular tags through their individual sequences and serve to identify the corresponding fractions and / or samples based on the sets they are members of. Tags, including barcodes, may be incorporated into adapters or otherwise attached to adapters. Tags may be incorporated by ligation, overlap extension PCR, among other methods.
[0263] Tagging strategies can be divided into unique tagging strategies and non-unique tagging strategies. In unique tagging, all or substantially all molecules in a sample have different tags, thereby allowing reads to be assigned to the original molecule based on a single tag information. Tags used in such methods are sometimes called "unique tags." In non-unique tagging, different molecules in the same sample can have the same tag, thereby assigning sequence reads to the original molecule using other information in addition to the tag information. Such information can include start and stop coordinates, coordinates that map the molecule, a single start or stop coordinate, etc. Tags used in such methods are sometimes called "non-unique tags." Therefore, it is not necessary to uniquely tag every molecule in a sample. This is sufficient to uniquely tag molecules within an identifiable class within a sample. In this way, molecules in different identifiable families can have the same tag without losing information about the identity of the tagged molecule.
[0264] In some embodiments, the adapter comprises a sufficient number of different tags such that the number of tag combinations results in a low probability, for example, 95, 99, or 99.9%, that two nucleic acids with the same start and stop points will receive the same combination of tags. Regardless of whether the adapter has the same or different tags, it can comprise the same or different primer binding sites. In some embodiments, the adapter comprises the same primer binding site.
[0265] In certain embodiments of non-unique tagging, the number of different tags used may be sufficient such that there is a very high probability (e.g., at least 99%, at least 99.9%, at least 99.99%, or at least 99.999%) that all molecules in a particular group have different tags. In some embodiments involving random barcode attachment, e.g., at both ends of the molecule, the combination of barcodes together constitutes a tag. This number in terms is a function of the number of molecules in the call. For example, a class may be all molecules mapping to the same start-stop position in a reference genome. A class may be all molecules mapping to a particular locus, e.g., a particular base or across a particular region (e.g., up to 100 bases or a gene or exon of a gene). In certain embodiments, the number of different tags used to uniquely identify the number z of molecules in a class is 2 * z, 3 * z, 4 * z, 5 * z, 6 * z, 7 * z, 8 * z, 9 * z, 10 * z, 11 * z, 12 * z, 13 * z, 14 * z, 15 * z, 16 * z, 17 * z, 18 * z, 19 * z, 20 * z or 100 * z (e.g., lower bound) and 100,000 * z, 10,000 * z, 1000 * z or 100 * z (e.g., upper limit).
[0266] For example, in a sample of about 5 ng to 30 ng of cell-free DNA, approximately 3,000 molecules are expected to map to a particular nucleotide coordinate, with about 3 to 10 molecules with any given starting coordinate expected to share the same stopping coordinate. Therefore, about 50 to about 50,000 different tags (e.g., about 6 to 220 barcode combinations) may be sufficient to uniquely tag all such molecules. To uniquely tag all 3,000 molecules mapping across nucleotide coordinates, about 1 million to about 20 million different tags would be required.
[0267] Generally, assignment of unique or non-unique tag barcodes in reactions follows the methods and systems described in U.S. Patent Applications Nos. 20010053519, 20030152490, 20110160078, and U.S. Patent Nos. 6,582,908, 7,537,898, and 9,598,731. Tags can be randomly or non-randomly linked to sample nucleic acids.
[0268] In some embodiments, tagged nucleic acid is loaded into microwell plate and then sequenced.Microwell plate can have 96, 384 or 1536 microwells.In some cases, they are introduced with the expected ratio of unique tags to microwells.For example, unique tags can be loaded so that each genome sample is loaded with more than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 10000, 50,000, 100,000, 500,000, 1,000,000, 10,000,000, 50,000,000 or 1,000,000,000 unique tags. In some cases, unique tags may be loaded such that less than about 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 10000, 50,000, 100,000, 500,000, 1,000,000, 10,000,000, 50,000,000, or 1,000,000,000 unique tags are loaded per genomic sample. In some cases, the average number of unique tags loaded per sample genome is about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 10000, 50,000, 100,000, 500,000, 1,000,000, 10,000,000, 50,000,000, or 1,000,000,000 1,000, or more than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 10000, 50,000, 100,000, 500,000, 1,000,000, 10,000,000, 50,000,000 or 1,000,000,000 unique tags per genome sample.
[0269] In some embodiments, 20 to 50 different tags (e.g., barcodes) are ligated to both ends of a target nucleic acid. For example, 35 different tags (e.g., barcodes) ligated to both ends of a target molecule create 35 x 35 permutations, which is equivalent to 1225 permutations for 35 tags. This number of tags is sufficient to ensure that different molecules with the same start and stop points have a high probability (e.g., at least 94%, 99.5%, 99.99%, 99.999%) of receiving different combinations of tags. Other barcode combinations include any number between 10 and 500, such as about 15 x 15, about 35 x 35, about 75 x 75, about 100 x 100, about 250 x 250, and about 500 x 500.
[0270] In some cases, the unique tag may be a predetermined or random or semi-random sequence oligonucleotide. In other cases, multiple barcodes may be used, so that the barcodes are not necessarily unique to each other among the multiple molecular barcodes. In this example, the barcode may be ligated to each molecule, so that the combination of barcode and sequence can be ligated to create a unique sequence that can be tracked individually. As described herein, the detection of a non-unique barcode in combination with the sequence data of the beginning (start) and end (stop) portions of the sequence read can allow for the assignment of a unique identity to a particular molecule. The length of each sequence read or its number of base pairs can also be used to assign a unique identity to such a molecule. As described herein, a fragment from a single strand of nucleic acid that has been assigned a unique identity can thereby allow for the subsequent identification of the fragment from the parent strand.
[0271] In some embodiments, two or more populations, samples, sub-samples, or fractions are differentially tagged, for example, by dividing the sub-samples and / or differentially degrading the sub-samples using one or more methylation-sensitive nucleases. Tags can be used to label individual DNA populations to correlate the tag (or tags) with a particular population or fraction. In some embodiments, a single tag can be used to label a particular population or fraction. In some embodiments, multiple different tags can be used to label a particular population or fraction. In embodiments using multiple different tags to label specific fractions, the set of tags used to label one fraction can be easily distinguished from the set of tags used to label other fractions. In some embodiments, the tag may have additional functionality, for example, the tag may be used to index the source of the sample, or may be used as a unique molecular identifier (which may be used to improve the quality of sequencing data by distinguishing sequencing errors from mutations, e.g., as described in Kinde et al., Proc Nat'l Acad Sci USA 108: 9530-9535 (2011), Kou et al., PLoS ONE,11: e0146638 (2016)), or as a non-unique molecular identifier, e.g., as described in U.S. Pat. No. 9,598,731. Similarly, in some embodiments, the tag may have additional functionality, for example, the tag may be used to index the source of the sample, or may be used as a non-unique molecular identifier (which may be used to improve the quality of sequencing data by distinguishing sequencing errors from mutations).
[0272] In some embodiments, tagging the fractions includes tagging the molecules in each fraction with a fraction tag. After the fractions are recombined (e.g., to reduce the number of required sequencing runs and avoid unnecessary costs) and the molecules are sequenced, the fraction tag identifies the source fraction. In another embodiment, different fractions are tagged with different molecular tag sets, including, for example, barcode pairs. In this way, each molecular barcode is useful for indicating the source fraction and distinguishing molecules within the fraction. For example, a first set of 35 barcodes can be used to tag molecules in a first fraction, and a second set of 35 barcodes can be used to tag molecules in a second fraction.
[0273] In some embodiments, after tagging, molecules can be pooled for sequencing in one run.In some embodiments, sample tag is added to molecule, for example, after adding other tags and in the step after pooling.Sample tag can facilitate the pooling of material generated from multiple samples for sequencing in one run.
[0274] In some embodiments, fraction tag can be associated with sample and fraction.As a simple example, the first tag can represent the first fraction of the first sample, the second tag can represent the second fraction of the first sample, the third tag can represent the first fraction of the second sample, and the fourth tag can represent the second fraction of the second sample.
[0275] Tags may be attached to molecules based on one or more characteristics, but the final tagged molecules in the library may no longer have those characteristics.For example, single-stranded DNA molecules may be distributed and / or tagged, but the final tagged molecules in the library will likely be double-stranded.Similarly, DNA may be distributed based on different methylation levels, but the tagged molecules derived from these molecules in the final library will likely be unmethylated.Therefore, the tags attached to molecules in the library typically represent the characteristics of the "parent molecule" from which the final tagged molecules are derived, and are not necessarily the characteristics of the tagged molecules themselves.
[0276] For example, use barcode 1, 2, 3, 4 etc. to tag and label the molecules in the first fraction; use barcode A, B, C, D etc. to tag and label the molecules in the second fraction; and use barcode a, b, c, d etc. to tag and label the molecules in the third fraction.Differentially tagged fractions can be pooled before sequencing.Differentially tagged fractions can be sequenced separately, or can be sequenced together simultaneously, for example, in the same flow cell of Illumina sequencer.
[0277] After sequencing, analysis of the reads can be performed at the fraction level as well as at the pooled DNA level. Tags are used to separate the reads from different fractions. Analysis can include in silico analysis to determine genetic and epigenetic variations (one or more of methylation, chromatin structure, etc.) using sequence information, genomic coordinate length, coverage, and / or copy number. D. Amplification
[0278] In some embodiments, DNA is amplified.For example, the DNA flanked by the adapter added to the DNA described herein can be amplified by PCR or other amplification methods.The amplification method used herein can include any suitable method, for example, the method known to those skilled in the art.In some embodiments, amplification is initiated by a primer that binds to the primer binding site in the adapter flanking the DNA molecule to be amplified. Amplification methods can involve cycles of denaturation, annealing, and extension due to thermal cycling, such as polymerase chain reaction (PCR), or can be isothermal, such as linear amplification, transcription-mediated amplification, recombinant polymerase amplification (RPA), helix-dependent amplification (HDA), loop-mediated isothermal amplification (LAMP) (Notomi et al., Nuc. Acids Res., 28, e63, 2000), rolling circle amplification (RCA) (Blanco et al., J. Biol. Chem., 264, 8935-8940, 1989), or hyperbranched rolling circle amplification (Lizard et al., Nat. Genetics, 19, 225-232, 1998). Other amplification methods include ligase chain reaction, strand displacement amplification, nucleic acid sequence-based amplification, and sequence-based self-sustained replication.
[0279] In some embodiments, detecting the presence or absence of one or more DNA sequences includes amplification, such as qPCR or digital PCR. Some such embodiments involving targeted detection of DNA sequences using qPCR or digital PCR do not include standard DNA library preparation steps, such as adapter ligation or tagging.
[0280] In some embodiments, dsDNA ligation using T-tailed and C-tailed adapters can be performed, which results in at least 50, 60, 70, or 80% amplification of the double-stranded nucleic acid before ligation to the adapters.
[0281] In some embodiments, the DNA is amplified before dividing the DNA into multiple sub-samples. In some embodiments, the DNA is amplified before enriching for one or more sets of target regions of the DNA from a first sub-sample. In some embodiments, the enriched DNA of the first sub-sample is amplified before combining the enriched DNA of the first sub-sample with the DNA of a second sub-sample. In some embodiments, the amplifying step is performed (a) after subjecting the DNA to a procedure that affects a first nucleobase of the DNA differently from a second nucleobase of the DNA, (b) after contacting at least the first and / or second sub-sample with at least one restriction enzyme, (c) after enriching for one or more sets of target regions of the DNA, or (d) any combination of (a)-(c). In some embodiments, the amplifying step is performed after subjecting the DNA to a procedure that affects a first nucleobase of the DNA differently from a second nucleobase of the DNA. In some embodiments, the amplifying step is performed after contacting at least the first and / or second sub-sample with at least one restriction enzyme. In some embodiments, the amplifying step is performed after enrichment for one or more sets of target regions of DNA. In certain embodiments, the amplification comprises thermocycling amplification. In other certain embodiments, the amplification comprises isothermal amplification. In some embodiments, an adapter comprising a barcode is ligated to the DNA before amplification. E. Dividing nucleic acids into multiple subsamples
[0282] In some embodiments, the methods herein include dividing the DNA into multiple sub-samples, wherein the multiple sub-samples include a first sub-sample and a second sub-sample. Dividing the DNA into multiple sub-samples may include physically separating the DNA sample into two, three, four, five, six, seven, eight, nine, ten, or more than ten sub-samples. All or a portion of the sub-samples, or all or a portion of the sub-samples, can be used in downstream analyses described herein.
[0283] In some embodiments, the plurality of divided sub-samples comprises two sub-samples: a first sub-sample and a second sub-sample, hi some embodiments, the plurality of divided sub-samples comprises three sub-samples: a first sub-sample, a second sub-sample, and a third sub-sample. In some embodiments, the method comprises a partitioning step that is performed after subjecting the DNA to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA, before subjecting the DNA to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA, after ligating an adapter comprising a barcode to the DNA, before ligating an adapter comprising a barcode to the DNA, after amplifying the DNA (e.g., after a first amplification of the DNA), before amplifying the DNA (e.g., before a first amplification of the DNA or before a second amplification of the DNA), before enriching for one or more sets of target regions of the DNA from one or more of the partitioned sub-samples, before contacting the DNA with a methylation-sensitive nuclease, before partitioning at least a portion of the DNA into a plurality of sub-samples (e.g., based on the epigenetic state, e.g., methylation state, of the DNA), before sequencing the DNA (e.g., before combining one or more of the partitioned sub-samples and sequencing the combined sub-samples), or any combination thereof.
[0284] In certain non-limiting examples, the method includes a dividing step performed after subjecting the DNA to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA. In another specific non-limiting example, the method includes a dividing step performed after ligating an adapter comprising a barcode to the DNA. In yet another specific non-limiting example, the method includes a dividing step performed after amplifying the DNA. In another specific non-limiting example, the method includes a dividing step performed after subjecting the DNA to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA, after ligating an adapter comprising a barcode to the DNA, and after amplifying the DNA (e.g., after a first amplification of the DNA). In some embodiments, the dividing step is performed after a first amplification of the DNA and before a second amplification of the DNA, where the second amplification comprises amplification of one or more DNAs of the divided sub-samples, or amplification of a combined sub-sample as described elsewhere herein.
[0285] Some embodiments include enriching for one or more sets of target regions of DNA from one or more of the divided sub-samples, e.g., the first sub-sample, the second sub-sample, or both the first and second sub-samples, thereby providing enriched DNA of one or more divided sub-samples, wherein the one or more sets of target regions comprise at least a set of epigenetic target regions and / or a set of sequence variable target regions.
[0286] Dividing the nucleic acid molecules in a sample into multiple sub-samples can, for example, allow region-specific adjustment of sequencing depth (sequencing coverage) in both epigenetic detection assays and gene detection assays.The sub-samples of the multiple sub-samples can be processed using one or more of the same or different methods described herein (e.g., subjecting one or more DNAs of the multiple sub-samples to a procedure that affects a first nucleic acid base in DNA differently from a second nucleic acid base in DNA, enriching one or more sub-samples for one or more target region sets, contacting one or more sub-samples with one or more methylation-sensitive or methylation-dependent nucleases, and / or dividing one or more of the multiple sub-samples based on, for example, methylation status).In some embodiments, the multiple sub-samples are then combined before sequencing to provide a combined sub-sample.The combined sub-samples can be sequenced in the same flow cell, which can reduce assay costs. In some embodiments, the DNA of the first and second sub-samples is not combined prior to sequencing. F. DNA enrichment; enriched portion; enriched set
[0287] In some embodiments, the methods herein include enriching (also known as "capturing") nucleic acid molecules comprising sequences present in the set of target regions for subsequent analysis. Such enrichment or capture may be performed on any sample or sub-sample described herein using any suitable approach known in the art. The enrichment may be performed on one or more sub-samples prepared during the methods disclosed herein. In some embodiments, DNA is enriched from at least a first sub-sample. In some embodiments, DNA is enriched from at least a first sub-sample or a second sub-sample, e.g., at least a first sub-sample and a second sub-sample. In some embodiments, DNA is enriched from a first sub-sample and / or a second sub-sample after dividing the DNA into multiple sub-samples. If the first sub-sample undergoes a separation step (e.g., a step to separate DNA that naturally contains the first nucleobase (e.g., hmC) from DNA that does not naturally contain the first nucleobase, e.g., hmC-seal), the capture step may be performed on any, any two, or all of the DNA that naturally contains the first nucleobase (e.g., hmC), the DNA that does not naturally contain the first nucleobase, and the DNA of the second sub-sample. In some embodiments, the sub-samples are differentially tagged (e.g., as described herein) and then pooled before undergoing capture.
[0288] In some embodiments, the capturing step includes contacting the DNA with a probe specific for such target region. In some embodiments, the probe includes an oligonucleotide and a capture moiety, such as biotin or one or more of the other examples described below. The probe may have a sequence, such as a gene, selected to span a panel of regions.
[0289] Methods involving DNA enrichment using probes containing a capture moiety, such as target-specific probes labeled with biotin, can also include a second moiety or binding partner that binds to the capture moiety, such as streptavidin. In some embodiments, the enrichment moiety and binding partner can have higher or lower capture yields for different sets of probes, such as those used to enrich for (capture of) a set of sequence-variable target regions and a set of epigenetic target regions, respectively, as discussed elsewhere herein. Methods involving capture moieties are further described, for example, in U.S. Patent No. 9,850,523, issued December 26, 2017, which is incorporated herein by reference.
[0290] Capture moieties include, without limitation, biotin, avidin, streptavidin, nucleic acids containing specific nucleotide sequences, haptens recognized by antibodies, and magnetically attractable particles. In some embodiments, the capture moiety bound to the analyte is captured by its binding partner bound to an isolatable moiety, such as a magnetically attractable particle or a large particle that can be sedimented by centrifugation. The capture moiety can be any type of molecule that allows affinity separation of nucleic acids bearing the capture moiety from nucleic acids lacking the capture moiety. Exemplary capture moieties are biotin, which allows affinity separation by binding to streptavidin that is or can be linked to a solid phase, or oligonucleotides, which allow affinity separation by binding to complementary oligonucleotides that are or can be linked to a solid phase.
[0291] In some embodiments, non-specifically bound DNA that does not contain the target region is washed away from the enriched DNA. In some embodiments, the DNA is then dissociated from the probe and eluted from the solid support using a buffer containing a salt wash or another DNA denaturing agent. In some embodiments, the probe is also eluted from the solid support, for example, by disrupting the biotin-streptavidin interaction. In some embodiments, the enriched DNA is amplified after elution from the solid support. In some such embodiments, DNA containing an adapter is amplified using PCR primers that anneal to the adapter. In some embodiments, the enriched DNA is amplified while bound to the solid support. In some such embodiments, amplification involves the use of a PCR primer that anneals to a sequence within the adapter and a PCR primer that anneals to a sequence within the probe that anneals to the target region of the DNA.
[0292] In some embodiments, the target region is enriched from an aliquot, portion, or sub-sample of a sample (e.g., a sample that has undergone adaptor binding and amplification), while the DNA partitioning step may be performed on a separate aliquot, portion, or sub-sample of the sample. Enriching or capturing DNA containing the target region may include contacting the DNA with a first or second set of target-specific probes. Such target-specific probes may have any of the features described herein with respect to target-specific probe sets, including, but not limited to, the embodiments described herein and the probe-related sections herein. The capturing step may be performed on one or more sub-samples prepared during the methods disclosed herein. In some embodiments, DNA is enriched from a first sub-sample or a second sub-sample, for example, after dividing the DNA into multiple sub-samples. In some embodiments, the sub-samples are differentially tagged (e.g., as described herein) and then pooled before undergoing enrichment. Exemplary methods for enriching DNA containing epigenetic target regions and / or sequence variable target regions can be found, for example, in WO2020 / 160414, which is incorporated herein by reference.
[0293] The enrichment step or steps may be carried out using conditions suitable for specific nucleic acid hybridization, which generally depend in part on characteristics of the probe, such as length, base composition, etc. Those skilled in the art will be familiar with appropriate conditions given their general knowledge in the art regarding nucleic acid hybridization.
[0294] In some embodiments, the methods described herein include enriching cfDNA obtained from a subject for a set of multiple target regions. The target regions may contain differences depending on whether they originate from a tumor, a healthy cell, or a specific cell type. For example, target regions containing epigenetic target regions may exhibit differences in methylation levels and / or fragmentation patterns depending on whether they originate from a tumor or a healthy cell. Similarly, target regions containing sequence-variable target regions may exhibit sequence differences depending on whether they originate from a tumor or a healthy cell. The enriching step generates an enriched set of cfDNA molecules. In some embodiments, cfDNA molecules corresponding to the set of sequence-variable target regions are enriched in the enriched set of cfDNA molecules with a higher capture yield than cfDNA molecules corresponding to the set of epigenetic target regions. In some embodiments, the methods described herein include contacting cfDNA obtained from a subject with a target-specific probe set, wherein the target-specific probe set is configured to capture cfDNA corresponding to the set of sequence-variable target regions with a higher capture yield than cfDNA corresponding to the set of epigenetic target regions.
[0295] Because analyzing sequence-variable target regions with sufficient reliability or accuracy may require a higher sequencing depth than that required for analyzing epigenetic target regions, it may be beneficial to enrich the cfDNA corresponding to a set of sequence-variable target regions with a higher capture yield than the cfDNA corresponding to a set of epigenetic target regions.The amount of data required to determine fragmentation patterns (for example, to test for perturbations in transcription start sites or CTCF binding sites) or fragment abundances (for example, in hypermethylated and hypomethylated fractions) is generally less than the amount of data required to determine the presence or absence of cancer-related sequence mutations.Capturing target region sets with different yields can facilitate sequencing target regions to different sequencing depths in the same sequencing run (for example, using pooled mixtures and / or in the same sequencing cell).Copy number variations such as local amplification are somatic mutations, but they can be detected by sequencing based on read frequency in a manner similar to the approach used to detect certain epigenetic changes, such as changes in methylation.
[0296] In some embodiments, the enriched DNA is amplified. In various embodiments, the method further comprises sequencing the enriched DNA to different sequencing depths, for example, for epigenetic and sequence variable target region sets, in accordance with the discussion herein. In some embodiments, an RNA probe is used. In some embodiments, a DNA probe is used. In some embodiments, a single-stranded probe is used. In some embodiments, a double-stranded probe is used. In some embodiments, a single-stranded RNA probe is used. In some embodiments, a double-stranded DNA probe is used.
[0297] In some embodiments, the enrichment step is performed simultaneously in the same vessel for the probes for the set of sequence variable target regions and the probes for the set of epigenetic target regions, e.g., the probes for the set of sequence variable target regions and the probes for the set of epigenetic target regions and the capture probe are in the same composition. This approach provides a relatively streamlined workflow.
[0298] Alternatively, the capturing step is carried out using a sequence variable target region probe set in a first container and an epigenetic target region probe set in a second container, or the contacting step is carried out using a sequence variable target region probe set at a first time and in the first container and an epigenetic target region probe set at a second time before or after the first time.This approach allows the first and second compositions to be prepared separately, each containing the captured DNA corresponding to the sequence variable target region set and the captured DNA corresponding to the epigenetic target region set.The compositions can be processed separately if desired (for example, to fractionate based on methylation as described elsewhere herein), and recombined at appropriate ratios as needed to provide material for further processing and analysis, such as sequencing.
[0299] In some embodiments, the DNA is amplified. In some embodiments, the amplification is performed before the capturing step. In some embodiments, the amplification is performed after the capturing step.
[0300] In some embodiments, the adapters are included in the DNA. This can be done simultaneously with the amplification procedure, for example, by providing the adapters at the 5' portion of the primers as described above. Alternatively, the adapters can be added by other approaches, such as ligation.
[0301] In some embodiments, a tag that can be or includes a barcode is included in the DNA. The tag can facilitate identification of the origin of the nucleic acid. For example, barcodes can be used to identify the source from which the DNA originates (e.g., a subject, a biological sample (e.g., samples collected at various time points), an enriched DNA sample (e.g., enriched DNA containing a set of epigenetic target regions or enriched DNA containing a set of sequence-variable target regions), a fraction, or the like) after pooling multiple samples for parallel sequencing. This can be done simultaneously with the amplification procedure, for example, by providing a barcode in the 5' portion of the primer as described above. In some embodiments, the adapter and tag / barcode are provided by the same primer or primer set. For example, the barcode can be located 3' of the adapter and 5' of the target-hybridizing portion of the primer. Alternatively, the barcode can be added by other approaches, such as ligation, optionally with the adapter in the same ligation substrate.
[0302] In some embodiments, a collection of target-specific probes is used in the methods described herein, including enriched DNA. In some embodiments, the collection of target-specific probes includes target binding probes specific for one or more sets of target regions. In some embodiments, the capture yield of the target binding probes specific for the set of sequence-variable target regions is higher (e.g., at least two-fold higher) than the capture yield of the target binding probes specific for the set of epigenetic target regions. In some embodiments, the collection of target-specific probes is configured to have a capture yield specific for the set of sequence-variable target regions that is higher (e.g., at least two-fold higher) than its capture yield specific for the set of epigenetic target regions. 1. Enriched Set
[0303] In some embodiments, provide enriched (also known as captured) DNA (for example, cfDNA) set.For example, in the disclosed method, after the dividing step described herein, enriched DNA set can be provided by carrying out enrichment step.Enriched set can comprise DNA corresponding to sequence variable target region set, epigenetic target region set, or combination thereof.
[0304] In some embodiments, a first set of target regions is enriched from a first subsample, the first set including at least epigenetic target regions. The epigenetic target regions enriched from the first subsample may include hypermethylated variable target regions. In some embodiments, hypermethylated variable target regions are CpG-containing regions that are unmethylated or hypomethylated (e.g., below average methylation compared to bulk cfDNA) in cfDNA from healthy subjects. In some embodiments, hypermethylated variable target regions are regions that show lower methylation in healthy cfDNA than in at least one other tissue type. Without wishing to be bound by any particular theory, cancer cells may shed more DNA into the bloodstream than healthy cells of the same tissue type. Therefore, the distribution of tissues of origin of cfDNA may change during carcinogenesis. Thus, an increased level of hypermethylated variable target regions in the first subsample may be an indicator of the presence of cancer (or recurrence, depending on the subject's medical history).
[0305] In some embodiments, a second set of target regions is enriched from a second subsample, the second set including at least epigenetic target regions. The epigenetic target regions may include hypomethylated variable target regions. In some embodiments, hypomethylated variable target regions are CpG-containing regions that are methylated or hypermethylated (e.g., above average methylation compared to bulk cfDNA) in cfDNA from healthy subjects. In some embodiments, hypomethylated variable target regions are regions that show higher methylation in healthy cfDNA than in at least one other tissue type. Without wishing to be bound by any particular theory, cancer cells may shed more DNA into the bloodstream than healthy cells of the same tissue type. Therefore, the distribution of tissues of origin of cfDNA may change during carcinogenesis. Thus, an increased level of hypomethylated variable target regions in the second subsample may be an indicator of the presence of cancer (or recurrence, depending on the subject's medical history).
[0306] In some embodiments, the amount of enriched sequence variable target region DNA is greater than the amount of enriched epigenetic target region DNA when normalized for differences in size (footprint size) of the regions of interest.
[0307] Alternatively, first and second enriched sets may be provided that contain DNA corresponding to the set of sequence variable target regions and DNA corresponding to the set of epigenetic target regions, respectively. The first and second enriched sets may be combined to provide a combined enriched set.
[0308] In some embodiments, where the enriched set comprising DNA corresponding to the set of sequence variable target regions and the set of epigenetic target regions comprises a combination of the enriched sets discussed above, the DNA corresponding to the set of sequence variable target regions is present at a higher concentration than the DNA corresponding to the set of epigenetic target regions, e.g., 1.1 to 1.2 fold higher, 1.2 to 1.4 fold higher, 1.4 to 1.6 fold higher, 1.6 to 1.8 fold higher, 1.8 to 2.0 fold higher, 2.0 to 2.2 fold higher, 2.2 to 2.4 fold higher, 2.4 to 2.6 fold higher, 2.6 to 2.8 fold higher, 2.8 to 3.0 fold higher, 3.0 to 3.5 fold higher, 3.5 to 4.0, 4.0 to 4.5 fold higher, 4.5 to 5.0 fold higher, 5.0 to 5.5 fold higher. High concentration, 5.5 to 6.0 times higher concentration, 6.0 to 6.5 times higher concentration, 6.5 to 7.0 times higher concentration, 7.0 to 7.5 times higher concentration, 7.5 to 8.0 times higher concentration, 8.0 to 8.5 times higher concentration, 8.5 to 9.0 times higher concentration, 9.0 to 9.5 times higher concentration, 9.5 to 10.0 times higher concentration, 10 to 11 times higher concentration, 11 to 12 times higher concentration, 12 to 13 times higher concentration, 13 to 14 times higher concentration, The concentration may be 14-15x higher, 15-16x higher, 16-17x higher, 17-18x higher, 18-19x higher, 19-20x higher, 20-30x higher, 30-40x higher, 40-50x higher, 50-60x higher, 60-70x higher, 70-80x higher, 80-90x higher, or 90-100x higher. The degree of concentration difference accounts for normalization with respect to the footprint size of the target region, as discussed in the definition section. G. Contacting the DNA with a methylation-sensitive or methylation-dependent nuclease
[0309] In some embodiments, DNA or a sub-sample thereof (e.g., a first, second, or third sub-sample prepared by partitioning the sample as described herein, such as by level of cytosine modification, such as methylation, e.g., 5-methylation) is contacted with a methylation-dependent nuclease or a methylation-sensitive nuclease. The contacting step can be performed using a sample divided into multiple sub-samples as disclosed herein and / or using a sample partitioned into multiple sub-samples as disclosed herein. Unless otherwise indicated, when the partitioning step is performed based on cytosine modification, the first sub-sample is the sub-sample with a higher level of modification, the second sub-sample is the sub-sample with a lower level of modification, and if present, the third sub-sample has a level of modification intermediate between the first and second sub-samples.
[0310] In some embodiments, the method herein comprises contacting DNA with methylation-sensitive nuclease, thereby degrading the DNA comprising unmethylated sequence or sequence with low methylation level.In some such embodiments, the methylation-sensitive nuclease is a methylation-sensitive restriction enzyme (MSRE), thereby degrading the DNA comprising the unmethylated recognition site of MSRE.In this way, the methylation-sensitive nuclease can be used in the method herein, comprising one or more steps of depleting unmodified or unmethylated sequence, for example, those present in cfDNA from subject.
[0311] In some embodiments, the method herein comprises contacting DNA with methylation-dependent nuclease, thereby degrading the DNA comprising methylated sequence or sequence with high methylation level.In some such embodiments, the methylation-dependent nuclease is a methylation-dependent restriction enzyme (MDRE), thereby degrading the DNA comprising the methylation recognition site of MSRE.In this way, the methylation-dependent nuclease can be used in the method herein, comprising one or more steps of depleting modified or methylated sequence, for example, those present in cfDNA from a subject.
[0312] As discussed above, the partitioning procedure can result in incomplete sorting of DNA molecules within the sub-sample. A methylation-dependent nuclease or a methylation-sensitive nuclease can be selected to degrade non-specifically partitioned DNA. For example, the second sub-sample can be contacted with a methylation-dependent nuclease, such as a methylation-dependent restriction enzyme. This can degrade non-specifically partitioned DNA (e.g., methylated DNA) in the second sub-sample to generate a processed second sub-sample. Alternatively or additionally, the first sub-sample can be contacted with a methylation-sensitive endonuclease, such as a methylation-sensitive restriction enzyme, thereby degrading non-specifically partitioned DNA in the first sub-sample to generate a processed first sub-sample. Degradation of non-specifically distributed DNA in either or both of the first or second subsamples is proposed as an improvement to the performance of methods that rely on accurate distribution of DNA based on cytosine modifications, for example, to detect the presence of abnormally modified DNA in a sample, to determine the tissue of origin of the DNA, and / or to determine whether a subject has cancer. For example, such degradation can provide improved sensitivity and / or simplify downstream analysis. Generally, if the non-specifically distributed DNA is hypermethylated, for example, in a hypomethylated fraction, a methylation-dependent nuclease, for example, a methylation-dependent restriction enzyme, should be used. Conversely, if the non-specifically distributed DNA is hypomethylated, for example, in a hypermethylated fraction, a methylation-sensitive nuclease, for example, a methylation-sensitive restriction enzyme, should be used. Methylation-dependent nucleases, e.g., methylation-dependent restriction enzymes, preferentially cleave methylated DNA compared to unmethylated DNA, whereas methylation-sensitive nucleases, e.g., methylation-sensitive restriction enzymes, preferentially cleave unmethylated DNA compared to methylated DNA.
[0313] The step of contacting the sub-sample with the nuclease(s) can use one or more nucleases. In some embodiments, the sub-sample is contacted with multiple nucleases. The sub-samples may be contacted with the nuclease(s) sequentially or simultaneously. The simultaneous use of nucleases can be advantageous to avoid unnecessary sample manipulation when the nucleases are active under similar conditions (e.g., buffer composition). Contacting the second sub-sample with more than one methylation-dependent restriction enzyme can more completely degrade non-specifically distributed hypermethylated DNA. Similarly, contacting the first sub-sample with more than one methylation-sensitive restriction enzyme can more completely degrade non-specifically distributed hypomethylated and / or unmethylated DNA.
[0314] In some embodiments, the methylation-dependent nuclease comprises one or more of MspJI, LpnPI, FspEI, or McrBC. In some embodiments, at least two methylation-dependent nucleases are used. In some embodiments, at least three methylation-dependent nucleases are used. In some embodiments, the methylation-dependent nuclease comprises FspEI. In some embodiments, the methylation-dependent nuclease comprises FspEI and MspJI, for example, used sequentially.
[0315] In some embodiments, the methylation-sensitive nucleases include one or more of AatII, AccII, AciI, Aor13HI, Aor15HI, BspT104I, BssHII, BstUI, Cfr10I, ClaI, CpoI, Eco52I, HaeII, HapII, HhaI, Hin6I, HpaII, HpyCH4IV, MluI, MspI, NaeI, NotI, NruI, NsbI, PmaCI, Pspl406I, PvuI, SacII, SalI, SmaI, and SnaBI. In some embodiments, at least two methylation-sensitive nucleases are used. In some embodiments, at least three methylation-sensitive nucleases are used. In some embodiments, the methylation-sensitive nuclease includes BstUI and HpaII. In some embodiments, the two methylation-sensitive nucleases include HhaI and AccII. In some embodiments, the methylation-sensitive nucleases include BstUI, HpaII, and Hin6I.
[0316] In some embodiments, FspEI is used to digest nucleic acid molecules in at least one sub-sample (e.g., a hypomethylated fraction). In some embodiments, BstUI, HpaII, and Hin6I are used to digest nucleic acid molecules in at least one sub-sample (e.g., a hypermethylated fraction), and FspEI is used to digest nucleic acid molecules in at least one other sub-sample (e.g., a hypomethylated fraction). In embodiments including an intermediate methylation fraction, the nucleic acid molecules may be digested with a methylation-sensitive or methylation-dependent nuclease. In some embodiments, the nucleic acid molecules in the intermediate methylation fraction are digested with the same nuclease as the hypermethylated fraction. For example, the intermediate methylation fraction may be pooled with the hypermethylated fraction, and the pooled fractions may then be subjected to digestion. In some embodiments, the nucleic acid molecules in the intermediate methylation fraction are digested with the same nuclease as the hypomethylated fraction. For example, the intermediately methylated fraction may be pooled with the lowly methylated fraction, and the pooled fractions may then be subjected to digestion.
[0317] In some embodiments, after tagging or binding adapters to both ends of DNA, the sub-sample is contacted with the above-mentioned nuclease. The tag or adapter can be resistant to cleavage by the nuclease using any of the above-mentioned approaches. In this approach, since the cleavage product lacks tags or adapters at both ends, cleavage can prevent non-specifically distributed molecules from being included in the analysis.
[0318] Alternatively, the step of tagging or attaching an adapter can be performed after the cleavage by the nuclease described above.The cleaved molecules can then be identified in the sequence reads based on having the end (the attachment point to the tag or adapter) corresponding to the nuclease recognition site.Processing molecules in this manner can also enable the acquisition of information from the cleaved molecules, for example, the observation of somatic mutations.When tagging or attaching an adapter after contacting a sub-sample with a nuclease, and when low-molecular-weight DNA, such as cfDNA, is analyzed, it may be desirable to remove high-molecular-weight DNA (e.g., contaminating genomic DNA) from the sample before the contacting step.It may also be desirable to use a nuclease that can be heat-inactivated at a relatively low temperature (e.g., 65°C or lower, or 60°C or lower) to avoid DNA denaturation, as denaturation may interfere with the subsequent ligation step.
[0319] When the sample is divided into three sub-samples, including a third sub-sample containing intermediate methylated molecules, the third sub-sample is, in some embodiments, contacted with a methylation-sensitive nuclease. Such a step may have any of the features described elsewhere herein in connection with the contacting step and may be performed before or after the adapter tagging or binding step, as discussed above. In some embodiments, the first and third sub-samples are combined before contacting with the methylation-sensitive nuclease. Such a step may have any of the features described elsewhere herein in connection with the contacting step and may be performed before or after the adapter tagging or binding step, as discussed above. In some embodiments, the first and third sub-samples are differentially tagged before being combined.
[0320] Alternatively, if the sample is divided into three sub-samples, including a third sub-sample containing intermediate methylated molecules, the third sub-sample, in some embodiments, is contacted with a methylation-dependent nuclease. Such a step may have any of the features described elsewhere herein in connection with the contacting step and may be performed before or after the adapter tagging or binding step, as discussed above. In some embodiments, the second and third sub-samples are combined before contacting with the methylation-dependent nuclease. Such a step may have any of the features described elsewhere herein in connection with the contacting step and may be performed before or after the adapter tagging or binding step, as discussed above. In some embodiments, the second and third sub-samples are differentially tagged before being combined.
[0321] In some embodiments, DNA is purified after contacting with a nuclease, for example, using SPRI beads. Such purification may be performed after heat inactivation of the nuclease. Alternatively, purification can be omitted, and subsequent steps, such as amplification, can be performed on a subsample containing the heat-inactivated nuclease. In another embodiment, the contacting step can be performed in the presence of a purification reagent, such as SPRI beads, to minimize losses associated with, for example, tube transfer. After cleavage and heat inactivation, the SPRI beads can be reused for purification by adding a molecular crowding reagent (e.g., PEG) and salt. H. Divide the sample into multiple subsamples
[0322] The methods disclosed herein include analyzing DNA in a sample. In such methods, different types of DNA (e.g., hypermethylated and hypomethylated DNA) can be physically partitioned into multiple sub-samples based on one or more characteristics of the DNA. This approach can be used to determine, for example, whether a particular sequence is hypermethylated or hypomethylated. Such partitioning can be performed before or after dividing the DNA into multiple sub-samples. Thus, the partitioning can be performed using an undivided DNA sample and / or using one or more sub-samples of a DNA sample divided as described herein. In the disclosed methods, the partitioning is not considered a form of enrichment for one or more sets of target regions of DNA.
[0323] In some embodiments, the partitioning step comprises contacting the DNA with an agent that recognizes a modification associated with (e.g., in) the DNA. In some embodiments, the agent that recognizes the modification is an antibody. In some embodiments, the agent is immobilized on a solid support. In some embodiments, the partitioning step comprises immunoprecipitation, e.g., using an antibody agent, such as an antibody, immobilized on a solid support.
[0324] In some embodiments, the modification is methylation, and in some such embodiments, the partitioning step comprises partitioning based on methylation level. In some such embodiments, the agent is a methyl-binding reagent. In some embodiments, the methyl-binding reagent specifically recognizes 5-methylcytosine. In some such embodiments, the agent is a hydroxymethyl-binding reagent. In some embodiments, the methyl-binding reagent specifically recognizes 5-hydroxymethylcytosine, biotinylated 5-hydroxymethylcytosine, glucosylated 5-hydroxymethylcytosine, or sulfonylated 5-hydroxymethylcytosine. In some embodiments, the partitioning step comprises partitioning based on binding to a protein, comprising contacting a sample containing DNA with a specific binding reagent. In some such embodiments, the binding reagent specifically binds to methylated proteins, acetylated proteins, e.g., methylated histones or acetylated histones. In some embodiments, the binding reagent specifically binds to unmethylated or unacetylated protein epitopes.
[0325] In some embodiments, the modification is hydroxymethylation, and in some such embodiments, the partitioning step comprises partitioning based on hydroxymethylation levels. In some such embodiments, the agent is a hydroxymethyl-binding reagent, e.g., an antibody. In some embodiments, the hydroxymethyl-binding reagent (e.g., an antibody) specifically recognizes 5-hydroxymethylcytosine (5-hmC). In some embodiments, the modification, such as hydroxymethylation, is labeled (e.g., biotinylated, glycosylated, or sulfonated) before contacting with an agent that recognizes the labeled form of the modification. For example, 5-hmC can be enzymatically glycosylated and then partitioned based on binding to J-binding protein 1. Exemplary methods for labeling and / or distributing 5-hmC are provided, for example, in Song et al., Nat. Biotech. 29:68-72 (2010); Ko et al., Nature 468:839-843 (2010); and Robertson et al., Nucleic Acids Res. 39:e55 (2011).
[0326] If immunoprecipitation is used and involves antibodies that recognize single-stranded DNA, the DNA can be converted to double-stranded form by complementary strand synthesis prior to subsequent steps. Such synthesis may use adapters as primer binding sites or may use random priming.
[0327] In some embodiments, a sample comprising DNA is partitioned into multiple sub-samples. In some embodiments, the multiple partitioned sub-samples include two sub-samples, a first sub-sample and a second sub-sample. In some embodiments, the multiple partitioned sub-samples include three sub-samples, a first sub-sample, a second sub-sample, and a third sub-sample. In some embodiments, the method includes a partitioning step performed (a) before enriching the DNA for one or more sets of epigenetic target regions and / or sequence variable target regions of the DNA; (b) after enriching the DNA for one or more sets of epigenetic target regions and / or sequence variable target regions of the DNA; (c) before subjecting the DNA of the first sub-sample to a procedure that affects a first nucleobase of the DNA differently from a second nucleobase of the DNA; (d) after subjecting the DNA of the first sub-sample to a procedure that affects a first nucleobase of the DNA differently from a second nucleobase of the DNA; or (e) any combination of (a)-(d).
[0328] In some embodiments comprising a third apportioned sub-sample, the third sub-sample comprises DNA that is associated with the modification at a higher rate than it is associated with DNA in the second sub-sample and at a lower rate than it is associated with DNA in the first sub-sample.
[0329] Partitioning nucleic acid molecules in a sample can increase rare signal levels, for example, by enriching for rare nucleic acid molecules that are more abundant in one fraction of the sample. For example, genetic variations present in hypermethylated DNA but less abundant (or absent) in hypomethylated DNA can be more easily detected by partitioning the sample into hypermethylated and hypomethylated nucleic acid molecules. By analyzing multiple fractions of a sample, multidimensional analysis of single molecules can be performed, thus achieving greater sensitivity. Partitioning can include physically partitioning nucleic acid molecules into fractions or subsamples based on the presence or absence of one or more methylated nucleic acid bases. Samples can be partitioned into fractions or subsamples based on features that are indicative of differential gene expression or disease state. Samples can be partitioned based on features, or combinations thereof, that provide a signal difference between normal and diseased states during the analysis of nucleic acids, such as cell-free DNA (cfDNA), non-cfDNA, tumor DNA, circulating tumor DNA (ctDNA), and cell-free nucleic acid (cfNA).
[0330] In some embodiments, hypermethylated and / or hypomethylated variable target regions are analyzed to determine whether they exhibit differential methylation signatures of tumor cells or cell types that do not normally contribute to the DNA sample (e.g., cfDNA) and / or particular immune cell type being analyzed.
[0331] In some embodiments, each fraction is differentially tagged.Then, the tagged fraction can be pooled together for collective sample preparation and / or sequencing.The step of dividing-tagging-pooling can occur more than once, and each dividing round occurs based on different characteristics (examples provided herein) and is tagged with a differential tag that distinguishes it from other fractions and dividing means.In other examples, the fractions that are differentially tagged are sequenced separately.
[0332] In some embodiments, sequence reads of differentially tagged and pooled DNA are obtained and analyzed in silico. Tags are used to sort reads from different fractions. Analysis to detect genetic variants can be performed at the level of each fraction and at the level of the entire nucleic acid population. For example, analysis can include in silico analysis to determine genetic variants, such as copy number variations (CNVs), single nucleotide variations (SNVs), insertions / deletions (indels), and / or fusions in the nucleic acids of each fraction. In some examples, in silico analysis can include determining chromatin structure. For example, the coverage of sequence reads can be used to determine nucleosome positions in chromatin. Higher coverage can be correlated with higher nucleosome occupancy in a genomic region, while lower coverage can be correlated with lower nucleosome occupancy or nucleosome-depleted regions (NDRs).
[0333] Examples of characteristics that can be used for partitioning include sequence length, methylation level, sequence mismatch, immunoprecipitation, and / or proteins that bind to DNA. The resulting fractions can contain one or more of the following nucleic acid types: single-stranded DNA (ssDNA), double-stranded DNA (dsDNA), short DNA fragments, and long DNA fragments. In some embodiments, partitioning based on cytosine modification (e.g., cytosine methylation) or methylation is generally performed, optionally combined with at least one additional partitioning step that can be based on any of the aforementioned DNA characteristics or types. In some embodiments, a heterogeneous nucleic acid population is partitioned into nucleic acids with one or more epigenetic modifications and nucleic acids without one or more epigenetic modifications. Examples of epigenetic modifications include the presence or absence of methylation; the level of methylation; the type of methylation (e.g., 5-methylcytosine vs. other types of methylation, such as adenine methylation and / or cytosine hydroxymethylation); and the association and level of association with one or more proteins, such as histones. Alternatively or additionally, heterogeneous nucleic acid populations can be divided into nucleosome-associated nucleic acid molecules and nucleosome-free nucleic acid molecules.Alternatively or additionally, heterogeneous nucleic acid populations can be divided into single-stranded DNA (ssDNA) and double-stranded DNA (dsDNA).Alternatively or additionally, heterogeneous nucleic acid populations can be divided based on nucleic acid length (for example, molecules up to 160 bp and molecules with a length greater than 160 bp).
[0334] The agent used to partition the nucleic acid population within the sample can be an affinity agent, such as an antibody with desired specificity, its natural binding partner or variant (Bock et al., Nat Biotech 28: 1106-1114 (2010); Song et al., Nat Biotech 29: 68-72 (2011)), or an artificial peptide selected to have specificity for a given target, for example, by phage display. In some embodiments, the agent used in the partitioning step is an agent that recognizes a modified nucleobase. In some embodiments, the modified nucleobase recognized by the agent is a modified cytosine, such as a methylcytosine (e.g., 5-methylcytosine). In some embodiments, the modified nucleobase recognized by the agent is the product of a procedure that affects a first nucleobase in DNA so that it differs from a second nucleobase in the DNA of the sample. In some embodiments, the modified nucleobase may be a "converted nucleobase," meaning that its base-pairing specificity has been altered by the procedure. For example, certain procedures convert unmethylated or unmodified cytosine to dihydrouracil, or more generally, at least one modified or unmodified form of cytosine undergoes deamination, resulting in uracil (considered a modified nucleobase in the context of DNA) or a further modified form of uracil. Examples of partitioning agents include antibodies, such as antibodies that recognize modified nucleobases, which may be modified cytosines, such as methylcytosines (e.g., 5-methylcytosine). In some embodiments, the partitioning agent is an antibody that recognizes modified cytosines other than 5-methylcytosine, such as 5-carboxylcytosine (5caC). Exemplary partitioning agents include proteins such as the methyl-binding domains (MBDs) and methyl-binding proteins (MBPs) described herein, including MeCP2, MBD2, and antibodies that preferentially bind to 5-methylcytosine. When antibodies are used to immunoprecipitate methylated DNA, the methylated DNA can be recovered in single-stranded form. In such embodiments, a second strand can be synthesized.The highly methylated (and optionally, intermediately methylated) sub-sample may then be contacted with a methylation-sensitive nuclease that does not cleave hemimethylated DNA, such as HpaII, BstUI, or Hin6i. Alternatively or additionally, the hypomethylated (and optionally, intermediately methylated) sub-sample may then be contacted with a methylation-dependent nuclease that cleaves hemimethylated DNA.
[0335] Additional non-limiting examples of partitioning agents or binding reagents are histone-binding proteins that can separate histone-bound nucleic acids from free or unbound nucleic acids. Examples of histone-binding proteins that can be used in the methods disclosed herein include RBBP4, RbAp48, and SANT domain peptides.
[0336] In some embodiments, the partitioning step can include both binary partitioning and partitioning based on the degree / level of modification. For example, methylated fragments can be partitioned by methylated DNA immunoprecipitation (MeDIP), or all methylated fragments can be partitioned from unmethylated fragments using a methyl-binding domain protein (e.g., MethylMinder Methylated DNA Enrichment Kit (ThermoFisher Scientific)). Additional partitioning can then involve eluting fragments with different levels of methylation by adjusting the salt concentration in the solution containing the methyl-binding domain and bound fragments. As the salt concentration increases, fragments with greater methylation levels are eluted.
[0337] In some cases, the final fraction is enriched for nucleic acids with different degrees of modification (over- or under-representation of the modification). Over- and under-representation can be defined by the number of modifications a nucleic acid has compared to the median number of modifications per strand in the population. For example, if the median number of 5-methylcytosine residues in nucleic acids in a sample is 2, nucleic acids containing more than two 5-methylcytosine residues will be over-represented in this modification, and nucleic acids with one or zero 5-methylcytosine residues will be under-represented. The effect of affinity separation is to enrich for nucleic acids that are over-represented in the modification in the binding phase and under-represented in the modification in the non-binding phase (i.e., in solution). The nucleic acids in the binding phase can be eluted prior to further processing.
[0338] When using MeDIP or MethylMiner® Methylated DNA Enrichment Kit (ThermoFisher Scientific), various methylation levels can be separated using sequential elution. For example, a low-methylated fraction (no methylation) can be separated from a methylated fraction by contacting the nucleic acid population with MBD from the kit bound to magnetic beads. The beads are used to separate methylated nucleic acids from unmethylated nucleic acids. One or more sequential elution steps are then performed to elute nucleic acids with different levels of methylation. For example, a first set of methylated nucleic acids can be eluted with a salt concentration of 160 mM or higher, e.g., at least 150 mM, at least 200 mM, 300 mM, 400 mM, 500 mM, 600 mM, 700 mM, 800 mM, 900 mM, 1000 mM, or 2000 mM. After elution of such methylated nucleic acids, magnetic separation is again used to separate highly methylated nucleic acids from nucleic acids with low levels of methylation. The elution and magnetic separation steps can be repeated to generate various fractions, such as a hypomethylated fraction (enriched in nucleic acids that do not contain methylation), a methylated fraction (enriched in nucleic acids that contain low methylation levels), and a hypermethylated fraction (enriched in nucleic acids that contain high methylation levels).
[0339] In some methods, nucleic acids bound to the agent used for affinity separation based on partitioning are subjected to a washing step. The washing step washes away nucleic acids that are weakly bound to the affinity agent. Such nucleic acids can be enriched for nucleic acids with a degree of modification closer to the mean or median (i.e., intermediate between nucleic acids that remain bound to the solid phase and nucleic acids that do not bind to the solid phase when the sample is first contacted with the agent).
[0340] Affinity separation results in at least two, sometimes three or more fractions of nucleic acids with different degrees of modification. Although the fractions are still separated, the nucleic acids of at least one fraction, usually two or three (or more) fractions, are usually linked to nucleic acid tags provided as components of adapters, and the nucleic acids in different fractions are tagged with different tags that distinguish members of one fraction from members of another fraction. The tags linked to nucleic acid molecules of the same fraction can be the same or different from each other. However, when different from each other, the tags can share part of their code so that the molecules to which they are linked are identified as molecules of a specific fraction.
[0341] For further details regarding partitioning nucleic acid samples based on characteristics such as methylation, see WO2018 / 119452, which is incorporated herein by reference.
[0342] In some embodiments, the partitioning step is performed after contacting the DNA with a methylation-sensitive restriction enzyme (MSRE) and / or a methylation-dependent restriction enzyme (MDRE). After treatment of the DNA with an MSRE or MDRE, the DNA may be partitioned based on size to generate hypermethylated (longest DNA molecules after MSRE treatment and shortest DNA fragments after MDRE treatment), intermediate (intermediate length DNA molecules after MSRE or MDRE treatment), and hypomethylated (shortest DNA molecules after MSRE treatment and longest DNA fragments after MDRE treatment) subsamples.
[0343] In some embodiments, the partitioning step is carried out by contacting the nucleic acid with a methyl-binding domain ("MBD") of a methyl-binding protein ("MBP"). In some such embodiments, the nucleic acid is contacted with the entire MBP. In some embodiments, the MBD binds to 5-methylcytosine (5mC), and the MBP comprises the MBD, and is referred to herein interchangeably as a methyl-binding protein or a methyl-binding domain protein. In some embodiments, the MBD is linked to paramagnetic beads, e.g., Dynabeads® M-280 streptavidin, via a biotin linker. Partitioning into fractions with different degrees of methylation can be carried out by eluting the fractions with increasing NaCl concentrations.
[0344] In some embodiments, the bound DNA is eluted by contacting the antibody or MBD with a protease, such as proteinase K. This may be performed instead of or in addition to the elution step using NaCl discussed herein.
[0345] Examples of agents that recognize modified nucleobases contemplated herein include, but are not limited to:
[0346] (a) MeCP2 and MBD2 are proteins that preferentially bind 5-methyl-cytosine compared to unmodified cytosine.
[0347] (b) RPL26, PRP8, and the DNA mismatch repair protein MHS6 preferentially bind 5-hydroxymethyl-cytosine compared to unmodified cytosine.
[0348] (c) FOXK1, FOXK2, FOXP1, FOXP4, and FOXI3 preferentially bind 5-formyl-cytosine compared to unmodified cytosine (Iurlaro et al., Genome Biol. 14: R119 (2013)).
[0349] (d) an antibody specific for one or more methylated or modified nucleobases, or their conversion products, e.g., 5mC, 5caC, or DHU; Examples include:
[0350] Typically, elution is a function of the number of modifications, such as the number of methylation sites per molecule, with molecules with more methylation eluting at increasing salt concentrations. A series of elution buffers with increasing NaCl concentrations can be used to elute DNA into distinct populations based on the degree of methylation. Salt concentrations can range from about 100 mM to about 2500 mM NaCl. In one embodiment, the process results in three fractions. The molecules are contacted with a solution at a first salt concentration containing molecules containing an agent that recognizes modified nucleobases, allowing the molecules to bind to a capture moiety such as streptavidin. At the first salt concentration, some population of molecules bind to the agent, while others remain unbound. The unbound population can be separated as a "hypomethylated" population. For example, the first fraction enriched in hypomethylated DNA is the fraction that remains unbound at low salt concentrations, e.g., 100 mM or 160 mM. A second fraction enriched in intermediately methylated DNA is eluted using an intermediate salt concentration, e.g., 100 mM to 2000 mM, and is also separated from the sample. A third fraction enriched in highly methylated DNA is eluted using a high salt concentration, e.g., at least about 2000 mM.
[0351] In some embodiments, a monoclonal antibody raised against 5-methylcytidine (5mC) is used to purify methylated DNA. The DNA is denatured, for example, at 95°C to generate single-stranded DNA fragments. Protein G is coupled to standard or magnetic beads, and the antibody-bound DNA is immunoprecipitated using a wash after incubation with the anti-5mC antibody. The DNA can then be eluted. The fractions can include unprecipitated DNA and one or more fractions eluted from the beads. In some embodiments, the DNA fractions are desalted and concentrated in preparation for the enzymatic step of library preparation. I. Pooling DNA from at least a first and a second subsample or portion thereof
[0352] In some embodiments, for example, after the partitioning step, the method includes preparing a pool containing at least a portion of the DNA of the second sub-sample (also referred to as the hypomethylated fraction) and at least a portion of the DNA of the first sub-sample (also referred to as the hypermethylated fraction). Target regions, e.g., target regions containing epigenetic target regions and / or sequence-variable target regions, can be enriched from the pool. The step of enriching for a set of target regions from at least a portion of the sub-samples described elsewhere herein encompasses an enrichment step performed on a pool containing DNA from the first and second sub-samples. A step of amplifying the DNA in the pool may be performed before enriching for target regions from the pool. The enrichment step can have any of the features described elsewhere herein.
[0353] Epigenetic target regions may exhibit differences in methylation levels and / or fragmentation patterns depending on whether they originate from tumors or healthy cells, or the type of tissue from which they originate, as discussed elsewhere herein. Sequence-variable target regions may exhibit sequence differences depending on whether they originate from tumors or healthy cells.
[0354] In some applications, analyzing epigenetic target regions from hypomethylated fractions may provide less information than analyzing sequence variable target regions from hypermethylated and hypomethylated fractions, and epigenetic target regions from hypermethylated fractions.Therefore, in methods in which sequence variable target regions and epigenetic target regions are enriched, the latter may be enriched to a lower degree than one or more of sequence variable target regions from hypermethylated and hypomethylated fractions, and epigenetic target regions from hypermethylated fractions.For example, sequence variable target regions may be enriched from a portion of hypomethylated fractions that are not pooled with hypermethylated fractions, and the pool may be prepared using a portion (e.g., a large proportion, substantially all, or all) of DNA from hypermethylated fractions and none or a portion (e.g., a small proportion) of DNA from hypomethylated fractions.This approach may reduce or eliminate sequencing of epigenetic target regions from hypomethylated fractions, thereby reducing the amount of sequencing data that is sufficient for further analysis.
[0355] In some embodiments, including a small percentage of DNA from the hypomethylated fraction in the pool facilitates quantification of one or more epigenetic traits (e.g., methylation, or other epigenetic traits discussed in detail elsewhere herein), e.g., in terms of relative bias.
[0356] In some embodiments, the pool contains a small proportion of hypomethylated DNA, e.g., less than about 50% of the hypomethylated DNA, e.g., less than or equal to about 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, or 5% of the hypomethylated DNA. In some embodiments, the pool contains about 5%-25% of the hypomethylated DNA. In some embodiments, the pool contains about 10%-20% of the hypomethylated DNA. In some embodiments, the pool contains about 10% of the hypomethylated DNA. In some embodiments, the pool contains about 15% of the hypomethylated DNA. In some embodiments, the pool contains about 20% of the hypomethylated DNA.
[0357] In some embodiments, the pool comprises a portion of the hypermethylated fraction, which may be at least about 50% of the DNA in the hypermethylated fraction. For example, the pool may comprise at least about 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, or 95% of the DNA in the hypermethylated fraction. In some embodiments, the pool comprises 50-55%, 55-60%, 60-65%, 65-70%, 70-75%, 75-80%, 80-85%, 85-90%, 90-95%, or 95-100% of the DNA in the hypermethylated fraction. In some embodiments, the second pool comprises all or substantially all of the hypermethylated fraction.
[0358] In some embodiments, the method includes preparing a first pool comprising at least a portion of DNA from a low methylation fraction. In some embodiments, the method includes preparing a second pool comprising at least a portion of DNA from a high methylation fraction. In some embodiments, the first pool further comprises a portion of DNA from a high methylation fraction. In some embodiments, the second pool further comprises a portion of DNA from a low methylation fraction. In some embodiments, the first pool comprises a large proportion of DNA from a low methylation fraction and, optionally, a small proportion of DNA from a high methylation fraction. In some embodiments, the second pool comprises a large proportion of DNA from a high methylation fraction and a small proportion of DNA from a low methylation fraction. In some embodiments comprising an intermediate methylation fraction, the second pool comprises at least a portion of DNA from an intermediate methylation fraction, e.g., a large proportion of DNA from an intermediate methylation fraction. In some embodiments, the first pool comprises a large proportion of DNA from a low methylation fraction and the second pool comprises a large proportion of DNA from a high methylation fraction and a large proportion of DNA from an intermediate methylation fraction.
[0359] In some embodiments, the method includes enriching for at least a first set of target regions from a first pool, e.g., the first pool is as described in any of the above embodiments. In some embodiments, the first set includes sequence variable target regions. In some embodiments, the first set includes hypomethylated variable target regions and / or fragmented variable target regions. In some embodiments, the first set includes sequence variable target regions and fragmented variable target regions. In some embodiments, the first set includes sequence variable target regions, hypomethylated variable target regions, and fragmented variable target regions. Amplifying the DNA in the first pool may be performed prior to this enrichment step. In some embodiments, enriching for the first set of target regions from the first pool includes contacting the DNA of the first pool with a first set of target-specific probes. In some embodiments, the first set of target-specific probes includes target-binding probes specific for sequence variable target regions. In some embodiments, the first set of target-specific probes includes target-binding probes specific for sequence variable target regions, hypomethylated variable target regions, and / or fragmented variable target regions.
[0360] In some embodiments, the method includes enriching for a second set of target regions or a plurality of sets of target regions from the second pool, e.g., the first pool is as described in any of the above embodiments. In some embodiments, the second plurality of sets includes epigenetic target regions, e.g., hypermethylated variable target regions and / or fragmented variable target regions. In some embodiments, the second plurality of sets includes sequence-variable target regions and epigenetic target regions, e.g., hypermethylated variable target regions and / or fragmented variable target regions. Amplifying the DNA in the second pool may be performed before this enrichment step. In some embodiments, enriching for the second plurality of sets of target regions from the second pool includes contacting the DNA of the first pool with a second set of target-specific probes, the second set of target-specific probes including target-binding probes specific for sequence-variable target regions and target-binding probes specific for epigenetic target regions. In some embodiments, the first set of target regions and the second set of target regions are not identical. For example, the first set of target regions may include one or more target regions that are not present in the second set of target regions. Alternatively or additionally, the second set of target regions may include one or more target regions that are not present in the first set of target regions. In some embodiments, at least one hypermethylated variable target region is enriched from the second pool but not from the first pool. In some embodiments, multiple hypermethylated variable target regions are enriched from the second pool but not from the first pool. In some embodiments, the first set of target regions includes a sequence variable target region and / or the second set of target regions includes an epigenetic target region. In some embodiments, the first set of target regions includes a sequence variable target region and a fragmentation variable target region, and the second set of target regions includes an epigenetic target region, e.g., a hypermethylated variable target region and a fragmentation variable target region.In some embodiments, the first set of target regions includes sequence variable target regions, fragmented variable target regions, and includes hypomethylated variable target regions, and the second set of target regions includes epigenetic target regions, e.g., hypermethylated variable target regions and fragmented variable target regions.
[0361] In some embodiments, the first pool contains a large proportion of DNA from the hypomethylated fraction and a portion (e.g., about half) of the DNA from the hypermethylated fraction, and the second pool contains a portion (e.g., about half) of the DNA from the hypermethylated fraction. In some such embodiments, the first set of target regions includes sequence variable target regions and / or the second set of target regions includes epigenetic target regions. The sequence variable target regions and / or epigenetic target regions may be as described in any of the embodiments described elsewhere herein. J. Providing Combined Subsamples
[0362] Some embodiments of the present disclosure include combining DNA from at least a first and a second sub-sample, wherein the at least a first and a second sub-sample are divided from a DNA sample. As described herein, dividing the DNA sample may include dividing the DNA into multiple sub-samples, which may include physically separating the DNA sample into two, three, four, five, six, seven, eight, nine, ten, or more than ten sub-samples. In some embodiments, at least a portion of the DNA from any combination of two, three, four, five, six, seven, eight, nine, ten, or more than ten sub-samples may be combined prior to sequencing to provide a combined sub-sample. In some embodiments, the combining step comprises physically combining at least a portion of the DNA of the first sub-sample, at least a portion of the DNA of the second sub-sample, at least a portion of the DNA of the third sub-sample, at least a portion of the DNA of the fourth sub-sample, at least a portion of the DNA of the fifth sub-sample, at least a portion of the DNA of the sixth sub-sample, at least a portion of the DNA of the seventh sub-sample, at least a portion of the DNA of the eighth sub-sample, at least a portion of the DNA of the ninth sub-sample, and / or at least a portion of the DNA of the tenth sub-sample, or any combination thereof, e.g., in the same container. The container (e.g., but not limited to, a pipette tip, a tube, a cuvette, a multi-well plate, a flow cell, or other container) may include any suitable container.
[0363] The DNA of the combined sub-samples can be sequenced, for example, in the same sequencing reaction, for example, in the same flow cell. In some embodiments, a first sub-sample that is partitioned, enriched, converted, and / or digested as described herein can be combined with a second sub-sample that (a) has not been partitioned, enriched, converted, and / or digested, or (b) has been partitioned, converted, and / or digested as described herein using the same or different partitioning, conversion, and / or digestion as the first sub-sample, thereby providing a combined sub-sample. For example, in some embodiments, before the sequencing step, the enriched DNA of the first sub-sample and the DNA of the second sub-sample are combined, and the DNA of the second sub-sample is not enriched for one or more sets of target regions of DNA, thereby providing a combined sub-sample. In certain embodiments, the DNA of a first sub-sample is subjected to a conversion step, e.g., any of the conversion steps described elsewhere herein, the converted DNA of the first sub-sample and the DNA of the second sub-sample are combined, and the DNA of the second sub-sample is not converted, thereby providing a combined sub-sample. In some embodiments, the DNA of the first sub-sample is subjected to a first conversion step, e.g., any of the conversion steps described elsewhere herein; the DNA of the second sub-sample is subjected to a second conversion step, e.g., any of the conversion steps described elsewhere herein (e.g., the second conversion step is different from the first conversion step), and the converted DNA of the first sub-sample and the converted DNA of the second sub-sample are combined, thereby providing a combined sub-sample.In further embodiments, the DNA of a first sub-sample is partitioned (e.g., to obtain a highly methylated fraction) and / or treated with a methylation-sensitive nuclease (e.g., a methylation-sensitive restriction enzyme), the partitioned and / or digested DNA of the first sub-sample is combined with the DNA of a second sub-sample, and optionally the DNA of the second sub-sample is converted, e.g., according to any of the conversion steps described elsewhere herein, and further optionally the DNA of the second sub-sample is not digested, thereby providing a combined sub-sample. In some embodiments, the DNA of the first sub-sample is digested, e.g., with a methylation-sensitive nuclease (e.g., a methylation-sensitive restriction enzyme), the DNA of the second sub-sample is digested, e.g., with a methylation-dependent nuclease (e.g., a methylation-dependent restriction enzyme), and the digested DNA of the first sub-sample and the digested DNA of the second sub-sample are combined, thereby providing a combined sub-sample. In further embodiments, the DNA of a first sub-sample is subjected to a partitioning step (e.g., to obtain a hypermethylated fraction), the partitioned DNA of the first sub-sample (e.g., the hypermethylated fraction) is combined with the DNA of a second sub-sample, and the DNA of the second sub-sample is not partitioned, thereby providing a combined sub-sample. In some embodiments, the DNA of the first sub-sample is subjected to a partitioning step (e.g., to obtain a hypermethylated fraction), and the DNA of the second sub-sample is subjected to a partitioning step (e.g., to obtain a hypomethylated fraction), and the partitioned DNA of the first sub-sample (e.g., the hypermethylated fraction) and the partitioned DNA of the second sub-sample (e.g., the hypomethylated fraction) are combined, thereby providing a combined sub-sample. In some embodiments, DNA of a first sub-sample is subjected to a partitioning step (e.g., to obtain a highly methylated fraction), and DNA of a second sample is subjected to a conversion step, where optionally the partitioning step includes a methylation-based partitioning step, and the converting step includes bisulfite conversion, and the partitioned DNA of the first sub-sample (e.g., the highly methylated fraction) is combined with the converted DNA of the second sub-sample, thereby providing a combined sub-sample.In some embodiments, the DNA of a first sub-sample is digested, e.g., digested with a methylation-sensitive nuclease (e.g., a methylation-sensitive restriction enzyme), the DNA of a second sub-sample is subjected to a conversion step, e.g., any of the conversion steps described elsewhere herein, e.g., bisulfite conversion, and the digested DNA of the first sub-sample and the converted DNA of the second sub-sample are combined, thereby providing a combined sub-sample. In certain embodiments, the DNA of the first sub-sample is subjected to a partitioning step (e.g., to obtain a highly methylated fraction) and digestion, e.g., digestion with a methylation-sensitive nuclease (e.g., a methylation-sensitive restriction enzyme), the DNA of the second sample is subjected to a conversion step, e.g., any of the conversion steps described elsewhere herein, e.g., bisulfite conversion, and the partitioned and digested DNA of the first sub-sample and the converted DNA of the second sub-sample are combined, thereby providing a combined sub-sample. In some embodiments, at least one sub-sample of the combined sub-samples is not enriched for one or more sets of target regions of DNA (and may or may not be subjected to another treatment described herein, e.g., a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA). The DNA of such a sub-sample may be sequenced (as part of the combined sub-sample) to assess, for example, one or more broad genome-wide signatures (e.g., global hypomethylation, somatic variation, and / or fragment mix signatures). Any of the above exemplary combinations may further include at least a portion of DNA from one or more additional sub-samples and / or DNA from additional samples or sources. In this manner, the methods disclosed herein may enable multiplexed sequencing analysis. Generally, prior to any combining step, the DNA of different sub-samples may be differentially tagged, e.g., by attaching an adapter to DNA containing a tag indicating the sub-sample from which the DNA was present. The adapter may further include additional elements, e.g., a barcode, as described elsewhere herein.
[0364] In some embodiments, the combined sub-sample comprises 80-100%, e.g., 80-95%, 85-100%, or 85-95%, of the DNA from the first sub-sample. In some embodiments, the combined sub-sample comprises 0%, 0.5-5%, 1-10%, 5-20%, 15-30%, 25-40%, 35-50%, 45-60%, 55-70%, 65-80%, 75-90%, or 85-100% of the DNA from the first sub-sample. In some embodiments, the combined sub-sample contains less than or equal to about 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, 5%, or 1% of the DNA from the first sub-sample. In some embodiments, the combined sub-sample comprises 0%, 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% of the DNA from the first sub-sample.
[0365] In some embodiments, the combined sub-sample comprises 0.5-20%, e.g., 0.5-15%, 5-20%, or 5-15%, of the DNA from the first sub-sample, hi some embodiments, the combined sub-sample comprises 0.5-5%, 1-10%, 5-20%, 15-30%, 25-40%, 35-50%, 45-60%, 55-70%, 65-80%, 75-90%, or 85-100% of the DNA from the second sub-sample. In some embodiments, the combined sub-sample contains less than or equal to about 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, 5%, or 1% of the DNA from the second sub-sample. In some embodiments, the combined sub-sample comprises 0%, 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% of the DNA from the second sub-sample.
[0366] In some embodiments, the combined sub-sample contains 0.5-5%, 1-10%, 5-20%, 15-30%, 25-40%, 35-50%, 45-60%, 55-70%, 65-80%, 75-90%, or 85-100% of the DNA from the third sub-sample. In some embodiments, the combined sub-sample contains less than or equal to about 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, 5%, or 1% of the DNA from the third sub-sample. In some embodiments, the combined sub-sample comprises 0%, 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% of the DNA from the third sub-sample.
[0367] In certain embodiments, the combined sub-sample comprises substantially all of the DNA from the first sub-sample, in certain embodiments, the combined sub-sample comprises substantially all of the DNA from the second sub-sample, in certain embodiments, the combined sub-sample comprises substantially all of the DNA from the third sub-sample.
[0368] In some embodiments, the DNA of at least the first and second sub-samples is not combined prior to sequencing. In some embodiments, the DNA of multiple sub-samples is not combined prior to sequencing. K. Detect; Sequencing
[0369] In some embodiments, detecting the presence or absence of DNA sequence and / or modification comprises sequencing.In some embodiments, DNA is sequenced in a manner that distinguishes first nucleic acid base from second nucleic acid base.Generally, the sample nucleic acid that comprises the nucleic acid flanked by adaptor can be subjected to sequencing with or without prior amplification. Sequencing methods include, for example, Sanger sequencing, high-throughput sequencing, pyrosequencing, sequencing-by-synthesis, long-read sequencing (also known as single-molecule sequencing or third-generation sequencing), nanopore sequencing (a type of long-read sequencing), five-letter or six-letter sequencing, semiconductor sequencing, sequencing-by-ligation, sequencing-by-hybridization, Digital Gene Expression (Helicos), next-generation sequencing (NGS), single-molecule sequencing-by-synthesis (SMSS) (Helicos), enzymatic methyl sequencing (EM-Seq), Tet-assisted pyridine borane sequencing (TAPS), massively parallel sequencing, Clonal Single Molecule Array (Solexa), shotgun sequencing, Ion Torrent, Oxford Nanopore, Roche Genia, Maxam-Gilbert sequencing, primer walking, and other sequencing methods available from PacBio, SOLiD, Ion Examples include sequencing using the Torrent or Nanopore platforms.
[0370] The sequencing reactions can be performed in a variety of sample processing units, which may include multiple lanes, multiple channels, multiple wells, or other means for processing multiple sets of samples substantially simultaneously. The sample processing units may also include multiple sample chambers capable of processing multiple runs simultaneously. In some embodiments, sequencing includes detecting and / or distinguishing between unmodified and modified nucleobases. For example, long-read sequencing (also referred to herein as third-generation sequencing) methods include methods that can generate longer sequencing reads, such as reads longer than 10 kilobases, compared to short-read sequencing methods, which generally generate reads up to about 600 bases in length. Compared to short reads, long reads can improve de novo assembly, identification of transcript isoforms, and detection and / or mapping of structural variants. Furthermore, long-read sequencing of native DNA or RNA molecules reduces amplification bias and preserves base modifications, such as methylation status. Long-read sequencing technologies useful herein may include, but are not limited to, any suitable long-read sequencing method, including Pacific Biosciences (PacBio) single-molecule real-time (SMRT) sequencing, Oxford Nanopore Technologies (ONT) nanopore sequencing, and synthetic long-read sequencing approaches, such as concatenated reads, proximal ligation strategies, and optical mapping. The synthetic long read approach involves the assembly of short reads from the same DNA molecule to generate synthetic long reads and can be used together with "true" long read sequencing technologies, such as SMRT and nanopore sequencing methods.
[0371] Single-molecule real-time (SMRT) sequencing, for example, facilitates the direct detection of 5-methylcytosine and 5-hydroxymethylcytosine, as well as unmodified cytosine (Weirather JL, et al., "Comprehensive comparison of Pacific Biosciences and Oxford Nanopore Technologies and their applications to transcriptome analysis," F1000Research, 6:100, 2017). While next-generation sequencing methods detect increased signal from clonal populations of amplified DNA fragments, SMRT sequencing captures single DNA molecules and preserves base modifications during sequencing. The error rate of raw PacBio SMRT sequencing-generated data is approximately 13–15%, due to the lack of a high signal-to-noise ratio from single DNA molecules. To increase accuracy, this platform uses a circular DNA template by ligating hairpin adapters to both ends of the target double-stranded DNA. As the polymerase repeatedly traverses and replicates the circular molecule, the DNA template is sequenced multiple times, generating continuous long reads (CLRs). CLRs can be split into multiple reads ("sub-reads") by removing adapter sequences, which generate circular consensus sequence ("CCS") reads with higher accuracy. The average length of a CLR is >10 kb and up to 60 kb, depending on the polymerase lifetime. Thus, the length and accuracy of the CCS read depend on the fragment size. PacBio sequencing has been utilized for genome (e.g., de novo assembly, structural variant detection, and haplotyping) and transcriptome (e.g., gene isoform reconstruction and novel gene / isoform discovery) studies.
[0372] ONT is a nanopore-based single-molecule sequencing technology (Weirather JL, et al., F1000Research, 6:100, 2017). ONT directly sequences native single-stranded DNA (ssDNA) molecules by measuring characteristic current changes as bases are threaded through a nanopore by molecular motor proteins. ONT uses a hairpin library structure similar to the PacBio circular DNA template, where the DNA template and its complement bind to a hairpin adapter. Thus, the DNA template passes through the nanopore, followed by the hairpin, and finally the complement. The raw read can be separated into two "1D" reads ("template" and "complement") by removing the adapter. The consensus sequence of the two "1D" reads is a "2D" read with higher accuracy.
[0373] Five-letter and six-letter sequencing methods include whole-genome sequencing methods that can sequence A, C, T, and G, in addition to 5mC and 5hmC, in a single workflow to provide a five-letter (A, C, T, G, and either 5mC or 5hmC) or six-letter (A, C, T, G, 5mC, and 5hmC) digital readout. DNA sample processing is entirely enzymatic, avoiding the DNA degradation and genome coverage bias of bisulfite treatment. In an exemplary five-letter sequencing method developed by Cambridge Epigenetix, sample DNA is first fragmented via sonication and then ligated to short synthetic DNA hairpin adapters at both ends (Fuellgrabe, et al. 2022, bioRxiv doi: https: / / doi.org / 10.1101 / 2022.07.08.499285). The construct is then split to separate the sense and antisense sample strands. For each original sample strand, a complementary copy strand is synthesized by DNA polymerase extension at the 3' end to create a hairpin construct with the original sample DNA strand connected to its complementary strand lacking epigenetic modifications via a synthesis loop. A sequencing adapter is then ligated to the end. Modified cytosines are enzymatically protected. Unprotected Cs are then deaminated to uracil, which is then read as thymine. In any such embodiment, the amplification method may include a uracil- and / or dihydrouracil-resistant amplification method, such as PCR, using a uracil- and / or dihydrouracil-resistant DNA polymerase (i.e., a DNA polymerase capable of reading and amplifying templates containing uracil and / or dihydrouracil bases). The deaminated construct is no longer fully complementary and has substantially reduced duplex stability; thus, the hairpin can be easily opened and amplified by PCR. The construct can be sequenced in a paired-end format, whereby read 1 (P1 primed) is the original strand and read 2 (P2 primed) is the copy strand. The reads are aligned pairwise, such that read 1 is aligned to its complementary read 2.Cognate residues from both reads are computationally decoded to generate a single genetic or epigenetic signature. Cognate base pairings that differ from the five allowed are the result of imperfect fidelity at some stage, including errors in base calling during sample preparation, amplification, or sequencing. These errors occur independently with cognate bases on each strand, resulting in substitutions that result in ineligible pairs. The ineligible pairs are masked (marked as N) within the decoded read, while the read itself is retained, resulting in minimal information loss and high accuracy at the read level. The decoded reads are aligned to the reference genome. Genetic variant and methylation counts are generated by read counting at the base level.
[0374] 5hmC has been shown to be valuable as a marker of biological states and diseases, including early cancer detection from cell-free DNA. In adapting five-letter sequences to six-letter sequencing, 5mC disambiguates 5hmC without compromising the calling of genetic bases within the same sample fragment. The first three steps of the workflow are identical to the five-letter sequencing described above, creating an adapter to which the sample fragment with the synthetic copy strand is ligated. Methylation at 5mC is enzymatically copied across the CpG units to the copy strand, while 5hmC is enzymatically protected from such copying. Thus, the 5mC and 5hmC of the unmodified C in each original CpG unit are distinguished by a unique two-base combination. Next, unmodified cytosine is deaminated to uracil, which is then read as thymine. The DNA is subjected to PCR amplification and sequencing as previously described. The reads are aligned pairwise and decoded using a two-base code. Since the three CpG units are separate sequencing environments of the two-base code, each of the unmodified C, 5mC, and 5hmC can be decoded.
[0375] In some embodiments, sequencing comprises targeted sequencing, in which one or more genomic regions of interest are sequenced.In some such embodiments, the genomic region of interest comprises a region present in one or more genes selected from Tables 1, 2, 3, 4, and / or 5.In some such embodiments, DNA sequences that do not contain a region of interest are not sequenced.Some embodiments comprise untargeted sequencing, for example, all genomic regions of the DNA in the processed sample or sub-sample are sequenced, or genomic regions are randomly selected for sequencing.In other embodiments, detecting the presence or absence of a sequence of DNA comprises sequencing DNA that is not enriched for the genomic region of interest (untargeted sequencing), for example, detectable sequences are obtained in a substantially unbiased manner.
[0376] Sequencing reactions can be performed on one or more types of nucleic acids, such as nucleic acids known to contain cancer or other disease markers.Sequencing reactions can also be performed on any nucleic acid fragments present in a sample.In some embodiments, the sequence coverage of the genome can be less than 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, 99.9%, or 100%.In some embodiments, sequencing reactions can provide at least 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, or 80% sequence coverage of the genome. Sequence coverage can be performed for at least 5, 10, 20, 70, 100, 200, or 500 different genes, or for at most 5000, 2500, 1000, 500, or 100 different genes.
[0377] Simultaneous sequencing reaction can be carried out using multiplex sequencing.In some embodiments, cell-free nucleic acid can be sequenced by at least 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, 100,000 sequencing reactions.In other embodiments, cell-free nucleic acid can be sequenced by less than 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, 100,000 sequencing reactions.Sequencing reaction can be carried out sequentially or simultaneously.Subsequent data analysis can be carried out for all or part of sequencing reaction. In some examples, data analysis can be performed for at least 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, or 100,000 sequencing reactions. In other examples, data analysis can be performed for less than 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, or 100,000 sequencing reactions. An exemplary read depth is 1000 to 50,000 reads per locus (base).
[0378] In some embodiments, sequences that are not cleaved during the degradation step are sequenced, hi some embodiments, less than 50%, 40%, 30%, 20%, 10%, 5%, 4%, 3%, 2%, or 1% of the sequences that are cleaved during the degradation step are sequenced. L. Subject
[0379] In some embodiments, DNA (e.g., cfDNA or DNA from a sample containing cells) is obtained from a subject (e.g., a test subject) with cancer or precancer. In some embodiments, the subject has stage I cancer, stage II cancer, stage III cancer, or stage IV cancer. In some embodiments, DNA from a subject is obtained and / or derived from a sample obtained from the subject. In some embodiments, DNA is obtained from a subject suspected of having cancer or precancer. In some embodiments, DNA is obtained from a subject with a tumor. In some embodiments, DNA is obtained from a subject suspected of having a tumor. In some embodiments, DNA is obtained from a subject with a neoplasm. In some embodiments, DNA is obtained from a subject suspected of having a neoplasm. In some embodiments, DNA is obtained from a subject in remission from a tumor, cancer, or neoplasm (e.g., after chemotherapy, surgical resection, radiation, or a combination thereof). In any of the foregoing embodiments, the precancer, cancer, tumor, or neoplasm, or suspected precancer, cancer, tumor, or neoplasm, may be a precancer, cancer, tumor, or neoplasm of the bladder, head and neck, lung, colon, rectum, kidney, breast, prostate, skin, or liver. In some embodiments, the precancer, cancer, tumor, or neoplasm, or suspected precancer, cancer, tumor, or neoplasm, is a lung precancer, cancer, tumor, or neoplasm. In some embodiments, the precancer, cancer, tumor, or neoplasm, or suspected precancer, cancer, tumor, or neoplasm, is a colon or rectal precancer, cancer, tumor, or neoplasm. In some embodiments, the precancer, cancer, tumor, or neoplasm, or suspected precancer, cancer, tumor, or neoplasm, is a breast precancer, cancer, tumor, or neoplasm. In some embodiments, the precancer, cancer, tumor, or neoplasm, or suspected precancer, cancer, tumor, or neoplasm, is a prostate precancer, cancer, tumor, or neoplasm. In any of the foregoing embodiments, the subject may be a human subject. In any of the foregoing embodiments, the subject may be a test subject. M. Sample
[0380] The sample may be any biological sample isolated from a subject. The sample may be a bodily sample. The sample may include bodily tissues or fluids, such as known or suspected solid tumors (e.g., carcinomas, adenocarcinomas, or sarcomas), whole blood, platelets, serum, plasma, feces, red blood cells, white blood cells or leucocytes, endothelial cells, tissue biopsies, cerebrospinal fluid, synovial fluid, lymphatic fluid, ascites, interstitial or extracellular fluid, interstitial space fluid, dental crevicular fluid, bone marrow, pleural effusion, pleural fluid, cerebrospinal fluid, saliva, mucus, sputum, semen, sweat, and urine. The sample is preferably a bodily fluid, particularly blood and its fractions, cerebrospinal fluid, pleural fluid, saliva, sputum, or urine. The sample may be in the original form isolated from the subject, or may have been further processed to remove or add components, such as cells, or to enrich one component for another.
[0381] In some embodiments, the nucleic acid population is obtained from serum, plasma, or blood samples from subjects suspected of having a neoplasm, tumor, precancer, or cancer, or from subjects already diagnosed with a neoplasm, tumor, precancer, or cancer. The population includes nucleic acids with various levels of sequence mutations, epigenetic mutations, post-translational chromatin modifications (PTM), and / or post-replicative or post-transcriptional modifications. Post-replicative modifications include cytosine modifications, particularly modifications at the 5th position of the nucleic acid base, such as 5-methylcytosine, 5-hydroxymethylcytosine, 5-formylcytosine, and 5-carboxylcytosine.
[0382] A sample can be isolated or obtained from a subject and transported to a site for sample analysis. The sample can be stored and shipped at a desired temperature, for example, room temperature, 4°C, -20°C, and / or -80°C. The sample can be isolated or obtained from a subject at a site for sample analysis. The subject can be a human, mammal, animal, companion animal, service animal, or pet. The subject can have cancer, precancer, infection, transplant rejection, or other disease or disorder associated with an altered immune system. The subject can have no cancer or detectable symptoms of cancer. The subject can have been treated with one or more cancer therapies, such as any one or more of chemotherapy, antibodies, vaccines, or biologics. The subject can be in remission. The subject may or may not be diagnosed with cancer or a predisposition to any cancer-associated genetic mutation / disorder.
[0383] In some embodiments, the sample comprises plasma. The volume of plasma obtained can depend on the desired read depth of the region being sequenced. Exemplary volumes are 0.4-40 mL, 5-20 mL, 10-20 mL, and 3-5 mL. For example, the volume can be 0.5 mL, 1 mL, 2 mL, 3 mL, 4 mL, 5 mL, 6 mL, 7 mL, 8 mL, 9 mL, 10 mL, 20 mL, 30 mL, or 40 mL. The volume of sampled plasma can be 5-20 mL. In some embodiments, the sample volume is 3-5 mL of plasma, e.g., 4 mL of plasma, per 10 mL of whole blood.
[0384] In some embodiments, the sample comprises whole blood. Exemplary volumes of sampled whole blood are 0.4 to 40 mL, 5 to 20 mL, 10 to 20 mL, 1 to 6 mL, 1 to 3 mL, and 3 to 5 mL. For example, the volume can be 0.5 mL, 1 mL, 2 mL, 3 mL, 4 mL, 5 mL, 6 mL, 7 mL, 8 mL, 9 mL, 10 mL, 20 mL, 30 mL, or 40 mL. The volume of sampled whole blood can be 5 to 20 mL. In some embodiments, the sample volume is 1 to 5 mL of whole blood, e.g., 2.5 mL of whole blood.
[0385] In some embodiments, the sample comprises a buffy coat separated from whole blood. Exemplary volumes of the sampled buffy coat are 0.1-20 mL, 1-10 mL, 1-5 mL, 0.2-0.6 mL, and 0.3-0.5 mL. For example, the volume can be 0.1 mL, 0.2 mL, 0.3 mL, 0.4 mL, 0.5 mL, 0.6 mL, 0.7 mL, 0.8 mL, 0.9 mL, 1 mL, 2 mL, 3 mL, 4 mL, 5 mL, 10 mL, or 20 mL. The volume of the sampled buffy coat can be 1-10 mL. In some embodiments, the sample volume is 0.1-0.5 mL of buffy coat, e.g., 0.3 mL of buffy coat, per 10 mL of whole blood.
[0386] In some embodiments, the sample contains PBMCs isolated from whole blood. Exemplary volumes of sampled PBMCs are 0.1-20 mL, 1-10 mL, 1-5 mL, 0.2-0.6 mL, and 0.3-0.5 mL. For example, volumes can be 0.1 mL, 0.2 mL, 0.3 mL, 0.4 mL, 0.5 mL, 0.6 mL, 0.7 mL, 0.8 mL, 0.9 mL, 1 mL, 2 mL, 3 mL, 4 mL, 5 mL, 10 mL, or 20 mL. The volume of sampled PBMCs can be 1-10 mL. In some embodiments, the sample volume is 0.1-0.5 mL of PBMCs, e.g., 0.3 mL of PBMCs, per 10 mL of whole blood.
[0387] In some embodiments, the sample contains leukocytes separated from the subject's blood using leukocyte reduction. Exemplary volumes of sampled leukocytes from leukocyte reduction are 0.1-20 mL, 1-10 mL, 1-5 mL, 0.2-0.6 mL, and 0.3-0.5 mL. For example, the volume can be 0.1 mL, 0.2 mL, 0.3 mL, 0.4 mL, 0.5 mL, 0.6 mL, 0.7 mL, 0.8 mL, 0.9 mL, 1 mL, 2 mL, 3 mL, 4 mL, 5 mL, 10 mL, or 20 mL. The volume of sampled leukocytes from leukocyte reduction can be 1-10 mL. In some embodiments, the sample volume is 0.1-0.6 mL of leukocytes from leukocyte reduction, e.g., 0.4 mL of leukocytes, per 10 mL of whole blood.
[0388] A sample can contain various amounts of nucleic acid, including genome equivalents. For example, a sample of about 30 ng of DNA contains about 10,000 (10 4 ) haploid human genome equivalents. Similarly, a sample of about 100 ng of DNA can contain about 30,000 haploid human genome equivalents.
[0389] A sample may contain nucleic acids from different sources, e.g., from cells of the same subject or from cells of different subjects. A sample may contain nucleic acids with mutations. For example, a sample may contain DNA with germline mutations and / or somatic mutations. A germline mutation refers to a mutation present in a subject's germline DNA. A somatic mutation refers to a mutation originating in a subject's somatic cells, e.g., precancerous or cancerous cells. A sample may contain DNA with a cancer-associated mutation (e.g., a cancer-associated somatic mutation). A sample may contain epigenetic variants (i.e., chemical or protein modifications), which are associated with the presence of a genetic variant, such as a cancer-associated mutation. In some embodiments, a sample that does not contain a genetic variant contains an epigenetic variant associated with the presence of the genetic variant.
[0390] Exemplary amounts of nucleic acid in a sample prior to amplification (e.g., DNA from a buffy coat sample or any other sample containing cells, such as a blood sample (e.g., a whole blood sample, a leukoreduction sample, or a PBMC sample)) range from about 1 fg to about 1 μg, e.g., 1 pg to 200 ng, 1 ng to 100 ng, or 10 ng to 1000 ng. For example, the amount may be up to about 600 ng, up to about 500 ng, up to about 400 ng, up to about 300 ng, up to about 200 ng, up to about 100 ng, up to about 50 ng, or up to about 20 ng of nucleic acid molecules. The amount may be at least 1 fg, at least 10 fg, at least 100 fg, at least 1 pg, at least 10 pg, at least 100 pg, at least 1 ng, at least 10 ng, at least 150 ng, or at least 200 ng of nucleic acid molecules. The amount may be up to 1 femtogram (fg), 10 fg, 100 fg, 1 picogram (pg), 10 pg, 100 pg, 1 ng, 10 ng, 100 ng, 150 ng, or 200 ng of nucleic acid molecule. The method may include obtaining between 1 femtogram (fg) and 200 ng.
[0391] Nucleic acids can be isolated from cells, e.g., cells in body fluids. The cells can be lysed, and the cellular nucleic acids can be processed. Generally, after adding a buffer and washing steps, the nucleic acids can be precipitated with alcohol. Additional purification steps, such as silica-based columns to remove contaminants or salts, can also be used. In some aspects of the procedure, e.g., to optimize yield, non-specific bulk carrier nucleic acids, such as C1 DNA, DNA, or proteins for bisulfite sequencing, hybridization, and / or ligation, can be added to the entire reaction.
[0392] After such processing, the sample may contain various forms of nucleic acids, including double-stranded DNA, single-stranded DNA, and single-stranded RNA. In some embodiments, single-stranded DNA and RNA are converted to double-stranded forms, which can be included in subsequent processing and analysis steps.
[0393] Reference or control molecule can be added or added to sample as control or normalization standard.For example, a certain amount of modified DNA from a species other than the target species from which sample is obtained or synthetic nucleic acid that contains a certain modification can be added to sample.In some embodiments, reference or control molecule can be distinguished from the molecule that originally exists in sample.In some embodiments, detected DNA sequence is normalized to reference or control molecule.
[0394] In some embodiments, multiple first sub-samples (e.g., from different subjects and / or tagged with distinct sample tags) are pooled prior to sequencing. This approach can reduce costs, for example, in that fewer reagents may be required per sub-sample processed.
[0395] DNA molecules can be ligated with adapters at either one or both ends. Typically, double-stranded molecules are blunt-ended by treatment with a polymerase containing a 5'-3' polymerase and a 3'-5' exonuclease (or proofreading function) in the presence of all four standard nucleotides. Klenow large fragment and T4 polymerase are examples of suitable polymerases. The blunt-ended DNA molecule can be ligated to an at least partially double-stranded adapter (e.g., a Y-shaped or bell-shaped adapter). Alternatively, complementary nucleotides can be added to the blunt ends of the sample nucleic acid and adapter to facilitate ligation. Both blunt-end and sticky-end ligation are contemplated herein. In blunt-end ligation, both the nucleic acid molecule and the adapter tag have blunt ends. In sticky-end ligation, the nucleic acid molecule typically has an "A" overhang and the adapter has a "T" overhang. N.Analysis
[0396] The present disclosure provides methods for analyzing combined sub-samples of DNA, such as those described herein. In some embodiments, the DNA of multiple sub-samples is not combined before sequencing. In some embodiments, the disclosed methods include analyzing DNA (e.g., DNA from a subject) to identify at least one cell type, cell cluster type, tissue type, and / or cancer type from which one or more type-specific epigenetic target regions and / or type-specific sequence variable target regions originated. In some embodiments, the methods include determining the level of one or more type-specific epigenetic target regions and / or type-specific sequence variable target regions originating from at least one cell type, cell cluster type, tissue type, and / or cancer type.
[0397] An exemplary method for analyzing DNA includes the following steps (e.g., in the order listed below), which is illustrated in FIG. 1A: 1. Preparing an extracted DNA sample (e.g., extracting DNA, e.g., cfDNA, from a human sample, e.g., a blood sample). 2. Performing either steps 2A-1 and 2A-2 or steps 2B-1 and 2B-2: 2A-1. Subjecting the extracted DNA to library preparation, including attaching tags that may include barcodes. 2A-2. After library preparation, subjecting the DNA to a procedure that affects a first nucleobase in the DNA so that it differs from a second nucleobase in the DNA, e.g., any form of epigenetic base conversion described elsewhere herein, e.g., bisulfite conversion. 2B-1. Subjecting the extracted DNA to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA, e.g., any form of epigenetic base conversion described elsewhere herein, e.g., bisulfite conversion. 2B-2. After the procedure that affects a first nucleobase in the DNA so that it is different from a second nucleobase in the DNA, the DNA is subjected to library preparation and a tag that may comprise a barcode is attached. 3. Dividing the DNA library into a plurality of sub-samples, including a first sub-sample and a second sub-sample. 4. Enriching for nucleic acid molecules containing sequences present in one or more target region sets (e.g., an epigenetic target region set and / or a sequence variable target region set, each of which may include any one or more of the epigenetic target region sets and sequence variable target region sets described elsewhere herein), and optionally amplifying the enriched DNA. 5. Combining at least a portion of the DNA of at least two of the sub-samples to provide a combined sub-sample. 6. Optionally, sequence the DNA of the combined sub-samples via NGS along with libraries from other samples (e.g., in the same NGS pool (e.g., flow cell)). 7. Performing bioinformatics analysis of the NGS data, including using one or more tags to identify unique molecules, molecules from specific sub-samples, and / or molecules from specific DNA samples.
[0398] Another exemplary method for analyzing DNA includes the following steps (e.g., in the order listed below), which is illustrated in FIG. 1B: 1. Preparing an extracted DNA sample (e.g., extracting DNA, e.g., cfDNA, from a human sample, e.g., a blood sample). 2. Partitioning the sample into a first sub-sample, eg, a hypermethylated fraction, and a second sub-sample, eg, a hypomethylated fraction. 3. Subjecting the DNA of the second sub-sample, and optionally the DNA of the first sub-sample, to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA, e.g., any form of epigenetic base conversion described elsewhere herein, e.g., bisulfite conversion. 3. Subjecting each of the DNA of the first sub-sample and the DNA of the second sub-sample to library preparation and attaching a tag, which may comprise a barcode. 4. Optionally, subjecting the DNA of the first sub-sample to digestion with a methylation-sensitive restriction enzyme. 5. Amplifying the DNA of the first sub-sample and the DNA of the second sub-sample. 6. Enriching the DNA of the first sub-sample for nucleic acid molecules comprising sequences present in a set of target regions, e.g., a set of epigenetic target regions (e.g., specific for hypermethylated and / or hypomethylated differentially methylated regions (DMRs)), using at least one of the plurality of sub-samples, and optionally amplifying the enriched DNA. 6. Combining at least a portion of the enriched DNA of the first sample and the DNA of the second sub-sample to provide a combined sub-sample. 7. Optionally, sequence the DNA of the combined sub-samples via NGS along with libraries from other samples (e.g., in the same NGS pool (e.g., flow cell)). 7. Performing bioinformatics analysis of the NGS data, including using one or more tags to identify unique molecules, molecules from specific sub-samples, and / or molecules from specific DNA samples.
[0399] In some embodiments, detecting the presence, level, or absence of DNA sequences and / or modifications facilitates the diagnosis of disease or the identification of appropriate treatments. In some embodiments, the presence or alteration of the level of one or more sequences and / or modifications indicates the presence of a disease or disorder in a subject, such as cancer or precancer, or other disorder that causes changes in nucleic acids compared to healthy subjects.
[0400] Furthermore, the present methods can be used to diagnose the presence of a condition, particularly cancer or precancer, in a subject, characterize the condition (e.g., stage the cancer or determine the heterogeneity of the cancer), monitor the response to treatment of the condition, and determine a prognostic risk of progression of the condition or subsequent progression of the condition. In some embodiments, the condition is cancer or precancer. In some embodiments, the condition is characterized (e.g., stage the cancer or determine the heterogeneity of the cancer), the response to treatment of the condition is monitored, or the prognostic risk of progression of the condition or subsequent progression of the condition is determined. The present disclosure can also be useful for determining the effectiveness of certain treatment options. A successful treatment option may reduce the amount of detected DNA sequences associated with cancer in the subject's blood, as fewer cancer cells may shed DNA. In other examples, this may not occur. In another example, certain treatment options may correlate with the genetic profile of the cancer over time. This correlation may be useful for selecting a therapy.
[0401] Additionally, if the cancer is observed to be in remission after treatment, the method can be used to monitor for residual disease or recurrence of the disease.
[0402] The types and number of cancers that can be detected can include blood cancer, brain cancer, lung cancer, skin cancer, nasal cancer, pharyngeal cancer, liver cancer, bone cancer, lymphoma, pancreatic cancer, skin cancer, intestinal cancer, rectal cancer, colon cancer, prostate cancer, thyroid cancer, bladder cancer, head and neck cancer, kidney cancer, oral cancer, stomach cancer, solid tumors, heterogeneous tumors, homogeneous tumors, and the like. The type and / or stage of cancer can be detected by genetic variations, including mutations, rare mutations, indels, copy number variations, transversions, translocations, recombinations, inversions, deletions, aneuploidy, partial aneuploidy, polyploidy, chromosomal instability, changes in chromosomal structure, gene fusions, chromosomal fusions, gene truncations, gene amplifications, gene duplications, chromosomal damage, DNA damage, abnormal changes in chemical modifications of nucleic acids, abnormal changes in epigenetic patterns, and abnormal changes in nucleic acid 5-methylcytosine.
[0403] In some embodiments, the methods described herein include detecting the presence or absence of nucleic acid, such as DNA, produced by a tumor (or neoplastic or cancerous cell) or by a precancerous cell.
[0404] The information and data generated by the methods disclosed herein can also be used to characterize specific forms of cancer. Cancers are often heterogeneous in both their composition and stage classification. The methods disclosed herein can enable characterization of specific subtypes of cancer, which can be important in diagnosing or treating that specific subtype. This information can also provide subjects or clinicians with clues regarding the prognosis of a particular type of cancer, allowing them to tailor treatment options as the disease progresses. Some cancers may progress to become more aggressive and genetically unstable. Other cancers may remain benign, inactive, or dormant. The systems and methods disclosed herein can be useful in determining disease progression.
[0405] Furthermore, the methods of the present disclosure can be used to characterize the heterogeneity of a condition in a subject. Such methods can include, for example, generating an aggregation profile of extracellular nucleic acids derived from a subject, where the aggregation profile includes multiple data obtained from various nucleic acid analyses. In some embodiments, the aggregation profile includes epigenetic and mutational analyses. In some embodiments, the aggregation profile includes the sum of information derived from different cells in a heterogeneous disease. This sum can include structural variation identifiers and levels, copy number variations, epigenetic variations, or other mutation analyses.
[0406] This method can be used to diagnose, prognose, monitor or observe pre-cancer, cancer or other diseases.In some embodiments, the method herein does not include the diagnosis, prognosis or monitoring of fetus, and therefore is not intended for non-invasive prenatal testing.In other embodiments, these methodologies can be used in pregnant subjects to diagnose, prognose, monitor or observe cancer or other diseases in unborn subjects whose DNA and other polynucleotides may circulate with maternal molecules. III. ADDITIONAL FEATURES OF CERTAIN DISCLOSED METHODS A. Target Region Set
[0407] In some embodiments, a specific genomic region of interest is detected and / or enriched. The genomic region of interest may comprise one or more target region sets. In some embodiments, the target region set comprises variations that are not present in DNA from healthy subjects or DNA obtained from healthy tissue regions. In some embodiments, the target region set comprises variations that are present in healthy cells but are not normally present in a sample type, such as a blood sample. In some embodiments, the variations are present in abnormal cells (e.g., hyperplastic, dysplastic, or neoplastic cells). Exemplary target region sets include sequence-variable target region sets and epigenetic target region sets.
[0408] In some embodiments, a first set of target regions is detected, the first set comprising at least epigenetic target regions. In some embodiments, the epigenetic target regions detected in the first subsample comprise hypermethylated variable target regions. In some embodiments, the hypermethylated variable target regions are CpG-containing regions that are unmethylated or hypomethylated (e.g., below average methylation compared to bulk cfDNA) in cfDNA from healthy subjects. In some embodiments, the hypermethylated variable target regions exhibit type-specific hypermethylation in healthy cfDNA from one or more related cell types or tissue types. Without wishing to be bound by any particular theory, the presence of cancer cells may increase the shedding of DNA into the bloodstream (e.g., from cancer and / or surrounding tissues). Therefore, the distribution of tissues of origin of cfDNA may change during carcinogenesis. Thus, an increased level of hypermethylated variable target regions in the first subsample may be an indicator of the presence of cancer (or recurrence, depending on the subject's medical history).
[0409] In some embodiments, the method herein includes detecting a second set of captured target regions from a sample or a second sub-sample containing at least epigenetic target regions. In some embodiments, the second set of epigenetic target regions includes hypomethylated variable target regions. In some embodiments, the hypomethylated variable target regions are CpG-containing regions that are methylated or have high methylation (e.g., above average methylation compared to bulk cfDNA) in cfDNA from healthy subjects. Without wishing to be bound by any particular theory, cancer cells may shed more DNA into the bloodstream than healthy cells of the same tissue type. Therefore, the distribution of tissues of origin of cfDNA may change during carcinogenesis. Thus, an increase in the level of hypomethylated variable target regions in the second sub-sample may be an indicator of the presence of cancer (or recurrence, depending on the subject's medical history).
[0410] In addition, the set of target regions can include DNA corresponding to a set of sequence-variable target regions. 1. Epigenetic target region set
[0411] In some embodiments, the target region set is or comprises an epigenetic target region set.The epigenetic target region set can comprise one or more types of target region that may distinguish the DNA from neoplastic (e.g., tumor or cancer) cells from the DNA from healthy cells, such as non-neoplastic circulating cells.Exemplary types of such regions are discussed in detail herein.The epigenetic target region set can also comprise one or more control regions, for example, as described herein.
[0412] In some embodiments, the set of epigenetic target regions has a footprint of at least 100 kbp, e.g., at least 200 kbp, at least 300 kbp, or at least 400 kbp. In some embodiments, the set of epigenetic target regions has a footprint in the range of 100-20 Mbp, e.g., 100-200 kbp, 200-300 kbp, 300-400 kbp, 400-500 kbp, 500-600 kbp, 600-700 kbp, 700-800 kbp, 800-900 kbp, 900-1,000 kbp, 1-1.5 Mbp, 1.5-2 Mbp, 2-3 Mbp, 3-4 Mbp, 4-5 Mbp, 5-6 Mbp, 6-7 Mbp, 7-8 Mbp, 8-9 Mbp, 9-10 Mbp, or 10-20 Mbp. In some embodiments, the set of epigenetic target regions has a footprint of at least 20 Mbp. a. Hypermethylated and hypomethylated variable target regions
[0413] In some embodiments, the epigenetic target region set comprises a hypermethylated variable target region.In some embodiments, the hypermethylated variable target region is differentially or exclusively hypermethylated in one or more related cell types or tissue types.Such hypermethylated variable target region may be hypermethylated in other cell types or tissue types, but not to the same extent as observed in one or more related cell types or tissue types.In some embodiments, the hypermethylated variable target region shows even higher methylation in the cfDNA from diseased cells of one or more related cell types or tissue types.
[0414] In some embodiments, a hypermethylated variable target region refers to a region, e.g., in a cfDNA sample, where an increased observed methylation level indicates an increased likelihood that the sample (e.g., a cfDNA sample) contains DNA produced by neoplastic cells, e.g., tumor or cancer cells. For example, hypermethylation of promoters of tumor suppressor genes has been repeatedly observed. See, e.g., Kang et al., Genome Biol. 18:53 (2017) and references cited therein. In another example, as discussed above, a hypermethylated variable target region may include a region that does not necessarily differ in methylation in cancerous tissue compared to DNA from healthy tissue of the same type, but differs in methylation (e.g., has more methylation) compared to cfDNA typical of healthy subjects. For example, if the presence of cancer results in increased cell death, such as apoptosis, of cells of the tissue type corresponding to the cancer, such cancer can be detected, at least in part, using such a hypermethylated variable target region.
[0415] An extensive discussion of methylation variable target regions in colorectal cancer is provided in Lam et al., Biochim Biophys Acta. 1866:106-20 (2016). These include VIM, SEPT9, ITGA4, OSM4, GATA4, and NDRG4. An exemplary set of hypermethylated variable target regions based on studies of colorectal cancer (CRC) is provided in Table 1. Many of these genes likely have relevance to cancers other than colorectal cancer; for example, TP53 is a critical tumor suppressor, and it is widely recognized that hypermethylation-based inactivation of this gene may be a common mechanism of tumorigenesis.
[0416] [Table 1-1] [Table 1-2]
[0417]
[0418] In some embodiments, the genomic region targeted for sequencing includes multiple loci listed in Table 1, for example, at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% of the loci listed in Table 1. In some embodiments, the genomic region is captured using probes. For example, for each locus included as a target region, there may be one or more probes with hybridization sites that bind between the transcription start site of the gene and the stop codon (or the final stop codon in the case of an alternatively spliced gene) or in the promoter region of the gene. In some embodiments, the one or more probes bind within 300 bp, for example, within 200 or 100 bp, of the transcription start site of the gene in Table 1.
[0419] Methylation-variable target regions in various types of lung cancer are described, for example, in Ooki et al., Clin. Cancer Res. 23:7141-52 (2017); Belinksy, Annu. Rev. Physiol. 77:453-74 (2015); Hulbert et al., Clin. Cancer Res. 23:1998-2005 (2017); Shi et al., BMC Genomics 18:901 (2017); Schneider et al., BMC Cancer. 11:102 (2011); Lissa et al., Transl Lung Cancer Res 5(5):492-504 (2016); Skvortsova et al., Br. J. Cancer. 94(10):1492-1495 (2006); Kim et al., Cancer Res. 61:3419-3424. (2001); Furonaka et al., Pathology International 55:303-309 (2005); Gomes et al., Rev. Port. Pneumol. 20:20-30 (2014); Kim et al., Oncogene. 20:1765-70 (2001); Hopkins-Donaldson et al., Cell Death Different. (2003); Kikuchi et al., Clin. Cancer Res. 11:2954-61 (2005); Heller et al., Oncogene 25:959-968 (2006); Licchesi et al., Carcinogenesis. 29:895-904 (2008); Guo et al., Clin. Cancer Res. 10:7917-24 (2004); Palmisano et al., Cancer Res. 63:4620-4625 (2003); and Toyooka et al., Cancer Res. 61:4556-4560, (2001).
[0420] An exemplary set of hypermethylated variable target regions based on lung cancer studies is provided in Table 2. Many of these genes may also have relevance to cancers other than lung cancer; for example, Casp8 (caspase 8) is a key enzyme in programmed cell death, and hypermethylation-based inactivation of this gene may be a common tumorigenesis mechanism not limited to lung cancer. In addition, several genes appear in both Tables 1 and 2, indicating generality.
[0421] [Table 2]
[0422] Any of the foregoing embodiments relating to target regions identified in Table 2 may be combined with any of the above embodiments relating to target regions identified in Table 1. In some embodiments, the genomic regions targeted for sequencing include multiple loci listed in Table 1 or Table 2, e.g., at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% of the loci listed in Table 1 or Table 2.
[0423] Additional hypermethylated target regions may be obtained, for example, from the Cancer Genome Atlas. Kang et al., Genome Biology 18:53 (2017) describe the construction of a probabilistic method called CancerLocator using hypermethylated target regions from breast, colon, kidney, liver, and lung. In some embodiments, the hypermethylated target regions may be specific to one or more types of cancer. Thus, in some embodiments, the hypermethylated target regions include one, two, three, four, or five subsets of hypermethylated target regions that collectively exhibit hypermethylation in one, two, three, four, or five of breast cancer, colon cancer, kidney cancer, liver cancer, and lung cancer.
[0424] In some embodiments, the set of epigenetic target regions comprises hypomethylated variable target regions. In some embodiments, the hypomethylated variable target regions are exclusively hypomethylated in one or more related cell types or tissue types. Such hypomethylated variable target regions may be hypomethylated in other cell types or tissue types, but not to the same extent as observed in one or more related cell types or tissue types.
[0425] In some embodiments, when different epigenetic target regions are captured, the epigenetic target regions include hypermethylated variable target regions and / or hypomethylated variable target regions.
[0426] Additionally, exemplary hypermethylated and hypomethylated variable target regions useful for distinguishing between various cell types have been identified by analyzing DNA from various cell types via whole-genome bisulfite sequencing, as described, for example, in Scott, CA, Duryea, JD, MacKay, H. et al., "Identification of cell type-specific methylation signals in bulk whole genome bisulfite sequencing data," Genome Biol 21, 156 (2020) (doi.org / 10.1186 / s13059-020-02065-5). Whole-genome bisulfite sequencing data is available from the Blueprint Consortium, available online at dcc.blueprint-epigenome.eu. b.CTCF binding region
[0427] In some embodiments, the epigenetic target region set comprises CTCF binding region. CTCF is a DNA binding protein that contributes to chromatin organization and often co-localizes with cohesin. Perturbation of CTCF binding site has been reported in a variety of different cancers. For example, see Katainen et al., Nature Genetics, doi:10.1038 / ng.3335; Guo et al., Nat. Commun. 9:1520 (2018), published online on June 8, 2015. CTCF binding results in a recognizable pattern of cfDNA, which can be detected by sequencing, for example, through fragment length analysis. Thus, perturbation of CTCF binding results in fluctuations in the fragmentation pattern of cfDNA. Therefore, CTCF binding site is a type of fragmentation variable target region.
[0428] There are many known CTCF binding sites.For example, see CTCFBSDB (CTCF Binding Site Database), available on the Internet at insulatordb.uthsc.edu / ; Cuddapah et al., Genome Res.19:24-32 (2009); Martin et al., Nat. Struct. Mol. Biol.18:708-14 (2011); Rhee et al., Cell.147:1408-19 (2011), each of which is incorporated herein by reference.Exemplary CTCF binding sites are nucleotides 56014955-56016161 on chromosome 8 and nucleotides 95359169-95360473 on chromosome 13.
[0429] In some embodiments, the CTCF binding regions include at least 10, 20, 50, 100, 200, or 500 CTCF binding regions, or 10-20, 20-50, 50-100, 100-200, 200-500, or 500-1000 CTCF binding regions, such as those listed above or in the CTCFBSDB or one or more of the above-cited articles by Cuddapah et al., Martin et al., or Rhee et al. In some embodiments, at least a portion of the CTCF sites can be methylated or unmethylated, and the methylation status correlates with whether the cell is a cancer cell. In some embodiments, the set of epigenetic target regions includes regions at least 100 bp, at least 200 bp, at least 300 bp, at least 400 bp, at least 500 bp, at least 750 bp, or at least 1000 bp upstream and downstream of the CTCF binding site. C transcription start site
[0430] In some embodiments, the epigenetic target region set comprises variable transcription start sites.Transcription start sites may show perturbations in neoplastic cells.For example, the nucleosome organization at various transcription start sites in healthy cells of hematopoietic lineage contributes substantially to cfDNA in healthy individuals, but may differ from the nucleosome organization at those transcription start sites in neoplastic cells.This results in different cfDNA patterns, which can be detected by sequencing, as generally discussed in, for example, Snyder et al., Cell 164:57-68 (2016); WO2018 / 009723; and US20170211143A1.In another example, transcription start sites are not necessarily epigenetically different in cancerous tissues compared with DNA from the same type of healthy tissue, but may be epigenetically different (for example, in terms of nucleosome organization) compared with DNA typical in healthy subjects.Transcription start site perturbations also result in variations in the fragmentation pattern of cfDNA. Therefore, the transcription start site is also one type of fragmentation variable target region.
[0431] Human transcription start sites are available from the Database of Human Transcription Start Sites (DBTSS), available online at dbtss.hgc.jp, and are described in Yamashita et al., Nucleic Acids Res. 34(Database issue): D86-D89 (2006), incorporated herein by reference. In some embodiments, the transcription start sites include at least 10, 20, 50, 100, 200, or 500 transcription start sites, or 10-20, 20-50, 50-100, 100-200, 200-500, or 500-1000 transcription start sites, e.g., transcription start sites described in DBTSS. In some embodiments, at least a portion of the transcription start sites can be methylated or unmethylated, and the methylation status correlates with whether the cell is a cancer cell. In some embodiments, the set of epigenetic target regions includes regions at least 100 bp, at least 200 bp, at least 300 bp, at least 400 bp, at least 500 bp, at least 750 bp, at least 1000 bp upstream and downstream of the transcription start site. d. Local amplification
[0432] Although local amplification is somatic mutation, it can be detected by sequencing based on read frequency in a similar manner to the approach of detecting certain epigenetic changes, such as methylation changes.Therefore, the region that may show local amplification in cancer can be included in epigenetic target region set, and it can include one or more of AR, BRAF, CCND1, CCND2, CCNE1, CDK4, CDK6, EGFR, ERBB2, FGFR1, FGFR2, KIT, KRAS, MET, MYC, PDGFRA, PIK3CA and RAF1. e. Methylation control region
[0433] It may be useful to include a control region to facilitate data validation. In some embodiments, the set of epigenetic target regions includes a control region that is expected to be methylated or unmethylated in essentially all samples, regardless of whether the DNA is derived from cancer cells or normal cells. In some embodiments, the set of epigenetic target regions includes a control hypomethylated region that is expected to be hypomethylated in essentially all samples. In some embodiments, the set of epigenetic target regions includes a control hypermethylated region that is expected to be hypermethylated in essentially all samples. 2. Sequence-variable target region set
[0434] In some embodiments, the target region set is or includes a sequence variable target region set. The sequence variable target region set may include one or more types of target regions that may distinguish DNA from neoplastic (e.g., tumor or cancer) cells from DNA from healthy cells, e.g., non-neoplastic circulating cells. Exemplary types of such regions are discussed in detail herein. The sequence variable target region set may also include one or more control regions, for example, as described herein. In some embodiments, the sequence variable target region set includes multiple regions known to undergo somatic mutations in cancer. In some aspects, the sequence variable target region set targets multiple different genes or genomic regions ("panels") selected so that a predetermined proportion of subjects with cancer exhibits genetic variants or tumor markers in one or more different genes or genomic regions in the panel. The panel may be selected to limit the sequencing region to a fixed number of base pairs. The panel may be selected to sequence a desired amount of DNA. The panel may also be selected to achieve a desired depth of sequence reads. A panel can be selected to achieve a desired sequence read depth or sequence read coverage in terms of the amount of sequenced base pairs. A panel can be selected to achieve a theoretical sensitivity, specificity, and / or accuracy for detecting one or more genetic variants in a sample.
[0435] Examples of lists of genomic locations of interest can be found, for example, in Tables 3 and 4 herein. In some embodiments, the set of sequence variable target regions includes a portion of at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, or 70 of the genes in Table 3. In some embodiments, the set of sequence variable target regions includes a portion of at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, or 70 of the SNVs in Table 3. In some embodiments, the set of sequence variable target regions includes a portion of at least 1, at least 2, at least 3, at least 4, at least 5, or 6 of the fusions in Table 3. In some embodiments, the set of sequence variable target regions comprises at least a portion of at least one, at least two, or three indels from Table 3. In some embodiments, the set of sequence variable target regions comprises a portion of at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, or 73 of the genes from Table 4. In some embodiments, the set of sequence variable target regions comprises at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, or 73 of the SNVs from Table 4. In some embodiments, the set of sequence variable target regions includes at least 1, at least 2, at least 3, at least 4, at least 5, or some of 6 of the fusions in Table 4.In some embodiments, the set of sequence variable target regions comprises at least a portion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, or 18 of the indels in Table 4. Each of these genomic locations of interest can be identified as a scaffold region or a hotspot region for a given panel. Table 5 shows an example list of hotspot genomic locations of interest. In some embodiments, the set of sequence variable target regions includes a portion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, or at least 20 of the genes in Table 5. Each hotspot genomic region is described along with the associated gene, the chromosome on which it resides, the start and stop positions in the genome representing the locus, the base pair length of the locus, the exons covered by the gene, and important features of the given genomic region of interest (e.g., type of mutation). [Table 3] [Table 4] [Table 5-1] [Table 5-2]
[0436] Examples of lists of target regions of interest can also be found in WO2020 / 160414, e.g., Table 4. Additional examples include the loci disclosed in Gale et al., PLoS One 13: e0194630 (2018), incorporated herein by reference, which describes a panel of 35 cancer-associated gene targets: AKT1, ALK, BRAF, CCND1, CDK2A, CTNNB1, EGFR, ERBB2, ESR1, FGFR1, FGFR2, FGFR3, FOXL2, GATA3, GNA11, GNAQ, GNAS, HRAS, IDH1, IDH2, KIT, KRAS, MED12, MET, MYC, NFE2L2, NRAS, PDGFRA, PIK3CA, PPP2R1A, PTEN, RET, STK11, TP53, and U2AF1. In some embodiments, the set of sequence variable target regions includes target regions from at least 10, 20, 30, or 35 genes associated with cancer, such as those described herein and in WO2020 / 160414.
[0437] In some embodiments, the set of sequence variable target regions has a footprint of at least 50 kbp, e.g., at least 100 kbp, at least 200 kbp, at least 300 kbp, or at least 400 kbp. In some embodiments, the set of sequence variable target regions has a footprint in the range of 100 to 2,000 kbp, e.g., 100 to 200 kbp, 200 to 300 kbp, 300 to 400 kbp, 400 to 500 kbp, 500 to 600 kbp, 600 to 700 kbp, 700 to 800 kbp, 800 to 900 kbp, 900 to 1,000 kbp, 1 to 1.5 Mbp, or 1.5 to 2 Mbp. In some embodiments, the set of sequence variable target regions has a footprint of at least 2 Mbp. B. Collection of target-specific probes
[0438] In some embodiments, a collection of target-specific probes is used in the methods described herein. In some embodiments, the collection of target-specific probes includes target-binding probes specific for a set of sequence-variable target regions and target-binding probes specific for a set of epigenetic target regions. In some embodiments, the capture yield of the target-binding probes specific for the set of sequence-variable target regions is higher (e.g., at least two-fold higher) than the capture yield of the target-binding probes specific for the set of epigenetic target regions. In some embodiments, the collection of target-specific probes is configured to have a capture yield specific for the set of sequence-variable target regions that is higher (e.g., at least two-fold higher) than its capture yield specific for the set of epigenetic target regions.
[0439] In some embodiments, the capture yield of target binding probes specific for the set of sequence variable target regions is at least 1.25x, 1.5x, 1.75x, 2x, 2.25x, 2.5x, 2.75x, 3x, 3.5x, 4x, 4.5x, 5x, 6x, 7x, 8x, 9x, 10x, 11x, 12x, 13x, 14x, or 15x higher than the capture yield of target binding probes specific for the set of epigenetic target regions. In some embodiments, the capture yield of target binding probes specific for the set of sequence variable target regions is 1.25-1.5 fold, 1.5-1.75 fold, 1.75-2 fold, 2-2.25 fold, 2.25-2.5 fold, 2.5-2.75 fold, 2.75-3 fold, 3-3.5 fold, 3.5-4 fold, 4-4.5 fold, 4.5-5 fold, 5-5.5 fold, 5.5-6 fold, 6-7 fold, 7-8 fold, 8-9 fold, 9-10 fold, 10-11 fold, 11-12 fold, 13-14 fold, or 14-15 fold higher than the capture yield of target binding probes specific for the set of epigenetic target regions.
[0440] In some embodiments, the collection of target-specific probes is configured to have a capture yield specific for a set of sequence variable target regions that is at least 1.25x, 1.5x, 1.75x, 2x, 2.25x, 2.5x, 2.75x, 3x, 3.5x, 4x, 4.5x, 5x, 6x, 7x, 8x, 9x, 10x, 11x, 12x, 13x, 14x, or 15x higher than the capture yield of the set of epigenetic target regions. In some embodiments, the collection of target-specific probes is configured to have a capture yield specific for a set of sequence variable target regions that is 1.25-1.5 times, 1.5-1.75 times, 1.75-2 times, 2-2.25 times, 2.25-2.5 times, 2.5-2.75 times, 2.75-3 times, 3-3.5 times, 3.5-4 times, 4-4.5 times, 4.5-5 times, 5-5.5 times, 5.5-6 times, 6-7 times, 7-8 times, 8-9 times, 9-10 times, 10-11 times, 11-12 times, 13-14 times, or 14-15 times higher than its capture yield specific for the set of epigenetic target regions.
[0441] A collection of probes can be configured to provide higher capture yields for a set of sequence-variable target regions in a variety of ways, including enrichment, varying length and / or chemistry (e.g., to affect affinity), and combinations thereof. Affinity can be modulated by adjusting the length of the probe and / or by including nucleotide modifications as discussed below.
[0442] In some embodiments, the target-specific probes specific for the set of sequence variable target regions are present at a higher concentration than the target-specific probes specific for the set of epigenetic target regions, hi some embodiments, the concentration of target-binding probes specific for the set of sequence variable target regions is at least 1.25-fold, 1.5-fold, 1.75-fold, 2-fold, 2.25-fold, 2.5-fold, 2.75-fold, 3-fold, 3.5-fold, 4-fold, 4.5-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 11-fold, 12-fold, 13-fold, 14-fold, or 15-fold higher than the concentration of target-binding probes specific for the set of epigenetic target regions. In some embodiments, the concentration of target-binding probes specific for the set of sequence-variable target regions is 1.25-1.5x, 1.5-1.75x, 1.75-2x, 2-2.25x, 2.25-2.5x, 2.5-2.75x, 2.75-3x, 3-3.5x, 3.5-4x, 4-4.5x, 4.5-5x, 5-5.5x, 5.5-6x, 6-7x, 7-8x, 8-9x, 9-10x, 10-11x, 11-12x, 13-14x, or 14-15x higher than the concentration of target-binding probes specific for the set of epigenetic target regions. In such embodiments, concentration may refer to the average mass / volume concentration of the individual probes in each set.
[0443] In some embodiments, target-specific probes specific to a set of sequence-variable target regions have higher affinity for the target than target-specific probes specific to a set of epigenetic target regions. Affinity can be modulated by any method known to those skilled in the art, including using different probe chemistries. For example, certain nucleotide modifications, such as cytosine 5-methylation (in the context of a particular sequence), modifications that introduce heteroatoms at the 2' sugar position, and LNA nucleotides, can increase the stability of double-stranded nucleic acids, and oligonucleotides with such modifications have been shown to have relatively high affinity for their complementary sequences. See, for example, Severin et al., Nucleic Acids Res. 39: 8740-8751 (2011); Freier et al., Nucleic Acids Res. 25: 4429-4443 (1997); U.S. Patent No. 9,738,894. Additionally, longer sequence lengths generally provide increased affinity. Other nucleotide modifications, such as substitution of guanine with the nucleobase hypoxanthine, reduce affinity by reducing the amount of hydrogen bonding between the oligonucleotide and its complementary sequence. In some embodiments, target-specific probes specific for a set of sequence-variable target regions have modifications that increase affinity for their targets. In some embodiments, instead or in addition, target-specific probes specific for a set of epigenetic target regions have modifications that decrease affinity for their targets. In some embodiments, target-specific probes specific for a set of sequence-variable target regions have a longer average length and / or a higher average melting temperature than target-specific probes specific for a set of epigenetic target regions. These embodiments can be combined with each other and / or with concentration differences as discussed above to achieve a desired fold difference in capture yield, for example, any of the fold differences or ranges described above.
[0444] In some embodiments, the target-specific probe comprises a capture moiety. The capture moiety can be any of the capture moieties described herein, for example, biotin. In some embodiments, the target-specific probe is linked to a solid support, for example, covalently or ...
Claims
1. 1. A method for analyzing DNA, comprising: (a) dividing the DNA into a plurality of sub-samples, the plurality of sub-samples comprising a first sub-sample and a second sub-sample; (b) enriching for one or more sets of target regions of DNA from said first sub-sample, thereby providing enriched DNA of said first sub-sample, said one or more sets of target regions comprising a set of sequence variable target regions; (c) combining the enriched DNA of the first sub-sample and the DNA of the second sub-sample, thereby providing a combined sub-sample, wherein the DNA of the second sub-sample is not enriched for one or more sets of target regions of DNA; and (d) sequencing the DNA of the combined sub-samples. A method comprising:
2. 1. A method for analyzing DNA, comprising: (a) dividing the DNA into a plurality of sub-samples, the plurality of sub-samples comprising a first sub-sample and a second sub-sample; (b) enriching for one or more sets of target regions of DNA from said first sub-sample, thereby providing enriched DNA of said first sub-sample, said one or more sets of target regions comprising a set of epigenetic target regions; (c) combining the enriched DNA of the first sub-sample and the DNA of the second sub-sample, thereby providing a combined sub-sample, wherein the DNA of the second sub-sample is not enriched for one or more sets of target regions of DNA; and (d) sequencing the DNA of the combined sub-samples. A method comprising:
3. 10. The method of claim 1, wherein the one or more sets of target regions further comprise one or more sets of epigenetic target regions.
4. 3. The method of claim 2, wherein said one or more sets of target regions further comprise one or more sets of sequence-variable target regions.
5. 10. The method of any one of the preceding claims, wherein the method further comprises subjecting the DNA or sub-sample thereof to a procedure that affects a first nucleobase of the DNA differently from a second nucleobase of the DNA, wherein the first nucleobase is a modified or unmodified nucleobase, and the second nucleobase is a modified or unmodified nucleobase that is different from the first nucleobase, and wherein the first nucleobase and the second nucleobase have the same base-pairing specificity, and optionally wherein the step of subjecting the DNA or sub-sample thereof to a procedure that affects a first nucleobase of the DNA differently from the second nucleobase of the DNA is performed before dividing the DNA into multiple sub-samples.
6. 1. A method for analyzing DNA, comprising: (a) subjecting the DNA to a procedure that affects a first nucleobase of the DNA differently from a second nucleobase of the DNA, wherein the first nucleobase is a modified or unmodified nucleobase, the second nucleobase is a modified or unmodified nucleobase that is different from the first nucleobase, and the first nucleobase and the second nucleobase have the same base-pairing specificity; (b) dividing the DNA into a plurality of sub-samples, the plurality of sub-samples comprising a first sub-sample and a second sub-sample; (c) enriching for one or more sets of target regions of DNA from said first sub-sample, thereby providing enriched DNA of said first sub-sample, said one or more sets of target regions comprising one or more of a set of sequence variable target regions and a set of epigenetic target regions; (d) combining the enriched DNA of the first sub-sample and the DNA of the second sub-sample, thereby providing a combined sub-sample, wherein the DNA of the second sub-sample is not enriched for one or more sets of target regions of DNA; and (e) sequencing the DNA of the combined sub-samples. A method comprising:
7. 7. The method of any one of claims 1 to 6, wherein dividing the DNA into a plurality of sub-samples comprises distributing the DNA into a plurality of sub-samples such that the first sub-sample contains a higher proportion of DNA with cytosine modifications than the second sub-sample, or wherein at least one of the first sub-sample and the second sub-sample is distributed into a plurality of further sub-samples comprising at least the first further sub-sample and the second further sub-sample such that the first further sub-sample contains a higher proportion of DNA with cytosine modifications than the second further sub-sample.
8. the distributing step is performed before the sequencing; and a) prior to said enriching step for one or more sets of epigenetic and / or sequence variable target regions of DNA from said DNA; b) after said enriching step for one or more sets of epigenetic and / or sequence variable target regions of DNA from said DNA; c) prior to said step of subjecting said DNA of said first sub-sample to a procedure that affects a first nucleobase of said DNA differently from a second nucleobase of said DNA; d) after the step of subjecting the DNA of the first sub-sample to a procedure that affects a first nucleobase of the DNA differently from a second nucleobase of the DNA; or e) Any combination of (a) to (d). The method of claim 7, wherein the method is carried out by
9. 9. The method of any one of claims 1 to 8, further comprising contacting at least one sub-sample with at least one restriction enzyme prior to said enriching or sequencing, optionally wherein said contacting is performed before performing said procedure that affects a first nucleobase of said DNA differently from a second nucleobase of said DNA and / or optionally contacting said first sub-sample with said at least one restriction enzyme.
10. 1. A method for analyzing DNA, comprising: (a) dividing the DNA into a plurality of sub-samples, the plurality of sub-samples comprising a first sub-sample and a second sub-sample; (b) subjecting said DNA of at least said first sub-sample to a procedure that affects a first nucleobase of said DNA differently from a second nucleobase of said DNA, thereby providing converted DNA of said first sub-sample, wherein said first nucleobase is a modified or unmodified nucleobase, said second nucleobase is a modified or unmodified nucleobase different from said first nucleobase, and said first nucleobase and said second nucleobase have the same base-pairing specificity; (c) combining the converted DNA of the first sub-sample and the DNA of the second sub-sample, thereby providing a combined sub-sample; and (d) sequencing the DNA of the combined sub-samples. A method comprising:
11. 1. A method for analyzing DNA, comprising: (a) dividing the DNA into a plurality of sub-samples, the plurality of sub-samples comprising a first sub-sample and a second sub-sample; (b) contacting at least said first sub-sample with at least one restriction enzyme, thereby providing digested DNA of said first sub-sample; (c) combining the digested DNA of the first sub-sample and the DNA of the second sub-sample, thereby providing a combined sub-sample; and (d) sequencing the DNA of the combined sub-samples. A method comprising:
12. 12. The method of claim 10 or claim 11, further comprising contacting the second sub-sample with at least one restriction enzyme prior to the combining step, thereby providing digested DNA of the second sub-sample.
13. 12. The method of claim 10 or claim 11, further comprising, prior to said combining step, subjecting the DNA of said second sub-sample to a procedure that affects a first nucleobase of said DNA differently from a second nucleobase of said DNA, thereby providing converted DNA of said second sub-sample, wherein said first nucleobase is a modified nucleobase or an unmodified nucleobase, said second nucleobase is a modified nucleobase or an unmodified nucleobase that is different from said first nucleobase, and wherein said first nucleobase and said second nucleobase have the same base-pairing specificity.
14. 14. The method of claim 10, wherein prior to the combining step, the first sub-sample is contacted with at least one methylation-sensitive restriction enzyme, thereby generating hypermethylated DNA of the first sub-sample.
15. 15. The method of any one of claims 10 to 14, wherein the second sub-sample is contacted with at least one methylation-dependent restriction enzyme, thereby generating hypomethylated DNA of the second sub-sample.
16. 1. A method for analyzing DNA, comprising: (a) distributing at least a portion of the DNA into a plurality of sub-samples, including a first sub-sample and a second sub-sample, wherein the first sub-sample contains a higher proportion of DNA with cytosine modifications than the second sub-sample; (b) subjecting said DNA of at least said first sub-sample to a procedure that affects a first nucleobase of said DNA differently from a second nucleobase of said DNA, thereby providing converted DNA of said first sub-sample, wherein said first nucleobase is a modified or unmodified nucleobase, said second nucleobase is a modified or unmodified nucleobase different from said first nucleobase, and said first nucleobase and said second nucleobase have the same base-pairing specificity; (c) combining the converted DNA of the first sub-sample and the DNA of the second sub-sample, thereby providing a combined sub-sample; and (d) sequencing the DNA of the combined sub-samples. A method comprising:
17. 1. A method for analyzing DNA, comprising: (a) distributing at least a portion of the DNA into a plurality of sub-samples, including a first sub-sample and a second sub-sample, wherein the first sub-sample contains a higher proportion of DNA with cytosine modifications than the second sub-sample; (b) contacting at least said first sub-sample with at least one restriction enzyme, thereby providing digested DNA of said first sub-sample; (c) combining the digested DNA of the first sub-sample and the DNA of the second sub-sample, thereby providing a combined sub-sample; and (d) sequencing the DNA of the combined sub-samples. A method comprising:
18. 18. The method of claim 16 or claim 17, further comprising contacting the second sub-sample with at least one nuclease prior to the combining step, thereby providing digested DNA of the second sub-sample, optionally wherein the nuclease is a methylation-dependent nuclease, and further optionally wherein the methylation-dependent nuclease is a methylation-dependent restriction enzyme.
19. 18. The method of claim 16 or claim 17, further comprising, prior to said combining step, subjecting the DNA of said second sub-sample to a procedure that affects a first nucleobase of said DNA differently from a second nucleobase of said DNA, thereby providing converted DNA of said second sub-sample, wherein said first nucleobase is a modified nucleobase or an unmodified nucleobase, said second nucleobase is a modified nucleobase or an unmodified nucleobase that is different from said first nucleobase, and wherein said first nucleobase and said second nucleobase have the same base-pairing specificity.
20. 20. The method of any one of claims 16 to 19, wherein prior to the combining step, the first sub-sample is contacted with at least one methylation-sensitive restriction enzyme, thereby generating hypermethylated DNA of the first sub-sample.
21. 20. The method of any one of claims 16 to 19, wherein the second sub-sample is contacted with at least one methylation-dependent restriction enzyme, thereby generating hypomethylated DNA of the second sub-sample.
22. a) amplifying the DNA prior to dividing the DNA into a plurality of sub-samples; b) amplifying the DNA from said first sub-sample prior to enrichment for one or more sets of target regions of DNA; and / or c) amplifying the enriched DNA of the first sub-sample before combining the enriched DNA of the first sub-sample and the DNA of the second sub-sample.
10. The method of any one of the preceding claims, further comprising:
23. 23. The method of claim 22, wherein the amplifying step is performed (a) after the step of subjecting the DNA to a procedure that affects a first nucleobase of the DNA differently from a second nucleobase of the DNA; (b) after the step of contacting the at least first and / or second sub-sample with at least one restriction enzyme; (c) after enriching for one or more sets of target regions of DNA; or (d) any combination of (a)-(c).
24. 24. The method of claim 22 or claim 23, wherein the amplifying step comprises one or more of polymerase chain reaction, linear amplification, rolling circle amplification, ligase chain reaction, strand displacement amplification, nucleic acid sequence-based amplification, and sequence-based self-sustained replication.
25. The method of any one of claims 22 to 24, wherein the amplification comprises thermocycling amplification.
26. The method of any one of claims 22 to 24, wherein the amplification comprises isothermal amplification.
27. 10. The method of any one of the preceding claims, wherein the DNA comprises a barcode.
28. 10. The method of any one of the preceding claims, wherein the method comprises ligating an adapter comprising a barcode to the DNA prior to the sequencing step.
29. 29. The method of any one of claims 22 to 28, wherein the method comprises ligating an adapter comprising a barcode to the DNA prior to the amplifying step.
30. 30. The method of any one of claims 7 to 9 or 16 to 29, wherein the partitioning step comprises partitioning based on methylation level.
31. 31. The method of any one of claims 7-9 or 16-30, wherein the partitioning step comprises contacting the DNA with an agent that recognizes modified cytosines in the DNA, and wherein the first sub-sample comprises DNA with a higher proportion of modified cytosines than the second sub-sample.
32. 32. The method of claim 31 , wherein the agent that recognizes a modified nucleobase in the DNA is a methyl-binding reagent.
33. 33. The method of claim 32, wherein the methyl-binding reagent is a methyl-binding domain (MBD) protein or antibody.
34. 34. The method of claim 32 or claim 33, wherein the methyl-binding reagent is specific for one or more methylated nucleotide bases, and optionally the one or more methylated nucleotide bases are 5-methylcytosine.
35. The method of any one of claims 32 to 34, wherein the methyl-binding reagent is immobilized on a solid support.
36. 36. The method of any one of claims 7 to 9 or 16 to 35, wherein the partitioning step comprises immunoprecipitation of methylated DNA.
37. 37. The method of any one of claims 7-9 or 16-36, wherein the partitioning step comprises partitioning based on binding to a protein, optionally the protein is a methylated protein, an acetylated protein, an unmethylated protein, an unacetylated protein; and / or optionally the protein is a histone.
38. 38. The method of claim 37, wherein said distributing step comprises contacting said DNA with a binding reagent specific for said protein and immobilized on a solid support.
39. 39. The method of any one of claims 7 to 9 or 16 to 38, wherein a first dispensed sub-sample of the plurality of dispensed sub-samples is differentially tagged from a second dispensed sub-sample of the plurality of dispensed sub-samples.
40. 40. The method of any one of claims 5-10, 13-16, or 19-39, wherein the first nucleobase is an unmodified cytosine and the second nucleobase is a modified cytosine, optionally wherein the modified cytosine is 5-methylcytosine or 5-hydroxymethylcytosine.
41. 41. The method of any one of claims 5-10, 13-16, or 19-40, wherein said procedure affecting a first nucleobase of said DNA differently from a second nucleobase of said DNA chemically converts said first or second nucleobase such that the base-pairing specificity of said converted nucleobase is altered.
42. 42. The method of any one of claims 5-10, 13-16, or 19-41, wherein said procedure that affects a first nucleobase of said DNA differently from a second nucleobase of said DNA is a methylation-sensitive conversion.
43. 43. The method of claim 42, wherein the methylation-sensitive conversion is bisulfite conversion, oxidative bisulfite (Ox-BS) conversion, Tet-assisted bisulfite (TAB) conversion, APOBEC-linked epigenetic (ACE) conversion, or enzymatic conversion.
44. 44. The method of claim 43, wherein the Tet-assisted conversion further comprises a substituted borane reducing agent, optionally wherein the substituted borane reducing agent is 2-picoline borane, borane pyridine, tert-butylamine borane, or ammonia borane.
45. 45. The method of any one of claims 9, 11-15, or 17-44, wherein the at least one restriction enzyme is a methylation-sensitive restriction enzyme (MSRE).
46. 46. The method of claim 45, wherein the first sub-sample is contacted with the MSRE.
47. 45. The method of any one of claims 9, 11-15, or 17-44, wherein the at least one restriction enzyme is a methylation-dependent restriction enzyme (MDRE).
48. 48. The method of claim 47, wherein the second sub-sample is contacted with the MDRE.
49. 49. The method of any one of claims 9, 11-15, or 17-48, comprising contacting at least one sub-sample with at least two restriction enzymes prior to the enriching step or prior to the sequencing step, optionally wherein said contacting occurs before performing said procedure that affects a first nucleobase of said DNA differently from a second nucleobase of said DNA.
50. 50. The method of claim 49, wherein the at least two restriction enzymes comprise or consist of two or three restriction enzymes.
51. 51. The method of any one of claims 9, 11-15, or 17-50, wherein the at least one restriction enzyme or the at least two restriction enzymes are selected from the group consisting of FspEI, LpnPI, MspJI, SgeI, AatII, AccII, AciI, Aor13HI, Aor15HI, BspT104I, BssHII, BstUI, Cfr10I, ClaI, CpoI, Eco52I, HaeII, HapII, HhaI, Hin6I, HpaII, HpyCH4IV, MluI, MspI, NaeI, NotI, NruI, NsbI, PmaCI, Psp1406I, PvuI, SacII, SalI, SmaI, and SnaBI.
52. 52. The method of any one of claims 9, 11-15, or 17-51, further comprising the step of attaching one or more adaptors to at least one end of at least some of the DNA molecules in the plurality of distributed sets prior to said step of contacting at least one sub-sample with at least one restriction enzyme or at least two restriction enzymes.
53. 53. The method of Claim 52, wherein the one or more adaptors comprise at least one tag.
54. 54. The method of Claim 53, wherein the at least one tag comprises a molecular barcode.
55. 55. The method of any one of claims 52 to 54, wherein the one or more adaptors are resistant to digestion by a methylation-sensitive or methylation-dependent restriction enzyme.
56. the one or more adapters that are resistant to digestion by a methylation-sensitive restriction enzyme a) one or more methylated nucleotides, optionally wherein said methylated nucleotides comprise 5-methylcytosine and / or 5-hydroxymethylcytosine; b) one or more nucleotide analogs that are resistant to methylation-sensitive restriction enzymes; or c) a nucleotide sequence that is not recognized by a methylation-sensitive restriction enzyme 56. The method of claim 55, comprising:
57. 57. The method of any one of claims 10-56, further comprising enriching for one or more sets of target regions of DNA from at least said first sub-sample, thereby providing enriched DNA of said first sub-sample, wherein said one or more sets of target regions comprise one or more of a set of sequence variable target regions and a set of epigenetic target regions.
58. 58. The method of any one of claims 1-9 or 22-57, wherein said enriching step comprises contacting said DNA with target-specific probes specific for said one or more sets of epigenetic target regions and / or said one or more sets of sequence variable target regions.
59. 59. The method of any one of claims 2 to 9 or 22 to 58, wherein the set of epigenetic target regions comprises a set of hypermethylated variable target regions and / or a set of hypomethylated variable target regions.
60. 60. The method of any one of claims 2-9 or 22-59, wherein the set of epigenetic target regions comprises a set of fragmented variable target regions.
61. 61. The method of claim 60, wherein the set of fragmented variable target regions comprises a transcription start site region.
62. 62. The method of claim 60 or claim 61, wherein the set of fragmented variable target regions comprises a CTCF binding region.
63. 63. The method of any one of claims 2-9 or 22-62, wherein the set of epigenetic target regions comprises one or more type-specific epigenetic target regions.
64. 64. The method of claim 63, wherein the one or more type-specific epigenetic target regions comprise type-specific differentially methylated regions and / or type-specific fragments.
65. 64. The method of claim 63, wherein the one or more type-specific epigenetic target regions comprise type-specific hypomethylated regions and / or type-specific hypermethylated regions.
66. 66. The method of any one of claims 63-65, wherein the one or more type-specific epigenetic target regions comprise cell type-specific, cell cluster type-specific, tissue type-specific, and / or cancer type-specific epigenetic target regions.
67. the one or more type-specific epigenetic target regions: hypermethylated in immune cells compared to non-immune cell types present in blood samples; differentially methylated in the colon compared with other tissue types; differentially methylated in breast compared with other tissue types; differentially methylated in liver compared with other tissue types; differentially methylated in kidney compared with other tissue types; differentially methylated in the pancreas compared with other tissue types; differentially methylated in the prostate compared with other tissue types; is differentially methylated in skin compared to other tissue types; or Differentially methylated in bladder compared with other tissue types 67. The method of any one of claims 63 to 66, comprising a target region.
68. 68. The method of any one of claims 65-67, wherein the hypermethylated target region is methylated to an extent that is at least 10%, 20%, 30%, or at least 40% greater than the average methylation of the target region in the sample or compared to other cell or tissue types.
69. the one or more type-specific epigenetic target regions: a target region that is hypomethylated in a non-immune cell type present in the sample compared to the methylation level of the target region in a different cell or tissue type in the sample; a fragment that is specific for immune cells relative to non-immune cell types present in said sample; or Fragments specific to colon, lung, breast, liver, kidney, pancreas, prostate, skin, or bladder compared to other tissue types 69. The method of any one of claims 63 to 68, comprising:
70. 70. The method of any one of claims 63 to 69, wherein the level of the one or more type-specific epigenetic target regions is determined based on cell type or tissue type of origin.
71. 71. The method of claims 63-70, wherein the level of the one or more type-specific epigenetic target regions originating from one or more immune cells, non-immune cell types present in a blood sample, and / or colon, lung, breast, liver, kidney, prostate, skin, bladder, or pancreatic cells is determined.
72. 72. The method of any one of claims 63-71, further comprising identifying at least one cell type, cell cluster type, tissue type, and / or cancer type from which said one or more type-specific epigenetic target regions originated.
73. 73. The method of any one of claims 63 to 72, comprising determining the methylation level of the type-specific epigenetic target region.
74. 10. The method of any one of the preceding claims, wherein the DNA of the first sub-sample and the DNA of the second sub-sample are differentially tagged.
75. 10. The method of any one of the preceding claims, wherein the combined sub-sample comprises (a) at least a portion of the converted DNA of the first sub-sample or at least a portion of the enriched DNA of the first sub-sample, and (b) at least a portion of the DNA of the second sub-sample.
76. 10. The method of any one of the preceding claims, wherein the combined sub-sample further comprises at least a portion of the DNA of a third sub-sample.
77. 10. The method of any one of the preceding claims, wherein the combined sub-sample comprises less than or equal to about 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, 5%, or 1% of the DNA of the first sub-sample.
78. 10. The method of any one of the preceding claims, wherein the combined sub-sample comprises about 50-70%, 70-90%, about 75-85%, or about 80% of the DNA of the first sub-sample.
79. 10. The method of any one of the preceding claims, wherein the combined sub-sample comprises less than or equal to about 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, or 5%, or 1% of the DNA of the second sub-sample.
80. 10. The method of any one of the preceding claims, wherein the combined sub-sample comprises about 50-70%, 70-90%, about 75-85%, or about 80% of the DNA of the second sub-sample.
81. 81. The method of any one of claims 76-80, wherein the combined sub-sample comprises less than or equal to about 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, or 5%, or 1% of the DNA of the third sub-sample.
82. 82. The method of any one of claims 76-81, wherein the combined sub-sample comprises about 50-70%, 70-90%, about 75-85%, or about 80% of the DNA of the third sub-sample.
83. 10. The method of any one of the preceding claims, wherein the combined sub-sample comprises substantially all of the DNA of the first sub-sample.
84. 10. The method of any one of the preceding claims, wherein the combined sub-sample comprises substantially all of the DNA of the second sub-sample.
85. 85. The method of any one of claims 76 to 84, wherein the combined sub-sample comprises substantially all of the DNA of the third sub-sample.
86. 10. The method of any one of the preceding claims, wherein said sequencing the DNA of the combined sub-samples comprises sequencing the DNA in a manner that distinguishes the first nucleobase from the second nucleobase.
87. 10. The method of any one of the preceding claims, wherein the sequencing of the DNA of the combined sub-samples comprises sequencing at least a portion of the DNA of at least the first and second sub-samples in the same sequencing cell.
88. 10. The method of any one of the preceding claims, wherein the step of sequencing the DNA of the combined sub-samples comprises sequencing the DNA in a manner that is sensitive to modifications.
89. 89. The method of Claim 88, wherein said sequencing in a manner sensitive to an alteration comprises long-read sequencing.
90. 90. The method of Claim 88 or Claim 89, wherein said sequencing in a manner sensitive to an alteration comprises nanopore sequencing.
91. 89. The method of claim 88, wherein said sequencing in a manner sensitive to a modification comprises five-letter or six-letter sequencing.
92. 10. The method of any one of the preceding claims, wherein the step of sequencing the DNA of the combined sub-samples comprises next generation sequencing.
93. 10. The method of any one of the preceding claims, wherein the sequencing of the DNA of the combined sub-samples comprises generating a plurality of sequencing reads, and the method further comprises mapping the plurality of sequence reads to one or more reference sequences to generate mapped sequence reads, and processing the mapped sequence reads corresponding to the set of sequence variable target regions and the set of epigenetic target regions.
94. at least one epigenetic state in the DNA of the first sub-sample; and at least one fragmentation feature and / or at least one somatic variant in the DNA of the second sub-sample.
10. The method of any one of the preceding claims, comprising detecting:
95. 10. The method of claim 1, wherein the at least one epigenetic state is the methylation state of a region or nucleotide, and optionally the region is a hypermethylated or hypomethylated variable target region, or the nucleotide is in a hypermethylated or hypomethylated variable target region.
96. 97. The method of Claim 95 or 96, wherein the at least one somatic variant comprises a single nucleotide variation (SNV), a copy number variation (CNV), a gene fusion, or an insertion or deletion (indel).
97. 10. The method of any one of the preceding claims, wherein the DNA is from a blood sample and / or a tissue sample.
98. 98. The method of claim 97, wherein the blood sample is a whole blood sample, a plasma sample, a buffy coat sample, a leukoreduced sample, or a PBMC sample.
99. 10. The method of any one of the preceding claims, wherein the DNA is cell-free DNA.
100. 100. The method of claim 99, wherein the cell-free DNA is in an amount of 1 ng to 500 ng.
101. 10. The method of any one of the preceding claims, wherein the DNA and / or the sample is from a subject.
102. 102. The method of claim 101, wherein the subject is an animal.
103. The method of claim 101 or claim 102, wherein the subject is a human.
104. 104. The method of any one of claims 97 to 103, wherein the blood sample is fractionated prior to enrichment for at least one set of epigenetic target regions of DNA.
105. 105. The method of any one of claims 101 to 104, wherein the subject has or is at risk of having cancer.
106. 106. The method of any one of claims 101 to 105, further comprising determining the presence or status of cancer in the subject.
107. 107. The method of any one of claims 101 to 106, further comprising determining the likelihood that the subject has an infection.
108. 108. The method of any one of claims 101 to 107, further comprising determining the likelihood that the subject will have transplant rejection.