Methods involving methylation-preserving amplification with error correction
Methylation-preserving amplification methods enhance the sensitivity of liquid biopsies by maintaining methylation states during DNA amplification and sequencing, thereby improving cancer diagnosis through non-invasive methods.
Patent Information
- Application Number
- JP2025536228
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-12-22
- Filing Date
- 2023-12-20
- Publication Date
- 2025-12-25
AI Technical Summary
Current methods for analyzing cell-free DNA in liquid biopsies lack sensitivity and specificity in detecting epigenetic modifications, such as methylation, which are crucial for early cancer detection, and often require invasive biopsies.
Methylation-preserving amplification methods that include DNA methyltransferase, followed by sequencing and enrichment for epigenetic target regions, allowing for single-base resolution and increased sensitivity in detecting epigenetic modifications.
Enhances the detection of epigenetic changes in cell-free DNA, improving cancer diagnosis through non-invasive liquid biopsies by maintaining methylation states during amplification and enabling precise sequencing.
Smart Images

Figure 2025542260000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 476,914, filed December 22, 2022, which is incorporated herein by reference for all purposes.
[0002] FIELD OF THE INVENTION The present disclosure provides compositions and methods related to analyzing DNA, for example, cell-free DNA. In some embodiments, the DNA is DNA from a subject who has or is suspected of having cancer, and / or the DNA includes DNA from cancer cells. In some embodiments, the DNA is amplified using methylation-preserving amplification before sequencing. In some embodiments, the methylation-preserving amplification includes DNA methyltransferase, for example, DNA methyltransferase 1 (DNMT1). [Background technology]
[0003] Introduction and Abstract Cancer causes millions of deaths worldwide each year. Early detection of cancer can lead to improved outcomes, as early-stage cancers tend to be more susceptible to treatment.
[0004] Improperly controlled cell growth is a hallmark of cancer, typically resulting from the accumulation of genetic and epigenetic alterations, such as copy number variations (CNVs), single nucleotide variations (SNVs), gene fusions, insertions and / or deletions (indels), cytosine modifications (e.g., 5-methylcytosine, 5-hydroxymethylcytosine, and other more oxidized forms), and epigenetic variations involving the association of DNA with chromatin proteins and transcription factors. Thus, cancer can be manifested by non-sequence alterations, such as methylation. Examples of methylation changes in cancer include localized gain of DNA methylation in CpG islands at the TSSs of genes involved in normal growth control, DNA repair, cell cycle regulation, and / or cell differentiation. Hypermethylation can be associated with abnormal loss of transcriptional capacity of the involved genes and occurs at least as frequently as point mutations and deletions as a cause of altered gene expression. Furthermore, without wishing to be bound by any particular theory, cells in or around cancer or neoplasia may shed more DNA than cells of the same tissue type in healthy subjects. The DNA from such cells may be epigenetically different from the shed DNA of healthy subjects. Therefore, the distribution of epigenetically modified (e.g., methylated) DNA in a certain DNA sample, such as cell-free DNA (cfDNA), may change during carcinogenesis. Thus, sufficiently sensitive epigenetic (e.g., DNA methylation) profiling can be used to detect abnormal methylation in the DNA of a sample.
[0005] Biopsy represents a traditional approach to detecting or diagnosing cancer in which cells or tissue are extracted from a potential cancer site and analyzed for relevant phenotypic and / or genotypic traits. Biopsies have the disadvantage of being invasive.
[0006] Cancer detection based on the analysis of bodily fluids, such as blood ("liquid biopsy"), is an interesting alternative based on the observation that DNA from cancer cells is released into bodily fluids. Liquid biopsies are non-invasive (sometimes requiring only a blood draw). However, due to the low concentration and heterogeneity of cell-free DNA, developing accurate and sensitive methods for analyzing liquid biopsy material that provide detailed information about nucleic acid base modifications has been a challenge. The contribution of DNA from cells in or around the cancer or neoplasm to a sample can be relatively small compared to the contribution from other cells, and the DNA contributed from other cells can be useless regarding the cancer state. Isolating and processing a useful fraction of cell-free DNA for further analysis in liquid biopsy procedures is an important part of these methods.
[0007] Furthermore, current methods for cancer diagnostic assays of cell-free nucleic acids (e.g., cell-free DNA or cell-free RNA) may focus on detecting tumor-associated somatic variants, including single nucleotide variations (SNVs), copy number variations (CNVs), fusions, and indels (i.e., insertions or deletions), all of which are mainstream targets for liquid biopsies. There is growing evidence that non-sequence modifications, such as methylation status and fragmentomic signals in cell-free DNA, can provide information about the origin and disease level of cell-free DNA. In addition, different types of modifications, such as 5-methylation and 5-hydroxymethylation, can have different implications regarding the presence or absence of disease. Detailed knowledge of non-sequence modifications in cell-free DNA (e.g., when combined with somatic mutation calling) can improve the assessment of tumor status. Summary of the Invention [Problem to be solved by the invention]
[0008] Thus, there is a continuing need for improved methods and compositions for analyzing DNA, including cell-free DNA, for example, in liquid biopsies. [Means for solving the problem]
[0009] The present disclosure aims to fulfill the need for improved analysis of DNA, such as cell-free DNA, and / or provide other advantages. In some embodiments, the present disclosure provides single-base resolution sequencing methods that exhibit increased sensitivity (e.g., the ability to reliably detect epigenetic features in samples containing fewer molecules) and / or reduced molecular loss compared to other sequencing methods, for example, methods that include a pre-amplification epigenetic base conversion step.
[0010] The following exemplary embodiments are provided:
[0011] Embodiment 1 is a method for analyzing DNA, comprising: (a) performing methylation-preserving amplification of DNA, wherein the DNA comprises a barcode; and (b) sequencing the DNA in a manner that is sensitive to the modification and determining an epigenetic consensus sequence of the DNA associated with at least a portion of the barcode.
[0012] Embodiment 2 is a method for analyzing DNA, comprising: (a) performing methylation-preserving amplification of DNA, wherein the DNA comprises a barcode; (b) subjecting the DNA to a procedure that affects a first nucleobase of the DNA differently from a second nucleobase of the DNA, wherein the first nucleobase is a modified or unmodified nucleobase, the second nucleobase is a modified or unmodified nucleobase that is different from the first nucleobase, and the first nucleobase and the second nucleobase have the same base-pairing specificity; and (c) sequencing the DNA and determining an epigenetic consensus sequence of the DNA associated with at least a portion of the barcode.
[0013] Embodiment 3 is a method for analyzing DNA, comprising: (a) performing methylation-preserving amplification of DNA, wherein the DNA comprises a barcode; (b) enriching the DNA for one or more sets of epigenetic target regions of DNA, thereby providing enriched DNA; (c) sequencing the enriched DNA in a manner sensitive to the modification and determining an epigenetic consensus sequence of the enriched DNA associated with at least a portion of the barcode.
[0014] Embodiment 4 is a method for analyzing DNA, comprising: (a) performing linear methylation-preserving amplification of DNA; and (b) sequencing the DNA in a manner that is sensitive to the modification, and determining an epigenetic consensus sequence of the DNA.
[0015] Embodiment 5 is a method for analyzing DNA, comprising: (a) performing linear methylation-preserving amplification of DNA; (b) subjecting the DNA to a procedure that affects a first nucleobase of the DNA differently from a second nucleobase of the DNA, wherein the first nucleobase is a modified or unmodified nucleobase, the second nucleobase is a modified or unmodified nucleobase that is different from the first nucleobase, and the first nucleobase and the second nucleobase have the same base-pairing specificity; and (c) sequencing the DNA and determining an epigenetic consensus sequence of the DNA.
[0016] Embodiment 6 is a method for analyzing DNA, comprising: (a) performing linear methylation-preserving amplification of DNA; (b) enriching the DNA for one or more sets of epigenetic target regions of DNA, thereby providing enriched DNA; (c) sequencing the enriched DNA in a manner sensitive to the modification and determining the epigenetic consensus sequence of the enriched DNA. The method includes:
[0017] Embodiment 7 is the method of any one of embodiments 1, 3, 4, or 6, comprising subjecting the DNA to a procedure that affects a first nucleobase of the DNA differently from a second nucleobase of the DNA, wherein the first nucleobase is a modified or unmodified nucleobase, the second nucleobase is a modified or unmodified nucleobase that is different from the first nucleobase, and wherein the first nucleobase and the second nucleobase have the same base-pairing specificity.
[0018] Embodiment 8 is the method of any one of the preceding embodiments, wherein the DNA comprises a barcode.
[0019] Embodiment 8.1 is the method of embodiment 8, wherein the barcode is a non-unique barcode.
[0020] Embodiment 9 is the method of any one of the preceding embodiments, wherein the method comprises ligating an adapter comprising a barcode to the DNA prior to sequencing.
[0021] Embodiment 10 is the method of any one of the preceding embodiments, wherein the method comprises ligating an adapter comprising a barcode to the DNA before amplifying the DNA.
[0022] Embodiment 11 is the method of any one of the preceding embodiments, wherein the method comprises ligating an adapter comprising a barcode to the DNA before performing said methylation-preserving amplification of the DNA.
[0023] Embodiment 12 is a method for analyzing DNA, comprising: (a) performing methylation-preserving amplification of DNA, thereby providing amplified DNA, wherein the DNA comprises an insert and adapters comprising barcodes, at least one of the adapters further comprising a restriction enzyme cleavage site between the barcode and a portion of the adapter, and the barcode is located between the insert and the restriction enzyme cleavage site; (b) contacting the amplified DNA with a restriction enzyme that recognizes and cleaves the DNA at the restriction enzyme cleavage site in the adapter; (c) before or after step (b), subjecting the amplified DNA to a procedure that affects a first nucleobase of the amplified DNA to be different from a second nucleobase of the amplified DNA, wherein the first nucleobase is a modified or unmodified nucleobase, the second nucleobase is a modified or unmodified nucleobase that is different from the first nucleobase, and the first nucleobase and the second nucleobase have the same base-pairing specificity; (d) after steps (b) and (c), ligating auxiliary adapters to the amplified DNA; (e) after step (d), performing uracil and / or dihydrouracil resistant amplification of the DNA; (f) after step (e), enriching the amplified DNA for one or more sets of epigenetic target regions of DNA, thereby providing enriched DNA; (g) optionally further amplifying the enriched DNA; and (h) sequencing the enriched DNA and determining an epigenetic consensus sequence of the DNA associated with at least a portion of the barcode.
[0024] Embodiment 13 is the method of embodiment 12, wherein the uracil- and / or dihydrouracil-resistant amplification of DNA comprises PCR using a uracil- and / or dihydrouracil-resistant DNA polymerase.
[0025] Embodiment 14 is the method of embodiment 12 or embodiment 13, wherein an optional step (step (g)) of amplifying the enriched DNA is performed.
[0026] Embodiment 15 is the method of embodiment 14, wherein amplifying the enriched DNA further comprises differentially tagging the enriched DNA.
[0027] Embodiment 16 is the method of embodiment 15, wherein differentially tagging the enriched DNA comprises attaching one or more sample indexes to the DNA.
[0028] Embodiment 17 is the method of any one of embodiments 1 to 3 or 7 to 16, wherein the barcode does not contain cytosines in a non-CpG context.
[0029] Embodiment 18 is the method of any one of embodiments 1 to 10 or 17, wherein the DNA comprises an insert and an adapter comprising a barcode, at least one of the adapters further comprising a restriction enzyme cleavage site between the barcode and a portion of the adapter, and the barcode is located between the insert and the enzyme cleavage site.
[0030] Embodiment 18.1 is the method of any one of the preceding embodiments, wherein the barcodes are molecular barcodes, and wherein the molecular barcodes distinguish between different DNA molecules in the same sample.
[0031] Embodiment 18.2 is the method of embodiment 18.1, wherein the molecular barcode is added to the DNA molecule, and optionally the molecular barcode is added by ligation.
[0032] Embodiment 19 is the method of any one of embodiments 18 to 18.2, further comprising contacting the amplified DNA with a restriction enzyme that recognizes and cleaves the DNA at a restriction site in the adapter.
[0033] Embodiment 20 is the method of embodiment 19, further comprising ligating an auxiliary adaptor to the amplified DNA after the steps of (a) contacting the amplified DNA with a restriction enzyme that recognizes and cleaves DNA at the restriction enzyme cleavage site in the adaptor, and before or after the steps of (b) subjecting the amplified DNA to a procedure that affects a first nucleobase of the amplified DNA to be different from a second nucleobase of the amplified DNA, wherein the first nucleobase is a modified or unmodified nucleobase, the second nucleobase is a modified or unmodified nucleobase that is different from the first nucleobase, and the first nucleobase and the second nucleobase have the same base-pairing specificity.
[0034] Embodiment 21 is the method of embodiment 20, further comprising performing uracil and / or dihydrouracil resistant amplification of DNA.
[0035] Embodiment 22 is the method of embodiment 21, wherein the uracil- and / or dihydrouracil-resistant amplification of DNA comprises PCR using a uracil- and / or dihydrouracil-resistant DNA polymerase.
[0036] Embodiment 23 is the method of any one of embodiments 11 to 17 or 20 to 22, wherein the auxiliary adapter does not comprise a barcode.
[0037] Embodiment 24 is the method of any one of embodiments 11 to 17 or 19 to 23, wherein cleavage of the DNA with a restriction enzyme results in overhangs.
[0038] Embodiment 25 is the method of embodiment 24, wherein the overhang is a single-base overhang.
[0039] Embodiment 26 is the method of embodiment 25, wherein the single-base overhang is a single-base 5'-overhang.
[0040] Embodiment 27 is the method of embodiment 25 or embodiment 26, wherein the single-base overhang is T.
[0041] Embodiment 28 is the method of embodiment 25 or embodiment 26, wherein the single-base overhang is A.
[0042] Embodiment 29 is the method of any one of the preceding embodiments, wherein the methylation-preserving amplification comprises contacting the DNA with a methyltransferase.
[0043] Embodiment 30 is the method of embodiment 29, wherein the methyltransferase preferentially methylates hemimethylated CpG and / or hemimethylated CpHpG.
[0044] Embodiment 31 is the method of embodiment 29 or embodiment 30, wherein the methyltransferase is DNMT1.
[0045] Embodiment 32 is the method of any one of the preceding embodiments, wherein the methylation-preserving amplification comprises one or more of polymerase chain reaction, linear amplification, rolling circle amplification, ligase chain reaction, strand displacement amplification, nucleic acid sequence-based amplification, and sequence-based self-sustained replication.
[0046] Embodiment 33 is the method of any one of the preceding embodiments, wherein the methylation-preserving amplification comprises thermocycling amplification.
[0047] Embodiment 34 is the method of any one of the preceding embodiments, wherein the methylation-preserving amplification comprises isothermal amplification.
[0048] Embodiment 35 is the method of embodiment 34, wherein the isothermal amplification comprises recombinant polymerase amplification (RPA), helix-dependent amplification (HDA), loop-mediated isothermal amplification (LAMP), or rolling circle amplification (RCA).
[0049] Embodiment 36 is the method of any one of the preceding embodiments, wherein the methylation-preserving amplification comprises linear amplification with thermal cycling.
[0050] Embodiment 37 is the method of any one of embodiments 2, 5, or 7-36, wherein the first nucleobase is an unmodified cytosine and the second nucleobase is a modified cytosine, and optionally the modified cytosine is 5-methylcytosine or 5-hydroxymethylcytosine.
[0051] Embodiment 38 is a method for determining whether or not a first nucleobase of DNA is affected differently from a second nucleobase of DNA, the method comprising: Before or after enrichment; and 38. The method of any one of embodiments 2, 5, or 7-37, performed before sequencing.
[0052] Embodiment 39 is the method of any one of embodiments 2, 5, or 7-38, wherein the procedure affecting a first nucleobase of the DNA differently from a second nucleobase of the DNA chemically converts the first or second nucleobase such that the base-pairing specificity of the converted nucleobase is altered.
[0053] Embodiment 40 is the method of any one of embodiments 2, 5, or 7 to 39, wherein the procedure that affects the first nucleobase of the DNA differently from the second nucleobase of the DNA is a methylation-sensitive conversion.
[0054] Embodiment 41 is the method of embodiment 40, wherein the methylation-sensitive conversion is bisulfite conversion, oxidative bisulfite (Ox-BS) conversion, Tet-assisted bisulfite (TAB) conversion, APOBEC-linked epigenetic (ACE) conversion, or enzymatic conversion.
[0055] Embodiment 42 is the method of embodiment 41, wherein the Tet-assisted conversion further comprises a substituted borane reducing agent, optionally wherein the substituted borane reducing agent is 2-picoline borane, borane pyridine, tert-butylamine borane, or ammonia borane.
[0056] Embodiment 43 is the method of any one of the preceding embodiments, wherein the method further comprises distributing at least a portion of the DNA into a plurality of sub-samples comprising a first sub-sample and a second sub-sample, wherein the first sub-sample comprises a higher proportion of DNA with cytosine modifications than the second sub-sample.
[0057] Embodiment 44 is the method of embodiment 43, wherein the step of distributing at least a portion of the DNA into a plurality of sub-samples comprises contacting the DNA with an agent that recognizes modified cytosines in the DNA, and wherein a first sub-sample comprises DNA with a higher proportion of modified cytosines than a second sub-sample.
[0058] Embodiment 45 is a method for distributing a sample of a liquid before the sequencing step, and i. Before performing DNA methylation-preserving amplification, ii. After performing methylation-preserving amplification of DNA, iii. before enriching the DNA for one or more sets of epigenetic target regions of DNA; and / or iv. The method of embodiment 43 or embodiment 44, which is performed after enriching the DNA for one or more sets of epigenetic target regions of DNA.
[0059] Embodiment 46 is the method of embodiment 44 or embodiment 45, wherein the agent that recognizes modified nucleobases in DNA is a methyl-binding reagent.
[0060] Embodiment 47 is the method of embodiment 46, wherein the methyl-binding reagent is a methyl-binding domain (MBD) protein or antibody.
[0061] Embodiment 48 is the method of embodiment 46 or embodiment 47, wherein the methyl-binding reagent is specific for one or more methylated nucleotide bases, and optionally, the one or more methylated nucleotide bases are 5-methylcytosine.
[0062] Embodiment 49 is the method of any one of embodiments 46 to 48, wherein the methyl-binding reagent is immobilized on a solid support.
[0063] Embodiment 50 is the method of any one of embodiments 46 to 49, wherein the partitioning step comprises immunoprecipitation of methylated DNA.
[0064] Embodiment 51 is the method of any one of embodiments 46 to 50, wherein the partitioning step comprises partitioning based on binding to a protein, optionally wherein the protein is a methylated protein, an acetylated protein, an unmethylated protein, or an unacetylated protein; and / or optionally wherein the protein is a histone.
[0065] Embodiment 52 is the method of any one of embodiments 46 to 51, wherein the distributing step comprises contacting the DNA with a binding reagent specific for the protein and immobilized on a solid support.
[0066] Embodiment 53 is the method of any one of embodiments 46 to 52, wherein a first dispensed sub-sample of the plurality of dispensed sub-samples is differentially tagged from a second dispensed sub-sample of the plurality of dispensed sub-samples.
[0067] Embodiment 54 is the method of any one of the preceding embodiments, comprising contacting the DNA or at least one sub-sample thereof with at least one nuclease prior to enrichment or sequencing, optionally wherein the at least one nuclease is at least one restriction enzyme.
[0068] Embodiment 55 is the method of embodiment 54, wherein the step of contacting the DNA or at least one sub-sample thereof with at least one nuclease is performed after dividing the sample into a plurality of sub-samples or before performing a procedure that affects a first nucleobase of the DNA differently from a second nucleobase of the DNA.
[0069] Embodiment 56 is the method of embodiment 54 or embodiment 55, wherein at least one restriction enzyme is a methylation-sensitive restriction enzyme (MSRE).
[0070] Embodiment 57 is the method of embodiment 54 or embodiment 55, wherein at least one restriction enzyme is a methylation-dependent restriction enzyme (MDRE).
[0071] Embodiment 58 is the method of any one of embodiments 54 to 56, wherein the DNA or at least one subsample thereof is contacted with at least one methylation-sensitive restriction enzyme, thereby generating hypermethylated DNA.
[0072] Embodiment 59 is the method of any one of embodiments 54, 55, or 57, wherein the DNA or at least one subsample thereof is contacted with at least one methylation-dependent restriction enzyme, thereby generating hypomethylated DNA.
[0073] Embodiment 60 is the method of any one of embodiments 54-56, 58, or 59, wherein the first sub-sample is contacted with an MSRE.
[0074] Embodiment 61 is the method of any one of embodiments 54, 55, or 57-60, wherein the second sub-sample is contacted with an MDRE.
[0075] Embodiment 62 is the method of any one of embodiments 54 to 61, comprising contacting at least one sub-sample with at least two restriction enzymes prior to enrichment or sequencing, optionally wherein the contacting occurs before performing a procedure that affects a first nucleobase of the DNA differently from a second nucleobase of the DNA.
[0076] Embodiment 63 is the method of embodiment 62, wherein the at least two restriction enzymes comprise or consist of two or three restriction enzymes.
[0077] Embodiment 64 is the method of any one of embodiments 54 to 63, wherein the at least one restriction enzyme is selected from the group consisting of FspEI, LpnPI, MspJI, SgeI, AatII, AccII, AciI, Aor13HI, Aor15HI, BspT104I, BssHII, BstUI, Cfr10I, ClaI, CpoI, Eco52I, HaeII, HapII, HhaI, Hin6I, HpaII, HpyCH4IV, MluI, MspI, NaeI, NotI, NruI, NsbI, PmaCI, Psp1406I, PvuI, SacII, SalI, SmaI, and SnaBI.
[0078] Embodiment 65 is the method of any one of embodiments 54 to 64, further comprising the step of attaching one or more adapters to at least one end of at least some of the DNA molecules in the plurality of distributed sets prior to digestion.
[0079] Embodiment 66 is the method of embodiment 65, wherein one or more adapters comprise at least one tag.
[0080] Embodiment 67 is the method of embodiment 66, wherein at least one tag comprises a molecular barcode.
[0081] Embodiment 68 is the method according to any one of embodiments 65 to 67, wherein one or more adapters are resistant to digestion by a methylation-sensitive or methylation-dependent restriction enzyme.
[0082] Embodiment 69 is a method for preparing a methylation-sensitive restriction enzyme-resistant nucleotide sequence comprising: a) one or more methylated nucleotides, optionally wherein the methylated nucleotides comprise 5-methylcytosine and / or 5-hydroxymethylcytosine; b) one or more nucleotide analogs that are resistant to methylation-sensitive restriction enzymes; or 69. The method of embodiment 68, wherein c) the nucleotide sequence is not recognized by a methylation-sensitive restriction enzyme.
[0083] Embodiment 70 is the method according to any one of embodiments 43 to 69, wherein the sub-samples are pooled before sequencing.
[0084] Embodiment 71 is the method of any one of embodiments 1, 2, 4, 5, 7-10, or 17-70, further comprising enriching the DNA for one or more sets of epigenetic target regions of the DNA, thereby providing enriched DNA.
[0085] Embodiment 72 is the method of any one of the preceding embodiments, further comprising enriching the DNA for one or more sets of sequence variable target regions of the DNA.
[0086] Embodiment 73 is the method of any one of the preceding embodiments, wherein the enriching step comprises contacting the DNA with target-specific probes specific for one or more sets of epigenetic target regions and / or one or more sets of sequence variable target regions.
[0087] Embodiment 74 is the method of any one of embodiments 3 or 6 to 73, wherein the set of epigenetic target regions comprises a set of hypermethylated variable target regions and / or a set of hypomethylated variable target regions.
[0088] Embodiment 75 is the method of any one of embodiments 3 or 6 to 74, wherein the set of epigenetic target regions comprises a set of fragmented variable target regions.
[0089] Embodiment 76 is the method of embodiment 75, wherein the set of fragmented variable target regions comprises a transcription start site region.
[0090] Embodiment 77 is the method of embodiment 75 or embodiment 76, wherein the set of fragmented variable target regions comprises a CTCF binding region.
[0091] Embodiment 78 is the method of any one of embodiments 3 or 6 to 77, wherein the set of epigenetic target regions comprises one or more type-specific epigenetic target regions.
[0092] Embodiment 79 is the method of embodiment 78, wherein the one or more type-specific epigenetic target regions comprise type-specific differentially methylated regions and / or type-specific fragments.
[0093] Embodiment 80 is the method of embodiment 78, wherein the one or more type-specific epigenetic target regions comprise type-specific hypomethylated regions and / or type-specific hypermethylated regions.
[0094] Embodiment 81 is the method of any one of embodiments 78 to 80, wherein the one or more type-specific epigenetic target regions comprise cell type-specific, cell cluster type-specific, tissue type-specific, and / or cancer type-specific epigenetic target regions.
[0095] Embodiment 82 is a method for determining whether the one or more type-specific epigenetic target regions are: a) hypermethylated in immune cells compared to non-immune cell types present in the blood sample; b) differentially methylated in the colon compared with other tissue types; c) differentially methylated in breast compared with other tissue types; d) differentially methylated in the liver compared with other tissue types; e) differentially methylated in kidney compared with other tissue types; f) differentially methylated in the pancreas compared with other tissue types; g) differentially methylated in the prostate compared with other tissue types; h) is differentially methylated in skin compared to other tissue types; or i) Differentially methylated in the bladder compared to other tissue types 82. The method of any one of embodiments 78-81, comprising a target region.
[0096] Embodiment 83 is the method of any one of embodiments 78 to 82, wherein the hypermethylated target region is methylated to an extent that is at least 10%, 20%, 30%, or at least 40% greater than the average methylation of the target region in the sample or compared to other cell or tissue types.
[0097] Embodiment 84 is a method for determining whether the one or more type-specific epigenetic target regions are: a) a target region that is hypomethylated in non-immune cell types present in the sample compared to the methylation level of the target region in different cell or tissue types in the sample; b) a fragment that is specific for immune cells compared to non-immune cell types present in the sample; or c) Fragments specific to colon, lung, breast, liver, kidney, pancreas, prostate, skin, or bladder compared to other tissue types 84. The method of any one of embodiments 78 to 83, comprising:
[0098] Embodiment 85 is the method of any one of embodiments 78 to 84, wherein the level of one or more type-specific epigenetic target regions of cell type or tissue type origin is determined.
[0099] Embodiment 86 is the method of embodiments 78-85, in which the level of one or more immune cells, non-immune cell types present in the blood sample, and / or one or more type-specific epigenetic target regions originating from colon, lung, breast, liver, kidney, prostate, skin, bladder, or pancreatic cells is determined.
[0100] Embodiment 87 is the method of any one of embodiments 78 to 86, further comprising identifying at least one cell type, cell cluster type, tissue type, and / or cancer type from which the one or more type-specific epigenetic target regions originated.
[0101] Embodiment 88 is the method of any one of embodiments 78 to 87, comprising determining the methylation level of the type-specific epigenetic target region.
[0102] Embodiment 89 is the method of any one of embodiments 43 to 88, wherein at least a portion of the DNA from the first sub-sample and at least a portion of the DNA from the second sub-sample are pooled, thereby providing a combined sub-sample.
[0103] Embodiment 90 is the method according to any one of embodiments 43 to 89, wherein the DNA of the first sub-sample and the DNA of the second sub-sample are differentially tagged.
[0104] Embodiment 91 is the method according to the immediately preceding embodiment, wherein the pool comprises less than or equal to about 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, or 5% of the DNA of the second sub-sample.
[0105] Embodiment 92 is the method of embodiment 90 or embodiment 91, wherein the pool comprises about 70-90%, about 75-85%, or about 80% of the DNA of the second sub-sample.
[0106] Embodiment 93 is the method according to any one of embodiments 90 to 92, wherein the pool comprises substantially all of the DNA of the first sub-sample.
[0107] Embodiment 94 is the method according to any one of embodiments 90 to 92, wherein the pool comprises substantially all of the DNA of the first sub-sample or the processed first sub-sample.
[0108] Embodiment 95 is the method of any one of embodiments 90 to 94, wherein the first set of target regions is captured from at least a portion of the first sub-sample after formation of the pool.
[0109] Embodiment 96 is the method according to any one of embodiments 90 to 95, wherein at least a portion of the DNA from the first sub-sample and at least a portion of the DNA from the second sub-sample are sequenced in the same sequencing cell.
[0110] Embodiment 97 is the method of any one of embodiments 90 to 96, wherein the plurality of sub-samples includes a third sub-sample that contains a higher proportion of DNA with epigenetic modifications than the second sub-sample but a lower proportion than the first sub-sample.
[0111] Embodiment 98 is the method of the immediately preceding embodiment, further comprising differentially tagging a third sub-sample.
[0112] Embodiment 99 is the method of embodiment 97 or embodiment 98, wherein the DNA from the first sub-sample, the DNA from the third sample, and the set of target regions are pooled, and, if necessary, the DNA from the first, second, and third sub-samples is sequenced in the same sequencing cell.
[0113] Embodiment 100 is the method of any one of embodiments 89 to 99, wherein sequencing the DNA of the combined sub-samples comprises sequencing the DNA in a modification-sensitive manner.
[0114] Embodiment 101 is the method of any one of embodiments 3, 6-10, or 17-96, wherein the DNA is amplified after the enrichment step.
[0115] Embodiment 102 is the method of embodiment 101, wherein the amplifying step further comprises differentially tagging the enriched DNA.
[0116] Embodiment 103 is the method of embodiment 102, wherein differentially tagging the enriched DNA comprises attaching one or more sample indexes to the DNA.
[0117] Embodiment 104 is the method of any one of the preceding embodiments, wherein sequencing the DNA comprises sequencing the DNA in a manner that distinguishes the first nucleobase from the second nucleobase.
[0118] Embodiment 105 is the method of any one of embodiments 1, 3, 4, or 6 to 104, wherein the step of sequencing in a manner sensitive to the modification comprises long-read sequencing.
[0119] Embodiment 106 is the method of any one of embodiments 1, 3, 4, or 6 to 105, wherein the step of sequencing in a manner sensitive to the modification comprises nanopore sequencing.
[0120] Embodiment 107 is the method of any one of embodiments 1, 3, 4, or 6 to 104, wherein the step of sequencing in a modification-sensitive manner comprises five-letter or six-letter sequencing.
[0121] Embodiment 108 is a method according to any one of the preceding embodiments, wherein the sequencing step comprises next generation sequencing.
[0122] Embodiment 109 is a method according to any one of the preceding embodiments, wherein the sequencing step comprises generating a plurality of sequencing reads, and the method further comprises mapping the plurality of sequence reads to one or more reference sequences to generate mapped sequence reads, and processing the mapped sequence reads corresponding to the set of sequence variable target regions and the set of epigenetic target regions.
[0123] Embodiment 110 is the method according to any one of embodiments 4 to 109, further comprising determining an epigenetic consensus sequence of DNA associated with at least a portion of the barcode.
[0124] Embodiment 111 is an embodiment of the present invention, wherein determining an epigenetic consensus sequence of DNA associated with at least a portion of the barcodes comprises comparing the sequence of at least a portion of the associated reads to (a) a unique barcode or a unique set of barcodes, (b) a unique genomic start and / or stop position, or (c) both (a) and (b), wherein the reads originate from a DNA molecule; and Determining the consensus epigenetic state of each nucleotide of the DNA molecule The method of any one of embodiments 1 to 3, 7 to 10, or 17 to 110, comprising:
[0125] Embodiment 112 is the method of any one of embodiments 1 to 3, 7 to 10, or 17 to 111, further comprising determining a consensus base sequence of DNA associated with at least a portion of the barcode.
[0126] Embodiment 113 is an embodiment of the present invention, wherein determining a consensus base sequence of DNA associated with at least a portion of the barcodes comprises comparing the sequence of at least a portion of the associated reads to (a) a unique barcode or a unique set of barcodes, (b) a unique genomic start and / or stop position, or (c) both (a) and (b), wherein the reads originate from a DNA molecule; and determining the identity of the consensus base for each nucleotide in the DNA molecule 113. The method of embodiment 112, comprising:
[0127] Embodiment 114 is the method of any one of the preceding embodiments, wherein the DNA is cell-free DNA.
[0128] Embodiment 115 is the method of embodiment 114, wherein the cell-free DNA is in an amount of 1 ng to 500 ng.
[0129] Embodiment 116 is the method of any one of the preceding embodiments, wherein the DNA is from a blood sample and / or a tissue sample.
[0130] Embodiment 117 is the method of embodiment 116, wherein the blood sample is a whole blood sample, a plasma sample, a buffy coat sample, a leukoreduced sample, or a PBMC sample.
[0131] Embodiment 118 is the method of any one of the preceding embodiments, wherein the DNA and / or sample is from a subject.
[0132] Embodiment 119 is the method of embodiment 118, wherein the subject is an animal.
[0133] Embodiment 120 is the method of embodiment 118 or embodiment 119, wherein the subject is a human.
[0134] Embodiment 121 is the method of any one of embodiments 94 to 120, wherein the blood sample is fractionated before enrichment for at least one set of epigenetic target regions of DNA.
[0135] Embodiment 122 is the method of any one of embodiments 118-121, wherein the subject has or is at risk of having cancer.
[0136] Embodiment 123 is the method of embodiments 118 to 122, further comprising determining the presence or status of cancer in the subject.
[0137] Embodiment 124 is the method of any one of embodiments 118 to 123, further comprising determining the likelihood that the subject has an infection.
[0138] Embodiment 125 is the method of any one of embodiments 118 to 124, further comprising determining the likelihood that the subject will have transplant rejection. [Brief explanation of the drawings]
[0139] [Figure 1-1] FIG. 1A illustrates an exemplary workflow according to certain embodiments disclosed herein. [Figure 1-2] Same as above.
[0140] [Figure 1-3] FIG. 1B illustrates an exemplary workflow according to certain embodiments disclosed herein. [Figure 1-4] Same as above.
[0141] [Figure 2] FIG. 2 is a schematic diagram of an example system suitable for use with some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0142] Detailed Description of Certain Embodiments Reference will now be made in detail to certain specific embodiments of the invention. While the invention will be described in conjunction with such embodiments, it will be understood that they are not intended to limit the invention to those embodiments. On the contrary, the invention is intended to cover all alternatives, modifications, and equivalents, which may be included within the scope of the present invention as defined by the appended claims.
[0143] Before describing the teachings of the present invention in detail, it is understood that the disclosure is not limited to specific compositions or process steps, as these may vary. It should be noted that as used in this specification and the appended claims, the singular forms "a," "an," and "the" include the plural forms unless the context clearly dictates otherwise. Thus, for example, a reference to "a nucleic acid" includes a plurality of nucleic acids, a reference to "a cell" includes a plurality of cells, etc.
[0144] Numerical ranges are inclusive of the numbers defining the range. Measured and measurable values are understood to be approximations, taking into account significant digits and errors associated with measurement. Similarly, the use of "comprise," "comprises," "comprising," "contain," "contains," "containing," "include," "includes," and "including" is not intended to be limiting. It is understood that both the foregoing general and detailed descriptions are exemplary and explanatory only and do not limit the present teachings.
[0145] Unless specifically noted in the specification above, embodiments herein that recite various components as "comprising" are also contemplated as "consisting of" or "consisting essentially of" the recited components, and embodiments herein that recite various components as "consisting essentially of" are also contemplated as "comprising" or "consisting essentially of" the recited components (this interchangeability does not apply to the use of these terms in the claims).
[0146] The section headings used herein are for organizational purposes only and are not to be construed as limiting the disclosed subject matter in any way. In the event that any document or other material incorporated by reference contradicts the express contents of this specification, including definitions, the present specification will control. I. Definition
[0147] As used herein, "amplify," "amplifying," or "amplification" refers to a process by which extra or multiple copies of a particular polynucleotide are formed. Amplification methods may include any suitable method known in the art. As used herein, nucleic acid molecules amplified using "methylation-preserving amplification" substantially maintain their methylation state after amplification.
[0148] "Buffy coat" refers to the portion of a blood (e.g., whole blood) or bone marrow sample that contains all or most of the sample's white blood cells and platelets. A buffy coat fraction of a sample can be prepared from the sample using centrifugation, which separates sample components by density. For example, after centrifugation of a whole blood sample, the buffy coat fraction is located between the plasma and erythrocyte (red blood cell) layers. Buffy coats can contain both mononuclear (e.g., T cells, B cells, NK cells, dendritic cells, and monocytes) and polymorphonuclear (e.g., granulocytes such as neutrophils and eosinophils) white blood cells.
[0149] "Cell-free DNA," "cfDNA molecules," or simply "cfDNA" includes DNA molecules naturally present in a subject in an extracellular form (e.g., in blood, serum, plasma, or other bodily fluids, such as lymph, cerebrospinal fluid, urine, or sputum). cfDNA was previously present in one or more cells of a large, complex biological organism, e.g., a mammal, but has been released from the cell(s) into fluids found in the organism, and can be obtained from a sample of the fluid without the need to perform an in vitro cell lysis step. cfDNA molecules may exist as DNA fragments.
[0150] As used herein, "methyltransferase" or "DNA methyltransferase" refers to any enzyme that methylates DNA, e.g., the 5-carbon of cytosine, to form 5'-methylcytosine, e.g., DNA methyltransferase of Enzyme Code (EC) 2.1.1.37. DNA methyltransferases that can be used in the methods described herein include methyltransferases (including modified, evolved, or engineered methyltransferases) that can act on hemimethylated CpG or hemimethylated CpHpG (where H = A, C, or T). Such DNA methyltransferases can include, but are not limited to, DNMT1 and modified DNMT3, which act on hemimethylated CpG or hemimethylated CpHpG. These enzymes can use S-adenosylmethionine as the methyl donor. Additional examples of methyltransferases are described elsewhere herein.
[0151] "DNA methyltransferase 1" or "DNMT1" refers to the enzyme encoded by the DNMT1 gene (UniProt Accession No. K7EP77). This enzyme transfers methyl groups to cytosine nucleotides in genomic DNA, preferentially methylating unmethylated cytosines in hemimethylated DNA, particularly hemimethylated CpG dinucleotides. DNMT1 is the enzyme primarily responsible for maintaining methylation patterns after DNA replication.
[0152] As used herein, "fragment" refers to a biological component, such as a nucleic acid molecule (e.g., DNA or RNA), that is broken or separated from one or more other pieces.Fragmentation, such as DNA fragmentation, can occur spontaneously (as in the cfDNA fragments that can be obtained from blood samples), or can be intentionally induced, for example, using standard laboratory procedures as described herein.DNA fragmentation can be carried out, for example, to prepare DNA (e.g., genomic DNA and / or DNA isolated from samples containing cells) for sequencing.In some samples, such as cfDNA samples, artificial fragmentation may be unnecessary.
[0153] As used herein, a modification or other feature is "present in a higher proportion" in a first sample or population of nucleic acids than in a second sample or population if the proportion of nucleotides having the modification or other feature is higher in the first sample or population than in the second population. For example, if one in ten nucleotides in a first sample are mC and one in twenty nucleotides in a second sample are mC, then the first sample contains a higher proportion of 5-methylated cytosine modifications than the second sample.
[0154] As used herein, "leukapheresis" refers to a procedure in which white blood cells (leukocytes) are isolated from a sample of blood collected from a subject. Leukapheresis can be performed, for example, to obtain cells, such as those described herein, for research, diagnostic, prognostic, or monitoring purposes. Thus, as used herein, a "leukapheresis sample" refers to a sample containing white blood cells collected from a subject using leukapheresis.
[0155] As used herein, " peripheral blood mononuclear cells " or " PBMC " refers to the immune cells that originate from bone marrow and have a single round nucleus, and are found in peripheral circulation. Such cells include, for example, lymphocytes (T cells, B cells, and NK cells) and monocytes, and are isolated from blood samples (for example, from the whole blood sample collected from a subject) using density gradient centrifugation.
[0156] As used herein, "without substantially changing the base-pairing specificity" of a given nucleobase means that the majority of molecules that comprise the nucleobase that can be sequenced have no change in the base-pairing specificity of the given nucleobase compared to its base-pairing specificity when present in the original isolated sample.In some embodiments, 75%, 90%, 95% or 99% of molecules that comprise the nucleobase that can be sequenced have no change in the base-pairing specificity when present in the original isolated sample.As used herein, "change in base-pairing specificity" of a given nucleobase means that the majority of molecules that comprise the nucleobase that can be sequenced have the base-pairing specificity of the nucleobase compared to its base-pairing specificity in the original isolated sample.
[0157] As used herein, "base-pairing specificity" refers to the standard DNA base (A, C, G, or T) with which a given base pairs most preferentially. For example, unmodified cytosine and 5-methylcytosine have the same base-pairing specificity (i.e., specificity for G), but uracil and cytosine have different base-pairing specificities, with uracil having base-pairing specificity for A and cytosine having base-pairing specificity for G. The ability of uracil to form a wobble base pair with G is irrelevant, since uracil nevertheless pairs most preferentially with A among the four standard DNA bases.
[0158] The "capture yield" of a collection of probes for a given target set refers to the amount of nucleic acid corresponding to the target set that the collection of probes captures under typical conditions (e.g., the amount or absolute amount relative to another target set). Exemplary typical capture conditions are incubating sample nucleic acid and probes at 65°C for 10-18 hours in a small reaction volume (approximately 20 μL) containing a stringent hybridization buffer. Capture yields can be expressed in absolute terms, or, in the case of a collection of multiple probes, in relative terms. When comparing capture yields for multiple sets of target regions, they are normalized with respect to the footprint size (e.g., on a per kilobase basis) of the target region set. Thus, for example, if the footprint sizes of the first and second target regions are 50 kb and 500 kb, respectively (using a normalization factor of 0.1), and the mass / volume concentration of the captured DNA corresponding to the first set of target regions is greater than 0.1 times the mass / volume concentration of the captured DNA corresponding to the second set of target regions, the DNA corresponding to the first set of target regions will be captured with a higher yield than the DNA corresponding to the second set of target regions. As a further example, using the same footprint size, if the captured DNA corresponding to the first set of target regions has a mass / volume concentration that is 0.2 times the mass / volume concentration of the captured DNA corresponding to the second set of target regions, the DNA corresponding to the first set of target regions will be captured with a capture yield that is 2 times higher than the DNA corresponding to the second set of target regions.
[0159] "Enriching" or "capturing" one or more target nucleic acids or one or more nucleic acids comprising at least one target region refers to preferentially isolating or separating one or more target nucleic acids or one or more nucleic acids comprising at least one target region from non-target nucleic acids or from nucleic acids that do not contain at least one target region.
[0160] An "enriched set" or "captured set" of nucleic acids, or an "enriched" or "captured" nucleic acid, refers to nucleic acids that have undergone capture.
[0161] As used herein, a "capture moiety" is a molecule that allows for affinity separation of a molecule, such as a nucleic acid, linked to the capture moiety from molecules that lack the capture moiety. Exemplary capture moieties include biotin, which allows for affinity separation by binding to streptavidin that is or can be linked to a solid phase, or oligonucleotides that allow for affinity separation through binding to complementary oligonucleotides that are or can be linked to a solid phase.
[0162] As used herein, a "cell type" is a set of cells that share common characteristics. For example, a cell type may include cells of different origins, differentiation types, activation types, or any combination of different origins, differentiation types, and activation types. In fact, the differentiation state and activation state may overlap and often change together in a given cell, such as an immune cell or cancer cell. For example, activation of an immune cell may induce differentiation of the cell. In some embodiments, cell types may be distinguished based on characteristics such as one or more cell surface markers, gene signatures (e.g., the expression (or expression level) of a particular gene or set of genes), and / or epigenetic signatures such as regions of DNA hypermethylation or hypomethylation.
[0163] As used herein, a "cell cluster" or "cluster" refers to a plurality of related cell types, e.g., immune cell types, tissue-specific cell types, and / or cancer cell types. In some embodiments, the cell types within a cluster have similar DNA methylation profiles, e.g., in multiple hypermethylated and / or hypomethylated variable target regions.
[0164] "Converted nucleobase" is a nucleobase that has a change in base pairing specificity, and the original base pairing specificity of the nucleobase has been changed by a procedure.For example, a certain procedure converts unmethylated or unmodified cytosine into dihydrouracil, or more generally, at least one modified or unmodified form of cytosine undergoes deamination, resulting in uracil (considered a modified nucleobase in the context of DNA) or a further modified form of uracil.As used herein, "converted sample" refers to a sample that contains DNA that contains at least one converted nucleobase.
[0165] As used herein, a "combination" of steps or other elements refers to the performance or presence of two or more steps or elements in a method or product; where appropriate, the elements may be together in a single composition, device, or the like, or may be in close proximity, e.g., in separate containers or compartments within a larger container, such as a multiwell plate, tube rack, refrigerator, freezer, incubator, water bath, ice bucket, machine, or other storage configuration. A combination, combinations, or combinations thereof refers to any and all permutations and combinations of the terms listed before the term "combination." For example, "A, B, C, or a combination thereof" is intended to include at least one of A, B, C, AB, AC, BC, or ABC, and also includes BA, CA, CB, ACB, CBA, BCA, BAC, or CAB if order is important in the particular context. Continuing with this example, combinations containing repeats of one or more items or terms are expressly included, e.g., BB, AAA, AAB, BBC, AAABCCCC, CBBAAA, CABABB, etc. Those skilled in the art will understand that there is typically no limit to the number of items or terms in any combination, unless otherwise clear from the context.
[0166] "Specifically binds" in the context of a primer, probe, or other oligonucleotide and a target sequence (e.g., a nucleic acid comprising a sequence partially or completely complementary to the primer, probe, or other oligonucleotide) means that, under appropriate hybridization conditions, the primer, probe, or other oligonucleotide hybridizes to its target sequence, or a copy thereof, to form a stable hybrid, while minimizing the formation of stable non-target hybrids. In this manner, the primer, probe, or other oligonucleotide hybridizes to the target sequence, or a copy thereof, to a sufficiently greater extent than non-target sequences to ultimately permit enrichment or detection of the target sequence. Appropriate hybridization conditions are well known in the art and can be predicted based on sequence composition or determined by using routine testing methods (e.g., Sambrook et al., Molecular Cloning, A Laboratory Manual, 2004, incorporated herein by reference). nd ed. (Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1989), see §§ 1.90-1.91, 7.37-7.57, 9.47-9.51, and 11.47-11.57, especially §§ 9.50-9.51, 11.12-11.13, 11.45-11.47, and 11.55-11.57).
[0167] A "target region" refers to a genomic locus that is targeted for identification and / or capture, e.g., by using a probe (e.g., through sequence complementarity). A "target region set" or "set of target regions" refers to multiple genomic loci that are targeted for identification and / or capture, e.g., by using a set of probes (e.g., through sequence complementarity). A "target region set" may include regions that share at least one common feature. In some embodiments, a target region set is identified by at least one common feature. For example, a hypermethylated variable target region set includes regions of DNA that are hypermethylated.
[0168] A "sequence variable target region" refers to a target region that may exhibit sequence changes, such as nucleotide substitutions (i.e., single-base mutations), insertions, deletions, or gene fusions or rearrangements, in neoplastic cells (e.g., tumor cells and cancer cells) compared to normal cells. A "set of sequence variable target regions" refers to a set of sequence variable target regions. In some embodiments, the sequence variable target regions are target regions that may exhibit changes affecting less than or equal to 50 consecutive nucleotides, e.g., less than or equal to 40, 30, 20, 10, 5, 4, 3, 2, or 1 nucleotide.
[0169] "Epigenetic target region" refers to a target region that may exhibit sequence-independent differences in different cell or tissue types (e.g., different types of immune cells) or abnormal cells, such as neoplastic cells (e.g., tumor cells and cancer cells), compared to normal cells, or in DNA, e.g., DNA from different cell types or from subjects with cancer, compared to DNA from healthy subjects, that may exhibit sequence-independent differences (i.e., differences in methylation, nucleosome distribution, or other epigenetic features, without changes to the nucleotide sequence). Examples of sequence-independent changes include, but are not limited to, changes in methylation (increase or decrease), nucleosome distribution, fragmentation patterns, CCCTC-binding factor ("CTCF") binding, transcription start sites (e.g., with respect to any one or more of the binding of RNA polymerase components, regulatory protein binding, fragmentation characteristics, and nucleosome distribution), and regulatory protein binding regions. Thus, epigenetic target region sets include, but are not limited to, hypermethylated variable target region sets, hypomethylated variable target region sets, and fragmented variable target region sets, such as CTCF binding sites and transcription start sites.For the purpose of the present invention, the loci that are prone to neoplasia, tumor, or cancer-related local amplification and / or gene fusion can also be included in epigenetic target region sets, because the detection of copy number changes by sequencing or fusion sequences that map to more than one locus in a reference genome tends to be more similar to the detection of the exemplary epigenetic changes discussed above than the detection of nucleotide substitutions, insertions, or deletions, for example, in that the detection of local amplification and / or gene fusion can be detected at a relatively low sequencing depth because it does not depend on the accuracy of base calls at one or a few individual positions." Epigenetic target region set " is a set of epigenetic target regions.
[0170] As used herein, a "differentially methylated region" refers to a region of DNA that has a detectably different degree of methylation in at least one cell or tissue type compared to the degree of methylation in the same region of DNA from at least one other cell or tissue type, or a region of DNA that has a detectably different degree of methylation in at least one cell or tissue type obtained from a subject with a disease or disorder compared to the degree of methylation in the same region of DNA in the same cell or tissue type obtained from a healthy subject. In some embodiments, a differentially methylated region has a detectably higher degree of methylation (e.g., a hypermethylated region) in at least one cell or tissue type, e.g., at least one immune cell type, compared to the degree of methylation in the same region of DNA from at least one other cell or tissue type, e.g., another immune cell type, or from the same cell or tissue type from a healthy subject. In some embodiments, a differentially methylated region has a detectably lower degree of methylation (e.g., a hypomethylated region) in at least one cell or tissue type, e.g., at least one immune cell type, compared to the degree of methylation in the same region of DNA from at least one other cell or tissue type, e.g., another immune cell type, or from the same cell or tissue type from a healthy subject.
[0171] A nucleic acid is "produced by a tumor" if it originates from a tumor cell. A tumor cell is a neoplastic cell that originates from a tumor, regardless of whether it remains within the tumor or leaves the tumor (e.g., in the case of metastatic cancer cells and circulating tumor cells). As used herein, a "precancer" or "precancerous condition" refers to an abnormality that has the potential to become cancerous, and the likelihood of becoming cancerous is higher than if the potential abnormality were not present, i.e., normal. Examples of precancer include, but are not limited to, adenoma, hyperplasia, dysplasia, dysplasia, benign neoplasm (benign tumor), premalignant intramucosal carcinoma, and polyp. It should be noted that certain types of intramucosal carcinoma are recognized in the art as cancerous, e.g., stage 0 cancer, as opposed to premalignant.
[0172] The term "methylation" or "DNA methylation" refers to the addition of a methyl group to a nucleic acid base in a nucleic acid molecule. In some embodiments, methylation refers to the addition of a methyl group to cytosine at a CpG site (cytosine-phosphate-guanine site (i.e., cytosine followed by guanine in the 5' to 3' direction of a nucleic acid sequence). In some embodiments, DNA methylation refers to the addition of a methyl group to an adenine, e.g., N6-methyladenine. In some embodiments, DNA methylation is 5-methylation (modification of the 5 carbon of the 6-membered ring of cytosine). In some embodiments, 5-methylation refers to the addition of a methyl group to the 5C position of cytosine to produce 5-methylcytosine (5mC). In some embodiments, methylation includes derivatives of 5mC, including, but not limited to, 5-hydroxymethylcytosine (5-hmC), 5-formylcytosine (5-fC), and 5-carboxylcytosine (5-caC). In some embodiments, DNA methylation is 3C methylation (modification of the nitrogen at the 3 position of the 6-membered ring of cytosine). In some embodiments, 3C methylation involves the addition of a methyl group to the 3C position of cytosine to generate 3-methylcytosine (3mC). Methylation can also occur at non-CpG sites, for example, methylation can occur at CpA, CpT, or CpC sites. DNA methylation can alter the activity of methylated DNA regions. For example, if DNA in a promoter region is methylated, gene transcription can be suppressed. DNA methylation is crucial for normal development, and abnormal methylation can disrupt epigenetic regulation. Disruption of epigenetic regulation, for example, suppression, can cause diseases such as cancer. Promoter methylation in DNA can indicate cancer.
[0173] The term "hypermethylation" refers to an increased level or degree of methylation of a nucleic acid molecule compared to other nucleic acid molecules containing the same genetic information within a population (e.g., sample) of nucleic acid molecules. In some embodiments, hypermethylated DNA can include DNA molecules containing at least one methylated residue, at least two methylated residues, at least three methylated residues, at least five methylated residues, or at least ten methylated residues.
[0174] The term "hypomethylation" refers to a decrease in the level or degree of methylation of a nucleic acid molecule compared to other nucleic acid molecules containing the same genetic information within a population (e.g., a sample) of nucleic acid molecules. In some embodiments, hypomethylated DNA includes unmethylated DNA molecules. In some embodiments, hypomethylated DNA can include DNA molecules containing zero methylated residues, at most one methylated residue, at most two methylated residues, at most three methylated residues, at most four methylated residues, or at most five methylated residues.
[0175] The term "agent that recognizes modified nucleobases in DNA," e.g., "agent that recognizes modified cytosines in DNA," refers to a molecule or reagent that binds to or detects one or more modified nucleobases in DNA, e.g., methylcytosines. A "modified nucleobase" is a nucleobase that comprises a difference in chemical structure from an unmodified nucleobase. In the case of DNA, the unmodified nucleobase is adenine, cytosine, guanine, or thymine. In some embodiments, the modified nucleobase is a modified cytosine. In some embodiments, the modified nucleobase is a methylated nucleobase. In some embodiments, the modified cytosine is a methylcytosine, e.g., 5-methylcytosine. In such embodiments, the cytosine modification is methyl. Agents that recognize methylcytosines in DNA include, but are not limited to, "methyl-binding reagents," which herein refer to reagents that bind to methylcytosines. Methyl-binding reagents include, but are not limited to, methyl-binding domains (MBDs) and methyl-binding proteins (MBPs), as well as antibodies specific for methylcytosines. In some embodiments, such antibody binds to 5-methylcytosine in DNA.In some such embodiments, DNA can be single-stranded or double-stranded.Suitable agents include those that recognize double-stranded DNA, single-stranded DNA, and modified nucleotides in both double-stranded and single-stranded DNA.
[0176] The term "epigenetic state" refers to a certain level or degree of sequence-independent variation that may exist in a DNA sequence. In some embodiments, the epigenetic state of a DNA sequence refers to the degree or level of sequence methylation, nucleosome distribution, cfDNA fragmentation pattern, CCCTC-binding factor ("CTCF") binding, transcription start site, or regulatory protein binding region. Thus, epigenetic states include, but are not limited to, hypermethylation, hypomethylation, and the presence or absence of CTCF binding site or transcription start site. The epigenetic state of a sequence may also be a "reference epigenetic state" that can be used to compare the epigenetic state of corresponding sequences in other DNA molecules. An example of a reference epigenetic state is a state that exists in a sample obtained from a healthy subject and is not associated with cancer.
[0177] As used herein, "methylation state" refers to the presence or absence of a methyl group on a DNA nucleobase (e.g., cytosine) at a particular genomic position in a nucleic acid, the degree of methylation of a nucleic acid (e.g., high, low, intermediate, or unmethylated), or the number of methylated nucleotides in a particular nucleic acid molecule. A "methylated form" of a nucleic acid is meant to include a sequence containing a methylated DNA nucleobase, e.g., a methylated cytosine in a CpG dinucleotide.
[0178] As used herein, "methylation-sensitive nuclease" refers to a nuclease that preferentially cleaves unmethylated DNA compared to methylated DNA. For example, a methylation-sensitive nuclease can cleave at or near a recognition sequence, such as a restriction site, in a manner that depends on the absence of methylation of at least one nucleic acid base in the recognition sequence, such as cytosine. In some embodiments, the nucleolytic activity of a methylation-sensitive nuclease is at least 10, 20, 50, or 100 times higher at an unmethylated recognition site compared to a methylated control in a standard nucleolytic assay. Methylation-sensitive nucleases include methylation-sensitive restriction enzymes.
[0179] As used herein, "methylation-sensitive restriction enzyme" or "MSRE" refers to a methylation-sensitive nuclease that is a restriction enzyme. MSREs are sensitive to the methylation state of DNA (e.g., cytosine methylation), i.e., the presence or absence of a methyl group in a nucleotide base in its recognition sequence alters the rate at which the enzyme cleaves DNA. In some embodiments, a methylation-sensitive restriction enzyme does not cleave DNA if a specific nucleotide base is methylated in the recognition sequence. For example, HpaII is a methylation-sensitive restriction enzyme with the recognition sequence "CCGG," and does not cleave DNA if the second cytosine in the recognition sequence is methylated.
[0180] As used herein, "methylation-dependent nuclease" refers to a nuclease that preferentially cleaves methylated DNA compared to unmethylated DNA. For example, a methylation-dependent nuclease can cleave at or near a recognition sequence, such as a restriction site, in a manner that depends on the methylation of at least one nucleic acid base in the recognition sequence, such as cytosine. In some embodiments, the nucleolytic activity of a methylation-dependent nuclease is at least 10, 20, 50, or 100 times higher at a methylated recognition site compared to an unmethylated control in a standard nucleolytic assay. Methylation-dependent nucleases include methylation-dependent restriction enzymes.
[0181] As used herein, "methylation-dependent restriction enzyme" or "MDRE" refers to a methylation-dependent nuclease that is a restriction enzyme. MDREs depend on DNA methylation (e.g., cytosine methylation), i.e., the presence or absence of a methyl group in a nucleotide base alters the rate at which the enzyme cleaves DNA. In some embodiments, a methylation-dependent restriction enzyme does not cleave DNA if a specific nucleotide base is not methylated in the recognition sequence. For example, MspJI is a methylation-dependent restriction enzyme with the recognition sequence "mCNNR(N9)" and does not cleave DNA if a methylated cytosine (mC) is not present in the recognition sequence.
[0182] As used herein, "digestion efficiency" or "cutting efficiency" refers to the efficiency of restriction enzyme digestion. Digestion efficiency can be calculated based on the number of control molecules observed upon digestion with a restriction enzyme and the number of control molecules observed in the absence of restriction enzyme digestion. MSRE digestion efficiency is calculated as follows: Efficiency = 1 - (negative control molecules) [MSRE] Number of negative control molecules [Mock] The MDRE digestion efficiency can be calculated by: Efficiency = 1 - (number of positive control molecules) [MDRE] Number of positive control molecules [Mock] It can be calculated by the number of
[0183] As used herein, "mutation" refers to a variation from a known reference sequence, including, for example, single nucleotide variations (SNVs) and mutations such as insertions or deletions (indels). Mutations can be germline mutations or somatic mutations. In some embodiments, the reference sequence for comparison purposes is the wild-type genomic sequence of the target species from which the test sample is provided, typically the human genome.
[0184] As used herein, the terms "neoplasm" and "tumor" are used interchangeably. They refer to an abnormal growth of cells in a subject. A neoplasm or tumor can be benign, potentially malignant, or malignant. A malignant tumor is called a cancer or cancerous tumor.
[0185] As used herein, "nucleic acid tag" refers to a short nucleic acid (e.g., less than about 500 nucleotides, about 100 nucleotides, about 50 nucleotides, or about 10 nucleotides in length) that is used to distinguish nucleic acids from different samples (e.g., representing a sample index), different types, different fractions in the same sample (e.g., representing a fraction tag), or different nucleic acid molecules (e.g., representing a molecular barcode), or that has undergone different processing. Nucleic acid tags comprise predetermined, fixed, non-random, random, or semi-random oligonucleotide sequences. Such nucleic acid tags may be used to label different nucleic acid molecules or different nucleic acid samples or sub-samples. Nucleic acid tags can be single-stranded, double-stranded, or at least partially double-stranded. Nucleic acid tags can have the same length or various lengths, as desired. Nucleic acid tags can also include double-stranded molecules with one or more blunt ends, include 5' or 3' single-stranded regions (e.g., overhangs), and / or include one or more other single-stranded regions elsewhere within a given molecule. Nucleic acid tags can be attached to one or both ends of other nucleic acids (e.g., sample nucleic acids to be amplified and / or sequenced). Nucleic acid tags can be decoded to reveal information such as the origin, morphology, or processing of a given nucleic acid. For example, nucleic acid tags can also be used to enable pooling and / or parallel processing of multiple samples containing nucleic acids with different molecular barcodes and / or sample indices, with the nucleic acids subsequently deconvoluted by detecting (e.g., reading) the nucleic acid tags. Nucleic acid tags can also be referred to as identifiers (e.g., molecular identifiers, sample identifiers). Additionally or alternatively, nucleic acid tags can be used as molecular identifiers (e.g., to distinguish between amplicons of different molecules or different parent molecules in the same sample or subsample). This includes, for example, unique tagging of different nucleic acid molecules in a given sample or non-unique tagging of such molecules.In the case of non-unique tagging applications, a limited number of tags (i.e., molecular barcodes) may be used to tag each nucleic acid molecule such that different molecules can be distinguished based on their intrinsic sequence information (e.g., start and / or stop positions where they map to a selected reference genome, subsequences at one or both ends of the sequence, and / or sequence length) in combination with at least one molecular barcode. Typically, a sufficient number of different molecular barcodes are used so that there is a low probability (e.g., less than about 10%, less than about 5%, less than about 1%, or less than about 0.1% chance) that any two molecules will have the same intrinsic sequence information (e.g., start and / or stop positions, subsequences at one or both ends of the sequence, and / or length) and also have the same molecular barcode.
[0186] As used herein, "partitioning" refers to physically separating, sorting, and / or fractionating a mixture of nucleic acid molecules in a sample into multiple subsamples or subpopulations of nucleic acids based on characteristics of the nucleic acid molecules. A sample or population can be partitioned into one or more partitioned subsamples or subpopulations based on characteristics indicative of genetic or epigenetic alterations or disease states. The partitioning step can be a physical partitioning of the molecules. The partitioning step can include separating nucleic acid molecules into groups or sets based on the level of an epigenetic trait (e.g., methylation). For example, nucleic acid molecules can be partitioned based on the level of methylation of the nucleic acid molecules. In other words, the partitioning step can include physically partitioning nucleic acid molecules based on the presence or absence of one or more methylated nucleic acid bases. In some embodiments, methods and systems used for partitioning can be found in PCT Patent Application No. PCT / US2017 / 068329, which is incorporated herein by reference in its entirety.
[0187] As used herein, a "distributed set" or "fraction" refers to a set of nucleic acid molecules distributed into sets or groups based on the different binding affinities of the nucleic acid molecules or proteins associated with the nucleic acid molecules for a binder. A distributed set may also be referred to as a subsample. A binder preferentially binds to nucleic acid molecules containing nucleotides with epigenetic modifications. For example, if the epigenetic modification is methylation, the binder may be a methyl-binding domain (MBD) protein. In some embodiments, a distributed set may include nucleic acid molecules that belong to a particular level or degree of epigenetic trait (e.g., methylation). For example, the nucleic acid molecules may be distributed into three sets: one set for highly methylated nucleic acid molecules (first sub-sample, high fraction, highly distributed set, or highly methylated distributed set), a second set for low methylated nucleic acid molecules (second sub-sample, low fraction, low distributed set, or low methylated distributed set), and a third set for intermediately methylated nucleic acid molecules (third sub-sample, intermediate distributed set, intermediate methylation distributed set, residual fraction, or residual distributed set). In another example, the nucleic acid molecules may be distributed based on the number of methylated nucleotides: one distributed set may have nucleic acid molecules with nine methylated nucleotides, and another distributed set may have unmethylated nucleic acid molecules (zero methylated nucleotides).
[0188] As used herein, "sample" means anything that can be analyzed by the methods and / or systems disclosed herein.
[0189] As used herein, "sequencing" refers to any of several techniques used to determine the sequence (e.g., the identity and order of monomeric units) of a biomolecule, e.g., a nucleic acid such as DNA or RNA. Examples of sequencing methods include, but are not limited to, targeted sequencing, single-molecule real-time sequencing, exon or exome sequencing, intron sequencing, electron microscope-based sequencing, panel sequencing, transistor-mediated sequencing, direct sequencing, random shotgun sequencing, Sanger dideoxy termination sequencing, whole genome sequencing, sequencing by hybridization, pyrosequencing, double-strand sequencing, cycle sequencing, single-base extension sequencing, solid-phase sequencing, and hybridization sequencing. These include throughput sequencing, massively parallel signature sequencing, emulsion PCR, co-amplification-PCR (COLD-PCR) at low denaturation temperature, multiplex PCR, reversible dye terminator sequencing, paired-end sequencing, near-term sequencing, exonuclease sequencing, ligation sequencing, short-read sequencing, single molecule sequencing, sequencing by synthesis, real-time sequencing, reverse terminator sequencing, long-read sequencing, nanopore sequencing, 454 sequencing, Solexa Genome analyzer sequencing, SOLiD sequencing, MS-PET sequencing, and combinations thereof.In some embodiments, sequencing can be carried out by genetic analyzer, such as the genetic analyzer commercially available from Illumina, Inc., Pacific Biosciences, Inc., or Applied Biosystems / Thermo Fisher Scientific, and many others.
[0190] As used herein, "long-read sequencing" refers to a sequencing method that can generate longer sequencing reads, such as reads exceeding 10 kilobases, compared to short-read sequencing methods (e.g., Illumina NovaSeq, HiSeq, NextSeq, and MiSeq instruments, BGI MGISEQ and BGISEQ models, or Thermo Fisher Ion Torrent sequencers), which typically generate reads up to approximately 600 bases in length. Long-read sequencing methods are sometimes referred to as "third-generation sequencing." Compared to short reads, long reads can improve the detection of de novo assembly, transcript isoform identification, and structural variant mapping confidence. Furthermore, long-read sequencing of native DNA or RNA molecules reduces amplification bias and preserves base modifications. Long-read sequencing technologies useful herein may include any suitable long-read sequencing method, such as, but not limited to, Pacific Biosciences (PacBio) single molecule real-time (SMRT) sequencing, Oxford Nanopore Technologies (ONT) nanopore sequencing, and synthetic long-read sequencing approaches, such as concatenated reads, proximal ligation strategies, and optical mapping.
[0191] As used herein, "nanopore sequencing" refers to a sequencing method that directly sequences single-stranded DNA molecules (e.g., native DNA molecules) by measuring characteristic signals (e.g., current changes) as bases are threaded through a membrane nanopore, for example, by a molecular motor protein. Nanopore sequencing methods can therefore generate significantly longer reads than those generated by next-generation (also known as second-generation) sequencing technologies.
[0192] As used herein, "5-letter sequencing" and "6-letter sequencing" refer to sequencing methods that can sequence A, C, T, and G in addition to 5mC and 5hmC in a single workflow to provide a 5-letter (A, C, T, G, and either 5mC or 5hmC) or 6-letter (A, C, T, G, 5mC, and 5hmC) digital readout, respectively. Such methods may use, for example, enzymatic treatment of DNA samples rather than bisulfite treatment.
[0193] As used herein, "next-generation sequencing" or "NGS" refers to a sequencing technology that has increased throughput compared to traditional Sanger and capillary electrophoresis-based approaches, for example, with the ability to generate hundreds of thousands of relatively small sequence reads at a time.Some examples of next-generation sequencing methods include, but are not limited to, sequencing by synthesis, sequencing by ligation, and sequencing by hybridization.In some embodiments, next-generation sequencing involves the use of an instrument that can sequence single molecules.Examples of commercially available instruments for performing next-generation sequencing include, but are not limited to, NextSeq, HiSeq, NovaSeq, MiSeq, Ion PGM, and Ion GeneStudio S5.
[0194] As used herein, the terms "somatic mutation" and "somatic variation" are used interchangeably. They refer to mutations in the genome that occur after conception. Somatic mutations can occur in any cell of the body except germ cells and are therefore not passed on to offspring.
[0195] As used herein, "subject" refers to an animal, such as a mammalian species (e.g., a human) or an avian (e.g., an avian) species, or other organism, such as a plant. More specifically, the subject may be a vertebrate, e.g., a mammal, such as a mouse, a primate, a monkey, or a human. Animals include livestock (e.g., beef cattle, dairy cattle, poultry, horses, pigs, and the like), sport animals, and companion animals (e.g., pets or support animals). A subject may be a healthy individual, an individual having or suspected of having a disease or predisposition to a disease, or an individual in need of treatment or suspected of needing treatment. The terms "individual" or "patient" are intended interchangeably with "subject." For example, a subject may be an individual who has been diagnosed with cancer, an individual who will undergo cancer treatment, and / or an individual who has undergone at least one cancer treatment. A subject may be in remission from cancer. As another example, a subject may be an individual who has been diagnosed with an autoimmune disease. As another example, the subject may be a female individual who may be diagnosed with or suspected of having a disease, e.g., cancer, an autoimmune disease, who is pregnant or planning to become pregnant.
[0196] "Or" is used in the inclusive sense, ie, equivalent to "and / or" unless the context requires otherwise. II. Exemplary Methods A. Overview
[0197] The formation and progression of cancer can result from both genetic alterations and epigenetic traits of deoxyribonucleic acid (DNA).The present disclosure provides a method and system for analyzing DNA, for example, cell-free DNA (cfDNA).The present disclosure provides a single-base resolution sequencing method that incorporates barcoding and methylation-specific amplification, and exhibits increased sensitivity (for example, the ability to reliably detect epigenetic traits in samples containing fewer molecules) and / or reduced molecular loss, compared to other sequencing methods that include a pre-amplification epigenetic base conversion step.
[0198] Without wishing to be bound by any particular theory, the cells in or around cancer or neoplasm may shed more DNA than the cells of the same tissue type of healthy subjects.Therefore, the distribution of the tissue of origin of a certain DNA sample, for example, cfDNA, may change during carcinogenesis.Thus, for example, the increased level of hypermethylated variable target region, which shows lower methylation in healthy cfDNA than in at least one other tissue type, can be an indication of the presence of cancer (or recurrence depending on the subject's medical history).Similarly, the increased level of hypomethylated variable target region in sample can be an indication of the presence of cancer (or recurrence depending on the subject's medical history).
[0199] In addition, cancer can be manifested by non-sequence alterations such as methylation. Examples of methylation changes in cancer include localized gain of DNA methylation in CpG islands at the TSSs of genes involved in normal growth control, DNA repair, cell cycle regulation, and / or cell differentiation. This hypermethylation can be associated with abnormal loss of transcriptional capacity of the involved genes, occurring at least as frequently as point mutations and deletions as a cause of altered gene expression.
[0200] In this way, DNA methylation profiling can be used to detect the abnormal methylation in the DNA of sample.For example, because of the abnormal increase in the contribution of tissue to sample type (for example, due to the increased DNA loss in or around neoplasia or cancer), and / or the abnormal increase in the contribution from the degree of genomic methylation that changes during development or is disrupted by disease, for example, cancer or any cancer-related disease, DNA can correspond to certain genomic regions (" differentially methylated regions " or " DMR ") that are usually hypermethylated or hypomethylated in given sample type (for example, cfDNA from bloodstream), but can show the abnormal degree of methylation that correlates with neoplasia or cancer.
[0201] In some embodiments, DNA methylation comprises the addition of a methyl group to a cytosine residue at a CpG site (cytosine-phosphate-guanine site (i.e., a cytosine followed by a guanine in the 5' to 3' direction of a nucleic acid sequence). In some embodiments, DNA methylation comprises the addition of a methyl group to an adenine residue, e.g., N6-methyladenine. In some embodiments, DNA methylation is 5-methylation (modification of the 5 carbon of the 6-membered ring of cytosine). In some embodiments, 5-methylation comprises the addition of a methyl group to the 5C position of a cytosine residue to produce 5-methylcytosine (m5c or 5-mC or 5mC). In some embodiments, methylation comprises derivatives of m5c, including, but not limited to, 5-hydroxymethylcytosine (5-hmC or 5hmC), 5-formylcytosine (5-fC), and 5-carboxyl Examples of DNA methylation include cytosine (5-caC). In some embodiments, DNA methylation is 3C methylation (modification of the nitrogen at the 3 position of the 6-membered ring of a cytosine residue). In some embodiments, 3C methylation involves the addition of a methyl group to the 3C position of a cytosine residue to generate 3-methylcytosine (3mC). Methylation can also occur at non-CpG sites, for example, methylation can occur at CpA, CpT, or CpC sites. DNA methylation can alter the activity of methylated DNA regions. For example, if DNA in a promoter region is methylated, gene transcription can be suppressed. DNA methylation is crucial for normal development, and abnormal methylation can disrupt epigenetic regulation. Disruption of epigenetic regulation, for example, suppression, can cause diseases such as cancer. Promoter methylation in DNA can indicate cancer.
[0202] Methylation profiling can involve determining methylation patterns across different regions of the genome, which can indicate regions of the genome that are more or less highly methylated compared to other regions, thus allowing genomic regions, as opposed to individual molecules, to differ in their degree of methylation.
[0203] However, current single-base resolution sequencing methods result in input DNA loss, resulting in the recovery of only approximately 1–25% of input DNA molecules. These base conversion steps are generally performed prior to amplification. To overcome this challenge, methods have been developed that attempt to amplify DNA while preserving the original DNA methylation pattern. These methods involve the use of the methyltransferase DNMT1, which effectively maintains the methylation state of bases during DNA replication and preferentially methylates the opposite strand of a hemimethylated substrate (CpG). However, in practice, DNMT1 generally exhibits incomplete preservation of methylation during in vitro amplification, with <100% efficiency in methylating across hemimethylated sites and / or >0% efficiency in methylating unmethylated dsDNA sites (i.e., de novo methylation). This limits the usefulness of current methylation-preserving amplification methods. While molecular loss can be reduced compared to standard methylation sequencing (or qPCR) methods, accuracy can be limited.
[0204] In certain aspects, the present disclosure provides a solution for overcoming the current limitations of methylation-preserving amplification by using molecular barcodes in methylation-preserving amplification, for example, to correct or eliminate errors due to methyltransferase activity. In other embodiments, barcodes are not used, and methylation-preserving amplification is a linear amplification method, meaning that errors generated by methyltransferases do not propagate exponentially in the resulting family of DNA copies. Without being bound by theory, the original template DNA molecule contains the "correct" epigenetic mark (i.e., the epigenetic mark of the sample DNA prior to the introduction of any errors during sample processing, e.g., during the amplification step). Thus, in some embodiments in which DNA is linearly amplified using methylation-preserving amplification, DNA copies are synthesized only from the original template molecule (e.g., using only a reverse primer), resulting in reduced errors compared to standard exponential amplification (e.g., PCR using forward and reverse primers, which amplifies both the original and copy strands, thereby potentially amplifying strands containing errors introduced during the amplification step).
[0205] The methods described herein, in which barcodes are used, result in a family of DNA copies generated from an original template with a shared molecular barcode that can be identified during sequencing. Sequence errors from DNA polymerases and methylation errors from DNMT1 infidelity are seen as variation between DNA copies (e.g., DNA molecules) from the same family, and DNA copies generated from the original template molecule have (a) a shared molecular barcode, (b) the same genomic start position, (c) the same genomic stop position, or a combination of (a), (b), and / or (c) at a given base. After epigenetic base conversion, an epigenetic consensus sequence can be determined from multiple reads of the family, e.g., from all copies / NGS reads of the family, thereby suppressing errors generated during methylation-preserving amplification.
[0206] This workflow also allows for the correction of sequence errors introduced in vitro by DNA polymerases. Base conversion workflows that convert methylated bases (e.g., TAPS, which converts methylated cytosine to thymine) can be used with the disclosed methylation-preserving amplification method to suppress both genetic and epigenetic errors.
[0207] In some embodiments, the method of analyzing DNA includes, in the following order (as indicated by lettering): (a) performing methylation-preserving amplification of DNA, wherein the DNA comprises a barcode; and (b) sequencing the DNA in a modification-sensitive manner and determining an epigenetic consensus sequence of the DNA associated with at least a portion of the barcode. In some embodiments, a method of analyzing DNA includes, in the following order (as indicated by lettering): (a) performing methylation-preserving amplification of DNA, wherein the DNA comprises a barcode; (b) subjecting the DNA to a procedure that affects a first nucleobase of the DNA differently from a second nucleobase of the DNA, wherein the first nucleobase is a modified nucleobase or an unmodified nucleobase, and the second nucleobase is a modified nucleobase or an unmodified nucleobase that is different from the first nucleobase, and wherein the first nucleobase and the second nucleobase have the same base-pairing specificity; and (c) sequencing the DNA and determining an epigenetic consensus sequence of the DNA associated with at least a portion of the barcode. In some embodiments, the method of analyzing DNA includes, in the following order (as indicated by lettering): (a) performing a methylation-preserving amplification of DNA, wherein the DNA comprises a barcode; (b) enriching the DNA for one or more sets of epigenetic target regions of the DNA, thereby providing enriched DNA; (c) sequencing the enriched DNA in a modification-sensitive manner and determining an epigenetic consensus sequence of the enriched DNA associated with at least a portion of the barcode. In some embodiments, the method of analyzing DNA includes, in the following order (as indicated by lettering): (a) performing a linear methylation-preserving amplification of DNA; and (b) sequencing the DNA in a modification-sensitive manner and determining an epigenetic consensus sequence of the DNA.In some embodiments, the method of analyzing DNA comprises, in the following order (as indicated by lettering): (a) performing linear methylation-preserving amplification of the DNA; (b) subjecting the DNA to a procedure that affects a first nucleobase of the DNA differently from a second nucleobase of the DNA, wherein the first nucleobase is a modified nucleobase or an unmodified nucleobase, and the second nucleobase is a modified nucleobase or an unmodified nucleobase that is different from the first nucleobase, and wherein the first nucleobase and the second nucleobase have the same base-pairing specificity; and (c) sequencing the DNA and determining an epigenetic consensus sequence of the DNA. In some embodiments, the method of analyzing DNA includes, in the following order (as indicated by lettering): (a) performing linear methylation-preserving amplification of DNA; (b) enriching the DNA for one or more sets of epigenetic target regions of DNA, thereby providing enriched DNA; (c) sequencing the enriched DNA in a modification-sensitive manner, and determining an epigenetic consensus sequence of the enriched DNA.
[0208] In certain embodiments where the conversion method yields unmethylated cytosines (e.g., bisulfite conversion, EM-Seq, or SEM-Seq, as described elsewhere herein), when an adapter is ligated to DNA, including cytosines in a non-CpG context, (1) if the cytosine is methylated, this methylation is not propagated by a DNA methyltransferase (e.g., DNMT1), and (2) if the cytosine is unmethylated (as they generally are after amplification), it is deaminated during base conversion using a deaminase enzyme or a reagent that acts on unmethylated cytosines, such as bisulfite. To address these limitations, in some embodiments of the disclosed methods, methylation-preserving amplification of DNA is performed, and the adapter attached to the DNA contains a barcode and a restriction enzyme cleavage site between the barcode and at least a portion of the adapter. In such embodiments, the barcode is positioned such that cleavage of the adapter at the restriction enzyme cleavage site by a restriction enzyme does not remove the barcode from the DNA, e.g., the barcode is positioned between the DNA sequence to be analyzed and the restriction site. For example, in an adapter ligated to the 5' end of the DNA sequence to be analyzed, the barcode is 3' of the restriction site and 5' of the DNA sequence to be analyzed. Thus, in some embodiments of the disclosed method, methylation-preserving amplification of DNA is performed, wherein the DNA comprises an insert and an adapter comprising a barcode, at least one of the adapters further comprising a restriction enzyme cleavage site between the barcode and a portion of the adapter, and the barcode is located between the insert and the restriction enzyme cleavage site, thereby providing amplified DNA. In such embodiments, the insert can comprise DNA to be analyzed, such as DNA from a sample (e.g., cfDNA or DNA from a sample containing cells (e.g., a sample from a subject)) or from its amplicon. Furthermore, in some embodiments, the barcode does not contain cytosines in non-CpG contexts.For example, in some embodiments, where the procedure affecting a first nucleobase in the DNA differently from a second nucleobase in the DNA of the first sub-sample comprises enzymatic methyl (EM) base conversion (EM-Seq) (or another base conversion method that converts unmethylated cytosines, such as bisulfite conversion or SEM-Seq described elsewhere herein), the barcode may be free of cytosines in non-CpG contexts. In such embodiments, if the barcode contains cytosines in non-CpG contexts, these cytosines will be converted to uracil (and read as thymine during sequencing); therefore, the design and sequencing of such barcodes should be considered as this avoids miscalling bases in the barcode during sequencing. In other embodiments, the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA of the first sub-sample includes Tet-assisted pyridine borane sequencing (TAPS) (or another base conversion method that does not convert unmethylated cytosines), the barcode may or may not include cytosines in a non-CpG context.
[0209] In some embodiments, such as those in which the methylation-preserving amplification step comprises rolling circle amplification, the DNA molecule being analyzed is circular. In such embodiments, "a restriction enzyme cleavage site between the barcode and a portion of the adapter" means that the restriction enzyme cleavage site is located on the shortest path along the molecule from the barcode to the portion of the adapter.
[0210] In some embodiments, the DNA is then contacted with a restriction enzyme (e.g., a restriction endonuclease) that cleaves the DNA at the restriction enzyme cleavage site in the adapter. In some embodiments, cleavage of the DNA by the restriction enzyme results in an overhang (e.g., an A or T overhang), which can be used, for example, as a cohesive end for subsequent ligation. Before or after the contacting step (e.g., after the contacting step), the DNA is subjected to a procedure that affects a first nucleobase of the DNA so that it is different from a second nucleobase of the DNA, where the first nucleobase is a modified or unmodified nucleobase, and the second nucleobase is a modified or unmodified nucleobase different from the first nucleobase, and the first nucleobase and the second nucleobase have the same base pairing specificity. In some such embodiments, the first nucleobase is an unmethylated cytosine, and the subjecting step (e.g., bisulfite conversion, EM-Seq, or SEM-Seq, as described elsewhere herein) affects the unmethylated cytosine. Auxiliary adapters (which in some embodiments do not include barcodes) are then ligated to the DNA, and uracil- and / or dihydrouracil-resistant amplification of the DNA is performed, e.g., PCR using a uracil- and / or dihydrouracil-resistant DNA polymerase (e.g., Q5U® Hot Start High-Fidelity DNA Polymerase from New England Biolabs, VeraSeq™ ULtra DNA Polymerase from Qiagen, or Phusion™ U Hot Start DNA Polymerase from Thermo Fisher). In some embodiments, the DNA is then enriched for one or more sets of epigenetic and / or sequence-variable target regions of the DNA, thereby providing enriched DNA. In some embodiments, the enriched DNA is amplified, where the amplification includes, for example, differentially tagging the enriched DNA for one or more sets of epigenetic and / or sequence-variable target regions.In some such embodiments, the differential tagging step comprises attaching the sample index to the DNA, for example by ligation or by incorporation using a PCR primer.
[0211] In certain embodiments, the step of contacting the DNA with a restriction enzyme that recognizes and cleaves the DNA at the restriction enzyme cleavage site in the adaptor is performed before the step of subjecting the DNA to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA, and the contacting step can also or alternatively be performed before or after the enrichment step. However, in embodiments in which the restriction enzyme cleavage site does not contain a cytosine in a non-CpG context, the step of contacting the DNA with a restriction enzyme that recognizes and cleaves the DNA at the restriction enzyme cleavage site in the adaptor can be performed (a) before or after the step of subjecting the DNA to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA, (b) before or after the enrichment step, and (c) before an optional step of amplifying the enriched DNA, where the amplifying step includes, for example, differentially tagging the enriched DNA for one or more sets of epigenetic target regions and / or one or more sets of sequence variable target regions.
[0212] The DNA is then sequenced, and the epigenetic consensus sequence of the DNA associated with at least a portion of the barcode is determined. In some embodiments, the epigenetic consensus sequence of the DNA associated with at least a portion of the barcode is determined by performing bioinformatics analysis of sequencing (e.g., NGS) data. In such embodiments, one or more components of each read (e.g., a barcode or barcode set combined with a genomic start site and / or a genomic stop site) are used to identify the read and group the read into a molecular family (e.g., DNA molecule), and the reads within the molecular family have (a) a common molecular barcode (i.e., a set of reads, each of which contains a barcode with the same sequence) and (b) the same genomic start position, (c) the same genomic stop position, or (d) both (b) and (c).
[0213] Restriction enzymes can recognize and bind to specific nucleotide sequences in DNA ("restriction sites," "restriction enzyme cleavage sites"), then cleave the DNA into fragments at locations within the restriction site via hydrolysis of the phosphodiester backbone. Restriction enzymes useful for cleaving the restriction enzyme cleavage sites in the adapters disclosed herein are known in the art and can generate either blunt or sticky ends after cleavage. The sticky ends resulting from cleavage typically have 3'- or 5'-overhangs of 1 to 4 (or more) nucleotides. In some embodiments, restriction enzymes that generate single-base overhangs, such as single-base 5'-overhangs, are used. In some embodiments, the single-base overhang (e.g., single-base 5'-overhang) is T. In some embodiments, the single-base overhang (e.g., single-base 5'-overhang) is A. In embodiments in which auxiliary adapters are ligated to the DNA to be analyzed after the conversion step and, optionally, the enrichment step, sticky-end or blunt-end ligation can be used. In certain embodiments, the restriction enzyme used to cleave the adapter at the restriction enzyme cleavage site generates a sticky end, and a secondary adapter (e.g., an adapter that does not contain a barcode) is ligated to the DNA using sticky end ligation after the enrichment step and before the optional step of amplifying the enriched DNA.
[0214] In some embodiments, a method of analyzing DNA comprises, in the following order (as indicated by lettering): (a) performing methylation-preserving amplification of said DNA, thereby providing amplified DNA, wherein the DNA comprises an insert and an adaptor comprising a barcode, at least one of the adaptors further comprising a restriction enzyme cleavage site between the barcode and a portion of the adaptor, the barcode being located between the insert and the restriction enzyme cleavage site; (b) contacting the amplified DNA with a restriction enzyme that recognizes and cleaves the DNA at the restriction enzyme cleavage site in the adaptor; (c) before or after step (b), subjecting the amplified DNA to a procedure that affects a first nucleobase of the amplified DNA so that it differs from a second nucleobase of the amplified DNA, wherein the first nucleobase is a modified nucleobase or an unmodified nucleobase; The method includes the steps of: (d) ligating auxiliary adapters to the amplified DNA after steps (b) and (c); (e) performing uracil and / or dihydrouracil-resistant amplification of DNA after step (d); (f) enriching the amplified DNA for one or more sets of epigenetic target regions of DNA after step (e), thereby obtaining enriched DNA; (g) optionally further amplifying the enriched DNA; and (h) sequencing the enriched DNA and determining the epigenetic consensus sequence of the DNA linked to at least a portion of the barcode. In some embodiments, the barcode does not contain cytosine in a non-CpG context. In some embodiments, the adapter ligated to DNA in step (d) does not contain a barcode. In some embodiments, an optional step of amplifying the enriched DNA is performed, further comprising differentially tagging the enriched DNA.In some such embodiments, differentially tagging the enriched DNA comprises attaching one or more sample indexes to the DNA, for example, by ligation or by incorporation using a PCR primer. In some embodiments, cleavage of the DNA with a restriction enzyme results in an overhang, for example, a single-base overhang, for example, a single-base 5'-overhang. In some such embodiments, the single-base overhang is T. In other such embodiments, the single-base overhang is A. B. Amplification
[0215] In some embodiments, DNA is amplified.For example, the DNA flanked by the adapter added to the DNA described herein can be amplified using methylation-preserving amplification method.The amplification method used herein can include any suitable method, for example, the method known to those skilled in the art.In some embodiments, amplification is initiated by a primer that binds to the primer binding site in the adapter flanking the DNA molecule to be amplified. Amplification methods can involve cycles of denaturation, annealing, and extension due to thermal cycling, e.g., polymerase chain reaction (PCR), or can be isothermal, e.g., linear amplification, transcription-mediated amplification, recombinant polymerase amplification (RPA), helix-dependent amplification (HDA), loop-mediated isothermal amplification (LAMP) (Notomi et al., Nuc. Acids Res., 28, e63, 2000), rolling circle amplification (RCA) (Blanco et al., J. Biol. Chem., 264, 8935-8940, 1989), or hyperbranched rolling circle amplification (Lizard et al., Nat. Genetics, 19, 225-232, 1998). In some embodiments, DNA is amplified using linear amplification with thermal cycling and DNMT1 (Chang, et al., "DNA 5-Methylcytosine-Specific Amplification and Sequencing," J. Am. Chem. Soc. 2020, 142(10):4539-4543). Other amplification methods used herein include ligase chain reaction, strand displacement amplification, nucleic acid sequence-based amplification, and sequence-based self-sustained replication. In some embodiments, detecting the presence or absence of one or more DNA sequences comprises methylation-preserving amplification, e.g., amplification performed in the presence of a methyltransferase.
[0216] The methylation agent used in the methylation-preserving amplification method described herein is known to those skilled in the art and may include, for example, any suitable methyltransferase. In some embodiments, the methylation agent is DNMT1. DNMT1 is the most abundant DNA methyltransferase in mammalian cells, and preferentially methylates hemimethylated CpG dinucleotides in mammalian genomes. For example, DNA molecules replicated using PCR amplification with DNMT1 incubation maintain their methylation state after amplification for use in further analysis, such as those described herein (e.g., epigenetic base conversion step and / or enrichment step).
[0217] Additional methylating agents useful herein include mammalian methyltransferases DNMT3a and DNMT3b, plant methyltransferase MET1, and CMT3. In some embodiments, DNMT1 or another suitable methyltransferase is used with a methyl donor, with or without cofactors known to those skilled in the art. DNMT1 works at 95% efficiency in vitro without cofactors; however, DNMT1 may also be used with cofactors such as NP95 (Uhrf1), as described in Bashtrykov PI, et al. "The UHRF1 protein stimulates the activity and specificity of the maintenance DNA methyltransferase DNMT1 by an allosteric mechanism," J. Biol. Chem. 2014. In some embodiments, DNMT1 is used at a concentration of about 50 to 10,000 U / mL, e.g., about 50 to 2,000, about 50 to 5,000, about 2,500 to 7,500, or about 5,000 to 10,000 U / mL. In some embodiments, DNMT1 is used at a concentration of about 100 to 500, about 500 to 1,000, about 100 to 1,000, about 1,000 to 1,500, about 500 to 1,500, about 600 to 1,400, about 700 to 1,300, about 800 to 1,200, about 900 to 1,100, or about 950 to 1,050 U / mL. In some embodiments, DNMT1 is used at a concentration of about 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250, 1300, 1350, 1400, 1450, 1500, 1550, 1600, 1650, 1700, 1750, 1800, 1850, 1900, 1950, or about 2000 U / mL. In some embodiments, DNMT1 is used at a concentration of about 1,000 U / ml.
[0218] Some PCR enzymes, such as archaeal polymerases, do not efficiently copy bisulfite-treated DNA due to stalling induced by template uracil. Thus, in embodiments where DNA is amplified after a conversion step that generates uracil or dihydroxyuracil (e.g., bisulfite conversion, SEM-Seq, or TAP), amplification can include uracil- and / or dihydrouracil-resistant amplification, such as PCR using a uracil- and / or dihydrouracil-resistant DNA polymerase. Exemplary such polymerases are known in the art and commercially available (e.g., Q5U® Hot Start High-Fidelity DNA Polymerase from New England Biolabs, VeraSeq™ ULtra DNA Polymerase from Qiagen, and Phusion™ U Hot Start DNA Polymerase from Thermo Fisher).
[0219] C. Adapter ligation or addition; tagging
[0220] In some embodiments, the disclosed methods include adding an adapter to DNA. In some embodiments, the adapter is added (e.g., "attached") to the DNA before or during an amplification step, such as a methylation-preserving amplification step, before or after subjecting the DNA to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA, and / or before or after enrichment of epigenetic and / or sequence-variable target regions of the DNA. In some embodiments, the adapter may be added to the DNA simultaneously with the amplification procedure (e.g., a methylation-preserving amplification step) before or after the amplification step, for example, by providing the adapter at the 5' portion of the primer (if PCR is used, this may be referred to as library prep-PCR or LP-PCR). In some embodiments, the adapter is added by other approaches. In some such methods, a first adapter is added to the nucleic acid by ligation to its 3' end, which may include ligation to single-stranded DNA. In other such methods, a first adapter is added to the nucleic acid by ligation to its 5' end, which may include ligation to single-stranded DNA. The adapter can be used as an initiation site for double-stranded synthesis, for example, using a universal primer and DNA polymerase. A second adapter can then be ligated to at least the 3' end of the second strand of the now double-stranded molecule. In some embodiments, the first adapter contains an affinity tag, such as biotin, and the nucleic acid ligated to the first adapter is attached to a solid support (e.g., beads) that can contain a binding partner for the affinity tag, such as streptavidin. For further discussion of related procedures, see Gansauge et al., Nature Protocols 8:737-748 (2013). Commercially available kits for sequencing library preparation compatible with single-stranded nucleic acids are available, such as the Accel-NGS® Methyl-Seq DNA Library Kit from Swift Biosciences. In some embodiments, after adapter ligation, the nucleic acid is amplified.In some embodiments, DNA end repair is performed prior to the addition of the adapters.
[0221] In certain embodiments, the first adaptor is added to the nucleic acid by ligation to its 3' end, which may include ligation to single-stranded DNA.Next, methyl-preserving linear amplification is carried out using the 3' adaptor as a primer binding site, followed by subjecting the DNA to a procedure that affects the first nucleobase of the DNA so that it is different from the second nucleobase of the DNA, where the first nucleobase is unmethylated cytosine, and subjecting the step (for example, bisulfite conversion, EM-Seq, or SEM-Seq, as described elsewhere herein) to unmethylated cytosine.The amplified converted DNA is then subjected to a single-stranded DNA library preparation step.
[0222] In some embodiments, single-stranded DNA library preparation is performed using a one-step phosphorylation / ligation reaction combination, as described, for example, in Troll et al., BMC Genomics, 20:1023 (2019), available at https: / / doi.org / 10.1186 / s12864-019-6355-0. This method, called Single-Reaction Single-Stranded LibrarY ("SRSLY"), can be performed without end polishing. SRSLY can be useful for converting short fragmented DNA molecules, such as cfDNA fragments, into a sequencing library while retaining their native length and ends. The SRSLY method can generate sequencing libraries (e.g., Illumina sequencing libraries) from fragmented or degraded template (input) DNA. In certain embodiments, the template DNA is first heat-denatured and then immediately subjected to a cold shock to render the template DNA molecules single-stranded. The DNA can be maintained as single-stranded throughout the ligation reaction by the inclusion of a thermostable single-strand binding protein (SSB). The template DNA, which is now single-stranded and may be coated with SSB, is then subjected to a phosphorylation / ligation duplex reaction using directional dsDNA NGS adapters containing single-stranded overhangs. Both forward and reverse sequencing adapters may share a similar structure, but differ in that the ends are unblocked to facilitate proper ligation. Both sequencing adapters may contain a dsDNA portion and single-stranded splint overhangs of random nucleotides that occur at the 3-prime end of the lower strand of the forward adapter and the 5-prime end of the lower strand of the reverse adapter. In this way, the forward adapter (e.g., (P5) Illumina adapter) can be delivered to the 5-prime end of the template molecule, and the reverse adapter (e.g., (P7) Illumina adapter) can be delivered to the 3-prime end of the template molecule. In this way, the native polarity of the input DNA molecule can be preserved.
[0223] During the dual phosphorylation / ligation reaction, T4 polynucleotide kinase (PNK) can be used to phosphorylate the 5-prime end and dephosphorylate the 3-prime end to prepare template DNA ends for ligation. T4 PNK works on both ssDNA and dsDNA molecules and has no activity against the phosphorylation state of proteins. Simultaneously, random nucleotides of the splint adapter can be annealed to the single-stranded template molecule. This creates a short, localized dsDNA molecule, allowing ligation of the template to the adapter by a ligase such as T4 DNA ligase, which has high ligation efficiency for dsDNA templates but low efficiency for ssDNA. After the single phosphorylation / ligation reaction is complete, the library DNA can be purified and directly subjected to standard NGS indexing PCR, compatible with both conventional single- and dual-index primers.
[0224] In some embodiments, after adapter binding, the nucleic acid is subjected to amplification, such as methylation-preserving amplification, which can use, for example, universal primers that recognize primer-binding sites in the adapters.
[0225] In some embodiments, the DNA is ligated at both ends to Y-shaped adapters that contain primer binding sites and tags. In some such embodiments, the DNA is amplified.
[0226] Tagging a DNA molecule is the process of attaching or associating a tag to a DNA molecule. Such a tag can be a molecule, such as a nucleic acid, that contains information that characterizes the molecule to which the tag is attached. The tag can enable distinguishing the molecule from which a sequence read originated. For example, a molecule can have a sample tag (which distinguishes molecules in one sample from molecules in a different sample) or a molecular tag / molecular barcode / barcode (which distinguishes different molecules from each other in both unique and non-unique tagging scenarios). For methods that include a partitioning step, a fraction tag (which distinguishes molecules in one fraction from molecules in a different fraction) may also be included. In some embodiments, the adapter added to the DNA molecule comprises a tag. In some such embodiments, the tag comprises a barcode or a combination of barcodes. As used herein, the term "barcode" refers to a nucleic acid molecule having a specific nucleotide sequence or the nucleotide sequence itself, depending on the context. A barcode can have, for example, between 10 and 100 nucleotides. A collection of barcodes may have degenerate sequences, or may have sequences with a certain Hamming distance if desired for a particular purpose. Thus, for example, a molecular barcode may be composed of one barcode or a combination of two barcodes, each attached to a different end of a molecule. Additionally or alternatively, different sets of molecular barcodes or molecular tags may be used for different fractions and / or samples, so that the barcodes serve as molecular tags through their individual sequences and serve to identify the corresponding fractions and / or samples based on the sets they are members of. Tags, including barcodes, may be incorporated into adapters or otherwise attached to adapters. Tags may be incorporated by ligation, overlap extension PCR, among other methods.
[0227] Tagging strategies can be divided into unique tagging strategies and non-unique tagging strategies. In unique tagging, all or substantially all molecules in a sample have different tags, thereby allowing reads to be assigned to original molecules based on a single tag information. Tags used in such methods are sometimes called "unique tags." In non-unique tagging, different molecules in the same sample can have the same tag, thereby assigning sequence reads to original molecules using other information in addition to tag information. Such information may include start and stop coordinates, coordinates that map the molecule, a single start or stop coordinate, etc. Tags used in such methods are sometimes called "non-unique tags." In some embodiments, non-unique tags include non-unique barcodes. Therefore, it is not necessary to uniquely tag every molecule in a sample. This is sufficient to uniquely tag molecules within an identifiable class within a sample. In this way, molecules in different identifiable families can have the same tag without losing information about the identity of the tagged molecule.
[0228] In some embodiments, the adapter comprises a sufficient number of different tags such that the number of tag combinations provides a low probability, for example, 95, 99, or 99.9%, that two nucleic acids with the same start and stop points will receive the same combination of tags. Regardless of whether the adapter has the same or different tags, it can comprise the same or different primer binding sites. In some embodiments, the adapter comprises the same primer binding site.
[0229] In certain embodiments of non-unique tagging, the number of different tags used may be sufficient such that there is a very high probability (e.g., at least 99%, at least 99.9%, at least 99.99%, or at least 99.999%) that all molecules in a particular group have different tags. In some embodiments involving random barcode attachment, e.g., at both ends of the molecule, the combination of barcodes together constitutes a tag. This number in terms is a function of the number of molecules in the call. For example, a class may be all molecules mapping to the same start-stop position in a reference genome. A class may be all molecules mapping to a particular locus, e.g., a particular base or across a particular region (e.g., up to 100 bases or a gene or exon of a gene). In certain embodiments, the number of different tags used to uniquely identify the number z of molecules in a class is 2 * z, 3 * z, 4 * z, 5 * z, 6 * z, 7 * z, 8 * z, 9 * z, 10 * z, 11 * z, 12 * z, 13 * z, 14 * z, 15 * z, 16 * z, 17 * z, 18 * z, 19 * z, 20 * z or 100 * z (e.g., lower bound) and 100,000 * z, 10,000 * z, 1000 * z or 100 * z (e.g., upper limit).
[0230] For example, in a sample of about 5 ng to 30 ng of cell-free DNA, approximately 3,000 molecules are expected to map to a particular nucleotide coordinate, with about 3 to 10 molecules with any given start coordinate expected to share the same stop coordinate. Therefore, about 50 to about 50,000 different tags (e.g., about 6 to 220 barcode combinations) may be sufficient to uniquely tag all such molecules. To uniquely tag all 3,000 molecules mapping across nucleotide coordinates, about 1 million to about 20 million different tags would be required.
[0231] Generally, the assignment of unique or non-unique tags (e.g., non-unique barcodes) in the reaction follows the methods and systems described in U.S. Patent Applications Nos. 20010053519, 20030152490, 20110160078, and U.S. Patent Nos. 6,582,908, 7,537,898, and 9,598,731. Tags can be randomly or non-randomly linked to the sample nucleic acids.
[0232] In some embodiments, tagged nucleic acid is loaded into microwell plate and then sequenced.Microwell plate can have 96, 384 or 1536 microwells.In some cases, they are introduced with the expected ratio of unique tags to microwells.For example, unique tags can be loaded so that each genome sample is loaded with more than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 10000, 50,000, 100,000, 500,000, 1,000,000, 10,000,000, 50,000,000 or 1,000,000,000 unique tags. In some cases, unique tags may be loaded such that less than about 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 10000, 50,000, 100,000, 500,000, 1,000,000, 10,000,000, 50,000,000, or 1,000,000,000 unique tags are loaded per genomic sample. In some cases, the average number of unique tags loaded per sample genome is about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 10000, 50,000, 100,000, 500,000, 1,000,000, 10,000,000, 50,000,000, or 1,000,000,000 1,000, or more than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 10000, 50,000, 100,000, 500,000, 1,000,000, 10,000,000, 50,000,000 or 1,000,000,000 unique tags per genome sample.
[0233] In some embodiments, 20 to 50 different tags (e.g., barcodes) are ligated to both ends of a target nucleic acid. For example, 35 different tags (e.g., barcodes) ligated to both ends of a target molecule create 35 x 35 permutations, which is equivalent to 1225 permutations for 35 tags. This number of tags is sufficient to ensure that different molecules with the same start and stop points have a high probability (e.g., at least 94%, 99.5%, 99.99%, 99.999%) of receiving different combinations of tags. Other barcode combinations include any number between 10 and 500, such as about 15 x 15, about 35 x 35, about 75 x 75, about 100 x 100, about 250 x 250, and about 500 x 500.
[0234] In some cases, the unique tag may be an oligonucleotide of predetermined or random or semi-random sequence. In other cases, multiple barcodes may be used, so that the barcodes are not necessarily unique to each other among the multiple molecular barcodes. In this example, the barcode may be ligated to each molecule, so that the combination of barcode and sequence can be ligated to create a unique sequence that can be tracked individually. As described herein, the detection of a non-unique barcode in combination with the sequence data of the beginning (start) and end (stop) portions of the sequence read can allow for the assignment of a unique identity to a particular molecule. The length of each sequence read or its number of base pairs can also be used to assign a unique identity to such a molecule. As described herein, a fragment from a single strand of nucleic acid that has been assigned a unique identity can thereby allow for the subsequent identification of the fragment from the parent strand.
[0235] In some embodiments, two or more populations, samples, sub-samples, or fractions are differentially tagged, for example, by dividing the sub-samples and / or differentially degrading the sub-samples using one or more methylation-sensitive nucleases. Tags can be used to label individual DNA populations to correlate the tag (or tags) with a particular population or fraction. In some embodiments, a single tag can be used to label a particular population or fraction. In some embodiments, multiple different tags can be used to label a particular population or fraction. In embodiments using multiple different tags to label specific fractions, the set of tags used to label one fraction can be easily distinguished from the set of tags used to label other fractions. In some embodiments, the tag may have additional functionality, for example, the tag may be used to index the source of the sample, or may be used as a unique molecular identifier (which may be used to improve the quality of sequencing data by distinguishing sequencing errors from mutations, e.g., as described in Kinde et al., Proc Nat'l Acad Sci USA 108: 9530-9535 (2011), Kou et al., PLoS ONE,11: e0146638 (2016)), or as a non-unique molecular identifier, e.g., as described in U.S. Pat. No. 9,598,731. Similarly, in some embodiments, the tag may have additional functionality, for example, the tag may be used to index the source of the sample, or may be used as a non-unique molecular identifier (which may be used to improve the quality of sequencing data by distinguishing sequencing errors from mutations).
[0236] In some embodiments, tagging the fractions includes tagging the molecules in each fraction with a fraction tag. After the fractions are recombined (e.g., to reduce the number of required sequencing runs and avoid unnecessary costs) and the molecules are sequenced, the fraction tag identifies the source fraction. In another embodiment, different fractions are tagged with different molecular tag sets, including, for example, barcode pairs. In this way, each molecular barcode is useful for indicating the source fraction and distinguishing molecules within the fraction. For example, a first set of 35 barcodes can be used to tag molecules in a first fraction, and a second set of 35 barcodes can be used to tag molecules in a second fraction.
[0237] In some embodiments, after tagging, molecules can be pooled for sequencing in one run.In some embodiments, sample tag is added to molecule, for example, after adding other tags and in the step after pooling.Sample tag can facilitate the pooling of material generated from multiple samples for sequencing in one run.
[0238] In some embodiments, fraction tag can be associated with sample and fraction.As a simple example, the first tag can represent the first fraction of the first sample, the second tag can represent the second fraction of the first sample, the third tag can represent the first fraction of the second sample, and the fourth tag can represent the second fraction of the second sample.
[0239] A tag may be attached to a molecule based on one or more characteristics, but the final tagged molecule in the library may no longer have that characteristic.For example, single-stranded DNA molecules may be distributed and / or tagged, but the final tagged molecule in the library will likely be double-stranded.Similarly, DNA may be distributed based on different methylation levels, but in the final library, the tagged molecules derived from these molecules will likely be unmethylated (e.g., after the conversion step and / or further amplification step disclosed herein).Therefore, the tag attached to the molecule in the library typically represents the characteristics of the "parent molecule" from which the final tagged molecule is derived, and is not necessarily a characteristic of the tagged molecule itself.
[0240] For example, use barcode 1, 2, 3, 4 etc. to tag and label the molecules in the first fraction; use barcode A, B, C, D etc. to tag and label the molecules in the second fraction; and use barcode a, b, c, d etc. to tag and label the molecules in the third fraction.Differentially tagged fractions can be pooled before sequencing.Differentially tagged fractions can be sequenced separately, or can be sequenced together simultaneously, for example, in the same flow cell of Illumina sequencer.
[0241] In some embodiments, the adapter used herein comprises a barcode and a restriction enzyme cleavage site. In some embodiments, the DNA analyzed in the disclosed method comprises an insert and an adapter comprising a barcode, wherein at least one of the adapters further comprises a restriction enzyme cleavage site between the barcode and a portion of the adapter, and the barcode is located between the insert and the restriction enzyme cleavage site. In such embodiments, cleavage of the adapter at the restriction enzyme cleavage site by a restriction enzyme does not remove the barcode from the DNA molecule to which the adapter is ligated. In some embodiments, the barcode does not contain a cytosine in a non-CpG context.
[0242] After sequencing, analysis of the reads can be performed at the fraction level as well as at the pooled DNA level. Tags are used to separate the reads from different fractions. Analysis can include in silico analysis to determine genetic and epigenetic variations (one or more of methylation, chromatin structure, etc.) using sequence information, genomic coordinate length, coverage, and / or copy number. D. Subjecting the DNA or a subsample thereof to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA
[0243] In some embodiments, the methods disclosed herein include subjecting DNA or a subsample thereof to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA, wherein the first nucleobase is a modified or unmodified nucleobase, the second nucleobase is a modified or unmodified nucleobase different from the first nucleobase, and the first nucleobase and the second nucleobase have the same base-pairing specificity. In some embodiments, the procedure chemically converts the first or second nucleobase so that the base-pairing specificity of the converted nucleobase changes. In some embodiments, the DNA is subjected to the procedure that affects the first nucleobase in the DNA differently from the second nucleobase in the DNA after methylation-preserving amplification and before sequencing. In certain embodiments, the DNA is subjected to the procedure before enriching the DNA for one or more epigenetic target regions and / or sequence-variable target regions, and / or before or after contacting the DNA with a methylation-sensitive nuclease.
[0244] In some embodiments, when the first nucleobase is modified or unmodified adenine, the second nucleobase is modified or unmodified adenine; when the first nucleobase is modified or unmodified cytosine, the second nucleobase is modified or unmodified cytosine; when the first nucleobase is modified or unmodified guanine, the second nucleobase is modified or unmodified guanine; when the first nucleobase is modified or unmodified thymine, the second nucleobase is modified or unmodified thymine (modified and unmodified uracil are encompassed within modified thymine for the purposes of this step).
[0245] In some embodiments, when the first nucleobase is modified or unmodified cytosine, the second nucleobase is modified or unmodified cytosine.For example, the first nucleobase can comprise unmodified cytosine (C), and the second nucleobase can comprise one or more of 5-methylcytosine (mC) and 5-hydroxymethylcytosine (hmC).Alternatively, the second nucleobase can comprise C, and the first nucleobase can comprise one or more of mC and hmC.Other combinations are also possible, such as when one of the first and second nucleobases comprises mC, and the other comprises hmC.
[0246] In some embodiments, the procedure that affects the first nucleic acid base in DNA differently from the second nucleic acid base in the DNA of the first sub-sample comprises bisulfite conversion.Bisulfite treatment converts unmodified cytosine and certain modified cytosine nucleotides (for example, 5-formylcytosine (fC) or 5-carboxylcytosine (caC)) to uracil, while other modified cytosines (for example, 5-methylcytosine, 5-hydroxymethylcytosine) are not converted.Thus, when using bisulfite conversion, the first nucleic acid base comprises one or more of unmodified cytosine, 5-formylcytosine, 5-carboxylcytosine, or other cytosine types that are affected by bisulfite, and the second nucleic acid base can comprise one or more of mC and hmC, for example, mC and optionally hmC.Sequencing of bisulfite-treated DNA identifies the position that is read as cytosine as mC or hmC position. On the other hand, positions that read as T are identified as T or bisulfite-sensitive forms of C, such as unmodified cytosine, 5-formylcytosine, or 5-carboxylcytosine. Thus, performing bisulfite conversion on a DNA sample as described herein facilitates identifying positions containing mC or hmC using sequence reads obtained from an exemplary sample. For an exemplary description of bisulfite conversion, see, e.g., Moss et al., Nat Commun. 2018; 9: 5068.
[0247] In some embodiments, the procedure that affects a first nucleobase in DNA differently from a second nucleobase in the DNA of a first subsample comprises oxidative bisulfite (Ox-BS) conversion. This procedure first converts hmC to fC, which is bisulfite-sensitive, followed by bisulfite conversion. Thus, when using oxidative bisulfite conversion, the first nucleobase comprises one or more of unmodified cytosine, fC, caC, hmC, or other cytosine forms that are affected by bisulfite, and the second nucleobase comprises mC. Sequencing of the Ox-BS-converted DNA identifies positions that are read as cytosine as mC positions. Meanwhile, positions that are read as T are identified as T, hmC, or bisulfite-sensitive forms of C, such as unmodified cytosine, fC, or hmC. Thus, performing Ox-BS conversion on a DNA sample as described herein facilitates identifying mC-containing positions using sequence reads obtained from the sample. For an exemplary description of oxidative bisulfite conversion, see, e.g., Booth et al., Science 2012; 336: 934-937.
[0248] In some embodiments, the procedure for affecting a first nucleobase in DNA differently from a second nucleobase in the DNA of a first sub-sample comprises Tet-assisted bisulfite (TAB) conversion. In TAB conversion, hmC is protected from conversion and mC is oxidized prior to bisulfite treatment, thereby converting the position originally occupied by mC to U and leaving the position originally occupied by hmC as a protected form of cytosine. For example, as described in Yu et al., Cell 2012; 149: 1368-80, after protecting hmC (forming 5-glucosylhydroxymethylcytosine (ghmC)) using β-glucosyltransferase, a TET protein such as mTet1 can be used to convert mC to caC, and then bisulfite treatment can be used to convert C and caC to U, while leaving ghmC unaffected. Alternatively, after protecting hmC using a carbamoyltransferase enzyme such as 5-hydroxymethylcytosine carbamoyltransferase (by converting hmC to 5-carbamoyloxymethylcytosine (5cmC)) as described in Yang et al., Bio-protocol, 2023; 12(17): e4496, a TET protein such as mTet1 can be used to convert mC to caC, followed by bisulfite treatment to convert C and caC to U, while leaving 5cmC unaffected. Thus, when using TAB conversion, the first nucleobase comprises one or more of unmodified cytosine, fC, caC, mC, or other types of cytosine affected by bisulfite, and the second nucleobase comprises hmC. Sequencing of the TAB-converted DNA identifies positions read as cytosine as hmC positions. On the other hand, positions read as T are identified as T, mC, or a bisulfite-sensitive form of C, such as unmodified cytosine, fC, or caC. Thus, performing TAB conversion on a DNA sample as described herein facilitates identifying positions containing hmC using sequence reads obtained from the sample.
[0249] In some embodiments, the procedure for affecting a first nucleobase in DNA differently from a second nucleobase in the DNA of the first subsample includes Tet-assisted conversion with a substituted borane reducing agent, optionally 2-picoline borane, borane pyridine, tert-butylamine borane, or ammonia borane. Tet-assisted pic-borane conversion with a substituted borane reducing agent converts mC and hmC to caC using a TET protein without affecting unmodified C. Then, caC and, if present, fC are converted to dihydrouracil (DHU) by treatment with 2-picoline borane (pic-borane) or another substituted borane reducing agent, such as borane pyridine, tert-butylamine borane, or ammonia borane, similarly without affecting unmodified C. See, for example, Liu et al., Nature Biotechnology 2019; 37:424-429 (e.g., Supplementary Figure 1 and Supplementary Note 7). DHU is read as T in sequencing. Thus, when using this type of conversion, the first nucleobase contains one or more of mC, fC, caC, or hmC, and the second nucleobase contains an unmodified cytosine. Sequencing the converted DNA identifies positions that read as cytosine as unmodified C positions, while positions that read as T are identified as T, mC, fC, caC, or hmC. Thus, performing TAP conversion on a DNA sample as described herein facilitates identifying positions containing unmodified C using sequence reads obtained from the sample. This procedure encompasses Tet-assisted pyridine borane sequencing (TAPS), as described in further detail in Liu et al. 2019, supra.
[0250] Alternatively, protection of hmC (e.g., using βGT or 5-hydroxymethylcytosine carbamoyltransferase) can be combined with Tet-assisted conversion with a substituted borane reducing agent. hmC can be protected as described above through glucosylation using βGT to form ghmC, or through carbamoylation using 5-hydroxymethylcytosine carbamoyltransferase to form 5cmC. Treatment with a TET protein, e.g., mTet1, then converts mC to caC but not C, ghmC, or 5cmC. caC is then converted to DHU by treatment with pic-borane or another substituted borane reducing agent, e.g., borane pyridine, tert-butylamine borane, or ammonia borane, similarly without affecting ghmC, 5cmC, or unmodified C. Thus, when using Tet-assisted conversion with a substituted borane reducing agent, the first nucleobase comprises mC, and the second nucleobase comprises unmodified cytosine or hmC, e.g., unmodified cytosine, and optionally one or more of hmC, fC, and / or caC. Sequencing of the converted DNA identifies positions that read as cytosine as hmC or unmodified C positions. Meanwhile, positions that read as T are identified as T, fC, caC, or mC. Thus, performing TAPSβ conversion on a DNA sample or the like as described herein facilitates distinguishing positions containing either unmodified C or hmC from positions containing mC using sequence reads from the sample. For an exemplary description of this type of conversion, see, e.g., Liu et al., Nature Biotechnology 2019; 37:424-429. 5-hydroxymethylcytosine carbamoyltransferase is described in Yang et al., Bio-protocol, 2023; 12(17): e4496.
[0251] In some embodiments, the procedure for affecting a first nucleobase in DNA differently from a second nucleobase in the DNA of the first sub-sample comprises chemical-assisted conversion with a substituted borane reducing agent, optionally 2-picoline borane, borane pyridine, tert-butylamine borane, or ammonia borane. In the chemical-assisted conversion with a substituted borane reducing agent, an oxidizing agent, such as potassium perruthenate (KRuO4) (also suitable for use in the ox-BS conversion), is used to specifically oxidize hmC to fC. Treatment with pic-borane or another substituted borane reducing agent, such as borane pyridine, tert-butylamine borane, or ammonia borane, converts fC and caC to DHU, but does not affect mC or unmodified C. Thus, when using this type of conversion, the first nucleobase comprises one or more of hmC, fC, and caC, and the second nucleobase comprises one or more of unmodified cytosine or mC, for example, unmodified cytosine and, optionally, mC. Sequencing the converted DNA identifies positions that read as cytosine as mC or unmodified C positions, while positions that read as T are identified as T, fC, caC, or hmC. Thus, performing this type of conversion on a DNA sample or the like as described herein facilitates distinguishing positions containing either unmodified C or mC from positions containing hmC using sequence reads from the sample. For an exemplary description of this type of conversion, see, e.g., Liu et al., Nature Biotechnology 2019; 37:424-429.
[0252] In some embodiments, the procedure for affecting a first nucleobase in DNA differently from a second nucleobase in the DNA of a first subsample comprises APOBEC-linked epigenetic (ACE) conversion. ACE conversion uses an AID / APOBEC family DNA deaminase enzyme, such as APOBEC3A (A3A), to deaminate unmodified cytosine and mC without deaminating hmC, fC, or caC. Thus, when using ACE conversion, the first nucleobase comprises unmodified C and / or mC (e.g., unmodified C and optionally mC), and the second nucleobase comprises hmC. Sequencing of the ACE-converted DNA identifies positions that read as cytosine as hmC, fC, or caC positions. Meanwhile, positions that read as T are identified as T, unmodified C, or mC. Thus, performing ACE conversion on a DNA sample as described herein facilitates using sequence reads from the sample to distinguish positions containing hmC from positions containing mC or unmodified C. For an exemplary description of ACE conversion, see, e.g., Schutsky et al., Nature Biotechnology 2018; 36: 1083-1090.
[0253] In some embodiments, the procedure affecting a first nucleobase in the DNA differently from a second nucleobase in the DNA of the first sub-sample comprises enzymatic conversion of the first nucleobase, e.g., enzymatic conversion in EM-Seq. See, e.g., Vaisvila R, et al. (2019) EM-seq: Detection of DNA methylation at single base resolution from picograms of DNA. bioRxiv; DOI: 10.1101 / 2019.12.20.884692, available at www.biorxiv.org / content / 10.1101 / 2019.12.20.884692v1. For example, TET2 and T4-βGT or 5-hydroxymethylcytosine carbamoyltransferase (described in Yang et al., Bio-protocol, 2023; 12(17): e4496) can be used to convert 5mC and 5hmC into substrates that cannot be deaminated by a deaminase (e.g., APOBEC3A), which can then be used to deaminate unmodified cytosines, converting them to uracil.
[0254] In some embodiments, the procedure of affecting a first nucleobase in DNA differently from a second nucleobase in the DNA comprises enzymatic conversion of the first nucleobase using a double-stranded DNA deaminase sensitive to non-specific modifications, e.g., enzymatic conversion in SEM-seq. See, e.g., Vaisvila et al. (2023) "Discovery of novel DNA cytosine deaminase activities enables a nondestructive single-enzyme methylation sequencing method for base resolution high-coverage methylome mapping of cell-free and ultra-low input DNA." bioRxiv; DOI: 10.1101 / 2023.06.29.547047, available at https: / / www.biorxiv.org / content / 10.1101 / 2023.06.29.547047v1. SEM-Seq uses a nonspecific modification-sensitive double-stranded DNA deaminase (MsddA) in a non-destructive, single-enzyme 5-methylcytosine sequencing (SEM-seq) method to deaminate unmodified cytosines. Therefore, SEM-seq does not require the denaturing step used in protocols based on TET2 and T4-βGT or 5-hydroxymethylcytosine carbamoyltransferase protection, as well as APOEC3A. Additionally, MsddA does not deaminate 5-formylated cytosine (5fC) or 5-carboxylated cytosine (5caC). In SEM-seq, unmodified cytosines in DNA are deaminated to uracil and read as "T" during sequencing. Modified cytosines (e.g., 5mC) are not converted and are read as "C" during sequencing. Cytosines that are read as thymine are identified in DNA as unmodified (e.g., unmethylated) cytosines or as thymine. Thus, performing SEM-seq conversion makes it easy to identify positions containing 5mC using the resulting sequence reads.In some embodiments, the step of affecting a first nucleobase in the DNA differently from a second nucleobase in the DNA comprises enzymatic conversion of the first nucleobase using MsddA.
[0255] In some embodiments, the procedure for affecting the first nucleobase in DNA so that it differs from the second nucleobase in the DNA of the first subsample comprises separating DNA that originally contains the first nucleobase from DNA that does not originally contain the first nucleobase. In some such embodiments, the first nucleobase is hmC. DNA that originally contains the first nucleobase can be separated from other DNA using a labeling procedure that includes a biotinylation site that originally contains the first nucleobase. In some embodiments, the first nucleobase is first derivatized with an azide-containing moiety, for example, a glucosyl-azide-containing moiety. The azide-containing moiety can then serve as a reagent for binding biotin, for example, through Huisgen cycloaddition chemistry. Next, DNA that originally contains the now biotinylated first nucleobase can be separated from DNA that does not originally contain the first nucleobase using a biotin-binding agent, such as avidin, neutravidin (deglycosylated avidin with an isoelectric point of about 6.3), or streptavidin. An example of a procedure for separating DNA that originally contains the first nucleobase from DNA that does not originally contain the first nucleobase is hmC-sealing, which involves labeling hmC to form β-6-azido-glucosyl-5-hydroxymethylcytosine, then attaching a biotin moiety via Huisgen cycloaddition, and then using a biotin-binding agent to separate the biotinylated DNA from other DNA. For an exemplary description of hmC-sealing, see, for example, Han et al., Mol. Cell 2016; 63: 711-719. This approach is useful for identifying fragments containing one or more hmC nucleobases.
[0256] In some embodiments, after such separation, the method further comprises the step of differentially tagging the DNA that originally comprises the first nucleobase and the DNA that does not originally comprise the first nucleobase.The method can further comprise the step of pooling the DNA that originally comprises the first nucleobase and the DNA that does not originally comprise the first nucleobase after differential tagging.The DNA that originally comprises the first nucleobase and the DNA that does not originally comprise the first nucleobase can then be used in downstream analysis.For example, the pooled DNA that originally comprises the first nucleobase and the DNA that does not originally comprise the first nucleobase can be sequenced in the same sequencing cell (for example, after being subjected to further processing such as the processing described herein), while retaining the ability to use differential tag to determine whether a given read originates from the molecule of DNA that originally comprises the first nucleobase or from the DNA that does not originally comprise the first nucleobase.
[0257] In some embodiments, the first nucleobase is a modified or unmodified adenine and the second nucleobase is a modified or unmodified adenine. In some embodiments, the modified adenine is N 6 In some embodiments, the modified adenine is N-methyladenine (mA). 6 -Methyladenine (mA), N 6 -hydroxymethyladenine (hmA), or N 6 -formyl adenine (fA).
[0258] Techniques involving partitioning based on methylation status or methylated DNA immunoprecipitation (MeDIP) can be used to separate DNA containing modified bases such as mC, mA, caC (e.g., which can be generated by oxidation of mC or hmC by Tet2, e.g., using a deaminase such as APOBEC3A, prior to enzymatic conversion of the unmodified C to U), or dihydrouracil, from other DNA. See, e.g., Kumar et al., Frontiers Genet. 2018;9:640; Greer et al., Cell 2015;161:868-878. An antibody specific for mA is described in Sun et al., Bioessays 2015;37:1155-62. Antibodies against various modified nucleobases, such as mC, caC, and forms of thymine / uracil, including halogenated forms such as dihydrouracil or 5-bromouracil, are commercially available. Various modified bases can also be detected based on changes in their base pairing specificity. For example, hypoxanthine is a modified form of adenine that can result from deamination and is read as G in sequencing. See, e.g., U.S. Patent No. 8,486,630; Brown, Genomes, 2002; nd Ed., John Wiley & Sons, Inc., New York, NY, 2002, chapter 14, "Mutation, Repair, and Recombination."
[0259] In some embodiments, the conversion procedure is an enzymatic conversion procedure that converts the base-pairing specificity of a modified nucleoside (e.g., a DM-seq conversion that involves adding a protecting group (e.g., a carboxymethyl group) to an unmodified cytosine and deaminating 5mC, e.g., using an APOBEC enzyme) or an enzymatic conversion procedure that converts the base-pairing specificity of an unmodified nucleoside (e.g., SEM-seq).
[0260] In some cases, the conversion procedure used in the method of the present disclosure is a procedure that changes the base pairing specificity of modified nucleosides (e.g., methylated cytosine), but does not change the base pairing specificity of corresponding unmodified nucleosides (e.g., cytosine), or does not change the base pairing specificity of any unmodified nucleosides (e.g., cytosine, adenosine, guanosine, and thymidine (or uracil)).The advantages of a method that does not change the base pairing specificity of unmodified nucleosides include reduced loss of sequence complexity, higher sequencing efficiency, and reduced alignment loss.In addition, methods such as DM-seq may in some cases be preferable to methods such as bisulfite sequencing and EM-seq because they are less destructive (particularly important for low-yield samples such as cfDNA), do not require denaturation, and non-conversion errors are theoretically more likely to be random.In methods that require denaturation for conversion, failure to denature DNA molecules results in the non-conversion of all bases in the DNA molecule. Because biological changes in methylation are preferentially coordinated to the localized region of interest, these non-random (localized) conversions may appear as false negatives (unmethylated regions). Random non-conversion methods can maximize the effect on low base percentages within a region, thus, by setting a threshold for the percentage of bases within a methylated / unmethylated region, the specificity of methylation change detection can be maximized (false positives can be reduced). Therefore, in some cases, conversion procedures that do not involve denaturation are preferred.
[0261] In other cases, the conversion procedure used in the disclosed methods is one that alters the base pairing specificity of an unmodified nucleoside (e.g., cytosine) but does not alter the base pairing specificity of the corresponding modified nucleoside (e.g., methylated cytosine).
[0262] Those skilled in the art can select a suitable method according to their needs, including which nucleoside modifications are to be detected and / or identified.
[0263] In some embodiments, the conversion procedure converts modified nucleosides. In some embodiments, the conversion procedure for converting modified nucleosides includes enzymatic conversion, such as DM-seq, as described in WO2023 / 288222A1. In DM-seq, unmodified cytosines in DNA are enzymatically protected from a subsequent deamination step, which converts 5mC in 5mCpG to T. Enzymatically protected unmodified (e.g., unmethylated) cytosines are not converted and are read as "C" during sequencing. Cytosines (in CpG contexts) that are read as thymine are identified as methylated cytosines in DNA.
[0264] Thus, when using this type of conversion, the first nucleobase contains an unmodified (e.g., unmethylated) cytosine, and the second nucleobase contains a modified (e.g., methylated) cytosine. Sequencing of the converted DNA identifies positions that are read as cytosine as unmodified C positions. Meanwhile, positions that are read as T are identified as T or 5mC. In this way, performing DM-seq conversion makes it easy to use the resulting sequence reads to identify positions containing 5mC.
[0265] Exemplary cytosine deaminase for use herein includes APOBEC enzyme, for example, APOBEC3A.Generally, AID / APOBEC family DNA deaminase enzyme, for example, APOBEC3A (A3A), is used to deaminate (unprotected) unmodified cytosine and 5mC.For exemplary explanation of APOBEC conversion, see, for example, Schutsky et al., Nature Biotechnology 2018; 36: 1083-1090.
[0266] Enzymatic protection of unmodified cytosines in DNA involves adding a protecting group to the unmodified cytosine. Such protecting groups can include alkyl groups, alkyne groups, carboxyl groups, carboxyalkyl groups, amino groups, hydroxymethyl groups, glucosyl groups, glucosylhydroxymethyl groups, isopropyl groups, or dyes. For example, DNA can be treated with a methyltransferase, such as a CpG-specific methyltransferase, which adds a protecting group to the unmodified cytosine. The term methyltransferase is used broadly herein to refer to an enzyme that can transfer methyl or a substituted methyl (e.g., carboxymethyl) to a substrate (e.g., cytosine in a nucleic acid). In some embodiments, DNA is contacted with a CpG-specific DNA methyltransferase (MTase), such as a CpG-specific carboxymethyltransferase (CxMTase), and a substituted methyl donor, such as a carboxymethyl donor (e.g., carboxymethyl-S-adenosyl-L-methionine). See, for example, WO2021 / 236778A2. In certain embodiments, CxMTase can facilitate the addition of a protective carboxymethyl group to unmethylated cytosine. In some embodiments, the unmethylated cytosine is unmodified cytosine. The carboxymethyl group can prevent deamination of cytosine during a deamination step (e.g., a deamination step using an APOBEC enzyme such as A3A). Substituted methyl or carboxymethyl donors useful in the disclosed methods include, but are not limited to, S-adenosyl-L-methionine (SAM) analogs, and optionally, the SAM analog is carboxy-S-adenosyl-L-methionine (CxSAM). SAM analogs are described, for example, in WO2022 / 197593A1. The MTase may be, for example, CpG methyltransferase from Spiroplasma sp. strain MQ1 (M.SssI), DNA-methyltransferase 1 (DNMT1), DNA-methyltransferase 3 alpha (DNMT3A), DNA-methyltransferase 3 beta (DNMT3B), or DNA adenine methyltransferase (Dam).The CxMTase may be a CpG methyltransferase from Mycoplasma penetrans (M.MpeI). In certain embodiments, the methyltransferase enzyme is a variant of M.MpeI having SEQ ID NO:1 or SEQ ID NO:2, or a sequence at least 90%, at least 92%, at least 94%, at least 96%, at least 97%, at least 98%, or at least 99% identical thereto, and optionally, the amino acid corresponding to position 374 is R or K.
[0267] In one embodiment, the methyltransferase enzyme is a variant of M.MpeI having an N374R or N374K substitution. The methyltransferase of SEQ ID NO:1 or SEQ ID NO:2 can further comprise one or more amino acid substitutions selected from: a) substitution of one or both of residues T300 and E305 with S, A, G, Q, D, or N; b) substitution of one or more of residues A323, N306, and Y299 with a positively charged amino acid selected from K, R, or H; and / or c) substitution of S323 with A, G, K, R, or H, which may enhance the activity of the enzyme.
[0268] Optionally, the conversion procedure further includes enzymatic protection of 5hmC in DNA, such as by glucosylation of 5hmC (e.g., using βGT) or by carbamoylation of 5hmC (e.g., using 5-hydroxymethylcytosine carbamoyltransferase) before deamination of unprotected modified cytosines. In this method, 5hmC can be protected from conversion through glucosylation using, for example, β-glucosyltransferase (βGT) to form (5-glucosylhydroxymethylcytosine) 5ghmC, or through carbamoylation using 5-hydroxymethylcytosine carbamoyltransferase to form 5cmC. This is described, for example, in Yu et al., Cell 2012; 149: 1368-80, and Yang et al., Bio-protocol, 2023; 12(17): e4496. Glucosylation or carbamoylation of 5hmC can reduce or eliminate deamination of 5hmC by deaminases such as APOBEC3A. Treatment with MTase or CxMTase then adds a protecting group to unmodified (unmethylated) cytosines in DNA. 5mC (but not the protected unmodified cytosine, not 5ghmC or 5cmC) is then deaminated (in the case of 5mC, converted to T) by treatment with a deaminase, e.g., an APOBEC enzyme (e.g., APOBEC3A). Sequencing of the converted DNA identifies positions that read as cytosine as 5hmC or unmodified C positions. Meanwhile, positions that read as T are identified as T or 5mC. Thus, performing DM-seq conversion of 5hmC glycosylation on a sample as described herein facilitates using the resulting sequence reads to distinguish positions containing either unmodified C or 5hmC from positions containing 5mC.
[0269] Also provided herein are methods that use alternative base conversion schemes, for example, unmethylated cytosine can remain intact, while methylated and hydroxymethyl cytosine are converted to a base that is read as thymine (e.g., uracil, thymine, or dihydrouracil).
[0270] In some embodiments, methylating cytosines in at least one of the first or second complementary strands comprises contacting the cytosines with a methyltransferase, such as DNMT1 or DNMT5. In such embodiments, oxidizing 5-hydroxymethylated cytosines to 5-formylcytosines (e.g., by contacting the 5-hydroxymethylcytosines in the first and second strands with KRuO4) can be optional.
[0271] In some embodiments, converting the modified cytosine in at least one of the first or second strands to thymine or a base that is read as thymine comprises oxidizing the hydroxymethylcytosine, e.g., oxidizing the hydroxymethylcytosine to formylcytosine. In some embodiments, oxidizing the hydroxymethylcytosine to formylcytosine comprises contacting the hydroxymethylcytosine with a ruthenate, e.g., potassium ruthenate (KRuO).
[0272] In some embodiments, the modified cytosine is converted to thymine, uracil, or dihydrouracil. In any such embodiment, the amplification method may comprise a uracil- and / or dihydrouracil-resistant amplification method, such as PCR using a uracil- and / or dihydrouracil-resistant DNA polymerase.
[0273] In some embodiments, the method includes converting formylcytosine and / or methylcytosine to carboxyl cytosine as part of converting at least one modified cytosine in the first or second strand to thymine or a base that is read as thymine. For example, converting formylcytosine and / or methylcytosine to carboxyl cytosine may include contacting the formylcytosine and / or methylcytosine with a TET enzyme, such as TET1, TET2, or TET3. In some embodiments, the method includes reducing the carboxyl cytosine as part of converting at least one modified cytosine in the first or second strand to thymine or a base that is read as thymine, and / or the carboxyl cytosine is reduced to dihydrouracil. In some embodiments, reducing the carboxyl cytosine includes contacting the carboxyl cytosine with a borane reducing agent or a borohydride reducing agent.
[0274] In some embodiments, the borane or borohydride reducing agent comprises pyridine borane, 2-picoline borane, borane, tert-butylamine borane, ammonia borane, sodium borohydride, sodium cyanoborohydride (NaBHCN), lithium borohydride (LiBH), ethylenediamine borane, dimethylamine borane, sodium triacetoxyborohydride, morpholine borane, 4-methylmorpholine borane, trimethylamine borane, dicyclohexylamine borane, or a salt thereof. In other embodiments, the reducing agent comprises lithium aluminum hydride, sodium amalgam, amalgam, sulfur dioxide, dithionate, thiosulfate, iodide, hydrogen peroxide, hydrazine, diisobutylaluminum hydride, oxalic acid, carbon monoxide, cyanide, ascorbic acid, formic acid, dithiothreitol, beta-mercaptoethanol, or any combination thereof.
[0275] Various TET enzymes can be used in the disclosed methods, as desired. In some embodiments, one or more TET enzymes include TETv. TETv is described in U.S. Patent No. 10,260,088, and its sequence is SEQ ID NO: 1 therein (SEQ ID NO: 3 herein). In some embodiments, one or more TET enzymes include TETcd. TETcd is described in U.S. Patent No. 10,260,088, and its sequence is SEQ ID NO: 3 therein (SEQ ID NO: 4 herein). In some embodiments, one or more TET enzymes include TET1. In some embodiments, one or more TET enzymes include TET2. TET2 can be expressed and used as a fragment comprising residues 1129-1480 of TET2 linked to residues 1844-1936 of TET2 by a linker (SEQ ID NO: 5 herein), e.g., as described in U.S. Patent No. 10,961,525. In some embodiments, one or more TET enzymes include TET1 and TET2. In some embodiments, the one or more TET enzymes comprise a V1900 TET mutant, such as a V1900A, V1900C, V1900G, V1900I, or V1900P TET mutant. In some embodiments, the one or more TET enzymes comprise a V1900 TET2 mutant, such as a V1900A, V1900C, V1900G, V1900I, or V1900P TET2 mutant. Examples of the V1900A, V1900C, V1900G, V1900I, and V1900P TET2 mutants are provided as SEQ ID NOs: 6-10. In some embodiments, the V1900 TET mutant has at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity to SEQ ID NOs: 6, 7, 8, 9, or 10. Position 1900 of the wild-type TET2 sequence corresponds to position 438 in each of SEQ ID NOs: 5-10. Because 5-caC is not a substrate for enzymatic deamination by APOBEC enzymes, such as APOBEC3A, it may be beneficial to use TET enzymes that maximize the formation of 5-carboxylcytosine (5-caC) relative to less oxidized modified cytosines, particularly 5-formylcytosine.Thus, maximizing the formation of 5-caC reduces the risk of false calls in which a base is identified as unmethylated because it underwent deamination even if it was methylated (or hydroxymethylated) in the original sample. Thus, in some embodiments, the TET enzyme contains a mutation that increases the formation of 5-caC. Exemplary mutations are described above. A "mutation that increases the formation of 5-caC" means that a TET enzyme with the mutation produces more 5-caC than a TET enzyme lacking the mutation, all else being equal. 5-caC production can be measured, for example, as described in Liu et al., Nat Chem Biol 13:181-187 (2017) (see the Online Methods section, TET reactions in vitro subsection, "driving" conditions). Any of the variants and / or mutants described in Liu et al. (2017) can be used in the disclosed methods, as desired. E. Enrich DNA; Enriched Portion; Enriched Set
[0276] In some embodiments, the methods herein include enriching (also known as "capturing") nucleic acid molecules comprising sequences present in the set of target regions for subsequent analysis. Such enrichment or capture can be performed on any sample or sub-sample described herein using any suitable approach known in the art. The enriching step can be performed on one or more sub-samples prepared during the methods disclosed herein. In some embodiments, DNA is enriched from at least a first sub-sample. In some embodiments, DNA is enriched from at least a first sub-sample or a second sub-sample, e.g., at least a first sub-sample and a second sub-sample. In some embodiments, the enriching step is performed after a methylation-preserving amplification step, before or after contacting the DNA with a methylation-sensitive nuclease and / or a methylation-dependent nuclease, after subjecting the DNA to a procedure that affects a first nucleobase in the DNA so that it differs from a second nucleobase in the DNA, or any combination thereof. If the first sub-sample is subjected to a separation step (e.g., a step to separate DNA that naturally contains the first nucleobase (e.g., hmC) from DNA that does not naturally contain the first nucleobase, e.g., hmC-seal), the enrichment step may be performed on any, any two, or all of the DNA that naturally contains the first nucleobase (e.g., hmC), the DNA that does not naturally contain the first nucleobase, and the DNA of the second sub-sample. In some embodiments, the sub-samples are differentially tagged (e.g., as described herein) and then pooled before undergoing enrichment.
[0277] In some embodiments, enriching comprises contacting the DNA with probes specific for such target regions. In some embodiments, the probes comprise an oligonucleotide and a capture moiety, such as biotin or one or more of the other examples described below. The probes may have sequences, such as genes, selected to span a panel of regions.
[0278] Methods involving DNA enrichment using probes containing a capture moiety, such as target-specific probes labeled with biotin, can also include a second moiety or binding partner that binds to the capture moiety, such as streptavidin. In some embodiments, the enrichment moiety and binding partner can have higher or lower capture yields for different sets of probes, such as those used to enrich for (capture of) a set of sequence-variable target regions and a set of epigenetic target regions, respectively, as discussed elsewhere herein. Methods involving capture moieties are further described, for example, in U.S. Patent No. 9,850,523, issued December 26, 2017, which is incorporated herein by reference.
[0279] Capture moieties include, without limitation, biotin, avidin, streptavidin, nucleic acids containing specific nucleotide sequences, haptens recognized by antibodies, and magnetically attractable particles. In some embodiments, the capture moiety bound to the analyte is captured by its binding partner bound to an isolatable moiety, such as a magnetically attractable particle or a large particle that can be sedimented by centrifugation. The capture moiety can be any type of molecule that allows affinity separation of nucleic acids bearing the capture moiety from nucleic acids lacking the capture moiety. Exemplary capture moieties are biotin, which allows affinity separation by binding to streptavidin that is or can be linked to a solid phase, or oligonucleotides, which allow affinity separation by binding to complementary oligonucleotides that are or can be linked to a solid phase.
[0280] In some embodiments, non-specifically bound DNA that does not contain the target region is washed away from the enriched DNA. In some embodiments, the DNA is then dissociated from the probe and eluted from the solid support using a buffer containing a salt wash or another DNA denaturing agent. In some embodiments, the probe is also eluted from the solid support, for example, by disrupting the biotin-streptavidin interaction. In some embodiments, the enriched DNA is amplified after elution from the solid support. In some such embodiments, DNA containing an adapter is amplified using PCR primers that anneal to the adapter. In some embodiments, the enriched DNA is amplified while bound to the solid support. In some such embodiments, amplification involves the use of a PCR primer that anneals to a sequence within the adapter and a PCR primer that anneals to a sequence within the probe that anneals to the target region of the DNA.
[0281] In some embodiments, the target region is enriched from an aliquot, portion, or sub-sample of a sample (e.g., a sample that has undergone adaptor binding and amplification), while the DNA partitioning step may be performed on a separate aliquot, portion, or sub-sample of the sample.
[0282] The step of enriching or capturing DNA containing a target region may include contacting the DNA with a first or second set of target-specific probes. Such target-specific probes may have any of the features described herein for target-specific probe sets, including, but not limited to, the embodiments described herein and the probe-related sections herein. The enrichment step may be performed on a DNA sample or one or more subsamples prepared during the methods disclosed herein. In some embodiments, DNA is enriched from a first subsample or a second subsample after methylation-preserving amplification or after the DNA or its subsample is subjected to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA. In some embodiments, the subsamples are differentially tagged (e.g., as described herein) and then pooled before undergoing enrichment. Exemplary methods for enriching DNA containing epigenetic target regions and / or sequence-variable target regions can be found, for example, in WO2020 / 160414, which is incorporated herein by reference.
[0283] The enrichment step or steps may be carried out using conditions suitable for specific nucleic acid hybridization, which generally depend in part on characteristics of the probe, such as length, base composition, etc. Those skilled in the art will be familiar with appropriate conditions given their general knowledge in the art regarding nucleic acid hybridization.
[0284] In some embodiments, the methods described herein include enriching cfDNA obtained from a subject for a set of multiple target regions. The target regions may contain differences depending on whether they originate from a tumor, a healthy cell, or a specific cell type. For example, target regions containing epigenetic target regions may exhibit differences in methylation levels and / or fragmentation patterns depending on whether they originate from a tumor or a healthy cell. Similarly, target regions containing sequence-variable target regions may exhibit sequence differences depending on whether they originate from a tumor or a healthy cell. The enriching step generates an enriched set of cfDNA molecules. In some embodiments, cfDNA molecules corresponding to the set of sequence-variable target regions are enriched in the enriched set of cfDNA molecules with a higher capture yield than cfDNA molecules corresponding to the set of epigenetic target regions. In some embodiments, the methods described herein include contacting cfDNA obtained from a subject with a target-specific probe set, wherein the target-specific probe set is configured to capture cfDNA corresponding to the set of sequence-variable target regions with a higher capture yield than cfDNA corresponding to the set of epigenetic target regions.
[0285] Because analyzing sequence-variable target regions with sufficient reliability or accuracy may require a higher sequencing depth than that required for analyzing epigenetic target regions, it may be beneficial to enrich the DNA corresponding to a set of sequence-variable target regions with a higher capture yield than the DNA corresponding to a set of epigenetic target regions.The amount of data required to determine fragmentation patterns (e.g., to test for perturbations in transcription start sites or CTCF binding sites) or fragment abundances (e.g., in hypermethylated and hypomethylated fractions) is generally less than the amount of data required to determine the presence or absence of sequence mutations associated with cancer.Capturing target region sets with different yields can facilitate sequencing target regions to different sequencing depths in the same sequencing run (e.g., using pooled mixtures and / or in the same sequencing cell).Copy number variations such as local amplifications are somatic mutations, but they can be detected by sequencing based on read frequency in a manner similar to the approach used to detect certain epigenetic changes, such as changes in methylation.
[0286] In some embodiments, the enriched DNA is amplified. In various embodiments, the method further comprises sequencing the enriched DNA to different sequencing depths, for example, for epigenetic and sequence variable target region sets, in accordance with the discussion herein. In some embodiments, an RNA probe is used. In some embodiments, a DNA probe is used. In some embodiments, a single-stranded probe is used. In some embodiments, a double-stranded probe is used. In some embodiments, a single-stranded RNA probe is used. In some embodiments, a double-stranded DNA probe is used.
[0287] In some embodiments, the enrichment step is performed simultaneously in the same vessel for the probes for the set of sequence variable target regions and the probes for the set of epigenetic target regions, e.g., the probes for the set of sequence variable target regions and the probes for the set of epigenetic target regions and the capture probe are in the same composition. This approach provides a relatively streamlined workflow.
[0288] Alternatively, the enrichment step is carried out using a sequence variable target region probe set in a first container and an epigenetic target region probe set in a second container, or the contacting step is carried out using a sequence variable target region probe set at a first time and in the first container and an epigenetic target region probe set at a second time before or after the first time.This approach allows for the preparation of separate first and second compositions containing enriched DNA corresponding to a sequence variable target region set and enriched DNA corresponding to an epigenetic target region set.The compositions can be processed separately if desired (for example, to fractionate based on methylation as described elsewhere herein), and recombined at appropriate ratios to provide material for further processing and analysis, such as sequencing.
[0289] In some embodiments, the DNA is amplified. In some embodiments, the amplification is performed before the capturing step. In some embodiments, the amplification is performed after the enrichment step.
[0290] In some embodiments, the adapters are included in the DNA. This can be done simultaneously with the amplification procedure, for example, by providing the adapters at the 5' portion of the primers as described above. Alternatively, the adapters can be added by other approaches, such as ligation.
[0291] In some embodiments, a tag that can be or include a barcode (e.g., a "sample index") is included in the enriched DNA, such as during a post-enrichment amplification step. In some embodiments, such tags are referred to as "auxiliary adapters." Tags can facilitate identification of the origin of nucleic acids. For example, barcodes can be used to identify the source from which DNA originates (e.g., a subject, a biological sample (e.g., samples collected at various time points), an enriched DNA sample (e.g., enriched DNA containing a set of epigenetic target regions or enriched DNA containing a set of sequence-variable target regions), a fraction, or the like) after pooling multiple samples for parallel sequencing. This can be done simultaneously with the amplification procedure, for example, by providing a barcode in the 5' portion of the primer as described above. In some embodiments, the adapter and tag / barcode are provided by the same primer or primer set. For example, the barcode can be located 3' of the adapter and 5' of the target-hybridizing portion of the primer. Alternatively, the barcode can be added by other approaches, such as ligation, optionally with the adapter in the same ligation substrate.
[0292] In some embodiments, a collection of target-specific probes is used in the methods described herein, including enriched DNA. In some embodiments, the collection of target-specific probes includes target binding probes specific for one or more sets of target regions. In some embodiments, the capture yield of the target binding probes specific for the set of sequence-variable target regions is higher (e.g., at least two-fold higher) than the capture yield of the target binding probes specific for the set of epigenetic target regions. In some embodiments, the collection of target-specific probes is configured to have a capture yield specific for the set of sequence-variable target regions that is higher (e.g., at least two-fold higher) than its capture yield specific for the set of epigenetic target regions. 1. Enriched Set
[0293] In some embodiments, provide enriched (also known as captured) DNA (for example, cfDNA) set.With regard to the disclosed method, enriched DNA set can be provided by, for example, carrying out enrichment step after methylation preservation amplification, or after subjecting DNA or its subsample to a procedure that affects the first nucleic acid base in DNA differently from the second nucleic acid base in DNA.Enriched set can comprise the DNA corresponding to sequence variable target region set, epigenetic target region set, or a combination thereof.
[0294] In some embodiments, the first set of target regions is enriched from DNA or a sub-sample thereof (e.g., a first sub-sample), and includes at least epigenetic target regions. The epigenetic target regions enriched from DNA or a sub-sample thereof may include hypermethylated variable target regions. In some embodiments, hypermethylated variable target regions are CpG-containing regions that are unmethylated or hypomethylated (e.g., below average methylation compared to bulk cfDNA) in cfDNA from healthy subjects. In some embodiments, hypermethylated variable target regions are regions that exhibit lower methylation in healthy cfDNA than in at least one other tissue type. Without wishing to be bound by any particular theory, cancer cells may shed more DNA into the bloodstream than healthy cells of the same tissue type. Therefore, the distribution of tissues of origin of cfDNA may change during carcinogenesis. Thus, increased levels of hypermethylated variable target regions in DNA or a sub-sample thereof may be indicative of the presence (or recurrence, depending on the subject's medical history) of cancer.
[0295] In some embodiments, the second set of target regions is enriched from DNA or a sub-sample thereof (e.g., a second sub-sample), and includes at least epigenetic target regions. The epigenetic target regions may include hypomethylated variable target regions. In some embodiments, hypomethylated variable target regions are CpG-containing regions that are methylated or have hypermethylation (e.g., above average methylation compared to bulk cfDNA) in cfDNA from healthy subjects. In some embodiments, hypomethylated variable target regions are regions that show higher methylation in healthy cfDNA than in at least one other tissue type. Without wishing to be bound by any particular theory, cancer cells may shed more DNA into the bloodstream than healthy cells of the same tissue type. Therefore, the distribution of tissues of origin of cfDNA may change during carcinogenesis. Thus, an increased level of hypomethylated variable target regions in DNA or a sub-sample thereof (e.g., a second sub-sample) may be indicative of the presence of cancer (or recurrence, depending on the subject's medical history).
[0296] In some embodiments, the amount of enriched sequence variable target region DNA is greater than the amount of enriched epigenetic target region DNA when normalized for differences in size (footprint size) of the regions of interest.
[0297] Alternatively, first and second enriched sets may be provided that contain DNA corresponding to the set of sequence variable target regions and DNA corresponding to the set of epigenetic target regions, respectively. The first and second enriched sets may be combined to provide a combined enriched set.
[0298] In some embodiments, where the enriched set comprising DNA corresponding to the set of sequence variable target regions and the set of epigenetic target regions comprises a combination of the enriched sets discussed above, the DNA corresponding to the set of sequence variable target regions is present at a higher concentration than the DNA corresponding to the set of epigenetic target regions, e.g., 1.1 to 1.2 fold higher, 1.2 to 1.4 fold higher, 1.4 to 1.6 fold higher, 1.6 to 1.8 fold higher, 1.8 to 2.0 fold higher, 2.0 to 2.2 fold higher, 2.2 to 2.4 fold higher, 2.4 to 2.6 fold higher, 2.6 to 2.8 fold higher, 2.8 to 3.0 fold higher, 3.0 to 3.5 fold higher, 3.5 to 4.0, 4.0 to 4.5 fold higher, 4.5 to 5.0 fold higher, 5.0 to 5.5 fold higher. High concentration, 5.5 to 6.0 times higher concentration, 6.0 to 6.5 times higher concentration, 6.5 to 7.0 times higher concentration, 7.0 to 7.5 times higher concentration, 7.5 to 8.0 times higher concentration, 8.0 to 8.5 times higher concentration, 8.5 to 9.0 times higher concentration, 9.0 to 9.5 times higher concentration, 9.5 to 10.0 times higher concentration, 10 to 11 times higher concentration, 11 to 12 times higher concentration, 12 to 13 times higher concentration, 13 to 14 times higher concentration, The concentration may be 14-15x higher, 15-16x higher, 16-17x higher, 17-18x higher, 18-19x higher, 19-20x higher, 20-30x higher, 30-40x higher, 40-50x higher, 50-60x higher, 60-70x higher, 70-80x higher, 80-90x higher, or 90-100x higher. The degree of concentration difference accounts for normalization with respect to the footprint size of the target region, as discussed in the definition section. F. Contacting the DNA with a methylation-sensitive or methylation-dependent nuclease
[0299] In some embodiments, DNA or a sub-sample thereof (e.g., a first, second, or third sub-sample prepared by partitioning the sample as described herein, such as by level of cytosine modification, such as methylation, e.g., 5-methylation) is contacted with a methylation-dependent nuclease or a methylation-sensitive nuclease. The contacting step can be performed, for example, using a sample partitioned into multiple sub-samples as disclosed herein. Unless otherwise indicated, when the partitioning step is performed based on cytosine modification, the first sub-sample is the sub-sample with a higher level of modification, the second sub-sample is the sub-sample with a lower level of modification, and if present, the third sub-sample has a level of modification intermediate between the first and second sub-samples.
[0300] In some embodiments, the method herein comprises contacting DNA with methylation-sensitive nuclease, thereby degrading the DNA comprising unmethylated sequence or sequence with low methylation level.In some such embodiments, the methylation-sensitive nuclease is a methylation-sensitive restriction enzyme (MSRE), thereby degrading the DNA comprising the unmethylated recognition site of MSRE.In this way, the methylation-sensitive nuclease can be used in the method herein, comprising one or more steps of depleting unmodified or unmethylated sequence, for example, those present in cfDNA from subject.
[0301] In some embodiments, the method herein comprises contacting DNA with methylation-dependent nuclease, thereby degrading the DNA comprising methylated sequence or sequence with high methylation level.In some such embodiments, the methylation-dependent nuclease is a methylation-dependent restriction enzyme (MDRE), thereby degrading the DNA comprising the methylation recognition site of MSRE.In this way, the methylation-dependent nuclease can be used in the method herein, comprising one or more steps of depleting modified or methylated sequence, for example, those present in cfDNA from a subject.
[0302] As discussed above, the partitioning procedure can result in incomplete sorting of DNA molecules within the sub-sample. A methylation-dependent nuclease or a methylation-sensitive nuclease can be selected to degrade non-specifically partitioned DNA. For example, the second sub-sample can be contacted with a methylation-dependent nuclease, such as a methylation-dependent restriction enzyme. This can degrade non-specifically partitioned DNA (e.g., methylated DNA) in the second sub-sample to generate a processed second sub-sample. Alternatively or additionally, the first sub-sample can be contacted with a methylation-sensitive endonuclease, such as a methylation-sensitive restriction enzyme, thereby degrading non-specifically partitioned DNA in the first sub-sample to generate a processed first sub-sample. Degradation of non-specifically distributed DNA in either or both of the first or second subsamples is proposed as an improvement to the performance of methods that rely on accurate distribution of DNA based on cytosine modifications, for example, to detect the presence of abnormally modified DNA in a sample, to determine the tissue of origin of the DNA, and / or to determine whether a subject has cancer. For example, such degradation can provide improved sensitivity and / or simplify downstream analysis. Generally, if the non-specifically distributed DNA is hypermethylated, for example, in a hypomethylated fraction, a methylation-dependent nuclease, for example, a methylation-dependent restriction enzyme, should be used. Conversely, if the non-specifically distributed DNA is hypomethylated, for example, in a hypermethylated fraction, a methylation-sensitive nuclease, for example, a methylation-sensitive restriction enzyme, should be used. Methylation-dependent nucleases, e.g., methylation-dependent restriction enzymes, preferentially cleave methylated DNA compared to unmethylated DNA, whereas methylation-sensitive nucleases, e.g., methylation-sensitive restriction enzymes, preferentially cleave unmethylated DNA compared to methylated DNA.
[0303] The step of contacting the sub-sample with the nuclease(s) can use one or more nucleases. In some embodiments, the sub-sample is contacted with multiple nucleases. The sub-samples may be contacted with the nuclease(s) sequentially or simultaneously. The simultaneous use of nucleases can be advantageous to avoid unnecessary sample manipulation when the nucleases are active under similar conditions (e.g., buffer composition). Contacting the second sub-sample with more than one methylation-dependent restriction enzyme can more completely degrade non-specifically distributed hypermethylated DNA. Similarly, contacting the first sub-sample with more than one methylation-sensitive restriction enzyme can more completely degrade non-specifically distributed hypomethylated and / or unmethylated DNA.
[0304] In some embodiments, the methylation-dependent nuclease comprises one or more of MspJI, LpnPI, FspEI, or McrBC. In some embodiments, at least two methylation-dependent nucleases are used. In some embodiments, at least three methylation-dependent nucleases are used. In some embodiments, the methylation-dependent nuclease comprises FspEI. In some embodiments, the methylation-dependent nuclease comprises FspEI and MspJI, for example, used sequentially.
[0305] In some embodiments, the methylation-sensitive nucleases include one or more of AatII, AccII, AciI, Aor13HI, Aor15HI, BspT104I, BssHII, BstUI, Cfr10I, ClaI, CpoI, Eco52I, HaeII, HapII, HhaI, Hin6I, HpaII, HpyCH4IV, MluI, MspI, NaeI, NotI, NruI, NsbI, PmaCI, Pspl406I, PvuI, SacII, SalI, SmaI, and SnaBI. In some embodiments, at least two methylation-sensitive nucleases are used. In some embodiments, at least three methylation-sensitive nucleases are used. In some embodiments, the methylation-sensitive nuclease includes BstUI and HpaII. In some embodiments, the two methylation-sensitive nucleases include HhaI and AccII. In some embodiments, the methylation-sensitive nucleases include BstUI, HpaII, and Hin6I.
[0306] In some embodiments, FspEI is used to digest nucleic acid molecules in at least one sub-sample (e.g., a hypomethylated fraction). In some embodiments, BstUI, HpaII, and Hin6I are used to digest nucleic acid molecules in at least one sub-sample (e.g., a hypermethylated fraction), and FspEI is used to digest nucleic acid molecules in at least one other sub-sample (e.g., a hypomethylated fraction). In embodiments including an intermediate methylation fraction, the nucleic acid molecules may be digested with a methylation-sensitive or methylation-dependent nuclease. In some embodiments, the nucleic acid molecules in the intermediate methylation fraction are digested with the same nuclease as the hypermethylated fraction. For example, the intermediate methylation fraction may be pooled with the hypermethylated fraction, and the pooled fractions may then be subjected to digestion. In some embodiments, the nucleic acid molecules in the intermediate methylation fraction are digested with the same nuclease as the hypomethylated fraction. For example, the intermediately methylated fraction may be pooled with the lowly methylated fraction, and the pooled fractions may then be subjected to digestion.
[0307] In some embodiments, after tagging or binding adapters to both ends of DNA, the sub-sample is contacted with the above-mentioned nuclease. The tag or adapter can be resistant to cleavage by the nuclease using any of the above-mentioned approaches. In this approach, since the cleavage product lacks tags or adapters at both ends, cleavage can prevent non-specifically distributed molecules from being included in the analysis.
[0308] Alternatively, the step of tagging or attaching an adapter can be performed after the cleavage by the nuclease described above.The cleaved molecules can then be identified in the sequence reads based on having the end (the attachment point to the tag or adapter) corresponding to the nuclease recognition site.Processing molecules in this manner can also enable the acquisition of information from the cleaved molecules, for example, the observation of somatic mutations.When tagging or attaching an adapter after contacting a sub-sample with a nuclease, and when low-molecular-weight DNA, such as cfDNA, is analyzed, it may be desirable to remove high-molecular-weight DNA (e.g., contaminating genomic DNA) from the sample before the contacting step.It may also be desirable to use a nuclease that can be heat-inactivated at a relatively low temperature (e.g., 65°C or lower, or 60°C or lower) to avoid DNA denaturation, as denaturation may interfere with the subsequent ligation step.
[0309] When the sample is divided into three sub-samples, including a third sub-sample containing intermediate methylated molecules, the third sub-sample is, in some embodiments, contacted with a methylation-sensitive nuclease. Such a step may have any of the features described elsewhere herein in connection with the contacting step and may be performed before or after the adapter tagging or binding step, as discussed above. In some embodiments, the first and third sub-samples are combined before contacting with the methylation-sensitive nuclease. Such a step may have any of the features described elsewhere herein in connection with the contacting step and may be performed before or after the adapter tagging or binding step, as discussed above. In some embodiments, the first and third sub-samples are differentially tagged before being combined.
[0310] Alternatively, if the sample is divided into three sub-samples, including a third sub-sample containing intermediate methylated molecules, the third sub-sample, in some embodiments, is contacted with a methylation-dependent nuclease. Such a step may have any of the features described elsewhere herein in connection with the contacting step and may be performed before or after the adapter tagging or binding step, as discussed above. In some embodiments, the second and third sub-samples are combined before contacting with the methylation-dependent nuclease. Such a step may have any of the features described elsewhere herein in connection with the contacting step and may be performed before or after the adapter tagging or binding step, as discussed above. In some embodiments, the second and third sub-samples are differentially tagged before being combined.
[0311] In some embodiments, DNA is purified after contacting with a nuclease, for example, using SPRI beads. Such purification may be performed after heat inactivation of the nuclease. Alternatively, purification can be omitted, and subsequent steps, such as amplification, can be performed on a subsample containing the heat-inactivated nuclease. In another embodiment, the contacting step can be performed in the presence of a purification reagent, such as SPRI beads, to minimize losses associated with, for example, tube transfer. After cleavage and heat inactivation, the SPRI beads can be reused for purification by adding a molecular crowding reagent (e.g., PEG) and salt. G. Dividing a sample into multiple subsamples
[0312] The methods disclosed herein include analyzing DNA in a sample. In such methods, different types of DNA (e.g., hypermethylated and hypomethylated DNA) can be physically divided into multiple sub-samples based on one or more characteristics of the DNA. This approach can be used to determine, for example, whether a particular sequence is hypermethylated or hypomethylated. Such a dividing step can be performed before or after a methylation-preserving amplification step and before sequencing. The dividing step can be performed using the DNA sample or one or more sub-samples thereof as described herein.
[0313] In some embodiments, the partitioning step comprises contacting the DNA with an agent that recognizes a modification associated with (e.g., in) the DNA. In some embodiments, the agent that recognizes the modification is an antibody. In some embodiments, the agent is immobilized on a solid support. In some embodiments, the partitioning step comprises immunoprecipitation, e.g., using an antibody agent, such as an antibody, immobilized on a solid support.
[0314] In some embodiments, the modification is methylation, and in some such embodiments, the partitioning step comprises partitioning based on methylation level. In some such embodiments, the agent is a methyl-binding reagent. In some embodiments, the methyl-binding reagent specifically recognizes 5-methylcytosine. In some such embodiments, the agent is a hydroxymethyl-binding reagent. In some embodiments, the methyl-binding reagent specifically recognizes 5-hydroxymethylcytosine, biotinylated 5-hydroxymethylcytosine, glucosylated 5-hydroxymethylcytosine, or sulfonylated 5-hydroxymethylcytosine. In some embodiments, the partitioning step comprises partitioning based on binding to a protein, comprising contacting a sample containing DNA with a binding reagent specific for the protein. In some such embodiments, the binding reagent specifically binds to methylated proteins, acetylated proteins, e.g., methylated histones or acetylated histones. In some embodiments, the binding reagent specifically binds to unmethylated or unacetylated protein epitopes.
[0315] In some embodiments, the modification is hydroxymethylation, and in some such embodiments, the partitioning step comprises partitioning based on hydroxymethylation levels. In some such embodiments, the agent is a hydroxymethyl-binding reagent, e.g., an antibody. In some embodiments, the hydroxymethyl-binding reagent (e.g., an antibody) specifically recognizes 5-hydroxymethylcytosine (5-hmC). In some embodiments, the modification, such as hydroxymethylation, is labeled (e.g., biotinylated, glycosylated, or sulfonated) before contacting with an agent that recognizes the labeled form of the modification. For example, 5-hmC can be enzymatically glycosylated and then partitioned based on binding to J-binding protein 1. Exemplary methods for labeling and / or distributing 5-hmC are provided, for example, in Song et al., Nat. Biotech. 29:68-72 (2010); Ko et al., Nature 468:839-843 (2010); and Robertson et al., Nucleic Acids Res. 39:e55 (2011).
[0316] If immunoprecipitation is used and includes an antibody that recognizes single-stranded DNA, the DNA can be converted to double-stranded form, for example, by complementary strand synthesis prior to the degradation step. Such synthesis can use adapters as primer binding sites or can use random priming.
[0317] In some embodiments, a sample containing DNA is divided into multiple sub-samples. In some embodiments, the multiple divided sub-samples include two sub-samples, namely, a first sub-sample and a second sub-sample. In some embodiments, the multiple divided sub-samples include three sub-samples, namely, a first sub-sample, a second sub-sample, and a third sub-sample. In some embodiments, the method includes a dividing step that is carried out before sequencing and (a) before carrying out methylation-preserving amplification of DNA, (b) after carrying out methylation-preserving amplification of DNA, (c) before enriching DNA for one or more sets of epigenetic target regions, and / or (d) after enriching DNA for one or more sets of epigenetic target regions.
[0318] In some embodiments comprising a third apportioned sub-sample, the third sub-sample comprises DNA that is associated with the modification at a higher rate than it is associated with DNA in the second sub-sample and at a lower rate than it is associated with DNA in the first sub-sample.
[0319] Partitioning nucleic acid molecules in a sample can increase rare signal levels, for example, by enriching for rare nucleic acid molecules that are more abundant in one fraction of the sample. For example, genetic variations present in hypermethylated DNA but less abundant (or absent) in hypomethylated DNA can be more easily detected by partitioning the sample into hypermethylated and hypomethylated nucleic acid molecules. By analyzing multiple fractions of a sample, multidimensional analysis of single molecules can be performed, thus achieving greater sensitivity. Partitioning can include physically partitioning nucleic acid molecules into fractions or subsamples based on the presence or absence of one or more methylated nucleic acid bases. Samples can be partitioned into fractions or subsamples based on features that are indicative of differential gene expression or disease state. Samples can be partitioned based on features, or combinations thereof, that provide a signal difference between normal and diseased states during the analysis of nucleic acids, such as cell-free DNA (cfDNA), non-cfDNA, tumor DNA, circulating tumor DNA (ctDNA), and cell-free nucleic acid (cfNA).
[0320] In some embodiments, hypermethylated and / or hypomethylated variable target regions are analyzed to determine whether they exhibit differential methylation signatures of tumor cells or cell types that do not normally contribute to the DNA sample (e.g., cfDNA) and / or particular immune cell type being analyzed.
[0321] In some embodiments, each fraction is differentially tagged.Then, the tagged fraction can be pooled together for collective sample preparation and / or sequencing.The step of dividing-tagging-pooling can occur more than once, and each dividing round occurs based on different characteristics (examples provided herein) and is tagged with a differential tag that distinguishes it from other fractions and dividing means.In other examples, the fractions that are differentially tagged are sequenced separately.
[0322] In some embodiments, sequence reads of differentially tagged and pooled DNA are obtained and analyzed in silico. Tags are used to sort reads from different fractions. Analysis to detect genetic variants can be performed at the level of each fraction and at the level of the entire nucleic acid population. For example, analysis can include in silico analysis to determine genetic variants, such as copy number variations (CNVs), single nucleotide variations (SNVs), insertions / deletions (indels), and / or fusions in the nucleic acids of each fraction. In some examples, in silico analysis can include determining chromatin structure. For example, the coverage of sequence reads can be used to determine nucleosome positions in chromatin. Higher coverage can be correlated with higher nucleosome occupancy in a genomic region, while lower coverage can be correlated with lower nucleosome occupancy or nucleosome-depleted regions (NDRs).
[0323] Examples of characteristics that can be used for partitioning include sequence length, methylation level, sequence mismatch, immunoprecipitation, and / or proteins that bind to DNA. The resulting fractions can contain one or more of the following nucleic acid types: single-stranded DNA (ssDNA), double-stranded DNA (dsDNA), short DNA fragments, and long DNA fragments. In some embodiments, partitioning based on cytosine modification (e.g., cytosine methylation) or methylation is generally performed, optionally combined with at least one additional partitioning step that can be based on any of the aforementioned DNA characteristics or types. In some embodiments, a heterogeneous nucleic acid population is partitioned into nucleic acids with one or more epigenetic modifications and nucleic acids without one or more epigenetic modifications. Examples of epigenetic modifications include the presence or absence of methylation; the level of methylation; the type of methylation (e.g., 5-methylcytosine vs. other types of methylation, such as adenine methylation and / or cytosine hydroxymethylation); and the association and level of association with one or more proteins, such as histones. Alternatively or additionally, heterogeneous nucleic acid populations can be divided into nucleosome-associated nucleic acid molecules and nucleosome-free nucleic acid molecules.Alternatively or additionally, heterogeneous nucleic acid populations can be divided into single-stranded DNA (ssDNA) and double-stranded DNA (dsDNA).Alternatively or additionally, heterogeneous nucleic acid populations can be divided based on nucleic acid length (for example, molecules up to 160 bp and molecules with a length greater than 160 bp).
[0324] The agent used to partition the nucleic acid population within the sample can be an affinity agent, such as an antibody with desired specificity, its natural binding partner or variant (Bock et al., Nat Biotech 28: 1106-1114 (2010); Song et al., Nat Biotech 29: 68-72 (2011)), or an artificial peptide selected to have specificity for a given target, for example, by phage display. In some embodiments, the agent used in the partitioning step is an agent that recognizes a modified nucleobase. In some embodiments, the modified nucleobase recognized by the agent is a modified cytosine, such as a methylcytosine (e.g., 5-methylcytosine). In some embodiments, the modified nucleobase recognized by the agent is the product of a procedure that affects a first nucleobase in DNA so that it differs from a second nucleobase in the DNA of the sample. In some embodiments, the modified nucleobase may be a "converted nucleobase," meaning that its base-pairing specificity has been altered by the procedure. For example, certain procedures convert unmethylated or unmodified cytosine to dihydrouracil, or more generally, at least one modified or unmodified form of cytosine undergoes deamination, resulting in uracil (considered a modified nucleobase in the context of DNA) or a further modified form of uracil. Examples of partitioning agents include antibodies, such as antibodies that recognize modified nucleobases, which may be modified cytosines, such as methylcytosines (e.g., 5-methylcytosine). In some embodiments, the partitioning agent is an antibody that recognizes modified cytosines other than 5-methylcytosine, such as 5-carboxylcytosine (5caC). Exemplary partitioning agents include proteins such as the methyl-binding domains (MBDs) and methyl-binding proteins (MBPs) described herein, including MeCP2, MBD2, and antibodies that preferentially bind to 5-methylcytosine. When antibodies are used to immunoprecipitate methylated DNA, the methylated DNA can be recovered in single-stranded form. In such embodiments, a second strand can be synthesized.The highly methylated (and optionally, intermediately methylated) sub-sample may then be contacted with a methylation-sensitive nuclease that does not cleave hemimethylated DNA, such as HpaII, BstUI, or Hin6i. Alternatively or additionally, the hypomethylated (and optionally, intermediately methylated) sub-sample may then be contacted with a methylation-dependent nuclease that cleaves hemimethylated DNA.
[0325] Additional non-limiting examples of partitioning agents or binding reagents are histone-binding proteins that can separate histone-bound nucleic acids from free or unbound nucleic acids. Examples of histone-binding proteins that can be used in the methods disclosed herein include RBBP4, RbAp48, and SANT domain peptides.
[0326] In some embodiments, the partitioning step can include both binary partitioning and partitioning based on the degree / level of modification. For example, methylated fragments can be partitioned by methylated DNA immunoprecipitation (MeDIP), or all methylated fragments can be partitioned from unmethylated fragments using a methyl-binding domain protein (e.g., MethylMinder Methylated DNA Enrichment Kit (ThermoFisher Scientific)). Additional partitioning can then involve eluting fragments with different methylation levels by adjusting the salt concentration in the solution containing the methyl-binding domain and bound fragments. As the salt concentration increases, fragments with greater methylation levels are eluted.
[0327] In some cases, the final fraction is enriched in nucleic acids with different degrees of modification (over- or under-representation of the modification). Over- and under-representation can be defined by the number of modifications a nucleic acid has compared to the median number of modifications per strand in the population. For example, if the median number of 5-methylcytosine residues in nucleic acids in a sample is 2, nucleic acids containing more than two 5-methylcytosine residues will be over-represented in this modification, and nucleic acids with one or zero 5-methylcytosine residues will be under-represented. The effect of affinity separation is to enrich nucleic acids that are over-represented in the modification in the binding phase and under-represented in the modification in the non-binding phase (i.e., in solution). The nucleic acids in the binding phase can be eluted before further processing.
[0328] When using MeDIP or MethylMiner® Methylated DNA Enrichment Kit (ThermoFisher Scientific), various methylation levels can be separated using sequential elution. For example, the low-methylated fraction (unmethylated) can be separated from the methylated fraction by contacting the nucleic acid population with MBD from the kit bound to magnetic beads. The beads are used to separate methylated nucleic acids from unmethylated nucleic acids. One or more sequential elution steps are then performed to elute nucleic acids with different methylation levels. For example, the first methylated nucleic acid set can be eluted at a salt concentration of 160 mM or higher, for example, at least 150 mM, at least 200 mM, 300 mM, 400 mM, 500 mM, 600 mM, 700 mM, 800 mM, 900 mM, 1000 mM, or 2000 mM. After eluting such methylated nucleic acids, magnetic separation is again used to separate the more highly methylated nucleic acids from nucleic acids with lower levels of methylation. The elution and magnetic separation steps can be repeated to create various fractions, such as a hypomethylated fraction (enriched in nucleic acids with no methylation), a methylated fraction (enriched in nucleic acids with low levels of methylation), and a hypermethylated fraction (enriched in nucleic acids with high levels of methylation).
[0329] In some methods, the nucleic acids bound to the agent used for the partitioning-based affinity separation are subjected to a washing step. The washing step washes away nucleic acids that are weakly bound to the affinity agent. Such nucleic acids can be enriched for nucleic acids with a degree of modification closer to the mean or median (i.e., intermediate between nucleic acids that remain bound to the solid phase and nucleic acids that do not bind to the solid phase when the sample is first contacted with the agent).
[0330] Affinity separation results in at least two, sometimes three or more fractions of nucleic acids with different degrees of modification. Although the fractions are still separated, the nucleic acids of at least one fraction, usually two or three (or more) fractions, are usually linked to nucleic acid tags provided as components of adapters, and the nucleic acids in different fractions are tagged with different tags that distinguish members of one fraction from members of another fraction. The tags linked to nucleic acid molecules of the same fraction can be the same or different from each other. However, when different from each other, the tags can share part of their code so that the molecules to which they are linked are identified as molecules of a specific fraction.
[0331] For further details regarding partitioning nucleic acid samples based on characteristics such as methylation, see WO2018 / 119452, which is incorporated herein by reference.
[0332] In some embodiments, the partitioning step is performed after contacting the DNA with a methylation-sensitive restriction enzyme (MSRE) and / or a methylation-dependent restriction enzyme (MDRE). After treatment of the DNA with an MSRE or MDRE, the DNA may be partitioned based on size to generate hypermethylated (longest DNA molecules after MSRE treatment and shortest DNA fragments after MDRE treatment), intermediate (intermediate length DNA molecules after MSRE or MDRE treatment), and hypomethylated (shortest DNA molecules after MSRE treatment and longest DNA fragments after MDRE treatment) subsamples.
[0333] In some embodiments, the partitioning step is carried out by contacting the nucleic acid with a methyl-binding domain ("MBD") of a methyl-binding protein ("MBP"). In some such embodiments, the nucleic acid is contacted with the entire MBP. In some embodiments, the MBD binds to 5-methylcytosine (5mC), and the MBP comprises the MBD, and is referred to herein interchangeably as a methyl-binding protein or a methyl-binding domain protein. In some embodiments, the MBD is linked to paramagnetic beads, e.g., Dynabeads® M-280 streptavidin, via a biotin linker. Partitioning into fractions with different degrees of methylation can be carried out by eluting the fractions with increasing NaCl concentrations.
[0334] In some embodiments, the bound DNA is eluted by contacting the antibody or MBD with a protease, such as proteinase K. This may be performed instead of or in addition to the elution step using NaCl discussed herein.
[0335] Examples of agents that recognize modified nucleobases contemplated herein include, but are not limited to:
[0336] (a) MeCP2 and MBD2 are proteins that preferentially bind 5-methyl-cytosine compared to unmodified cytosine.
[0337] (b) RPL26, PRP8, and the DNA mismatch repair protein MHS6 preferentially bind 5-hydroxymethyl-cytosine compared to unmodified cytosine.
[0338] (c) FOXK1, FOXK2, FOXP1, FOXP4, and FOXI3 preferentially bind 5-formyl-cytosine compared to unmodified cytosine (Iurlaro et al., Genome Biol. 14: R119 (2013)).
[0339] (d) an antibody specific for one or more methylated or modified nucleobases, or their conversion products, e.g., 5mC, 5caC, or DHU; Examples include:
[0340] Typically, elution is a function of the number of modifications, such as the number of methylation sites per molecule, with molecules with more methylation eluting at increasing salt concentrations. A series of elution buffers with increasing NaCl concentrations can be used to elute DNA into distinct populations based on the degree of methylation. Salt concentrations can range from about 100 mM to about 2500 mM NaCl. In one embodiment, the process results in three fractions. The molecules are contacted with a solution at a first salt concentration containing molecules containing an agent that recognizes modified nucleobases, allowing the molecules to bind to a capture moiety such as streptavidin. At the first salt concentration, some population of molecules bind to the agent, while others remain unbound. The unbound population can be separated as a "hypomethylated" population. For example, the first fraction enriched in hypomethylated DNA is the fraction that remains unbound at low salt concentrations, e.g., 100 mM or 160 mM. A second fraction enriched in intermediately methylated DNA is eluted using an intermediate salt concentration, e.g., 100 mM to 2000 mM, and is also separated from the sample. A third fraction enriched in highly methylated DNA is eluted using a high salt concentration, e.g., at least about 2000 mM.
[0341] In some embodiments, a monoclonal antibody raised against 5-methylcytidine (5mC) is used to purify methylated DNA. The DNA is denatured, for example, at 95°C to generate single-stranded DNA fragments. Protein G is coupled to standard or magnetic beads, and the antibody-bound DNA is immunoprecipitated using a wash after incubation with the anti-5mC antibody. The DNA can then be eluted. The fractions can include unprecipitated DNA and one or more fractions eluted from the beads. In some embodiments, the DNA fractions are desalted and concentrated in preparation for the enzymatic step of library preparation. H. Pooling DNA from at least the first and second subsamples or portions thereof
[0342] In some embodiments, for example, after the partitioning step, the method includes preparing a pool containing at least a portion of the DNA of the second sub-sample (also referred to as the hypomethylated fraction) and at least a portion of the DNA of the first sub-sample (also referred to as the hypermethylated fraction). Target regions, e.g., target regions containing epigenetic target regions and / or sequence-variable target regions, can be enriched from the pool. The step of enriching for a set of target regions from at least a portion of the sub-samples described elsewhere herein encompasses an enrichment step performed on a pool containing DNA from the first and second sub-samples. A step of amplifying the DNA in the pool may be performed before enriching for target regions from the pool. The enrichment step can have any of the features described elsewhere herein.
[0343] Epigenetic target regions may exhibit differences in methylation levels and / or fragmentation patterns depending on whether they originate from tumors or healthy cells, or the type of tissue from which they originate, as discussed elsewhere herein. Sequence-variable target regions may exhibit sequence differences depending on whether they originate from tumors or healthy cells.
[0344] In some applications, analyzing epigenetic target regions from hypomethylated fractions may provide less information than analyzing sequence variable target regions from hypermethylated and hypomethylated fractions, and epigenetic target regions from hypermethylated fractions.Therefore, in methods in which sequence variable target regions and epigenetic target regions are enriched, the latter may be enriched to a lower degree than one or more of sequence variable target regions from hypermethylated and hypomethylated fractions, and epigenetic target regions from hypermethylated fractions.For example, sequence variable target regions may be enriched from a portion of hypomethylated fractions that are not pooled with hypermethylated fractions, and the pool may be prepared using a portion (e.g., a large proportion, substantially all, or all) of DNA from hypermethylated fractions and none or a portion (e.g., a small proportion) of DNA from hypomethylated fractions.This approach may reduce or eliminate sequencing of epigenetic target regions from hypomethylated fractions, thereby reducing the amount of sequencing data that is sufficient for further analysis.
[0345] In some embodiments, including a small percentage of DNA from the hypomethylated fraction in the pool facilitates quantification of one or more epigenetic traits (e.g., methylation, or other epigenetic traits discussed in detail elsewhere herein), e.g., in terms of relative bias.
[0346] In some embodiments, the pool contains a small proportion of hypomethylated DNA, e.g., less than about 50% of the hypomethylated DNA, e.g., less than or equal to about 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, or 5% of the hypomethylated DNA. In some embodiments, the pool contains about 5%-25% of the hypomethylated DNA. In some embodiments, the pool contains about 10%-20% of the hypomethylated DNA. In some embodiments, the pool contains about 10% of the hypomethylated DNA. In some embodiments, the pool contains about 15% of the hypomethylated DNA. In some embodiments, the pool contains about 20% of the hypomethylated DNA.
[0347] In some embodiments, the pool comprises a portion of the hypermethylated fraction, which may be at least about 50% of the DNA in the hypermethylated fraction. For example, the pool may comprise at least about 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, or 95% of the DNA in the hypermethylated fraction. In some embodiments, the pool comprises 50-55%, 55-60%, 60-65%, 65-70%, 70-75%, 75-80%, 80-85%, 85-90%, 90-95%, or 95-100% of the DNA in the hypermethylated fraction. In some embodiments, the second pool comprises all or substantially all of the hypermethylated fraction.
[0348] In some embodiments, the method includes preparing a first pool comprising at least a portion of DNA from a low methylation fraction. In some embodiments, the method includes preparing a second pool comprising at least a portion of DNA from a high methylation fraction. In some embodiments, the first pool further comprises a portion of DNA from a high methylation fraction. In some embodiments, the second pool further comprises a portion of DNA from a low methylation fraction. In some embodiments, the first pool comprises a large proportion of DNA from a low methylation fraction and, optionally, a small proportion of DNA from a high methylation fraction. In some embodiments, the second pool comprises a large proportion of DNA from a high methylation fraction and a small proportion of DNA from a low methylation fraction. In some embodiments comprising an intermediate methylation fraction, the second pool comprises at least a portion of DNA from an intermediate methylation fraction, e.g., a large proportion of DNA from an intermediate methylation fraction. In some embodiments, the first pool comprises a large proportion of DNA from a low methylation fraction and the second pool comprises a large proportion of DNA from a high methylation fraction and a large proportion of DNA from an intermediate methylation fraction.
[0349] In some embodiments, the method includes enriching for at least a first set of target regions from a first pool, e.g., the first pool is as described in any of the above embodiments. In some embodiments, the first set includes sequence variable target regions. In some embodiments, the first set includes hypomethylated variable target regions and / or fragmented variable target regions. In some embodiments, the first set includes sequence variable target regions and fragmented variable target regions. In some embodiments, the first set includes sequence variable target regions, hypomethylated variable target regions, and fragmented variable target regions. Amplifying the DNA in the first pool may be performed prior to this enrichment step. In some embodiments, enriching for the first set of target regions from the first pool includes contacting the DNA of the first pool with a first set of target-specific probes. In some embodiments, the first set of target-specific probes includes target-binding probes specific for sequence variable target regions. In some embodiments, the first set of target-specific probes includes target-binding probes specific for sequence variable target regions, hypomethylated variable target regions, and / or fragmented variable target regions.
[0350] In some embodiments, the method includes enriching for a second set of target regions or a plurality of sets of target regions from the second pool, e.g., the first pool is as described in any of the above embodiments. In some embodiments, the second plurality of sets includes epigenetic target regions, e.g., hypermethylated variable target regions and / or fragmented variable target regions. In some embodiments, the second plurality of sets includes sequence-variable target regions and epigenetic target regions, e.g., hypermethylated variable target regions and / or fragmented variable target regions. Amplifying the DNA in the second pool may be performed before this enrichment step. In some embodiments, enriching for the second plurality of sets of target regions from the second pool includes contacting the DNA of the first pool with a second set of target-specific probes, the second set of target-specific probes including target-binding probes specific for sequence-variable target regions and target-binding probes specific for epigenetic target regions. In some embodiments, the first set of target regions and the second set of target regions are not identical. For example, the first set of target regions may include one or more target regions that are not present in the second set of target regions. Alternatively or additionally, the second set of target regions may include one or more target regions that are not present in the first set of target regions. In some embodiments, at least one hypermethylated variable target region is enriched from the second pool but not from the first pool. In some embodiments, multiple hypermethylated variable target regions are enriched from the second pool but not from the first pool. In some embodiments, the first set of target regions includes a sequence variable target region and / or the second set of target regions includes an epigenetic target region. In some embodiments, the first set of target regions includes a sequence variable target region and a fragmentation variable target region, and the second set of target regions includes an epigenetic target region, e.g., a hypermethylated variable target region and a fragmentation variable target region.In some embodiments, the first set of target regions includes sequence variable target regions, fragmented variable target regions, and includes hypomethylated variable target regions, and the second set of target regions includes epigenetic target regions, e.g., hypermethylated variable target regions and fragmented variable target regions.
[0351] In some embodiments, the first pool contains a large proportion of DNA from the hypomethylated fraction and a portion (e.g., about half) of the DNA from the hypermethylated fraction, and the second pool contains a portion (e.g., about half) of the DNA from the hypermethylated fraction. In some such embodiments, the first set of target regions includes sequence variable target regions and / or the second set of target regions includes epigenetic target regions. The sequence variable target regions and / or epigenetic target regions may be as described in any of the embodiments described elsewhere herein. I. Detect; Sequencing
[0352] In some embodiments, detecting the presence or absence of DNA sequence and / or modification comprises sequencing.In some embodiments, DNA is sequenced in a manner that distinguishes first nucleic acid base from second nucleic acid base.Generally, the sample nucleic acid that comprises the nucleic acid flanked by adaptor can be subjected to sequencing with or without prior amplification. Sequencing methods include, for example, Sanger sequencing, high-throughput sequencing, pyrosequencing, sequencing-by-synthesis, long-read sequencing (also known as single-molecule sequencing or third-generation sequencing), nanopore sequencing (a type of long-read sequencing), five-letter or six-letter sequencing, semiconductor sequencing, sequencing-by-ligation, sequencing-by-hybridization, Digital Gene Expression (Helicos), next-generation sequencing (NGS), single-molecule sequencing-by-synthesis (SMSS) (Helicos), enzymatic methyl sequencing (EM-Seq), Tet-assisted pyridine borane sequencing (TAPS), massively parallel sequencing, Clonal Single Molecule Array (Solexa), shotgun sequencing, Ion Torrent, Oxford Nanopore, Roche Genia, Maxam-Gilbert sequencing, primer walking, and other sequencing methods available from PacBio, SOLiD, Ion Examples of sequencing include sequencing using the Torrent or Nanopore platform. Sequencing reactions can be carried out in various sample processing units, which can include multiple lanes, multiple channels, multiple wells, or other means for processing multiple sets of samples substantially simultaneously. Sample processing units can also include multiple sample chambers that can process multiple runs simultaneously.
[0353] In some embodiments, sequencing includes detecting and / or distinguishing between unmodified and modified nucleobases. For example, long-read sequencing (also referred to herein as third-generation sequencing) methods include methods that can generate longer sequencing reads, such as reads longer than 10 kilobases, compared to short-read sequencing methods, which generally generate reads up to about 600 bases in length. Compared to short reads, long reads can improve de novo assembly, identification of transcript isoforms, and detection and / or mapping of structural variants. Furthermore, long-read sequencing of native DNA or RNA molecules reduces amplification bias and preserves base modifications, such as methylation status. Long-read sequencing technologies useful herein may include, but are not limited to, any suitable long-read sequencing method, including Pacific Biosciences (PacBio) single-molecule real-time (SMRT) sequencing, Oxford Nanopore Technologies (ONT) nanopore sequencing, and synthetic long-read sequencing approaches, such as concatenated reads, proximal ligation strategies, and optical mapping. The synthetic long read approach involves the assembly of short reads from the same DNA molecule to generate synthetic long reads and can be used together with "true" long read sequencing technologies, such as SMRT and nanopore sequencing methods.
[0354] Single-molecule real-time (SMRT) sequencing, for example, facilitates the direct detection of 5-methylcytosine and 5-hydroxymethylcytosine, as well as unmodified cytosine (Weirather JL, et al., "Comprehensive Comparison of Pacific Biosciences and Oxford Nanopore Technologies and Their Applications to Transcriptome Analysis," F1000Research, 6:100, 2017). While next-generation sequencing methods detect increased signal from clonal populations of amplified DNA fragments, SMRT sequencing captures single DNA molecules and preserves base modifications during sequencing. The error rate of raw PacBio SMRT sequencing-generated data is approximately 13–15%, due to the lack of a high signal-to-noise ratio from single DNA molecules. To increase accuracy, this platform uses a circular DNA template by ligating hairpin adapters to both ends of the target double-stranded DNA. As the polymerase repeatedly traverses and replicates the circular molecule, the DNA template is sequenced multiple times, generating continuous long reads (CLRs). CLRs can be split into multiple reads ("sub-reads") by removing adapter sequences, which generate circular consensus sequence ("CCS") reads with higher accuracy. The average length of a CLR is >10 kb and up to 60 kb, depending on the polymerase lifetime. Thus, the length and accuracy of the CCS read depend on the fragment size. PacBio sequencing has been utilized for genome (e.g., de novo assembly, structural variant detection, and haplotyping) and transcriptome (e.g., gene isoform reconstruction and novel gene / isoform discovery) studies.
[0355] ONT is a nanopore-based single-molecule sequencing technology (Weirather JL, et al., F1000Research, 6:100, 2017). ONT directly sequences native single-stranded DNA (ssDNA) molecules by measuring characteristic current changes as bases are threaded through a nanopore by molecular motor proteins. ONT uses a hairpin library structure similar to the PacBio circular DNA template, where the DNA template and its complement bind to a hairpin adapter. Thus, the DNA template passes through the nanopore, followed by the hairpin, and finally the complement. The raw read can be separated into two "1D" reads ("template" and "complement") by removing the adapter. The consensus sequence of the two "1D" reads is a "2D" read with higher accuracy.
[0356] Five-letter and six-letter sequencing methods include whole-genome sequencing methods that can sequence A, C, T, and G, in addition to 5mC and 5hmC, in a single workflow to provide a five-letter (A, C, T, G, and either 5mC or 5hmC) or six-letter (A, C, T, G, 5mC, and 5hmC) digital readout. DNA sample processing is entirely enzymatic, avoiding the DNA degradation and genome coverage bias of bisulfite treatment. In an exemplary five-letter sequencing method developed by Cambridge Epigenetix, sample DNA is first fragmented via sonication and then ligated to short synthetic DNA hairpin adapters at both ends (Fuellgrabe, et al. 2022, bioRxiv doi: https: / / doi.org / 10.1101 / 2022.07.08.499285). The construct is then split to separate the sense and antisense sample strands. For each original sample strand, a complementary copy strand is synthesized by DNA polymerase extension at the 3' end to create a hairpin construct with the original sample DNA strand connected to its complementary strand lacking epigenetic modifications via a synthesis loop. A sequencing adapter is then ligated to the end. Modified cytosines are enzymatically protected. Unprotected Cs are then deaminated to uracil, which is then read as thymine. In any such embodiment, the amplification method may include a uracil- and / or dihydrouracil-resistant amplification method, such as PCR, using a uracil- and / or dihydrouracil-resistant DNA polymerase (i.e., a DNA polymerase capable of reading and amplifying templates containing uracil and / or dihydrouracil bases). The deaminated construct is no longer fully complementary and has substantially reduced duplex stability; thus, the hairpin can be easily opened and amplified by PCR. The construct can be sequenced in a paired-end format, whereby read 1 (P1 primed) is the original strand and read 2 (P2 primed) is the copy strand. The reads are aligned pairwise, such that read 1 is aligned to its complementary read 2.Cognate residues from both reads are computationally decoded to generate a single genetic or epigenetic signature. Cognate base pairings that differ from the five allowed are the result of imperfect fidelity at some stage, including errors in base calling during sample preparation, amplification, or sequencing. These errors occur independently with cognate bases on each strand, resulting in substitutions that result in ineligible pairs. The ineligible pairs are masked (marked as N) within the decoded read, while the read itself is retained, resulting in minimal information loss and high accuracy at the read level. The decoded reads are aligned to the reference genome. Genetic variant and methylation counts are generated by read counting at the base level.
[0357] 5hmC has been shown to be valuable as a marker of biological states and diseases, including early cancer detection from cell-free DNA. In adapting five-letter sequences to six-letter sequencing, 5mC disambiguates 5hmC without compromising the calling of genetic bases within the same sample fragment. The first three steps of the workflow are identical to the five-letter sequencing described above, creating an adapter to which the sample fragment with the synthetic copy strand is ligated. Methylation at 5mC is enzymatically copied across the CpG units to the copy strand, while 5hmC is enzymatically protected from such copying. Thus, the 5mC and 5hmC of the unmodified C in each original CpG unit are distinguished by a unique two-base combination. Next, unmodified cytosine is deaminated to uracil, which is then read as thymine. The DNA is subjected to PCR amplification and sequencing as previously described. The reads are aligned pairwise and decoded using a two-base code. Since the three CpG units are separate sequencing environments of the two-base code, each of the unmodified C, 5mC, and 5hmC can be decoded.
[0358] In some embodiments, sequencing comprises targeted sequencing, in which one or more genomic regions of interest are sequenced.In some such embodiments, the genomic region of interest comprises a region present in one or more genes selected from Tables 1, 2, 3, 4, and / or 5.In some such embodiments, DNA sequences that do not contain a region of interest are not sequenced.Some embodiments comprise untargeted sequencing, for example, all genomic regions of the DNA in the processed sample or sub-sample are sequenced, or genomic regions are randomly selected for sequencing.In other embodiments, detecting the presence or absence of a sequence of DNA comprises sequencing DNA that is not enriched for the genomic region of interest (untargeted sequencing), for example, detectable sequences are obtained in a substantially unbiased manner.
[0359] Sequencing reactions can be performed on one or more types of nucleic acids, such as nucleic acids known to contain cancer or other disease markers.Sequencing reactions can also be performed on any nucleic acid fragments present in a sample.In some embodiments, the sequence coverage of the genome can be less than 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, 99.9%, or 100%.In some embodiments, sequencing reactions can provide at least 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, or 80% sequence coverage of the genome. Sequence coverage can be performed for at least 5, 10, 20, 70, 100, 200, or 500 different genes, or for at most 5000, 2500, 1000, 500, or 100 different genes.
[0360] Simultaneous sequencing reaction can be carried out using multiplex sequencing.In some embodiments, cell-free nucleic acid can be sequenced by at least 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, 100,000 sequencing reactions.In other embodiments, cell-free nucleic acid can be sequenced by less than 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, 100,000 sequencing reactions.Sequencing reaction can be carried out sequentially or simultaneously.Subsequent data analysis can be carried out for all or part of sequencing reaction. In some examples, data analysis can be performed for at least 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, or 100,000 sequencing reactions. In other examples, data analysis can be performed for less than 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, or 100,000 sequencing reactions. An exemplary read depth is 1000 to 50,000 reads per locus (base).
[0361] In some embodiments, sequences that are not cleaved during the degradation step are sequenced, hi some embodiments, less than 50%, 40%, 30%, 20%, 10%, 5%, 4%, 3%, 2%, or 1% of the sequences that are cleaved during the degradation step are sequenced. J. Subject
[0362] In some embodiments, DNA (e.g., cfDNA or DNA from a sample containing cells) is obtained from a subject (e.g., a test subject) with cancer or precancer. In some embodiments, the subject has stage I cancer, stage II cancer, stage III cancer, or stage IV cancer. In some embodiments, DNA from a subject is obtained and / or derived from a sample obtained from the subject. In some embodiments, DNA is obtained from a subject suspected of having cancer or precancer. In some embodiments, DNA is obtained from a subject with a tumor. In some embodiments, DNA is obtained from a subject suspected of having a tumor. In some embodiments, DNA is obtained from a subject with a neoplasm. In some embodiments, DNA is obtained from a subject suspected of having a neoplasm. In some embodiments, DNA is obtained from a subject in remission from a tumor, cancer, or neoplasm (e.g., after chemotherapy, surgical resection, radiation, or a combination thereof). In any of the foregoing embodiments, the precancer, cancer, tumor, or neoplasm, or suspected precancer, cancer, tumor, or neoplasm, may be a precancer, cancer, tumor, or neoplasm of the bladder, head and neck, lung, colon, rectum, kidney, breast, prostate, skin, or liver. In some embodiments, the precancer, cancer, tumor, or neoplasm, or suspected precancer, cancer, tumor, or neoplasm, is a lung precancer, cancer, tumor, or neoplasm. In some embodiments, the precancer, cancer, tumor, or neoplasm, or suspected precancer, cancer, tumor, or neoplasm, is a colon or rectal precancer, cancer, tumor, or neoplasm. In some embodiments, the precancer, cancer, tumor, or neoplasm, or suspected precancer, cancer, tumor, or neoplasm, is a breast precancer, cancer, tumor, or neoplasm. In some embodiments, the precancer, cancer, tumor, or neoplasm, or suspected precancer, cancer, tumor, or neoplasm, is a prostate precancer, cancer, tumor, or neoplasm. In any of the foregoing embodiments, the subject may be a human subject. In any of the foregoing embodiments, the subject may be a test subject. K. Sample
[0363] The sample may be any biological sample isolated from a subject. The sample may be a bodily sample. The sample may include bodily tissues or fluids, such as known or suspected solid tumors (e.g., carcinomas, adenocarcinomas, or sarcomas), whole blood, platelets, serum, plasma, feces, red blood cells, white blood cells or leucocytes, endothelial cells, tissue biopsies, cerebrospinal fluid, synovial fluid, lymphatic fluid, ascites, interstitial or extracellular fluid, interstitial space fluid, dental crevicular fluid, bone marrow, pleural effusion, pleural fluid, cerebrospinal fluid, saliva, mucus, sputum, semen, sweat, and urine. The sample is preferably a bodily fluid, particularly blood and its fractions, cerebrospinal fluid, pleural fluid, saliva, sputum, or urine. The sample may be in the original form isolated from the subject, or may have been further processed to remove or add components, such as cells, or to enrich one component for another.
[0364] In some embodiments, the nucleic acid population is obtained from serum, plasma, or blood samples from subjects suspected of having a neoplasm, tumor, precancer, or cancer, or from subjects already diagnosed with a neoplasm, tumor, precancer, or cancer. The population includes nucleic acids with various levels of sequence mutations, epigenetic mutations, post-translational chromatin modifications (PTM), and / or post-replicative or post-transcriptional modifications. Post-replicative modifications include cytosine modifications, particularly modifications at the 5th position of the nucleic acid base, such as 5-methylcytosine, 5-hydroxymethylcytosine, 5-formylcytosine, and 5-carboxylcytosine.
[0365] A sample can be isolated or obtained from a subject and transported to a site for sample analysis. The sample can be stored and shipped at a desired temperature, for example, room temperature, 4°C, -20°C, and / or -80°C. The sample can be isolated or obtained from a subject at a site for sample analysis. The subject can be a human, mammal, animal, companion animal, service animal, or pet. The subject can have cancer, precancer, infection, transplant rejection, or other disease or disorder associated with an altered immune system. The subject can have no cancer or detectable symptoms of cancer. The subject can have been treated with one or more cancer therapies, such as any one or more of chemotherapy, antibodies, vaccines, or biologics. The subject can be in remission. The subject may or may not be diagnosed with cancer or a predisposition to any cancer-associated genetic mutation / disorder.
[0366] In some embodiments, the sample comprises plasma. The volume of plasma obtained can depend on the desired read depth of the region being sequenced. Exemplary volumes are 0.4-40 mL, 5-20 mL, 10-20 mL, and 3-5 mL. For example, the volume can be 0.5 mL, 1 mL, 2 mL, 3 mL, 4 mL, 5 mL, 6 mL, 7 mL, 8 mL, 9 mL, 10 mL, 20 mL, 30 mL, or 40 mL. The volume of sampled plasma can be 5-20 mL. In some embodiments, the sample volume is 3-5 mL of plasma, e.g., 4 mL of plasma, per 10 mL of whole blood.
[0367] In some embodiments, the sample comprises whole blood. Exemplary volumes of sampled whole blood are 0.4 to 40 mL, 5 to 20 mL, 10 to 20 mL, 1 to 6 mL, 1 to 3 mL, and 3 to 5 mL. For example, the volume can be 0.5 mL, 1 mL, 2 mL, 3 mL, 4 mL, 5 mL, 6 mL, 7 mL, 8 mL, 9 mL, 10 mL, 20 mL, 30 mL, or 40 mL. The volume of sampled whole blood can be 5 to 20 mL. In some embodiments, the sample volume is 1 to 5 mL of whole blood, e.g., 2.5 mL of whole blood.
[0368] In some embodiments, the sample comprises a buffy coat separated from whole blood. Exemplary volumes of the sampled buffy coat are 0.1-20 mL, 1-10 mL, 1-5 mL, 0.2-0.6 mL, and 0.3-0.5 mL. For example, the volume can be 0.1 mL, 0.2 mL, 0.3 mL, 0.4 mL, 0.5 mL, 0.6 mL, 0.7 mL, 0.8 mL, 0.9 mL, 1 mL, 2 mL, 3 mL, 4 mL, 5 mL, 10 mL, or 20 mL. The volume of the sampled buffy coat can be 1-10 mL. In some embodiments, the sample volume is 0.1-0.5 mL of buffy coat, e.g., 0.3 mL of buffy coat, per 10 mL of whole blood.
[0369] In some embodiments, the sample contains PBMCs isolated from whole blood. Exemplary volumes of sampled PBMCs are 0.1-20 mL, 1-10 mL, 1-5 mL, 0.2-0.6 mL, and 0.3-0.5 mL. For example, volumes can be 0.1 mL, 0.2 mL, 0.3 mL, 0.4 mL, 0.5 mL, 0.6 mL, 0.7 mL, 0.8 mL, 0.9 mL, 1 mL, 2 mL, 3 mL, 4 mL, 5 mL, 10 mL, or 20 mL. The volume of sampled PBMCs can be 1-10 mL. In some embodiments, the sample volume is 0.1-0.5 mL of PBMCs, e.g., 0.3 mL of PBMCs, per 10 mL of whole blood.
[0370] In some embodiments, the sample contains leukocytes separated from the subject's blood using leukocyte reduction. Exemplary volumes of sampled leukocytes from leukocyte reduction are 0.1-20 mL, 1-10 mL, 1-5 mL, 0.2-0.6 mL, and 0.3-0.5 mL. For example, the volume can be 0.1 mL, 0.2 mL, 0.3 mL, 0.4 mL, 0.5 mL, 0.6 mL, 0.7 mL, 0.8 mL, 0.9 mL, 1 mL, 2 mL, 3 mL, 4 mL, 5 mL, 10 mL, or 20 mL. The volume of sampled leukocytes from leukocyte reduction can be 1-10 mL. In some embodiments, the sample volume is 0.1-0.6 mL of leukocytes from leukocyte reduction, e.g., 0.4 mL of leukocytes, per 10 mL of whole blood.
[0371] A sample can contain various amounts of nucleic acid, including genome equivalents. For example, a sample of about 30 ng of DNA contains about 10,000 (10 4 ) haploid human genome equivalents. Similarly, a sample of about 100 ng of DNA can contain about 30,000 haploid human genome equivalents.
[0372] A sample may contain nucleic acids from different sources, e.g., from cells of the same subject or from cells of different subjects. A sample may contain nucleic acids with mutations. For example, a sample may contain DNA with germline mutations and / or somatic mutations. A germline mutation refers to a mutation present in a subject's germline DNA. A somatic mutation refers to a mutation originating in a subject's somatic cells, e.g., precancerous or cancerous cells. A sample may contain DNA with a cancer-associated mutation (e.g., a cancer-associated somatic mutation). A sample may contain epigenetic variants (i.e., chemical or protein modifications), which are associated with the presence of a genetic variant, such as a cancer-associated mutation. In some embodiments, a sample that does not contain a genetic variant contains an epigenetic variant associated with the presence of the genetic variant.
[0373] Exemplary amounts of nucleic acid in a sample prior to amplification (e.g., DNA from a buffy coat sample or any other sample containing cells, such as a blood sample (e.g., a whole blood sample, a leukoreduction sample, or a PBMC sample)) range from about 1 fg to about 1 μg, e.g., 1 pg to 200 ng, 1 ng to 100 ng, or 10 ng to 1000 ng. For example, the amount may be up to about 600 ng, up to about 500 ng, up to about 400 ng, up to about 300 ng, up to about 200 ng, up to about 100 ng, up to about 50 ng, or up to about 20 ng of nucleic acid molecules. The amount may be at least 1 fg, at least 10 fg, at least 100 fg, at least 1 pg, at least 10 pg, at least 100 pg, at least 1 ng, at least 10 ng, at least 150 ng, or at least 200 ng of nucleic acid molecules. The amount may be up to 1 femtogram (fg), 10 fg, 100 fg, 1 picogram (pg), 10 pg, 100 pg, 1 ng, 10 ng, 100 ng, 150 ng, or 200 ng of nucleic acid molecule. The method may include obtaining between 1 femtogram (fg) and 200 ng.
[0374] Nucleic acids can be isolated from cells, e.g., cells in body fluids. The cells can be lysed, and the cellular nucleic acids can be processed. Generally, after adding a buffer and washing steps, the nucleic acids can be precipitated with alcohol. Additional purification steps, such as silica-based columns to remove contaminants or salts, can also be used. In some aspects of the procedure, e.g., to optimize yield, non-specific bulk carrier nucleic acids, such as C1 DNA, DNA, or proteins for bisulfite sequencing, hybridization, and / or ligation, can be added to the entire reaction.
[0375] After such processing, the sample may contain various forms of nucleic acids, including double-stranded DNA, single-stranded DNA, and single-stranded RNA. In some embodiments, single-stranded DNA and RNA are converted to double-stranded forms, which can be included in subsequent processing and analysis steps.
[0376] Reference or control molecule can be added or added to sample as control or normalization standard.For example, a certain amount of modified DNA from a species other than the target species from which sample is obtained or synthetic nucleic acid that contains a certain modification can be added to sample.In some embodiments, reference or control molecule can be distinguished from the molecule that originally exists in sample.In some embodiments, detected DNA sequence is normalized to reference or control molecule.
[0377] DNA molecules can be ligated with adapters at either one or both ends. Typically, double-stranded molecules are blunt-ended by treatment with a polymerase containing a 5'-3' polymerase and a 3'-5' exonuclease (or proofreading function) in the presence of all four standard nucleotides. Klenow large fragment and T4 polymerase are examples of suitable polymerases. The blunt-ended DNA molecule can be ligated to an at least partially double-stranded adapter (e.g., a Y-shaped or bell-shaped adapter). Alternatively, complementary nucleotides can be added to the blunt ends of the sample nucleic acid and adapter to facilitate ligation. Both blunt-end and sticky-end ligation are contemplated herein. In blunt-end ligation, both the nucleic acid molecule and the adapter tag have blunt ends. In sticky-end ligation, the nucleic acid molecule typically has an "A" overhang and the adapter has a "T" overhang. L.Analysis
[0378] The present disclosure provides methods for analyzing DNA using a methylation-preserving amplification step. In some embodiments, the disclosed methods include analyzing DNA (e.g., DNA from a subject) to identify at least one cell type, cell cluster type, tissue type, and / or cancer type from which one or more type-specific epigenetic target regions and / or type-specific sequence variable target regions originated. In some embodiments, the methods include determining the level of one or more type-specific epigenetic target regions and / or type-specific sequence variable target regions originating from the at least one cell type, cell cluster type, tissue type, and / or cancer type.
[0379] An exemplary method for analyzing DNA includes the following steps (e.g., in the order listed below), which is illustrated in FIG. 1A: 1. Preparing an extracted DNA sample (e.g., extracting DNA, e.g., cfDNA, from a human sample, e.g., a blood sample). 2. Subjecting the extracted DNA to library preparation, including attaching tags (e.g., adapters) containing molecular barcodes, which are optionally methylated. 3. Performing methylation-preserving amplification (also called methylation-maintaining amplification). 4. Performing either steps 4A-1 and 4A-2 or steps 4B-1 and 4B-2: 4A-1. Subjecting the DNA from step 3 to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA, e.g., any form of epigenetic base conversion described elsewhere herein, e.g., bisulfite conversion. 4A-2. Enriching the DNA from step 4A-1 for nucleic acid molecules containing sequences present in one or more target region sets (e.g., an epigenetic target region set and / or a sequence variable target region set, each of which may include any one or more of the epigenetic target region sets and sequence variable target region sets described elsewhere herein), and optionally amplifying the enriched DNA. 4B-1. Enriching the DNA from step 3 for nucleic acid molecules containing sequences present in one or more target region sets (e.g., an epigenetic target region set and / or a sequence variable target region set, each of which may include any one or more of the epigenetic target region sets and sequence variable target region sets described elsewhere herein), and optionally amplifying the enriched DNA. 4B-2. Subjecting the enriched DNA to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA, e.g., any form of epigenetic base conversion described elsewhere herein, e.g., bisulfite conversion. 5. Sequencing the DNA via NGS. 6. Performing bioinformatics analysis of the NGS data, including using one or more components of each read (e.g., a barcode or barcode set in combination with a genomic start site and / or a genomic stop site) to identify reads and group reads into molecular families (e.g., DNA molecules), where reads within a molecular family have (a) a common molecular barcode and (b) the same genomic start position, (c) the same genomic stop position, or (d) both (b) and (c). 7. Determining a consensus epigenetic sequence (methylation state) for each cytosine in the sequence of a given molecular family.
[0380] Another exemplary method for analyzing DNA includes the following steps (e.g., in the order listed below), of which steps 3, 5, and 8-11 are illustrated in FIG. 1A: 1. Preparing an extracted DNA sample (e.g., extracting DNA, e.g., cfDNA, from a human sample, e.g., a blood sample). 2. Subjecting the extracted DNA to library preparation, including ligating an adapter containing a molecular barcode, optionally methylated. The adapter further comprises a restriction enzyme cleavage site between the barcode and at least a portion of the remainder of the adapter. In such cases, the barcode is located between the DNA sequence to be analyzed and the restriction site. For example, in an adapter ligated to the 5' end of a DNA sequence, the barcode is 3' of the restriction site and 5' of the DNA sequence to be analyzed. Furthermore, the barcode in this example does not contain cytosines in a non-CpG context. 3. Performing methylation-preserving amplification (also called methylation-maintaining amplification). 4. Contacting the amplified DNA with a restriction enzyme that recognizes a restriction site in the adapter to remove at least a portion of the adapter that does not contain a barcode. In some embodiments, cleavage of the DNA with the restriction enzyme results in an overhang (e.g., an A or T overhang) that can be used, for example, as a cohesive end for subsequent ligation. 5. Subjecting the DNA from step 3 to a procedure that affects a first nucleobase in the DNA differently than a second nucleobase in the DNA, wherein the first nucleobase is an unmethylated cytosine and the subjecting affects the unmethylated cytosine (e.g., bisulfite conversion, EM-Seq, or SEM-Seq). 6. Attaching an auxiliary adapter to the DNA, wherein the auxiliary adapter does not include a barcode. 7. Performing uracil and / or dihydrouracil resistant amplification. 8. Enriching the DNA from step 8 for nucleic acid molecules comprising sequences present in one or more target region sets (e.g., an epigenetic target region set and / or a sequence variable target region set, each of which may include any one or more of the epigenetic target region sets and sequence variable target region sets described elsewhere herein), and optionally amplifying the enriched DNA, wherein the amplification includes, for example, differentially tagging the enriched DNA for one or more epigenetic target region sets or one or more sequence variable target region sets. The differential tagging may include attaching a sample index to the DNA. 9. Sequencing the DNA via NGS. 10. Performing bioinformatics analysis of the NGS data, including using one or more components of each read (e.g., a barcode or barcode set in combination with a genomic start site and / or a genomic stop site) to identify reads and group reads into molecular families (e.g., DNA molecules), where reads within a molecular family have (a) a common molecular barcode and (b) the same genomic start position, (c) the same genomic stop position, or (d) both (b) and (c). 11. Determining a consensus epigenetic sequence (methylation state) for each cytosine in the sequence of a given molecular family.
[0381] Another exemplary method for analyzing DNA includes the following steps (e.g., in the order listed below), which is illustrated in FIG. 1B: 1. Preparing an extracted DNA sample (e.g., extracting DNA, e.g., cfDNA, from a human sample, e.g., a blood sample). 2. Subjecting the extracted DNA to library preparation, including optionally attaching tags (e.g., adapters) comprising molecular barcodes, which are optionally methylated. 3. Performing linear methylation-preserving amplification (also called methylation-maintaining amplification), for example by providing a reverse primer as the only primer. 4. Performing either steps 4A-1 and 4A-2 or steps 4B-1 and 4B-2: 4A-1. Subjecting the extracted DNA to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA, e.g., any form of epigenetic base conversion described elsewhere herein, e.g., bisulfite conversion. 4A-2. Enriching the DNA from step 4A-1 for nucleic acid molecules containing sequences present in one or more target region sets (e.g., an epigenetic target region set and / or a sequence variable target region set, each of which may include any one or more of the epigenetic target region sets and sequence variable target region sets described elsewhere herein), and optionally amplifying the enriched DNA. 4B-1. Enriching the DNA from step 3 for nucleic acid molecules containing sequences present in one or more target region sets (e.g., an epigenetic target region set and / or a sequence variable target region set, each of which may include any one or more of the epigenetic target region sets and sequence variable target region sets described elsewhere herein), and optionally amplifying the enriched DNA. 4B-2. Subjecting the enriched DNA to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA, e.g., any form of epigenetic base conversion described elsewhere herein, e.g., bisulfite conversion. 5. Sequencing the DNA via NGS. 6. Performing bioinformatics analysis of the NGS data, including using one or more components of each read (e.g., a barcode or barcode set in combination with a genomic start site and / or a genomic stop site) to identify reads and group reads into molecular families (e.g., DNA molecules), where reads within a molecular family have (a) a common molecular barcode and (b) the same genomic start position, (c) the same genomic stop position, or (d) both (b) and (c). 7. Determining a consensus epigenetic sequence (methylation state) for each cytosine in the sequence of a given molecular family.
[0382] Another exemplary method for analyzing DNA includes the following steps (e.g., in the order listed below), which is illustrated in FIG. 1B: 1. Preparing an extracted DNA sample (e.g., extracting DNA, e.g., cfDNA, from a human sample, e.g., a blood sample). 2. Subjecting the extracted DNA to library preparation, which includes attaching adapters containing molecular barcodes. 3. Performing methylation-preserving linear amplification. 4. Performing either steps 4A-1 and 4A-2 or steps 4B-1 and 4B-2: 4A-1. Subjecting the DNA from step 3 to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA, e.g., any form of epigenetic base conversion described elsewhere herein, e.g., bisulfite conversion. 4A-2. Enriching the DNA from step 4A-1 for nucleic acid molecules containing sequences present in one or more target region sets (e.g., an epigenetic target region set and / or a sequence variable target region set, each of which may include any one or more of the epigenetic target region sets and sequence variable target region sets described elsewhere herein), and optionally amplifying the enriched DNA. 4B-1. Enriching the DNA from step 3 for nucleic acid molecules containing sequences present in one or more target region sets (e.g., an epigenetic target region set and / or a sequence variable target region set, each of which may include any one or more of the epigenetic target region sets and sequence variable target region sets described elsewhere herein), and optionally amplifying the enriched DNA. 4B-2. Subjecting the enriched DNA to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA, e.g., any form of epigenetic base conversion described elsewhere herein, e.g., bisulfite conversion. 5. Sequencing the DNA via NGS. 6. Performing bioinformatics analysis of the NGS data, including using one or more components of each read (e.g., a barcode or barcode set, a genomic start site, and / or a genomic stop site) to identify NGS reads and group the NGS reads into molecular families (e.g., DNA molecules), where reads within a molecular family have (a) a common molecular barcode, (b) the same genomic start position, (c) the same genomic stop position, or a combination of (a), (b), and / or (c). 7. Determining a consensus epigenetic sequence (methylation state) for each cytosine / base in the sequence of a given molecular family.
[0383] Another exemplary method for analyzing DNA includes the following steps (e.g., in the order listed below), of which steps 3, 5, and 8-11 are illustrated in FIG. 1B: 1. Preparing an extracted DNA sample (e.g., extracting DNA, e.g., cfDNA, from a human sample, e.g., a blood sample). 2. Subjecting the extracted DNA to library preparation, including ligating an adapter containing a molecular barcode. The adapter further comprises a restriction enzyme cleavage site between the barcode and at least a portion of the remainder of the adapter. In such a case, the barcode is located between the DNA sequence to be analyzed and the restriction site. For example, in an adapter ligated to the 5' end of a DNA sequence, the barcode is 3' of the restriction site and 5' of the DNA sequence to be analyzed. Furthermore, the barcode in this example does not contain a cytosine in a non-CpG context. 3. Performing methylation-preserving linear amplification. 4. Contacting the amplified DNA with a restriction enzyme that recognizes a restriction site in the adapter to remove at least a portion of the adapter that does not contain a barcode. In some embodiments, cleavage of the DNA with the restriction enzyme results in an overhang (e.g., an A or T overhang) that can be used, for example, as a cohesive end for subsequent ligation. 5. Subjecting the DNA from step 3 to a procedure that affects a first nucleobase in the DNA differently than a second nucleobase in the DNA, wherein the first nucleobase is an unmethylated cytosine and the subjecting affects the unmethylated cytosine (e.g., bisulfite conversion, EM-Seq, or SEM-Seq). 6. Attaching an auxiliary adapter to the DNA, wherein the auxiliary adapter does not include a barcode. 7. Performing uracil and / or dihydrouracil resistant amplification. 8. Enriching the DNA from step 8 for nucleic acid molecules comprising sequences present in one or more target region sets (e.g., an epigenetic target region set and / or a sequence variable target region set, each of which may include any one or more of the epigenetic target region sets and sequence variable target region sets described elsewhere herein), and optionally amplifying the enriched DNA, wherein the amplification includes, for example, differentially tagging the enriched DNA for one or more epigenetic target region sets or one or more sequence variable target region sets. The differential tagging may include attaching a sample index to the DNA. 9. Sequencing the DNA via NGS. 10. Performing bioinformatics analysis of the NGS data, including using one or more components of each read (e.g., a barcode or barcode set, a genomic start site, and / or a genomic stop site) to identify NGS reads and group the NGS reads into molecular families (e.g., DNA molecules), where reads within a molecular family have (a) a common molecular barcode, (b) the same genomic start position, (c) the same genomic stop position, or a combination of (a), (b), and / or (c). 11. Determining a consensus epigenetic sequence (methylation state) for each cytosine / base in the sequence of a given molecular family.
[0384] In some embodiments, the consensus epigenetic sequence (methylation state) for each cytosine / base in a sequence of a given molecular family can be determined based on a quantitative measure obtained from reads (within the family) that indicate whether it is methylated or not. In some embodiments, a base can be considered methylated if the quantitative measure exceeds the methylation level threshold for the family. The methylation level threshold for the family to call a given base methylated should be informed by the methyltransferase fidelity (i.e., sensitivity and / or specificity) and the number of amplification cycles if the amplification is PCR, or more generally, exponential / quasi-exponential in nature (DNA copies can be templates for subsequent DNA amplification).
[0385] In some embodiments, detecting the presence, level, or absence of DNA sequences and / or modifications facilitates the diagnosis of disease or the identification of appropriate treatments. In some embodiments, the presence or alteration of the level of one or more sequences and / or modifications indicates the presence of a disease or disorder in a subject, such as cancer or precancer, or other disorder that causes changes in nucleic acids compared to healthy subjects.
[0386] Furthermore, the present methods can be used to diagnose the presence of a condition, particularly cancer or precancer, in a subject, characterize the condition (e.g., stage the cancer or determine the heterogeneity of the cancer), monitor the response to treatment of the condition, and determine a prognostic risk of progression of the condition or subsequent progression of the condition. In some embodiments, the condition is cancer or precancer. In some embodiments, the condition is characterized (e.g., stage the cancer or determine the heterogeneity of the cancer), the response to treatment of the condition is monitored, or the prognostic risk of progression of the condition or subsequent progression of the condition is determined. The present disclosure can also be useful for determining the effectiveness of certain treatment options. A successful treatment option may reduce the amount of detected DNA sequences associated with cancer in the subject's blood, as fewer cancer cells may shed DNA. In other examples, this may not occur. In another example, certain treatment options may correlate with the genetic profile of the cancer over time. This correlation may be useful for selecting a therapy.
[0387] Additionally, if the cancer is observed to be in remission after treatment, the method can be used to monitor for residual disease or recurrence of the disease.
[0388] The types and number of cancers that can be detected can include blood cancer, brain cancer, lung cancer, skin cancer, nasal cancer, pharyngeal cancer, liver cancer, bone cancer, lymphoma, pancreatic cancer, skin cancer, intestinal cancer, rectal cancer, colon cancer, prostate cancer, thyroid cancer, bladder cancer, head and neck cancer, kidney cancer, oral cancer, stomach cancer, solid tumors, heterogeneous tumors, homogeneous tumors, and the like. The type and / or stage of cancer can be detected by genetic variations, including mutations, rare mutations, indels, copy number variations, transversions, translocations, recombinations, inversions, deletions, aneuploidy, partial aneuploidy, polyploidy, chromosomal instability, changes in chromosomal structure, gene fusions, chromosomal fusions, gene truncations, gene amplifications, gene duplications, chromosomal damage, DNA damage, abnormal changes in chemical modifications of nucleic acids, abnormal changes in epigenetic patterns, and abnormal changes in nucleic acid 5-methylcytosine.
[0389] In some embodiments, the methods described herein include detecting the presence or absence of nucleic acid, such as DNA, produced by a tumor (or neoplastic or cancerous cell) or by a precancerous cell.
[0390] The information and data generated by the methods disclosed herein can also be used to characterize specific forms of cancer. Cancers are often heterogeneous in both their composition and stage classification. The methods disclosed herein can enable characterization of specific subtypes of cancer, which can be important in diagnosing or treating that specific subtype. This information can also provide subjects or clinicians with clues regarding the prognosis of a particular type of cancer, allowing them to tailor treatment options as the disease progresses. Some cancers may progress to become more aggressive and genetically unstable. Other cancers may remain benign, inactive, or dormant. The systems and methods disclosed herein can be useful in determining disease progression.
[0391] Furthermore, the methods of the present disclosure can be used to characterize the heterogeneity of a condition in a subject. Such methods can include, for example, generating an aggregation profile of extracellular nucleic acids derived from a subject, where the aggregation profile includes multiple data obtained from various nucleic acid analyses. In some embodiments, the aggregation profile includes epigenetic and mutational analyses. In some embodiments, the aggregation profile includes the sum of information derived from different cells in a heterogeneous disease. This sum can include structural variation identifiers and levels, copy number variations, epigenetic variations, or other mutation analyses.
[0392] This method can be used to diagnose, prognose, monitor or observe pre-cancer, cancer or other diseases.In some embodiments, the method herein does not include the diagnosis, prognosis or monitoring of fetus, and therefore is not intended for non-invasive prenatal testing.In other embodiments, these methodologies can be used in pregnant subjects to diagnose, prognose, monitor or observe cancer or other diseases in unborn subjects whose DNA and other polynucleotides may circulate with maternal molecules. III. ADDITIONAL FEATURES OF CERTAIN DISCLOSED METHODS A. Target Region Set
[0393] In some embodiments, a specific genomic region of interest is detected and / or enriched. The genomic region of interest may comprise one or more target region sets. In some embodiments, the target region set comprises variations that are not present in DNA from healthy subjects or DNA obtained from healthy tissue regions. In some embodiments, the target region set comprises variations that are present in healthy cells but are not normally present in a sample type, such as a blood sample. In some embodiments, the variations are present in abnormal cells (e.g., hyperplastic, dysplastic, or neoplastic cells). Exemplary target region sets include sequence-variable target region sets and epigenetic target region sets.
[0394] In some embodiments, a first set of target regions is detected, the first set comprising at least epigenetic target regions. In some embodiments, the epigenetic target regions detected in the first subsample comprise hypermethylated variable target regions. In some embodiments, the hypermethylated variable target regions are CpG-containing regions that are unmethylated or hypomethylated (e.g., below average methylation compared to bulk cfDNA) in cfDNA from healthy subjects. In some embodiments, the hypermethylated variable target regions exhibit type-specific hypermethylation in healthy cfDNA from one or more related cell types or tissue types. Without wishing to be bound by any particular theory, the presence of cancer cells may increase the shedding of DNA into the bloodstream (e.g., from cancer and / or surrounding tissues). Therefore, the distribution of tissues of origin of cfDNA may change during carcinogenesis. Thus, an increased level of hypermethylated variable target regions in the first subsample may be an indicator of the presence of cancer (or recurrence, depending on the subject's medical history).
[0395] In some embodiments, the method herein includes detecting a second set of captured target regions from a sample or a second sub-sample containing at least epigenetic target regions. In some embodiments, the second set of epigenetic target regions includes hypomethylated variable target regions. In some embodiments, the hypomethylated variable target regions are CpG-containing regions that are methylated or have high methylation (e.g., above average methylation compared to bulk cfDNA) in cfDNA from healthy subjects. Without wishing to be bound by any particular theory, cancer cells may shed more DNA into the bloodstream than healthy cells of the same tissue type. Therefore, the distribution of tissues of origin of cfDNA may change during carcinogenesis. Thus, an increase in the level of hypomethylated variable target regions in the second sub-sample may be an indicator of the presence of cancer (or recurrence, depending on the subject's medical history).
[0396] In addition, the set of target regions can include DNA corresponding to a set of sequence-variable target regions. 1. Epigenetic target region set
[0397] In some embodiments, the target region set is or comprises an epigenetic target region set.The epigenetic target region set can comprise one or more types of target region that may distinguish the DNA from neoplastic (e.g., tumor or cancer) cells from the DNA from healthy cells, such as non-neoplastic circulating cells.Exemplary types of such regions are discussed in detail herein.The epigenetic target region set can also comprise one or more control regions, for example, as described herein.
[0398] In some embodiments, the set of epigenetic target regions has a footprint of at least 100 kbp, e.g., at least 200 kbp, at least 300 kbp, or at least 400 kbp. In some embodiments, the set of epigenetic target regions has a footprint in the range of 100-20 Mbp, e.g., 100-200 kbp, 200-300 kbp, 300-400 kbp, 400-500 kbp, 500-600 kbp, 600-700 kbp, 700-800 kbp, 800-900 kbp, 900-1,000 kbp, 1-1.5 Mbp, 1.5-2 Mbp, 2-3 Mbp, 3-4 Mbp, 4-5 Mbp, 5-6 Mbp, 6-7 Mbp, 7-8 Mbp, 8-9 Mbp, 9-10 Mbp, or 10-20 Mbp. In some embodiments, the set of epigenetic target regions has a footprint of at least 20 Mbp. a. Hypermethylated and hypomethylated variable target regions
[0399] In some embodiments, the epigenetic target region set comprises a hypermethylated variable target region.In some embodiments, the hypermethylated variable target region is differentially or exclusively hypermethylated in one or more related cell types or tissue types.Such hypermethylated variable target region may be hypermethylated in other cell types or tissue types, but not to the same extent as observed in one or more related cell types or tissue types.In some embodiments, the hypermethylated variable target region shows even higher methylation in the cfDNA from diseased cells of one or more related cell types or tissue types.
[0400] In some embodiments, a hypermethylated variable target region refers to a region, e.g., in a cfDNA sample, where an increased observed methylation level indicates an increased likelihood that the sample (e.g., a cfDNA sample) contains DNA produced by neoplastic cells, e.g., tumor or cancer cells. For example, hypermethylation of promoters of tumor suppressor genes has been repeatedly observed. See, e.g., Kang et al., Genome Biol. 18:53 (2017) and the references cited therein. In another example, as discussed above, a hypermethylated variable target region may include a region that does not necessarily differ in methylation in cancerous tissue compared to DNA from healthy tissue of the same type, but differs in methylation (e.g., has more methylation) compared to cfDNA typical of healthy subjects. For example, if the presence of cancer results in increased cell death, such as apoptosis, of cells of the tissue type corresponding to the cancer, such cancer can be detected, at least in part, using such a hypermethylated variable target region.
[0401] An extensive discussion of methylation variable target regions in colorectal cancer is provided in Lam et al., Biochim Biophys Acta. 1866:106-20 (2016). These include VIM, SEPT9, ITGA4, OSM4, GATA4, and NDRG4. An exemplary set of hypermethylated variable target regions based on studies of colorectal cancer (CRC) is provided in Table 1. Many of these genes likely have relevance to cancers other than colorectal cancer; for example, TP53 is a critical tumor suppressor, and it is widely recognized that hypermethylation-based inactivation of this gene may be a common mechanism of tumorigenesis. [Table 1-1] [Table 1-2]
[0402] In some embodiments, the genomic region targeted for sequencing includes multiple loci listed in Table 1, for example, at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% of the loci listed in Table 1. In some embodiments, the genomic region is captured using probes. For example, for each locus included as a target region, there may be one or more probes with hybridization sites that bind between the transcription start site of the gene and the stop codon (or the final stop codon in the case of an alternatively spliced gene) or in the promoter region of the gene. In some embodiments, the one or more probes bind within 300 bp, for example, within 200 or 100 bp, of the transcription start site of the gene in Table 1.
[0403] Methylation variable target regions in various lung cancer types are described, for example, in Ooki et al., Clin. Cancer Res. 23:7141-52 (2017); Belinksy, Annu. Rev. Physiol. 77:453-74 (2015); Hulbert et al., Clin. Cancer Res. 23:1998-2005 (2017); Shi et al., BMC Genomics 18:901 (2017); Schneider et al., BMC Cancer. 11:102 (2011); Lissa et al., Transl Lung Cancer Res 5(5):492-504 (2016); Skvortsova et al., Br. J. Cancer. 94(10):1492-1495 (2006); Kim et al., Cancer Res. 61:3419-3424 (2001);Furonaka et al., Pathology International 55:303-309 (2005);Gomes et al., Rev. Port. Pneumol. 20:20-30 (2014);Kim et al., Oncogene. 20:1765-70 (2001);Hopkins-Donaldson et al., Cell Death Differ. 10:356-64 (2003);Kikuchi et al., Clin. Cancer Res. 11:2954-61 (2005);Heller et al., Oncogene 25:959-968 (2006);Licchesi et al., Carcinogenesis. 29:895-904 (2008);Guo et al., Clin. Cancer Res. 10:7917-24 (2004);Palmisano et al., Cancer Res. 63:4620-4625 (2003); and Toyooka et al., Cancer Res. 61:4556-4560, (2001).
[0404] An exemplary set of hypermethylated variable target regions based on lung cancer studies is provided in Table 2. Many of these genes may also have relevance to cancers other than lung cancer; for example, Casp8 (caspase 8) is a key enzyme in programmed cell death, and hypermethylation-based inactivation of this gene may be a common tumorigenesis mechanism not limited to lung cancer. In addition, several genes appear in both Tables 1 and 2, indicating generality. [Table 2]
[0405] Any of the foregoing embodiments relating to target regions identified in Table 2 may be combined with any of the above embodiments relating to target regions identified in Table 1. In some embodiments, the genomic regions targeted for sequencing include multiple loci listed in Table 1 or Table 2, e.g., at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% of the loci listed in Table 1 or Table 2.
[0406] Additional hypermethylated target regions may be obtained, for example, from the Cancer Genome Atlas. Kang et al., Genome Biology 18:53 (2017) describe the construction of a probabilistic method called CancerLocator using hypermethylated target regions from breast, colon, kidney, liver, and lung. In some embodiments, the hypermethylated target regions may be specific to one or more types of cancer. Thus, in some embodiments, the hypermethylated target regions include one, two, three, four, or five subsets of hypermethylated target regions that collectively exhibit hypermethylation in one, two, three, four, or five of breast cancer, colon cancer, kidney cancer, liver cancer, and lung cancer.
[0407] In some embodiments, the set of epigenetic target regions comprises hypomethylated variable target regions. In some embodiments, the hypomethylated variable target regions are exclusively hypomethylated in one or more related cell types or tissue types. Such hypomethylated variable target regions may be hypomethylated in other cell types or tissue types, but not to the same extent as observed in one or more related cell types or tissue types.
[0408] In some embodiments, when different epigenetic target regions are captured, the epigenetic target regions include hypermethylated variable target regions and / or hypomethylated variable target regions.
[0409] Additionally, exemplary hypermethylated and hypomethylated variable target regions useful for distinguishing between various cell types have been identified by analyzing DNA from various cell types via whole-genome bisulfite sequencing, as described, for example, in Scott, CA, Duryea, JD, MacKay, H. et al., "Identification of cell type-specific methylation signals in bulk whole genome bisulfite sequencing data," Genome Biol, 21, 156 (2020) (doi.org / 10.1186 / s13059-020-02065-5). Whole-genome bisulfite sequencing data is available from the Blueprint Consortium, available online at dcc.blueprint-epigenome.eu. b.CTCF binding region
[0410] In some embodiments, the epigenetic target region set comprises CTCF binding region. CTCF is a DNA binding protein that contributes to chromatin organization and often co-localizes with cohesin. Perturbation of CTCF binding site has been reported in a variety of different cancers. For example, see Katainen et al., Nature Genetics, doi:10.1038 / ng.3335; Guo et al., Nat. Commun. 9:1520 (2018), published online on June 8, 2015. CTCF binding results in a recognizable pattern of cfDNA, which can be detected by sequencing, for example, through fragment length analysis. Thus, perturbation of CTCF binding results in fluctuations in the fragmentation pattern of cfDNA. Therefore, CTCF binding site is a type of fragmentation variable target region.
[0411] There are many known CTCF binding sites.See, for example, CTCFBSDB (CTCF Binding Site Database), available on the Internet at insulatordb.uthsc.edu / ; Cuddapah et al., Genome Res. 19:24-32 (2009); Martin et al., Nat. Struct. Mol. Biol. 18:708-14 (2011); Rhee et al., Cell. 147:1408-19 (2011), each of which is incorporated herein by reference.Exemplary CTCF binding sites are nucleotides 56014955-56016161 on chromosome 8 and nucleotides 95359169-95360473 on chromosome 13.
[0412] In some embodiments, the CTCF binding regions include at least 10, 20, 50, 100, 200, or 500 CTCF binding regions, or 10-20, 20-50, 50-100, 100-200, 200-500, or 500-1000 CTCF binding regions, such as those listed above or in the CTCFBSDB or one or more of the above-cited articles by Cuddapah et al., Martin et al., or Rhee et al. In some embodiments, at least a portion of the CTCF sites can be methylated or unmethylated, and the methylation status correlates with whether the cell is a cancer cell. In some embodiments, the set of epigenetic target regions includes regions at least 100 bp, at least 200 bp, at least 300 bp, at least 400 bp, at least 500 bp, at least 750 bp, or at least 1000 bp upstream and downstream of the CTCF binding site. C transcription start site
[0413] In some embodiments, the epigenetic target region set comprises variable transcription start sites.Transcription start sites may show perturbations in neoplastic cells.For example, the nucleosome organization at various transcription start sites in healthy cells of hematopoietic lineage contributes substantially to cfDNA in healthy individuals, but may differ from the nucleosome organization at those transcription start sites in neoplastic cells.This results in different cfDNA patterns, which can be detected by sequencing, as generally discussed in, for example, Snyder et al., Cell 164:57-68 (2016); WO2018 / 009723; and US20170211143A1.In another example, transcription start sites are not necessarily epigenetically different in cancerous tissues compared with DNA from the same type of healthy tissue, but may be epigenetically different (for example, in terms of nucleosome organization) compared with DNA typical in healthy subjects.Transcription start site perturbations also result in variations in the fragmentation pattern of cfDNA. Therefore, the transcription start site is also one type of fragmentation variable target region.
[0414] Human transcription start sites are available from the Database of Human Transcription Start Sites (DBTSS), available online at dbtss.hgc.jp, and are described in Yamashita et al., Nucleic Acids Res. 34(Database issue): D86-D89 (2006), incorporated herein by reference. In some embodiments, the transcription start sites include at least 10, 20, 50, 100, 200, or 500 transcription start sites, or 10-20, 20-50, 50-100, 100-200, 200-500, or 500-1000 transcription start sites, e.g., transcription start sites described in DBTSS. In some embodiments, at least a portion of the transcription start sites can be methylated or unmethylated, and the methylation status correlates with whether the cell is a cancer cell. In some embodiments, the set of epigenetic target regions includes regions at least 100 bp, at least 200 bp, at least 300 bp, at least 400 bp, at least 500 bp, at least 750 bp, at least 1000 bp upstream and downstream of the transcription start site. d. Local amplification
[0415] Although local amplification is somatic mutation, it can be detected by sequencing based on read frequency in a similar manner to the approach of detecting certain epigenetic changes, such as methylation changes.Therefore, the region that may show local amplification in cancer can be included in epigenetic target region set, and it can include one or more of AR, BRAF, CCND1, CCND2, CCNE1, CDK4, CDK6, EGFR, ERBB2, FGFR1, FGFR2, KIT, KRAS, MET, MYC, PDGFRA, PIK3CA and RAF1. e. Methylation control region
[0416] It may be useful to include a control region to facilitate data validation. In some embodiments, the set of epigenetic target regions includes a control region that is expected to be methylated or unmethylated in essentially all samples, regardless of whether the DNA is derived from cancer cells or normal cells. In some embodiments, the set of epigenetic target regions includes a control hypomethylated region that is expected to be hypomethylated in essentially all samples. In some embodiments, the set of epigenetic target regions includes a control hypermethylated region that is expected to be hypermethylated in essentially all samples. 2. Sequence-variable target region set
[0417] In some embodiments, the target region set is or includes a sequence variable target region set. The sequence variable target region set may include one or more types of target regions that may distinguish DNA from neoplastic (e.g., tumor or cancer) cells from DNA from healthy cells, e.g., non-neoplastic circulating cells. Exemplary types of such regions are discussed in detail herein. The sequence variable target region set may also include one or more control regions, for example, as described herein. In some embodiments, the sequence variable target region set includes multiple regions known to undergo somatic mutations in cancer. In some aspects, the sequence variable target region set targets multiple different genes or genomic regions ("panels") selected so that a predetermined proportion of subjects with cancer exhibits genetic variants or tumor markers in one or more different genes or genomic regions in the panel. The panel may be selected to limit the sequencing region to a fixed number of base pairs. The panel may be selected to sequence a desired amount of DNA. The panel may also be selected to achieve a desired depth of sequence reads. A panel can be selected to achieve a desired sequence read depth or sequence read coverage in terms of the amount of sequenced base pairs. A panel can be selected to achieve a theoretical sensitivity, specificity, and / or accuracy for detecting one or more genetic variants in a sample.
[0418] Examples of lists of genomic locations of interest can be found, for example, in Tables 3 and 4 herein. In some embodiments, the set of sequence variable target regions includes a portion of at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, or 70 of the genes in Table 3. In some embodiments, the set of sequence variable target regions includes a portion of at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, or 70 of the SNVs in Table 3. In some embodiments, the set of sequence variable target regions includes a portion of at least 1, at least 2, at least 3, at least 4, at least 5, or 6 of the fusions in Table 3. In some embodiments, the set of sequence variable target regions comprises at least a portion of at least one, at least two, or three indels from Table 3. In some embodiments, the set of sequence variable target regions comprises a portion of at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, or 73 of the genes from Table 4. In some embodiments, the set of sequence variable target regions comprises at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, or 73 of the SNVs from Table 4. In some embodiments, the set of sequence variable target regions includes at least 1, at least 2, at least 3, at least 4, at least 5, or some of 6 of the fusions in Table 4.In some embodiments, the set of sequence variable target regions comprises at least a portion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, or 18 of the indels in Table 4. Each of these genomic locations of interest can be identified as a scaffold region or a hotspot region for a given panel. Table 5 shows an example list of hotspot genomic locations of interest. In some embodiments, the set of sequence variable target regions includes a portion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, or at least 20 of the genes in Table 5. Each hotspot genomic region is described along with the associated gene, the chromosome on which it resides, the genomic start and stop positions representing the locus, the base pair length of the locus, the exons covered by the gene, and important features of the given genomic region of interest (e.g., type of mutation). [Table 3] [Table 4-1] [Table 4-2] [Table 5-1] [Table 5-2] [Table 5-3]
[0419] Examples of lists of target regions of interest can also be found in WO2020 / 160414, e.g., Table 4. Additional examples include the loci disclosed in Gale et al., PLoS One 13: e0194630 (2018), incorporated herein by reference, which describes a panel of 35 cancer-associated gene targets: AKT1, ALK, BRAF, CCND1, CDK2A, CTNNB1, EGFR, ERBB2, ESR1, FGFR1, FGFR2, FGFR3, FOXL2, GATA3, GNA11, GNAQ, GNAS, HRAS, IDH1, IDH2, KIT, KRAS, MED12, MET, MYC, NFE2L2, NRAS, PDGFRA, PIK3CA, PPP2R1A, PTEN, RET, STK11, TP53, and U2AF1. In some embodiments, the set of sequence variable target regions includes target regions from at least 10, 20, 30, or 35 genes associated with cancer, such as those described herein and in WO2020 / 160414.
[0420] In some embodiments, the set of sequence variable target regions has a footprint of at least 50 kbp, e.g., at least 100 kbp, at least 200 kbp, at least 300 kbp, or at least 400 kbp. In some embodiments, the set of sequence variable target regions has a footprint in the range of 100 to 2,000 kbp, e.g., 100 to 200 kbp, 200 to 300 kbp, 300 to 400 kbp, 400 to 500 kbp, 500 to 600 kbp, 600 to 700 kbp, 700 to 800 kbp, 800 to 900 kbp, 900 to 1,000 kbp, 1 to 1.5 Mbp, or 1.5 to 2 Mbp. In some embodiments, the set of sequence variable target regions has a footprint of at least 2 Mbp. B. Collection of target-specific probes
[0421] In some embodiments, a collection of target-specific probes is used in the methods described herein. In some embodiments, the collection of target-specific probes includes target-binding probes specific for a set of sequence-variable target regions and target-binding probes specific for a set of epigenetic target regions. In some embodiments, the capture yield of the target-binding probes specific for the set of sequence-variable target regions is higher (e.g., at least two-fold higher) than the capture yield of the target-binding probes specific for the set of epigenetic target regions. In some embodiments, the collection of target-specific probes is configured to have a capture yield specific for the set of sequence-variable target regions that is higher (e.g., at least two-fold higher) than its capture yield specific for the set of epigenetic target regions.
[0422] In some embodiments, the capture yield of target binding probes specific for the set of sequence variable target regions is at least 1.25x, 1.5x, 1.75x, 2x, 2.25x, 2.5x, 2.75x, 3x, 3.5x, 4x, 4.5x, 5x, 6x, 7x, 8x, 9x, 10x, 11x, 12x, 13x, 14x, or 15x higher than the capture yield of target binding probes specific for the set of epigenetic target regions. In some embodiments, the capture yield of target binding probes specific for the set of sequence variable target regions is 1.25-1.5 fold, 1.5-1.75 fold, 1.75-2 fold, 2-2.25 fold, 2.25-2.5 fold, 2.5-2.75 fold, 2.75-3 fold, 3-3.5 fold, 3.5-4 fold, 4-4.5 fold, 4.5-5 fold, 5-5.5 fold, 5.5-6 fold, 6-7 fold, 7-8 fold, 8-9 fold, 9-10 fold, 10-11 fold, 11-12 fold, 13-14 fold, or 14-15 fold higher than the capture yield of target binding probes specific for the set of epigenetic target regions.
[0423] In some embodiments, the collection of target-specific probes is configured to have a capture yield specific for a set of sequence variable target regions that is at least 1.25x, 1.5x, 1.75x, 2x, 2.25x, 2.5x, 2.75x, 3x, 3.5x, 4x, 4.5x, 5x, 6x, 7x, 8x, 9x, 10x, 11x, 12x, 13x, 14x, or 15x higher than the capture yield of the set of epigenetic target regions. In some embodiments, the collection of target-specific probes is configured to have a capture yield specific for a set of sequence variable target regions that is 1.25-1.5 times, 1.5-1.75 times, 1.75-2 times, 2-2.25 times, 2.25-2.5 times, 2.5-2.75 times, 2.75-3 times, 3-3.5 times, 3.5-4 times, 4-4.5 times, 4.5-5 times, 5-5.5 times, 5.5-6 times, 6-7 times, 7-8 times, 8-9 times, 9-10 times, 10-11 times, 11-12 times, 13-14 times, or 14-15 times higher than its capture yield specific for the set of epigenetic target regions.
[0424] A collection of probes can be configured to provide higher capture yields for a set of sequence-variable target regions in a variety of ways, including enrichment, varying length and / or chemistry (e.g., to affect affinity), and combinations thereof. Affinity can be modulated by adjusting the length of the probe and / or by including nucleotide modifications, as discussed below.
[0425] In some embodiments, the target-specific probes specific for the set of sequence variable target regions are present at a higher concentration than the target-specific probes specific for the set of epigenetic target regions, hi some embodiments, the concentration of target-binding probes specific for the set of sequence variable target regions is at least 1.25-fold, 1.5-fold, 1.75-fold, 2-fold, 2.25-fold, 2.5-fold, 2.75-fold, 3-fold, 3.5-fold, 4-fold, 4.5-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 11-fold, 12-fold, 13-fold, 14-fold, or 15-fold higher than the concentration of target-binding probes specific for the set of epigenetic target regions. In some embodiments, the concentration of target-binding probes specific for the set of sequence-variable target regions is 1.25-1.5x, 1.5-1.75x, 1.75-2x, 2-2.25x, 2.25-2.5x, 2.5-2.75x, 2.75-3x, 3-3.5x, 3.5-4x, 4-4.5x, 4.5-5x, 5-5.5x, 5.5-6x, 6-7x, 7-8x, 8-9x, 9-10x, 10-11x, 11-12x, 13-14x, or 14-15x higher than the concentration of target-binding probes specific for the set of epigenetic target regions. In such embodiments, concentration may refer to the average mass / volume concentration of the individual probes in each set.
[0426] In some embodiments, target-specific probes specific to a set of sequence-variable target regions have higher affinity for the target than target-specific probes specific to a set of epigenetic target regions. Affinity can be modulated by any method known to those skilled in the art, including using different probe chemistries. For example, certain nucleotide modifications, such as cytosine 5-methylation (in the context of a particular sequence), modifications that introduce heteroatoms into the 2' sugar position, and LNA nucleotides, can increase the stability of double-stranded nucleic acids, and oligonucleotides with such modifications have been shown to have relatively high affinity for their complementary sequences. See, for example, Severin et al., Nucleic Acids Res. 39: 8740-8751 (2011); Freier et al., Nucleic Acids Res. 25: 4429-4443 (1997); U.S. Patent No. 9,738,894. Furthermore, longer sequence lengths generally provide increased affinity. Other nucleotide modifications, such as substitution of guanine with the nucleobase hypoxanthine, reduce affinity by reducing the amount of hydrogen bonding between the oligonucleotide and its complementary sequence. In some embodiments, target-specific probes specific for a set of sequence-variable target regions have modifications that increase affinity for their targets. In some embodiments, target-specific probes specific for a set of epigenetic target regions have modifications that decrease affinity for their targets. In some embodiments, target-specific probes specific for a set of sequence-variable target regions have a longer average length and / or a higher average melting temperature than target-specific probes specific for a set of epigenetic target regions. These embodiments can be combined with each other and / or with concentration differences as discussed above to achieve a desired fold difference in capture yield, for example, any of the fold differences or ranges described above.
[0427] In some embodiments, the target-specific probe comprises a capture moiety. The capture moiety can be any of the capture moieties described herein, for example, biotin. In some embodiments, the target-specific probe is linked to a solid support, for example, covalently or non-covalently, such as by interaction of the capture moiety's binding pair. In some embodiments, the solid support is a bead, for example, a magnetic bead.
[0428] In some embodiments, the target-specific probes specific for the set of sequence-variable target regions and / or the target-specific probes specific for the set of epigenetic target regions are probes that comprise capture moieties and sequences selected to span a panel of regions, such as the bait set discussed above, e.g., genes.
[0429] In some embodiments, the target-specific probes are provided in a single composition. The single composition may be in solution (liquid or frozen). Alternatively, the composition may be lyophilized.
[0430] Alternatively, target-specific probes can be provided as multiple compositions, including, for example, a first composition containing probes specific to a set of epigenetic target regions and a second composition containing probes specific to a set of sequence-variable target regions. These probes can be mixed in appropriate ratios to provide a combined probe composition having any of the above-mentioned fold differences in concentration and / or capture yield. Alternatively, these probes can be used in separate capture procedures (e.g., on aliquots of a sample or sequentially on the same sample) to provide first and second compositions containing captured epigenetic and sequence-variable target regions, respectively. 1. Probes specific to epigenetic target regions
[0431] The probes for the set of epigenetic target regions may include probes specific for one or more types of target regions that may distinguish DNA from neoplastic (e.g., tumor or cancer) cells from healthy cells, e.g., non-neoplastic circulating cells. Exemplary types of such regions are discussed in detail h...
Claims
1. 1. A method for analyzing DNA, comprising: (a) performing methylation-preserving amplification of the DNA, wherein the DNA comprises a barcode; and (b) sequencing the DNA in a manner sensitive to alterations and determining an epigenetic consensus sequence of the DNA associated with at least a portion of the barcode. A method comprising:
2. 1. A method for analyzing DNA, comprising: (a) performing methylation-preserving amplification of the DNA, wherein the DNA comprises a barcode; (b) subjecting the DNA to a procedure that affects a first nucleobase of the DNA differently from a second nucleobase of the DNA, wherein the first nucleobase is a modified or unmodified nucleobase, the second nucleobase is a modified or unmodified nucleobase that is different from the first nucleobase, and the first nucleobase and the second nucleobase have the same base-pairing specificity; and (c) sequencing the DNA and determining an epigenetic consensus sequence of the DNA associated with at least a portion of the barcode. A method comprising:
3. 1. A method for analyzing DNA, comprising: (a) performing methylation-preserving amplification of the DNA, wherein the DNA comprises a barcode; (b) enriching said DNA for one or more sets of epigenetic target regions of DNA, thereby providing enriched DNA; (c) sequencing the enriched DNA in a manner sensitive to the modification and determining an epigenetic consensus sequence of the enriched DNA associated with at least a portion of the barcode. A method comprising:
4. 1. A method for analyzing DNA, comprising: (a) performing linear methylation-preserving amplification of the DNA; and (b) sequencing the DNA in a manner sensitive to the modification and determining the epigenetic consensus sequence of the DNA. A method comprising:
5. 1. A method for analyzing DNA, comprising: (a) performing a linear methylation-preserving amplification of said DNA; (b) subjecting the DNA to a procedure that affects a first nucleobase of the DNA differently from a second nucleobase of the DNA, wherein the first nucleobase is a modified or unmodified nucleobase, the second nucleobase is a modified or unmodified nucleobase that is different from the first nucleobase, and the first nucleobase and the second nucleobase have the same base-pairing specificity; and (c) sequencing the DNA and determining an epigenetic consensus sequence of the DNA. A method comprising:
6. 1. A method for analyzing DNA, comprising: (a) performing a linear methylation-preserving amplification of said DNA; (b) enriching said DNA for one or more sets of epigenetic target regions of DNA, thereby providing enriched DNA; (c) sequencing the enriched DNA in a manner sensitive to the modification and determining the epigenetic consensus sequence of the enriched DNA. A method comprising:
7. 10. The method of any one of Claims 1, 3, 4, or 6, wherein said sequencing in a modification-sensitive manner comprises subjecting the DNA to a procedure that affects a first nucleobase of the DNA differently from a second nucleobase of the DNA, wherein the first nucleobase is a modified or unmodified nucleobase, the second nucleobase is a modified or unmodified nucleobase that is different from the first nucleobase, and the first nucleobase and the second nucleobase have the same base-pairing specificity.
8. 10. The method of any one of the preceding claims, wherein the DNA comprises a barcode.
9. 10. The method of any one of the preceding claims, wherein the method comprises ligating an adapter comprising a barcode to the DNA prior to the sequencing.
10. 10. The method of any one of the preceding claims, wherein the method comprises ligating an adapter comprising a barcode to the DNA prior to amplifying the DNA.
11. 10. The method of any one of the preceding claims, wherein the method comprises ligating an adapter comprising a barcode to the DNA prior to performing the methylation-preserving amplification of the DNA.
12. 1. A method for analyzing DNA, comprising: (a) performing methylation-preserving amplification of the DNA, thereby providing amplified DNA, wherein the DNA comprises an insert and an adapter comprising a barcode, at least one of the adapters further comprising a restriction enzyme cleavage site between the barcode and a portion of the adapter, and the barcode is located between the insert and the restriction enzyme cleavage site; (b) contacting the amplified DNA with a restriction enzyme that recognizes and cleaves the DNA at the restriction enzyme cleavage site in the adapter; (c) before or after step (b), subjecting the amplified DNA to a procedure that affects a first nucleobase of the amplified DNA to be different from a second nucleobase of the amplified DNA, wherein the first nucleobase is a modified or unmodified nucleobase, the second nucleobase is a modified or unmodified nucleobase that is different from the first nucleobase, and the first nucleobase and the second nucleobase have the same base-pairing specificity; (d) after steps (b) and (c), ligating auxiliary adapters to the amplified DNA; (e) after step (d), performing uracil and / or dihydrouracil resistant amplification of said DNA; (f) after step (e), enriching the amplified DNA for one or more sets of epigenetic target regions of DNA, thereby providing enriched DNA; (g) optionally further amplifying the enriched DNA; and (h) sequencing the enriched DNA and determining an epigenetic consensus sequence of the DNA associated with at least a portion of the barcode. A method comprising:
13. 13. The method of claim 12, wherein the uracil- and / or dihydrouracil-resistant amplification of the DNA comprises PCR using a uracil- and / or dihydrouracil-resistant DNA polymerase.
14. 14. The method of claim 12 or claim 13, wherein the optional step (step (g)) of amplifying the enriched DNA is performed.
15. 15. The method of claim 14, wherein said step of amplifying said enriched DNA further comprises the step of differentially tagging said enriched DNA.
16. 16. The method of claim 15, wherein differentially tagging the enriched DNA comprises attaching one or more sample indexes to the DNA.
17. 17. The method of any one of claims 1 to 3 or 7 to 16, wherein the barcode does not contain cytosines in a non-CpG context.
18. 18. The method of any one of claims 1 to 10 or 17, wherein the DNA comprises an insert and an adaptor comprising a barcode, at least one of the adaptors further comprising a restriction enzyme cleavage site between the barcode and a portion of the adaptor, and the barcode is located between the insert and the restriction enzyme cleavage site.
19. 20. The method of claim 18, further comprising contacting the amplified DNA with a restriction enzyme that recognizes and cleaves the DNA at the restriction enzyme cleavage site in the adapter.
20. 20. The method of claim 19, further comprising the steps of: ligating an auxiliary adaptor to the amplified DNA after the step of: (a) contacting the amplified DNA with a restriction enzyme that recognizes and cleaves the DNA at the restriction enzyme cleavage site in the adaptor; and (b) subjecting the amplified DNA to a procedure that affects a first nucleobase of the amplified DNA to be different from a second nucleobase of the amplified DNA, wherein the first nucleobase is a modified or unmodified nucleobase, the second nucleobase is a modified or unmodified nucleobase that is different from the first nucleobase, and the first nucleobase and the second nucleobase have the same base-pairing specificity.
21. 21. The method of claim 20, further comprising performing uracil and / or dihydrouracil resistant amplification of the DNA.
22. 22. The method of claim 21, wherein the uracil- and / or dihydrouracil-resistant amplification of the DNA comprises PCR using a uracil- and / or dihydrouracil-resistant DNA polymerase.
23. 23. The method of any one of claims 11 to 17 or 20 to 22, wherein the auxiliary adapter does not comprise a barcode.
24. 24. The method of any one of claims 11 to 17 or 19 to 23, wherein cleavage of the DNA with the restriction enzyme results in overhangs.
25. 25. The method of claim 24, wherein the overhang is a single-base overhang.
26. 26. The method of claim 25, wherein the single-base overhang is a single-base 5'-overhang.
27. 27. The method of claim 25 or claim 26, wherein the single-base overhang is a T.
28. 27. The method of claim 25 or claim 26, wherein the single-base overhang is A.
29. 10. The method of any one of the preceding claims, wherein the methylation-preserving amplification comprises contacting the DNA with a methyltransferase.
30. 30. The method of claim 29, wherein the methyltransferase preferentially methylates hemimethylated CpG and / or hemimethylated CpHpG.
31. 31. The method of claim 29 or claim 30, wherein the methyltransferase is DNMT1.
32. 10. The method of any one of the preceding claims, wherein the methylation-preserving amplification comprises one or more of polymerase chain reaction, linear amplification, rolling circle amplification, ligase chain reaction, strand displacement amplification, nucleic acid sequence-based amplification, and sequence-based self-sustained replication.
33. 10. The method of any one of the preceding claims, wherein the methylation-preserving amplification comprises thermocycling amplification.
34. 10. The method of any one of the preceding claims, wherein the methylation-preserving amplification comprises isothermal amplification.
35. 35. The method of claim 34, wherein the isothermal amplification comprises recombinant polymerase amplification (RPA), helix-dependent amplification (HDA), loop-mediated isothermal amplification (LAMP), or rolling circle amplification (RCA).
36. 10. The method of any one of the preceding claims, wherein the methylation-preserving amplification comprises linear amplification with thermal cycling.
37. 37. The method of any one of claims 2, 5, or 7-36, wherein the first nucleobase is an unmodified cytosine and the second nucleobase is a modified cytosine, optionally wherein the modified cytosine is 5-methylcytosine or 5-hydroxymethylcytosine.
38. said step of affecting a first nucleobase of said DNA differently from a second nucleobase of said DNA comprising: before said enrichment or after said enrichment; and Before sequencing 38. The method of any one of claims 2, 5, or 7-37, wherein the method is carried out in a
39. 39. The method of any one of Claims 2, 5, or 7-38, wherein said procedure affecting a first nucleobase of said DNA differently from a second nucleobase of said DNA chemically converts said first or second nucleobase such that the base-pairing specificity of the converted nucleobase is altered.
40. 40. The method of any one of claims 2, 5, or 7-39, wherein said procedure that affects a first nucleobase of said DNA differently from a second nucleobase of said DNA is a methylation-sensitive conversion.
41. 41. The method of claim 40, wherein the methylation-sensitive conversion is bisulfite conversion, oxidative bisulfite (Ox-BS) conversion, Tet-assisted bisulfite (TAB) conversion, APOBEC-linked epigenetic (ACE) conversion, or enzymatic conversion.
42. 42. The method of claim 41, wherein the Tet-assisted conversion further comprises a substituted borane reducing agent, optionally wherein the substituted borane reducing agent is 2-picoline borane, borane pyridine, tert-butylamine borane, or ammonia borane.
43. 10. The method of any one of the preceding claims, wherein the method further comprises distributing at least a portion of the DNA into a plurality of sub-samples comprising a first sub-sample and a second sub-sample, wherein the first sub-sample comprises a higher proportion of DNA with cytosine modifications than the second sub-sample.
44. 44. The method of Claim 43, wherein said distributing at least a portion of said DNA into a plurality of sub-samples comprises contacting said DNA with an agent that recognizes modified cytosines in said DNA, said first sub-sample comprising DNA with a higher proportion of said modified cytosines than said second sub-sample.
45. the distributing step is performed before the sequencing; and i. before performing said methylation-preserving amplification of said DNA; ii. after performing the methylation-preserving amplification of the DNA, iii. before enriching said DNA for one or more sets of epigenetic target regions of DNA; and / or iv. After said enrichment for one or more sets of epigenetic target regions of DNA from said DNA, 45. The method of claim 43 or claim 44, wherein the method is performed on
46. 46. The method of claim 44 or claim 45, wherein the agent that recognizes a modified nucleobase in the DNA is a methyl-binding reagent.
47. 47. The method of claim 46, wherein the methyl-binding reagent is a methyl-binding domain (MBD) protein or antibody.
48. 48. The method of claim 46 or claim 47, wherein the methyl-binding reagent is specific for one or more methylated nucleotide bases, and optionally the one or more methylated nucleotide bases are 5-methylcytosine.
49. 49. The method of any one of claims 46 to 48, wherein the methyl-binding reagent is immobilized on a solid support.
50. 50. The method of any one of claims 46 to 49, wherein the partitioning step comprises immunoprecipitation of methylated DNA.
51. 51. The method of any one of claims 46 to 50, wherein the partitioning step comprises partitioning based on binding to a protein, optionally the protein is a methylated protein, an acetylated protein, an unmethylated protein, an unacetylated protein; and / or optionally the protein is a histone.
52. 52. The method of any one of claims 46 to 51, wherein the partitioning step comprises contacting the DNA with a binding reagent specific for the protein and immobilized on a solid support.
53. 53. The method of any one of claims 46 to 52, wherein a first dispensed sub-sample of the plurality of dispensed sub-samples is differentially tagged from a second dispensed sub-sample of the plurality of dispensed sub-samples.
54. 10. The method of any one of the preceding claims, comprising contacting the DNA or at least one sub-sample thereof with at least one nuclease prior to said enriching or said sequencing, optionally wherein said at least one nuclease is at least one restriction enzyme.
55. 55. The method of Claim 54, wherein said contacting said DNA or at least one sub-sample thereof with at least one nuclease occurs after partitioning said sample into said plurality of sub-samples or before performing a procedure that affects a first nucleobase of said DNA differently from a second nucleobase of said DNA.
56. 56. The method of claim 54 or claim 55, wherein the at least one restriction enzyme is a methylation-sensitive restriction enzyme (MSRE).
57. 56. The method of claim 54 or claim 55, wherein the at least one restriction enzyme is a methylation-dependent restriction enzyme (MDRE).
58. 57. The method of any one of claims 54 to 56, wherein the DNA or at least one sub-sample thereof is contacted with the at least one methylation-sensitive restriction enzyme, thereby generating hypermethylated DNA.
59. 58. The method of any one of claims 54, 55, or 57, wherein the DNA or at least one sub-sample thereof is contacted with the at least one methylation-dependent restriction enzyme, thereby generating hypomethylated DNA.
60. 60. The method of any one of claims 54 to 56, 58, or 59, wherein the first sub-sample is contacted with the MSRE.
61. 61. The method of any one of claims 54, 55, or 57-60, wherein the second sub-sample is contacted with the MDRE.
62. 62. The method of any one of claims 54 to 61, comprising contacting at least one sub-sample with at least two restriction enzymes prior to said enriching or prior to sequencing, optionally wherein said contacting occurs before performing said procedure that affects a first nucleobase of said DNA differently from a second nucleobase of said DNA.
63. 63. The method of claim 62, wherein the at least two restriction enzymes comprise or consist of two or three restriction enzymes.
64. 64. The method of any one of claims 54 to 63, wherein the at least one restriction enzyme is selected from the group consisting of FspEI, LpnPI, MspJI, SgeI, AatII, AccII, AciI, Aor13HI, Aor15HI, BspT104I, BssHII, BstUI, Cfr10I, ClaI, CpoI, Eco52I, HaeII, HapII, HhaI, Hin6I, HpaII, HpyCH4IV, MluI, MspI, NaeI, NotI, NruI, NsbI, PmaCI, Psp1406I, PvuI, SacII, SalI, SmaI, and SnaBI.
65. 65. The method of any one of claims 54 to 64, further comprising the step of attaching one or more adaptors to at least one end of at least some of the DNA molecules in said plurality of distributed sets prior to digestion.
66. 66. The method of Claim 65, wherein the one or more adaptors comprise at least one tag.
67. 67. The method of Claim 66, wherein the at least one tag comprises a molecular barcode.
68. 68. The method of any one of claims 65 to 67, wherein the one or more adaptors are resistant to digestion by a methylation-sensitive or methylation-dependent restriction enzyme.
69. the one or more adapters that are resistant to digestion by a methylation-sensitive restriction enzyme a) one or more methylated nucleotides, optionally wherein said methylated nucleotides comprise 5-methylcytosine and / or 5-hydroxymethylcytosine; b) one or more nucleotide analogs that are resistant to methylation-sensitive restriction enzymes; or c) a nucleotide sequence that is not recognized by a methylation-sensitive restriction enzyme 69. The method of claim 68, comprising:
70. 70. The method of any one of claims 43 to 69, wherein the sub-samples are pooled prior to said sequencing.
71. 71. The method of any one of claims 1, 2, 4, 5, 7-10, or 17-70, further comprising enriching the DNA for one or more sets of epigenetic target regions of DNA, thereby providing enriched DNA.
72. 10. The method of any one of the preceding claims, further comprising enriching said DNA for one or more sets of sequence variable target regions of DNA.
73. 10. The method of any one of the preceding claims, wherein the enriching step comprises contacting the DNA with target-specific probes specific for the one or more sets of epigenetic target regions and / or the one or more sets of sequence variable target regions.
74. 74. The method of any one of claims 3 or 6 to 73, wherein the set of epigenetic target regions comprises a set of hypermethylated variable target regions and / or a set of hypomethylated variable target regions.
75. 75. The method of any one of claims 3 or 6-74, wherein the set of epigenetic target regions comprises a set of fragmented variable target regions.
76. 76. The method of claim 75, wherein the set of fragmented variable target regions comprises a transcription start site region.
77. 77. The method of claim 75 or claim 76, wherein the set of fragmented variable target regions comprises a CTCF binding region.
78. 78. The method of any one of claims 3 or 6-77, wherein the set of epigenetic target regions comprises one or more type-specific epigenetic target regions.
79. 79. The method of Claim 78, wherein said one or more type-specific epigenetic target regions comprise type-specific differentially methylated regions and / or type-specific fragments.
80. 79. The method of claim 78, wherein the one or more type-specific epigenetic target regions comprise type-specific hypomethylated regions and / or type-specific hypermethylated regions.
81. 81. The method of any one of claims 78-80, wherein the one or more type-specific epigenetic target regions comprise cell type-specific, cell cluster type-specific, tissue type-specific, and / or cancer type-specific epigenetic target regions.
82. the one or more type-specific epigenetic target regions: a) are hypermethylated in immune cells compared to non-immune cell types present in blood samples; b) differentially methylated in the colon compared to other tissue types; c) differentially methylated in breast compared to other tissue types; d) differentially methylated in the liver compared to other tissue types; e) differentially methylated in kidney compared to other tissue types; f) differentially methylated in the pancreas compared to other tissue types; g) differentially methylated in the prostate compared to other tissue types; h) is differentially methylated in skin compared to other tissue types; or i) Differentially methylated in the bladder compared to other tissue types 82. The method of any one of claims 78 to 81, comprising a target region.
83. 83. The method of any one of claims 78-82, wherein the hypermethylated target region is methylated to an extent that is at least 10%, 20%, 30%, or at least 40% greater than the average methylation of the target region in the sample or compared to other cell or tissue types.
84. the one or more type-specific epigenetic target regions: a) a target region that is hypomethylated in non-immune cell types present in the sample compared to the methylation level of the target region in different cell or tissue types in the sample; b) a fragment that is specific for immune cells relative to non-immune cell types present in said sample; or c) fragments specific to colon, lung, breast, liver, kidney, pancreas, prostate, skin, or bladder compared to other tissue types 84. The method of any one of claims 78 to 83, comprising:
85. 85. The method of any one of claims 78 to 84, wherein the level of the one or more type-specific epigenetic target regions is determined based on cell type or tissue type of origin.
86. 86. The method of claims 78-85, wherein said level of said one or more type-specific epigenetic target regions originating from one or more immune cells, non-immune cell types present in a blood sample, and / or colon, lung, breast, liver, kidney, prostate, skin, bladder, or pancreatic cells is determined.
87. 87. The method of any one of claims 78-86, further comprising identifying at least one cell type, cell cluster type, tissue type, and / or cancer type from which said one or more type-specific epigenetic target regions originated.
88. 88. The method of any one of claims 78 to 87, comprising determining the methylation level of the type-specific epigenetic target region.
89. 89. The method of any one of claims 43-88, wherein at least a portion of the DNA from the first sub-sample and at least a portion of the DNA from the second sub-sample are pooled, thereby providing a combined sub-sample.
90. 90. The method of any one of claims 43 to 89, wherein the DNA of the first sub-sample and the DNA of the second sub-sample are differentially tagged.
91. 10. The method of claim 1, wherein the pool comprises less than or equal to about 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, or 5% of the DNA of the second sub-sample.
92. 92. The method of claim 90 or claim 91, wherein the pool comprises about 70-90%, about 75-85%, or about 80% of the DNA of the second sub-sample.
93. 93. The method of any one of claims 90 to 92, wherein the pool comprises substantially all of the DNA of the first sub-sample.
94. 93. The method of any one of claims 90 to 92, wherein the pool comprises substantially all of the DNA of the first sub-sample or processed first sub-sample.
95. 95. The method of any one of claims 90 to 94, wherein the first set of target regions is captured from at least a portion of the first sub-sample after forming the pool.
96. 96. The method of any one of claims 90 to 95, wherein at least a portion of the DNA from the first sub-sample and at least a portion of the DNA from the second sub-sample are sequenced in the same sequencing cell.
97. 97. The method of any one of claims 90-96, wherein the plurality of sub-samples comprises a third sub-sample comprising a higher proportion of DNA with epigenetic modifications than the second sub-sample but a lower proportion than the first sub-sample.
98. 10. The method of claim 1, wherein the method further comprises the step of differentially tagging the third sub-sample.
99. 99. The method of claim 97 or claim 98, wherein the DNA from the first sub-sample, the DNA from a third sample, and the set of target regions are pooled, and optionally the DNA from the first, second, and third sub-samples is sequenced in the same sequencing cell.
100. 100. The method of any one of claims 89 to 99, wherein said sequencing the DNA of the combined sub-samples comprises sequencing the DNA in a manner that is sensitive to modification.
101. 97. The method of any one of claims 3, 6-10, or 17-96, wherein the DNA is amplified after the enriching step.
102. 102. The method of Claim 101, wherein said amplifying step further comprises differentially tagging said enriched DNA.
103. 103. The method of Claim 102, wherein differentially tagging the enriched DNA comprises attaching one or more sample indexes to the DNA.
104. 10. The method of any one of the preceding claims, wherein said sequencing said DNA comprises sequencing said DNA in a manner that distinguishes said first nucleobase from said second nucleobase.
105. 105. The method of any one of claims 1, 3, 4, or 6-104, wherein said sequencing in a manner sensitive to an alteration comprises long-read sequencing.
106. 106. The method of any one of claims 1, 3, 4, or 6-105, wherein said sequencing in a manner sensitive to an alteration comprises nanopore sequencing.
107. 105. The method of any one of claims 1, 3, 4, or 6-104, wherein said sequencing in a manner sensitive to modification comprises five-letter or six-letter sequencing.
108. 10. The method of any one of the preceding claims, wherein the sequencing step comprises next generation sequencing.
109. 10. The method of any one of the preceding claims, wherein the sequencing step comprises generating a plurality of sequencing reads, and the method further comprises mapping the plurality of sequence reads to one or more reference sequences to generate mapped sequence reads, and processing the mapped sequence reads corresponding to the set of sequence variable target regions and the set of epigenetic target regions.
110. 110. The method of any one of claims 4 to 109, further comprising determining an epigenetic consensus sequence of the DNA associated with at least a portion of the barcode.
111. determining an epigenetic consensus sequence of the DNA associated with at least a portion of the barcodes by comparing a sequence of at least a portion of the associated reads to (a) a unique barcode or a unique set of barcodes, (b) a unique genomic start and / or stop position, or (c) both (a) and (b), wherein the reads originate from a DNA molecule; and determining the consensus epigenetic state of each nucleotide of said DNA molecule 111. The method of any one of claims 1 to 3, 7 to 10, or 17 to 110, comprising:
112. 112. The method of any one of claims 1-3, 7-10, or 17-111, further comprising determining a consensus base sequence of the DNA associated with at least a portion of the barcode.
113. determining a consensus base sequence of the DNA associated with at least a portion of the barcode by comparing the sequence of at least a portion of the associated read to (a) a unique barcode or a unique set of barcodes, (b) a unique genomic start and / or stop position, or (c) both (a) and (b), wherein the read originates from a DNA molecule; and determining the identity of the consensus base for each nucleotide of said DNA molecule 113. The method of claim 112, comprising:
114. 10. The method of any one of the preceding claims, wherein the DNA is from a blood sample and / or a tissue sample.
115. 115. The method of claim 114, wherein the blood sample is a whole blood sample, a plasma sample, a buffy coat sample, a leukocyte-reduced sample, or a PBMC sample.
116. 114. The method of any one of claims 1 to 113, wherein the DNA is cell-free DNA.
117. 117. The method of claim 116, wherein the cell-free DNA is in an amount of 1 ng to 500 ng.
118. 10. The method of any one of the preceding claims, wherein the DNA and / or the sample is from a subject.
119. 119. The method of claim 118, wherein the subject is an animal.
120. The method of claim 118 or claim 119, wherein the subject is a human.
121. 121. The method of any one of claims 94 to 120, wherein the blood sample is fractionated prior to enrichment for at least one set of epigenetic target regions of DNA.
122. 122. The method of any one of claims 118-121, wherein the subject has or is at risk of having cancer.
123. 123. The method of any one of claims 118 to 122, further comprising determining the presence or status of cancer in the subject.
124. 124. The method of any one of claims 118 to 123, further comprising determining the likelihood that the subject has an infection.
125. 125. The method of any one of claims 118 to 124, further comprising determining the likelihood that the subject will have transplant rejection.