Methods for double stranded DNA sequencing

The method of attaching hairpin adaptors to dsDNA for sequencing both strands addresses the limitations of conventional sequencing, enabling accurate and quantitative analysis of cytosine variants and epigenetic modifications, enhancing error correction and mutation detection.

WO2025257103A1PCT designated stage Publication Date: 2025-12-18BIOMODAL LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/065973
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-10
Filing Date
2025-06-09
Publication Date
2025-12-18

AI Technical Summary

Technical Problem

Conventional sequencing approaches are limited in their ability to accurately and quantitatively sequence both strands of double-stranded DNA, particularly in distinguishing cytosine variants like C, mC, and hmC, which hampers the understanding of epigenetic regulation and the detection of low-frequency genetic mutations.

Method used

A method involving attaching hairpin adaptors to dsDNA, introducing a single strand break, extending nucleotide strands, and sequencing both original and copy strands to determine the sequences of the dsDNA molecule, allowing for the identification of cytosine variants and epigenetic modifications.

Benefits of technology

Enables accurate, quantitative sequencing of both strands of dsDNA, improving error correction, mutation detection, and epigenetic modification mapping, with high confidence in base calling and sensitivity to low-frequency mutations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025065973_18122025_PF_FP_ABST
    Figure EP2025065973_18122025_PF_FP_ABST
Patent Text Reader

Abstract

This invention relates for method of sequencing double stranded DNA (dsDNA) molecules. A first hairpin adaptor is attached to a first end of a dsDNA molecule and a second hairpin adaptor to a second end of the dsDNA molecule to produce a double hairpin construct. The first hairpin adaptor is cleaved to introduce a single strand break into the double hairpin construct and the 3' end of the single strand break is extended along the strand of the double hairpin construct to produce an extended hairpin construct that comprises first and second original strands and first and second copy strands. The first and second original strands correspond to the first and second strands of the dsDNA molecule and the first and second copy strands are complementary to the first and second original strands. A sequencing adaptor is attached to a free end of the extended hairpin construct. The first and second original and copy strands of the extended hairpin construct are then sequenced. The sequences of the first and second strands of the dsDNA molecule are then determined from the sequences of the first and second original strands and / or the first and second copy strands of the extended hairpin construct. Modified cytosines at CpG sites in the first and second strands of the dsDNA molecule and / or hemi-hydroxymethylated CpG sites, hemi-hydroxymethylated CpG sites may be identified. Methods and kits are provided.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] METHODS FOR DOUBLE STRANDED DNA SEQUENCING

[0002] Field

[0003] This invention relates to methods for the simultaneous sequencing both strands of double-stranded DNA molecules.

[0004] All higher eukaryotes started life as a single cell, with a single copy of genetic instructions. During development, these same instructions are used selectively to create every other cell in the body. This remarkable feat is achieved by epigenetic gene regulation — the specific, stable, and heritable control of gene expression. In mammals, cytosine methylation plays a critical role in epigenetic gene regulation, maintaining cell identity through its involvement in stable and heritable transcriptional repression.1While 5- methylcytosine (mC) deposited during embryogenesis can persist throughout one’s entire lifetime1, it can also be enzymatically converted to 5-hydroxymethylcytosine (hmC)2. This can lead to demethylation by either a replication-dependent (passive) or independent (active) pathway — where hmC is further oxidised to 5-formylcytosine (fC) or 5-carboxylcytosine (caC) before being excised and replaced with unmodified cytosine (C)3.

[0005] Beyond being a transient demethylation intermediate, hmC can itself be a stable mark45. Unlike its oxidised derivatives, hmC cannot be directly removed from the genome6, and across different cell types its abundance is orders of magnitude greater3. Hydroxymethylation at gene bodies positively correlates with transcription7-10and is associated with maintaining transcriptional fidelity in multiple cell types11 12. The mechanistic basis of these findings remains unclear, but hmC may play a causal role10-12. Notably, hmC’s genomic distribution is highly cell-type specific78, and its dysregulation is linked to diseases ranging from Alzheimer’s13to cancer6 14— where a global depletion of hmC has been observed in virtually every tumour type tested6. Consequently, hmC signatures in circulating cell-free DNA constitute a clinically promising biomarker14. Elucidating the distribution and role of hmC — and its complex interplay with mC — is thus essential to our understanding of fundamental biology and disease. mC and hmC occur most frequently at CpG sites15, which contain a cytosine residue on each DNA strand. A CpG site can therefore contain nine possible combinations of C, mC and hmC, which we call ‘CpG states’ (Fig. 1 a). Within a CpG site, the modification sites of the two cytosines are close (~7 A) and occupy the major groove (Fig. 1 a). Thus, each CpG state presents a unique structural signature for potential protein recognition. Indeed, crystal structures of many known mC and hmC reader proteins show that they directly interact with both modification sites (Fig. 1 b). However, limitations in conventional sequencing approaches have so far prevented the quantitative, simultaneous sequencing of C, mC, and hmC on both DNA strands. Thus, how these epigenetic cytosine states are distributed in the double-stranded epigenome remains poorly understood.

[0006] In addition to distinguishing between cytosine variants, improving the accuracy of base-calling during sequencing remains a central goal in the development of sequencing technologies, regardless of the sequencing platform used. Furthermore, identifying genetic mutations — notably those occurring at a low frequency — and differentiating these from errors arising from sample preparation or sequencing is critical for many areas of fundamental scientific research, as well as for clinical applications (e.g. cancer diagnostics and screening, advanced detection of cancer recurrence and prenatal diagnostics)57.

[0007] The present inventors have developed an accurate, quantitative method for simultaneously sequencing both strands of a double-stranded (ds) DNA molecule. This may be useful for example in allowing sequencingplatform independent error correction, identification of low-frequency genetic mutations, and / or the identification of cytosine variants.

[0008] A first aspect of the invention provides a method of sequencing a double stranded DNA (dsDNA) molecule comprising;

[0009] (i) providing a dsDNA molecule having a first and a second strand,

[0010] (ii) attaching a first hairpin adaptor to a first end and a second hairpin adaptor to a second end of the dsDNA molecule to produce a double hairpin construct, wherein the first hairpin adaptor comprises a cleavage site,

[0011] (iii) cleaving the first hairpin adaptor at the cleavage site to introduce a single strand break (SSB) into the double hairpin construct,

[0012] (iv) extending a nucleotide strand from the 3’ end of the SSB along the strand of the double hairpin construct to produce an extended hairpin construct that comprises first and second original strands and first and second copy strands, wherein the first and second original strands correspond to the first and second strands of the dsDNA molecule and the first and second copy strands are complementary to the first and second original strands,

[0013] (v) attaching a sequencing adaptor to a free end of the extended hairpin construct,

[0014] (vi) sequencing the first and second original and copy strands of the extended hairpin construct and

[0015] (vii) determining the sequences of the first and second strands of the dsDNA molecule from the sequences of the first and second original and copy strands original and copy strands of the extended hairpin construct.

[0016] In some embodiments, methods of the first aspect may comprise deaminating cytosines (C) in the first and second original and copy strands before sequencing.

[0017] In some embodiments, methods of the first aspect may be useful in sequencing, identifying or mapping epigenetically modified cytosines, such as 5-methylcytosine (mC), and 5-hydroxymethylcytosine (hmC). The methods may be suitable for use with sequencing platforms that natively detect cytosine modification or with sequencing platforms that do not natively detect cytosine modification. For example, a method of the first aspect may further comprise prior to step (vi);

[0018] (a) protecting 5-hydroxymethylcytosines in the first and second original strands of the extended hairpin construct, (b) methylating cytosines (C) in the second and / or first copy strands of the extended hairpin construct at CpG sites that comprise a 5-methylcytosine (mC) in the first and / or second original strands,

[0019] (c) protecting 5-methylcytosines (mC) in the first and second original and copy strands and

[0020] (d) deaminating cytosines (C) in the first and second original and copy strands

[0021] Deamination converts cytosines (C) to uracil (U) in the first and second original and copy strands. Protected 5-methylcytosines and 5-hydroxymethylcytosines in the first and second original and copy strands are resistant to deamination and are not converted to uracil. 5-hydroxymethylcytosines may be protected by glycosylation. 5-methylcytosines may be protected by oxidation and glycosylation.

[0022] In some embodiments, a method of the first aspect may further comprise;

[0023] (e) amplifying the first and second original and copy strands.

[0024] Amplification converts uracil (U) in the deaminated original and copy strands to thymine (T) i.e. T is present in the amplified original and strands at positions at which U is present in the deaminated original and copy strands.

[0025] The first and second original and copy strands of the extended hairpin construct may be sequenced following steps (a) to (d) and optionally (e), to produce a sequence read. The presence of U or T at a position in the sequence read is indicative of a C at that position in the sequence of the first and second original and copy strands of the extended hairpin construct. The presence of a C at a position in the sequence read is indicative of the presence of a modified cytosine (hmC or mC) at that position in the sequence of the first and second original and copy strands. The sequences of the first and second original and copy strands of the extended hairpin may be determined from the sequence reads.

[0026] The sequences of the first and second original and copy strands of the extended hairpin construct may be compared to identify cytosines and epigenetically modified cytosines in the first and second strands of the dsDNA molecule. For example, the two cytosines at a CpG site in the double stranded DNA molecule may each be identified as unmodified cytosine (C), 5-methylcytosine (mC) or 5-hydroxymethylcytosine (hmC) from the sequences of the sequences of the first and second original and copy strands of the extended hairpin construct. The combinations of nucleotides at a CpG site in the first and second original and copy strands that are indicative of the state of the CpG site in the dsDNA molecule (i.e. the combination of C, mC or hmC that are present at the site on the two strands of the dsDNA molecule) are shown in Table 1 .

[0027] For example, T / U at the CpG site in the first original strand sequence, C at the CpG site in the second original strand sequence, T / U at the CpG site in the second copy strand sequence, and T / U at the CpG site in the first copy strand sequence may be indicative of the presence of C in the first strand and hmC in the second strand at the CpG site in the double stranded DNA.

[0028] C at the CpG site in the first original strand sequence, T / U at the CpG site in the second original strand sequence, T / U at the CpG site in the second copy strand sequence, and T / U at the CpG site in the first copy strand sequence, may be indicative of the presence of hmC in the first strand and C in the second strand at the CpG site in the double stranded DNA.

[0029] T / U at the CpG site in the first original strand sequence, T / U at the CpG site in the second original strand sequence, T / U at the CpG site in the second copy strand sequence, and T / U at the CpG site in the first copy strand sequence may be indicative of the presence of C in the first strand and C in the second strand at the CpG site in the double stranded DNA.

[0030] T / U at the CpG site in the first original strand sequence, C at the CpG site in the second original strand sequence, C at the CpG site in the second copy strand sequence, and T / U at the CpG site in the first copy strand sequence may be indicative of the presence of C in the first strand and mC in the second strand at the CpG site in the double stranded DNA.

[0031] C at the CpG site in the first original strand sequence, T / U at the CpG site in the second original strand sequence, T / U at the CpG site in the second copy strand sequence, and C at the CpG site in the first copy strand sequence may be indicative of the presence of mC in the first strand and C in the second strand at the CpG site in the double stranded DNA.

[0032] C at the CpG site in the first original strand sequence, C at the CpG site in the second original strand sequence, C at the CpG site in the second copy strand sequence, and C at the CpG site in the first copy strand sequence may be indicative of the presence of mC in the first strand and mC in the second strand at the CpG site in the double stranded DNA.

[0033] C at the CpG site in the first original strand sequence, C at the CpG site in the second original strand sequence, T / U at the CpG site in the second copy strand sequence, and T / U at the CpG site in the first copy strand sequence may be indicative of the presence of hmC in the first strand and hmC in the second strand at the CpG site in the double stranded DNA.

[0034] C at the CpG site in the first original strand sequence, C at the CpG site in the second original strand sequence, T / U at the CpG site in the second copy strand sequence, and C at the CpG site in the first copy strand sequence may be indicative of the presence of mC in the first strand and hmC in the second strand at the CpG site in the double stranded DNA.

[0035] C at the CpG site in the first original strand sequence, C at the CpG site in the second original strand sequence, C at the CpG site in the second copy strand sequence, and T / U at the CpG site in the first copy strand sequence may be indicative of the presence of hmC in the first strand and mC in the second strand at the CpG site in the double stranded DNA.

[0036] Methods of the first aspect may be useful in sequencing a population of dsDNA molecules, for example a population of genomic DNA fragments. The present inventors have also recognised that transcriptionally active genes are characterised by CpG sites that include hmC (H) on one strand and C on the other (hemi-hydroxymethylated or CH / HC CpG sites). This finding allows for example, the identification and mapping of transcriptionally active genes within genomic DNA.

[0037] A second aspect of the invention provides a method of mapping transcriptionally active genes in genomic DNA comprising; providing a population of fragments of genomic DNA, and identifying hemi-hydroxymethylated CpG sites in the fragments, wherein hemi-hydroxymethylated CpG sites in the fragments are indicative of transcriptionally active genes in the genomic DNA.

[0038] A hemi-hydroxymethylated CpG site in a genomic DNA fragment may comprise hmC (H) on one strand of the genomic DNA fragment and C on the other strand i.e. a hemi-hydroxymethylated CpG site may be a HC or a CH CpG site). Suitable hemi-hydroxymethylated CpG sites may be identified in the exons or introns (i.e. the gene body) of a transcriptionally active gene.

[0039] Hemi-hydroxymethylated CpG sites may be identified in methods of the second aspect by sequencing the first and second strands of genomic DNA fragments and identifying CpG sites with a hmC on the first strand and a C on the second strand or a C on the first strand and a hmC on the second strand. The identified hemi-hydroxymethylated CpG sites may be mapped onto the genomic DNA to identify genes within the genomic DNA that are transcriptionally active. Suitable methods of sequencing the first and second strands of genomic DNA fragments include methods of the first aspect.

[0040] Other aspects and embodiments of the invention are described in more detail below.

[0041] Brief Description of the Figures

[0042] Figure 1 shows the quantitative sequencing of C, mC and hmC in double-stranded DNA. (a) CpG sites contain a cytosine residue on each DNA strand, resulting in nine possible combinations of C, mC, and hmC. Within an individual CpG site, both cytosine modification sites are close (~7 A). Each CpG state thus produces a unique structural signature in the major groove for protein recognition, and many readers of cytosine modifications interrogate both modification sites simultaneously, (b) Schematic of several known readers of modified cytosines that interact simultaneously with the modifications on both strands of a CpG site (A & B), forming close contacts (< 3.6 A).

[0043] Figure 2 shows a schematic of examples of methods with and without duet / lllumina™ sequencing, (a) Schematic of method coupled with duet / lllumina sequencing. 1. Ligation of combination of cleavable and non-cleavable hairpin adaptors; 2. Cleavage of cleavable hairpin adapter, generating self-priming construct; 3. Synthesis of the two copy strands (grey) from their corresponding original strands (black) and simultaneous glucosylation of hmC; 4. Ligation of Y-shaped sequencing adaptor; 5. TET-oxidation of mC and protection of resulting hmC by glucosylation; 6. Deamination of C with APOBEC3A and UvrD helicase, library amplification and sequencing; 7. Reconstruction of CpG states using custom bioinformatics pipeline. (b) All possible CpG states are encoded by C or mC calls (hollow and black circles, respectively), (c) Schematic of method for direct sequencing of DNA modifications alongside genetic bases. Steps 1-3 are the same as a, 4. Ligation of sequencing adapter, e.g. one compatible with Oxford Nanopore (shown) or PacBio (not shown) sequencers; 5. Sequencing (full or partial duplex denaturation of the hairpin construct occurs during sequencing). 6. Reconstruction of CpG states using direct readout of modified nucleotides in sequences A & B; sequences C & D, which are unmodified copies of A & B, provide an inbuilt negative control. Information from all four sequences (A, B, C, D) may be combined to generate a consensus sequence to reduce error rates associated with base-calling, (d) The nine CpG states comprising C, mC, and hmC are reconstructed following direct readout of the cytosine variants in the sequences derived from the two original strands (A & B).

[0044] Figure 3 shows that all nine CpG states are accurately and simultaneously resolved, (a) Synthetic spike-in containing known CpG states (double-stranded combinations of C, mC, and hmC at CpG sites) with its corresponding CpG-state readout, viewed in IGV. Each CpG state is correctly and unambiguously called, (b) Confusion matrix showing the rate at which each CpG state in the synthetic spike-ins was correctly called (yellow) versus the rate it was miscalled as each other state (green). All states are called with an accuracy exceeding 94%.

[0045] Figure 4 shows the global abundance and local distribution of CpG states, (a) Genome-wide abundance of the nine CpG states comprising C, mC, and hmC, whose distributions depend strongly on their doublestranded context. 18% of all CpG sites are asymmetrically modified. This includes 98% of all hydroxymethylated sites, where, in addition to hmC, the CpG site contains either C (denoted ‘HC’ and ‘CH’, with hmC on the plus and minus strand, respectively) or mC (‘HM’ and ‘MH’), (b) Illustrative example of the local distribution of CpG states in the bivalent Pax3 gene, which encodes a transcription factor critical during embryogenesis and whose aberrant expression is linked to several forms of cancer22. Data are scaled from 0-100%, except for hmC states which are scaled from 0-50% for clarity (viewed in IGV19).

[0046] Figure 5 shows the distribution of hmC states at active, primed, and poised enhancers, (a) Total levels of hmC across each enhancer type (crimson), alongside background hmC levels (grey, randomly selected genomic regions of the same size). Characteristic features of each enhancer type are also shown26. Levels are the mean percentage of hydroxymethylated CpG states for each enhancer. Primed enhancers contain high levels of hmC, whereas at active and poised enhancers hmC is somewhat depleted, (b) Distribution of hmC states throughout enhancers (200-bp bin size). Shaded regions correspond to the mean length of each enhancer type. Primed enhancers are characterised by a sharp increase in HC / CH and a modest increase in HM / MH, with levels highest at the centre of enhancers. Conversely, in active and poised enhancers levels of HC / CH and HM / MH are lowest at the centre, though HC / CH levels are elevated at regions directly flanking the enhancers, (c) Mean levels of hmC states across each enhancer. hmC at enhancers is overwhelmingly asymmetric and is almost entirely in the HC / CH context, except at primed enhancers, which also contain substantial but lower levels of HM / MH. Median levels of HC / CH at primed enhancers are more than double that at active and poised enhancers. Box-plots: centre line, median; box limits, upper and lower quartiles; whiskers, 1.5* interquartile range. Figure 6 shows that hmC in the HC / CH context dominates hmC’s relationship with transcription at gene bodies and marks transcriptionally active regions, (a) Spearman’s correlation coefficients between transcription and CpG-state levels at promoters, introns, and exons of protein-coding genes (P < 2.2 x 10-16 in all cases, n = 16,357, 15,564, and 15,265 genes, respectively), (b) Mean levels of hmC states across introns and exons of protein-coding genes, grouped by transcription level: silent (n = 4,271), low (n = 3,458), medium (n = 4,004), and high (n = 4,258). Intronic HC / CH levels positively correlate with transcription, with a magnitude close to promoter methylation (p = 0.38, P < 2.2 x 10-16 in all cases). The other hmC states, HM / MH and HH exhibit much weaker, positive correlations, (c) Distributions of HC / CH and HM / MH across protein-coding genes grouped by transcription level: silent (yellow, n = 6,306), low (light green, n = 3,778), medium (dark green, n = 4,276), and high (blue, n = 4,485). In silent genes, the highest levels of HC / CH occur at the TSS, where there is a sharp inversion even for lowly transcribed genes. For all active genes, HC / CH levels are considerably higher than HM / MH levels either side of the TSS region, around the TTS, throughout gene bodies, and even several kilobases upstream of promoters. Box-plots: centre line, median; box limits, upper and lower quartiles; whiskers, 1.5* interquartile range.

[0047] Figure 7 shows that proteomics screen reveals that Dppa2 / 4 are readers of hmC and bind selectively to the symmetric form, (a) Design of our biotinylated probes, which share an identical 68-bp sequence. Each of the five probes contained a different CpG state (shown above). The sequence flanking each CpG site was varied to maximise sequence space, (b) Volcano plot showing proteins enriched in the HH probe (symmetric hmC) relative to its unmodified counterpart. Dppa4 and Dppa2 exhibited the strongest enrichment. The modified probes were also enriched in Uhrfl — a known reader of mC and hmC27,28. (c) Specific binding of Dppa2 / 4 determined by automated Western blotting (Wes), which, along with our MS data, showed selective binding to the symmetric hmC probe. The plot was produced using an ordinary one-way ANOVA and Tukey’s multiple comparison tests (n = 3, F (5,12) = 23.37, error bars show standard deviation), (d) Levels of hydroxymethylated CpG states across bivalent promoters, grouped by their Dppa2 / 4 regulation status (dependent: both H3K4me3 and H3K27me3 marks regulated by Dppa2 / 4; sensitive: Dppa2 / 4 regulate only H3K4me3; independent: bivalent promoter not regulated by Dppa2 / 434. Box-plots: centre line, median; box limits, upper and lower quartiles; whiskers, 1 .5* interquartile range.) Levels of all hmC states are considerably higher in bivalent promoters regulated by Dppa2 / 4 compared to those that are not. (e) CpG sites containing the highest levels of HH (>95th percentile) are significantly enriched in bivalent promoters regulated by Dppa2 / 4 when compared to the rest of the genome (Fisher’s two-tailed exact test).

[0048] Detailed Description

[0049] This invention relates to methods of simultaneously sequencing both strands of a double stranded DNA (dsDNA) molecule or a population of dsDNA molecules. This may be useful for a range of sequencing applications and may for example, provide improved error correction compared to conventional methods.

[0050] Methods of the invention may for example, be useful in distinguishing true mutations in genomic DNA from sequencing or amplification errors. In conventional, single-stranded sequencing approaches, a given locus from an original (single-stranded) DNA fragment is sequenced only once. At the level of an individual sequence read, a true mutation cannot be distinguished from an error introduced during amplification or sequencing, since both mutations and errors appear as a difference in sequence between the sequence read and the reference genome. Methods of the invention comprise simultaneously sequencing both strands of a dsDNA along with copies of both strands. A true mutation has a corresponding nucleotide change in one or both of the copy strands, whereas a sequence error introduced during amplification or sequencing does not have a corresponding nucleotide change. This allows mutations in genomic DNA to be distinguished from sequencing or amplification errors.

[0051] Methods of the invention may also be useful in providing increased confidence in base calling. The presence of both original strands of the dsDNA molecule, and their unmodified copy strands increases confidence in base-calling because the uncertainty of each base-call is reduced exponentially by combining each measurement. For example, each genomic site in the dsDNA molecule may be encoded by four corresponding base calls (one on each of the original strands, and one on each of the copy strands). An A, for example, in the first strand of the dsDNA molecule would, when correctly sequenced, produce the following base calls: A (Original first), T (Original second), A (Copy second), T (Copy First), which may be abbreviated to the tetramer sequence AT AT. Sequencing platforms that discriminate only between genetic bases may in principle produce 256 possible such tetramer sequences. Given that, when applied to resolving the genetic bases, only 4 of the possible tetramers correspond to real states (i.e. A, C, G, T), 252 of the possible tetramers would represent readily identifiable sequencing errors. Thus, the probability of a tetramer corresponding to a real state occurring by chance is low. Coupling this information with quality scores for each of the four base calls in the tetramer allow for sequence determination with extremely high confidence.

[0052] Methods of the invention may also be useful for sequencing platforms that natively sequence DNA modifications (e.g. Nanopore™ platforms). For dsDNA molecules that have been subjected to ligation of the two different hairpin adapters and copy strand synthesis (but not subsequent chemical conversions), sequence reads of the original duplex strands would contain the signal of the ds DNA molecule and any modifications (e.g. mC or hmC), while the sequence reads of the copy strands contain the same sequences but without any modifications. This provides a control experiment performed at the level of each read.

[0053] Methods of the invention may also be useful in the detection of mutations with high accuracy and sensitivity, and the simultaneous resolution of mutations and epigenetic cytosine modifications. For example, methods of the invention may be useful in mapping or identifying the positions of epigenetically modified cytosines in a double-stranded DNA molecule or a population of dsDNA molecules.

[0054] A dsDNA molecule comprises two complementary deoxyribonucleic acid strands (a first strand and a second strand) that hybridise through hydrogen bonding between the bases of each strand in accordance with Watson-Crick base pairing rules to form a double helical structure.

[0055] Double stranded DNA may include genomic DNA. For example, a genomic DNA molecule or a population of genomic DNA molecules or fragments may be sequenced as described herein. Genomic DNA is double stranded DNA from the genome of a cell. For example, genomic DNA may be part of a chromosome, a minichromosome, a whole chromosome or more than one chromosome. Suitable genomic DNA may also include plasmid or viral DNA inserted into a chromosome or genome. Genomic DNA may be from cell organelle, such as a nucleus or mitochondrion. Genomic DNA may comprise all or part of the sequence of one or more genes, including exons, introns or upstream or downstream regulatory elements, and / or the sequences may comprise a genomic sequence that is not associated with a gene. In some embodiments, a genomic DNA may represent the whole genome or a specific genomic locus of an organism or a population or organisms.

[0056] Suitable genomic DNA may include prokaryotic DNA, such as bacterial dsDNA. For example, dsDNA for use as described herein may include part or all of the genome of a prokaryotic cell, such as a bacterial cell, more preferably the whole genome of a prokaryotic cell.

[0057] Preferably, genomic DNA is eukaryotic genomic DNA, for example mammalian genomic DNA, such as human DNA. For example, genomic DNA for use as described herein may include part or all of the nuclear and / or mitochondrial genome of a eukaryotic cell, more preferably the whole genome of a eukaryotic cell.

[0058] Genomic DNA may be obtained or isolated from a sample of cells, for example, mammalian cells, preferably human cells. Suitable samples include isolated cells and tissue samples, such as biopsies. In some embodiments, the dsDNA may be obtained from a frozen cell or tissue sample.

[0059] In some embodiments, genomic DNA for use as described herein may be obtained from a eukaryotic cell. Eukaryotic cells may be isolated, for example as immortalised cell lines or primary cells obtained from an individual or may be in the form of tissues or organoids. For example, a cell nucleus may be obtained from eukaryotic cells within a sample obtained from an individual, such as a biopsy or xenograft sample.

[0060] Suitable eukaryotic cells may include mammalian cells, preferably human cells. For example, eukaryotic cells may include somatic and germ-line cells and may be at any stage of development, including fully or partially differentiated cells or non-differentiated or pluripotent cells, including stem cells, such as adult or somatic stem cells, foetal stem cells or embryonic stem cells. Suitable eukaryotic cells also include induced pluripotent stem cells (iPSCs), which may be derived from any type of somatic cell in accordance with standard techniques. Eukaryotic cells may also include neural cells, including neurons and glial cells; contractile muscle cells; smooth muscle cells; liver cells; hormone synthesising cells; sebaceous cells; pancreatic islet cells; adrenal cortex cells; fibroblasts; mesenchymal cells; epithelial cells; keratinocytes; endothelial cells; urothelial cells; osteocytes; chondrocytes; immune cells; such as leukocytes; mesothelial cells and adipocytes; bone marrow cells or samples.

[0061] Suitable eukaryotic cells also include normal cells and disease cells, for example cells associated with disease conditions, including cancer cells, such as carcinoma, sarcoma, lymphoma, blastoma or germ-line tumour cells, and cells with the genotype of a genetic disorder, such as Huntington’s disease, cystic fibrosis, sickle cell disease, phenylketonuria, Down syndrome or Marfan syndrome.

[0062] Suitable eukaryotic cells also include embryonic cells and cell culture models.

[0063] In some embodiments, genomic DNA for use as described herein may from a disease cell, for example a cell obtained from an individual with a disease, such as a cancer cell. The location or distribution of transcriptionally active regions within the genomic DNA may be indicative of the diagnosis, severity, penetrance grade or prognosis of a disease, such as cancer, in the individual from whom the genomic DNA is obtained.

[0064] In other embodiments, the genomic DNA may from a cell at a specific developmental stage, for example a cell obtained from an embryo; an age-associated cell, for example a cell obtained from an individual of a specific age, such as an elderly person, or a treated cell. Treated cells may include cells treated with reprogramming factors, such as Oct3 / 4, Sox2, Klf4 and c-Myc (the “Yamanaka factors” or “OSKM”), radiation, chemical compounds, and drugs, such as chemotherapy agents.

[0065] Genomic DNA may be obtained from a population of cells or an individual cell (e.g. single cell genomics). The analysis of nucleic acids from a single cell may, for example, allow transcriptional activity in individual cells and cell-types to be determined.

[0066] For example, a sample of genomic DNA for use as described herein may contain 1 or more, 10 or more 100 or more or 1000 or 10,000, or 100,000 or more cell genomes. A sample of genomic DNA for use as described herein may contain 6pg or more, 60pg or more, 600pg or more, or 6ng or 60 ng or more of DNA.

[0067] Methods of extracting and isolating genomic DNA, from an individual cell, a sample of cells, organoids or tissue are well-known in the art. For example, genomic DNA or RNA may be isolated using any convenient isolation techniques, such as phenol / chloroform extraction and alcohol precipitation, caesium chloride density gradient centrifugation, solid-phase anion-exchange chromatography and silica gel-based techniques.

[0068] Following isolation, a sample of genomic DNA may be fragmented to produce genomic DNA fragments. Fragmentation may reduce the size of the genomic DNA. Suitable genomic DNA fragment sizes may depend on the read length capabilities of the platform to be used for sequencing. For example, following fragmentation, the genomic DNA fragments may be 10bp to 10OOObp, preferably 50 to 5000bp, 20bp to 2000bp or 30bp to 10OObp.

[0069] Suitable fragmentation methods are well-known in the art and include nebulization, sonication or acoustic shearing, mechanical shearing, tagmentation, and endonuclease digestion. The whole or a fraction of the fragmented genomic DNA sample may be used as described herein. Suitable fractions of genomic DNA may be based on size or other criteria. For example, a fraction of genomic DNA which is enriched for CpG islands (CGIs) may be used as described herein.

[0070] In some embodiments, following fragmentation, the ends of the dsDNA molecules may be repaired to produce a population of blunt-ended dsDNA molecules suitable for the attachment of hairpin adaptors. This may be useful for example, if the DNA molecules are fragmented by physical methods, such as nebulization, sonication or acoustic shearing, mechanical shearing. Suitable methods of repairing nucleic acid ends are well-known in the art. For example, fragmented nucleic acids may be converted into blunt-ended molecules by filling in 5’ overhangs using a 5’^3’ polymerase and removing 3' overhangs using a 3' to 5' exonuclease, in accordance with standard techniques. Suitable polymerases include T4 DNA polymerase. Suitable techniques and kits for the end-repair of DNA molecules are widely available from commercial suppliers (e.g. End-lt™ , Epicentre; NEBNext™ end repair Module, New England Biolabs; Fast DNA End Repair, Thermo Fisher Scientific; DNA End Repair Mix, Life Technologies; Paired-End Sample Prep Kit, Illumina Inc).

[0071] Following fragmentation and end-repair, the ends of the dsDNA molecules in the population may be blunt and may, in some preferred embodiments, comprise free hydroxyls at both the 5’ and 3’ ends of each strand.

[0072] In some embodiments, the ends of the dsDNA molecule may be blunt. The ends may for example comprise free hydroxyls at both the 5’ and 3’ ends of each strand.

[0073] In other embodiments, the blunt-ends of the dsDNA molecule may be modified before ligation of the hairpin adaptors. For example, if Klenow fragment is used as a polymerase to repair the the ends of the dsDNA molecules, a one-base overhang consisting of an adenine residue (A-tail) may be added to the 3’ strands of the blunt ends. This 3’ overhang may facilitate ligation of the hairpin adaptors. Other suitable modifications to facilitate ligation are well-known in the art. For example, the 5’ ends of the dsDNA molecule may be phosphorylated. Suitable techniques are available in the art and include treatment with a kinase, such as T4 polynucleotide kinase.

[0074] In some embodiments, the sample of genomic DNA may be fragmented using a restriction endonuclease. The restriction endonuclease may generate genomic DNA fragments with specific overhanging ends (so- called “sticky” ends). Hairpin adaptors may be designed to be directly attachable to these specific overhanging ends by ligation without end repair or other treatments. In some embodiments, depending on the choice of restriction endonuclease, the 5’ ends of the genomic DNA fragments may be phosphorylated using a kinase before the ligation step.

[0075] In other embodiments, the sample of genomic DNA may be fragmented and the hairpin adaptors attached by tagmentation. For example, a transposase may be pre-loaded with the first and second hairpin adaptors and used to tagment a sample genomic DNA. Suitable first and second hairpin adaptors for use in tagmentation may for example comprise a Mosaic End (ME) sequence for Tn5 compatibility. The gaps in the DNA strands of the resultant construct following tagmentation may be repaired by conventional techniques to generate the double hairpin construct.

[0076] A dsDNA molecule suitable for sequencing as described herein is linear and has two ends (a first end and a second end). In the methods described herein, hairpin adaptors are attached to the ends of the dsDNA molecule to form a double hairpin construct. The double hairpin construct comprises a dsDNA molecule having hairpin adaptors attached to both ends. The strands of the dsDNA molecule are therefore covalently connected at both ends through the hairpin adaptors.

[0077] A hairpin adaptor is an oligonucleotide adaptor that forms a hairpin or stem-loop structure in solution. A hairpin adaptor may comprise complementary nucleic acid sequences at its ends. These complementary sequences hybridise to each other to form a double-stranded stem region. For example, the melting temperature of the complementary sequences of the hairpin adaptors may be sufficient to allow stable stem formation at 25°C. The complementary end sequences of the hairpin adaptor flank a central sequence which is not complementary either to itself or to the complementary nucleic acid sequences. The central sequence forms a single-stranded loop region. The loop region may be located at an end of the stem region formed by the hybridisation of the complementary nucleic acid sequences. The other end of the stem region may be a free end and may comprise the 5’ and 3’ termini of the hairpin adaptor. The hairpin adaptor may be attached to DNA molecules through the free end of the stem region. The design of hairpin adaptors is well- known in the art.

[0078] Suitable hairpin adaptors may be 10 to 100 nucleotides in length, preferably 10 to 40 or more preferably 15 to 30 nucleotides.

[0079] The nucleotide sequences of the hairpin adaptors may be sufficiently distinctive and distinguishable from the reference genome to allow the hairpin adaptors to be unambiguously identified within a sequence read.

[0080] A first hairpin adaptor may be attached to the first end of the dsDNA molecule and a second hairpin adaptor may be attached to the second end of the dsDNA molecule.

[0081] The first hairpin adaptor is cleaved in the methods described herein to introduce a single strand break into the double hairpin construct. The first hairpin adaptor is cleavable. For example, the first hairpin adaptor may comprise one or more cleavage sites. A cleavage site is a site within the first hairpin adaptor that can be specifically cleavable by enzymatic, chemical, or other means. Phosphodiester bonds at a cleavage site may be cleaved to break the DNA strand of the first hairpin adaptor and introduce 5’ and 3’ ends (i.e. a single strand break is introduced at the cleavage sites in the first hairpin adaptor). The 3’ end may be extendable with a DNA polymerase and the 5’ end may be ligatable to a sequencing adaptor. The one or more cleavage sites are preferably located in the stem region of the first hairpin adaptor i.e. in a complementary sequence of the first hairpin adaptor.

[0082] Suitable cleavage sites are well known in the art. and include restriction endonuclease recognition sites and modified nucleotides, such as 8-oxoguanine or 8-oxoadenine, which are cleavable by formamidopyrimidine [fapy]-DNA glycosylase (Fpg) . In some preferred embodiments, the cleavage may be a 2’-deoxyuridine residue. Suitable techniques for cleaving nucleic acid sequences at 2’-deoxyuridine residues are well known in the art and include enzymatic cleavage, for example using Uracil DNA glycosylase and DNA glycosylase- lyase (e.g. USER™ enzyme mix; NEB, USA).

[0083] The second hairpin adaptor is not cleaved in the methods described herein and provides a covalent link between the original strands and the copy strands in the extended hairpin construct. A suitable second hairpin adaptor may be non-cleavable and may lack cleavage sites.

[0084] In some embodiments, the first hairpin adaptor may comprise a 2’-deoxyuridine residue such that it is cleavable with Uracil DNA glycosylase and DNA glycosylase-lyase and the second hairpin adaptor may lack a 2’-deoxyuridine residue, such that it is not cleavable with Uracil DNA glycosylase and DNA glycosylase- lyase. Preferably, the first and the second hairpin adaptors have different nucleotide sequences. This may facilitate identification during downstream analysis.

[0085] The first hairpin adaptor may be cleaved at the cleavage site in the first hairpin adaptor to introduce a break in the strand of the double hairpin construct. Following cleavage, the double hairpin construct may comprise 5’ and 3’ ends within the sequence of the first hairpin adaptor, preferably within the stem region of the first hairpin adaptor.

[0086] The 3’ end generated by cleavage of the first hairpin adaptor may be extended using the strand of the double hairpin construct as a template to produce an extended hairpin construct. Suitable methods of 5’ to 3’ extension are well-known in the art and include extension using a DNA polymerase, such as T4 polymerase or Klenow fragment.

[0087] The extended hairpin construct may comprise first and second original strands that correspond to the first and second strands of the dsDNA molecule. The extended hairpin construct may further comprise first and second copy strands that are complementary to the first and second original strands, respectively, and are generated by the polymerase extension. The first copy strand (Figure 2 “D”) is complementary to the first original strand (Figure 2 “A”) and the second copy strand (Figure 2 “C”) is complementary to the second original strand (Figure 2 “B”).

[0088] The extended hairpin construct may have a free end and a closed end. The original and copy strands may be linked at the closed end by a hairpin. The hairpin may correspond to the first hairpin adaptor of the double hairpin construct.

[0089] A sequencing adaptor may be attached to the free end of the extended hairpin construct.

[0090] A sequencing adaptor is an oligonucleotide molecule that facilitates sequencing of the target nucleic acid molecule to which it is attached.

[0091] Sequencing adaptors suitable for use in sequencing nucleic acids are well-known in the art and may be designed and produced using known techniques or obtained from commercial sources. Sequencing adaptors are generally specific for a sequencing platform and the choice of sequencing adaptor therefore depends on the specific sequencing method to be employed. Suitable sequencing platforms are described below.

[0092] Sequencing adaptors suitable for any of these sequencing platforms may be used in the methods described herein. In some embodiments, adaptors may include a region that is complementary to the universal primers on a solid support (e.g. a flowcell or bead) and a region that is complementary to universal sequencing primers (i.e. which when annealed to the adaptor sequence and extended allows the sequence of the nucleic acid molecule to be read). The sequence and design of the sequencing adaptor depends on the sequencing platform used. In some embodiments, for example when using Illumina™ sequencing, the sequencing adaptor may be a Y shaped adaptor that comprises two oligonucleotide strands. The sequences of the strands are partially complementary such that they hybridise together at a first end to form a double-stranded stem region and non-complementary and non-hybridised at a second end to form single stranded arms. The first end of the sequencing adaptor is attached to the free end of the extended hairpin construct. The single stranded arms at the second end hybridise to the two non-complementary flow-cell oligonucleotides of the sequencing platform during sequencing.

[0093] In some embodiments, a sequencing adaptor may comprise a unique barcode or index nucleotide sequence that identifies the source of the nucleic acid (e.g. the sample) and allows the multiple samples to be sequenced in a multiplex sequencing reaction. A suitable barcode or index may consist of 6-10 nucleotides. Outside the index, the sequences of the sequencing adaptors or may be the same for all the double stranded DNA molecules being sequenced.

[0094] In other embodiments, the barcode or index may be contained within amplification primers and introduced during library amplification or more preferably, may be contained in the loop region of the non-cleavable second hairpin adaptor.

[0095] Following attachment of the sequencing adaptor, the extended hairpin constructs may be sequenced to determine the sequences of the original and copy strands. The extended hairpin constructs may be sequenced using standard sequencing techniques. Suitable techniques include using any convenient low or high throughput sequencing technique or platform, including Sanger sequencing, Solexa-lllumina sequencing (for example Hiseq™, MiSeq™ or NextSeq™), ligation-based sequencing (SOLiD™; Life technologies), pyrosequencing (e.g. 454 Sequencing; Roche 454); single molecule real-time sequencing (SMRT™; PacBioscience); or semiconductor array sequencing (Ion Torrent™), nanopore sequencing (Oxford Nanopore e.g. MinlON™, GridlON™, PromethlON™), or sequencing platforms from Element (e.g. AVITI™), Singular Genomics (e.g. G4), Complete Genomics (e.g. T7), Qiagen, or Ultima Genomics (e.g. UG 100).

[0096] Preferably, sequencing is performed by a next-generation sequencing technique. Suitable protocols, reagents and apparatus for nucleic acid sequencing are well-known in the art and are available commercially.

[0097] The sequences of the first and second strands of the dsDNA molecule may be determined from the sequences of the first and second original and copy strands of the extended hairpin construct.

[0098] The presence within the extended hairpin construct of the first and second original strands and the first and second copy strands allows accurate base-calling with reduced uncertainty. For example, when applying this principle to sequencing the genetic bases (A, C, G, and T), each sequence position in the dsDNA molecule is encoded by four corresponding base calls (one on each of the first and second original strand sequences, and one on each of the first and second copy strand sequences). An A, for example, in the first strand of the dsDNA molecule (and a T in the corresponding position of the second strand) would, when correctly sequenced, produce the following base calls: A (Original first), T (Original second), A (Copy second), T (Copy First), which may be abbreviated to the tetramer sequence ATAT. Sequencing platforms that discriminate only between genetic bases may in principle produce 256 possible such tetramer sequences. However, because only 4 of the possible tetramers correspond to real states (i.e. A, C, G, and T), 252 of the possible tetramers would correspond to sequencing errors. The probability of a tetramer corresponding to a real state occurring by chance is low. Using the four base calls in the tetramer to identify the nucleotide at each sequence position in the dsDNA molecule allows for sequence determination with extremely high confidence.

[0099] The presence within the extended hairpin construct of the first and second original strands and the first and second copy strands may also allow the identification of sequencing or amplification errors. A nucleotide change that is present in all four of the first and second original strands and the first and second copy strands may be identified as a mutation in the dsDNA molecule at the position. A nucleotide change that is not present at a position in all four of the first and second original strands and the first and second copy strands mayube identified as a sequence error introduced during amplification or sequencing.. This allows sequencing or amplification errors to be distinguished from mutations in the dsDNA molecule.

[0100] The methods of sequencing described herein may be useful in identifying and / or mapping cytosines that are epigenetically modified, for example 5-methylcytosine (mC) or 5-hydroxymethylcytosine (hmC).

[0101] In some embodiments, epigenetically modified cytosines at CpG sites in the dsDNA molecule may be identified and / or mapped. A CpG site is a site in a sequence of genomic DNA which consists of cytosine immediately followed in a 5’ to 3’ direction by guanine (5’-CpG-3’) i.e. cytosine and guanine bases are directly linked in a 5’ to 3’ direction by a phosphodiester bond. Genomic regions with a high frequency of CpG sites are called CpG islands.

[0102] In some embodiments, sequencing may be performed using a platform that directly identifies epigenetically modified cytosines.

[0103] In other embodiments, the extended hairpin construct may be further treated before sequencing in order to identify and / or map epigenetically modified cytosines in double-stranded DNA molecules.

[0104] Following attachment of the sequencing adaptor, 5-hydroxymethylcytosines in the first and / or second original strand of the extended hairpin construct may be protected. For example, 5-hydroxymethylcytosines may be glycosylated. Suitable methods of glycosylation include glucosylation, for example by treatment with p- glucosyltransferase. Protection of a 5-hydroxymethylcytosine at a CpG site in the first and second original strand protects the cytosine that is present at the CpG site in the first and second copy strands from methylation by DNA methyltransferase. Protection, for example by glycosylation, also protects 5- hydroxymethylcytosine from APOBEC3A deamination. In other embodiments, deamination may be bisulphite mediated deamination and 5-hydroxymethylcytosine (hmC) may be resistant to the bisulphite mediated deamination without glycosylation or other protection. Following protection of 5-hydroxymethylated cytosines, for example by glycosylation, the extended hairpin construct may then be treated with a DNA methyltransferase. Cytosines (C) in the copy strands of the extended hairpin construct at CpG sites that comprise methylated cytosines (mC) in the original strands are methylated by the DNA methylase. For example, CpG sites in the extended hairpin construct at CpG sites that comprise a mC on the first original strand and a C on the first copy strand (i.e. CpG sites with mC / C and C / mC states) are converted into CpG sites containing mC on both the first original and first copy strands (mC / mC) sites. CpG sites in the extended hairpin construct at CpG sites that comprise a mC on the second original strand and a C on the second copy strand (i.e. CpG sites with C / mC and mC / C states) are converted into CpG sites containing mC on both the second original and second copy strands (mC / mC) sites. Suitable DNA methyltransferases are available in the art and include DNMT1 and DNMT5. Preferably, the DNA methyltransferase is DNMT5.

[0105] Following DNA methyltransferase treatment, 5-methylcytosines (mC) in the first and second original and copy strands may be be protected. For example, 5-methylcytosines may be oxidised and glycosylated. In some embodiments, 5-methylcytosine (mC) may be oxidised to 5-hydroxymethylcytosine (hmC) and glycosylated to produced glycosyl-5-hydroxymethylcytosine. Suitable oxidation methods include treatment with TET2. Suitable glycosylation methods include glucosylation, for example treatment with p- glucosyltransferase to produce p-glucosyl-5-hydroxymethylcytosine.. 5-methylcytosine (mC) may also be oxidised to 5-formylcytosine or 5-carboxylcytosine. This protects methylated cytosines in the original strands and copy strands from deamination, for example APOBEC3A mediated deamination. In other embodiments, deamination may be bisulphite mediated deamination and 5-methylcytosine (mC) may be resistant to the bisulphite mediated deamination without oxidation and glycosylation or other protection.

[0106] The first and second original and copy strands may be deaminated to produce deaminated strands. Cytosines (C) in the original strands and copy strands may then be deaminated into uracils (U). Suitable deamination methods include treatment with APOBEC3A and UvrD helicase. UvrD helicase denatures duplex DNA to produce single stranded DNA that can be deaminated by APOBEC3A. In other embodiments, bisulphite mediated deamination may be employed.

[0107] Following deamination, the extended hairpin construct may be amplified before sequencing to produce amplified strands. Suitable amplification methods are well-known in the art.

[0108] In some embodiments, amplification may replace uracils (U) in the deaminated strands with by thymines (T).

[0109] Following deamination and optional amplification, the extended hairpin constructs may be sequenced to determine the sequences of the first and second original and copy strands.

[0110] From the sequences of the first and second original strands and the first and second copy strands the cytosine or modified cytosine residues at CpG sites may be identified. These identities may be indicative of the identities of the cytosine or modified cytosine residues at the corresponding CpG sites in the first and second strands of the double stranded DNA molecule. For example, the presence of U / T in a CpG site in the sequence of the first original strand, U / T in the CpG site in the sequence of the second original strand, U / T in the CpG site in the second copy strand and U / T in the CpG site in the first copy strand, may be indicative of the presence of C in both strands at the CpG site in the double stranded DNA molecule (CC).

[0111] The presence of U / T in a CpG site in the sequence of the first original strand, C in the CpG site in the sequence of the second original strand, C in the CpG site in the sequence of the second copy strand and U / T in the CpG site in the sequence of the first copy strand, may be indicative of the presence of C in the first strand and mC in the second strand at the CpG site in the double stranded DNA (CM).

[0112] The presence of C in a CpG site in the sequence of the first original strand, U / T in the CpG site in the sequence of the second original strand, U / T in the CpG site in the sequence of the second copy strand and C in the CpG site in the sequence of the first copy strand, may be indicative of the presence of mC in the first strand and C in the second strand at the CpG site in the double stranded DNA (MC).

[0113] The presence of C in a CpG site in the sequence of the first original strand, C in the CpG site in the sequence of the second original strand, C in the CpG site in the sequence of the second copy strand and C in the CpG site in the sequence of the first copy strand, may be indicative of the presence of mC in the first strand and mC in the second strand at the CpG site in the double stranded DNA (MM).

[0114] The presence of U / T in a CpG site in the sequence of the first original strand, C in the CpG site in the sequence of the second original strand, U / T in the CpG site in the sequence of the second copy strand and U / T in the CpG site in the sequence of the first copy strand, may be indicative of the presence of C in the first strand and hmC in the second strand at the CpG site in the double stranded DNA (CH).

[0115] The presence of C in a CpG site in the sequence of the first original strand, U / T in the CpG site in the sequence of the second original strand, U / T in the CpG site in the sequence of the second copy strand and U / T in the CpG site in the sequence of the first copy strand, may be indicative of the presence of hmC in the first strand and C in the second strand at the CpG site in the double stranded DNA (HC).

[0116] The presence of C in a CpG site in the sequence of the first original strand, C in the CpG site in the sequence of the second original strand, U / T in the CpG site in the sequence of the second copy strand and U / T in the CpG site in the sequence of the first copy strand, may be indicative of the presence of hmC in the first strand and hmC in the second strand at the CpG site in the double stranded DNA (HH).

[0117] The presence of C in a CpG site in the sequence of the first original strand, C in the CpG site in the sequence of the second original strand, U / T in the CpG site in the sequence of the second copy strand and C in the CpG site in the sequence of the first copy strand, may be indicative of the presence of mC in the first strand and hmC in the second strand at the CpG site in the double stranded DNA (MH). The presence of C in a CpG site in the sequence of the first original strand, C in the CpG site in the sequence of the second original strand, C in the CpG site in the sequence of the second copy strand and U / T in the CpG site in the sequence of the first copy strand, may be indicative of the presence of hmC in the first strand and mC in the second strand at the CpG site in the double stranded DNA (HM).

[0118] The methods described above may be useful in identifying the positions of hemi-hydroxymethylated CpG. For example, the presence of U / T in a CpG site in the sequence of the first original strand, C in the CpG site in the sequence of the second original strand, U / T in the CpG site in the sequence of the second copy strand and U / T in the CpG site in the sequence of the first copy strand, may be indicative of an hemi- hydroxymethylated CpG site with C in the first strand and hmC in the second strand at the CpG site in the double stranded DNA (CH).

[0119] The presence of C in a CpG site in the sequence of the first original strand, U / T in the CpG site in the sequence of the second original strand, U / T in the CpG site in the sequence of the second copy strand and U / T in the CpG site in the sequence of the first copy strand, may be indicative of a hemi-hydroxymethylated CpG site with hmC in the first strand and C in the second strand at the CpG site in the double stranded DNA (HC).

[0120] A set of sequence reads of extended hairpin constructs may be generated by sequencing. For example 1000 or more, 10,000 or more, 100,000 or more, 1000,000, 10,000,000 or more, or 100,000,000 or more, or 1000, 000,000 or more sequence reads may be generated. The sequence reads may be analysed by bioinformatic techniques. Suitable techniques are described below.

[0121] The sequence reads in the set may be analysed for the presence of hemi-hydroxymethylated CpG sites, the frequency of hemi- hydroxymethylated CpG sites and / or patterns of hemi- hydroxymethylated CpG sites within the genomic DNA. The sequence reads in the set may further be analysed for the presence of other features, such as mutations, epigenetic modifications or sequence motifs, that are associated with hemi- hydroxymethylated CpG sites.

[0122] In some embodiments, duplicate reads, low quality sequence reads and reads arising only from sequencing adaptors may be removed.

[0123] The fragment sequence reads in the set may be mapped to one or more locations in a reference genome. Suitable reference genomes are available in the art. For example, human genomic fragment sequence reads in the set may be mapped to locations in the sequence of the human genome. In some embodiments, the reference genome may be matched to the gender, ethnicity and / or other characteristics of the individual from whom the genomic DNA is obtained. The fragment sequence reads in the set may be mapped by aligning sequence reads in the set with the sequence of the reference genome, for example the human genome. The location of the sequence reads within the reference genome may be identified. This may allow CpG sites in which the cytosine in one or both strands is epigenetically modified to be mapped to one or more locations in the reference genome. For example, human genomic fragment sequence reads in the set may be mapped to locations in the sequence of the human genome. Suitable software tools for mapping populations of sequence reads within the genome are readily available in the art.

[0124] The distribution of the sequence reads of fragments in the set within the genome or at a set of sites or loci within the genome (i.e. the number of sequence reads that map to each location within the genome or set of sites or loci) may be determined from the locations of the sequence reads in the set. The distribution of epigenetically modified CpG sites, including symmetrically and asymmetrically modified CpG sites, at these sites or loci may be determined. Optionally, the distribution of the fragment sequence reads may be subjected to mathematical transformation.

[0125] A sample profile of CpG sites in which the cytosine in one or both strands is epigenetically modified, including symmetrically and asymmetrically modified CpG sites, may be generated from the distribution or transformed distribution of the genomic fragment sequence reads. For example, a sample profile of one or more of the 9 CpG states described above may be generated. The sample profile may comprise a set of scores or values indicative of the number or density of nucleic acid fragment sequence reads containing CpG sites in which the cytosine in one or both strands is epigenetically modified that map to each location or position within the genome (i.e. a genome wide plot) or set of sites or loci within the genome. The number or density of fragment sequence reads containing CpG sites in which the cytosine in one or both strands is epigenetically modified that map to a location or position in the genome may be indicative of the frequency of these CpG sites at that location or position. The sample profile may therefore reflect the distribution of CpG sites containing CpG sites in which the cytosine in one or both strands is epigenetically modified in the genome or in target loci within the genome. For example, the sample profile may reflect the distribution of one or more of the 9 CpG states described above. The sample profile may be expressed in any convenient format, for example numerically or graphically.

[0126] In some embodiments, the sample profile may be used to identify CpG sites at which the epigenetic modification of the cytosine in one or both strands is associated with biological responses to a drug or other compound, such as cytotoxic responses, and may be useful in determining or predicting the response of an individual to treatment with the drug or other compound. The sample profile may be used for therapeutic stratification, the optimization of combination therapies, diagnosis of a disease condition, the determination of side-effects of treatment with a drug or other compound, cytotoxicity assays, the determination of the off- target effects of nucleases, assisted reproductive technology methods, or the study of ageing or development. The sample profile may also be used for epigenetics research, as a proxy for other genetic or epigenetic features, or to determine cellular / tissue origin of dsDNA fragments.

[0127] This invention also relates to methods for mapping the locations of transcriptionally active genes in genomic DNA by identifying genes containing CpG sites that are hemi-hydroxymethylated (i.e. one strand is C and the other strand is hmC). The presence of hemi-hydroxymethylated CpG sites in a gene is indicative that the gene is transcriptionally active. Enhancers within the genomic DNA may also be identified or mapped from the positions of the hemi hydroxymethylated CpG sites. A transcriptionally active gene is a gene that is actively expressed in the source cells from which the genomic DNA is obtained. A transcriptionally active gene may for example be located within active euchromatin in the source cells. A hemi-hydroxymethylated CpG site may be located in the gene body of a transcriptionally active gene i.e. within an exon or an intron of the gene. In some embodiments, a hemi- hydroxymethylated CpG site may be located between the transcription start site (TSS) and the transcription termination site (TTS) of a gene. In other embodiments, a hemi-hydroxymethylated CpG site may be located several kilobases upstream of the TSS or downstream of the TTS, for example up to 1 kb, 2kb, 3kb, 4kb or 5kb or more upstream of the TSS or downstream of the TTS.

[0128] In some embodiments, a method may further comprise identifying CpG sites that are unmodified (i.e. both strands are C) in a promoter that is operably linked to a gene identified as containing hemi- hydroxymethylated CpG sites. The presence of CpG sites with unmodified cytosines in the promoter is further indicative that the gene is transcriptionally active

[0129] Mapping transcriptionally active genes may be useful for example in studying development the function of genes and regulatory elements, and epigenetic regulation), diagnostics, prognostics, and other healthcare applications, for example determining epigenetic status of disease cells, response to medication, and / or drug discovery.

[0130] As described above, a hemi-hydroxymethylated CpG site is a CpG site that comprises hydroxymethyl cytosine (hmC) on one strand of a genomic DNA molecule and unmodified cytosine (C) on the complementary strand. For example, a hemi hydroxymethylated CpG site in genomic DNA having a first and a second strand may comprise hmC on the first DNA strand and C on the second DNA strand or hmC on the second DNA strand and C on the first DNA strand (hmC / C or C / hmC).

[0131] The presence of a hemi hydroxymethylated CpG site is shown herein to be indicative of a transcriptionally active gene in a genomic DNA molecule. In some embodiments, the sequences of a population of genomic DNA fragments, may be analysed for the presence of hemi-hydroxymethylated CpG sites.

[0132] The positions of hemi- hydroxymethylated CpG sites may be identified by sequencing the first and second strands of a genomic DNA fragment.

[0133] In some embodiments, a hemi-hydroxymethylated CpG site identified within a genomic DNA fragment may be mapped to a position in the genomic DNA. For example, the CpG site may be mapped to a intron or exon of a gene in the genomic DNA. In other embodiments, a CpG site mapped to a position in the genomic DNA may be identified as being hemi-hydroxymethylated.

[0134] The first and second strands of a genomic DNA fragment or a population of genomic DNA fragments may be sequenced by any suitable method. In preferred embodiments, the first and second strands are sequenced by a method described above. The present invention also relates to the identification and / or mapping of enhancers, including primed enhancers, in genomic DNA by the presence of hemi-hydroxymethylated CpG sites. The presence of these CpG sites may be indicative of the priming of developmental pathways in samples of genomic DNA.

[0135] A method of mapping enhancers in genomic DNA comprising; providing a population of fragments of genomic DNA, and identifying hemi-hydroxymethylated CpG sites in the fragments, wherein hemi-hydroxymethylated CpG sites in the fragments are indicative of enhancers in the genomic DNA.

[0136] The mapping of HC and CH CpG sites as described herein may also be useful in the diagnosis, prognosis or study of diseases associated with hmC dysregulation, such as cancer, Alzheimer's and Parkinsons and as a marker for ageing.

[0137] The present invention also relates to the finding that CpG sites that include hmC on both strands (HH CpG sites) are recognised by the pluripotency factors Dppa2 and Dppa4 and are enriched in bivalent promoters regulated by Dppa2 and / or Dppa4. HH CpG sites may therefore be useful in the study of bivalent promoters in genomic DNA.

[0138] A method of mapping or characterising bivalent Dppa2 / 4-regulated promoters in genomic DNA may comprise; providing a population of fragments of genomic DNA, and identifying CpG sites in the population that comprise hmC on both strands, wherein the presence of CpG sites comprise hmC on both strands is indicative of a bivalent promoter in the genomic DNA.

[0139] The presence of CpG sites that include mC on one strand and hmC on the other (MH or HM CpG sites) within promoters has also been found to be associated with transcriptional repression.

[0140] Method of identifying repressed promoters in genomic DNA comprising; providing a population of fragments of genomic DNA, and identifying one or more promoters in the genomic DNA that include a CpG site comprising hmC on one strand and mC on the other, wherein the presence of the CpG site is indicative that the promoter is repressed.

[0141] The present invention also relates to the finding that CpG sites that include C on both strands, for example within promoters, are characteristic of active transcription and may be useful as a marker for mapping transcription or expression of genes in genomic DNA.

[0142] Method of identifying transcriptionally active promoters in genomic DNA comprising; providing a population of fragments of genomic DNA, and identifying one or more promoters in the genomic DNA that include a CpG site comprising C on both strands, wherein the presence of the CpG site is indicative that the promoter is transcriptionally active.

[0143] The presence of CpG sites that include mC on one strand and hmC on the other (MH or HM CpG sites); hmC on one strand and C on the other (HC or CH CpG sites); mC on one strand and C on the other (MC or CM CpG sites); or mC on both strands (MM sites), for example within promoters, has also been found to be associated with transcriptional repression.

[0144] Also provided herein is a kit suitable for use in a sequencing method described herein. The kit may comprise a first hairpin adaptor and a second hairpin adaptor, wherein the first hairpin adaptor comprises a cleavage site, optionally a 2’-deoxyuridine residue.

[0145] The kit may further comprise one or more of; a sequencing adaptor, a DNA ligase, a uracil DNA glycosylase (UDG) a DNA glycosylase-lyase endonuclease a DNA polymerase, a DNA methytransferase, a TET methylcytosine dioxygenase, a APOBEC3A cytosine deaminase (A3A), and a UvrD helicase.

[0146] The kit may further comprise sequencing primers and / or other reagents. The kit may further comprise suitable buffers and washing solutions; end-repair and A-tailing reagents, fixing agents and permeabilising agents.

[0147] The kit may further comprise a solid support. Suitable solid supports may comprise a capture molecule that binds to eukaryotic cells, such as a lectin; and / or the binding member.

[0148] The kit may further comprise nucleic acid extraction and purification reagents. Suitable reagents are well- known in the art and include spin-chromatography columns.

[0149] The kit may further comprise amplification reagents. Suitable reagents are well-known in the art and include primers, dNTPs, and thermostable polymerases. In some embodiments, a kit may comprise amplification primers for amplification prior to sequencing.

[0150] The kit may include instructions for use in a method described herein. Other aspects and embodiments of the invention provide the aspects and embodiments described above with the term “comprising” replaced by the term “consisting of’ and the aspects and embodiments described above with the term “comprising” replaced by the term” consisting essentially of’.

[0151] It is to be understood that the application discloses all combinations of any of the above aspects and embodiments described above with each other, unless the context demands otherwise. Similarly, the application discloses all combinations of the preferred and / or optional features either singly or together with any of the other aspects, unless the context demands otherwise.

[0152] Modifications of the above embodiments, further embodiments and modifications thereof will be apparent to the skilled person on reading this disclosure, and as such, these are within the scope of the present invention.

[0153] All documents and sequence database entries mentioned in this specification are incorporated herein by reference in their entirety for all purposes.

[0154] “and / or” where used herein is to be taken as specific disclosure of each of the two specified features or components with or without the other. For example, “A and / or B” is to be taken as specific disclosure of each of (i) A, (ii) B and (iii) A and B, just as if each is set out individually herein.

[0155] Materials and Methods

[0156] Cell culture

[0157] Basal medium for ESC cultures was DMEM high glucose (Sigma, D6546-500mL) containing 10% FBS (Gibco, 16141079), 2 mM GlutaMax-l (Gibco, 35050-038), 1 x NEAA (Gibco, 11140-035) and 0.1 mM p- mercaptoethanol (Sigma, M3148 -25 mL diluted to 50 mM in 50 uM EDTA / PBS solution). mESCs were regularly maintained on gelatin (G9391-100G)-coated plates by supplementing basal medium with 1 ,000 U / mL LIF (PeproTech, murine LIF, 250-02-25UG). Cultures were routinely monitored for Mycoplasma contamination.

[0158] Isolation of gDN A

[0159] Isolation of gDNA samples was performed as described previously42with the following modifications: Cultures were washed in PBS and lysed by adding RLT buffer (Qiagen) containing 1 :100 p-mercaptoethanol. Isolation of genomic DNA was performed with Zymo-Spin IIC-XL columns according to the instruction of the ZR-Duet DNA / RNA MiniPrep Kit (Zymo Research). Lysates were applied onto spin columns and bound material was incubated for 10 min with Genomic Lysis Buffer (Zymo Research) supplemented with 0.2 mg / mL RNaseA (Qiagen). After washing genomic DNA was eluted in water.

[0160] Preparation of sequencing libraries

[0161] Genomic DNA was sonicated using a Covaris M220 (150-bp target size). Fragments of ~100 bp were then isolated by automated pulsed-field gel electrophoresis (BluePippin, 3% agarose, Sage Science, catalogue number BDQ3010) according to the manufacturer’s protocol. 100 ng of the purified fragments was combined with the ‘H1 ’ and ‘H2’ spike-in controls (0.38 ng each, ATDBio Ltd), 1 uL of spike-in control (duet evoC, biomodal, catalogue number 6101), and adjusted to 50 uL with low-EDTA TE buffer (VWR, catalogue number A8569). Samples were end-repaired and A-tailed (NEB Ultra II End Repair / dA-Tailing Module, catalogue number E7546S) and adapter-ligated (NEB Ultra II Ligation Module, catalogue number E7595S) according to the manufacturer’s guidelines, except that 3.75 uL of a 1 :1 mixture of custom, pre-annealed hairpin adapters (obtained from ATDBio Ltd) were used from a stock of 5 uM per adapter. Prior to their mixing, each of these hairpin adapters was annealed separately in a buffer of 10 mM Tris HCI, 0.1 mM EDTA, and 100 mM NaCI, pH 8.0.

[0162] Following a clean-up (1 .6* SPRISelect, Beckman Coulter, catalogue number B23317), eluting in 23.75 uL nuclease-free water, samples were combined with rCutSmart buffer (10*, NEB) and 3.25 uL USER enzyme mix (NEB, catalogue number M5505S), incubated (37°C, 30 min), then purified by spin column (DNA Clean & Concentrator 5, Zymo Research, catalogue number D4014), eluting in 13 uL DNA Elution Buffer. 12 uL of each DNA sample was combined with 2.5 uL nuclease-free water, 2.5 uL NEB Buffer 4 (10*), 0.5 uL dNTPs (10 mM stock, NEB, catalogue number N0447S), 2.5 uL UDP-glucose (Thermo Scientific, catalogue number EO0831), 2 uL Polynucleotide Kinase (NEB, catalogue number M0201S), 2 uL Klenow Fragment exo- (Thermo Scientific, catalogue number EP0422), and 1 uL T4 p-glucosyltransferase (Thermo Scientific, catalogue number EO0831). Following incubation (37 °C, 30 min) a 1 .2* SPRISelect clean-up was performed, eluting in 14 uL Tris.HCI (10 mM, pH 8.0). Samples were then denatured and reannealed (95 °C, 2 min, then cooling to 10 °C at a rate of 0.1 °C / s). To ensure maximal ligation efficiency in the following step, 13.5 uL of each reannealed sample was combined with low-TE buffer (11.5 uL), NEBNext Ultra II End Prep Reaction Buffer (3.5 uL) and NEBNext Ultra II End Prep Enzyme Mix (1.5 uL), and incubated (30 min at 20 °C, then 30 min at 65 °C).

[0163] Ligation of the constructs to an Illumina-compatible Y-shaped adapter was performed by premixing the 30-uL DNA sample with EM-seq adapter (1 .25 uL, 15 uM stock, NEB, catalogue number E7140S) followed by NEB Ultra II Ligation Master Mix (15 uL) and NEB Ultra II Ligation Enhancer (0.5 uL). Following incubation (20 °C, 15 min), samples were immediately purified (1 * SPRISelect), eluting with 16 uL nuclease-free water. Next, the construct was treated with a series of enzymatic reactions using an early access duet evoC kit from biomodal. The MT Enzyme (2 uL, biomodal) was diluted with nuclease-free water (51.3 uL). To each 15-uL DNA sample was then added MT Buffer (6 uL, biomodal), MT Additive 3 (0.4 uL, biomodal), diluted MT Enzyme (2 uL, biomodal), MT Additive 2 (0.8 uL, biomodal), MT Additive 1 (1.5 uL, biomodal) and nuclease- free water (4.3 uL). Samples were then incubated (23 °C, 1 hour). Next, the oxidation reaction was assembled by combining the DNA sample (30 uL) with Ox Additive 1 (10 uL, biomodal, reconstituted according to the manufacturer’s instructions), DTT (1 uL, biomodal), hmCP Enzyme Additive (1 uL, biomodal), hmCP Enzyme (1 uL, biomodal), and Ox Enzyme (2 uL, biomodal). Ox Additive 2 (1 uL, biomodal) was then diluted in nuclease-free water (1 ,249 uL) and 5 uL of this solution was added to the DNA-containing reaction mixture.

[0164] Following incubation of the samples (37 °C, 1 hour), PRK (1 uL, biomodal) was added, followed by further incubation (10 mins at 37 °C, then 10 mins at 55 °C, lid heated to 85 °C). Samples were then purified (1 .8* SPRISelect); each was eluted in 32 uL nuclease-free water, of which 31 uL was collected. 1 uL of each sample was used for quantitation (Qubit dsDNA HS). For the deamination reaction, to each 30 uL DNA sample was added nuclease-free water (12.7 uL), DA Buffer (17.5 uL, biomodal), ATP (1.8 uL, biomodal), MgCI2 (3.5 uL, biomodal), DA Enzyme 1 (2 uL, biomodal), and DA Enzyme 2 (2.5 uL, biomodal) and the samples incubated (90 mins at 37 °C, lid heated to 45 °C).

[0165] To each sample was added DA Cleanup (6.5 uL, biomodal), followed by SPRI purification (1 x SPRISelect, eluting in 21 uL nuclease-free water). 20 uL of each DNA sample was then combined with Q5U Master Mix (25 uL, NEB, catalogue number M0597S) and an EM-seq UDI primer pair (5 uL, NEB, catalogue number E7140S), and amplified by 8 cycles of PCR (30 seconds at 98 °C for initial denaturation, 10 seconds at 98 °C for denaturation, 30 seconds at 62 °C for annealing, and 60 seconds at 65 °C for extension, and five minutes at 65 °C for the final extension). The amplified constructs were purified by a final SPRI clean-up (0.55x SPRISelect), eluting in 16 uL low-EDTA TE buffer.

[0166] Libraries were quantified by TapeStation D5000 (Agilent) and diluted to 1 .2-1 .8 nM prior to sequencing. To compensate for the reduced sequence diversity of the deaminated libraries, 10% PhiX (Illumina #FC-1 I Q- 3001) was added according to the manufacturer’s guidelines. All libraries were sequenced on an Illumina NovaSeq 6000 using the SP Standard workflow with 500-cycle kits in a 251-8-8-251 base-reads setup.

[0167] Sequencing data analysis

[0168] CpG-state calling from raw reads

[0169] CpG states were extracted by processing raw reads using a custom bioinformatics pipeline. First, target reads were filtered (based on their containing the correct configuration of adapter sequences) and trimmed using Cutadapt43. As different filtering steps were applied to reads 1 & 2, any resulting unpaired reads were removed using SeqKit Pair44. Next, inserts were separated from flanking hairpin-adapter sequences using Cutadapt, such that each read pair produced a set of four insert sequences (A & B from read 1 , C & D from read 2, Fig. 2a). Read pairs that did not yield all four corresponding insert sequences were discarded using SeqKit Pair. The resulting inserts were aligned in paired-end mode (A & D, B & C) using BSBolt45. Data from lanes 1 & 2 of each NovaSeq run were then merged with SAMtools46 and deduplicated in paired-end mode with Picard47. Next, the paired, deduplicated bam files were split back into the four corresponding insert sequences (A-D) using BamTools48. Methylation calls were extracted at the single-read level using a custom Python script. Using a custom R49 script, the four methylation calls (from inserts A, B, C, D) associated with each CpG site from each read pair were compiled and used to call the corresponding CpG state (Fig. 2a). The genomic coordinates and count of each called CpG state were then compiled. At this stage, data from the two independent experiments were either merged (for the main analysis) or analysed separately to assess agreement between datasets. Using a custom R script, CpG sites whose doublestranded depth was below 5x were discarded, as were those which overlapped with blacklisted regions50 or containing multiple CpG-state calling errors. After filtering, 15.8 million CpG sites were retained. The percentage of each individual CpG state at each CpG site was then calculated, as were total percentages of corresponding asymmetric CpG states (MC + CM, HC + CH, HM + MH). For the analysis comparing the two replicates, individual processing and filtering of the replicates yielded 13.7 million CpG sites (Replicate 1) and 4.9 million CpG sites (Replicate 2). Confusion matrix

[0170] Data used in the call-rate matrix were derived from all CpG sites of the synthetic spike-ins and analysed using a custom R script.

[0171] Global analysis

[0172] To infer the true rate of each CpG state in the genome, we used the signal obtained from the oligonucleotide controls to correct the observed levels of each CpG state. Specifically, each observed CpG state was modelled with a system of 9 linear equations that capture the additive contribution of all the other detected states levels (signal) observed on the oligonucleotide control sequences. where:

[0173] Op. is the observed ith CpG state across the genome,

[0174] Pij-. is the combination of all 9 measured CpG states’ levels observed across the oligonucleotide controls, and

[0175] Ty. is the true level of each CpG state.

[0176] With a custom R script, the system of linear equations was instructed. The coefficients were estimated from the sequencing data as well as the observed genome levels 0^. The system of linear equations was resolved using the R base function solveO; any resulting negative values were set to zero. The results obtained from this were visualized in the barplot of Fig. 3a.

[0177] Enhancers

[0178] The enhancer analysis was performed using a custom R script. Briefly, all CpG sites overlapping with enhancer peaks (primed, poised, or active, as defined by Cruz-Molina et al.26) were identified with BEDTools51 , and mean levels of each CpG state (or grouped, corresponding CpG states) at each enhancer were calculated. This was also done for randomly shuffled enhancer peaks (BEDTools shuffle) to enable the background levels of each state to be determined. Profiles of CpG states at each enhancer type were generated using deepTools52(200-bp bin size).

[0179] Transcription

[0180] RNA-seq data from E14-mESCs cultured in serum / LIF were obtained from a publicly available dataset33and, using a custom R script, grouped by transcriptional output (silent, and the tertiles low, medium, and high). CpG coordinates were annotated using ChlPseeker53(promoter-TSS region defined as 1 kb either side of the TSS) and filtered for protein-coding genes. Mean CpG-state levels were calculated for each gene and annotation (e.g. introns, exons). Spearman’s correlation coefficients were then determined for each CpG state (or grouped corresponded states) for each annotation feature using the R stats function cor.testQ.

[0181] Multiple states and annotations were combined by factoring pairwise combinations of CpG-state levels from different annotations. Plots of Fig. 5c were produced using deepTools52from input data binned to 200 bp. For gene-body plots, genes were scaled to 5 kb and additional binning (100 bp) was applied for clarity. For the TSS / TTS plots, no scaling or additional binning was applied. Dppa2 / 4

[0182] Coordinates of bivalent promoters grouped by Dppa2 / 4 regulation status were obtained from ref. 34. Using a custom script, CpG sites overlapping these promoters were determined (bedtools51intersect) and the average level of each CpG state across each promoter was calculated. To determine the enrichment of hmC states in bivalent promoters grouped by Dppa2 / 4 status, coordinates of CpG sites containing the highest 5% of modification levels were identified (R stats package quantileQ) and their enrichment or depletion in each group of bivalent promoters vs the entire genome was determined using Fisher’s exact test (bedtools51fisher). mESC nuclear extraction

[0183] All procedures were performed on ice using ice-cold buffers, all centrifugations were carried out at 4 °C unless stated and all rotations were carried out at 20 rpm unless stated. From each of six 150 mm plates of 80% confluent E14 mES cells per biological replicate (four biological replicates were used for mass spectrometry experiments), media was removed and cells were washed with DPBS (Gibco, 4190250), then scraped in DPBS and pelleted by centrifugation at 500 g for 20 minutes. Pellets were resuspended in three volumes of low salt buffer (20 mM HEPES pH 7.4, 10 mM NaCI, 3 mM MgCb, 0.2 mM EDTA, 1x HALT protease inhibitor cocktail EDTA-free, 1 mM DTT) and incubated for 15 minutes on ice. NP-40 (10% v / v in water) was added to a final concentration of 0.5%, the sample vortexed for 1 minute then centrifuged at 900 g for 15 minutes and the supernatant collected as cytoplasmic extract. Pelleted nuclei were resuspended in three volumes of high salt buffer (20 mM HEPES pH 7.4, 500 mM NaCI, 3 mM MgCh, 0.2 mM EDTA, 0.5% NP-40, 1x HALT protease inhibitor cocktail EDTA-free, 1 mM DTT) and incubated on ice with intermittent vortexing for 30 minutes. Sonication (4 x [30 sec on, 1 min off]) using a Bioruptor Plus (Diagenode, catalogue number B01020001) was used to break up genomic DNA before centrifugation at 21000 g for 10 minutes. The supernatant was collected as nuclear extract and the protein concentration quantified using a Pierce detergent-compatible Bradford Assay kit (Thermo, 23246) as per manufacturer’s instructions. 100 pL magnetic beads (Dynabeads Streptavidin M280, Thermo, 11205D) per biological replicate were washed as per manufacturer’s instructions, 750 pg nuclear lysate per DNA probe from each replicate added and the volume made up to 1 mL with incubation buffer (50 mM Tris-HCI pH 8, 150 mM NaCI, 0.5% NP-40, 1 mM DTT, 1x HALT protease inhibitor cocktail EDTA-free). The mixture was rotated for 1 hour at 4 °C, the beads collected by magnet and the supernatant kept as pre-cleared nuclear lysate.

[0184] DNA-protein pulldown

[0185] DNA oligonucleotides were obtained from Merck or ATDBio and were annealed in annealing buffer (100 mM NaCI, 1.1 mM Tris-HCI pH 7.5, 0.1 mM EDTA) at 20 pM and 30 pM for forward (biotinylated) and reverse strands respectively. Annealing was performed by heating samples to 95 °C for 3 minutes then cooling the samples at 2 °C / min until at 4 °C in a BioRad T100 Thermal Cycler (1861096). 50 pL magnetic beads (Dynabeads Streptavidin M280, Thermo, 11205D) per sample were washed and annealed with 10 pg per replicate of the relevant DNA duplex as per manufacturer’s instructions (for no-DNA control samples the same volume of DNA annealing buffer was used). Beads were rotated at room temperature for 15 minutes then washed as per manufacturer’s instructions and resuspended in 50 pL zero-NaCI incubation buffer per sample. To each bead / DNA sample was added 5 pg each of poly-dldC (Sigma, P4929) and poly-dAdT (Sigma, P0883) competitor DNA, 750 pg pre-cleared nuclear lysate and the volume made up to 1 mL with incubation buffer (final NaCI concentration corrected to 150 mM). The samples were agitated by rotation for 2 hours at 4 °C then the supernatant removed. The beads were washed twice in incubation buffer (1 mL) then once in PBS (400 pL), at which point samples were split for mass spectrometry, or silver staining and Western blotting. Samples intended for staining and blotting had the PBS removed and were snap-frozen before storing at -80 °C for further use. Samples intended for mass spectrometry analysis had the PBS removed and were washed twice in 100 mM ice-cold ammonium bicarbonate (1 mL), before the supernatant was removed and the beads snap-frozen before storing at -80 °C until analysis.

[0186] Beads for silver staining and Western blotting were resuspended in 30 pL Pierce RIPA buffer (Thermo, 89900) and shaken at 98 °C at 750 rpm for 10 minutes. The supernatant was separated from the beads and used as described.

[0187] Note that pulldowns used for MS analysis used BO, CC, CM, HC, HM and HH probes. Subsequent pulldowns to generate material used only for Western blotting also included the MM probe. Due to technical issues, the HC probe used in the initial MS experiment was reprepared and the pulldown and MS analysis repeated for this condition alongside CC for comparison (new samples denoted “CC_r2” and “HC_r2” in the raw data).

[0188] Silver staining.

[0189] To 20 pL of lysate from protein-DNA pulldowns was added 7 pL 4x LDS loading buffer (Thermo, NP0007) and 3 pL 10x reducing agent (Thermo, NP0009). 10 pL of the resulting mixture was loaded on a 1 .0 mm 4-12 % Bis-Tris polyacrylamide gel (Thermo, NP0321) and the gel run in 1x MOPS buffer (Thermo, NP0001) for 35 minutes at 200 V (constant voltage) before silver staining according to kit instructions (Thermo, 24612). Briefly, following electrophoresis gels were washed twice with water then fixed by two 15 min washes in 3:1 :6 ethanol:acetic acid:water solution. Gels were washed twice in 10 % ethanol then twice in water before washing in sensitizer solution for 1 min and washing twice with water. Gels were stained for 30 min, washed twice with water then developed for 150 seconds before stopping development with two washes in 5 % acetic acid.

[0190] Automatic Western blotting

[0191] Validation of the relative enrichment of Dppa2 and Dppa4 as identified by mass spectrometry was carried out using a Wes automatic Western blot system (Protein Simple 004-600) following the manufacturer’s instructions. Briefly, 1 pL of each pulldown sample (for the input lane, 1 uL of 1 :50-diluted mESC nuclear lysate was used) was mixed with 2 pL 0.1x sample buffer and 0.75 pL 5x fluorescent master mix (Protein Simple SM-W001), then samples were heated at 95 °C for 5 min before cooling on ice. 3 pL of sample was loaded per lane into a Wes microplate along with the relevant primary antibody (Rabbit anti-Dppa4, Thermo PA5-81373 diluted 1 :500; or Mouse anti-Dppa2, Merck MAB4356 diluted 1 :500), secondary conjugate (DM- 001 or DM-002, respectively), luminol-peroxide and wash buffer. The plate was centrifuged at 2500 rpm for 5 minutes at room temperature and the run initiated. Run files were analysed using Compass for SW software as follows: Peak Fit threshold was reduced from the default value of 10.0 to 1 .0, to allow quantification of low-intensity peaks. The Area of selected peaks were normalised by their Width and for each sample the Width-normalised Area value for the bead-only sample from the corresponding replicate was subtracted. Plots were produced using Prism 9 following an ordinary one-way ANOVA and Tukey’s multiple comparison tests (n = 3).

[0192] LC-MS / MS Analysis

[0193] Samples were prepared as described previously54. In brief, after on-bead tryptic digestion, the peptides were desalted using Pierce C18 Spin Tips (Thermo, 84850) equilibrated with 50% acetonitrile (Fisher Scientific) (plus 0.1 % formic acid, Fisher Scientific) and 0.1 % formic acid respectively. The acidified peptides were loaded, and the peptide-loaded cartridges were washed with 0.1 % formic acid and eluted with 60% acetonitrile (plus 0.1 % formic acid). Dried peptides were reconstituted in 0.1 % formic acid, for further LC- MS / MS analysis.

[0194] Reconstituted peptides were analysed on a Dionex Ultimate 3000 system coupled with the nano-ESI source Q-Exactive / Q-Exactive-HF instruments (Thermo Scientific). Peptides were loaded and separated on a reverse-phase trap column (2 cm x 100 pm) and analytical column (25 cm x 75 pm) respectively with 5-45% acetonitrile gradient in 0.1 % formic acid at 300 nL / min flow rate. In each data collection cycle, one full MS scan (400-1 ,600 m / z) was acquired in the Orbitrap (60K resolution, automatic gain control (AGC) setting of 3x106 and Maximum Injection Time (MIT) of 100 ms). The most abundant ions with top 10 settings were selected for fragmentation by High-energy Collision induced dissociation (HCD). HCD was performed with a collision energy of 28%, an AGC setting of 2x104, an isolation window of 2.0 Da and a MIT of 100ms. Previously analysed precursor ions were dynamically excluded for 30s.

[0195] Spectral .raw files from data dependent acquisition were processed with the SequestHT search engine on Proteome Discoverer 2.4 software (Thermo Scientific). Data were searched using Uniprot Database Mus musculus fasta file (taxon ID 1009 - Version October 2023). The node for SequestHT included the following parameters: Precursor Mass Tolerance 20 ppm, Fragment Mass Tolerance 0.02 Da. Dynamic Modifications were Oxidation of M (+15.995 Da) and Deamidation of N, Q (+0.984 Da). The Precursor Ion Quantifier node (Minora Feature Detector) used for label-free quantification included a Minimum Trace Length of 5, Max. ART of Isotope Pattern 0.2 min. The consensus workflow included peptide validator, protein filter and scorer. For calculation of Precursor ion intensities, Feature mapper was set True for RT alignment, with the mass tolerance of 10 ppm. Precursor abundance was quantified based on intensity and the level of confidence for peptide identifications was estimated using the Percolator node with a Strict FDR at q-value < 0.01 .

[0196] Data processing, normalisation, statistical analysis and plotting were carried out in R using the qPLEXanalyzer package54. Bead-only control groups (referred to as “BO” in deposited data) were removed, as were peptides missing in all remaining samples. Peptides present in at least two of the four biological replicates for any DNA probe were retained, before peptide intensities were normalised using median scaling. Peptides missing in three or four replicates for any DNA probe were replaced with minimum values from that DNA probe, and random missing values were imputed within each DNA probe group using nearest neighbour averaging. Protein level quantification was obtained by the summation of the normalized peptide intensities. A limma-based analysis was performed comparing modified probes (CM, HC, HM, HH) to the unmodified control (CC): a linear model is fitted for each protein including variables for each group and MS run, then Iog2-fold changes between comparisons are estimated. Multiple testing correction of p-values was applied using the Benjamini-Hochberg method55to control the FDR. Proteins differentially enriched between pairs of probes were regarded as significant when the adjusted P value was <0.05, absolute Iog2-fold change was >1 .0 and >3 unique peptides for a given protein were detected.

[0197] Results

[0198] Quantitative, simultaneous sequencing of C, mC, and hmC in double-stranded DNA

[0199] To determine the double-stranded epigenetic states comprising C, mC, and hmC, we have developed a sequencing approach that quantitatively resolves all three cytosine states simultaneously on both strands of the DNA double helix. In this method, the two strands of each double-stranded DNA fragment are physically tethered throughout the experiment by a hairpin adapter (Fig. 2a). This prevents separation of the strands — and thus loss of duplex information — which would otherwise occur before sequencing20. Unmodified copy strands of both original DNA strands are then produced within each construct. A series of enzymatic conversions of the resulting constructs enables C, mC, and hmC on each strand to be sequenced at singlebase resolution to simultaneously sequence C, mC and hmC in single-stranded DNA.

[0200] In the workflow, genomic DNA is first fragmented by sonication. Fragments of -100 bp are then size-selected by automated pulsed-field electrophoresis. These fragments are 5’-phosphorylated and 3’-A-tailed, followed by ligation to a 1 :1 mixture of two short, synthetic hairpin adapters — one containing 2’-deoxyuridine (dU), the other only canonical nucleotides. Target constructs are capped with each of the hairpin adapters. The dU from the adapter at one end of the construct is enzymatically cleaved. The resulting self-priming construct is extended by Klenow exo-, forming a larger hairpin in which the two original DNA strands are paired with newly synthesised ‘copy’ strands (Fig. 2a). The extended target construct is then separated from other constructs by size-selection and ligated to a standard, Illumina-compatible adapter. At this stage, CpG sites may be (i) fully unmodified on both strands, (ii) hemimethylated, or (iii) hemihydroxymethylated (protected by glucosylation). For subsequent base-conversion steps, we use the established method described previously17. Briefly, DNMT5-mediated copy-methylation enzymatically converts hemimethylated CpG to symmetrically methylated CpG sites but leaves other CpG sites (containing either C or hmC) unchanged. Next, mC and hmC are protected from deamination by TET2 oxidation and glucosylation, and all unmodified cytosines are then enzymatically deaminated using APOBEC3A and UvrD helicase. The resulting loss of complementarity within the hairpin enables efficient PCR amplification and Illumina sequencing. Both original strands of the DNA fragment are sequenced in Read 1 , their copy strands in Read 2. Thus, each CpG site has four associated cytosines, whose methylation calls are extracted by a custom bioinformatics pipeline. In this pipeline, target reads are selected based on their hairpin-adapter content, after which the four insert sequences are extracted from flanking hairpin sequences. The two resulting pairs (A & D, and B & C in Fig. 2a) of complementary sequences are aligned in paired-end mode, followed by deduplication and methylation calling. Each CpG site (numbered 1-3 in Fig. 2a) in an original double-stranded fragment is thus described by four methylation calls, which encode all nine possible CpG states comprising C, mC, and hmC (Fig. 2b). Seven other combinations of methylation calls are also possible and do not correspond to plausible CpG states. These rarely observed calls (0.31 % of all CpG calls) are flagged as errors downstream and discarded. Note that this workflow may be used with any other sequencing technology that distinguishes between canonical DNA bases by a straightforward substitution of the Illumina-compatible adapter. We performed our workflow on genomic DNA isolated from two different passages of E14 mouse-embryonic stem cells (mESCs). The E14 cell line is a well-studied model system with a wealth of genomic, epigenetic and transcriptomic datasets available21. To quantify the accuracy of our sequencing method, synthetic control oligonucleotide duplexes containing all possible permutations of CpG states comprising C, mC, and hmC (Fig. 3a) were added to the sonicated samples following size-selection, along with fully CpG-methylated bacteriophage I DNA, and fully unmethylated pUC19. We sequenced the two independent replicates on a NovaSeq 6000 (SP flow cell, 500 cycles), producing an average of 762 million indexed, paired-end reads per run. The average PCR and cluster duplication rate was 12.2%. Evaluation of the spike-in controls showed that our method accurately resolves all nine CpG states comprising C, mC, and hmC, with a call-rate accuracy of over 94% in each case (Fig. 3b). Call-rate accuracies (Pearson’s R = 0.999, P < 2.2 x 10-16) and bioinformatics analysis show close agreement between the two independent datasets. The E14 datasets were thus merged to maximise depth and coverage, resulting in 94.4% of the genome being covered by at least one read — including 18.7 million CpG sites at an average depth of 12.9x. Because both original DNA strands are covered in each read pair, this corresponds to a depth of 25.8x when compared to conventional, single-stranded methods. hmC in mESCs exists in all forms and is mostly asymmetric

[0201] Having established the accuracy of our method, we performed a global analysis of the mESC epigenome, correcting for the low-levels of false positive and negative rates associated with each CpG-state call, as determined from the spike-in data. Our data reveal that all nine possible CpG states comprising C, mC, and hmC exist at substantial levels in the mESC epigenome. Symmetric methylation (MM) is the most common state globally (60% of all CpG state calls), followed by symmetric unmodified CpG (CC, 22%, Fig. 4a). The next most abundant states are the two forms of hemimethylation (MC and CM denoting mC on the plus and minus strands, respectively), which together constitute 13% of all CpG states. The remaining 6% of CpG states are hydroxymethylated. Strikingly, the vast majority (98%) of hmC is asymmetric, with hmC on one DNA strand and either C or mC on the other (denoted HC / CH and HM / MH, respectively). Only 2% of all hydroxymethylated CpG sites are symmetrically modified, constituting just 0.1 % of all CpG states. The two forms of asymmetric hmC (HC / CH and HM / MH) occur at different rates, with levels of the HC / CH states being almost double those of HM / MH (3.6% and 2% of all CpG states, respectively, Fig. 4a). The genomic distributions of different hmC states are distinct (Figs. 5-7).

[0202] Primed enhancers are characterised by high levels of HC / CH

[0203] Enhancers are genetic elements that regulate the transcription of distal, cognate genes. They play a central role in cell differentiation by establishing cell-type-specific gene expression programmes23. The oxidation of mC to hmC is key to modulating enhancer activity during the early stages of differentiation24, with enhancers exhibiting among the highest enrichment of hmC in every cell type examined25. We determined the distribution of all nine CpG states comprising C, mC, and hmC, at active, primed, and poised enhancers, using the enhancer classifications of Cruz-Molina et al26(whose primary characteristics are outlined in Fig. 5a). Contrary to previous reports1524we find that hmC at active and poised enhancers is somewhat lower than at randomly sampled regions of the genome (Fig. 5a). Primed enhancers, however, contain much higher overall levels of hmC than do other enhancer types and randomly sampled regions (Fig. 5a). We next looked at how individual hmC states are distributed around enhancers. The HC / CH states are the most abundant form of hmC across all enhancer types, followed by HM / MH (Figure 5b, c). Notably, primed enhancers exhibit a sharp increase in HC / CH relative to background levels (Figure 5b). By contrast, we observe only a small increase in the HM / MH and HH states. Unlike primed enhancers, the centres of active and poised enhancers exhibit the lowest levels of both HC / CH and HM / MH. Notably, though, HC / CH levels are elevated at the flanking regions of active and poised enhancers, relative to background levels (Fig. 5b). The HM / MH states do not exhibit this trend, with levels depleted throughout active and poised enhancers (Fig. 5b, c).

[0204] Primed enhancers are also distinct with respect to methylation, containing considerably higher levels of both symmetric and hemimethylation than both poised and active enhancers. Uniquely, total methylation levels at primed enhancers exceed those of unmodified cytosine. Overall, our findings reveal a unique relationship between hydroxymethylation, methylation and primed enhancers, and that the distribution of hmC at enhancers strongly depends on its double-stranded context. hmC in the HC / CH context marks active genes.

[0205] Along with enhancers, gene bodies of the most highly expressed genes are enriched in hmC across mammalian cell types25. Gene-body hmC positively correlates with transcription across a broad range of mammalian tissues7-9and relates to maintaining transcriptional fidelity in early mouse embryogenesis11and smooth muscle cells12. The molecular basis of these relationships remains to be fully established, but hmC is suggested to play a causal role10-12. Strikingly, our results reveal that the relationship between hmC and transcription strongly depends on its double-stranded context (HC / CH, HM / MH, or HH). First, we determined the distributions of hmC states at gene bodies, transcription start sites (TSSs) and transcription termination sites (TTSs) for protein-coding genes grouped by transcript level. All hmC states exhibit unique distributions, with HC / CH showing the strongest relationship with transcription (Fig. 6). At gene bodies, HC / CH levels increase substantially with increasing transcription (Fig. 6). This effect is much stronger than for HM / MH (Fig. 6) and is negligible for HH (Fig. 6a). HC / CH levels also increase markedly with transcription around the TTS — the highest levels occurring just downstream of the most highly transcribed genes (Fig. 6c). By contrast, we see the opposite trend at the TSS, where an inverse relationship exists between transcription and HC / CH levels. Within silent genes, HC / CH levels at the TSS are higher than in any other region, whereas within active genes they are at their minimum, decreasing with increasing transcriptional activity (Fig. 6c). Notably, even several kilobases upstream and downstream of the TSS and TTS, HC / CH levels differ considerably according to transcript level, while HM / MH levels do not (Fig. 6c).

[0206] Next, we determined the correlation between the levels of each CpG state and transcription at promoters (±1 kb from the TSS), introns, and exons (Fig. 6a). We find that this relationship is strongly affected by both the CpG state and its genomic context. At promoters, all modified CpG states negatively correlate with transcription. Notably, this relationship is strongest for hemi-methylation (p = -0.48, P < 2.2 x 10-16), closely followed by symmetric methylation (p = -0.46, P < 2.2 x 10-16).

[0207] Of the hydroxymethylated states, HM / MH exhibits the strongest relationship at promoters (p = -0.35, P < 2.2 x 10-16), followed by HC / CH (p = -0.30, P < 2.2 x 10-16) and HH (p = -0.16, P < 2.2 x 10-16). At promoters, the unmodified CC state is the only CpG state that correlates positively with transcription, and whose magnitude is comparable to hemi- and symmetric methylation (p = 0.49, P < 2.2 x 1 O~16).

[0208] At gene bodies, our results reveal that the relationship between all hmC states and transcription is inverted, becoming positive in all cases. Furthermore, HC / CH dominates this relationship, exhibiting the strongest correlation with transcription of any CpG state (Fig. 6a, b). Notably, the strength of this relationship (p = 0.38, P < 2.2 x 1 O~16, Fig. 6a) is comparable to that of symmetric methylation at promoters (p = -0.46, P < 2.2 x 10~16). We find that methylation at gene bodies (both hemi- and symmetric) negatively correlates with transcription. The two forms of methylation exhibit similar strengths at introns (p = -0.26 and -0.24, respectively, P < 2.2 x 10-16) and at exons (p = -0.10 and -0.11 , respectively, P < 2.2 x 10-16). Both forms of methylation exhibit a stronger correlation at the first intron than at all other introns, with the effect being greater for symmetric methylation.

[0209] Our data reveal a key relationship between asymmetric CpG states and transcription. Thus, we next explored whether this relationship is affected by the orientation of these asymmetric states with respect to a given gene (e.g. whether hemi-methylated CpG sites correlate more strongly with transcription when the mC residue is on the coding strand). We observe no appreciable strand-dependent effect of either mC or hmC on transcription.

[0210] We then looked at whether combining information from multiple CpG states can be a more powerful predictor of a gene’s transcription level than individual states. We factored together the levels of CpG states at different genic regions (promoters, introns, or exons) in a pairwise manner and determined the resulting correlation with transcription. Of all such pairwise combinations, we find that levels of CC at promoters combined with intronic HC / CH correlate most strongly with transcription (p = 0.54, P < 2.2 x 1 O~16) — more than either state does separately, and more than any other CpG state does individually, regardless of region. Together, these data show that different forms of hmC have distinct relationships with transcription and can exhibit coherence with other CpG states.

[0211] Dppa2 / 4 are readers of symmetric hmC

[0212] Understanding how mC and hmC alter DNA-protein interactions is essential to elucidating their function. So far, few readers of hmC have been identified and well characterised27 28, leading to speculation that hmC may actually function by blocking proteins that would otherwise bind29. However, proteome-wide screens for hmC readers have not considered the double-stranded context of hmC2728. We performed a proteome-wide screen on mESC nuclear lysate, using synthetic DNA probes containing each hmC state (HC, HM, HH), along with an unmodified probe to determine relative enrichments. Additionally, we included a hemimethylated probe to validate our screen by enrichment of Uhrfl , a DNA-binding protein expressed in mESCs that binds both mC and hmC, with a preference for hemimethylated CpG sites30. All probes shared the same 68-bp sequence, with the four CpG sites within each probe containing the same CpG state in the same orientation (Fig. 7a,). As readers of cytosine modifications can be highly sensitive to sequence31, the CpG sites were flanked by differing sequences to increase sequence diversity around the modification sites. Following incubation with the mESC lysate, the bead-bound probes were washed to remove non-specific binders, treated with trypsin, and quantitatively analysed by liquid chromatography-tandem mass spectrometry (LC-MS / MS) with label-free quantitation (LFQ). MS data from each probe were averaged across four biological replicates. Specific binders of each modified probe were identified by their enrichment relative to the unmodified probe.

[0213] We first confirmed binding of Uhrfl , which in our MS data was enriched in all modified probes, most strongly and significantly for hemimethylation (Iog2 fold change = 2.6, P = 3.4 x 10-5, Fig. 7b). Having validated our screen by capturing the binding activity of a known mC and hmC reader28, we then turned to proteins enriched in the hmC probes, where we discovered new readers of hmC. Proteins most strongly enriched in the symmetric hmC probe were the epigenetic priming factors Dppa4 (Iog2 fold change = 2.2, P = 4.6 x 10-3) and Dppa2 (Iog2 fold change = 1.9, P = 3.2 x 10~2). In the HM probe, we also observed Dppa4, which was enriched to a lesser extent than for symmetric hmC (Iog2 fold change = 1 .1 , P = 1 .8 x 10-2), and the well- characterised mC / hmC reader MeCP2 (Iog2 fold change = 1.9, P = 1.5 x 10-2). We observed no enrichment of Dppa2 / 4 in the HC probe, where the only DNA binding protein enriched was Mapls (Iog2 fold change = 2.4, P = 4.0 x 10'2), a suppressor of epithelial-to-mesenchymal transition and putative tumour suppressor.

[0214] To our knowledge, this is the first time Dppa2 / 4 have been shown to recognise hmC in DNA. Dppa2 / 4 are DNA-binding proteins32that are highly expressed in mESCs33. They prime developmental promoters for future, lineage-specific activation34and play a critical role in zygotic genome activation32. Recently, two groups independently found that Dppa2 / 4 directly regulate a subset of bivalent genes in mESCs — recruiting the COMPASS and PRC2 complexes that deposit the activating (H3K4me3) and repressive (H3K27me3) marks, whose colocalisation defines the bivalent state3435. Given their essential role in embryonic development, Dppa2 / 4 became our primary focus. We next validated our MS data by Western / Wes analysis of enrichments from lysate (Fig. 7c), which confirmed Dppa2 / 4 recognises HH, and to a lesser extent, HM (and MM), as seen in the MS data. The Western / Wes data indicate that both Dppa2 and Dppa4 exhibit similar binding affinities towards different CpG states (Fig. 7c), consistent with observations that the two proteins form a heterodimer3435.

[0215] Having confirmed Dppa2 / 4 as readers of hmC, we turned to our sequencing data to determine whether bivalent genes regulated by Dppa2 / 4 are enriched in hmC. We grouped bivalent promoters according to how their bivalency depends on Dppa2 / 4, as determined in previous work34: (i) ‘Dependent’, in which Dppa2 / 4 regulate both bivalent marks, (ii) ‘Sensitive’, in which Dppa2 / 4 regulate only the H3K4me3 mark, and (iii) ‘Independent’, whose bivalency not regulated by Dppa2 / 434. We find that bivalent promoters regulated by Dppa2 / 4 contain considerably higher levels of hydroxymethylation (Fig. 7d), with mean total hmC levels of 6.8%, 5.1 %, and 2.8% across Dppa2 / 4-dependent, sensitive and independent promoters, respectively. We also find that CpG sites containing the highest levels of HH — the proteins’ principal target — are strongly enriched in Dppa2 / 4-dependent, and, to a lesser extent, in Dppa2 / 4-sensitive promoters (Fig. 7e). This is also the case for the other hmC states.

[0216] For the first time, we have decoded the mESC genome and DNA epigenome simultaneously, resolving the four genetic states (A, C, G, T) and nine epigenetic states of C, mC and hmC at CpG sites across both strands. Until now, no technique could accurately and quantitatively sequence hmC in both strands of double-stranded DNA. Thus, whether hmC at CpG sites is predominantly symmetric has remained controversial and unresolved. Some studies estimated that symmetric hmC predominates in mESCs3637— with relative levels exceeding 90%36— whereas others have suggested a higher prevalence of asymmetric hmC15’38. Our results show that the vast majority (-98%) of hydroxymethylation at CpG sites is asymmetric. Furthermore, our ability to sequence C and mC simultaneously alongside hmC has revealed that both forms of asymmetric hmC (where the CpG site contains hmC and either C or mC on the opposite strand) are relatively common and have distinct genomic distributions and relationships with transcription. Our observation that hmC is overwhelmingly asymmetric would preclude a maintenance mechanism for hmC that is analogous to mC. In maintenance methylation, mC symmetry at CpG sites is crucial; upon replication, each daughter strand carries one of the two marks at a given locus, enabling direct transmission of methylation information from the parent cell to both daughter cells. Our findings raise key questions about the fate of hmC upon replication, and how this impacts the epigenetic states of daughter cells.

[0217] We observe major differences in the behaviour of hmC, depending on whether C, mC or hmC occurs on the other strand of the CpG site. Notably, though, for a given asymmetric form of hmC (e.g. HC / CH), we observed no overall strand bias for one orientation versus the other, relative to gene orientation (i.e. whether the hmC is on the sense or antisense strand). This may reflect some equivalence between the two orientations, or perhaps that any differences are resolvable only with longer read lengths or at the single-cell level. Further work is needed to establish this.

[0218] Strikingly, we find that total hmC levels at primed enhancers — for which hmC levels have not previously been reported — vastly exceed those at poised and active enhancers. Furthermore, poised and active enhancers are somewhat depleted in hmC, though with elevated levels of HC / CH at the flanking regions. This seemingly contradicts earlier reports of strong hmC enrichment at these enhancer types1524. However, these studies defined poised enhancers differently — by the presence of H3K4me1 and absence of H3K27ac (rather than the presence of H3K27me3). This may therefore not have differentiated between poised and primed enhancers, as defined by Cruz-Molina et al26. For active enhancers, a similar depletion of hmC has indeed been observed by others24. Our data also reveal the double-stranded context of C, mC, and hmC at enhancers. We observe major differences in the distributions of all CpG states — both within enhancers, and between enhancer types. The predominant form of hmC at active, poised and primed enhancers is the HC / CH states. Notably, primed enhancers are unique in also having elevated levels of the HM / MH states — and in having comparatively higher levels of symmetric and hemi-methylation, which has not previously been identified. As we cannot determine the lifetime of CpG states, further work is needed to establish whether the abundance of CpG asymmetry at primed enhancers reflects continual cycles of methylation and active demethylation, or whether these marks are stable. At active and poised enhancers, we observe very similar distributions of CpG states, which may reflect that they are alike in most of their features (e.g., p300 and BRG1 binding, and comparable levels of nucleosomal depletion)39

[0219] As well as being the major form of hmC at enhancers, our results reveal that the HC / CH states alone largely account for the widely reported correlation7-10between gene-body hydroxymethylation and transcription. Notably, the combined levels of intronic HC / CH and promoter CC correlate more strongly with transcription than either one does separately. This indicates a relationship between CpG states in different genic regions and transcription and highlights the importance of simultaneously resolving double-stranded CpG states. Further characterisation of this relationship using machine-learning approaches may enhance its predictive power for transcription. Unlike other forms of hmC, the HC / CH states also exhibit clear differences several kilobases either side of the TSS and TTS, where levels are considerably higher even for lowly transcribed genes, indicating that these states mark broader regions of active transcription.

[0220] Our discovery that the epigenetic priming factors Dppa2 / 4 are readers of symmetric hmC indicates a direct functional role of hydroxymethylation in regulating chromatin bivalency. Dppa2 / 4 preferentially bind to GC- rich DNA, but how they are guided to specific genomic loci is unknown32. Together, our proteomics and sequencing data suggest a mechanism by which hmC (in the HH form) directly recruits Dppa2 / 4 to specific promoters, whose bivalency they regulate by recruiting the COMPASS and PRC2 complexes. Further work, beyond the scope of this study, is needed to confirm the proposed mechanism. Dppa2 / 4 expression occurs almost exclusively in early embryos32, but Dppa2 / 4 upregulation in many forms of cancer drives cell proliferation, leading to their designation as oncogenes in both mice and humans32 40. Mechanistic studies into the relationship between hmC symmetry and Dppa2 / 4 in cancer cells may therefore provide new insights into oncogenesis. More generally, our results demonstrate the importance of the double-stranded CpG site in determining the specificity of hmC readers, indicating that different forms of hmC have divergent functions. Given our observation that the asymmetric forms of hmC vastly outnumber the symmetric state, it is perhaps surprising that the strongest hmC readers we identified specifically recognise symmetric hmC. We anticipate that continued exploration of the double-stranded context of hmC — coupled with expanding the sequence space — will lead to the discovery of further readers and functional insights.

[0221] Overall, our work demonstrates that resolving C, mC or hmC on single-stranded DNA is insufficient to describe a CpG site’s epigenetic status. The behaviour of epigenetic bases is strongly dependent on whether C, mC, or hmC occurs on the opposite strand. Thus, different forms of hmC are not equivalent, and different combinations of these states at CpG sites constitute distinct units of epigenetic information.

[0222]

[0223] Table 1

[0224] References

[0225] 1 . Dor, Y. & Cedar, H. Principles of DNA methylation and their implications for biology and medicine. The Lancet 392, 777-786 (2018).

[0226] 2. Tahiliani, M. et al. Conversion of 5-Methylcytosine to 5-Hydroxymethylcytosine in Mammalian DNA by MLL Partner TET1 . Science 324, 930-935 (2009).

[0227] 3. Carell, T., Kurz, M. Q., Muller, M., Rossa, M. & Spada, F. Non-canonical Bases in the Genome: The Regulatory Information Layer in DNA. Angewandte Chemie International Edition 57, 4296-4312 (2018).

[0228] 4. Bachman, M. et al. 5-Hydroxymethylcytosine is a predominantly stable DNA modification. Nature Chemistry 6, 1049-1055 (2014).

[0229] 5. Yan, R. et al. Dynamics of DNA hydroxymethylation and methylation during mouse embryonic and germline development. Nat Genet 55, 130-143 (2023).

[0230] 6. Pfeifer, G. P., Xiong, W., Hahn, M. A. & Jin, S.-G. The role of 5-hydroxymethylcytosine in human cancer. Cell and Tissue Research 356, 631-641 (2014).

[0231] 7. He, B. et al. Tissue-specific 5-hydroxymethylcytosine landscape of the human genome. Nature Communications 12, 4249 (2021).

[0232] 8. Cui, X.-L. et al. A human tissue map of 5-hydroxymethylcytosines exhibits tissue specificity through gene and enhancer modulation. Nat Commun 11 , 6161 (2020).

[0233] 9. Lin, l.-H., Chen, Y.-F. & Hsu, M.-T. Correlated 5-Hydroxymethylcytosine (5hmC) and Gene Expression Profiles Underpin Gene and Organ-Specific Epigenetic Regulation in Adult Mouse Brain and Liver. PLOS ONE 12, e0170779 (2017).

[0234] 10. Colquitt, B. M., Allen, W. E., Barnea, G. & Lomvardas, S. Alteration of genic 5- hydroxymethylcytosine patterning in olfactory neurons correlates with changes in gene expression and cell identity. Proceedings of the National Academy of Sciences of the United States of America 110, 14682- 14687 (2013).

[0235] 11. Kang, J. et al. Simultaneous deletion of the methylcytosine oxidases Tet1 and Tet3 increases transcriptome variability in early embryogenesis. Proc. Natl. Acad. Sci. U.S.A. 112, (2015).

[0236] 12. Wu, F. et al. Spurious transcription causing innate immune responses is prevented by 5- hydroxymethylcytosine. Nat Genet 55, 100-111 (2023).

[0237] 13. Kuehner, J. N. et al. 5-hydroxymethylcytosine is dynamically regulated during forebrain organoid development and aberrantly altered in Alzheimer’s disease. Cell Reports 35, 109042 (2021).

[0238] 14. Li, W. et al. 5-Hydroxymethylcytosine signatures in circulating cell-free DNA as diagnostic biomarkers for human cancers. Cell Res 27, 1243-1257 (2017).

[0239] 15. Yu, M. et al. Base-Resolution Analysis of 5-Hydroxymethylcytosine in the Mammalian Genome. Cell 149, 1368-1380 (2012).

[0240] 16. Wu, X. & Zhang, Y. TET-mediated active DNA demethylation: mechanism, function and beyond. Nature Reviews Genetics 18, 517-534 (2017).

[0241] 17. Fullgrabe, J. et al. Simultaneous sequencing of genetic and epigenetic bases in DNA. Nat Biotechnol 41 , 1457-1464 (2023).

[0242] 18. Zheng, G., Lu, X. jun & Olson, W. K. Web 3DNA - A web server for the analysis, reconstruction, and visualization of three-dimensional nucleic-acid structures. Nucleic Acids Research 37, 240-246 (2009).

[0243] 19. Robinson, J. T. et al. Integrative genomics viewer. Nature Biotechnology 29, 24-26 (2011). 20. Laird, C. D. et al. Hairpin-bisulfite PCR: Assessing epigenetic methylation patterns on complementary strands of individual DNA molecules. Proceedings of the National Academy of Sciences of the United States of America 101 , 204-209 (2004).

[0244] 21. Incarnato, D., Krepelova, A. & Neri, F. High-throughput single nucleotide variant discovery in E14 mouse embryonic stem cells provides a new reference genome assembly. Genomics 104, 121-127 (2014).

[0245] 22. Wang, Q. et al. Pax genes in embryogenesis and oncogenesis. Journal of Cellular and Molecular Medicine 12, 2281-2294 (2008).

[0246] 23. Heinz, S., Romanoski, C. E., Benner, C. & Glass, C. K. The selection and function of cell typespecific enhancers. Nat Rev Mol Cell Biol 16, 144-154 (2015).

[0247] 24. Hon, G. C. et al. 5mC Oxidation by Tet2 Modulates Enhancer Activity and Timing of Transcriptome Reprogramming during Differentiation. Molecular Cell 56, 286-297 (2014).

[0248] 25. Lio, C.-W. J. et al. TET methylcytosine oxidases: new insights from a decade of research. J Biosci 45, 21 (2020).

[0249] 26. Cruz-Molina, S. et al. PRC2 Facilitates the Regulatory Topology Required for Poised Enhancer Function during Pluripotent Stem Cell Differentiation. Cell Stem Cell 20, 689-705. e9 (2017).

[0250] 27. Spruijt, C. G. et al. Dynamic Readers for 5-(Hydroxy)Methylcytosine and Its Oxidized Derivatives. Cell 152, 1146-1159 (2013).

[0251] 28. lurlaro, M. et al. A screen for hydroxymethylcytosine and formylcytosine binding proteins suggests functions in transcription and chromatin regulation. Genome Biology 14, R119 (2013).

[0252] 29. Pfeifer, G. P. 5-hydroxymethylcytosine stabilizes transcription by preventing aberrant initiation in gene bodies. Nat Genet 55, 2-3 (2023).

[0253] 30. Hashimoto, H. et al. Recognition and potential mechanisms for replication and erasure of cytosine hydroxymethylation. Nucleic Acids Research 40, 4841-4849 (2012).

[0254] 31. Klose, R. J. et al. DNA binding selectivity of MeCP2 due to a requirement for A / T sequences adjacent to methyl-CpG. Molecular Cell 19, 667-678 (2005).

[0255] 32. Eckersley-Maslin, M. A. Keeping your options open: insights from Dppa2 / 4 into how epigenetic priming factors promote cell plasticity. Biochemical Society Transactions 48, 2891-2902 (2020).

[0256] 33. Alda-Catalinas, C. et al. A Single-Cell Transcriptomics CRISPR-Activation Screen Identifies Epigenetic Regulators of the Zygotic Genome Activation Program. Cell Systems 11 , 25-41 .e9 (2020).

[0257] 34. Eckersley-Maslin, M. A. et al. Epigenetic priming by Dppa2 and 4 in pluripotency facilitates multilineage commitment. Nature Structural and Molecular Biology 27, 696-705 (2020).

[0258] 35. Gretarsson, K. H. & Hackett, J. A. Dppa2 and Dppa4 counteract de novo methylation to establish a permissive epigenome for development. Nature Structural and Molecular Biology 27, 706-716 (2020).

[0259] 36. Sun, Z. et al. High-Resolution Enzymatic Mapping of Genomic 5-Hydroxymethylcytosine in Mouse Embryonic Stem Cells. Cell Reports 3, 567-576 (2013).

[0260] 37. Ficz, G. et al. Dynamic regulation of 5-hydroxymethylcytosine in mouse ES cells and during differentiation. Nature 473, 398-404 (2011).

[0261] 38. Song, C.-X., Diao, J., Brunger, A. T. & Quake, S. R. Simultaneous single-molecule epigenetic imaging of DNA methylation and hydroxymethylation. Proceedings of the National Academy of Sciences 113, 4338-4343 (2016).

[0262] 39. Calo, E. & Wysocka, J. Modification of Enhancer Chromatin: What, How, and Why? Molecular Cell 49, 825-837 (2013). 40. Tung, P.-Y., Varlakhanova, N. V. & Knoepfler, P. S. Identification of DPPA4 and DPPA2 as a novel family of pluripotency-related oncogenes. Stem Cells 31 , 2330-2342 (2013).

[0263] 41. Handyside, A. H., O’Neill, G. T., Jones, M. & Hooper, M. L. Use of BRL-conditioned medium in combination with feeder layers to isolate a diploid embryonal stem cell line. Roux’s Arch Dev Biol 198, 48-56 (1989).

[0264] 42. Pfaffeneder, T. et al. Tet oxidizes thymine to 5-hydroxymethyluracil in mouse embryonic stem cell DNA. Nature Chemical Biology 10, 574-81 (2014).

[0265] 43. Martin, M. Cutadapt removes adapter sequences from high-throughput sequencing reads. EMBnet.journal; Vol 17, No 1 : Next Generation Sequencing Data Analysis (2011) doi:10.14806 / ej.17.1 .200.

[0266] 44. Shen, W., Le, S., Li, Y. & Hu, F. SeqKit: A Cross-Platform and Ultrafast Toolkit for FASTA / Q File Manipulation. PLOS ONE 11 , e0163962 (2016).

[0267] 45. Farrell, C., Thompson, M., Tosevska, A., Oyetunde, A. & Pellegrini, M. BiSulfite Bolt: A bisulfite sequencing analysis platform. GigaScience 10, giab033 (2021).

[0268] 46. Li, H. et al. The Sequence Alignment / Map format and SAMtools. Bioinformatics 25, 2078-2079 (2009).

[0269] 47. Picard, http: / / broadinstitute.github.io / picard / .

[0270] 48. Barnett, D. W., Garrison, E. K., Quinlan, A. R., Stromberg, M. P. & Marth, G. T. BamTools: a C++ API and toolkit for analyzing and managing BAM files. Bioinformatics 27, 1691-1692 (2011).

[0271] 49. R Core Team (2023). R: A Language and Environment for Statistical Computing. R Foundation for Statistical Computing, Vienna, Austria. <https: / / www.R-project.org / >.

[0272] 50. Amemiya, H. M., Kundaje, A. & Boyle, A. P. The ENCODE Blacklist: Identification of Problematic Regions of the Genome. Scientific Reports 9, 9354 (2019).

[0273] 51. Quinlan, A. R. & Hall, I. M. BEDTools: a flexible suite of utilities for comparing genomic features. Bioinformatics 26, 841-842 (2010).

[0274] 52. Ramirez, F. et al. deepTools2: a next generation web server for deep-sequencing data analysis. Nucleic Acids Research 44, W160-W165 (2016).

[0275] 53. Yu, G., Wang, L.-G. & He, Q.-Y. ChlPseeker: an R / Bioconductor package for ChIP peak annotation, comparison and visualization. Bioinformatics 31 , 2382-2383 (2015).

[0276] 54. Papachristou, E. K. et al. A quantitative mass spectrometry-based approach to monitor the dynamics of endogenous chromatin-associated protein complexes. Nature Communications 9, 2311 (2018).

[0277] 55. Benjamini, Y. & Hochberg, Y. Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing. Journal of the Royal Statistical Society: Series B (Methodological) 57, 289-300 (1995).

[0278] 56. Perez-Riverol, Y. et al. The PRIDE database resources in 2022: a hub for mass spectrometry-based proteomics evidences. Nucleic Acids Research 50, D543-D552 (2022).

[0279] 57. Salk, J., Schmitt, M. & Loeb, L. Enhancing the accuracy of next-generation sequencing for detecting rare and subclonal mutations. Nat Rev Genet 19, 269-285 (2018

Claims

1. Claims:

1. A method of sequencing a double stranded DNA molecule comprising;(i) providing a double-stranded DNA molecule having a first strand and a second strand,(ii) attaching a first hairpin adaptor to a first end and a second hairpin adaptor to a second end of the double-stranded DNA molecule to produce a double hairpin construct,(iii) cleaving the first hairpin adaptor to introduce a single strand break into the double hairpin construct,(iv) extending the 3’ end of the single strand break along the strand of the double hairpin construct to produce an extended hairpin construct that comprises first and second original strands and first and second copy strands, wherein the first and second original strands correspond to the first and second strands of the dsDNA molecule and the first and second copy strands are complementary to the first and second original strands,(v) attaching a sequencing adaptor to a free end of the extended hairpin construct,(vi) sequencing the first and second original and copy strands of the extended hairpin construct and(vii) determining the sequences of the first and second strands of the dsDNA molecule from the sequences of the first and second original strands and / or the first and second copy strands of the extended hairpin construct.

2. A method according to claim 1 wherein the method comprises; identifying the nucleotide at a position in the first or the second strand of the dsDNA molecule from the nucleotides at the corresponding positions in the sequences of (a) the first original strand (b) the second original strand (c) the first copy strand and (d) the second copy strand of the extended hairpin construct.

3. A method according to one of claim 1 or claim 2 wherein the first hairpin adaptor is cleaved at a cleavage site.

4. A method according to claim 3 wherein the cleavable site is a 2’-deoxyuridine residue, and the first hairpin adaptor is cleaved using a uracil DNA glycosylase (UDG) and a DNA glycosylase-lyase endonuclease.

5. A method according to any one of claims 1 to 4 comprising identifying the nucleotides at a CpG site in the first and second strands of the dsDNA molecule.

6. A method according to claim 5 wherein comprising identifying the cytosines at the CpG site in the first and second strands of the dsDNA molecule as C, mC or hmC.

7. A method according to claim 5 or 6 further comprising prior to sequencing the first and second original and copy strands of the extended hairpin construct;(a) protecting 5-hydroxymethylcytosines in the first and second original strands of the extended hairpin construct,(b) methylating cytosines (C) at CpG sites in the first and second copy strands that comprise methylated cytosines (mC) in the first and second original strands using a DNA methyltransferase,(c) protecting 5-methylcytosines (mC) in the first and second original and copy strands and(d) deaminating cytosines (C) in the first and second original and copy strands to convert the cytosines into uracils (U) to produce deaminated original and copy strands.

8. A method according to claim 7 wherein the DNA methylase is DNMT5.

9. A method according to one of claim 7 or claim 8 wherein the 5-hydroxymethylcytosines are protected by glycosylation.

10. .A method according to claim 9 wherein the 5-hydroxymethylcytosines are protected by a method comprising treating the 5-hydroxymethylcytosines with a p-glycosyltransferase to produce glycosyl- 5-hydroxymethylcytosines.

11. A method according to any one of claims 7 to 10 wherein the 5-methylcytosines are protected by oxidation and glycosylation.

12. A method according to claim 11 wherein the 5-methylcytosines are protected by a method comprising treating the 5-methylcytosines with TET2 to produce 5-hydroxymethylcytosines and treating the 5-hydroxymethylcytosines with a p-glycosyltransferase to produce glycosyl-5- hydroxymethylcytosines..

13. A method according to any one of claims 6 to 12 further comprising amplifying the deaminated original and copy strands to produce amplified original and copy strands.

14. A method according to any one of claims 6 to 13 wherein the method comprises sequencing the deaminated or amplified original and copy strands.

15. A merthod according to claim 14 comprising identifying the nucleotides at a CpG site in the first and second strands of the dsDNA molecule from the nucleotides at the CpG site in sequences of the deaminated or amplified original and copy strands.

16. A method according to claim 15 wherein;U or T at a CpG site in the first original strand sequence, C at the CpG site in the second original strand sequence, U or T at the CpG site in the second copy strand sequence, and;U or T at the CpG site in the first copy strand sequence, is indicative of the presence of C in the first strand and hmC in the second strand at the CpG site in the dsDNA.

17. A method according to any one of claims 15 to 16 wherein;C at the CpG site in the first original strand sequence,U or T at the CpG site in the second original strand sequence,U or T at the CpG site in the second copy strand sequence, and;U or T at the CpG site in the first copy strand sequence, is indicative of the presence of hmC in the first strand and C in the second strand at the CpG site in the double stranded DNA.

18. A method according to any one of claims 15 to 17 wherein;U or T at the CpG site in the first original strand sequence,U or T at the CpG site in the second original strand sequence,U or T at the CpG site in the second copy strand sequence, and;U or T at the CpG site in the first copy strand sequence, is indicative of the presence of C in the first strand and C in the second strand at the CpG site in the double stranded DNA.

19. A method according to any one of claims 15 to 18 wherein;U or T at the CpG site in the first original strand sequence,C at the CpG site in the second original strand sequence,C at the CpG site in the second copy strand sequence, and;U or T at the CpG site in the first copy strand sequence, is indicative of the presence of C in the first strand and mC in the second strand at the CpG site in the double stranded DNA.

20. A method according to any one of claims 15 to 19 wherein;C at the CpG site in the first original strand sequence,U or T at the CpG site in the second original strand sequence,U or T at the CpG site in the second copy strand sequence, and;C at the CpG site in the first copy strand sequence, is indicative of the presence of mC in the first strand and C in the second strand at the CpG site in the double stranded DNA.

21. A method according to any one of claims 15 to 20 whereinC at the CpG site in the first original strand sequence,C at the CpG site in the second original strand sequence,C at the CpG site in the second copy strand sequence, and;C at the CpG site in the first copy strand sequence,is indicative of the presence of mC in the first strand and mC in the second strand at the CpG site in the double stranded DNA.

22. A method according to any one of claims 15 to 21 whereinC at the CpG site in the first original strand sequence, C at the CpG site in the second original strand sequence, U or T at the CpG site in the second copy strand sequence, and;U or T at the CpG site in the first copy strand sequence, is indicative of the presence of hmC in the first strand and hmC in the second strand at the CpG site in the double stranded DNA.

23. A method according to any one of claims 15 to 22 whereinC at the CpG site in the first original strand sequence, C at the CpG site in the second original strand sequence, U or T at the CpG site in the second copy strand sequence, and;C at the CpG site in the first copy strand sequence, is indicative of the presence of mC in the first strand and hmC in the second strand at the CpG site in the double stranded DNA.

24. A method according to any one of claims 15 to 23 wherein;C at the CpG site in the first original strand sequence,C at the CpG site in the second original strand sequence,C at the CpG site in the second copy strand sequence, and;U or T at the CpG site in the first copy strand sequence, is indicative of the presence of hmC in the first strand and mC in the second strand at the CpG site in the double stranded DNA.

25. A method of mapping transcriptionally active genes in double stranded genomic DNA comprising; providing a population of fragments of genomic DNA, and identifying one or more hemi-hydroxymethylated CpG sites in the fragments, wherein the hemi-hydroxymethylated CpG sites identified in the fragments are indicative of transcriptionally active genes in the genomic DNA.

26. A method according to claim 25 wherein the one or more hemi hydroxymethylated CpG sites are identified in gene bodies in the fragments.

27. A method according to claim 25 or 26 wherein the genomic DNA fragments comprise a first strand and a second strand and an hemi hydroxymethylated CpG site comprises;(i) 5-hydroxymethylcytosine (hmC) on the first strand and cytosine (C) on the second strand or(ii) cytosine (C) on the first strand and 5-hydroxymethylcytosine (hmC) on the second strand.

28. A method according to claim 27 wherein the one or more hemi hydroxymethylated CpG sites are identified by sequencing the first strand and the second strand of a genomic DNA fragment in the population.

29. A method according to any one of claims 25 to 28 wherein the genomic DNA is eukaryotic genomic DNA.

30. A method according to any one of claims 25 to 29 wherein the population of fragments is provided by fragmenting a genomic DNA sample.

31. A method according to claim 30 wherein the sample of genomic DNA is isolated from a cell, organoid, tissue or biological fluid.

32. A method according to claim 31 wherein the sample of genomic DNA is cfDNA.

33. A method according to any one of claims 25 to 32 comprising mapping the identified hemihydroxymethylated CpG sites onto the genomic DNA.

34. A method according to claim 33 comprises identifying genes in which the hemi hydroxymethylated CpG sites are located as transcriptionally active.

35. A method according to any one of claims 25 to 34 wherein the fragments of genomic DNA are sequenced by a method according to any one of claims 1 to 19.

36. A kit comprising a first hairpin adaptor and a second hairpin adaptor wherein the first hairpin adaptor comprises a cleavage site, optionally a 2’-deoxyuridine residue.

37. A kit according to claim 36 further comprising one or more of; a sequencing adaptor, a DNA ligase, a uracil DNA glycosylase (UDG) a DNA glycosylase-lyase endonuclease a DNA polymerase, a DNA methylase, a TET methylcytosine dioxygenase, a APOBEC3A cytosine deaminase (A3A), and a UvrD helicase.

Citation Information

Patent Citations

  • Methods for differentiating modified nucleobases

    WO2023034814A1