Methods for nucleic acid analysis

The method generates a DNA library from single cells using compartmentalization and barcoded hairpin polynucleotides to simultaneously read genetic and epigenetic bases, addressing the limitations of existing sequencing technologies in capturing comprehensive DNA information.

WO2026093508A1PCT designated stage Publication Date: 2026-05-07BIOMODAL LTD
View PDF 10 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
BIOMODAL LTD
Filing Date
2025-10-31
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Existing sequencing methods fail to accurately capture both genetic and epigenetic information from DNA, particularly in single cells, due to ambiguity in base conversions and inability to distinguish 5mC from 5hmC, leading to false-positive matches and increased computational complexity.

Method used

A method involving compartmentalization of single cells, fragmentation of DNA, and attachment of barcoded hairpin polynucleotides to generate a DNA library, allowing simultaneous readout of genetic and epigenetic bases by self-annealing and extension of 3'-end hairpin polynucleotides to create complementary strands.

Benefits of technology

Enables high-accuracy detection of both genetic and epigenetic features from single cells, preserving epigenetic modifications in the original strand and allowing accurate mapping of genetic sequences, with improved sensitivity and specificity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025081496_07052026_PF_FP_ABST
    Figure EP2025081496_07052026_PF_FP_ABST
Patent Text Reader

Abstract

The invention provides methods for generating a DNA library from single cells or single nuclei. The methods comprise compartmentalising single cells or single nuclei, lysing the cells or nuclei to release genomic DNA, fragmenting the genomic DNA, introducing first and second hairpin polynucleotides to the DNA fragments, and extending a hairpin polynucleotide along the DNA strand to generate a library of barcoded DNA. The invention also provides methods for DNA sequencing and methods for mapping the location of modified cytosine residues, as well as kits for use with the methods.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] METHODS FOR NUCLEIC ACID ANALYSIS

[0002] Related Application

[0003] This present case is related to, and claims the benefit of, GB 2416167.1 filed on 01 November 2024 (01.11.2024), the contents of which are hereby incorporated by reference in their entirety.

[0004] Technical Field

[0005] This invention relates to methods for generating a DNA library from single cells. The invention also provides a method for sequencing DNA from a single cell, and methods for mapping the location of modified cytosine residues from a single cell.

[0006] Background

[0007] Information encoded in nucleic acids is fundamental to the biology of living systems. There are multiple dimensions of information stored within DNA. Genetic sequencing of the DNA bases G, C, T and A has been transformed by high-throughput sequencing approaches in the past two decades. Epigenetic information in DNA provides insights into dynamic changes in biology that are closely associated with transcriptional programs (He et al., 2022) and cell fate (Mazid et al., 2022). The combination of genetic and epigenetic information provides a more comprehensive view of biology. More recently, 5-hydroxymethylcytosine (5hmC) has emerged as an important base modification that can provide information that goes beyond 5mC and genetics (Sprujit et al., 2013, Mellen et al., 2017). Hitherto, researchers have accessed either genetic or epigenetic information, without resolving 5mC from 5hmC.

[0008] Commonly used sequencing approaches do not capture full information from both genetics and epigenetics. Next-generation sequencing directly captures the canonical bases G, C, T and A in its readout (Bentley et al., 2008). A number of base-conversion chemistries have been developed to help differentiate unmodified C from its epigenetic variants, 5mC or 5hmC. These include bisulfite-based approaches such as whole-genome bisulfite sequencing (WGBS) (Frommer et al., 1992) and bisulfite-free approaches such as enzymatic-methyl sequencing (EM-seq) (Vaisvila et al., 2021) and TET-assisted pyridine borane sequencing (Liu et al., 2019). An important shortfall of all such methods is that conversion of either the C base, or one of its epigenetic derivatives, to a II (read as T) compromises the direct detection of genetic C-to-T changes, which is the most common mutation in the mammalian genome (Cagan et al., 2022) and in cancer (Alexandrov et al., 2020). Furthermore, the ambiguity caused by C-to-T conversions in the sequenced reads being mapped against either C or T in the reference genome increases false-positive matches in the search space, consequently making computational alignment and mapping of

[0009] 008856551 converted reads slower, more expensive and less accurate (Xi et al., 2009). Also, these existing methods cannot distinguish 5mC from 5hmC in a single workflow.

[0010] Methods to distinguish 5hmC from 5mC by exclusively converting only one base have been developed for example oxidative bisulfite sequencing (Booth et al., 2013), TET-assisted pyridine borane sequencing-beta (Liu et al., 2021), Tet-assisted bisulfite sequencing (Yu et al., 2012) and APOBEC-coupled epigenetic sequencing (Schutsky et al., 2018) or by selectively copying 5mC across strands of DNA (WO 2013 / 090588; Kawasaki et al., 2017). However, some of these can involve separate, parallel workflows and sequencing to yield full information, which may increase sample requirement, cost and time taken and / or yield data that lack phased information. Combining separate datasets is fraught with difficulties that lead to additive measurement error and coverage gaps across workflows.

[0011] There is a need for methods to detect epigenetic modifications in nucleic acids, particularly for single cell applications. Accordingly, the present inventors have developed new methods for generating DNA libraries from single cells that allow for the detection of epigenetic modifications with high sensitivity.

[0012] Summary of the Invention

[0013] At its most general, the present invention relates to a method of generating a DNA library from a single cell by compartmentalising single cells or nuclei from single cells, fragmenting the DNA, and attaching a barcode to each DNA fragment, as well as a 3’-end hairpin polynucleotide. The method comprises self-annealing of the 3’-end hairpin polynucleotide, which is then extended along the DNA fragment, to provide a library of barcoded, hairpin- tagged DNA fragments comprising an original strand and a copy strand.

[0014] The barcode can be used to identify DNA from each cell. The methods of the invention are thus useful in single-cell sequencing. The methods are particularly useful for identifying both genetic and epigenetic features from single-cell DNA, since epigenetic modifications are preserved in the original DNA strand, while the copy strand allows the genetic sequence to be accurately mapped. The decoding of bases across an original strand and a copy strand provides a simultaneous readout of genetic and epigenetic bases with high accuracy.

[0015] In a first aspect, the invention provides a method for generating a DNA library, the method comprising the steps of:

[0016] (a) providing a plurality of compartments, each compartment comprising a single cell, or a nucleus from a single cell;

[0017] (b) lysing the compartmentalised cells or nuclei to release genomic DNA into each compartment;

[0018] (c) fragmenting the genomic DNA to produce double-stranded DNA fragments, and introducing a first hairpin polynucleotide to the 5’-end of each DNA strand in the

[0019] 008856557 fragments and a second hairpin polynucleotide to the 3’-end of each DNA strand in the fragments, wherein at least one of the first and second hairpin polynucleotides comprises a barcode sequence that differs between each compartment, thereby producing uniquely labelled DNA in each compartment;

[0020] (d) cleaving the first hairpin polynucleotide that is attached to the 5’-end of each DNA strand; and

[0021] (e) allowing the second hairpin polynucleotide at the 3’-end of each DNA strand to self-anneal, and extending the second hairpin polynucleotide along the DNA strand, to generate a library of barcoded DNA strands each having complementary regions covalently linked by the second hairpin polynucleotide.

[0022] The method allows the DNA within each compartment, which are derived from a single cell, or single nucleus thereof, to be uniquely labelled and distinguished from DNA derived from a different cell. At least one of the first hairpin polynucleotide and the second hairpin polynucleotide comprises a barcode sequence. This barcode sequence differs between each compartment in the plurality of compartments. In this way, the DNA in each compartment is uniquely labelled by one or more barcode sequences. By “uniquely labelled”, it is meant that the DNA in each compartment is distinguishable from the DNA in each other compartment in the plurality of compartments through one or more barcode sequences.

[0023] In some embodiments the first hairpin polynucleotide comprises the barcode sequence. In some embodiments the second hairpin polynucleotide comprises the barcode sequence. In some embodiments both the first and second hairpin polynucleotides comprise a barcode sequence, and the combination of barcode sequences allows for unique DNA labelling.

[0024] The first hairpin polynucleotide is cleaved during step (d). Typically, cleavage is selective for the first hairpin polynucleotide and the second hairpin polynucleotide is not cleaved.

[0025] Preferably, the first hairpin polynucleotide lacks the ability to fold into a hairpin structure after cleavage. Cleavage of the first hairpin polynucleotide may be anywhere within the first hairpin polynucleotide, such as away from the 3’-end of the first hairpin polynucleotide, or at the first 3’-end of the hairpin polynucleotide. Cleaving the first hairpin polynucleotide that is attached to the 5’-end of each DNA strand may be performed by removing all or a portion of the first hairpin polynucleotide that is attached to the 5’-end of each DNA strand. In some preferred embodiments, a portion of the first hairpin polynucleotide is removed from the 5’- end of each DNA strand, such as where a barcode sequence contained in the first hairpin polynucleotide remains linked to the DNA strand. In other embodiments, all of the first hairpin polynucleotide may be removed from each DNA strand, and the second hairpin polynucleotide may comprise the barcode sequence.

[0026] The second hairpin polynucleotide is covalently linked to the 3’-end of each DNA fragment and is used as a primer to synthesise a copy strand that is complementary to the genomic

[0027] 008856557 DNA. This generates hairpin-tagged DNA strands comprising an original portion where epigenetic information is retained, and a copy portion comprising genetic information (i.e. primary sequence). The hairpin-tagged DNA strands can be sequenced to accurately determine both the genetic and epigenetic bases in the genomic DNA. The simultaneous decoding of genetic and epigenetic bases within a hairpin is described in Fullgrabe et al. and WO 2022 / 02375, for example.

[0028] Preferably, the first hairpin polynucleotide comprises a first barcode sequence, wherein the first barcode sequence in a compartment is different to the first barcode sequence in each other compartment. In these embodiments, step (d) may comprise selectively cleaving the first hairpin polynucleotide at a site 5’- to the first barcode sequence. In this way, a portion of the first hairpin polynucleotide comprising the barcode sequence remains linked to the DNA fragment. The barcode sequence is useful for identifying the DNA from individual cells, such as when the DNA from multiple compartments are pooled together and sequenced, and therefore the method is useful in single-cell sequencing.

[0029] The method may comprise pooling the DNA fragments from two or more compartments after introducing the first hairpin polynucleotide, such as prior to introducing the second hairpin polynucleotide to the 3’-end of each fragment. By pooling DNA from two or more compartments, the downstream workflow can be simplified, which helps to improve throughput and productivity of the method. Efficiency is increased since the barcoded DNA from multiple cells can be processed together in a single reaction vessel.

[0030] The first hairpin polynucleotide may comprise a non-canonical DNA nucleotide 5’- to the barcode sequence, such as a deoxyuridine residue. In these embodiments, step (d) may comprise contacting the DNA strands with a DNA glycosylase, such as a DNA uracil glycosylase, and optionally an endonuclease. The non-canonical DNA nucleobase may be excised, and the resultant abasic site may be cleaved, thereby cleaving the DNA at a locus that is 5’- to the barcode sequence. In this way a portion of the first hairpin polynucleotide is removed from the 5’-end of each DNA strand, whilst retaining the barcode sequence on the DNA strand. Typically, the second hairpin polynucleotide is not cleaved or removed from the DNA strands during the methods described herein.

[0031] Each second hairpin polynucleotide may comprise a barcode sequence. In some embodiments, each first hairpin polynucleotide comprises a first barcode sequence, and each second hairpin polynucleotide comprises a second barcode sequence. Preferably, where the second hairpin polynucleotide comprises a barcode sequence, said barcode sequence differs between each compartment. When combined with a barcoded first hairpin polynucleotide, this allows each DNA strand to be labelled with two barcode sequences, which allows for superior error correction, such as when the libraries are sequenced and the sequencing data are analysed. Within each compartment, the respective barcode sequences in the first and second polynucleotides may be the same, or they may be

[0032] 008856557 different. In some embodiments, the first and second polynucleotides each comprise a barcode sequence, wherein the barcode sequences are the same.

[0033] Step (b) may comprise contacting the cells, or nuclei, or lysate thereof, with proteinase K. In this way, the genomic DNA can be released from nucleosomes.

[0034] In some embodiments, step (c) comprises introducing the first hairpin polynucleotide to each DNA strand in a fragment and extending the 3’-end of the complementary DNA strand in the fragment to introduce the second hairpin polynucleotide to the complementary DNA strand. In this way, the first hairpin polynucleotide of a strand is used as a template for DNA synthesis to introduce the second hairpin polynucleotide to the complement of said strand. This can improve the yield of hairpin labelling, such as compared to DNA ligation. The extension may be performed with a strand displacing polymerase to displace a base-paired region in the first hairpin polynucleotide of the complementary strand, such as a region that is self-annealed to form a hairpin.

[0035] Step (c) may comprise cleaving the genomic DNA with a complex comprising a transposase bound to a barcoded hairpin polynucleotide, to give double-stranded DNA fragments wherein a first hairpin polynucleotide is covalently linked to the 5’-end of each DNA strand. The transposase fragments the genomic DNA and ligates a respective first hairpin polynucleotide to each fragment in a single step, which can have higher efficiency than fragmenting and ligating DNA in two steps, for example. Preferably, the transposase is a Tn5 transposase, such as wild-type Tn5 transposase, or a variant thereof.

[0036] In some embodiments, step (c) comprises ligating a barcoded hairpin polynucleotide to the 5’-end of the DNA fragments, to give double-stranded DNA fragments wherein a first hairpin polynucleotide is covalently linked to the 5’-end of each DNA strand. In these embodiments the DNA is typically fragmented prior to ligation. Fragmentation may be by any known method, such as sonication or enzymatic fragmentation.

[0037] The plurality of compartments may be a plurality of total compartments, or a portion of a total compartments. The plurality of compartments may be the wells in a multi-well plate, such as a 96-well plate or a 384-well plate. The plurality of compartments may be all of the wells in a multi-well plate, or a portion of the wells in a multi-well plate. A portion of wells in a multiwell plate may be a row or a column of wells, for example.

[0038] In a second aspect, there is provided a method for sequencing DNA from a sample single cell, or the nucleus thereof, the method comprising the steps of:

[0039] (i) providing a plurality of single cells comprising the sample single cell, or a plurality of nuclei comprising the nucleus of the sample single cell;

[0040] (ii) distributing each single cell or single nucleus into respective compartments to provide a plurality of compartmentalised single cells or single nuclei;

[0041] 008856557 (iii) generating a DNA library from the compartmentalised cells or nuclei by a method according to the first aspect, wherein DNA from the sample single cell is uniquely labelled with one or more barcode sequences; and

[0042] (iv) sequencing the library, and identifying the DNA from the sample single cell by the one or more barcode sequences.

[0043] The method allows for single-cell sequencing with high coverage.

[0044] Preferences for the first aspect also apply to the second aspect.

[0045] Step (iii) may comprise introducing sequencing adapters to the DNA libraries, such as by ligation, such as after steps (a) to (e).

[0046] In a third aspect, there is provided a method for mapping the location of a modified cytosine residue from a sample single cell, the method comprising the steps of:

[0047] (i) providing a plurality of single cells comprising the sample single cell, or a plurality of nuclei comprising the nucleus of the sample single cell;

[0048] (ii) distributing each single cell or single nucleus into respective compartments to provide a plurality of compartmentalised single cells or single nuclei;

[0049] (iii) generating a DNA library from the compartmentalised cells or nuclei by a method according to the first aspect, wherein DNA from the sample single cell is uniquely labelled with one or more barcode sequences;

[0050] (iv) deaminating cytosine residues in the DNA library to form a treated DNA library; and

[0051] (v) sequencing the treated library, and identifying the DNA from the sample single cell by the one or more barcode sequences.

[0052] Advantageously, the method allows for the detection of modified cytosines in genomic DNA with high sensitivity and specificity.

[0053] Step (iv) may comprise deaminating canonical (unmodified) and / or modified cytosine residues. Preferably, step (iv) comprises deaminating canonical (unmodified) cytosine residues in the DNA library, such as selectively deaminating canonical (unmodified) cytosine residues over one or more modified cytosine residues.

[0054] Preferences for the first and second aspects also apply to the third aspect.

[0055] The deamination in step (iv) may be performed by contacting the DNA with a deaminase enzyme, such as an apolipoprotein B mRNA editing enzyme catalytic polypeptide (APOBEC) or a fragment thereof, or with a deaminating agent, such as bisulfite.

[0056] 008856557 The method may further comprise protecting one or more modified cytosine residues from deamination. Typically, this step is performed prior to step (iv). The protecting step may comprise contacting the DNA with an oxidising agent and / or a glucosylation agent. The protecting step may comprise contacting the DNA library with a methylcytosine dioxygenase and a p-glucosyltransferase, such as contacting the DNA library simultaneously with the methylcytosine dioxygenase and a p-glucosyltransferase.

[0057] The modified cytosine residues may be selected from 5 methylcytosine (5mc), 5-hydroxymethylcyotsine (5hmC), 5 formylcytosine (5fC), and 5-carboxylcytosine (5caC). Preferably, the modified cytosine residues are 5-methylcytosine and / or 5-hydroxymethylcyotsine.

[0058] The sequence of the treated DNA library may be indicative of the location of the modified cytosine and / or canonical (non-modified) cytosine residues in the double-stranded DNA sample. Thus, in some embodiments the method comprises determining the location of the modified cytosine residues in the double-stranded DNA based on the sequence of the treated library.

[0059] The method may comprise determining the location of the modified cytosine based on the sequence of complementary regions of the strands within the DNA library. The method may comprise identifying a cytosine:guanine base pair in complementary regions of the DNA sequences as the location of a modified cytosine residue. The method may comprise identifying a uracikguanine and / or thymine:guanine base pair in complementary regions of the DNA sequences as the location of a (canonical) cytosine residue. A modified cytosine residue in the input DNA may be distinguishable from a canonical cytosine residue in this way.

[0060] In some embodiments, the cytosine residues in the first hairpin polynucleotide, the second hairpin polynucleotide, or both, are modified cytosine residues. These residues may be protected from deamination and thereby read as cytosine during sequencing. In some embodiments, the cytosine residues in the first hairpin polynucleotide, the second hairpin polynucleotide, or both, are canonical (unmodified) cytosine residues.

[0061] In a fourth aspect, there is provided a kit for use in a method as described herein, comprising:

[0062] (a) a transposase;

[0063] (b) a DNA hairpin polynucleotide comprising a barcode sequence and a non- canonical DNA nucleotide; and

[0064] (c) proteinase K.

[0065] Preferences for the first to third aspects also apply to the fourth aspect.

[0066] 008856557 The non-canonical DNA nucleotide may be a deoxyuridine residue.

[0067] Summary of the Figures

[0068] The present invention is described with reference to the figures listed below.

[0069] Figure 1 shows a schematic for a single-cell sequencing method according to an embodiment of the invention.

[0070] Figure 2 shows the workflow in a library preparation method according to an embodiment of the invention.

[0071] Figure 3 shows LIMAP clustering projection of mouse ES cells (E14) grown under two different conditions LIF & LI F / 2i using 5mC & 5hmC information determined by a sequencing method according to an embodiment of the invention.

[0072] Figure 4 shows a genome view (IGV) of 5mC and 5hmC base calls from single cell clusters and bulk sequencing.

[0073] Figure 5 shows a LIMAP plot using binned and combined 5mC & 5hmC data (5modC) determined by a sequencing method according to an embodiment of the invention from single nuclei of -3,500 mouse cortex cells (single nuclei duet). The data is mapped onto publicly available data from bisulphite obtained from Liu et al., 2023 and Yao et al., 2023.

[0074] Figure 6 shows a reprojected LIMAP plot using binned 5hmC data and 5mC data from single nuclei of -3,500 mouse cortex cells, as determined by a sequencing method according to an embodiment of the invention. The cell types are annotated in the figure.

[0075] Figure 7 shows a multimodal LIMAP plot of neuronal and non-neuronal cell types and the differences in global levels of 5mC and 5hmC in the difference cell types, as determined by a sequencing method according to an embodiment of the invention.

[0076] Detailed Description of the Invention

[0077] In a general aspect, the invention relates to a method of generating a DNA library from a single cell by compartmentalising single cells or nuclei from single cells, fragmenting the DNA, and attaching a barcode to each DNA fragment, as well as a 3’-end hairpin polynucleotide. The method comprises self-annealing of the 3’-end hairpin polynucleotide, which is then extended along the DNA fragment, to provide a library of barcoded, hairpin- tagged DNA fragments comprising an original strand and a copy strand.

[0078] 008856557 A barcode may be a barcode sequence, such as a DNA barcode. By attaching a barcode to the DNA strands, DNA from each single cell can be identified accordingly, and the methods are therefore useful in single-cell sequencing. The methods are particularly useful for identifying both genetic and epigenetic features from single-cell DNA, since epigenetic modifications are preserved in the original DNA strand within the libraries. The decoding of bases across the original and copy strand provides a readout that allows a simultaneous readout of genetic and epigenetic bases with high accuracy.

[0079] Fullgrabe et al. and WO 2022 / 023753 describe methods of simultaneously sequencing genetic and epigenetic bases in DNA. The genetic and epigenetic information in an original DNA strand is decoded by sonicating the DNA into fragments, followed by ligation of a hairpin at both ends of each strand. The ligated hairpin is used to synthesise a copy strand that is complementary to the original DNA strand, and coupled decoding of bases across the original and copy strand provides a phased digital readout which detects genetic and epigenetic bases with high accuracy. The methods are demonstrated on bulk DNA samples.

[0080] WO 2023 / 288222 describes methods for profiling DNA methylation patterns on DNA in solution or DNA that is affixed to a solid support. The method is applied to high-input cell- free DNA, for example. US 2017 / 0002404 describes methods for identifying methylated cytosine describes the use of hairpin adapters, which is demonstrated on 30-mer dsDNA fragments. These methods are not specific to single-cell analysis, and do not allow DNA from single cells to be uniquely identified.

[0081] Chen et al., 2017 describe a method of single-cell whole-genome analysis. DNA is tagmented using Tn5, the tagmented DNA is linearly amplified into multiple copies of RNAs, and DNA is synthesised from the RNA copies by reverse transcription. The method is not suitable for analysing epigenetic bases within the genomic DNA, since the epigenetic information is lost during transcription and amplification. The method does not involve a step of partially removing a 5’-end hairpin from the tagmented DNA, or the use of a 3’-end hairpin as a primer for DNA polymerase extension.

[0082] WO 2022 / 161294 describes a method of constructing a single-cell copy number library by sorting cells and fragmenting genomic DNA using Tn5 transposase. The method uses a double-stranded adapter for Tn5 tagmentation, which is followed by PCR amplification of the tagmented DNA. The method is not suitable for analysing epigenetic bases within the genomic, since these are lost during the PCR amplification step.

[0083] US 2017 / 114390 and WO 2021 / 155057 describe methods of partitioning single cells and barcoding the nucleic acids from the single cells. These methods are not suitable for accurately decoding the epigenetic as well as genetic information from single cells.

[0084] 008856557 The present invention provides a method for generating sequencing libraries from single cells by compartmentalising individual cells or nuclei and barcoding the genomic DNA. At least one barcode is introduced via a hairpin polynucleotide that is covalently linked to an end of the DNA fragment. The barcoded DNA is covalently linked via its 3’-end to a hairpin polynucleotide, which is used as a primer to synthesise a complementary or copy strand for each original piece of DNA. By barcoding and copying the DNA within a hairpin in this way, the methods can be used to decode the genetic and epigenetic information from individual cells. The libraries can be used to identify epigenetic bases, such as modified cytosine residues, with high sensitivity and high specificity. As shown in Example 1-4, the methods of the invention may be used to sequence single cells, and in particular can simultaneously read the genetic and epigenetic sequences from single cells to reveal information that cannot be detected from bulk samples.

[0085] Methods for Generating Sequencing Library

[0086] In an aspect of the invention, there is provided a method for generating a DNA library, the method comprising the steps of:

[0087] (a) providing a plurality of compartments, each compartment comprising a single cell, or a nucleus from a single cell;

[0088] (b) lysing the compartmentalised cells or nuclei to release genomic DNA into each compartment;

[0089] (c) fragmenting the genomic DNA to produce double-stranded DNA fragments, and introducing a first hairpin polynucleotide to the 5’-end of each DNA strand in the fragments and a second hairpin polynucleotide to the 3’-end of each DNA strand in the fragments, wherein at least one of the first and second hairpin polynucleotides comprises a barcode sequence that differs between each compartment, thereby producing uniquely labelled DNA in each compartment;

[0090] (d) cleaving the first hairpin polynucleotide that is attached to the 5’-end of each DNA strand; and

[0091] (e) allowing the second hairpin polynucleotide at the 3’-end of each DNA strand to self-anneal, and extending the second hairpin polynucleotide along the DNA strand, to generate a library of barcoded DNA strands each having complementary regions covalently linked by the second hairpin polynucleotide.

[0092] Step (d) may be performed prior to step (e), or step (d) may be performed after step (e). Preferably, step (d) is performed prior to step (e). In some embodiments, steps (a) to (e) are performed in order.

[0093] 008856557 Compartmentalisation

[0094] Step (a) of the method comprises providing a plurality of compartments that each comprise a single cell or a nucleus from a single cell.

[0095] Cells for use in the method may be obtained, or obtainable from, a cell line or a tissue or a biopsy sample. Cells from a tissue may be dissociated from one another, such as by enzymatic digestion or by homogenization, for example. This may be performed prior to step (a).

[0096] Nuclei may be isolated from whole cells by any method known in the art. Whole cells may be incubated in a suitable buffer, such as a buffer comprising a detergent to permeabilise the plasma membrane. In some embodiments, whole cells are incubated in a nonionic surfactant, such as a polysorbate or 2-[4-(2,4,4-trimethylpentan-2-yl)phenoxy]ethanol (Triton X-100). The polysorbate may be polysorbate 20 (Tween 20) or polysorbate 80 (Tween 80), such as polysorbate 20. Nuclei isolation may comprise centrifugation or flow cytometry or nuclei sorting. These steps may be performed prior to step (a).

[0097] A plurality of compartments is two or more compartments, such as three or more compartments, such as four or more compartments, such as 6 or more compartments, such as 12 or more compartments.

[0098] A plurality of compartments may be a plurality of wells in a multi-well plate, or a plurality of tubes, such as microtubes.

[0099] In some embodiments, a plurality of total compartments may be provided, and the plurality of compartments may be a portion of the total compartments. For example, a multi-well plate may be provided comprising a total number of compartments or wells, and the plurality of compartments used in the method may be a fraction of the total number of compartments or wells.

[0100] A multi-well plate may also be referred to as a microplate, which has a plurality of sample wells such as 6, 12, 24, 48, 96, 384 or 1 ,536 wells. In some embodiments, the plurality of compartments is a 96-well plate or a 384-well plate, or a portion of the wells in said plates.

[0101] In some preferred embodiments the plurality of compartments is a plurality of wells in a multi-well plate, or a portion thereof. In some of these embodiments the plurality of compartments is all of the wells in the multi-well plate, and in other embodiments the plurality of compartments is a portion of the wells in a multi-well plate, such as a row or a column or another grouping of wells within the multi-well plate.

[0102] 008856557 The method may comprise distributing single cells or single nuclei into respective compartments, such as prior to step (a). The method of distribution is not particularly limited, and may be by cell sorting, such as plate sorting, or by flow cytometry or microfluidics.

[0103] Cell or Nucleus Lysis

[0104] Step (b) of the method comprises lysing the compartmentalised cells or nuclei. Lysis is performed in situ in the compartments. Genomic DNA from a single cell is released into each compartment in this step. By “releasing” genomic DNA, it is meant that the genomic DNA is made accessible for downstream processing, such as accessible for use in step (c) as described herein.

[0105] In some embodiments, step (a) comprises providing single cells in each compartment, and step (b) may comprise lysing each cell and each nucleus thereof to release the genomic DNA. In other embodiments, step (a) comprises providing a single nucleus in each compartment, and step (b) comprises lysing the nuclei to release the genomic DNA.

[0106] The cells may be permeabilised, such as to release nuclei and / or genomic DNA into each compartment, by incubating the cells in a suitable buffer. The buffer may comprise 2-[4-(2,4,4-trimethylpentan-2-yl)phenoxy]ethanol (Triton X-100), sodium dodecyl sulfate or a polysorbate. The polysorbate may be polysorbate 20 or polysorbate 80, such as polysorbate 20.

[0107] Genomic DNA may be released from isolated nuclei by incubating the nuclei in a suitable buffer, such as a buffer comprising proteinase K, and optionally in the presence of a nonionic surfactant, such as digitonin.

[0108] The genomic DNA released into each compartment in step (c) may be free, or may be substantially free, of nucleosomes. The genomic DNA may be contacted with a protease. In preferred embodiments, step (c) comprises contacting the cells or nuclei with proteinase K. The nuclei may be incubated with proteinase K in a suitable buffer, such as at 50 °C.

[0109] DNA Fragmentation and Hairpin Tagging

[0110] Step (c) of the method comprises fragmentation of the genomic DNA. Fragmentation may be enzymatic or by sonication, for example. This produces double-stranded DNA fragments. “Double-stranded DNA” as described herein comprises two strands of DNA which are complementary to one another. Unless stated otherwise, the two strands of a doublestranded DNA are not covalently linked to one another. By contrast, a single strand of DNA may comprise self-complementary regions, but such regions are not provided on separate strands of DNA.

[0111] 008856557 Step (c) further comprises introducing a first and second hairpin polynucleotide, to the 5’-end and the 3’-end of each DNA strand of each fragment, respectively. By “introducing” a hairpin polynucleotide to the DNA fragments, it is meant the hairpin polynucleotide is attached, or covalently linked, to the DNA fragment during step (c).

[0112] Introducing the first and second hairpin polynucleotide to each DNA strand may be stepwise. In some embodiments, the first hairpin polynucleotide is introduced to each DNA strand, and subsequently the second hairpin polynucleotide is introduced to each DNA strand. In some embodiments, the first hairpin polynucleotide of a DNA strand is used as a template and copied to the complementary strand within a duplex by polymerase extension to provide the second hairpin polynucleotide for the complementary strand.

[0113] A hairpin polynucleotide is a polynucleotide that is capable of folding into a hairpin structure, also known as a stem-loop structure. A hairpin has self-complementary portions that can base-pair to one another to form a duplex, also known as a stem region. The two self-complementary portions of a hairpin polynucleotide are covalently connected by an unpaired loop region. The end of the stem region that is away from the loop region may be the two ends (5’- and 3’-) of the hairpin polynucleotide, or there may be an overhang of one or more nucleotides on one strand that is unpaired when the hairpin is folded. The hairpin first and second polynucleotides are DNA hairpins. A hairpin polynucleotide may be self-annealed, such as to form a hairpin structure, or a hairpin polynucleotide may be denatured.

[0114] A hairpin polynucleotide is preferably a DNA hairpin polynucleotide.

[0115] In some embodiments, fragmentation is performed together with introducing the first hairpin polynucleotides to the fragments, such as simultaneously in a single stage. This may comprise contacting the genomic DNA with a transposase, such that the transposase cleaves the genomic DNA into fragments and tags (or covalently links) the 5’-end of each strand to a first hairpin polynucleotide. This may be referred to as tagmentation.

[0116] During tagmentation, the genomic DNA may be contacted with a complex comprising a transposase bound to a hairpin polynucleotide. The transposase may be provided as a dimer complex, wherein each transposase within the dimer is loaded with a first hairpin polynucleotide. The two hairpin polynucleotides within a transposase dimer complex may be the same, or the two hairpin polynucleotides may each have one or more constant regions that are the same, and one or more variable regions, such as one or more barcodes, that are unique to each hairpin polynucleotide.

[0117] Typically, each DNA strand that is cleaved by a transposase receives a single hairpin polynucleotide that is covalently linked to the 5’-end of said DNA strand by the transposase. The 3’-end of a DNA strand cleaved by the transposase is generally not linked to a hairpin

[0118] 008856557 polynucleotide. The transposase may generate duplex DNA fragments having a top strand and a bottom strand. A hairpin polynucleotide may be covalently linked to the 5’-end of the top and bottom strands, respectively. A gap may be present between the 3’-end of each strand and a folded hairpin polynucleotide that is linked to the complementary strand, such as a gap between the 3’-end of the top strand and 5’-end of the hairpin polynucleotide that is covalently linked to the bottom strand. The gap may be from 5 to 20 nucleotides in length, such as from 5 to 15 nucleotides, such as about 9 nucleotides.

[0119] The transposase may be Tn5 transposase. The transposase may be a wild-type Tn5 transposase, or may be a variant of that Tn5 transposase. Suitable variants are described in Hennig et al., 2018 (including a E54K and L372P double mutant), and in Kia et al., 2017 (including mutant Tn5-059, see Mutant protein expression and purification), the contents of which are incorporated by reference herein. The transposase may be an engineered transposase having reduced sequence bias compared to wild-type enzyme. Suitable engineered transposases are described in Riggs et al., 2020, the contents of which are incorporated by reference herein.

[0120] In other embodiments, fragmentation may be performed prior to introducing either the first or second hairpin polynucleotides. Fragmentation may comprise sonication of the DNA, or fragmentation may comprise enzymatic DNA digestion. Fragmentation may be followed by ligating the first and / or second hairpin polynucleotides to each DNA fragment, such as by use of a DNA ligase. In these embodiments, the first hairpin polynucleotide may be selectively ligated to the 5’-end of each strand.

[0121] A hairpin polynucleotide may have a stem region, or self-complementary region, that is 5 base pairs or more, such as 10 base pairs or more, such as 15 base pairs or more, such as 18 base pairs or more, such as 19 base pairs or more, such as 20 base pairs or more, such as 25 base pairs or more, such as 30 base pairs or more. The stem region may be 100 base pairs or less, such as 50 base pairs or less, such as 30 base pairs or less, such as 25 base pairs or less, such as 20 base pairs or less, such as 19 base pairs or less. The stem region may have a length in a range with upper and lower values as described herein, such as from 5 to 100 base pairs, including from 5 to 50 base pairs, such as from 10 to 25 base pairs, such as 15 to 20 base pairs. The stem region may be about 19 base pairs.

[0122] The stem region may comprise a transposase binding site. The transposase binding site is a sequence that facilitates loading of the hairpin polynucleotide onto a transposase. The transposase binding site may be a Tn5 transposase binding site, such as a Tn5 mosaic end. A Tn5 transposase binding site may be a duplex of 19 base pairs.

[0123] In some preferred embodiments, the first and / or second hairpin polynucleotides each comprise a 19-base pair Tn5 binding site.

[0124] 008856557 A hairpin polynucleotide may have a loop region that is 5 nucleotides or more, such as 6 nucleotides or more, such as 7 nucleotides or more, such as 8 nucleotides or more, such as 9 nucleotides or more. The loop region may be 40 nucleotides or less, such as 30 nucleotides or less, such as 20 nucleotides or less, such as 15 nucleotides or less, such as 12 nucleotides or less, such as 10 nucleotides or less, such as 9 nucleotides or less. The loop region may have a length in a range with upper and lower values as described herein, such as from 5 to 40 nucleotides, such as from 5 to 20 nucleotides, including from 6 to 15 nucleotides, such as from 8 to 12 nucleotides. The loop region may be about 9 nucleotides.

[0125] The total length of a hairpin polynucleotide may be from 20 to 150 nucleotides, such as from 20 to 100 nucleotides, such as from 20 to 60 nucleotides, such as from 30 to 60 nucleotides, such as about 47 nucleotides.

[0126] The first hairpin polynucleotide, or the second hairpin polynucleotide, or both, comprises a barcode sequence. Preferably, the first hairpin polynucleotide comprises a barcode sequence, which may be referred to as the first barcode sequence. More preferably, the first hairpin polynucleotide comprises a first barcode sequence and the second hairpin polynucleotide comprises a second barcode sequence. The first and second barcode sequences may be the same or they may be different. Preferably, the first and second barcode sequences are the same.

[0127] Each hairpin polynucleotide may comprise two or more barcode sequences. At least one of the first and second hairpin polynucleotides comprise a barcode sequence, wherein the barcode sequence differs between each compartment. Preferably, each barcode sequence in a hairpin polynucleotide is different between each compartment within the plurality of compartments.

[0128] A barcode sequence may be a nucleotide sequence for identifying DNA from a given compartment. For example, a barcode sequence may be a known (or predetermined) sequence used to label DNA within a compartment, thus allowing DNA from multiple compartments to be pooled together and sequenced in parallel. A barcode sequence may be for uniquely labelling each DNA fragment, such as where the barcode sequence is a randomised sequence. Typically, each randomised sequence is unique such that the barcode sequences in each compartment are different to the barcode sequences in another compartment, and each barcode sequence within a compartment is also different to the other sequences in the same compartment. This allows each fragmented DNA strand in a sample to be labelled with a unique barcode sequence, so that the labelled DNA strand can be distinguished from other strands and any duplicates can be removed during sequencing data analysis. The randomised barcode sequence may also be referred to as a unique molecular index (UMI).

[0129] 008856557 At least one of the first hairpin polynucleotide and the second hairpin polynucleotide comprises a barcode sequence that differs between each compartment. This produces uniquely labelled DNA in each compartment. By “uniquely labelled”, it is meant that the DNA in any compartment is distinguishable from the DNA in each other compartment in the plurality of compartments, through the one or more barcode sequences. Typically, a hairpin polynucleotide, such as the first hairpin polynucleotide is provided in each compartment having a barcode sequence that differs from the hairpin polynucleotide in each other compartment in the plurality of compartments. Preferably, a barcode that is unique to each compartment is introduced to the DNA fragments in step (c), and optionally two or more barcodes that are unique to each compartment is introduced to the DNA fragments.

[0130] In some embodiments, a barcode sequence comprises or consists of a known (or predetermined sequence) where the barcode sequence of a hairpin polynucleotide in any two compartments within the plurality of compartments is different to each other. A known barcode sequence may be from 5 to 15 nucleotides in length, such as from 5 to 10 nucleotides, such as from 6 to 8 nucleotides, such as 6 nucleotides or 8 nucleotides. Preferably, the barcode sequence is about 8 nucleotides.

[0131] In some embodiments, a barcode sequence comprises or consists of a randomised sequence, which is typically different between individual hairpin polynucleotides. A randomised barcode sequence may be from 5 to 15 nucleotides, such as from 5 to 12 nucleotides, such as from 6 to 12 nucleotides, such as 6, 9 or 12 nucleotides. A randomised barcode sequence may be about 9 nucleotides.

[0132] Preferably, each first hairpin polynucleotide comprises a barcode sequence that is a known (or predetermined) sequence used to label all of the DNA fragments within a compartment (and thereby, the DNA fragments originating from a single cell). Each first hairpin polynucleotide and / or second polynucleotide may comprise a further barcode sequence. The further barcode sequences within a compartment may have the same sequence as one another, which is different to the barcode sequences in other compartments, thus providing a second barcode for identifying DNA from said compartment or the corresponding cell. The further barcode sequences within a compartment may be different from one another, such as where the further barcode sequence comprises a unique molecular identifier (UMI) to uniquely label each DNA fragment within a compartment.

[0133] The first or second hairpin polynucleotides may comprise two or more barcode sequences, such as two barcode sequences. Each barcode may be independently selected from a pre-determined sequence and a randomised sequence. Two barcode sequences may be adjacent one another, or they may be separated by one or more nucleotides. Two barcode sequences provided within a compartment, such as in a first hairpin polynucleotide and a second hairpin polynucleotide may be the same or they may be difference.

[0134] 008856557 One or more barcode sequences in the second hairpin polynucleotide, where present, may be complementary to the barcode sequence(s) in the first hairpin polynucleotide within the same compartment. In some embodiments, the first hairpin polynucleotide comprises one or more, such as two or more barcode sequences, and the second hairpin polynucleotide comprises the complement of each barcode sequence in the first hairpin polynucleotide. Complementary barcode sequences may be read as the same sequence upon DNA sequencing.

[0135] A hairpin polynucleotide may comprise a sequencing primer region that is suitable for use with known sequencing platforms. Suitable sequencing platforms are known in the art, including low or high throughput sequencing platforms such as Sanger sequencing, Solexa- lllumina sequencing, Ligation-based sequencing (SOLiD™), pyrosequencing; strobe sequencing (SMRT™); semiconductor array sequencing (Ion Torrent™); and nanopore sequencing (ION). The sequencing primer region may be for sequencing a barcode sequence, where this is present in the hairpin polynucleotide.

[0136] A barcode sequence and / or a sequencing primer, where present, may be in the loop region or the stem region of a hairpin polynucleotide. Each of these regions may be adjacent to another of these regions, or they may be separated by one or more nucleotides.

[0137] The first hairpin polynucleotide may comprise a cleavage site for removing at least a portion of the hairpin polynucleotide from the covalently linked DNA strand. In other embodiments, the first hairpin polynucleotide may be degraded non-specifically, such as by an exonuclease.

[0138] Preferably, the first hairpin polynucleotide comprises a cleavage site, and the second hairpin polynucleotide does not comprise a cleavage site.

[0139] A cleavage site may be a non-canonical DNA nucleotide or a restriction site.

[0140] Restriction sites are DNA sequences that bind to a restriction enzyme, and which can be cleaved by the restriction enzyme. A restriction enzyme may cleave the DNA leaving blunt ends or sticky ends. The restriction site is not particularly limited, and examples are well known in the art, such as EcoRI and BamHI.

[0141] Preferably, the first hairpin polynucleotide comprises a non-canonical DNA nucleotide. The non-canonical DNA nucleotide may be recognised by a DNA glycosylase and the corresponding nucleobase may be excised from the hairpin polynucleotide by the glycosylase, such as to generate an abasic site and / or a single-strand break. Examples of suitable nucleotides include 2’-deoxyuridine, 2’-deoxy-5-hydroxymethyluridine, 2’-deoxy-5- formyluridine, and 2’-deoxy-8-oxoguanosine residues. In some embodiments, the first hairpin polynucleotide comprises a 2’-deoxyuridine residue.

[0142] 008856557 The first hairpin polynucleotide may comprise one or more non-canonical DNA nucleotides as described herein, such as one or more deoxyuridine residues. In some embodiments the first hairpin polynucleotide comprises two or more non-canonical DNA nucleotides, such as two or more deoxyuridine residues, such as two non-canonical DNA nucleotides, such as two deoxyuridine residues. References herein to a deoxynucleoside residue are to a 2’-deoxynucleoside residue unless stated otherwise. Where a hairpin polynucleotide comprises two or more non-canonical DNA nucleotides these may be contiguous, or they may be separated by one or more canonical DNA nucleotide.

[0143] In some embodiments the first hairpin polynucleotide comprises no more than two non-canonical DNA nucleotides. In some embodiments the first hairpin polynucleotide comprises no more than two deoxyuridine residues.

[0144] A cleavage site may be in the stem region or the loop region of a hairpin polynucleotide. In some preferred embodiments, the hairpin polynucleotide, such as the first hairpin polynucleotide, comprises one or more non-canonical DNA nucleotides, such as one or more deoxyuridine residues, in the loop region. Preferably, the first hairpin polynucleotide comprises a non-canonical DNA nucleotide, such as a deoxyuridine residue, and a barcode sequence, both of which are in the loop region. More preferably, the cleavage site is 5’- to the barcode sequence.

[0145] In other embodiments, the hairpin polynucleotide comprises a non-canonical DNA nucleotide in the stem region, such as in a mosaic end sequence of the hairpin polynucleotide. The deoxyuridine residue may replace a thymidine residue in the mosaic end sequence. In some embodiments, the hairpin polynucleotide comprises one or more non-canonical DNA nucleotides, such as one or more deoxyuridine residues, in the loop region.

[0146] The, or each, non-canonical DNA nucleotide of a hairpin polynucleotide may be within the 30 nucleotides at the 3’-end of the hairpin polynucleotide, such as within the 25 nucleotides at the 3’-end, such as within the 20 nucleotides at the 3’-end, such as within the 10 nucleotides at the 3’-end, such as within the 5 nucleotides at the 3’-end. In some embodiments, the non-canonical DNA nucleotide is the 3’-end nucleotide of the hairpin polynucleotide. Preferably, the non-canonical DNA nucleotide is not the 3’-end nucleotide of the hairpin polynucleotide.

[0147] A cleavage site is preferably positioned 5’- to a barcode sequence in the first hairpin polynucleotide. In these embodiments, following cleavage in step (d) of the method, the barcode sequence of the first hairpin polynucleotide remains attached to 5’-end of the covalently linked DNA strand, which may be present in addition to a second barcode sequence present on the DNA strand, such as contained within the second hairpin polynucleotide. Where the first polynucleotide comprises two or more barcode sequences, a

[0148] 008856557 cleavage site is preferably positioned 5’- to at least one of the barcode sequences. The cleavage site may be positioned between two barcode sequences.

[0149] The first hairpin polynucleotide may be cleaved within a loop region, or up to 5 nucleotides away from the loop region. The first hairpin polynucleotide may have a cleavage site, such as a non-canonical DNA nucleotide, in a loop region, or up to 5 nucleotides away from the loop region. Preferably, the first hairpin polynucleotide is cleaved within the loop region, such as where the first hairpin polynucleotide has a cleavage site at the 5’-end of the loop region. The first hairpin polynucleotide may lack the ability to self-anneal following the cleavage in step (d).

[0150] In some preferred embodiments, the first hairpin polynucleotide comprises the following regions, from 5’-end to 3’-end: a first portion of a mosaic end sequence, a cleavage site, such as a deoxyuridine residue, a barcode sequence, and a second portion of a mosaic end sequence. The first and second portions of the mosaic end sequence are complementary to one another and are capable of forming the stem region of a hairpin structure. The cleavage site and the barcode sequence are preferably in the loop region, optionally wherein the loop region consists of one or more cleavage sites and one or more barcode sequences, such as where the loop region consists of one cleavage site and a one barcode sequence.

[0151] Each hairpin polynucleotide may comprise one or more modified cytosine residues which are protected from deamination, such as one or more residues selected from 5-methylcytosine, 5-hydroxymethylcytosine, 5-formylcytosine and 5-carboxycytosine. In some embodiments, all of the cytosine residues in the hairpin polynucleotide are modified cytosine residues. In some embodiments, the hairpin polynucleotide does not contain canonical (or unmodified) cytosine residues. In other embodiments, all of the cytosine residues in the hairpin polynucleotide are canonical (unmodified) cytosine residues.

[0152] Step (c) of the methods comprise introducing a second hairpin polynucleotide to the 3’-end of each DNA strand within a fragment. The second hairpin polynucleotide may be introduced by polymerase extension, such as by extending the 3’-end of a DNA fragment along the first hairpin polynucleotide of the complementary strand within a duplex fragment. The first hairpin polynucleotide that is attached to the complementary DNA strand may thus be used as a template during polymerase extension, thereby copying the sequence of the first hairpin polynucleotide to the complementary strand. In these embodiments, the first and second hairpin polynucleotides within a compartment may be complementary to one another. In these embodiments, the second hairpin polynucleotide comprises a corresponding region that is complementary to each barcode sequence in the first hairpin polynucleotide, where the corresponding region may itself be suitable as a barcode sequence. Preferably, in these embodiments the labelled DNA produced in step (c) are blunt-ended DNA fragments, for example wherein the second hairpin polynucleotide on one

[0153] 008856557 strand within a fragment is complementary to the first hairpin polynucleotide on the other strand of the fragment.

[0154] Where fragmentation of the genomic DNA is performed using a transposase, the transposase may leave a gap at the end of the 3’-ends of each strand after tagmentation, which may be filled in during polymerase extension using the complementary strand as a template, followed by the first hairpin polynucleotide that is linked to the complementary strand as the template. Extension of the 3’-ends is typically performed along the whole of the first hairpin polynucleotide. Extension may terminate at the end of the first hairpin polynucleotide to produce blunt-ended DNA fragments. In some embodiments, a tail, such as a dA tail, may be added to the DNA strand before termination of the 3’-end extension step. Preferably, step (c) produced blunt-ended DNA fragments.

[0155] The 3’-end extension step is a polymerase extension step. The 3’-end extension step preferably comprises displacing a base-paired region in the first hairpin polynucleotide of the complementary strand. Preferably, 3’-end extension comprises unfolding of the first hairpin polynucleotide of the complementary strand, such that a duplex region is generated comprising the first hairpin polynucleotide on one strand of the duplex, and the complementary polynucleotide (which is a second hairpin polynucleotide) on the other strand.

[0156] Typically, 3’-end extension is carried out using a DNA polymerase. The extension is preferably performed using a strand displacing DNA polymerase.

[0157] A strand displacing polymerase is a polymerase having strand displacing activity. Examples include phi29 DNA polymerase and variants thereof, Klenow Fragment of DNA Polymerase I (Klenow) and Bacillus stearothermophilus DNA Polymerase I (Bst DNA polymerase) and variants thereof. Preferably, extension is performed using phi29 DNA polymerase.

[0158] The extension may be performed using a non-strand displacing polymerase. In alternative embodiments, the method does not use a non-strand displacing polymerase for extending the 3’-ends of each strand in step (c).

[0159] 3’-end extension may be carried out using deoxynucleoside triphosphates. The second hairpin polynucleotide is typically DNA hairpin.

[0160] A 3’-end extension step may be carried out with deoxynucleoside triphosphates, such as a mixture of 2’-deoxycytidine triphosphate, 2’-deoxyadenosine triphosphate, 2’-deoxythymidine triphosphate and 2’-deoxyguanidine triphosphate. A 3’-end extension step may be carried out with modified deoxynucleoside triphosphates, such as 2’-deoxy-5-methylcytidine triphosphate, 2’-deoxy-5-hydroxymethylcytidine triphosphate, 2’-deoxy-5-formylcytidine triphosphate or 2’-deoxy-5-carboxylcytidine triphosphate, such as in addition to

[0161] 008856557 2’-deoxyadenosine triphosphate, 2’-deoxythymidine triphosphate and 2’-deoxyguanidine triphosphate. In some embodiments, 3’-end extension is performed in the absence of 2’-deoxyuridine triphosphate, and / or in the absence of 2’-deoxycytidine triphosphate.

[0162] Preferably, the polymerase used for 3’-end extension is capable of DNA synthesis beyond a non-canonical DNA nucleobase, such as uracil or a modified cytosine residue. In some preferred embodiments, the polymerase is a strand displacing polymerase that is uracil tolerant.

[0163] In some embodiments, the second hairpin polynucleotide may be introduced to DNA strands by ligation or by tagmentation.

[0164] Some or all of the stages during step (c) are carried out in situ in the plurality of compartments. Preferably, fragmentation and barcoding, such as introducing at least one of the first or second hairpin polynucleotides, is carried in the plurality of compartments. More preferably, fragmentation and introducing the first hairpin polynucleotide, which is optionally performed by tagmentation, is carried out in situ in the plurality of compartments.

[0165] In some embodiments, DNA from two or more compartmentalised cells is pooled during or after step (c). Thus, the DNA may be removed from the compartments. Preferably, the DNA from two or more compartmentalised cells are pooled during step (c). The DNA may be pooled after a barcode sequence is introduced.

[0166] In some preferred embodiments, the first hairpin polynucleotide comprises a barcode sequence and the DNA from two or more compartmentalised cells are pooled together after introducing the first hairpin polynucleotide. More preferably, pooling is performed between introducing the first and second hairpin polynucleotides.

[0167] In other embodiments, DNA may be pooled at the end of step (c).

[0168] Pooling refers to combining the DNA from two or more compartments from the plurality of compartments. DNA from three or more compartments, such as four or more, such as 6 or more, such as 12 or more, may be pooled together. In some embodiments, the plurality of compartments is a plurality of wells in a multi-well plate, and DNA from a row of wells or a column of wells may be pooled together, or the DNA from the entire plate may be pooled together.

[0169] The methods described herein may be used to prepare a plurality of DNA libraries in parallel. For example, each row or each column in a multi-well plate may be provided as a plurality of compartments in step (a) of the methods, where the DNA in each well within a row or column is uniquely barcoded. The DNA within each row of wells may be pooled together, and multiple rows of wells may be processed in parallel and sequenced. In these embodiments,

[0170] 008856557 the DNA within two wells may receive the same barcode, provided they are sequenced separately. Each library may be labelled with a further barcode to distinguish between different sequencing runs or sequencing libraries.

[0171] Strand Cleavage

[0172] The methods of the present invention comprise cleaving the first hairpin polynucleotide. Cleavage may be anywhere within the first hairpin polynucleotide, including at the 3’-end of the polynucleotide and away from the 3’-end. Thus, at least a portion of the first hairpin polynucleotide is removed from the 5’-ends of each DNA strand.

[0173] Step (d) may comprise removing all or a portion of the first hairpin polynucleotide from the 5’-end of each DNA strand.

[0174] Where the first hairpin polynucleotide is used as a template for extension of the complementary strand, cleaving the first hairpin polynucleotide is preferably performed after extension. That is, step (c) of the method is preferably performed prior to step (d). In this way, the first hairpin polynucleotide may be used to generate the second hairpin polynucleotide, and the first hairpin polynucleotide is subsequently fully or partially removed.

[0175] The cleavage in step (d) is selective for the first hairpin polynucleotide, and typically the second hairpin polynucleotide is not cleaved during this step. Preferably, the second hairpin polynucleotide is not cleaved during any step in the method.

[0176] A cleaved portion of the first hairpin polynucleotide may remain hybridized to its complementary region within a duplex fragment. A cleaved portion may be removed from the DNA fragments by denaturation.

[0177] The cleavage preferably occurs at a site that is located 5’- to a barcode sequence within the first hairpin polynucleotide. In this way, a portion of the first hairpin polynucleotide is selectively removed from each DNA strand. A residual portion of the hairpin polynucleotide remains attached to the 5’-end of the DNA strands, which comprises the barcode sequence(s). Where the first hairpin polynucleotide comprises two or more barcode sequences, at least one barcode sequence is 3’- to the cleavage site. The cleavage site may be between two barcode sequences, such that one barcode sequence is removed and one barcode sequence is retained after strand cleavage. In other embodiments, the cleavage site is 5’- to all of the barcode sequences in the first hairpin polynucleotide.

[0178] In other embodiments, the first hairpin polynucleotide is cleaved at the 3’-end, thereby removing all of the first hairpin polynucleotide from each DNA strand.

[0179] 008856557 The cleavage step may comprise contacting the DNA fragments attached to a hairpin polynucleotide with an enzyme, such as a glycosylase or a restriction enzyme.

[0180] Preferably, the cleavage step comprises excising a non-canonical DNA nucleobase from the hairpin polynucleotide, such as using a glycosylase. The DNA fragments generated in step (c) may be double-stranded DNA fragments which are directly subjected to glycosylase treatment, or the double-stranded DNA fragments may be denatured and then treated with a glycosylase.

[0181] In some preferred embodiments, step (c) of the method comprises contacting the DNA strands with a glycosylase, such that a non-canonical DNA nucleobase is excised from the DNA. The glycosylase is typically a DNA glycosylase, which may be capable of cleaving the / V-glycosidic bond of a non-canonical deoxynucleotide.

[0182] The hairpin polynucleotide may be cleaved at a non-canonical DNA nucleotide. A non- canonical DNA nucleotide is any nucleotide residue other than a 2’-deoxyadenosine, 2’-deoxyguanosine, 2’-deoxycytidine or thymidine residue. The non-canonical DNA nucleotide may comprise a nucleobase, such as a modified nucleobase, that is other than canonical adenine, guanine, cytidine or thymine. The non-canonical DNA nucleotide may comprise a sugar modification. Cleaving DNA “at a non-canonical nucleotide” includes cleaving the DNA between the non-canonical DNA nucleotide and an adjacent nucleotide, such as at a phosphodiester bond between the non-canonical DNA nucleotide and a neighbouring nucleotide that is immediately 5’- or 3’- to the non-canonical DNA nucleotide.

[0183] Suitable non-canonical DNA nucleotides include those comprising the following residues: deoxyuridine (2’-deoxyuridine), 5-hydroxymethyluridine (2’-deoxy-5-hydroxymethyluridine), 5-formyluridine (2’-deoxy-5-formyluridine) and 8-oxoguanine (2’-deoxy-8-oxoguanidine).

[0184] The glycosylase may be a monofunctional glycosylase. Examples include uracil DNA glycosylase (including UDG and UNG), and single-strand selective monofunctional uracil DNA glycosylase (SMUG1).

[0185] The glycosylase may be a bifunctional glycosylase that has glycosylase activity and AP lyase activity. Examples include oxoguanine glycosylases, such as OGG1 and OGG2.

[0186] The DNA may be contacted with an endonuclease, such as in addition to a glycosylase, such as in addition to a monofunctional glycosylase. The endonuclease may be an AP endonuclease. An “AP endonuclease” as referred to herein is an enzyme having endonuclease or lyase activity at an abasic site. Examples include Endonuclease VIII, APE1, and APE2.

[0187] 008856557 The DNA may be contacted with a glycosylase and an endonuclease in a single step, such the DNA is contacted with an enzyme mixture. The DNA may be contacted sequentially with a glycosylase and then an endonuclease, and optionally with purification of the DNA in between these steps.

[0188] In some embodiments, the DNA is contacted with uracil DNA glycosylase, such as UDG or UNG, and endonuclease VIII. The DNA may be contacted with a mixture of uracil DNA glycosylase (such as UDG or UNG) and endonuclease VIII, or the DNA may be contacted with uracil DNA glycosylase followed by endonuclease VIII.

[0189] Hairpin Extension

[0190] The methods described herein comprise a step of allowing the second hairpin polynucleotide at the 3’-ends each DNA strand to self-anneal, and subsequently extending the self-annealed hairpin long the covalently attached DNA (which may be referred to as a template strand or original strand) to generate tagged DNA strands each having an original region, a complementary region, which regions are covalently linked by the second hairpin polynucleotide. By “extending the second polynucleotide”, it is meant that polymerase extension is performed using the second hairpin polynucleotide as a primer.

[0191] The DNA strands may be denatured prior to hairpin self-annealing. Denaturation may be full denaturation to provide single-stranded DNA, or denaturation may be partial, provided that self-annealing of the second hairpin polynucleotide is possible.

[0192] Denaturing the DNA strands may be performed after cleaving the first hairpin polynucleotide (i.e. step (d) is performed prior to step (e)), or denaturing may be performed before cleaving the first hairpin polynucleotide. Preferably, step (d) is performed before step (e).

[0193] Methods of denaturing DNA are known in the art. DNA may be denatured with heat, or DNA may be denatured by treatment with a denaturing agent. Suitable denaturing agents include sodium hydroxide and formamide.

[0194] The second hairpin polynucleotide is self-annealed, such as after denaturation. The hairpin may be self-annealed to form a hairpin structure (stem-loop structure). The DNA may be subjected to conditions to facilitate self-annealing, or self-annealing may be spontaneous. Conditions for facilitating self-annealing may depend on the conditions used for denaturation. For example, DNA may be cooled following heat denaturation to promote self-annealing, or where a denaturing agent is used this may be quenched or removed from the DNA to promote self-annealing. Where DNA is denatured using sodium hydroxide, the sodium hydroxide may be neutralised with an acid to promote self-annealing.

[0195] 008856557 Following self-annealing of the second hairpin polynucleotide, step (e) of the methods includes a stage of extending the second hairpin polynucleotide along the covalently linked DNA strand. The DNA is typically extended from the stem region of the self-annealed hairpin, using the fragmented DNA strand as a template. This generates tagged DNA strands having two complementary regions linked at one end by the second hairpin polynucleotide.

[0196] Hairpin extension may be performed by a DNA polymerase. The DNA polymerase may be a high-fidelity polymerase. The polymerase may be a strand displacing polymerase or a nonstrand displacing DNA polymerase. Suitable polymerases include Klenow Fragment of DNA Polymerase I (Klenow), phi29 DNA polymerase, Taq Polymerase, Bst Polymerase, Q5 polymerase, Phusion polymerase, T4 DNA polymerase and KAPA HiFi polymerase. Preferably, the DNA polymerase is Klenow Fragment of DNA Polymerase I or phi29 DNA polymerase. More preferably, the DNA polymerase is Klenow Fragment of DNA Polymerase I, such as Klenow fragment 3’-5’ exo-.

[0197] Step (e) of the method generates a library of barcoded DNA strands. The DNA strands are labelled with one or more barcode sequences, such that the DNA from each compartment (which originate from a single cell or nucleus) can be distinguished from the DNA from each other compartment (which originates from a different single cell or nucleus). In other words, the DNA strands from each compartment are uniquely labelled. Preferably, the barcode is provided as a portion of the first hairpin polynucleotide, wherein the first hairpin polynucleotide is partially removed during step (d), such as by cleavage at a site that is 5’- to the barcode sequence. In other embodiments, the barcode for the DNA strands within a library may be provided by the second hairpin polynucleotide. In some embodiments, the DNA strands are uniquely labelled by a combination of barcodes in the first and second hairpin polynucleotide.

[0198] Further Preferences

[0199] The DNA used as an input for the methods described herein is a genomic sequence. The sequence may comprise all or part of the sequence of a gene, including exons, introns and / or upstream or downstream regulatory elements, or the sequence may comprise genomic sequence that is not associated with a gene. In some embodiments, the nucleic acid sequence may comprise one or more CpG islands.

[0200] The methods described herein may further comprise a step of ligating a sequencing adapter to the barcoded and hairpin-tagged DNA, such as the end away from the second hairpin polynucleotide. This step is preferably performed after step (e).

[0201] A sequencing adapter may comprise a double-stranded portion which is ligated to the free ends of the hairpin-tagged DNA. The sequencing adapter is preferably a Y-shaped adapter.

[0202] 008856557 Y-shaped adapters, or forkhead adapters, typically comprise a double-stranded region and non-complementary regions. In other embodiments, the method may comprise ligating two single-stranded adapters, to respective free ends of the hairpin-tagged DNA.

[0203] The cells used in the method may be provided within a sample. The cells may be mammalian cells, preferably human cells.

[0204] Suitable samples include isolated samples of cells and tissue samples, such as biopsies, as well as blood samples. A sample may be a blood sample, from which circulating free DNA (cfDNA) or circulating tumour DNA (ctDNA) may be extracted.

[0205] Modified cytosine residues, including 5mC, have been detected in a range of cell types including embryonic stem cells (ESCS) and neural cells. Suitable cells include somatic and germ-line cells.

[0206] Suitable cells may be at any stage of development, including fully or partially differentiated cells or non-differentiated or pluripotent cells, including stem cells, such as adult or somatic stem cells, fetal stem cells or embryonic stem cells. In some embodiments, the cells are embryonic stem cells.

[0207] Suitable cells also include induced pluripotent stem cells (iPSCs), which may be derived from any type of somatic cell in accordance with standard techniques.

[0208] For example, the cells may be neural cells, including neurons and glial cells, contractile muscle cells, smooth muscle cells, liver cells, hormone synthesising cells, sebaceous cells, pancreatic islet cells, adrenal cortex cells, fibroblasts, keratinocytes, endothelial and urothelial cells, osteocytes, and chondrocytes. In some embodiments, the cells are neural cells.

[0209] Suitable cells include disease-associated cells, for example cancer cells, such as carcinoma, sarcoma, lymphoma, blastoma or germ line tumour cells. Suitable cells include cells with the genotype of a genetic disorder such as Huntington’s disease, cystic fibrosis, sickle cell disease, phenylketonuria, Down syndrome or Marfan syndrome.

[0210] Methods of extracting and isolating genomic DNA from samples of cells are well-known in the art. For example, genomic DNA may be isolated using any convenient isolation technique, such as phenol / chloroform extraction and alcohol precipitation, caesium chloride density gradient centrifugation, solid-phase anion-exchange chromatography and silica gelbased techniques.

[0211] 008856557 In some embodiments, whole genomic DNA isolated from cells may be used directly for downstream processing after isolation. In other embodiments, the isolated genomic DNA may be subjected to further preparation steps, such as prior to performing step (c).

[0212] In some embodiments, a fraction of the genomic DNA released from each may be used as described herein. Suitable fractions of genomic DNA may be based on size or other criteria. In some embodiments, a fraction of genomic DNA which is enriched for CpG islands (CGIs) may be used as described herein.

[0213] Following fractionation and / or other preparation steps, the genomic DNA may be purified by any convenient technique.

[0214] Following preparation, the DNA may be provided in a suitable form for further treatment as described herein. For example, the DNA may be in aqueous solution in the absence of buffers before treatment as described herein.

[0215] The DNA sample may be divided into two, three, four or more separate portions, each of which may be independently treated and sequenced, such as described herein.

[0216] In some aspects, the invention provides a method for labelling DNA from single cells, the method comprising:

[0217] (a) providing a plurality of compartments, each compartment comprising a single cell, or a nucleus from a single cell;

[0218] (b) lysing the compartmentalised cells or nuclei to release genomic DNA into each compartment;

[0219] (c) fragmenting the genomic DNA to produce double-stranded DNA fragments, and introducing a first hairpin polynucleotide to the 5’-end of each DNA strand in a fragment wherein the first hairpin polynucleotide comprises a barcode sequence that differs between each compartment, thereby producing uniquely labelled DNA in each compartment; and

[0220] (d) pooling the 5’-end labelled DNA fragments to provide a library of labelled DNA.

[0221] The method may comprise one or more steps in a method described herein.

[0222] Step (c) may comprise contacting the genomic DNA with a transposase as described herein.

[0223] Step (c) may further comprise introducing a second hairpin polynucleotide to the 3’-end of each DNA strand in a fragment, such as by extending the first hairpin polynucleotide of the complementary strand within the fragment. The method may further comprise selectively cleaving the first hairpin polynucleotide at a site 5’- to the barcode sequence, such as after 3’-end extension.

[0224] 008856557 In some aspects, there is provided a DNA library obtainable from a plurality of single cells, wherein each DNA strand in the library comprises a genomic region flanked by a first hairpin polynucleotide and a second hairpin polynucleotide, and wherein the first hairpin polynucleotide and optionally the second hairpin polynucleotide comprises a barcode sequence such that DNA from each single cell is uniquely labelled. The DNA in the library may be an intermediate that is obtainable by the methods described herein.

[0225] A genomic region may be a genomic DNA sample, or a fragment thereof. Preferably, the DNA strands each comprise a fragment of genomic DNA, such as a fragment generated by tag mentation.

[0226] The first hairpin polynucleotide may be obtained, or may be obtainable, by transposition. Preferably, the first hairpin polynucleotide comprises a cleavage site. The cleavage site may be a non-canonical DNA nucleotide, such as a deoxyuridine residue.

[0227] The second hairpin polynucleotide may be obtained, or may be obtainable, by 3’-end extension of the genomic region using a first hairpin polynucleotide of the complementary strand as a template. Preferably, the second hairpin polynucleotide does not comprise a cleavage site. For example, the second hairpin polynucleotide may not comprise a deoxyuridine residue.

[0228] In some embodiments the library is single-stranded library. In some embodiments the library is a double-stranded library. A library of double-stranded DNA may be a library of blunt-ended DNA.

[0229] Typically, the first and second hairpin polynucleotides are each capable of folding into a stem-loop structure.

[0230] In some aspects, there is provided a DNA library obtainable from DNA from a plurality of single cells, wherein each DNA strand in the library comprises a genomic region flanked by a 5’-end barcode sequence and a 3’-end hairpin polynucleotide, wherein the DNA from each single cell is uniquely labelled by the 5’-end barcode sequence. The DNA in the library may be an intermediate that is obtainable by the methods described herein.

[0231] The 3’-end hairpin polynucleotide may be the second hairpin polynucleotide described herein.

[0232] The genomic region may be a genomic DNA sample, or a fragment thereof. Preferably, the DNA strands each comprise a fragment of genomic DNA, such as a fragment generated by tagmentation. In some embodiments the library is single-stranded library. In some embodiments the library is a double-stranded library. A library of double-stranded DNA may be a library of blunt-ended DNA.

[0233] 008856557 Methods for DNA Sequencing

[0234] The DNA libraries generated by a method described herein may be a sequencing library that is compatible with a suitable sequencing platform. Thus, in some aspects there is provided a method of sequencing DNA from a sample single cell, or the nucleus thereof, the method comprising the steps of:

[0235] (i) providing a plurality of single cells comprising the sample single cell, or a plurality of nuclei comprising the nucleus of the sample single cell;

[0236] (ii) distributing each single cell or single nucleus into respective compartments to provide a plurality of compartmentalised single cells or single nuclei;

[0237] (iii) generating a DNA library from the compartmentalised cells or nuclei by a method described herein, wherein DNA from the sample single cell is uniquely labelled with one or more barcode sequences; and

[0238] (iv) sequencing the library, and identifying the DNA from the sample single cell by the one or more barcode sequences.

[0239] Preferably, step (iii) comprises introducing sequencing adapters to the DNA within the library. A sequencing adapter may be a Y-shaped or a forkhead adapter, which may be introduced by ligation, for example.

[0240] The sample single cell, or nucleus thereof, is provided within a plurality of single cells or single nuclei. The plurality of single cells may be a collection of cells within a total cell sample, such as a portion thereof, wherein the DNA from each single cell in the portion of cells is labelled with a different barcode. Within the total cell sample, two or more cells may be labelled with same barcode, provided that said cells are sequenced separately.

[0241] A library is optionally amplified before sequencing. In some embodiments, the method comprises no more than 24 PCR cycles prior to sequencing, such as no more than 18 cycles, such as no more than 16 cycles, such as no more than 12 cycles. Preferably, PCR amplification is performed after pooling DNA from two or more compartments together.

[0242] The sequencing libraries generated by the methods described herein may be compatible with a suitable sequencing platform.

[0243] Sequencing may be carried out by any suitable sequencing method or platform, such as Sanger sequencing, Solexa-lllumina sequencing, Ligation-based sequencing (SOLiD™), pyrosequencing; strobe sequencing (SMRTTM); semiconductor array sequencing (Ion Torrent™); and nanopore sequencing (ION).

[0244] A library preparation method as described herein generates a library of hairpin-tagged DNA whereby the epigenetic information is retained in DNA strands. These libraries can be used

[0245] 008856557 to map the location of one or more epigenetic bases, such as a modified cytosine residue. Thus, in some aspects, there is provided a method of mapping the location of a modified cytosine residue from a sample single cell, the method comprising the steps of:

[0246] (i) providing a plurality of single cells comprising the sample single cell, or a plurality of nuclei comprising the nucleus of the sample single cell;

[0247] (ii) distributing each single cell nucleus into respective compartments to provide a plurality of compartmentalised single cells or single nuclei;

[0248] (iii) generating a DNA library from the compartmentalised cells or nuclei by a method described herein, wherein DNA from the sample single cell is uniquely labelled with one or more barcode sequences;

[0249] (iv) deaminating cytosine residues in the DNA library to form a treated DNA library; and

[0250] (v) sequencing the treated library, and identifying the DNA from the sample single cell by the one or more barcode sequences.

[0251] The modified cytosine residues may be selected from 5-methylcytosine (5mC), 5-hydroxymethylcytosine (5hmC), 5-formylcytosine (5fC) and 5-carboxycytosine (5caC). Preferably, the modified cytosine residue is selected from 5mC and 5hmC. In some embodiments the modified cytosine residue is 5mC. In some embodiments the modified cytosine residue is 5hmC.

[0252] The method comprises deaminating cytosine residues. The modified (canonical) and / or unmodified cytosine residues in the DNA may be deaminated during step (iv). The deamination may be selective for cytosine residues (unmodified, or canonical cytosine residues) over a modified cytosine residue, or the deamination may be selective for a modified cytosine residue, such as over another modified cytosine residue and / or over an unmodified (canonical) cytosine residue. The sequence of the DNA library may thus be indicative of the location of the modified cytosine residues in the double-stranded DNA sample. Selectivity may be achieved by converting one or more modified cytosine residues as described herein, typically prior to deamination.

[0253] The treated library is sequenced. Within the sequence of each hairpin-tagged DNA strand, a complementary region, or a base pair, can be used to determine the identity of the DNA base in the original DNA strand. In some embodiments, a cytosine:guanine base pair in complementary regions of the sequenced DNA (corresponding to a base pair in a hairpin- tagged DNA strand) is indicative of a modified cytosine residue, such as 5mC and / or 5hmC. In some embodiments, a uracikguanine or thymine:guanine base pair in complementary regions of the sequenced DNA is indicative of a (canonical) cytosine residue.

[0254] Unless stated otherwise, references herein to a “cytosine residue” refer to an unmodified, or canonical cytosine residue. A cytosine residue may be distinguished from a modified cytosine residue, such as 5mC, 5hmC, 5fC or 5caC.

[0255] 008856557 Step (iv) in the mapping method may comprise deaminating unmodified cytosine residues selectively over one or more modified cytosine residues, such as selectively over 5mC and 5hmC. The selectivity for cytosine over a modified cytosine residue may be 2-fold or more, such as 5-fold or more, such as 10-fold or more, such as 20-fold or more, such as 50-fold or more, such as 100-fold or more.

[0256] Deamination may be performed by contacting the DNA strands in the library with a deaminase enzyme. The deaminase enzyme may be a cytosine deaminase or a variant thereof, such as an apolipoprotein B mRNA editing enzyme catalytic polypeptide (APOBEC) or a variant or a fragment thereof, such as APOBEC3A or a variant or a fragment thereof. A cytosine deaminase variant may selectively deaminate an unmodified (canonical) cytosine residue, or a modified cytosine residue, such as a 5mC or 5hmC residue.

[0257] Deamination may be performed by contacting the DNA strands in a library with a deamination agent, such as bisulfite.

[0258] Deamination may be carried out in the presence of a denaturing agent, such as a helicase, sodium hydroxide, or formamide. In some embodiments, denaturation, such as using sodium hydroxide or formamide, is carried out prior to deamination.

[0259] The method may comprise protecting modified cytosine residues from deamination, such as by contacting the DNA strands with an oxidising agent and a glucosylation agent. Protection in this way may be performed prior to cytosine deamination in step (iv). The oxidising agent may be an agent that is capable of oxidising one or more of 5-methylcytosine, 5-hydroxymethylcytosine and 5-formylcytosine.

[0260] The oxidizing agent may be a methylcytosine dioxygenase, such as a ten-eleven translocation (TET) enzyme, such as TET 1 , TET2 or TET3. The oxidizing agent may be a metal oxide, such as a ruthenate, such as potassium ruthenate.

[0261] The glucosylating agent may be an agent that glycosylates 5-hydroxymethylcytosine (5hmC), such as beta-glucosyltransferase (PGT).

[0262] In some preferred embodiments, the DNA strands are contacted with a mixture comprising the oxidising agent and a glucosylating agent, such as a mixture comprising a TET enzyme and beta-glucosyltransferase (PGT).

[0263] In some embodiments, the method comprises: protecting modified cytosine (5mC and / or 5hmC) residues from deamination, such as by contacting the DNA strand with a methylcytosine dioxygenase and a beta-glucosyltransferase (PGT); and

[0264] 008856557 deaminating unmodified cytosine residues in the DNA, such as by contacting the DNA strand with a deaminase enzyme, optionally in the presence of a denaturing agent such as a helicase.

[0265] The method may comprise contacting the DNA strands with a DNA methyltransferase. A methyltransferase may be DNMT1 or DNMT3. In this way, a methylcytosine residue in the template DNA is marked by a methylcytosine residue in the synthesised complementary DNA at a proximal (adjacent) position, which can be detected upon sequencing and used for improved error correction. Preferably, the methylation step is performed prior to cytosine deamination in step (iv). The methylcytosine residue in the synthesised complementary strand may be distinguished from a canonical cytosine residue, such as by deaminating cytosine residues after treatment with the DNA methyltransferase.

[0266] 5hmC residues may be protected from DNA methyltransferase activity, such as by contacting the DNA strands with a beta-glucosyltransferase (PGT) prior to the DNA methyltransferase. Subsequently, the methylated bases (in the original strand and the copied methylation sites) may be protected from deamination, such as by oxidation and glucosylation as described herein.

[0267] In some embodiments, the method comprises: glucosylating a 5hmC residue in a DNA strand, such as by contacting the DNA strand with a beta-glucosyltransferase (PGT); copying a methylated cytosine residue across the hairpin DNA strand, such as by contacting the DNA strand with a DNA methyltransferase; protecting modified cytosine (5mC and / or 5hmC) residues from deamination, such as by contacting the DNA strand with a methylcytosine dioxygenase and a beta-glucosyltransferase (PGT); and deaminating cytosine residues in the DNA, such as by contacting the DNA strand with a deaminase enzyme, optionally in the presence of a denaturing agent such as a helicase.

[0268] In some embodiments, a portion or all of the cytosine residues in a hairpin polynucleotide, and / or any adapters, where present, are methylated or modified (as 5mC, 5hmC, 5fC or 5caC, for example). This step may protect these residues within these regions from deamination, such as by a cytosine deaminase enzyme.

[0269] A DNA library prepared according by a method herein may be sequenced using any convenient low or high throughput sequencing technique or platform. Suitable protocols, reagents and apparatus for nucleic acid sequencing are well known in the art and are available commercially.

[0270] 008856557 In some embodiments, the DNA library or a portion thereof may be amplified before sequencing. Preferably, the library is amplified after tagging with the hairpin polynucleotide, such as after extension of the second hairpin polynucleotide and after sequencing adapter ligation, and more preferably after any deamination and / or protection steps as described herein. Suitable methods for the amplification of nucleic acids are well known in the art. Following amplification, the amplified portions of the population of nucleic acids may be sequenced.

[0271] Nucleotide sequences obtained through sequencing may be compared and the residues the sequences fragments may be identified using computer-based sequence analysis.

[0272] Computer-based sequence analysis may be performed using any convenient computer system and software. A typical computer system comprises a central processing unit (CPU), input means, output means and data storage means (such as RAM). A monitor or other image display is preferably provided. The computer system may be operably linked to a DNA and / or RNA sequencer.

[0273] The methods of the invention allow for sequencing results obtained from an original (or template) strand to be compared against a nucleotide sequence obtained from the corresponding copy strand. A comparison between these sequences can show where modified nucleotides are present, such as modified cytosine residues. In some embodiments, sequencing results obtained from the template strand may be compared against a reference nucleotide sequence. A reference nucleotide sequence may be obtained from a database, or may be obtained by sequencing an untreated sample that has not undergone one or more treatment steps described herein. A DNA sample may be divided into two or more portions. The DNA in each portion may be sequenced and compared against each other to allow for identification of base modification sites in the treated portion.

[0274] In some embodiments, the sequencing methods described herein further comprises: determining the identity of a first base in the template strand, and the identity of a second base in the copy strand, wherein the second base is at or proximal to the corresponding locus of the first base; and determining a value of a true base at the locus of the first base in the template strand.

[0275] A proximal locus may be an adjacent nucleotide.

[0276] The methods can be used to sequence DNA with high accuracy. Sequencing and analysis may be carried out as described in WO 2022 / 023753 or Fullgrabe et al., the contents of which are incorporated by reference, such as to leverage internal logic comparisons of two-base sequencing methods and systems. The method may also be used in 4-base

[0277] 008856557 genome contexts and expanded 5- and 6-base genome contexts, as described in WO 2022 / 023753.

[0278] Determining the value of a true base may be carried out using a computer comprising a processor, a memory and instructions stored thereupon that, when executed, determines the value of the true base.

[0279] Kits

[0280] In some aspects, the invention provides a kit comprising:

[0281] (a) a transposase;

[0282] (b) a DNA hairpin polynucleotide comprising a barcode sequence and a non- canonical DNA nucleotide; and

[0283] (c) proteinase K.

[0284] The transposase and hairpin polynucleotide may be as described herein. The transposase may be a Tn5 transposase. The hairpin polynucleotide may be a first hairpin polynucleotide described herein, and preferably comprises one or more of a non-canonical DNA nucleotide and a barcode sequence.

[0285] The kit may comprise a plurality of hairpin polynucleotides, such as hairpin polynucleotides with different barcode sequences.

[0286] The kit may further comprise a DNA glycosylase as described herein, such as uracil DNA glycosylase. The kit may comprise an endonuclease as described herein, such as an AP endonuclease, such as endonuclease VIII.

[0287] The kit may comprise one or more polymerases, such as a strand-displacing polymerase. The polymerase may be a DNA polymerase or an RNA polymerase. The polymerase may be a thermostable polymerase, for example a high discrimination polymerase. Preferably, the polymerase is capable of DNA synthesis past a labelled cytosine residue and / or a uracil residue. The polymerase may be Klenow Fragment of DNA Polymerase I (such as exo-).

[0288] The kit may be provided in a suitable container and / or with suitable packaging.

[0289] The kit may include instructions for use, e.g., written instructions on how to use the kit in a method of detecting 5mC in a nucleic acid sample.

[0290] A kit may further comprise a population of control nucleic acids comprising one or canonical or modified residues, for example adenine (A), guanine (G), thymine (T), uracil (II), cytosine (C), 5-methylcytosine (5mC) and 5-hydroxymethylcytosine (5hmC). In some embodiments,

[0291] 008856557 the population of control nucleic acids may be divided into one or more portions, each portion comprising a different control residue.

[0292] The kit may include instructions for use in a method of generating a DNA library as described herein.

[0293] A kit may include one or more other reagents required for the method, such as buffer solutions, deoxynucleotide triphosphates (dNTPs), primers, and sequencing and other reagents. A kit for use in generating tagged DNA may include one or more articles and / or reagents for performance of the method, such as means for providing the test sample itself, including DNA isolation and purification reagents, and sample handling containers (such components generally being sterile).

[0294] A kit may include sequencing adapters and one or more reagents for the attachment of sequencing adapters to the ends of isolated nucleic acids, such as T4 ligase.

[0295] A kit may include one or more reagents for the amplification of a population of nucleic acids using the amplification primers. Suitable reagents may include dNTPs and an appropriate buffer.

[0296] Other Embodiments

[0297] Each and every compatible combination of the embodiments described above is explicitly disclosed herein, as if each and every combination was individually and explicitly recited. Various further aspects and embodiments of the present invention will be apparent to those skilled in the art in view of the present disclosure.

[0298] “and / or” where used herein is to be taken as specific disclosure of each of the two specified features or components with or without the other. For example, “A and / or B” is to be taken as specific disclosure of each of (i) A, (ii) B and (iii) A and B, just as if each is set out individually herein.

[0299] Unless context dictates otherwise, the descriptions and definitions of the features set out above are not limited to any particular aspect or embodiment of the invention and apply equally to all aspects and embodiments which are described.

[0300] Certain aspects and embodiments of the invention will now be illustrated by way of example and with reference to the figures described above.

[0301] 008856557 Examples

[0302] The methods of the invention are exemplified in the following examples.

[0303] Example 1 - Cell Sorting and Tagmentation

[0304] A library of barcoded DNA may be generated from a plurality of single cells as shown in Figure 1.

[0305] In step 1, nuclei isolated from single cells are compartmentalised by a method known in the art, such as by microfluidics, on a multi-well chip, by DLP spotting or by plate sorting. See for instance Conte et al. and Adey. After compartmentalisation, the nuclei membrane is permeabilised and nucleosomes are disrupted to release genomic DNA into each compartment. This step may comprise contacting the DNA with proteinase K.

[0306] In step 2, the genomic DNA is tagmented (see e.g. Example 2) using a barcoded first hairpin polynucleotide.

[0307] The tagmented DNA is pooled in step 3, and a second hairpin polynucleotide is introduced to the DNA, which is used for copy strand synthesis.

[0308] Finally, in step 4 a DNA library is generated and sequenced, and the barcode sequence associated with each DNA read is used to map the read to a single cell.

[0309] Example 2 - DNA Library Preparation

[0310] A schematic for library preparation is shown in Figure 2. This may be combined with the compartmentalisation as described in Example 1.

[0311] In a first step, a first hairpin polynucleotide is provided comprising a barcode sequence, and a non-canonical DNA nucleotide, namely a deoxyuridine residue. The hairpin polynucleotide comprises a mosaic end sequence (labelled ME seq). This polynucleotide is loaded onto a transposase, to provide a transposase dimer complex.

[0312] The assembled complex is contacted with genomic DNA in step 2. This generates doublestranded genomic DNA fragments, wherein each strand in a fragment is covalently linked to the first hairpin polynucleotide at its 5’-end. The 3’-end are free and are not covalently linked to a hairpin. For compartmentalised samples, such as compartmentalised single cell or single nuclei, the barcode sequence differs between each compartment such that DNA within each compartment (and therefore from each cell or nucleus) is unique labelled by the first hairpin polynucleotide via the barcode.

[0313] 008856557 In step 3, the free 3’-ends of the fragment DNA are extended, using the complementary strand within the duplex, followed by the first hairpin polynucleotide attached to the complementary strand, as a template. In this way, the first hairpin polynucleotide is copied across the strand to provide the second hairpin polynucleotide.

[0314] In step 4, the DNA is digested with uracil DNA glycosylase and endonuclease VIII (USER, NEB) to release a portion of the first hairpin polynucleotide. The portion of the first hairpin polynucleotide that remains attached to the DNA fragment comprises the barcode, and thus the DNA is labelled after cleavage. The second hairpin polynucleotide lacks any deoxyuridine residues and remains intact.

[0315] The resultant DNA is denatured and the second hairpin polynucleotide is self-annealed, which forms a hairpin and extended using Klenow to create a barcoded, hairpin-tagged DNA strand.

[0316] In alternative embodiments, the first hairpin polynucleotide comprises a barcode sequence, and a non-canonical DNA nucleotide, such as a deoxyuridine residue, is provided at the 3’-end of the first hairpin polynucleotide. In these embodiments, the barcode sequence is copied during step 3 to the complementary step, thus labelling each DNA strand at the 3’-end. During USER digestion in step 4, the barcode sequence in the first hairpin polynucleotide is removed from each DNA strand. The resultant DNA is denatured and selfannealed as described above. This forms a hairpin-tagged strand having a barcode sequence at the 3’-end hairpin, which can be used to identify the origin of the DNA.

[0317] General Protocol A: DNA Library Preparation from Single Cells

[0318] Fresh cell cultures were harvested in a conical centrifuge tube (15 mL or 50 mL) at room temperature and 100,000 cell aliquots were collected. The aliquots were centrifuged for 5 min at 500 x g at 4 °C, then washed with 500 pL of cold PBS and 0.1% BSA. Cells were centrifuged at 500 x g at 4 °C and resuspended in 500 pL of nuclei isolation buffer (10 mM Tris-HCI, pH 7.5; 10 mM NaCI; 3 mM MgCh; 0.03% Tween 20). Samples were then centrifuged at 1 ,300 x g at 4 °C for 4 minutes and resuspended in 500 pL of PBS with 0.1 % BSA, and this process was repeated.

[0319] A 0.2 mg / mL solution of proteinase K was freshly prepared and 4 pL was dispensed into wells of a 96-well plate. Cells were diluted with PBS with 0.1% BSA to a final concentration of 2,000 nuclei / mL and 10 pL of 7-AAD solution was added to 1 mL of diluted cells. The diluted cells were then sorted into the wells using a Paia cell dispensing. The plate was incubated at 50 °C for 1 hour, followed by 80 °C for 30 minutes.

[0320] Hairpin sequence: / 5Phos / CTG TCT CTT ATA CAC ATC T / ideoxyU / N NNN NNN NAG ATG

[0321] 008856557 TGT ATA AGA GAC AG, where N NNN NNN N is a barcode sequence (pre-determined barcodes 1-96 used).

[0322] Hairpin adapters were prepared as 25 pM stock solutions and a master mix with 0.7 pL of hairpin adapter / reaction was combined with TPS buffer (1X), Tn5 (1.51 pg / pL,

[0323] 8 pmol / reaction) and TE buffer. The master mix was incubated at 25 °C for 30 minutes for transposome assembly.

[0324] Tagmentation buffer was prepared by combining 100 pL of HEPES buffer (1 M, pH 7.5), 300 pL of 5 M NaCI, 1.25 pL of 2M spermidine, 200 pL of protease inhibitor, and made up to 2.5 mL. For every 250 pL of the mixture, 0.5 pL of 1 % digitonin, 5 pL of 1 M MgCh, 150 pL of 40% PEG8000 and 94 pL of water was added, to provide 500 pL of tagmentation buffer.

[0325] The assembled transposomes were diluted 10* in the tagmentation buffer, and 5 pL of the mixture was added to each well containing the sorted cells. The plate was incubated at 55 °C for 10 minutes, after which 2 pL of 0.2% SDS was added per well and the plate was incubated at 55 °C for 10 minutes. Samples were then equilibrated to room temperature. Each row of the plate (12 nuclei) were pooled into a single PCR tube, and the DNA was purified purified by SPRIselect beads, and eluted in 17 pL of nuclease-free water.

[0326] The DNA was then combined with phi29 DNA polymerase (0.8 pL, 10 U / pL) and phi29 DNA polymerase buffer (1 x), 200 pM dNTPs, and 100 pg / mL albumin in a reaction volume of 25 pL. Samples were incubated at 30 °C for 30 minutes followed by 65 °C for 10 minutes. The DNA was purified by SPRIselect beads, and eluted in 24 pL of 10 mM Tris-HCI pH 8.0.

[0327] Next, USER enzyme (3.3 pL, NEB) and rCutSmart Buffer (1X, NEB) were added to the samples which was then incubated at 37 °C for 30 minutes. The DNA was purified by SPRIselect beads, and eluted in 13 pL of 10 mM Tris-HCI pH 8.0. Samples were then denatured by heating to 95 °C for 2 minutes and cooled down to 10 °C at a rate of 0.1 °C / second and held for 10 minutes.

[0328] DNA extension was carried out by combining the DNA with copy strand buffer (20 mM Tris, pH 8.0, 10 mM Magnesium acetate, 50 mM Potassium acetate, 1 mM DTT), dNTPs (1 mM), Klenow and T4 PNK (ThermoScientific) in 20 pL reaction volume and incubated at 37 °C for 30 minutes. Samples were then denatured by heating to 95 °C for 2 minutes and cooled down to 10 °C at a rate of 0.1 °C / second and held for 10 minutes.

[0329] Adapter ligation was carried out by combining DNA samples with adapters (forward: ACACTCTTTCC CTACACGACGCTCTTCCGATC*T, *indicates phosphorothioate, reverse: GATCGGAAGAGCACACGTCTGAACTCCAGTCA, all Cs are mC, Biomers.net GmbH;

[0330] 750 nM), ligation master mix and ligation enhancer (NEB Ultra II) in 50 pL reaction volume,

[0331] 008856557 and the solution was incubated at 20 °C for 15 minutes. The DNA was purified by SPRIselect beads and eluted in 30 pL of 10 mM Tris-HCI pH 8.0.

[0332] To the ligated DNA, TET2 (2 pL), pGT (1 pL, ThermoScientific), DTT (Sigma, 2 mM), UDP-Glucose (1 pL, ThermoScientific) and TET Buffer (10 mM a-ketogluterate, Sigma in 0.25 M Tris-HCI pH 8.0, 10 mM ATP) was added, in a total reaction volume of 45 pL. Next, a solution of 500 mM Fe(ll) sulfate hexahydrate-50mM DTT was diluted 1,250X and 5 pL of the diluted solution was added to the sample and incubated at 37 °C for 1 hour. The DNA was purified by SPRIselect beads and eluted in 31 pL of 10 mM Tris-HCI pH 8.0.

[0333] Samples were combined with 17.5 pL of 4X APOBEC buffer (200 mM BisT ris pH 6.1 , 0.4% Tween), 1.75 pL of 100 mM ATP (Sigma), 3.5 pL of 100 mM MgCI2 (Sigma), 2 pL of APOBEC-A3A (Cambridge Epigenetix) and 2.5 pL of UvrD Helicase (Cambridge Epigenetix). The reaction was incubated for 90 minutes at 37 °C and the DNA was purified by SPRIselect beads and eluted in 20 pL of nuclease-free water.

[0334] Library amplification was carried out using 5 pL of ABclonal Unique Dual Index Primers for Illumina (ABclonal no. RK21624_SetA) and 25 pL of 2X KAPA HiFi U+ Polymerase (no. KK2802). The PCR program was as follows: 30 seconds at 98 °C for initial denaturation, 10 seconds at 98 °C for denaturation, 30 seconds at 62 °C for annealing, 60 seconds at 65 °C for extension and five minutes at 65 °C for final extension. After PCR the final libraries were purified by SPRIselect beads, eluted in 15 pl of 10 mM Tris-HCI pH 8.0 and quantified using TapeStation D5000 reagents (Agilent).

[0335] Libraries were quantified using a Qubit dsDNA HS kit according to the manufacturer’s instructions.

[0336] The protocol above was followed for 5-letter sequencing where 5mC and 5hmC are both read as cytosine:guanine base pair. For 6-letter sequencing where 5hmC is distinguished from 5mC, DNA was additionally treated with DNMT5 as described in Fullgrabe et al.

[0337] General Protocol B: Library Preparation for Bulk DNA

[0338] The sequence of the hairpin adapters used are shown in Table 1.

[0339] 008856557 Table 1 : Hairpin adapter sequences, wherein NNNNNNNNN represents a barcode index.

[0340] Hairpin adapters were prepared as 25 pM stock solutions. Hairpin adapters (either Hairpin-1 or Hairpin-2) and Tn5 were diluted to 20 pM solutions before use using nuclease-free water, and mixed together in a 1 :1 ratio for transposome assembly.

[0341] For tagmentation, the assembled transposomes were added to 80 ng DNA samples (diluted using 10 mM Tris-HCI pH 8.0 as necessary) and incubated at 56 °C for 10 minutes.

[0342] Samples were then incubated with SDS (0.05% final concentration) at 55 °C for 15 minutes to inactivate the Tn5 enzyme. The tagmented DNA was purified by SPRIselect beads according to the manufacturer’s instructions, and eluted in 17.5 pL of nuclease-free water. Tagmentation was verified using Tapestation D5000 reagents (Agilent) according to the manufacturer’s instructions.

[0343] The DNA was then combined with phi29 DNA polymerase (0.8 pL, 10 U / pL) and phi29 DNA polymerase buffer (1X), 200 pM dNTPs, and 100 pg / mL albumin in a reaction volume of 25 pL. Samples were incubated at 30 °C for 30 minutes followed by 65 °C for 10 minutes. The DNA was purified by SPRIselect beads, and eluted in 25.7 pL of 10 mM Tris-HCI pH 8.0.

[0344] Next, USER enzyme (3.3 pL, NEB) and rCutSmart Buffer (1X, NEB) were added to the samples which was then incubated at 37 °C for 30 minutes. The DNA was purified by SPRIselect beads, and eluted in 13 pL of 10 mM Tris-HCI pH 8.0. Samples were then denatured by heating to 95 °C for 2 minutes and cooled down to 10 °C at a rate of 0.1 °C / second and held for 10 minutes.

[0345] DNA extension was carried out by combining the DNA with copy strand buffer (20 mM Tris, pH8.0, 10 mM Magnesium acetate, 50 mM Potassium acetate, 1 mM DTT), dNTPs (1 mM), Klenow and T4 PNK (ThermoScientific) in 20 pL reaction volume and incubated at 37 °C for 30 minutes. Samples were then denatured by heating to 95 °C for 2 minutes and cooled down to 10 °C at a rate of 0.1 °C / second and held for 10 minutes.

[0346] 008856557 Adapter ligation was carried out by combining DNA samples with adapters (forward: ACA CTC TTT CCC TAC ACG ACG CTC TTC CGA TC*T, indicates phosphorothioate, reverse: GAT CGG AAG AGO ACA CGT CTG AAC TCC AGT CA, all Cs are mC, Biomers.net GmbH; 750 nM), ligation master mix and ligation enhancer (NEB Ultra II) in 50 pL reaction volume, and the solution was incubated at 20 °C for 15 minutes. The DNA was purified by SPRIselect beads and eluted in 30 pL of 10 mM Tris-HCI pH 8.0.

[0347] To the ligated DNA, TET2 (2 pL), pGT (1 pL, ThermoScientific), DTT (Sigma, 2 mM), UDP- Glucose (1 pL, ThermoScientific) and TET Buffer (10 mM a-ketogluterate, Sigma in 0.25 M Tris-HCI pH 8.0, 10 mM ATP) was added, in a total reaction volume of 45 pL. Next, a solution of 500 mM Fe(ll) sulfate hexahydrate-50mM DTT was diluted 1,250X and 5 pL of the diluted solution was added to the sample and incubated at 37 °C for 1 hour. The DNA was purified by SPRIselect beads and eluted in 31 pL of 10 mM Tris-HCI pH 8.0.

[0348] Samples were combined with 17.5 pL of 4X APOBEC buffer (200 mM BisT ris pH 6.1 , 0.4% Tween), 1.75 pL of 100 mM ATP (Sigma), 3.5 pL of 100 mM MgCI2 (Sigma), 2 pL of APOBEC-A3A (Cambridge Epigenetix) and 2.5 pL of UvrD Helicase (Cambridge Epigenetix). The reaction was incubated for 90 minutes at 37 °C and the DNA was purified by SPRIselect beads and eluted in 20 pL of nuclease-free water.

[0349] Library amplification was carried out using 5 pL of ABclonal Unique Dual Index Primers for Illumina (ABclonal no. RK21624_SetA) and 25 pL of 2X KAPA HiFi U+ Polymerase (no. KK2802). The PCR program was as follows: 30 seconds at 98 °C for initial denaturation, 10 seconds at 98 °C for denaturation, 30 seconds at 62 °C for annealing, 60 seconds at 65 °C for extension and five minutes at 65 °C for final extension. After PCR the final libraries were purified by SPRIselect beads, eluted in 15 pl of 10 mM Tris-HCI pH 8.0 and quantified using TapeStation D5000 reagents (Agilent).

[0350] The protocol above was followed for 5-letter sequencing where 5mC and 5hmC are both read as a cytosine:guanine base pair and distinguished from canonical cytosine which is read as a thymine:guanine base pair. For 6-letter sequencing where 5hmC is distinguished from 5mC, DNA was additionally treated with DNMT5 as described in Fullgrabe et al.

[0351] Libraries were quantified using a Qubit dsDNA HS kit according to the manufacturer’s instructions.

[0352] General Protocol C: Library Preparation using Adapter Ligation for Bulk DNA

[0353] DNA (80 ng) was fragmented and hairpin adapter ligation (ATG ACG ATG CGT TCG AGC ATC GUC AUT, all Cs are methylated, Biomers.net GmbH) was carried using a NEBNext Ultra II DNA library Prep Kit (NEB) as described in Fullgrabe et al. DNA was then subjected to the five-letter seq protocol as described in Fullgrabe et al.

[0354] 008856551 Sequencing and Data Processing

[0355] Sequencing was performed using an Illumina sequencer as paired-end sequencing runs. 2 x 151 bp paired-end sequencing analysis was carried using, using 2 x index reads for sample pooling and deconvolution. An additional sequencing read was performed using a custom primer corresponding to the hairpin adapter sequence used, according to the manufacturer’s instructions. Sequencing runs included a spike-in of a base balanced library, such as PhiX (5-15% PhiX).

[0356] Data processing and genetic accuracy metrics were carried out as described in Fullgrabe et al. Sequencing reads that do not begin with an expected were removed.

[0357] Example 3 - Single Cell Sequencing of Mouse Embryonic Stem Cells

[0358] To demonstrate that the methods of the invention are able to read single-cell cytosine modifications, DNA libraries were prepared using mouse embryonic stem cells, E14, grown under two different conditions. These conditions affect cytosine modifications, with LI F / 2i, leading to hypomethylation.

[0359] DNA from single cells were prepared and sequenced according to General Protocol A. Extracted nuclei from each condition were mixed and then sorted as to contain an equal amount from each condition. Each cell was tagmented with a unique barcode, and multiple plates of 96 cells were sequenced after libraries were prepared.

[0360] As a control, bulk DNA was sequenced by extracting DNA from cells in each condition and sequencing libraries were prepared from bulk according to either General Protocol B or a ligation-based method (General Protocol C, as described in Fullgrabe et al.).

[0361] The single-cell sequencing method was used to produce a multimodal LIMAP based on 5hmC and 5mC, which was found to form two clusters (Figure 3). The profile for 5mC and 5hmC from each pseudo-cluster were found to match bulk sequencing data of cells grown under LIF and LI F / 2i conditions, which were sequenced separately as a reference (Figure 4). This comparison shows that the sequences arising from “Cluster 1” as shown in Figure 3 are from the cells grown under LI F / 2i condition and “Cluster 2” sequences as shown in Figure 3 are the cells grown under LIF.

[0362] Example 4 - Single Cell Sequencing of Mouse Cortex Tissue

[0363] In this example, DNA from -3,500 single cells in mouse cortex tissue (>40 x 96 cells) was sequenced according to General Protocol A.

[0364] 008856557 Figure 5 shows a LIMAP plot using binned and combined 5mC and 5hmC data. The sequences obtained were assigned a cell type, with 5mC and 5hmC data being merged into a “5modC” signal. The merged data (single nuclei duet) was then jointly mapped using publicly available data (Liu et al. 2023 and Yao et al., 2023) using the Harmony tool as described in Korunsky et al., 2019.

[0365] The cell type assignments were then transferred to 5mC and 5hmC data, revealing distinct clusters of cells. The resulting LIMAP is presented in Figure 6. This represents the most complete map of both 5mC and 5hmC at single cell level of the mouse cortex to date. As a first validation that this LIMAP represents known cell-type specific epigenetic information, the neuronal cells in this LIMAP were found to be particularly enriched in 5hmC (Figure 7), which is consistent with reports in the literature.

[0366] Using the assigned UMAP containing the complete predominate cytosine modification profiles for each cell and cell type, a number of genes were found to be differentially methylated or hydroxymethlated, and certain genes were found to have a pattern of hypomethylation across gene bodies that are cell-type specific. To examine this in more detail, one gene from this dataset, Foxp2, which is highly expressed in certain glutamatergic neurons such as L6-CT-CTX-Glut and NP-CT-L6b-Glut was studied further. A pattern of low 5mC (hypomethylation) and high levels of 5hmC (hyperhydroxymethylation) across the gene body was observed for these two cell types in the single-cell sequencing dataset. However, the signal is attenuated when observed without discrimination for which cytosine modification is present (e.g. combined 5modC), due to the opposing effect of 5mC and 5hmC. 5modC in CH context is also less clear.

[0367] To determined where this is a global phenomenon in the mouse cortex, patterns across expressed and non-expressed genes were explored. The results showed that the pattern of reduced 5mC and increased levels of 5hmC was present at expressed genes for different cell types (both neuronal and non-neuronal cell types). The pattern of 5modC is attenuated in comparison.

[0368] In conclusion, the examples above demonstrate that the methods described herein are capable of reading the four canonical bases in DNA together with complete epigenetic information encoded in DNA, as applied to single cells. Applying the methods to nuclei isolated from mouse cortex, a UMAP was generated for mouse cortex which provides insights into genome-wide methylome and hydroxymethlome patterns across the genome for different cells types at a resolution not previously achieved.

[0369] The results also show that cell-type specific genes are marked by low methylation and higher hydroxymethylation across the gene body. Since current methods using bisulphite often do not make the distinction between these two cytosine modifications, dynamic changes in cell- specific genes are radically reduced by a combined 5modC signal. This

[0370] 008856557 demonstrates the power of reading all six bases as a new lens to examine the dynamic information encoded in DNA, particularly in the context of single-cell sequencing.

[0371] Oligonucleotide Sequences

[0372] Table 2: Sequences of oligonucleotides used.

[0373] References

[0374] A number of publications are cited above in order to more fully describe and disclose the invention and the state of the art to which the invention pertains. Full citations for these references are provided below. The entirety of each of these references is incorporated herein.

[0375] Adey, Genome Res. 31 , 1693-1705 (2021)

[0376] Alexandrov et al., Nature 578, 94-101 (2020)

[0377] Bentley et al., Nature 456, 53-59 (2008)

[0378] 008856557 Booth etal., Nat. Protoc. 8, 1841-1851 (2013)

[0379] Cagan et al., Nature 604, 517-524 (2022)

[0380] Chen et al., Science 356, 189-194 (2017)

[0381] Conte et al., Trends in Genetics, 40, 83-93 (2024)

[0382] Frommer et al., PNAS 89, 1827-1831 (1992)

[0383] Fullgrabe et al. Nat. Biotech. 41 1457-1464 (2023)

[0384] He et al., Nat. Commun. 13, 1335, (2022)

[0385] Hennig et al., G3 (Bethesda) 8, 79-89 (2018)

[0386] Kawasaki et al., Nucleic Acids Res. 45, e24 (2017)

[0387] Kia et al., BMC Biotechnol. 17, 6 (2017)

[0388] Korsunsky et al., Nat. Methods 16, 1289-1296 (2019)

[0389] Liu et al., Nat. Biotechnol. 37, 424-429 (2019)

[0390] Liu et al., Nat. Commun. 12, 618 (2021)

[0391] Liu et al., Nature 624, 366-377 (2023)

[0392] Mazid et al., Nature 605, 315-324 (2022)

[0393] Mellen et al., PNAS 114, E7812-E7821 (2017)

[0394] Riggs et al., Front. Mol. Biosci. 8, 734154 (2021)

[0395] Schutsky et al., Nat. Biotechnol. 36, 1083-1090 (2018)

[0396] Sprujit et al., Cell 152, 1146-1159 (2013)

[0397] Vaisvila et al., Genome Res. Doi 10.1101 / ge.266551.120 (2021)

[0398] Xi et al., BMC Bioinf. 10, 232 (2009)

[0399] Yao et al. Nature 624, 317-332 (2023)

[0400] Yu et al., Nat. Protoc. 7, 2159-2170 (2012)

[0401] WO 2013 / 090588

[0402] WO 2022 / 023753

[0403] WO 2022 / 161294

[0404] WO 2023 / 288222

[0405] US 2017 / 0002404

[0406] 008856557

Claims

Claims:

1. A method for generating a DNA library, the method comprising the steps of:(a) providing a plurality of compartments, each compartment comprising a single cell, or a nucleus from a single cell;(b) lysing the compartmentalised cells or nuclei to release genomic DNA into each compartment;(c) fragmenting the genomic DNA to produce double-stranded DNA fragments, and introducing a first hairpin polynucleotide to the 5’-end of each DNA strand in the fragments and a second hairpin polynucleotide to the 3’-end of each DNA strand in the fragments, wherein at least one of the first and second hairpin polynucleotides comprises a barcode sequence that differs between each compartment, thereby producing uniquely labelled DNA in each compartment;(d) cleaving the first hairpin polynucleotide that is attached to the 5’-end of each DNA strand; and(e) allowing the second hairpin polynucleotide at the 3’-end of each DNA strand to self-anneal, and extending the second hairpin polynucleotide along the DNA strand, to generate a library of barcoded DNA strands each having complementary regions covalently linked by the second hairpin polynucleotide.

2. The method of claim 1 , wherein each first hairpin polynucleotide comprises a first barcode sequence, wherein the first barcode sequence is different between each compartment, and wherein step (d) comprises selectively cleaving the first hairpin polynucleotide at a site 5’- to the first barcode sequence.

3. The method of claim 2, comprising pooling the DNA fragments from two or more compartments after introducing the first hairpin polynucleotide.

4. The method of claim 3, wherein the pooling is performed prior to introducing the second hairpin polynucleotide.

5. The method of any one of claims 2 to 4, wherein each second hairpin polynucleotide comprises a second barcode sequence.0088565576. The method of any one of claims 2 to 5, wherein the first hairpin polynucleotide comprises a non-canonical DNA nucleotide 5’- to the first barcode sequence.

7. The method of claim 6, wherein the non-canonical DNA nucleotide is a deoxyuridine residue.

8. The method of claim 6 or claim 7, wherein step (d) comprises contacting the DNA with a DNA glycosylase and optionally an endonuclease.

9. The method of any one of claims 1 to 8, wherein step (b) comprises contacting the cells or nuclei or lysate thereof with proteinase K.

10. The method of any one of claims 1 to 9, wherein step (c) comprising introducing the first hairpin polynucleotide to each DNA strand in a fragment and extending the 3’-end of the complementary DNA strand in said fragment to introduce the second hairpin polynucleotide to the complementary DNA strand.

11. The method of claim 10, wherein extending the 3’-ends in step is performed with a strand displacing polymerase to displace a base-paired region in the first hairpin polynucleotide of the complementary DNA strand.

12. The method of claim 10 or claim 11, wherein step (c) comprises extending the 3’-end of the complementary DNA strand to produce blunt-ended DNA fragments.

13. The method of any one of claims 1 to 12, wherein step (c) comprises cleaving the genomic DNA with a complex comprising a transposase bound to a barcoded hairpin polynucleotide, to give double-stranded DNA fragments wherein the first hairpin polynucleotide is covalently linked to the 5’-end of each DNA strand.

14. The method of claim 13, wherein the transposase is a Tn5 transposase, such as a wild-type Tn5 transposase or a variant thereof.

15. The method of claim 13 or claim 14, wherein step (c) comprises loading a transposase with the hairpin primer to produce the complex.

16. The method of any one of claims 1 to 15, wherein step (c) comprises ligating a barcoded hairpin polynucleotide to the 5’-end of the DNA fragments, to give double-stranded008856557DNA fragments wherein the first hairpin polynucleotide is covalently linked to the 5’-end of each DNA strand.

17. The method of any one of claims 1 to 16, wherein the plurality of compartments is a plurality of wells in a multi-well plate, such as a 96-well plate or a 384-well plate.

18. The method of any of claims 1 to 17, wherein step (e) comprises denaturing the DNA fragments prior to self-annealing of the second hairpin polynucleotide.

19. A method for sequencing DNA from a sample single cell, or the nucleus thereof, the method comprising the steps of:(i) providing a plurality of single cells comprising the sample single cell, or a plurality of nuclei comprising the nucleus of the sample single cell;(ii) distributing each single cell or single nucleus into respective compartments to provide a plurality of compartmentalised single cells or single nuclei;(iii) generating a DNA library from the compartmentalised cells or nuclei by a method according to any one of claims 1 to 18, wherein DNA from the sample single cell is uniquely labelled with one or more barcode sequences; and(iv) sequencing the library, and identifying the DNA from the sample single cell by the one or more barcode sequences.

20. A method for mapping the location of a modified cytosine residue from a sample single cell, the method comprising the steps of:(i) providing a plurality of single cells comprising the sample single cell, or a plurality of nuclei comprising the nucleus of the sample single cell;(ii) distributing each single cell or single nucleus into respective compartments to provide a plurality of compartmentalised single cells or single nuclei;(iii) generating a DNA library from the compartmentalised cells or nuclei by a method according to any one of claims 1 to 18, wherein DNA from the sample single cell is uniquely labelled with one or more barcode sequences;(iv) deaminating cytosine residues in the DNA library to form a treated DNA library; and(v) sequencing the treated library, and identifying the DNA from the sample single cell by the one or more barcode sequences.00885655721. The method of claim 20, comprising protecting one or more modified cytosine residues from deamination.

22. The method of claim 20 or claim 21, comprising identifying a cytosine:guanine base pair in the complementary regions of the DNA in the treated library as the location of a modified cytosine residue.

23. The method of any one of claims 20 to 22, comprising identifying a uracikguanine or a thymine:guanine base pair in the complementary regions of the DNA in the treated library as the location of a cytosine residue.

24. The method of any one of claims 20 to 23, wherein the modified cytosine residues are selected from 5 methylcytosine (5mc), 5-hydroxymethylcyotsine (5hmC), 5-formylcytosine (5fC), and 5-carboxylcytosine (5caC).

25. The method of any one of claims 20 to 24, wherein the cytosine residues in the first hairpin polynucleotide, the second hairpin polynucleotide, or both, are modified cytosine residues.

26. The method of claim 25, wherein the cytosine residues in the first hairpin polynucleotide, the second hairpin polynucleotide, or both, are selected from 5-methylcytosine residues, 5-hydroxymethylcytosine residues, 5-formylcytosine residues, and 5-carboxycytosine (5-caC) residues.

27. A kit for use in a method according to any one of claims 1 to 26, comprising:(a) a transposase;(b) a DNA hairpin polynucleotide comprising a barcode sequence and a non- canonical DNA nucleotide; and(c) proteinase K.

28. The kit of claim 27, wherein the non-canonical DNA nucleotide is a deoxyuridine residue.

29. The kit of claim 27 or claim 28, comprising a plurality of the DNA hairpin polynucleotides.00885655730. The kit of claim 29, wherein the plurality of DNA hairpin polynucleotides comprise different barcode sequences.008856557

Citation Information

Patent Citations

  • Methods and systems for processing polynucleotides

    US20170114390A1

  • Methods and kits for detection of methylation status

    WO2013090588A1

  • Barcoded wells for spatial mapping of single cells through sequencing

    WO2021155057A1

  • Distributed networks having a plurality of subnets

    WO2022002375A1

  • Modified adapters for enzymatic DNA deamination and methods of use thereof for epigenetic sequencing of free and immobilized DNA

    WO2023288222A1