Methods for nucleic acid analysis and for generating a library of tagged DNA
The method generates DNA libraries using a transposase to link hairpin primers to DNA fragments, addressing the challenge of distinguishing 5mC and 5hmC, thereby improving sequencing accuracy and efficiency.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- BIOMODAL LTD
- Filing Date
- 2025-10-31
- Publication Date
- 2026-05-07
AI Technical Summary
Existing sequencing methods fail to accurately distinguish between 5-methylcytosine (5mC) and 5-hydroxymethylcytosine (5hmC) and often result in false-positive matches and computational inefficiencies due to C-to-T conversions, compromising the detection of genetic and epigenetic information.
A method for generating DNA libraries using a transposase to covalently link a hairpin primer to DNA fragments, allowing for high-yield production of tagged DNA strands that preserve modified nucleobases, enabling simultaneous decoding of genetic and epigenetic sequences with high accuracy.
The method achieves high sequencing accuracy by minimizing PCR bias and efficiently differentiating modified nucleobases from canonical ones, enhancing sensitivity and specificity in detecting cytosine residues.
Smart Images

Figure EP2025081498_07052026_PF_FP_ABST
Abstract
Description
[0001] METHODS FOR NUCLEIC ACID ANALYSIS
[0002] Related Application
[0003] This present case is related to, and claims the benefit of, GB 2416165.5 filed on 01 November 2024 (01.11.2024), the contents of which are hereby incorporated by reference in their entirety.
[0004] Technical Field
[0005] This invention relates to methods for generating DNA libraries using a transposase, and methods of mapping the location of modified cytosine residues in a DNA sample using the DNA libraries. The invention also provides a DNA library, and also a transposase complex and a kit.
[0006] Background
[0007] Information encoded in nucleic acids is fundamental to the biology of living systems. There are multiple dimensions of information stored within DNA. Genetic sequencing of the DNA bases, G, C, T and A, has been transformed by high-throughput sequencing approaches in the past two decades. Epigenetic information in DNA provides insights into dynamic changes in biology that are closely associated with transcriptional programs (He et al., 2022) and cell fate (Mazid et al., 2022). The combination of genetic and epigenetic information provides a more comprehensive view of biology. More recently, 5-hydroxymethylcytosine (5hmC) has emerged as an important base modification that can provide information that goes beyond 5-methylcytosine (5mC) and genetics (Sprujit et al., 2013, Mellen et al., 2017). Hitherto, researchers have accessed either genetic or epigenetic information, without resolving 5mC from 5hmC.
[0008] Commonly used sequencing approaches do not capture full information from both genetics and epigenetics. Next-generation sequencing directly captures the canonical bases G, C, T and A in its readout (Bentley et al., 2008). A number of base-conversion chemistries have been developed to help differentiate unmodified C from its epigenetic variants, 5mC or 5hmC. These include bisulfite-based approaches such as whole-genome bisulfite sequencing (WGBS) (Frommer et al., 1992) and bisulfite-free approaches such as enzymatic-methyl sequencing (EM-seq) (Vaisvila et al., 2021) and TET-assisted pyridine borane sequencing (Liu ef a / .,2019). An important shortfall of all such methods is that conversion of either the C base, or one of its epigenetic derivatives, to a U (read as T) compromises the direct detection of genetic C-to-T changes, which is the most common mutation in the mammalian genome (Cagan et al., 2022) and in cancer (Alexandrov et al., 2020). Furthermore, the ambiguity caused by C-to-T conversions in the sequenced reads being mapped against either C or T in the reference genome increases false-positive
[0009] 008856544 matches in the search space, consequently making computational alignment and mapping of converted reads slower, more expensive and less accurate (Xi et al., 2009). Also, these existing methods cannot distinguish 5mC from 5hmC in a single workflow.
[0010] Methods to distinguish 5hmC from 5mC by exclusively converting only one base have been developed for example oxidative bisulfite sequencing (Booth et al., 2013), TET-assisted pyridine borane sequencing-beta (Liu et al., 2021), Tet-assisted bisulfite sequencing (Yu et al., 2012) and APOBEC-coupled epigenetic sequencing (Schutsky et al., 2018) or by selectively copying 5mC across strands of DNA (WO 2013 / 090588; Kawasaki et al., 2017). However, some of these can involve separate, parallel workflows and sequencing to yield full information, which may increase sample requirement, cost and time taken and / or yield data that lack phased information. Combining separate datasets is fraught with difficulties that lead to additive measurement error and coverage gaps across workflows.
[0011] There is a need for methods to detect epigenetic modifications in nucleic acids with high accuracy. Accordingly, the present inventors have developed new methods for generating DNA libraries in high yield that allow for the detection of epigenetic modifications with high sensitivity.
[0012] Summary of the Invention
[0013] At its most general, the present invention relates to a method for generating a library of tagged DNA strands from a DNA sample, the tagged DNA strands each having a template strand region and a complementary strand region, which regions are covalently linked together by a hairpin tag.
[0014] Libraries generated in this way are useful for simultaneously decoding the genetic and the epigenetic sequences from a DNA sample. Any modified nucleobases that are present in the template strand are preserved in the tagged DNA, and the resultant libraries can be used to differentiate modified nucleobases from canonical nucleobases in the template strand with high accuracy.
[0015] The method includes fragmenting a piece of DNA and covalently linking a hairpin primer to the 5’-end of each DNA fragment using a transposase enzyme. The 3’-ends of the resultant fragments are extended, using the hairpin primer that is covalently linked to the complementary strand within the same fragment as a template. This generates DNA strands each having a 3’-end hairpin primer, with the primer being able to self-anneal into a hairpin structure. Extension of the self-annealed 3’-end hairpin as a primer for DNA synthesis, along the original template DNA, provides tagged DNA comprising a template DNA and a copy DNA, which are linked together by the synthesised 3’-end hairpin primer.
[0016] 008856544 Advantageously, the methods of the invention can be used to generate tagged DNA libraries with high efficiency. The methods can be useful for PCR-free sequencing, which is particularly useful for increasing sequencing accuracy.
[0017] In a first aspect, there is provided a method for generating a library of tagged DNA, the method comprising the steps of:
[0018] (i) cleaving a double-stranded DNA with a complex comprising a transposase bound to a hairpin primer, to give double-stranded DNA fragments where the hairpin primer is covalently linked to the 5’-end of each strand in each double-stranded DNA fragment;
[0019] (ii) extending the 3’-ends of each strand in the double-stranded DNA fragments from step (i) to produce DNA fragments comprising a complementary hairpin primer at the 3’-end of each DNA strand;
[0020] (iii) removing all or a portion of the hairpin primer from the 5’-end of each DNA strand;
[0021] (iv) allowing the complementary hairpin primer to self-anneal; and
[0022] (v) extending the complementary hairpin primer to generate a library of tagged DNA having complementary regions covalently linked at an end by the complementary hairpin primer.
[0023] The method generates tagged DNA having a template DNA region and a synthesised region that is complementary to the template DNA, with these regions being joined by a hairpin polynucleotide. A hairpin is covalently linked to the input DNA at the 3’-end of each fragment, without the need for DNA ligation by a ligase enzyme. The yield of the tagged DNA strands generated in this way is high, such as higher than the yield obtainable from methods utilizing DNA fragmentation and enzymatic ligation. DNA libraries can be generated by the methods of the invention using minimal PCR amplification, and PCR-free libraries can also be obtained in this way. This is particularly advantageous for nextgeneration sequencing since PCR bias is minimised.
[0024] In step (ii) of the method, extending the 3’-ends of each DNA strand comprises extension of the DNA along the hairpin primer of the complementary strand within the same DNA fragment. The hairpin primer of the complementary strand is used as the template during extension. The hairpin primer of the complementary strand may be folded, such as in a hairpin structure, and extending the 3’-end may comprise displacing a base-paired region in the hairpin primer of the complementary strand. Preferably, extending the 3’-ends of each strand of the fragments in step (ii) is performed using a strand displacing polymerase. In this way, each hairpin primer that is linked to a DNA strand in step (i) is copied to its complementary strand within the duplex fragment, thus generating DNA strands having respective hairpin primers attached to the 5’ and 3’ ends.
[0025] 008856544 Additionally or alternatively, the 3’-end extension step may be referred to as a nucleic acid extension reaction, or a polymerase extension step. The 3’-end extension step may comprise extension of the DNA along the hairpin primer of the complementary strand, such as along substantially all of the hairpin primer. Typically, the DNA fragments produced in step (ii) are double-stranded DNA fragments, which comprise two strands of DNA which are not covalently linked to one another.
[0026] The DNA fragments produced in step (ii) may be blunt-ended DNA fragments.
[0027] The hairpin primer may comprise a cleavage site for removing all or a portion of the hairpin primer from the 5’-ends of the DNA strands, such as after 3’-end extension. The cleavage site may be a non-canonical DNA nucleotide, or may be a restriction site. Preferably, the hairpin primer comprises a non-canonical DNA nucleotide. The hairpin primer may be cleaved at the non-canonical DNA nucleotide, thereby releasing at least a portion of the hairpin primer from the 5’-ends of the DNA strands.
[0028] In some embodiments, removing all or a portion of the hairpin primer from the 5’-end of each strand (step (iii) of the method) comprises contacting the DNA strands with a glycosylase and optionally an endonuclease, such that the DNA strands are cleaved at a non-canonical DNA nucleotide.
[0029] A non-canonical DNA nucleotide may comprise a deoxyuridine residue. Uracil is readily excised from DNA by uracil DNA glycosylase (such as UDG or UNG) to generate an abasic site, and the DNA strand can be cleaved at the resultant abasic site, such as by an AP endonuclease.
[0030] The hairpin primer may comprise a barcode sequence. A barcode sequence may be a polynucleotide for identifying the tagged DNA strand and / or the library. A barcode sequence may have a predetermined sequence (e.g. an index), or a barcode sequence may have a unique sequence, such as a randomised sequence, for unique molecular labelling. The hairpin primer may comprise one or more barcode sequences, such as two or more barcode sequences. Two barcodes in a hairpin primer may be the same or they may be different.
[0031] In some embodiments, the hairpin primer comprises a cleavage site that is 3’- to a barcode sequence, such as 3’- to one or more barcode sequences. This allows the barcode sequence to be removed from the 5’-end of the DNA strands, such as after the barcode sequence has been used as a template during 3’-end extension in step (ii).
[0032] The hairpin primer may comprise a cleavage site that is positioned 5’- to a barcode sequence. In this way, tagged DNA strands that are generated by the method may have a first barcode that is 5’- to the template DNA region (which remains linked to the DNA after hairpin removal in step (iii) of the method), and a second barcode that is 3’- to the template
[0033] 008856544 DNA region (which is generated from 3’-end extension in step (ii) of the method). DNA strands tagged with a dual barcode in this way allows for superior error correction, such as when the libraries are sequenced and the sequencing data are analysed.
[0034] Preferably, the transposase is a Tn5 transposase. The hairpin primer may comprise a transposase binding site for loading onto the transposase, such as a Tn5 transposase binding site. This helps to improve the yield of double-stranded fragments comprising the hairpin primer.
[0035] In some embodiments, step (i) may comprise loading a transposase with the hairpin primer to produce the complex. In other embodiments, the complex may be provided pre-formed. The complex may be a transposase dimer comprise two transposase enzymes, wherein in transposase is loaded with a respective hairpin primer.
[0036] Step (iv) may comprise denaturing the DNA fragments, such as prior to self-annealing of the complementary hairpin primer. Denaturation may be partial, or the DNA fragments may be fully denatured to give single DNA strands comprising the complementary hairpin primer at the 3’-end of each strand.
[0037] The double-stranded DNA used as an input for step (i) may be genomic DNA. The double-stranded DNA may be from a tissue sample or a biopsy.
[0038] In a second aspect, there is provided a method for generating a sequencing library, the method comprising the steps of:
[0039] (a) generating a library of tagged DNA by a method according to the first aspect; and
[0040] (b) ligating a sequencing adapter to the free ends of the tagged DNA.
[0041] Preferences for the first aspect also apply to the second aspect.
[0042] In a third aspect, there is provided a method of mapping the location of one or more modified cytosine residues in a double-stranded DNA sample, comprising:
[0043] (a) generating a DNA library by a method according to the first or second aspect;
[0044] (b) deaminating cytosine residues in the DNA library to form a treated DNA library; and
[0045] (c) sequencing the treated DNA library.
[0046] Preferences for the first and second aspects also apply to the third aspect.
[0047] Advantageously, the mapping method allows modified cytosine residues to be detected with high sensitivity.
[0048] 008856544 Preferably, step (b) comprises selectively deaminating the non-modified (or canonical) cytosine residues in the DNA library, such as selectively over modified cytosine residues.
[0049] In some embodiments, the method may comprise protecting one or more modified cytosine residues from deamination. Protecting a modified cytosine residue may be by contacting the DNA with an oxidising agent or a glucosylating agent, or both.
[0050] Preferably, the modified cytosine residues are 5-methylcytosine and / or 5-hydroxymethylcyotsine.
[0051] The sequence of the library of tagged DNA strands may be indicative of the location of the modified cytosine and / or canonical (non-modified) cytosine residues in the double-stranded DNA sample. Thus, in some embodiments the method comprises determining the location of the modified cytosine residues in the double-stranded DNA based on the sequence of the library of tagged DNA strands.
[0052] The method may comprise determining the location of the modified cytosine based on the sequence of complementary regions of the tagged DNA within the treated library. The method may comprise identifying a cytosine:guanine base pair in complementary regions of the DNA sequences as the location of a modified cytosine residue. The method may comprise identifying a uracil:guanine and / or thymine:guanine base pair in complementary regions of the DNA sequences as the location of a (canonical) cytosine residue. A modified cytosine residue in the input DNA may be distinguished from a canonical cytosine residue in this way.
[0053] In a fourth aspect, there is provided an isolated transposase complex comprising a transposase dimer wherein each transposase is loaded with a respective hairpin primer, wherein each hairpin primer comprises a cleavage site comprising a non-canonical DNA nucleotide.
[0054] Preferences and embodiments for the transposase and the hairpin primers are as described for the first aspect. Each hairpin primer may comprise a deoxyuridine residue.
[0055] In a fifth aspect, there is provided a DNA library, wherein each DNA strand in the library comprises a genomic region flanked by a 5’-end hairpin primer and a 3’-end hairpin primer, and wherein the 5’-end hairpin primer comprises a non-canonical DNA nucleotide, such as wherein the 5’-end hairpin primer comprises a deoxyuridine residue. Preferably, the 3’-end hairpin primer does not comprise a deoxyuridine residue.
[0056] Preferences for the DNA and the hairpin primer are as described for the first to fourth aspects.
[0057] 008856544 In a sixth aspect, there is provided a kit comprising: a transposase; a hairpin primer comprising a transposase binding site; and a strand-displacing polymerase.
[0058] Preferences for the transposase and the hairpin primer are as described for the first to fifth aspects.
[0059] In some embodiments, the hairpin primer comprises a cleavage site, optionally wherein the cleavage site comprises a non-canonical DNA nucleotide, such as a deoxyuridine residue. The hairpin primer may comprise two cleavage sites, such as two non-canonical DNA nucleotides, such as two deoxyuridine residues.
[0060] Summary of the Figures
[0061] The present invention is described with reference to the figures listed below.
[0062] Figure 1 shows a workflow of a method according to the invention.
[0063] Figure 2 shows a workflow of a method according to the present invention using a hairpin primer comprising a barcode that is a unique molecular identifier (UM I) and a deoxyuridine cleavage site positioned 5’-to the UM I, thereby generating DNA libraries having two UM Is associated with each fragment.
[0064] Figure 3 shows the final concentration (top panel) and percentage of resolved bases (bottom panel) for sequencing libraries prepared by a method according to the invention (right-hand side) compared to sequencing libraries prepared by a control method using hairpin ligation (left-hand side).
[0065] Figure 4 shows the sensitivity (top panel) and specificity (bottom panel) metrics of sequencing libraries prepared by a method according to the invention (right-hand side) compared to sequencing libraries prepared by control method using hairpin ligation (left-hand side).
[0066] Figure 5 shows the coverage (top panel) and GC bias (bottom panel) of sequencing libraries prepared by a method according to the invention (right-hand side) compared to sequencing libraries prepared by control method using hairpin ligation (left-hand side).
[0067] Detailed Description of the Invention
[0068] At its most general, the present invention relates to a method for generating a library of tagged DNA strands from a DNA sample, the tagged DNA strands each having a template
[0069] 008856544 strand region and a complementary strand region, which regions are covalently linked together by a hairpin.
[0070] In the methods of the invention, a DNA sample (or input DNA) is fragmented and a hairpin is covalently attached to the 5’-end of each DNA strand in a single step. Subsequently, the 3’-ends of the DNA strands are extended, using the hairpin primer that is covalently linked to the complementary strand within the same fragment as a template. This generates DNA strands each having 3’-end hairpin primers, which are able to self-anneal into a hairpin structure. Extension of a self-annealed hairpin as a primer along the original DNA provides a tagged DNA strand comprising a template DNA and a copy DNA, which are linked together by the 3’-end hairpin.
[0071] A tagged DNA produced by the methods of the invention is a single polynucleotide. This single polynucleotide may comprise a duplex region, as the template DNA and copy DNA are complementary. The tagged DNA may be denatured, or the tagged DNA may be selfannealed such as to form a hairpin structure.
[0072] The methods of the invention provide hairpin-tagged DNA in high yield, and the recovery of DNA is excellent. The tagged DNA can be processed for DNA sequencing without the need for extensive DNA amplification, such as with minimal PCR cycles or without any PCR amplification. The DNA libraries are useful for detecting both genetic and epigenetic bases in a DNA sample.
[0073] Chen et al., 2017 describe a method of whole-genome analysis by tagging genomic DNA with a T7 promoter, using Tn5 transposase. The tagmented DNA is linearly amplified into multiple copies of RNAs, and DNA is synthesised from the RNA copies by reverse transcription. This method is not suitable for analysing epigenetic bases within the genomic DNA, since the epigenetic information is lost during transcription and amplification. The method does not involve a step of removing the hairpin from the tagmented DNA, or use of a 3’-end hairpin as a primer for DNA polymerase extension.
[0074] Fullgrabe et al. describe a method of simultaneously sequencing genetic and epigenetic bases in DNA. The genetic and epigenetic information in an original DNA strand is decoded by sonicating the DNA into fragments, followed by ligation of a hairpin at both ends of each strand. The ligated hairpin is used to synthesise a copy strand that is complementary to the original DNA strand, and coupled decoding of bases across the original and copy strand provides a phased digital readout which detects genetic and epigenetic bases with high accuracy. WO 2023 / 288222 describes a method for profiling DNA methylation by shearing sample DNA, end-repairing and introducing uracil-containing hairpin linkers to both ends of the DNA. Fragmentation and hairpin ligation are carried out in two separate steps in these approaches, which can limit the overall yield of DNA libraries prepared in this way.
[0075] 008856544 \N0 2022 / 023753 describes tagmentation of genomic DNA using Tn5 transposase loaded with a hairpin polynucleotide to introduce the hairpin polynucleotide to the 5’ end of fragmented DNA strands. The DNA fragmentation and hairpin ligation is carried out in a single step. The Tn5 transposase leaves a gap at the end of the 3’-ends of each strand, and WO 2022 / 023753 describes gap filling between the 3’-end of each DNA strand and the 5’-end of the hairpin primer in the complementary strand. Gap filling typically includes use of a non-strand displacing polymerase to extend the DNA across the gap, followed by a ligase to seal the nick in the DNA. WO 2022 / 023753 also describes ligating a polynucleotide to the 3’ end of each strand to generate sequencing libraries. These methods require the use of a ligase enzyme, which can have low efficiency and thus the yield and recovery of DNA using these methods is limited.
[0076] The present inventors have found that a DNA hairpin can be introduced to the 3’-end of a DNA fragment without the need for DNA ligation. A 3’-end hairpin is covalently attached to the DNA strands by polymerase extension, using a hairpin primer that is covalently linked to the complementary strand within a duplex as the template. During the polymerase extension step, the hairpin primer on the complementary strand is unfolded and the sequence is copied to the complementary strand. The synthesised hairpin primer can self-hybridise, such as into a hairpin structure, which allows the DNA that is tagged with a hairpin at the 3’-end to be obtained in high yield. The resultant DNA strands can be processed to identify genetic and epigenetic bases in the template DNA, by decoding the bases across the original DNA and the copy region. The methods described herein are particularly advantageous for preparing libraries with high efficiency, such that minimal PCR amplification is required before sequencing. As shown in Examples 1-3, a method according to the invention may be used to prepare a DNA library suitable for decoding the genetic and epigenetic information at substantially higher concentrations compared to a method that uses ligation, whilst also allowing for excellent sensitivity and specificity, as well as maintaining CpG:GW coverage and GC bias.
[0077] Methods for Generating DNA Library
[0078] In an aspect of the invention, there is provided a method for generating a library of tagged DNA, the method comprising the steps of:
[0079] (i) cleaving a double-stranded DNA with a complex comprising a transposase bound to a hairpin primer, to give double-stranded DNA fragments where the hairpin primer is covalently linked to the 5’-end of each strand in each double-stranded DNA fragment;
[0080] (ii) extending the 3’-ends of each strand in the double-stranded DNA fragments from step (i) to produce DNA fragments comprising a complementary hairpin primer at the 3’-end of each DNA strand;
[0081] (iii) removing all or a portion of the hairpin primer from the 5’-end of each DNA strand;
[0082] 008856544 (iv) allowing the complementary hairpin primer to self-anneal; and
[0083] (v) extending the complementary hairpin primer to generate a library of tagged DNA having complementary regions covalently linked at an end by the complementary hairpin primer.
[0084] Steps (i) to (v) may be performed in order. In some embodiments, the removal of the hairpin primer in step (iii) is performed prior to the self-annealing in step (iv). In other embodiments, step (iii) follows step (iv).
[0085] Tagmentation
[0086] Step (i) of the methods described herein comprises contacting a double-stranded DNA with a transposase, such that the transposase cleaves the DNA into fragments and tags (or covalently links) the 5’-end of each strand within a fragment to a hairpin primer. This step may be referred to as tagmentation.
[0087] The double-stranded DNA may be provided within a sample, and there may be a plurality of double-stranded DNA provided in step (i).
[0088] A “double-stranded DNA” as described herein comprises two strands of DNA which are complementary to one another. Unless stated otherwise, the two strands of a doublestranded DNA are not covalently linked to one another. By contrast, a single strand of DNA may comprise self-complementary regions, but such regions are not provided on separate strands of DNA.
[0089] During step (i), the double-stranded DNA may be contacted with a complex comprising a transposase bound to a hairpin primer. The transposase may be provided as a dimer, wherein each transposase within the dimer is loaded with a hairpin primer. Two hairpin primers within a transposase dimer complex may be the same, or they may be different. The two hairpin primers may each have a constant region that is the same, and a variable region, such as one or more barcodes, that are unique to each hairpin primer.
[0090] Typically, in step (i) each DNA strand that is cleaved by the transposase receives a single hairpin primer that is covalently linked to the 5’-end of said DNA strand by the transposase. The 3’-end of each DNA strand cleaved by the transposase is generally not linked to a hairpin primer. The transposase may generate duplex DNA fragments having a first polynucleotide and a second polynucleotide. A first hairpin primer may be covalently linked to the 5’-end of the first polynucleotide, and a second hairpin primer may be covalently linked to the 5’-end of the second polynucleotide. A gap may be present between the 3’-end of each polynucleotide and a folded hairpin primer that is linked to the complementary polynucleotide, such as a gap between the 3’-end of the first polynucleotide and 5’-end of the hairpin primer that is covalently linked to the second polynucleotide. The gap may be
[0091] 008856544 from 5 to 20 nucleotides in length, such as from 5 to 15 nucleotides, such as about 9 nucleotides.
[0092] The transposase may be a Tn5 transposase. The transposase may be a wild-type Tn5 transposase, or may be a variant of Tn5 transposase. Suitable variants are described in Hennig et al., 2018 (including a E54K and L372P double mutant), and in Kia et al., 2017 (including mutant Tn5-059, see Mutant protein expression and purification), the contents of which are incorporated by reference herein.
[0093] The transposase may be an engineered transposase having reduced sequence bias compared to wild-type enzyme. Suitable engineered transposases are described in Riggs et al., 2020, the contents of which is incorporated by reference herein.
[0094] The hairpin primer is a polynucleotide that is capable of folding into a hairpin structure, also known as a stem-loop structure. Hairpins have self-complementary portions that can base-pair to one another to form a duplex, also known as a stem region. The two self-complementary portions of a hairpin primer are covalently connected by an unpaired loop region. The end of the stem region that is away from the loop region may be the two ends (5’- and 3’-) of the hairpin primer, or there may be an overhang of one or more nucleotides on one strand that is unpaired when the hairpin is folded. The hairpin primer is typically a DNA hairpin primer.
[0095] The hairpin primer may have a stem region, or self-complementary region, that is 5 base pairs or more, such as 10 base pairs or more, such as 15 base pairs or more, such as 18 base pairs or more, such as 19 base pairs or more, such as 20 base pairs or more, such as 25 base pairs or more, such as 30 base pairs or more. The stem region may be 100 base pairs or less, such as 50 base pairs or less, such as 30 base pairs or less, such as 25 base pairs or less, such as 20 base pairs or less, such as 19 base pairs or less. The stem region may have a length in a range with upper and lower values as described herein, such as from 5 to 100 base pairs, including from 5 to 50 base pairs, such as from 10 to 25 base pairs, such as 15 to 20 base pairs. The stem region may be about 19 base pairs.
[0096] The stem region may comprise a transposase binding site. The transposase binding site is a sequence that facilitates loading of the hairpin primer onto a transposase. The transposase binding site may be a Tn5 transposase binding site, such as a Tn5 mosaic end. A Tn5 transposase binding site may be a duplex of 19 base pairs.
[0097] In some preferred embodiments, the hairpin primer comprises a 19-base pair Tn5 binding site.
[0098] The hairpin primer may have a loop region that is 10 nucleotides or more, such as 15 nucleotides or more, such as 20 nucleotides or more, such as 25 nucleotides or more,
[0099] 008856544 such as 35 nucleotides or more, such as 45 nucleotides or more, such as 50 nucleotides or more, such as 60 nucleotides or more, such as 70 nucleotides or more. The loop region may be 150 nucleotides or less, such as 100 nucleotides or less, such as 90 nucleotides or less, such as 80 nucleotides or less, such as 70 nucleotides or less, such as 60 nucleotides or less, such as 50 nucleotides or less, such as 40 nucleotides or less. The loop region may have a length in a range with upper and lower values as described herein, such as from 10 to 150 nucleotides, such as from 10 to 100 nucleotides, including from 20 to 60 nucleotides, such as from 30 to 60 nucleotides. The loop region may be about 30 nucleotides, or about 50 nucleotides.
[0100] The total length of the hairpin may be from 20 to 150 nucleotides, such as from 50 to 120 nucleotides, such as from 60 to 100 nucleotides.
[0101] The hairpin primer may comprise a barcode sequence, which may be a nucleotide sequence for identifying DNA from a given library, or for uniquely labelling a DNA fragment. The barcode sequence may be a known (or predetermined) sequence used to label a library, thus allowing multiple libraries to be pooled together and sequenced in parallel. The barcode sequence may be a randomised sequence, which may be used to uniquely label each DNA fragment.
[0102] In some embodiments, the same barcode sequence is used to tag each DNA fragment in a single library. In some embodiments, a different sequence is used to tag each DNA fragment in a single library.
[0103] A barcode sequence may be from 5 to 15 nucleotides in length, such as from 5 to 10 nucleotides, such as from 6 to 8 nucleotides, such as 6 nucleotides, 8 or 9 nucleotides. Preferably, the barcode sequence is 8 nucleotides or 9 nucleotides.
[0104] A barcode sequence may comprise or consist of a randomised sequence, which is typically different between individual hairpin primers. This allows each DNA strand in a sample to be labelled with a unique sequence, so that the tagged DNA strand can be distinguished from other strands and any duplicates can be removed during sequencing data analysis. The randomised barcode sequence may also be referred to as a unique molecular index (UMI). The randomised barcode sequence may be from 5 to 15 nucleotides, such as from 5 to 12 nucleotides, such as from 6 to 12 nucleotides, such as 6, 9 or 12 nucleotides. Preferably, the randomised barcode sequence is about 9 nucleotides.
[0105] The hairpin primer may comprise two or more barcode sequences, such as two barcode sequences. Each barcode may be independently selected from a pre-determined sequence and a randomised sequence. Two barcode sequences may be adjacent one another, or they may be separated by one or more nucleotides.
[0106] 008856544 The hairpin primer may comprise a sequencing primer region that is suitable for use with known sequencing platforms. Suitable sequencing platforms are known in the art, including low or high throughput sequencing platforms such as Sanger sequencing, Solexa-lllumina sequencing, Ligation-based sequencing (SOLiD™), pyrosequencing; strobe sequencing (SMRT™); semiconductor array sequencing (Ion Torrent™); and nanopore sequencing (ION). The sequencing primer region may be for sequencing a barcode sequence, where this is present in the hairpin primer. The sequencing primer region may be 5’- to one or more barcode sequence, such as each barcode sequence, in the hairpin primer.
[0107] One or more barcode sequences and / or a sequencing primer, where present, may be in the loop region of the hairpin primer. Each of these regions may be adjacent to another of these regions, or they may be separated by one or more nucleotides.
[0108] The hairpin primer may comprise a cleavage site that is suitable for cleaving the covalently linked hairpin primer from a DNA strand. In other embodiments, the hairpin primer may be degraded non-specifically, such as by an exonuclease.
[0109] Preferably, the hairpin primer comprises a cleavage site for removing all or a portion of the hairpin primer from a covalently linked DNA. The cleavage site may be a non-canonical DNA nucleotide or a restriction site.
[0110] Restriction sites are DNA sequences that bind to a restriction enzyme, and which can be cleaved by the restriction enzyme. A restriction enzyme may cleave the DNA leaving blunt ends or sticky ends. The restriction site is not particularly limited, and examples are well known in the art, such as EcoRI and BamHI.
[0111] Preferably, the hairpin primer comprises a non-canonical DNA nucleotide. The non-canonical DNA nucleotide may be recognised by a DNA glycosylase and the corresponding nucleobase may be excised from the hairpin primer by the glycosylase, such as to generate an abasic site and / or a single-strand break. Examples of suitable nucleotides include 2’-deoxyuridine, 2’-deoxy-5-hydroxymethyluridine, 2’-deoxy-5-formyluridine, and 2’-deoxy-8-oxoguanosine residues. In some embodiments, the hairpin primer comprises a 2’-deoxyuridine residue.
[0112] The hairpin primer may comprise one or more non-canonical DNA nucleotides as described herein, such as one or more deoxyuridine residues. In some embodiments the hairpin primer comprises two or more non-canonical DNA nucleotides, such as two or more deoxyuridine residues, such as two non-canonical DNA nucleotides, such as two deoxyuridine residues. References herein to a deoxynucleoside residue are to a 2’- deoxynucleoside residue unless stated otherwise.
[0113] 008856544 In some embodiments the hairpin primer comprises no more than one non-canonical DNA nucleotide. In some embodiments the hairpin primer comprises no more than one deoxyuridine residue.
[0114] A cleavage site, where present, may be in the stem region or the loop region of the hairpin primer. In some embodiments, the hairpin primer comprises one or more non-canonical DNA nucleotides, such as one or more deoxyuridine residues, in the stem region, such as in a mosaic end sequence of the hairpin primer. The deoxyuridine residue may replace a thymidine residue in the mosaic end sequence. In some embodiments, the hairpin primer comprises one or more non-canonical DNA nucleotides, such as one or more deoxyuridine residues, in the loop region.
[0115] The, or each, non-canonical DNA nucleotide of a hairpin primer may be within the 20 nucleotides at the 3’-end of the hairpin primer, such as within the 15 nucleotides at the 3’-end, such as within the 10 nucleotides at the 3’-end. In some embodiments, the non- canonical DNA nucleotide is the nucleotide at the 3’-end of the hairpin primer. In other embodiments, the non-canonical DNA nucleotide is not at the 3’-end of the hairpin primer. Where the hairpin primer comprises two or more non-canonical DNA nucleotides these may be contiguous, or they may be separated by one or more canonical DNA nucleotide.
[0116] The hairpin primer may comprise a non-canonical DNA nucleotide, such as a deoxyuridine residue, and one or more of a barcode sequence as described herein.
[0117] The cleavage site, where present, may be positioned 3’- to a barcode sequence. In these embodiments, cleavage of the hairpin primer at the cleavage site results in loss of the barcode sequence from the 5’-end of the DNA strand.
[0118] The cleavage site, where present, may be positioned 5’- to a barcode sequence. In these embodiments, upon cleavage of the hairpin primer at the cleavage site, a first barcode sequence remains at the 5’-end of the DNA strand, which is typically in addition to a second barcode sequence present at the 3’-end of the DNA strand.
[0119] The hairpin primer may comprise one or more modified cytosine residues that are protected from deamination, such as one or more residues selected from 5-methylcytosine, 5-hydroxymethylcytosine, 5-formylcytosine and 5-carboxycytosine. In some embodiments, all of the cytosine residues in the hairpin primer are modified cytosine residues. In some embodiments, the hairpin primer does not contain canonical (or unmodified) cytosine residues.
[0120] 008856544 3’-end Extension
[0121] The methods of the invention comprise extending the 3’-end of a DNA fragment generated by tagmentation. DNA extension is carried out from the 3’-end of each fragment generated by the transposase, using the complementary DNA strand and the hairpin primer attached to the complementary strand as a template, thereby generating duplex fragments with the 3’-end of each strand comprising a complementary hairpin primer. A transposase may leave a gap at the end of the 3’-ends of each strand after tagmentation, which may be filled in during 3’-end extension using the complementary strand as a template, and then the complementary hairpin is copied during 3’-end extension. Typically, the DNA fragments produced in step (ii) are double-stranded DNA fragments, which comprise two strands of DNA which are not covalently linked to one another. In additional or alternative embodiments, the DNA fragments produced in step (ii) are not circular DNA strands.
[0122] A “complementary hairpin primer” is a polynucleotide or portion thereof that is complementary to a hairpin primer as described herein. A complementary hairpin primer may itself be a hairpin primer, and typically is capable of folding into a hairpin structure. Where a hairpin primer comprises a barcode sequence as described herein, the complementary hairpin primer comprises a corresponding region which is complementary to the barcode sequence in the hairpin primer, which may itself be suitable as a barcode sequence.
[0123] 3’-end extension is typically performed along the whole of the hairpin primer. Extension may terminate at the end of the hairpin primer on the complementary strand to produce blunt- ended DNA fragments. In some embodiments, a tail, such as a dA tail, may be added to the DNA strand before termination of the 3’end extension step. Preferably, step (ii) produces blunt-ended DNA fragments.
[0124] Additionally or alternatively, 3’-end extension comprises extension along the hairpin primer for 10 or more bases, such as 12 or more bases, such as 15 or more bases, such as 20 or more bases.
[0125] Step (ii) of the method may comprise extending the 3’-ends of each strand in the doublestranded DNA fragments generated in step (i), such as plurality of double-stranded DNA fragments. Extending the 3’-ends of each DNA strand comprises extension of the DNA along the hairpin primer of the complementary strand within the same DNA fragment, thereby copying the hairpin primer of the complementary strand.
[0126] The 3’-end extension step is a polymerase extension step. The 3’-end extension step preferably comprises displacing a base-paired region in the hairpin primer of the complementary strand. Preferably, 3’-end extension comprises unfolding of the hairpin primer of the complementary strand, such that a duplex region is generated comprising a
[0127] 008856544 hairpin primer on one strand of the duplex, and the corresponding complementary hairpin primer on the other strand.
[0128] Typically, 3’-end extension is carried out using a DNA polymerase. Additionally or alternatively, the 3’-end extension may be referred to as a nucleic acid extension reaction, or a polymerase extension step. The extension is preferably performed using a strand displacing DNA polymerase.
[0129] A strand displacing polymerase is a polymerase having strand displacing activity. Examples include phi29 DNA polymerase and variants thereof, Klenow Fragment of DNA Polymerase I (Klenow) and variants thereof, and Bacillus stearothermophilus DNA Polymerase I (Bst DNA polymerase) and variants thereof. Preferably, extension is performed using phi29 DNA polymerase.
[0130] The extension may be performed using a non-strand displacing polymerase. In alternative embodiments, the method does not use a non-strand displacing polymerase for extending the 3’-ends of each strand in step (ii).
[0131] A 3’-end extension may be carried out using deoxynucleoside triphosphates. The 3’-end hairpin (or complementary hairpin) is a DNA hairpin.
[0132] A 3’-end extension may be carried out with modified deoxynucleoside triphosphates, such as 2’-deoxy-5-methylcytidine triphosphate, 2’-deoxy-5-hydroxymethylcytidine triphosphate, 2’- deoxy-5-formylcytidine triphosphate or 2’-deoxy-5-carboxylcytidine triphosphate, such as in addition to 2’-deoxyadenosine triphosphate, 2’-deoxythymidine triphosphate and 2’-deoxyguanidine triphosphate. In some embodiments, 3’-end extension is performed in the absence of 2’-deoxyuridine triphosphate, and / or in the absence of 2’-deoxycytidine triphosphate.
[0133] Preferably, the polymerase used for 3’-end extension is capable of DNA synthesis beyond a non-canonical DNA nucleobase, such as uracil or a modified cytosine residue. In some preferred embodiments, the polymerase is a strand displacing polymerase that is uracil tolerant.
[0134] Strand Cleavage
[0135] The methods of the present invention comprise removing all or a portion of the hairpin primer from the 5’-ends of each strand of the DNA fragments. Typically, this is performed after extending the 3’-ends as described herein (i.e. step (ii) is performed prior to step (iii) of the method). This generates double-stranded fragments comprising a complementary hairpin primer at the 3’-ends of each strand. After strand cleavage, the complementary hairpin primer or a portion thereof may be present as a 3’-end overhang.
[0136] 008856544 Removing the hairpin primer may comprise cleaving the hairpin primer from the 5’-end of each DNA strand. The cleaved hairpin primer may remain hybridized to its complementary region in a blunt-ended fragment, and this may be removed from the DNA fragments by denaturation.
[0137] The removal step may comprise contacting the fragments having a hairpin primer with an enzyme to cleavage the hairpin at a cleavage site as described herein. In some embodiments, removal of the hairpin primer comprises contacting the DNA fragments generated in step (ii) to an enzyme, such as a glycosylase or a restriction enzyme.
[0138] Preferably, hairpin removal comprises excising a non-canonical DNA nucleobase from the hairpin primer, such as using a glycosylase. The fragments generated in step (ii) may be directly subjected to glycosylase treatment, or the fragments may be denatured and then treated with a glycosylase.
[0139] In some embodiments, step (iii) of the method comprises contacting the DNA strands with a glycosylase, such that a non-canonical DNA nucleobase is excised from the DNA. The glycosylase is typically a DNA glycosylase, which may be capable of cleaving the / V-glycosidic bond of a non-canonical deoxynucleotide.
[0140] The hairpin primer may be cleaved at a non-canonical DNA nucleotide. A non-canonical DNA nucleotide is any nucleotide residue other than a 2’-deoxyadenosine, 2’-deoxyguanosine, 2’-deoxycytidine or thymidine residue. The non-canonical DNA nucleotide may comprise a nucleobase, such as a modified nucleobase, that is other than canonical adenine, guanine, cytidine or thymine. The non-canonical DNA nucleotide may comprise a sugar modification. Cleaving DNA “at a non-canonical nucleotide” includes cleaving the DNA between the non-canonical DNA nucleotide and an adjacent nucleotide, such as at a phosphodiester bond between the non-canonical DNA nucleotide and a neighbouring nucleotide that is immediately 5’- or 3’- to the non-canonical DNA nucleotide.
[0141] Suitable non-canonical DNA nucleotides include those comprising the following residues: deoxyuridine (2’-deoxyuridine), 5-hydroxymethyluridine (2’-deoxy-5-hydroxymethyluridine), 5-formyluridine (2’-deoxy-5-formyluridine) and 8-oxoguanine (2’-deoxy-8-oxoguanidine).
[0142] The glycosylase may be a monofunctional glycosylase. Examples include uracil DNA glycosylase (including UDG and UNG), and single-strand selective monofunctional uracil DNA glycosylase (SMUG1).
[0143] The glycosylase may be a bifunctional glycosylase that has glycosylase activity and AP lyase activity. Examples include oxoguanine glycosylases, such as OGG1 and OGG2.
[0144] 008856544 The DNA may be contacted with an endonuclease, such as in addition to a glycosylase, such as in addition to a monofunctional glycosylase. The endonuclease may be an AP endonuclease. An “AP endonuclease” as referred to herein is an enzyme having endonuclease or lyase activity at an abasic site. Examples include Endonuclease VIII, APE1 , and APE2.
[0145] The DNA may be contacted with a glycosylase and an endonuclease in a single step, such the DNA is contacted with an enzyme mixture. DNA may be contacted sequentially with a glycosylase and then an endonuclease, and optionally with purification of the DNA in between these steps.
[0146] In some embodiments, the DNA is contacted with uracil DNA glycosylase, such as UDG or UNG, and endonuclease VIII. The DNA may be contacted with a mixture of uracil DNA glycosylase (such as UDG or UNG) and endonuclease VIII, or the DNA may be contacted with uracil DNA glycosylase followed by endonuclease VIII.
[0147] In some embodiments, all of the 5’-end hairpin primer is removed from each DNA strand. In these embodiments, the hairpin primer may comprise a cleavage site, such as a non- canonical DNA nucleotide, at the 3’-end of the 5’-end hairpin primer. Typically, the 3’-end complementary hairpin primer that is synthesised during 3’-end extension is not removed from the DNA strands.
[0148] In some embodiments, a portion of the 5’-end hairpin primer is selectively removed from each strand of the tagmented DNA fragments. In these embodiments, a residual portion of the hairpin primer remains attached to the 5’-end of the DNA strands. The 5’-end hairpin primer may thus comprise a cleavage site, such as a non-canonical DNA nucleotide, away from the 3’-end of the 5’-end hairpin primer. The non-canonical DNA nucleotide may be one nucleotide or more away from the 3’-end of the hairpin primer, such as two nucleotides or more, such as three nucleotides or more, such as 5 nucleotides or more, such as 10 nucleotides or more, such as 20 nucleotides or more. The residual portion of the hairpin primer may comprise one or more barcode sequence as described herein, such as one barcode sequence.
[0149] Typically, the 3’-end complementary hairpin primer does not comprise a cleavage site, such as a deoxyuridine residue. Preferably, the 5’-end hairpin primer is selectively cleaved, such as where the complementary hairpin primer on the 3’-end of the DNA is not cleaved. Preferably, the 3’-end complementary hairpin primer is not removed from the DNA strands in step (iii), or in any other step in the method.
[0150] In some embodiments, the method comprises use of a ligase. Preferably, the method does not comprise use of a ligase, such as where the method does not comprise use of a ligase before removing all or a portion of the hairpin primer from the 5’-end of DNA strands (i.e.
[0151] 008856544 step (iii) of the method). Preferably, the method does not comprise ligating a polynucleotide to the 3’-end of a DNA strand that is cleaved by a transposase. Preferably, the method does not comprise using a ligase for nick sealing.
[0152] Hairpin Extension
[0153] The methods described herein comprises a step of allowing the complementary hairpin primer at the 3’-ends of the single strands to self-anneal, and subsequently extending the self-annealed complementary hairpin long the covalently attached DNA (which may be referred to as the template strand or original strand) to generate hairpin-tagged DNA strands. Hairpin extension is by polymerase extension, in particular DNA polymerase extension.
[0154] The DNA strands may be denatured prior to hairpin self-annealing. Denaturation may be full denaturation to provide single-stranded DNA, or denaturation may be partial, provided that self-annealing of the complementary hairpin primer is possible.
[0155] Denaturing the DNA strands may be performed after at least a portion of the 5’-end hairpin primer is removed from the strands (i.e. step (iii) is performed prior to step (iv)), or denaturing may be performed before removal of the 5’-end hairpin primer. Preferably, step (iii) is performed before step (iv).
[0156] Methods of denaturing DNA are known in the art. DNA may be denatured with heat, or DNA may be denatured by treatment with a denaturing agent. Suitable denaturing agents include sodium hydroxide and formamide.
[0157] The 3’-end complementary hairpin may be self-annealed, such as after denaturation. The hairpin may be self-annealed to form a hairpin structure (stem-loop structure). The DNA may be subjected to conditions to facilitate self-annealing, or self-annealing may be spontaneous. Conditions for facilitating self-annealing may depend on the conditions used for denaturation. For example, DNA may be cooled following heat denaturation to promote self-annealing, or where a denaturing agent is used this may be quenched or removed from the DNA to promote self-annealing. Where DNA is denatured using sodium hydroxide, the sodium hydroxide may be neutralised with an acid to promote self-annealing.
[0158] Following self-annealing of the complementary hairpin primer, the method includes a step of extending the complementary hairpin primer along the covalently linked DNA strand. The DNA is typically extended from the stem region of the self-annealed 3’-end hairpin, using the fragmented DNA strand as a template. This generates tagged DNA strands having two complementary regions linked at one end by a DNA hairpin.
[0159] 008856544 Hairpin extension may be performed by a DNA polymerase. The DNA polymerase may be a high-fidelity polymerase. The polymerase may be a strand displacing polymerase or a nonstrand displacing DNA polymerase. Suitable polymerases include Klenow Fragment of DNA Polymerase I (Klenow), phi29 DNA polymerase, Taq Polymerase, Bst Polymerase, Q5 polymerase, Phusion polymerase, T4 DNA polymerase and KAPA HiFi polymerase. Preferably, the DNA polymerase is Klenow Fragment of DNA Polymerase I or phi29 DNA polymerase. More preferably, the DNA polymerase is Klenow Fragment of DNA Polymerase I, such as Klenow fragment 3’-5’ exo-.
[0160] A DNA sample, such as a double-stranded DNA cleaved in step (i) of the methods described herein, may be a genomic sequence. For example, the sequence may comprise all or part of the sequence of a gene, including exons, introns and / or upstream or downstream regulatory elements, or the sequence may comprise genomic sequence that is not associated with a gene. In some embodiments, the nucleic acid sequence may comprise one or more CpG islands.
[0161] The DNA sample may be obtained or isolated from a sample of cells, for example, mammalian cells, preferably human cells.
[0162] Suitable samples include isolated samples of cells and tissue samples, such as biopsies, as well as blood samples.
[0163] Modified cytosine residues, including 5mC, have been detected in a range of cell types including embryonic stem cells (ESCS) and neural cells. Suitable cells include somatic and germ-line cells.
[0164] Suitable cells may be at any stage of development, including fully or partially differentiated cells or non-differentiated or pluripotent cells, including stem cells, such as adult or somatic stem cells, fetal stem cells or embryonic stem cells.
[0165] Suitable cells also include induced pluripotent stem cells (iPSCs), which may be derived from any type of somatic cell in accordance with standard techniques.
[0166] For example, DNA may be obtained or isolated from neural cells, including neurons and glial cells, contractile muscle cells, smooth muscle cells, liver cells, hormone synthesising cells, sebaceous cells, pancreatic islet cells, adrenal cortex cells, fibroblasts, keratinocytes, endothelial and urothelial cells, osteocytes, and chondrocytes.
[0167] Suitable cells include disease-associated cells, for example cancer cells, such as carcinoma, sarcoma, lymphoma, blastoma or germ line tumour cells. Suitable cells include cells with the genotype of a genetic disorder such as Huntington’s disease, cystic fibrosis, sickle cell disease, phenylketonuria, Down syndrome or Marfan syndrome.
[0168] 008856544 Methods of extracting and isolating genomic DNA from samples of cells are well-known in the art. For example, genomic DNA may be isolated using any convenient isolation technique, such as phenol / chloroform extraction and alcohol precipitation, caesium chloride density gradient centrifugation, solid-phase anion-exchange chromatography and silica gelbased techniques.
[0169] In some embodiments, whole genomic DNA isolated from cells may be used directly as a DNA sample as described herein after isolation. In other embodiments, the isolated genomic DNA may be subjected to further preparation steps.
[0170] A sample may also be a blood sample, from which circulating free DNA (cfDNA) or circulating tumour DNA (ctDNA) may be extracted.
[0171] The genomic DNA may be fragmented prior to transposition, for example by sonication, shearing or endonuclease digestion, to produce genomic DNA fragments. A fraction of the genomic DNA may be used as described herein. Suitable fractions of genomic DNA may be based on size or other criteria. In some embodiments, a fraction of genomic DNA which is enriched for CpG islands (CGIs) may be used as described herein.
[0172] Following fractionation and / or other preparation steps, the genomic DNA may be purified by any convenient technique.
[0173] Following preparation, the DNA may be provided in a suitable form for further treatment as described herein. For example, the DNA may be in aqueous solution in the absence of buffers before treatment as described herein.
[0174] The DNA sample may be divided into two, three, four or more separate portions, each of which may be independently treated and sequenced, such as described herein.
[0175] Further Embodiments
[0176] In some aspects, the invention provides a method comprising the steps of:
[0177] (a) providing a double-stranded DNA fragment comprising a 5’-end hairpin-forming region and a 3’-end hairpin-forming region on each strand;
[0178] (b) removing all or a portion of the 5’-end hairpin-forming region from each end of the DNA fragment;
[0179] (c) allowing the 3’-end hairpin-forming region from step (b) to self-anneal; and
[0180] (d) extending the 3’-end hairpin to generate a library of tagged DNA having complementary regions covalently linked at an end by a hairpin.
[0181] 008856544 The hairpin-forming regions may each be a hairpin primer as described herein. Preferably, the 5’-end hairpin-forming region is a hairpin primer as described herein, which may comprise one or more cleavage sites, such as one or more non-canonical DNA nucleotides. Preferably, the 3’-end hairpin-forming region is a complementary hairpin primer as described herein, such as where the 3’-end hairpin-forming region is obtained, or is obtainable, by 3’-end extension as described herein.
[0182] Step (c) may comprise denaturing the fragment from step (b) to give single strands comprising a hairpin-forming region covalently linked to the 3’-end of each strand, such as prior to allowing the hairpin-forming region to self-anneal.
[0183] In some aspects, the invention provides a method for generating a library of tagged DNA, the method comprising the steps of:
[0184] (a) providing a population of double-stranded DNA fragments, where the 5’-end of each strand in each double-stranded fragment is covalently linked to a hairpin primer;
[0185] (b) extending the 3’-ends of each strand in a fragment to produce a DNA fragment comprising a complementary hairpin primer at the 3’-end of each strand;
[0186] (c) removing all or a portion of the hairpin primer from the 5’-end of each DNA strand to give DNA fragments comprising the complementary hairpin primer at the 3’- end of each strand;
[0187] Preferences for the hairpin primer are as described herein. Preferably, the hairpin primer comprises a cleavage site for removing all or a portion of the hairpin primer from the DNA fragment.
[0188] The method may comprise producing the population of double-stranded DNA fragments prior to step (a), such as using a transposase complex as described herein.
[0189] The method may comprise one or more subsequent steps as described herein after step (c), such as to denature the fragment, and / or self-annealing of the complementary hairpin, and / or extending the hairpin to generate a library of tagged DNA.
[0190] Methods for DNA Sequencing
[0191] The libraries generated by the methods described herein may be sequencing libraries that are compatible with a suitable sequencing platform. Thus, in some aspects there is provided a method for generating a sequencing library, the method comprising the steps of:
[0192] (a) generating a library of tagged DNA by a method as described herein; and
[0193] (b) ligating a sequencing adapter to the free ends of the tagged DNA.
[0194] 008856544 The free DNA ends of the tagged DNA is typically at an opposite end of the complementary regions of DNA that that is covalently attached to the complementary hairpin primer, for example the 5’-end of a strand where the hairpin primer is located at the 3’-end of said strand.
[0195] The sequencing adapter may comprise a double-stranded portion which is ligated to the free ends of the tagged DNA. The sequencing adapter is preferably a Y-shaped adapter. Y-shaped adapters, or forkhead adapters, typically comprise a double-stranded region and non-complementary regions. In other embodiments, the method may comprise ligating two single-stranded adapters, to respective free ends of the tagged DNA.
[0196] The library may be sequenced. Thus, in some aspects there is provided a method for sequencing comprising the steps of:
[0197] (a) generating a library of tagged DNA by a method as described herein; and
[0198] (b) ligating a sequencing adapter to the free ends of the tagged DNA; and
[0199] (c) sequencing the library.
[0200] The library is optionally amplified before sequencing. Preferably, the method comprises no more than 12 PCR cycles prior to sequencing, such as no more than 10 cycles, such as no more than 8 cycles, such as no more than 6 cycles, such as no more than 2 cycles. In some embodiments, the method is PCR-free.
[0201] Sequencing may be carried out by any suitable sequencing method or platform, such as Sanger sequencing, Solexa-lllumina sequencing, Ligation-based sequencing (SOLiD™), pyrosequencing; strobe sequencing (SMRT™); semiconductor array sequencing (Ion Torrent™); and nanopore sequencing (ION).
[0202] In some aspects, the invention provides a method comprising:
[0203] (a) generating a library of tagged DNA strands by a method as described herein, the tagged DNA strands each comprising a template strand and a copy strand;
[0204] (b) sequencing the library to determine the identity of a first base in the template strand, and the identity of a second base in the copy strand, wherein the second base is at or proximal to the corresponding locus of the first base; and
[0205] (c) determining a value of a true base at the locus of the first base in the template strand.
[0206] A proximal locus may be an adjacent nucleotide.
[0207] The methods can be used to sequence DNA with high accuracy. Sequencing and analysis may be carried out as described in WO 2022 / 023753 or Fullgrabe et al., the contents of which are incorporated by reference, such as to leverage internal logic comparisons of two-
[0208] 008856544 base sequencing methods and systems. The method may also be used in 4-base genome contexts and expanded 5- and 6-base genome contexts, as described in WO 2022 / 023753.
[0209] Step (c) may be carried out using a computer comprising a processor, a memory and instructions stored thereupon that, when executed, determines the value of the true base.
[0210] In some aspects, the invention provides a method of mapping the location of one or more modified cytosine residues in a double-stranded DNA sample, comprising:
[0211] (a) generating a library by a method according described herein;
[0212] (b) deaminating cytosine residues in the DNA library to form a treated DNA library; and
[0213] (c) sequencing the treated library.
[0214] The library generated in step (a) may be a library of tagged DNA, or may be a sequencing library. Preferably, the library generated in step (a) is a sequencing library.
[0215] In some embodiments, the library generated in step (a) is a library of tagged DNA, by a method comprising steps (i)-(v) as described herein, and a step of ligating sequencing adapters is performed between steps (b) and (c).
[0216] The one or more modified cytosine residues may be selected from 5-methylcytosine (5mC), 5-hydroxymethylcytosine (5hmC), 5-formylcytosine (5fC) and 5-carboxycytosine (5caC). Preferably, the modified cytosine residue is selected from 5mC and 5hmC. In some embodiments the modified cytosine residue is 5mC. In some embodiments the modified cytosine residue is 5hmC.
[0217] The deamination may be selective for cytosine residues (unmodified, or canonical cytosine residues) over the one or more modified cytosine residues. The sequence of the DNA library may thus be indicative of the location of the modified cytosine residues in the doublestranded DNA sample. Selectivity may be achieved by converting one or more modified cytosine residues as described herein, typically prior to deamination.
[0218] The treated library is sequenced. Within the sequence of each hairpin-tagged DNA strand, a complementary region, or a base pair, can be used to determine the identity of the DNA base in the original DNA strand. In some embodiments, a cytosine:guanine base pair in complementary regions of the sequenced DNA (corresponding to a base pair in a hairpin- tagged DNA strand) is indicative of a modified cytosine residue, such as 5mC and / or 5hmC. In some embodiments, a uracil:guanine or thymine:guanine base pair in complementary regions of the sequenced DNA is indicative of a (canonical) cytosine residue.
[0219] 008856544 Unless stated otherwise, a “cytosine residue” as referred to herein is an unmodified, or canonical cytosine residue. A cytosine residue may be distinguished from a modified cytosine residue, such as 5mC, 5hmC, 5fC or 5caC.
[0220] Step (b) in the mapping method may comprise deaminating cytosine residues selectively over one or more modified cytosine residues, such as 5mC and 5hmC. The selectivity for cytosine over a modified cytosine residue may be 2-fold or more, such as 5-fold or more, such as 10-fold or more, such as 20-fold or more, such as 50-fold or more, such as 100-fold or more.
[0221] Deamination may be performed by contacting the DNA strands in the library with a deaminase enzyme. The deaminase enzyme may be a cytosine deaminase or a variant thereof, such as an apolipoprotein B mRNA editing enzyme catalytic polypeptide (APOBEC) or a variant or a fragment thereof, such as APOBEC3A or a variant or a fragment thereof. A cytosine deaminase variant may selectively deaminate an unmodified (canonical) cytosine residue, or a modified cytosine residue, such as a 5mC or 5hmC residue.
[0222] Deamination may be performed by contacting the tagged DNA strands with a deamination agent, such as bisulfite.
[0223] Deamination may be carried out in the presence of a denaturing agent, such as a helicase, sodium hydroxide, or formamide. In some embodiments, denaturation, such as using sodium hydroxide or formamide, is carried out prior to deamination.
[0224] The method may comprise protecting modified cytosine residues from deamination, such as by contacting the DNA strands with an oxidising agent and / or a glucosylation agent. Protection in this way may be performed prior to cytosine deamination in step (b).
[0225] The oxidising agent may be an agent that is capable of oxidising one or more of 5-methylcytosine, 5-hydroxymethylcytosine and 5-formylcytosine.
[0226] The oxidizing agent may be a methylcytosine dioxygenase, such as a ten-eleven translocation (TET) enzyme, such as TET 1 , TET2 or TET3. The oxidizing agent may be a metal oxide, such as a ruthenate, such as potassium ruthenate.
[0227] The glucosylating agent may be an agent that glycosylates 5-hydroxymethylcytosine (5hmC), such as beta-glucosyltransferase (|3GT).
[0228] In some preferred embodiments, the DNA strands are contacted with a mixture comprising the oxidising agent and a glucosylating agent, such as a mixture comprising a TET enzyme and beta-glucosyltransferase (|3GT).
[0229] 008856544 In some embodiments, the method comprises: protecting modified cytosine (5mC and / or 5hmC) residues from deamination, such as by contacting the DNA strand with a methylcytosine dioxygenase and a beta-glucosyltransferase (|3GT); and deaminating unmodified cytosine residues in the DNA, such as by contacting the DNA strand with a deaminase enzyme, optionally in the presence of a denaturing agent such as a helicase.
[0230] The method may comprise contacting the DNA strands with a DNA methyltransferase. A methyltransferase may be DNMT1 or DNMT3. In this way, a methylcytosine residue in the template DNA is marked by a methylcytosine residue in the synthesised complementary DNA at a proximal (adjacent) position, which can be detected upon sequencing and used for improved error correction. Preferably, the methylation step is performed prior to cytosine deamination in step (b). The methylcytosine residue in the synthesised complementary strand may be distinguished from a canonical cytosine residue, such as by deaminating cytosine residues after treatment with the DNA methyltransferase.
[0231] 5hmC residues may be protected from DNA methyltransferase activity, such as by contacting the DNA strands with a beta-glucosyltransferase (|3GT) prior to the DNA methyltransferase. Subsequently, the methylated bases (in the original strand and the copied methylation sites) may be protected from deamination, such as by oxidation and glucosylation as described herein.
[0232] In some embodiments, the method comprises: glucosylating a 5hmC residue in a DNA strand, such as by contacting the DNA strand with a beta-glucosyltransferase (|3GT); copying a methylated cytosine residue across the hairpin DNA strand, such as by contacting the DNA strand with a DNA methyltransferase; protecting modified cytosine (5mC and / or 5hmC) residues from deamination, such as by contacting the DNA strand with a methylcytosine dioxygenase and a beta-glucosyltransferase (|3GT); and deaminating cytosine residues in the DNA, such as by contacting the DNA strand with a deaminase enzyme, optionally in the presence of a denaturing agent such as a helicase.
[0233] In some embodiments, a portion or all of the cytosine residues in a hairpin primer, complementary hairpin primer, and / or any adapters, where present, are methylated (as 5-methylcytosine). This step may protect these residues within these regions from deamination, such as by a cytosine deaminase enzyme.
[0234] DNA that is tagged by a method of the present invention may be sequenced using any convenient low or high throughput sequencing technique or platform. Suitable protocols,
[0235] 008856544 reagents and apparatus for nucleic acid sequencing are well known in the art and are available commercially.
[0236] In some embodiments, the DNA library or a portion thereof may be amplified before sequencing. Preferably, the library is amplified after tagging with the hairpin, such as after extension of the complementary hairpin primer, or after sequencing adapter ligation. Suitable methods for the amplification of nucleic acids are well known in the art. Following amplification, the amplified portions of the population of nucleic acids may be sequenced.
[0237] Nucleotide sequences obtained through sequencing may be compared and the residues the sequences fragments may be identified using computer-based sequence analysis.
[0238] Computer-based sequence analysis may be performed using any convenient computer system and software. A typical computer system comprises a central processing unit (CPU), input means, output means and data storage means (such as RAM). A monitor or other image display is preferably provided. The computer system may be operably linked to a DNA and / or RNA sequencer.
[0239] The methods of the invention allow for sequencing results obtained from the template strand to be compared against a nucleotide sequence obtained from the corresponding copy strand. A comparison between these sequences can show where modified nucleotides are present, such as modified cytosine residues. In some embodiments, sequencing results obtained from the template strand may be compared against a reference nucleotide sequence. A reference nucleotide sequence may be obtained from a database, or may be obtained by sequencing an untreated sample that has not undergone one or more treatment steps described herein. A DNA sample may be divided into two or more portions. The DNA in each portion may be sequenced and compared against each other to allow for identification of base modification sites in the treated portion.
[0240] DNA Library
[0241] In some aspects, there is provided a DNA library, wherein each DNA strand in the library comprises a genomic region flanked by a 5’-end hairpin primer and a 3’-end hairpin primer, and wherein at least one hairpin primer on each strand comprises a cleavage site for removing said hairpin primer from the DNA strand. The DNA may be an intermediate that is obtainable by the methods described herein.
[0242] Preferably, the 5’-end hairpin primer comprises a cleavage site.
[0243] The cleavage site may be a non-canonical DNA nucleotide, such as a deoxyuridine residue.
[0244] 008856544 In some embodiments, there is provided a DNA library, wherein each DNA strand in the library comprises a genomic region flanked by a 5’-end hairpin primer and a 3’-end hairpin primer, wherein the 5’-end hairpin primer comprises a non-canonical DNA nucleotide, such as a deoxyuridine residue.
[0245] Preference for the hairpin primers is as described herein.
[0246] A genomic region may be a genomic DNA sample, or a fragment thereof. Preferably, the DNA strands each comprise a fragment of genomic DNA, such as a fragment generated by tagmentation.
[0247] The 5’-end hairpin primer may be obtained, or may be obtainable, by transposition. The 5’-end hairpin primer may comprise a cleavage site for removing the hairpin primer from the DNA, such as a non-canonical DNA nucleotide, such as a deoxyuridine residue.
[0248] The 3’-end hairpin primer may be obtained, or may be obtainable, by 3’-end extension of the genomic region using a hairpin primer of a complementary strand as a template. Preferably, the 3’-end hairpin primer does not comprise a cleavage site that is cleaved when the 5’-end hairpin primer is removed. In preferred embodiments, the 3’-end hairpin primer does not comprise a deoxyuridine residue.
[0249] In some embodiments the library is single-stranded library. In some embodiments the library is a double-stranded library. A library of double-stranded DNA may be a library of blunt-ended DNA.
[0250] Typically, the two hairpin primers of each DNA strand are each capable of folding into a stem-loop structure.
[0251] Transposase Complex
[0252] The methods of the invention utilise a transposase complex wherein one or more hairpin primers, as described herein, are loaded onto the transposase complex.
[0253] Provided herein is an isolated transposase complex comprising a transposase dimer loaded with two hairpin primers. The two hairpin primers within a transposase complex may be the same, and optionally they may comprise one or more unique barcodes.
[0254] An “isolated” species as described herein may have a purity of 50% or higher, such as 60% or higher, such as 70% or higher, such as 80% or higher, such as 90% or higher, such as 95% or higher, such as 98% or higher, such as 99% a higher. An isolated enzyme may be substantially free of other enzymes.
[0255] 008856544 The transposase may be a wild-type transposase. The transposase may be an engineered transposase having reduced sequence bias compared to wild-type enzyme, such as described in Riggs et al., 2020.
[0256] The transposase may be a Tn5 transposase. The transposases may be wild-type Tn5 transposase, or may be a variant of Tn5 transposase. Examples of variants are described in Hennig et al., 2018 and Kia et al., 2017.
[0257] In some aspects, the invention provides an isolated transposase complex comprising a transposase dimer wherein each transposase is loaded with a respective hairpin primer, wherein each hairpin primer comprises a cleavage site comprising a non-canonical DNA nucleotide. The non-canonical DNA nucleotide may be as described herein, such as a deoxyuridine residue.
[0258] Each hairpin primer in the complex may be as described herein, such as comprising one or more of a non-canonical DNA nucleotide, a barcode sequence, and a sequencing primer, or a part thereof. In some embodiments, each hairpin primer in the complex comprises one or more non-canonical DNA nucleotide and one or more barcode sequences, such as a non- canonical DNA nucleotide and a barcode sequence.
[0259] Kits
[0260] In some aspects, the invention provides a kit comprising: a transposase; a hairpin primer comprising a transposase binding site; and a strand-displacing polymerase.
[0261] The transposase, hairpin primer and strand-displacing polymerase may be as described herein. The transposase may be a Tn5 transposase. The strand-displacing polymerase may be Phi29 DNA polymerase, or Bacillus stearothermophilus DNA polymerase I (Bst polymerase), or a homologue thereof. The hairpin primer may comprise one or more of a non-canonical DNA nucleotide, a barcode sequence, and a sequencing primer, or a part thereof.
[0262] The kit may comprise a plurality of hairpin primers, such as hairpin primers with different barcode sequences.
[0263] The kit may further comprise a DNA glycosylase as described herein, such as uracil DNA glycosylase. The kit may comprise an endonuclease as described herein, such as an AP endonuclease, such as endonuclease VIII.
[0264] 008856544 The kit may further comprise one or more polymerases, in addition to the strand-displacing polymerase. The polymerase may be a DNA polymerase or an RNA polymerase. The polymerase may be a thermostable polymerase, for example a high discrimination polymerase. Preferably, the polymerase is capable of DNA synthesis past a labelled cytosine residue and / or a uracil residue. The polymerase may be Klenow Fragment of DNA Polymerase I (such as exo-).
[0265] The kit may be provided in a suitable container and / or with suitable packaging.
[0266] The kit may include instructions for use, e.g., written instructions on how to use the kit in a method of detecting 5mC in a nucleic acid sample.
[0267] A kit may further comprise a population of control nucleic acids comprising one or canonical or modified residues, for example adenine (A), guanine (G), thymine (T), uracil (U), cytosine (C), 5-methylcytosine (5mC) and 5-hydroxymethylcytosine (5hmC). In some embodiments, the population of control nucleic acids may be divided into one or more portions, each portion comprising a different control residue.
[0268] The kit may include instructions for use in a method of tagging DNA as described herein.
[0269] A kit may include one or more other reagents required for the method, such as buffer solutions, deoxynucleotide triphosphates (dNTPs), primers, and sequencing and other reagents. A kit for use in generating tagged DNA may include one or more articles and / or reagents for performance of the method, such as means for providing the test sample itself, including DNA isolation and purification reagents, and sample handling containers (such components generally being sterile).
[0270] A kit may include sequencing adapters and one or more reagents for the attachment of sequencing adapters to the ends of isolated nucleic acids, such as T4 ligase.
[0271] A kit may include one or more reagents for the amplification of a population of nucleic acids using the amplification primers. Suitable reagents may include dNTPs and an appropriate buffer.
[0272] Other Embodiments
[0273] Each and every compatible combination of the embodiments described above is explicitly disclosed herein, as if each and every combination was individually and explicitly recited. Various further aspects and embodiments of the present invention will be apparent to those skilled in the art in view of the present disclosure.
[0274] 008856544 “and / or” where used herein is to be taken as specific disclosure of each of the two specified features or components with or without the other. For example “A and / or B” is to be taken as specific disclosure of each of (i) A, (ii) B and (iii) A and B, just as if each is set out individually herein.
[0275] Unless context dictates otherwise, the descriptions and definitions of the features set out above are not limited to any particular aspect or embodiment of the invention and apply equally to all aspects and embodiments which are described.
[0276] Certain aspects and embodiments of the invention will now be illustrated by way of example and with reference to the figures described above.
[0277] Examples
[0278] The methods of the invention are exemplified in the following examples.
[0279] Example 1 - Generating Hairpin Tagged DNA Strands
[0280] A library of tagged DNA may be generated as shown in Figure 1. In this example, a hairpin primer is loaded onto a transposase, such as Tn5 transposase, to form a dimer complex (step 1). Each Tn5 unit within the dimer is loaded with a respective hairpin primer. Each hairpin primer has a transposase binding site to facilitate transposase loading.
[0281] In step 2, a genomic DNA sample is contacted with the transposase complex. The transposase cleaves the genomic DNA and generates duplex DNA fragments, with each strand in the duplex being covalently linked at its 5’-end to a hairpin primer. At the 3’-end of each strand there is a gap between the DNA fragment and the 5’-end of the folded hairpin primer from the complementary strand.
[0282] In step 3, the free 3’-ends of the DNA fragments generated by the transposase are extended with a polymerase, which synthesises past the gap that is generated by the transposase. During this step the hairpin primer on the complementary strand is unfolded (i.e. displaced). The polymerase creates a duplex DNA fragment with a hairpin primer copied on to the complementary strand (i.e. a 3’-end complementary hairpin primer). Within each fragment, the DNA comprises the template strand which is a fragment from the original DNA sample, a 5’-end hairpin primer, and a complementary 3’-end hairpin primer.
[0283] The hairpin primers on the 5’-ends each have a cleavage site, namely a uracil residue. This is cleaved in step 4, for example using a mixture of uracil DNA glycosylase and endonuclease VIII (USER enzyme, NEB) which removes the hairpin primer from the 5’-ends. The 3’-end hairpin does not contain uracil and therefore remains intact.
[0284] 008856544 The strands within the duplex are then separated, and the complementary hairpin primer on the 3’-end of the resulting single stranded DNA is allowed to self-anneal into a hairpin.
[0285] In step 5, the hairpin on the 3’-end is extended using e.g. Klenow fragment to generate a tagged DNA strand having a template region (OS in Figure 1), and a copy region (CS in Figure 1), which are covalently linked by a hairpin. A sequencing adapter is ligated to the free end of the tagged DNA, thus generating a DNA strand that is suitable for sequencing.
[0286] Example 2 - Dual UMI Labelling
[0287] Figure 2 shows an embodiment of the invention where DNA is labelled with two unique molecular identifiers (UMIs). In this embodiment, a hairpin primer containing a uracil residue and a UMI (e.g. a randomised barcode sequence) is used, and the uracil is positioned 5’ to the UMI in the hairpin primer.
[0288] DNA is tagmented with the hairpin as described in Example 1 (see steps 1 and 2 of Example 1), thereby incorporating a first UMI at the 5’-end of each DNA strand. The 3’-end of the generated strands are extended as described in Example 1 (see step 3), which incorporates a second UMI at the 3’-end (labelled as de novo hairpin). In this example, Tn5 transposase is used and 3’-end extension is performed with phi29 DNA polymerase.
[0289] The double-stranded DNA fragment is treated with USER enzyme (NEB) as described in Example 1 . As the uracil residue is positioned 5’ to the first UMI, when the hairpin is cleaved by the USER enzyme a residual portion of the hairpin primer that contains the UMI remains attached to the 5’-end of each strand. The 3’-end hairpin is unchanged by the USER enzyme. Thus, each DNA strand generated by the method has two UMIs that can be read during sequencing. This allows for improved duplex error correction.
[0290] The DNA strands are then denatured and annealed, and can be further processed as described in Example 1.
[0291] Libraries prepared in this way have a lower error rate compared to a library prepared by standard Illumina sequencing.
[0292] Example 3 - Library Preparation and Analysis
[0293] General Protocol: DNA library preparation using Tn5 transposase
[0294] The sequence of the hairpin adapters used are shown in Table 1.
[0295] 008856544 Table 1 : Hairpin adapter sequences, wherein NNNNNNNNN represents a barcode index.
[0296] Hairpin adapters were prepared as 25 pM stock solutions. Hairpin adapters (either Hairpin-1 or Hairpin-2) and Tn5 were diluted to 20 pM solutions before use using nuclease-free water, and mixed together in a 1 :1 ratio for transposome assembly.
[0297] For tagmentation, the assembled transposomes were added to 80 ng DNA samples (diluted using 10 mM Tris-HCI pH 8.0 as necessary) and incubated at 56 °C for 10 minutes.
[0298] Samples were then incubated with SDS (0.05% final concentration) at 55 °C for 15 minutes to inactivate the Tn5 enzyme. The tagmented DNA was purified by SPRIselect beads according to the manufacturer’s instructions, and eluted in 17.5 pL of nuclease-free water. Tagmentation was verified using Tapestation D5000 reagents (Agilent) according to the manufacturer’s instructions.
[0299] The DNA was then combined with phi29 DNA polymerase (0.8 pL, 10 U / pL) and phi29 DNA polymerase buffer (1X), 200 pM dNTPs, and 100 pg / mL albumin in a reaction volume of 25 pL. Samples were incubated at 30 °C for 30 minutes followed by 65 °C for 10 minutes. The DNA was purified by SPRIselect beads, and eluted in 25.7 pL of 10 mM Tris-HCI pH 8.0.
[0300] Next, USER enzyme (3.3 pL, NEB) and rCutSmart Buffer (1X, NEB) were added to the samples which was then incubated at 37 °C for 30 minutes. The DNA was purified by SPRIselect beads, and eluted in 13 pL of 10 mM Tris-HCI pH 8.0. Samples were then denatured by heating to 95 °C for 2 minutes and cooled down to 10 °C at a rate of 0.1 °C / second and held for 10 minutes.
[0301] DNA extension was carried out by combining the DNA with copy strand buffer (20 mM Tris, pH8.0, 10 mM Magnesium acetate, 50 mM Potassium acetate, 1 mM DTT), dNTPs (1 mM), Klenow and T4 PNK (ThermoScientific) in 20 pL reaction volume and incubated at 37 °C for 30 minutes. Samples were then denatured by heating to 95 °C for 2 minutes and cooled down to 10 °C at a rate of 0.1 °C / second and held for 10 minutes.
[0302] 008856544 Adapter ligation was carried out by combining DNA samples with adapters (forward: ACA CTC TTT CCC TAC ACG ACG CTC TTC CGA TC*T, "Indicates phosphorothioate, reverse: GAT CGG AAG AGC ACA CGT CTG AAC TCC AGT CA, all Cs are mC, Biomers.net GmbH; 750 nM), ligation master mix and ligation enhancer (NEB Ultra II) in 50 pL reaction volume, and the solution was incubated at 20 °C for 15 minutes. The DNA was purified by SPRIselect beads and eluted in 30 pL of 10 mM Tris-HCI pH 8.0.
[0303] To the ligated DNA, TET2 (2 pL), |3GT (1 pL, ThermoScientific), DTT (Sigma, 2 mM), UDP- Glucose (1 pL, ThermoScientific) and TET Buffer (10 mM a-ketogluterate, Sigma in 0.25 M Tris-HCI pH 8.0, 10 mM ATP) was added, in a total reaction volume of 45 pL. Next, a solution of 500 mM Fe(ll) sulfate hexahydrate-50mM DTT was diluted 1 ,250X and 5 pL of the diluted solution was added to the sample and incubated at 37 °C for 1 hour. The DNA was purified by SPRIselect beads and eluted in 31 pL of 10 mM Tris-HCI pH 8.0.
[0304] Samples were combined with 17.5 pL of 4X APOBEC buffer (200 mM BisTris pH 6.1 , 0.4% Tween), 1.75 pL of 100 mM ATP (Sigma), 3.5 pL of 100 mM MgCI2 (Sigma), 2 pL of APOBEC-A3A (Cambridge Epigenetix) and 2.5 pL of UvrD Helicase (Cambridge Epigenetix). The reaction was incubated for 90 minutes at 37 °C and the DNA was purified by SPRIselect beads and eluted in 20 pL of nuclease-free water.
[0305] Library amplification was carried out using 5 pL of ABclonal Unique Dual Index Primers for Illumina (ABclonal no. RK21624_SetA) and 25 pL of 2X KAPA HiFi U+ Polymerase (no. KK2802). The PCR program was as follows: 30 seconds at 98 °C for initial denaturation, 10 seconds at 98 °C for denaturation, 30 seconds at 62 °C for annealing, 60 seconds at 65 °C for extension and five minutes at 65 °C for final extension. After PCR the final libraries were purified by SPRIselect beads, eluted in 15 pl of 10 mM Tris-HCI pH 8.0 and quantified using Tapestation D5000 reagents (Agilent).
[0306] The protocol above was followed for 5-letter sequencing where 5mC and 5hmC are both read as a cytosine:guanine base pair and distinguished from canonical cytosine which is read as a thymine:guanine base pair. For 6-letter sequencing where 5hmC is distinguished from 5mC, DNA was additionally treated with DNMT5 as described in Fullgrabe et al.
[0307] Libraries were quantified using a Qubit dsDNA HS kit according to the manufacturer’s instructions.
[0308] Sequencing was performed using an Illumina sequencer as paired-end sequencing runs. 2 x 151 bp paired-end sequencing analysis was carried using, using 2 x index reads for sample pooling and deconvolution. An additional sequencing read was performed using a custom primers corresponding to the hairpin adapter sequence used, according to the manufacturer’s instructions. Sequencing runs included a spike-in of a base balanced library,
[0309] 008856544 such as PhiX (5-15% PhiX). Data processing and genetic accuracy metrics were carried out as described in Fullgrabe et al.
[0310] DNA sample preparation using adapter ligation (control libraries)
[0311] DNA (80 ng) was fragmented and hairpin adapter ligation (ATG ACG ATG CGT TCG AGC ATC GUC AUT, all Cs are methylated, Biomers.net GmbH) was carried using a NEBNext Ultra II DNA library Prep Kit (NEB) as described in Fullgrabe et al. DNA was then subjected to the five-letter seq protocol as described in Fullgrabe et al.
[0312] Results
[0313] DNA libraries were prepared using Tn5 tagmentation and complementary hairpin primer synthesis, followed by copy strand synthesis and then subjected to the five-letter sequencing protocol according to Fullgrabe et al. Control libraries were prepared using sonication and adapter ligation according to Fullgrabe et al. Unmethylated pUC19 DNA and fully methylated lambda DNA was used as controls to test specificity and sensitivity, as described in Fullgrabe et al.
[0314] The final concentration of the libraries obtained was substantially higher using the Tn5 approach compared to ligation, as shown in Fig. 3 (top panel). Other metrics including resolved bases, lambda sensitivity, pUC19 specificity, CpG:GW coverage and GC bias was comparable between the libraries, as shown in Fig. 3, 4 and 5. These results demonstrate that the Tn5 tagmentation approach can be used to prepare DNA libraries in excellent yield while enabling good sequencing metrics.
[0315] Oligonucleotide Sequences
[0316] Table 2: Sequence of oligonucleotides
[0317] 008856544
[0318] References
[0319] A number of publications are cited above in order to more fully describe and disclose the invention and the state of the art to which the invention pertains. Full citations for these references are provided below. The entirety of each of these references is incorporated herein.
[0320] Alexandrov et al., Nature 578, 94-101 (2020)
[0321] Bentley et al., Nature 456, 53-59 (2008)
[0322] Booth et al., Nat. Protoc. 8, 1841-1851 (2013)
[0323] Cagan et al., Nature 604, 517-524 (2022)
[0324] Chen et al., Science 356, 189-194 (2017)
[0325] Frommer ef al., PNAS 89, 1827-1831 (1992)
[0326] Fullgrabe et al. Nat. Biotech. 41 1457-1464 (2023)
[0327] He et al., Nat. Commun. 13, 1335, (2022)
[0328] Hennig et al., G3 (Bethesda) 8, 79-89 (2018)
[0329] Kawasaki et al., Nucleic Acids Res. 45, e24 (2017)
[0330] Kia et al., BMC Biotechnol. 17, 6 (2017)
[0331] Liu et al., Nat. Biotechnol. 37, 424-429 (2019)
[0332] Liu et al., Nat. Commun. 12, 618 (2021)
[0333] Mazid et al., Nature 605, 315-324 (2022)
[0334] Mellen et al., PNAS 114, E7812-E7821 (2017)
[0335] Riggs et al., Front. Mol. Biosci. 8, 734154 (2021)
[0336] Schutsky et al., Nat. Biotechnol. 36, 1083-1090 (2018)
[0337] Sprujit ef al., Cell 152, 1146-1159 (2013)
[0338] 008856544 Vaisvila et al., Genome Res. Doi 10.11O1 / ge.266551.120 (2021)
[0339] Xi et al., BMC Bioin f. 10, 232 (2009)
[0340] Yu et al., Nat. Protoc. 7, 2159-2170 (2012)
[0341] WO 2013 / 090588
[0342] WO 2022 / 023753
[0343] \NO 2023 / 288222
[0344] 008856544
Claims
1. Claims:1 . A method for generating a library of tagged DNA, the method comprising the steps of:(i) cleaving a double-stranded DNA with a complex comprising a transposase bound to a hairpin primer, to give double-stranded DNA fragments where the hairpin primer is covalently linked to the 5’-end of each strand in each double-stranded DNA fragment;(ii) extending the 3’-ends of each strand in the double-stranded DNA fragments from step (i) to produce DNA fragments comprising a complementary hairpin primer at the 3’-end of each DNA strand;(iii) removing all or a portion of the hairpin primer from the 5’-end of each DNA strand;(iv) allowing the complementary hairpin primer to self-anneal; and(v) extending the complementary hairpin primer to generate a library of tagged DNA having complementary regions covalently linked at an end by the complementary hairpin primer.
2. The method of claim 1 , wherein extending the 3’-ends in step (ii) comprises displacing a base-paired region in the hairpin primer of the complementary strand.
3. The method of claim 1 or claim 2, wherein extending the 3’-ends in step (ii) is performed with a strand displacing polymerase to displace a base-paired region in the hairpin primer of the complementary strand.
4. The method of any one of claims 1 to 3, wherein the hairpin primer comprises a cleavage site for removing all or a portion of the hairpin primer from the 5’-end of each strand.
5. The method of any one of claims 1 to 4, wherein the cleavage site is a non-canonical DNA nucleotide, such as a deoxyuridine residue.
6. The method of claim 5, wherein step (iii) comprises contacting the DNA strands with a glycosylase and optionally an endonuclease.
7. The method of claim 6, wherein the hairpin primer comprises a deoxyuridine residue, and wherein step (iii) comprises contacting the DNA strands with a uracil glycosylase and optionally an endonuclease.0088565448. The method of any one of claims 4 to 7, wherein the hairpin primer comprises a barcode sequence.
9. The method of any one of claims 4 to 8, wherein the cleavage site is positioned 3’- to the barcode sequence.
10. The method of any one of claims 4 to 8, wherein the cleavage site is positioned 5’- to the barcode sequence.11 . The method of any one of claims 4 to 10, wherein the hairpin primer comprises two or more barcode sequences.
12. The method of any one of claims 1 to 11 , wherein the transposase is a Tn5 transposase, such as a wild-type Tn5 transposase or a variant thereof.
13. The method of any one of claims 1 to 12, wherein step (i) comprises loading a transposase with the hairpin primer to produce the complex.
14. The method of any one of claims 1 to 13, wherein the double-stranded DNA is genomic DNA.
15. The method of any one of claims 1 to 14, wherein the DNA fragments produced in step (ii) are blunt-ended DNA fragments.
16. The method of any of claims 1 to 15, wherein step (iv) comprises denaturing the DNA fragments prior to self-annealing, such as denaturing complementary regions between the hairpin primer and the complementary hairpin primer.
17. A method for generating a sequencing library, the method comprising the steps of:(a) generating a library of tagged DNA by a method according to any one of claims 1 to 16; and(b) ligating a sequencing adapter to the free ends of the tagged DNA.
18. A method of mapping the location of one or more modified cytosine residues in a double-stranded DNA sample, comprising:(a) generating a DNA library by a method according to any one of claims 1 to 16;008856544(b) deaminating cytosine residues in the DNA library to form a treated DNA library; and(c) sequencing the treated DNA library.
19. The method of claim 18, comprising protecting one or more modified cytosine residues from deamination.
20. The method of claim 18 or claim 19, comprising: identifying a cytosine:guanine base pair in the complementary regions of the DNA in the treated library as the location of a modified cytosine residue; and / or identifying a uracil:guanine or a thymine:guanine base pair in the complementary regions of the DNA in the treated library as the location of a cytosine residue.21 . The method of any one of claims 18 to 20, wherein the modified cytosine residues are selected from 5-methylcytosine (5mc), 5-hydroxymethylcyotsine (5hmC), 5-formylcytosine (5fC), and 5-carboxylcytosine (5caC).
22. The method of any one of claims 1 to 21 , wherein the method is a PCR-free method.
23. A DNA library, wherein each DNA strand in the library comprises a genomic region flanked by a 5’-end hairpin primer and a 3’-end hairpin primer, wherein the 5’-end hairpin primer comprises a non-canonical DNA nucleotide, such as a deoxyuridine residue.
24. An isolated transposase complex comprising a transposase dimer wherein each transposase is loaded with a respective hairpin primer, wherein each hairpin primer comprises a cleavage site comprising a non-canonical DNA nucleotide.
25. The isolated transposase complex of claim 24, wherein the non-canonical DNA nucleotide is a deoxyuridine residue.
26. The isolated transposase complex of claim 24 or 25, wherein each hairpin primer comprises a barcode sequence.
27. A kit comprising: a transposase; a hairpin primer comprising a transposase binding site; and a strand-displacing polymerase.00885654428. The kit of claim 27, wherein the strand-displacing polymerase is selected from Phi29 DNA polymerase, Bacillus stearothermophilus DNA polymerase I (Bst polymerase), and a homologue thereof.008856544
Citation Information
Patent Citations
Methods and kits for detection of methylation status
WO2013090588A1
Compositions and methods for nucleic acid analysis
WO2022023753A1
Modified adapters for enzymatic DNA deamination and methods of use thereof for epigenetic sequencing of free and immobilized DNA
WO2023288222A1
Transposon end compositions and methods for modifying nucleic acids
US20150353925A1