Profiling Method
The enzymatic 'unmethylome' profiling method addresses limitations of existing epigenetic analysis techniques by tagging unmodified nucleotides for unbiased genome-wide profiling, facilitating diagnostic tools and personalized medicine applications.
Patent Information
- Application Number
- GB2023017422
- Authority / Receiving Office
- GB · GB
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-14
- Publication Date
- 2025-10-01
AI Technical Summary
Current methods for studying epigenetic modifications, such as methylation and hydroxymethylation of cytosine, are limited by DNA degradation during bisulfite conversion, low conversion efficiencies in bisulfite-free approaches, and biased analysis of genomic loci, particularly in low DNA samples like circulating cell-free DNA, limiting their application in personalized medicine and diagnostic tools.
An enzymatic 'unmethylome' profiling method that uses a methyltransferase to tag unmodified nucleotides, allowing for unbiased genome-wide profiling through a one-pot process compatible with low DNA quantities, enabling efficient sequencing and fractionation into modified and unmodified fractions.
The method provides reproducible, unbiased epigenetic profiling suitable for diagnostic tools and personalized medicine, capable of analyzing low DNA samples with minimal sample loss and enabling single-cell analysis.
Smart Images

Figure 00000001_0000 
Figure 00000001_0001 
Figure 00000002_0000
Abstract
Description
Field The present application relates to methods of determining polynucleotide modification. 5 Introduction Epigenetic modifications of polynucleotides, such as methylation and hydroxymethylation of cytosine, play an important role in determining the activity of a gene or a much more extended region of the genome. For example, methylation of DNA w is critical in embryogenesis, early development and is known to change predictably in correlation with biological ageing of an organism. On the other hand, aberrant modification of DNA can be an important driver of tumourigenesis, and the broader dysregulation of genes is likely to play a key role in many diseases. 15 Despite the critical role of cytosine modification in the regulation of gene expression, current methods for studying epigenetic modifications fundamentally limit the scope of current studies of the epigenome. Methods comprising bisulfite conversion of cytosine to uracil may be used for 20 epigenetic analysis, but the treatment of DNA with bisulfite can lead to DNA degradation. This limits the application of bisulfite in samples where the DNA quantity is low, as is typical for circulating cell-free DNA (cfDNA) in blood. Bisulfite-free approaches have been demonstrated that employ pyridine borane base 25 conversion or enzymatic deamination of unmethylated cytosine. However, such approaches can suffer from low conversion efficiencies, relative to bisulfite conversion, and are inherently focussed on the analysis of individual cytosine bases, necessitating (comparative) whole-genome sequencing for biomarker discovery. This invariably leads to reduction of the test to a panel of genomic loci for application in the clinic, which 30 effectively limits the diagnosis to a population that is similar to that profiled during the biomarker discovery phase of the test development. Enrichment-based approaches for epigenetic profiling allow cost-effective, whole genome profiling that is more suited to the challenges of delivering personalised 35 medicine. Until recently, the amount of available cfDNA in a blood sample limited the application of either antibodies (“methyl-DNA Immunoprecipitation”, also referred to as “MeDIP-Seq”) or methyl-binding domain protein (referred to as “MBD-Seq”) in the analysis of cfDNA. A further limitation of approaches involving MeDIP-Seq or MBD-Seq is that these proteins preferentially bind heavily methylated regions of the genome, which leads to an over-representation of these loci in the sequencing experiment. 5 There is, therefore, a need for improved epigenetic profiling methods, that are capable of profiling small quantities of DNA, such as may be obtained from peripheral blood samples, that are capable of application to large portions of the genome without bias arising from differences in sequence, such as CpG density, and that may be applied cost 10 effectively at large scales for use as diagnostic tools and to inform personalised medicine approaches. The inventors have developed an epigenetic profiling technique which meets these requirements and overcomes various limitations of existing methods, including those 15 discussed above. The disclosed method is an enzymatic “unmethylome” profiling approach, in which unmodified nucleotides, such as unmodified CpG dinucleotides, in a DNA sample are derivatised in such a way that they can then be isolated without bias and subsequently sequenced. The disclosed technique is an advantageously simple process that can be performed in a one-pot approach, that can be integrated readily 20 with standard high throughput sequencing platforms, and may be used with lower quantities of input DNA than has previously been possible, to generate genome-wide epigenetic profiles. The disclosed method combines a procedure of DNA library preparation for next 25 generation sequencing and a method for labelling unmodified nucleotides. Using the label, the DNA library is subsequently fractionated into modified and unmodified fractions. The method advantageously minimises the number of DNA purification steps required and is highly efficient. As a result, the disclosed enzymatic platform can be applied at DNA concentrations that are compatible with single-cell analysis (picogram 30 inputs). The approach has been found to be highly reproducible and unbiased, and may be used as a platform for the diagnosis of disease and the identification of tissue of origin in a sample. The underlying chemistry requires no a priori assumptions to be made about the sample, making the platform ideally suited for the discovery of novel biomarkers of disease. 35 Hence, in a first aspect, there is provided a method of determining the modification status of nucleotide residues in a polynucleotide sample, the method comprising the steps of: 1. using a methyltransferase enzyme configured to modify a nucleotide residue in 5 a target position to apply a tag to each unmodified nucleotide residue in a polynucleotide of the sample, wherein each unmodified nucleotide residue is unmodified in the target position; 2. inactivating the methyltransferase; 3. preparing the polynucleotide sample into a sequencing library; io 4. binding an affinity label to each tag; 5. fractionating the sequencing library into first and second fractions, wherein the first fraction is enriched for polynucleotides comprising an affinity label, and wherein the second fraction is enriched for polynucleotides lacking an affinity label; and 6. sequencing the polynucleotides of the first and / or second fraction. 15 In some embodiments, the method comprises sequencing the polynucleotides of the first fraction only. As used herein, unless otherwise stated, the terms “fractionating” and “fractionation” of 20 the sequencing library may also be described and / or referred to as “enriching” and / or “enrichment” of the sequencing library for labelled polynucleotides or unlabelled polynucleotides, as appropriate. The method may further comprise a step of amplifying the polynucleotides. In some 25 embodiments, the method may comprise the amplification of the polynucleotides after inactivation of the methyltransferase (i.e. after step 2 or step 3). In some embodiments, the method may comprise the amplification of the polynucleotides of the sequencing library after binding of the affinity label (i.e. after step 4). In some embodiments, the method may comprise the amplification of the polynucleotides of the first and / or 30 second fraction (i.e. after step 5). The polynucleotide may be DNA. Accordingly, the method may comprise determining the modification status of nucleotide residues in a DNA sample. As such, the method may comprise the steps of: 35 1. using a methyltransferase enzyme configured to modify a nucleotide residue in a target position to apply a tag to each unmodified nucleotide residue in a DNA molecule of the sample wherein each unmodified nucleotide residue is unmodified in the target position; 2. inactivating the methyltransferase; 3. preparing the DNA sample into a sequencing library; 5 4. binding an affinity label to each tag; 5. fractionating the sequencing library into first and second fractions, wherein the first fraction is enriched for DNA molecules comprising an affinity label, and wherein the second fraction is enriched for DNA molecules lacking an affinity label; and 10 6. sequencing the D NA of the first and / or second fraction. In some embodiments, the method may comprise sequencing the DNA of the first fraction only. 15 The method may further comprise a step of amplifying the DNA. In some embodiments, the method may comprise the amplification of the DNA after inactivation of the methyltransferase (i.e. after step 2 or step 3). In some embodiments, the method may comprise the amplification of the sequencing library7, after binding of the affinity label (i.e. after step 4). In some embodiments, the method may comprise 20 the amplification of the DNA of the first and / or second fraction (i.e. after step 5). The terms “enriched”, “enriching”, and “enrichment” of polynucleotides as used herein, unless otherwise stated, refer to a polynucleotide concentration and / or proportion that is greater than the corresponding polynucleotide concentration and / or proportion in 25 the initial (unenriched) sample. References to “enriching the labelled polynucleotide library” and similar terms refer to fractionating the sequencing library / polynucleotide sample into first and second fractions, wherein the first fraction is enriched for polynucleotides comprising an affinity7 label, and wherein the second fraction is enriched for polynucleotides lacking an affinity label. Likewise, references to "enriching 30 the labelled DNA library” and similar terms refer to fractionating the sequencing library / DNA sample into first and second fractions, wherein the first fraction is enriched for DNA molecules comprising an affinity label, and wherein the second fraction is enriched for DNA molecules lacking an affinity label. 35 The library preparation step (step 3) is performed before the fractionation step (step 5). In some embodiments, the steps of the method are performed in the numerical sequence in ascending order (i.e. in sequence from step i to step 6). In some embodiments, the library preparation step (step 3) may be performed before the affinity labelling and fraction steps (i.e. before steps 4 and 5). In other embodiments, step 3 may be performed after step 4, and before step 5. 5 In some embodiments, one or more aspect of the library preparation step, such as end repair, A-tailing, and / or adapter ligation, maybe performed before step 1. In such embodiments, the remaining step(s) of library preparation may be performed subsequently, such as after step 4. 10 The inventors have found that the disclosed method provides significant advantages over previous methods of epigenetic analysis. The non-destructive nature of the disclosed method means that epigenetic analysis of the sample may be performed in parallel with other analytical approaches, such as nucleotide sequencing. 15 In addition, the inventors have found that the disclosed method is highly efficient, al lowing analysis of very low levels of input sample. Moreover, the inventors have identified that previous “unmethylome” profiling approaches introduce bias in relation to densely modified polynucleotides, and the disclosed method avoids this detrimental 20 bias. As discussed herein, these advantages have been made possible by minimising the loss of sample, by performing various operations in specific sequences and combinations. The method may be a “one pot” method, wherein all of the steps are performed in a 25 single container, thereby providing significant efficiencies in terms of time and reagents, and advantages in terms of automation. This approach has been found to maximise sensitivity, time, reagents and the overall yield of the polynucleotide enrichment. 30 The “nucleotide modification status” and “modification status of nucleotide residues” as used herein, unless otherwise stated, refer to the presence (modified) or absence (unmodified) of any chemical modification in a target position on a nucleotide residue that may be catalysed by a methyltransferase enzyme. Thus, the disclosed method may be used to determine the modification status of any position within a nucleotide that 35 may be chemically modified by a methyltransferase enzyme. Such methyltransferase catalysed modifications include, for example, the modification of cytosine (at the C5 or N4 position), adenine (at the N6 position). The nucleotide residues may be cytosine residues and / or adenine residues. 5 Accordingly, the method may be a method of determining the modification status at target positions of cytosine and / or adenine residues in a polynucleotide. The method may be a method of determining the modification status at the cytosine C5 position of each CpG dinucleotide of a DNA sample, the method comprising the steps 10 of: 1. using a methyltransferase enzyme configured to modify the cytosine C5 position of a CpG dinucleotide motif to apply a tag to each unmodified cytosine residue of the sample, wherein each unmodified cytosine residue is the cytosine of a CpG dinucleotide motif that is unmodified in the C5 position; 15 2. inactiyating the methyltransferase; 3. preparing the DNA sample into a sequencing library; 4. binding an affinity label to each tag; 5. fractionating the sequencing library into first and second fractions, wherein the first fraction is enriched for DNA molecules comprising an affinity label, and wherein 20 the second fraction is enriched for DNA molecules lacking an affinity label; and 6. sequencing the DNA of the first / and or second fraction. In some embodiments, the method comprises sequencing the DNA of the first fraction only. 25 The method may further comprise a step of amplifying the DNA. In some embodiments, the method may comprise the amplification of the DNA after inactivation of the methyltransferase (i.e. after step 2 or step 3). In some embodiments, the method may comprise the amplification of the sequencing library, after binding of 30 the affinity- label (i.e. after step 4). In some embodiments, the method may comprise the amplification of the DNA of the first and / or second fraction (i.e. after step 5). The “modification status of cytosine residues” as used herein, unless otherwise stated, refers to the presence (modified) or absence (unmodified) of any methyltransferase 35 catalysed chemical modification of cytosine. In particular, the modification may comprise modification at the C5 position of cytosine. A cytosine residue may be understood to have the following structure: In an unmodified cytosine residue R1 may be understood to be H. R1 may also be 5 referred to herein as the “C5 position”. In a modified cytosine residue R* may be anything other than H. Thus, “modified cytosine” refers to any cytosine residue that has been modified in any way at the C5 position, including, in particular, 5-methylcytosine (5-mC) and its oxidized products 5-10 hydroxymethylcytosine (5-hmC), 5-formylcytosine (5-fC) and 5-carboxylcytosine (5-caC). It may be therefore appreciated that in modified cytosine, R1 may be methyl, CH20H, COH or COOH. Alternatively, the modification may comprise modification at the N4 position of 15 cytosine. Accordingly, a modified cytosine may have the following structure: where R1 maybe anything other than H. Thus, “modified cytosine” refers to any cytosine residue that has been modified in any way at the N4 position. It may be 20 therefore appreciated that in modified cytosine, R1 may be methyl, CH20H, COH or COOH. Thus, the method may be a method of determining the modification status of cytosine residues in a DNA sample, the method comprising the steps of: 25 1. using a methyltransferase enzyme configured to modify a cytosine residue in the N4 position to apply a tag to each unmodified cytosine residue of the sample, wherein each unmodified cytosine residue is unmodified in the N4 position; 2. inactivating the methyltransferase; 3. preparing the DNA sample into a sequencing library; 4- binding an affinity label to each tag; 5. fractionating the sequencing library into first and second fractions, wherein the first fraction is enriched for DNA molecules comprising an affinity label, and wherein the second fraction is enriched for DNA molecules lacking an affinity label; and 6. sequencing the DNA of the first and / or second fractions. The “modification status of adenine residues” as used herein, unless otherwise stated, refers to the presence (modified) or absence (unmodified) of any chemical modification at the N6 position of adenine. An adenine residue maybe understood to have the following structure: R1 HN' In an unmodified adenine residue R1 maybe understood to be H. In a modified adenine residue R1 may be anything other than H. Thus, “modified adenine” refers to any adenine residue that has been modified in any way at the N6 position, including, in particular, IV6-methyladenine (m6A). It may be therefore appreciated that in modified adenine, R1 may be methyl, CH20H, COH or COOH. Thus, the method may be a method of determining the modification status of adenine residues in a polynucleotide sample, the method comprising the steps of: 1. using a methyltransferase enzyme configured to modify an adenine residue in the N6 position to apply a tag to each unmodified adenine residue of the sample, wherein each unmodified adenine residue is unmodified in the N6 position; 2. inactivating the methyltransferase; 3. preparing the polynucleotide into a sequencing library: 4. binding an affinity label to each tag; 5. fractionating the sequencing library into first and second fractions, wherein the first fraction is enriched for polynucleotides comprising an affinity label, and wherein the second fraction is enriched for polynucleotides lacking an affinity label; and 6. sequencing the first and / or second fraction. As used herein, “unmodified” and “unmethylated” refer to all nucleotides that are unmodified in any way (such as methylated, hydroxymethylated, carboxylated, acylated). 5 The method comprises the detection of unmodified nucleotide residues in a polynucleotide. In embodiments in which the method comprises determining, for a plurality of nucleotides, the presence (modified) or absence (unmodified) of any chemical 10 modification catalysed by a methyltransferase, “modification status” may also be referred to as the “profile”. Thus, the disclosed method may be used to determine a profile of methyltransferase catalysed modifications within a polynucleotide sample. The method may comprise the detection of unmodified cytosine residues in CpG 15 dinucleotides of a DNA sample. The terms “CpG”, “CpG site”, and “CpG dinucleotide”, are used interchangeably herein to refer to a cytosine-phosphate-guanine sequence in a 5’ to 3’ direction in the backbone of a nucleic acid. The terms “CpG modification status” and “modification status” as used interchangeably 20 herein, unless otherwise stated, refer to the presence (modified) or absence (unmodified) of any chemical modification at the C5 position of cytosine within one or a plurality of CpG dinucleotides. In embodiments in which the method comprises determining, for a plurality of CpG 25 dinucleotides, the presence (modified) or absence (unmodified) of any chemical modification at the C5 position of each of the plurality of cytosines, the “CpG modification status” and “modification status” may also be referred to as the “profile”. In embodiments in which the profile corresponds to the entire genome, the profile may 30 be referred to as the “unmethylome profile” or “unmethylome”. The polynucleotide may be a DNA sample, or may be a mixed sample, comprising DNA and RNA. The sample maybe a DNA sample comprising an epigenome. The terms “epigenome” and “epigenetic” as used herein, unless otherwise specified, refer to the 35 chemical modification of a polynucleotide or genome in such a way that gene expression is regulated. Thus, the method maybe a method of determining the epigenetic profile of a genomic DNA sample, the method further comprising determining the epigenetic profile based on the modification status of nucleotide residues in the sample. For example, the 5 method may comprise determining the epigenetic profile based on the modification status of cytosine residues in CpG dinucleotides of the sample, i.e. the CpG modification status. The method may be a method of analysing a polynucleotide sample, such as a DNA 10 sample, from a subject. The method may be a method for determining the modification status of one or more specific nucleotides in a polynucleotide sample from a subject. 15 The method may be a method for determining the modification status of the cytosine residue in one or more specific CpG dinucleotides of a biomarker in a sample from a subject. The method may be a method for determining the modification status of the cytosine 20 residues in one or more CpG dinucleotides of a plurality of genomic regions in a sample from a subject. The method maybe a method for determining the modification status of one or more specific adenine nucleotides of a biomarker in a sample from a subject. 25 The method may be a method for determining the modification status of one or more adenine residues in a plurality of genomic regions in a sample from a subject. The method maybe a method for determining the modification status of one or more 30 specific adenine nucleotides of one or more biomarkers in a sample from a subject. The method may be an in vitro method performed on a polynucleotide sample that has previously been obtained from a subject. 35 The term “subject”, as used herein, may refer to any type of organism, including for example, a mammalian species (such as a human or domesticated animal), other animal species, a plant such as a crop, or other type of organism, including single celled organisms, and viruses. The subject may be a developing organism, such as an embryo or foetus. The subject may be a healthy individual. The subject may be an individual that has, or is suspected of having, a disease or predisposition to a disease. The subject 5 may be an individual in need of therapy or suspected of needing therapy. The method may be a method of determining the disease status of a subject. Accordingly, the method may comprise determining the modification status of nucleotide residues in a polynucleotide in a sample from the subject using the disclosed 10 method, and determining the disease status based on the modification status. For example, the method may comprise determining the modification status of cytosine residues in CpG dinucleotides of a polynucleotide in a sample from the subject using the disclosed method, and determining the disease status based on the CpG modification status. 15 The method maybe a method of diagnosing a disease in a subject. Accordingly, the method may comprise determining the modification status of nucleotide residues in a polynucleotide in a sample from the subject using the disclosed method, and diagnosing the disease based on the modification status. For example, the method may 20 comprise determining the modification status of cytosine residues in CpG dinucleotides of a polynucleotide in a sample from the subject using the disclosed method, and diagnosing the disease based on the CpG modification status. The method may be a method of making a disease prognosis in a subject. Accordingly, 25 the method may comprise determining the modification status of nucleotide residues in a polynucleotide in a sample from the subject using the disclosed method, and making a disease prognosis based on the modification status. For example, the method may comprise determining the modification status of cytosine residues in CpG dinucleotides of a polynucleotide in a sample from the subject using the disclosed method, and 30 making a disease prognosis based on the CpG modification status. Methods comprising making a determination based on the nucleotide modification status may comprise comparing the modification status of specific nucleotide residues in a polynucleotide from the subject to the modification status of the corresponding 35 residues in a reference sample. Accordingly, methods comprising making a determination based on the CpG modification status may comprise comparing the modification status of cytosine residues in CpG dinucleotides of a polynucleotide from the subject to the CpG modification status of the corresponding residues in a reference sample. 5 The reference sample may comprise a polynucleotide from a healthy subject. The reference sample may comprise a polynucleotide from a diseased subject. The reference sample may comprise a polynucleotide from the same subject as the test sample, taken at a different time point and / or from a different location in the body. Differences in the nucleotide modification status between the test and reference samples may be 10 indicative of the presence or absence of a particular phenotype or clinical feature. The method may be a method of treating a subject. Accordingly, the method may comprise determining the modification status of nucleotide residues in a polynucleotide in a sample from the subject using the disclosed method, diagnosing a 15 disease based on the nucleotide modification status, and providing a therapeutic composition to the subject to treat the disease based on the diagnosis. For example, the method may comprise determining the modification status of cytosine residues in CpG dinucleotides of a polynucleotide in a sample from the subject using the disclosed method, diagnosing a disease based on the CpG modification status, and providing a 20 therapeutic composition to the subject to treat the disease based on the diagnosis. The method may be a method of determining a personalised or precision method of treatment for a subject. Accordingly, the method may comprise determining the modification status of nucleotide residues in a polynucleotide in a sample from the 25 subject using the disclosed method, determining the disease status of the subject based on the nucleotide modification status, and determining a personalised medical treatment for the subject based on the disease status. For example, the method may comprise determining the modification status of specific cytosine residues in CpG dinucleotides of a polynucleotide in a sample from the subject using the disclosed 30 method, determining the disease status of the subject based on the CpG modification status, and determining a personalised medical treatment for the subject based on the disease status. The method may be a personalised or precision method of treating a subject. 35 Accordingly, the method may comprise determining the modification status of nucleotide residues in a polynucleotide in a sample from the subject using the disclosed method, determining the disease status of the subject based on the nucleotide modification status, and providing a personalised medical treatment to the subject based on the disease status. For example, the method may comprise determining the modification status of specific cytosi ne residues in CpG dinucleotides of a 5 polynucleotide in a sample from the subject using the disclosed method, determining the disease status of the subject based on the CpG modification status, and providing a personalised medical treatment to the subject based on the disease status. The subject maybe an individual that has been diagnosed with having a disease. The 10 subject may be an individual that has been identified as being predisposed to, or at risk of having, a disease. The subject may be an individual that has not been diagnosed with having a disease. The subject may be an individual that has been diagnosed with cancer. The subject may 15 be pending or undergoing treatment such as a cancer therapy. The subject can be in remission of a cancer. Cancer can be identified on the basis of epigenetic variations. Cancer may be associated with both DNA hypomethylation and hypermethylation, but these two types of 20 epigenetic abnormalities may affect different DNA sequences, and occur at different stages of cancer progression. For example, genomic hypermethylation in cancer may be seen in CpG islands in gene regions, whereas hypomethylation may be observed in repeated DNA sequences in cancer, including heterochromatic DNA repeats, retrotransposons, and endogenous retroviral elements. In addition, unique sequences, 25 such as transcription control sequences, are often subject to cancer-associated hypomethylation. These epigenetic changes may be detected using the disclosed method. Thus, the method may be a method of diagnosing cancer in subject. Accordingly, the 30 method may comprise determining the modification status of nucleotide residues in a polynucleotide in a sample from the subject using the disclosed method, and making a cancer diagnosis based on the nucleotide modification status. For example, the method may comprise determining the modification status of cytosine residues in CpG dinucleotides of a polynucleotide in a sample from the subject using the disclosed 35 method, and making a cancer diagnosis based on the CpG modification status. The frequency of cancer-linked DNA hypomethylation, the nature of the affected sequences, and the absence of associations with DNA hypermethylation are belived to suggest a role for DNA hypomethylation early in carcinogenesis and cancer formation, but can also be associated with tumor progression. 5 Thus, the method may be a method of detecting cancer in a subject. Accordingly, the method may comprise determining the modification status of nucleotide residues in a polynucleotide in a sample from the subject using the disclosed method, and determining the presence or absence of cancer based on the modification status of the 10 nucelotide residues. For example, the method may comprise determining the modification status of cytosine residues in CpG dinucleotides of a polynucleotide in a sample from the subject using the disclosed method, and determining the presence or absence of cancer based on the modification status of the cytosine residues. The method may comprise determining the modification status of adenine residues in a 15 sample from the subject using the disclosed method, and determining the presence or absence of cancer based on the modification status of the adenine residues. The method may comprise the analysis of a plurality of genomic regions, and detecting the presence or absence of cancer from the modification status of nucleotide residues in 20 the plurality of genomic regions. The method may be a method of detecting any type of cancer. Different types of cancer may be preferentially detected and / or analysed using different sampling approaches based on the disclosed method. 25 The method may be a method of treating cancer in a subject. Accordingly, the method may comprise determining the modification status of nucleotide residues in a polynucleotide in a sample from the subject using the disclosed method, diagnosing a cancer based on the nucleotide modification status, and providing a therapeutic 30 composition to the subject to treat the cancer based on the diagnosis. For example, the method may comprise determining the modification status of cytosine residues in CpG dinucleotides of a polynucleotide in a sample from the subject using the disclosed method, diagnosing a cancer based on the CpG modification status, and providing a therapeutic composition to the subject to treat the cancer based on the diagnosis. The method may be a method of determining a personalised or precision method of cancer treatment for a subject. Accordingly, the method may comprise determining the modification status of nucleotide residues in a polynucleotide in a sample from the subject using the disclosed method, determining the genetic profile of the cancer based 5 on the nucleotide modification status, and determining a personalised medical treatment for the subject based on the genetic profile of the cancer. For example, the method may comprise determining the modification status of specific cytosine residues in CpG dinucleotides of a polynucleotide in a sample from the subject using the disclosed method, determining the genetic profile of the cancer based on the CpG 10 modification status, and determining a personalised medical treatment for the subject based on the genetic profile of the cancer. The method may be a personalised or precision method of treating cancer in a subject. Accordingly, the method may comprise determining the modification status of 15 nucleotide residues in a polynucleotide in a sample from the subject using the disclosed method, determining the genetic profile of the cancer based on the nucleotide modification status, and providing a personalised medical treatment to the subject based on the genetic profile of the cancer. For example, the method may comprise determining the modification status of specific cytosine residues in CpG dinucleotides 20 of a polynucleotide in a sample from the subject using the disclosed method, determining the genetic profile of the cancer based on the CpG modification status, and providing a personalised medical treatment to the subject based on the genetic profile of the cancer. 25 The method may be an in vitro method performed on a DNA sample that has previously been obtained from a tissue biopsy. Biopsy is a diagnostic procedure for cancers and other diseases. For example, tissue biopsy may provide material for cancer genotyping, which may assist in the design of targeted therapeutic approaches. 30 The method may be a method of genotyping a cancerous or otherwise diseased tissue. Accordingly, the method may comprise determining the modification status of nucleotide residues in a polynucleotide from a biopsy of the tissue using the disclosed method. For example, the method may comprise determining the modification status of cytosine residues in CpG dinucleotides of a polynucleotide from a biopsy of the tissue 35 using the disclosed method. The method may further comprise designing a targeted therapeutic approach based on the nucleotide modification status. The method may be performed on a DNA sample that has previously been obtained from a biopsy of any type of tissue from a subject. 5 Existing tissue biopsy-based cancer diagnostic procedures may have limitations in relation to the analysis of the development and progression of certain types of cancers, due to tumor heterogeneity and evolution. Liquid biopsy, which has the advantage of minimal invasiveness, has shown potential in 10 detecting cancers, including early stage cancers and pre-cancerous lesions. The analysis of cell-free DNA (“cfDNA” or “circulating cfDNA”), which refers to DNA present at very low concentration in various bodily fluids, comprises extracellular nucleic acid fragments, for example, released by damaged cells during apoptosis, necrosis, or secretion. cfDNA has been found to exhibit the genetic and epigenetic alterations of 15 cancers, including mutations, copy number alterations, chromosomal rearrangements, hypermethylation, and hypomethylation. The analysis of cfDNA has the potential to revolutionise the detection of early stage cancers and other diseases. In samples from cancer patients, cfDNA may comprise circulating tumor DNA (“ctDNA”), which is cell free tumor-derived fragmented DNA in a bodily fluid. Thus, cfDNA may comprise 20 ctDNA. As demonstrated in the enclosed examples, the disclosed method has advantageously been found to be capable of providing high quality and consistently reproducible results from the very low concentrations of nucleic acid that are typically present in liquid biopsy (such as circulating cfDNA) samples. Thus, the sample for use in the disclosed method may be a cfDNA sample. The sample may consist of or 25 comprise ctDNA. Various liquid biopsy samples may be used for the analysis of cfDNA, including blood, plasma, urine, and spinal fluid. Preferably the liquid biopsy sample may be a blood or plasma sample. 30 Thus, the method may be an in vitro method of diagnosing disease in a cfDNA sample from a subject. Accordingly, the method may comprise determining the modification status of nucleotide residues in a cfDNA sample from the subject, and diagnosing the disease based on the nucleotide modification status. For example, the method may 35 comprise determining the modification status of cytosine residues in CpG dinucleotides of a cfDNA sample from the subject, and diagnosing the disease based on the CpG modification status. The method may also be used to determine the tissue of origin of the nucleic acid 5 present in a liquid biopsy sample. Thus, the method may be a method of identifying the cellular origin of cfDNA in a sample from a subject. Accordingly, the method may comprise determining the modification status of nucleotide residues in a polynucleotide in a sample from the 10 subject using the disclosed method, and identifying the cellular origin of the DNA based on the modification status of nucleotide residues in the sample. The method maybe a method of diagnosing the recurrence of cancer in a subject. Accordingly, the method may comprise determining the modification status of 15 nucleotide residues in a cfDNA sample from the subject, comparing the nucleotide modification status to the nucleotide modification status of a tumour sample from the subject that has previously been determined using the disclosed method, and diagnosing the recurrence of the cancer in the subject based on the comparison. The cfDNA sample from the subject may be a blood sample. The tumour sample from the 20 subject may be a sample of a solid tumour, for example previously obtained from the subject in a surgical procedure. The method may be a method of sequencing polynucleotides, the method comprising sequencing the polynucleotides of the first and / or second fraction to generate a 25 plurality7 of sequencing reads. The method may further comprise comparing the sequences of the sequencing reads to a reference sequence or genome to determine the genomic location of the sequencing reads. The method maybe a method of preparing a profile of the modification status of 30 nucleotide residues in a reference sequence, such as one or more regions of a genome. Accordingly, step 6 of the method may comprise sequencing the polynucleotides of the first and / or second fraction to generate a plurality of sequencing reads, comparing the sequences of the sequencing reads to the reference sequence, such as a reference genome or genomic region, to determine the location, such as the genomic location, of 35 the sequencing reads within the reference sequence and thereby the presence or otherwise of unmodified nucleotide residues at specific locations w ithin the reference sequence, such as the one or more regions of the genome. For example, the method may be a method of preparing a profile of the modification 5 status at the cytosine C5 position of each CpG dinucleotide in a reference sequence, such as one or more regions of a genome. Accordingly, step 6 of the method may comprise sequencing the polynucleotides of the first and / or second fraction, to generate a plurality of sequencing reads, comparing the sequences of the sequencing reads to the reference sequence, such as a reference genome or genomic region, to 10 determine the location, such as the genomic location, of the sequencing reads within the reference sequence and thereby the presence or otherwise of unmodified cytosine residues in specific CpG dinucleotides within the reference sequence, such as across the genome or across one or more regions of the genome. 15 In another example, the method may be a method of preparing a profile of the modification status of cytosine residues at the N4 position, in a reference sequence, such as one or more regions of a genome. Accordingly, step 6 of the method may comprise sequencing the polynucleotides of the first and / or second fraction to generate a plurality of sequencing reads, comparing the sequences of the sequencing reads to the 20 reference sequence, such as a reference genome or genomic region, to determine the location, such as the genomic location, of the sequencing reads within the reference sequence and thereby the presence or otherwise of cytosine residues unmodified at the N4 position, at specific locations within the reference sequence, such as the across the genome or across one or more regions of the genome. 25 In another example, the method may be a method of preparing a profile of the modification status of adenine nucleotides at the N6 position, in a reference sequence, such as one or more regions of a genome. Accordingly, step 6 of the method may comprise sequencing the polynucleotides of the first and / or second fraction to generate 30 a plurality of sequencing reads, comparing the sequences of the sequencing reads to the reference sequence, such as a reference genome or genomic region, to determine the location, such as the genomic location, of the sequencing reads within the reference sequence and thereby the presence or otherwise of adenine residues unmodified at the N6 position, at specific locations within the reference sequence, such as the across the 35 genome or across one or more regions of the genome. Sample The sample for use in the disclosed method may be obtained from any type of cell or tissue. For example, the sample may be obtained from tissue, blood, plasma, serum, urine, saliva, stool, cerebrospinal fluid, buccal swab, pleural tap, etc.. The sample may 5 be obtained from tissue. The sample may be obtained from blood. The sample may be a DNA sample. The DNA sample may be a cfDNA sample, which may comprise ctDNA. For example, the DNA sample may be a cfDNA sample from peripheral blood. A “cell-free” sample as used herein, refers to nucleic acids not contained within or otherwise bound to a cell or, remaining in a sample following the removal of intact cells. Cell-free 10 nucleic acids can include, for example, all non-encapsulated nucleic acids sourced from a bodily fluid (e.g., blood, plasma, serum, cerebrospinal fluid, etc.) from a subject. The cfDNA may be released into bodily fluid through secretion or a cell death process. The cfDNA may comprise DNA released into bodily fluid from cancer cells, and may be referred to as comprising circulating tumor DNA (ctDNA). The cfDNA may be released 15 from healthy cells. Methods of preparing samples for use in the disclosed method, such as DNA samples, comprising, for example, extracting and purifying nucleic acids such as DNA from cells or tissues will be known to the skilled person. Any method that is suitable for preparing 20 a polynucleotide sample for analysis, such as sequencing, may be used. The sample may comprise fragmented DNA. Fragmentation may be performed using any method used in the analysis of DNA, such as any fragmentation method used in the preparation of a DNA sample for genetic sequencing. For example, the DNA may be 25 fragmented enzymatically, chemically by acoustic shearing, mechanical shearing (example, French pressure cells), sonicating, hydrodynamic shearing or chemically (for example, heat and divalent metal cation). In some implementations of the disclosed method, fragmentation of the DNA sample is 30 not required. In such embodiments, the method does not include fragmentation of the DNA sample. For example, in embodiments in which the DNA sample is degraded or fragmented, such as when cfDNA is extracted from blood, the DNA sample may be used directly in the disclosed method. Typically, cfDNA extracted from blood substantially comprises DNA fragment sizes of 50-300 base pairs (bp) in length. The method may comprise the selection of polynucleotides, such as DNA fragments, of a desired length. Thus, in some embodiments, the method may comprise, before the use of a 5 methyltransferase (step 1), a step of fragmenting the polynucleotide and / or selecting polynucleotides of a desired length. The method may comprise the use of polynucleotides, such as DNA fragments, substantially or predominantly having a length between to and 500 bp, such as 10 between 30 and 400 bp, and preferably between 50 and 300 bp in length. In some embodiments, the method may comprise the use of polynucleotides, such as DNA fragments, substantially or predominantly having a length in the region of between too and 250 bp, preferably between about 150 and 180 bp, to match the sample to the DNA sequencing read length. 15 The method may comprise the use of polynucleotides corresponding to an amount of DNA in the range of about 1 fg to about 1 pg, such as about 10 fg to about too ng, about too fg to about 10 ng, about 1 pg to about 1 ng. The sample may comprise a quantity of DNA in the picogram range. The sample may comprise less than lpg of DNA, such as 20 less than soong, less than toong, or less than tong of DNA. Preferably, the sample comprises between ing and toong of DNA. Derivatisation of DNA - Tag The polynucleotide may be derivatised using a methyltransferase enzyme to apply a tag 25 to unmodified nucleotide residues. For example, the polynucleotide may be derivatised using an appropriate methyltransferase enzyme to apply a tag to specific nucleotide residues that are unmodified in target positions. Thus, to determine the modification status of a specific nucleotide residue at a specific target position, a methyltransferase enzyme may be used that is configured to apply a tag to the target nucleotide residue in 30 the target position. The same tag maybe used to derivatise different unmodified nucleotides and / or different target positions. 35 In other embodiments, different tags may be used to derivatise different unmodified nucleotides and / or different target positions. In some embodiments, the polynucleotide may be derivatised with different tags (e.g. on different unmodified nucleotides) sequentially, for example, with the inactivation of the first methyltransferase before the addition of a second, different methyltransferase. 5 The use of a plurality of different tags advantageously allows the tags to be independently functionalised, for example, to provide selective enrichment / fractionation. 10 As used herein, unless otherwise indicated, references to the “target position” in which an unmodified nucleotide residue is unmodified refer to a specific position within the chemical structure of the nucleotide. The tag may be applied in the cytosine C5 position. Accordingly, the fragmented DNA 15 sample will comprise tagged residues. A tagged cytosine residue may be understood to have the following structure: I ^vw wherein R2 is the tag. 20 The tag may be applied in the N4 position in cytosine. Accordingly, the fragmented DNA sample will comprise tagged residues. A tagged cytosine residue may be understood to have the following structure: wherein R2 is the tag. 25 Similarly, a tag maybe added at the N6 position of adenine, the N2 or N7 position of guanine or at the 2’-0H position of ribose. In some embodiments, a tag at the 2’-0H position of ribose is a tag at the 2’-0H position of a terminal ribose. Accordingly, a tagged adenine residue may be understood to haw the following structure: R2 HN uvw wherein R2 is the tag. 5 The term “tag”, which may also be referred to as a “linker”, a “functional linker” or “DNA tag”, as used herein, unless otherwise specified, refers to a reactive moiety that is applied site-specifically to the polynucleotide, such as fragmented DNA. Polynucleotides that have been tagged in this way may be referred to as “derivatised”. 10 The disclosed method may comprise the use of a methyltransferase cofactor analogue, such as a synthetic methyltransferase cofactor analogue, comprising the tag and a methyltransferase-binding moiety. Thus, the method may comprise the use of a methyltransferase enzyme to catalyse the transfer of the tag from the methyltransferase 15 cofactor analog to an unmodified nucleotide residue, such as to the C5 position of a cytosine base of an unmodified CpG dinucleotide, in a polynucleotide sample. The presence of a modification, such as a methyl group or other chemical modification of the nucleotide residue, such as in the C5 position within a CpG dinucleotide, prevents the transfer of the tag. Thus, in the disclosed method, only nucleotides, such as CpG 20 dinucleotides, that are unmodified (such as unmethylated) in this position may be labelled with a tag. The methyltransferase cofactor may be an ion of formula (I): 25 wherein, X is S or Se: LUs -CH2- or -CH2CH2-; R2 is the tag; R3 and R1 are independently H or an optionally substituted C r f> alkyl an optionally substituted C2-6 alkenyl or an optionally substituted C2-6 alkynyl: or R3 and R4 together 5 with the nitrogen to which they are attached, form an optionally substituted 5- or 6-membered heterocyclyl ring; and R5 is NH2, NHBoc or H; or a salt, solvate or tautomer thereof. 10 The ion of formula (I) may be provided together with a counterion. The counterion may be an organic or inorganic anion carrying one or more negative charges. The counterion may be formate or acetate. R2 maybe -CH2-U-[L3]m-[HM]n-[L2]p-[R6]q, wherein: 15 m, n, p and q are each independently selected from 0 and 1; L2 is a linker; HM is a hydrolysable moiety; L3 is a linker; U is an unsaturated group selected from an alkene, an alkyne, an aromatic group (e.g. 20 aryl), a carbonyl group, SO and S02; R6 is a heavy atom or a heavy atom cluster suitable for phasing of X-ray diffraction data, a radioactive or stable rare isotope, a fluorophore, a fluorescence quencher, an affinity tag, a crosslinking agent, a nucleic acid cleaving reagent, a spin label, a chromophore, a protein, peptide or amino acid which may optionally be modified a 25 nucleotide, nucleoside or nucleic acid which may optionally be modified, a carbohydrate, a lipid, a transfection reagent, an intercalating agent, a nanoparticle or bead, or a functional group, wherein the functional group is selected from the group consisting of: an amino group (including a protected amino), a thiol group, a 1,2-diol group, a hydrazino group, a 30 hydroxyamino group, a haloacetamide group, a maleimide group, a cyanide group, a cyclic hydrocarbon (such as a bridged cyclic hydrocarbon (e.g. norbornene) or a cycloalkyl group (e.g. a C3-6 cycloalkyl), a halo group (e.g. -F, -Cl, -Br, -I), an aldehyde group, a ketone group, a 1,2-aminothiol group, a azido group, an isothiocyanate or thiocyanate group, an alkene group, such as a terminal alkene, an alkyne group, such as 35 a terminal alkyne group, a 1,3-diene function, a dienophilic function (e.g. an activated carbon-carbon double bond), an arylhalide group, an arylboronic acid group, a terminal haloalkyne group, a terminal silylalkyne group, -N=C=O; -N=C=S, -0-C(0)NH2, a protected amino, a group comprising a sterically strained alkyne or alkene (such as norbornene or DBCO), a nitrone, a tetrazine, a tetrazole, and 1,2-a mi nothiol group. 5 In embodiments where R4 is an optionally substituted Ci-4 alkyl an optionally substituted C2-4 alkenyl or an optionally substituted C2-4 alkynyl, the alkyl, alkenyl or alkynyl may be unsubstituted or substituted with one or more substituents selected from the group consisting of: -NR?R8; -OH; -SH; -CN; -C(0)0R7; -C(O)R7; C(O)NR7R8; N3; and halo, wherein R' and R8 are independently H or a Cw alkyl. Halo 10 may be F, Cl, Br or I. Similarly, in embodiments where R3 and R4 together with the nitrogen to which they are attached, form an optionally substituted 5- or 6-membered heterocyclyl ring, the 5-or 6-membered heterocyclyl ring may be unsubstituted, or substituted with one or 15 more substituents selected from the group consisting of: -NR7R8; -OH; -SH; -CN; -C(0)0R7; -C(O)R7; C(O)NR R8; N3; and halo, wherein R and R8 are independently H or a Ci 4 alkyl. Halo may be F, Cl, Br or I. Synthetic methyltransferase cofactors are described in more detail in 20 PCT / GB2022 / 052438, EP3186266B1 and US8008007B2. It may be appreciated that preferred embodiments of the X, L1, R2, R3 and R4 groups in the compound of formula (I) may be as defined for the equivalent groups in these applications. X maybe S. 25 L1 may be -CH2CH2-. R3 may be H. Alternatively, R3 may be an optionally substituted Ci-4 alkyl an optionally substituted C2 4 alkenyl or an optionally substituted C2 4 alkynyl, more preferably an optionally substituted methyl or an optionally substituted ethyl. The alkyl, alkenyl or 30 alkynyl may be unsubstituted or substituted with an OH. Accordingly, R ■ may be - CH2CH20H. R4 maybe H. 35 R5 may be NH2. In some embodiments, q is 1. In some embodiments, R6 is -N3. p may be i. 5 L2 maybe a linker comprising a backbone of between 1 and 50 atoms, between 2 and 40 atoms, between 3 and 30 atoms, between 4 and 20 or between 5 and 15 atoms. The backbone may be made up of carbon, oxygen and / or nitrogen atoms. In embodiments where the linker comprises a cyclic group, the backbone may be understood to consist of the atoms which define the shortest possible route between the two ends of the linker 10 group. In some embodiments, L2 comprises between 1 and 5 groups selected from an optionally substituted hydrocarbon, an optionally substituted polyether chain, an arylene moiety and a (C=O)NH group. 15 The hydrocarbon may be an optionally substituted alkylene, preferably an optionally substituted CM0 alkylene and more preferably a C1-5 alkylene. The optionally substituted polyether chain may be an optionally substituted polyethylene glycol chain. The polyethylene glycol chain may comprise up to 15 monomers, up to 10 monomers or 20 up to 5 monomers of ethylene glycol. In some embodiments, the polyethylene glycol chain consists of between 1 and 5 or between 2 and 3 monomers of ethylene glycol. The arylene moiety may be a CeH4 phenylene ring. Accordingly, in some embodiments, L2 may be: O wherein w is an integer from between 1 and 15, e.g. between 2 and 10 or between 3 and 5. In some embodiments, w is 2 or 3. Alternatively, in some embodiments, p is 0. 30 In some embodiments, n is 1. The hydrolysable moiety may be a Schiff base, for example, an imine moiety, an oxime moiety and / or a hydrazone moiety. 10 In some embodiments, the hydrolysable moiety comprises a disulphide (S-S) bond. In some embodiments, the hydrolysable moiety is 15 In some embodiments, n is 0. In some embodiments, m is 1. L3 may be a linker comprising a linear chain of from 1 to 20, from 2 to 15, from 3 to 10 or from 4 to 9 atoms. The atoms may be carbon, oxygen and / or nitrogen atoms). The linker may be substituted or unsubstituted. In some 20 embodiments, L3 comprises an optionally substituted hydrocarbon (e.g. an alkyl) chain. In some embodiments, L3 comprises an optionally substituted linear CHn alkyl chain, e.g. an optionally substituted C2-g or an optionally substituted C4-6 alkyl chain. In some embodiments the alkyd chain is unsubstituted. In some embodiments the alkyl chain is substituted. In some embodiments, 13 is a linear, unsubstituted C2, C3 or C4 alkyl chain. In some embodiments, U is In some embodiments, the synthetic methyltransferase cofactor may be: OH OH , where R is H. 70 The above compound may be called ETA-AdoHcy-N3 Derivatisation ofDNA - Methyltransferase The methyltransferase may be any methyltransferase that is capable of using S- 15 adenosyl methionine as a cofactor. Thus, the methyltransferase may be an S-adenosylmethionine-dependent methyltransferase, such as an S-adenosyl-L-methionine-dependent methyltransferase. The methyltransferase may be a cytosine-5 (C5) methyltransferase, such as a bacterial 20 cytosine C5 methyltransferase. The methyltransferase may be an adenine methyltransferase, such as a bacterial adenine methyltransferase. For example, the methyltransferase maybe M.TaqI, which is a DNA adenine methyltransferase. 25 The methyltransferase maybe a methyltransferase from Mycoplasma. The methyltransferase maybe a constitutively active methyltransferase. The methyltransferase may be one of the enzymes described in US 2017 / 0283453. The methyltransferase maybe M.Mpel, M.Hhal, M.SssI, M.AccII, M.MspI or M.TaqI. 5 The methyltransferase may be an active mutant, variant, and / or fragment of M.Mpel, M.Hhal, M.SssI, M.AccII, M.MspI or M.TaqI. M.Mpel has been found to be particularly advantageous for use in the disclosed method, in part due to being particularly non-selective in terms of target locus. Thus, 10 the methyltransferase may be M.Mpel or an active mutant, variant, and / or fragment thereof. Although methyltransferase enzymes share a relatively low level of sequence similarity, they do share a highly conserved structural fold. This conserved fold is known as the 15 Rossmann fold and comprises a series of beta strand and alpha helical segments, in which the beta strands are hydrogen bonded to form a beta-sheet. The cofactor binding pocket of the methyltransferase enzyme may be modified w ithin the Rossman fold, for example, by substitution of one or more amino acids, to improve 20 the suitability of the enzyme for use in the disclosed method, such as, for example, by improving cofactor compatibility. For example, one or more amino acids within the Rossman fold of the methyltransferase enzyme may be substituted, for example, to reduce or relieve potential steric interaction with the cofactor analogue. 25 The methyltransferase may be modified such that an amino acid having a relatively large side chain, such as, for example, glutamine or asparagine may be substituted for an amino acid comprising a shorter side chain, such as, for example, alanine. The use of a methyltransferase enzyme that has been modified in this way may be particularly desirable when larger cofactor analogues are used, such as cofactor analogues 30 comprising transferrable groups with longer alkyl-chains than those with shorter chains. The methyltransferase may be any bacterial cytosine C5 methyltransferase enzyme comprising one or more, such as 2,3, 4, 5, 6, or 7, amino acid substitutions in the 35 Rossman fold. The methyltransferase may be any bacterial cytosine C5 methyltransferase enzyme comprising an amino acid substitution in the position of the amino acid residue of the Rossman fold corresponding to the residue that is Gln82, Tyr254, and / or Asn3O4 in the w ild type sequence of the M.Hhal methyltransferase (i.e. the sequence having the NCBI 5 accession number P05102). The methyltransferase may be any bacterial cytosine C5 methyltransferase enzyme comprising an alanine residue in the position of the amino acid residue of the Rossman fold corresponding to the residue that is Gln82, Tyr254, and / or Asn3O4 in the wild type 10 sequence of the M.Hhal methyltransferase (i.e. the sequence having the NCBI accession number P05102). The methyltransferase maybe an M.Mpel methyltransferase. The M.Mpel methyltransferase enzyme has been found to be particularly advantageous for use in the 15 disclosed method due to non-selectively targeting any and all CpG dinucleotides for modification. The methyltransferase maybe, or may comprise, a variant, and / or fragment of the wild type M.Mpel sequence, which is defined as the sequence having the NCBI accession 20 number BAC44284. The methyltransferase may be, or may comprise, a variant, and / or fragment of the wild type M.Mpel sequence, comprising at least 80% sequence identity, such as at least 85%, 90%, or 95% sequence identity7 to the wild type M.Mpel sequence. 25 The methyltransferase may comprise one or more, such as 2,3,4, 5, 6 or 7 amino acid substitutions relative to the wild type M.Mpel sequence having the NCBI accession number BAC44284. 30 The use of the methyltransferase enzyme to apply the tag to unmodified nucleotides, such as unmodified CpG dinucleotides, may be carried out under conditions w hich enable the methyltransferase to transfer the tag from the methyltransferase cofactor analogue to the target DNA. The reaction mixture may be incubated at a temperature of from to to 6o°C, from 20 to 50°C, or from 30 to 40°C. Preferably, the reaction mixture may be incubated at a temperature of about 37°C. 5 The incubation may be performed for a time sufficient to enable transfer of the tag to all of the available unmodified nucleotides, such as unmodified CpG dinucleotides, in the fragmented DNA sample. The incubation may be performed for a period of 5 minutes to 5 hours, to minutes to 4 hours, 15 minutes to 3 hours, 30 minutes to 2 hours, or 40 to 90 minutes. Preferably the incubation is performed for a period of about 1 hour. 10 The incubation may be performed in a suitable buffer at a pH that is selected based on the methyltransferase that is being used. For example, the pH may be between 7.5 and 8.5, such as between 7.8 and 8.2, or about pH 8. 15 Methyltransferase inactivation After an appropriate incubation to label the unmodified nucleotides, such as unmodified CpG dinucleotides, in the sample with a tag, the presence of methyltransferase in the subsequent processing of the sample has been found to reduce the efficiency of the method. Thus, the method comprises the inactivation of the 20 methyltransferase enzyme. In some embodiments, methyltransferase enzymes have been found to bind tightly to DNA, thereby inhibiting downstream processing of the DNA. Methods comprising removal of the methyltransferase or purification of the sample have been found to 25 reduce the efficiency of the process due to the additional time and reagents required and due to the loss of sample. Therefore the methyltransferase enzyme may be inactivated in the reaction mixture. The inactivation of the methyltransferase in this way has surprisingly been found to 30 provide significant processing efficiencies in the disclosed method. The terms “inactivated” and “inactivation” as used herein, unless otherwise specified, refer to any alteration in the structure and / or function of the methyltransferase that prevents further activity of the methyltransferase on the target polynucleotide. Thus, the terms “inactive” and “inactivated” as used herein, unless otherwise specified, refer to an 35 enzyme that has less than 10%, such as less than 5%, less than 2%. or preferably less than 1% of its maximum activity7. Alterations in the structure and / or function of the methyltransferase that prevent further activity of the methyltransferase on the target polynucleotide may include, for example, denaturation, modification, inhibition, and / or fragmentation of the 5 methyltransferase. Thus, in some embodiments, inactivation of the methyltransferase may comprise denaturation of the methyltransferase. In some embodiments, inactivation of the methyltransferase may comprise modification of the methyltransferase. In some embodiments, inactivation of the methyltransferase may comprise inhibition of the methyltransferase. In some embodiments, inactivation of the 10 methyltransferase may comprise fragmentation of the methyltransferase. The methyltransferase may be inactivated by any suitable method. Suitable methods include changing the environmental conditions of the methyltransferase, and targeted inactivation of the methyltransferase. 15 Changing the environmental conditions may consist of or comprise, for example, changing the temperature and / or pH of the reaction mixture. Thus, in some embodiments, inactivation of the methyltransferase may comprise 20 incubation of the reaction mixture at a temperature of from 55 to 8s°C, from 60 to 8o°C, or from 65 to 75°C. In some embodiments, inactivation of the methyltransferase may comprise incubation of the reaction mixture at a temperature of from 55 to 6s°C, such as at or about 6o°C, or from 75 to 8s°C, such as at or about 8o°C. 25 Inactivation of the methyltransferase may comprise incubation at an elevated temperature for a period of 5 minutes to 1 hour, or 10-30 minutes. Preferably inactivation of the methyltransferase may comprise incubation at an elevated temperature for about 15 minutes. 30 Targeted inactivation of the methyltransferase may comprise the addition of an agent to alter the structure and / or function of the methyltransferase. Such an agent may comprise, for example, a methyltransferase inhibitor. Any suitable methyltransferase inhibitor maybe used, including, for example, 5-azacitidine, decitabine, clofarabine, arsenic trioxide, guadecitabine, RX-3117,5-fluoro-2’-deoxycytidine, 5,6-diliydro-,5- 35 azacytidine, cladribine, fludarabine, fazarabine, procaine, EGCG, hydralazine, genistein, equol, curcumin, disulfiram, resveratrol, and / or caffeic acid. The methyltransferase inhibitor may be a S-Adenyl-l-methionine (SAM) analogue, such as sinefungin or S-adenosyl-l-homocysteine (SAH). The methyltransferase may be inactivated in the reaction mixture after an appropriate 5 incubation to label the unmodified nucleotides, such as unmodified CpG dinucleotides, with a tag, thereby terminating the derivatisation reaction. Advantageously, the presence of the inactivated methyltransferase has not been found to be detrimental to subsequent processing. On the contrary, the inactivation of the methyltransferase enzyme in the reaction mixture this way, rather than by removal or dilution, has been 10 found to provide increased efficiencies and significantly improved yields in subsequent steps of the process. The inactivation of the methyltransferase may, therefore, provide significant advantages by removing the requirement for purification of the polynucleotide at this 15 stage, and permitting the efficient combination of the methyltransferase and library preparation processes in a single reaction mixture. These efficiency advantages are shomi in the Examples. Library preparation 20 After inactivation of the methyltransferase, the polynucleotide sample may be modified into a form that is compatible for high throughput sequencing. This process may be referred to as “preparing a sequencing library” or “library preparation”. As used herein, unless otherwise indicated, a “sequencing library” refers to a plurality 25 of polynucleotides, each comprising a sequencing adaptor, such as a sequencing adaptor arranged for use in next generation sequencing. Accordingly, “preparing a polynucleotide sample into a sequencing library”, as used herein, unless otherwise indicated, refers to the addition of one or more sequencing adaptors to the polynucleotides of the sample. 30 In addition to adaptor ligation, the process of preparing a polynucleotide sample into a sequencing library may comprise end repair and / or A-tailing of the polynucleotides. Preferably the process of preparing a sequencing library does not comprise the 35 combining the polynucleotides together to form an extended ligated polynucleotide for use, for example, in a sequencing method comprising nanopore technology. The reaction mixture comprises a buffer mixture such as the labelling buffer, together with inactive methyltransferase, excess cofactor analogue, and the polynucleotide sample. In general for sequencing applications, library7 preparation is typically 5 conducted with purified DNA. It has surprisingly been found by the inventors, however, that the library preparation process may be performed directly in the reaction mixture following methyltransferase inactivation and that the efficiency of the library preparation process is not compromised by the use of a different buffer, or the presence of unpurified sample, such as DNA, and / or residual enzyme and cofactor components 10 in the mixture. This finding provides a significant advantage over previous methods, offering significant efficiencies in terms of savings of time and reagents. In particular, the finding that any washing procedure may be avoided significantly preserves the level of polynucleotide sample present in the reaction mixture. 15 Thus, the method comprises the inactivation of the methyltransferase followed by library preparation without any intervening steps or clean-up process, for example, comprising removal of inactivated enzymes, exchange of reaction buffer, or isolation or purification of the polynucleotide sample. Performing library preparation at this stage, for example, prior to any enrichment process, and without the requirement for any 20 washing steps or clean-up of the sample, surprisingly provides significant processing efficiencies, including significantly reducing any loss of polynucleotide sample. The sample may be subjected to a library preparation process comprising end repair of the polynucleotide sample. 25 The end repair process may comprise removal of 3' overhangs, for example using a Klenow fragment-based enzyme. The end repair process may also comprise modifying 3’ ends as necessary to comprise a hydroxyl group. 30 The end repair process may additionally or alternatively fill 5' overhangs, for example, using a T4 DNA polymerase. The end repair process may also comprise phosphorylation of 5' ends where necessary, for example, using of a T4 polunucleotide kinase (PNK). 35 The polynucleotide sample may be subjected to a library preparation process comprising A-tailing. The A-tailing process may comprise the addition of an adenosine residue to the 3' ends of the polynucleotide sample. This process may reduce the possibility of the polynucleotides in the sample ligating to each other. The A-tailing process may also 5 increase the rate of adapter ligation, particularly in embodiments in which the adapters comprise a thymine overhang. The A-tailing process may comprise the use of an “exo-Klenow” enzyme. In some embodiments, the polynucleotide sample may be subjected to end repair and 10 A-tailing processes simultaneously. For example, an end repair and A-tailing buffer comprising end repair and A-tailing enzymes may be used. The polynucleotide sample may be subjected to a li brary preparation process comprising one or more “adapter ligation” processes comprising the ligation of 15 sequencing adapters to the polynucleotides in the sample. The term “adapter” as used herein, unless otherwise specified, refers to a short nucleic acid (such as less than about 500, less than 100, or less than about 50 nucleotides in length) that is typically at least partially double-stranded and is attached to either or 20 both ends of a nucleic acid molecule. The adapters may include a primer binding site for amplification of the sample. The adapters may include a primer binding site for sequencing applications, such as next-generation sequencing (NGS) applications. The adapters may include a binding site for capture probes, such as an oligonucleotide attached to a flow cell support. A plurality of adapters of the same or different 25 sequences may be attached to the polynucleotides in the sample. The ligated adapters may include a nucleic acid tag. The nucleic acid tag may be positioned relative to an amplification primer and / or sequencing primer binding site, such that the tag sequence is included in subsequent amplicons and sequence reads. In 30 some embodiments, a plurality- of adapters having the same sequence apart from different nucleic acid tags may be attached to the polynucleotides in the sample. The ligated adapters may include a barcode that can be introduced at one or both ends of the sample DNA molecule. A “barcode”, “indexing barcode”, or “molecular barcode” 35 as used herein, unless otherwise stated, refers to a nucleic acid molecule comprising a sequence that can serve as a molecular identifier. A barcode may be a type of nucleic acid tag. For example, individual "barcode" sequences may be added to the polynucleotides in the sample for use in next-generation sequencing (NGS) so that the sequencing read can be identified and sorted before the final data analysis. 5 The adapter ligation process may comprise the ligation of sequencing adapters to the polynucleotides in the sample. The adapter ligation process may comprise the use of a T4 DNA ligase. Advantageously, any sequencing adapters maybe used. Sequencing adapters that have been found to be particularly suitable for use in the disclosed method include, for example, any sequencing adapters suitable for use with high 10 throughout sequencing methods, such as sequencing applications on the Illumina platform. The method may comprise the use of a double-stranded indexing and unique dual indexing (UDI) adapter that enables efficient ligation and identification of PCR amplification replicates in the sequencing dataset. Other adapters may also be used, such as hairpin adapters. 15 In some embodiments, the adapter ligation processes may comprise the ligation of adaptors that do not comprise indexing barcodes, and such adaptors may be referred to herein as “stubby adapters”. 20 Barcodes may be applied to one or both ends of the polynucleotides as part of the library preparation process. In addition, or alternatively, barcodes may be added to one or both ends of the polynucleotides in a separate amplification step. Labelling of derivatised DNA 25 The method comprises the labelling of derivatised polynucleotides. Tags on the polynucleotides in the sample may be modified by the addition of an affinity label. The affinity label may be referred to as an “affinity label” or “label” w hen bound to the tag and an “affinity label precursor” beforehand. 30 It has been found that an affinity label may be added to the tag by the addition of the affinity label precursor to the reaction mixture. The finding that an affinity label maybe applied to the tag in this technically simple and efficient manner is advantageous in view of the fact that the reaction mixture comprises various components including , inactive methyltransferase, excess cofactor analogue, and the reagents and enzymes 35 required for 1 i brarx preparation. The finding that the affinity label may be added to the tag in this way provides a significant advantage over previous methods, by avoiding the requirement for a washing step, thereby providing efficiency savings in terms of time and reagents and avoiding any loss of sample. Thus, the method comprises the addition of an affinity label to the tag after library 5 preparation, without any intervening steps or clean-up process, for example, comprising removal of peptides or enzymes, exchange of reaction buffer, or isolation or purification of the sample. The affinity label may comprise biotin. 10 The affinity label precursor may be a compound of formula (II): R9-L4-R1O (ID 15 wherein: RQ is a reactive moiety configured to react with a group in the tag and to thereby form a bond therebetween; L* is a linker; and R10 comprises or consists of biotin. 20 R9 may be an optionally substituted 5 to 30 membered heterocyclyl, an optionally substituted 5 to 30 membered heteroaryl, an optionally substituted Ce-3o membered aryl or an optionally substituted C3 3() cycloalkv1. A multicyclic group may be understood to be a group comprising two or more fused 25 rings. Accordingly, a multicyclic group may have 2 or 3 fused rings. As used herein, a “heterocyclyl”, “heterocyclic” or “heterocycle” group includes nonaromatic saturated or partially saturated mono and multicyclic groups. A heterocyclic ring contains 1 or more heteroatoms in the ring, which may independently selected 30 from nitrogen, oxygen or sulfur. A multicyclic group may be understood to be multicyclic heterocyclyl group if it contains at least one heteroatom and at least one ring which is a non-aromatic saturated or partially saturated ring. As used herein, a “cycloalkyl” group includes non-aromatic saturated or partially 35 saturated mono and multicyclic groups. A multicyclic group may be understood to be multicyclic cycloalkyl group if it only contains carbon atoms in the rings and it contains at least one ring which is a non-aromatic saturated or partially saturated ring. As used herein, a “heteroaryl” group includes aromatic mono and multicyclic groups. A 5 heteroaryl ring contains 1 or more heteroatoms in the ring, which may independently selected from nitrogen, oxygen or sulfur. A multicyclic group may be understood to be multicyclic heteroaryl group if it contains at least one heteroatom and every ring is aromatic. 10 Preferably, R9 contains a triple bond. Preferably, Ry is an optionally substituted io to 20 membered multicyclic heterocyclyl, an optionally substituted 10 to 20 membered multicyclic heteroaryl or an optionally substituted C10-20 multicyclic cycloalkyl. R9 may be a 14 to 18 membered multicyclic 15 heterocyclyl, an optionally substituted 14 to 18 membered multicyclic heteroaryl or an optionally substituted Ci-18 multicyclic cycloalkyl. In a preferred embodiment, Ry is . wherein X2 is N or CH. Preferably, X2 is N. 20 L* may comprise between 1 and 12 groups, each group selected from an optionally substituted hydrocarbon, an optionally substituted polyether chain, NH, 0, S or S-S. The hydrocarbon may be an optionally substituted alkylene, preferably an optionally 25 substituted CmO alkylene and more preferably a Ct5 alkylene. The alkylene may be substituted with an OH or oxo group. Preferably, the alkylene is substituted with an oxo group. The optionally substituted polyether chain may be an optionally substituted 30 polyethylene glycol chain. The polyethylene glycol chain may comprise up to 15 monomers, up to 10 monomers or up to 5 monomers of ethylene glycol. Accordingly, L* may have the structure -L5-L6-L -L8-*, wherein L5 to L8 are each independently absent or an optionally substituted hydrocarbon, an optionally substituted polyether chain, an NH, 0, S or S-S; and 5 an asterisk indicates a point of bonding to R10. In some embodiments, L5 is an optionally substituted hydrocarbon. Accordingly, L5 may be C0CH2CH2. io In some embodiments, L6 is NH. In some embodiments, 17 is an optionally substituted hydrocarbon. Accordingly, 17 may be C0CH2CH2. 15 In some embodiments, L8 is an optionally substituted polyether chain. The optionally substituted polyether chain may be an optionally substituted polyethylene glycol chain. The polyethylene glycol chain may comprise up to 15 monomers, up to 10 monomers or up to 5 monomers of ethylene glycol. Accordingly, L8 may be (0CH2CH2)r, where r is an integer between 1 and 15, more preferably between 2 and 10 or between 3 and 5. In 20 some embodiments, r is 4. Accordingly, 17 may be 17 may have no charge. 25 Negatively charged linkers have been found to react poorly with the tag. Preferably, 17 is not negatively charged. R10 may have the following formula: 30 RU-(CH2)S-L9- wherein R11 is biotin s is an integer between 1 and 8; and 17 is absent or is COO or CONH. Sulfonated linkers have been found to react particularly poorly with the tag. Preferably, L4 and R10 are not sulfonated. 5 Preferably, the affinity label precursor is not DBCO-SS-biotin. Preferably, the affinity label precursor is not NHS-SS-biotin. A modified tagged cytosine residue, which comprises the affinity label, may be understood to have the following structure: JO < / wv wherein L4 and R10 are as defined above; and L10 is a linker. Li° maybe understood to be -CHa-U-tLsJm-fHMJn-EIAlp-L11-, wherein U, L2, L3, HM, m, 15 n and p are as defined above and L11 is a linker formed due to a reaction between the R6 and R9 groups. Accordingly, L11 may be , where an asterisk indicates a point of bonding to L+ and X2 is as defined above. 20 The present inventors have surprisingly found that in previous methods, such as that described by Kriukiene et al. (Nature Communications 2013 4:2190), DNA fragments having a significant density' of CpG sites, such as, for example, 5 or more CpG sites per 100 bp, may be underrepresented in the sequencing reads, thereby introducing bias to 25 the results. An advantage of the disclosed method is that if a purification process, such as DNA isolation, is performed at this point, it is the only clean-up step for the entire process, and this has been found to dramatically improve the efficiency and sensitivity of the process. This is made possible, firstly, by the inactivation of the methyltransferase and, secondly, the surprising finding that the enzymes used for library preparation exhibit high levels of activity in the resulting buffers, which are significantly different to the buffer mixtures designed for use in library preparation. 5 Thus, in some embodiments, the method may comprise, after the affinity labelling step (step 4), and before the fractionation step (step 5), a step of purifying the polynucleotide. Preferably the method involves no more than one step of purifying the polynucleotide. 10 Any suitable method for purifying the polynucleotide may be used. For example, DNA may be purified using a DNA purification kit. The DNA may be washed, for example, using ethanol, such as 80% ethanol, or other DNA washing buffer. After washing, the DNA may be eluted, for example, using a suitable elution buffer, such as phosphate 15 buffer. Fractionation The method comprises, after the labelling of derivatised polynucleotides (step 4), fractionating the polynucleotides into first and second fractions. The first fraction is 20 enriched for polynucleotides comprising an affinity label, and the second fraction is enriched for polynucleotides lacking an affinity label. The fractionation may comprise selectively isolating the labelled polynucleotides using the affinity label. For example, fractionation of the polynucleotides may comprise 25 binding of the affinity label to a capture probe that specifically binds to the affinity label. In embodiments in which the affinity label comprises biotin, fractionation of the polynucleotides may comprise selectively isolatingthe labelled polynucleotides using a 30 biotin-binding protein. The biotin-binding protein may comprise, for example, streptavidin, avidin, and / or a biotin-specific antibody. 35 The biotin-binding protein may comprise streptavidin, or a functional analogue or derivative of streptavidin. The biotin-binding protein may comprise a separation medium or substrate. For example, the biotin-binding protein may be conjugated to a surface. The surface may comprise a plurality of microbeads, such as paramagnetic microbeads. 5 In embodiments in which the separation medium or substrate comprises a plurality of microbeads coated with biotin-binding protein, the method may comprise binding the labelled polynucleotides to the biotin-binding protein on the coated microbeads and then isolating the coated microbeads. Isolation of the microbeads may be performed by 10 centrifugation. In embodiments comprising the use of paramagnetic microbeads, isolation of the microbeads may comprise the application of a magnetic field to the reaction mixture to separate the beads from the remainder of the suspension. The biotin-binding protein may comprise streptavidin conjugated to the surface of 15 microbeads. Preferably the streptavidin-coated microbeads may be streptavidin-coated paramagnetic microbeads. After an appropriate incubation to bind the labelled polynucleotides to the capture probe, the probe may be washed to remove unbound and non-specifically bound 20 polynucleotides. In embodiments in which the affinity label comprises biotin, the biotin-binding protein may be washed by’ any suitable method to remove unbound and non-specifically bound polynucleotides. After the selective isolation of the labelled polynucleotides, the polynucleotides are 25 separated from the capture probe. In previous methods, such as that described by Kriukiene et al. (Nature Communications 2013 4:2190), DNA fragments are released from a streptavidin capture agent using oxidative cleavage of a disulfide bond within the affinity label. 30 However, this method has been found by the present inventors to be inconsistently reproducible and to have poor efficiency. In the disclosed method, the polynucleotides are preferably not separated from the capture probe by a method comprising oxidative cleavage of the tag or affinity label. Thus, in some embodiments, the method does not comprise the separation of the polynucleotides from the capture probe by cleavage, such as oxidative cleavage or hydrolysis, of the tag or affinity label. 5 In some embodiments, the method may comprise the denaturation of the capture probe. For example, in embodiments in which the capture probe comprises a biotinbinding protein, the method may comprise the denaturation of the the biotin-binding protein. This method has been found to be particularly advantageous due to the consistent release of DNA fragments regardless of the number of the attached affinity 10 labels. The ability of streptavidin to bind to biotin is dependent on both a sterically defined binding pocket and the highly polar residues within it. Any agent that induces a conformational change of streptavidin may, therefore, be used to release the labelled 15 polynucleotides. The inventors have found that, in embodiments in which the biotinbinding protein comprises streptavidin, the labelled polynucleotides maybe released from the streptavidin by any method that denatures streptavidin without damaging the polynucleotides. 20 In some embodiments, the labelled polynucleotides may be released from the streptavidin by incubation in pure water at a temperature of about 7O°C. In some embodiments, the labelled polynucleotides maybe released from the streptavidin by incubation in 12-15% (v / v) phenol at room temperature. 25 In some embodiments, streptavidin may be denatured using a denaturing reagent, such as 1% sodium dodecyl sulphate and heating the sample to 9O°C. Advantageously, because the affinity label is not damaged by this method comprising 30 the denaturation of streptavidin, the first fraction may be further enriched for polynucleotides comprising an affinity label by repeating the selective isolation (affinity purification) step in one or more further cycles. Seauencina Following fractionation, the polynucleotides of the first and / or second fraction may be amplified. Amplification may be used to simultaneously introduce indexing barcodes to the polynucleotides. 5 Sequencing the polynucleotides of the first and / or second fraction may comprise pooling the first and second fractions and sequencing them together. In such embodiments, the first and second fraction may be distinguished using indexing barcodes. 10 The inventors have surprisingly found that standard DNA polymerases are able to amplify densely modified DNA comprising the disclosed affinity labels, prior to sequencing. Moreover, this has advantageously been found to be possible under the conditions employed for DNA release. This negates the need for DNA purification prior 15 to amplification. The polynucleotides may be amplified by PCR or qPCR, for example, using primers designed to anneal within the ligated sequencing adapters, and an appropriate PCR program. 20 The amplification method may be arranged to introduce indexing barcodes into the polynucleotides. For example, in some embodiments, the method may comprise amplification before fractionation to introduce indexing barcodes into the polynucleotides. In some embodiments, the method may comprise amplification after 25 fractionation to introduce indexing barcodes into the polynucleotides of the first and / or second fraction. The amplified polynucleotides may be purified prior to sequencing by any suitable method for cleaning up PCR products for use in a sequencing platform. 30 The amplified polynucleotides may be used directly for sequencing without further purification. The term “sequencing” as used herein, unless otherwise indicated, refers to any method 35 that may be used to determine the sequence (i.e. the order of nucleotides) in a nucleic acid such as DNA or RNA. Any type of sequencing platform may be used to determine the sequences of the polynucleotides, in combination with the appropriately ligated sequencing adapter. Thus, sequencing approaches that may be suitable for use in the disclosed method 5 include, but are not limited to, Sanger sequencing, high-throughput sequencing, pyrosequencing, sequencing-by-synthesis, single-molecule sequencing, nanoporebased sequencing, semiconductor sequencing, sequencing-by-ligation, sequencing-by-hybridization, RNA-Seq (Illumina), Digital Gene Expression (Helicos), next generation sequencing (NGS), Single Molecule Sequencing by Synthesis (SMSS) (Helicos), 10 massh ely-parallel sequencing, Clonal Single Molecule Array (Solexa), shotgun sequencing, Ion Torrent, Oxford Nanopore, Roche Genia, Maxim-Gilbert sequencing, primer walking, sequencing using Singular Genomics, Ultima Genomics, Element Biosciences, PacBio, SOLiD, Ion Torrent, or Nanopore platforms. Sequencing reactions can be performed in a variety of sample processing units, which may include multiple 15 lanes, multiple channels, multiple wells, or other means of processing multiple sample sets substantially simultaneously. The method may comprise a high throughput sequencing method. The terms “next generation sequencing”, “NGS”, and “high throughput sequencing” as used herein, refer 20 to sequencing technologies having increased throughput as compared to traditional Sanger- and capillary electrophoresis-based approaches. The high throughput sequencing method may be capable of generating hundreds of thousands of sequence reads in parallel. The method may comprise a multiplex sequencing technique. The high throughput sequencing methods that may be used include, but are not limited to, 25 sequencing by synthesis, sequencing by ligation, and sequencing by hybridization. The sequencing method may be capable of sequencing single molecules. In some embodiments, the method comprises sequencing the polynucleotides using the Illumina platform. 30 The method may comprise comparing the sequencing reads to a reference sequence. The method may comprise comparing the sequencing reads to a reference genome to determine the genomic location of the enriched polynucleotides. The reference sequence may be a human genome, or may comprise one or more portions thereof, 35 such as one or more chromosomes and / or chromosomal regions. Thus, the method may comprise aligning the sequencing reads to the reference sequence. Any suitable alignment method compatible with high throughput sequencing data may be used, for example, using the Burrows Wheeler Alignment algorithm. Aligned reads may be normalised. Any method of normalisation used in the art may 5 may be used. For example, normalising the sequencing reads may comprise reporting the number of reads in an aligned region (bin) as a fraction of reads per million reads of the sequencing output. The method may comprise comparing the normalised read counts from two or more 10 sequencing experiments. Such an approach may be advantageous in methods further comprising the step of diagnosing a disease based on the modification status of the nucleotide residues in the sample. Methods comprising making a determination based on the modification status of 15 nucleotide residues in a polynucleotide may comprise the production of a profile. For example, the profile may reflect the position of modified and / or unmodified residues within a polynucleotide such as a portion of a genome or an entire genome. Accordingly, methods comprising making a determination based on the CpG modification status of a plurality of CpG dinucleotides in a polynucleotide may 20 comprise the production of a profile, wherein the profile may reflect the position of modified and / or unmodified CpG dinucleotides within a polynucleotide such as a portion of a genome or an entire genome. Thus, in some methods, comparing the modification status of nucleotide residues in a 25 polynucleotide from the subject to the modification status of the corresponding residues in a reference sample may comprise comparing the profile obtained from the sample with the profile of a reference sample. For example, comparing the modification status of cytosine residues in CpG dinucleotides of a polynucleotide from the subject to the CpG modification status of the corresponding residues in a reference sample may 30 comprise comparing the profile obtained from the sample with the profile of a reference sample. The profile from the reference sample may comprise a profile that is representative of a healthy individual. In other embodiments, the profile from the reference sample may 35 comprise a profile that is obtained from, or indicative of a particular disease, such as, for example, a cancer. The method may comprise comparing the profile obtained from the sample with a database of profiles. The database of profiles may comprise a plurality of profiles relating to a single disease, wherein the disease may be diagnosed in the subject from 5 which the test sample was derived based on similarities between the profile of the test sample and the database of profiles. In other embodiments, the database of profiles may comprise a plurality of profiles representative of different diseases, wherein a disease may be diagnosed in the subject 10 from which the test sample was derived based on similarities between the profile of the test sample and one of more of the profiles within the the database. A comparison between profiles may be made using any suitable method. For example, a comparison may be made statistically, using an appropriate metric, for example, a p-15 value. A comparison may also be made using a machine-learning platform. In accordance with a second aspect there is provided a kit for determining the modification status of nucleotide residues in a polynucleotide sample, the kit comprising: 20 1. a methyltransferase enzyme configured to modify a nucleotide residue in a target position to apply a tag from a cofactor analogue to each unmodified nucleotide residue in a polynucleotide of the sample, and a cofactor analogue comprising the tag precursor, wherein each unmodified nucleotide residue is unmodified in the target position; 25 2. sequencing adaptors and enzymes for preparing the polynucleotide into a sequencing library; 3. an affinity label precursor suitable for binding an affinity label to the tags, wherein the affinity label comprises biotin; 4. a capture agent comprising a biotin-binding protein for fractionating the 30 sequencing library into first and second fractions, wherein the first fraction is enriched for polynucleotides comprising an affinity label, and wherein the second fraction is enriched for polynucleotides lacking an affinity label; and 5. a releasing agent for releasing the affinity label from the capture agent by denaturation of the biotin-binding protein. For use with polynucleotide samples comprising DNA, the kit may comprise: 1. a methyltransferase enzyme configured to modify a nucleotide residue in a target position to apply a tag from a cofactor analogue to each unmodified nucleotide residue in the DNA sample, and a cofactor analogue comprising the tag precursor, wherein each unmodified nucleotide residue is unmodified in the target position; 5 2. sequencing adaptors and enzymes for preparing the DNA into a sequencing library; 3. an affinity label precursor suitable for binding an affinity label comprising a biotin to the tags 4. a capture agent comprising a biotin-binding protein for fractionating the 10 sequencing library into first and second fractions, wherein the first fraction is enriched for DNA molecules comprising an affinity label, and wherein the second fraction is enriched for DNA molecules lacking an affinity label; and 5. a releasing agent for releasing the affinity label from the capture agent by denaturation of the biotin-binding protein. 15 The kit may be a kit for determining the modification status of cytosine residues in CpG dinucleotides of a DNA sample, the kit comprising: 1. a methyltransferase enzy me configured to modify the cytosine C5 position of a CpG dinucleotide to apply a tag from a cofactor analogue to each unmodified cytosine 20 residue in the DNA sample, and a cofactor analogue comprising the tag precursor, wherein each unmodified cytosine residue is the cytosine of a CpG dinucleotide that is unmodified in the C5 position; 2. sequencing adaptors and enzymes for preparing the DNA into a sequencing library; 25 3. an affinity label precursor suitable for binding an affinity label to the tags, wherein the affinity label comprises biotin ; 4. a capture agent comprising a biotin-binding protein forfractionating the sequencing library- into first and second fractions, wherein the first fraction is enriched for DNA molecules comprising an affinity label, and wherein the second fraction is 30 enriched for DNA molecules lacking an affinity label; and 5. a releasing agent for releasing the affinity label from the capture agent by denaturation of the biotin-binding protein The kit may be suitable for use in or as a one pot method as described in accordance 35 with the first aspect. The methyltransferase enzyme may be a methyltransferase as described in accordance with the first aspect. The cofactor analogue may be a cofactor analogue as described in accordance with the 5 first aspect. The cofactor analogue may be synthetic methyltransferase cofactor analogue. The cofactor analogue may comprise a compound of formula (I). The sequencing adaptors and enzymes for preparing the polynucleotide into a sequencing library may comprise enzymes and reagents for reverse transcription, end 10 repair, A-tailing, and / or adapter ligation, as described in accordance with the first aspect. The affinity label precursor may be an affinity label precursor as described in accordance with the first aspect. The affinity label precursor may comprise a compound 15 of formula (II). The capture agent comprising a biotin-binding protein for fractionating the sequencing library maybe as described in accordance with the first aspect. 20 The releasing agent for releasing the affinity label from the capture agent by denaturation of the biotin-binding protein. All features described herein (including any accompanying claims and drawings), and / or all of the steps of any method or process so disclosed, may be combined with 25 any of the above aspects in any combination, except combinations where at least some of such features and / or steps are mutually exclusive. The invention will now be illustrated by reference to specific Examples showing how embodiments maybe carried into effect, which are not intended to be limiting. Data 30 from the Examples is presented in the Figures, in which: Figure 1A is a flow chart showing the disclosed method. Figure 1B is a diagram providing an overview of an embodiment of the disclosed 35 method. Purified fragmented DNA (cfDNA or fragmented genomic DNA) is derivatised by treatment with a CpG-targeting methyltransferase (in this case M.Mpel) and a synthetic cofactor analogue (in this case ETA-AdoHcy-N3) that results in the addition of tags at unmodified CpG sites. The methyltransferase is inactivated and the DNA fragments are then end-repaired and ligated to sequencing adapters. Fragments comprising a tag are subsequently labelled by attaching an affinity label (for example, 5 biotin) to the tag, and isolated (for example, using streptavidin-coated magnetic beads). Tagged (unmodified CpG sites) and untagged (predominantly smCpG and shmCpG sites) can be separately amplified and sequenced. Figure 2A is a graph showing sequencing adapter ligation efficiencies using unlabelled 10 DNA (control; left-hand cluster), methyltransferase-labelled DNA without methyltransferase inactivation (centre cluster), and methyltransferase-labelled DNA with methyltransferase inactivation (right-hand cluster), prior to adapter ligation. Figure 2B shows the raw data for this experiment. 15 Figure 3A is a bar chart showing efficiencies for capture (light bars) and capture / release (dark bars) of target DNA from solution, as a function of target CpG site density. Figures 3B and 3C are reproduced from Kriukiene et al. (Nature Communications 20 2013 4:2190). Figure 3D shows the enrichment efficiency of the present method for a (target) DNA molecule with a high density of CpG sites (10 sites per -150 bp). The target DNA is mixed with 24 ng of non-target DNA and is selectively purified with high efficiency at a 25 range of concentrations. Final DNA concentration was quantified using spectrophotometry (Qubit). Figure 3E shows enrichment of unmethylated DNA as a function CpG density using three different enrichment chemistries. The current approach, using a single pot 30 reaction and one purification step (dark grey bars) shows over three times improvement in the retention of DNA throughout the Tag-Seq enrichment process, compared to the approach described by Kruikiene et al. (light grey) and an approach using two DNA purification steps (grey). Density of CpG sites is shown as number of sites per 300 bp genomic window. Figure 4 is a graph showing threshold eyrie versus target DNA concentration (ng) showing a linear response from 1.25 ng target DNA down to 1.25 pg target in a background of 24 ng (over 19000X excess) of DNA containing no target sites. 5 Figure 5 is a bar chart showing a comparison of sequencing coverage at CpG sites in the enriched (unmodified CpG) fraction (light grey) of a DNA sample, and the unenriched (modified CpG) fraction (dark grey) of the sample. Note that 1.4% and 44.7% of reads did not contain a CpG site for the enriched and unenriched fractions, respectively. 10 Figures 6A and 6B are bar charts showing sequencing coverage of enriched (light grey) and unenriched (dark grey) fractions at unmodified CpG sites (Figure 6A) and methylated CpG sites (Figure 6B). Note the different y-axis scales in the two plots. 15 Figure 7A is a bar chart showing enrichment using the disclosed method (NRPM) across a range of unmodified CpG densities (75 bp window) (groups 1-4, light grey bars) compared to similar enrichment using the known MeDIP-Seq method (group 5, hashed bars), with baseline for whole-genome sequencing provided for context (group 6, white bars). 20 Figure 7B is a bar chart corresponding to Figure 7A showing the corresponding enrichment profiles for methylated CpG sites. Figure 8A shows example profiles obtained using the disclosed method (dark blue) 25 compared to MeDIP-Seq (green) and WGBS (yellow) profiles for (top) the KRAS gene and (bottom) a megabase-scale region of chromosome 1. Figure 8B shows the emrichment of read counts for two technical replicates of the profile obtained using the disclosed method (blue) MeDIP-Seq (yellow) and shallow 30 whole genome sequencing (red) at gene transcription start sites (TSS). Figure 8C shows the profile obtained using the disclosed method (blue) correlates inversely with the MeDIP-Seq profile (purple) and with chromatin domain organisation identified in Hi-C experiments (red) on the megabase scale. Figure 9A shows enrichment of genomic DNA at regions corresponding to H3K4 monomethylation. Comparison of two experiments using the disclosed method (technical repeats, different users) (light and dark blue), MeDIP Seq (green) and shallow whole genome sequencing (control, no enrichment) (orange). 5 Figure 9B shows enrichment of genomic DNA at regions corresponding to H3K4 trimethylation. Comparison of two experiments using the disclosed method (technical repeats, different users) (light and dark blue), MeDIP Seq (green) and shallow whole genome sequencing (control, no enrichment) (orange). 10 Figure 9C shows enrichment of genomic DNA at regions corresponding to H3K27 acetylation. Comparison of two experiments using the disclosed method (technical repeats, different users) (light and dark blue), MeDIP Seq (green) and shallow whole genome sequencing (control, no enrichment) (orange). 15 Figure 10 shows a Spearmann correlation analysis to assess the similarity of the profiles obtained using the disclosed method across six different cell lines and for three technical repeats of each sample. Dark blue indicates a high degree of similarity7 of the profiles. Notably, DNA from each cell line has a clearly distinct profile when compared 20 to the sample technical repeats, consistent with the known utility of methylation profiles for the identification of tissues. Figure 11 shows volcano plots showing the comparison of tumour and normal adjacent tissue profiles obtained using the disclosed method, across the genome for six 25 patients with a range of different cancers. Differentially methylated windows are defined as those with an adjusted p-value of less than 5% and a log-fold change in signal of greater than 0.58 (i-5x). Red lines indicate the locations of these thresholds in the volcano plots. Blue markers are windows that show hypermethylation in cancer, red markers are for windows that are hypomethylated in cancer. 30 Figures 12A and 12B show DNA Agilent TapeStation (Cell-free DNA ScreenTape) traces showing profiles for the cell free DNA (cfDNA) that was input for the disclosed profiling method (top) and the output from the enriched, amplified libraries (bottom) of the disclosed method, for a healthy patient sample (Figure 12A) and for a sample 35 from a patient with Stage 1 non-small cell lung cancer (Figure 12B). The mono- and dinucleosomal pattern of the input cfDNA is maintained in the final libraries, the size of which corresponds to the duplicated original strand plus the Illumina P5 / P7 sequencing adaptors. Figure 13 shows example profiles using the disclosed method for a genomic region 5 (SHOX2 gene, a known methylation biomarker for lung cancer) in the healthy (blue) and lung cancer (grey) patients, compared to genomic DNA, extracted from the healthy patient’s buffy coat (yellow). Black traces show the duplicate profiles in both cases, dark blue tick marks show the known CpG site density across the gene. Traces are based on normalised read counts for all profiles and the cfDNA samples are displayed on the 10 same scale for direct comparison. Figure 14 shows a summary of data derived from triplicate repeat experiments using DNA isolated from FFPE (formalin-fixed, paraffin-embedded) samples. A) Enrichment (normalised read count) as a function of CpG site density (75 bp windows) showing 15 steadily increasing levels of enrichment with increasing CpG density. B) Plots showing normalised read counts across the APC gene transcription start site. Consistent with the enrichment profiles in (A), enrichment of DNA at the CpG-dense gene promoter is more marked for Patient A (pink) than Patient B (blue) than Patient C (green). Profile of the HT-29 cell line (colorectal cancer) shown in grey for comparison. C) PCA plot 20 showing excellent consistency of the technical replicates of these samples. Examples The disclosed method may be used to provide a uniquely straightforward and robust approach to epigenetic profiling that can be readily adapted for concurrent or 25 simultaneous readout of genetic features of the genome of interest. The method is an enzymatic technology that enriches for unmodified nucleotide residues such as, in particular, CpG sites, across the polynucleotide sample, such as across the whole genome. The method requires no de novo knowledge of the polynucleotide sequence for its application in epigenetic profiling. It does not damage the nucleic acid sample nor 30 rely on base conversion and therefore allows concurrent analysis of other genetic features, such as mutations. The method may also provide a profile whose signal correlates with markers of active genomic regions. The enzymatic chemistry enables unbiased fractionation of modified and unmodified nucleic acids from a sample for subsequent analysis. The Examples below show that the disclosed method is a uniquely sensitive approach, relative to other available methods for epigenetic analysis. The method can be applied at polynucleotide (e.g. DNA) concentrations that are compatible with single-cell analysis (picogram inputs). The workflow can be adapted to enable concurrent readout 5 of a genome’s genetic and epigenetic features. As a result of the simplicity of the approach, comprising a single enzymatic step, followed by fractionation using a capture probe, technical repeats of the experiment show excellent consistency. Enrichment of unmodified nucleic acids leads to an 10 epigenetic profile that achieves saturation at around 70M, 150 bp reads. Studies in cell lines and in patient tissue samples demonstrate the potential of the disclosed method as a platform for the diagnosis of disease and the identification of tissue of origin in a sample, for example. The underlying chemistry requires no a priori assumptions to be made about the sample, making the platform ideally suited as a research tool, for 15 example, for the discovery of novel biomarkers of disease. A flow chart of the disclosed method is shown in Figure 1A and an overview of an embodiment of the disclosed method is shown schematically in Figure 1B. 20 In the embodiment shown in Figure 1B, a bacterial DNA methyltransferase enzyme (M.Mpel) is used to target unmodified CpG sites for modification w ith an unnatural cofactor analogue of S-adenosyl-L-methionine referred to herein as ETA-AdoHcy-N3. Incubation of the methyltransferase, DNA and cofactor for one-hour results in 25 complete modification of the target DNA, which is functionalised w ith azideterminating tags. These tags can be further modified (for example, biotinylated) to enable fractionation of modified and unmodified DNA (where ‘unmodified’ refers to all genomic DNA fragments containing one or more CpG dinucleotide that is unmodified (for example, that is not methylated, hydroxymethylated, carboxylated, or acylated) at 30 the C5-position). The inventors have developed a modified library preparation that integrates this labelling step, thereby minimising handling and purification steps, and as a result, improving robustness and maximising sensitivity. Example 1 A significant advantage of the disclosed method is that it employs a single clean-up step for the entire process, dramatically improving the efficiency of the fractionation. This is made possible by the use of, firstly, a step to inactivate the methyltransferase and, 5 secondly, the surprising activity of the enzymes for library preparation in the resulting buffer mixtures. Figure 2 shows that the efficiency of adapter ligation is significantly inhibited in the absence of inactivation of the methyltransferaseafter labelling and before adapter io ligation. This surprising result is due to the high binding affinity of the methyltransferase enzyme to the DNA molecule, which has been found to limit the activity of DNA-targeting enzymes in subsequent steps of the procedure. This activity can be recovered byinactivating the methyltransferase enzyme. 15 Example 2 Figure 3 shows the results of example experiments investigating the recovery of DNA samples with different affinity-' labels and comprising different CpG densities. In the experiment of Figure 3A, a mixture containing 100 ng of DNA carrying a known 20 number of CpG sites (o, 1, 2, 4 or 10) was incubated with M.Mpel (0.0274 pg / pL) and ETA-AdoHCy-N3 (100 pM). The reaction was incubated at 37 °C for 1 hour. The DNA was purified using AMPure beads (Beckman Coulter), followed by conjugation of an affinity label comprising biotin using click chemsitry. Finally, DNA was purified using a standard PCR clean-up kit (Zymo Clean and Concentrate). 25 Purified DNA was fractionated using DynaBeads MyOne Streptavidin-coated beads. The beads were then washed twice with 150 pL of PBST. Finally, captured DNA was released. 30 As show n in Figure 3A, capture efficiency^ is improved by the current method (grey bars), relative to the method reported by Kriukiene et al. (Nature Communications 2013 4:2190), Figure 3C. 35 Figure 3A (blue bars) shows the release efficiency of DNA in the current workflow (Active-Seq). The disclosed method provides a significant improvement in capture efficiency relative to the method described in Kriukiene et al. (Nature Communications 2013 4:2190). As shown in Figure 3B (reproduced from Figure 2b of Kriukiene et al.), Kriukiene at el. 5 report capture efficiencies in the 20-30% range using a method comprising an azide-DBCO label. This is significantly lower than the capture efficiencies that may be obtained using the method disclosed herein, as shown, for example, in Figure 3A. The method described by Kriukiene et al. shows (in Figure 2c, reproduced herein as 10 Figure 3C) capture of around 30-40% of target DNA containing 2 CG sites using an azide-DBCO affinity label and streptavidin-coated magnetic beads. In contrast, while the method disclosed herein is able to isolate DNA at similar input levels to those of the method described by Kriukiene et al., the captured DNA may be recovered from the capture agent in much more significant proportions, at least in part due to the efficient 15 release of the labelledDNA molecules (see Figure 3D). Kriukiene et al. only includes data on the level of DNA capture, and there is no discussion or data in Kriukiene et al., on the efficiency of release of the sample from the magnetic beads. The present inventors have found that using the method disclosed by Kriukiene et al., the release of enriched DNA fragments is highly inefficient and inconsistently reproducible. 20 The efficient enrichment of DNA that is rich in CpG sites is critical for even representation of the (enriched) genome in the sequencing experiment. CpG-rich regions often lie in important regulatory regions of the genome. Figure 3E shows a plot of mean normalised read count per million reads (NRPM) for samples prepared using 25 the method described in Kriukiene et al. (left-hand bars) or the method disclosed herein (right-hand bars), as a function of the number of CpG sites in a given read. The plot is generated for CpG-rich regions of the genome (CG islands). Figure 3E clearly shows higher read densities across CpG rich regions, demonstrating 30 the significantly improved enrichment of CpG-rich DNA using the disclosed method. The overall effect of the method disclosed herein is to enable efficient enrichment of unmodified DNA from as little as a few picograms of input DNA. This is particularly critical for samples where the DNA concentration is limited, such as liquid biopsy 35 (blood, urine, saliva, spinal fluid) samples. Example 3 Experiments were performed to establish how the enrichment / fractionation platform performs as a function of DNA concentration, at input amounts consistent with cfDNA and single cell analyses; and the linearity of the enrichment efficiency across DNA 5 molecules with a range of (unmodified) CpG site densities. The inventors have found that this latter issue was a particular limitation of the method described by Kriukiene et al., which resorted to dilution of the methyltransferase enzyme in the labelling reaction to limit the number of (relatively insoluble) DNA modifications introduced to a single DNA molecule. 10 Between 1.25 ng and 1.25 pg (equivalent to between 200 and 1 / 3 of a copy of the human genome) of target DNA (153 bp, containing 10 unmodified CpG sites) were spiked into a background of 24 ng of non-target DNA (142 bp containing no CpG sites). 15 The target DNA was tagged and thereby enriched using streptavidin coated beads for analysis by qPCR, the results of which are shown in Figure 4. Enrichment efficiencies in excess of 80% were obtained for all of the spike-in samples, clearly demonstrating the compatibility of the approach for enrichment of DNA at input 20 levels consistent with single-cell analysis. Tagged DNA is compatible with PCR and can be amplified using a standard polymerase, following enrichment. As shown in Figure 3 A, the initial step of enrichment (capture of labelled DNA, for example, by streptavidin-coated beads) shows only a very minor dependence on the 25 number of CpG sites available on a DNA molecule. Light grey bars show capture efficiencies and dark grey bars showcapture / release of target DNA from solution. Example 4 Having demonstrated the performance of the biochemical approach on simple DNA 30 fragments, the utility of the platform disclosed herein on genomic DNA was investigated by generating genome-wide epigenetic profiles from DNA extracted from a range of cell lines. Extracted DNA was fragmented by sonication (-150 bp) and subject to enrichment of the DNA fragments lacking CpG modification. 35 Samples were sequenced using an Illumina NovaSeq platform (Source Biosciences) to approximately 120M reads per sample. For the enriched, unmodified DNA fraction, saturation analysis shows that data reaches 90% saturation between 65M and 90M reads. Initial quality control using MultiQC showed low levels of read duplicates (~15%) and an average enrichment in GC-content of the genome, from 41% in the nascent human genome to an average of ~47% in enriched samples, consistent with enrichment 5 at regions rich in CpG dinucleotides, such as CG islands (lung cancer derived cell lines showed higher average GC content, with a mean of approximately 54%). Example 5 Successful enrichment at CpG sites was assessed by comparison of the enriched 10 (unmodified CpG) and unenriched (modified CpG) fractions of the genome by sequencing. This was done by examining the fraction of reads containing a CpG site and the sequencing coverage at each CpG site. In the enriched fraction, 98.6% of the reads contain a CpG site. By contrast, in the unenriched fraction only 55.3% of reads contain a CpG site, indicating effective enrichment at CpG sites. Furthermore, in the enriched 15 fraction, a majority of the CpG sites of the genome (54.0%) are covered by greater than 5 reads, whereas in the unenriched fraction, this figure is just 6.3% of the CpG sites. In the human genome, 70-80% of the CpG sites are modified (e.g. methylated) and, hence, in the disclosed method the enriched fraction of the sample might be expected to 20 be focussed only on the remaining 20-30% of the genome’s CpG sites. Despite this, over half of the genomic CpG sites were found to have ‘high’ coverage (>5-fold) in the enriched (unmodified) fraction (light grey bars), as shown in Figure 5. This is likely due, at least in part, to the high efficiency of the disclosed method. 25 For enriched DNA, typically between 1 and 5% of reads were found to contain no CpG sites. The source of these reads is likely varied but will include non-specifically enriched DNA, as well as reads that do not cover a motif but that originate from a molecule that does (specifically enriched but CpG not sequenced). This read fraction is denoted as the ‘background’ for the enriched sample. 30 To further understand the composition of the enriched DNA fraction, sequencing coverage of CpG sites was compared at known modified and unmodified CpG sites (as determined by whole genome bisulfite sequencing (WGBS)). The results are shown in Figures 6A and 6B. As defined herein, an unmodified (e.g. ‘unmethylated’) site is a site 35 having a modification (e.g. methylation) level (P-value) of less than 0.05 by whole genome bisulfite sequencing. A ‘modified’ (e.g. ‘methylated’) site has a modification (e.g. methylation) level (0-value) of greater than 0.95 by whole genome bisulfite sequencing. A total of 430,245 CpG sites met the definition of an ‘unmodified CpG site’ (< 5 % 5 modified by WGBS). Significant enrichment was observed at these sites, as judged by the high coverage of sites in the enriched fraction (70% of sites (-301,000 CpG sites) have greater than 5-fold coverage) as compared to the unenriched fraction of the sample (less than 1% (-1500 sites) have greater than 5-fold coverage). This is in good agreement with the initial validation of the approach, confirming that where CpG sites 10 are unmodified, efficient enrichment is seen using the disclosed method. There are 53,081 CpG sites in the genome defined as ‘modified’ (> 95 % modified by WGBS). Similar coverage of these sites is observed in both enriched and unenriched fractions of the sample, indicating that little enrichment of DNA occurs at these highly 15 modified sites. Hence, where CpG sites are modified, little enrichment of these sites is seen using the disclosed method. Example 6 As further validation of the disclosed method, the inventors sought to understand the 20 enrichment of genomic DNA, as a function of unmodified (as determined by whole genome bisulfite sequencing) CpG site density. This analysis mirrors the validation experiments using DNA molecules of known sequence and CpG site density (as shown in Figure 4) but for enrichment using genomic DNA. The results are shown in Figures 7A and 7B. 25 A remarkably similar enrichment profile was observed for genomic DNA, as for the model DNA fragments with known CpG site densities with efficient enrichment of DNA, even where only one or two unmodified CpG sites are available. The enrichment profile for the disclosed method shows consistency across the range of unmodified site 30 densities (Figure 7A). This is in stark contrast to the analogous experiment for MeDIP-Seq for highly modified DNA molecules, Figure 7B. For MeDIP-Seq, no significant enrichment of DNA molecules was observed (relative to the WGS baseline) with a CpG density of less than 3 sites in a 75 bp genomic window; and a significant bias towards enrichment of densely modified regions of the genome (> 6 CpG sites per 75 bp). In all, these results are consistent with the initial validation of the disclosed approach, which demonstrates exceptionally high enrichment efficiencies for unmodified CpG sites; near uniform efficiency across a range of CpG densities and no significant off-target enrichment of DNA in the disclosed method. 5 Example 7 A number of studies were conducted to compare the disclosed method to other (epi)genomic analyses. 10 An advantage of the disclosed method is provided by the enzymatic targeting of unmodified CpG sites. The approach is well-suited to the enrichment of hypomethylated DNA from tumour cells in the blood. A key genomic feature that were hypothesised to be prominent in the profiles produced by the disclosed method are extended regions of unmodified DNA, which are epigenetically-stable and conserved in 15 mammals, with consistently low unmodified levels on length scales of 5-20 kbp. The term ‘non-modified island’ (NM1) is used herein since such unmodified regions are rather more island-like in the profiles produced by the disclosed method. Read count peaks in profiles of unmodified DNA produced by the disclosed method 20 (bottom) were found to anticorrelate to those observed in MeDIP-Seq (top) and correlate with regions of low modification, identified in whole genome bisulfite sequencing (yellow), as shown in Figure 8. At the gene-level, read counts of unmodified DNA using the disclosed method show 25 peaks centred at CG islands, that span the broader, regulatory regions of genes and are consistent with NMIs, as shown in Figure 8A. NMIs are thought to exist to reduce mutation rates in functionally-important genomic regions. Deamination of methylated cytosine, which converts to thymine, has been shown to be 30 the most frequent mutation in human cancers. NMIs play a central role in regulation of gene expression and their methylation levels are regulated by the TET (demethylating) enzymes, via the polycomb protein complex. In agreement with this, genome-wide analysis shows significant enrichment of unmodified DNA around transcription start sites, relative to the analogous MeDIP-Seq experiment, as shown in Figure 8A. On the scale of hundreds of kbp-to-Mbp, anticorrelation of the profile of unmodified DNA produced by the disclosed method (bottom) to both MeDIP-Seq (top) and WGBS (middle) is retained, as shown in Figure 8A. A particularly striking aspect of the profile produced by the disclosed method (bottom) is the presence of clear domains of 5 modified and unmodified genomic regions, that are consistent with the expected correlation between genomic modification (e.g. methylation) levels and genome organisation (Figure 8A). Example 8 10 For further validation, the disclosed method was compared to established sequencing approaches that are known to correlate with DNA methylation levels, as show n in Figure 8. Clear regions of highly unmethylated DNA, also evident in the MeDIP-Seq and WGBS profiles (Figure 8A). Enrichment in the profile produced with the disclosed method at transcription start sites and markers of active chromatin (HsKqMet, 15 HsKqMes and H3K27ac) anticorrelates with loss of MeDIP-Seq signal at these regions (Figure 8B and Figure 9). These regions of low / high methylation are correlated w ith defined structural domains of the genome identified in Hi-C experiments, as shown in Figure 8C. 20 Example 9 Having demonstrated the potential of the disclosed method to generate meaningful genome-wide epigenetic profiles, the approach was applied to nine DNA samples derived from cultured cell lines for a range of cancers (breast (MCF7, HCC1937), 25 colorectal (HT29, SW48, C0I0201, RKO), liver (HepG2) and lung (SW1271, NCI- H 2170)). Genome-wide correlation analysis of this dataset shows excellent correlation of a series of three technical repeats for each of the samples. Each of the cell lines examined forms 30 a distinct cluster of correlated data, with cell lines from similar tissues broadly clustering together, consistent with the expectation that the epigenetic profile can be employed for the identification of tissue of origin for a sample. These distinct cell-line-specific profiles result in part from the robustness of the method and the remarkable consistency of the disclosed method across sequencing runs and operators. In order to establish the potential of the disclosed method as a method for the diagnosis of cancer, a series of experiments were performed using tumour tissue and normal adjacent tissue from six patients with different cancers. Here, whole genome correlation analysis provides a simple overview that clearly highlights the 5 discriminative ability of the disclosed method for both disease diagnosis and derivation of tissue of origin from a patient sample (Figure 10). This approach was extended to better understand the specific regions of the genome that give rise to differences in epigenetic profiles and their link to known biological 10 function. The profiles produced by the disclosed method of tumour and normal adjacent tissue were compared for a patient in 75 bp windows, across the whole genome. This analysis requires no a priori knowledge of a patient’s genomic sequence and makes no assumptions about regions of interest. By doing so, tens or hundreds of thousands of differentially modified (e.g. methylated) regions were identified in the 15 profiles produced by the disclosed method for each patient. These are summarised in a series of volcano plots, shown in Figure 12. The results show the relative statistical significance and, critically, the population of modified (e.g. hypo- and hypermethylated) regions identified in each patient (the disclosed method does not solely focus on discovery of unmodified regions of the genome). Examples of individual 20 profiles produced using the disclosed method for the tissue lung cancer sample are shown at genes that have been implicated in relevant cancer pathways in the literature. Example 10 DNA shed from tumour cells can be isolated in the blood of cancer patients. However, 25 the technical challenge associated with its analysis is two-fold; DNA it is typically present in healthy and early-stage cancer patients at less than to ng per millilitre of plasma; and cell free DNA isolated from plasma can contain less than 1% tumour fraction (ctDNA). 30 The disclosed method is ideally suited to the analysis of ctDNA because it is performant with input DNA orders of magnitude less than one nanogram and provides genomewide analysis. To demonstrate the suitability of the disclosed method for the analysis of ctDNA, two 35 patient samples were prepared for analysis, one a healthy patient and one from a patient diagnosed with stage 1 lung cancer (non-small cell lung cancer). DNA was - 62 - extracted from 3mL of plasma using an automated platform (Informed Genomics), which returned 47UL of DNA at a concentration of 0.50 ng / pL and 0.55 ng / pL for the healthy and cancer patient, respectively (Figure 12A and 12B). 5 The method was performed in duplicate (separate preparations and sequencing runs) with 8.9 ng and 7.2 ng input DNA for the healthy patient and 10 ng input on both occasions for the lung cancer patient. Both input DNA samples and the output of the enriched library maintain the fragment size distribution that is characteristic of nucleosomal cell-free DNA, Figure 12. 10 Duplication rates for the sequencing data were 9.2% and 9% for the larger sequencing run, with coverage of 100M reads for both samples (Figure 13). This represents a significant improvement on duplication rates typical for approaches using base conversion, which can reach 30-40%. The background, defined as the percentage of 15 reads lacking a CpG site in the dataset, was 4.5% and 3.6% for the healthy and cancer samples, respectively. Example 11 Formalin fixed, paraffin-embedded (FFPE) treatment typically leads to extensive 20 damage (depurination, depyrimidation and deamination) of the genome. The ability to generate a meaningful epigenomic profile from DNA preserved in these samples using the disclosed method was investigated. Genome-wide profiles were generated using the disclosed method for three FFPE 25 embedded samples in triplicate, sourced from the Welsh Cancer Bank, derived from patients with colorectal cancer. Consistent with other sample types, sequencing reached 90% saturation by 80M (150 bp, paired-end) reads for all samples. The resultant datasets show good overall coverage of the genome and excellent consistency for the technical repeats (Figure 14). 30 For two of the three samples, high levels of relative enrichment of CpG dense regions of the genome were observed, (Figure 14). However, comparison of the three FFPE datasets to similar data for the HT-29 cell line shows good consistency of the profile and the observed ‘background’ of the sequencing dataset (reads lacking CpG sites) is 35 consistently below 5% for all samples. Such an increase in the relative enrichment of regions that are dense in CpG sites is consistent with the expected damage of CpG sites by the FFPE treatment. This likely stems from a relative reduction of the concentration of enrichable DNA molecules with few CpG sites in the sample. Conversely, those molecules with many CpG sites retain a few taggable CpG sites, post FFPE treatment. 5 Materials and Methods A typical workflow is set out below. In DNA samples requiring DNA fragmentation, DNA was sheared to an average of 180 bp. 10 A mixture of DNA (<tong), M.Mpel and ETA-AdoHCy-N3 was prepared on ice. This solution was incubated at 37 °C for 1 hour. Following incubation, the methyltransferase enzyme was inactivated by heating. 15 Without purification, the sample was cooled to to°C. End Repair &A-Tailing Master Mix was added (Kapa Biosystems). The mixture was mixed thoroughly by pipette aspiration and incubated at 2O°C for 30 mins followed by a 6sQC incubation for a further 30 mins. 20 The sample was cooled to to°C, and sequencing adapters were ligated. Without purification, biotin-PEGq-DBCO (Jena Biosciences) was added and the mixture was incubated at 37 °C for 1 hour with shaking at 500 rpm. The DNA was subsequently purified from the reaction mixture. 25 5 pL Dynabeads MyOne Streptavidin Cl beads (ThermoFisher) were washed with 150 pL of PBST. The DNA was added to the beads and the mixture was further incubated at 23 °C for 15 minutes, with shaking at 1000 rpm. Once completed, the supernatant was removed (as the “second fraction”), and the beads were washed twice with 150 pL of 30 PBST. Finally, the bound DNA was released from the beads (as the “first fraction”) by denaturation of streptavidin. Amplified libraries were pooled together with 0.1% PhiX and sequenced on a S4 Flow Cell using an Illumina NovaSeq Sequencer (Source Biosciences). qPCR was performed on the Azure Cielo 6 thermocycler (Azure Biosystems) with the following conditions: initial denaturation at 98°C for 30 seconds, then 40 cycles at 95°C for 10 seconds and 6o°C for 60 seconds with fluorescence detection. Analysis of the acquired fluorescence intensity and subsequent quantification of DNA in the samples 5 was performed using Azure Cielo Manager Analysis Software (V1.0.4). After sequencing, adaptors were removed from the reads using BBTools and then aligned to human reference genome HG38 using BWA-MEM2. Ambiguously aligned reads and those with low mapping scores (MAPQ score <40) were removed using 10 SamTools. Duplicates were removed with Sambamba (PM1D: 25697820) and reads hard-clipped using jvarkit (https: / / github.com / lindenb / jvarkit). Spearman correlation plots were generated from the processed bam files with deepTools using a binsize of tooobp and RPGC normalisation. Saturation figures and CpG density plots were generated for Chri-22 using the QSEA and Repitools R packages. To allow direct 15 comparison of enriched and unenriched samples, Bam files were down sampled to the same sequencing depth using SamTools. High confidence methylated and unmethylated CpG sites used for comparison were taken from the consensus of two whole genome shotgun bisulphite sequencing (WGBS) datasets performed on cell line NA12878 by the same lab (www.encodeproject.org / experiments / ENCSR89oUQO / ). 20 where less than 5% methylation in both datasets was considered to be unmethylated and greater than 95% methylation in both datasets was considered to be methylated.
Claims
1. A method of determining the modification status of nucleotide residues in a polynucleotide sample, the method comprising the steps of:5 (i) using a methyltransferase enzyme configured to modify a nucleotide residue ina target position to apply a tag to each unmodified nucleotide residue in a polynucleotide of the sample, wherein each unmodified nucleotide residue is unmodified in the target position;(ii) inactivating the methyltransferase;io (iii) preparing the polynucleotide sample into a sequencing library;(iv) binding an affinity label to each tag;(v) fractionating the sequencing library into first and second fractions, wherein the first fraction is enriched for polynucleotides comprising an affinity label, and wherein the second fraction is enriched for polynucleotides lacking an affinity label; and15 (vi) sequencing the polynucleotides of the first and / or second fraction.
2. The method of claim 1, wherein the method is a method of preparing a profile of the modification status of nucleotide residues in a reference sequence, and wherein the method further comprises comparing the sequences of the sequencing reads of the first 20 and / or second fraction to the reference sequence to determine the location of the sequencing reads within the reference sequence and thereby the presence or otherwise of unmodified nucleotide residues at specific locations within the reference sequence.
3. The method of either of claim 1 or claim 2, wherein the method is a one pot 25 method.
4. The method of any of claims 1-3, wherein the polynucleotide sample comprises DNA, and wherein the target position of the methyltransferase enzyme consists of the cytosine C5 position, the cytosine N4 position, or the adenine N6 position.
305. The method of claim 4, wherein the method is a method of determining the modification status at the cytosine C5 position of each CpG dinucleotide of the DNA sample, the method comprising the steps of:(i) using a methyltransferase enzyme configured to modify the cytosine C5 position 35 of a CpG dinucleotide to apply a tag to each unmodified cytosine residue of the sample,wherein each unmodified cytosine residue is the cytosine of a CpG dinucleotide that is unmodified in the C5 position;(ii) inactivating the methyltransferase;(iii) preparing the DNA sample into a sequencing library;5 (iv) binding an affinity label to each tag;(v) fractionating the sequencing library into first and second fractions, wherein the first fraction is enriched for DNA molecules comprising an affinity label, and wherein the second fraction is enriched for DNA molecules lacking an affinity7 label; and(vi) sequencing the DNA of the first / and or second fraction.
106. The method of claim 5, wherein the method is a method of preparing a profile of the modification status at the cytosine C5 position of CpG dinucleotides in a reference sequence, wherein the method further comprises comparing the sequences of the sequencing reads of the first / and or second fraction to the reference sequence to15 determine the location of the sequencing reads within the reference sequence andthereby the presence or otherwise of unmodified cytosine residues in specific CpG dinucleotides within the reference sequence.
7. The method of any of claims 1-6, wherein the polynucleotide sample comprises 20 polynucleotides substantially or predominantly having a length between 10 and 500 bp.
8. The method of any of claims 1-7, wherein using a methyltransferase enzyme to apply a tag to each unmodified nucleotide residue comprises the use of a methyltransferase cofactor analogue.
259. The method of claim 8, wherein the methyltransferase enzyme is a C5 methyltransferase, and wherein the methyltransferase cofactor analogue is ETA-AdoHcy-N3.30 10. The method of claim 8 or claim 9, wherein the methyltransferase enzymeconsists of or comprises a variant or fragment of the wild type M.Mpel sequence, comprising at least 80% sequence identity to the wild type M.Mpel sequence haying the NCBI accession number BAC44284.
11. The method of any of claims 1-10, wherein preparing the polynucleotide into a sequencing library comprises end repair, A-tailing, and adapter ligation of the polynucleotides in the sample.5 12. The method of any of claims 1-11, wherein binding an affinity label to each tagcomprises adding an affinity label precursor directly into the sequencing library preparation mixture, without a washing step.
13. The method of claim 12, wherein the affinity label comprises biotin.1014. The method of claim 13, wherein fractionating the sequencing library into first and second fractions comprises fractionation using a capture agent comprising a biotinbinding protein.15 15. The method of any of claims 1-14, wherein the method is an in vitro method ofdiagnosing disease in a subject, the method further comprising diagnosing the disease based on the modification status of nucleotide residues in a polynucleotide sample from the subject.20 16. The method of any of claims 1-15, wherein the polynucleotide sample is a cfDNAsample.
17. A kit for determining the modification status of nucleotide residues in a polynucleotide sample, the kit comprising:25 (i) a methyltransferase enzyme configured to modify a nucleotide residue in atarget position to apply a tag from a cofactor analogue to each unmodified nucleotide residue in a polynucleotide of the sample, and a cofactor analogue comprising the tag precursor, wherein each unmodified nucleotide residue is unmodified in the target position;30 (iii) sequencing adaptors and enzymes for preparing the polynucleotide into a sequencing library;(iv) an affinity label precursor suitable for binding an affinity label to the tags, wherein the affinity label comprises a biotin analogue;(v) a capture agent comprising a biotin-binding protein for fractionating the35 sequencing library into first and second fractions, wherein the first fraction is enriched for polynucleotides comprising an affinity label, and wherein the second fraction isenriched for polynucleotides lacking an affinity label; and(vi) a releasing agent for releasing the affinity label from the capture agent bydenaturation of the biotin-binding protein.
Citation Information
Patent Citations
Analysis of methylation sites
EP2594651A1
Optical mapping of genomic DNA
US20130130255A1
Epigenetic profiling method
WO2021053346A1
METHODS OF DETECTING METHYLATED CpG
WO2021250677A1