Profiling method

The method addresses limitations in profiling epigenetic modifications by labeling nucleotides for simultaneous mutation and residue analysis, enabling efficient, cost-effective, and unbiased profiling of small DNA samples for diagnostic and personalized medicine applications.

WO2026022472A1PCT designated stage Publication Date: 2026-01-29TAGOMICS LTD

Patent Information

Application Number
PCT/GB2025/051633
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-23
Filing Date
2025-07-22
Publication Date
2026-01-29

AI Technical Summary

Technical Problem

Current methods for studying epigenetic modifications in polynucleotides, such as DNA methylation, are limited by DNA degradation in low quantity samples, low conversion efficiencies, and inability to simultaneously profile nucleotide residues and genetic mutations, particularly in small samples like circulating cell-free DNA, which hinders personalized medicine applications.

Method used

A profiling method that labels target nucleotides in a DNA sample for identification, allowing simultaneous profiling of genetic mutations and nucleotide residues, compatible with low DNA inputs, and integrates with high-throughput sequencing platforms, minimizing purification steps and enabling cost-effective, unbiased analysis.

Benefits of technology

The method provides a reproducible, non-destructive approach for profiling nucleotide residues and genetic mutations in small samples, suitable for diagnostic tools and personalized medicine, with increased throughput and reduced sequencing costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure GB2025051633_29012026_PF_FP_ABST
    Figure GB2025051633_29012026_PF_FP_ABST
Patent Text Reader

Abstract

A method for preparing a polynucleotide sample for sequencing is disclosed. The method involves preparing a fractionated amplified sequencing library in which each polynucleotide comprises a first indexing barcode at a first end but not a second end.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Profiling Method

[0002] Field

[0003] The present application relates to methods of determining the status of nucleotide residues and the presence of genetic mutations in polynucleotides.

[0004] Introduction

[0005] Epigenetic modifications of polynucleotides, such as methylation and hydroxymethylation of cytosine, play an important role in determining the activity of a gene or a much more extended region of the genome. For example, methylation of DNA is critical in embryogenesis, early development, and is known to change predictably in correlation with biological ageing of an organism. On the other hand, aberrant modification of DNA can be an important driver of tumourigenesis, and the broader dysregulation of genes is likely to play a key role in many diseases.

[0006] Despite the critical role of nucleotide, and in particular cytosine, modification in the regulation of gene expression, current methods for studying epigenetic modifications fundamentally limit the scope of current studies of the epigenome. Methods comprising bisulfite conversion of cytosine to uracil may be used for epigenetic analysis, but the treatment of DNA with bisulfite can lead to DNA degradation. This limits the application of bisulfite in samples where the DNA quantity is low, as is typical for circulating cell-free DNA (cfDNA) in blood. Bisulfite-free approaches have been demonstrated that employ pyridine borane base conversion or enzymatic deamination of unmethylated cytosine. However, such approaches can suffer from low conversion efficiencies, relative to bisulfite conversion, and are inherently focussed on the analysis of individual cytosine bases, necessitating (comparative) whole-genome sequencing for biomarker discovery. This invariably leads to reduction of the test to a panel of genomic loci for application in the clinic, which effectively limits the diagnosis to a population that is similar to that profiled during the biomarker discovery phase of the test development.

[0007] Enrichment-based approaches for epigenetic profiling allow cost-effective, whole genome profiling that is more suited to the challenges of delivering personalised medicine. Recent advances have enabled the application of both antibodies (“methyl - DNA Immunoprecipitation”, also referred to as “MeDIP-Seq”) and methyl-binding domain protein (referred to as “MBD-Seq”) for the analysis of cfDNA. These approaches have been used successfully for targeted analysis of methylated genomic regions.

[0008] Moreover, current methods of determining the status of nucleotide residues in a polynucleotide, such as epigenetic profiling, do not permit the simultaneous profiling of genetic mutations in small samples of the polynucleotide of interest. This is particularly challenging for those approaches employing base conversion, which render the unambiguous identification of mutations against a background of converted bases challenging. Knowing the rates of genetic mutations associated with epigenetically modified regions of the genome may be advantageous, for example, for the diagnosis of disease. For example, it is broadly understood that the genomes of many cancers are both genetically mutated and hypomethylated, relative to a healthy genome.

[0009] There is, therefore, a need for improved profiling methods, that are capable of providing a profile of both the status of nucleotide residues and genetic mutations, in small quantities of the polynucleotide of interest, such as may be obtained from peripheral blood samples, that are capable of application to large portions of the polynucleotide sample without bias arising from differences in sequence, such as CpG density, and that may be applied cost effectively at large scales for use as diagnostic tools and to inform personalised medicine approaches.

[0010] The inventors have developed a profiling technique which meets these requirements and overcomes various limitations of existing methods, including those discussed above.

[0011] The disclosed method is based on an approach in which target nucleotides, such as epigenetically modified nucleotides, or unmodified CpG dinucleotides, in a DNA sample are labelled for identification, and a regional or global profile of genetic mutations in the polynucleotide sample may also be obtained in parallel from the same sample.

[0012] The disclosed method maybe used, for example, to provide in parallel both a profile of the status of nucleotide residues and a genetic mutation profile in a target region of the genome or across the whole genome. The disclosed technique is an advantageously simple process that can be integrated readily with standard high throughput sequencing platforms, and may be used with lower quantities of input DNA than has previously been possible, to generate genome-wide nucleotide status and genetic mutation profiles.

[0013] The disclosed method combines a procedure of DNA library preparation for next generation sequencing and a method for labelling nucleotides based on their specific modification status. The method advantageously minimises the number of DNA purification steps required and is highly efficient. As a result, the disclosed approach may be used with DNA input concentrations that are compatible with single-cell analysis (picogram inputs). The approach has been found to be highly reproducible and unbiased, and may be used as a platform for the diagnosis of disease and the identification of tissue of origin in a sample. The underlying chemistry requires no a priori assumptions to be made about the sample, making the platform ideally suited for the discovery of novel biomarkers of disease.

[0014] The inventors have found that the disclosed method provides significant advantages over previous methods of determining the status of nucleotide residues and epigenetic analysis. The non-destructive nature of the disclosed method means approach maybe used in parallel with other analytical approaches, such as nucleotide sequencing.

[0015] In the disclosed method, polynucleotides in a sample are labelled site-specifically on the basis of a particular nucleotide status, which may be, for example, a specific modification (such as a cytosine that is methylated in the C5 position), or absence of a specific modification (such as a cytosine that is unmethylated in the C5 position), before or after sequencing library preparation and amplification to incorporate an indexing barcode at a first end, but not a second end, of the polynucleotides in the sample. The sample is subsequently processed to incorporate, at the second end, a second or third indexing barcode based on the presence or absence of the label. In subsequent processing of the sequencing data, the combination of indexing barcodes maybe used to identify and distinguish polynucleotides comprising residues having the particular nucleotide status from polynucleotides that are reflective of the sample as a whole. The inventors have thus developed a unique method for indexing the sequencing library, which allows the labelled and unlabelled polynucleotides to be pooled and sequenced in a single sequencing run, thereby providing information for genetic analysis and nucleotide modification status from a single sequencing operation, thus increasing throughput and minimising sequencing costs. Central to the method is the production of a subset of polynucleotides, within the wider processed sample, that comprise an indexing barcode incorporated at a first end of the polynucleotide but not a second end, and a label bound site-specifically to a nucleotide residue of the polynucleotide. Thus, in a first aspect, there is provided a polynucleotide comprising: an indexing barcode incorporated at a first end of the polynucleotide but not a second end; and a label bound site-specifically to a nucleotide residue of the polynucleotide. Thus, the polynucleotide comprises: an indexing barcode incorporated at a first end of the polynucleotide only and not at a second end; and a label bound site-specifically to a nucleotide residue of the polynucleotide. The polynucleotide comprises: a first end that comprises an incorporated indexing barcode and a second end that does not comprise an incorporated indexing barcode; and a label bound site-specifically to a nucleotide residue of the polynucleotide. The polynucleotide comprises: a single indexing barcode; and a label bound site-specifically to a nucleotide residue of the polynucleotide.

[0016] In some embodiments, the polynucleotide is double-stranded. In some embodiments, the polynucleotide is double-stranded DNA. In some embodiments, the polynucleotide does not comprise double-stranded RNA.

[0017] In some embodiments, the polynucleotide is formed by the incorporation of an indexing barcode at a first end but not a second end of a polynucleotide comprising a label that is bound site-specifically to a nucleotide residue. In some embodiments, the polynucleotide may have a length of 100-500 nucleotides, such as 125-450 nucleotides, or 150-400 nucleotides.

[0018] In some embodiments, the polynucleotide may have a length corresponding to the DNA sequencing read length. In such embodiments, the polynucleotide may have a length of 100-250 nucleotides, preferably 150-180 nucleotides.

[0019] As used herein, unless otherwise stated, the terms “label”, “labelled”, “labelling” and similar terms refer to the targeted binding (covalent or otherwise, either directly or indirectly) of a compound that facilitates selective enrichment of the targeted region of the polynucleotide.

[0020] In some embodiments, the label is not a nucleotide. In some embodiments, the label is not bound to the nucleotide residue of the polynucleotide by hydrogen-bonding between complementary nucleotide bases. Thus, the label is not bound to the nucleotide residue of the polynucleotide by base-pairing.

[0021] In some embodiments, the label is not, or does not comprise, a fluorophore, a quantum dot, a dendrimer, a nanowire, a bead, a radiolabel, or an electromagnetic label.

[0022] In some embodiments, the label does not emit a signal. In some embodiments, the label does not alter a signal delivered to the label. In some embodiments, the label is not a fluorescent label. In some embodiments, the tag is not a fluorescent tag. In some embodiments, neither the tag nor the label is fluorescent. Thus, in some embodiments, the label and / or tag does not consist of or comprise a fluorophore or fluorophore derivative that is capable of emitting light when excited, such as re-emitting light upon light excitation.

[0023] The label may be suitable for use for enriching the polynucleotides from a mixture comprising the polynucleotide of the first aspect and an unlabelled polynucleotide.

[0024] Thus, in some embodiments, the label may comprise a tag that is bound to a nucleotide residue that is unmodified in a target position. In some embodiments, the label may comprise an antibody that specifically binds to a nucleotide residue comprising a target modification but does not significantly bind to the corresponding nucleotide that does not comprise the target modification. In some embodiments, the label may comprise a DNA binding protein such as a transcription factor or histone that has been crosslinked to a nucleotide residue.

[0025] In some embodiments, the label may be suitable for use directly for enriching the labelled polynucleotides. In other embodiments, a secondary compound that specifically binds the label, preferably with a high affinity, may be used for enrichment of the labelled polynucleotides.

[0026] In some embodiments, the label may comprise a tag from a cofactor analogue. The tag may be bound to a nucleotide residue that is unmodified in a target position.

[0027] In some embodiments, the label may comprise a high affinity binding protein or antibody that is bound site-specifically to a target modified nucleotide.

[0028] Thus, the label may comprise a tag from a cofactor analogue that is bound to a nucleotide residue that is unmodified in a target position, or a high affinity binding protein or antibody that is bound site-specifically to a target modified nucleotide.

[0029] As used herein, unless otherwise stated, the terms “site-specific”, “site-specifically”, and similar terms, refer to the application of a label to a target atom within a nucleotide residue having a particular configuration. In some embodiments, the particular configuration may be the presence of a specific modification, such as a cytosine that is methylated in the C5 position, and the label may bind to an atom within the modification. In some embodiments, the particular configuration may be the absence of a specific modification, such as a cytosine that is unmethylated in the C5 position, and the label may bind to an atom that would otherwise have been bound by the modification.

[0030] In some embodiments, the indexing barcode may comprise a unique dual nucleotide index, and / or a unique molecular identifier.

[0031] In some embodiments, the polynucleotide may comprise a sequencing adapter. For the avoidance of doubt, any of the statements herein describing or referring to embodiments may relate to any of the disclosed aspects as applicable.

[0032] In a second aspect, there is provided an amplified sequencing library comprising: - a first indexing barcode incorporated at a first end but not a second end of each polynucleotide in the amplified sequencing libraiy; a first subset of polynucleotides comprising a site-specifically bound label; and a second subset of polynucleotides that do not comprise a label. The first subset of polynucleotides of the sequencing library of the second aspect may comprise a polynucleotide of the first aspect.

[0033] In some embodiments, the amplified sequencing library may comprise first and second fractions, wherein the first fraction of the amplified sequencing library is enriched for polynucleotides comprising a site-specifically bound label, and wherein the second fraction is enriched for polynucleotides lacking a label.

[0034] In such embodiments, the polynucleotides of the first fraction may comprise a second indexing barcode incorporated at a second end but not a first end of each polynucleotide. The polynucleotides of the second fraction may comprise a third indexing barcode incorporated at a second end but not a first end of each polynucleotide.

[0035] In some embodiments, the amplified sequencing library may be a fractionated amplified sequencing library in which all polynucleotides comprise a first indexing barcode, the amplified sequencing library comprising: a first fraction of polynucleotides comprising a site-specifically bound label and a second indexing barcode; and a second fraction of polynucleotides comprising a third indexing barcode, that do not comprise a label.

[0036] In some embodiments, the amplified sequencing library may be a fractionated amplified sequencing library in which: a first fraction of polynucleotides comprises a site-specifically bound label and first and second indexing barcodes; and a second fraction of polynucleotides comprises first and third indexing barcodes, and do not comprise a label.

[0037] As used herein, unless otherwise stated, the terms “fractionating”, “fractionated”, “fractionation”, and similar terms, in relation to the sequencing library refer to the separation of the polynucleotides of the sequencing library into different groups or “fractions” on the basis of the presence or absence of one or more specific features. In particular, polynucleotides maybe fractionated based on the presence or absence of an associated affinity label, thereby forming a first fraction that is enriched for polynucleotides comprising an affinity label and a second fraction that is enriched for polynucleotides that do not comprise an affinity label. As such, the fractionation of the sequencing library may also be described and / or referred to as “enriching” and / or “enrichment” of the sequencing library for labelled polynucleotides or unlabelled polynucleotides, as appropriate.

[0038] The terms “enriched”, “enriching”, and “enrichment” of polynucleotides as used herein, unless otherwise stated, refer to a polynucleotide concentration and / or proportion that is greater than the corresponding polynucleotide concentration and / or proportion in the initial (unenriched) sample. References to “enriching the amplified sequencing library” and similar terms refer to fractionating the sequencing library / polynucleotide sample into first and second fractions, wherein the first fraction is enriched for polynucleotides comprising a label, and wherein the second fraction is enriched for polynucleotides lacking a label. Likewise, references to "enriching the amplified DNA library”, “enriching the labelled sequencing library” and similar terms refer to fractionating the sequencing library / polynucleotide / DNA sample into first and second fractions, wherein the first fraction is enriched for polynucleotides / DNA molecules comprising a label, and wherein the second fraction is enriched for polynucleotides / DNA molecules lacking a label. In a third aspect, there is provided a method for preparing a polynucleotide sample for sequencing, wherein the prepared polynucleotide sample is suitable for determining both the status of nucleotide residues and the presence of a genetic mutation in the polynucleotide sample, the method comprising:

[0039] (i) preparing a fractionated amplified sequencing library, wherein each polynucleotide of the fractionated amplified sequencing library comprises a first indexing barcode at a first end but not a second end, and wherein a first fraction of the fractionated amplified sequencing library is enriched for polynucleotides comprising a site-specifically bound label, and wherein the second fraction is enriched for polynucleotides lacking a label;

[0040] (ii) amplifying the first fraction and incorporating a second indexing barcode to a second end of each polynucleotide in the first fraction;

[0041] (iii) amplifying the second fraction and incorporating a third indexing barcode to a second end of each polynucleotide in the second fraction; and

[0042] (iv) pooling the amplified first and second fractions to obtain the prepared polynucleotide sample.

[0043] In some embodiments, step (i) of the method may comprise preparing an amplified sequencing library of the second aspect, and fractionating the amplified sequencing library to obtain a first fraction that is enriched for the first subset of polynucleotides and a second fraction that is enriched for the second subset of polynucleotides.

[0044] In some embodiments, determining the status of nucleotide residues in the polynucleotide sample may comprise determining the epigenetic modification status. In such embodiments, the site-specifically bound label may comprise:

[0045] (i) a tag from a cofactor analogue that is bound to a cytosine residue that is epigenetically unmodified in the C5 position;

[0046] (ii) a tag from a cofactor analogue that is bound to a cytosine residue that is epigenetically unmodified in the N4 position;

[0047] (iii) a tag from a cofactor analogue that is bound to an adenine residue that is epigenetically unmodified in the N6 position; (iv) a methyl-CpG binding domain protein;

[0048] (v) a high affinity antibody specific for 5-methylcytosine;

[0049] (vi) a high affinity antibody specific for 5-hydroxymethylcytosine;

[0050] (vii) a high affinity antibody specific for Nq-methylcytosine;

[0051] (viii) a high affinity antibody specific for N6-methyladenine; or (ix) a high affinity antibody specific for 5-methylcytosine.

[0052] Thus, in some embodiments, the site-specifically bound label may comprise a methyl - CpG binding domain protein or an antibody specific for a target modified nucleotide residue. Thus, in some embodiments, the site-specifically bound label may comprise:

[0053] (i) a methyl-CpG binding domain protein;

[0054] (ii) a high affinity antibody specific for 5-methylcytosine;

[0055] (iii) a high affinity antibody specific for 5-hydroxymethylcytosine; (iv) a high affinity antibody specific for N4-methylcytosine;

[0056] (v) a high affinity antibody specific for N6-methyladenine; or

[0057] (vi) a high affinity antibody specific for 5-methylcytosine.

[0058] In some embodiments, the site-specifically bound label may comprise: (i) a tag from a cofactor analogue that is bound to a cytosine residue that is epigenetically unmodified in the C5 position;

[0059] (ii) a tag from a cofactor analogue that is bound to a cytosine residue that is epigenetically unmodified in the N4 position; or

[0060] (iii) a tag from a cofactor analogue that is bound to an adenine residue that is epigenetically unmodified in the N6 position.

[0061] In some embodiments, the site-specifically bound label may comprise a tag from a cofactor analogue that is bound to a cytosine residue that is epigenetically unmodified in the C5 position.

[0062] As used herein, unless otherwise stated, the terms “genetic mutation” and “genetic sequence mutation” refer to any change in the nucleotide sequence relative to a control, native, or wild-type sequence. The control sequence may be, for example, a healthy genome or polynucleotide obtained from the same individual as the test sample. For example, in some embodiments, the control sequence may be obtained from a control sample from the same individual as the test sample, wherein the control and test samples are obtained from a tissue biopsy. In such embodiments, the test sample may be obtained from diseased tissue within the biopsy, and the control sample may be obtained from tissue adjacent to the diseased tissue.

[0063] As used herein, unless otherwise stated, the terms “method of’ and “method for”, such as “method of determining” and “method for determining” are intended to be interpreted interchangeably, to encompass methods “suitable for” the described purpose. The “status of nucleotide residues”, “modification status of nucleotide residues”, and “modification status” as used interchangeably herein, unless otherwise stated, refer to the presence (modified) or absence (unmodified) of any modification in a target position on a nucleotide residue. Such modifications may include, for example, the modification of cytosine (at the C5 or N4 position), or adenine (at the N6 position).

[0064] In some embodiments, the modification status of nucleotide residues may be the epigenetic modification status. In some embodiments, the nucleotide residues may be cytosine residues and / or adenine residues. Accordingly, the method maybe a method for determining both the modification status at target positions of cytosine and / or adenine residues and the presence of a genetic mutation in the polynucleotide sample. The modification may be a chemical modification in a target position on a nucleotide residue that maybe catalysed by a methyltransferase enzyme. Thus, the disclosed method may be used to determine the modification status of any position within a nucleotide that maybe chemically modified by a methyltransferase enzyme. Such methyltransferase catalysed modifications include, for example, the modification of cytosine (at the C5 or N4 position), adenine (at the N6 position).

[0065] The “modification status of cytosine residues” as used herein, unless otherwise stated, refers to the presence (modified) or absence (unmodified) of any methyltransferase catalysed chemical modification of cytosine.

[0066] In particular, the modification may comprise modification at the C5 position of cytosine. A cytosine residue may be understood to have the following structure:

[0067] In an unmodified cytosine residue R1maybe understood to be H. R1may also be referred to herein as the “C5 position”.

[0068] In a modified cytosine residue R1maybe anything other than H. Thus, “modified cytosine” refers to any cytosine residue that has been modified in any way at the C5 position, including, in particular, methylation (5-methylcytosine (5-mC)), and its oxidized products 5-hydroxymethylcytosine (5-hmC), 5-formylcytosine (5-fC) and 5- carboxylcytosine (5-caC). It maybe therefore appreciated that in modified cytosine, R1maybe methyl, CH20H, COH or COOH.

[0069] Alternatively, the modification may comprise modification at the N4 position of cytosine. Accordingly, a modified cytosine may have the following structure: where R1maybe anything other than H. Thus, “modified cytosine” refers to any cytosine residue that has been modified in anyway at the N4 position. It maybe therefore appreciated that in modified cytosine, R1maybe methyl, CH20H, COH or COOH. In some embodiments, the method may comprise the detection of modified cytosine residues in a polynucleotide.

[0070] Thus, in some embodiments, the method may be a method for preparing a polynucleotide sample for sequencing, wherein the prepared polynucleotide sample is suitable for determining both the modification status of cytosine residues and the presence of a genetic mutation in the polynucleotide sample, the method comprising:

[0071] (i) preparing a fractionated amplified sequencing library, wherein each polynucleotide of the fractionated amplified sequencing library comprises a first indexing barcode at a first end but not a second end, and wherein a first fraction of the fractionated amplified sequencing library is enriched for polynucleotides comprising a label bound site-specifically to modified cytosine, and wherein the second fraction is enriched for polynucleotides lacking a label;

[0072] (ii) amplifying the first fraction and incorporating a second indexing barcode to a second end of each polynucleotide in the first fraction;

[0073] (iii) amplifying the second fraction and incorporating a third indexing barcode to a second end of each polynucleotide in the second fraction; and (iv) pooling the amplified first and second fractions to obtain the prepared polynucleotide sample.

[0074] In some embodiments, the method may comprise the detection of unmodified cytosine residues in a polynucleotide.

[0075] Thus, in some embodiments, the method may be a method for preparing a polynucleotide sample for sequencing, wherein the prepared polynucleotide sample is suitable for determining both the modification status of cytosine residues and the presence of a genetic mutation in the polynucleotide sample, the method comprising:

[0076] (i) preparing a fractionated amplified sequencing library, wherein each polynucleotide of the fractionated amplified sequencing library comprises a first indexing barcode at a first end but not a second end, and wherein a first fraction of the fractionated amplified sequencing library is enriched for polynucleotides comprising a label bound site-specifically to unmodified cytosine, and wherein the second fraction is enriched for polynucleotides lacking a label;

[0077] (ii) amplifying the first fraction and incorporating a second indexing barcode to a second end of each polynucleotide in the first fraction; (iii) amplifying the second fraction and incorporating a third indexing barcode to a second end of each polynucleotide in the second fraction; and

[0078] (iv) pooling the amplified first and second fractions to obtain the prepared polynucleotide sample. The “modification status of adenine residues” as used herein, unless otherwise stated, refers to the presence (modified) or absence (unmodified) of any chemical modification at the N6 position of adenine. An adenine residue maybe understood to have the following structure: In an unmodified adenine residue R1may be understood to be H. In a modified adenine residue R1maybe anything other than H. Thus, “modified adenine” refers to any adenine residue that has been modified in anyway at the N6 position, including, in particular,6-methyladenine (m6A). It maybe therefore appreciated that in modified adenine, R1maybe methyl, CH20H, COH or COOH.

[0079] In some embodiments, the method may comprise the detection of modified adenine residues in a polynucleotide.

[0080] Thus, in some embodiments, in some embodiments, the method may be a method for preparing a polynucleotide sample for sequencing, wherein the prepared polynucleotide sample is suitable for determining both the modification status of adenine residues and the presence of a genetic mutation in the polynucleotide sample, the method comprising:

[0081] (i) preparing a fractionated amplified sequencing library, wherein each polynucleotide of the fractionated amplified sequencing library comprises a first indexing barcode at a first end but not a second end, and wherein a first fraction of the fractionated amplified sequencing library is enriched for polynucleotides comprising a label bound site-specifically to modified adenine, and wherein the second fraction is enriched for polynucleotides lacking a label;

[0082] (ii) amplifying the first fraction and incorporating a second indexing barcode to a second end of each polynucleotide in the first fraction;

[0083] (iii) amplifying the second fraction and incorporating a third indexing barcode to a second end of each polynucleotide in the second fraction; and (iv) pooling the amplified first and second fractions to obtain the prepared polynucleotide sample.

[0084] As used herein, “unmodified” and “unmethylated” refer to all nucleotides that are unmodified in any way (such as methylated, hydroxymethylated, carboxylated, acylated).

[0085] In some embodiments, the method comprises the site-specific labelling and amplification of the polynucleotides prior to fractionation of the sequencing library. In embodiments in which the method comprises the site-specific labelling and amplification of the polynucleotides prior to fractionation of the sequencing library, a pool of polynucleotides is created comprising a small proportion of polynucleotides that comprise the label, and a significant excess of unlabelled polynucleotides reflecting an unmodified copy of the sample as a whole. Despite the significantly disproportionate sizes of the labelled and unlabelled polynucleotide populations in the sequencing library after amplification, the inventors have surprisingly found that the polynucleotides may be efficiently fractionated using the label to provide a first fraction of labelled polynucleotides and a second, much larger, fraction of unlabelled polynucleotides. The modification status of nucleotide residues in the original polynucleotide sample may be determined by sequencing the labelled polynucleotides, and in parallel, the unlabelled polynucleotides are representative of the sample (for example a whole genome) and, when sequenced, may be used for genetic analysis, for example, of the whole genome, or specific genomic regions of interest by targeted enrichment using a panel of bait oligonucleotides. Thus, in some embodiments, step (i) may comprise:

[0086] (a) preparing an amplified sequencing library, wherein each polynucleotide of the amplified sequencing library comprises a first indexing barcode at a first end but not a second end, and wherein the amplified sequencing library further comprises a label bound site- specifically to a first subset of the polynucleotides but not a second subset; and

[0087] (b) fractionating the amplified sequencing library into first and second fractions, wherein the first fraction is enriched for polynucleotides comprising a label, and wherein the second fraction is enriched for polynucleotides lacking a label. Thus, in some embodiments, the method may comprise:

[0088] (i)(a) preparing an amplified sequencing library, wherein each polynucleotide of the amplified sequencing library comprises a first indexing barcode at a first end but not a second end, and wherein the amplified sequencing library further comprises a label bound site- specifically to a first subset of the polynucleotides but not a second subset;

[0089] (i)(b) fractionating the amplified sequencing library into first and second fractions, wherein the first fraction is enriched for polynucleotides comprising a label, and wherein the second fraction is enriched for polynucleotides lacking a label;

[0090] (ii) amplifying the first fraction and incorporating a second indexing barcode to a second end of each polynucleotide in the first fraction; (iii) amplifying the second fraction and incorporating a third indexing barcode to a second end of each polynucleotide in the second fraction; and

[0091] (iv) pooling the amplified first and second fractions to obtain the prepared polynucleotide sample.

[0092] In some embodiments, the method comprises the site-specific labelling of the polynucleotides prior to amplification and subsequent fractionation of the sequencing library. Thus, in some embodiments, step (i) part (a) may comprise: preparing the polynucleotides into a sequencing library and additionally:

[0093] (i) binding a label site-specifically to the polynucleotides to form a first subset of polynucleotides, that comprise the site-specifically bound label, and a second subset of polynucleotides, that do not comprise the site-specifically bound label; and (2) amplifying the polynucleotides and incorporating a first indexing barcode at a first end of the polynucleotides but not a second end.

[0094] Thus, in some embodiments, the method may comprise:

[0095] (i)(a) preparing the polynucleotides into a sequencing library and additionally: (1) binding a label site-specifically to the polynucleotides to form a first subset of polynucleotides, that comprise the site-specifically bound label, and a second subset of polynucleotides, that do not comprise the site-specifically bound label; and

[0096] (2) amplifying the polynucleotides and incorporating a first indexing barcode at a first end of the polynucleotides but not a second end; (i)(b) fractionating the amplified sequencing library into first and second fractions, wherein the first fraction is enriched for polynucleotides comprising a label, and wherein the second fraction is enriched for polynucleotides lacking a label;

[0097] (ii) amplifying the first fraction and incorporating a second indexing barcode to a second end of each polynucleotide in the first fraction; (iii) amplifying the second fraction and incorporating a third indexing barcode to a second end of each polynucleotide in the second fraction; and

[0098] (iv) pooling the amplified first and second fractions to obtain the prepared polynucleotide sample. In some embodiments, the method comprises the amplification of the polynucleotides prior to site-specific labelling and subsequent fractionation of the sequencing library. This approach creates a pool of ‘filler DNA’ which has been found to advantageously enable the efficient enrichment and analysis of extremely low quantities of input polynucleotide sample, such as, for example, as little as a few picograms of input DNA.

[0099] Thus, in some embodiments, step (i) part (a) may comprise: preparing the polynucleotides into a sequencing library and additionally:

[0100] (1) amplifying the polynucleotides and incorporating a first indexing barcode at a first end of the polynucleotides but not a second end; and (2) binding a label site-specifically to the polynucleotides to form a first subset of polynucleotides, that comprise the site-specifically bound label, and a second subset of polynucleotides, that do not comprise the site-specifically bound label.

[0101] Thus, in some embodiments, the method may comprise: (i)(a) preparing the polynucleotides into a sequencing library and additionally:

[0102] (1) amplifying the polynucleotides and incorporating a first indexing barcode at a first end of the polynucleotides but not a second end; and

[0103] (2) binding a label site-specifically to the polynucleotides to form a first subset of polynucleotides, that comprise the site-specifically bound label, and a second subset of polynucleotides, that do not comprise the site-specifically bound label;

[0104] (i)(b) fractionating the amplified sequencing library into first and second fractions, wherein the first fraction is enriched for polynucleotides comprising a label, and wherein the second fraction is enriched for polynucleotides lacking a label;

[0105] (ii) amplifying the first fraction and incorporating a second indexing barcode to a second end of each polynucleotide in the first fraction;

[0106] (iii) amplifying the second fraction and incorporating a third indexing barcode to a second end of each polynucleotide in the second fraction; and

[0107] (iv) pooling the amplified first and second fractions to obtain the prepared polynucleotide sample.

[0108] Library preparation

[0109] The polynucleotide sample maybe modified into a form that is compatible for high throughput sequencing. This process maybe referred to as “preparing a sequencing library” or “library preparation”. As used herein, unless otherwise indicated, a “sequencing library” refers to a plurality of polynucleotides, each comprising a sequencing adapter, such as a sequencing adapter arranged for use in next generation sequencing. Accordingly, “preparing a polynucleotide sample into a sequencing library”, as used herein, unless otherwise indicated, refers to the addition of one or more sequencing adapters to the polynucleotides of the sample. Thus, the process of preparing a polynucleotide sample into a sequencing library may comprise the ligation of one or more sequencing adapters to the polynucleotides. In addition to adapter ligation, the process of preparing a polynucleotide sample into a sequencing library may comprise end repair and / or A-tailing of the polynucleotides.

[0110] Preferably the process of preparing a sequencing library does not comprise combining the polynucleotides together to form an extended ligated polynucleotide for use, for example, in a sequencing method comprising nanopore technology.

[0111] In some embodiments, the library preparation process may comprise end repair of the polynucleotide sample. In some embodiments, the end repair process may comprise removal of 3' overhangs, for example using a Klenow fragment-based enzyme. The end repair process may also comprise modifying 3’ ends as necessary to comprise a hydroxyl group.

[0112] In some embodiments, the end repair process may additionally or alternatively fill 5' overhangs, for example, using a T4 DNA polymerase. The end repair process may also comprise phosphorylation of 5' ends where necessary, for example, using of a T4 polynucleotide kinase (PNK).

[0113] In some embodiments, the library preparation process may comprise A-tailing.

[0114] The A-tailing process may comprise the addition of an adenine residue to the 3' ends of the polynucleotide sample. This process may reduce the possibility of the polynucleotides in the sample ligating to each other. The A-tailing process may also increase the rate of adapter ligation, particularly in embodiments in which the adapters comprise a thymine overhang. In some embodiments, the A-tailing process may comprise the use of an “exo-Klenow” enzyme. In some embodiments, the method may comprise simultaneous end repair and A- tailing. For example, an end repair and A-tailing buffer comprising end repair and A- tailing enzymes may be used.

[0115] In some embodiments, the library preparation process may comprise one or more “adapter ligation” processes comprising the ligation of sequencing adapters to the polynucleotides in the sample. The term “adapter” as used herein, unless otherwise specified, refers to a short nucleic acid (such as less than about 500, less than too, or less than about 50 nucleotides in length) that is typically at least partially double-stranded and is attached to either or both ends of a nucleic acid molecule. The adapters may include a primer binding site for amplification of the sample. The adapters may include a primer binding site for sequencing applications, such as next-generation sequencing (NGS) applications. The adapters may include a binding site for capture probes, such as an oligonucleotide attached to a flow cell support. A plurality of adapters of the same or different sequences may be attached to the polynucleotides in the sample. In some embodiments, the ligated adapters may include a nucleic acid tag. The nucleic acid tag may be positioned relative to an amplification primer and / or sequencing primer binding site, such that the tag sequence is included in subsequent amplicons and sequence reads. In some embodiments, a plurality of adapters having the same sequence apart from different nucleic acid tags may be attached to the polynucleotides in the sample.

[0116] In some embodiments, the adapter ligation process may comprise the ligation of sequencing adapters to the polynucleotides in the sample. The adapter ligation process may comprise the use of a T4 DNA ligase. Advantageously, any sequencing adapters may be used. Sequencing adapters that have been found to be particularly suitable for use in the disclosed method include, for example, any sequencing adapters suitable for use with high throughout sequencing methods, such as sequencing applications on the Illumina platform. The method may comprise the use of a double-stranded indexing and unique dual indexing (UDI) adapter that enables efficient ligation and identification of PCR amplification replicates in the sequencing dataset. Other adapters may also be used, such as hairpin adapters. In some embodiments, the ligated adapters may include a barcode that can be introduced at one or both ends of the sample DNA molecule. For the avoidance of doubt however, the method comprises an amplification step which incorporates a first indexing barcode at a first end but not a second end of the polynucleotides, and this amplification step is additional to the process of sequencing library preparation, and the first indexing barcode applied in the amplification step is additional to any barcodes that may be applied to one or both ends of the polynucleotides as part of the library preparation process.

[0117] In some embodiments, the adapter ligation processes may comprise the ligation of adapters that do not comprise indexing barcodes, and such adapters may be referred to herein as “stubby adapters”. In some embodiments, the adapter ligation processes may comprise the ligation of Y adapters.

[0118] As used herein, unless otherwise stated, the term “Y adapter” refers to a polynucleotide sequence that maybe annealed to the 5' and / or 3' end of a polynucleotide in a sequencing library. When annealed to both the 5’ and 3’ ends of the polynucleotide, the

[0119] Y adapter allows different, noncomplementary, sequences to be added to the 5' and 3' ends of the library. The arms of the Y adapter comprise different sequences that maybe non-complementary, and the stem of the Y adapter, that is arranged to be ligated to the polynucleotide of interest, comprises double-stranded (i.e., complementary) DNA.

[0120] Amplification and incorporation of an indexing barcode

[0121] A “barcode”, “indexing barcode”, or “molecular barcode” as used herein, unless otherwise stated, refers to a nucleic acid molecule comprising a sequence that can serve as a molecular identifier. A barcode maybe a type of nucleic acid tag. For example, individual "barcode" sequences may be added to the polynucleotides in the sample for use in next-generation sequencing (NGS) so that the sequencing read can be identified and sorted before the final data analysis.

[0122] Prior to amplification, the polynucleotide sample comprises a mixture of both labelled and unlabelled polynucleotides. As the skilled person would understand, the amplification process will not replicate in the amplification products a nucleotide modification that is present in the polynucleotide sample. Thus, all of the polynucleotides that are generated in the amplification step will be unlabelled. The amplification process will, therefore, significantly reduce the proportion of polynucleotides in the sample that comprise a nucleotide modification. Indeed, following amplification, the number of polynucleotides in the sample that comprise a nucleotide modification may be many orders of magnitude lower than the number of unmodified and / or unlabelled polynucleotides present.

[0123] The inventors have surprisingly found, however, that not only is it possible to isolate polynucleotides that comprise a nucleotide modification from the sample following an amplification step, but there are, in fact, significant surprising advantages provided by performing the method in this way.

[0124] For example, by performing an amplification step prior to fractionation, the population of unlabelled polynucleotides has been found to be representative of the initial sample as a whole (e.g. the whole genome), as demonstrated, for example, in Figure 19. Thus, while the polynucleotides that comprise a nucleotide modification maybe removed and used to provide information about, for example, epigenetic modification of the polynucleotide sample, sequencing the unlabelled polynucleotides has been found to provide an accurate representation of mutation rates of the polynucleotide sample.

[0125] Amplification may be performed by any suitable method. For example, in some embodiments, the adapters that are ligated to the polynucleotides during library preparation may comprise a primer binding site that maybe used for the binding of an amplification primer. The polynucleotides may be amplified by PCR or qPCR, for example, using primers designed to anneal within the ligated sequencing adapters, and an appropriate PCR program.

[0126] In the disclosed method, amplification is used to incorporate a first indexing barcode at a first end of the polynucleotides but not a second end. The incorporation of a first indexing barcode at a first end of the polynucleotides but not a second end involves linear amplification as the polynucleotides are copied in one direction only. As discussed below, indexing barcodes are added to the polynucleotides for use in nextgeneration sequencing (NGS) so that the sequencing reads can be identified and sorted before the final data analysis. An example method by which a first indexing barcode maybe incorporated at a first end of the polynucleotides but not a second end is described in the Examples below and shown in Figure 1.

[0127] The use of indexing in Next Generation Sequencing, such as Illumina sequencing, allows DNA to be tracked at the single molecule level. This can be particularly advantageous when working at low DNA concentrations or when using a targeted sequencing approach, such as a liquid panel; and using PCR-based amplification of the genome, which can lead to the over-representation of individual DNA molecules in the final sequencing experiment. Indexes allow reads to be properly quantified and for replicated reads to be removed or combined bioinformatically.

[0128] In some embodiments, the indexing barcodes comprise unique molecular identifiers (UMIs). UMIs comprise a unique barcode comprising, for example, a random, short (e.g. 6-15) nucleotide sequence. UMIs may provide error correction and increased accuracy during sequencing. By incorporating individual barcodes on each original

[0129] DNA fragment, variant alleles present in the original sample (true variants) can be distinguished from errors introduced, for example, by PCR methods, during library preparation, target enrichment, or sequencing. Advantageously, the use of UMIs can reduce the rate of false-positive variant calls and increase the sensitivity of variant detection. The proportion of different variants in the sample, reflecting, for example, the mutation rate, can thus be accurately quantified. Since each nucleic acid in the starting material is tagged with a unique molecular barcode, bioinformatics software can filter out duplicate reads and PCR errors with a high level of accuracy and report unique reads, removing the identified errors before final data analysis.

[0130] The use of UMIs has been found to be particularly advantageous in embodiments in which target enrichment is performed on the second fraction, as this allows identification of an individual polynucleotide in the output of the sequencing experiment. This provides reliable quantification of reads and the prevention / removal of spurious results by allowing identification of PCR duplicates in the sequencing reads.

[0131] Labelling Approaches In some embodiments, the site-specific labelling of a target nucleotide residue may comprise binding a high affinity binding protein or antibody specific for a target modified nucleotide to each target modified nucleotide in a polynucleotide of the sample. A target modified nucleotide may be any nucleotide that is epigenetically modified in a target position, such as, for example, 5-methylcytosine. Site-specific labelling - methyl-CpG binding domain protein

[0132] In some embodiments, the high affinity binding protein may comprise a protein comprising a methyl-CpG binding domain protein (an MBD protein). In such embodiments, the site-specific labelling of a target nucleotide residue may comprise binding an MBD protein to each 5-methylcytosine nucleotide in a polynucleotide of the sample.

[0133] Thus, in some embodiments, the method may comprise:

[0134] (i)(a) preparing the polynucleotides into a sequencing library and additionally:

[0135] (1) amplifying the polynucleotides and incorporating a first indexing barcode at a first end of the polynucleotides but not a second end; and

[0136] (2) binding a methyl-CpG binding domain protein to 5-methylcytosine residues in the polynucleotides to form a first subset of polynucleotides, that comprise a bound methyl-CpG binding domain protein, and a second subset of polynucleotides, that do not comprise a bound methyl-CpG binding domain protein; (i)(b) fractionating the amplified sequencing library into first and second fractions, wherein the first fraction is enriched for polynucleotides comprising a bound methyl- CpG binding domain protein, and wherein the second fraction is enriched for polynucleotides lacking a bound methyl-CpG binding domain protein;

[0137] (ii) amplifying the first fraction and incorporating a second indexing barcode to a second end of each polynucleotide in the first fraction;

[0138] (iii) amplifying the second fraction and incorporating a third indexing barcode to a second end of each polynucleotide in the second fraction; and

[0139] (iv) pooling the amplified first and second fractions to obtain the prepared polynucleotide sample.

[0140] In some embodiments, binding a methyl-CpG binding domain protein to 5- methylcytosine residues in the polynucleotides may comprise a method as previously described in relation to approaches referred to as “MBD-Seq”. Site-specific labelling - antibody specific for - ethy Icy tosine In some embodiments, the antibody maybe an antibody specific for 5-methyl cytosine. In such embodiments, the site-specific labelling of a target nucleotide residue may comprise binding the antibody specific for 5-methyl cytosine to each 5-methyl cytosine nucleotide in a polynucleotide of the sample.

[0141] In some embodiments, binding an antibody specific for 5-methylcytosine to 5- methylcytosine residues in the polynucleotides may comprise a method as previously described in relation to approaches referred to as “MeDIP-Seq”. Any antibody used in previously described MeDIP-Seq methods maybe used in the disclosed method.

[0142] Site-specific labelling - antibody specific for -hydroxymethylcytosine

[0143] In some embodiments, the antibody maybe an antibody specific for 5- hydroxymethylcytosine. In such embodiments the site-specific labelling of a target nucleotide residue may comprise binding an antibody specific for 5- hydroxymethylcytosine to each 5-hydroxymethylcytosine nucleotide in a polynucleotide of the sample.

[0144] In some embodiments, binding an antibody specific for 5-hydroxymethylcytosine to 5- hydroxymethylcytosine residues in the polynucleotides may comprise a method as previously described in relation to approaches referred to as “HmeDIP-Seq”. Any antibody used in previously described HmeDIP-Seq methods may be used in the disclosed method.

[0145] Site-specific labelling - antibody specific for Na-methy Icy tosine In some embodiments, the antibody may be an antibody specific for N4- methylcytosine. In such embodiments, the site-specific labelling of a target nucleotide residue may comprise binding an antibody specific for Nq-methylcytosine to each N4- methylcytosine nucleotide in a polynucleotide of the sample. In some embodiments, binding an antibody specific for N4-methylcytosine to N4- methylcytosine residues in the polynucleotides may comprise a method as previously described in relation to approaches referred to as “m4C-IP-Seq”. Any antibody used in previously described m4C-IP-Seq methods may be used in the disclosed method.

[0146] Site-specific labelling - antibody specific for N6-methyladenine In some embodiments, the antibody maybe an antibody specific for N6-methyladenine. In such embodiments, the site-specific labelling of a target nucleotide residue may comprise binding an antibody specific for N6-methyladenine to each N6- methyladenine nucleotide in a polynucleotide of the sample.

[0147] In some embodiments, binding an antibody specific for N6-methyladenine to N6- methyladenine residues in the polynucleotides may comprise a method as previously described in relation to approaches referred to as “m6A-IP-Seq. Any antibody used in previously described m6A-IP-Seq methods may be used in the disclosed method.

[0148] In embodiments in which the site-specific label comprises an antibody specific for a target modified nucleotide residue, the method may comprise:

[0149] (i)(a) preparing the polynucleotides into a sequencing library and additionally:

[0150] (1) amplifying the polynucleotides and incorporating a first indexing barcode at a first end of the polynucleotides but not a second end; and

[0151] (2) binding the antibody to target modified nucleotide residues in the polynucleotides to form a first subset of polynucleotides, that comprise a site- specifically bound antibody, and a second subset of polynucleotides, that do not comprise a bound antibody; (i)(b) fractionating the amplified sequencing library into first and second fractions, wherein the first fraction is enriched for polynucleotides comprising a site-specifically bound antibody, and wherein the second fraction is enriched for polynucleotides lacking a bound antibody;

[0152] (ii) amplifying the first fraction and incorporating a second indexing barcode to a second end of each polynucleotide in the first fraction;

[0153] (iii) amplifying the second fraction and incorporating a third indexing barcode to a second end of each polynucleotide in the second fraction; and

[0154] (iv) pooling the amplified first and second fractions to obtain the prepared polynucleotide sample.

[0155] In some embodiments, the method may comprise dissociating double stranded polynucleotides into single strands (such as melting DNA) prior to binding the antibody to target modified nucleotide residues in the polynucleotides. Site-specific labelling - covalent bond with DNA binding protein

[0156] In some embodiments, the site-specific labelling of a target nucleotide residue may comprise forming a covalent bond between a nucleotide in a polynucleotide of the sample and an associated DNA-binding protein.

[0157] In some embodiments, forming a covalent bond between a nucleotide in a polynucleotide of the sample and an associated DNA-binding protein may comprise a method as previously described in relation to approaches referred to as “ChlP-Seq”.

[0158] Covalent bond formation in this context may be referred to as “cross-linking”. Thus, in some embodiments, the method may comprise:

[0159] (i)(a) preparing the polynucleotides into a sequencing library and additionally:

[0160] (1) forming a covalent bond between a nucleotide in a polynucleotide of the sample and an associated DNA-binding protein to form a first subset of polynucleotides, that comprise the cross-linked DNA-binding protein, and a second subset of polynucleotides, that do not comprise a cross-linked DNA-binding protein; and

[0161] (2) amplifying the polynucleotides and incorporating a first indexing barcode at a first end of the polynucleotides but not a second end;

[0162] (i)(b) fractionating the amplified sequencing library into first and second fractions, wherein the first fraction is enriched for polynucleotides comprising a cross-linked

[0163] DNA-binding protein, and wherein the second fraction is enriched for polynucleotides lacking a cross-linked DNA-binding protein;

[0164] (ii) amplifying the first fraction and incorporating a second indexing barcode to a second end of each polynucleotide in the first fraction; (iii) amplifying the second fraction and incorporating a third indexing barcode to a second end of each polynucleotide in the second fraction; and

[0165] (iv) pooling the amplified first and second fractions to obtain the prepared polynucleotide sample. In embodiments involving the formation of a covalent bond with a DNA-binding protein, the method may further comprise exonuclease treatment to remove nucleotides that are unshielded by the DNA-binding protein from the polynucleotides.

[0166] In some embodiments, the DNA-binding protein maybe a histone protein or a transcription factor. Site-specific labelling - Derivatisation of unmodified nucleotides

[0167] In some embodiments, labelling a target nucleotide residue may comprise binding a label or tag to an unmodified target nucleotide in a polynucleotide of the sample. An unmodified target nucleotide may be any nucleotide that is unmodified in a target position, that may be an epigenetically relevant position, such as, for example, cytosine that is unmethylated in the C5 position.

[0168] In some embodiments, the method may comprise an enzymatic “unmethylome” profiling approach, in which unmodified nucleotides, such as unmodified CpG dinucleotides, in a DNA sample are derivatised in such a way that they can then be isolated without bias and subsequently sequenced. Thus, the disclosed method may be used, for example, to provide a profile of genetic mutation rates in methylated and unmethylated regions of the genome. Thus, in some embodiments, the site-specific labelling of a target nucleotide residue may comprise:

[0169] (a) using a methyltransferase enzyme configured to modify a nucleotide residue in a target position to apply a tag to each unmodified nucleotide residue in a polynucleotide of the sample, wherein each unmodified nucleotide residue is unmodified in the target position;

[0170] (b) inactivating the methyltransferase;

[0171] (c) preparing the polynucleotide sample into a sequencing library;

[0172] (d) binding an affinity label to each tag. In preferred embodiments, the steps (a)-(d) maybe performed as a “one pot” method, wherein all of the steps are performed in a single container, thereby providing significant efficiencies in terms of time and reagents, and advantages in terms of automation. This approach has been found to maximise sensitivity, time, reagents and the overall yield of the polynucleotide enrichment.

[0173] In some embodiments, preparing an amplified sequencing library may involve at most one sample purification step.

[0174] In some embodiments, the method may comprise: (i)(a) preparing an amplified sequencing library comprising:

[0175] (1) - using a methyltransferase enzyme configured to modify a nucleotide residue in a target position to apply a tag to each unmodified nucleotide residue in a polynucleotide of the sample, wherein each unmodified nucleotide residue is unmodified in the target position, and subsequently inactivating the methyltransferase;

[0176] (2) - preparing the polynucleotide sample into a sequencing library; (3) - binding an affinity label to each tag; and

[0177] (4) - amplifying the polynucleotides and incorporating a first indexing barcode at a first end of the polynucleotides but not a second end, wherein steps (2), (3), and (4) are performed in any order or combination after step (1); (i)(b) fractionating the amplified sequencing library into first and second fractions, wherein the first fraction is enriched for polynucleotides comprising a label, and wherein the second fraction is enriched for polynucleotides lacking a label;

[0178] (ii) amplifying the first fraction and incorporating a second indexing barcode to a second end of each polynucleotide in the first fraction; (iii) amplifying the second fraction and incorporating a third indexing barcode to a second end of each polynucleotide in the second fraction; and

[0179] (iv) pooling the amplified first and second fractions to obtain the prepared polynucleotide sample. In preferred embodiments, step (2) is performed before step (4).

[0180] Thus, in some embodiments, step (2) is performed before steps (3) and (4). In such embodiments, steps (3) and (4) may be performed in any order. Thus, in some embodiments, the steps (2), (3), and (4) are performed in the sequence (2), then (4), then (3). In some embodiments, the steps (2), (3), and (4) are performed in numerical order, i.e. in the sequence (2), then (3), then (4).

[0181] Preferably, the steps (2), (3), and (4) are performed in the sequence (2), then (4), then (3)-

[0182] The inventors have identified that the disclosed method is particularly suitable for use with the disclosed “unmethylome” profiling approach. Previous unmethylome profiling approaches have been found to introduce bias in relation to densely modified polynucleotides, and the disclosed unmethylome profiling approach avoids this detrimental bias and is particularly suitable for use in combination with genetic analysis as described herein.

[0183] As discussed herein, these advantages have been made possible by minimising the loss of sample, by performing various operations in specific sequences and combinations.

[0184] In some embodiments in which the polynucleotide sample comprises DNA, the method maybe a method for preparing a DNA sample for sequencing, for determining both the modification status at the cytosine C5 position of each CpG dinucleotide of the DNA sample and the presence of a genetic mutation in the DNA sample, and the method may comprise the steps of:

[0185] (i)(a) preparing an amplified sequencing library comprising:

[0186] (1) - using a methyltransferase enzyme configured to modify the cytosine C5 position of a CpG dinucleotide to apply a tag to each unmodified cytosine residue of the sample, wherein each unmodified cytosine residue is the cytosine of a CpG dinucleotide that is unmodified in the C5 position, and subsequently inactivating the methyltransferase ;

[0187] (2) - preparing the DNA sample into a sequencing library;

[0188] (3) - binding an affinity label to each tag; (4) - amplifying the sequencing library and incorporating a first indexing barcode to a first end of each polynucleotide in the sequencing library; wherein steps (2), (3), and (4) are performed in any order or combination after step (1);

[0189] (i)(b) fractionating the amplified sequencing library into first and second fractions, wherein the first fraction is enriched for polynucleotides comprising an affinity label, and wherein the second fraction is enriched for polynucleotides lacking an affinity label;

[0190] (ii) amplifying the first fraction and incorporating a second indexing barcode to a second end of each polynucleotide in the first fraction; (iii) amplifying the second fraction and incorporating a third indexing barcode to a second end of each polynucleotide in the second fraction; and

[0191] (iv) pooling the amplified first and second fractions to obtain the prepared polynucleotide sample. In preferred embodiments, step (2) is performed before step (4). Thus, in some embodiments, step (2) is performed before steps (3) and (4). In such embodiments, steps (3) and (4) may be performed in any order.

[0192] Thus, in some embodiments, the steps (2), (3), and (4) are performed in the sequence (2), then (4), then (3). In some embodiments, the steps (2), (3), and (4) are performed in numerical order, i.e. in the sequence (2), then (3), then (4).

[0193] Preferably, the steps (2), (3), and (4) are performed in the sequence (2), then (4), then (3).

[0194] Derivatisation ofDNA - Tag

[0195] The polynucleotide may be derivatised using a methyltransferase enzyme to apply a tag to unmodified nucleotide residues. For example, the polynucleotide maybe derivatised using an appropriate methyltransferase enzyme to apply a tag to specific nucleotide residues that are unmodified in target positions. Thus, to determine the modification status of a specific nucleotide residue at a specific target position, a methyltransferase enzyme may be used that is configured to apply a tag to the target nucleotide residue in the target position. In some embodiments, the same tag may be used to derivatise different unmodified nucleotides and / or different target positions.

[0196] In other embodiments, different tags may be used to derivatise different unmodified nucleotides and / or different target positions.

[0197] In some embodiments, the polynucleotide maybe derivatised with different tags (e.g. on different unmodified nucleotides) sequentially, for example, with the inactivation of the first methyltransferase before the addition of a second, different methyltransferase. The use of a plurality of different tags advantageously allows the tags to be independently functionalised, for example, to provide selective enrichment / fractionation.

[0198] As used herein, unless otherwise indicated, references to the “target position” in which an unmodified nucleotide residue is unmodified refer to a specific position within the chemical structure of the nucleotide. The tag may be applied in the cytosine C5 position. Accordingly, the fragmented DNA sample will comprise tagged residues. A tagged cytosine residue may be understood to have the following structure: wherein R2is the tag.

[0199] The tag maybe applied in the N4 position in cytosine. Accordingly, the fragmented DNA sample will comprise tagged residues. A tagged cytosine residue maybe understood to have the following structure: wherein R2is the tag.

[0200] Similarly, a tag maybe added at the N6 position of adenine, the N2 or N7 position of guanine or at the 2’-0H position of ribose. In some embodiments, a tag at the 2’-0H position of ribose is a tag at the 2’-0H position of a terminal ribose.

[0201] Accordingly, a tagged adenine residue may be understood to have the following structure: wherein R2is the tag.

[0202] The term “tag”, which may also be referred to as a “linker”, a “functional linker” or “DNA tag”, as used herein, unless otherwise specified, refers to a reactive moiety that is applied site-specifically to the polynucleotide, such as fragmented DNA. Polynucleotides that have been tagged in this way may be referred to as “derivatised”.

[0203] The disclosed method may comprise the use of a methyltransferase cofactor analogue, such as a synthetic methyltransferase cofactor analogue, comprising the tag and a methyltransferase-binding moiety. Thus, using a methyltransferase enzyme to apply a tag to each unmodified nucleotide residue may comprise the use of a methyltransferase cofactor analogue. Thus, the method may comprise the use of a methyltransferase enzyme to catalyse the transfer of the tag from the methyltransferase cofactor analog to an unmodified nucleotide residue, such as to the C5 position of a cytosine base of an unmodified CpG dinucleotide, in a polynucleotide sample. The presence of a modification, such as a methyl group or other chemical modification of the nucleotide residue, such as in the C5 position within a CpG dinucleotide, prevents the transfer of the tag. Thus, in the disclosed method, only nucleotides, such as CpG dinucleotides, that are unmodified (such as unmethylated) in this position may be labelled with a tag.

[0204] The methyltransferase cofactor maybe an ion of formula (I): wherein, X is S or Se;

[0205] L1is -CH2- or -CH2CH2-;

[0206] R2is the tag;

[0207] R3 and R4are independently H or an optionally substituted C1-6 alkyl an optionally substituted C2-6 alkenyl or an optionally substituted C2-6 alkynyl; or R3 and R4together with the nitrogen to which they are attached, form an optionally substituted 5- or 6- membered heterocyclyl ring; and

[0208] Rs is NH2, NHBOC or H; or a salt, solvate or tautomer thereof. The ion of formula (I) may be provided together with a counterion. The counterion maybe an organic or inorganic anion carrying one or more negative charges. The counterion may be formate or acetate. R2may be -CH2-U-[L3]m-[HM]n-[L2]p-[R6]q, wherein: m, n, p and q are each independently selected from o and 1;

[0209] L2is a linker;

[0210] HM is a hydrolysable moiety;

[0211] L3 is a linker; U is an unsaturated group selected from an alkene, an alkyne, an aromatic group (e.g. aryl), a carbonyl group, SO and S02;

[0212] R6is a heavy atom or a heavy atom cluster suitable for phasing of X-ray diffraction data, a radioactive or stable rare isotope, a fluorophore, a fluorescence quencher, an affinity tag, a crosslinking agent, a nucleic acid cleaving reagent, a spin label, a chromophore, a protein, peptide or amino acid which may optionally be modified a nucleotide, nucleoside or nucleic acid which may optionally be modified, a carbohydrate, a lipid, a transfection reagent, an intercalating agent, a nanoparticle or bead, or a functional group, wherein the functional group is selected from the group consisting of: an amino group (including a protected amino), a thiol group, a 1,2-diol group, a hydrazino group, a hydroxyamino group, a haloacetamide group, a maleimide group, a cyanide group, a cyclic hydrocarbon (such as a bridged cyclic hydrocarbon (e.g. norbornene) or a cycloalkyl group (e.g. a C3-6 cycloalkyl), a halo group (e.g. -F, -Cl, -Br, -I), an aldehyde group, a ketone group, a 1,2-aminothiol group, a azido group, an isothiocyanate or thiocyanate group, an alkene group, such as a terminal alkene, an alkyne group, such as a terminal alkyne group, a 1,3-diene function, a dienophilic function (e.g. an activated carbon-carbon double bond), an arylhalide group, an arylboronic acid group, a terminal haloalkyne group, a terminal silylalkyne group, -N=C=O; -N=C=S, -0-C(0)NH2, a protected amino, a group comprising a sterically strained alkyne or alkene (such as norbornene or DBCO), a nitrone, a tetrazine, a tetrazole, and 1,2-aminothiol group.

[0213] In embodiments where R4is an optionally substituted Ci-4alkyl an optionally substituted C2.4alkenyl or an optionally substituted C2.4alkynyl, the alkyl, alkenyl or alkynyl may be unsubstituted or substituted with one or more substituents selected from the group consisting of: -NR7R8; -OH; -SH; -CN; -C(O)OR7; -C(O)R7; C(O)NR7R8; N3; and halo, wherein R7 and R8are independently H or a Ci-4alkyl. Halo may be F, Cl, Br or I.

[0214] Similarly, in embodiments where R3and R4together with the nitrogen to which they are attached, form an optionally substituted 5- or 6-membered heterocyclyl ring, the 5- or 6-membered heterocyclyl ring may be unsubstituted, or substituted with one or more substituents selected from the group consisting of: -NR7R8; -OH; -SH; -CN; - C(O)OR7; -C(O)R7; C(O)NR7R8; N3; and halo, wherein R7and R8are independently H or a C1-4 alkyl. Halo may be F, Cl, Br or I.

[0215] Synthetic methyltransferase cofactors are described in more detail in PCT / GB2022 / 052438, EP3186266B1 and US8008007B2. It maybe appreciated that preferred embodiments of the X, L1, R2, R3and R4groups in the compound of formula (I) may be as defined for the equivalent groups in these applications. X may be S.

[0216] L1may be -CH2CH2-.

[0217] R3may be H. Alternatively, R3may be an optionally substituted Ci-4alkyl an optionally substituted C2.4alkenyl or an optionally substituted C2.4alkynyl, more preferably an optionally substituted methyl or an optionally substituted ethyl. The alkyl, alkenyl or alkynyl may be unsubstituted or substituted with an OH. Accordingly, R3may be - CH2CH20H. R4maybe H.

[0218] R5may be NH2.

[0219] In some embodiments, q is 1. In some embodiments, R6is -N3. p may be 1.

[0220] L2maybe a linker comprising a backbone of between 1 and 50 atoms, between 2 and 40 atoms, between 3 and 30 atoms, between 4 and 20 or between 5 and 15 atoms. The backbone maybe made up of carbon, oxygen and / or nitrogen atoms. In embodiments where the linker comprises a cyclic group, the backbone may be understood to consist of the atoms which define the shortest possible route between the two ends of the linker group.

[0221] In some embodiments, L2comprises between i and 5 groups selected from an optionally substituted hydrocarbon, an optionally substituted polyether chain, an arylene moiety and a (C=O)NH group.

[0222] The hydrocarbon may be an optionally substituted alkylene, preferably an optionally substituted C1-10 alkylene and more preferably a Ci-5alkylene. The optionally substituted polyether chain may be an optionally substituted polyethylene glycol chain.

[0223] The polyethylene glycol chain may comprise up to 15 monomers, up to 10 monomers or up to 5 monomers of ethylene glycol. In some embodiments, the polyethylene glycol chain consists of between 1 and 5 or between 2 and 3 monomers of ethylene glycol. The arylene moiety may be a CeH4phenylene ring.

[0224] Accordingly, in some embodiments, L2maybe: wherein w is an integer from between 1 and 15, e.g. between 2 and 10 or between 3 and 5. In some embodiments, w is 2 or 3.

[0225] Alternatively, in some embodiments, p is o.

[0226] In some embodiments, n is 1.

[0227] The hydrolysable moiety may be , wherein Rxis hydrogen, deuterium or a Ci-4alkyl. The Ci-4alkyl maybe methyl. The hydrolysable moiety may be a Schiff base, for example, an imine moiety, an oxime moiety and / or a hydrazone moiety.

[0228] In some embodiments, the hydrolysable moiety comprises a disulphide (S-S) bond.

[0229] In some embodiments, the hydrolysable moiety is

[0230] In some embodiments, n is o.

[0231] In some embodiments, m is 1. L3 may be a linker comprising a linear chain of from 1 to 20, from 2 to 15, from 3 to 10 or from 4 to 9 atoms. The atoms may be carbon, oxygen and / or nitrogen atoms). The linker maybe substituted or unsubstituted. In some embodiments, L3comprises an optionally substituted hydrocarbon (e.g. an alkyl) chain.

[0232] In some embodiments, L3comprises an optionally substituted linear C1-10 alkyl chain, e.g. an optionally substituted C2-s or an optionally substituted C4-6 alkyl chain. In some embodiments the alkyl chain is unsubstituted. In some embodiments the alkyl chain is substituted. In some embodiments, L3is a linear, unsubstituted C2, C3or C4alkyl chain.

[0233] In some embodiments,

[0234] Accordingly, in some embodiments, R2is . In alternative embodiments, In some embodiments, the synthetic methyltransferase cofactor maybe:

[0235] , where R is H.

[0236] The above compound maybe called ETA-AdoHcy-N3

[0237] Derivatisation ofDNA - Methyltransferase

[0238] The methyltransferase may be any methyltransferase that is capable of using S- adenosyl methionine as a cofactor. Thus, the methyltransferase may be an S- adenosylmethionine-dependent methyltransferase, such as an S-adenosyl-L- methionine-dependent methyltransferase.

[0239] In some embodiments, the methyltransferase may be a cytosine-5 (C5) methyltransferase, such as a bacterial cytosine C5 methyltransferase.

[0240] In some embodiments, the methyltransferase may be an adenine methyltransferase, such as a bacterial adenine methyltransferase. For example, the methyltransferase may be M.TaqI, which is a DNA adenine methyltransferase.

[0241] In some embodiments, the methyltransferase may be a methyltransferase from Mycoplasma.

[0242] In some embodiments, the methyltransferase may be a constitutively active methyltransferase.

[0243] In some embodiments, the methyltransferase maybe one of the enzymes described in US 2017 / 0283453. In some embodiments, the methyltransferase maybe M.Mpel, M.Hhal, M.SssI, M.AccII, M.MspI or M.TaqI. The methyltransferase may be an active mutant, variant, and / or fragment of M.Mpel, M.Hhal, M.SssI, M.AccII, M.MspI or M.TaqI. M.Mpel has been found to be particularly advantageous for use in the disclosed method, in part due to being particularly non-selective in terms of target locus. Thus, the methyltransferase maybe M.Mpel or an active mutant, variant, and / or fragment thereof.

[0244] Although methyltransferase enzymes share a relatively low level of sequence similarity, they do share a highly conserved structural fold. This conserved fold is known as the Rossmann fold and comprises a series of beta strand and alpha helical segments, in which the beta strands are hydrogen bonded to form a beta-sheet.

[0245] In some embodiments, the cofactor binding pocket of the methyltransferase enzyme may be modified within the Rossman fold, for example, by substitution of one or more amino acids, to improve the suitability of the enzyme for use in the disclosed method, such as, for example, by improving cofactor compatibility. For example, one or more amino acids within the Rossman fold of the methyltransferase enzyme may be substituted, for example, to reduce or relieve potential steric interaction with the cofactor analogue. In some embodiments, the methyltransferase may be modified such that an amino acid having a relatively large side chain, such as, for example, glutamine or asparagine may be substituted for an amino acid comprising a shorter side chain, such as, for example, alanine. The use of a methyltransferase enzyme that has been modified in this way may be particularly desirable when larger cofactor analogues are used, such as cofactor analogues comprising transferrable groups with longer alkyl-chains than those with shorter chains.

[0246] In some embodiments, the methyltransferase may be any bacterial cytosine C5 methyltransferase enzyme comprising one or more, such as 2, 3, 4, 5, 6, or 7, amino acid substitutions in the Rossman fold.

[0247] In some embodiments, the methyltransferase may be any bacterial cytosine C5 methyltransferase enzyme comprising an amino acid substitution in the position of the amino acid residue of the Rossman fold corresponding to the residue that is Gln82, Tyr254, and / or Asn3O4 in the wild type sequence of the M.Hhal methyltransferase (i.e. the sequence having the NCBI accession number P05102). In some embodiments, the methyltransferase may be any bacterial cytosine C5 methyltransferase enzyme comprising an alanine residue in the position of the amino acid residue of the Rossman fold corresponding to the residue that is Gln82, Tyr254, and / or Asn3O4 in the wild type sequence of the M.Hhal methyltransferase (i.e. the sequence having the NCBI accession number P05102).

[0248] In some embodiments, the methyltransferase maybe an M.Mpel methyltransferase.

[0249] The M.Mpel methyltransferase enzyme has been found to be particularly advantageous for use in the disclosed method due to non-selectively targeting any and all CpG dinucleotides for modification.

[0250] In some embodiments, the methyltransferase maybe, or may comprise, a variant, and / or fragment of the wild type M.Mpel sequence, which is defined as the sequence having the NCBI accession number BAC44284.

[0251] In some embodiments, the methyltransferase maybe, or may comprise, a variant, and / or fragment of the wild type M.Mpel sequence, comprising at least 80% sequence identity, such as at least 85%, 90%, or 95% sequence identity to the wild type M.Mpel sequence having the NCBI accession number BAC44284.

[0252] In some embodiments, the methyltransferase may comprise one or more, such as 2, 3, 4, 5, 6 or 7 amino acid substitutions relative to the wild type M.Mpel sequence having the NCBI accession number BAC44284.

[0253] In some embodiments, the use of the methyltransferase enzyme to apply the tag to unmodified nucleotides, such as unmodified CpG dinucleotides, may be carried out under conditions which enable the methyltransferase to transfer the tag from the methyltransferase cofactor analogue to the target DNA.

[0254] In some embodiments, the reaction mixture may be incubated at a temperature of from 10 to 6o°C, from 20 to 5O°C, or from 30 to 4O°C. Preferably, the reaction mixture may be incubated at a temperature of about 37°C. In some embodiments, the incubation may be performed for a time sufficient to enable transfer of the tag to all of the available unmodified nucleotides, such as unmodified CpG dinucleotides, in the fragmented DNA sample. The incubation may be performed for a period of 5 minutes to 5 hours, 10 minutes to 4 hours, 15 minutes to 3 hours, 30 minutes to 2 hours, or 40 to 90 minutes. Preferably the incubation is performed for a period of about 1 hour.

[0255] In some embodiments, the incubation may be performed in a suitable buffer at a pH that is selected based on the methyltransferase that is being used. For example, the pH maybe between 7.5 and 8.5, such as between 7.8 and 8.2, or about pH 8. Methyltransferase inactivation

[0256] After an appropriate incubation to label the unmodified nucleotides, such as unmodified CpG dinucleotides, in the sample with a tag, the presence of methyltransferase in the subsequent processing of the sample has been found to reduce the efficiency of the method. Thus, the method may comprise the inactivation of the methyltransferase enzyme.

[0257] In some embodiments, methyltransferase enzymes have been found to bind tightly to DNA, thereby inhibiting downstream processing of the DNA. Methods comprising removal of the methyltransferase or purification of the sample have been found to reduce the efficiency of the process due to the additional time and reagents required and due to the loss of sample.

[0258] Thus, in some embodiments, the methyltransferase enzyme may be inactivated in the reaction mixture. The inactivation of the methyltransferase in this way has surprisingly been found to provide significant processing efficiencies in the disclosed method.

[0259] The terms “inactivated” and “inactivation” as used herein, unless otherwise specified, refer to any alteration in the structure and / or function of the methyltransferase that prevents further activity of the methyltransferase on the target polynucleotide. Thus, the terms “inactive” and “inactivated” as used herein, unless otherwise specified, refer to an enzyme that has less than 10%, such as less than 5%, less than 2%, or preferably less than 1% of its maximum activity.

[0260] Alterations in the structure and / or function of the methyltransferase that prevent further activity of the methyltransferase on the target polynucleotide may include, for example, denaturation, modification, inhibition, and / or fragmentation of the methyltransferase. Thus, in some embodiments, inactivation of the methyltransferase may comprise denaturation of the methyltransferase. In some embodiments, inactivation of the methyltransferase may comprise modification of the methyltransferase. In some embodiments, inactivation of the methyltransferase may comprise inhibition of the methyltransferase. In some embodiments, inactivation of the methyltransferase may comprise fragmentation of the methyltransferase.

[0261] The methyltransferase may be inactivated by any suitable method. Suitable methods include changing the environmental conditions of the methyltransferase, and targeted inactivation of the methyltransferase.

[0262] Changing the environmental conditions may consist of or comprise, for example, changing the temperature and / or pH of the reaction mixture. Thus, in some embodiments, inactivation of the methyltransferase may comprise incubation of the reaction mixture at a temperature of from 55 to 85°C, from 60 to 8o°C, or from 65 to 75°C. In some embodiments, inactivation of the methyltransferase may comprise incubation of the reaction mixture at a temperature of from 55 to 65°C, such as at or about 6o°C, or from 75 to 85°C, such as at or about 8o°C.

[0263] In some embodiments, inactivation of the methyltransferase may comprise incubation at an elevated temperature for a period of 5 minutes to 1 hour, or 10-30 minutes. Preferably inactivation of the methyltransferase may comprise incubation at an elevated temperature for about 15 minutes.

[0264] Targeted inactivation of the methyltransferase may comprise the addition of an agent to alter the structure and / or function of the methyltransferase. Such an agent may comprise, for example, a methyltransferase inhibitor. Any suitable methyltransferase inhibitor may be used, including, for example, 5-azacitidine, decitabine, clofarabine, arsenic trioxide, guadecitabine, RX-3117, 5-fluoro-2’-deoxycytidine, 5,6-dihydro-5- azacytidine, cladribine, fludarabine, fazarabine, procaine, EGCG, hydralazine, genistein, equol, curcumin, disulfiram, resveratrol, and / or caffeic acid. The methyltransferase inhibitor may be a S-Adenyl-l-methionine (SAM) analogue, such as sinefungin or S-adenosyl-l-homocysteine (SAH). In some embodiments, the methyltransferase may be inactivated in the reaction mixture after an appropriate incubation to label the unmodified nucleotides, such as unmodified CpG dinucleotides, with a tag, thereby terminating the derivatisation reaction. Advantageously, the presence of the inactivated methyltransferase has not been found to be detrimental to subsequent processing. On the contrary, the inactivation of the methyltransferase enzyme in the reaction mixture this way, rather than by removal or dilution, has been found to provide increased efficiencies and significantly improved yields in subsequent steps of the process. The inactivation of the methyltransferase may, therefore, provide significant advantages by removing the requirement for purification of the polynucleotide at this stage, and permitting the efficient combination of the methyltransferase and library preparation processes in a single reaction mixture. These efficiency advantages are shown in the Examples.

[0265] Library preparation ofderivatised DNA

[0266] In embodiments comprising the use of a methyltransferase to bind a label or tag to an unmodified target nucleotide the reaction mixture comprises a buffer mixture such as the labelling buffer, together with inactive methyltransferase, excess cofactor analogue, and the polynucleotide sample. In general for sequencing applications, library preparation is typically conducted with purified DNA. It has surprisingly been found by the inventors, however, that the library preparation process maybe performed directly in the reaction mixture following methyltransferase inactivation and that the efficiency of the library preparation process is not compromised by the use of a different buffer, or the presence of unpurified sample, such as DNA, and / or residual enzyme and cofactor components in the mixture. This finding provides a significant advantage over previous methods, offering significant efficiencies in terms of savings of time and reagents. In particular, the finding that any washing procedure maybe avoided significantly preserves the level of polynucleotide sample present in the reaction mixture. Thus, in such embodiments, the preparation of an amplified sequencing library is preferably performed in a one-pot approach.

[0267] Thus, in some embodiments, the method comprises the inactivation of the methyltransferase followed by library preparation without any intervening steps or clean-up process, for example, comprising removal of inactivated enzymes, exchange of reaction buffer, or isolation or purification of the polynucleotide sample. Performing library preparation at this stage, for example, prior to any enrichment process, and without the requirement for any washing steps or clean-up of the sample, surprisingly provides significant processing efficiencies, including significantly reducing any loss of polynucleotide sample.

[0268] In some embodiments, preparing the polynucleotide into a sequencing library may comprise end repair, A-tailing, and adapter ligation of the polynucleotides in the sample. Affinity Labelling ofderivatised DNA

[0269] In embodiments comprising the use of a methyltransferase to bind a label or tag to an unmodified target nucleotide the method comprises the affinity labelling of the derivatised polynucleotides. Tags on the polynucleotides in the sample maybe modified by the addition of an affinity label. The affinity label may be referred to as an “affinity label” when bound to the tag and an “affinity label precursor” beforehand.

[0270] It has been found that an affinity label may be added to the tag by the addition of the affinity label precursor to the reaction mixture. The finding that an affinity label may be applied to the tag in this technically simple and efficient manner is advantageous in view of the fact that the reaction mixture comprises various components including, inactive methyltransferase, excess cofactor analogue, and the reagents and enzymes required for library preparation. The finding that the affinity label may be added to the tag in this way provides a significant advantage over previous methods, by avoiding the requirement for a washing step, thereby providing efficiency savings in terms of time and reagents and avoiding any loss of sample.

[0271] Thus, in some embodiments, the method may comprise the addition of an affinity label to the tag after library preparation, without any intervening steps or clean-up process, for example, comprising removal of peptides or enzymes, exchange of reaction buffer, or isolation or purification of the sample.

[0272] In some embodiments, binding an affinity label to each tag may comprise adding an affinity label precursor directly into the sequencing library preparation mixture, without a washing step.

[0273] In some embodiments, the affinity label may comprise biotin. In some embodiments, the affinity label precursor maybe a compound of formula (II):

[0274] R9-L4-R10

[0275] (ID wherein: R9is a reactive moiety configured to react with a group in the tag and to thereby form a bond therebetween;

[0276] L4 is a linker; and

[0277] R10comprises or consists of biotin.

[0278] R9may be an optionally substituted 5 to 30 membered heterocyclyl, an optionally substituted 5 to 30 membered heteroaryl, an optionally substituted C6-3Omembered aryl or an optionally substituted C3-3Ocycloalkyl. A multicyclic group may be understood to be a group comprising two or more fused rings. Accordingly, a multicyclic group may have 2 or 3 fused rings.

[0279] As used herein, a “heterocyclyl”, “heterocyclic” or “heterocycle” group includes nonaromatic saturated or partially saturated mono and multicyclic groups. A heterocyclic ring contains 1 or more heteroatoms in the ring, which may independently selected from nitrogen, oxygen or sulfur. A multicyclic group may be understood to be multicyclic heterocyclyl group if it contains at least one heteroatom and at least one ring which is a non-aromatic saturated or partially saturated ring. As used herein, a “cycloalkyl” group includes non-aromatic saturated or partially saturated mono and multicyclic groups. A multicyclic group may be understood to be multicyclic cycloalkyl group if it only contains carbon atoms in the rings and it contains at least one ring which is a non-aromatic saturated or partially saturated ring. As used herein, a “heteroaryl” group includes aromatic mono and multicyclic groups. A heteroaryl ring contains 1 or more heteroatoms in the ring, which may independently selected from nitrogen, oxygen or sulfur. A multicyclic group may be understood to be multicyclic heteroaryl group if it contains at least one heteroatom and every ring is aromatic.

[0280] Preferably, R9contains a triple bond. Preferably, R9is an optionally substituted to to 20 membered multicyclic heterocyclyl, an optionally substituted 10 to 20 membered multicyclic heteroaryl or an optionally substituted Cw-20 multicyclic cycloalkyl. R9maybe a 14 to 18 membered multicyclic heterocyclyl, an optionally substituted 14 to 18 membered multicyclic heteroaryl or an optionally substituted C13-18 multicyclic cycloalkyl.

[0281] In a preferred embodiment, , wherein X2is N or CH. Preferably, X2is

[0282] N.

[0283] L4may comprise between 1 and 12 groups, each group selected from an optionally substituted hydrocarbon, an optionally substituted polyether chain, NH, O, S or S-S.

[0284] The hydrocarbon may be an optionally substituted alkylene, preferably an optionally substituted C1-10 alkylene and more preferably a Ci-5alkylene. The alkylene may be substituted with an OH or oxo group. Preferably, the alkylene is substituted with an oxo group.

[0285] The optionally substituted polyether chain may be an optionally substituted polyethylene glycol chain. The polyethylene glycol chain may comprise up to 15 monomers, up to 10 monomers or up to 5 monomers of ethylene glycol.

[0286] Accordingly, L4may have the structure

[0287] -L5-L6-I -L8-*, wherein Ls to L8are each independently absent or an optionally substituted hydrocarbon, an optionally substituted polyether chain, an NH, O, S or S-S; and an asterisk indicates a point of bonding to R10.

[0288] In some embodiments, Ls is an optionally substituted hydrocarbon. Accordingly, Ls may be C0CH2CH2.

[0289] In some embodiments, L6is NH. In some embodiments, 17 is an optionally substituted hydrocarbon. Accordingly, 15 may be C0CH2CH2. In some embodiments, L8is an optionally substituted polyether chain. The optionally substituted polyether chain may be an optionally substituted polyethylene glycol chain. The polyethylene glycol chain may comprise up to 15 monomers, up to 10 monomers or up to 5 monomers of ethylene glycol. Accordingly, L8maybe (0CH2CH2)r, where r is an integer between 1 and 15, more preferably between 2 and 10 or between 3 and 5. In some embodiments, r is 4.

[0290] Accordingly,

[0291] L4 may have no charge.

[0292] Negatively charged linkers have been found to react poorly with the tag. Preferably, L3 is not negatively charged.

[0293] R10may have the following formula:

[0294] RU-(CH2)S-L9- wherein R11is biotin s is an integer between 1 and 8; and

[0295] L9is absent or is COO or CONH.

[0296] Sulfonated linkers have been found to react particularly poorly with the tag. Preferably, L4and R10are not sulfonated.

[0297] Preferably, the affinity label precursor is not DBCO-SS-biotin. Preferably, the affinity label precursor is not NHS-SS-biotin.

[0298] A modified tagged cytosine residue, which comprises the affinity label, may be understood to have the following structure: wherein L4and R10are as defined above; and

[0299] L10is a linker. L10may be understood to be -CHs-U-CLoJm-fHMJn-EL^p-L11-, wherein U, L2, L3, HM, m, n and p are as defined above and L11is a linker formed due to a reaction between the R6and R9groups.

[0300] Accordingly, L11may asterisk indicates a point of bonding to L4and X2is as defined above.

[0301] The present inventors have surprisingly found that in previous methods, such as that described by Kriukiene et al. (Nature Communications 2013 4:2190), DNA fragments having a significant density of CpG sites, such as, for example, 5 or more CpG sites per 100 bp, may be underrepresented in the sequencing reads, thereby introducing bias to the results.

[0302] An advantage of the disclosed method is that if a purification process, such as DNA isolation, is performed at this point, it is the only clean-up step for the entire process, and this has been found to dramatically improve the efficiency and sensitivity of the process. This is made possible, firstly, by the inactivation of the methyltransferase and, secondly, the surprising finding that the enzymes used for library preparation exhibit high levels of activity in the resulting buffers, which are significantly different to the buffer mixtures designed for use in library preparation.

[0303] Thus, in some embodiments, the method may comprise, after the affinity labelling step, and before the fractionation step, a step of purifying the polynucleotide. Preferably the method involves no more than one step of purifying the polynucleotide. Any suitable method for purifying the polynucleotide maybe used. For example, DNA may be purified using a DNA purification kit. The DNA may be washed, for example, using ethanol, such as 80% ethanol, or other DNA washing buffer. After washing, the DNA may be eluted, for example, using a suitable elution buffer, such as phosphate buffer.

[0304] Amplification ofderivatised DNA

[0305] In embodiments comprising the use of a methyltransferase to bind a label or tag to an unmodified target nucleotide the polynucleotides are preferably amplified prior to fractionation.

[0306] The inventors have surprisingly found that standard DNA polymerases are able to amplify polynucleotides that have been tagged and, in some embodiments, affinity labelled using the disclosed approach. Indeed, it has advantageously been found that DNA polymerases are able to amplify DNA comprising a plurality of affinity labels at a high density, without bias. The derivatised polynucleotides may, therefore, be amplified and sequenced without further modification.

[0307] Moreover, amplification has advantageously been found to be possible under the conditions employed in the preceding steps of the disclosed method. This negates the need for DNA purification prior to amplification and provides the possibility of performing the method as a one-pot approach, and automating the method.

[0308] Significantly, prior to amplification, the polynucleotide sample comprises a mixture of both labelled and unlabelled polynucleotides. As discussed above, the amplification process will not generate a tag or affinity label on the amplification products. Thus, all of the polynucleotides that are generated in the amplification step will be unlabelled. The amplification process will, therefore, significantly reduce the proportion of polynucleotides in the sample that comprise an affinity label. Indeed, following amplification, the number of labelled polynucleotides in the sample may be many orders of magnitude lower than the number of unlabelled polynucleotides present.

[0309] The inventors have surprisingly found, however, that not only is it possible to isolate tagged polynucleotides from the sample following an amplification step, but performing the method in this way provides an effective and highly sensitive approach for determining both the status of nucleotide residues and the presence of a genetic mutation in the polynucleotide sample.

[0310] Fractionation The method comprises fractionating the polynucleotides into first and second fractions, wherein the first fraction is enriched for polynucleotides comprising a label, and the second fraction is enriched for polynucleotides lacking a label.

[0311] Fractionation is performed after amplification of the sequencing library. After amplification, the proportion of labelled polynucleotides in the sequencing library may be many orders of magnitude smaller than the proportion of unlabelled polynucleotides. The inventors have surprisingly found, however, that the labelled polynucleotides maybe isolated and separated from the unlabelled polynucleotides with high efficiency.

[0312] In some embodiments, fractionation may comprise the use of the affinity label, such that polynucleotides comprising an affinity label are separated, using the affinity label, from polynucleotides lacking an affinity label. As a result, the first fraction is enriched for polynucleotides comprising an affinity label, in the sense that substantially or entirely all of the polynucleotides in the first fraction comprise an affinity label. Similarly, the second fraction is enriched for polynucleotides lacking an affinity label, in the sense that substantially or entirely all of the polynucleotides in the second fraction lack an affinity label. The fractionation may comprise selectively isolating the labelled polynucleotides using the affinity label. For example, fractionation of the polynucleotides may comprise binding of the affinity label to a capture agent that specifically binds to the affinity label. Any suitable affinity label may be used, and any suitable capture agent may be used, the affinity label and capture agent being selected in combination based on high affinity selective binding.

[0313] In some embodiments, the affinity label and / or capture agent may comprise an antibody. In some embodiments, the affinity label comprises biotin, and fractionating the sequencing library into first and second fractions may comprise fractionation using a capture agent comprising a biotin-binding protein. In embodiments in which the affinity label comprises biotin, fractionation of the polynucleotides may comprise selectively isolating the labelled polynucleotides using a biotin-binding protein.

[0314] The biotin-binding protein may comprise, for example, streptavidin, avidin, and / or a biotin-specific antibody.

[0315] The biotin-binding protein may comprise streptavidin, or a functional analogue or derivative of streptavidin. In some embodiments, fractionation may comprise the use a separation medium or substrate. For example, the capture agent maybe conjugated to a surface. The surface may comprise a plurality of microbeads, such as paramagnetic microbeads.

[0316] In embodiments in which the separation medium or substrate comprises a plurality of microbeads coated with the capture agent, the method may comprise binding the labelled polynucleotides to the capture agent on the coated microbeads and then isolating the coated microbeads. Isolation of the microbeads may be performed by centrifugation. In embodiments comprising the use of paramagnetic microbeads, isolation of the microbeads may comprise the application of a magnetic field to the reaction mixture to separate the beads from the remainder of the suspension.

[0317] For example, in some embodiments, the capture agent may comprise streptavidin conjugated to the surface of microbeads. Preferably the streptavidin-coated microbeads maybe streptavidin-coated paramagnetic microbeads.

[0318] After an appropriate incubation to bind the labelled polynucleotides to the capture agent, the agent may be washed to remove unbound and non-specifically bound polynucleotides. In embodiments in which the affinity label comprises biotin, the biotin-binding protein may be washed by any suitable method to remove unbound and non-specifically bound polynucleotides. After the selective isolation of the labelled polynucleotides, the polynucleotides are separated from the capture agent.

[0319] In previous methods, such as that described by Kriukiene et al. (Nature Communications 20134:2190), DNA fragments are released from a streptavidin capture agent using oxidative cleavage of a disulfide bond within the affinity label. However, this method has been found by the present inventors to be inconsistently reproducible and to have poor efficiency. In the disclosed method, the polynucleotides are preferably not separated from the capture agent by a method comprising oxidative cleavage of the tag or affinity label. Thus, in some embodiments, the method does not comprise the separation of the polynucleotides from the capture agent by cleavage, such as oxidative cleavage or hydrolysis, of the tag or affinity label.

[0320] In some embodiments, the method may comprise the denaturation of the capture agent. For example, in embodiments in which the capture probe comprises a biotinbinding protein, the method may comprise the denaturation of the biotin-binding protein. This method has been found to be particularly advantageous due to the consistent release of DNA fragments regardless of the number of the attached affinity labels.

[0321] The ability of streptavidin to bind to biotin is dependent on both a sterically defined binding pocket and the highly polar residues within it. Any agent that induces a conformational change of streptavidin may, therefore, be used to release the labelled polynucleotides. The inventors have found that, in embodiments in which the biotinbinding protein comprises streptavidin, the labelled polynucleotides maybe released from the streptavidin by any method that denatures streptavidin without damaging the polynucleotides.

[0322] In some embodiments, the labelled polynucleotides may be released from the streptavidin by incubation in pure water at a temperature of about 7O°C.

[0323] In some embodiments, the labelled polynucleotides may be released from the streptavidin by incubation in 12-15% (v / v) phenol at room temperature. In some embodiments, streptavidin may be denatured using a denaturing reagent, such as 1% sodium dodecyl sulphate and heating the sample to 9O°C.

[0324] Advantageously, because the affinity label is not damaged by this method comprising the denaturation of streptavidin, the first fraction may be further enriched for polynucleotides comprising an affinity label by repeating the selective isolation (affinity purification) step in one or more further cycles.

[0325] Target enrichment Particular diseases and conditions may be associated with specific genomic regions. For this and other reasons, in some embodiments, target enrichment of the second fraction maybe performed.

[0326] Target enrichment is used to describe a variety of strategies to selectively isolate specific genomic regions of interest for sequencing analysis. Any suitable method for target enrichment may be used in the disclosed method. The most suitable approach may depend on the specific aim of the study. For example, embodiments in which the aim is to enrich genomic regions that have clinical relevance may require a more focused enrichment strategy. On the other hand, embodiments in which the aim is to discover novel variants that may be associated with a given phenotype may require an enrichment strategy that provides a balance between sequencing costs and target coverage. For example, the diagnosis of specific cancers may require the identification of somatic variants present at extremely low abundance in cfDNA or in mixtures of malignant and stromal cells, which may necessitate an increased depth of sequencing coverage rather than a broader genomic approach which may be economically impractical.

[0327] Thus, in some embodiments, step (iii) may further comprise target enrichment of the second fraction for one or more genome regions of interest. In some embodiments, step (iii) may further comprise target enrichment of the second fraction for one or more genome regions of interest prior to amplifying and incorporating a third indexing barcode.

[0328] Thus, in some embodiments, the method may comprise: (iii) enriching the second fraction for one or more genome regions of interest, amplifying the enriched second fraction and incorporating a third indexing barcode to a second end of each polynucleotide in the second fraction.

[0329] Thus, in some embodiments, the method may comprise: (i) preparing a fractionated amplified sequencing library, wherein each polynucleotide of the fractionated amplified sequencing library comprises a first indexing barcode at a first end but not a second end, and wherein a first fraction of the fractionated amplified sequencing library is enriched for polynucleotides comprising a site-specifically bound label, and wherein the second fraction is enriched for polynucleotides lacking a label;

[0330] (ii) amplifying the first fraction and incorporating a second indexing barcode to a second end of each polynucleotide in the first fraction;

[0331] (iii) enriching the second fraction for one or more genome regions of interest, amplifying the enriched second fraction and incorporating a third indexing barcode to a second end of each polynucleotide in the second fraction; and

[0332] (iv) pooling the amplified first and second fractions to obtain the prepared polynucleotide sample.

[0333] Thus, in some embodiments, the method may comprise: (i)(a) preparing an amplified sequencing library, wherein each polynucleotide of the amplified sequencing library comprises a first indexing barcode at a first end but not a second end, and wherein the amplified sequencing library further comprises a label bound site- specifically to a first subset of the polynucleotides but not a second subset; (i)(b) fractionating the amplified sequencing library into first and second fractions, wherein the first fraction is enriched for polynucleotides comprising a label, and wherein the second fraction is enriched for polynucleotides lacking a label;

[0334] (ii) amplifying the first fraction and incorporating a second indexing barcode to a second end of each polynucleotide in the first fraction; (iii) enriching the second fraction for one or more genome regions of interest, amplifying the enriched second fraction and incorporating a third indexing barcode to a second end of each polynucleotide in the second fraction; and

[0335] (iv) pooling the amplified first and second fractions to obtain the prepared polynucleotide sample.

[0336] Any suitable method for target enrichment may be used. In some embodiments, the target enrichment may comprise an in-solution hybridization-based approach. For example, in some embodiments, target enrichment may comprise the use of biotinylated oligonucleotide “bait” probes to capture genomic regions of interest, for example, using streptavidin-coated magnetic beads. Suitably, bait probes may comprise 50-150 nucleotides. Thus, in some embodiments, target enrichment of the second fraction may comprise in-solution hybridization using oligonucleotide “bait” probes specific for genomic regions of interest. In some embodiments, the target enrichment may comprise a PCR-based enrichment method. For example, in some embodiments, target enrichment may comprise the use of specifically designed primers to amplify in parallel up to 250 target regions using PCR. Other target enrichment strategies that may suitably be used include multiplex extension ligation, molecular inversion probes (MIPS) / padlock probes, nested patch PCR, and selector probes.

[0337] In a fourth aspect, there is provided a polynucleotide sample for sequencing, wherein the polynucleotide sample is obtained or obtainable by the method for the third aspect.

[0338] In a fifth aspect, there is provided a method for obtaining sequencing information of a polynucleotide sample, the method comprising: - obtaining a polynucleotide sample prepared by the method for the third aspect; and sequencing the polynucleotide sample to obtain sequencing information.

[0339] Thus, the method for the fifth aspect may comprise: (i) preparing a fractionated amplified sequencing library, wherein each polynucleotide of the fractionated amplified sequencing library comprises a first indexing barcode at a first end but not a second end, and wherein a first fraction of the fractionated amplified sequencing library is enriched for polynucleotides comprising a site-specifically bound label, and wherein the second fraction is enriched for polynucleotides lacking a label; (ii) amplifying the first fraction and incorporating a second indexing barcode to a second end of each polynucleotide in the first fraction;

[0340] (iii) amplifying the second fraction and incorporating a third indexing barcode to a second end of each polynucleotide in the second fraction; (iv) pooling the amplified first and second fractions to obtain the prepared polynucleotide sample; and

[0341] (v) sequencing the polynucleotide sample to obtain sequencing information.

[0342] Thus, in some embodiments, the method may comprise: (i) preparing a fractionated amplified sequencing library, wherein each polynucleotide of the fractionated amplified sequencing library comprises a first indexing barcode at a first end but not a second end, and wherein a first fraction of the fractionated amplified sequencing library is enriched for polynucleotides comprising a site-specifically bound label, and wherein the second fraction is enriched for polynucleotides lacking a label;

[0343] (ii) amplifying the first fraction and incorporating a second indexing barcode to a second end of each polynucleotide in the first fraction;

[0344] (iii) enriching the second fraction for one or more genome regions of interest, amplifying the enriched second fraction and incorporating a third indexing barcode to a second end of each polynucleotide in the second fraction; and

[0345] (iv) pooling the amplified first and second fractions to obtain the prepared polynucleotide sample; and

[0346] (v) sequencing the polynucleotide sample to obtain sequencing information. Thus, in some embodiments, the method may comprise:

[0347] (i)(a) preparing an amplified sequencing library, wherein each polynucleotide of the amplified sequencing library comprises a first indexing barcode at a first end but not a second end, and wherein the amplified sequencing library further comprises a label bound site- specifically to a first subset of the polynucleotides but not a second subset;

[0348] (i)(b) fractionating the amplified sequencing library into first and second fractions, wherein the first fraction is enriched for polynucleotides comprising a label, and wherein the second fraction is enriched for polynucleotides lacking a label;

[0349] (ii) amplifying the first fraction and incorporating a second indexing barcode to a second end of each polynucleotide in the first fraction; (iii) amplifying the second fraction and incorporating a third indexing barcode to a second end of each polynucleotide in the second fraction;

[0350] (iv) pooling the amplified first and second fractions to obtain the prepared polynucleotide sample; and (v) sequencing the polynucleotide sample to obtain sequencing information.

[0351] Sequencing

[0352] The polynucleotides of the first and second fractions are pooled and the resulting polynucleotide sample is sequenced. In some embodiments, the polynucleotides of the first and second fractions may be amplified and pooled prior to sequencing of the resulting polynucleotide sample.

[0353] Substantially or entirely all of the polynucleotides in the first fraction comprise a target modified or unmodified nucleotide. The sequence data obtained from the first fraction may be used to determine the modification status of nucleotide residues in the polynucleotide sample. For example, the sequence data obtained from the first fraction may be used to generate a methylome or unmethylome profile of the polynucleotide sample. Using the disclosed method, and as demonstrated in the Examples, target modified or unmodified nucleotides, such as unmodified CpG dinucleotides, in the polynucleotide sample are derivatised in such a way that they may be isolated and subsequently sequenced without bias. Thus, the first fraction maybe used with extremely low sample quantities to provide a highly accurate epigenetic profile of the polynucleotide sample.

[0354] The inventors have surprisingly found that the presence of the disclosed affinity label on the polynucleotides advantageously does not interfere with the action of DNA polymerases. The derivatised polynucleotides may, therefore, be amplified and sequenced without further modification.

[0355] Based on this surprising finding, and the non-destructive nature of the disclosed method, the inventors have developed the disclosed method for determining both the status of nucleotide residues and the presence of a genetic mutation in a polynucleotide sample. The second fraction, that is enriched for unlabelled polynucleotides, is representative of the initial sample as a whole (e.g. the whole genome). Thus, sequencing the polynucleotides of the second (unlabelled) fraction, has been found to provide an accurate representation of mutation rates of the polynucleotide sample.

[0356] The sequence data obtained from the second fraction may be used to produce a genetic mutation profile of the polynucleotide sample.

[0357] The sequence data obtained from the second fraction is representative of the initial polynucleotide sample as a whole and thus may be used to provide a genetic mutation profile of the modified and unmodified genomic fractions.

[0358] In some embodiments, the polynucleotides of the first and / or second fraction maybe purified prior to sequencing by any suitable method for cleaning up PCR products for use in a sequencing platform.

[0359] In some embodiments, the polynucleotides of the first and / or second fraction maybe used directly for sequencing without further purification. The term “sequencing” as used herein, unless otherwise indicated, refers to any method that maybe used to determine the sequence (i.e. the order of nucleotides) in a nucleic acid such as DNA or RNA.

[0360] Any type of sequencing platform may be used to determine the sequences of the polynucleotides, in combination with the appropriately ligated sequencing adapter.

[0361] Thus, sequencing approaches that maybe suitable for use in the disclosed method include, but are not limited to, Sanger sequencing, high-throughput sequencing, pyrosequencing, sequencing-by-synthesis, single-molecule sequencing, nanoporebased sequencing, semiconductor sequencing, sequencing-by-ligation, sequencing-by- hybridization, RNA-Seq (Illumina), Digital Gene Expression (Helicos), next generation sequencing (NGS), Single Molecule Sequencing by Synthesis (SMSS) (Helicos), massively-parallel sequencing, Clonal Single Molecule Array (Solexa), shotgun sequencing, Ion Torrent, Oxford Nanopore, Roche Genia, Maxim-Gilbert sequencing, primer walking, sequencing using Singular Genomics, Ultima Genomics, Element Biosciences, PacBio, SOLiD, Ion Torrent, or Nanopore platforms. Sequencing reactions can be performed in a variety of sample processing units, which may include multiple lanes, multiple channels, multiple wells, or other means of processing multiple sample sets substantially simultaneously.

[0362] The method may comprise a high throughput sequencing method. The terms “next generation sequencing”, “NGS”, and “high throughput sequencing” as used herein, refer to sequencing technologies having increased throughput as compared to traditional Sanger- and capillary electrophoresis-based approaches. The high throughput sequencing method maybe capable of generating hundreds of thousands of sequence reads in parallel. The method may comprise a multiplex sequencing technique. The high throughput sequencing methods that may be used include, but are not limited to, sequencing by synthesis, sequencing by ligation, and sequencing by hybridization. The sequencing method maybe capable of sequencing single molecules.

[0363] In some embodiments, the method comprises sequencing the polynucleotides using the Illumina platform.

[0364] In a sixth aspect, there is provided sequencing information, wherein the sequencing information is obtained or obtainable by the method for the fifth aspect. In a seventh aspect, there is provided a method for determining both the status of nucleotide residues and the presence of a genetic mutation in a polynucleotide sample, the method comprising: obtaining sequencing information using the method for the fifth aspect; using the indexing barcodes to distinguish the sequencing information of the polynucleotides of the first and second fractions; and using the sequencing information of the polynucleotides of the first fraction to determine the status of nucleotide residues in the polynucleotide sample, and using the sequencing information of the polynucleotides of the second fraction to determine the genetic sequence of a region of the polynucleotide sample and thereby the presence of a genetic mutation in the polynucleotide sample.

[0365] Thus, in some embodiments, the method for the seventh aspect may comprise:

[0366] (i) preparing a fractionated amplified sequencing library, wherein each polynucleotide of the fractionated amplified sequencing library comprises a first indexing barcode at a first end but not a second end, and wherein a first fraction of the fractionated amplified sequencing library is enriched for polynucleotides comprising a site-specifically bound label, and wherein the second fraction is enriched for polynucleotides lacking a label;

[0367] (ii) amplifying the first fraction and incorporating a second indexing barcode to a second end of each polynucleotide in the first fraction; (iii) amplifying the second fraction and incorporating a third indexing barcode to a second end of each polynucleotide in the second fraction;

[0368] (iv) pooling the amplified first and second fractions to obtain the prepared polynucleotide sample;

[0369] (v) sequencing the polynucleotide sample to obtain sequencing information; (vi) using the indexing barcodes to distinguish the sequencing information of the polynucleotides of the first and second fractions; and

[0370] (vii) using the sequencing information of the polynucleotides of the first fraction to determine the status of nucleotide residues in the polynucleotide sample, and using the sequencing information of the polynucleotides of the second fraction to determine the genetic sequence of a region of the polynucleotide sample and thereby the presence of a genetic mutation in the polynucleotide sample.

[0371] Thus, in some embodiments, the method may comprise:

[0372] (i)(a) preparing an amplified sequencing library, wherein each polynucleotide of the amplified sequencing library comprises a first indexing barcode at a first end but not a second end, and wherein the amplified sequencing library further comprises a label bound site- specifically to a first subset of the polynucleotides but not a second subset;

[0373] (i)(b) fractionating the amplified sequencing library into first and second fractions, wherein the first fraction is enriched for polynucleotides comprising a label, and wherein the second fraction is enriched for polynucleotides lacking a label;

[0374] (ii) amplifying the first fraction and incorporating a second indexing barcode to a second end of each polynucleotide in the first fraction;

[0375] (iii) amplifying the second fraction and incorporating a third indexing barcode to a second end of each polynucleotide in the second fraction; and

[0376] (iv) pooling the amplified first and second fractions to obtain the prepared polynucleotide sample; and

[0377] (v) sequencing the polynucleotide sample to obtain sequencing information;

[0378] (vi) using the indexing barcodes to distinguish the sequencing information of the polynucleotides of the first and second fractions; and (vii) using the sequencing information of the polynucleotides of the first fraction to determine the status of nucleotide residues in the polynucleotide sample, and using the sequencing information of the polynucleotides of the second fraction to determine the genetic sequence of a region of the polynucleotide sample and thereby the presence of a genetic mutation in the polynucleotide sample.

[0379] In some embodiments, the method may comprise comparing the sequencing reads of the polynucleotide sample to a reference sequence. In some embodiments, the method may comprise comparing the sequencing reads of the polynucleotide sample to a reference genome to determine the genomic location of the sequencing reads.

[0380] In some embodiments, the reference sequence may be a human genome, or may comprise one or more portions thereof, such as one or more chromosomes and / or chromosomal regions. Thus, in some embodiments, the method may comprise aligning the sequencing reads of the polynucleotide sample to the reference sequence. Any suitable alignment method compatible with high throughput sequencing data may be used, for example, using the Burrows Wheeler Alignment algorithm.

[0381] Aligned reads may be normalised. Any method for normalisation used in the art may maybe used. For example, normalising the sequencing reads may comprise reporting the number of reads in an aligned region (bin) as a fraction of reads per million reads of the sequencing output.

[0382] In some embodiments, the method may comprise comparing the normalised read counts from two or more sequencing experiments. Such an approach maybe advantageous in methods further comprising the step of diagnosing a disease based on the modification status of the nucleotide residues in the sample.

[0383] In an eighth aspect, there is provided a method for preparing a profile of a region of a polynucleotide sample, the profile comprising both the status of nucleotide residues and any genetic mutations in the region of the polynucleotide sample, the method comprising: - obtaining the status of nucleotide residues in the polynucleotide sample and the genetic sequence of a region of the polynucleotide sample using the method for the seventh aspect; comparing the sequences of the sequencing reads of the first and second fractions to a reference sequence to determine the location of the sequencing reads within the reference sequence and thereby the status of specific nucleotides and / or genetic mutations at specific locations within the reference sequence.

[0384] Thus, in some embodiments, the method for the eighth aspect may comprise:

[0385] (i) preparing a fractionated amplified sequencing library, wherein each polynucleotide of the fractionated amplified sequencing library comprises a first indexing barcode at a first end but not a second end, and wherein a first fraction of the fractionated amplified sequencing library is enriched for polynucleotides comprising a site-specifically bound label, and wherein the second fraction is enriched for polynucleotides lacking a label; (ii) amplifying the first fraction and incorporating a second indexing barcode to a second end of each polynucleotide in the first fraction;

[0386] (iii) amplifying the second fraction and incorporating a third indexing barcode to a second end of each polynucleotide in the second fraction;

[0387] (iv) pooling the amplified first and second fractions to obtain the prepared polynucleotide sample;

[0388] (v) sequencing the polynucleotide sample to obtain sequencing information;

[0389] (vi) using the indexing barcodes to distinguish the sequencing information of the polynucleotides of the first and second fractions;

[0390] (vii) using the sequencing information of the polynucleotides of the first fraction to determine the status of nucleotide residues in the polynucleotide sample, and using the sequencing information of the polynucleotides of the second fraction to determine the genetic sequence of a region of the polynucleotide sample and thereby the presence of a genetic mutation in the polynucleotide sample; and

[0391] (viii) comparing the sequences of the sequencing reads to a reference sequence to determine the location of the sequencing reads within the reference sequence and thereby the status of specific nucleotides and / or genetic mutations at specific locations within the reference sequence.

[0392] Thus, in some embodiments, the method may comprise: (i)(a) preparing an amplified sequencing library, wherein each polynucleotide of the amplified sequencing library comprises a first indexing barcode at a first end but not a second end, and wherein the amplified sequencing library further comprises a label bound site- specifically to a first subset of the polynucleotides but not a second subset;

[0393] (i)(b) fractionating the amplified sequencing library into first and second fractions, wherein the first fraction is enriched for polynucleotides comprising a label, and wherein the second fraction is enriched for polynucleotides lacking a label;

[0394] (ii) amplifying the first fraction and incorporating a second indexing barcode to a second end of each polynucleotide in the first fraction;

[0395] (iii) amplifying the second fraction and incorporating a third indexing barcode to a second end of each polynucleotide in the second fraction; and

[0396] (iv) pooling the amplified first and second fractions to obtain the prepared polynucleotide sample; and

[0397] (v) sequencing the polynucleotide sample to obtain sequencing information;

[0398] (vi) using the indexing barcodes to distinguish the sequencing information of the polynucleotides of the first and second fractions; and

[0399] (vii) using the sequencing information of the polynucleotides of the first fraction to determine the status of nucleotide residues in the polynucleotide sample, and using the sequencing information of the polynucleotides of the second fraction to determine the genetic sequence of a region of the polynucleotide sample and thereby the presence of a genetic mutation in the polynucleotide sample; and

[0400] (viii) comparing the sequences of the sequencing reads to a reference sequence to determine the location of the sequencing reads within the reference sequence and thereby the status of specific nucleotides and / or genetic mutations at specific locations within the reference sequence.

[0401] In some embodiments, the method may be an in vitro method for diagnosing disease in a subject, the method comprising diagnosing the disease based on a profile obtained by the method for the eighth aspect, using a polynucleotide sample obtained from the subject.

[0402] In a ninth aspect, there is provided a profile of a region of a polynucleotide sample, comprising both the status of nucleotide residues and any genetic mutations in the region of the polynucleotide sample, wherein the profile is obtained or obtainable by the method for the eighth aspect. In some embodiments, the region may consist of or comprise one or more specific portions of the polynucleotide sample. In such embodiments, the method may comprise target enrichment of the second fraction for one or more genome regions of interest prior to amplifying and incorporating a third indexing barcode.

[0403] In some embodiments, the region may consist of or comprise the entire polynucleotide sample, which may be, for example, a whole genome.

[0404] Methods comprising making a determination based on the modification status of nucleotide residues and the presence of any genetic mutations in a polynucleotide may comprise the production of a profile. For example, the profile may reflect the position of modified and / or unmodified residues and genetic mutations within a polynucleotide such as a portion of a genome or an entire genome. Accordingly, methods comprising making a determination based on the CpG modification status of a plurality of CpG dinucleotides in a polynucleotide and the presence of any genetic mutations may comprise the production of a profile, wherein the profile may reflect the position of modified and / or unmodified CpG dinucleotides and any genetic mutations within a polynucleotide such as a portion of a genome or an entire genome. Thus, in some embodiments, comparing the modification status of nucleotide residues and the presence of any genetic mutations in a polynucleotide from the subject to the modification status and genetic sequence of the corresponding residues in a reference sample may comprise comparing the profile obtained from the sample with the profile of a reference sample. For example, comparing the modification status of cytosine residues in CpG dinucleotides of a polynucleotide and the presence of any genetic mutations from the subject to the CpG modification status and genetic sequence of the corresponding residues in a reference sample may comprise comparing the profile obtained from the sample with the profile of a reference sample. In some embodiments, the profile from the reference sample may comprise a profile that is representative of a healthy individual. In other embodiments, the profile from the reference sample may comprise a profile that is obtained from, or indicative of a particular disease, such as, for example, a cancer. In some embodiments, the method may comprise comparing the profile obtained from the sample with a database of profiles. The database of profiles may comprise a plurality of profiles relating to a single disease, wherein the disease may be diagnosed in the subject from which the test sample was derived based on similarities between the profile of the test sample and the database of profiles. In other embodiments, the database of profiles may comprise a plurality of profiles representative of different diseases, wherein a disease maybe diagnosed in the subject from which the test sample was derived based on similarities between the profile of the test sample and one of more of the profiles within the database. A comparison between profiles may be made using any suitable method. For example, a comparison may be made statistically, using an appropriate metric, for example, a p- value. A comparison may also be made using a machine-learning platform.

[0405] In embodiments in which the method comprises determining, for a plurality of nucleotides, the presence (modified) or absence (unmodified) of any chemical modification catalysed by a methyltransferase, “modification status” may also be referred to as the “profile”. Thus, the disclosed method maybe used to determine a profile of methyltransferase catalysed modifications within a polynucleotide sample. In some embodiments, the method may comprise the detection of unmodified cytosine residues in CpG dinucleotides of a DNA sample. The terms “CpG”, “CpG site”, and “CpG dinucleotide”, are used interchangeably herein to refer to a cytosine-phosphate-guanine sequence in a 5’ to 3’ direction in the backbone of a nucleic acid. The terms “CpG modification status” and “modification status” as used interchangeably herein, unless otherwise stated, refer to the presence (modified) or absence (unmodified) of any chemical modification at the C5 position of cytosine within one or a plurality of CpG dinucleotides. In some embodiments, the method may be a method for preparing both a profile of the modification status at the cytosine C5 position of CpG dinucleotides and a genetic mutation profile, wherein the method further comprises comparing the sequences of the sequencing reads of the prepared polynucleotide sample to a reference sequence to determine the location of the sequencing reads within the reference sequence and thereby the presence or otherwise of unmodified cytosine residues in specific CpG dinucleotides and / or genetic mutations within the reference sequence. In embodiments in which the method comprises determining, for a plurality of CpG dinucleotides, the presence (modified) or absence (unmodified) of any chemical modification at the C5 position of each of the plurality of cytosines, the “CpG modification status” and “modification status” may also be referred to as the “profile”.

[0406] In embodiments in which the profile corresponds to the entire genome, the profile may be referred to as the “unmethylome profile” or “unmethylome”. The terms “genetic profile”, and “genetic mutation profile” as used herein, unless otherwise indicated, may be used to refer to the presence or absence of a genetic mutation at each position within the polynucleotide sequence of interest. Thus, the disclosed method may be used to determine a profile of genetic mutations within a polynucleotide sample.

[0407] A profile may comprise both the modification status of nucleotide residues and the genetic profile of a polynucleotide from the subject.

[0408] In some embodiments, the polynucleotide maybe a DNA sample, or maybe a mixed sample, comprising DNA and RNA. In some embodiments, the sample maybe a DNA sample comprising an epigenome.

[0409] The terms “epigenome” and “epigenetic” as used herein, unless otherwise specified, refer to the chemical modification of a polynucleotide or genome in such a way that gene expression is regulated.

[0410] Thus, in some embodiments, the method may be a method for determining both the epigenetic profile and genetic mutation profile of a genomic DNA sample, the method further comprising determining the epigenetic profile based on the modification status of nucleotide residues in the sample. For example, the method may comprise determining the epigenetic profile based on the modification status of cytosine residues in CpG dinucleotides of the sample, i.e. the CpG modification status.

[0411] In some embodiments, the method may be a method for analysing a polynucleotide sample, such as a DNA sample, from a subject. The method preferably does not comprise the use of bisulfite, such as bisulfite conversion of cytosine to uracil.

[0412] The method preferably does not comprise pyridine borane base conversion or enzymatic deamination of unmethylated cytosine.

[0413] The sequencing library may be suitable for use with high throughout sequencing methods. Sequencing the polynucleotides preferably comprises the use of next generation sequencing, such as sequencing applications on the Illumina platform. The method preferably does not comprise the use of a nanopore-based sequencing method.

[0414] In some embodiments, the method may be a method for determining both the modification status of one or more specific nucleotides and any genetic mutations in a polynucleotide sample from a subject.

[0415] In some embodiments, the method may be a method for determining both the modification status of the cytosine residue in one or more specific CpG dinucleotides and any genetic mutations and / or any specific genotype in one or more biomarkers in a sample from a subject.

[0416] In some embodiments, the method may be a method for determining both the modification status of the cytosine residues in one or more CpG dinucleotides and any genetic mutations in a plurality of genomic regions in a sample from a subject.

[0417] In some embodiments, the method may be a method for determining both the modification status of one or more specific adenine nucleotides and any genetic mutations and / or any specific genotype in one or more biomarkers in a sample from a subject.

[0418] In some embodiments, the method may be a method for determining both the modification status of one or more adenine residues and any genetic mutations in a plurality of genomic regions in a sample from a subject. In some embodiments, the method may be an in vitro method performed on a polynucleotide sample that has previously been obtained from a subject. The term “subject”, as used herein, may refer to any type of organism, including for example, a mammalian species (such as a human or domesticated animal), other animal species, a plant such as a crop, or other type of organism, including single celled organisms, and viruses. The subject may be a developing organism, such as an embryo or foetus. The subject maybe a healthy individual. The subject maybe an individual that has, or is suspected of having, a disease or predisposition to a disease. The subject may be an individual in need of therapy or suspected of needing therapy. In some embodiments, the method may be a method for determining the disease status of a subject. Accordingly, the method may comprise determining both the modification status of nucleotide residues and the presence of a genetic mutation in a polynucleotide in a sample from the subject using the disclosed method, and determining the disease status based on the modification status and any genetic mutations present. For example, the method may comprise determining both the modification status of cytosine residues in CpG dinucleotides of a polynucleotide in a sample from the subject and the presence of a genetic mutation in the sample using the disclosed method, and determining the disease status based on the CpG modification status and any genetic mutations present.

[0419] In some embodiments, the method may be a method for diagnosing a disease in a subject. Accordingly, the method may comprise determining both the modification status of nucleotide residues and the presence of a genetic mutation in a polynucleotide in a sample from the subject using the disclosed method, and diagnosing the disease based on the modification status and any genetic mutations present. For example, the method may comprise determining both the modification status of cytosine residues in CpG dinucleotides of a polynucleotide in a sample from the subject and the presence of a genetic mutation in the sample using the disclosed method, and diagnosing the disease based on the CpG modification status and any genetic mutations present.

[0420] In some embodiments, the method may be a method for making a disease prognosis in a subject. Accordingly, the method may comprise determining both the modification status of nucleotide residues and the presence of a genetic mutation in a polynucleotide in a sample from the subject using the disclosed method, and making a disease prognosis based on the modification status and any genetic mutations present. For example, the method may comprise determining both the modification status of cytosine residues in CpG dinucleotides of a polynucleotide in a sample from the subject and the presence of a genetic mutation in the sample using the disclosed method, and making a disease prognosis based on the CpG modification status and any genetic mutations present.

[0421] Methods comprising making a determination based on the nucleotide modification status may comprise comparing the modification status of specific nucleotide residues in a polynucleotide from the subject to the modification status of the corresponding residues in a reference sample. Accordingly, methods comprising making a determination based on the CpG modification status may comprise comparing the modification status of cytosine residues in CpG dinucleotides of a polynucleotide from the subject to the CpG modification status of the corresponding residues in a reference sample. Likewise, methods comprising making a determination based on the presence of a genetic mutation in a polynucleotide in a sample may comprise comparing a nucleotide sequence from a polynucleotide in the sample to the nucleotide sequence of the corresponding residues in a reference sample. In some embodiments, the reference sample may comprise a polynucleotide from a healthy subject. The reference sample may comprise a polynucleotide from a diseased subject. The reference sample may comprise a polynucleotide from the same subject as the test sample, taken at a different time point and / or from a different location in the body. Differences in the nucleotide modification status between the test and reference samples, and / or the presence of a genetic mutation, may be indicative of the presence or absence of a particular phenotype or clinical feature.

[0422] In some embodiments, the method may be a method for treating a subject. Accordingly, the method may comprise determining both the modification status of nucleotide residues and the presence of a genetic mutation in a polynucleotide in a sample from the subject using the disclosed method, diagnosing a disease based on the nucleotide modification status and any genetic mutations present, and providing a therapeutic composition to the subject to treat the disease based on the diagnosis. For example, the method may comprise determining both the modification status of cytosine residues in CpG dinucleotides of a polynucleotide in a sample from the subject and the presence of a genetic mutation in the sample using the disclosed method, diagnosing a disease based on the CpG modification status and any genetic mutations present, and providing a therapeutic composition to the subject to treat the disease based on the diagnosis.

[0423] In some embodiments, the method may be a method for determining a personalised or precision method for treatment for a subject. Accordingly, the method may comprise determining both the modification status of nucleotide residues and the presence of a genetic mutation in a polynucleotide in a sample from the subject using the disclosed method, determining the disease status of the subject based on the nucleotide modification status and any genetic mutations present, and determining a personalised medical treatment for the subject based on the disease status. For example, the method may comprise determining both the modification status of specific cytosine residues in CpG dinucleotides of a polynucleotide in a sample from the subject and the presence of a genetic mutation in the sample using the disclosed method, determining the disease status of the subject based on the CpG modification status and any genetic mutations present, and determining a personalised medical treatment for the subject based on the disease status.

[0424] In some embodiments, the method may be a personalised or precision method for treating a subject. Accordingly, the method may comprise determining both the modification status of nucleotide residues and the presence of a genetic mutation in a polynucleotide in a sample from the subject using the disclosed method, determining the disease status of the subject based on the nucleotide modification status and any genetic mutations present, and providing a personalised medical treatment to the subject based on the disease status. For example, the method may comprise determining both the modification status of specific cytosine residues in CpG dinucleotides of a polynucleotide in a sample from the subject and the presence of a genetic mutation in the sample using the disclosed method, determining the disease status of the subject based on the CpG modification status and any genetic mutations present, and providing a personalised medical treatment to the subject based on the disease status.

[0425] In some embodiments, the subject maybe an individual that has been diagnosed with having a disease. The subject maybe an individual that has been identified as being predisposed to, or at risk of having, a disease. The subject may be an individual that has not been diagnosed with having a disease. In some embodiments, the subject maybe an individual that has been diagnosed with cancer. The subject maybe pending or undergoing treatment such as a cancer therapy. The subject can be in remission of a cancer. Cancer can be identified on the basis of epigenetic variations. Cancer may be associated with both DNA hypomethylation and hypermethylation, but these two types of epigenetic abnormalities may affect different DNA sequences, and occur at different stages of cancer progression. For example, genomic hypermethylation in cancer maybe seen in CpG islands in gene regions, whereas hypomethylation may be observed in repeated DNA sequences in cancer, including heterochromatic DNA repeats, retrotransposons, and endogenous retroviral elements. In addition, unique sequences, such as transcription control sequences, are often subject to cancer-associated hypomethylation. These epigenetic changes may be detected using the disclosed method.

[0426] Cancer is also associated with the presence of genetic mutations. Hundreds of different genetic mutations, including sequence alterations, insertions, and deletions, have been found to be associated with different cancers. Genetic mutations associated with certain cancers may occur in specific genomic regions, for example, spontaneously from exposure to a carcinogen, or as an inherited genetic variant. Cancer-associated genetic mutations are known to frequently occur in the transcriptional regulatory regions of genes, including epigentically regulated regions, and it is broadly understood that the cancer genome is hypomethylated, relative to a healthy genome. The disclosed method advantageously allows profiling of the mutation rates in epigenetically modified and unmodified regions of the genome, providing an improved method for the diagnosis of disease. In this context, references to “genetic mutations” encompass tumour- associated genotypes.

[0427] Thus, the method maybe a method for diagnosing cancer in subject. Accordingly, the method may comprise determining both the modification status of nucleotide residues and the presence of a genetic mutation in a polynucleotide in a sample from the subject using the disclosed method, and making a cancer diagnosis based on the nucleotide modification status and any genetic mutations present. For example, the method may comprise determining both the modification status of cytosine residues in CpG dinucleotides of a polynucleotide in a sample from the subject and the presence of a genetic mutation in the sample using the disclosed method, and making a cancer diagnosis based on the CpG modification status and any genetic mutations present.

[0428] The frequency of cancer-linked DNA hypomethylation, the nature of the affected sequences, and the absence of associations with DNA hypermethylation are believed to suggest a role for DNA hypomethylation early in carcinogenesis and cancer formation, but can also be associated with tumour progression.

[0429] Thus, in some embodiments, the method may be a method for detecting cancer in a subject. Accordingly, the method may comprise determining both the modification status of nucleotide residues and the presence of a genetic mutation in a polynucleotide in a sample from the subject using the disclosed method, and determining the presence or absence of cancer based on the modification status of the nucleotide residues and any genetic mutations present. For example, the method may comprise determining both the modification status of cytosine residues in CpG dinucleotides of a polynucleotide in a sample from the subject and the presence of a genetic mutation in the sample using the disclosed method, and determining the presence or absence of cancer based on the modification status of the cytosine residues and any genetic mutations present. The method may comprise determining both the modification status of adenine residues in a sample from the subject and the presence of a genetic mutation in the sample using the disclosed method, and determining the presence or absence of cancer based on the modification status of the adenine residues and any genetic mutations present. In some embodiments, the method may comprise the analysis of a plurality of genomic regions, and detecting the presence or absence of cancer from the modification status of nucleotide residues and the presence of any genetic mutations in the plurality of genomic regions. In some embodiments, the method may be a method for detecting any type of cancer.

[0430] Different types of cancer may be preferentially detected and / or analysed using different sampling approaches based on the disclosed method.

[0431] In some embodiments, the method may be a method for treating cancer in a subject. Accordingly, the method may comprise determining both the modification status of nucleotide residues and the presence of a genetic mutation in a polynucleotide in a sample from the subject using the disclosed method, diagnosing a cancer based on the nucleotide modification status and any genetic mutations present, and providing a therapeutic composition to the subject to treat the cancer based on the diagnosis. For example, the method may comprise determining both the modification status of cytosine residues in CpG dinucleotides of a polynucleotide in a sample from the subject and the presence of a genetic mutation in the sample using the disclosed method, diagnosing a cancer based on the CpG modification status and any genetic mutations present, and providing a therapeutic composition to the subject to treat the cancer based on the diagnosis.

[0432] In some embodiments, the method may be a method for determining a personalised or precision method for cancer treatment for a subject. Accordingly, the method may comprise determining both the modification status of nucleotide residues and the presence of a genetic mutation in a polynucleotide in a sample from the subject using the disclosed method, determining the genetic profile of the cancer based on the nucleotide modification status and any genetic mutations present, and determining a personalised medical treatment for the subject based on the genetic profile of the cancer. For example, the method may comprise determining both the modification status of specific cytosine residues in CpG dinucleotides of a polynucleotide in a sample from the subject and the presence of a genetic mutation in the sample using the disclosed method, determining the genetic profile of the cancer based on the CpG modification status and any genetic mutations present, and determining a personalised medical treatment for the subject based on the genetic profile of the cancer. In some embodiments, the method may be a personalised or precision method for treating cancer in a subject. Accordingly, the method may comprise determining both the modification status of nucleotide residues and the presence of a genetic mutation in a polynucleotide in a sample from the subject using the disclosed method, determining the genetic profile of the cancer based on the nucleotide modification status and any genetic mutations present, and providing a personalised medical treatment to the subject based on the genetic profile of the cancer. For example, the method may comprise determining both the modification status of specific cytosine residues in CpG dinucleotides of a polynucleotide in a sample from the subject and the presence of a genetic mutation in the sample using the disclosed method, determining the genetic profile of the cancer based on the CpG modification status and any genetic modifications present, and providing a personalised medical treatment to the subject based on the genetic profile of the cancer.

[0433] In some embodiments, the method may be an in vitro method performed on a DNA sample that has previously been obtained from a tissue biopsy. Biopsy is a diagnostic procedure for cancers and other diseases. For example, tissue biopsy may provide material for cancer genotyping, which may assist in the design of targeted therapeutic approaches. In some embodiments, the method may be a method for genotyping a cancerous or otherwise diseased tissue. Accordingly, the method may comprise determining both the modification status of nucleotide residues and the presence of a genetic mutation in a polynucleotide from a biopsy of the tissue using the disclosed method. For example, the method may comprise determining both the modification status of cytosine residues in CpG dinucleotides of a polynucleotide and the presence of a genetic mutation in the polynucleotide from a biopsy of the tissue using the disclosed method. The method may further comprise designing a targeted therapeutic approach based on the nucleotide modification status and any genetic mutations present. In some embodiments, the method may be performed on a DNA sample that has previously been obtained from a biopsy of any type of tissue from a subject.

[0434] Existing tissue biopsy-based cancer diagnostic procedures may have limitations in relation to the analysis of the development and progression of certain types of cancers, due to tumor heterogeneity and evolution.

[0435] Liquid biopsy, which has the advantage of minimal invasiveness, has shown potential in detecting cancers, including early stage cancers and pre-cancerous lesions. The analysis of cell-free DNA (“cfDNA” or “circulating cfDNA”), which refers to DNA present at very low concentration in various bodily fluids, comprises extracellular nucleic acid fragments, for example, released by damaged cells during apoptosis, necrosis, or secretion. cfDNA has been found to exhibit the genetic and epigenetic alterations of cancers, including mutations, copy number alterations, chromosomal rearrangements, hypermethylation, and hypomethylation. The analysis of cfDNA has the potential to revolutionise the detection of early stage cancers and other diseases. In samples from cancer patients, cfDNA may comprise circulating tumor DNA (“ctDNA”), which is cell free tumor-derived fragmented DNA in a bodily fluid. Thus, cfDNA may comprise ctDNA. As demonstrated in the enclosed examples, the disclosed method has advantageously been found to be capable of providing high quality and consistently reproducible results from the very low concentrations of nucleic acid that are typically present in liquid biopsy (such as circulating cfDNA) samples. Thus, the sample for use in the disclosed method may be a cfDNA sample. The sample may consist of or comprise ctDNA.

[0436] Various liquid biopsy samples may be used for the analysis of cfDNA, including blood, plasma, urine, and spinal fluid. Preferably the liquid biopsy sample may be a blood or plasma sample.

[0437] Thus, in some embodiments, the method may be an in vitro method for diagnosing disease in a cfDNA sample from a subject. Accordingly, the method may comprise determining both the modification status of nucleotide residues and the presence of a genetic mutation in a cfDNA sample from the subject, and diagnosing the disease based on the nucleotide modification status and any genetic mutations present. For example, the method may comprise determining both the modification status of cytosine residues in CpG dinucleotides of a cfDNA sample from the subject and the presence of a genetic mutation in the sample, and diagnosing the disease based on the CpG modification status and any genetic mutations present.

[0438] In some embodiments, the method may also be used to determine the tissue of origin of the nucleic acid present in a liquid biopsy sample.

[0439] Thus, in some embodiments, the method may be a method for identifying the cellular origin of cfDNA in a sample from a subject. Accordingly, the method may comprise determining both the modification status of nucleotide residues and the presence of a genetic mutation in a polynucleotide in a sample from the subject using the disclosed method, and identifying the cellular origin of the DNA based on the modification status of nucleotide residues and any genetic mutations present in the sample.

[0440] In some embodiments, the method may be a method for diagnosing the recurrence of cancer in a subject. Accordingly, the method may comprise determining both the modification status of nucleotide residues and the presence of a genetic mutation in a cfDNA sample from the subject, comparing the nucleotide modification status and genetic mutation profile to the nucleotide modification status and genetic mutation profile of a tumour sample from the subject that has previously been determined using the disclosed method, and diagnosing the recurrence of the cancer in the subject based on the comparison. The cfDNA sample from the subject maybe a blood sample. The tumour sample from the subject may be a sample of a solid tumour, for example previously obtained from the subject in a surgical procedure.

[0441] In some embodiments, the method may be a method for sequencing polynucleotides, the method comprising sequencing the prepared polynucleotide sample to generate a plurality of sequencing reads. In some embodiments, the method may further comprise enriching the second fraction for one or more genome regions of interest prior to sequencing. The method may further comprise comparing the sequences of the sequencing reads to a reference sequence or genome to determine the genomic location of the sequencing reads.

[0442] In some embodiments, the method may be a method for preparing a profile of the modification status of nucleotide residues and any genetic mutations in a reference sequence, such as one or more regions of a genome. Accordingly, the method may comprise sequencing the prepared polynucleotide sample to generate a plurality of sequencing reads, comparing the sequences of the sequencing reads to the reference sequence, such as a reference genome or genomic region, to determine the location, such as the genomic location, of the sequencing reads within the reference sequence and thereby the presence or otherwise of unmodified nucleotide residues and any genetic mutations at specific locations within the reference sequence, such as the one or more regions of the genome.

[0443] For example, the method may be a method for preparing a profile of the modification status at the cytosine C5 position of each CpG dinucleotide and any genetic mutations in a reference sequence, such as one or more regions of a genome. Accordingly, the method may comprise sequencing the prepared polynucleotide sample, to generate a plurality of sequencing reads, comparing the sequences of the sequencing reads to the reference sequence, such as a reference genome or genomic region, to determine the location, such as the genomic location, of the sequencing reads within the reference sequence and thereby the presence or otherwise of unmodified cytosine residues in specific CpG dinucleotides and any genetic mutations within the reference sequence, such as across the genome or across one or more regions of the genome. In another example, the method may be a method for preparing a profile of the modification status of cytosine residues at the N4 position and any genetic mutations in a reference sequence, such as one or more regions of a genome. Accordingly, the method may comprise sequencing the prepared polynucleotide sample to generate a plurality of sequencing reads, comparing the sequences of the sequencing reads to the reference sequence, such as a reference genome or genomic region, to determine the location, such as the genomic location, of the sequencing reads within the reference sequence and thereby the presence or otherwise of cytosine residues unmodified at the N4 position, and any genetic mutations, at specific locations within the reference sequence, such as the across the genome or across one or more regions of the genome.

[0444] In another example, the method may be a method for preparing a profile of the modification status of adenine nucleotides at the N6 position and any genetic mutations in a reference sequence, such as one or more regions of a genome.

[0445] Accordingly, the method may comprise sequencing the prepared polynucleotide sample to generate a plurality of sequencing reads, comparing the sequences of the sequencing reads to the reference sequence, such as a reference genome or genomic region, to determine the location, such as the genomic location, of the sequencing reads within the reference sequence and thereby the presence or otherwise of adenine residues unmodified at the N6 position, and any genetic mutations, at specific locations within the reference sequence, such as the across the genome or across one or more regions of the genome. Sample

[0446] The sample for use in the disclosed method may be obtained from any type of cell or tissue. For example, the sample maybe obtained from tissue, blood, plasma, serum, urine, saliva, stool, cerebrospinal fluid, buccal swab, pleural tap, etc.. The sample may be obtained from tissue. The sample may be obtained from blood. The sample may be a DNA sample. The DNA sample may be a cfDNA sample, which may comprise ctDNA.

[0447] For example, the DNA sample maybe a cfDNA sample from peripheral blood. A “cell- free” sample as used herein, refers to nucleic acids not contained within or otherwise bound to a cell or, remaining in a sample following the removal of intact cells. Cell-free nucleic acids can include, for example, all non-encapsulated nucleic acids sourced from a bodily fluid (e.g., blood, plasma, serum, cerebrospinal fluid, etc.) from a subject. The cfDNA may be released into bodily fluid through secretion or a cell death process. The cfDNA may comprise DNA released into bodily fluid from cancer cells, and may be referred to as comprising circulating tumor DNA (ctDNA). The cfDNA may be released from healthy cells. Methods of preparing samples for use in the disclosed method, such as DNA samples, comprising, for example, extracting and purifying nucleic acids such as DNA from cells or tissues, will be known to the skilled person. Any method that is suitable for preparing a polynucleotide sample for analysis, such as sequencing, may be used. In some embodiments, the sample may comprise fragmented DNA. Fragmentation may be performed using any method used in the analysis of DNA, such as any fragmentation method used in the preparation of a DNA sample for genetic sequencing. For example, the DNA may be fragmented enzymatically, chemically by acoustic shearing, mechanical shearing (example, French pressure cells), sonicating, hydrodynamic shearing or chemically (for example, heat and divalent metal cation).

[0448] In some embodiments, fragmentation of the DNA sample is not required. In such embodiments, the method does not include fragmentation of the DNA sample. For example, in embodiments in which the DNA sample is degraded or fragmented, such as when cfDNA is extracted from blood, the DNA sample may be used directly in the disclosed method. Typically, cfDNA extracted from blood substantially comprises DNA fragment sizes of 50-300 base pairs (bp) in length.

[0449] In some embodiments, the method may comprise the selection of polynucleotides, such as DNA fragments, of a desired length.

[0450] Thus, in some embodiments, the method may comprise, as an initial step, a step of fragmenting the polynucleotide and / or selecting polynucleotides of a desired length. In some embodiments, the method may comprise the use of polynucleotides, such as DNA fragments, substantially or predominantly having a length between 10 and 500 bp, such as between 30 and 400 bp, and preferably between 50 and 300 bp in length. In some embodiments, the method may comprise the use of polynucleotides, such as DNA fragments, substantially or predominantly having a length in the region of between too and 250 bp, preferably between about 150 and 180 bp, to match the sample to the DNA sequencing read length. In some embodiments, the method may comprise the use of polynucleotides corresponding to an amount of DNA in the range of about i fg to about i pg, such as about to fg to about 100 ng, about 100 fg to about 10 ng, about i pg to about i ng. The sample may comprise a quantity of DNA in the picogram range. The sample may comprise less than ipg of DNA, such as less than 500ng, less than loong, or less than long of DNA. Preferably, the sample comprises between ing and loong of DNA.

[0451] In a tenth aspect there is provided a kit for determining both the status of nucleotide residues and the presence of a genetic mutation in a polynucleotide sample, the kit comprising:

[0452] (i) a labelling component suitable for site-specifically labelling a nucleotide residue of a polynucleotide in the sample;

[0453] (ii) a capture agent suitable for binding to the labelling component and for fractionating the polynucleotides into first and second fractions, wherein the first fraction is enriched for labelled polynucleotides, and wherein the second fraction is enriched for polynucleotides that do not comprise a label, optionally wherein the capture agent is tethered to a solid-phase support; and optionally

[0454] (iii) a releasing agent suitable for releasing the polynucleotides from the capture agent.

[0455] Thus, in some embodiments, the kit may comprise:

[0456] (i) a labelling component comprising a methyl-CpG binding domain protein;

[0457] (ii) a capture agent comprising an antibody that specifically binds to methyl-CpG binding domain protein, optionally wherein the capture agent is tethered to a solidphase support such as a magnetic bead or nanoparticle; and optionally

[0458] (iii) a releasing agent suitable for releasing the polynucleotides from the capture agent. In some embodiments, the kit may comprise:

[0459] (i) a labelling component comprising an antibody specific for a modified nucleotide residue. The antibody may be:

[0460] - a high affinity antibody specific for 5-methyl cytosine;

[0461] - a high affinity antibody specific for 5-hydroxymethylcytosine; - a high affinity antibody specific for Nq-methylcytosine;

[0462] - a high affinity antibody specific for N6-methyladenine; - a high affinity antibody specific for 5-methyl cytosine;

[0463] (ii) a capture agent comprising an agent, such as an antibody, that specifically binds to the antibody of the labelling component, optionally wherein the capture agent is tethered to a solid-phase support such as a magnetic bead or nanoparticle; and optionally

[0464] (iii) a releasing agent suitable for releasing the polynucleotides from the capture agent.

[0465] In some embodiments, the kit may comprise: (i) a labelling component comprising a methyltransferase enzyme configured to modify a nucleotide residue in a target position to apply a tag from a cofactor analogue to each unmodified nucleotide residue in a polynucleotide of the sample, and a cofactor analogue comprising the tag precursor, wherein each unmodified nucleotide residue is unmodified in the target position, and an affinity label precursor suitable for binding an affinity label to the tags, wherein the affinity label comprises biotin;

[0466] (ii) a capture agent comprising a biotin-binding protein for fractionating the sequencing library into first and second fractions, wherein the first fraction is enriched for polynucleotides comprising an affinity label, and wherein the second fraction is enriched for polynucleotides lacking an affinity label, optionally wherein the capture agent is tethered to a solid-phase support such as a magnetic bead or nanoparticle; and optionally

[0467] (iii) a releasing agent suitable for releasing the polynucleotides from the capture agent. In some embodiments, the kit may comprise:

[0468] (i) a labelling component comprising a cross-linking agent suitable for covalently binding a DNA-binding protein that is non-covalently bound to the polynucleotide sample to the polynucleotide, optionally an antibody suitable for binding to the DNA- binding protein, and optionally an exonuclease suitable for removing nucleotides that are unshielded by the DNA-binding protein from the polynucleotide;

[0469] (ii) a capture agent comprising an agent, such as an antibody, that specifically binds to the DNA-binding protein or antibody of the labelling component, optionally wherein the capture agent is tethered to a solid-phase support such as a magnetic bead or nanoparticle; and optionally (iii) a releasing agent suitable for releasing the polynucleotides from the capture agent. All features described herein (including any accompanying claims and drawings), and / or all of the steps of any method or process so disclosed, may be combined with any of the above aspects in any combination, except combinations where at least some of such features and / or steps are mutually exclusive.

[0470] The invention will now be illustrated by reference to specific Examples showing how embodiments may be carried into effect, which are not intended to be limiting. Data from the Examples is presented in the Figures, in which:

[0471] Figure 1A is a flow chart showing an overview of one embodiment of the disclosed method for preparing a fractionated amplified sequencing library for determining both the status of nucleotide residues and the presence of a genetic mutation in a polynucleotide sample.

[0472] Figure 1B is a diagram outlining the method of Figure 1A. In the top row, an amplified sequencing library is prepared comprising the steps of

[0473] - site-specifically labelling target nucleotides (e.g. methylated or unmethylated CpG residues) (top row left); - preparing the nucleotides into a sequencing library (top row centre); and

[0474] - incorporating a first indexing barcode at a first end but not a second end of each polynucleotide (top row right).

[0475] The result (shown middle row middle) is a sequencing library comprising polynucleotides having an indexing barcode at a first end but not a second end, including a subset of polynucleotides that comprise a label bound site-specifically to a nucleotide residue. The amplified sequencing library is then fractionated using the label (middle row) to provide a first fraction enriched for polynucleotides comprising a site- specifically bound label (middle row left) and a second fraction enriched for polynucleotides that do not comprise a site-specifically bound label (middle row right). A second indexing barcode is incorporated at a second end of the polynucleotides of the first fraction (bottom row left), and a third indexing barcode is incorporated at a second end of the polynucleotides of the second fraction (bottom row right). Optionally, target enrichment of the second fraction may be performed. The first and second fractions are then pooled and sequenced, for example, in a single sequencing run. Figure 1C is a diagram showing a method for preparation of polynucleotides for sequencing in accordance with one embodiment: i. Labelled (i.e. with a “modification”) polynucleotides are ligated with short, y- shaped adapters containing a unique molecular index (MID); 2. PCR (for example, up to about five cycles) is performed with a single primer composed of a priming region complementary to the y-shaped adapter, a first barcoding index (BC1), and additional priming regions for amplification using standard sequencing primers and during the sequencing experiment;

[0476] 3. The labelled polynucleotides are enriched, creating two pools that are enriched for labelled and unlabelled DNA molecules, respectively;

[0477] 4. A second PCR is performed to attach the second barcoding index (BC2 or BC3) and to amplify the polynucleotides;

[0478] 5. The fractions are then pooled and sequenced. Figure 1D is a flow chart and Figure 1E is a corresponding diagram showing one embodiment of a method for labelling polynucleotides, in this case unmethylated CpG, for use in the disclosed method for determining both the status of nucleotide residues and the presence of a genetic mutation in a polynucleotide sample. A polynucleotide sample, such as purified fragmented DNA (cfDNA or fragmented genomic DNA) is derivatised by treatment with a CpG-targeting methyltransferase (in this case M.Mpel) and a synthetic cofactor analogue (in this case ETA-AdoHcy-N3) that results in the addition of tags at unmodified CpG sites. The methyltransferase is inactivated and the DNA fragments are then end-repaired and ligated to sequencing adapters. Prior to, during, or after the polynucleotides have been labelled, the polynucleotides are amplified and a first indexing barcode is incorporated at a first end but not a second end of each polynucleotide to form the amplified sequencing library (not shown). Polynucleotides comprising a tag are subsequently labelled by attaching an affinity label (for example, biotin) to the tag, and isolated (for example, using streptavidin- coated magnetic beads). The resulting fractions, enriched for tagged (unmodified CpG sites) and untagged (predominantly smCpG and shmCpG sites) polynucleotides, can be separately amplified to incorporate indexing barcodes at a second end of polynucleotide in each fraction (not shown). The fractions can then be pooled and sequenced.

[0479] Figure 2A is a graph showing sequencing adapter ligation efficiencies using unlabelled DNA (control; left-hand cluster), methyltransferase-labelled DNA without methyltransferase inactivation (centre cluster), and methyltransferase-labelled DNA with methyltransferase inactivation (right-hand cluster), prior to adapter ligation. Figure 2B shows the raw data for this experiment.

[0480] Figure 3A is a bar chart showing efficiencies for capture (light bars) and capture / release (dark bars) of target DNA from solution, as a function of target CpG site density.

[0481] Figures 3B and 3C are reproduced from Kriukiene et al. (Nature Communications 20134:2190).

[0482] Figure 3D shows the enrichment efficiency of the present method for a (target) DNA molecule with a high density of CpG sites (10 sites per -150 bp). The target DNA is mixed with 24 ng of non-target DNA and is selectively purified with high efficiency at a range of concentrations. Final DNA concentration was quantified using spectrophotometry (Qubit).

[0483] Figure 3E shows enrichment of unmethylated DNA as a function CpG density using three different enrichment chemistries. The current approach, using a single pot reaction and one purification step (dark grey bars) shows over three times improvement in the retention of DNA throughout the Tag-Seq enrichment process, compared to the approach described by Kruikiene et al. (light grey) and an approach using two DNA purification steps (grey). Density of CpG sites is shown as number of sites per 300 bp genomic window. Figure 4 is a graph showing threshold cycle versus target DNA concentration (ng) showing a linear response from 1.25 ng target DNA down to 1.25 pg target in a background of 24 ng (over 19OOOX excess) of DNA containing no target sites.

[0484] Figure 5 is a bar chart showing a comparison of sequencing coverage at CpG sites in the enriched (unmodified CpG) fraction (light grey) of a DNA sample, and the unenriched (modified CpG) fraction (dark grey) of the sample. Note that 1.4% and 44.7% of reads did not contain a CpG site for the enriched and unenriched fractions, respectively. Figures 6A and 6B are bar charts showing sequencing coverage of enriched (light grey) and unenriched (dark grey) fractions at unmodified CpG sites (Figure 6A) and methylated CpG sites (Figure 6B). Note the different y-axis scales in the two plots. Figure 7A is a bar chart showing enrichment using the disclosed method (NRPM) across a range of unmodified CpG densities (75 bp window) (groups 1-4, light grey bars) compared to similar enrichment using the known MeDIP-Seq method (group 5, hashed bars), with baseline for whole-genome sequencing provided for context (group 6, white bars).

[0485] Figure 7B is a bar chart corresponding to Figure 7A showing the corresponding enrichment profiles for methylated CpG sites.

[0486] Figure 8A shows example profiles obtained using the disclosed method (dark blue) compared to MeDIP-Seq (green) and WGBS (yellow) profiles for (top) the KRAS gene and (bottom) a megabase-scale region of chromosome 1.

[0487] Figure 8B shows the enrichment of read counts for two technical replicates of the profile obtained using the disclosed method (blue) MeDIP-Seq (yellow) and shallow whole genome sequencing (red) at gene transcription start sites (TSS).

[0488] Figure 8C shows the profile obtained using the disclosed method (blue) correlates inversely with the MeDIP-Seq profile (purple) and with chromatin domain organisation identified in Hi-C experiments (red) on the megabase scale.

[0489] Figure 9A shows enrichment of genomic DNA at regions corresponding to H3K4 monomethylation. Comparison of two experiments using the disclosed method (technical repeats, different users) (light and dark blue), MeDIP-Seq (green) and shallow whole genome sequencing (control, no enrichment) (orange).

[0490] Figure 9B shows enrichment of genomic DNA at regions corresponding to H3K4 trimethylation. Comparison of two experiments using the disclosed method (technical repeats, different users) (light and dark blue), MeDIP-Seq (green) and shallow whole genome sequencing (control, no enrichment) (orange). Figure 9C shows enrichment of genomic DNA at regions corresponding to H3K27 acetylation. Comparison of two experiments using the disclosed method (technical repeats, different users) (light and dark blue), MeDIP-Seq (green) and shallow whole genome sequencing (control, no enrichment) (orange).

[0491] Figure 10 shows a Spearmann correlation analysis to assess the similarity of the profiles obtained using the disclosed method across six different cell lines and for three technical repeats of each sample. Dark blue indicates a high degree of similarity of the profiles. Notably, DNA from each cell line has a clearly distinct profile when compared to the sample technical repeats, consistent with the known utility of methylation profiles for the identification of tissues.

[0492] Figure 11 shows volcano plots showing the comparison of tumour and normal adjacent tissue profiles obtained using the disclosed method, across the genome for six patients with a range of different cancers. Differentially methylated windows are defined as those with an adjusted p-value of less than 5% and a log-fold change in signal of greater than 0.58 (1.5X). Red lines indicate the locations of these thresholds in the volcano plots. Blue markers are windows that show hypermethylation in cancer, red markers are for windows that are hypomethylated in cancer.

[0493] Figures 12A and 12B show DNA Agilent TapeStation (Cell-free DNA ScreenTape) traces showing profiles for the cell free DNA (cfDNA) that was input for the disclosed profiling method (top) and the output from the enriched, amplified libraries (bottom) of the disclosed method, for a healthy patient sample (Figure 12A) and for a sample from a patient with Stage 1 non-small cell lung cancer (Figure 12B). The mono- and dinucleosomal pattern of the input cfDNA is maintained in the final libraries, the size of which corresponds to the duplicated original strand plus the Illumina P5 / P7 sequencing adapters. Figure 13 shows example profiles using the disclosed method for a genomic region (SHOX2 gene, a known methylation biomarker for lung cancer) in the healthy (blue) and lung cancer (grey) patients, compared to genomic DNA, extracted from the healthy patient’s buffy coat (yellow). Black traces show the duplicate profiles in both cases, dark blue tick marks show the known CpG site density across the gene. Traces are based on normalised read counts for all profiles and the cfDNA samples are displayed on the same scale for direct comparison. Figure 14 shows a summary of data derived from triplicate repeat experiments using DNA isolated from FFPE (formalin-fixed, paraffin-embedded) samples. A) Enrichment (normalised read count) as a function of CpG site density (75 bp windows) showing steadily increasing levels of enrichment with increasing CpG density. B) Plots showing normalised read counts across the APC gene transcription start site. Consistent with the enrichment profiles in (A), enrichment of DNA at the CpG-dense gene promoter is more marked for Patient A (pink) than Patient B (blue) than Patient C (green). Profile of the HT-29 cell line (colorectal cancer) shown in grey for comparison. C) PCA plot showing excellent consistency of the technical replicates of these samples.

[0494] Figure 15 shows an example profile for a genomic region (ALK gene) produced using the disclosed method. A polynucleotide sample was prepared into first and second fractions as described. The top row shows an epigenetic profile produced from the first fraction, and the bottom row shows a profile produced from the second fraction of DNA enriched using a targeted sequencing panel.

[0495] Figure 16 shows the output of the sequencing experiment of Example 12, for four differentially indexed samples, sequenced using the same sequencing run. Lanes 1-4 represent samples 1-4.

[0496] Figure 17 shows the output of a sequencing experiment of Example 13, for differentially indexed sample number 4 (as an example of the six samples used in the experiment), sequenced using the same sequencing run. Lane 1: Sample 4 gDNA Genome pool (second fraction); Lane 2: Sample 4 gDNA Enriched pool (first fraction); Lane 3: Sample 4 cfDNA Genome pool (second fraction); and Lane 4: Sample 4 cfDNA Enriched pool (first fraction).

[0497] Figure 18 shows the output of the sequencing experiment of Example 14 for one differentially indexed sample, sequenced using the same sequencing run. The two different pools (fractions) are shown, Lane 1: Sample 1 gDNA Genome pool (second fraction); and Lane 2: Sample 1 gDNA Enriched pool (first fraction).

[0498] Figure 19 shows the variant allele fraction called (Measured VAF) for a genomic DNA standard (Oncospan gDNA) for the unenriched, amplified fraction of the genome, compared with the known variant allele fraction for mutations of this sample. The measured VAF is directly proportional to the known VAF, indicating no skewing of the variant allele fraction as a result of the sample processing.

[0499] Examples Performing multiple tests on a single patient sample can improve the sensitivity and specificity of a diagnostic test. However, this is typically challenging, particularly using liquid biopsy, where the amount of available genomic material can be severely limited (less than to ng per millilitre of blood). To overcome this, the disclosed method provides a flexible approach to sample indexing for next-generation sequencing. Single indexing barcodes are added to the sample in discrete preparation steps. By doing so, multiple analytical approaches may be applied in parallel to a single sample, such as, for example, both whole-genome amplification of the sample, and targeted enrichment of, for example, unmethylated genomic material. The readout of these analytical approaches may be obtained in a single, pooled sequencing experiment. Reads can be readily de-duplexed, post-sequencing, enabling simultaneous and comparative analysis of the same sample.

[0500] To allow a single sample to be characterised, with a multiplexed sequencing experiment, the disclosed method employs a novel solution that adds distinct molecular indexes to the sample at three different stages of the workflow. This results in uniquely indexed DNA for two sample types, as shown in Figure 1A.

[0501] As shown in Figure 1B, in a first step (shown in the top row), polynucleotides are site- specifically labelled, prepared into a sequencing library (comprising, for example, end repair, A-tailing, adapter ligation) and amplified by PCR using a single indexing primer.

[0502] This approach creates a pool of polynucleotides having an indexing barcode at a first end but not a second end. The pool of polynucleotides comprising an unmodified copy of the whole sample (such as a whole genome), site-specifically labelled polynucleotides, and polynucleotides derived from the original (e.g. genomic) sample (carrying native modifications such as methylation).

[0503] The labelled, amplified sequencing library is then enriched for target polynucleotides (shown in the middle row), using one or more enrichment approaches, thereby creating at least two fractions of polynucleotides; a first fraction that is enriched for the labelled polynucleotides, and a second fraction that is enriched for unlabelled polynucleotides and represents an unenriched copy of the sample as a whole. The different fractions are then amplified using PCR with a second indexing primer (shown in the bottom row), thereby incorporating different indexing barcodes to the second end of the polynucleotides in each fraction. The different fractions may be pooled and multiplexed in a single sequencing run and the whole sample (e.g. whole genome) copy can be used for genetic analysis (either by whole genome sequencing or by targeted enrichment using a panel of bait oligonucleotides to enrich e.g. the exome).

[0504] The treatment of the polynucleotides in an embodiment of the disclosed method is shown in Figure 1C.

[0505] This approach maybe applied for simultaneous epigenetic and genetic profiling of the same sample. In some embodiments, site-specific labelling may comprise the addition of a tag to unmodified nucleotides, such as unmethylated CpG dinucleotides, using a methyltransferase. An example of a suitable method is shown in Figures 1D and 1E. This approach, such as shown in Figures 1D and 1E, may be used to provide a uniquely straightforward and robust approach to epigenetic profiling that provides concurrent or simultaneous readout of genetic features of the genome of interest. The method is an enzymatic technology that enriches for unmodified nucleotide residues such as, in particular, CpG sites, across the polynucleotide sample, such as across the whole genome. The method requires no de novo knowledge of the polynucleotide sequence for its application in epigenetic profiling. It does not damage the nucleic acid sample nor rely on base conversion and therefore allows concurrent analysis of other genetic features, such as mutations. The method may also provide a profile whose signal correlates with markers of active genomic regions. The enzymatic chemistry enables unbiased fractionation of modified and unmodified nucleic acids from a sample for subsequent analysis.

[0506] The Examples below, and Figures 2-16, show that the epigenetic aspect of the disclosed method is a uniquely sensitive approach, relative to other available methods for epigenetic analysis. The method can be applied at polynucleotide (e.g. DNA) concentrations that are compatible with single-cell analysis (picogram inputs). The workflow also enables concurrent readout of a genome’s genetic and epigenetic features. As a result of the simplicity of the approach, comprising a single enzymatic step, followed by fractionation using a capture probe, technical repeats of the experiment show excellent consistency. Enrichment of unmodified nucleic acids leads to an epigenetic profile that achieves saturation at around 70M, 150 bp reads. Studies in cell lines and in patient tissue samples demonstrate the potential of the disclosed method as a platform for the diagnosis of disease and the identification of tissue of origin in a sample, for example. The underlying chemistry requires no a priori assumptions to be made about the sample, making the platform ideally suited as a research tool, for example, for the discovery of novel biomarkers of disease.

[0507] Enrichment of polynucleotides using antibodies towards modified nucleotide residues (such as antibodies specific for 5-methylcytosine) or DNA binding proteins significantly reduces the cost of DNA sequencing, relative to an untargeted, whole-genome approach. Moreover, these approaches require no a priori assumptions to be made about the sample, making them well suited as research tools, for example, for the discovery of novel biomarkers of disease.

[0508] A flow chart and diagram providing an overview of part of an embodiment of the disclosed method comprising labelling unmodified nucleotides is shown in Figures 1D and 1E. In the embodiment shown, a bacterial DNA methyltransferase enzyme (M.Mpel) is used to target unmodified CpG sites for modification with an unnatural cofactor analogue of S-adenosyl-L-methionine referred to herein as ETA-AdoHcy-N3. This approach comprising the use of a methyltransferase to apply a tag to unmodified nucleotides in a target position is demonstrated in Examples 1-12 below. The use of this approach in the disclosed method in which the labelled polynucleotides are subsequently fractionated using the tag, amplified, and pooled prior to sequencing is demonstrated in Example 14. Incubation of the methyltransferase, DNA and cofactor for one-hour results in complete modification of the target DNA, which is functionalised with azide- terminating tags. These tags can be further modified (for example, biotinylated) to enable fractionation of modified and unmodified DNA (where ‘unmodified’ refers to all genomic DNA fragments containing one or more CpG dinucleotide that is unmodified (for example, that is not methylated, hydroxymethylated, carboxylated, or acylated) at the C5-position). The inventors have developed a modified library preparation that integrates this labelling step, thereby minimising handling and purification steps, and as a result, improving robustness and maximising sensitivity.

[0509] Example 1

[0510] A significant advantage of the disclosed method is that it employs a single clean-up step for the entire process, dramatically improving the efficiency of the fractionation. This is made possible by the use of, firstly, a step to inactivate the methyltransferase and, secondly, the surprising activity of the enzymes for library preparation in the resulting buffer mixtures.

[0511] Figure 2 shows that the efficiency of adapter ligation is significantly inhibited in the absence of inactivation of the methyltransferase after labelling and before adapter ligation. This surprising result is due to the high binding affinity of the methyltransferase enzyme to the DNA molecule, which has been found to limit the activity of DNA-targeting enzymes in subsequent steps of the procedure. This activity can be recovered by inactivating the methyltransferase enzyme. Example 2

[0512] Figure 3 shows the results of example experiments investigating the recovery of DNA samples with different affinity labels and comprising different CpG densities.

[0513] In the experiment of Figure 3A, a mixture containing too ng of DNA cariying a known number of CpG sites (o, 1, 2, 4 or 10) was incubated with M.Mpel (0.0274 pg / L) and

[0514] ETA-AdoHCy-N3 (too pM) . The reaction was incubated at 37 °C for 1 hour. The DNA was purified using AMPure beads (Beckman Coulter), followed by conjugation of an affinity label comprising biotin using click chemistry. Finally, DNA was purified using a standard PCR clean-up kit (Zymo Clean and Concentrate).

[0515] Purified DNA was fractionated using DynaBeads MyOne Streptavidin-coated beads.

[0516] The beads were then washed twice with 150 pL of PBST. Finally, captured DNA was released. As shown in Figure 3A, capture efficiency is improved by the current method (grey bars), relative to the method reported by Kriukiene et al. (Nature Communications 20134:2190), Figure 3C. Figure 3A (blue bars) shows the release efficiency of DNA in the current workflow (Active-Seq).

[0517] The disclosed method provides a significant improvement in capture efficiency relative to the method described in Kriukiene et al. (Nature Communications 20134:2190). As shown in Figure 3B (reproduced from Figure 2b of Kriukiene et al.), Kriukiene at el. report capture efficiencies in the 20-30% range using a method comprising an azide- DBCO label. This is significantly lower than the capture efficiencies that maybe obtained using the method disclosed herein, as shown, for example, in Figure 3A. The method described by Kriukiene et al. shows (in Figure 2c, reproduced herein as Figure 3C) capture of around 30-40% of target DNA containing 2 CG sites using an azide-DBCO affinity label and streptavidin-coated magnetic beads. In contrast, while the method disclosed herein is able to isolate DNA at similar input levels to those of the method described by Kriukiene et al., the captured DNA maybe recovered from the capture agent in much more significant proportions, at least in part due to the efficient release of the labelled DNA molecules (see Figure 3D). Kriukiene et al. only includes data on the level of DNA capture, and there is no discussion or data in Kriukiene et al., on the efficiency of release of the sample from the magnetic beads. The present inventors have found that using the method disclosed by Kriukiene et al., the release of enriched DNA fragments is highly inefficient and inconsistently reproducible.

[0518] The efficient enrichment of DNA that is rich in CpG sites is critical for even representation of the (enriched) genome in the sequencing experiment. CpG-rich regions often lie in important regulatory regions of the genome. Figure 3E shows a plot of mean normalised read count per million reads (NRPM) for samples prepared using the method described in Kriukiene et al. (left-hand bars) or the method disclosed herein (right-hand bars), as a function of the number of CpG sites in a given read. The plot is generated for CpG-rich regions of the genome (CG islands). Figure 3E clearly shows higher read densities across CpG rich regions, demonstrating the significantly improved enrichment of CpG-rich DNA using the disclosed method. The overall effect of the method disclosed herein is to enable efficient enrichment of unmodified DNA from as little as a few picograms of input DNA. This is particularly critical for samples where the DNA concentration is limited, such as liquid biopsy (blood, urine, saliva, spinal fluid) samples.

[0519] Example 3

[0520] Experiments were performed to establish how the enrichment / fractionation platform performs as a function of DNA concentration, at input amounts consistent with cfDNA and single cell analyses; and the linearity of the enrichment efficiency across DNA molecules with a range of (unmodified) CpG site densities. The inventors have found that this latter issue was a particular limitation of the method described by Kriukiene et al., which resorted to dilution of the methyltransferase enzyme in the labelling reaction to limit the number of (relatively insoluble) DNA modifications introduced to a single DNA molecule.

[0521] Between 1.25 ng and 1.25 pg (equivalent to between 200 and 1 / 3 of a copy of the human genome) of target DNA (153 bp, containing 10 unmodified CpG sites) were spiked into a background of 24 ng of non-target DNA (142 bp containing no CpG sites).

[0522] The target DNA was tagged and thereby enriched using streptavidin coated beads for analysis by qPCR, the results of which are shown in Figure 4.

[0523] Enrichment efficiencies in excess of 80% were obtained for all of the spike-in samples, clearly demonstrating the compatibility of the approach for enrichment of DNA at input levels consistent with single-cell analysis. Tagged DNA is compatible with PCR and can be amplified using a standard polymerase, following enrichment.

[0524] As shown in Figure 3A, the initial step of enrichment (capture of labelled DNA, for example, by streptavidin-coated beads) shows only a very minor dependence on the number of CpG sites available on a DNA molecule. Light grey bars show capture efficiencies and dark grey bars show capture / release of target DNA from solution.

[0525] Example 4 Having demonstrated the performance of the biochemical approach on simple DNA fragments, the utility of the platform disclosed herein on genomic DNA was investigated by generating genome-wide epigenetic profiles from DNA extracted from a range of cell lines. Extracted DNA was fragmented by sonication (-150 bp) and subject to enrichment of the DNA fragments lacking CpG modification. Samples were sequenced using an Illumina NovaSeq platform (Source Biosciences) to approximately 120M reads per sample. For the enriched, unmodified DNA fraction, saturation analysis shows that data reaches 90% saturation between 65M and 90M reads. Initial quality control using MultiQC showed low levels of read duplicates (—15%) and an average enrichment in GC-content of the genome, from 41% in the nascent human genome to an average of ~47% in enriched samples, consistent with enrichment at regions rich in CpG dinucleotides, such as CG islands (lung cancer derived cell lines showed higher average GC content, with a mean of approximately 54%).

[0526] Example 5 Successful enrichment at CpG sites was assessed by comparison of the enriched

[0527] (unmodified CpG) and unenriched (modified CpG) fractions of the genome by sequencing. This was done by examining the fraction of reads containing a CpG site and the sequencing coverage at each CpG site. In the enriched fraction, 98.6% of the reads contain a CpG site. By contrast, in the unenriched fraction only 55.3% of reads contain a CpG site, indicating effective enrichment at CpG sites. Furthermore, in the enriched fraction, a majority of the CpG sites of the genome (54.0%) are covered by greater than 5 reads, whereas in the unenriched fraction, this figure is just 6.3% of the CpG sites.

[0528] In the human genome, 70-80% of the CpG sites are modified (e.g. methylated) and, hence, in the disclosed method the enriched fraction of the sample might be expected to be focussed only on the remaining 20-30% of the genome’s CpG sites. Despite this, over half of the genomic CpG sites were found to have ‘high’ coverage (> 5-fold) in the enriched (unmodified) fraction (light grey bars), as shown in Figure 5. This is likely due, at least in part, to the high efficiency of the disclosed method.

[0529] For enriched DNA, typically between 1 and 5% of reads were found to contain no CpG sites. The source of these reads is likely varied but will include non-specifically enriched DNA, as well as reads that do not cover a motif but that originate from a molecule that does (specifically enriched but CpG not sequenced). This read fraction is denoted as the ‘background’ for the enriched sample. To further understand the composition of the enriched DNA fraction, sequencing coverage of CpG sites was compared at known modified and unmodified CpG sites (as determined by whole genome bisulfite sequencing (WGBS)). The results are shown in Figures 6A and 6B. As defined herein, an unmodified (e.g. ‘unmethylated’) site is a site having a modification (e.g. methylation) level (P-value) of less than 0.05 by whole genome bisulfite sequencing. A ‘modified’ (e.g. ‘methylated’) site has a modification (e.g. methylation) level (P-value) of greater than 0.95 by whole genome bisulfite sequencing. A total of 430,245 CpG sites met the definition of an ‘unmodified CpG site’ (< 5 % modified by WGBS). Significant enrichment was observed at these sites, as judged by the high coverage of sites in the enriched fraction (70% of sites (~3oi,ooo CpG sites) have greater than 5-fold coverage) as compared to the unenriched fraction of the sample (less than 1% (-1500 sites) have greater than 5-fold coverage). This is in good agreement with the initial validation of the approach, confirming that where CpG sites are unmodified, efficient enrichment is seen using the disclosed method.

[0530] There are 53,081 CpG sites in the genome defined as ‘modified’ (> 95 % modified by WGBS). Similar coverage of these sites is observed in both enriched and unenriched fractions of the sample, indicating that little enrichment of DNA occurs at these highly modified sites. Hence, where CpG sites are modified, little enrichment of these sites is seen using the disclosed method.

[0531] Example 6 As further validation of the disclosed method, the inventors sought to understand the enrichment of genomic DNA, as a function of unmodified (as determined by whole genome bisulfite sequencing) CpG site density. This analysis mirrors the validation experiments using DNA molecules of known sequence and CpG site density (as shown in Figure 4) but for enrichment using genomic DNA. The results are shown in Figures A and 7B.

[0532] A remarkably similar enrichment profile was observed for genomic DNA, as for the model DNA fragments with known CpG site densities with efficient enrichment of DNA, even where only one or two unmodified CpG sites are available. The enrichment profile for the disclosed method shows consistency across the range of unmodified site densities (Figure 7A). This is in stark contrast to the analogous experiment for MeDIP- Seq for highly modified DNA molecules, Figure 7B. For MeDIP-Seq, no significant enrichment of DNA molecules was observed (relative to the WGS baseline) with a CpG density of less than 3 sites in a 75 bp genomic window; and a significant bias towards enrichment of densely modified regions of the genome (> 6 CpG sites per 75 bp).

[0533] In all, these results are consistent with the initial validation of the disclosed approach, which demonstrates exceptionally high enrichment efficiencies for unmodified CpG sites; near uniform efficiency across a range of CpG densities and no significant off- target enrichment of DNA in the disclosed method.

[0534] Example 7

[0535] A number of studies were conducted to compare the disclosed method to other (epi)genomic analyses. An advantage of the disclosed method is provided by the enzymatic targeting of unmodified CpG sites. The approach is well-suited to the enrichment of hypomethylated DNA from tumour cells in the blood. A key genomic feature that were hypothesised to be prominent in the profiles produced by the disclosed method are extended regions of unmodified DNA, which are epigenetically-stable and conserved in mammals, with consistently low unmodified levels on length scales of 5-20 kbp. The term ‘non-modified island’ (NMI) is used herein since such unmodified regions are rather more island-like in the profiles produced by the disclosed method.

[0536] Read count peaks in profiles of unmodified DNA produced by the disclosed method (bottom) were found to anticorrelate to those observed in MeDIP-Seq (top) and correlate with regions of low modification, identified in whole genome bisulfite sequencing (yellow), as shown in Figure 8.

[0537] At the gene-level, read counts of unmodified DNA using the disclosed method show peaks centred at CG islands, that span the broader, regulatory regions of genes and are consistent with NMIs, as shown in Figure 8A. NMIs are thought to exist to reduce mutation rates in functionally-important genomic regions.

[0538] Deamination of methylated cytosine, which converts to thymine, has been shown to be the most frequent mutation in human cancers. NMIs play a central role in regulation of gene expression and their methylation levels are regulated by the TET (demethylating) enzymes, via the polycomb protein complex. In agreement with this, genome-wide analysis shows significant enrichment of unmodified DNA around transcription start sites, relative to the analogous MeDIP-Seq experiment, as shown in Figure 8A. On the scale of hundreds of kbp-to-Mbp, anticorrelation of the profile of unmodified DNA produced by the disclosed method (bottom) to both MeDIP-Seq (top) and WGBS (middle) is retained, as shown in Figure 8A. A particularly striking aspect of the profile produced by the disclosed method (bottom) is the presence of clear domains of modified and unmodified genomic regions, that are consistent with the expected correlation between genomic modification (e.g. methylation) levels and genome organisation (Figure 8A).

[0539] Example 8

[0540] For further validation, the disclosed method was compared to established sequencing approaches that are known to correlate with DNA methylation levels, as shown in Figure 8. Clear regions of highly unmethylated DNA, also evident in the MeDIP-Seq and WGBS profiles (Figure 8A). Enrichment in the profile produced with the disclosed method at transcription start sites and markers of active chromatin (HsKqMei, HsKqMes and H3K27ac) anticorrelates with loss of MeDIP-Seq signal at these regions (Figure 8B and Figure 9).

[0541] These regions of low / high methylation are correlated with defined structural domains of the genome identified in Hi-C experiments, as shown in Figure 8C. Example 9

[0542] Having demonstrated the potential of the disclosed method to generate meaningful genome-wide epigenetic profiles, the approach was applied to nine DNA samples derived from cultured cell lines for a range of cancers (breast (MCF7, HCC1937), colorectal (HT29, SW48, C0I0201, RKO), liver (HepG2) and lung (SW1271, NCI- H2170)).

[0543] Genome-wide correlation analysis of this dataset shows excellent correlation of a series of three technical repeats for each of the samples. Each of the cell lines examined forms a distinct cluster of correlated data, with cell lines from similar tissues broadly clustering together, consistent with the expectation that the epigenetic profile can be employed for the identification of tissue of origin for a sample. These distinct cell-line- specific profiles result in part from the robustness of the method and the remarkable consistency of the disclosed method across sequencing runs and operators.

[0544] In order to establish the potential of the disclosed method as a method for the diagnosis of cancer, a series of experiments were performed using tumour tissue and normal adjacent tissue from six patients with different cancers. Here, whole genome correlation analysis provides a simple overview that clearly highlights the discriminative ability of the disclosed method for both disease diagnosis and derivation of tissue of origin from a patient sample (Figure to).

[0545] This approach was extended to better understand the specific regions of the genome that give rise to differences in epigenetic profiles and their link to known biological function. The profiles produced by the disclosed method of tumour and normal adjacent tissue were compared for a patient in 75 bp windows, across the whole genome. This analysis requires no a priori knowledge of a patient’s genomic sequence and makes no assumptions about regions of interest. By doing so, tens or hundreds of thousands of differentially modified (e.g. methylated) regions were identified in the profiles produced by the disclosed method for each patient. These are summarised in a series of volcano plots, shown in Figure 12. The results show the relative statistical significance and, critically, the population of modified (e.g. hypo- and hypermethylated) regions identified in each patient (the disclosed method does not solely focus on discovery of unmodified regions of the genome). Examples of individual profiles produced using the disclosed method for the tissue lung cancer sample are shown at genes that have been implicated in relevant cancer pathways in the literature.

[0546] Example 10

[0547] DNA shed from tumour cells can be isolated in the blood of cancer patients. However, the technical challenge associated with its analysis is two-fold; DNA it is typically present in healthy and early-stage cancer patients at less than 10 ng per millilitre of plasma; and cell free DNA isolated from plasma can contain less than 1% tumour fraction (ctDNA).

[0548] The disclosed method is ideally suited to the analysis of ctDNA because it is performant with input DNA orders of magnitude less than one nanogram and provides genome- wide analysis. To demonstrate the suitability of the disclosed method for the analysis of ctDNA, two patient samples were prepared for analysis, one a healthy patient and one from a patient diagnosed with stage 1 lung cancer (non-small cell lung cancer). DNA was extracted from 3mL of plasma using an automated platform (Informed Genomics), which returned 4?uL of DNA at a concentration of 0.50 ng / pL and 0.55 ng / pL for the healthy and cancer patient, respectively (Figure 12A and 12B).

[0549] The method was performed in duplicate (separate preparations and sequencing runs) with 8.9 ng and 7.2 ng input DNA for the healthy patient and 10 ng input on both occasions for the lung cancer patient. Both input DNA samples and the output of the enriched library maintain the fragment size distribution that is characteristic of nucleosomal cell-free DNA, Figure 12.

[0550] Duplication rates for the sequencing data were 9.2% and 9% for the larger sequencing run, with coverage of 100M reads for both samples (Figure 13). This represents a significant improvement on duplication rates typical for approaches using base conversion, which can reach 30-40%. The background, defined as the percentage of reads lacking a CpG site in the dataset, was 4.5% and 3.6% for the healthy and cancer samples, respectively.

[0551] Example 11

[0552] Formalin fixed, paraffin-embedded (FFPE) treatment typically leads to extensive damage (depurination, depyrimidation and deamination) of the genome. The ability to generate a meaningful epigenomic profile from DNA preserved in these samples using the disclosed method was investigated.

[0553] Genome-wide profiles were generated using the disclosed method for three FFPE embedded samples in triplicate, sourced from the Welsh Cancer Bank, derived from patients with colorectal cancer. Consistent with other sample types, sequencing reached 90% saturation by 80M (150 bp, paired-end) reads for all samples. The resultant datasets show good overall coverage of the genome and excellent consistency for the technical repeats (Figure 14).

[0554] For two of the three samples, high levels of relative enrichment of CpG dense regions of the genome were observed, (Figure 14). However, comparison of the three FFPE datasets to similar data for the HT-29 cell line shows good consistency of the profile and the observed ‘background’ of the sequencing dataset (reads lacking CpG sites) is consistently below 5% for all samples. Such an increase in the relative enrichment of regions that are dense in CpG sites is consistent with the expected damage of CpG sites by the FFPE treatment. This likely stems from a relative reduction of the concentration of enrichable DNA molecules with few CpG sites in the sample. Conversely, those molecules with many CpG sites retain a few taggable CpG sites, post FFPE treatment.

[0555] Example 12

[0556] In DNA samples requiring DNA fragmentation, DNA was sheared to an average of 180 bp.

[0557] A mixture of DNA (<iong), M.Mpel and ETA-AdoHCy-N3 was prepared on ice. This solution was incubated at 37 °C for 1 hour. Following incubation, the methyltransferase enzyme was inactivated by heating.

[0558] Without purification, the sample was cooled to io°C. End Repair & A-Tailing Master Mix was added (Kapa Biosystems). The mixture was mixed thoroughly by pipette aspiration and incubated at 20°C for 30 mins followed by a 65°C incubation for a further 30 mins.

[0559] The sample was cooled to io°C, and sequencing adapters were ligated.

[0560] Without purification, biotin-PEGq-DBCO (Jena Biosciences) was added and the mixture was incubated at 37 °C for 1 hour with shaking at 500 rpm. The DNA was subsequently purified from the reaction mixture.

[0561] The purified DNA was amplified using a standard PCR master mix. Separately, 5 pL Dynabeads MyOne Streptavidin Ci beads (ThermoFisher) were washed with 150 pL of PBST. The amplified DNA was added to the beads and the mixture was incubated at 23 °C for 15 minutes, with shaking at 1000 rpm. Once completed, the supernatant was removed (as the “second fraction”), and the beads were washed twice with 150 pL of PBST. Finally, the bound DNA was released from the beads (as the “first fraction”) by denaturation of streptavidin. The second fraction was enriched for target sequences using a targeted genomic panel

[0562] (Agilent SureSelect CGP), according to the manufacturer’s instructions. DNA libraries from fraction one and fraction two were pooled separately, along with 0.1% PhiX and both pools were sequenced using a S4 Flow Cell using an Illumina NovaSeq Sequencer.

[0563] An example genomic profile from the sequenced datasets for fraction one and fraction two can be found in Figure 15.

[0564] Materials and Methods A typical workflow for embodiments comprising the use of a methyltransferase to apply a tag to unmodified nucleotides in a target position is set out below.

[0565] In DNA samples requiring DNA fragmentation, DNA was sheared to an average of 180 bp.

[0566] A mixture of DNA (<iong), M.Mpel and ETA-AdoHCy-N3 was prepared on ice. This solution was incubated at 37 °C for 1 hour. Following incubation, the methyltransferase enzyme was inactivated by heating. Without purification, the sample was cooled to io°C. End Repair & A-Tailing Master Mix was added (Kapa Biosystems). The mixture was mixed thoroughly by pipette aspiration and incubated at 20°C for 30 mins followed by a 65°C incubation for a further 30 mins. The sample was cooled to io°C, and sequencing adapters were ligated.

[0567] Without purification, biotin-PEGq-DBCO (Jena Biosciences) was added and the mixture was incubated at 37 °C for 1 hour with shaking at 500 rpm. The DNA was subsequently purified from the reaction mixture.

[0568] 5 pL Dynabeads MyOne Streptavidin Ci beads (ThermoFisher) were washed with 150 pL of PBST. The DNA was added to the beads and the mixture was further incubated at 23 °C for 15 minutes, with shaking at 1000 rpm. Once completed, the supernatant was removed (as the “second fraction”), and the beads were washed twice with 150 pL of PBST. Finally, the bound DNA was released from the beads (as the “first fraction”) by denaturation of streptavidin. - too -

[0569] Amplified libraries were pooled together with 0.1% PhiX and sequenced on a S4 Flow Cell using an Illumina NovaSeq Sequencer (Source Biosciences). qPCR was performed on the Azure Cielo 6 thermocycler (Azure Biosystems) with the following conditions: initial denaturation at 98°C for 30 seconds, then 40 cycles at 95°C for 10 seconds and 6o°C for 60 seconds with fluorescence detection. Analysis of the acquired fluorescence intensity and subsequent quantification of DNA in the samples was performed using Azure Cielo Manager Analysis Software (V1.0.4).

[0570] After sequencing, adapters were removed from the reads using BBTools and then aligned to human reference genome HG38 using BWA-MEM2. Ambiguously aligned reads and those with low mapping scores (MAPQ score < 40) were removed using SamTools. Duplicates were removed with Sambamba (PMID: 25697820) and reads hard-clipped using jvarkit (https: / / github.com / lindenb / jvarkit). Spearman correlation plots were generated from the processed bam files with deepTools using a binsize of tooobp and RPGC normalisation. Saturation figures and CpG density plots were generated for Chri-22 using the QSEA and Repitools R packages. To allow direct comparison of enriched and unenriched samples, Bam files were down sampled to the same sequencing depth using SamTools. High confidence methylated and unmethylated CpG sites used for comparison were taken from the consensus of two whole genome shotgun bisulphite sequencing (WGBS) datasets performed on cell line NA12878 by the same lab (www.encodeproject.org / experiments / ENCSR89oUQO / ), where less than 5% methylation in both datasets was considered to be unmethylated and greater than 95% methylation in both datasets was considered to be methylated.

[0571] Example 13 - Differential indexing of four samples

[0572] Shearing of genomic DNA

[0573] Where necessary, genomic DNA was sheared to an average of 180 bp in 8 microTube- 50 AFA Fiber V2 Strips using a Covaris E220 evolution sonicator. The fragmented DNA size was assessed on a D1000 Tapestation assay (Agilent).

[0574] End-repair andA-tailing and KAPA adapter Ligation

[0575] 5pL of End Repair & A-Tailing Master Mix was added to 25pL of sample containing 50ng of sheared NA12878 gDNA. The mixture was mixed thoroughly and incubated at 20°C for 30 mins followed by a 65°C incubation for a further 30 mins. 2.5pL of KAPA adapters were ligated by adding 22.5pL of Ligation Master Mix and incubating the mixture for 30 min at 20°C.

[0576] DNA purification The DNA was purified from the reaction mixture using 0.8X AMPure Beads (Beckman Coulter). Briefly, DNA:AMPure beads mixture was incubated at 23°C for 5 min with shacking at tooorpm; after immobilising the beads on a magnetic rack, two washes with 150 pL of 80% Ethanol solution were performed; the DNA was eluted in 20 pL of elution buffer (lomM Tris-HCl, 0.01% Tween-20, pH 8.5) at 23°C for 5 min with shacking at tooorpm; finally, after immobilising the beads on a magnetic rack, the purified DNA was recovered in a new 8-tube strip.

[0577] First indexing PCR (PCR index 1) i8ul of purified DNA was added to a PCR mixture comprising 25 pl of 2X KAPA HiFi PCR Mix (Roche), o.5pl of toouM indexing primer 1, o.5pl of toouM adapter complimentary primer and 6pl of water. Following 8 PCR cycles, the DNA was purified using 1X AMPure beads, as described above, and eluted in 30pl of elution buffer (lomM Tris-HCl, 0.01% Tween-20, pH 8.5). This first PCR product was checked for purity on a D1000 Tapestation assay (Agilent and quantified with QubitTM dsDNA HS assay (ThermoFisher Scientific).

[0578] Second indexing PCR (PCR index 2)

[0579] The single-indexed DNA stock was diluted in elution buffer (lomM Tris-HCl, 0.01% Tween-20, pH 8.5) and subdivided in 4 PCR reactions, comprising around 3Ong of DNA, 25pl of 2X KAPA HiFi PCR Mix, o.5pl of toouM indexing primer 2 (different for each PCR reaction), 5pl of 10X adapter complimentary primers and water up to 50pL.

[0580] The mixture underwent 8 cycles of PCR and DNA was purified using 1X AMPure beads as described above and eluted in 20pl of elution buffer (lomM Tris-HCl, 0.01% Tween- 20, pH 8.5).

[0581] Pooling Libraries and MiSeq Sequencing

[0582] The 4 dual-indexed libraries were checked for purity on a D1000 Tapestation assay (Agilent) and quantified with QubitTM dsDNA HS assay (ThermoFisher Scientific), and then pooled to 4nM multiplexed library. A 2nM diluted and denatured library was finally sequenced using a V2 MiSeq 300 Cycle Kit (2x150 bp). Results

[0583] Table 1 below provides a summary of output from MultiQC report for sequencing of one genomic sample (DNA extracted from the NA12878 cell line), split into four technical repeats. As discussed above, the polynucleotides were indexed with a first indexing barcode in a first round of PCR. The sample was then split into four pools, and then each pool was labelled with a different second indexing barcode. The samples were then combined and sequenced simultaneously on the same MiSeq run (V2 reagent kit). The results are shown in Table 1.

[0584] Table 1

[0585] The output of the sequencing experiment for the four differentially indexed samples, sequenced using the same sequencing run is shown in Figure 16. In Table 1 and Figure 16, the sequencing quality metrics are similar across all four samples, indicating that the second barcoding reaction has given rise to a similar DNA library in each of the four samples. Coverage is even and similar across the chromosome. Example 14 - Differential indexing with enrichment of unmethylated DNA

[0586] In this Example, unmethylated CpG nucleotides in gDNA and cfDNA samples were tagged using a methyltransferase using the approach demonstrated in Examples 1-12. Each sample was then prepared into a sequencing library and amplified to incorporate a first indexing barcode at a first end of the polynucleotides but not a second end. Each of the amplified sequencing libraries were then fractionated to obtain a first fraction enriched for the labelled polynucleotides and a second fraction enriched for unlabelled polynucleotides (which is representative of the original sample as a whole). Each of the first and second fractions were then separately amplified to apply a further indexing barcode at the second end of the polynucleotides. All of the fractions from all of the gDNA and cfDNA samples were then pooled and sequenced in a single sequencing run.

[0587] Shearing of genomic DNA

[0588] Where necessary, genomic DNA was sheared to an average of 180 bp in 8 microTube- 50 AFA Fiber V2 Strips using a Covaris E220 evolution sonicator. Sheared DNA was further purified with Zymo DNA Clean&Concetrator Kit following manufactures’ protocol, and the fragmented DNA size was assessed on a D1000 Tapestation assay (Agilent).

[0589] DNA labelling reaction

[0590] 5ng or tong of input DNA (6 cell-free DNA [cfDNA] or 6 genomic DNA [gDNA] samples , respectively) were labelled in a 20 pL reaction using M.Mpel methyltransferase (0.0274 pg / pL) and a cofactor analogue (ETA-AdoHCy-N3 (100 pM)). The reaction was incubated at 37 °C for 1 hour.

[0591] End-repair andA-tailing and KAPA adapter Ligation 9pL of End Repair & A-Tailing Master Mix was added to 20pL of sample containing 5 or long of labelled cfDNA or gDNA, respectively. The mixture was mixed thoroughly and incubated at 20°C for 30 mins followed by a 65°C incubation for a further 30 mins. 2.5pL of KAPA adapters were ligated by adding 22.5pL of Ligation Master Mix and incubating the mixture for 30 min at 20°C.

[0592] DNA capture tagging

[0593] DNA was mixed with 6.i2pL of smM biotin-DBCO and the mixture was incubated at 37 °C for 1 hour with shaking at 500 rpm. DNA purification

[0594] The DNA was purified from the reaction mixture using 1.8X AMPure Beads (Beckman Coulter). Briefly, DNA:AMPure beads mixture was incubated at 23°C for 5 min with shacking at looorpm; after immobilising the beads on a magnetic rack, two washes with 150 pL of 80% Ethanol solution were performed; the DNA was eluted in 20 pL of elution buffer (lomM Tris-HCl, 0.01% Tween-20, pH 8.5) at 23°C for 5 min with shacking at tooorpm; finally, after immobilising the beads on a magnetic rack, the purified DNA was recovered in a new 8-tube strip.

[0595] First indexing PCR 20ul of purified DNA was added to a PCR mixture comprising 25pl of 2X KAPA HiFi

[0596] PCR Mix (Roche) and 5pl of primers mixture (topM indexing primer 1 and topM adapter complimentary primers). Following 2 PCR cycles, the DNA was purified using 2X AMPure beads, as described above. Purified DNA was quantified with QubitTM assay (ThermoFisher Scientific).

[0597] Enrichment of tagged DNA

[0598] Tagged DNA was enriched using 5 pL Dynabeads MyOne Streptavidin Ci beads (ThermoFisher). The DNA: Beads mixture was incubated at 23°C for 15 min with shacking at tooorpm and, after immobilising the beads on a magnetic rack, the supernatant was recovered in a new tube: this fraction contains the unenriched DNA, which comprises the untagged DNA and the DNA products from the first indexing PCR (copy of the input DNA). This unbound fraction was further cleaned-up using 1X AMPure beads, as described above, and eluted in 20pL of elution buffer (lomM Tris- HC1, 0.01% Tween-20, pH 8.5). The beads enriched with the tagged DNA were washed twice with 150 pL of PBS-Tween buffer. For DNA release, the beads were suspended in

[0599] 20 pL of Release Buffer (0.1% sodium dodecylsulfate) and incubated at 9O°C for 30 min with shacking at tooorpm. After immobilising the beads on a magnetic rack, released DNA (tagged) was transferred to a new tube. Second indexing PCR (PCR index 2)

[0600] The pools of enriched and unenriched DNA were further amplified in two separate PCR reactions, comprising 25pl of 2X KAPA HiFi PCR Mix, 5pl of different primers mixture for each DNA pool (topM indexing primer 2 or topM of indexing primer 3, in combination with topM adapter complimentary primers). The unenriched pool underwent 6 cycles of PCR, whereas the enriched pool underwent 13 cycles of PCR, and then the amplified DNA was purified using 1X AMPure beads, as described above, and eluted in 20pl of elution buffer (lomM Tris-HCl, 0.01% Tween-20, pH 8.5).

[0601] Pooling Libraries and NextSeq Sequencing Finally, the 16 dual indexed libraries were checked for purity on a D1000 Tapestation assay (Agilent) and quantified with QubitTM dsDNA HS assay (ThermoFisher Scientific), then pooled to 4nM multiplexed library and sequenced using a P3 NextSeq 2000 Cycle Kit (2xi5obp).

[0602] Results Table 2 below provides a summary of output from MultiQC report for sequencing of 6 gDNA and 6 cfDNA samples.

[0603] The two different pools (or fractions) indicated are the first fraction comprising tag- enriched polynucleotide (the “enriched pool”), and the second fraction that is enriched for untagged polynucleotides (the “genome pool”, which is representative of the original sample as a whole and composed of untagged polynucleotides and the copy of the input DNA produced by amplification). Each sample was indexed in a first round of PCR, and then labelled with a different second indexing barcode, allowing sequencing of all samples’ pools simultaneously on the same NextSeq run (P3 reagent kit).

[0604]

[0605] The output of the sequencing experiment for differentially indexed sample number 4 (as an example of the six samples used in the experiment), sequenced using the same sequencing run is shown in Figure 17.

[0606] For both the gDNA and cfDNA samples, the two different pools (fractions) are shown; a fraction enriched for unlabelled sample (“Genome pool”), comprising unlabelled polynucleotides and the copy of the DNA input, and a second fraction (“Enriched pool”) enriched for labelled nucleotides. (Lane 1: Sample 4 gDNA Genome pool (second fraction); Lane 2: Sample 4 gDNA Enriched pool (first fraction); Lane 3: Sample 4 cfDNA Genome pool (second fraction); and Lane 4: Sample 4 cfDNA Enriched pool (first fraction)).

[0607] It is clear from Table 2 and Figure 17 that the results are highly consistent and reproducible. The coverage is even and similar across the chromosome, as evidenced by the “Genome pool” samples (Lanes 1 and 3), whereas the profile of the “Enriched pool” samples (Lanes 2 and 4) shows clear peaks in correspondence between the gDNA and cfDNA samples of the regions of the chromosome comprising labelled nucleotides (unmethylated CpG).

[0608] Example 15 - Differential indexing with enrichment of methylated DNA In this Example, an amplified sequencing library was obtained comprising polynucleotides having a first indexing barcode at a first, but not a second end. The amplified fraction of the DNA molecules in this pool are generated by PCR and are unmethylated. This pool of polynucleotides is denatured and those DNA molecules containing methylated CpG nucleotides are tagged using an antibody specific for 5- methylcytosine using an approach known as MeDIP (methylated DNA immunoprecipitation). Thus, the sample is fractionated, wherein a first fraction of the sequencing library was enriched for the labelled polynucleotides, and a second fraction was enriched for unlabelled polynucleotides (which is representative of the original sample as a whole). Each of the first and second fractions were then separately amplified to apply a further indexing barcode at the second end of the polynucleotides. All of the fractions were then pooled and sequenced in a single sequencing run. Shearing of genomic DNA Where necessary, genomic DNA was sheared to an average of 180 bp in 8 microTube- 50 AFA Fiber V2 Strips using a Covaris E220 evolution sonicator. The fragmented DNA size was assessed on a D1000 Tapestation assay (Agilent)). End-repair andA-tailing and KAPA adapter Ligation

[0609] 5pL of End Repair & A-Tailing Master Mix was added to 25pL of sample containing 5Ong of sheared NA12878 gDNA. The mixture was mixed thoroughly and incubated at 20°C for 30 mins followed by a 65°C incubation for a further 30 mins. 2.5pL of KAPA adapters were ligated by adding 22.5pL of Ligation Master Mix and incubating the mixture for 30 min at 20°C.

[0610] DNA purification

[0611] The DNA was purified from the reaction mixture using 0.8X AMPure Beads (Beckman Coulter). Briefly, DNA:AMPure beads mixture was incubated at 23°C for 5 min with shacking at tooorpm; after immobilising the beads on a magnetic rack, two washes with 150 pL of 80% Ethanol solution were performed; the DNA was eluted in 20 pL of elution buffer (lomM Tris-HCl, 0.01% Tween-20, pH 8.5) at 23°C for 5 min with shacking at tooorpm; finally, after immobilising the beads on a magnetic rack, the purified DNA was recovered in a new tube.

[0612] First indexing PCR (PCR index 1)

[0613] 20ul of purified DNA was added to a PCR mixture comprising 25pl of 2X KAPA HiFi PCR Mix (Roche) and 5pl of primers mixture (topM indexing primer 1 and topM adapter complimentary primers). Following 2 PCR cycles, the DNA was purified using 1X AMPure beads, as described above, and eluted in 20pl of MGB Water.

[0614] DNA denaturation

[0615] I2ong of single-indexed and purified DNA were mixed with topL of DNA denaturing buffer and water up to 50pL, following manufactures’ instruction (Methylated-DNA IP kit from Zymo Research), and incubated at 98°C for 5 min.

[0616] Methylated DNA Immune Precipitation (MeDIP)

[0617] During the DNA denaturation step, the following IP reaction was prepared: 250pl MIP Buffer, I5pl of ZymoMag Protein A beads and o.8pl Anti-5-Methyl cytosine antibody. Immediately after the denaturation has finished, the denatured DNA was added to the IP mix. The reaction was then incubated at 37°C for 1 hour with shaking at 300rpm. After immobilising the beads on a magnetic rack, the beads were washed twice with 5OO|al of MIP buffer, and once with 500( 1 of DNA elution buffer. Beads were finally resuspended in 15J1I of DNA elution buffer and incubated at 75°C for 5 min. After spinning the sample for 2 min in a micro-centrifuge, the supernatant was transferred to a new tube. This fraction corresponds to the captured methylated-DNA.

[0618] Second indexing PCR (PCR index 2)

[0619] The pools of enriched (MeDIP fraction) and unenriched DNA (single-indexed input DNA) were further amplified in two separate PCR reactions, comprising 25pl of 2X KAPA HiFi PCR Mix, 5pl of different primers mixture for each DNA pool (topM indexing primer 2 or IOJIM of indexing primer 3, in combination with IOJIM adapter complimentary primers). The unenriched pool underwent 7 cycles of PCR, whereas the enriched pool underwent 14 cycles of PCR, and then the amplified DNA was purified using 1X AMPure beads as described above and eluted in 20jil of elution buffer (lomM Tris-HCl, o.imM EDTA, pH 8.5).

[0620] Pooling Libraries and NextSeq Sequencing

[0621] Finally, the 2 dual indexed libraries were checked for purity on a D1000 Tapestation assay (Agilent) and quantified with QubitTM dsDNA HS assay (ThermoFisher Scientific), then pooled to qnM multiplexed library (in combination with libraries from experiment 2) and sequenced using a P3 NextSeq 2000 Cycle Kit (2xi5obp).

[0622] Results

[0623] Table 3 below provides a summary of output from MultiQC report for sequencing of the genomic sample (DNA extracted from the NA12878 cell line).

[0624] The two different pools (or fractions) indicated are the first fraction comprising antibody-capture enriched polynucleotide (the “enriched pool”), and the second fraction that is enriched for unlabelled polynucleotides (the “genome pool”, which is representative of the original sample as a whole and composed of unlabelled polynucleotides and the copy of the input DNA produced by amplification). Each sample was indexed in a first round of PCR, and then labelled with a different second indexing barcode, allowing sequencing of all of the samples simultaneously on the same NextSeq run (P3 reagent kit).

[0625] Table 3

[0626] The output of the sequencing experiment for the differentially indexed sample, sequenced using the same sequencing run is shown in Figure 18. The two different pools (fractions) are shown; a fraction enriched for unlabelled sample (“Genome pool”; Lane 1), comprising the DNA input, and a second fraction (“Enriched pool”; Lane 2) enriched for methylated-DNA enriched by antibody capture (MeDIP).

[0627] It is clear from Table 3 and Figure 18 that the coverage is even and similar across the chromosome, as evidenced by the “Genome pool” sample (Lane 1), whereas the profile of the “Enriched pool” sample (Lane 2) shows clear peaks corresponding to the regions of the chromosome comprising labelled nucleotides (methylated CpG).

Claims

Claims1. A polynucleotide comprising: an indexing barcode incorporated at a first end of the polynucleotide but not a second end; and a label bound site-specifically to a nucleotide residue of the polynucleotide.

2. A polynucleotide as claimed in claim 1, wherein the polynucleotide has a length of 100-500 nucleotides.

3. A polynucleotide as claimed in claim 1 or 2, wherein the label is not a fluorescent label.

4. A polynucleotide as claimed in any of claims 1-3, wherein the label comprises: (i) a tag from a cofactor analogue that is bound to a nucleotide residue that is unmodified in a target position;(ii) a high affinity binding protein or antibody that is bound site-specifically to a target modified nucleotide; or(iii) a DNA binding protein that has been crosslinked to a nucleotide residue.

5. A polynucleotide as claimed in any of claims 1-4, wherein the indexing barcode comprises a unique dual nucleotide index, and / or a unique molecular identifier.

6. A polynucleotide as claimed in any of claims 1-5, wherein the label is suitable for use for enriching the polynucleotide from a mixture comprising the polynucleotide and an unlabelled polynucleotide.

7. An amplified sequencing library comprising: a first indexing barcode incorporated at a first end but not a second end of each polynucleotide in the amplified sequencing libraiy; a first subset of polynucleotides comprising a site-specifically bound label; and a second subset of polynucleotides that do not comprise a label.

8. The amplified sequencing library of claim 7, wherein the first subset of polynucleotides of the sequencing library comprises a polynucleotide as claimed in any of claims 1-6.

9. A method for preparing a polynucleotide sample for sequencing, wherein the prepared polynucleotide sample is suitable for determining both the status of nucleotide residues and the presence of a genetic mutation in the polynucleotide sample, the method comprising:(i) preparing a fractionated amplified sequencing library, wherein each polynucleotide of the fractionated amplified sequencing library comprises a first indexing barcode at a first end but not a second end, and wherein a first fraction of the fractionated amplified sequencing library is enriched for polynucleotides comprising a site-specifically bound label, and wherein the second fraction is enriched for polynucleotides lacking a label;(ii) amplifying the first fraction and incorporating a second indexing barcode to a second end of each polynucleotide in the first fraction;(iii) amplifying the second fraction and incorporating a third indexing barcode to a second end of each polynucleotide in the second fraction; and(iv) pooling the amplified first and second fractions to obtain the prepared polynucleotide sample.

10. The method of claim 9, wherein step (i) comprises preparing an amplified sequencing library as claimed in claim 7 or 8, and fractionating the amplified sequencing library to obtain a first fraction that is enriched for the first subset of polynucleotides and a second fraction that is enriched for the second subset of polynucleotides.

11. The method of claim 9 or 10, wherein determining the status of nucleotide residues in the polynucleotide sample comprises determining the epigenetic modification status, and wherein the site-specifically bound label comprises:- a methyl-CpG binding domain protein;- a high affinity antibody specific for 5-methyl cytosine; - a high affinity antibody specific for 5-hydroxymethylcytosine;- a high affinity antibody specific for Nq-methylcytosine;- a high affinity antibody specific for N6-methyladenine; or- a high affinity antibody specific for 5-methyl cytosine.

12. The method of claim 11, wherein step (i) comprises:(a) preparing an amplified sequencing library, wherein each polynucleotide of the amplified sequencing library comprises a first indexing barcode at a first end but not a second end, and wherein the amplified sequencing library further comprises a label bound site- specifically to a first subset of the polynucleotides but not a second subset; and(b) fractionating the amplified sequencing library into first and second fractions, wherein the first fraction is enriched for polynucleotides comprising a label, and wherein the second fraction is enriched for polynucleotides lacking a label.

13. The method of claim 12, wherein part (a) comprises preparing the polynucleotides into a sequencing library and additionally:(1) binding a label site-specifically to the polynucleotides to form the first subset of polynucleotides that comprises the site-specifically bound label and the second subset of polynucleotides that does not comprise the site-specifically bound label; and (2) amplifying the polynucleotides and incorporating a first indexing barcode at a first end of the polynucleotides but not a second end, wherein step (1) may be performed before step (2), or step (2) may be performed before step (1).

14. The method of claim 9 or 10, wherein determining the status of nucleotide residues in the polynucleotide sample comprises determining the epigenetic modification status, and wherein the site-specifically bound label comprises:- a tag from a cofactor analogue that is bound to a cytosine residue that is epigenetically unmodified in the C5 position; - a tag from a cofactor analogue that is bound to a cytosine residue that is epigenetically unmodified in the N4 position; or- a tag from a cofactor analogue that is bound to an adenine residue that is epigenetically unmodified in the N6 position.

15. The method of claim 14, wherein step (i) comprises:(a) preparing an amplified sequencing library comprising:(1) - using a methyltransferase enzyme configured to modify a nucleotide residue in a target position to apply a tag to each unmodified nucleotide residue in a polynucleotide of the sample, wherein each unmodified nucleotide residue is unmodified in the target position, and subsequently inactivating the methyltransferase;(2) - preparing the polynucleotide sample into a sequencing library;(3) - binding an affinity label to each tag; and(4) - amplifying the polynucleotides and incorporating a first indexing barcode at a first end of the polynucleotides but not a second end, wherein steps (2), (3), and (4) are performed in any order or combination after step (1);(b) fractionating the amplified sequencing library into first and second fractions, wherein the first fraction is enriched for polynucleotides comprising a label, and wherein the second fraction is enriched for polynucleotides lacking a label.

16. A method as claimed in claim 15, wherein preparing an amplified sequencing library is performed in a one-pot approach.

17. A method as claimed in claim 15 or 16, wherein preparing an amplified sequencing library involves at most one sample purification step.

18. A method as claimed in any of claims 15-17, wherein using a methyltransferase enzyme to apply a tag to each unmodified nucleotide residue comprises the use of a methyltransferase cofactor analogue.

19. A method as claimed in any of claims 15-18, wherein the methyltransferase enzyme is a C5 methyltransferase, and wherein the methyltransferase cofactor analogue is ETA-AdoHcy-N3.

20. A method as claimed in any of claims 15-19, wherein the methyltransferase enzyme consists of or comprises a variant or fragment of the wild type M.Mpel sequence, comprising at least 80% sequence identity to the wild type M.Mpel sequence having the NCBI accession number BAC44284.

21. A method as claimed in any of claims 15-20, wherein preparing the polynucleotide sample into a sequencing library comprises end repair, A-tailing, and adapter ligation of the polynucleotides in the sample.

22. A method as claimed in any of claims 15-21, wherein binding an affinity label to each tag comprises adding an affinity label precursor directly into the sequencing library preparation mixture, without a washing step.

23. A method as claimed in any of claims 15-22, wherein the affinity label comprises biotin, and wherein fractionating the sequencing library into first and second fractions comprises fractionation using a capture agent comprising a biotin-binding protein.

24. A method as claimed in any of claims 9-23, wherein step (iii) further comprises target enrichment of the second fraction for one or more genome regions of interest.

25. A method as claimed in claim 24, wherein the method further comprises amplification of the polynucleotides of the second fraction after target enrichment.

26. A method as claimed in claim 24 or 25, wherein target enrichment of the second fraction comprises:(a) in-solution hybridization using oligonucleotide “bait” probes specific for genomic regions of interest; and / or (b) a PCR-based enrichment method.

27. A method as claimed in any of claims 9-26, wherein the polynucleotide sample is a cfDNA sample.

28. A polynucleotide sample for sequencing, wherein the polynucleotide sample is obtained or obtainable by a method as claimed in any of claims 9-27.

29. A method for obtaining sequencing information of a polynucleotide sample, the method comprising: - obtaining a polynucleotide sample prepared by a method as claimed in any of claims 9-27; and sequencing the polynucleotide sample to obtain sequencing information.

30. A method for determining both the status of nucleotide residues and the presence of a genetic mutation in a polynucleotide sample, the method comprising: obtaining sequencing information using a method as claimed in claim 29; using the indexing barcodes to distinguish the sequencing information of the polynucleotides of the first and second fractions; and using the sequencing information of the polynucleotides of the first fraction to determine the status of nucleotide residues in the polynucleotide sample, and using the sequencing information of the polynucleotides of the second fraction to determine thegenetic sequence of a region of the polynucleotide sample and thereby the presence of a genetic mutation in the polynucleotide sample.

31. A method for preparing a profile of a region of a polynucleotide sample, the profile comprising both the status of nucleotide residues and any genetic mutations in the region of the polynucleotide sample, the method comprising: obtaining the status of nucleotide residues in the polynucleotide sample and the genetic sequence of a region of the polynucleotide sample using a method as claimed in claim 30; and - comparing the sequences of the sequencing reads of the first and second fractions to a reference sequence to determine the location of the sequencing reads within the reference sequence and thereby the status of specific nucleotides and / or genetic mutations at specific locations within the reference sequence.

32. An in vitro method for diagnosing disease in a subject, the method comprising diagnosing the disease based on a profile obtained by a method as claimed in claim 31, using a polynucleotide sample obtained from the subject.

33. A profile of a region of a polynucleotide sample, comprising both the status of nucleotide residues and any genetic mutations in the region of the polynucleotide sample, wherein the profile is obtained or obtainable by a method as claimed in claim 31-34. A method as claimed in claims 31 or 32, or a profile as claimed in claim 33, wherein the region may consist of one or more specific portions of the polynucleotide sample or the entire polynucleotide sample.

35. A kit for determining both the status of nucleotide residues and the presence of a genetic mutation in a polynucleotide sample, the kit comprising: (i) a labelling component suitable for site-specifically labelling a nucleotide residue of a polynucleotide in the sample;(ii) a capture agent suitable for binding to the labelling component and for fractionating the polynucleotides into first and second fractions, wherein the first fraction is enriched for labelled polynucleotides, and wherein the second fraction is enriched for polynucleotides that do not comprise a label, optionally wherein the capture agent is tethered to a solid-phase support; and optionally(iii) a releasing agent suitable for releasing the polynucleotides from the capture agent.

Citation Information

Patent Citations

  • S-adenosyl-l-cysteine analogues as cofactors for methyltransferases

    EP3186266B1

  • S-adenosyl-L-methionine analogs with extended activated groups for transfer by methyltransferases

    US8008007B2

  • Methods and systems for analyzing nucleic acid molecules

    WO2018119452A2

  • Compositions and methods for enriching methylated polynucleotides

    WO2022115810A1

  • Methods involving methylation preserving amplification with error correction

    WO2024137880A2

Cited By

  • Epigenetic profiling method of nucleotide residues in cell-free DNA

    WO2026125880A1