A method for sequencing chromatin

By characterizing polynucleotides with bound proteins using a transmembrane pore, the method enhances stability and read length, simplifies preparation, and allows for protein location identification, addressing the challenges of protein removal in nanopore sensing.

WO2025242713A1PCT designated stage Publication Date: 2025-11-27OXFORD NANOPORE TECH LTD
View PDF 42 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/063943
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-21
Filing Date
2025-05-21
Publication Date
2025-11-27

AI Technical Summary

Technical Problem

Existing nanopore sensing methods require the removal of proteins bound to polynucleotides before analysis, which complicates the preparation process and can degrade the polynucleotides, limiting read length and stability.

Method used

A method that allows polynucleotides with bound proteins to be characterized through a transmembrane pore by retaining the proteins until the polynucleotide is contacted with the pore, stripping them during movement, and taking measurements to characterize the polynucleotide.

Benefits of technology

This approach simplifies the preparation protocol, increases polynucleotide stability and read length, and enables long-term storage without protein removal steps, while providing insights into protein location through signal changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000045_0001
    Figure IMGF000045_0001
  • Figure 00000050_0000
    Figure 00000050_0000
  • Figure 00000051_0000
    Figure 00000051_0000
Patent Text Reader

Abstract

The invention relates to a method of preparing a target polynucleotide comprising one or more bound proteins, such as one or more histones, for characterisation using a transmembrane pore. The method also relates to the characterisation of such polynucleotides.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] A METHOD FOR SEQUENCING CHROMATIN

[0002] TECHNICAL FIELD

[0003] The invention relates to a method of preparing a target polynucleotide comprising one or more bound proteins, such as one or more histones, for characterisation using a transmembrane pore. The method also relates to the characterisation of such polynucleotides.

[0004] INTRODUCTION

[0005] Nanopore sensing is an approach to analyte detection and characterization that relies on the observation of individual binding or interaction events between the analyte molecule and an ion conducting channel. Two of the essential components of analyte characterisation using nanopore sensing are (1) the control of analyte movement through the pore and (2) the discrimination of the composing building blocks as the analyte is moved through the pore. The movement of the analyte through the pore is typically controlled using an enzyme, such as a helicase. During nanopore sensing, the narrowest part of the pore forms the most discriminating part of the nanopore with respect to the current signatures as a function of the passing analyte.

[0006] For polynucleotide analytes, nucleotide discrimination is achieved by measuring the current as the polynucleotide passes through the pore. Multiple nucleotides contribute to the observed current.

[0007] Polynucleotides are typically prepared for nanopore sensing by removing proteins bound to the polynucleotide. This allows the polynucleotide to move effectively and efficiently through the nanopore. Protein removal is usually carried out during the polynucleotide purification and enrichment steps before the movement control protein is added to the polynucleotide. There is a need to identify novel ways to improve nanopore sensing, especially of polynucleotides.

[0008] SUMMARY OF THE INVENTION

[0009] The present inventors have surprisingly shown that target polynucleotides can move through a transmembrane pore and be characterised using the pore even if proteins bound to the polynucleotides are not removed before the polynucleotides are contacted with the pore. The movement of the polynucleotide with respect to, such as through, the transmembrane pore strips the bound proteins from the polynucleotide, especially when a movement control protein used.

[0010] The inventors have also shown that the one or more proteins bound to a target polynucleotide increase the stability of the target polynucleotide and, in particular, increase the read length of the target polynucleotide using nanopore sensing. The invention therefore provides a method of characterising a target polynucleotide comprising one or more bound proteins, the method comprising: (a) retaining the one or more bound proteins on the target polynucleotide until the target polynucleotide is contacted with a transmembrane pore, (b) contacting the target polynucleotide with the transmembrane pore such that one or more of the bound proteins are stripped from the target polynucleotide as the target polynucleotide moves with respect to the pore; and (c) taking one or more measurements as the target polynucleotide moves with respect to the pore wherein the measurements are indicative of one or more characteristics of the target polynucleotide and thereby characterising the target polynucleotide.

[0011] The invention also provides a method of increasing the read length of a target polynucleotide using a transmembrane pore, wherein the method comprises binding one or more proteins to the target polynucleotide or retaining one or more bound proteins on the target polynucleotide.

[0012] The invention also provides: a method of characterising chromatin, the method comprising: (a) retaining one or more bound proteins on the chromatin until the chromatin is contacted with a transmembrane pore, (b) contacting the chromatin with the transmembrane pore such that one or more of the bound proteins are stripped from the chromatin as the chromatin moves with respect to the pore; and (c) taking one or more measurements as the chromatin moves with respect to the pore wherein the measurements are indicative of one or more characteristics of the chromatin and thereby characterising the chromatin;

[0013] - a method of preparing a target polynucleotide comprising one or more bound proteins for characterisation using a transmembrane pore, the method comprising: (a) retaining the one or more bound proteins on the target polynucleotide until the target polynucleotide is contacted with the transmembrane pore and (b) attaching a movement control protein to the target polynucleotide before the target polynucleotide is contacted with the transmembrane pore;

[0014] - a method of preparing a target polynucleotide for characterisation using a transmembrane pore, the method comprising: (a) binding one or more proteins to the target polynucleotide before the target polynucleotide is contacted with the transmembrane pore and (b) attaching a movement control protein to the target polynucleotide before the target polynucleotide is contacted with the transmembrane pore; - a target polynucleotide comprising one or more bound proteins and further comprising a movement control protein;

[0015] - a target polynucleotide comprising one or more selectively retained nucleosomes and a sequencing adaptor; and

[0016] - an apparatus produced by a method comprising (a) obtaining a target polynucleotide comprising one or more bound proteins, (b) attaching a movement control protein and / or a sequencing adaptor to the target polynucleotide and (c) contacting the target polynucleotide obtained in (b) with a transmembrane pore.

[0017] The methods of the invention have several key advantages. For instance, by removing one or more protein removal steps from the protocol for preparing a target polynucleotide for nanopore sensing, the protocol is simplified, and involves fewer reagents. Genomic material, i.e., chromatin, can simply be extracted from cells and characterised without having to remove the histones and other proteins from the chromatin.

[0018] The presence of one or more proteins bound to the target polynucleotide also increases the stability of the target polynucleotide. In particular, the presence of one or more proteins bound to the target polynucleotide increases the read length of the target polynucleotide using a transmembrane pore. The increased stability of the target polynucleotide means the manner in which the target polynucleotide is handled before it is contacted with the transmembrane pore is less important. The target polynucleotide can be handled in ways which might degrade or damage other polynucleotides without one or more bound proteins. The stabilised target polynucleotide can also be stored long term.

[0019] In addition, changes in pore signal when the one or more proteins are stripped from the target polynucleotide may be used to identify the location of the one or more proteins on the target polynucleotide, for instance in foot printing methods. This is discussed in more detail below. Additional advantages of the method of the invention are discussed below with reference to specific embodiments.

[0020] DESCRIPTION OF THE FIGURES

[0021] It is to be understood that Figures are for the purpose of illustrating particular embodiments of the invention only and are not intended to be limiting.

[0022] Figure 1 : Read length distributions are shown for reconstituted chromatin versus the same batch of non-reconstituted input DNA.

[0023] Figure 2: Products of the MNase digestion assay are run on an agarose gel, chromatin was extracted from cultured HG002 cells and sheared using sonication. From left to right, lane 1 and 2 are ladders (NEB: N3200L), lanes 3-5 show the digestion products of extracted chromatin in a 10-fold dilution series of MNase without ProteinaseK digestion, lanes 6-8 show the digestion products of extracted chromatin in a 10-fold dilution series of MNase with ProteinaseK digestion, lane 9 shows a 150mer amplicon without MNase digestion or ProteinaseK and lane 10 is a lOObp ladder (NEB: N3231L). At low concentrations of MNase, or for no digestion control wells, the DNA remains in the well. At higher concentrations of MNase we see a band at around 150bps which indicates chromatin has been properly formed.

[0024] Figure 3: A basecalled and aligned raw current trace generated by passing chromatin, extracted from cultured HG002 cells, through a nanopore under the control of a motor protein. The trace contains several recognisable components, including: A - the RA sequencing adapter, B - the FRA tagmentation adapter, C - extracted DNA.

[0025] DETAILED DESCRIPTION

[0026] It is to be understood that different applications of the disclosed products and methods may be tailored to the specific needs in the art. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments of the invention only and is not intended to be limiting.

[0027] All publications, patents and patent applications cited herein, whether supra or infra, are hereby incorporated by reference in their entirety. All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent publications and patents or patent applications incorporated by reference contradict the disclosure contained in the specification, the specification is intended to supersede and / or take precedence over any such contradictory material.

[0028] The invention, both as to organization and method of operation, together with features and advantages thereof, may best be understood by reference to the following detailed description when read in conjunction with the accompanying Figures. The aspects and advantages of the invention will be apparent from and elucidated with reference to the embodiment(s) described hereinafter. Reference throughout this specification to "embodiment(s)" means that a particular feature, structure, or characteristic described in connection with the embodiment(s) is included in at least one embodiment of the present invention. Thus, appearances of the phrases "in one embodiment", "in another embodiment" or "in a preferred embodiment" in various places throughout this specification are not necessarily all referring to the same embodiment, but may do so. Similarly, it should be appreciated that in the description of exemplary embodiments of the invention, various features of the invention are sometimes grouped together in a single embodiment, or description thereof for the purpose of streamlining the disclosure and aiding in the understanding of one or more of the various inventive aspects. This method of disclosure, however, is not to be interpreted as reflecting an intention that the claimed invention requires more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive aspects lie in less than all features of a single foregoing disclosed embodiment.

[0029] Definitions

[0030] Where an indefinite or definite article is used when referring to a singular noun e.g., "a" or "an", "the", this includes a plural of that noun unless something else is specifically stated. Where the term "comprising" is used in the present description and claims, it does not exclude other elements or steps. The term "comprising" is interchangeable with "consisting of" or "consisting essentially of".

[0031] Furthermore, the terms first, second, third and the like in the description and in the claims, are used for distinguishing between similar elements and not necessarily for describing a sequential or chronological order. It is to be understood that the terms so used are interchangeable under appropriate circumstances and that the embodiments of the invention described herein are capable of operation in other sequences than described or illustrated herein.

[0032] The following terms or definitions are provided solely to aid in the understanding of the invention. Unless specifically defined herein, all terms used herein have the same meaning as they would to one skilled in the art of the invention. Practitioners are particularly directed to Sambrook et al., Molecular Cloning: A Laboratory Manual, 4thed., Cold Spring Harbor Press, Plainsview, New York (2012); and Ausubel et al., Current Protocols in Molecular Biology (Supplement 114), John Wiley & Sons, New York (2016), for definitions and terms of the art. The definitions provided herein should not be construed to have a scope less than understood by a person of ordinary skill in the art.

[0033] "About” as used herein when referring to a measurable value such as an amount, percentage, a temporal duration, and the like, is meant to encompass variations of ± 20 % or ± 10 %, more preferably ± 5 %, even more preferably ± 1 %, and still more preferably ± 0.1 % from the specified value, as such variations are appropriate to perform the disclosed methods. Any embodiment containing the term "about" includes the same feature without the term. For example, about 10 nucleotides or more includes 10 nucleotides or more.

[0034] "Polynucleotide" as used herein refers to a polymeric form of nucleotides of any length, either ribonucleotides or deoxyribonucleotides. This term refers only to the primary structure of the molecule. Thus, this term includes double- and single-stranded DNA, and RNA. The term "polynucleotide" as used herein, is a single or double stranded covalently linked sequence of nucleotides in which the 3' and 5' ends on each nucleotide are joined by phosphodiester bonds. The polynucleotide may be made up of deoxyribonucleotide bases or ribonucleotide bases. Polynucleotides may be manufactured synthetically in vitro or isolated from natural sources. Polynucleotides may further include modified DNA or RIMA, for example DNA or RNA that has been methylated, or RNA that has been subject to post- translational modification, for example 5'-capping with 7-methylguanosine, 3'-processing such as cleavage and polyadenylation, and splicing. Polynucleotides may also include synthetic nucleic acids (XNA), such as hexitol nucleic acid (HNA), cyclohexene nucleic acid (CeNA), threose nucleic acid (TNA), glycerol nucleic acid (GNA), locked nucleic acid (LNA), peptide nucleic acid (PNA) and (poly)ADPribose modified DNA. Sizes of polynucleotides are typically expressed as the number of base pairs (bp) or nucleotide pairs for double stranded polynucleotides, or in the case of single stranded polynucleotides as the number of nucleotides (nt). One thousand bp or nt equal a kilobase (kb). Polynucleotides of less than around 40 nucleotides in length are typically called "oligonucleotides" and may comprise primers for use in manipulation of DNA such as via polymerase chain reaction (PCR). The term "polynucleotide" is interchangeable with "polynucleotide sequence", "nucleotide sequence", "DNA sequence", "nucleic acid", or "nucleic acid molecule(s)".

[0035] The term "amino acid" in the context of the present disclosure is used in its broadest sense and is meant to include organic compounds containing amine (NH2) and carboxyl (COOH) functional groups, along with a side chain (e.g., a R group) specific to each amino acid. The amino acids typically refer to naturally occurring L o-amino acids or residues. The commonly used one and three letter abbreviations for naturally occurring amino acids are used herein: A=Ala; C=Cys; D=Asp; E=Glu; F=Phe; G=Gly; H=His; I=Ile; K=Lys; L=Leu; M = Met; N=Asn; P=Pro; Q=Gln; R=Arg; S=Ser; T=Thr; V=Val; W=Trp; and Y=Tyr (Lehninger, A. L., (1975) Biochemistry, 2d ed., pp. 71-92, Worth Publishers, New York). The general term "amino acid" further includes D-amino acids, retro-inverso amino acids as well as chemically modified amino acids such as amino acid analogues, naturally occurring amino acids that are not usually incorporated into proteins such as norleucine, and chemically synthesised compounds having properties known in the art to be characteristic of an amino acid, such as P-amino acids. For example, analogues or mimetics of phenylalanine or proline, which allow the same conformational restriction of the peptide compounds as do natural Phe or Pro, are included within the definition of amino acid. Such analogues and mimetics are referred to herein as "functional equivalents" of the respective amino acid. Other examples of amino acids are listed by Roberts and Vellaccio, The Peptides: Analysis, Synthesis, Biology, Gross and Meiehofer, eds., Vol. 5 p. 341, Academic Press, Inc., N.Y. 1983, which is incorporated herein by reference.

[0036] The terms "polypeptide", and "peptide" are interchangeably used herein to refer to a polymer of amino acid residues and to variants and synthetic analogues of the same. Thus, these terms apply to amino acid polymers in which one or more amino acid residues is a synthetic non-naturally occurring amino acid, such as a chemical analogue of a corresponding naturally occurring amino acid, as well as to naturally occurring amino acid polymers. Polypeptides can also undergo maturation or post-translational modification processes that may include, but are not limited to glycosylation, proteolytic cleavage, lipidization, signal peptide cleavage, propeptide cleavage, phosphorylation, and such like. A peptide can be made using recombinant techniques, e.g., through the expression of a recombinant or synthetic polynucleotide. A recombinantly produced peptide it typically substantially free of culture medium, e.g., culture medium represents less than about 20 %, more preferably less than about 10 %, and most preferably less than about 5 % of the volume of the protein preparation.

[0037] The term "protein" is used to describe a folded polypeptide having a secondary or tertiary structure. The protein may be composed of a single polypeptide or may comprise multiple polypepties that are assembled to form a multimer. The multimer may be a homooligomer, or a heterooligmer. The protein may be a naturally occurring or wild type protein, or a modified, or non-naturally, occurring protein. The protein may, for example, differ from a wild type protein by the addition, substitution, or deletion of one or more amino acids.

[0038] A "variant" of a protein encompasses peptides, oligopeptides, polypeptides, proteins, and enzymes having amino acid substitutions, deletions and / or insertions relative to the unmodified or wild-type protein in question and having similar biological and functional activity as the unmodified protein from which they are derived. The term "amino acid identity" as used herein refers to the extent that sequences are identical on an amino acid- by-amino acid basis over a window of comparison. Thus, a "percentage of sequence identity" is calculated by comparing two optimally aligned sequences over the window of comparison, determining the number of positions at which the identical amino acid residue (e.g., Ala, Pro, Ser, Thr, Gly, Vai, Leu, He, Phe, Tyr, Trp, Lys, Arg, His, Asp, Glu, Asn, Gin, Cys and Met) occurs in both sequences to yield the number of matched positions, dividing the number of matched positions by the total number of positions in the window of comparison ( / .e., the window size), and multiplying the result by 100 to yield the percentage of sequence identity.

[0039] For all aspects and embodiments of the invention, a "variant" has at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98% or at least about 99% complete sequence identity or homology to the amino acid sequence of the corresponding wild-type protein. Sequence identity or homology can also be to a fragment or portion of the full-length polynucleotide or polypeptide. Hence, a sequence may have only at least about 40% overall sequence identity or homology with a full-length reference sequence, but a sequence of a particular region, domain or subunit could share at least about 80%, at least about 90%, or as much as at least about 99% sequence identity or homology (or any of the %s discussed above) with the reference sequence.

[0040] Standard methods in the art may be used to determine homology. For example, the UWGCG Package provides the BESTFIT program which can be used to calculate homology, for example used on its default settings (Devereux et al (1984) Nucleic Acids Research 12, p387-395). The PILEUP and BLAST algorithms can be used to calculate homology or line up sequences (such as identifying equivalent residues or corresponding sequences (typically on their default settings)), for example as described in Altschul S. F. (1993) J Mol Evol 36:290- 300; Altschul, S.F et al (1990) J Mol Biol 215:403-10. Software for performing BLAST analyses is publicly available through the National Center for Biotechnology Information (http: / / www.ncbi.nlm.nih.gov / ).

[0041] The term "wild-type" refers to a gene or gene product isolated from a naturally occurring source. A wild-type gene is that which is most frequently observed in a population and is thus arbitrarily designed the "normal" or "wild-type" form of the gene. In contrast, the term "modified", "mutant" or "variant" refers to a gene or gene product that displays modifications in sequence (e.g., substitutions, truncations, or insertions), post-translational modifications and / or functional properties (e.g., altered characteristics) when compared to the wild-type gene or gene product. It is noted that naturally occurring mutants can be isolated; these are identified by the fact that they have altered characteristics when compared to the wild-type gene or gene product. Methods for introducing or substituting naturally occurring amino acids are well known in the art. For instance, methionine (M) may be substituted with arginine (R) by replacing the codon for methionine (ATG) with a codon for arginine (CGT) at the relevant position in a polynucleotide encoding the mutant monomer. Methods for introducing or substituting non-naturally occurring amino acids are also well known in the art. For instance, non-naturally occurring amino acids may be introduced by including synthetic aminoacyl-tRNAs in the IVTT system used to express the mutant monomer. Alternatively, they may be introduced by expressing the mutant monomer in E. coli that are auxotrophic for specific amino acids in the presence of synthetic ( / .e., non-naturally occurring) analogues of those specific amino acids. They may also be produced by naked ligation if the mutant monomer is produced using partial peptide synthesis. Conservative substitutions replace amino acids with other amino acids of similar chemical structure, similar chemical properties, or similar side-chain volume. The amino acids introduced may have similar polarity, hydrophilicity, hydrophobicity, basicity, acidity, neutrality, or charge to the amino acids they replace. Alternatively, the conservative substitution may introduce another amino acid that is aromatic or aliphatic in the place of a pre-existing aromatic or aliphatic amino acid. Conservative amino acid changes are well- known in the art. A mutant or variant protein can also be chemically modified in any way and at any site. A mutant or variant protein is preferably chemically modified by attachment of a molecule to one or more cysteines (cysteine linkage), attachment of a molecule to one or more lysines, attachment of a molecule to one or more non-natural amino acids, enzyme modification of an epitope or modification of a terminus. Suitable methods for carrying out such modifications are well-known in the art. The mutant or variant protein may be chemically modified by the attachment of any molecule. For instance, the mutant or variant protein may be chemically modified by attachment of a dye or a fluorophore.

[0042] Method of preparing a target polynucleotide

[0043] The invention provides a method of a preparing a target polynucleotide for characterisation using a transmembrane pore. Characterisation and transmembrane pores are described in more detail below. The target polynucleotide comprises one or more bound proteins. The target polynucleotide may comprise any number and type of bound proteins. Examples of numbers and types of proteins are provided below.

[0044] The target polynucleotide may naturally comprise one or more bound proteins. The target polynucleotide may comprise one or more bound proteins in its natural state, such as in a cell. Examples of suitable cells are discussed below.

[0045] The one or more bound proteins may be added to the target polynucleotide. The one or more bound proteins may be added to the target polynucleotide artificially and / or as part of the method of the invention. In other words, the target polynucleotide may not comprise one or more bound proteins in its natural state. Alternatively, the target polynucleotide itself may be artificial and be produced without proteins before the one or more bound proteins are added for instance as part of the method of the invention.

[0046] Step (a)

[0047] The method comprises (a) retaining the one or more bound proteins on the target polynucleotide until the target polynucleotide is contacted with the transmembrane pore. The term "retaining" is interchangeable with "keeping" or "maintaining". The method preferably comprises selectively retaining or intentionally retaining the one or more bound proteins on the target polynucleotide until the target polynucleotide is contacted with the transmembrane pore. In all instances herein, "retaining" is interchangeable with "selectively retaining" or "intentionally retaining".

[0048] The method comprises taking one or more steps and / or avoiding one or more steps to ensure the one or more proteins are bound to the target polynucleotide until the target polynucleotide is contacted with the transmembrane pore. The method comprises "selectively" retaining the one or more bound proteins because it may comprise removing other proteins as discussed in more detail below. The invention also comprises a method of preparing a sample comprising a target polynucleotide, wherein the target polynucleotide comprises one or more bound proteins for characterisation using a transmembrane pore, the method comprising: (a) retaining the one or more bound proteins on the target polynucleotide until the target polynucleotide is contacted with the transmembrane pore and (b) attaching a movement control protein to the target polynucleotide before the target polynucleotide is contacted with the transmembrane pore. The sample may comprise chromatin.

[0049] In one embodiment, the method comprises retaining the one or more bound proteins by removing one or more proteins from the target polynucleotide other than the one or more bound proteins. The method preferably comprises selectively removing all proteins from the target polynucleotide other than the one or more bound proteins. This can be achieved using (a) a salt concentration, (b) a temperature and / or (c) a pH, such as (a), (b), (c), (a) and (b), (b) and (c), (a) and (c) or (a), (b) and (c), for removing one or more, or all, proteins other than the one or more bound proteins. Antibodies or tags specific for the proteins other than the one or more bound proteins may be used to remove them from the target polynucleotide.

[0050] In another embodiment, the method comprises retaining the one or more bound proteins by not comprising a step for removing the one or more bound proteins from the target polynucleotide. The method may comprise retaining the one or more bound proteins by not comprising a step for removing the one or more bound proteins from the sample. The method preferably does not comprise a step for removing the one or more bound proteins from the target polynucleotide. The method preferably does not comprise a step for removing the one or more bound proteins from the sample. The method may not comprise a step for removing protein or any protein from the target polynucleotide. The method may not comprise a step for removing protein or any protein from the sample.

[0051] Steps for removing proteins from target polynucleotides are known in the art. DNA and RIMA extraction kits are commercially available (such as from Qiagen, NEB, Thermo, etc.). The method may not comprise any of these steps. The step for removing the one or more bound proteins is preferably a phenol or chloroform step, a precipitation step and / or a proteinase step. The step may be a phenol or chloroform step. The step may be a precipitation step. The step may be a proteinase step. The step may be a phenol or chloroform step and a precipitation step. The step may comprise a precipitation step and a proteinase step. The step may be a phenol or chloroform step and a proteinase step. The step may be a phenol or chloroform step, a precipitation step, and a proteinase step. The method preferably does not comprise a phenol or chloroform step, a precipitation step and / or a proteinase step. The method preferably does not comprise a phenol or chloroform step, a precipitation step and / or a proteinase step for removing the one or more bound proteins from the target polynucleotide. The method preferably does not comprise a phenol or chloroform step, a precipitation step and / or a proteinase step for removing the one or more bound proteins from the sample. The method preferably does not comprise a phenol or chloroform step, a precipitation step and / or a proteinase step for removing protein or any protein from the target polynucleotide. The method preferably does not comprise a phenol or chloroform step, a precipitation step and / or a proteinase step for removing protein or any protein from the sample. The proteinase is preferably proteinase K, trypsin, chymotrypsin, and / or papain.

[0052] Although the method comprises retaining the one or more bound proteins on the target polynucleotide, the method may comprise removing other proteins from the target polynucleotide or from the sample comprising the target polynucleotide. The method preferably comprises removing one or more non-bound proteins. The method preferably comprises removing one or more non-bound proteins from the sample. The one or more non-bound proteins may be any proteins that may be present with the target polynucleotide or present in the sample. As discussed below, the target polynucleotide may be isolated from cells. The one or more non-bound proteins may be one or more proteins derived from the cells. The one or more non-bound proteins derived from cells include, but are not limited to, membrane proteins, cytosolic proteins, organelle proteins, including nuclear envelope proteins, and exogenous or viral proteins expressed from transfected plasmids or viruses.

[0053] Any number of individual non-bound proteins may be removed, such as about 1 or more, about 2 or more, about 3 or more, about 4 or more, about 5 or more, about 6 or more, about 7 or more, about 8 or more, about 9 or more, about 10 or more, about 20 or more, about 30 or more, about 40 or more, about 50 or more, about 100 or more, about 200 or more, about 500 or more, about 1000 or more, about 104or more, about 105or more, about 106or more, about 107or more, about 108or more, about 109or more, about 1010or more, about 1011or more, about 1012or more individual proteins. Any number of different types of non-bound proteins may be removed, such as such as about 1 or more, about 2 or more, about 3 or more, about 4 or more, about 5 or more, about 6 or more, about 7 or more, about 8 or more, about 9 or more, about 10 or more, about 20 or more, about 30 or more, about 40 or more, about 50 or more, about 100 or more, about 200 or more, about 500 or more or about 1000 or more different types of protein.

[0054] Step (b)

[0055] The method also comprises (b) attaching a movement control protein to the target polynucleotide before the target polynucleotide is contacted with the transmembrane pore. Movement control proteins and their attachment to target polynucleotides are known in the art. The movement control protein may be attached to the target polynucleotide using one or more sequencing adaptors. This is described in more detail below. The movement control protein may be any of those described in more detail below.

[0056] Target polynucleotide

[0057] The target polynucleotide may comprise any combination of any nucleotides. The nucleotides can be naturally occurring or artificial. A nucleotide typically contains a nucleobase, a sugar and at least one phosphate group. The nucleobase and sugar form a nucleoside. The nucleobase is typically heterocyclic. Nucleobases include, but are not limited to, purines and pyrimidines and more specifically adenine (A), guanine (G), thymine (T), uracil (U) and cytosine (C).

[0058] The sugar is typically a pentose sugar. Nucleotide sugars include, but are not limited to, ribose and deoxyribose. The sugar is preferably a deoxyribose. The polynucleotide preferably comprises the following nucleosides: deoxyadenosine (dA), deoxyuridine (dU) and / or thymidine (dT), deoxyguanosine (dG) and deoxycytidine (dC).

[0059] The nucleotide is typically a ribonucleotide or deoxyribonucleotide. The nucleotide typically contains a monophosphate, diphosphate, or triphosphate. The nucleotide may comprise more than three phosphates, such as 4 or 5 phosphates. Phosphates may be attached on the 5' or 3' side of a nucleotide. Nucleotides include, but are not limited to, adenosine monophosphate (AMP), guanosine monophosphate (GMP), thymidine monophosphate (TMP), uridine monophosphate (UMP), 5-methylcytidine monophosphate, 5- hydroxy methylcytidine monophosphate, cytidine monophosphate (CMP), cyclic adenosine monophosphate (cAMP), cyclic guanosine monophosphate (cGMP), deoxyadenosine monophosphate (dAMP), deoxyguanosine monophosphate (dGMP), deoxythymidine monophosphate (dTMP), deoxyuridine monophosphate (dUMP), deoxycytidine monophosphate (dCMP) and deoxymethylcytidine monophosphate. The nucleotides are preferably selected from AMP, TMP, GMP, CMP, UMP, dAMP, dTMP, dGMP, dCMP and dUMP.

[0060] A nucleotide may be abasic ( / .e., lack a nucleobase). A nucleotide may also lack a nucleobase and a sugar ( / .e., is a C3 spacer).

[0061] The nucleotides in the polynucleotide may be attached to each other in any manner. The nucleotides are typically attached by their sugar and phosphate groups as in nucleic acids. The nucleotides may be connected via their nucleobases as in pyrimidine dimers.

[0062] The polynucleotide can be a nucleic acid, such as deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). The polynucleotide can comprise one strand of RNA hybridized to one strand of DNA. The polynucleotide may be any synthetic nucleic acid known in the art, such as peptide nucleic acid (PNA), glycerol nucleic acid (GNA), threose nucleic acid (TNA), locked nucleic acid (LNA), bridged nucleic acid (BNA) or other synthetic polymers with nucleotide side chains. The PNA backbone is composed of repeating N-(2-aminoethyl)-glycine units linked by peptide bonds. The GNA backbone is composed of repeating glycol units linked by phosphodiester bonds. The TNA backbone is composed of repeating threose sugars linked together by phosphodiester bonds. LNA is formed from ribonucleotides as discussed above having an extra bridge connecting the 2' oxygen and 4' carbon in the ribose moiety.

[0063] The polynucleotide is preferably DNA, RIMA or a DNA or RNA hybrid, most preferably DNA. A DNA / RNA hybrid may comprise DNA and RNA on the same strand. Preferably, the DNA / RNA hybrid comprises one DNA strand hybridized to an RNA strand.

[0064] The target polynucleotide can be single stranded, double stranded or a mixture of both.

[0065] The target polynucleotide can be any length. For example, the target polynucleotide can be at least about 10, at least about 50, at least about 100, at least about 150, at least about 200, at least about 250, at least about 300, at least about 400 or at least about 500 nucleotides or nucleotide pairs in length. The target polynucleotide can be 1000 or more nucleotides or nucleotide pairs in length, such as about 1.00 x 104or more, about 1.00 x 105or more, about 1.00 x 106or more, about 1.50 x 106or more, or about 2.00 x 106or more nucleotides or nucleotide pairs in length.

[0066] The target polynucleotide can be derived from any source. The target polynucleotide may derived from one or more cells. The one or more cells may be any type of cells. The one or more cells may be one or more prokaryotic cells. The one or more cells may be bacterial or archaeal cells. The one or more cells are typically eukaryotic cells. The one or more cells may be protozoan, algal, fungal, plant or animal cells.

[0067] The one or more eukaryotic cells may be any type of cells. The one or more eukaryotic cells may be protozoan, algal, fungal, plant or animal cells. For example, the one or more eukaryotic cells may be mammalian cells, such as a human, dog, cat, primate, horse, murine, rat, rodent, bovine, murine, porcine, or ovine cells. The one or more eukaryotic cells are preferably human cells. The cells may be plant cells, such as cereal, legume, fruit, or vegetable cells. Examples include, but are not limited to, wheat, barley, oat, canola, maize, soya, rice, banana, apple, tomato, potato, grape, tobacco, bean, lentil, sugar cane, cocoa, cotton, tea, or coffee cells.

[0068] The animal cells may be derived from the ectoderm, endoderm, or mesoderm. The one or more eukaryotic cells may be stem cells, such as embryonic stem cells, induced pluripotent stem cells or mesenchymal stem cells, bone cells, such as osteoclasts, osteoblasts or osteocytes, tendon cells, such as tenoblasts or tenocytes, chondrocytes, synovial cells, vascular cells, blood cells, such as red blood cells, immune cells, platelet, neutrophils or basophils, muscle cells, such as skeletal muscle cells, cardiac muscle cells or smooth muscle cells, reproductive cells, such as sperm, oocytes, duct cells or epididymal cells, secretory cells, adipocytes, liver lipocytes, epithelial cells, odontoblasts, cementoblasts, hormone- secreting cells, barrier cells, exocrine secretory epithelial cells, nerve cells, astrocytes, oligodendrocytes, or neurons.

[0069] The immune cells may be neutrophil granulocyte and precursors, such as myeloblasts, promyelocytes, myelocytes, or metamyelocytes, eosinophil granulocyte and precursors, basophil granulocyte and precursors, mast cells, leukocytes, lymphocytes, helper T cells, regulatory T cells, cytotoxic T cells, natural killer T cells, B cells, macrophages, dendritic cells, plasma cells, neutrophils, or monocytes.

[0070] The cells may be wild-type or naturally occurring. The cells may be genetically modified or genetically engineered. For instance, the immune cells may be genetically engineered to express a recombinant chimeric antigen receptor (CAR) or T cell receptor (TCR). The cells may be genetically modified or genetically engineered using transduction or transfection or any other common techniques known to those skilled in the art.

[0071] The one or more eukaryotic cells may be healthy cells or obtained from a healthy donor or source. The one or more eukaryotic cells may be diseased or damaged, associated with a disease or damage or obtained from a diseased or damaged donor or source.

[0072] The target polynucleotide may be derived from a pathogenic agent. The target polynucleotide may be derived from one or more bacterial cells, one or more archaeon cells, one or more fungal cells, or one or more viral particles.

[0073] The one or more bacterial cells may be Gram negative or Gram positive. The one or more Gram-positive bacterial cells are preferably from the genus Bacillus, Clostridium, Enterococcus, Mycobacterium, Staphylococcus or Streptococcus. The one or more Grampositive bacterial cells may be from the genus Pasteurella or Nocardia.

[0074] The one or more Gram-negative bacterial cells are preferably from the genus Aggregatibacter, Bacteroides, Bartonella, Brucella, Campylobacter, Chylamidia, Enterbacter, Francisella, Haemophilus, Heliobacter, Klebsiella, Legionella, Moraxella, Neisseria, Porphyromonas, Pseudomonas, Salmonella, Serratia, Stenotrophomonas, Vibrio or Yersinia. The one or more Gram-negative bacterial cells may be from the genus Escherichia or Pseudomonas.

[0075] The one or more bacterial cells may be from the genus Borrelia, Chlamydophila, Listeria, Mycoplasma, Proteus, or Treponema. The one or more bacterial cells are preferably Aggregatibacter actinomycetemcomitans, Bacillus anthracis, Bacillus licheniformis, Bacteroides fragilis, Bartonella henselae, Bordetella pertussis, Borrelia burgdorferi, Brucella abortus, Campylobacter jejuni, Chlamydia trachomatis, Chlamydophila pneumoniae, Clostridium difficile, Clostridium perfringens, Enterobacter aerogenes, Enterococcus faecalis, Enterococcus faecium, Francisella tularensis, Haemophilus influenzae, Helicobacter pylori, Klebsiella oxytoca, Legionella pneumophila, Listeria monocytogenes, Moraxella catarrhalis, Mycobacterium avium, Mycobacterium bovis, Mycoplasma genitalium, Mycoplasma pneumoniae, Neisseria gonorrhoeae, Neisseria meningitidis, Porphyromonas gingivalis, Proteus mirabilis, Pseudomonas aeruginosa, Salmonella enter ica, Serratia marcescens, Staphylococcus aureus, Staphylococcus epidermidis, Staphylococcus haemolyticus, Stenotrophomonas maltophilia, Streptococcus mutans, Streptococcus pyogenes, Streptococcus salivarius, Streptococcus sanguinis, Treponema pallidum, Vibrio cholera, Vibrio parahaemolyticus or Yersinia enterocolitica.

[0076] Other specific examples of one or more bacterial cells one or more include, but are not limited, to Mycobacterium tuberculosis, Mycobacterium intracellilare, Mycobacterium kansaii, Mycobacterium gordonae, Streptococcus agalactiae, Streptococcus viridans group, Streptococcus faecalis, Streptococcus bovis, Streptococcus pneumoniae, Corynebacterium diptheriae, Erysipelothrix rhusiopathie, Clostridium tetani, Klebsiella pneumoniae, Pasteurella multocida, Fusobacterium nucleatum, Streptobacillus moniliformis, Treponema pertenue and Actinomyces israelii.

[0077] The one or more fungal cells are preferably from the genus Absidia, Acremonium, Aspergillus, Aureobasidium, Basidiobolus, Blastomyces, Blastoschizomyces, Candida, Cladosporium, Coccidioides, Cryptococcus, Cunninghamella, Curvularia, Debaryomyces, Exophiala, Exserohilum, Fonsecea, Fusarium, Geotrichum, Histoplasma, Issatchenkia, Kluyveromyces, Malezzesia, Mucor, Paracoccidioides, Paecilomyces, Penicillium, Pichia, Pneumocystis, Rhizomucor, Rhizopus, Rhodotorula, Saccharomyces, Scedosporium, Schizophyllum, Scopulariopsis, Sporothrix, Trichoderma, Trichophyton or Trichosporon. The fungus is preferably Aspergillus fumigatus, Aspergillus flavus, Aspergillus lentulus, Aspergillus terreus, Aspergillus nidulans, Aspergillus oryzae, Aspergillus niger, Candida albicans, Candida caribbica (Candida fermentati), Candida dubliniensis, Candida famata (Debaryomyces hansenii), Candida fukuyamaensis (Candida xestobii or Candida carpophila), Candida guilliermondii, Candida kef yr (Kluyveromyces marxianus), Candida krusei (Issatchenkia orientalis), Candida metapsilosis, Candida orthopsilosis, Candida parapsilosis, Candida parapsilosis, Candida pelliculosa, Candida psychrophila, Candida rugosa, Candida smithsonii, Candida tropicalis, Candida utilis, Coccidioides immitis , Cryptococcus bacillisporus, Cryptococcus gattii, Cryptococcus grubii, Cryptococcus neoformans, Debaryomyces coudertii, Debaryomyces maramus, Debaryomyces nepalensis, Debaryomyces prosopidis, Debaryomyces robertsiae, Debaryomyces udenii, Histoplasma capsulatum, Kluyveromyces lactis, Pichia cecembensis, Rhodotorula araucariae, Rhodotorula babjevae, Rhodotorula dairensis, Rhodotorula diobovatum, Rhodotorula glutinis, Rhodotorula kratochvilovae, Rhodotorula paludigenum, Rhodotorula sphaerocarpum, Rhodotorula toruloides, Rhodotorula mucliaginosa, Saccharomyces 'sensu stricto', Saccharomyces bayanus, Saccharomyces boulardii, Saccharomyces cariocanus, Saccharomyces kudriavzevii, Saccharomyces mikatae, Saccharomyces paradoxus, Saccharomyces pastorianus, Saccharomyces uvarum, Saccharomyces cerevisiae or Tsuchiyaea wingfieldii.

[0078] The one or more viral particles may belong to the family Retroviridae, such as human deficiency viruses, such as HIV-I (also referred to as HTLV- III), HIV-II, LAC, IDLV-III / LAV, HIV-III or other isolates such as HIV-LP, the family Picornaviridae, such as poliovirus, hepatitis A, enteroviruses, human Coxsackie viruses, rhinoviruses, echoviruses, the family Calciviridae, such as viruses that cause gastroenteritis, the family Togaviridae, such as equine encephalitis viruses and rubella viruses, the family Flaviviridae, such as dengue viruses, encephalitis viruses and yellow fever viruses, the family Coronaviridae, such as coronaviruses, including SARS-Cov-2 (COVID-19), the family Rhabdoviridae, such as vesicular stomata viruses and rabies viruses, the family Filoviridae, such as Ebola viruses, the family Paramyxoviridae, such as parainfluenza viruses, mumps viruses, measles virus and respiratory syncytial virus, the family Orthomyxoviridae, such as influenza viruses, the family Bungaviridae, such as Hataan viruses, bunga viruses, phleoboviruses and Nairo viruses, the family Arena viridae, such as hemorrhagic fever viruses, the family Reoviridae, such as reoviruses, orbiviruses and rotaviruses, the family Bimaviridae, the family Hepadnaviridae, such as hepatitis B virus, the family Parvoviridae, such as parvoviruses, the Papovaviridae, such as papilloma viruses and polyoma viruses, the family Adenoviridae, such as adenoviruses, the family Herpesviridae, such as herpes simplex virus (HSV) I and II, varicella zoster virus and pox viruses, or the family Iridoviridae, such as African swine fever virus). The one or more viral particles may be an unclassified virus, such as the etiologic agents of Spongiform encephalopathies, the agent of delta hepatitis, the agents of non-A, non-B hepatitis (class 1 enterally transmitted; class 2 parenterally transmitted such as Hepatitis C), Norwalk and related viruses and astroviruses.

[0079] The target polynucleotide may be derived from any number of one or more cells or particles, such as at least about 1.00 x 104, at least about 1.00 x 105, at least about 1.00 x 106, at least about 1.00 x 107, at least about 1.00 x 108, at least about 1.00 x 109, or at least about 1.00 x 1010cells or particles. The target polynucleotide may be derived from even more cells or particles, such as at least about 1011or at least about 1012cells or particles.

[0080] Chromatin

[0081] In a preferred embodiment, the target polynucleotide comprises chromatin. Chromatin is a mixture of DNA and proteins in a variety of cells. Chromatin forms the chromosomes in eukaryotic cells. Bacteria cells also comprise chromatin that is compacted into a membrane- free region known as the nucleoid that changes shape and composition depending on the bacterial state. Chromatin and chromosomes of fungi are highly diverse and dynamic, even within a species. Histones have also been found to be encoded in some DNA viruses and have been shown to form nucleosome-like structures in viruses in vitro.

[0082] The chromatin is typically eukaryotic. It may be derived from any of the one or more eukaryotic cells described above. The chromatin may be bacterial, fungal or viral. It may be derived from any of the one or more bacterial cells, one or more fungal cells or one or more viral particles. These may be any of the bacteria, fungi or viruses described above.

[0083] The one or more bound proteins are typically the one or more bound proteins naturally found in chromatin. The one or more bound proteins are typically one or more bound histones and / or one or more bound nucleosomes. These are discussed in more detail below.

[0084] The chromatin can be any length. For example, the chromatin can be at least about 10, at least about 50, at least about 100, at least about 150, at least about 200, at least about 250, at least about 300, at least about 400 or at least about 500 nucleotide pairs in length. The chromatin can be 1000 or more nucleotide pairs in length, such as about 1.00 x 104or more, about 1.00 x 105or more, about 1.00 x 106or more, about 1.50 x 106or more, or about 2.00 x 106or more nucleotide pairs in length.

[0085] The method of the invention comprises retaining one or more bound proteins on the chromatin until the chromatin is contacted with the transmembrane pore. Any amount of one or more bound proteins may be retained on the chromatin after its extraction from one or more cells, such as at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, or at least about 95% of the one or more proteins bound to the chromatin in the one or more cells. The one or more proteins preferably comprise one or more histones and / or one or more nucleosomes.

[0086] Deriving or extracting chromatin from one or more cells is routine in the art and can be conducted using known methods. For instance, suitable methods are described in Kuznetsov VI, Haws SA, Fox CA, Denu JM. General method for rapid purification of native chromatin fragments. J Biol Chem. 2018 Aug 3;293(31): 12271-12282 (doi:

[0087] 10.1074 / jbc.RA118.002984. Epub 2018 May 24. PMID: 29794135; PMCID: PMC6078465); Sullivan AE, Santos SDM. An Optimized Protocol for ChlP-Seq from Human Embryonic Stem Cell Cultures. STAR Protoc. 2020 Sep 18;l(2): 100062 (doi: 10.1016 / j.xpro.2020.100062. PMID: 33000002; PMCID: PMC7501726); and Thorne AW, Myers FA, Hebbes TR. Native chromatin immunoprecipitation. Methods Mol Biol. 2004;287:21-44 (doi: 10.1385 / 1-59259- 828-5:021. PMID: 15273401) (incorporated herein by reference in their entireties). The method of the invention preferably comprises deriving or extracting chromatin from one or more cells and retaining one or more bound proteins on the chromatin until the chromatin is contacted with the transmembrane pore. The one or more cells may be any of those described above. The one or more bound proteins are typically one or more bound histones and / or one or more bound nucleosomes.

[0088] Chromatin derivation typically requires the one or more cells to be lysed. A method for doing this is described in Kuznetsov et al., supra. The method of the invention preferably comprises lysing one or more cells, deriving or extracting chromatin from the one or more lysed cells and retaining one or more bound proteins on the chromatin until the chromatin is contacted with the transmembrane pore. The one or more cells may be any of those described above. The one or more bound proteins are typically one or more bound histones and / or one or more bound nucleosomes.

[0089] The invention also provides a method of preparing chromatin for characterisation using a transmembrane pore, wherein the method comprises retaining one or more bound proteins on the chromatin until the chromatin is contacted with the transmembrane pore. The invention also provides a method of preparing a sample comprising chromatin for characterisation using a transmembrane pore, wherein the method comprises retaining one or more bound proteins on the chromatin until the chromatin is contacted with the transmembrane pore. The one or more bound proteins are typically one or more bound histones and / or one or more bound nucleosomes.

[0090] Binding one or more proteins

[0091] In one embodiment, the one or more bound proteins are added or bound to the target polynucleotide before the method is conducted. The one or more bound proteins may be added or bound to the target polynucleotide outside of the context of one or more cells. The one or more bound proteins preferably do not naturally bind to the target polynucleotide and are bound to the target polynucleotide as part of its preparation. The one or more bound proteins preferably do not bind to the target polynucleotide in its natural state and are bound to the target polynucleotide as part of its preparation. The one or more bound proteins are preferably added or bound to the target polynucleotide to stabilise the target polynucleotide. The one or more bound proteins may replace one or more bound proteins that appear on the target polynucleotide in nature. Stabilisation of target polynucleotides is discussed in more detail below. The method preferably comprises the initial step of binding the one or more proteins to the target polynucleotide.

[0092] In this embodiment, the one or more bound proteins preferably do not comprise a movement control protein. Movement control proteins are known in the art and are discussed in more detail below. In this embodiment, the target polynucleotide is preferably DNA, RIMA or an artificial or nonnatural polynucleotide. The target polynucleotide may be any of the polynucleotides discussed in more detail below. The target polynucleotide may be derived from one or more cells as discussed above and additional and / or different one or more bound proteins are bound to the target polynucleotide.

[0093] The invention provides a method of preparing a target polynucleotide for characterisation using a transmembrane pore, the method comprising: (a) binding one or more proteins to the target polynucleotide before the target polynucleotide is contacted with the transmembrane pore and (b) attaching a movement control protein to the target polynucleotide before the target polynucleotide is contacted with the transmembrane pore. The invention also comprises a method of preparing a sample comprising a target polynucleotide for characterisation using a transmembrane pore, the method comprising:

[0094] (a) binding one or more proteins to the target polynucleotide before the target polynucleotide is contacted with the transmembrane pore and (b) attaching a movement control protein to the target polynucleotide before the target polynucleotide is contacted with the transmembrane pore.

[0095] In some embodiments, the method may comprise removing one or more bound proteins, such as one or more histones and / or one or more nucleosomes, from the target polynucleotide, such as chromatin, before binding the one or more proteins to the target polynucleotide. The one or more proteins bound to the target polynucleotide are typically different from the one or more proteins removed from the target polynucleotide.

[0096] The invention provides a method of preparing a target polynucleotide, such as chromatin, for characterisation using a transmembrane pore, the method comprising: (a) removing one or more bound proteins from the target polynucleotide, (b) binding one or more proteins, such as one or more different proteins, to the target polynucleotide before the target polynucleotide is contacted with the transmembrane pore and (b) attaching a movement control protein to the target polynucleotide before the target polynucleotide is contacted with the transmembrane pore. The invention also comprises a method of preparing a sample comprising a target polynucleotide for characterisation using a transmembrane pore, the method comprising: (a) removing one or more proteins from the target polynucleotide,

[0097] (b) binding one or more proteins, such as one or more different proteins, to the target polynucleotide before the target polynucleotide is contacted with the transmembrane pore and (c) attaching a movement control protein to the target polynucleotide before the target polynucleotide is contacted with the transmembrane pore. The different one or more proteins may be any of those discussed below. In these embodiments, the one or more proteins removed from the target polynucleotide may be one or more histones and / or one or more nucleosomes. Bound protein(s)

[0098] The target polynucleotide may comprise any number of bound proteins. The target polynucleotide may comprise any number of individual bound proteins and / or different types of bound proteins.

[0099] The target polynucleotide may comprise any number of individual bound proteins, such as about 1 or more, about 2 or more, about 3 or more, about 4 or more, about 5 or more, about 6 or more, about 7 or more, about 8 or more, about 9 or more, about 10 or more, about 20 or more, about 30 or more, about 40 or more, about 50 or more, about 100 or more, about 200 or more, about 500 or more, about 1000 or more, about 104or more, about 105or more, about 106or more, about 107or more, about 108or more, about 109or more, about 1010or more, about 1011or more, or about 1012or more individual proteins. For instance, the target polynucleotide may comprise any of these numbers of histone H2A, including all of the subtypes described below.

[0100] The target polynucleotide may comprise any number of different types of bound proteins, such as such as about 1 or more, about 2 or more, about 3 or more, about 4 or more, about 5 or more, about 6 or more, about 7 or more, about 8 or more, about 9 or more, about 10 or more, about 20 or more, about 30 or more, about 40 or more, about 50 or more, about 100 or more, about 200 or more, about 500 or more or about 1000 or more different types of protein. For instance, the target polynucleotide may comprise any of these numbers of different histones selected from the subtypes described below and / or other bound proteins described below.

[0101] The one or more bound proteins may bind the target polynucleotide in any way. The one or more bound proteins may comprise one or more polynucleotide binding domains. The one or more polynucleotide binding domains may be selected from zinc finger (ZF) domains, helix-turn-helix domains, leucine zippers, transcription activator like effectors (TALEs) and antibody domains. The one or more bound proteins may comprise any number of polynucleotide binding domains, such as about 1 or more, about 2 or more, about 3 or more, about 4 or more, about 5 or more, about 6 or more, about 7 or more, about 8 or more, about 9 or more or about 10 or more.

[0102] The one or more bound proteins may be one or more naturally-occurring proteins. The one or more bound proteins may be one or more artificial proteins. The one or more bound proteins may be one or more mutants or one or more variants of naturally-occurring proteins. Mutants and variants are defined above. The one or more bound proteins may be a mixture of one or more naturally-occurring proteins and one or more artificial proteins, mutants, or variants. The one or more bound proteins may comprise one or more polynucleotide antibodies or fragments thereof. The term "antibody" includes whole antibodies. Naturally occurring antibodies typically comprise a tetramer which is usually composed of at least two heavy (H) chains and at least two light (L) chains. Each heavy chain is comprised of a heavy chain variable region (abbreviated herein as VH) and a heavy chain constant region, usually comprised of three domains (CHI, CH2 ad CH3). Heavy chains can be of any isotype, including IgG (IgGl, IgG2, IgG3 and IgG4 subtypes), IgA (IgAl and IgA2 subtypes), IgM and IgE. Each light chain is comprised of a light chain variable region (abbreviated herein as VL) and a light chain constant region (CL). Light chain includes kappa (K) chains and lambda (A) chains. The heavy and light chain variable region is typically responsible for antigen recognition, whilst the heavy and light chain constant region may mediate the binding of the immunoglobulin to host tissues or factors, including various cells of the immune system (e.g., effector cells) and the first component (Clq) of the classical complement system. The VH and VL regions can be further subdivided into regions of hypervariability, termed complementarity determining regions (CDR), interspersed with regions that are more conserved, termed framework regions (FR). Each VH and VL is composed of three CDRs and four FRs arranged from amino-terminus to carboxy-terminus in the following order: FR1, CDR1, FR2, CDR2, FR3, CDR3, FR4. The variable regions of the heavy and light chains contain a binding domain that interacts with an antigen.

[0103] The term "functional fragment" refers to a fragment of an intact antibody that retains the ability to specifically bind to a given molecule or target analyte. Such fragments include Fab fragments, Fab' fragments, monovalent fragments consisting of the VL, VH, CL and CHI domains; F(ab')2 fragments, bivalent fragments comprising two Fab fragments linked by a disulfide bridge at the hinge region; Fd fragment consisting of the VH and CHI domains; Fv fragments consisting of the VL and VH domains of a single arm of an antibody; a dAb fragment (Ward et al., 1989 Nature 341 :544-546), which consists of a VH domain; and an isolated complementarity determining region (CDR). Furthermore, although the two domains of the Fv fragment, VL and VH, are coded for by separate genes, they can be joined, using recombinant methods, by a synthetic linker that enables them to be made as a single protein chain in which the VL and VH regions pair to form monovalent molecules (known as single chain Fv (scFv); see e.g., Bird et al., 1988 Science 242:423-426; and Huston et al., 1988 Proc. Natl. Acad. Sci. 85:5879-5883). Such single chain antibodies are also intended to be encompassed within the term "functional fragment" of an antibody. These antibody fragments are obtained using conventional techniques known to those of skill in the art, and the fragments are screened for utility in the same manner as are intact antibodies.

[0104] The one or more bound proteins are preferably one or more polynucleotide binding proteins. A polynucleotide binding protein is any protein that is capable of binding to a polynucleotide. Such proteins are known in the art. The one or more bound proteins preferably do not comprise a movement control protein. The one or more bound proteins preferably do not comprise a polymerase, an exonuclease, a translocase, a helicase, or a topoisomerase. In other words, the one or more bound proteins preferably do not comprise any of a polymerase, an exonuclease, a translocase, a helicase, and a topoisomerase.

[0105] The one or more bound polynucleotide binding proteins preferably comprise (a) one or more bound histones, (b) one or more bound transcription factors, (c) one or more bound proteins involved in DNA modelling or remodelling, replication, transcription and / or repair, or (d) combinations thereof, such as (a); (b); (c); (d); (a) and (b); (a) and (c); (a) and (d);

[0106] (b) and (c); (b) and (d); (c) and (d); (a), (b) and (c); (a), (b) and (d); (a), (c) and (d); (b),

[0107] (c) and (d); or (a), (b), (c) and (d).

[0108] The one or more bound proteins preferably comprise one or more bound histones and / or one or more bound transcription factors.

[0109] The target polynucleotide preferably comprises one or more histones. The chromatin typically comprises one or more histones, preferably a plurality of histones. The target polynucleotide or chromatin may comprise any number of histones, such as about 1 or more, about 2 or more, about 3 or more, about 4 or more, about 5 or more, about 6 or more, about 7 or more, about 8 or more, about 9 or more, about 10 or more, about 20 or more, about 30 or more, about 40 or more, about 50 or more, about 100 or more, about 200 or more, about 500 or more, about 1000 or more, about 104or more, about 105or more, about 106or more, about 107or more, about 108or more, about 109or more, about 1010, about more 1011or more, or about 1012or more histones.

[0110] The one or more histones may be any histones. The one or more histones may be selected the families Hl, H2A, H2B, H3 and H4. The one or more histones are preferably selected the families H2A, H2B, H3 and H4. The one or more histones may be selected from Hl-1, Hl-2, Hl-3, Hl-4, Hl-5, Hl-6, Hl-0, Hl-7, Hl-8, Hl-10, H2AC1, H2AC4, H2AC6, H2AC7, H2AC8, H2AC11, H2AC12, H2AC13, H2AC14, H2AC15, H2AC16, H2AC17, H2AC18, H2AC19, H2AC20, H2AC21, H2AC25, H2AZ1, H2AZ2, MACROH2A1, MACROH2A2, H2AX, gamma H2AX, H2AJ, H2AB1, H2AB2, H2AB3, H2AP, H2AL1Q, H2AL3, H2AC2P, H2AC3P, H2AC5P, H2AC9P, H2AC10P, H2AQ1P, H2AL1MP, H2BC1, H2BC3, H2BC4, H2BC5, H2BC6, H2BC7, H2BC8, H2BC9, H2BC10, H2BC11, H2BC12, H2BC13, H2BC14, H2BC15, H2BC17, H2BC18, H2BC21, H2BC26, H2BC12L, H2BK1, H2BW1, H2BW2, H2BW3P, H2BN1, H2BC2P, H2BC16P, H2BC19P, H2BC20P, H2BC27P, H2BL1P, H2BW3P, H2BW4P, H3C1, H3C2, H3C3, H3C4, H3C6, H3C7, H3C8, H3C10, H3C11, H3C12, H3C13, H3C14, H3C15, H3-4, H3-3A, H3-3B, H3-5, H3-7, H3Y1, H3Y2, CENPA, H3C5P, H3C9P, H3P16, H3P44, H4C1, H4C2, H4C3, H4C4, H4C5, H4C6, H4C7, H4C8, H4C9, H4C11, H4C12, H4C13, H4C14, H4C15, H4C16 and H4C10P. The one or more histones may be multimers, such as dimers, trimers, or tetramers. Any of the histones may be modified, for instance by acetylation, methylation, phosphorylation, and / or ubiquitylation.

[0111] The target polynucleotide may comprise one or more bound nucleosomes. The chromatin typically comprises one or more bound nucleosomes. The target polynucleotide or chromatin may comprise any number of bound nucleosomes, such as about 1 or more, about 2 or more, about 3 or more, about 4 or more, about 5 or more, about 6 or more, about 7 or more, about 8 or more, about 9 or more, about 10 or more, about 20 or more, about 30 or more, about 40 or more, about 50 or more, about 100 or more, about 200 or more, about 500 or more, about 1000 or more, about 104or more, about 105or more, about 106or more, about 107or more, about 108or more, about 109or more, about 1010or more, about 1011or more, or about 1012or more bound nucleosomes. Each bound nucleosome typically comprises two H2A-H2B dimers and one H3-H4 tetramer. The H2A histones in a bound nucleosome are preferably selected from H2AC1, H2AC4, H2AC6, H2AC7, H2AC8, H2AC11, H2AC12, H2AC13, H2AC14, H2AC15, H2AC16, H2AC17, H2AC18, H2AC19, H2AC20, H2AC21, H2AC25, H2AZ1, H2AZ2, MACROH2A1, MACROH2A2, H2AX, gamma H2AX, H2AJ, H2AB1, H2AB2, H2AB3, H2AP, H2AL1Q, H2AL3, H2AC2P, H2AC3P, H2AC5P, H2AC9P, H2AC10P, H2AQ1P, and H2AL1MP. The H2B histones in a bound nucleosome are preferably selected from H2BC1, H2BC3, H2BC4, H2BC5, H2BC6, H2BC7, H2BC8, H2BC9, H2BC10, H2BC11, H2BC12, H2BC13, H2BC14, H2BC15, H2BC17, H2BC18, H2BC21, H2BC26, H2BC12L, H2BK1, H2BW1, H2BW2, H2BW3P, H2BN 1, H2BC2P, H2BC16P, H2BC19P, H2BC20P, H2BC27P, H2BL1P, H2BW3P, and H2BW4P. The H3 histones in a bound nucleosome are preferably selected from H3C1, H3C2, H3C3, H3C4, H3C6, H3C7, H3C8, H3C10, H3C11, H3C12, H3C13, H3C14, H3C15, H3-4, H3-3A, H3-3B, H3-5, H3-7, H3Y1, H3Y2, CENPA, H3C5P, H3C9P, H3P16, and H3P44. The H4 histones in a bound nucleosome are preferably selected from H4C1, H4C2, H4C3, H4C4, H4C5, H4C6, H4C7, H4C8, H4C9, H4C11, H4C12, H4C13, H4C14, H4C15, H4C16 and H4C10P. Any of the histones may be modified, for instance by acetylation, methylation, phosphorylation, and / or ubiquitylation.

[0112] The one or more transcription factors may be selected from CTCF, FOX, FOS / JUN, NFKB, CEBP, Cas9, dCas9, RecA, Rad51, p53, PhiC31 integrase, sigma 70, sigma 54, Cohesin, Condesin, Hu, Fis, IHF, Tus, and H-NS.

[0113] Sequencing adaptor(s)

[0114] The method preferably further comprises attaching one or more sequencing adaptors to the target polynucleotide. The method preferably further comprises attaching one or more sequencing adaptors to the chromatin. Sequencing adaptors facilitate characterisation or sequencing, especially using transmembrane pores. Any number of sequencing adaptors may be attached to the target polynucleotide or chromatin, such as such as about 1 or more, about 2 or more, about 3 or more, about 4 or more, about 5 or more, about 6 or more, about 7 or more, about 8 or more, about 9 or more or about 10 or more.

[0115] A sequencing adaptor typically comprises a polynucleotide strand capable of being attached to the end of the target polynucleotide or chromatin. A sequencing adaptor may be added to both ends of the target polynucleotide or chromatin. Alternatively, different adaptors may be added to the two ends of the target polynucleotide or chromatin. An adaptor may be added to just one end of the target polynucleotide or chromatin. Methods of adding adaptors to polynucleotides are known in the art. Adaptors may be attached to polynucleotides, for example, by ligation, by click chemistry, by tagmentation, by topoisomerisation or by any other suitable method.

[0116] An adaptor may be synthetic or artificial. Typically, an adaptor comprises a polymer. The adaptor preferably comprises a polynucleotide. An adaptor may comprise a single-stranded polynucleotide strand. An adaptor may comprise a double-stranded polynucleotide. A sequencing adaptor may comprise any of the polynucleotide discussed above with to the target polynucleotide and includes DNA, RIMA, modified DNA (such as a basic DNA), RNA, PNA, LNA, BNA and / or PEG. Usually, the adaptor comprises single stranded and / or double stranded DNA or RNA.

[0117] The one or more sequencing adaptors may be one or more Y adaptors. Y adaptors are typically double stranded and comprise (a) at one end, a region where the two strands are hybridised together and (b), at the other end, a region where the two strands are not complementary. The non-complementary parts of the strands form overhangs. The hybridised stem of the adaptors typically attaches to the 5' end of a first strand of a doublestranded polynucleotide and the 3' end of a second strand of a double-stranded polynucleotide; or to the 3' end of a first strand of a double-stranded polynucleotide and the 5' end of a second strand of a double-stranded polynucleotide. The presence of a non- complementary region in the Y adaptors gives them their Y shape since the two strands typically do not hybridise to each other unlike the double stranded portion. The hybridised stem end of the Y adaptors may also comprise a short overhang that allows them to specifically hybridise to and be attached to the first adaptors, RT primers, second adaptors and barcoded constructs.

[0118] Some of the methods of the invention use a movement control protein to control the movement of the target polynucleotide or chromatin with respect to a transmembrane pore. A movement control protein may bind to an overhang of an adaptor such as a Y adaptor. A movement control protein may bind to the double stranded region. A movement control protein may bind to a single-stranded and / or a double-stranded region of the adaptor. A first movement control protein may bind to the single-stranded region of such an adaptor and a second movement control may bind to the double-stranded region of the adaptor. The one or more sequencing adaptors preferably comprise a membrane anchor and / or a pore anchor. The anchors may be attached to polynucleotides that are complementary to and hence hybridised to the overhangs to which a movement control protein is bound.

[0119] One of the non-complementary strands of the one or more sequencing adaptors, such as Y adaptors, may comprise leader sequences, which when contacted with a transmembrane pore are capable of threading into the pore.

[0120] The leader sequences typically comprise a polymer such as a polynucleotide, for instance DNA or RIMA, a modified polynucleotide (such as abasic DNA), PNA, LNA, polyethylene glycol (PEG) or a polypeptide. The leader sequences preferably comprise a single strand of DNA, such as a poly dT section. The leader sequences can be any length, but are typically 10 to 150 nucleotides in length, such as from 20 to 120, 30 to 100, 40 to 80 or 50 to 70 nucleotides in length.

[0121] The one or more sequencing adaptors may be hairpin loop adaptors. Hairpin loop adaptors are adaptors comprising a single polynucleotide strand, wherein the ends of the polynucleotide strand are capable of hybridising to each other, or are hybridized to each other, and wherein the middle section of the polynucleotide forms a loop. Suitable hairpin loop adaptors can be designed using methods known in the art. Typically, the 3' end of a hairpin loop adaptor attaches to the 5' end of a first strand of a double-stranded polynucleotide and the 5' end of the hairpin loop adaptor attaches to the 3' end of a second strand of a double-stranded polynucleotide; or the 5' end of a hairpin loop adaptor attaches to the 3' end of a first strand of a double-stranded polynucleotide and the 3' end of the hairpin loop adaptor attaches to the 5' end of a second strand of a double-stranded polynucleotide.

[0122] Those skilled in the art will also appreciate that when the one or more sequencing adaptors comprise a polynucleotide strand, the sequences of the adaptors are typically not determinative and can be controlled or chosen according to the movement control protein and other experimental conditions such as the target polynucleotide or chromatin to be characterised. Exemplary sequences are provided solely by way of illustration in the examples. For example, the one or more sequencing adaptors may comprise a sequence such as one or more of SEQ ID NOs: 21-26 or 28-33 in WO 2021 / 255476 (incorporated herein by reference in its entirety) or polynucleotide sequences having at least about 20%, such as at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95% at least about 96%, at least about 97%, at least about 98% or at least about 99% sequence similarity or identity to one or more of SEQ ID NOs: 21-26 or 28-33 in WO 2021 / 255476 (incorporated herein by reference in its entirety). The sequences of the one or more sequencing adaptors can typically be altered without negatively affecting the efficacy of the method of the invention.

[0123] The one or more sequencing adaptors may comprise a loading site for loading the movement control protein. The loading site may be for instance a single-stranded region which can targeted by the movement control protein. The loading site may be a region of the one or more sequencing adaptors to which an exogenous polynucleotide strand comprising the movement control protein can bind in order to transfer the movement control protein to the target polynucleotide or chromatin.

[0124] The movement control protein, if present, may be provided on one or more sequencing adaptors. WO 2015 / 110813 and WO 2020 / 234612 describe the loading of movement control proteins onto a target polynucleotide such as an adaptor and are hereby incorporated by reference in their entireties.

[0125] The one or more sequencing adaptors may comprise a membrane anchor. The anchor typically assists in the characterisation of the target polynucleotide or chromatin. For example, a membrane anchor may promote localisation of the target polynucleotide or chromatin around or near the transmembrane pore.

[0126] The anchor may be a polypeptide anchor and / or a hydrophobic anchor that can be inserted into the membrane. The hydrophobic anchor is preferably a lipid, fatty acid, sterol, carbon nanotube, polypeptide, protein, or amino acid, for example cholesterol, palmitate, or tocopherol. The anchor may comprise thiol, biotin, or a surfactant.

[0127] The anchor may be biotin (for binding to streptavidin), amylose (for binding to maltose binding protein or a fusion protein), Ni-NTA (for binding to poly-histidine or poly-histidine tagged proteins) or peptides (such as an antigen).

[0128] The anchor preferably comprises a linker, or 2, 3, 4 or more linkers. Preferred linkers include, but are not limited to, polymers, such as polynucleotides, polyethylene glycols (PEGs), polysaccharides and polypeptides. These linkers may be linear, branched, or circular. For instance, the linker may be a circular polynucleotide. The adaptor may hybridise to a complementary sequence on a circular polynucleotide linker. The one or more anchors or one or more linkers may comprise a component that can be cut or broken down, such as a restriction site or a photolabile group. The linker may be functionalised with maleimide groups to attach to cysteine residues in proteins. Suitable linkers are described in WO 2010 / 086602 (incorporated herein by reference in its entirety).

[0129] The anchor is preferably cholesterol or a fatty acyl chain. For example, any fatty acyl chain having a length of from 6 to 30 carbon atom, such as hexadecanoic acid, may be used. Examples of suitable anchors and methods of attaching anchors to adaptors are disclosed in WO 2012 / 164270 and WO 2015 / 150786 (incorporated herein by reference in their entireties).

[0130] The anchor may consist or comprise a hydrophobic modification to the polynucleotide or sequencing adaptor. The hydrophobic modification may comprise a modified phosphate group comprised within the polynucleotide or polynucleotide anchor. The hydrophobic modification may for example comprise a phosphorothioate such as a charge-neutralized alkyl-phosphorothioate (PPT) as described in Jones et al, J. Am. Chem. Soc. 2021, 143, 22, 8305, the entire contents of which are hereby incorporated by reference. Suitable alkyl groups include for example Ci-Ci0alkyl groups such as C2-C6alkyl groups, e.g., methyl, ethyl, propyl, butyl, pentyl and hexyl groups. Incorporation of the charge-neutralized alkyl- phosphorothioate into a polynucleotide allows for the polynucleotide to anchor to a hydrophobic region such as a membrane or lipid bilayer.

[0131] Movement control proteins

[0132] In step (b), the preparation method comprises attaching a movement control protein to the target polynucleotide or chromatin. The movement control protein can be attached to the target polynucleotide or chromatin via one or more sequencing adaptors. This is described above.

[0133] The invention also provides a method of preparing chromatin for characterisation using a transmembrane pore, wherein the method comprises (a) retaining one or more bound proteins on the chromatin until the chromatin is contacted with the transmembrane pore and (b) attaching a movement control protein to the chromatin before the chromatin is contacted with the transmembrane pore.

[0134] The invention also comprises a method of preparing a sample comprising chromatin, wherein the chromatin, wherein the method comprises (a) retaining one or more bound proteins on the chromatin until the chromatin is contacted with the transmembrane pore and (b) attaching a movement control protein to the chromatin before the chromatin is contacted with the transmembrane pore.

[0135] As those skilled in the art will appreciate, any suitable movement control protein can be used in the method of the invention. The movement control protein may be any protein that is capable of binding to a polynucleotide and controlling its movement with respect to a detector, e.g., a transmembrane pore.

[0136] In more detail, movement control proteins such as helicases can typically control the movement of polynucleotides in at least two active modes of operation (when is provided with all the necessary components to facilitate movement e.g., ATP and Mg2+) and one inactive mode of operation (when not provided with the necessary components to facilitate movement; or when the movement control protein is modified in order to prevent the active mode).

[0137] When provided with all the necessary components to facilitate movement, a movement control protein may move along a polynucleotide, such as DNA, in either a 5'-3' direction or a 3'-5' direction. Many movement control proteins process polynucleotides, such as DNA, in a 5'-3' direction. Movement control proteins which control the movement of polynucleotides in this manner are typically suitable for use in the method of the invention.

[0138] However, when a movement control protein is not provided with the necessary components to facilitate movement or is modified in order to prevent it from actively controlling the movement of the polynucleotide with respect to the transmembrane pore, it can still passively control the movement of the polynucleotide with respect to the transmembrane pore. For example, the movement control protein can bind to the polynucleotide and act as a brake slowing the movement of the polynucleotide when it is pulled into the pore by an applied field (e.g., by the first force in the method of the invention). In the "inactive" mode it typically does not matter whether the polynucleotide is captured either 3' or 5' down ( / .e., moves through the transmembrane pore in a 5'-3' direction or in a 3'-5' direction), as the applied force provides the impetus to move the polynucleotide through the transmembrane pore. However, in such embodiments, the movement control protein may still control the movement of the polynucleotide with respect to the transmembrane pore e.g., by acting as a brake. When in the inactive mode the movement control of a polynucleotide by a movement control protein can be described in a number of ways including ratcheting, sliding, and braking. Typically the method of the invention do not comprise the use of a movement control protein operating in the passive mode. However, when a movement control protein the movement control protein is used, it may be a movement control protein operating in the passive mode.

[0139] Some methods of the invention may comprise use of a movement control protein as a pausing moiety to impede the movement of the polynucleotide strand through the transmembrane pore. The movement control protein may be a protein which binds to polynucleotides but which does not have polynucleotide processing capacity, i.e., it is not a polynucleotide handling enzyme.

[0140] A polynucleotide-handling enzyme is a polypeptide that is capable of interacting with a polynucleotide. The enzyme may modify the polynucleotide by cleaving it to form individual nucleotides or shorter chains of nucleotides, such as di- or trinucleotides. The enzyme may modify the polynucleotide by orienting it or moving it to a specific position. A movement control protein as used herein may be, or may be derived from a polynucleotide handling enzyme. A movement control protein may be, or may be derived from a polynucleotide- handling enzyme.

[0141] The movement control protein may be derived from a member of any of the Enzyme Classification (EC) groups 3.1.11, 3.1.13, 3.1.14, 3.1.15, 3.1.16, 3.1.21, 3.1.22, 3.1.25, 3.1.26, 3.1.27, 3.1.30 and 3.1.31.

[0142] Typically, the movement control protein is a helicase, a polymerase, an exonuclease, a topoisomerase, or a variant thereof.

[0143] The movement control protein may be or may be derived from an exonuclease. Suitable enzymes include, but are not limited to, exonuclease I from E. coli, exonuclease III enzyme from E. coli, RecJ from T. thermophilus and bacteriophage lambda exonuclease, TatD exonuclease and variants thereof.

[0144] The movement control protein may be a polymerase. The polymerase may be PyroPhage® 3173 DNA Polymerase (which is commercially available from Lucigen® Corporation), SD Polymerase (commercially available from Bioron®), Klenow from NEB or variants thereof. In one embodiment, the enzyme is Phi29 DNA polymerase or a variant thereof. Modified versions of Phi29 polymerase that may be used in the invention are disclosed in US Patent No. 5,576,204.

[0145] The movement control protein may be a topoisomerase. In one embodiment, the topoisomerase is a member of any of the Moiety Classification (EC) groups 5.99.1.2 and 5.99.1.3. The topoisomerase may be a reverse transcriptase, which are enzymes capable of catalysing the formation of cDNA from a RNA template. They are commercially available from, for instance, New England Biolabs® and Invitrogen®.

[0146] The movement control protein is preferably a helicase. Any suitable helicase can be used in accordance with the method of the invention. For example, the or each enzyme used in accordance with the present disclosure may be independently selected from a Hel308 helicase, a RecD helicase, a Tral helicase, a TrwC helicase, an XPD helicase, and a Dda helicase, or a variant thereof. Monomeric helicases may comprise several domains attached together. For instance, Tral helicases and Tral subgroup helicases may contain two RecD helicase domains, a relaxase domain and a C-terminal domain. The domains typically form a monomeric helicase that is capable of functioning without forming oligomers. Particular examples of suitable helicases include Hel308, NS3, Dda, UvrD, Rep, PcrA, Pifl and Tral.

[0147] These helicases typically work on single stranded DNA. Examples of helicases that can move along both strands of a double stranded DNA include FtsK and hexameric enzyme complexes, or multisubunit complexes such as RecBCD. The movement control protein is preferably a Dda (DNA-dependent ATPase) helicase. Hel308 helicases are described in publications such as WO 2013 / 057495, the entire contents of which are incorporated by reference. RecD helicases are described in publications such as WO 2013 / 098562, the entire contents of which are incorporated by reference. XPD helicases are described in publications such as WO 2013 / 098561, the entire contents of which are incorporated by reference. Dda helicases are described in publications such as WO 2015 / 055981 and WO 2016 / 055777, the entire contents of each of which are incorporated by reference.

[0148] The helicase may be Trwc Cba or a variant thereof, Hel308 Mbu or a variant thereof or Dda or a variant thereof. Variants may differ from the native sequences in any of the ways discussed herein. An example variant of Dda comprises E94C / A360C. A further example variant of Dda comprises E94C / A360C and then (AM1)G1G2 ( / .e., deletion of Ml and then addition of G1 and G2).

[0149] Suitable movement control proteins are disclosed in WO 2013 / 057495, WO 2013 / 098562, WO2013098561, WO 2014 / 013260, WO 2014 / 013259, WO 2014 / 013262, WO 2015 / 055981, and PCT / EP2024 / 063118. All of these are incorporated by reference in their entirety.

[0150] Additional method steps

[0151] The method may further comprise any additional steps for preparing the target polynucleotide or chromatin for characterisation using a transmembrane pore.

[0152] The method preferably further comprises lysing one or more cells, such as one or more eukaryotic cells, one or more bacterial cells or one or more fungal cells. Methods for doing this are described above.

[0153] The method preferably further comprises one or more of (a) binding the target polynucleotide to a surface, (b) washing the target polynucleotide, and (c) eluting the target polynucleotide, such as (a), (b), (c), (a) and (b), (a) and (c), (b) and (c) or (a), (b) or (c). These step are routine in the art. The surface may be a membrane or a bead. Suitable beads include, but are not limited to, NEB Monarch beads (available as part of the NEB DNA extraction kit). Other methods and surfaces are described in Kuznetsov et al., Sullivan et al. and Thorne et al., supra (incorporated herein by reference in their entireties).

[0154] Method of stabilising a target polynucleotide

[0155] The invention provides a method of preparing a stabilised target polynucleotide. The method comprises binding one or more proteins to the target polynucleotide or retaining one or more bound proteins on the target polynucleotide. The method is preferably for preparing a stabilised target polynucleotide for characterisation using a transmembrane pore. The stabilised target polynucleotide has all of the advantages discussed above. The stabilised target polynucleotide may be stored long term. For instance, the stabilised target polynucleotide may be stored for about at least 1 week, at least about 2 weeks, at least about 3 weeks, at least about 1 month, at least about 3 months, at least about 6 months, at least about 12 months, at least about 2 years, at least about 3 years, or at least about 5 years. The stabilised target polynucleotide may be stored for longer. The stabilised target polynucleotide may be stored using standard storage conditions. For instance, the stabilised target polynucleotide may be stored at about -20°C or about -80°C and / or may undergo multiple freeze-thaw cycles. The stabilised target polynucleotide can stored at about 4°C or at even room temperature.

[0156] The invention also provides a method of increasing the read length of a target polynucleotide using a transmembrane pore, wherein the method comprises binding one or more proteins to the target polynucleotide or retaining one or more bound proteins on the target polynucleotide. Read length is the length of the target polynucleotide that can be characterised or sequenced by the transmembrane pore. The read length may be increased by any amount, such as by at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about a factor of 2, at least about a factor of 3, at least about a factor of 4, at least about a factor of 5, at least about a factor of 6, at least about a factor of 7, at least about a factor of 8, at least about a factor of 9, or at least about a factor of 10. The read length is increased compared with the target polynucleotide without the one or more proteins bound to it.

[0157] The target polynucleotide may be any of those discussed above. Any of the embodiments relating to one or more proteins or one or more bound proteins discussed above equally apply to these embodiments. Preferably, the one or more proteins or one or more bound proteins are one or more histones, one or more transcription factors, or one or more proteins involved in DNA modelling or remodelling, replication, transcription and / or repair. The one or more proteins are preferably one or more histones and / or one or more nucleosomes. The one or more bound proteins are preferably bound to the target polynucleotide before the target polynucleotide is contacted with a transmembrane pore.

[0158] The one or more proteins preferably do not comprise a movement control protein. The one or more proteins preferably do not comprise a polymerase, an exonuclease, a translocase, a helicase, or a topoisomerase. In other words, the one or more proteins preferably do not comprise any of a polymerase, an exonuclease, a translocase, a helicase, and a topoisomerase.

[0159] The method preferably further comprises attaching one or more sequencing adaptors to the target polynucleotide and / or further comprises attaching a movement control protein to the target polynucleotide. Any of the embodiments discussed above equally apply to this embodiment.

[0160] Characterisation

[0161] The method of preparing a target polynucleotide of the invention preferably further comprises characterising or sequencing the target polynucleotide or chromatin. Any method may be used for sequencing or characterising the target polynucleotide or chromatin. The target polynucleotide or chromatin is preferably characterised or sequenced using the transmembrane pore. This is discussed in more detail below.

[0162] The invention also comprises a method of characterising a target polynucleotide comprising one or more bound proteins. The method comprises (a) retaining the one or more bound proteins on the target polynucleotide until the target polynucleotide is contacted with a transmembrane pore. The method comprises preparing the target polynucleotide using a method of the invention and characterising the prepared target polynucleotide. Any of the embodiments relating to the target polynucleotide, the one or more bound proteins and retaining the one or more bound proteins discussed above equally apply to this embodiment (including in the absence of the movement control protein). This includes the one or more bound proteins being naturally bound to the target polynucleotide or the one or bound proteins being artificially bound to the target polynucleotide, for instance to replace one or more removed proteins.

[0163] In a preferred embodiment, step (a) comprises preparing the target polynucleotide using a method of the invention. Any method of the invention discussed above may be used in step (a).

[0164] The characterisation method also comprises (b) contacting the target polynucleotide with the transmembrane pore such that one or more of the bound proteins are stripped from the target polynucleotide as the target polynucleotide moves with respect to the pore. The one or more proteins are typically stripped from the target polynucleotide as the target polynucleotide moves through the transmembrane pore but the one or more proteins are too large to pass through the transmembrane pore. All of the one or more bound proteins are preferably stripped from the target polynucleotide as the target polynucleotide moves with respect to, such as through, the pore. The term "stripped" typically means the one or more proteins are unbound from the target polynucleotide and do not re-bind to the target polynucleotide. In all instances herein, the term "stripped" is interchangeable with "irreversibly stripped".

[0165] The characterisation method also comprises (c) taking one or more measurements as the target polynucleotide moves with respect to the pore wherein the measurements are indicative of one or more characteristics of the target polynucleotide and thereby characterising the target polynucleotide. The invention also provides a method of characterising chromatin. The method comprises (a) retaining one or more bound proteins on the chromatin until the chromatin is contacted with a transmembrane pore. Any of the embodiments relating to chromatin, the one or more bound proteins, including one or more bound histones and / or one or more nucleosomes, and retaining the one or more bound proteins discussed above equally apply to this embodiment (including in the absence of a movement control protein).

[0166] The method preferably comprises preparing the chromatin using a method of the invention and characterising the prepared chromatin. Any method of preparing chromatin of the invention discussed above may be used in step (a).

[0167] The characterisation method also comprises (b) contacting the chromatin with the transmembrane pore such that one or more of the bound proteins are stripped from the chromatin as the chromatin moves with respect to the pore. The one or more bound proteins such as one or more histones and / or one or more nucleosomes, are typically stripped from the target polynucleotide as the target polynucleotide moves through the transmembrane pore but the one or more proteins are too large to pass through the transmembrane pore. All of the one or more bound proteins are preferably stripped from the chromatin tide as the chromatin moves with respect to, such as through, the pore.

[0168] The method also comprises (c) taking one or more measurements as the chromatin moves with respect to the pore wherein the measurements are indicative of one or more characteristics of the chromatin and thereby characterising the chromatin.

[0169] Steps (a) and (b) preferably comprise contacting the chromatin with a transmembrane pore without removing one or more bound histones and / or one or more bound nucleosomes from the chromatin such that the chromatin moves with respect to the pore. Steps (a) and (b) more preferably comprise contacting the chromatin with a transmembrane pore without removing the histones and / or the nucleosomes from the chromatin such that the chromatin moves with respect to the pore.

[0170] The target polynucleotide or chromatin may be characterised in the method of the invention in any suitable manner. The target polynucleotide or chromatin is preferably characterised by detecting an ionic current or optical signal as it moves with respect to the transmembrane pore. This is described in more detail below. The method is amenable to these and other methods of characterising polynucleotides.

[0171] The one or more characteristics are preferably selected from (i) the length of the target polynucleotide or chromatin, (ii) the identity of the target polynucleotide or chromatin, (iii) the sequence of the target polynucleotide or chromatin, (iv) the secondary structure of the target polynucleotide or chromatin and (v) whether or not the target polynucleotide or chromatin is modified. The target polynucleotide or chromatin may be modified by methylation, by oxidation, by damage, with one or more proteins or with one or more labels, tags, or spacers. The method may comprise foot printing the target polynucleotide. In this embodiment, the signal provided by the transmembrane pore may be used to identify locations where the one or more proteins bind to the target polynucleotide and those locations are then characterised or sequenced using the transmembrane pore. This method is capable of deriving information about the protein binding sites on the target polynucleotide.

[0172] The one or more characteristics of the target polynucleotide or chromatin are preferably measured by electrical measurement and / or optical measurement. The electrical measurement is preferably a current measurement, an impedance measurement, a tunnelling measurement, or a field effect transistor (FET) measurement.

[0173] The method more preferably comprises (i) contacting the target polynucleotide or chromatin with a transmembrane pore such that one or more of the bound proteins are stripped from the target polynucleotide as the target polynucleotide or chromatin moves through the transmembrane pore and (ii) measuring the current moving through the transmembrane pore as the target polynucleotide or chromatin moves through the transmembrane pore wherein the current is indicative of one or more characteristics of the target polynucleotide or chromatin and thereby characterising the target polynucleotide or chromatin. The one or more characteristics may be any of those described above.

[0174] The movement of the target polynucleotide or chromatin with respect to the transmembrane pore or through the transmembrane pore is preferably controlled by a movement control protein. The use of such proteins in transmembrane pore sequencing is known. Examples of suitable proteins are discussed in more detail above.

[0175] The one or more bound proteins, such as the one or more, or all, bound histones and / or one or more, or all, bound nucleosomes, are stripped from the target polynucleotide or chromatin as it moves with respect to the transmembrane pore. As the proteins are stripped from chromatin, it becomes DNA rather than a mixture of DNA and proteins.

[0176] Any suitable transmembrane pore can be used. A transmembrane pore is a structure that crosses the membrane to some degree. It permits hydrated ions driven by an applied potential to flow across or within the membrane. The transmembrane pore typically crosses the entire membrane so that hydrated ions may flow from one side of the membrane to the other side of the membrane. However, the transmembrane pore does not have to cross the membrane. It may be closed at one end. For instance, the pore may be a well, gap, channel, trench or slit in the membrane along which or into which hydrated ions may flow.

[0177] The transmembrane pore typically has a first opening and a second opening. The first opening is typically the cis opening and the second opening is typically the trans opening. However, the first opening may be the trans opening and the second opening may be the cis opening. Any polynucleotide binding protein used in the method of the invention is typically provided at the first opening of the transmembrane pore and thus controls the movement of the target polynucleotide in the direction from the second opening of the transmembrane pore towards the first opening of the transmembrane pore.

[0178] Any transmembrane pore may be used in the method of the invention. The pore may be biological or artificial. Suitable pores include, but are not limited to, protein pores, polynucleotide pores and solid-state pores. The pore may be a DNA origami pore (Langecker et al., Science, 2012; 338: 932-936). Suitable DNA origami pores are disclosed in WO2013 / 083983 (incorporated by reference herein in its entirety).

[0179] The transmembrane pore is preferably a transmembrane protein pore. A transmembrane protein pore is a polypeptide or a collection of polypeptides that permits hydrated ions, such as polynucleotide, to flow from one side of a membrane to the other side of the membrane. In the method of the invention, the transmembrane protein pore is capable of forming a pore that permits hydrated ions driven by an applied potential to flow from one side of the membrane to the other. The transmembrane protein pore preferably permits polynucleotides to flow from one side of the membrane, such as a triblock copolymer membrane, to the other. The transmembrane protein pore allows a polynucleotide to be moved through the pore.

[0180] The transmembrane pore may be a transmembrane protein pore which is a monomer or an oligomer. The pore is preferably made up of several repeating subunits, such as at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 11, at least about 12, at least about 13, at least about 14, at least about 15, or at least about 16 subunits. The pore is preferably a hexameric, heptameric, octameric or nonameric pore. The pore may be a homo-oligomer or a hetero-oligomer.

[0181] The transmembrane protein pore may comprise a barrel or channel through which the ions may flow. The subunits of the pore typically surround a central axis and contribute strands to a transmembrane [3-barrel or channel or a transmembrane a-helix bundle or channel.

[0182] Typically, the barrel or channel of the transmembrane protein pore comprises amino acids that facilitate interaction with an analyte, such as the target polynucleotide or chromatin. These amino acids are preferably located near a constriction of the barrel or channel. The transmembrane protein pore typically comprises one or more positively charged amino acids, such as arginine, lysine or histidine, or aromatic amino acids, such as tyrosine or tryptophan. These amino acids typically facilitate the interaction between the pore and nucleotides, polynucleotides, or nucleic acids. The transmembrane protein pore may be from or derived from Wza, Iota toxin, Anthrax protective antigen, Vibrio cholerae cytolysin, Cytotoxin K (CytK), CELIII, CsgG, CsgF, CsgG- CsgF, Aerolysin, alpha hemolysin, MspA, MspB, MspC, PorARr, PorBRr, PorARc, PilQ, necrotic enteritis B-like toxin (NetB), FraC, portal proteins including G20c, P23_45, T4, SPP1, P22 and Phi29, gamma hemolysin, Monalysin, Lysenin, ClyA, an actinoporin, Clostridium perfringens beta toxin, parasporin-2, epsilon toxin, lectin from the parasitic mushroom Laetiporus sulphureus (LSL), volvatoxin, Cry toxins, CytlAa, Cyt2Aa, Complement component 9 (C9), Perfringolysin O, Pleurotolysin, Listeriolysin, Perforin-2, Gasdermin-A3, L-, P- and M-ring protein, Type II secretion system protein D, GspD, InvG, VirB7, SpoIIIAG, Cag8, Cag3, Cag or other proteins in the Type IV secretion system apparatus protein CagY, WzzB, Pentraxin, Afp2, Major vault protein, Thioredoxin-dependent peroxidase reductase, Arf-GAP, Respiratory syncytial virus ribonucleoprotein, Chikungunya virus nonstructural protein 1, PRC, YaxA, XaxA, HfaB, NfpAB, leukocidin and PrgH.

[0183] Suitable transmembrane pores for use in the invention include those described in WO 2016 / 034591, WO 2017 / 149316, WO 2017 / 149317, WO 2017 / 149318, WO 2018 / 211241, WO 2019 / 002893, WO 2024 / 033421, WO 2024 / 033422, and WO 2024 / 033443 (all incorporated by reference herein in their entirety).

[0184] The transmembrane may be formed from a chimeric pore monomer comprising two or more regions, wherein at least two of the two or more regions are from at least two different pores. The chimeric pore monomer may comprise any number of regions, such as three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more or ten or more regions, from different pores. The chimeric pore monomer may comprise two or three regions. The regions may be any of those discussed above. The regions are preferably selected from a cap region, a constriction region, and a transmembrane region. The regions may be a cap region and a constriction region. The regions may be a cap region, a constriction region, and a transmembrane region. The at least two different pores are typically at least two different pores that appear in nature. The at least two different pores are typically at least two different wild-type or naturally occurring pores. The at least two different pores are preferably different before any artificial or synthetic modifications, such as additions, deletions and / or substitutions, are made to them. The at least two different pores are preferably homologues, for example structural homologues. A structural homologue refers to a protein or molecule that shares a similar three-dimensional structure with another protein or molecule. This can be determined using standard methods in the art (e.g., AlphaFold or PSIPRED). Structural homologues typically have similar sequences. Structural homologues are normally identified in similar species. The at least two different pores may be selected from any of the pores listed above. The at least two different pores may be two different PorARc pores or three different PorARc pores. The at least two different pores may be two different CsgG pores or three different CsgG pores. The chimeric pore monomer may be any of those described in PCT / EP2023 / 080135 (incorporated by reference herein in its entirety).

[0185] The transmembrane pore may be formed from a pore monomer comprising (a) a CsgG monomer and (b) a fusion polypeptide comprising a first portion comprising a CsgF peptide and a second portion comprising a helix-forming auxiliary protein, wherein the fusion protein is attached to the pore monomer. The pore monomer may be derived from a protein transmembrane pore complex comprising (a) a CsgG transmembrane pore comprising a lumen and (b) a fusion polypeptide comprising a first portion comprising a CsgF protein and a second portion comprising a helix-forming auxiliary protein, wherein the fusion protein is attached to the transmembrane pore. The auxiliary protein can be designed de novo using computer-based structural analysis tools to confer certain desirable features to the CsgG monomer (e.g., modulation of pore width, lengthening of pore lumen, formation of one or more additional constrictions, etc.). The de novo designed auxiliary protein may form one or more additional constrictions in the lumen of a CsgG pore formed from the monomer, and improve discrimination of polymer units as an analyte moves through the pore. The pore monomer may be any of the pore monomers described in WO 2024 / 033447 (incorporated by reference herein in its entirety).

[0186] Membrane

[0187] The transmembrane pore is typically present in a membrane. Any suitable membrane may be used.

[0188] The membrane is preferably an amphiphilic layer. An amphiphilic layer is a layer formed from amphiphilic molecules, such as phospholipids, which have both hydrophilic and lipophilic properties. The amphiphilic molecules may be synthetic or naturally occurring. Non-naturally occurring amphiphiles and amphiphiles which form a monolayer are known in the art and include, for example, block copolymers (Gonzalez-Perez et al., Langmuir, 2009, 25, 10447-10450). Block copolymers are polymeric materials in which two or more monomer sub-units that are polymerized together to create a single polymer chain. Block copolymers typically have properties that are contributed by each monomer sub-unit. However, a block copolymer may have unique properties that polymers formed from the individual sub-units do not possess. Block copolymers can be engineered such that one of the monomer sub-units is hydrophobic ( / .e., lipophilic), whilst the other sub-unit(s) are hydrophilic whilst in aqueous media. In this case, the block copolymer may possess amphiphilic properties and may form a structure that mimics a biological membrane. The block copolymer may be a diblock (consisting of two monomer sub-units) but may also be constructed from more than two monomer sub-units to form more complex arrangements that behave as amphipiles. The copolymer may be a triblock, tetrablock or pentablock copolymer. The membrane may be a triblock copolymer membrane. Archaebacterial bipolar tetraether lipids are naturally occurring lipids that are constructed such that the lipid forms a monolayer membrane. These lipids are generally found in extremophiles that survive in harsh biological environments, thermophiles, halophiles and acidophiles. Their stability is believed to derive from the fused nature of the final bilayer. It is straightforward to construct block copolymer materials that mimic these biological entities by creating a triblock polymer that has the general motif hydrophilic-hydrophobic- hydrophilic. This material may form monomeric membranes that behave similarly to lipid bilayers and encompass a range of phase behaviours from vesicles through to laminar membranes. Membranes formed from these triblock copolymers hold several advantages over biological lipid membranes. Because the triblock copolymer is synthesised, the exact construction can be carefully controlled to provide the correct chain lengths and properties required to form membranes and to interact with pores and other proteins.

[0189] Block copolymers may also be constructed from sub-units that are not classed as lipid submaterials; for example, a hydrophobic polymer may be made from siloxane or other non- hydrocarbon-based monomers. The hydrophilic sub-section of block copolymer can also possess low protein binding properties, which allows the creation of a membrane that is highly resistant when exposed to raw biological samples. This head group unit may also be derived from non-classical lipid head-groups.

[0190] Triblock copolymer membranes also have increased mechanical and environmental stability compared with biological lipid membranes, for example a much higher operational temperature or pH range. The synthetic nature of the block copolymers provides a platform to customise polymer-based membranes for a wide range of applications.

[0191] The membrane may be one of the membranes disclosed in International Application No. WO2014 / 064443 or WO2014 / 064444 (both of which are incorporated herein by reference in their entireties).

[0192] The amphiphilic molecules may be chemically modified or functionalised to facilitate coupling of the polynucleotide. The amphiphilic layer may be a monolayer or a bilayer. The amphiphilic layer is typically planar. The amphiphilic layer may be curved. The amphiphilic layer may be supported.

[0193] Amphiphilic membranes are typically naturally mobile, essentially acting as two-dimensional fluids with lipid diffusion rates of approximately IO-8cm s4. This means that the pore and coupled polynucleotide can typically move within an amphiphilic membrane.

[0194] The membrane may be a lipid bilayer. Lipid bilayers are models of cell membranes and serve as excellent platforms for a range of experimental studies. For example, lipid bilayers can be used for in vitro investigation of membrane proteins by single-channel recording. Alternatively, lipid bilayers can be used as biosensors to detect the presence of a range of substances. The lipid bilayer may be any lipid bilayer. Suitable lipid bilayers include, but are not limited to, a planar lipid bilayer, a supported bilayer, or a liposome. The lipid bilayer is preferably a planar lipid bilayer. Suitable lipid bilayers are disclosed in WO 2008 / 102121, WO 2009 / 077734, and WO 2006 / 100484 (incorporated herein by reference in their entireties).

[0195] Methods for forming lipid bilayers are known in the art. Lipid bilayers are commonly formed by the method of Montal and Mueller (Proc. Natl. Acad. Sci. USA., 1972; 69: 3561-3566).

[0196] A lipid bilayer may be formed as described in WO 2009 / 077734 (incorporated herein by reference in its entirety). In this method, the lipid bilayer is formed from dried lipids. A lipid bilayer may be formed across an opening as described in W02009 / 077734.

[0197] The membrane may comprise a solid-state layer. Solid-state layers can be formed from both organic and inorganic materials including, but not limited to, microelectronic materials, insulating materials such as Si3N4, A12O3, and SiO, organic and inorganic polymers such as polyamide, plastics such as Teflon® or elastomers such as two-component addition-cure silicone rubber, and glasses. The solid-state layer may be formed from graphene. Suitable graphene layers are disclosed in WO 2009 / 035647 (incorporated herein by reference in its entirety). If the membrane comprises a solid-state layer, the pore is typically present in an amphiphilic membrane or layer contained within the solid-state layer, for instance within a hole, well, gap, channel, trench or slit within the solid-state layer. The skilled person can prepare suitable solid-state / amphiphilic hybrid systems. Suitable systems are disclosed in WO 2009 / 020682 and WO 2012 / 005857 (incorporated herein by reference in their entireties). Any of the amphiphilic membranes or layers discussed above may be used.

[0198] The methods disclosed herein are typically carried out using (i) an artificial amphiphilic layer comprising a pore, (ii) an isolated, naturally occurring lipid bilayer comprising a pore, or (iii) a cell having a pore inserted therein. The methods are typically carried out using an artificial amphiphilic layer, such as an artificial triblock copolymer layer. The layer may comprise other transmembrane and / or intramembrane proteins as well as other molecules in addition to the pore. Suitable apparatus and conditions are discussed below. The method of the invention is typically carried out in vitro.

[0199] General methods

[0200] The method of the invention may be carried out using any apparatus that is suitable for nanopore sensing. For example, the apparatus may comprise a chamber comprising an aqueous solution and a barrier that separates the chamber into two sections. The barrier may have an aperture in which a membrane containing a transmembrane pore is formed. The methods may be carried out using the apparatus described in WO 2008 / 102120, WO 2010 / 122293, or WO 00 / 28312 (incorporated herein by reference in their entireties). In brief, the binding of a target polynucleotide or chromatin in the channel of a pore will have an effect on the open-channel ion flow through the pore, which is the essence of "molecular sensing" of pore channels. Variation in the open-channel ion flow can be measured using suitable measurement techniques by the change in electrical current. The degree of reduction in ion flow, as measured by the reduction in electrical current, is related to the size of the obstruction within, or in the vicinity of, the pore. Binding of the target polynucleotide or chromatin in or near the pore therefore provides a detectable and measurable event, thereby forming the basis of a "biological sensor".

[0201] When used to characterise the target polynucleotide or chromatin, the presence, absence or one or more characteristics of the target polynucleotide or chromatin are determined. The methods may be for determining the presence, absence or one or more characteristics of at least one target polynucleotide or chromatin. The methods may concern determining the presence, absence or one or more characteristics of two or more target polynucleotides or two more stretches of chromatin. The methods may comprise determining the presence, absence or one or more characteristics of any number of target polynucleotides or stretches of chromatin, such as about 2, about 5, about 10, about 15, about 20, about 30, about 40, about 50, about 100 or more target polynucleotides or stretches of chromatin. Any number of characteristics of the one or more target polynucleotides or stretches of chromatin may be determined, such as about 1, about 2, about 3, about 4, about 5, about 10 or more characteristics. Characteristics amenable to being detected in the methods provide herein include the identity or sequence of the target polynucleotide or chromatin, the length of the target polynucleotide or chromatin, whether or not the target polynucleotide or chromatin is modified, etc. In some embodiments the method of the invention are methods of sequencing the target polynucleotide or chromatin. In some embodiments the sequences of the target polynucleotide or chromatin may be determined in real-time by aligning real-time signal or basecalling to known references. Exemplary methods of determining a polynucleotide sequence are described in WO 2016 / 059427 (incorporated herein by reference in its entirety).

[0202] When used to characterize the target polynucleotide or chromatin, the methods may involve measuring the ion current flow through the pore, typically by measurement of a current. Alternatively, the ion flow through the pore may be measured optically, such as disclosed by Heron et al: J. Am. Chem. Soc. 9 Vol. 131, No. 5, 2009. Therefore, the apparatus may also comprise an electrical circuit capable of applying a potential and measuring an electrical signal across the membrane and pore. The characterisation methods may be carried out using a patch clamp or a voltage clamp. The characterisation methods preferably involve the use of a voltage clamp. The methods may involve measuring an optical signal as described in Chen et al, Nature Communications (2018)9: 1733, the entire contents of which are hereby incorporated by reference. For example, a transmembrane pore such as an optically engineered nanopore structure (e.g., a plasmonic nanoslit) may be used to locally enable single-molecule surface enhanced Raman spectroscopy (SERS) to allow the characterisation of the polynucleotide through direct Raman spectroscopic detection.

[0203] The methods may be carried out on a silicon-based array of wells where each array comprises 128, 256, 512, 1024, 2000, 3000, 4000, 6000, 10000, 12000, 15000 or more wells.

[0204] The methods may involve the measuring of a current flowing through the transmembrane pore. The method is typically carried out with a voltage applied across the membrane and pore. The voltage used is typically from +2 V to -2 V, typically -400 mV to +400mV. The voltage used is preferably in a range having a lower limit selected from -400 mV, -300 mV, -200 mV, -150 mV, -100 mV, -50 mV, -20mV and 0 mV and an upper limit independently selected from +10 mV, + 20 mV, +50 mV, +100 mV, +150 mV, +200 mV, +300 mV and +400 mV. The voltage used is more preferably in the range 100 mV to 240mV and most preferably in the range of 120 mV to 220 mV. It is possible to increase discrimination between different nucleotides by a pore by using an increased applied potential.

[0205] The methods are typically carried out in the presence of any charge carriers, such as metal salts, for example alkali metal salts, halide salts, for example chloride salts, such as alkali metal chloride salt. Charge carriers may include ionic liquids or organic salts, for example tetramethyl ammonium chloride, trimethylphenyl ammonium chloride, phenyltrimethyl ammonium chloride, or l-ethyl-3-methyl imidazolium chloride. In the exemplary apparatus discussed above, the salt is present in the aqueous solution in the chamber. Potassium chloride (KCI), sodium chloride (NaCI) or caesium chloride (CsCI) is typically used. KCI is preferred. The salt may be an alkaline earth metal salt such as calcium chloride (CaCI2).

[0206] The salt concentration may be at saturation. The salt concentration may be 3M or lower and is typically from 0.1 to 2.5 M, from 0.3 to 1.9 M, from 0.5 to 1.8 M, from 0.7 to 1.7 M, from 0.9 to 1.6 M or from 1 M to 1.4 M. The salt concentration is preferably from 150 mM to 1 M. The method is preferably carried out using a salt concentration of at least 0.3 M, such as at least 0.4 M, at least 0.5 M, at least 0.6 M, at least 0.8 M, at least 1.0 M, at least 1.5 M, at least 2.0 M, at least 2.5 M or at least 3.0 M. High salt concentrations provide a high signal to noise ratio and allow for currents indicative of binding / no binding to be identified against the background of normal current fluctuations.

[0207] The methods are typically carried out in the presence of a buffer. In the exemplary apparatus discussed above, the buffer is present in the aqueous solution in the chamber. Any suitable buffer may be used. Typically, the buffer is HEPES. Another suitable buffer is Tris-HCI buffer. The methods are typically carried out at a pH of from 4.0 to 12.0, from 4.5 to 10.0, from 5.0 to 9.0, from 5.5 to 8.8, from 6.0 to 8.7 or from 7.0 to 8.8 or 7.5 to 8.5. The pH used is preferably about 7.5.

[0208] The methods may be carried out at from 0 °C to 100 °C, from 15 °C to 95 °C, from 16 °C to 90 °C, from 17 °C to 85 °C, from 18 °C to 80 °C, 19 °C to 70 °C, or from 20 °C to 60 °C. The methods are typically carried out at room temperature. The methods are optionally carried out at a temperature that supports enzyme function, such as about 37 °C.

[0209] Any of the proteins described herein may be made synthetically or by recombinant means. For example, the protein may be synthesised by in vitro translation and transcription (IVTT). The amino acid sequence of the protein may be modified to include non-naturally occurring amino acids or to increase the stability of the protein. When a protein is produced by synthetic means, such amino acids may be introduced during production. The protein may also be altered following either synthetic or recombinant production.

[0210] Any of the proteins described herein can be produced using standard methods known in the art. Polynucleotide sequences encoding protein may be derived and replicated using standard methods in the art. Polynucleotide sequences encoding a protein may be expressed in a bacterial host cell using standard techniques in the art. The protein may be produced in a cell by in situ expression from a recombinant expression vector. The expression vector optionally carries an inducible promoter to control the expression of the protein. These methods are described in Sambrook, J. and Russell, D. (2001). Molecular Cloning: A Laboratory Manual, 3rd Edition. Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY.

[0211] The protein may be produced in large scale following purification by any protein liquid chromatography system from protein producing organisms or after recombinant expression. Typical protein liquid chromatography systems include FPLC, AKTA systems, the Bio-Cad system, the Bio-Rad BioLogic system, and the Gilson HPLC system.

[0212] Polynucleotides

[0213] The invention also provides a target polynucleotide comprising one or more bound proteins and further comprising a movement control protein. The target polynucleotide preferably further comprises a sequencing adaptor. Any of the specific embodiments discussed above with reference to the method of the invention are equally applicable to the target polynucleotides of the invention. For instance, the target polynucleotide may be any of the lengths described above and / or may comprise any of the proteins discussed above.

[0214] The invention also provides a target polynucleotide comprising one or more selectively retained nucleosomes and a sequencing adaptor. The target polynucleotide preferably further comprises a movement control protein. Any of the specific embodiments discussed above with reference to the method of the invention are equally applicable to the target polynucleotide of the invention. For instance, the target polynucleotide may be any of the lengths described above and / or may comprise any of the proteins discussed above.

[0215] Apparatus

[0216] The invention also provides an apparatus produced by a method comprising (a) obtaining a target polynucleotide comprising one or more bound proteins, (b) attaching a movement control protein and / or a sequencing adaptor to the target polynucleotide and (c) contacting the target polynucleotide obtained in (b) with a transmembrane pore. Any of the specific embodiments discussed above with reference to the method of the invention are equally applicable to the apparatus of the invention. The one or more bound proteins preferably do not comprise a polymerase, an exonuclease, a translocase, a helicase, or a topoisomerase. In other words, the one or more bound proteins preferably do not comprise any of a polymerase, an exonuclease, a translocase, a helicase, and a topoisomerase. The one or more bound proteins preferably comprise one or more bound histones and / or one or more bound nucleosomes. The target polynucleotide comprising one or more bound proteins is preferably chromatin.

[0217] The following Example illustrates the invention. It is to be understood that although particular embodiments, specific configurations as well as materials and / or molecules, have been discussed herein for methods according to the invention, various changes or modifications in form and detail may be made without departing from the scope and spirit of this invention. The following examples are provided to better illustrate particular embodiments, and they should not be considered limiting the application. The application is limited only by the claims.

[0218] EXAMPLE 1

[0219] This example describes a method of reconstituting nucleosomes, using purified DNA and recombinant histones, prior to sequencing on a nanopore.

[0220] Materials and methods

[0221] In-vitro reconstitution of chromatin

[0222] Firstly, lOpg of the target polynucleotides were mixed with 9.45pl of recombinant human histone octamer protein (from Abeam: ab216230) in lx assembly buffer (10 mM HEPES- NaOH (pH 7.6), 2M NaCI, 0.5 mM EDTA, 0.05% IGEPAL CA-630) in a final volume of 800pl. The resulting mixture was gradually dialysed using Sigma-Aldrich Pur-A-Lyzer Midi IkDa MWCO dialysis kit (Sigma: PURD10005) in 500ml of a low salt buffer (lOmM HEPES-NaOH (pH 7.6), 50mM NaCI, 0.5 mM EDTA, 0.05% IGEPAL CA-630 ) over the course of 72 hours at room temperature with mixing on a magnetic stirrer plate, the wash buffer was changed every 4 hours. Library preparation and sequencing lOpI of reconstituted chromatin (250ng) mixed with Ipl of ONT FRA transposase mix (ONT: SQK-RAD114) for tagmentation, the reaction was incubated at 30°C for 2 minutes followed by 80°C for 3 minutes. Ipl of ONT RA (ONT: SQK-RAD114) was added and incubated at 21°C for 5 minutes. 12pl of the resulting library was then made up to 75p I with 37.5pl of SB (ONT: SQK-RAD114) and 25.5pl of LB (ONT: SQK-RAD114). MinlON flow cells were then flushed twice with pre-prepared flush mixture (1170pl FCF and 30pl FCT, ONT: SQK-RAD 114), 5 minutes apart. Libraries were then loaded to the flow cell and data collected for 12 hours using the MinKNOW software (ONT).

[0223] Results

[0224] Reconstituted chromatin sequencing data was basecalled and aligned to the reference (HG002 T2T vl.01 : https: / / github.com / marbl / hg002). Results in Table 1 show that compared to the control, the reconstituted chromatin sample gave higher throughput, longer read length (as measured by mean or N50) and similar basecall accuracy. Figure 1 shows the differences in read length distributions between reconstituted chromatin and the same batch of non-reconstituted input DNA.

[0225] Table 1: Throughputs, median basecalled accuracies and read length N50 and means are shown for reconstituted chromatin versus the same batch of non-reconstituted input DNA.

[0226] EXAMPLE 2

[0227] This example describes a method of extracting chromatin from mammalian cells and preparing it for nanopore sequencing.

[0228] Materials and methods

[0229] Chromatin extraction from mammalian cells

[0230] Firstly, cultured mammalian cells were counted using Propidium Iodide / Acridine Orange stain and 6 x 106viable cells were pelleted by centrifugation (500g for 5 minutes), washed with 500pl lx PBS (NaCI: 137mM, KCI: 2.7mM, Na2HPO4: lOmM, KH2PO4: 1.8mM) and pelleted again by centrifugation. Chromatin was extracted from mammalian cells by resuspending the pellet in 1ml of ice-cold Radio Immunoprecipitation Assay buffer (Thermofisher: 89901) with 10|_il proteinase inhibitor cocktail freshly added (Thermofisher: 78441). Cells / nuclei were left to lyse for 15 minutes on ice. Extracted chromatin was then sonicated using a bath sonicator (Diagenode Bioruptor Pico) pre-chilled to 4°C, 30 / 30 seconds on / off for 60 cycles.

[0231] MNase digestion

[0232] Extracted chromatin quality was assessed by digestion using a micrococcal nuclease digestion assay. Briefly, chromatin was added to five tubes containing MNase assay buffer (NEB: B0247SVIAL), Ipl of Recombinant BSA (NEB: B9200SVIAL) and a 3-fold dilution series of MNase enzyme (NEB: M0247SVIAL). The reaction was incubated for 1 hour at 37°C followed by addition of 2pl of ProteinaseK (NEB: P8107S) and a further incubation of 20 minutes at 56°C. ProteinaseK was heat inactivated by incubating at 95°C for 5 minutes.

[0233] Library preparation and sequencing

[0234] Extracted chromatin was diluted 1 / 100 and 95ul were mixed with 5p I of ONT FRA transposase mix (ONT: SQK-RAD114) for tagmentation, the reaction was incubated at 30°C for 1 minute followed by 80°C for 1 minute. Ipl of ONT RA (ONT: SQK-RAD114) was added and incubated at 21°C for 5 minutes. 12 I of the resulting library was then made up to 75p I with 37.5pl of SB (ONT: SQK-RAD 114) and 25.5pl of LB (ONT: SQK-RAD114). MinlON flow cells (ONT: FLO-MIN114 R10.4.1) were then flushed twice with pre-prepared flush mixture (1170pl FCF and 30pl FCT, ONT: SQK-RAD 114), 5 minutes apart. Libraries were then loaded to the flow cell and data collected using the MinKNOW software (ONT).

[0235] Results

[0236] Extracted chromatin formed in this example was shown to be of high molecular weight as seen for lanes with low MNase concentration (Figure 2). The MNase digestion assay reported a band at 150bp confirming the presence of mono-nucleosomes and indicating histones were bound correctly to the extracted polynucleotides (Figure 2).

[0237] Sequencing produced data which was basecalled and aligned to the reference (HG002 T2T vl.01 : https: / / github.com / marbl / hg002) using Dorado v0.6.0 with the dna_rl0.4.1_e8.2_400bps_sup @v4.2.0 model. Figure 3 shows a raw current trace for a typical read which has been successfully basecalled and aligned to the reference. The current trace contains expected features which show that sequencing is not impeded by the presence of histones or other bound proteins.

Claims

CLAIMS1. A method of characterising a target polynucleotide comprising one or more bound proteins, the method comprising: (a) retaining the one or more bound proteins on the target polynucleotide until the target polynucleotide is contacted with a transmembrane pore, (b) contacting the target polynucleotide with the transmembrane pore such that one or more of the bound proteins are stripped from the target polynucleotide as the target polynucleotide moves with respect to the pore; and (c) taking one or more measurements as the target polynucleotide moves with respect to the pore wherein the measurements are indicative of one or more characteristics of the target polynucleotide and thereby characterising the target polynucleotide.

2. A method according to claim 1, wherein all of the one or more bound proteins are stripped from the target polynucleotide as the target polynucleotide moves with respect to the pore.

3. A method according to claim 1 or 2, wherein step (a) does not comprise a step for removing the one or more bound proteins from the target polynucleotide.

4. A method according to any one of claims 1-3, wherein step (a) does not comprise a step for removing protein from the target polynucleotide.

5. A method according to claim 3 or 4, wherein the step for removing the one or more bound proteins is a phenol or chloroform step, a precipitation step and / or a proteinase step.

6. A method according to claim 5, wherein the proteinase is proteinase K, trypsin, chymotrypsin, and / or papain.

7. A method according to any one of the preceding claims, wherein the target polynucleotide comprises chromatin.

8. A method according to claim 7, wherein the chromatin is derived from one or more eukaryotic cells.

9. A method according to claim 8, wherein the one or more eukaryotic cells are human.

10. A method according to any one of claims 1-6, wherein the one or more bound proteins are added to the target polynucleotide before the method is conducted.

11. A method according to claim 10, wherein the one or more bound proteins are added to the target polynucleotide to stabilise the target polynucleotide.

12. A method according to claim 10 or 11, wherein the method comprises the initial step of binding the one or more proteins to the target polynucleotide.

13. A method according to any one of claims 10-12, wherein the target polynucleotide is DNA, RIMA or an artificial or non-natural polynucleotide.

14. A method according to any one of the preceding claims, wherein the one or more bound proteins are one or more bound polynucleotide binding proteins.

15. A method according to claim 14, wherein the one or more bound polynucleotide binding proteins comprise (a) one or more bound histones, (b) one or more bound transcription factors, (c) one or more bound proteins involved in DNA modelling or remodelling, replication, transcription and / or repair, or (d) combinations thereof.

16. A method according to any one of the preceding claims, wherein the one or more bound proteins comprise one or more bound histones and / or one or more bound transcription factors.

17. A method according to any one of the preceding claims, wherein the one or more bound proteins do not comprise a polymerase, an exonuclease, a translocase, a helicase, or a topoisomerase.

18. A method according to any one of the preceding claims, wherein the method further comprises attaching one or more sequencing adaptors to the target polynucleotide.

19. A method according to any one of the preceding claims, wherein the method further comprises attaching a movement control protein to the target polynucleotide.

20. A method according to claim 19, wherein the movement control protein is a polymerase, exonuclease, translocase, helicase, or topoisomerase.

21. A method according to any one of the preceding claims, wherein the method further comprises before step (a) lysing one or more cells and / or further comprises before step (a) one or more of binding the target polynucleotide to a surface, washing the target polynucleotide, and eluting the target polynucleotide.

22. A method according to claim 21, wherein the method comprises removing one or more non-bound proteins.

23. A method according to any one of the preceding claims, wherein the method comprises determining one or more characteristics selected from (i) the length of the target polynucleotide, (ii) the identity of the target polynucleotide, (iii) the sequence of the target polynucleotide, (iv) the secondary structure of the target polynucleotide and (v) whether or not the target polynucleotide is modified.

24. A method according to any one of the preceding claims, wherein the movement of the target polynucleotide with respect to the pore is controlled by a movement control protein.

25. A method of increasing the read length of a target polynucleotide using a transmembrane pore, wherein the method comprises binding one or more proteins to the target polynucleotide or retaining one or more bound proteins on the target polynucleotide.

26. A method according to claim 25, wherein the one or more proteins are as defined in any one of claims 15-17.

27. A method according to any one of the preceding claims, wherein the transmembrane pore is a transmembrane protein pore.

28. A method according to claim 27, wherein the transmembrane protein pore is from or derived from Wza, Iota toxin, Anthrax protective antigen, Vibrio cholerae cytolysin, Cytotoxin K (CytK), CELIII, CsgG, CsgF, CsgG-CsgF, Aerolysin, alpha hemolysin, MspA, MspB, MspC, PorARr, PorBRr, PorARc, PilQ, necrotic enteritis B-like toxin (NetB), FraC, portal proteins including G20c, P23_45, T4, SPP1, P22 and Phi29, gamma hemolysin, Monalysin, Lysenin, ClyA, an actinoporin, Clostridium perfringens beta toxin, parasporin- 2, epsilon toxin, lectin from the parasitic mushroom Laetiporus sulphureus (LSL), volvatoxin, Cry toxins, CytlAa, Cyt2Aa, Complement component 9 (C9), Perfringolysin O, Pleurotolysin, Listeriolysin, Perforin-2, Gasdermin-A3, L-, P- and M-ring protein, Type II secretion system protein D, GspD, InvG, VirB7, SpoIIIAG, Cag8, Cag3, Cag or other proteins in the Type IV secretion system apparatus protein CagY, WzzB, Pentraxin, Afp2, Major vault protein, Thioredoxin-dependent peroxidase reductase, Arf-GAP, Respiratory syncytial virus ribonucleoprotein, Chikungunya virus nonstructural protein 1, PRC, YaxA, XaxA, HfaB, NfpAB, leukocidin and PrgH.

29. A method according to any one of claims 1-27, wherein the transmembrane pore is a solid-state pore.

30. A method of characterising chromatin, the method comprising: (a) retaining one or more bound proteins on the chromatin until the chromatin is contacted with a transmembrane pore, (b) contacting the chromatin with the transmembrane pore such that one or more of the bound proteins are stripped from the chromatin as the chromatin moves with respect to the pore; and (c) taking one or more measurements as the chromatin moves with respect to the pore wherein the measurements are indicative of one or more characteristics of the chromatin and thereby characterising the chromatin.

31. A method of preparing a target polynucleotide comprising one or more bound proteins for characterisation using a transmembrane pore, the method comprising: (a) retaining the one or more bound proteins on the target polynucleotide until the target polynucleotide is contacted with the transmembrane pore and (b) attaching a movement control protein to the target polynucleotide before the target polynucleotide is contacted with the transmembrane pore.

32. A method according to claim 31, wherein the target polynucleotide is chromatin.

33. A method according to claim 32, wherein the one or more bound proteins are one or more bound histones and / or one or more bound nucleosomes.

34. A method of preparing a target polynucleotide for characterisation using a transmembrane pore, the method comprising: (a) binding one or more proteins to the target polynucleotide before the target polynucleotide is contacted with the transmembrane pore and (b) attaching a movement control protein to the target polynucleotide before the target polynucleotide is contacted with the transmembrane pore.

35. A target polynucleotide comprising one or more bound proteins and further comprising a movement control protein.

36. A target polynucleotide according to claim 35, wherein the target polynucleotide further comprises a sequencing adaptor.

37. A target polynucleotide comprising one or more selectively retained nucleosomes and a sequencing adaptor.

38. A target polynucleotide according to claim 37, wherein the stretch of chromatin further comprises a movement control protein.

39. An apparatus produced by a method comprising (a) obtaining a target polynucleotide comprising one or more bound proteins, (b) attaching a movement control protein and / or a sequencing adaptor to the target polynucleotide and (c) contacting the target polynucleotide obtained in (b) with a transmembrane pore.

40. An apparatus according to claim 39, wherein the target polynucleotide comprises chromatin.

Citation Information

Patent Citations

  • phi 29 DNA polymerase

    US5576204A

  • A miniature support for thin films containing single channels or nanopores and methods for using same

    WO2000028312A1

  • Deliver of molecules to a li id bila

    WO2006100484A2

  • Lipid bilayer sensor system

    WO2008102120A1

  • Formation of lipid bilayers

    WO2008102121A1