Nanopores

JP2024535120A5Pending Publication Date: 2025-09-02OXFORD NANOPORE TECH LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024536538
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-08-26
Filing Date
2022-08-26
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

Current methods for nucleic acid sequencing are time-consuming and expensive, and techniques for characterizing polypeptides are less advanced despite their importance, with existing methods like mass spectrometry and Edman degradation being inadequate for single-molecule analysis and prone to contamination.

Method used

Development of mutant cytotoxin K monomers that form pores for nanopore sensing, allowing for rapid and inexpensive characterization of target analytes by altering the interaction between the monomer and analyte through modifications in specific regions of the cytotoxin K sequence, particularly between positions S100 and K170, enhancing current range and discrimination capabilities.

Benefits of technology

The mutant cytotoxin K monomers enable improved discrimination and characterization of polypeptides and polynucleotides, offering high sensitivity and specificity in nanopore-based methods, facilitating rapid and cost-effective sequencing and analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The present invention relates to mutant forms of cytotoxin K. The present invention also relates to methods of analyte detection and characterization using cytotoxin K, together with devices and kits for carrying out such methods.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to mutant forms of cytotoxin K. The present invention also relates to methods of analyte detection and characterization using cytotoxin K, together with devices and kits for carrying out such methods. [Background technology]

[0002] Nanopore sensing is an approach to sensing that relies on the observation of individual binding or interaction events between analyte molecules and a detector. Nanopore sensors can be made by placing a single pore of nanometer dimensions in an insulating film and measuring the voltage-driven ionic transport through the pore in the presence of analyte molecules. The identity of the analyte is revealed through its unique current signature, in particular the duration and extent of the current block, as well as the discrepancy in current levels. Such nanopore sensors are commercially available, such as the MinION™ device sold by Oxford Nanopore Technologies Ltd, which contains an array of nanopores integrated with an electronic chip.

[0003] There is currently a need for a fast and inexpensive nucleic acid (e.g., DNA or RNA) sequencing technology that can be used for a wide range of applications. Existing technologies are time-consuming and expensive, mainly because they rely on amplification techniques to generate large amounts of nucleic acids and require large amounts of specialized fluorescent reagents for signal detection. Nanopore sensing has the potential to provide fast and inexpensive nucleic acid sequencing by reducing the amount of nucleotides and reagents required.

[0004] Furthermore, there is currently a need for new techniques for characterizing polypeptides, especially at the single molecule level. Single molecule techniques for characterizing biomolecules such as polynucleotides have proven particularly attractive due to their high fidelity and avoidance of amplification bias.

[0005] While techniques for characterizing (e.g., sequencing) polynucleotides have been widely developed, techniques for characterizing polypeptides have made less progress, despite their enormous biotechnological importance. For example, knowledge of protein sequences allows the establishment of structure-activity relationships, which has influenced rational drug development strategies for developing ligands for specific receptors. Identification of post-translational modifications is also key to understanding the functional properties of many proteins. For example, in eukaryotes, typically 30-50% of protein species are phosphorylated. Some proteins may have multiple phosphorylation sites that serve to activate or inactivate the protein, promote its degradation, or regulate interactions with protein partners. Thus, there is a pressing need for methods to characterize proteins and other polypeptides.

[0006] Known methods for characterizing polypeptides include mass spectrometry and Edman degradation.

[0007] Protein mass spectrometry involves characterizing whole proteins or fragments thereof in ionized form. Known methods of protein mass spectrometry include electrospray ionization (ESI) and matrix-assisted laser desorption / ionization (MALDI). Although mass spectrometry has several advantages, the results obtained can be affected by the presence of contaminants and fragile molecules can be difficult to process without fragmentation. Furthermore, mass spectrometry is not a single molecule technique and provides only bulk information about the investigated sample. Mass spectrometry is inappropriate for characterizing differences within a population of polypeptide samples and is cumbersome when trying to distinguish adjacent residues.

[0008] Edman degradation is an alternative to mass spectrometry that allows for residue-by-residue sequencing of polypeptides. Edman degradation sequences polypeptides by sequentially cleaving the N-terminal amino acids and characterizing the individually cleaved residues using chromatography or electrophoresis. However, Edman sequencing is slow, requires the use of expensive reagents, and is not a single-molecule technique like mass spectrometry.

[0009] One attractive method for single molecule characterization of biomolecules such as polypeptides is nanopore sensing. Nanopore sensing is an approach for analyte detection and characterization that relies on the observation of individual binding or interaction events between analyte molecules and ion-conducting channels. Nanopore sensors can be made by placing a single pore of nanometer dimensions in an electrically insulating membrane and measuring the voltage-driven ionic current through the pore in the presence of analyte molecules. The presence of an analyte inside or near the nanopore alters the ionic flow through the pore, resulting in a change in the ionic or electrical current measured on the channel. The identity of the analyte is revealed through its unique current signature, particularly the duration and extent of the current blockade, and the discrepancy in the current level during its interaction time with the pore. Nanopore sensing has the potential to enable rapid and inexpensive characterization of polypeptides.

[0010] Nanopore sensing and characterization of polypeptides has been proposed in the art, for example in WO2013 / 123379 and WO2021 / 111125. However, there remains a need for alternative and / or improved methods of characterizing polypeptides.

[0011] Two of the key components of using nanopore sensing to characterize analytes such as nucleic acids and amino acids are (1) control of the movement of the analyte through the pore, and (2) identification of the analyte as it moves through the pore. In the past, to achieve analyte identification, the analyte has been passed through a mutant of hemolysin. This provided the current signature that has been shown to be analyte dependent.

[0012] Although the current range for analyte discrimination has been improved through the modification of the hemolysin pore, new nanopore-based systems would have higher performance if the current difference between analytes could be further improved. Furthermore, it would be of great benefit to the proteomics field to provide new and / or alternative systems that can be used for the characterization of polypeptide analytes. Summary of the Invention

[0013] The present disclosure relates to mutant cytotoxin K monomers capable of forming pores for use in methods for characterization of target analytes.

[0014] Accordingly, the present invention provides a method for characterizing a target analyte, comprising the steps of: (a) contacting a target analyte with a pore comprising at least one mutant cytotoxin K monomer comprising a variant of the amino acid sequence of SEQ ID NO:1 such that the target analyte moves relative to the pore; contacting, wherein the variant comprises one or more modifications at one or more positions within the region of SEQ ID NO:1 between about S100 and about K170 that alter the ability of the monomer to interact with an analyte; (b) taking one or more measurements characteristic of the analyte as it migrates relative to the pore; and a method for characterizing a target analyte.

[0015] The present invention also provides mutant cytotoxin K monomers comprising a variant of the amino acid sequence of SEQ ID NO:1, wherein the monomer is capable of forming a pore, and the variant comprises one or more modifications at one or more positions within the region of SEQ ID NO:1 between about S100 and about K170 that alter the ability of the monomer to interact with an analyte.

[0016] The present invention also provides a construct comprising two or more covalently linked monomers derived from cytotoxin K, wherein at least one of the monomers is a mutant cytotoxin K monomer as defined in accordance with the present invention.

[0017] The present invention also provides a polynucleotide encoding a mutant cytotoxin K monomer according to the invention or a construct according to the invention.

[0018] The present invention also provides a homo-oligomeric pore comprising a plurality of mutant monomers according to the invention, said pore preferably being a heptameric pore.

[0019] The present invention also provides a hetero-oligomeric pore comprising at least one mutant monomer according to the invention, said pore preferably being a heptameric pore.

[0020] The present invention also provides a pore comprising at least one structure according to the present invention.

[0021] The present invention also provides a membrane comprising a pore according to the present invention.

[0022] The invention also provides an array comprising a plurality of membranes according to the invention.

[0023] The invention also provides a device comprising an array of the invention, means for applying a potential across the membrane, and means for detecting an electrical or optical signal across the membrane.

[0024] The present invention also provides a method for characterizing a target analyte, comprising the steps of: (a) contacting a target analyte with a pore according to the invention such that the target analyte migrates relative to the pore; (b) taking one or more measurements characteristic of the analyte as it migrates relative to the pore; Methods are also provided whereby a target analyte is characterised.

[0025] The present invention also provides the use of a pore according to the present invention for characterising a target analyte.

[0026] The present invention also provides a method for characterizing a target polypeptide, comprising the steps of: (a) contacting a target polypeptide with a cytotoxin K pore such that the target analyte translocates relative to the pore; (b) taking one or more measurements characteristic of the polypeptide as it moves relative to the pore; Methods are also provided whereby target polypeptides are characterised.

[0027] The present invention also provides the use of the cytotoxin K pore to characterize a target polypeptide.

[0028] The present invention also provides a kit for characterising a target analyte comprising (a) a pore according to the present invention and (b) a polynucleotide binding protein or a polypeptide handling enzyme. [Brief description of the drawings]

[0029] [Figure 1] Pairwise sequence alignment of CytK and aHL performed using Clustalx version 2.1. The transmembrane beta barrel of aHL is indicated by three boxes. Sp|P09616|HLA_STAAU is aHL and tr|A7GM18|A7GM18_BACCN is CytK. [Diagram 2] Structural model of the CytK pore. The model was made using the aHL structure as a template for CytK, which was obtained from the Protein Data Bank (accession code 7AHL). The CytK model was made using Modeller software. The top row shows a cartoon representation of the CytK model, while the bottom row shows a surface representation. The left image of the bottom row shows a cross section through the pore. [Diagram 3]Figure 1. Predicted amino acid sequence of the CytK transmembrane beta barrel. The predicted central regions of the three major constrictions are indicated by dashed boxes. Any residue with a number corresponds to a residue predicted to point into the cavity of the pore. Any residue without a number corresponds to a residue predicted to point towards the membrane. [Figure 4] Comparison of radial profiles of CytK and aHL channels generated using HOLE mapping software. The CytK model was generated using the aHL structure as a template, which was obtained from the Protein Data Bank (accession code 7AHL). [Diagram 5] Ionic current profiles through aHL wild-type and CytK wild-type and mutants as the voltage was gradually increased in 25 mV steps every 30 s in both the negative and positive directions from (-)25 mV to (-)200 mV. The applied voltage is indicated by the dashed line (blue line in the original color image), the raw current trace is indicated by the gray line (black line in the original color image), and the event detection signal is indicated by the black line (red line in the original color image). [Figure 6] Average ionic current profiles through aHL wild type and CytK wild type as the voltage is gradually increased in 25 mV steps every 30 s in both the negative and positive directions from (-)25 mV to (-)200 mV. The top row shows the average current within the voltage step grouped by either run (left) or pore batch (right). The bottom row shows the average current for the first 100 ms within the voltage step grouped by either run (left) or pore batch (right). Plotting the average current for the first 100 ms reduces the effect of pore gating on the measured current. Pore batch A = aHL-(WT), pore batch B = CytK-(WT-H6), pore batch C = CytK-(WT-H6), pore batch D = CytK-(WT-H6-D8), pore batch E = CytK-(WT-H6-D8). [Figure 7]Average ionic current profiles through CytK wild type and CytK mutants as the voltage is gradually increased in 25 mV steps in both the negative and positive directions from (-)25 mV to (-)200 mV. Panels 1 and 3 (top row of original images) show the average currents within voltage steps grouped by either run (panel 1) or pore batch (panel 3). Panels 2 and 4 (bottom row of original images) show the average currents for the first 100 ms within voltage steps grouped by either run (panel 2) or pore batch (panel 4). Plotting the average currents for the first 100 ms reduces the effect of pore gating on the measured currents. Pore ​​batch B = CytK-(WT-H6), Pore batch C = CytK-(WT-H6), Pore batch D = CytK-(WT-H6-D8), Pore batch E = CytK-(WT-H6-D8), Pore batch F = CytK-(WT-E113S / K156S-D8), Pore batch G = CytK-(WT-Q123S / Q146S-D8), Pore batch H = CytK-(WT-K129S / E140S-D8), Pore batch I = CytK-(WT-Q123S / Q146S / K129S / E140S-D8), pore batch J=CytK-(WT-Q123S / Q146S / K129S / E140S-D8), pore batch K=CytK-(WT-E113S / K156S / Q123S / Q146S / K129S / E140S), pore batch L=CytK-(WT-E113N / K156S / Q123S / Q146S / K129S / E140S-D8). [Figure 8] Current versus time traces as DNA translocates through aHL wild type and CytK wild type and mutants. Raw current traces are shown as grey lines (black lines in original color images) and event detection signals are shown as black lines (red lines in original color images). For each pore, the top row shows the full DNA current trace, the middle row shows the first section of the current trace, and the bottom row shows a zoomed-in view of the first section of the current trace. [Figure 9]1 is a table summarizing the pore characteristics of CytK wild type and mutants. SNR is the signal-to-noise ratio, which is the range of the signal when DNA translocates through the pore divided by the noise. Median current is the median current of the signal when DNA translocates through the pore. [Figure 10] Box plots showing pore characteristics of CytK wild type and mutants. SNR is the signal-to-noise ratio, which is the range of signal when DNA translocates through the pore divided by the noise. Median current is the median current of the signal when DNA translocates through the pore. [Figure 11] 1 is a bar graph showing pore properties of CytK wild type and mutants under condition 7, which is 1 mM ATP, 10 mM MgCl2, 100 nM Hel308 mutant, 1 M NaCl, pH 8, 100 mM HEPES, 10 mM potassium ferrocyanide, 10 mM potassium ferricyanide, 180 mV. [Figure 12] 1 is a bar graph showing pore properties of CytK wild type and mutants under condition 9, which is 1 mM ATP, 10 mM MgCl2, 100 nM He1308 mutant, 625 mM KCl, pH 8, 100 mM HEPES, 75 mM potassium ferrocyanide, 25 mM potassium ferricyanide. [Figure 13] Polynucleotide-polypeptide conjugates used to translocate peptides through nanopores. [Figure 14] 1 is an exemplary current versus time trace when a polynucleotide-polypeptide conjugate is translocated through CytK wild type and mutant, where the polypeptide section comprises GGSGRRSGSG. The irregularly curved peptide section is highlighted with a box (red box in original color image). The trace begins with a long flat section corresponding to capture of the C3 leader on the adaptor. [Figure 15]1 is an exemplary current versus time trace when a polynucleotide-polypeptide conjugate is translocated through the CytK mutant CytK-(WT-Q123S / Q146S / K129S / E140S), where the polypeptide section comprises either GGSGRRSGSG, GGSGYYSGSG, or GGSGDDSGSG. The irregularly curvilinear peptide section is highlighted with a box (red box in original color image). [Figure 16] DNA sequencing Y-adapter used to translocate ssDNA through the nanopore. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0030] The present invention has been described with respect to certain embodiments and with reference to certain drawings, but the present invention is not limited thereto, but only by the claims. Any reference signs in the claims should not be interpreted as limiting the scope. Of course, it should be understood that not necessarily all aspects or advantages can be achieved according to any particular embodiment of the present invention. Thus, for example, a person skilled in the art will recognize that the present invention can be embodied or performed in a manner that achieves or optimizes one advantage or group of advantages taught herein, but does not necessarily achieve other aspects or advantages that may be taught or suggested herein.

[0031] The present invention, both as to its construction and method of operation, together with its features and advantages, may be best understood by reference to the following detailed description when read in conjunction with the accompanying drawings. Aspects and advantages of the present invention will become apparent from and be elucidated with reference to the embodiment(s) described below. Throughout this specification, reference to "one embodiment" or "an embodiment" means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present invention. Thus, the appearance of the phrases "in one embodiment" or "in an embodiment" in various places throughout this specification do not necessarily all refer to the same embodiment, but may. Similarly, in describing exemplary embodiments of the present invention, it should be understood that various features of the invention may be grouped together in a single embodiment, figure, or description thereof for the purpose of streamlining the disclosure and aiding in the understanding of one or more of the various aspects of the present invention. However, this method of disclosure is not to be interpreted as reflecting an intention that the claimed invention requires more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive aspects lie in less than all features of a single foregoing disclosed embodiment.

[0032] It is to be understood that "embodiments" of the present disclosure may be specifically combined together, unless the context indicates otherwise. Any specific combination of the disclosed embodiments is a further disclosed embodiment of the claimed invention (unless the context implies otherwise).

[0033] Additionally, as used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to a "polynucleotide" includes two or more polynucleotides, reference to a "helicase" includes two or more helicases, reference to a "monomer" refers to two or more monomers, reference to a "pore" includes two or more pores, etc.

[0034] All publications, patents, and patent applications cited herein, whether supra or infra, are hereby incorporated by reference in their entirety.

[0035] definition When an indefinite or definite article is used when referring to a singular noun, e.g., "a" or "an" or "the," this includes the plural of that noun, unless otherwise specified. When the term "comprising" is used in the present specification and claims, it does not exclude other elements or steps. Furthermore, the terms first, second, third, etc. in the description and claims are used to distinguish between similar elements and are not necessarily used to describe a sequential or chronological order. It is to be understood that the terms used in this manner are interchangeable under appropriate circumstances and that the embodiments of the invention described herein can operate in sequences other than those described or illustrated herein. The following terms or definitions are provided solely to aid in the understanding of the present invention. Unless otherwise specifically defined herein, all terms used herein have the same meaning as they would to one skilled in the art of the present invention. Skilled artisans are particularly accustomed to the definitions and technical terms of the art in question, as described in Sambrook et al., Molecular Cloning: A Laboratory Manual, 4 thed., Cold Spring Harbor Press, Plainsview, New York (2012), and Ausubel et al., Current Protocols in Molecular Biology (Supplement 114), John Wiley & Sons, New York (2016). The definitions provided herein should not be construed to have a scope less than understood by a person of ordinary skill in the art.

[0036] As used herein, "about" when referring to a measurable value, such as an amount, duration, and the like, is meant to encompass variations of ±20% or ±10%, more preferably ±5%, even more preferably ±1%, and even more preferably ±0.1% from the particular value, such variations being appropriate for carrying out the disclosed methods.

[0037] As used herein, "nucleotide sequence," "DNA sequence," or "nucleic acid molecule(s)" refers to a polymeric form of nucleotides of any length, either ribonucleotides or deoxyribonucleotides. The term refers only to the primary structure of the molecule. Thus, the term includes double- and single-stranded DNA, as well as RNA. As used herein, the term "nucleic acid" is a single- or double-stranded covalently linked nucleotide sequence in which the 3' and 5' ends on each nucleotide are linked by a phosphodiester bond. Polynucleotides can be composed of deoxyribonucleotide or ribonucleotide bases. Nucleic acids can be produced synthetically in vitro or isolated from natural sources. Nucleic acids can further include modified DNA or RNA, such as methylated DNA or RNA, or RNA that has undergone post-translational modifications, such as 5'-capping with 7-methylguanosine, 3'-processing such as truncation and polyadenylation, and splicing. Nucleic acids can also include synthetic nucleic acids (XNA), such as hexitol nucleic acid (HNA), cyclohexene nucleic acid (CeNA), threose nucleic acid (TNA), glycerol nucleic acid (GNA), locked nucleic acid (LNA) and peptide nucleic acid (PNA). The size of a nucleic acid, also referred to herein as a "polynucleotide", is typically expressed as the number of base pairs (bp) for double-stranded polynucleotides or the number of nucleotides (nt) for single-stranded polynucleotides. 1000 bp or nt is equivalent to a kilobase (kb). Polynucleotides less than about 40 nucleotides in length are typically referred to as "oligonucleotides" and can include primers for use in manipulating DNA, such as via polymerase chain reaction (PCR).

[0038] The term "amino acid" in the context of this disclosure is used in its broadest sense and is meant to include organic compounds containing an amine (NH2) functional group and a carboxyl (COOH) functional group, along with a side chain (e.g., R group) specific to each amino acid. In some embodiments, amino acid refers to a naturally occurring Lα-amino acid or residue. Commonly used one-letter and three-letter abbreviations for naturally occurring amino acids are used herein: A=Ala, C=Cys, D=Asp, E=Glu, F=Phe, G=Gly, H=His, I=Ile, K=Lys, L=Leu, M=Met, N=Asn, P=Pro, Q=Gln, R=Arg, S=Ser, T=Thr, V=Val, W=Trp, and Y=Tyr (Lehninger, AL, (1975) Biochemistry, 2d ed., pp. 71-92, Worth Publishers, New York). The general term "amino acid" further includes D-amino acids, retro-inverso amino acids, and chemically modified amino acids such as amino acid analogs, naturally occurring amino acids that are not normally incorporated into proteins, such as norleucine, and chemically synthesized compounds that have properties known in the art to be characteristic of amino acids, such as β-amino acids. For example, analogs or mimetics of phenylalanine or proline that allow the same conformational restriction of peptide compounds as natural Phe or Pro are included within the definition of amino acids. Such analogs and mimetics are referred to herein as "functional equivalents" of the respective amino acids. Other examples of amino acids are described in Roberts and Vellaccio, The Peptides: Analysis, Synthesis, Biology, Gross and Meiehofer, eds., Vol. 5 p. 341, Academic Press, Inc., NY 1983, which is incorporated herein by reference.

[0039] The terms "polypeptide" and "peptide" are used interchangeably herein to refer to a polymer of amino acid residues, as well as variants and synthetic analogs thereof. Thus, these terms apply to amino acid polymers in which one or more amino acid residues are synthetic, non-naturally occurring amino acids, such as chemical analogs of the corresponding naturally occurring amino acids, as well as to naturally occurring amino acid polymers. Polypeptides may also undergo maturation or post-translational modification processes, including, but not limited to, glycosylation, proteolytic cleavage, lipidation, signal peptide cleavage, propeptide cleavage, phosphorylation, and the like. Peptides may be produced using recombinant techniques, for example, through expression of recombinant or synthetic polynucleotides. Recombinantly produced peptides are typically substantially free of culture medium, e.g., culture medium represents less than about 20% of the volume of the protein preparation, more preferably less than about 10%, and most preferably less than about 5%.

[0040] The term "protein" is used to describe a folded polypeptide having secondary or tertiary structure. A protein may consist of a single polypeptide or may contain multiple polypeptides that assemble to form a multimer. A multimer may be a homo- or hetero-oligomer. A protein may be a naturally occurring or wild-type protein, or a modified or non-naturally occurring protein. A protein may differ from a wild-type protein, for example, by the addition, substitution, or deletion of one or more amino acids.

[0041] A "variant" of a protein includes peptides, oligopeptides, polypeptides, proteins, and enzymes that have amino acid substitutions, deletions, and / or insertions compared to the unmodified or wild-type protein in question and have biological and functional activities similar to the unmodified protein from which they are derived. The term "amino acid identity" as used herein refers to the degree to which sequences are identical amino acid-by-amino acid across a comparison window. Thus, "percentage of sequence identity" is calculated by comparing two optimally aligned sequences across a comparison window, determining the number of positions at which identical amino acid residues (e.g., Ala, Pro, Ser, Thr, Gly, Val, Leu, Ile, Phe, Tyr, Trp, Lys, Arg, His, Asp, Glu, Asn, Gln, Cys, and Met) appear in both sequences to obtain the number of matched positions, dividing the number of matched positions by the total number of positions in the comparison window (i.e., window size), and multiplying the result by 100 to obtain the percentage of sequence identity.

[0042] In all aspects and embodiments of the invention, a "variant" has at least 50%, 60%, 70%, 80%, 90%, 95%, or 99% complete sequence identity with the amino acid sequence of the corresponding wild-type protein. Sequence identity may also be to a fragment or portion of a full-length polynucleotide or polypeptide. Thus, a sequence may have only 50% overall sequence identity with a full-length reference sequence, but the sequence of a particular region, domain, or subunit may share 80%, 90%, or even 99% sequence identity with the reference sequence.

[0043] The term "wild type" refers to a gene or gene product isolated from a natural source. A wild type gene is the gene most frequently observed in a population, and is therefore arbitrarily designated the "normal" or "wild type" form of the gene. In contrast, the terms "modified," "mutant," or "variant" refer to a gene or gene product that exhibits modifications in sequence (e.g., substitutions, truncations, or insertions), post-translational modifications, and / or functional properties (e.g., altered characteristics) when compared to the wild type gene or gene product. It should be noted that naturally occurring mutants can be isolated, which are identified by the fact that they have altered characteristics compared to the wild type gene or gene product. Methods for introducing or substituting naturally occurring amino acids are well known in the art. For example, methionine (M) can be substituted with arginine (R) by replacing the codon for methionine (ATG) with the codon for arginine (CGT) at the relevant position in the polynucleotide encoding the mutant monomer. Methods for introducing or substituting non-naturally occurring amino acids are also well known in the art. For example, non-naturally occurring amino acids can be introduced by including synthetic aminoacyl-tRNA in the IVTT system used to express the mutant monomers. Alternatively, they can be introduced by expressing mutant monomers in E. coli that are auxotrophic for specific amino acids in the presence of synthetic (i.e., non-naturally occurring) analogs of those specific amino acids. They can also be generated by naked ligation when the mutant monomers are generated using partial peptide synthesis. Conservative substitutions replace amino acids with other amino acids of similar chemical structure, similar chemical properties, or similar side chain volume. The amino acids introduced can have similar polarity, hydrophilicity, hydrophobicity, basic, acidic, neutral, or charge as the amino acids they replace. Alternatively, conservative substitutions may introduce another amino acid that is aromatic or aliphatic in place of an existing aromatic or aliphatic amino acid. Conservative amino acid changes are well known in the art and can be selected according to the properties of the 20 major amino acids defined in Table 1 below.If the amino acids have a similar polarity, this can also be determined by reference to the hydrophobicity scale for the amino acid side chains in Table 2. [Table 1] [Table 2]

[0044] The mutant or modified protein, monomer, or peptide can also be chemically modified in any manner and at any site. The mutant or modified monomer or peptide is preferably chemically modified by binding a molecule to one or more cysteines (cysteine ​​binding), binding a molecule to one or more lysines, binding a molecule to one or more unnatural amino acids, enzymatic modification of an epitope, or modification of a terminal. Suitable methods for carrying out such modifications are well known in the art. The modified protein, monomer, or peptide variant can be chemically modified by binding any molecule. For example, the modified protein, monomer, or peptide variant can be chemically modified by binding a dye or fluorophore.

[0045] Mutant cytotoxin K monomer The present invention provides a method for characterising an analyte using a pore comprising at least one mutant cytotoxin K (CytK) monomer.

[0046] The present invention also provides mutant cytotoxin K (CytK) monomers. Mutant CytK monomers can be used to form the pores of the present invention. Mutant CytK monomers are monomers whose sequence differs from wild-type CytK monomer (SEQ ID NO: 1) and retain the ability to form pores. Methods for confirming the ability of mutant monomers to form pores are well known in the art and are described in more detail below. For example, the ability of mutant monomers to form pores can be determined as described in the examples.

[0047] The pores comprising the mutant monomers of the present invention have an increased current range when subjected to a potential applied in a nanopore-based analyte characterization method, compared to a pore composed of a wild-type CytK monomer. The increased current range facilitates identifying and characterizing the target analyte, and in particular facilitates discrimination between components of the target analyte. For example, when the target analyte is a polypeptide, the increased current range facilitates discrimination between amino acids within the polypeptide.

[0048] Pores comprising mutant CytK monomers of the invention can be used to characterize any suitable analyte. Suitable analytes are further described herein. The increased current range in particular makes pores comprising mutant CytK monomers of the invention particularly applicable to the nanopore-based methods of characterizing polypeptide analytes described herein. Techniques for characterizing polypeptides are of significant biotechnological importance. For example, knowledge of protein sequences allows structure-activity relationships to be established, which has influenced rational drug development strategies for developing ligands for specific receptors. Identification of post-translational modifications is also key to understanding the functional properties of many proteins. For example, in eukaryotes, typically 30-50% of protein species are phosphorylated. Some proteins may have multiple phosphorylation sites that serve to activate or inactivate the protein, promote its degradation, or regulate interactions with protein partners. Described herein is the successful utilization of pores comprising mutant CytK monomers in nanopore-based methods of characterizing target polypeptides. Thus, the inventors have surprisingly identified a novel means for characterizing polypeptide analytes.

[0049] The inventors have surprisingly identified a region within the CytK monomer that can be modified to alter the interaction between the monomer and an analyte, such as when the analyte is characterized using a nanopore-based method of analyte characterization described herein, including the use of a pore comprising a CytK mutant monomer of the invention. With reference to the wild-type polypeptide sequence of the CytK monomer defined by SEQ ID NO:1, the region is from about position S100 to about position K170 of SEQ ID NO:1. At least a portion of this region typically contributes to the transmembrane region of CytK. At least a portion of this region typically contributes to the barrel or channel of CytK. At least a portion of this region typically contributes to the inner wall or inner layer of CytK.

[0050] Improved analyte properties of the CytK mutant monomer are achieved by introducing one or more modifications at one or more positions within the region of SEQ ID NO: 1 between about S100 and about K170 that alter the ability of the monomer to interact with an analyte. Preferred mutations are further described herein. Thus, there is provided a mutant CytK monomer comprising a variant of the amino acid sequence of SEQ ID NO: 1, the monomer capable of forming a pore, the variant comprising one or more modifications at one or more positions within the region of SEQ ID NO: 1 between about S100 and about K170 that alter the ability of the monomer to interact with an analyte.

[0051] According to the invention, the variants include one or more modifications at one or more positions within the region of SEQ ID NO:1 between about S100 and K170 that alter the ability of the monomer, or preferably the region, to interact with an analyte. The interaction between the monomer and the analyte can be increased or decreased. An increase in the interaction between the monomer and the analyte, for example, facilitates capture of the analyte by a pore comprising the mutant monomer. A decrease in the interaction between the monomer and the analyte, for example, improves recognition or discrimination of the analyte. The recognition or discrimination of the analyte can be improved by increasing the current range by modifications to the CytK monomer between about S100 and K170 of SEQ ID NO:1 described herein. The improvement in the recognition or discrimination of the analyte can be particularly improved and achieved by five main mechanisms, namely, the following independent modifications: conformation (e.g., increasing or decreasing the size of amino acid residues), the net charge of the amino acid residue at the modification position (e.g. introducing or removing a negative (-ve) charge and / or introducing or removing a positive (+ve) charge); the hydrogen bonding properties of the amino acid residues at the modified positions (e.g., introducing amino acids capable of hydrogen bonding to the analyte); π-stacking (e.g., introducing into or removing from an amino acid residue at a modified position one or more chemical groups that interact via a delocalized electron π-system), and / or · An amino acid residue at a modified position, thereby changing the structure of the pore (e.g., introducing an amino acid that increases or decreases the size of the barrel or channel).

[0052] Thus, the one or more modifications can each independently (a) change the size of the amino acid residue at the modified position, (b) change the net charge of the amino acid residue at the modified position, (c) change the hydrogen bonding properties of the amino acid residue at the modified position, (d) introduce or remove one or more chemical groups that interact through a delocalized electron π-system into or from the amino acid residue at the modified position, and / or (e) change the structure of the amino acid residue at the modified position.

[0053] Any one or more of these mechanisms of independent change may be responsible for the improved properties of pores formed from mutant monomers of the invention For example, pores comprising mutant monomers of the invention may exhibit improved polypeptide and / or polynucleotide reading properties as a result of conformational changes, hydrogen bonding changes, and structural changes.

[0054] Thus, provided herein is a method for characterizing a target analyte, comprising: (a) contacting a target analyte with a pore comprising at least one mutant cytotoxin K monomer comprising a variant of the amino acid sequence of SEQ ID NO:1 such that the target analyte moves relative to the pore; contacting, wherein the variant comprises one or more modifications at one or more positions within the region of SEQ ID NO:1 between about S100 and about K170 that alter the ability of the monomer to interact with an analyte; (b) taking one or more measurements characteristic of the analyte as it migrates relative to the pore; thereby characterizing the target analyte.

[0055] Also provided is a mutant CytK monomer comprising a variant of the amino acid sequence of SEQ ID NO:1, wherein the monomer is capable of forming a pore, and the variant comprises one or more modifications at one or more positions within the region of SEQ ID NO:1 between about S100 and about K170 that alter the ability of the monomer to interact with an analyte.

[0056] The ability of the monomer to interact with the target analyte and interact with the analyte can be determined using methods well known in the art. The monomer can interact with the analyte in any manner, for example, by non-covalent interactions such as hydrophobic interactions, hydrogen bonds, van der Waals forces, pi(π)-cation interactions, or electrostatic forces. For example, the ability of the region to bind to the analyte can be measured using a conventional binding assay. Suitable assays include, but are not limited to, fluorescence-based binding assays, nuclear magnetic resonance (NMR), isothermal titration calorimetry (ITC), or electron spin resonance (ESR) spectroscopy. Alternatively, the ability of a pore comprising one or more of the mutant monomers to interact with the analyte can be determined using any of the methods discussed above or below. A preferred assay is described in the Examples.

[0057] The one or more modifications are in the region of SEQ ID NO:1 from about position 100 to about position 170. The one or more modifications are preferably in the region of SEQ ID NO:1 from about position 110 to about position 160. Even more preferably, the one or more modifications are in the region of SEQ ID NO:1 from about position 113 to about position 156.

[0058] Modifications of protein nanopores that change their ability to interact with analytes, and in particular improve their current range, are well documented in the art. For example, such modifications are disclosed in WO2010 / 034018, WO2010 / 055307, WO2013 / 153359, and WO2016 / 034591. Similar modifications can be made to the CytK monomer according to the present invention.

[0059] Any number of modifications can be made, such as 1, 2, 5, 10, 15, 20, 30 or more modifications. Any modification(s) can be made, as long as the ability of the monomer to interact with the polynucleotide is changed and the monomer remains capable of forming a pore. Suitable modifications include, but are not limited to, amino acid substitutions, amino acid additions, and amino acid deletions. The one or more modifications are preferably one or more substitutions. This will be discussed in more detail below.

[0060] The one or more modifications preferably (a) change the steric effect of the monomer, or preferably change the steric effect of the region, (b) change the net charge of the monomer, or preferably change the net charge of the region, (c) change the ability of the monomer, or preferably the region, to hydrogen bond with the analyte, (d) introduce or remove chemical groups that interact via a delocalized electron π-system, and / or (e) change the structure of the monomer, or preferably the region. The one or more modifications are more preferably (a) to (e), for example, (a) and (b); (a) and (c); (a) and (d); (a) and (e); (b) and (c); (b) and (d); (b) and (e); (c) and (d); (c) and (e); (d) and (e); (a), (b) and (c); (a), (b) and (d); (a), (b) and (e); (a), (c) and (d); (a), (c and (e); (a), (d) and (e); (b), (c) and (d); (b), (c) and (e); (b), (d) and (e); (c), (d) and (e); (a), (b), (c) and d; (a), (b), (c) and (e); (a), (b), (c), (d) and (e); (a), (b), (d) and (e); (a), (c), (d) and (e); (b), (c), (d) and (e); and (a), (b), (c) and (d).

[0061] For (a), the steric effect of the monomer can be increased or decreased. Any method of changing the steric effect can be used according to the present invention. The introduction of bulky residues such as phenylalanine (F), tryptophan (W), tyrosine (Y), or histidine (H) increases the conformation of the monomer. The one or more modifications are preferably the introduction of one or more of F, W, Y, and H. Any combination of F, W, Y, and H can be introduced. One or more of F, W, Y, and H can be introduced by addition. One or more of F, W, Y, and H are preferably introduced by substitution. Suitable positions for the introduction of such residues are discussed in more detail below.

[0062] Removal of bulky residues such as phenylalanine (F), tryptophan (W), tyrosine (Y), or histidine (H) conversely reduces the conformation of the monomer. The one or more modifications are preferably removal of one or more of F, W, Y, and H. Any combination of F, W, Y, and H may be removed. One or more of F, W, Y, and H may be removed by deletion. One or more of F, W, Y, and H are preferably removed by substituting with a residue having a smaller side group such as serine (S), threonine (T), alanine (A), and valine (V).

[0063] Regarding (b), the net charge can be altered in any manner. The net positive charge is preferably increased or decreased. The net positive charge can be increased in any manner. The net positive charge is preferably increased by introducing one or more positively charged amino acids, preferably by substitution, and / or by neutralizing one or more negative charges, preferably by substitution.

[0064] The net positive charge is preferably increased by introducing one or more positively charged amino acids. The one or more positively charged amino acids may be introduced by addition. The one or more positively charged amino acids are preferably introduced by substitution. A positively charged amino acid is an amino acid that has a net positive charge. The positively charged amino acid(s) may be naturally occurring or non-naturally occurring. The positively charged amino acid may be synthetic or modified. For example, modified amino acids with a net positive charge may be specifically designed for use in the present invention. Several different types of modifications to amino acids are well known in the art. The one or more modifications, including the introduction of one or more positively charged amino acids, preferably include the introduction of one or more of histidine (H), lysine (K), and arginine (R) by substitution or addition, but most preferably by substitution. Suitable positions for the introduction of such residues are discussed in more detail below.

[0065] Methods for adding or substituting naturally occurring amino acids are well known in the art. For example, the nucleotides that make up a codon contained within a polynucleotide coding sequence may be modified to change the nucleotide content of the codon, thereby resulting in a different amino acid being encoded by the codon. Such polynucleotides can then be expressed as discussed below.

[0066] Methods for adding or substituting non-naturally occurring amino acids are also well known in the art. For example, non-naturally occurring amino acids can be introduced by including synthetic aminoacyl-tRNA in the IVTT system used to express the pore. Alternatively, they can be introduced by expressing monomers in E. coli that are auxotrophic for specific amino acids in the presence of synthetic (i.e., non-naturally occurring) analogs of those specific amino acids. They can also be generated by naked ligation when the pore is generated using partial peptide synthesis.

[0067] In one or more modifications, any amino acid may be replaced with a positively charged amino acid. In one or more modifications, one or more uncharged, nonpolar, and / or aromatic amino acids may be replaced with one or more positively charged amino acids. An uncharged amino acid has no net charge. Suitable uncharged amino acids include, but are not limited to, cysteine ​​(C), serine (S), threonine (T), methionine (M), asparagine (N), and glutamine (Q). A nonpolar amino acid has a nonpolar side chain. Suitable nonpolar amino acids include, but are not limited to, glycine (G), alanine (A), proline (P), isoleucine (I), leucine (L), and valine (V). Aromatic amino acids have an aromatic side chain. Suitable aromatic amino acids include, but are not limited to, histidine (H), phenylalanine (F), tryptophan (W), and tyrosine (Y). Preferably, one or more modifications involve the replacement of one or more negatively charged amino acids with one or more positively charged amino acids. Suitable negatively charged amino acids include, but are not limited to, aspartic acid (D) and glutamic acid (E).

[0068] Preferred introductions of one or more modifications include, but are not limited to, substitution of E with K, substitution of M with R, substitution of M with H, substitution of M with K, substitution of D with R, substitution of D with H, substitution of D with K, substitution of E with R, substitution of E with H, substitution of N with R, substitution of T with R, and substitution of G with R. Most preferably, E is substituted with K.

[0069] In the one or more modifications, any number of positively charged amino acids can be introduced or substituted, for example, 1, 2, 5, 10, 15, 20, 25, 30 or more positively charged amino acids can be introduced or substituted.

[0070] The net positive charge is more preferably increased by neutralizing one or more negative charges. One or more negative charges can be neutralized by substituting one or more negatively charged amino acids with one or more uncharged, non-polar, and / or aromatic amino acids. Removal of the negative charges increases the net positive charge. The uncharged, non-polar, and / or aromatic amino acids may be naturally occurring or non-naturally occurring. They may be synthetic or modified. Suitable uncharged, non-polar, and aromatic amino acids are discussed above. Preferred substitutions include, but are not limited to, E with Q, E with S, E with A, D with Q, E with N, D with N, D with G, and D with S.

[0071] Any number and combination of uncharged, nonpolar, and / or aromatic amino acids can be substituted in one or more modifications. For example, 1, 2, 5, 10, 15, 20, 25, or 30 or more uncharged, nonpolar, and / or aromatic amino acids can be substituted. A negatively charged amino acid can be substituted with (1) an uncharged amino acid, (2) a nonpolar amino acid, (3) an aromatic amino acid, (4) an uncharged and nonpolar amino acid, (5) an uncharged and aromatic amino acid, and (5) a nonpolar and aromatic amino acid, or (6) an uncharged, nonpolar, and aromatic amino acid.

[0072] One or more negative charges can be neutralized by introducing one or more positively charged amino acids close to, such as within 1, 2, 3, or 4 amino acids, or adjacent to, one or more negatively charged amino acids. Examples of positively and negatively charged amino acids are discussed above. Positively charged amino acids can be introduced in any manner discussed above, for example by substitution.

[0073] The net positive charge is preferably reduced by introducing one or more negatively charged amino acids and / or neutralizing one or more positive charges. How this can be done will be clear from the above description with reference to increasing the net positive charge. All of the embodiments discussed above in relation to increasing the net positive charge apply equally to reducing the net positive charge, except that the charge is changed in the opposite manner. In particular, the one or more positive charges are preferably neutralized by substituting one or more positively charged amino acids with one or more uncharged amino acids, nonpolar amino acids, and / or aromatic amino acids, or by introducing one or more negatively charged amino acids close to the one or more negatively charged amino acids, for example within 1, 2, 3, or 4 amino acids, or adjacent to the one or more negatively charged amino acids.

[0074] The net negative charge is preferably increased or decreased. All of the above embodiments discussed above with reference to increasing or decreasing the net positive charge apply equally to decreasing or increasing the net negative charge, respectively.

[0075] Regarding (c), the hydrogen bonding ability of the monomer may be altered in any suitable manner. For example, the one or more modifications may include introducing one or more of serine (S), threonine (T), asparagine (N), glutamine (Q), tyrosine (Y) or histidine (H) by addition or substitution, thereby increasing the hydrogen bonding ability of the monomer. The one or more modifications preferably include introducing one or more of S, T, N, Q, Y, and H in any suitable combination, preferably the introduction is by substitution. Suitable locations for the introduction of such residues are discussed in more detail below.

[0076] Removal of serine (S), threonine (T), asparagine (N), glutamine (Q), tyrosine (Y) or histidine (H) reduces the hydrogen bonding ability of the monomer. For example, the one or more modifications can include removal of one or more of S, T, N, Q, Y, and H. The one or more modifications preferably include removal of any combination of S, T, N, Q, Y, and H by deletion or by substitution in any suitable combination, thereby reducing the hydrogen bonding ability of the monomer. The one or more modifications preferably include substitution with other amino acids that are less good hydrogen bonders, such as alanine (A), valine (V), isoleucine (I), and leucine (L).

[0077] For (d), introduction of aromatic residues such as phenylalanine (F), tryptophan (W), tyrosine (Y) or histidine (H) also increases pi-stacking in the monomer. Removal of aromatic residues such as phenylalanine (F), tryptophan (W), tyrosine (Y) or histidine (H) also increases pi-stacking in the monomer. Such amino acids may be introduced or removed as discussed above with reference to (a).

[0078] Regarding (e), one or more modifications made according to the present invention that change the structure of the monomer. For example, one or more loop regions can be removed, shortened, or extended. This typically facilitates the entry and exit of the polynucleotide into the pore. One or more loop regions can be on the cis side of the pore, the trans side of the pore, or on both sides of the pore. Alternatively, one or more regions at the amino and / or carboxy termini of the pore can be extended or deleted. This typically changes the size and / or charge of the pore.

[0079] From the above discussion, it will be apparent that the introduction of certain amino acids enhances the ability of the monomer to interact with the analyte via more than one mechanism. For example, the substitution of E with H not only increases the net positive charge (by neutralizing the negative charge) according to (b), but also increases the hydrogen bonding ability of the monomer according to (c).

[0080] The inventors have surprisingly identified three constrictions within a pore consisting of wild-type CytK monomers. A constriction is typically a narrowing of the channel through the nanopore that may determine or control the signal obtained in any of the known nanopore-based methods of analyte characterization or any of the methods of analyte characterization described herein when an analyte translocates relative to the nanopore. The structure of each CytK monomer within the pore results in the formation of three constrictions in the barrel region of the pore. The amino acids involved in the formation of the three constrictions are included between about S100 and K170.

[0081] Thus, the mutant CytK monomers of the invention may contain one or more modifications at one or more positions within the region of SEQ ID NO:1 between about S100 and K170 that alter the ability of the monomer to interact with an analyte, the modifications altering one or more of the three constrictions in a pore comprising a CytK monomer of the invention compared to a pore comprised of a wild-type CytK monomer. The modifications may therefore alter the interaction of the analyte with the constriction as the analyte moves through the pore. Preferably, the monomers of the invention are capable of forming a pore having a solvent-accessible channel from a first opening to a second opening of the pore, the solvent-accessible channel comprising at least one constriction, and the one or more modifications are made to amino acids within the constriction. Thus, by modifying the regions of the CytK monomer involved in forming the three constrictions in a wild-type CytK pore, the interaction between the CytK monomer and the analyte may be altered, for example, when the analyte is characterized using a nanopore-based method of analyte characterization described herein.

[0082] The amino acids involved in the formation of the three constrictions are included between about S100 and K170 of SEQ ID NO: 1, which defines the CytK monomer, and preferably face inward in the channel region when the CytK monomer assembles to form the CytK pore. Thus, preferably, one or more modifications that change the characteristics of the constriction region of the CytK monomer of the present invention, compared to the wild-type CytK monomer, are made to amino acids that face inwardly to the channel region where the CytK monomer assembles to form the CytK pore. The amino acids involved in the contribution of a single CytK monomer to a constriction in the CytK pore typically comprise a pair of amino acids in the CytK monomer. Thus, the one or more modifications to the amino acids involved in the formation of the three constrictions are included between about S100 and K170 of SEQ ID NO: 1, and preferably are modifications to a pair of amino acids. Such pairs of amino acids are further described herein.

[0083] Thus, the one or more modifications that alter the constriction portion can each independently (a) alter the size of the constriction portion (e.g., by increasing or decreasing the size of the amino acid residue at the modified position), (b) alter the net charge of the constriction portion (e.g., by changing the net charge of the amino acid residue at the modified position), (c) alter the hydrogen bonding properties of the amino acid residue at the constriction portion (e.g., by changing the hydrogen bonding properties of the amino acid residue at the modified position), (d) introduce or remove one or more chemical groups that interact through a delocalized electron π-system into or from the constriction portion (e.g., by introducing or removing one or more chemical groups that interact through a delocalized electron π-system into or from the amino acid residue at the modified position), and / or (e) alter the structure of the constriction portion (e.g., by altering the structure of the amino acid residue at the modified position). The one or more modifications that alter (a)-(e) with respect to the constriction portion can generally be the same as those described herein with respect to the monomer.

[0084] As described herein (see particularly the Examples), the inventors have identified three constrictions in a wild-type CytK pore (i.e., a CytK pore composed of wild-type CytK monomers). The loop region of each CytK monomer comprises amino acids that define the three constrictions of the wild-type CytK pore. Each constriction is defined by amino acids on opposite sides of the loop region. The upper constriction (closest to the cap region of the pore) is preferably defined by the region of SEQ ID NO: 1 between about X109 to about T117, more preferably V111 to T115, and about S152 to about X160, preferably S154 to X158. The lower constriction (furthest from the cap region of the pore) is preferably defined by the region of SEQ ID NO: 1 between about G126 and about V132, preferably between S127 and S131, and between about P137 and about A143, preferably between S138 and G142. The middle constriction (furthest from the cap region of the pore) is preferably defined by the region of SEQ ID NO: 1 between about S119 to about G126, preferably between S121 to G125, and between about A143 to about S150, preferably between T144 to T148.

[0085] In wild-type CytK, amino acids from about V111 to about S131 of SEQ ID NO: 1 and from about S138 to about T158 of SEQ ID NO: 1 form a loop region of the pore that includes three constrictions (see FIG. 3). Preferably, the amino acids that form the three constrictions include amino acids in the loop region that faces inward in the channel of the pore. More preferably, amino acids between about V111 to about S131 of SEQ ID NO: 1 are paired with amino acids between S138 to about T158 of SEQ ID NO: 1 to form a constriction in the channel of the pore. Each amino acid in the pair is on the opposite side of the loop region from the other. Thus, in the monomers of the invention described herein, variants may include one or more modifications in the region of SEQ ID NO: 1 between about V111 to about S131 and / or between about S135 to about T158. Preferably, in the monomers of the invention described herein, the variants may include one or more modifications within the regions of SEQ ID NO: 1 between about V111 and about S131, and between about S135 and about T158. In another aspect, in the monomers of the invention described herein, the variants may include 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more modifications between about V111 and about S131 of SEQ ID NO: 1, and 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more modifications between about S135 and about T158 of SEQ ID NO: 1. Most preferably, the same number of modifications are made within the regions of SEQ ID NO: 1 between about V111 and about S131, and between about S135 and about T158.

[0086] In the CytK monomers of the invention, variants may include one or more modifications within the regions of SEQ ID NO: 1 between about S119 and about G126, preferably between S121 and G125, and / or between about A143 and about S150, preferably between T144 and T148. Preferably, in the monomers of the invention described herein, variants may include one or more modifications within the regions of SEQ ID NO: 1 between about S119 and about G126, preferably between S121 and G125, and between about A143 and about S150, preferably between T144 and T148. In another aspect, in the monomers of the invention described herein, variants may include 1, 2, 3, 4, or 5 or more modifications between about S119 and about G126, preferably between S121 and G125 of SEQ ID NO: 1, and 1, 2, 3, 4, or 5 or more modifications between about A143 and about S150, preferably between T144 and T148 of SEQ ID NO: 1. Most preferably, the same number of modifications are made within the region of SEQ ID NO: 1 between about S119 and about G126, preferably between S121 and G125 of SEQ ID NO: 1, and between about A143 and about S150, preferably between T144 and T148 of SEQ ID NO: 1.

[0087] In the CytK monomers of the invention, variants may include one or more modifications within the region of SEQ ID NO: 1 between about G126 and about V132, preferably between S127 and S131, and / or between about P137 and about A143, preferably between S138 and G142. Preferably, in the monomers of the invention described herein, variants may include one or more modifications within the region of SEQ ID NO: 1 between about G126 and about V132, preferably between S127 and S131, and between about P137 and about A143, preferably between S138 and G142. In another aspect, in the monomers of the invention described herein, variants may include 1, 2, 3, 4, or 5 or more modifications between about G126 to about V132, preferably between S127 to S131 of SEQ ID NO: 1, and 1, 2, 3, 4, or 5 or more modifications between about P137 to about A143, preferably between S138 to G142 of SEQ ID NO: 1. Most preferably, the same number of modifications are made within the region of SEQ ID NO: 1 between about G126 to about V132, preferably between S127 to S131 of SEQ ID NO: 1, and between about P137 to about A143, preferably between S138 to G142 of SEQ ID NO: 1.

[0088] In the CytK monomers of the invention, variants may include one or more modifications within the regions of SEQ ID NO: 1 between about N109 and about T117, preferably between V111 and T115, and / or between about S152 and about Y160, preferably between S154 and T158. Preferably, in the monomers of the invention described herein, variants may include one or more modifications within the regions of SEQ ID NO: 1 between about N109 and about T117, preferably between V111 and T115, and between about S152 and about Y160, preferably between S154 and T158. In another aspect, in the monomers of the invention described herein, variants may include 1, 2, 3, 4, or 5 or more modifications between about N109 to about T117, preferably between V111 and T115, of SEQ ID NO: 1, and 1, 2, 3, 4, or 5 or more modifications between about S152 to about Y160, preferably between S154 and T158, of SEQ ID NO: 1. Most preferably, the same number of modifications are made within the region of SEQ ID NO: 1 between about N109 to about T117, preferably between V111 and T115, of SEQ ID NO: 1, and between about S152 to about Y160, preferably between S154 and T158, of SEQ ID NO: 1.

[0089] The variants preferably contain modifications at one or more of the following positions of SEQ ID NO: 1: E113, T115, T117, S119, S121, Q123, G125, S127, K129, S131, V132, T133, P134, S135, G136, P137, S138, E140, G142, T144, Q146, T148, S150, S152, S154 and K156. The variants preferably contain modifications at 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 or more of these positions. The variants may independently contain one or more amino acid substitutions, additions, and / or deletions at one or more of the positions. The amino acids substituted in the variants may be naturally occurring or non-naturally occurring derivatives thereof. The amino acids substituted in the variants may be D-amino acids. In particular, the variants may contain one or more amino acid substitutions at the positions listed above, where the amino acid(s) substituted in the variants are selected from aspartate, glutamate, serine, threonine, asparagine, glutamine, glycine, alanine, valine, leucine, isoleucine, cysteine, arginine, lysine, and phenylalanine.

[0090] The variant preferably includes one or more of the following modifications of SEQ ID NO:1: a)E113S / T / N / Q / G / A / V / L / I / C / R / K / F / Y; b)T115S / N / Q / G / A / V / L / I / C / R / K / F; c)T117S / N / Q / G / A / V / L / I / C / R / K / F; d)S119T / N / Q / G / A / V / L / I / C / R / K / F; e)S121T / N / Q / G / A / V / L / I / C / R / K / F; f)Q123S / T / N / G / A / V / L / I / C / R / K / F / M / Y; g)G125S / T / N / Q / A / V / L / I / C / R / K / F; h)S127T / N / Q / G / A / V / L / I / C / R / K / F; i)K129S / T / N / Q / G / A / V / L / I / C / R / F / Y; j)S131T / N / Q / G / A / V / L / I / C / R / K / F; k)V132S / T / N / Q / G / A / L / I / C / R / K / F; l)T133S / N / Q / G / A / V / L / I / C / R / K / F; m)P134S / T / N / Q / G / A / V / L / I / C / R / K / F; n)S135T / N / Q / G / A / V / L / I / C / R / K / F; o)G136S / T / N / Q / A / V / L / I / C / R / K / F; p)P137S / T / N / Q / G / A / V / L / I / C / R / K / F; q)S138T / N / Q / G / A / V / L / I / C / R / K / F; r)E140S / T / N / Q / G / A / V / L / I / C / R / K / F; s)G142S / T / N / Q / A / V / L / I / C / R / K / F; t)T144S / N / Q / G / A / V / L / I / C / R / K / F; u)Q146S / T / N / G / A / V / L / I / C / R / K / F / M / Y; v)T148S / N / Q / G / A / V / L / I / C / R / K / F; w)S150T / N / Q / G / A / V / L / I / C / R / K / F; x)S152T / N / Q / G / A / V / L / I / C / R / K / F; y)S154T / N / Q / G / A / V / L / I / C / R / K / F; and z)K156S / T / N / Q / G / A / V / L / I / C / R / F.

[0091] The inventors have specifically identified six amino acids that form three pairs in the loop region of wild-type CytK that are believed to be the amino acids involved in the three constrictions in the wild-type CytK pore. Thus, variants may contain modifications to any one or more of the six amino acids, as follows: a) E113; b) Q123; c) K129; d) E140; e) Q146; and f) K156.

[0092] The variants may in particular include modifications of SEQ ID NO:1 at Q123 and / or Q146. The variants may in particular include modifications of SEQ ID NO:1 at Q123 and Q146.

[0093] The variant may in particular include modifications of SEQ ID NO:1 at K129 and / or E140. The variant may in particular include modifications of SEQ ID NO:1 at K129 and E140.

[0094] The variant may in particular include modifications of SEQ ID NO:1 at E113 and / or K156. The variant may in particular include modifications of SEQ ID NO:1 at E113 and K156.

[0095] The variant may contain one or more modifications within two or three of the constrictions of CytK. Thus, the variant may include: -(i) Q123 and / or Q146; and (ii) K129 and / or E140. -(i) E113 and / or K156; and (ii) Q123 and / or Q146; or - may include modifications of SEQ ID NO: 1 at (i) E113 and / or K156; and (ii) K129 and / or E140.

[0096] More preferably, the variant may include one or more modifications within the middle and lower constrictions. Thus, the variant may include modifications at Q123 and / or Q146, and (ii) K129 and / or E140 of SEQ ID NO: 1, and even more preferably, modifications at all of Q123, Q146, K129, and E140:

[0097] The variant may include one or more of the following modifications in SEQ ID NO:1: a) E113S / N / Y / K / R; b) Q123S / A / N / M / Y / G / K / R; c) K129S / N / Y; d) E140S / N / K / R; e) Q146S / A / N / M / K / R / G / Y; and f) K156S / N.

[0098] The variant may include any of the following modification pairs in SEQ ID NO:1: a)E113S / T / N / Q / G / A / V / L / I / C / R / K / F and K156S / T / N / Q / G / A / V / L / I / C / R / F; b) Q123S / T / N / G / A / V / L / I / C / R / K / F and Q146S / T / N / G / A / V / L / I / C / R / K / F; or c)K129S / T / N / Q / G / A / V / L / I / C / R / F and E140S / T / N / Q / G / A / V / L / I / C / R / K / F.

[0099] The variant may even more preferably comprise any of the following pairs of two or more mutations in SEQ ID NO:1: a) E113S / T / N / Q / G / A / V / L / I / C / R / K / F and K156S / T / N / Q / G / A / V / L / I / C / R / F and Q123S / T / N / G / A / V / L / I / C / R / K / F and Q146S / T / N / G / A / V / L / I / C / R / K / F; b) E113S / T / N / Q / G / A / V / L / I / C / R / K / F and K156S / T / N / Q / G / A / V / L / I / C / R / F and K129S / T / N / Q / G / A / V / L / I / C / R / F and E140S / T / N / Q / G / A / V / L / I / C / R / K / F; c) Q123S / T / N / G / A / V / L / I / C / R / K / F and Q146S / T / N / G / A / V / L / I / C / R / K / F and K129S / T / N / Q / G / A / V / L / I / C / R / F and E140S / T / N / Q / G / A / V / L / I / C / R / K / F; or d) E113S / T / N / Q / G / A / V / L / I / C / R / K / F and K156S / T / N / Q / G / A / V / L / I / C / R / F and Q123S / T / N / G / A / V / L / I / C / R / K / F and Q146S / T / N / G / A / V / L / I / C / R / K / F and K129S / T / N / Q / G / A / V / L / I / C / R / F and E140S / T / N / Q / G / A / V / L / I / C / R / K / F.

[0100] The monomers of the invention may in particular comprise variants of the sequence of SEQ ID NO:1, the variants comprising the following modifications: a) E113S and K156S; b) Q123S and Q146S; c) K129S and E140S; d) Q123S, Q146S, K129S and E140S; or e) E113S, K156S, Q123S, Q146S, K129S and E140S.

[0101] In addition to the specific mutations discussed above, the variant may include other mutations. Over the entire length of the amino acid sequence of SEQ ID NO: 1, the variant is preferably at least 50% homologous to the sequence based on amino acid identity. More preferably, the variant may be at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, and more preferably at least 95%, 97%, or 99% homologous to the amino acid sequence of SEQ ID NO: 1 over the entire sequence based on amino acid identity. There may be at least 80%, for example at least 85%, 90%, or 95% amino acid identity ("hard homology") over a stretch of 100 or more, for example 125, 150, 175, or 200 or more consecutive amino acids.

[0102] Standard methods in the art can be used to determine homology. For example, UWGCG Package provides the BESTFIT program, which can be used to calculate homology, for example, using its default settings (Devereux et al (1984) Nucleic Acids Research 12, p387-395). For example, PILEUP and BLAST algorithms can be used to calculate homology or align sequences (identify equivalent residues or corresponding sequences, typically at their default settings), as described in Altschul SF (1993) J Mol Evol 36:290-300, Altschul, SF et al (1990) J Mol Biol 215:403-10. Software for performing BLAST analysis is publicly available through the National Center for Biotechnology Information (http: / / www.ncbi.nlm.nih.gov / ).

[0103] The mutant monomers of the present invention may be chemically modified. In particular, the monomers may be chemically modified in any manner and at any site. The mutant monomers are preferably chemically modified by attachment of a molecule to one or more cysteines (cysteine ​​attachment), attachment of a molecule to one or more lysines, attachment of a molecule to one or more unnatural amino acids, enzymatic modification of an epitope, or modification of the termini. Suitable methods for performing such modifications are well known in the art. Suitable unnatural amino acids include, but are not limited to, 4-azido-L-phenylalanine (Faz), and any one of the amino acids numbered 1-71 in Figure 1 of Liu CC and Schultz PG, Annu. Rev. Biochem., 2010, 79, 413-444. The mutant monomers may be chemically modified by attachment of any molecule. For example, the mutant monomers may be chemically modified by attachment of a nucleic acid such as polyethylene glycol (PEG), DNA, a dye, a fluorophore, or a chromophore.

[0104] In some embodiments, the mutant monomer is chemically modified with a molecular adaptor that facilitates the interaction between a pore containing the monomer and a target analyte, target nucleotide or target polynucleotide. The presence of the adaptor improves the host-guest chemistry of the pore and the nucleotide or polynucleotide, thereby improving the sequencing ability of the pore formed from the mutant monomer. The principles of host-guest chemistry are well known in the art. The adaptor affects the physical or chemical properties of the pore to improve the interaction with the nucleotide or polynucleotide. The adaptor may change the charge of the barrel or channel of the pore or may specifically interact or bind with the nucleotide or polynucleotide, thereby facilitating its interaction with the pore.

[0105] Molecular adaptors are preferably circular molecules such as cyclodextrins, hybridizable species, DNA binding or interchelating agents, peptides or peptide analogs, synthetic polymers, aromatic planar molecules, small positively charged molecules, or small molecules capable of hydrogen bonding.

[0106] The adaptor may be annular. The annular adaptor preferably has the same symmetry as the pore.

[0107] The adaptor typically interacts with the analyte, nucleotide or polynucleotide through host-guest chemistry. The adaptor typically can interact with the nucleotide or polynucleotide. The adaptor comprises one or more chemical groups capable of interacting with the nucleotide or polynucleotide. The one or more chemical groups preferably interact with the nucleotide or polynucleotide through non-covalent interactions such as hydrophobic interactions, hydrogen bonds, van der Waals forces, π-cation interactions, and / or electrostatic forces. The one or more chemical groups capable of interacting with the nucleotide or polynucleotide are preferably positively charged. The one or more chemical groups capable of interacting with the nucleotide or polynucleotide more preferably comprise an amino group. The amino group may be attached to a primary, secondary, or tertiary carbon atom. The adaptor even more preferably comprises a ring of amino groups, such as a ring of 6, 7, 8, or 9 amino groups. The adaptor most preferably comprises a ring of 6 or 9 amino groups. The protonated amino group ring may interact with a negatively charged phosphate group in the nucleotide or polynucleotide.

[0108] The correct positioning of the adaptor in the pore can be facilitated by host-guest chemistry between the adaptor and the pore containing the mutant monomer. The adaptor preferably comprises one or more chemical groups capable of interacting with one or more amino acids in the pore. The adaptor more preferably comprises one or more chemical groups capable of interacting with one or more amino acids in the pore via non-covalent interactions such as hydrophobic interactions, hydrogen bonds, van der Waals forces, π-cation interactions, and / or electrostatic forces. The chemical groups capable of interacting with one or more amino acids in the pore are typically hydroxyl or amine. The hydroxyl group may be attached to a primary, secondary, or tertiary carbon atom. The hydroxyl group may form hydrogen bonds with an uncharged amino acid of the pore. Any adaptor that facilitates the interaction between the pore and the nucleotide or polynucleotide may be used.

[0109] Suitable adaptors include, but are not limited to, cyclodextrin, cyclic peptides, and cucurbituril. The adaptor is preferably a cyclodextrin or a derivative thereof. The cyclodextrin or a derivative thereof can be any of those disclosed in Eliseev, AV, and Schneider, HJ. (1994) J. Am. Chem. Soc. 116, 6081-6088. The adaptor is more preferably heptakis-6-amino-β-cyclodextrin (am7-βCD), 6-monodeoxy-6-monoamino-β-cyclodextrin (am1-βCD), or heptakis-(6-deoxy-6-guanidino)-cyclodextrin (gu7-βCD). The guanidino group of gu7-βCD has a much higher pKa than the primary amine of am7-βCD, and is therefore more positively charged. This gu7-βCD adaptor can be used to increase the residence time of the nucleotide within the pore, increasing the precision of the measured residual current and increasing the rate of base detection at high temperatures or low data acquisition rates.

[0110] When succinimidyl 3-(2-pyridyldithio)propionate (SPDP) crosslinker is used as described in more detail below, the adapter is preferably heptakis(6-deoxy-6-amino)-6-N-mono(2-pyridyl)dithiopropanoyl-β-cyclodextrin (am6amPDP1-βCD).

[0111] More suitable adaptors include γ-cyclodextrin, which contains eight sugar units (and therefore has eight-fold symmetry), which may contain a linker molecule or may be modified to include all or most of the modified sugar units used in the β-cyclodextrin example above.

[0112] The molecular adaptor is preferably covalently linked to the mutant monomer. The adaptor can be covalently linked to the pore using any method known in the art. The adaptor is typically linked via chemical bonds. If the molecular adaptor is linked via a cysteine ​​bond, one or more cysteines are preferably introduced into the mutant by substitution. The mutant monomer of the present invention can of course contain a cysteine ​​residue at one or both of positions 272 and 283. The mutant monomer can be chemically modified by binding of a molecular adaptor to one or both of these cysteines. Alternatively, the mutant monomer can be chemically modified by binding of a molecule to one or more cysteines or to a non-natural amino acid such as FAz introduced at other positions.

[0113] The reactivity of cysteine ​​residues can be enhanced by modification of adjacent residues. For example, the basic group of an adjacent arginine, histidine, or lysine residue changes the pKa of the cysteine ​​thiol group to that of the more reactive S-group. The reactivity of cysteine ​​residues can be protected by thiol protecting groups such as dTNB. These can be reacted with one or more cysteine ​​residues of the mutant monomer before the linker is attached.

[0114] The molecule may be attached directly to the mutant monomer. The molecule is preferably attached to the mutant monomer using a linker, such as a chemical crosslinker or a peptide linker.

[0115] Suitable chemical crosslinkers are well known in the art. Preferred crosslinkers include 2,5-dioxopyrrolidin-1-yl 3-(pyridin-2-yldisulfanyl)propanoate, 2,5-dioxopyrrolidin-1-yl 4-(pyridin-2-yldisulfanyl)butanoate, and 2,5-dioxopyrrolidin-1-yl 8-(pyridin-2-yldisulfanyl)octananoate. The most preferred crosslinker is succinimidyl 3-(2-pyridyldithio)propionate (SPDP). Typically, the molecule is covalently attached to the bifunctional crosslinker before the molecule / crosslinker complex is covalently attached to the mutant monomer, although it is also possible to covalently attach the bifunctional crosslinker to the monomer before the bifunctional crosslinker / monomer complex is attached to the molecule.

[0116] The linker is preferably resistant to dithiothreitol (DTT). Suitable linkers include, but are not limited to, iodoacetamide-based and maleimide-based linkers.

[0117] In other embodiments, the monomers may be attached to polynucleotide binding proteins, forming a modular sequencing system that can be used in the methods of the invention. Polynucleotide binding proteins are discussed below.

[0118] The polynucleotide binding protein may be covalently linked to the mutant monomer. The protein may be covalently linked to the pore using any method known in the art. The monomer and protein may be chemically fused or genetically fused. The monomer and protein are genetically fused when the entire construct is expressed from a single polynucleotide sequence. Genetic fusion of the pore to a polynucleotide binding protein is described in International Application No. PCT / GB09 / 001679 (published as WO2010 / 004265).

[0119] The polynucleotide binding protein may be attached to the mutant monomer directly or via one or more linkers. The polynucleotide binding protein may be attached to the mutant monomer using a hybridization linker as described in International Application No. PCT / GB10 / 000132 (published as WO2010 / 086602). Alternatively, a peptide linker may be used. A peptide linker is an amino acid sequence. The length, flexibility, and hydrophilicity of the peptide linker are typically designed so that it does not interfere with the function of the monomer and molecule. A preferred flexible peptide linker is a series of 2-20 serine and / or glycine amino acids, such as 4, 6, 8, 10, or 16. More preferred flexible linkers include (SG)1, (SG)2, (SG)3, (SG)4, (SG)5, and (SG)8, where S is serine and G is glycine. Preferred rigid linkers are a stretch of 2 to 30 proline amino acids, such as 4, 6, 8, 16, or 24. More preferred rigid linkers include (P)12, where P is proline.

[0120] The mutant monomers can be chemically modified with molecular adaptors and polynucleotide binding proteins.

[0121] Polynucleotides The present invention also provides polynucleotide sequences encoding mutant monomers of the present invention. The mutant monomers can be any of those discussed above. The polynucleotide sequences preferably comprise a sequence that is at least 50%, 60%, 70%, 80%, 90%, or 95% homologous based on nucleotide identity to the sequence of SEQ ID NO:2 over the entire sequence. There may be at least 80%, such as at least 85%, 90%, or 95% nucleotide identity ("hard homology") over a stretch of 300 or more, such as 375, 450, 525, or 600 or more consecutive nucleotides. Homology may be calculated as described above. The polynucleotide sequence may comprise a sequence that differs from SEQ ID NO:2 based on the degeneracy of the genetic code.

[0122] The present invention also provides polynucleotide sequences encoding any of the genetically fused constructs of the present invention. The polynucleotide preferably comprises two or more variants of the sequence shown in SEQ ID NO:2. The polynucleotide sequence preferably comprises two or more sequences having at least 50%, 60%, 70%, 80%, 90%, or 95% homology with SEQ ID NO:2 based on nucleotide identity over the entire sequence. There may be at least 80%, such as at least 85%, 90%, or 95% nucleotide identity ("hard homology") over a stretch of 600 or more, such as 750, 900, 1050, or 1200 or more consecutive nucleotides. Homology may be calculated as described above.

[0123] Polynucleotide sequences can be obtained and replicated using standard methods in the art. Chromosomal DNA encoding wild-type CytK can be extracted from pore-producing organisms such as Bacillus cereus. Genes encoding pore subunits can be amplified using PCR with specific primers. The amplified sequences can then be subjected to site-directed mutagenesis. Suitable methods for site-directed mutagenesis are known in the art, including, for example, combinatorial chain reaction. Polynucleotides encoding constructs of the present invention can be made using well-known techniques, such as those described in Sambrook, J. and Russell, D. (2001). Molecular Cloning: A Laboratory Manual, 3rd Edition. Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY.

[0124] The resulting polynucleotide sequence can then be incorporated into a recombinant replicable vector, such as a cloning vector. The vector can be used to replicate the polynucleotide in a compatible host cell. Thus, a polynucleotide sequence can be produced by introducing the polynucleotide into a replicable vector, introducing the vector into a compatible host cell, and growing the host cell under conditions that allow replication of the vector. The vector can be recovered from the host cell. Host cells suitable for cloning polynucleotides are known in the art and are described in more detail below.

[0125] The polynucleotide sequence can be cloned into a suitable expression vector. In the expression vector, the polynucleotide sequence is typically operably linked to a control sequence that can provide the expression of the coding sequence by the host cell. Such an expression vector can be used to express the pore subunit.

[0126] The term "operably linked" refers to a juxtaposition where the components described are in a relationship permitting them to function in their intended manner. A control sequence "operably linked" to a coding sequence is ligated in such a way that expression of the coding sequence is achieved under conditions compatible with the control sequences. Multiple copies of the same or different polynucleotide sequences can be introduced into the vector.

[0127] The expression vector can then be introduced into a suitable host cell. Thus, the mutant monomers or constructs of the present invention can be produced by inserting a polynucleotide sequence into an expression vector, introducing the vector into a compatible bacterial host cell, and growing the host cell under conditions that result in expression of the polynucleotide sequence. The recombinantly expressed monomers or constructs can self-assemble into pores in the host cell membrane. Alternatively, the recombinant pores produced in this manner can be removed from the host cell and inserted into another membrane. When making pores that contain at least two different monomers or constructs, the different monomers or constructs can be expressed separately in different host cells as described above, removed from the host cell, and assembled into pores in separate membranes, such as rabbit cell membranes or synthetic membranes.

[0128] The vector may be, for example, a plasmid, virus or phage vector provided with an origin of replication, optionally a promoter for the expression of the polynucleotide sequence, and optionally a regulator of the promoter. The vector may contain one or more selectable marker genes, for example, the tetracycline resistance gene. The promoter and other expression regulation signals may be selected to be compatible with the host cell for which the expression vector is designed. Typically, T7, trc, lac, ara or lambda L promoters are used.

[0129] The host cell typically expresses the monomer or construct at high levels. The host cell transformed with the polynucleotide sequence is selected to be compatible with the expression vector used to transform the cell. The host cell is typically a bacterium, preferably Escherichia coli. Any cell with a λDE3 lysogen, such as C41(DE3), BL21(DE3), JM109(DE3), B834(DE3), TUNER, Origami and Origami B, can express vectors containing the T7 promoter.

[0130] The invention also includes a method of making a mutant monomer of the invention or a construct of the invention, the method comprising expressing a polynucleotide of the invention in a suitable host cell. The polynucleotide is preferably part of a vector, and is preferably operably linked to a promoter.

[0131] Generation of mutant CytK The invention also provides a method for improving the ability of a CytK monomer comprising the sequence set forth in SEQ ID NO: 1 to characterize a target analyte, the method comprising making one or more modifications between about S100 and about K170 of SEQ ID NO: 1 that alter the ability of the monomer to interact with a polynucleotide, but do not affect the ability of the monomer to form a pore. Any of the embodiments discussed above in connection with mutant CytK monomers in connection with characterizing polynucleotides apply equally to this method of the invention.

[0132] pore The present invention also provides various pores. The pores of the present invention are ideal for characterizing analytes. Such pores can be used in the methods provided herein. The pores of the present invention are particularly ideal for characterizing polynucleotides, such as sequencing, because they can distinguish different nucleotides with high sensitivity. The pores can be used to characterize nucleic acids, such as DNA and RNA, including sequencing nucleic acids and identifying single base changes. The pores of the present invention can even distinguish between methylated and unmethylated nucleotides. The base resolution of the pores of the present invention is surprisingly high. The pores show nearly complete separation of all four DNA nucleotides. The pores can also be used to distinguish between deoxycytidine monophosphate (dCMP) and methyl-dCMP based on the residence time in the pore and the current flowing through the pore.

[0133] The pores of the present invention can also distinguish different nucleotides under various conditions. In particular, the pores distinguish nucleotides under conditions favorable for characterization, such as sequencing, of polynucleotides. The degree to which the pores of the present invention can distinguish different nucleotides can be controlled by varying the applied potential, salt concentration, buffer, temperature, and the presence of additives such as urea, betaine, and DTT. This allows the function of the pores to be fine-tuned, especially when sequencing. This will be discussed in more detail below. The pores of the present invention can also be used to identify polynucleotide polymers from interactions with one or more monomers, rather than on a nucleotide-by-nucleotide basis.

[0134] The pore of the present invention can be isolated, substantially isolated, purified, or substantially purified. The pore of the present invention is isolated or purified when it does not contain any other components such as lipids or other pores. The pore is substantially isolated when it is mixed with a carrier or diluent that does not interfere with its intended use. For example, the pore is substantially isolated or substantially purified when it exists in a form that contains less than 10%, less than 5%, less than 2%, or less than 1% of other components, such as lipids or other pores. Alternatively, the pore of the present invention can exist in a lipid bilayer.

[0135] The pores of the present invention may be present as individual or single pores, or alternatively, the pores of the present invention may be present in a homogeneous or heterogeneous population or plurality of two or more pores.

[0136] Homo-oligomer pores The present invention also provides homo-oligomeric pores derived from CytK that contain the same mutant monomers of the present invention. The monomers are identical in terms of their amino acid sequence. The homo-oligomeric pores of the present invention are ideal for characterizing polynucleotides, such as sequencing. Such pores can be used in the methods provided herein. The homo-oligomeric pores of the present invention can have any of the advantages discussed above. The advantages of certain homo-oligomeric pores of the present invention are shown in the examples.

[0137] The homo-oligomeric pore may contain any number of mutant monomers. The pore typically contains two or more mutant monomers, but typically contains at least 7, at least 8, at least 9, or at least 10 identical mutant monomers, for example, 7, 8, 9 or 10 mutant monomers. Most preferably, the homo-oligomeric pore is a heptameric pore.

[0138] One or more of the mutant monomers are preferably chemically modified as discussed above. In other words, the fact that one or more of the monomers are chemically modified (and the other monomers that are not chemically modified) does not prevent the pore from being a homo-oligomer, so long as the amino acid sequence of each of the monomers is identical.

[0139] Hetero-oligomeric pores The present invention also provides a hetero-oligomeric pore derived from CytK, comprising at least one mutant monomer of the present invention, and at least one of the monomers is different from other monomers. The monomer is different from other monomers in terms of its amino acid sequence. The hetero-oligomeric pore of the present invention is ideal for characterizing polynucleotides, such as sequencing. Such pores can be used in the methods provided herein. The hetero-oligomeric pore can be produced using methods known in the art (e.g., Protein Sci.2002 Jul;11(7):1813-24).

[0140] The hetero-oligomeric pore contains sufficient monomers to form a pore. The pore typically contains two or more mutant monomers, but typically contains at least 7, at least 8, at least 9, or at least 10 identical mutant monomers, for example, 7, 8, 9, or 10 mutant monomers. Most preferably, the hetero-oligomeric pore is a heptameric pore.

[0141] In a preferred embodiment, all of the monomers (e.g. 10, 9, 8, or 7 of the monomers) are mutant monomers of the invention, at least one of which is different from the other monomers. In a more preferred embodiment, the pore comprises 8 or 9 mutant monomers of the invention, at least one of which is different from the other monomers. They may all be different from each other.

[0142] The mutant monomers of the invention within the pore are preferably of approximately the same length or of the same length. The barrels of the mutant monomers of the invention within the pore are preferably of approximately the same length or of the same length. Length may be measured in number of amino acids and / or in length units.

[0143] In another preferred embodiment, at least one of the mutant monomers is not a mutant monomer of the present invention. In this embodiment, the remaining monomers are preferably mutant monomers of the present invention. Thus, the pore may contain 9, 8, 7, 6, 5, 4, 3, 2, or 1 mutant monomer of the present invention. Any number of monomers in the pore may not be mutant monomers of the present invention. The pore preferably contains 7 or 8 mutant monomers of the present invention and a monomer that is not a monomer of the present invention. The mutant monomers of the present invention may be the same or different.

[0144] The mutant monomers of the invention within a construct are preferably about the same length or are the same length. The barrels of the mutant monomers of the invention within a construct are preferably about the same length or are the same length. Length may be measured in number of amino acids and / or in length units.

[0145] The pore may comprise one or more monomers which are not mutant monomers of the invention.

[0146] Methods for creating pores are discussed in more detail below.

[0147] Constructs The present invention also provides a construct comprising two or more covalently linked monomers derived from CytK, at least one of the monomers being a mutant monomer of the present invention. The construct of the present invention retains the ability to form a pore. This can be determined as discussed above. One or more constructs of the present invention can be used to form a pore for characterizing a polypeptide or polynucleotide, such as sequencing. Such a pore can be used in the methods provided herein. The construct can comprise at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10 monomers. The construct preferably comprises two monomers. The two or more monomers can be the same or different.

[0148] At least one monomer in the construct is a mutant monomer of the invention. Two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, or ten or more monomers in the construct can be mutant monomers of the invention. All of the monomers in the construct are preferably mutant monomers of the invention. The mutant monomers may be the same or different. In a preferred embodiment, the construct comprises two mutant monomers of the invention.

[0149] The monomers within a construct are preferably genetically fused. The monomers are genetically fused if the entire construct is expressed from a single polynucleotide sequence. The coding sequences of the monomers can be combined in any way to form a single polynucleotide sequence that encodes the construct.

[0150] Monomers may be genetically fused in any configuration. Monomers may be fused via their terminal amino acids. For example, the amino terminus of one monomer may be fused to the carboxy terminus of another monomer. The second and subsequent monomers (amino to carboxy direction) in the construct may contain methionines at their amino termini (each of which is fused to the carboxy terminus of the previous monomer). For example, if M is a monomer (not containing an amino terminal methionine) and mM is a monomer with an amino terminal methionine, the construct may contain the sequence M-mM, M-mM-mM, or M-mM-mM-mM. The presence of these methionines typically results from the expression of a start codon (i.e., ATG) at the 5' end of the polynucleotide encoding the second or subsequent monomer within the polynucleotide encoding the entire construct. The first monomer (amino to carboxy direction) in the construct may also contain a methionine (e.g., mM-mM, mM-mM-mM, or mM-mM-mM-mM).

[0151] Two or more monomers can be fused together directly. The monomers are preferably fused using a linker. The linker can be designed to limit the mobility of the monomers. A preferred linker is an amino acid sequence (i.e., a peptide linker). Any of the peptide linkers discussed above can be used.

[0152] In another preferred embodiment, the monomer is chemically fused. The two monomers are chemically fused, for example, through a chemical crosslinker, when the two moieties are chemically bonded. Any of the chemical crosslinkers discussed above can be used. The linker can be attached to one or more cysteine ​​residues introduced into the mutant monomer of the present invention. Alternatively, the linker can be attached to the end of one of the monomers in the construct.

[0153] When the construct contains different monomers, crosslinking of the monomer to itself can be prevented by keeping the concentration of the linker at a large amount of monomer. Alternatively, a "lock and key" configuration can be used in which two linkers are used. Only one end of each linker can react together to form a longer linker, and the other end of the linker reacts with a different monomer each. Such linkers are described in International Application No. PCT / GB10 / 000132 (published as WO2010 / 086602).

[0154] Pore ​​containing construct The present invention also provides a pore comprising at least one construct of the present invention. Such a pore can be used in the method provided herein. The construct of the present invention comprises two or more covalently linked monomers derived from CytK, and at least one of the monomers is a mutant CytK monomer of the present invention. In other words, the construct must contain two or more monomers. At least two of the monomers in the pore are in the form of the construct of the present invention. The monomers can be of any type.

[0155] The pore typically comprises (a) one construct containing two monomers, and (b) a sufficient number of monomers to form the pore. The construct can be any of those discussed above. The monomers can be any of those discussed above, including the mutant monomers of the invention.

[0156] Another exemplary pore comprises two or more constructs of the present invention, such as two, three, or four constructs of the present invention. Such pores further comprise a sufficient number of monomers to form the pore. The monomers can be any of those discussed above. Further pores of the present invention comprise only constructs that comprise two monomers. A particular pore according to the present invention comprises several constructs, each of which comprises two monomers. The constructs may oligomerize into a pore having a structure such that only one monomer from each construct contributes to the pore. Typically, the other monomers of the construct (i.e., the monomers that do not form the pore) are outside the pore.

[0157] As described above, mutations can be introduced into the construct. The mutations can be alternating, i.e., the mutations are different for each monomer in the two monomer construct, and the construct is assembled as a homo-oligomer resulting in alternating modifications. In other words, monomers containing MutA and MutB are fused and assembled to form an AB:AB:AB:AB pore. Alternatively, the mutations can be adjacent, i.e., the same mutation is introduced into two monomers in the construct, which are then oligomerized with different mutant monomers. In other words, a monomer containing MutA is fused and subsequently oligomerized with a MutB-containing monomer to form AA:B:B:B:B:B:B.

[0158] One or more of the monomers of the invention within the pore containing the construct may be chemically modified as discussed above.

[0159] Pore ​​Creation of the Invention The present invention also provides a method of producing a pore of the present invention. The method comprises allowing at least one mutant monomer of the present invention or at least one construct of the present invention to oligomerize with a sufficient number of mutant CytK monomers of the present invention, constructs of the present invention, or monomers derived from CytK to form a pore. When the method relates to producing a homo-oligomeric pore of the present invention, all of the monomers used in the method are mutant CytK monomers of the present invention having the same amino acid sequence. When the method relates to producing a hetero-oligomeric pore of the present invention, at least one of the monomers is different from the other monomers. Any of the embodiments discussed above in relation to the pores of the present invention apply equally to the method of producing a pore.

[0160] A preferred method for making the pores of the present invention is disclosed in Example 1.

[0161] film The pores of the invention may be present in a membrane. Thus, the invention provides a membrane comprising a pore of the invention.

[0162] In the method of the present invention, the polynucleotide typically contacts the pore of the present invention in a membrane. Any membrane can be used according to the present invention. Suitable membranes are well known in the art. The membrane is preferably an amphiphilic layer. An amphiphilic layer is a layer formed from amphiphilic molecules, such as phospholipids, that have both hydrophilic and lipophilic properties. The amphiphilic molecules can be synthetic or naturally occurring. Non-naturally occurring amphiphiles and amphiphiles that form monolayers are known in the art, including, for example, block copolymers (Gonzalez-Perez et al., Langmuir, 2009, 25, 10447-10450). Block copolymers are polymeric materials in which two or more monomer subunits are polymerized together to make a single polymer chain. Block copolymers typically have properties contributed by each monomer subunit. However, block copolymers can have unique properties that polymers formed from individual subunits do not have. Block copolymers can be engineered such that in aqueous media, one of the monomer subunits is hydrophobic (i.e., lipophilic) and the other subunit(s) is hydrophilic. In this case, the block copolymer can have amphiphilic properties and form structures that mimic biological membranes. Block copolymers can be diblock (composed of two monomer subunits), but can also be built from more than two monomer subunits to form more complex arrangements that behave as amphiphiles. The copolymers can be triblock, tetrablock, or pentablock copolymers. The membrane is preferably a triblock copolymer membrane.

[0163] Archaeal bipolar tetraether lipids are naturally occurring lipids that are constructed such that the lipids form monolayer membranes. These lipids are commonly found in extremophiles, thermophiles, halophiles, and acidophiles that live in harsh biological environments. Their stability is believed to derive from the fused nature of the final bilayer. It is straightforward to construct block copolymer materials that mimic these biological entities by creating triblock polymers with a general motif of hydrophilic-hydrophobic-hydrophilic. This material can form monomeric membranes that behave similarly to lipid bilayers and encompass a range of phase behaviors from vesicles to lamellar membranes. Membranes formed from these triblock copolymers hold several advantages over biological lipid membranes. As the triblock copolymers are synthesized, the exact structure can be carefully controlled to provide the correct chain length and properties required to form membranes and interact with pores and other proteins.

[0164] Block copolymers may also be constructed from subunits that are not classified as lipid submaterials, for example, hydrophobic polymers may be made from siloxanes or other non-hydrocarbon monomers. The hydrophilic subcompartments of block copolymers may also have low protein binding properties, allowing for the creation of membranes that are highly resistant when exposed to untreated biological samples. The head group units may be derived from non-classical lipid head groups.

[0165] Triblock copolymer membranes also have increased mechanical and environmental stability, e.g., much higher operating temperature or pH ranges, compared to biological lipid membranes. The synthetic nature of block copolymers provides a platform for customizing polymer-based membranes for a wide range of applications.

[0166] The membrane is most preferably one of the membranes disclosed in International Application No. PCT / GB2013 / 052766 or PCT / GB2013 / 052767.

[0167] The amphipathic molecules can be chemically modified or functionalized to facilitate attachment of the polynucleotide.

[0168] The amphiphilic layer may be a monolayer or a bilayer. The amphiphilic layer is typically planar. The amphiphilic layer may be curved. The amphiphilic layer may be supported.

[0169] Amphiphilic membranes are typically mobile in nature and essentially act as two-dimensional fluids with lipid diffusion rates of approximately 10-8 cms-1. This means that pores and bound polynucleotides can typically move within the amphiphilic membrane.

[0170] The membrane may be a lipid bilayer. Lipid bilayers are models of cell membranes and serve as an excellent platform for a range of experimental studies. For example, lipid bilayers can be used for in vitro investigation of membrane proteins by single channel recording. Alternatively, lipid bilayers can be used as biosensors to detect the presence of a range of substances. The lipid bilayer may be any lipid bilayer. Suitable lipid bilayers include, but are not limited to, planar lipid bilayers, supported bilayers or liposomes. The lipid bilayer is preferably a planar lipid bilayer. Suitable lipid bilayers are disclosed in International Application No. PCT / GB08 / 000563 (published as WO2008 / 102121), International Application No. PCT / GB08 / 004127 (published as WO2009 / 077734), and International Application No. PCT / GB2006 / 001057 (published as WO2006 / 100484).

[0171] Methods for forming lipid bilayers are known in the art. Lipid bilayers are commonly formed by the method of Montal and Mueller (Proc. Natl. Acad. Sci. USA., 1972; 69: 3561-3566), in which a lipid monolayer is supported on the aqueous solution / air interface past both sides of a hole perpendicular to the interface. Lipids are usually first dissolved in an organic solvent and then added to the surface of an aqueous electrolyte solution by evaporating a drop of the solvent on the surface of the aqueous solution on either side of the opening. As the organic solvent evaporates, the solution / air interface on either side of the opening physically moves up and down through the opening until a bilayer is formed. Planar lipid bilayers can be formed across the opening of the membrane or across the opening into a recess.

[0172] The Montal & Mueller method is popular because it is a cost-effective and relatively simple method for forming good quality lipid bilayers suitable for protein pore insertion. Other common methods of bilayer formation include tip dipping, painting of bilayers, and patch clamping of liposomal bilayers.

[0173] Tip-dipping bilayer formation involves contacting an aperture surface (e.g., a pipette tip) onto the surface of a test solution carrying a lipid monolayer. Again, a lipid monolayer is first generated at the solution / air interface by evaporating a droplet of lipid dissolved in an organic solvent at the solution surface. The bilayer is then formed by the Langmuir-Schaefer process, which requires mechanical automation to move the aperture relative to the solution surface.

[0174] In coated bilayers, a drop of lipid dissolved in an organic solvent is applied directly to the aperture, which is then immersed in an aqueous test solution. A paintbrush or equivalent is used to spread the lipid solution thinly over the aperture. Thinning the solvent results in the formation of a lipid bilayer. However, it is difficult to completely remove the solvent from the bilayer, and as a result, bilayers formed by this method are less stable and more prone to noise during electrochemical measurements.

[0175] Patch clamping is commonly used in the study of biological cell membranes. The cell membrane is fixed to the end of a pipette by suction and a patch of membrane is attached over the opening. The method has been adapted to generate lipid bilayers by fixing liposomes, which are then ruptured leaving a lipid bilayer that covers and seals the opening of the pipette. The method requires the production of stable giant unilamellar liposomes and a small opening in a material with a glass surface.

[0176] Liposomes can be formed by sonication, extrusion, or the Mozafari method (Colas et al. (2007) Micron 38:841-847).

[0177] In a preferred embodiment, the lipid bilayer is formed as described in International Application No. PCT / GB08 / 004127 (published as WO2009 / 077734). Advantageously, in this method, the lipid bilayer is formed from dry lipids. In a most preferred embodiment, the lipid bilayer is formed across the opening as described in WO2009 / 077734 (PCT / GB08 / 004127).

[0178] A lipid bilayer is formed from two opposing layers of lipids. The two layers of lipids are arranged so that their hydrophobic tail groups face each other to form a hydrophobic interior. The hydrophilic head groups of the lipids face outward toward the aqueous environment on each side of the bilayer. Bilayers can exist in a number of lipid phases, including but not limited to liquid disordered phases (fluid lamella), liquid ordered phases, solid ordered phases (lamellar gel phase, interdigitated gel phase), and planar bilayer crystals (lamellar subgel phase, lamellar crystal phase).

[0179] Any lipid composition that forms a lipid bilayer may be used. The lipid composition is selected so that a lipid bilayer is formed with the required properties, such as surface charge, ability to support membrane proteins, packing density, or mechanical properties. The lipid composition may contain one or more different lipids. For example, the lipid composition may contain up to 100 lipids. The lipid composition preferably contains 1 to 10 lipids. The lipid composition may contain naturally occurring lipids and / or artificial lipids.

[0180] A lipid typically comprises a head group, an interface moiety, and two hydrophobic tail groups, which may be the same or different. Suitable head groups include, but are not limited to, neutral head groups, such as diacylglyceride (DG) and ceramide (CM), zwitterionic head groups, such as phosphatidylcholine (PC), phosphatidylethanolamine (PE), and sphingomyelin (SM), negatively charged head groups, such as phosphatidylglycerol (PG), phosphatidylserine (PS), phosphatidylinositol (PI), phosphatidic acid (PA), and cardiolipin (CA), and positively charged head groups, such as trimethylammonium-propane (TAP). Suitable interface moieties include, but are not limited to, naturally occurring interface moieties, such as glycerol-based moieties or ceramide-based moieties. Suitable hydrophobic tail groups include, but are not limited to, saturated hydrocarbon chains such as lauric acid (n-Dodecanolic acid), myristic acid (n-Tetradecononic acid), palmitic acid (n-Hexadecanoic acid), stearic acid (n-Octadecanoic acid), and arachidic acid (n-Eicosanoic acid), unsaturated hydrocarbon chains such as oleic acid (cis-9-Octadecanoic acid), and branched hydrocarbon chains such as phytanoyl. The chain length and the position and number of double bonds of the unsaturated hydrocarbon chains can vary. The chain length and the position and number of branches such as methyl groups of the branched hydrocarbon chains can vary. The hydrophobic tail group can be linked to the interface moiety as an ether or ester. The lipid can be a mycolic acid.

[0181] Lipids can also be chemically modified. Head group or tail group of lipids can be chemically modified. Suitable lipids with chemically modified head groups include, but are not limited to, PEG-modified lipids, such as 1,2-diacyl-sn-glycero-3-phosphoethanolamine-N-[methoxy(polyethylene glycol)-2000], functionalized PEG lipids, such as 1,2-distearoyl-sn-glycero-3-phosphoethanolamine-N-[biotinyl(polyethylene glycol)2000], and lipids modified for conjugation, such as 1,2-dioleoyl-sn-glycero-3-phosphoethanolamine-N-(succinyl) and 1,2-dipalmitoyl-sn-glycero-3-phosphoethanolamine-N-(biotinyl). Suitable lipids with chemically modified tail groups include, but are not limited to, polymerizable lipids such as 1,2-bis(10,12-tricosadiynoyl)-sn-glycero-3-phosphocholine, fluorinated lipids such as 1-palmitoyl-2-(16-fluoropalmitoyl)-sn-glycero-3-phosphocholine, deuterated lipids such as 1,2-dipalmitoyl-D62-sn-glycero-3-phosphocholine, and ether-linked lipids such as 1,2-di-O-phytanyl-sn-glycero-3-phosphocholine. Lipids can be chemically modified or functionalized to facilitate attachment of polynucleotides.

[0182] Amphiphilic layer, for example, lipid composition, typically contains one or more additives that will affect the properties of the layer.Suitable additives include, but are not limited to, fatty acid, for example, palmitic acid, myristic acid, and oleic acid, fatty alcohol, for example, palmitic alcohol, myristic alcohol, and oleic alcohol, sterol, for example, cholesterol, ergosterol, lanosterol, sitosterol, and stigmasterol, lysophospholipid, for example, 1-acyl-2-hydroxy-sn-glycero-3-phosphocholine, and ceramide.

[0183] In another preferred embodiment, the membrane comprises a solid layer. The solid layer can be formed from both organic and inorganic materials, including but not limited to microelectronic materials, insulating materials such as Si3N4, A12O3 and SiO, organic and inorganic polymers such as polyamides, plastics such as Teflon or elastomers such as two-component addition-cured silicone rubber, and glass. The solid-state layer can be formed from graphene. A suitable graphene layer is disclosed in International Application No. PCT / US2008 / 010637 (published as WO2009 / 035647). When the membrane comprises a solid-state layer, the pores are typically present in an amphiphilic membrane or layer contained within the solid-state layer, for example, in holes, wells, gaps, channels, trenches, or slits within the solid-state layer. Those skilled in the art can prepare suitable solid-state / amphiphilic hybrid systems. Suitable systems are disclosed in WO2009 / 020682 and WO2012 / 005857. Any of the amphiphilic membranes or layers described above may be used.

[0184] The methods described herein are typically carried out using (i) an artificial amphiphilic layer containing a pore, (ii) an isolated naturally occurring lipid bilayer containing a pore, or (iii) a cell with a pore inserted therein. The methods are typically carried out using an artificial amphiphilic layer, such as an artificial triblock copolymer layer. In addition to the pore, the layer may contain other transmembrane and / or intramembrane proteins, as well as other molecules. Suitable equipment and conditions are discussed below. The methods of the present invention are typically carried out in vitro.

[0185] array The invention also provides an array comprising a plurality of the membranes of the invention, hi a preferred embodiment, each membrane in the array contains one pore of the invention.

[0186] Preferably, the array is configured to carry out the method of characterizing an analyte described herein.For example, the array may form part of a device that includes a chamber further containing an aqueous solution and a barrier that separates the chamber into two compartments.The barrier typically has an opening in which a membrane containing pores is formed.Alternatively, the barrier forms a membrane in which pores are present.

[0187] device The invention also provides a device comprising an array of the invention, means for applying an electrical potential across the membrane, and means for detecting an electrical or optical signal across the membrane. The device of the invention is preferably configured to carry out the methods of characterizing an analyte described herein.

[0188] Preferably, the device comprises an electrical circuit capable of applying an electrical potential and measuring the electrical signal across the membrane and pore.

[0189] The device is preferably capable of supporting a plurality of pores and membranes and is operable to perform analyte characterization using the pores and membranes according to the methods of characterizing an analyte described herein. The device may in particular comprise at least one reservoir for holding material for performing the characterization, a fluidics system configured to controllably deliver material from the at least one reservoir to the sensor device, and one or more containers for receiving respective samples, the fluidics system configured to selectively deliver samples from the one or more containers to the device.

[0190] Methods for characterizing an analyte The present invention provides a method for determining the presence, absence, or one or more characteristics of a target analyte. In particular, the method is for characterizing a target analyte. The method for characterizing a target analyte comprises: (a) contacting a target analyte with a pore according to the invention such that the target analyte migrates relative to the pore; (b) taking one or more measurements characteristic of the analyte as it migrates relative to the pore; The target analyte is thereby characterized.

[0191] Steps (a) and (b) of the method are preferably carried out with an applied potential across the pore. As discussed in more detail below, the applied potential typically results in the formation of a complex between the pore and the polynucleotide binding protein. The applied potential may be a voltage potential. Alternatively, the applied potential may be a chemical potential. An example of this uses a salt gradient across the amphiphilic layer. Salt gradients are disclosed in Holden et al., J Am Chem Soc. 2007 Jul 11;129(27):8650-5.

[0192] The method is for determining the presence, absence, or one or more characteristics of a target analyte. The method may be for determining the presence, absence, or one or more characteristics of at least one analyte. The method may relate to determining the presence, absence, or one or more characteristics of two or more analytes. The method may include determining the presence, absence, or one or more characteristics of any number of analytes, for example 2, 5, 10, 15, 20, 30, 40, 50, 100 or more analytes. Any number of characteristics of one or more analytes may be determined, for example 1, 2, 3, 4, 5, 10 or more characteristics.

[0193] The target analyte is preferably a metal ion, an inorganic salt, a polymer, an amino acid, a peptide, a polypeptide, a protein, a nucleotide, an oligonucleotide, a polynucleotide, an oligosaccharide. The method may relate to determining the presence, absence, or one or more characteristics of two or more analytes of the same type, such as two or more proteins, two or more nucleotides, or two or more pharmaceuticals. Alternatively, the method may relate to determining the presence, absence, or one or more characteristics of two or more different types of analytes, such as one or more proteins, one or more nucleotides, and one or more pharmaceuticals.

[0194] The target analyte may be secreted from the cell, or it may be an analyte that is present inside the cell, and so the analyte must be extracted from the cell before the present invention can be performed.

[0195] The analyte is preferably an amino acid, a peptide, a polypeptide, and / or a protein. The amino acid, peptide, polypeptide, or protein may be naturally occurring or non-naturally occurring. The polypeptide or protein may include synthetic or modified amino acids therein. Several different types of modifications to amino acids are known in the art. Suitable amino acids and their modifications are described above. For the purposes of the present invention, it is understood that the target analyte may be modified by any method available in the art.

[0196] The protein may be an enzyme, an antibody, a hormone, a growth factor, or a growth regulatory protein such as a cytokine. The cytokine may be selected from interleukins, preferably IFN-1, IL-1, IL-2, IL-4, IL-5, IL-6, IL-10, IL-12, and IL-13, interferons, preferably IL-g, and other cytokines such as TNF-a. The protein may be a bacterial protein, a fungal protein, a viral protein, or a protein from a parasite.

[0197] The target analyte is preferably a nucleotide, an oligonucleotide, or a polynucleotide. Nucleotides and polynucleotides are discussed below. An oligonucleotide is typically a short nucleotide polymer having 50 or less nucleotides, such as 40 or less, 30 or less, 20 or less, 10 or less, or 5 or less nucleotides. An oligonucleotide may include any of the nucleotides discussed below, including abasic and modified nucleotides.

[0198] The target analyte, such as a target polynucleotide, may be present in any of the suitable samples discussed below.

[0199] The pores are typically present in a membrane, as discussed below. The target analytes can be bound or delivered to the membrane using the methods discussed below.

[0200] Any of the measurements discussed below can be used to determine the presence, absence, or one or more properties of the target analyte. The method preferably includes contacting the target analyte with the pore such that the analyte moves relative to, e.g., moves through, the pore, and measuring a current passing through the pore as the analyte moves relative to the pore, thereby determining the presence, absence, or one or more properties of the analyte.

[0201] The target analyte is present if current flows through the pore in an analyte-specific manner (i.e., a unique current associated with the analyte is detected flowing through the pore). If current does not flow through the pore in a nucleotide-specific manner, the analyte is not present. Control experiments can be performed in the presence of the analyte to determine whether the method affects the current flowing through the pore.

[0202] The present invention can be used to distinguish analytes of similar structure based on the different effects they have on the current passing through the pore. Individual analytes can be identified at the single molecule level from the current amplitude when interacting with the pore. The present invention can also be used to determine whether a particular analyte is present in a sample. The present invention can also be used to measure the concentration of a particular analyte in a sample. Characterization of analytes using pores other than CytK is known in the art.

[0203] Polynucleotide characterization The method of the present invention can be used to characterize target polynucleotide.Therefore, the present invention can provide a method of characterizing target polynucleotide, for example, a method of sequencing polynucleotide.There are two main strategies for characterizing or sequencing polynucleotide using nanopores, namely, strand characterization / sequencing and exonuclease characterization / sequencing.The method of the present invention can be either method.

[0204] In strand sequencing, DNA is translocated through the pore with or against an applied potential. An exonuclease that acts processively or processively on double-stranded DNA can be used on the cis side of the pore to thread the remaining single strand under the applied potential or on the trans side under a reverse potential. Similarly, a helicase that unwinds double-stranded DNA can be used in a similar manner. Polymerases can also be used. There are potential sequencing applications that require strand translocation against an applied potential, but the DNA must first be "captured" by the enzyme under a reverse or no potential. Then, when the potential is returned following binding, the strand will pass through the pore from cis to trans and be held in an extended conformation by the current. Single-stranded DNA exonuclease or single-stranded DNA-dependent polymerase can act as a molecular motor to pull the newly translocated single strand back through the pore from trans to cis in a controlled manner against the applied potential.

[0205] In one embodiment, the method of characterizing a target polynucleotide involves contacting a target sequence with a pore and a helicase enzyme of the invention. Any helicase can be used in the method. Suitable helicases are discussed below. The helicase can work in two modes with respect to the pore. First, the method is preferably carried out using a helicase to control the movement of the target sequence through the pore with an electric field resulting from an applied potential. In this mode, the 5' end of the DNA is first captured in the pore, and the enzyme controls the movement of the DNA into the pore so that it passes through the pore with an electric field until the target analyte is finally translocated to the trans side of the bilayer. Alternatively, the method is preferably carried out such that the helicase enzyme controls the movement of the target sequence through the pore against an electric field resulting from an applied electric potential. In this mode, the 3' end of the DNA is first captured in the pore, and the enzyme controls the movement of the DNA through the pore so that it is pulled out of the pore against an applied electric field until the target sequence is finally pushed out to the cis side of the bilayer.

[0206] In exonuclease sequencing, an exonuclease releases individual nucleotides from one end of a target polynucleotide, and these individual nucleotides are identified as discussed below. In another embodiment, a method of characterizing a target polynucleotide involves contacting a target sequence with a pore and an exonuclease enzyme. Any of the exonuclease enzymes discussed below can be used in this method. The enzyme can be covalently attached to the pore as discussed below.

[0207] An exonuclease is an enzyme that typically latches onto one end of a polynucleotide and digests a sequence of one nucleotide at a time from that end. An exonuclease can digest a polynucleotide in the 5' to 3' direction or the 3' to 5' direction. The end of a polynucleotide to which an exonuclease binds is typically determined by the choice of enzyme used and / or by using methods known in the art. Hydroxyl groups or cap structures at either end of a polynucleotide can typically be used to prevent or promote binding of an exonuclease to a particular end of a polynucleotide.

[0208] The method includes contacting a polynucleotide with an exonuclease such that nucleotides are digested from the end of the polynucleotide at a rate that allows for characterization or identification of the proportion of nucleotides as described above. Methods for doing this are well known in the art. For example, Edman degradation is used to sequentially digest single amino acids from the end of a polypeptide so that they can be identified using high performance liquid chromatography (HPLC). Homologous methods may also be used in the present invention.

[0209] The rate at which an exonuclease functions is typically slower than the optimal rate of the wild-type exonuclease. Suitable activity rates of the exonuclease in the method of the invention involve digestion of 0.5-1000 nucleotides per second, 0.6-500 nucleotides per second, 0.7-200 nucleotides per second, 0.8-100 nucleotides per second, 0.9-50 nucleotides per second, or 1-20 or 10 nucleotides per second. The rate is preferably 1, 10, 100, 500, or 1000 nucleotides per second. A suitable rate of exonuclease activity can be achieved in various ways. For example, variant exonucleases with reduced optimal activity rates can be used in accordance with the invention.

[0210] In strand characterization embodiments, the method comprises contacting a polynucleotide with a pore of the invention such that the polynucleotide translocates relative to, e.g., through, the pore, and taking one or more measurements as the polynucleotide translocates relative to the pore, the measurements being indicative of one or more properties of the polynucleotide, thereby characterizing the target polynucleotide.

[0211] In an embodiment of exonucleotide characterization, the method includes contacting a polynucleotide with a pore and an exonuclease of the invention, such that the exonuclease digests individual nucleotides from one end of the target polynucleotide and the individual nucleotides translocate relative to, e.g., through, the pore, and taking one or more measurements as the individual nucleotides translocate relative to the pore, the measurements being indicative of one or more properties of the individual nucleotides, thereby characterizing the target polynucleotide.

[0212] An individual nucleotide is a single nucleotide. An individual nucleotide is one that is not bound to another nucleotide or polynucleotide by a nucleotide bond. A nucleotide bond includes one of the phosphate groups of a nucleotide that is bound to the sugar group of another nucleotide. An individual nucleotide is typically one that is not bound to another polynucleotide of at least 5, at least 10, at least 20, at least 50, at least 100, at least 200, at least 500, at least 1000, or at least 5000 nucleotides by a nucleotide bond. For example, an individual nucleotide is digested from a target polynucleotide sequence, such as a DNA or RNA strand. The nucleotide can be any of those discussed below.

[0213] Individual nucleotides may interact with the pore in any manner and at any site. The nucleotides preferably reversibly bind to the pore via or in conjunction with the adaptors discussed above. The nucleotides most preferably reversibly bind to the pore via or in conjunction with the adaptors when passing through the pore across the membrane. The nucleotides can also reversibly bind to the barrel or channel of the pore via or in conjunction with the adaptors when passing through the pore across the membrane.

[0214] During the interaction between an individual nucleotide and the pore, the nucleotide typically affects the current flowing through the pore in a manner specific to that nucleotide. For example, a particular nucleotide will reduce the current flowing through the pore for a particular average period and to a particular extent. In other words, the current flowing through the pore is unique to a particular nucleotide. A control experiment can be performed to determine the effect of a particular nucleotide on the current passing through the pore. The results from the implementation of the method of the present invention on a test sample can then be compared to the results derived from such a control experiment to identify a particular nucleotide in the sample or to determine whether a particular nucleotide is present in the sample. The frequency with which the current flowing through the pore is affected in a manner indicative of a particular nucleotide can be used to determine the concentration of that nucleotide in the sample. The ratio of different nucleotides in a sample can also be calculated. For example, the ratio of dCMP to methyl-dCMP can be calculated.

[0215] The methods involve measuring one or more properties of a target polynucleotide, which may also be referred to as a template polynucleotide or a polynucleotide of interest.

[0216] This embodiment also makes use of the pores of the present invention. Any of the pores and embodiments discussed above with reference to the target analyte can be used.

[0217] Polynucleotide Analytes A polynucleotide, such as a nucleic acid, is a polymer that contains two or more nucleotides. A polynucleotide or a nucleic acid can contain any combination of any nucleotides. Nucleotides can be naturally occurring or artificial. One or more nucleotides in a polynucleotide can be oxidized or methylated. One or more nucleotides of a polynucleotide can be damaged. For example, a polynucleotide can contain pyrimidine dimers. Such dimers are typically associated with UV damage and are the main cause of skin melanoma. One or more nucleotides of a polynucleotide can be modified, for example, with a label or tag. Suitable labels are described below. A polynucleotide can contain one or more spacers.

[0218] A nucleotide typically contains a nucleobase, a sugar, and at least one phosphate group. The nucleobase and sugar form a nucleoside.

[0219] Nucleobases are typically heterocyclic. Nucleobases include, but are not limited to, purines and pyrimidines, more specifically adenine (A), guanine (G), thymine (T), uracil (U), and cytosine (C).

[0220] The sugar is typically a pentose sugar. Nucleotide sugars include, but are not limited to, ribose and deoxyribose. The sugar is preferably deoxyribose.

[0221] The polynucleotide preferably comprises the following nucleosides: deoxyadenosine (dA), deoxyuridine (dU) and / or thymidine (dT), deoxyguanosine (dG), and deoxycytidine (dC). The nucleotides are typically ribonucleotides or deoxyribonucleotides. The nucleotides typically contain monophosphates, diphosphates, or triphosphates. The nucleotides may contain four or more phosphates, such as four or five phosphates. The phosphates may be attached to the 5' or 3' side of the nucleotide. Nucleotides include, but are not limited to, adenosine monophosphate (AMP), guanosine monophosphate (GMP), thymidine monophosphate (TMP), uridine monophosphate (UMP), 5-methylcytidine monophosphate, 5-hydroxymethylcytidine monophosphate, cytidine monophosphate (CMP), cyclic adenosine monophosphate (cAMP), cyclic guanosine monophosphate (cGMP), deoxyadenosine monophosphate (dAMP), deoxyguanosine monophosphate (dGMP), deoxythymidine monophosphate (dTMP), deoxyuridine monophosphate (dUMP), deoxycytidine monophosphate (dCMP), and deoxymethylcytidine monophosphate. The nucleotide is preferably selected from AMP, TMP, GMP, CMP, UMP, dAMP, dTMP, dGMP, dCMP, and dUMP.

[0222] A nucleotide may be abasic (i.e., lacking a nucleobase). A nucleotide may also lack a nucleobase and a sugar (i.e., is a C3 spacer).

[0223] The nucleotides of a polynucleotide may be linked together in any manner. The nucleotides are typically linked by their sugar and phosphate groups, as in nucleic acids. The nucleotides may be connected through their nucleobases, as in pyrimidine dimers. The polynucleotide may be single-stranded or double-stranded. The polynucleotide is preferably single-stranded. Single-stranded polynucleotide characterization is referred to as 1D in the examples. At least a portion of the polynucleotide may be double-stranded. The polynucleotide may be a nucleic acid, such as deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). The polynucleotide may comprise one strand of RNA hybridized to one strand of DNA. The polynucleotide may be any synthetic nucleic acid known in the art, such as peptide nucleic acid (PNA), glycerol nucleic acid (GNA), threose nucleic acid (TNA), locked nucleic acid (LNA), or other synthetic polymers with nucleotide side chains. The PNA backbone is composed of repeating N-(2-aminoethyl)-glycine units linked by peptide bonds. The GNA backbone is composed of repeating glycol units linked by phosphodiester bonds. The TNA backbone is composed of repeating threose sugars linked together by phosphodiester bonds. LNAs are formed from the ribonucleotides mentioned above with an extra bridge connecting the 2' oxygen and the 4' carbon of the ribose moiety. Bridged Nucleic Acids (BNAs) are modified RNA nucleotides. They may also be called constrained RNA or inaccessible RNA. BNA monomers can contain 5-, 6-, or even 7-membered bridge structures with a "fixed" C3'-terminal sugar puckering. Bridges are synthetically incorporated at the 2',4' positions of the ribose to generate 2',4'-BNA monomers.

[0224] The polynucleotide is most preferably ribonucleic acid (RNA) or deoxyribonucleic acid (DNA).

[0225] A polynucleotide can be of any length. For example, a polynucleotide can be at least 10, at least 50, at least 100, at least 150, at least 200, at least 250, at least 300, at least 400, or at least 500 nucleotides or nucleotide pairs in length. A polynucleotide can be 1000 nucleotides or nucleotide pairs or more in length, 5000 nucleotides or nucleotide pairs or more in length, or 100000 nucleotides or nucleotide pairs or more in length.

[0226] Any number of polynucleotides can be investigated. For example, the methods of the invention can involve the characterization of 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 50, 100, or more polynucleotides. When more than one polynucleotide is characterized, they can be different polynucleotides or two instances of the same polynucleotide.

[0227] The polynucleotides may be naturally occurring or artificial. For example, the method can be used to verify the sequence of a manufactured oligonucleotide. The method is typically carried out in vitro.

[0228] sample Polynucleotides are typically present in any suitable sample. The invention is typically performed on a sample known or suspected to contain a polynucleotide. Alternatively, the invention can be performed on a sample to confirm the identity of a polynucleotide known or expected to be present in the sample.

[0229] The sample may be a biological sample. The invention may be performed in vitro using a sample obtained or extracted from any organism or microorganism. The organism or microorganism is typically an archaea, a prokaryote, or a eukaryote, typically belonging to one of the five kingdoms Plantae, Animalia, Fungi, Monera, and Protista. The invention may be performed in vitro on a sample obtained or extracted from any virus. The sample is preferably a fluid sample. The sample typically comprises a body fluid of a patient. The sample may be urine, lymph, saliva, mucus, or amniotic fluid, but is preferably blood, plasma, or serum.

[0230] Typically the sample is from a human, but may alternatively be from another mammal, for example from a commercially farmed animal such as a horse, cow, sheep, fish, chicken or pig, or a pet such as a cat or dog. Alternatively the sample may be from a plant, for example a sample obtained from a commercial crop such as a cereal, legume, fruit or vegetable, for example wheat, barley, oats, rapeseed, corn, soybean, rice, rhubarb, banana, apple, tomato, potato, grape, tobacco, bean, lentil, sugar cane, cocoa, cotton.

[0231] The sample may be a non-biological sample. The non-biological sample is preferably a liquid sample. Examples of non-biological samples include surgical fluids, water, such as drinking water, sea water, or river water, and reagents for clinical testing. The sample is typically processed before being used in the present invention, for example, by centrifugation or by passing through a membrane that filters out unwanted molecules or cells, such as red blood cells. It may be measured immediately after being collected. The sample may also typically be stored before assay, preferably below -70°C.

[0232] Polynucleotide characterization The method may involve measuring two, three, four, five or more characteristics of the polynucleotide, wherein the one or more characteristics are preferably selected from: (i) the length of the polynucleotide, (ii) the identity of the polynucleotide, (iii) the sequence of the polynucleotide, (iv) the secondary structure of the polynucleotide, and (v) whether the polynucleotide is modified. Any combination of (i)-(v) such as {i}, {ii}, {iii}, {iv}, {v}, {i,ii}, {i,iii}, {i,iv}, {i,v}, {ii,iii}, {ii,iv}, {ii,v}, {iii,iv}, {iii,v}, {iv,v}, {i,ii,iii}, {i,ii,iv}, {i,ii,v}, {i,iii,iv}, {i,iii,v}, {i,iv,v}, {ii,iii,iv}, {ii,iii,v}, {ii,iv,v}, {iii,iv,v}, {i,ii,iii,iv}, {i,ii,iii,v}, {i,ii,iii,iv}, {i,ii,iii,v}, {i,ii,iv ... or {i,ii,iii,iv,v} may be measured in accordance with the present invention. Different combinations of (i) through (v) can be measured for a first polynucleotide compared to a second polynucleotide, including any of those combinations listed above.

[0233] With regard to (i), the length of the polynucleotide may be measured, for example, by determining the number of interactions between the polynucleotide and the pore or the duration of an interaction between the polynucleotide and the pore.

[0234] Regarding (ii), the identity of a polynucleotide can be measured in several ways. The identity of a polynucleotide can be measured in conjunction with or without measuring the sequence of the polynucleotide. The former is straightforward, the polynucleotide is sequenced and thereby identified. The latter can be done in several ways. For example, the presence of a particular motif in the polynucleotide can be measured (without measuring the sequence of the rest of the polynucleotide). Alternatively, the measurement of a particular electrical and / or optical signal in the method can identify the polynucleotide as being obtained from a particular source.

[0235] For (iii), the sequence of the polynucleotide can be determined as described above. Suitable sequencing methods, particularly those using electrical measurements, are described in Stoddart D et al., Proc Natl Acad Sci, 12; 106(19): 7702-7, Lieberman KR et al., J Am Chem Soc. 2010; 132(50): 17961-72, and International Application WO2000 / 28312.

[0236] Regarding (iv), the secondary structure can be measured in a variety of ways. For example, if the method involves electrical measurements, the secondary structure can be measured using changes in residence time or changes in the current passing through the pore. This makes it possible to distinguish between regions of single-stranded and double-stranded polynucleotides.

[0237] For (v), the presence or absence of any modification can be measured. The method preferably includes determining whether the polynucleotide is modified by methylation, by oxidation, by damage, by one or more proteins, or by one or more labels, tags, or spacers. Specific modifications will result in specific interactions with the pore, which can be measured using the methods described below. For example, methylcyotsine can be distinguished from cytosine based on the current that passes through the pore during its interaction with each nucleotide.

[0238] The target polynucleotide contacts the pore of the present invention. The pore is typically present in a membrane. Suitable membranes are discussed below. The method may be carried out using any device suitable for investigating a membrane / pore system in which the pore is present in the membrane. The method may be carried out using any device suitable for transmembrane pore sensing. For example, the device comprises a chamber containing an aqueous solution and a barrier dividing the chamber into two compartments. The barrier typically has an opening in which the membrane containing the pore is formed. Alternatively, the barrier forms a membrane in which the pore is present.

[0239] The method may be carried out using the apparatus described in International Application No. PCT / GB08 / 000562 (WO2008 / 102120).

[0240] A variety of different types of measurements may be performed, including but not limited to electrical and optical measurements. Possible electrical measurements include current measurements, impedance measurements, tunneling measurements (Ivanov AP et al., Nano Lett. 2011 Jan 12; 11(1): 279-85), and FET measurements (International Application No. WO2005 / 124888). Optical measurements may be combined with electrical measurements (Soni GV et al., Rev Sci Instrum. 2010 Jan; 81(1): 014301, Chen C., et al “High spatial resolution nanoslit SERS for single-molecule nucleobase sensing.” Nat. Comm. (2018) 9: 1733). The measurement may be a transmembrane current measurement, for example, a measurement of the ionic current passing through the pore.

[0241] Electrical measurements can be performed using standard single channel recording devices as described in Stoddart D et al., Proc Natl Acad Sci, 12; 106(19): 7702-7, Lieberman KR et al, J Am Chem Soc. 2010; 132(50): 17961-72, and International Application No. WO2000 / 28312. Alternatively, electrical measurements can be performed using multi-channel systems as described, for example, in International Application No. WO2009 / 077734 and International Application No. WO2011 / 067559.

[0242] The method is preferably carried out with a potential applied across the membrane. The applied potential can be a voltage potential. Alternatively, the applied potential can be a chemical potential. An example of this is the use of a salt gradient across a membrane such as an amphiphilic layer. Salt gradients are disclosed in Holden et al., J Am Chem Soc. 2007 Jul 11;129(27):8650-5. In some instances, the current passing through the pore during the movement of the polynucleotide relative to the pore is used to infer or determine the sequence of the polynucleotide. This is strand sequencing.

[0243] The method may involve measuring the current passing through the pore as the polynucleotide moves relative to the pore. Thus, the device used in the method may also include an electrical circuit capable of applying a potential and measuring the electrical signal across the membrane and the pore. The method may be performed using a patch clamp or a voltage clamp. The method preferably involves the use of a voltage clamp.

[0244] The method of the invention may involve measuring the current passing through the pore as the polynucleotide translocates relative to the pore. Suitable conditions for measuring the ionic current through a transmembrane protein pore are known in the art and disclosed in the Examples. The method is typically carried out with a voltage applied across the membrane and pore. The voltage used is typically +5V to -5V, e.g. +4V to -4V, +3V to -3V, or +2V to -2V. The voltage used is typically -600mV to +600mV or -400mV to +400mV. The voltage used is preferably in a range having a lower limit selected from -400mV, -300mV, -200mV, -150mV, -100mV, -50mV, -20mV, and 0mV, and an upper limit independently selected from +10mV, +20mV, +50mV, +100mV, +150mV, +200mV, +300mV, and +400mV. The voltage used is more preferably in the range of 100mV to 240mV, most preferably in the range of 120mV to 220mV. By using an increased applied potential, it is possible to increase the discrimination between different nucleotides by the pore.

[0245] The method is typically carried out in the presence of any charge carrier, such as a metal salt, e.g. an alkali metal salt, a halogen salt, e.g. a chloride salt, e.g. an alkali metal chloride salt. The charge carrier may include an ionic liquid or an organic salt, e.g. tetramethylammonium chloride, trimethylphenylammonium chloride, phenyltrimethylammonium chloride, or 1-ethyl-3-methylimidazolium chloride. In the exemplary device described above, the salt is present in an aqueous solution in the chamber. Potassium chloride (KCl), sodium chloride (NaCl), cesium chloride (CsCl), or a mixture of potassium ferrocyanide and potassium ferricyanide are typically used. KCl, NaCl, and a mixture of potassium ferrocyanide and potassium ferricyanide are preferred. The charge carriers may be asymmetric across the membrane. For example, the type and / or concentration of the charge carriers may be different on each side of the membrane. The salt concentration may be saturated. The salt concentration may be 3M or less, typically 0.1-2.5M, 0.3-1.9M, 0.5-1.8M, 0.7-1.7M, 0.9-1.6M, or 1M-1.4M. The salt concentration is preferably 150mM-1M. The method is preferably carried out using a salt concentration of at least 0.3M, such as at least 0.4M, at least 0.5M, at least 0.6M, at least 0.8M, at least 1.0M, at least 1.5M, at least 2.0M, at least 2.5M, or at least 3.0M. A high salt concentration provides a high signal-to-noise ratio, allowing identification of a current indicative of the presence of a nucleotide against a background of normal current fluctuations. The method is typically carried out in the presence of a buffer. In the exemplary device described above, the buffer is present in the aqueous solution in the chamber. Any buffer can be used in the method of the invention. Typically, the buffer is a phosphate buffer. Other suitable buffers are HEPES and Tris-HCl buffers. The method is typically carried out at a pH of 4.0-12.0, 4.5-10.0, 5.0-9.0, 5.5-8.8, 6.0-8.7, or 7.0-8.8, or 7.5-8.5. The pH used is preferably about 7.5.

[0246] The method may be carried out at 0° C. to 100° C., 15° C. to 95° C., 16° C. to 90° C., 17° C. to 85° C., 18° C. to 80° C., 19° C. to 70° C., or 20° C. to 60° C. The method is typically carried out at room temperature. The method is optionally carried out at a temperature that supports enzyme function, for example, about 37° C.

[0247] Polynucleotide-binding proteins The strand characterization method preferably involves contacting a polynucleotide with a polynucleotide binding protein such that the protein controls the movement of the polynucleotide relative to, for example through, the pore.

[0248] More preferably, the method comprises (a) contacting a polynucleotide with a pore and a polynucleotide-binding protein of the invention such that the protein controls movement of the polynucleotide relative to, e.g., through, the pore; and (b) taking one or more measurements as the polynucleotide moves relative to the pore, the measurements being indicative of one or more properties of the polynucleotide, thereby characterizing the polynucleotide.

[0249] More preferably, the method comprises: (a) contacting a polynucleotide with a pore and a polynucleotide-binding protein of the invention such that the protein controls movement of the polynucleotide relative to, e.g., through, the pore; and (b) measuring a current through the pore as the polynucleotide moves relative to the pore, wherein the current is indicative of one or more properties of the polynucleotide, thereby characterizing the polynucleotide.

[0250] A polynucleotide-binding protein can be any protein that can bind to a polynucleotide and control its movement through a pore. It is easy in the art to determine whether a protein binds to a polynucleotide. A protein typically interacts with or modifies at least one property of a polynucleotide. A protein can modify a polynucleotide by cleaving it to form individual nucleotides or shorter chains of nucleotides, such as dinucleotides or trinucleotides. A protein can modify a polynucleotide by orienting it to a specific position or moving it, i.e., controlling its movement.

[0251] The polynucleotide binding protein is preferably derived from a polynucleotide handling enzyme. A polynucleotide handling enzyme is a polypeptide capable of interacting with a polynucleotide and modifying at least one property of the polynucleotide. An enzyme can modify a polynucleotide by cleaving the polynucleotide to form individual nucleotides or shorter nucleotide chains such as di- or tri-nucleotides. An enzyme can modify a polynucleotide by orienting or moving the polynucleotide to a specific position. A polynucleotide handling enzyme does not need to exhibit enzymatic activity as long as it can bind to a polynucleotide and control its movement through a pore. For example, the enzyme may be modified to remove its enzymatic activity or may be used under conditions that prevent it from acting as an enzyme. Such conditions are discussed in more detail below.

[0252] The polynucleotide handling enzyme is preferably derived from a nuclease. The polynucleotide handling enzyme used in the construction of the enzyme is more preferably derived from any member of Enzyme Classification (EC) groups 3.1.11, 3.1.13, 3.1.14, 3.1.15, 3.1.16, 3.1.21, 3.1.22, 3.1.25, 3.1.26, 3.1.27, 3.1.30, and 3.1.31. The enzyme may be any of those disclosed in International Application No. PCT / GB10 / 000133 (published as WO2010 / 086603).

[0253] Preferred enzymes are polymerases, exonucleases, helicases, and topoisomerases, such as gyrases. Suitable enzymes include, but are not limited to, exonuclease I from E. coli (SEQ ID NO: 3), exonuclease III enzyme from E. coli (SEQ ID NO: 4), RecJ from T. thermophilus (SEQ ID NO: 5) and bacteriophage lambda exonuclease (SEQ ID NO: 6), TatD exonuclease, and variants thereof. Three subunits comprising the sequence shown in SEQ ID NO: 5 or variants thereof interact to form a trimeric exonuclease. These exonucleases can also be used in the exonuclease method of the present invention. The polymerase can be PyroPhage® 3173 DNA polymerase (commercially available from Lucigen® Corporation), SD polymerase (commercially available from Bioron®), or variants thereof. The enzyme is preferably Phi29 DNA polymerase (SEQ ID NO: 7) or a variant thereof. The topoisomerase is preferably a member of any of the subclassification (EC) groups 5.99.1.2 and 5.99.1.3.

[0254] The enzyme is most preferably derived from a helicase such as Hel308 Mbu (SEQ ID NO: 8), Hel308 Csy (SEQ ID NO: 9), Hel308 Tga (SEQ ID NO: 10), Hel308 Mhu (SEQ ID NO: 11), TraI Eco (SEQ ID NO: 12), XPD Mbu (SEQ ID NO: 13), or variants thereof. Any helicase can be used in the present invention. The helicase can be or is derived from Hel308 helicase, RecD helicase, such as TraI helicase or TrwC helicase, XPD helicase, or Dda helicase. Helicases are disclosed in International Application Nos. PCT / GB2012 / 052579 (published as WO2013 / 057495), PCT / GB2012 / 053274 (published as WO2013 / 098562), PCT / GB2012 / 053273 (published as WO2013 / 098561), PCT / GB2013 / 051925 (published as WO2014 The helicase may be any of the helicases, modified helicases or helicase constructs disclosed in PCT / GB2013 / 051924 (published as WO2014 / 013260), PCT / GB2013 / 051924 (published as WO2014 / 013259), PCT / GB2013 / 051928 (published as WO2014 / 013262), and PCT / GB2014 / 052736.

[0255] The helicase preferably comprises the sequence shown in SEQ ID NO: 15 (Trwc Cba) or a variant thereof, the sequence shown in SEQ ID NO: 8 (Hel308 Mbu) or a variant thereof, or the sequence shown in SEQ ID NO: 14 (Dda) or a variant thereof. The variant may differ from the native sequence in any of the ways discussed below for the transmembrane pore. A preferred variant of SEQ ID NO: 14 comprises (a) E94C and A360C, or (b) E94C, A360C, C109A, and C136A, optionally followed by (ΔM1)G1G2 (i.e., deletion of M1 followed by addition of G1 and G2).

[0256] Any number of helicases can be used according to the invention, for example 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more helicases can be used, in some embodiments different numbers of helicases can be used.

[0257] The method of the invention preferably comprises contacting a polynucleotide with two or more helicases. The two or more helicases are typically the same helicase. The two or more helicases may be different helicases.

[0258] The two or more helicases may be any combination of the above-mentioned helicases. The two or more helicases may be two or more Dda helicases. The two or more helicases may be one or more Dda helicases and one or more TrwC helicases. The two or more helicases may be different variants of the same helicase.

[0259] The two or more helicases are preferably linked to each other. The two or more helicases are more preferably covalently linked to each other. The helicases may be linked in any order and using any method. Preferred helicase constructs for use in the present invention are described in International Application Nos. PCT / GB2013 / 051925 (published as WO2014 / 013260), PCT / GB2013 / 051924 (published as WO2014 / 013259), PCT / GB2013 / 051928 (published as WO2014 / 013262), and PCT / GB2014 / 052736.

[0260] A variant of SEQ ID NO: 7, 3, 4, 5, 16, 8, 9, 10, 11, 12, 13, 14, or 15 is an enzyme that has an amino acid sequence different from the sequence of SEQ ID NO: 7, 3, 4, 5, 16, 8, 9, 10, 11, 12, 13, 14, or 15 and retains polynucleotide binding ability. This can be measured using any method known in the art. For example, a variant can be contacted with a polynucleotide and its ability to bind and translocate along the polynucleotide can be measured. A variant can include modifications that enhance polynucleotide binding and / or enhance its activity at high salt concentration and / or room temperature. A variant can be modified to bind to a polynucleotide (i.e., retain polynucleotide binding ability) but not function as a helicase (i.e., do not translocate along a polypeptide when provided with all necessary components to facilitate translocation, e.g., ATP and Mg2+). Such modifications are known in the art. For example, modifications of the Mg2+-binding domain in a helicase typically result in variants that do not function as helicases. These types of variants may act as molecular brakes (see below).

[0261] The variants are preferably at least 50% homologous to the amino acid sequence of SEQ ID NO: 7, 3, 4, 5, 16, 8, 9, 10, 11, 12, 13, 14, or 15 based on amino acid identity over the entire length of the sequence. More preferably, the variant polypeptides may be at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, more preferably at least 95%, 97%, or 99% homologous to the amino acid sequence of SEQ ID NO: 7, 3, 4, 5, 16, 8, 9, 10, 11, 12, 13, 14, or 15 based on amino acid identity over the entire sequence. There may be at least 80%, for example at least 85%, 90%, or 95% amino acid identity ("hard homology") over a stretch of 200 or more, for example 230, 250, 270, 280, 300, 400, 500, 600, 700, 800, 900, or 1000 or more consecutive amino acids. Homology is determined as described above. Variants may differ from the wild-type sequence in any of the ways discussed above with reference to SEQ ID NO: 1 above. The enzyme may be covalently linked to the pore. Any method may be used to covalently link the enzyme to the pore.

[0262] A preferred molecular brake is TrwC Cba-Q594A (SEQ ID NO: 15 with the Q594A mutation), which does not function as a helicase (i.e., it binds to a polynucleotide but does not translocate along the polynucleotide when provided with all the necessary components to facilitate translocation, e.g., ATP and Mg2+).

[0263] In strand sequencing, polynucleotides are translocated through the pore with or against an applied potential. Exonucleases that act processively or processively on double-stranded polynucleotides can be used on the cis side of the pore to feed the remaining single strand under the applied potential or on the trans side under a reversed potential. Similarly, helicases that unwind double-stranded DNA can be used in a similar manner. Polymerases can also be used. There are potential sequencing applications that require strand translocation against an applied potential, but the DNA must first be "captured" by the enzyme under a reversed or no potential. Then, when the potential is returned following binding, the strand will pass through the pore from cis to trans and be held in an extended conformation by the current. Single-stranded DNA exonucleases or single-stranded DNA-dependent polymerases can act as molecular motors to pull the newly translocated single strand back through the pore from trans to cis in a controlled manner against the applied potential.

[0264] Any helicase can be used in the method. The helicase can work in two modes with respect to the pore. First, the method is preferably carried out using a helicase such that it moves the polynucleotide through the pore with the electric field resulting from the applied potential. In this mode, the 5' end of the polynucleotide is first captured in the pore, and the helicase moves the polynucleotide into the pore so that it passes through the pore with the electric field until it is finally translocated to the trans side of the membrane. Alternatively, the method is preferably carried out such that the helicase moves the polynucleotide through the pore with the electric field resulting from the applied electric potential. In this mode, the 3' end of the polynucleotide is first captured in the pore, and the helicase moves the polynucleotide through the pore so that it is pulled out of the pore with the electric field until it is finally pushed to the cis side of the membrane.

[0265] The method may also be performed in the reverse direction: the 3' end of the polynucleotide may first be trapped in the pore, and the helicase may translocate the polynucleotide into the pore so that it passes through the pore with an electric field until it is finally translocated to the trans side of the membrane.

[0266] If the helicase does not contain the necessary components to facilitate the translocation or is modified to hinder or prevent its translocation, it binds to the polynucleotide and acts as a brake to slow the translocation of the polynucleotide when it is drawn into the pore by the applied electric field. In the inactive mode, whether the polynucleotide is captured 3' or 5' down, it is the applied electric field that pulls the polynucleotide into the pore toward the trans side, with the enzyme acting as a brake. When in the inactive mode, the control of the translocation of the polynucleotide by the helicase can be described in several ways, including ratcheting, sliding, and glaking. Helicase variants lacking helicase activity can also be used in this method.

[0267] The polynucleotide can be contacted with the polynucleotide-binding protein and the pore in any order.When the polynucleotide is contacted with the polynucleotide-binding protein and the pore, such as a helicase, the polynucleotide is preferably first complexed with the protein.When a voltage is applied across the pore, the polynucleotide / protein complex then complexes with the pore and controls the movement of the polynucleotide through the pore.

[0268] Any step in the method of using polynucleotide binding proteins is typically carried out in the presence of free nucleotides or free nucleotide analogs and enzyme cofactors that facilitate the action of the polynucleotide binding proteins. The free nucleotides can be any one or more of the individual nucleotides discussed above. Free nucleotides include adenosine monophosphate (AMP), adenosine diphosphate (ADP), adenosine triphosphate (ATP), guanosine monophosphate (GMP), guanosine diphosphate (GDP), guanosine triphosphate (GTP), thymidine monophosphate (TMP), thymidine diphosphate (TDP), thymidine triphosphate (TTP), uridine monophosphate (UMP), uridine diphosphate (UDP), uridine triphosphate (UTP), cytidine monophosphate (CMP), cytidine diphosphate (CDP), cytidine triphosphate (CTP), cyclic adenosine monophosphate (cAMP), cyclic guanosine monophosphate (cGMP), deoxyadenosine monophosphate (dAMP), deoxyadenosine monophosphate (dAMP), deoxyadenosine tri ... The free nucleotides include, but are not limited to, adenosine diphosphate (dADP), deoxyadenosine triphosphate (dATP), deoxyguanosine monophosphate (dGMP), deoxyguanosine diphosphate (dGDP), deoxyguanosine triphosphate (dGTP), deoxythymidine monophosphate (dTMP), deoxythymidine diphosphate (dTDP), deoxythymidine triphosphate (dTTP), deoxyuridine monophosphate (dUMP), deoxyuridine diphosphate (dUDP), deoxyuridine triphosphate (dUTP), deoxycytidine monophosphate (dCMP), deoxycytidine diphosphate (dCDP), and deoxycytidine triphosphate (dCTP). The free nucleotide is preferably selected from AMP, TMP, GMP, CMP, UMP, dAMP, dTMP, dGMP, or dCMP. The free nucleotide is preferably adenosine triphosphate (ATP). An enzyme cofactor is a factor that allows a structure to function. The enzyme cofactor is preferably a divalent metal cation. The divalent metal cation is preferably Mg2+, Mn2+, Ca2+, or Co2+. The enzyme cofactor is most preferably Mg2+.

[0269] Helicase(s) and molecular brake(s) The method may include providing the target analyte with one or more helicases and one or more molecular brakes bound to the target polynucleotide, particularly where the target analyte is a polynucleotide. For example, the method of analyte characterization may include: (a) providing a polynucleotide with one or more helicases and one or more molecular brakes bound to the polynucleotide; (b) contacting a polynucleotide with a pore of the invention and applying a potential across the pore such that the one or more helicases and the one or more molecular brakes together control the movement of the polynucleotide relative to, e.g., through, the pore; (c) taking one or more measurements as the polynucleotide moves relative to the pore, the measurements being indicative of one or more characteristics of the polynucleotide, thereby characterizing the polynucleotide.

[0270] This type of method is discussed in detail in international application PCT / GB2014 / 052737.

[0271] A helicase can be any of those discussed above. The one or more molecular brakes can be any compound or molecule that binds to a polynucleotide and retards the movement of the polynucleotide through the pore. The one or more molecular brakes preferably include one or more compounds that bind to a polynucleotide. The one or more compounds are preferably one or more macrocycles. Suitable macrocycles include, but are not limited to, cyclodextrins, calixarenes, cyclic peptides, crown ethers, cucurbiturils, pyrararenes, derivatives thereof, or combinations thereof. The cyclodextrin or derivatives thereof can be any of those disclosed in Eliseev, AV, and Schneider, HJ. (1994) J. Am. Chem. Soc. 116, 6081-6088. The agent is more preferably heptakis-6-amino-β-cyclodextrin (am7-βCD), 6-monodeoxy-6-monoamino-β-cyclodextrin (am1-βCD), or heptakis-(6-deoxy-6-guanidino)-cyclodextrin (gu7-βCD). The one or more molecular brakes are preferably one or more single-stranded binding proteins (SSBs). The one or more molecular brakes are more preferably single-stranded binding proteins (SSBs) that include a carboxy-terminal (C-terminal) region that has no net negative charge, or (ii) a modified SSB that includes one or more modifications in its C-terminal region that reduce the net negative charge of the C-terminal region. The one or more molecular brakes are most preferably one of the SSBs disclosed in International Application No. PCT / GB2013 / 051924 (published as WO2014 / 013259). The one or more molecular brakes are preferably one or more polynucleotide binding proteins. A polynucleotide binding protein can be any protein that can bind to a polynucleotide and control its movement through a pore. It is straightforward in the art to determine whether a protein binds to a polynucleotide. A protein typically interacts with or modifies at least one property of a polynucleotide.A protein can modify a polynucleotide by cleaving it to form individual nucleotides or shorter chains of nucleotides, such as dinucleotides or trinucleotides. A moiety can modify a polynucleotide by orienting it to a particular location or translocating it, i.e., controlling its movement.

[0272] The polynucleotide binding protein is preferably derived from a polynucleotide handling enzyme. The one or more molecular brakes may be derived from any of the polynucleotide handling enzymes discussed above. A modified version of Phi29 polymerase (SEQ ID NO: 16) that functions as a molecular brake is disclosed in U.S. Pat. No. 5,576,204. The one or more molecular brakes are preferably derived from a helicase. Any number of molecular brakes derived from a helicase may be used. For example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more helicases may be used as molecular brakes. When two or more helicases are used as molecular brakes, the two or more helicases are typically the same helicase. The two or more helicases may be different helicases.

[0273] The two or more helicases may be any combination of the above-mentioned helicases. The two or more helicases may be two or more Dda helicases. The two or more helicases may be one or more Dda helicases and one or more TrwC helicases. The two or more helicases may be different variants of the same helicase.

[0274] The two or more helicases are preferably linked to each other. The two or more helicases are more preferably covalently linked to each other. The helicases can be linked in any order and using any method. The one or more molecular brakes derived from the helicase are preferably modified to reduce the size of the opening in the polynucleotide binding domain, through which the polynucleotide can be unbound from the helicase in at least one conformational state. This is disclosed in WO2014 / 013260. Preferred helicase constructs for use in the present invention are described in International Application Nos. PCT / GB2013 / 051925 (published as WO2014 / 013260), PCT / GB2013 / 051924 (published as WO2014 / 013259), PCT / GB2013 / 051928 (published as WO2014 / 013262), and PCT / GB2014 / 052736.

[0275] When one or more helicases are used in active mode (i.e., when the one or more helicases are provided with all the components necessary to facilitate translocation, e.g., ATP and Mg2+), the one or more molecular brakes are preferably used in either (a) an inactive mode (i.e., when the components necessary to facilitate translocation are absent or active translocation is not possible), (b) an active mode in which the one or more molecular brakes move in the opposite direction to the one or more helicases, or (c) an active mode in which the one or more molecular brakes move in the same direction as the one or more helicases and move more slowly than the one or more helicases.

[0276] When one or more helicases are used in an inactive mode (i.e. when one or more helicases do not have all the components required to facilitate translocation, e.g. ATP and Mg2+, or when active translocation is not possible), the one or more molecular brakes are preferably either (a) used in an inactive mode (i.e. when the components required to facilitate translocation are absent or when active translocation is not possible), or (b) used in an active mode, where the one or more molecular brakes translocate along the polynucleotide through the pore in the same direction as the polynucleotide.

[0277] The one or more helicases and the one or more molecular brakes may be attached to the polynucleotide at any position such that together they control the movement of the polynucleotide through the pore. The one or more helicases and the one or more molecular brakes are at least one nucleotide apart, such as at least 5, at least 10, at least 50, at least 100, at least 500, at least 1000, at least 5000, at least 10,000, at least 50,000 nucleotides apart or more. When the method relates to characterizing a double-stranded polynucleotide with a Y adaptor at one end and a hairpin loop adaptor at the other end, the one or more helicases are preferably attached to the Y adaptor and the one or more molecular brakes are preferably attached to the hairpin loop adaptor. In this embodiment, the one or more molecular brakes are preferably one or more helicases that bind to the polynucleotide but are modified so as not to function as helicases. The one or more helicases attached to the Y adaptor are preferably inhibited with a spacer, as discussed in more detail below. The one or more molecular brakes attached to the hairpin loop adaptor are preferably not inhibited by a spacer. The one or more helicases and the one or more molecular brakes preferably come together when the one or more helicases reach the hairpin loop. The one or more helicases may be attached to the Y adaptor before the Y adaptor is attached to the polynucleotide or after the Y adaptor is attached to the polynucleotide. The one or more molecular brakes may be attached to the hairpin loop adaptor before the hairpin loop adaptor is attached to the polynucleotide or after the hairpin loop adaptor is attached to the polynucleotide.

[0278] The one or more helicases and the one or more molecular brakes are preferably not bound to each other. The one or more helicases and the one or more molecular brakes are more preferably not covalently bound to each other. The one or more helicases and the one or more molecular brakes are preferably not bound as described in International Application Nos. PCT / GB2013 / 051925 (published as WO2014 / 013260), PCT / GB2013 / 051924 (published as WO2014 / 013259), PCT / GB2013 / 051928 (published as WO2014 / 013262), and PCT / GB2014 / 052736.

[0279] Spacer The one or more helicases may be inhibited with one or more spacers, as described in International Application No. PCT / GB2014 / 050175. Any configuration of one or more helicases and one or more spacers disclosed in the International Application may be used in the present invention.

[0280] When a portion of the polynucleotide enters the pore and moves through the pore along the electric field resulting from the applied potential, the one or more helicases are moved by the pore through the spacers as the polynucleotide moves through the pore, as the polynucleotide (including the one or more spacers) moves through the pore and the one or more helicases remain at the top of the pore.

[0281] The one or more spacers are preferably part of the polynucleotide, e.g., they interrupt the polynucleotide sequence. The one or more spacers are preferably not part of one or more blocking molecules, e.g., speed bumps, hybridized to the polynucleotide.

[0282] Any number of spacers may be present in a polynucleotide, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more spacers. Preferably, 2, 4, or 6 spacers are present in a polynucleotide. There may be one or more spacers in different regions of the polynucleotide, such as one or more spacers in a Y adaptor and / or hairpin loop adaptor.

[0283] The one or more spacers each provide an energy barrier that the one or more helicases cannot overcome even in active mode. The one or more spacers may inhibit the one or more helicases by reducing the pulling force of the helicase (e.g., by removing bases from nucleotides of the polynucleotide) or by physically blocking the movement of the one or more helicases (e.g., using bulky chemical groups).

[0284] The one or more spacers may comprise any molecule or combination of molecules that inhibit one or more helicases. The one or more spacers may comprise any molecule or combination of molecules that prevent one or more helicases from moving along the polynucleotide. In the absence of a transmembrane pore and an applied potential, it is easy to determine whether one or more helicases are inhibited by one or more spacers. For example, the ability of the helicase to move through the spacer and dissociate the complementary strand of DNA can be measured by PAGE.

[0285] The one or more spacers typically comprise linear molecules such as polymers. The one or more spacers typically have a structure different from that of a polynucleotide. For example, if the polynucleotide is DNA, the one or more spacers are typically not DNA. In particular, if the polynucleotide is deoxyribonucleic acid (DNA) or ribonucleic acid (RNA), the one or more spacers preferably comprise peptide nucleic acid (PNA), glycerol nucleic acid (GNA), threose nucleic acid (TNA), locked nucleic acid (LNA), or synthetic polymers with nucleotide side chains. The one or more spacers may comprise one or more nucleotides in the opposite direction to the polynucleotide. For example, if the polynucleotide is in the 5' to 3' direction, the one or more spacers may comprise one or more nucleotides in the 3' to 5' direction. The nucleotides may be any of those discussed above.

[0286] The one or more spacers preferably comprise one or more nitroindoles, e.g., one or more 5-nitroindoles, one or more inosines, one or more acridines, one or more 2-aminopurines, one or more 2-6-diaminopurines, one or more 5-bromo-deoxyuridines, one or more inverted thymidines (inverted dT), one or more inverted dideoxy-thymidines (ddT), one or more dideoxy-cytidines (ddC), one or more 5-methylcytidines, one or more 5-hydroxymethylcytidines, one or more 2'-O-methyl RNA bases, one or more iso-deoxycytidines (Iso-dC), one or more iso-deoxyguanosines (Iso-dG), one or more iSpC3 groups (i.e., nucleotides lacking a sugar and a base), one or more photocleavable (PC) groups, one or more hexanediol groups, one or more spacer 9 (iSp9) groups, one or more spacer 18 (iSp18) groups, a polymer, or one or more thiol linkages. One or more of the spacers may contain any combination of these groups, many of which are commercially available from IDT® (Integrated DNA Technologies®).

[0287] The one or more spacers may contain any number of these groups. For example, for 2-aminopurine, 2-6-diaminopurine, 5-bromo-deoxyuridine, inverted dT, ddT, ddC, 5-methylcytidine, 5-hydroxymethylcytidine, 2'-O-methyl RNA bases, iso-dC, iso-dG, iSpC3 groups, PC groups, hexanediol groups, and thiol bonds, the one or more spacers preferably contain 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12 or more. The one or more spacers preferably contain 2, 3, 4, 5, 6, 7, 8 or more iSp9 groups. The one or more spacers preferably contain 2, 3, 4, 5, or 6 or more iSp18 groups. The most preferred spacer is 4 iSp18 groups.

[0288] The polymer is preferably a polypeptide or polyethylene glycol (PEG). A polypeptide preferably comprises 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12 or more amino acids. A PEG preferably comprises 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12 or more monomer units.

[0289] The one or more spacers preferably contain one or more abasic nucleotides (i.e., nucleotides lacking a nucleobase), for example, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, or more abasic nucleotides. The nucleobase may be replaced by -H (idSp) or -OH in the abasic nucleotide. The abasic spacer may be inserted into the polynucleotide by removing the nucleobase from one or more adjacent nucleotides. For example, the polynucleotide may be modified to contain 3-methyladenine, 7-methylguanine, 1,N6-ethenoadenine inosine, or hypoxanthine, and the nucleobase may be removed from these nucleotides using human alkyladenine DNA glycosylase (hAAG). Alternatively, the polynucleotide may be modified to contain uracil, and the nucleobase removed with uracil-DNA glycosylase (UDG). In one embodiment, the one or more spacers do not contain any abasic nucleotides.

[0290] One or more helicases may be inhibited by (i.e., before) or on each linear molecular spacer. When linear molecular spacers are used, the polynucleotide is preferably provided with a double-stranded region of the polynucleotide adjacent to the end of each spacer, along which one or more helicases move. The double-stranded region typically serves to inhibit one or more helicases on the adjacent spacer. The presence of a double-stranded region(s) is particularly preferred when the method is performed at a salt concentration of about 100 mM or less. Each double-stranded region is typically at least 10 nucleotides long, for example at least 12 nucleotides long. When the polynucleotide used in the present invention is single-stranded, the double-stranded region may be formed by hybridizing a shorter polynucleotide to a region adjacent to the spacer. The shorter polynucleotide is typically formed from the same nucleotides as the polynucleotide, but may also be formed from different nucleotides. For example, the shorter polynucleotide may be formed from LNA.

[0291] When linear molecular spacers are used, the polynucleotide is preferably provided with a blocking molecule at the end of each spacer opposite the end where the one or more helicases move. This can help ensure that the one or more helicases remain inhibited on each spacer. This can also help to retain the one or more helicases on the polynucleotide when the polynucleotide diffuses in solution. The blocking molecule can be any of the chemical groups discussed below that cause physical inhibition of the one or more helicases. The blocking molecule can be a double-stranded region of the polynucleotide.

[0292] The one or more spacers preferably comprise one or more chemical groups that physically inhibit one or more helicases. The one or more chemical groups are preferably one or more pendant chemical groups. The one or more chemical groups may be attached to one or more nucleobases in the polynucleotide. The one or more chemical groups may be attached to the backbone of the polynucleotide. There may be any number of these chemical groups, such as 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12 or more. Suitable groups include, but are not limited to, fluorophores, streptavidin and / or biotin, cholesterol, methylene blue, dinitrophenol (DNPs), digoxigenin and / or antidigoxigenin, and dibenzylcyclooctyne groups.

[0293] Different spacers in a polynucleotide may contain different inhibitor molecules. For example, one spacer may contain one of the linear molecules discussed above, and another spacer may contain one or more chemical groups that physically inhibit one or more helicases. A spacer may contain any of the linear molecules discussed above and one or more chemical groups that physically inhibit one or more helicases, such as one or more abasics and fluorophores.

[0294] Suitable spacers can be designed depending on the type of polynucleotide and the conditions under which the method of the invention is carried out. Most helicases bind to and move along DNA, and therefore can be inhibited using anything other than DNA. Suitable molecules are discussed above.

[0295] The method of the present invention is preferably carried out in the presence of free nucleotides and / or in the presence of a helicase cofactor, which is discussed in more detail below. In the absence of a transmembrane pore and an applied potential, the one or more spacers are preferably capable of inhibiting the one or more helicases in the presence of free nucleotides and / or in the presence of a helicase cofactor.

[0296] When the methods of the invention are performed in the presence of free nucleotides and a helicase cofactor (such that one of the helicases is in active mode) as discussed below, one or more longer spacers are typically used to ensure that one or more helicases are inhibited on the polynucleotide before contacting the transmembrane pore and applying a potential. One or more shorter spacers may be used in the absence of free nucleotides and a helicase cofactor (such that one or more helicases are in inactive mode).

[0297] Salt concentration also affects the ability of the one or more spacers to inhibit one or more helicases. In the absence of a transmembrane pore and an applied potential, the one or more spacers are preferably capable of inhibiting one or more helicases at a salt concentration of about 100 mM or less. The higher the salt concentration used in the method of the present invention, the shorter the one or more spacers typically used, and vice versa.

[0298] Preferred combinations of features are shown in Table 3 below. [Table 3]

[0299] The method may involve translocating two or more helicases through a spacer. In such instances, the length of the spacer is typically increased to prevent the subsequent helicase from passing the preceding helicase through the spacer in the absence of a pore and an applied potential. When the method involves translocating two or more helicases through one or more spacers, the length of the spacer discussed above may be increased by at least 1.5-fold, such as 2-fold, 2.5-fold, or 3-fold.

[0300] Polypeptide characterization The methods of the invention may also be utilized to characterize a target polypeptide. Thus, the invention provides a method of characterizing a target polypeptide, comprising: (a) contacting a target polypeptide with a cytotoxin K pore such that the target analyte translocates relative to the pore; (b) taking one or more measurements characteristic of the polypeptide as it moves relative to the pore; thereby characterizing the target polynucleotide.

[0301] The cytotoxin K pore may be a wild-type pore or a pore comprising a mutant CytK monomer of the invention as described herein.

[0302] The method of polypeptide characterization described herein may include, and the invention may include, (i) contacting a polypeptide with a polypeptide handling enzyme capable of controlling the movement of the polypeptide relative to the pore, and (ii) taking one or more measurements characteristic of the polypeptide as the polypeptide moves relative to the pore. More preferably, the method of characterizing a target analyte includes characterizing a target polypeptide, but the method preferably includes forming a conjugate with a polynucleotide and controlling the movement of the conjugate relative to the nanopore using a polynucleotide handling protein, such as a polynucleotide handling enzyme. The method of the present disclosure may also involve controlling the movement of the polypeptide relative to the nanopore using a polypeptide handling enzyme. Such methods involving the use of a polypeptide or polynucleotide binding protein are described in more detail in WO2021 / 111125 and are applicable to the method of polypeptide characterization involving the use of the mutant CytK monomer of the invention described herein.

[0303] The methods disclosed herein utilize the ability of polynucleotide handling proteins to control the movement of conjugates that do not only contain polynucleotides.In particular, conjugates that contain polypeptides can be moved in a controlled manner using polynucleotide handling proteins as described herein.Polynucleotide handling proteins suitable for use in the disclosed methods are described in more detail herein.

[0304] Thus, the method for characterizing a target polypeptide preferably comprises: - conjugating a target polypeptide to a polynucleotide to form a polynucleotide-polypeptide conjugate; - contacting the conjugate with a polynucleotide handling protein capable of controlling the translocation of the polynucleotide relative to the nanopore; - taking one or more measurements characteristic of the polypeptide as the conjugate translocates relative to the nanopore; The polypeptide is thereby characterized.

[0305] Any suitable polypeptide can be characterized using the methods disclosed herein. In some embodiments, the target polypeptide is a protein or a naturally occurring polypeptide. In some embodiments, the polypeptide is a synthetic polypeptide. Polypeptides that can be characterized according to the disclosed methods are described in more detail herein.

[0306] Any suitable polynucleotide can be used in forming the conjugate for use in the methods disclosed herein.In some embodiments, the polynucleotide has at least the same length as the portion of the target polypeptide to be characterized.In some embodiments, the polynucleotide has a length longer than the portion of the target polypeptide to be characterized.This will be discussed in more detail below.Polynucleotides suitable for use in the methods disclosed herein are disclosed in more detail herein.

[0307] In the disclosed methods, the target polypeptide can be conjugated to the polynucleotide using any suitable means, several exemplary means are described in more detail herein.

[0308] The conjugate formed by the disclosed methods is contacted with a polynucleotide handling protein capable of controlling the translocation of the polynucleotide relative to the nanopore. Exemplary polynucleotide handling proteins are described in more detail herein.

[0309] The polynucleotide handling protein controls the translocation of a polynucleotide to a nanopore comprising a mutant CytK monomer of the invention.Any pore of the invention is suitable for use in the methods of polypeptide characterization described herein.

[0310] The disclosed method includes taking one or more measurements characteristic of the polypeptide as the conjugate moves relative to the nanopore. The one or more measurements can be any suitable measurements. Typically, the one or more measurements are electrical measurements, such as current measurements, and / or one or more optical measurements. Apparatus for recording suitable measurements, and information such measurements can provide, are described in more detail herein.

[0311] As disclosed herein, polynucleotide can be used to control the movement of polypeptide to the nanopore comprising the CytK monomer of the present invention.The movement of polynucleotide is controlled by polynucleotide handling protein.Since polynucleotide is conjugated to the polypeptide in the conjugate, the movement of polynucleotide drives the movement of polypeptide.

[0312] The use of polynucleotide handling proteins to control the movement of polynucleotides, and thus the movement of polypeptides, can be associated with advantages compared to the methods for characterizing polypeptides known in the art.For example, polynucleotide handling proteins can process the handling of polynucleotides with a higher turnover rate compared to polypeptide handling enzymes.This means that characterization data can be obtained more quickly for the polypeptides characterized according to the disclosed methods compared to previously known methods.

[0313] These and other advantages will become apparent throughout this disclosure.

[0314] The polynucleotide handling protein is preferably located on the cis side of the nanopore and translocates the conjugate into the pore, i.e., from the cis side to the trans side. The opposite setup can also be used.

[0315] In other words, in some embodiments, the polynucleotide handling protein is located on the cis side of the nanopore, and the polynucleotide handling protein controls the movement of the conjugate from the cis side of the nanopore to the trans side of the nanopore. Thus, in some embodiments, the polynucleotide handling protein is located on the cis side of the nanopore, and the polynucleotide handling protein controls the movement of the polynucleotide from the cis side of the nanopore to the trans side of the nanopore, thereby controlling the movement of the polypeptide through the nanopore.

[0316] In other embodiments, the polynucleotide handling protein is located on the trans side of the nanopore, and the polynucleotide handling protein controls the movement of the conjugate from the trans side of the nanopore to the cis side of the nanopore. Thus, in some embodiments, the polynucleotide handling protein is located on the trans side of the nanopore, and the polynucleotide handling protein controls the movement of the polynucleotide from the trans side of the nanopore to the cis side of the nanopore, thereby controlling the movement of the polypeptide through the nanopore.

[0317] As described herein, the conjugate may include a leader. As described herein, any suitable leader may be used. Optionally, the leader may be a polynucleotide. The leader may be the same as or different from the polynucleotide in the conjugate. As described above, the leader may facilitate threading of the conjugate through the nanopore.

[0318] In other words, in some embodiments, the conjugate is L-{PN}-P m and wherein the structure comprises one or more structures of the form - L is a leader, L is optionally an N moiety, -P is a polypeptide, -N comprises a polynucleotide, and - m is 0 or 1, The method may include threading a leader (L) through the nanopore, thereby contacting the polypeptide (P) with the nanopore.

[0319] In some such embodiments, the polynucleotide handling protein is located on the cis side of the nanopore and the method comprises enabling the polynucleotide handling protein to control the movement of the polynucleotide portion (N) from the cis side of the nanopore to the trans side of the nanopore, thereby controlling the movement of the polypeptide (P) through the nanopore. In other embodiments, the polynucleotide handling protein is located on the trans side of the nanopore and the method comprises enabling the polynucleotide handling protein to control the movement of the polynucleotide portion (N) from the trans side of the nanopore to the cis side of the nanopore, thereby controlling the movement of the polypeptide (P) through the nanopore.

[0320] As described in more detail herein, a conjugate may include one or more adaptors and / or anchors.

[0321] As described in more detail herein, in some embodiments, the conjugate comprises multiple polynucleotides and polypeptides. In such embodiments, the polynucleotide handling protein sequentially controls the movement of the polynucleotide relative to the nanopore, thus sequentially moving the polypeptide relative to the nanopore. In this manner, each polypeptide in the conjugate can be sequentially characterized in the disclosed methods.

[0322] For example, the conjugate is L-P1-N-{PN} n -P m and wherein the structure is one or more of the form: -n is a positive integer, - L is a leader, L is optionally an N moiety, - each P, which may be the same or different, is a polypeptide; - each N, which may be the same or different, comprises a polynucleotide; and - m is 0 or 1, The method may include threading a leader (L) through the nanopore, thereby contacting the polypeptide (P1) with the nanopore.

[0323] Typically, in such embodiments, n is from 1 to about 1000, such as from 2 to about 100, for example from about 3 to about 10, such as 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10.

[0324] In some such embodiments, the polynucleotide handling protein is located on the cis side of the nanopore and the method comprises enabling the polynucleotide handling protein to sequentially control the movement of each polynucleotide (N) from the cis side of the nanopore to the trans side of the nanopore, thereby controlling the movement of each polypeptide (P) sequentially through the nanopore. In other such embodiments, the polynucleotide handling protein is located on the trans side of the nanopore and the method comprises enabling the polynucleotide handling protein to sequentially control the movement of each polynucleotide (N) from the trans side of the nanopore to the cis side of the nanopore, thereby controlling the movement of each polypeptide (P) sequentially through the nanopore.

[0325] Those skilled in the art will appreciate that where a conjugate comprises more than one polypeptide, it may be advantageous for the polynucleotide handling protein (as described in more detail herein) to remain bound to the conjugate when contacting the polypeptide, rather than dissociating. In particular, this allows the polynucleotide handling protein to move over successive portions of the polynucleotide, passing over them as it contacts the polypeptide portions in the conjugate, in order to control movement of the conjugate relative to the nanopore.

[0326] The conjugate may comprise a polynucleotide and a polypeptide, and is contacted with a polynucleotide handling protein such that the polypeptide threads the nanopore. In the illustrated embodiment, a leader (optionally an additional polynucleotide) is used to facilitate threading of the polypeptide through the nanopore. While such use is within the scope of the disclosed methods, it is not required.

[0327] The polynucleotide handling protein processes the polynucleotide conjugated to the polypeptide.When the polynucleotide handling protein processes the polynucleotide, the conjugate passes through the nanopore, and thus the polypeptide passes through the nanopore.When the polypeptide passes through the nanopore, the polypeptide is characterized.

[0328] The polynucleotide handling protein moves the conjugate "out" of the pore from the "viewpoint" of the polynucleotide handling protein. For example, as shown, the polynucleotide handling protein is located on the cis side of the nanopore and moves the conjugate into the pore, i.e., from the trans side to the cis side. The opposite setup can also be used.

[0329] In other words, in some embodiments, the polynucleotide handling protein is located on the cis side of the nanopore, and the polynucleotide handling protein controls the movement of the conjugate from the trans side of the nanopore to the cis side of the nanopore. Thus, in some embodiments, the polynucleotide handling protein is located on the cis side of the nanopore, and the polynucleotide handling protein controls the movement of the polynucleotide from the trans side of the nanopore to the cis side of the nanopore, thereby controlling the movement of the polypeptide through the nanopore.

[0330] In other embodiments, the polynucleotide handling protein is located on the trans side of the nanopore, and the polynucleotide handling protein controls the movement of the conjugate from the cis side of the nanopore to the trans side of the nanopore. Thus, in some embodiments, the polynucleotide handling protein is located on the trans side of the nanopore, and the polynucleotide handling protein controls the movement of the polynucleotide from the cis side of the nanopore to the trans side of the nanopore, thereby controlling the movement of the polypeptide through the nanopore.

[0331] Using similar notation as above, in some embodiments, the conjugate is L-{PN}-P m and wherein the structure comprises one or more structures of the form - L is a leader, L is optionally an N moiety, -P is a polypeptide, -N comprises a polynucleotide, - m is 0 or 1, The method may include threading a leader (L) through the nanopore, thereby contacting the polypeptide (P) with the nanopore.

[0332] In some such embodiments, the polynucleotide handling protein is located on the cis side of the nanopore and the method comprises enabling the polynucleotide handling protein to control the movement of a polynucleotide (N) from the trans side of the nanopore to the cis side of the nanopore, thereby controlling the movement of a polypeptide (P) through the nanopore. In other such embodiments, the polynucleotide handling protein is located on the trans side of the nanopore and the method comprises enabling the polynucleotide handling protein to control the movement of a polynucleotide (N) from the cis side of the nanopore to the trans side of the nanopore, thereby controlling the movement of a polypeptide (P) through the nanopore.

[0333] In some embodiments, particularly those in which the polynucleotide handling protein controls the movement of the conjugate "out" of the nanopore as discussed above, the conjugate may include a blocking moiety attached to the polypeptide via an optional linker. The blocking moiety is typically too large to pass through the nanopore, so that movement of the conjugate relative to the nanopore brings the blocking moiety into contact with the nanopore, preventing further movement of the conjugate through the nanopore. At such times, the polynucleotide handling protein may be able to temporarily disassociate from the conjugate. In embodiments of the disclosed methods in which the conjugate moves relative to the nanopore under an applied force (e.g., electric potential or chemical potential), the conjugate may "move back" through the pore in a direction opposite to the movement controlled by the polynucleotide handling protein. Movement of the conjugate back through the pore allows the polypeptide portion of the conjugate to be recharacterized.

[0334] This process can be repeated multiple times by successively binding and rebinding the polynucleotide handling protein to the conjugate. In this manner, the conjugate can be oscillated through the pore (i.e., "flossed" through the nanopore). This "flossing" allows the polypeptide portion of the conjugate to be repeatedly characterized by the nanopore. In some embodiments, this can increase the accuracy of the characterization information.

[0335] In such embodiments, any suitable blocking moiety can be used. For example, the conjugate can be modified with biotin and the blocking moiety can be, for example, streptavidin, avidin, or neutravidin. The blocking moiety can be a large chemical group, such as a dendrimer. The blocking moiety can be a nanoparticle or a bead. Other suitable blocking moieties will be apparent to those skilled in the art.

[0336] Thus, in some embodiments, the method comprises: i) contacting the conjugate with the nanopore such that the blocking moiety is on the opposite side of the nanopore to the polynucleotide handling protein; ii) contacting the conjugated polynucleotide with a polynucleotide handling protein; iii) enabling the polynucleotide handling protein to control the translocation of the polynucleotide relative to the nanopore, thereby controlling the translocation of the polypeptide through the nanopore; iv) when the blocking moiety contacts the nanopore, thereby preventing further translocation of the conjugate through the nanopore, the polynucleotide handling protein temporarily unbinds from the polynucleotide, thereby allowing the conjugate to translocate through the nanopore under an applied force in a direction opposite to the translocation direction controlled by the polynucleotide handling protein; v) optionally repeating steps (ii)-(iv) to oscillate the polypeptide through the nanopore.

[0337] Polypeptides As explained above, the disclosed methods can include characterizing the target polypeptide within the conjugate as the conjugate translocates relative to the nanopore.

[0338] Any suitable polypeptide can be characterized by the disclosed methods.

[0339] In some embodiments, the target polypeptide is an unmodified protein or portion thereof, or a naturally occurring polypeptide or portion thereof.

[0340] In some embodiments, the target polypeptide is secreted from the cell. Alternatively, the target polypeptide may be produced intracellularly, such that it must be extracted from the cell for characterization by the disclosed methods. The polypeptide may be contained within a plasmid, e.g., Sambrook et al., Molecular Cloning: A Laboratory Manual, 4 th ed., Cold Spring Harbor Press, Plainsview, New York (2012), and Ausubel et al., Current Protocols in Molecular Biology (Supplement 114), John Wiley & Sons, New York (2016).

[0341] Polypeptides may be obtained or extracted from any organism or microorganism. Polypeptides may be obtained from humans or animals, for example, from urine, lymph, saliva, mucus, semen, or amniotic fluid, or from whole blood, plasma, or serum. Polypeptides may be obtained from plants, for example, from cereals, legumes, fruits, or vegetables.

[0342] The target polypeptide may be provided as an impure mixture of one or more polypeptides and one or more impurities. The impurities may include truncated forms of the target polypeptide that are different from the "target polypeptide" for characterization in the disclosed methods. For example, the target polypeptide may be a full-length protein and the impurities may include fractions of the protein. The impurities may also include proteins other than the target protein, which may be, for example, co-purified from a cell culture or obtained from a sample.

[0343] A polypeptide can contain any combination of amino acids, amino acid analogs, and modified amino acids (i.e., amino acid derivatives). The amino acids (and derivatives, analogs, etc.) in a polypeptide can be distinguished by their physical size and charge.

[0344] The amino acids / derivatives / analogs may be naturally occurring or artificial.

[0345] In some embodiments, the polypeptide may comprise any naturally occurring amino acid. Twenty amino acids are encoded by the universal genetic code. These are alanine (A), arginine (R), asparagine (N), aspartic acid (D), cysteine ​​(C), glutamic acid / glutamate (E), glutamine (Q), glycine (G), histidine (H), isoleucine (I), leucine (L), lysine (K), methionine (M), phenylalanine (F), proline (P), serine (S), threonine (T), tryptophan (W), tyrosine (Y), and valine (V). Other naturally occurring amino acids include selenocysteine ​​and pyrrolysine.

[0346] In some embodiments, the polypeptide is modified. In some embodiments, the polypeptide is modified for detection using the disclosed methods. In some embodiments, the disclosed methods are for characterizing modifications in a target polypeptide.

[0347] In some embodiments, one or more of the amino acids / derivatives / analogs in the polypeptide are modified. In some embodiments, one or more of the amino acids / derivatives / analogs in the polypeptide are post-translationally modified. Thus, the methods disclosed herein can be used to detect the presence, absence, and number of positions of post-translational modifications in a polypeptide. The methods disclosed can be used to characterize the degree to which a polypeptide is post-translationally modified.

[0348] Any one or more post-translational modifications may be present on a polypeptide. Typical post-translational modifications include modification with hydrophobic groups, modification with cofactors, addition of chemical groups, glycosylation (non-enzymatic attachment of sugars), biotinylation and pegylation. Post-translational modifications may also be non-natural, such as chemical modifications carried out in a laboratory for biotechnological or biomedical purposes. This may allow monitoring the levels of laboratory-produced peptides, polypeptides or proteins as opposed to their natural counterparts.

[0349] Examples of post-translational modifications by hydrophobic groups include myristoylation, myristic acid, C 14 Attachment of saturated acids; palmitoylation, palmitic acid, C 16 These include attachment of a saturated acid; isoprenylation or prenylation, the attachment of an isoprenoid group; farnesylation, the attachment of a farnesol group; geranylgeranylation, the attachment of a geranylgeraniol group; and glypiation, and the formation of a glycosylphosphatidylinositol (GPI) anchor via an amide bond.

[0350] Examples of post-translational modifications by cofactors include lipoylation, attachment of a lipoate (C8) functional group; flavinylation, attachment of a flavin moiety (e.g., flavin mononucleotide (FMN) or flavin adenine dinucleotide (FAD)); attachment of heme C, e.g., via a thioether bond with cysteine; phosphopantetheinylation, attachment of a 4'-phosphopantetheinyl group; and formation of a retinylidene Schiff base.

[0351] Examples of post-translational modifications by addition of chemical groups include acylation, e.g., O-acylation (ester), N-acylation (amide) or S-acylation (thioester); acetylation, e.g., attachment of an acetyl group to the N-terminus or lysine; formylation; alkylation, addition of an alkyl group such as methyl or ethyl; methylation, e.g., addition of a methyl group to lysine or arginine; amidation; butyration; gamma carboxylation; glycosylation, e.g., enzymatic attachment of a glycosyl group to arginine, asparagine, cysteine, hydroxylysine, serine, threonine, tyrosine or tryptophan; polysialylation, attachment of polysialic acid; malonylation; hydroxylation; iodination; bromination; citrulline. phosphorylation, e.g., the attachment of a phosphate group to serine, threonine or tyrosine (O-linked) or histidine (N-linked); adenylylation, e.g., the attachment of an adenylyl moiety to tyrosine (O-linked) or histidine or lysine (N-linked); propionation; the formation of pyroglutamate; S-glutathionylation; sumoylation; S-nitrosylation; succinylation, e.g., the attachment of a succinyl group to lysine; selenoylation, the incorporation of selenium; and ubiquitination, the addition of a ubiquitin subunit (N-linked).

[0352] It is within the scope of the methods provided herein that the polypeptide is labeled with molecular label.Molecular label can be a modification to the polypeptide that facilitates the detection of the polypeptide in the methods provided herein.For example, label can be a modification to the polypeptide that changes the signal obtained when the conjugate is characterized.For example, label can interfere with the flow of ions through the nanopore.In this way, label can improve the sensitivity of the method.

[0353] In some embodiments, the polypeptide contains one or more cross-linked sections, e.g., CC bridges. In some embodiments, the polypeptide is not cross-linked before being characterized using the disclosed methods.

[0354] In some embodiments, the polypeptide contains sulfide-containing amino acids and therefore has the potential to form disulfide bonds. Typically, in such embodiments, the polypeptide is reduced using a reagent such as DTT (dithiothreitol) or TCEP (tris(2-carboxyethyl)phosphine) before being characterized using the disclosed methods.

[0355] In some embodiments, the polypeptide is a full-length protein or a naturally occurring polypeptide. In some embodiments, the protein or naturally occurring polypeptide is fragmented before being conjugated to the polynucleotide. In some embodiments, the protein or polypeptide is chemically or enzymatically fragmented. In some embodiments, the polypeptide or polypeptide fragment may be conjugated to form a longer target polypeptide.

[0356] The polypeptide may be of any suitable length. In some embodiments, the polypeptide has a length of about 2 to about 300 peptide units. In some embodiments, the polypeptide has a length of about 2 to about 100 peptide units, such as about 2 to about 50 peptide units, such as about 3 to about 50 peptide units, such as about 5 to about 25 peptide units, such as about 7 to about 16 peptide units, such as about 9 to about 12 peptide units.

[0357] Any number of polypeptides can be characterized with the disclosed methods. For example, the methods can involve the characterization of 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 50, 100, or more polynucleotides. When more than one polypeptide is used, they can be different polypeptides or two or more instances of the same polypeptide.

[0358] It will thus be appreciated that measurements taken in the disclosed methods are typically characteristic of one or more properties of a polypeptide selected from (i) the length of the polypeptide, (ii) the identity of the polypeptide, (iii) the sequence of the polypeptide, (iv) the secondary structure of the polypeptide, and (v) whether the polypeptide is modified. In typical embodiments, the measurements are characteristic of the sequence of the polypeptide, or whether the polypeptide is modified, for example, by one or more post-translational modifications. In some embodiments, the measurements are characteristic of the sequence of the polypeptide.

[0359] In some embodiments, the polypeptide is in a relaxed form. In some embodiments, the polypeptide is maintained in a linearized form. Maintaining the polypeptide in a linearized form may facilitate characterization of the polypeptide on a residue-by-residue basis, since "bunching" of the polypeptide within the nanopore is prevented.

[0360] The polypeptide may be maintained in a linearized form using any suitable means.

[0361] For example, if the polypeptide is charged, applying a voltage can cause the polypeptide to be held in a linearized form.

[0362] If the polypeptide is uncharged or only slightly charged, the charge can be changed or controlled by adjusting the pH. For example, a high pH can be used to increase the relative negative charge of the polypeptide, thereby holding it in a linearized form. Increasing the negative charge of the polypeptide can hold it in a linearized form, for example, under a positive voltage. Alternatively, a low pH can be used to increase the relative positive charge of the polypeptide, thereby holding it in a linearized form. Increasing the positive charge of the polypeptide can hold it in a linearized form, for example, under a negative voltage. In the disclosed method, a polynucleotide handling protein is used to control the movement of the polynucleotide relative to the nanopore. Since polynucleotides are typically negatively charged, it is generally most suitable to increase the linearization of the polypeptide, as with polynucleotides, by increasing the pH, thereby making the polypeptide more negatively charged. In this way, the conjugate holds an overall negative charge and can be easily moved under an applied voltage.

[0363] The polypeptide can be maintained in a linearized form by using suitable denaturing conditions. Suitable denaturing conditions include, for example, the presence of a suitable concentration of a denaturing agent, such as guanidine HCl and / or urea. The concentration of such a denaturing agent used in the disclosed method depends on the target polypeptide to be characterized in the method and can be easily selected by a person skilled in the art.

[0364] The polypeptide can be maintained in a linearized form by using a suitable detergent. Suitable detergents for use in the disclosed methods include SDS (sodium dodecyl sulfate).

[0365] The polypeptide can be maintained in a linearized form by performing the disclosed methods at elevated temperatures: increasing the temperature overcomes the intrachain bonds and allows the polypeptide to assume a linearized form.

[0366] Polypeptides can be held in a linearized form by performing the disclosed methods under strong electroosmotic forces. Such forces can be provided by using asymmetric salt conditions and / or by providing a suitable charge on the nanopore channel. The charge of the protein nanopore channel can be altered, for example, by mutagenesis. Varying the charge of the nanopore is within the capabilities of one of ordinary skill in the art. Varying the charge of the nanopore generates a strong electroosmotic force from the unbalanced flow of cations and anions through the nanopore when a potential is applied across the nanopore.

[0367] A polypeptide can be held in a linearized form by passing through a structure such as an array of nanopillars, a nanoslit, or across a nanogap, in some embodiments, the physical constraints of such structures can force the polypeptide to adopt a linearized form.

[0368] Conjugate formation As described in more detail herein, a conjugate comprises a polynucleotide conjugated to a target polypeptide.

[0369] The target polypeptide can be conjugated to the polynucleotide at any suitable position. For example, the polypeptide can be conjugated to the polynucleotide at the N-terminus or C-terminus of the polypeptide. The polypeptide can be conjugated to the polynucleotide through the side group of a residue (e.g., an amino acid residue) in the polypeptide.

[0370] In some embodiments, a target polypeptide has naturally occurring reactive functional groups that can be used to facilitate conjugation to a polynucleotide, for example, cysteine ​​residues can be used to form disulfide bonds to a polynucleotide or modifying groups thereon.

[0371] In some embodiments, the target polypeptide is modified to facilitate its conjugation to a polynucleotide. For example, in some embodiments, the polypeptide is modified by attaching a moiety that includes a reactive functional group for attachment to a polynucleotide. For example, in some embodiments, the polypeptide can be extended at the N-terminus or C-terminus by one or more residues (e.g., amino acid residues) that include one or more reactive functional groups for reacting with corresponding reactive functional groups on a polynucleotide. For example, in some embodiments, the polypeptide can be extended at the N-terminus and / or C-terminus by one or more cysteine ​​residues. Such residues can be used for attachment to the polynucleotide portion of the conjugate, for example, by maleimide chemistry (e.g., reaction of cysteine ​​with an azido-maleimide compound such as azido-[Pol]-maleimide, where [Pol] is typically a short-chain polymer such as PEG, e.g., PEG2, PEG3, or PEG4; followed by coupling to an appropriately functionalized polynucleotide, e.g., a polynucleotide bearing a BCN group for reaction with an azide). Such chemistries are described in Example 2. For the avoidance of doubt, where a polypeptide comprises suitable naturally occurring residues at the N-terminus and / or C-terminus (e.g. naturally occurring cysteine ​​residues at the N-terminus and / or C-terminus), such residue(s) may be used for attachment to the polynucleotide.

[0372] In some embodiments, residues in the target polypeptide are modified to facilitate attachment of the target polypeptide to a polynucleotide. In some embodiments, residues (e.g., amino acid residues) in the polypeptide are chemically modified for attachment to a polynucleotide. In some embodiments, residues (e.g., amino acid residues) in the polypeptide are enzymatically modified for attachment to a polynucleotide.

[0373] The conjugation chemistry between the polynucleotide and the polypeptide in the conjugate is not particularly limited. Any suitable combination of reactive functional groups can be used. Many suitable reactive groups and their chemical targets are known in the art. Some exemplary reactive groups and their corresponding targets include arylazides that can react with amines, carbodiimides that can react with amines and carboxyl groups, hydrazides that can react with carbohydrates, hydroxymethylphosphines that can react with amines, imide esters that can react with amines, isocyanates that can react with hydroxyl groups, carbonyls that can react with hydrazines, maleimides that can react with sulfhydryl groups, NHS-esters that can react with amines, PFP-esters that can react with amines, psoralens that can react with thymine, pyridyl disulfides that can react with sulfhydryl groups, vinyl sulfones that can react with sulfhydrylamines and hydroxyl groups, vinyl sulfonamides, and the like.

[0374] Other suitable chemistries for conjugating polypeptides to polynucleotides include click chemistry. Many suitable click chemistry reagents are known in the art. Suitable examples of click chemistry include, but are not limited to: (a) Copper(I)-catalyzed azide-alkyne cycloaddition (azide-alkyne Huisgen cycloaddition), (b) Strain-promoted azide-alkyne cycloaddition; [3+2] cycloaddition of alkenes and azides; inverse demand Diels-Alder reaction of alkenes and tetrazines; and photoclick reaction of alkenes and tetrazoles. (c) Copper-free variants of the 1,3 dipolar cycloaddition reaction in which an azide reacts with a strained alkyne, e.g., in a cyclooctane ring, e.g., in bicyclic [6.1.0]nonyne (BCN); (d) reaction of an oxygen nucleophile on one linker with an epoxide or aziridine reactive moiety on the other; and (e) Staudinger ligation, in which the alkyne moiety can be replaced by an aryl phosphine, resulting in a specific reaction with an azide to give an amide bond.

[0375] Any reactive group can be used to form the conjugate. Some suitable reactive groups include [1,4-bis[3-(2-pyridyldithio)propionamido]butane; 1,1 1-bis-maleimidotriethylene glycol; 3,3'-dithiodipropionic acid di(N-hydroxysuccinimide ester); ethylene glycol-bis(succinic acid N-hydroxysuccinimide ester); 4,4'-diisothiocyanatostilbene-2,2'-disulfonic acid disodium salt; bis[2-(4-azidosalicylamido)ethyl]disulfide; 3-(2-pyridyldithio)propionic acid N-hydroxysuccinimide ester; 4-maleimidobutyric acid N-hydroxysuccinimide ester; iodoacetic acid N-hydroxysuccinimide ester; S-acetylthioglycolic acid N-hydroxysuccinimide ester; azido-PEG-maleimide; and alkyne-PEG-maleimide. The reactive group may be any of those disclosed in WO 2010 / 086602, in particular Table 3 of that application.

[0376] In some embodiments, the reactive functional group is included in the polynucleotide and the targeting functional group is included in the polypeptide prior to the conjugation step. In other embodiments, the reactive functional group is included in the polypeptide and the targeting functional group is included in the polynucleotide prior to the conjugation step. In some embodiments, the reactive functional group is attached directly to the polypeptide. In some embodiments, the reactive functional group is attached to the polypeptide via a spacer. Any suitable spacer can be used. Suitable spacers include, for example, alkyl diamines such as ethyl diamine, and the like.

[0377] As is evident from the above discussion, in some embodiments, the conjugate comprises multiple polypeptide sections and / or multiple polynucleotide sections. For example, the conjugate may comprise a structure of the form ...-PNPNPN..., where P is a polypeptide and N is a polynucleotide. In such embodiments, the polynucleotide handling protein sequentially controls the N portion of the conjugate relative to the nanopore, thus sequentially controlling the movement of the P section relative to the nanopore, thus allowing sequential characterization of the P section. In such embodiments, multiple polynucleotides and polypeptides may be conjugated together by the same or different chemistries.

[0378] As described herein, the conjugate may include a leader. As described herein, any suitable leader may be used. In some embodiments, the leader is a polynucleotide. In embodiments where the leader is a polynucleotide, the leader may be the same type of polynucleotide as the polynucleotide used in the conjugate, or the leader may be a different type of polynucleotide. For example, the polynucleotide in the conjugate may be DNA and the leader may be RNA, or vice versa.

[0379] In some embodiments, the leader is a charged polymer, e.g., a negatively charged polymer. In some embodiments, the leader comprises a polymer such as PEG or a polysaccharide. In such embodiments, the leader may be 10-150 monomer units (e.g., ethylene glycol or saccharide units) long, such as 20-120, such as 30-100, such as 40-80, such as 50-70 monomer units (e.g., ethylene glycol or saccharide units) long.

[0380] The disclosed methods of characterizing the target polypeptides described herein include conjugating the polypeptide to a polynucleotide and controlling the translocation of the conjugate relative to the nanopore using a polynucleotide handling protein.

[0381] Any suitable polynucleotide can be used in the disclosed methods, such polynucleotides being further described herein in connection with methods of polynucleotide characterization.

[0382] Concatenation Preferably, the target analyte, where the analyte is a polynucleotide or a polypeptide, can be linked to a membrane comprising a pore in the method of the present invention described herein. The method can include linking the analyte to a membrane comprising a pore. The polynucleotide is preferably linked to the membrane using one or more anchors. The polynucleotide can be linked to the membrane using any known method.

[0383] Each anchor comprises a group that links (or binds) to the analyte and a group that links (or binds) to the membrane. Each anchor can be covalently bound (or bound) to the analyte and / or the membrane.

[0384] Where the analyte is a polynucleotide, a Y adaptor and / or a hairpin loop adaptor(s) (both such adaptors are known in the art) may be used, and the polynucleotide is preferably linked to the membrane using the adaptors.

[0385] The analyte may be linked to the membrane using any number of anchors, such as 2, 3, 4 or more anchors, etc. For example, the analyte may be linked to the membrane using two anchors, each of which separately links (or binds) to both the analyte and the membrane.

[0386] The one or more anchors may comprise one or more helicases and / or one or more molecular brakes.

[0387] When the membrane is an amphiphilic layer such as a copolymer membrane or lipid bilayer, the one or more anchors preferably comprise a polypeptide anchor present in the membrane and / or a hydrophobic anchor present in the membrane.The hydrophobic anchor is preferably a lipid, a fatty acid, a sterol, a carbon nanotube, a polypeptide, a protein or an amino acid, such as cholesterol, palmitate, tocopherol, or charge-neutralized alkyl-phosphorothioate.In a preferred embodiment, the one or more anchors are not pores.

[0388] Components of the membrane, such as amphiphilic molecules, copolymers, or lipids, can be chemically modified or functionalized to form one or more anchors. Suitable chemical modifications and methods of functionalizing the components of the membrane are discussed in more detail below. Any percentage of the membrane components can be functionalized, for example, at least 0.01%, at least 0.1%, at least 1%, at least 10%, at least 25%, at least 50%, or 100%.

[0389] The analyte may be directly linked to the membrane. The one or more anchors used to link the analyte to the membrane preferably include a linker. The one or more anchors may include one or more, e.g., two, three, four, or more linkers. One linker may be used to link more than one, such as two, three, four, or more analytes to the membrane.

[0390] Preferred linkers include, but are not limited to, polymers such as polynucleotides, polyethylene glycol (PEG), polysaccharides, and polypeptides. These linkers can be linear, branched, or cyclic. For example, the linker can be a cyclic polynucleotide. The polynucleotide can hybridize to a complementary sequence on the cyclic polynucleotide linker.

[0391] One or more of the anchors or one or more of the linkers may contain a moiety that can be cleaved or degraded, such as a restriction site or a photolabile group.

[0392] Functionalized linkers and the methods by which they can be attached to molecules are known in the art. For example, a linker functionalized with a maleimide group will react with and bind to a cysteine ​​residue in a protein. In the context of the present invention, the protein may be present in a membrane or may be used to link (or bind) to an analyte. This is discussed in more detail below.

[0393] Cross-linking of analytes can be avoided using a "lock and key" arrangement: only one end of each linker can react together to form a longer linker, the other end of the linker reacting with a polynucleotide or a membrane, respectively. Such linkers are described in International Application No. PCT / GB10 / 000132 (published as WO2010 / 086602).

[0394] In the sequencing embodiment discussed herein, the use of a linker is preferred. If the polynucleotide or polypeptide is permanently linked directly to the membrane, in the sense that it does not detach when interacting with the pore (i.e., does not detach in step (b) or (e)), some sequence data will be lost because the sequencing run cannot continue to the end of the analyte due to the distance between the membrane and the pore. If a linker is used, the polynucleotide or polypeptide can be processed to completion.

[0395] The linkage can be permanent or stable, in other words, the linkage can be such that the analyte remains linked to the membrane when it interacts with the pore.

[0396] The linkage can be temporary, that is to say, the linkage can be such that the polynucleotide can dissociate from the membrane upon interaction with the pore.

[0397] For certain applications, such as aptamer detection, linkages of a temporary nature are preferred. For example, if a permanent or stable linker is attached directly to either the 5' or 3' end of the polynucleotide target analyte and the linker is shorter than the distance between the membrane and the channel of the transmembrane pore, some sequence data will be lost since the sequencing run cannot continue to the end of the polynucleotide. If the attachment is temporary, the polynucleotide can be processed to completion once the attached end is randomly separated from the membrane. Chemical groups that form permanent / stable or temporary bonds are discussed in more detail below. Polynucleotides can be temporarily attached to the amphiphilic layer or triblock copolymer membrane using cholesterol or fatty acyl chains. Any fatty acyl chain with a length of 6 to 30 carbon atoms, such as hexadecanoic acid, can be used.

[0398] In a preferred embodiment, a target analyte, such as a polypeptide or polynucleotide, is linked to an amphiphilic layer, such as a triblock copolymer membrane or lipid bilayer. Linking of nucleic acids to synthetic lipid bilayers has been previously performed using a variety of different tethering methods, which are summarized in Table 4 below. [Table 4]

[0399] The synthetic polynucleotides and / or linkers can be functionalized during the synthesis reaction using modified phosphoramidites, which are easily adaptable for direct addition of suitable anchor groups such as cholesterol, tocopherol, palmitic acid, thiol, lipid, and biotin groups. These different linking chemistries provide a set of options for binding to polynucleotides. The different modified groups bind to polynucleotides in slightly different ways, and the binding is not necessarily permanent, thus providing different residence times for the polynucleotides in the membrane. The advantages of temporary linkages have been discussed above.

[0400] Attachment of polynucleotides to linkers or functionalized membranes can also be accomplished by several other means, provided that complementary reactive or anchor groups can be added to the polynucleotide. Addition of reactive groups to either end of a polynucleotide has been reported previously. Using T4 polynucleotide kinase and ATPγS, a thiol group can be added to the 5′ of ssDNA or dsDNA (Grant, GP and PZQin (2007). “A facile method for attaching nitroxide spin labels at the 5′ terminus of nucleic acids.” Nucleic Acids Res 35(10):e77). Using T4 polynucleotide kinase and γ-[2-azidoethyl]-ATP or γ-[6-azidohexyl]-ATP, an azide group can be added to the 5′-phosphate of dsDNA or ssDNA. Using thiol or click chemistry, tethers containing either thiols, iodoacetamide OPSS, or maleimide groups (reactive towards thiols), or DIBO (dibenzocyclooxtyne), or alkyne groups (reactive towards azides) can be covalently attached to polynucleotides. A more diverse selection of chemical groups such as biotin, thiols, and fluorophores can be added using terminal transferase to incorporate modified oligonucleotides 3' into ssDNA (Kumar, A., P. Tchen, et al. (1988). "Nonradioactive labeling of synthetic oligonucleotide probes with terminal deoxynucleotidyl transferase." Anal Biochem 169(2):376-82). For any other polynucleotide, streptavidin / biotin and / or streptavidin / desthiobiotin linkages can be used.The following examples illustrate how polynucleotides can be linked to membranes using streptavidin / biotin and streptavidin / desthiobiotin. It may also be possible that anchors can be added directly to polynucleotides using terminal transferase with appropriately modified nucleotides (e.g., cholesterol or palmitate).

[0401] The one or more anchors preferably link the polynucleotide target analyte to the membrane via hybridization. Hybridization at the one or more anchors allows for binding in a temporary manner as described above. Hybridization can be at any part of the one or more anchors, such as between the one or more anchors and the polynucleotide, within the one or more anchors, or between the one or more anchors and the membrane. For example, a linker can include two or more polynucleotides, e.g., three, four, or five polynucleotides, hybridized together. The one or more anchors can hybridize to the polynucleotide. The one or more anchors can hybridize directly to the polynucleotide, or directly to a Y adaptor and / or leader sequence attached to the polynucleotide, or directly to a hairpin loop adaptor attached to the polynucleotide (as discussed below). Alternatively, the one or more anchors can hybridize to one or more, e.g., two or three, intermediate polynucleotides (or "sprints") that are hybridized to the polynucleotide, to a Y adaptor and / or leader sequence attached to the polynucleotide, or to a hairpin loop adaptor attached to the polynucleotide (as discussed below).

[0402] One or more anchors may comprise single-stranded or double-stranded polynucleotides. A portion of the anchor may be ligated to a single-stranded or double-stranded polynucleotide. Ligation of a short piece of ssDNA has been reported using T4 RNA ligase I (Troutt, AB, MG McHeyzer-Williams, et al. (1992). "Ligation-anchored PCR: a simple amplification technique with single-sided specificity." Proc Natl Acad Sci USA 89(20):9823-5). Alternatively, either a single-stranded or double-stranded polynucleotide can be ligated to a double-stranded polynucleotide, and the two strands can then be separated by thermal or chemical denaturation. A piece of single-stranded polynucleotide can be added to a double-stranded polynucleotide at one or both ends of the double strand, or a double-stranded polynucleotide can be added to one or both ends. For the addition of a single-stranded polynucleotide to a double-stranded polynucleotide, this can be accomplished using T4 RNA ligase I for ligation to another region of the single-stranded polynucleotide. For the addition of a double-stranded polynucleotide to a double-stranded polynucleotide, the ligation can then be "blunt-ended" with complementary 3' dA / dT tails on each polynucleotide and the added polynucleotide (as is routinely done for many sample preparation applications to prevent concatemer or dimer formation), or using "sticky ends" generated by restriction digestion of the polynucleotide and ligation of compatible adaptors. Then, when the duplex is melted, each single strand will have either a 5' or 3' modification if a single-stranded polynucleotide was used for the ligation, or a modification at the 5' end, 3' end, or both if a double-stranded polynucleotide was used for the ligation.

[0403] If the polynucleotide is a synthetic strand, one or more anchors may be incorporated during chemical synthesis of the polynucleotide. For example, the polynucleotide may be synthesized using a primer with a reactive group attached. The adenylated polynucleotide is an intermediate in the ligation reaction, in which adenosine monophosphate is attached to the 5'-phosphate of the polynucleotide. Various kits are available for the generation of this intermediate, such as NEB's 5'DNA Adenylation Kit. Replacing modified nucleotide triphosphates with ATP during the reaction may allow the addition of reactive groups (e.g., thiol, amine, biotin, azide, etc.) to the 5' of the polynucleotide. It may also be possible that the anchor can be added directly to the polynucleotide using the 5'DNA Adenylation Kit with a suitably modified nucleotide (e.g., cholesterol or palmitate).

[0404] A common technique for the amplification of a portion of genomic DNA is to use the polymerase chain reaction (PCR). Here, two synthetic oligonucleotide primers can be used to generate several copies of the same section of DNA, with each copy 5' of each strand in the duplex being a synthetic polynucleotide. By using a polymerase, single or multiple nucleotides can be added to the 3' end of single or double stranded DNA. Examples of polymerases that can be used include, but are not limited to, terminal transferase, Klenow, and E. coli Poly(A) polymerase. Anchors such as cholesterol, thiol, amine, azide, biotin, or lipids can be incorporated into the double stranded polynucleotide by replacing modified nucleotide triphosphates with ATP during the reaction. Thus, each copy of the amplified polynucleotide will contain an anchor.

[0405] Ideally, polynucleotides are bound to membranes without the need to functionalize polynucleotides. This can be achieved by binding one or more anchors, such as polynucleotide-binding proteins or chemical groups, to the membrane and allowing the one or more anchors to interact with polynucleotides, or by functionalizing the membrane. The one or more anchors can be bound to the membrane by any of the methods described herein. In particular, the one or more anchors can include one or more linkers, such as maleimide-functionalized linkers.

[0406] In this embodiment, the polynucleotide is typically RNA, DNA, PNA, TNA, or LNA and may be double-stranded or single-stranded. This embodiment is particularly suitable for genomic DNA polynucleotides.

[0407] The one or more anchors may comprise any group that couples to, binds to, or interacts with a single- or double-stranded polynucleotide, a specific sequence of nucleotides within a polynucleotide, or a pattern of modified nucleotides within a polynucleotide, or any other ligand present on a polynucleotide.

[0408] Binding proteins suitable for use in the anchor include, but are not limited to, E. coli single-stranded binding protein, P5 single-stranded binding protein, T4 gp32 single-stranded binding protein, TOPO V dsDNA binding region, human histone proteins, E. coli HU DNA binding protein, and other archaeal, prokaryotic, or eukaryotic single- or double-stranded polynucleotide (or nucleic acid) binding proteins, including those listed below.

[0409] The specific nucleotide sequence may be a sequence recognized by a transcription factor, a ribosome, an endonuclease, a topoisomerase, or a replication initiation factor. The pattern of modified nucleotides may be a pattern of methylation or damage.

[0410] The one or more anchors may include any group that couples, binds, intercalates, or interacts with a polynucleotide. The group intercalates or interacts with a polynucleotide through electrostatic, hydrogen bonding, or van der Waals interactions. Such groups include lysine monomers, poly-lysine (which will interact with ssDNA or dsDNA), ethidium bromide (which will intercalate with dsDNA), universal bases or nucleotides (which can hybridize to any polynucleotide), and osmium complexes (which can react with methylated bases). A polynucleotide may thus be attached to a membrane using one or more universal nucleotides attached to the membrane. Each universal nucleotide may be attached to the membrane using one or more linkers. The universal nucleotide preferably comprises one of the following nucleobases: hypoxanthine, 4-nitroindole, 5-nitroindole, 6-nitroindole, formyl indole, 3-nitropyrrole, nitroimidazole, 4-nitropyrazole, 4-nitrobenzimidazole, 5-nitroindazole, 4-aminobenzimidazole, or phenyl (C6 aromatic ring).The universal nucleotide is more preferably the following nucleoside: 2'-deoxyinosine, inosine, 7-deaza-2'-deoxyinosine, 7-deaza-inosine, 2-aza-deoxyinosine, 2-aza-inosine, 2-O'methylinosine, 4-nitroindole 2'-deoxyribonucleoside, 4-nitroindole ribonucleoside, 5-nitroindole 2'deoxyribonucleoside, 5-nitroindole 2'deoxyribonucleoside, 6-nitroindole 2'deoxyribonucleoside, 6-nitroindole ribonucleoside, 3-nitropyrrole 2'deoxyribonucleoside, 3-nitropyrrole ribonucleoside, acyclic sugar analogues of hypoxanthine, nitroimidazole 2'deoxyribonucleoside, The universal nucleotide may comprise one of the following: 2'-deoxyribonucleoside, nitroimidazole ribonucleoside, 4-nitropyrazole 2' deoxyribonucleoside, 4-nitropyrazole ribonucleoside, 4-nitrobenzimidazole 2' deoxyribonucleoside, 4-nitrobenzimidazole ribonucleoside, 5-nitroindazole 2' deoxyribonucleoside, 5-nitroindazole ribonucleoside, 4-aminobenzimidazole 2' deoxyribonucleoside, 4-aminobenzimidazole ribonucleoside, phenyl C-ribonucleoside, phenyl C-2'-deoxyribosyl nucleoside, 2'-deoxynebularine, 2'-deoxyisoguanosine, K-2'-deoxyribose, P-2'-deoxyribose, and pyrrolidine. The universal nucleotide more preferably comprises 2'-deoxyinosine. The universal nucleotide is more preferably IMP or dIMP. The universal nucleotide is most preferably dPMP (2'-deoxy-P-nucleoside monophosphate) or dKMP (N6-methoxy-2,6-diaminopurine monophosphate).

[0411] One or more anchors may couple (or bind) to a polynucleotide via Hoogsteen hydrogen bonds (two nucleobases are held together by hydrogen bonds) or reversed Hoogsteen hydrogen bonds (one nucleobase is rotated 180° relative to the other nucleobase). For example, one or more anchors may include one or more nucleotides, one or more oligonucleotides, or one or more polynucleotides that form Hoogsteen or reversed Hoogsteen hydrogen bonds with the polynucleotide. These types of hydrogen bonds allow a third polynucleotide strand to wrap around the double-stranded helix and form a triplex. One or more anchors may couple (or bind) to a double-stranded polynucleotide by forming a triplex with the double-stranded structure.

[0412] In this embodiment, at least 1%, at least 10%, at least 25%, at least 50%, or 100% of the membrane components may be functionalized.

[0413] When one or more anchors comprise proteins, they may be able to directly anchor in the membrane without further functionalization, for example, if they already have an external hydrophobic region that is compatible with the membrane. Examples of such proteins include, but are not limited to, transmembrane proteins, intramembrane proteins, and membrane proteins. Alternatively, proteins may be expressed with genetically fused hydrophobic regions that are compatible with the membrane. Such hydrophobic protein regions are known in the art.

[0414] The one or more anchors are preferably mixed with the polynucleotides prior to contact with the membrane, although the one or more anchors can be contacted with the membrane and then contacted with the polynucleotides.

[0415] In another aspect, the polynucleotide may be functionalized using the methods described above so that it can be recognized by a specific binding group. Specifically, the analyte may be functionalized with a ligand such as biotin (for binding to streptavidin), amylose (for binding to maltose binding protein or fusion proteins), Ni-NTA (for binding to poly-histidine or poly-histidine tagged proteins), or a peptide (such as an antigen).

[0416] According to a preferred embodiment, one or more anchors can be used to bind a polynucleotide to a membrane when the polynucleotide is bound to a leader sequence that preferentially enters the pore. Leader sequences are discussed in more detail below. Preferably, the polynucleotide is bound (e.g., ligated) to a leader sequence that preferentially enters the pore. Such a leader sequence may comprise a homopolymer polynucleotide or an abasic region. The leader sequence is typically designed to hybridize to one or more anchors, either directly or via one or more intermediate polynucleotides (or splints). In such an example, the one or more anchors typically comprise a polynucleotide sequence complementary to a sequence in the leader sequence or a sequence in one or more intermediate polynucleotides (or splints). In such an example, the one or more splints typically comprise a polynucleotide sequence complementary to a sequence in the leader sequence.

[0417] An example of a molecule used for chemical conjugation is EDC (1-ethyl-3-[3-dimethylaminopropyl]carbodiimide hydrochloride). Reactive groups can also be added to the 5' end of polynucleotides using commercially available kits (Thermo Pierce, part number 22980). Suitable methods include, but are not limited to, temporary affinity binding using histidine residues and Ni-NTA, and stronger covalent binding through reactive cysteine, lysine, or unnatural amino acids.

[0418] kit In addition, the following is provided: a pore according to the invention; - a polynucleotide binding protein or a polypeptide handling enzyme.

[0419] In some embodiments, the pore is modified according to a variant described herein to alter the ability of the monomer to interact with an analyte. Most preferably, one or more constrictions within the pore are modified according to a variant described herein, thereby altering the ability of the one or more constrictions to interact with an analyte.

[0420] The kit may be configured for use with an algorithm, also provided herein, adapted to run on a computer system. The algorithm may be adapted to detect information characteristic of the polypeptide (e.g., characteristic of the sequence of the polypeptide and / or whether the polypeptide is modified) and selectively process a signal obtained when a conjugate comprising the polypeptide conjugated to a polynucleotide moves relative to the nanopore. Also provided is a system comprising a computing means configured to detect information characteristic of the polypeptide (e.g., characteristic of the sequence of the polypeptide and / or whether the polypeptide is modified) and selectively process a signal obtained when a conjugate comprising the polypeptide conjugated to a polynucleotide moves relative to the nanopore. In some embodiments, the system comprises a receiving means for receiving data from the detection of the polypeptide, a processing means for processing a signal obtained when the conjugate moves relative to the nanopore, and an output means for outputting the characterization information thus obtained.

[0421] Although specific embodiments of the method, specific configurations, and materials and / or molecules according to the present invention have been discussed herein, it is understood that various changes or modifications in form and details may be made without departing from the scope and spirit of the present invention. The foregoing embodiments and the following examples are provided for illustrative purposes only and should not be considered as limiting the application. This application is limited only by the claims. EXAMPLES

[0422] Examples - Materials and Methods The experimental results are detailed in the figure legends.

[0423] Calculation tools Pairwise sequence alignments (see especially FIG. 1) were performed using the publicly available software, Clustalx (http: / / www.clustal.org / clustal2 / ).

[0424] A structural model of CytK (see especially FIG. 2) was generated using Modeller software (https: / / salilab.org / modeller / ).

[0425] Pore ​​radial profiles (see especially FIG. 4) were generated using publicly available software, HOLE (http: / / www.holeprogram.org / ).

[0426] E. coli pore formation See especially Figures 5 to 16 and their legends.

[0427] DNA encoding the mature form of CytK protein was synthesized by GenScript USA Inc. and cloned into a pT7 vector containing an ampicillin resistance gene. The DNA concentration was adjusted to 400 ng / μL.

[0428] Plasmid DNA was thawed at room temperature and mixed by slowly pipetting up and down. Chemically competent BL21(DE3) E. coli cells were thawed on ice. 1 μl of DNA at 400 ng / μl was added to the cells and mixed by slowly pipetting up and down. This was then left on ice for 25 minutes before the cells were heat shocked at 42°C for 45 seconds. The cells were then left on ice for 2 minutes. 250 μl of SOC (Sigma, S1797) medium pre-warmed to 37°C was added to the cells and left at 37°C for 1 hour with shaking. Half of the cells were then plated on a large LB agar plate containing 50 μg / ml ampicillin and then incubated at 37°C overnight.

[0429] A single colony of transformed BL21(DE3) cells was inoculated with 100 μg / ml carbenicillin in 100 ml LB medium. This starter culture was incubated overnight at 37° C. and 250 rpm in a 500 ml flask. 500 ml LB medium containing 100 μg / ml carbenicillin was added to a 2.5 1 flask. This was then added to 5 ml of the starter culture (diluted 1:100) and the cells were split at 37° C. and 250 rpm until an OD of 0.6 was reached. Once an OD of 0.6 was reached, the incubator temperature was reduced to 18° C. and the cells were induced with 0.2 mM IPTG (final concentration in the medium). The cells were incubated overnight at 18° C. and 250 rpm. Finally, the cells were harvested by spinning at 6000 g for 30 min at 4° C.

[0430] The cell paste was weighed to calculate the appropriate volume of functional lysis buffer to be prepared (cells are resuspended in 100 ml lysis buffer per 10 g paste). The required amount of functional lysis buffer was prepared by adding Benzonase (10 μl / 100 ml) and 4 tablets of protease inhibitor cocktail without EDTA to a buffer containing 50 mM Tris / HCl, 0.5 M NaCl, pH 8.0 at room temperature. The cells were resuspended in functional lysis buffer and mixed for 1 hour at room temperature using a magnetic stirrer. The cell suspension was frozen at -80°C and thawed at room temperature. DDM was added to the cell suspension at a final concentration of 1% and mixed again for 1 hour at 37°C using a magnetic stirrer. The cell extract was transferred to a 40 ml Beckman tube and spun at 50,000 g rpm for 30 minutes at room temperature. The supernatant was then filtered through a 0.22 μM PES syringe filter.

[0431] The supernatant was then loaded onto 2 x 5 mL His Trap FF columns (Fisher, 10571680). The column was washed with 50 mM Tris, 0.5 M NaCl, 5 mM imidazole, 0.1% DDM, pH 8.0 (mobile phase A) until a stable baseline was maintained for 10 column volumes (CV). The column was then washed with 50 mM Tris, 2 M NaCl, 5 mM imidazole, 0.1% DDM, pH 8.0, and then returned to 150 mM buffer. Elution was performed with a gradient of 0-100% over 20 CV with 0.5 M imidazole, mobile phase B contained 50 mM Tris, 0.5 M NaCl, 0.5 M imidazole, 0.1% DDM, pH 8.0.

[0432] The fractions of interest from HisTrap purification were identified by SDS-PAGE. The peaks were pooled and then concentrated to approximately 1 ml using a 50 kDa MWCO (Millipore, UFC905024). The concentrated remaining supernatant was subjected to gel filtration on 320 ml of Superdex200 (Fisher, 11390342) in 50 mM Tris, 0.25 M NaCl, 0.1% DDM, pH 8.0. The fractions identified as containing CytK were collected and pooled. After this, the pooled supernatant was diluted five times with 50 mM Tris / HCl, 0.1% DDM, pH 9.0. It was then loaded onto a POROS HQ10 column pre-equilibrated with 50 mM Tris / HCl, 0.1% DDM, pH 9.0. Prior to starting the gradient, the column was washed with 50 mM Tris / HCl, 0.1% DDM, pH 9.0 until a stable baseline over 10 CV was achieved. A gradient of 50 mM Tris / HCl, 0.1% DDM, pH 9.0 to 100% 50 mM Tris / HCl, 0.1% DDM, 1 M NaCl, pH 9.0 was achieved over 25 CV. Fractions of interest from the POROS HQ10 purification were identified by SDS-PAGE, collected, and then assayed with electrophysiological recordings.

[0433] In vitro transcription-translation (IVTT) pore generation See especially Figures 5 to 16 and their legends.

[0434] For a single 25 μL reaction, prepare the following: [Table 5]

[0435] The above components were mixed and incubated at 30°C and 700 rpm on a Thermo Shaker. The samples were then spun down at 21,000g for 10 min at room temperature. The supernatant was carefully removed and discarded, while the pellet was resuspended in 1x Laemmli buffer (BioRad, 1610737) by pipetting up and down. The resuspension was then loaded onto a 7.5% Tris-HCl, pH 8.0 slab gel and electrophoresis at 55V was carried out overnight (16 h) in 1x TGS running buffer (Sigma, T7777). The gel was then dried under vacuum for 5 h at 50°C. X-ray film (Sigma, Z370371) was exposed to the gel for 2 h and developed using a combination of Devalex (Champion, 120102) and Fixaplus (Champion, 120202X) solutions in X-ray film developer. The film was then placed on top of the dried gel and the relevant bands were extracted using the film as a reference. Each extracted band was rehydrated in 100 μL of 50 mM Tris / HCl, 2 mM EDTA, pH 8.0 buffer and ground with a pestle until a homogenous slurry was obtained. The slurry was incubated overnight at room temperature, loaded onto a 0.45 μm CoStar column (Sigma, CLS8162) and spun at 21,000 g for 10 min. The supernatant was collected and assayed with electrophysiological recordings.

[0436] IV curve and irregular curve of DNA See especially Figures 5 to 12 and their legends.

[0437] Electrical measurements were taken from aHL and CytK wild-type and mutant nanopores inserted into a MinION flow cell. After inserting a single nanopore into the block copolymer membrane, 2 mL of buffer containing 25 mM calcium phosphate, 150 mM potassium ferrocyanide (II), 150 mM potassium ferricyanide (III), pH 8.0 was flowed through the system to remove any excess CytK nanopore. Ion current profiles through the nanopore were then obtained as the voltage was gradually increased in 25 mV steps every 30 seconds in both the negative and positive directions from (-)25 mV to (-)200 mV.

[0438] The Y-adapter was prepared by annealing the DNA oligonucleotides shown in Figure 16. A DNA motor (Dda helicase) was loaded into the adapter and closed. The material was then HPLC purified. The Y-adapter contains 30 C3 leader sections to facilitate capture by the nanopore and a side arm for tethering to the membrane.

[0439] The analyte used to assess DNA irregular curvature was a 3.6 kilobase ssDNA segment from the 3' end of the lambda genome. Analyte preparation, ligation of the analyte to the Y-adapter, SPRI-bead cleanup of the ligated analyte, and loading onto the MinION flow cell were performed using the Oxford Nanopore Technologies Q-SQK-LSK109 protocol.

[0440] Electrical measurements were acquired using a MinION Mk1b from Oxford Nanopore Technologies. Standard sequencing scripts were run at -180 mV for 1-6 h with a static flick every 5 min to remove elongation nanopore blocks. Raw data were collected into bulk FAST5 files using MinKNOW software (Oxford Nanopore Technologies).

[0441] Irregular curves in peptides See especially Figures 15 and 16 and their legends.

[0442] Exemplary current versus time traces for translocating a peptide through CytK wild type and mutants were obtained by using a conjugate containing two pieces of polynucleotide; a dsDNA Y-adapter (DNA1) and a polypeptide flanked by a dsDNA tail (DNA2). A polynucleotide handling protein on the cis side of the nanopore controls the movement of the conjugate by first unwinding DNA1, translocating it 5'-3' onto the ssDNA, then sliding across the polypeptide section and finally unwinding the DNA2 segment. As this construct moves from the cis to the trans side of the nanopore, the DNA and polypeptide sections can be visualized in a current versus time plot.

[0443] Y-adapters were prepared by annealing DNA oligonucleotides (Figure 13). A DNA motor (Dda helicase) was loaded into the adapter and closed. The subsequent material was HPLC purified. The Y-adapters contain 30 C3 leader sections to facilitate capture by the nanopore, and a side arm for tethering to the membrane. The DNA tails were made by annealing two DNA oligonucleotides, which also contain a side arm for tethering resulting in two tethering sites per construct to increase the efficiency of capture.

[0444] Polypeptide analytes were obtained containing an azide moiety at the N-terminus and immediately after the C-terminus using an ethyldiamine spacer along the peptide backbone. Each analyte was then conjugated to a Y-adapter and DNA tail via a copper-free click chemistry reaction between the azide and BCN (bicyclo[6.1.0]nonyne) moiety. Samples were purified using Agencourt AMPure XP (Beckman Coulter) beads, washed twice in 28% PEG 8K, 2.5 M NaCl, 25 mM Tris (pH 8.0) buffer, and eluted in 10 mM Tris-Cl, 50 mM NaCl (pH 8.0).

[0445] Electrical measurements were taken using a MinION Mk1b from Oxford Nanopore Technologies and a custom MinION flow cell with either a CytK wild type or a CytK mutant pore inserted. The flow cell was flushed with tether mix containing 50 nM DNA tether and SQB buffer lacking ATP. 800 μL of tether mix was added first for 5 min, then another 200 μL of mix was run through the system with the SpotON port open. DNA-peptide constructs were prepared at a concentration of 0.5 nM in buffers such as SQB from the Oxford Nanopore Technologies Sequencing Kit (SQK-LSK109) lacking ATP, and LB from the Oxford Nanopore Technologies Sequencing Kit (SQK-LSK109) lacking ATP, to obtain the "sequencing mix". 75 μL of sequencing mix was added to the MinION flow cell via the SpotON flow cell port. The mixture was incubated on the flow cell for 5–10 min to allow tethering of the construct and its subsequent capture by the nanopore. In the absence of ATP, the DNA motor remains stalled in the spacer region of the Y-adapter and the conjugate is captured into the nanopore but does not translocate. After incubation, 200 µL of SQB from the Oxford Nanopore Technologies sequencing kit (SQK-LSK109) was added and, in the presence of ATP, the captured DNA-peptide conjugate is translocated across the nanopore by the helicase, resulting in reproducible current footprints.

[0446] Standard sequencing scripts were run at -180 mV for 1–6 h with static flicks every 1 min to remove elongation nanopore blocks. Raw data were collected into bulk FAST5 files using MinKNOW software (Oxford Nanopore Technologies).

[0447] Array Enumeration Description SEQ ID NO:1 shows the wild-type amino acid sequence of the cytotoxin K monomer. SEQ ID NO:2 shows the polynucleotide sequence encoding the wild-type cytotoxin K monomer. SEQ ID NO: 3 shows the amino acid sequence of the exonuclease I enzyme from E. coli (EcoExo I). SEQ ID NO: 4 shows the amino acid sequence of the Exonuclease III enzyme from E. coli. This enzyme performs a 3'-5' distributed digestion of 5' monophosphate nucleosides from one strand of double-stranded DNA (dsDNA). The enzyme starts on the strand and requires a 5' overhang of approximately 4 nucleotides. SEQ ID NO: 5 shows the amino acid sequence of the RecJ enzyme from T. thermophilus (TthRecJ-cd). This enzyme performs processive digestion of 5' monophosphate nucleosides from ssDNA in the 5'-3' direction. Enzyme initiation on the chain requires at least four nucleotides. SEQ ID NO:6 shows the amino acid sequence of bacteriophage lambda exonuclease. The sequence is one of three identical subunits that assemble into a trimer. The enzyme performs highly processive digestion of nucleotides from one strand of dsDNA in the 5'-3' direction (http: / / www.neb.com / nebecomm / products / productM0262.asp). Enzyme initiation on the strand preferentially requires a 5' overhang of approximately 4 nucleotides with a 5' phosphate. SEQ ID NO: 7 shows the amino acid sequence of Phi29 DNA polymerase. SEQ ID NO: 8 shows the amino acid sequence of Hel308 Mbu. SEQ ID NO: 9 shows the amino acid sequence of Hel308 Csy. SEQ ID NO: 10 shows the amino acid sequence of Hel308 Tga. SEQ ID NO: 11 shows the amino acid sequence of Hel308 Mhu. SEQ ID NO: 12 shows the amino acid sequence of TraI Eco. SEQ ID NO: 13 shows the amino acid sequence of XPD Mbu. SEQ ID NO: 14 shows the amino acid sequence of Dda 1993. SEQ ID NO: 15 shows the amino acid sequence of Trwc Cba. SEQ ID NO: 16 shows the polynucleotide sequence encoding Phi29 DNA polymerase.

[0448] Sequence Listing SEQ ID NO:1 MQTTSQVVTDIGQNAKTHTSYNTFNNEQADNMTMSLKVTFIDDPSADKQIAVINTTGSFMKANPTLSDAPVDGYPIPGASVTLRYPSQYDIAMNLQDNTSRFFHVAPTNAVEETTVTSSVSYQLGGSIKASVTPSGPSGESGATGQVTWSDSV SYKQTSYKTNLIDQTNKHVKWNVFFNGYNNQNWGIYTRDSYHALYGNQLFMYSRTYPHETDARGNLVPMNDLPALTNSGFSPGMIAVVISEKDTEQSSIQVAYTKHADDYTLRPGFTFGTGNWVGNNIKDVDQKTFNKSFVLDWKNKKLVEKK SEQ ID NO:2 ATGCAAACCACCTCCCAAGTCGTCACGGACATCGGTCAGAACGCTAAAACCCATACCAGCTACAATACCTTCAATAACGAACAAGCAGATAACATGACCATGAGCCTGAAAGTCACGTTTATTGATGACCCGTCTGCAGATAAGCAGATTGCTGTTATCAACACCACGGGCTCATTCATGAAAGCAAATCCGACGCTGTCGGATGCTCCGGTGGACGGTTATCCGATTCCGGGTGCTAGTGTTACCCTGCGTTATCCGTCCCAGTACGATATCGCGATGAACCTGCAAGACAATACCAGTCGCTTTTTCCATGTGGCGCCGACGAATGCCGTTGAAGAAACCACGGTCACCAGCTCTGTGAGCTATCAGCTGGGCGGTAGCATCAAAGCCTCTGTGACCCCGTCTGGTCCGAGTGGTGAATCCGGTGCAACCGGTCAAGTCACGTGGTCAGATAGCGTGAGCTATAAACAGACCAGCTACAAGACGAACCTGATTGACCAAACCAATAAACACGTTAAGTGGAACGTCTTTTTCAATGGCTATAACAATCAGAACTGGGGTATCTACACCCGTGATAGTTATCATGCCCTGTACGGCAATCAACTGTTTATGTATTCCCGTACCTACCCGCACGAAACGGATGCGCGCGGTAACCTGGTGCCGATGAATGACCTGCCGGCCCTGACCAACTCAGGCTTCTCGCCGGGTATGATTGCAGTGGTTATCTCTGAAAAAGATACCGAACAGAGTTCCATTCAAGTTGCGTATACCAAGCATGCCGATGACTACACGCTGCGTCCGGGTTTTACCTTCGGTACGGGTAATTGGGTTGGTAACAATATCAAAGATGTCGACCAGAAAACCTTCAATAAATCGTTCGTGCTGGACTGGAAAAATAAGAAACTGGTGGAAAAGAAATAATGA Sequence number 3

Table 6

Table 10-1

Table 10-2

Table 10-3

[0449] The following are numbered aspects of the present invention. 1. A mutant cytotoxin K monomer comprising a variant of the amino acid sequence of SEQ ID NO:1, wherein the monomer is capable of forming a pore; A mutant cytotoxin K monomer, wherein the variant comprises one or more modifications at one or more positions within the region of SEQ ID NO:1 between about S100 and about K170 that alter the ability of the monomer to interact with an analyte. 2. The monomer of embodiment 1, wherein the variant has at least 70% identity with the amino acid sequence of SEQ ID NO:1. 3. The monomer of aspect 1 or 2, wherein the one or more modifications each independently (a) change the size of the amino acid residue at the modified position, (b) change the net charge of the amino acid residue at the modified position, (c) change the hydrogen bonding properties of the amino acid residue at the modified position, (d) introduce or remove one or more chemical groups that interact through a delocalized electron π-system into or from the amino acid residue at the modified position, and / or (e) change the structure of the amino acid residue at the modified position. 4. The monomer of any one of the preceding aspects, wherein the monomer is capable of forming a pore having a solvent-accessible channel from a first opening to a second opening of the pore, the solvent-accessible channel comprising at least one constriction, and wherein one or more modifications are made to amino acids within the constriction. 5. The monomer of embodiment 4, wherein the modification alters the interaction of the analyte with the constriction as the analyte translocates through the pore. 6. The monomer of aspect 4 or 5, wherein the one or more modifications (a) change the size of the constriction, (b) change the net charge of the constriction, (c) change the hydrogen bonding properties of amino acid residues within the constriction, (d) introduce or remove one or more chemical groups into or from the constriction that interact via a delocalized electron π-system, and / or (e) change the structure of the constriction. 7. The monomer of any one of the preceding aspects, wherein the variant comprises one or more modifications at one or more positions within the region of SEQ ID NO:1 between about V111 and about T158. 8. The monomer of any one of the preceding aspects, wherein the variant comprises one or more modifications within the region of SEQ ID NO:1 between about V111 and about S131, and / or between about S135 and about T158. 9. The monomer of any one of the preceding aspects, wherein the variant comprises one or more modifications within a region of SEQ ID NO:1 between about S119 and about G126, preferably between S121 and G125, and / or between about A143 and about S150, preferably between T144 and T148. 10. The monomer of any one of the preceding aspects, wherein the variant comprises one or more modifications within a region of SEQ ID NO:1 between about G126 and about V132, preferably between S127 and S131, and / or between about P137 and about A143, preferably between S138 and G142. 11. The monomer according to any one of the preceding aspects, wherein the variant comprises one or more modifications within a region of SEQ ID NO:1 between about N109 and about T117, preferably between V111 and T115, and / or between about S152 and about Y160, preferably between S154 and T158. 12. The monomer of any one of the preceding aspects, comprising a modification at any one of the following positions in SEQ ID NO:1: E113, T115, T117, S119, S121, Q123, G125, S127, K129, S131, V132, T133, P134, S135, G136, P137, S138, E140, G142, T144, Q146, T148, S150, S152, S154, and K156. 13. The monomer of any one of the preceding aspects, wherein the variant comprises, independently, one or more amino acid substitutions, additions, and / or deletions at said one or more positions. 14. The monomer of any one of the preceding aspects, wherein the variant comprises one or more amino acid substitutions, wherein the substituted amino acid(s) in the variant is selected from aspartate, glutamate, serine, threonine, asparagine, glutamine, glycine, alanine, valine, leucine, isoleucine, cysteine, arginine, lysine, and phenylalanine. 15. E113S / T / N / Q / G / A / V / L / I / C / R / K / F / Y T115S / N / Q / G / A / V / L / I / C / R / K / F T117S / N / Q / G / A / V / L / I / C / R / K / F S119T / N / Q / G / A / V / L / I / C / R / K / F S121T / N / Q / G / A / V / L / I / C / R / K / F Q123S / T / N / G / A / V / L / I / C / R / K / F / M / Y G125S / T / N / Q / A / V / L / I / C / R / K / F S127T / N / Q / G / A / V / L / I / C / R / K / F K129S / T / N / Q / G / A / V / L / I / C / R / F / Y S131T / N / Q / G / A / V / L / I / C / R / K / F V132S / T / N / Q / G / A / L / I / C / R / K / F T133S / N / Q / G / A / V / L / I / C / R / K / F P134S / T / N / Q / G / A / V / L / I / C / R / K / F S135T / N / Q / G / A / V / L / I / C / R / K / F G136S / T / N / Q / A / V / L / I / C / R / K / F P137S / T / N / Q / G / A / V / L / I / C / R / K / F S138T / N / Q / G / A / V / L / I / C / R / K / F E140S / T / N / Q / G / A / V / L / I / C / R / K / F G142S / T / N / Q / A / V / L / I / C / R / K / F T144S / N / Q / G / A / V / L / I / C / R / K / F Q146S / T / N / G / A / V / L / I / C / R / K / F / M / Y T148S / N / Q / G / A / V / L / I / C / R / K / F S150T / N / Q / G / A / V / L / I / C / R / K / F S152T / N / Q / G / A / V / L / I / C / R / K / F S154T / N / Q / G / A / V / L / I / C / R / K / F, and The monomer of any one of the preceding aspects, comprising one or more modifications selected from K156S / T / N / Q / G / A / V / L / I / C / R / F. 16. The monomer of any one of the preceding aspects, comprising a modification at one or more of E113, Q123, K129, E140, Q146, and K156. 17. A monomer according to any one of the preceding aspects, comprising a modification in Q123 and / or Q146. 18. The monomer according to any one of the preceding aspects, comprising a modification at K129 and / or E140. 19. The monomer according to any one of the preceding aspects, comprising a modification at E113 and / or K156. 20. -(i) Q123 and / or Q146, and (ii) K129 and / or E140. -(i) E113 and / or K156, and (ii) Q123 and / or Q146, or - the monomer according to any one of the preceding aspects, comprising modifications at (i) E113 and / or K156, and (ii) K129 and / or E140. 21. The monomer according to any one of the preceding aspects, comprising modifications at (i) E113 and / or K156, (ii) Q123 and / or Q146, and (iii) K129 and / or E140. 22. The monomer of any one of the preceding aspects, containing one or more of E113S / N / Y / K / R, Q123S / A / N / M / Y / G / K / R, K129S / N / Y, E140S / N / K / R, Q146S / A / N / M / K / R / G / Y and K156S / N. 23. The monomer of any one of the preceding aspects, wherein the monomer is chemically modified. 24. The monomer according to aspect 23, wherein the monomer is chemically modified by attachment of the molecule to one or more cysteines, attachment of the molecule to one or more lysines, attachment of the molecule to one or more unnatural amino acids, enzymatic modification of an epitope, or a terminal modification. 25. The monomer of any one of the preceding aspects, wherein the monomer is capable of forming a heptameric pore. 26. A construct comprising two or more covalently linked monomers derived from cytotoxin K, wherein at least one of the monomers is a mutant cytotoxin K monomer as defined in any one of the preceding aspects. 27. The construct according to aspect 26, wherein the monomers are genetically fused or linked via a linker. 28. A polynucleotide encoding a mutant cytotoxin K monomer according to any one of aspects 1 to 25, or a construct according to aspect 26 or 27. 29. A homo-oligomeric pore comprising a plurality of mutant monomers according to any one of aspects 1 to 25, said pore preferably being a heptameric pore. 30. A hetero-oligomeric pore comprising at least one mutant monomer according to any one of aspects 1 to 25, said pore preferably being a heptameric pore. 31. A pore comprising at least one construct according to embodiment 26 or 27. 32. A construct according to aspect 26 or 27, or a pore according to any one of aspects 29 to 31, wherein at least one monomer in the construct or pore is a monomer of SEQ ID NO:1. 33. A membrane comprising a pore according to any one of embodiments 29 to 31. 34. An array comprising a plurality of membranes according to embodiment 33. 35. A device comprising an array according to embodiment 34, means for applying a potential across the membrane, and means for detecting an electrical or optical signal across the membrane. 36. A method for characterizing a target analyte, comprising: (a) contacting a target analyte with a pore according to any one of aspects 29 to 31 such that the target analyte migrates relative to the pore; (b) taking one or more measurements characteristic of the analyte as it migrates relative to the pore; thereby characterizing the target analyte. 37. The method according to embodiment 36, wherein the target analyte is a metal ion, an inorganic salt, a polymer, an amino acid, a peptide, a polypeptide, a protein, a nucleotide, an oligonucleotide, a polynucleotide, an oligosaccharide. 38. The method of embodiment 37, wherein the target analyte is or comprises a polypeptide or polynucleotide. 39. The method of embodiment 37 or 38, wherein the target analyte comprises a polynucleotide, the method comprising: (i) contacting the polynucleotide with a polynucleotide-binding protein capable of controlling movement of the polynucleotide relative to the pore; and (ii) taking one or more measurements characteristic of the polynucleotide as the polynucleotide moves relative to the pore. 40. Use of a pore according to any one of aspects 29 to 31 for characterising a target analyte. 41. A method for characterizing a target polypeptide, comprising: (a) contacting a target polypeptide with a cytotoxin K pore such that the target analyte translocates relative to the pore; (b) taking one or more measurements characteristic of the polypeptide as it moves relative to the pore; thereby characterizing the target polypeptide. 42. The method of embodiment 41, wherein the method comprises: (i) contacting the polypeptide with a polypeptide handling enzyme capable of controlling the movement of the polypeptide relative to the pore; and (ii) taking one or more measurements characteristic of the polypeptide as it moves relative to the pore. 43. The method of embodiment 41 or 42, wherein the target analyte comprises a polynucleotide-polypeptide conjugate, the method comprising: (i) contacting the conjugate with a polynucleotide binding protein capable of controlling movement of the polynucleotide of the conjugate relative to the pore; and (ii) taking one or more measurements characteristic of the polypeptide as the conjugate moves relative to the pore. 44. The method according to aspect 43, wherein the cytotoxin K pore is a pore according to any one of aspects 29 to 31. 45. Use of the cytotoxin K pore to characterize target polypeptides. 46. ​​Use of a cytotoxin K pore according to aspect 45, wherein the cytotoxin K pore comprises a mutant cytotoxin K monomer according to any one of aspects 1 to 25. 47. Use of a cytotoxin K pore according to aspect 45 or 46, wherein the cytotoxin K pore is a pore according to any one of aspects 29 to 31. 48. A kit for characterizing a target analyte, comprising: (a) a pore according to any one of aspects 29 to 31; and (b) a polynucleotide binding protein or a polypeptide handling enzyme.

Claims

1. 1. A method for characterizing a target analyte, comprising: (a) contacting the target analyte with a pore comprising at least one mutant cytotoxin K monomer comprising a variant of the amino acid sequence of SEQ ID NO: 1 such that the target analyte moves relative to the pore; the variant comprises a modification at one or more positions selected from E113, Q123, K129, E140, Q146 and K156; The modification at Q123 is Q123S / T / N / G / A / V / L / I / C / R / K / M and the modification at K129 is selected from K129S / T / N / Q / G / A / V / L / I / C / R; contacting, wherein the modification alters the ability of the monomer to interact with the analyte; (b) taking one or more measurements characteristic of the analyte as it moves relative to the pore; thereby characterizing said target analyte.

2. A: The variant has at least 70% identity with the amino acid sequence of SEQ ID NO: 1; and / or B: the one or more modifications each independently (a) change the size of the amino acid residue at the modified position, (b) change the net charge of the amino acid residue at the modified position, (c) change the hydrogen bonding properties of the amino acid residue at the modified position, (d) introduce or remove one or more chemical groups that interact via a delocalized electron π-system into or from the amino acid residue at the modified position, and / or (e) change the structure of the amino acid residue at the modified position; and / or C: The method of claim 1, wherein the pore has a solvent-accessible channel from a first opening to a second opening of the pore, the solvent-accessible channel comprising at least one constriction, and the one or more modifications are made to amino acids within the constriction to (a) alter the interaction of the analyte with the constriction as the analyte moves through the pore, (b) alter the size of the constriction, (c) alter the net charge of the constriction, (d) alter the hydrogen-bonding properties of the amino acid residues within the constriction, (e) introduce or remove one or more chemical groups within the constriction that interact via a delocalized electron π-system, and / or (f) alter the structure of the constriction.

3. wherein the variant is i) comprising one or more modifications at one or more positions within the region of SEQ ID NO: 1 between about V111 and about T158; ii) comprises one or more modifications within the region of SEQ ID NO: 1 between about V111 and about S131, and / or between about S135 and about T158; iii) comprises one or more modifications within the region of SEQ ID NO: 1 between about S119 and about G126, preferably between S121 and G125, and / or between about A143 and about S150, preferably between T144 and T148; iv) comprises one or more modifications within the region of SEQ ID NO: 1 between about G126 and about V132, preferably between S127 and S131, and / or between about P137 and about A143, preferably between S138 and G142; v) the method of claim 1, comprising one or more modifications within the region of SEQ ID NO: 1 between about N109 and about T117, preferably between V111 and T115, and / or between about S152 and about Y160, preferably between S154 and T158.

4. 2. The method of claim 1, wherein the monomer comprises a modification at one or more of the following positions of SEQ ID NO:1: T115, T117, S119, S121, G125, S127, S131, V132, T133, P134, S135, G136, P137, S138, G142, T144, T148, S150, S152 and S154.

5. 2. The method of claim 1, wherein the variants independently comprise one or more amino acid substitutions, additions and / or deletions at the one or more positions, preferably wherein the variants comprise one or more amino acid substitutions, and the amino acids substituted in the variants are selected from aspartate, glutamate, serine, threonine, asparagine, glutamine, glycine, alanine, valine, leucine, isoleucine, cysteine, arginine, lysine and phenylalanine.

6. The monomer is E113S / T / N / Q / G / A / V / L / I / C / R / K / F / Y T115S / N / Q / G / A / V / L / I / C / R / K / F T117S / N / Q / G / A / V / L / I / C / R / K / F S119T / N / Q / G / A / V / L / I / C / R / K / F S121T / N / Q / G / A / V / L / I / C / R / K / F Q123S / T / N / G / A / V / L / I / C / R / K / F / M / Y G125S / T / N / Q / A / V / L / I / C / R / K / F S127T / N / Q / G / A / V / L / I / C / R / K / F K129S / T / N / Q / G / A / V / L / I / C / R / F / Y S131T / N / Q / G / A / V / L / I / C / R / K / F V132S / T / N / Q / G / A / L / I / C / R / K / F T133S / N / Q / G / A / V / L / I / C / R / K / F P134S / T / N / Q / G / A / V / L / I / C / R / K / F S135T / N / Q / G / A / V / L / I / C / R / K / F G136S / T / N / Q / A / V / L / I / C / R / K / F P137S / T / N / Q / G / A / V / L / I / C / R / K / F S138T / N / Q / G / A / V / L / I / C / R / K / F E140S / T / N / Q / G / A / V / L / I / C / R / K / F G142S / T / N / Q / A / V / L / I / C / R / K / F T144S / N / Q / G / A / V / L / I / C / R / K / F Q146S / T / N / G / A / V / L / I / C / R / K / F / M / Y T148S / N / Q / G / A / V / L / I / C / R / K / F S150T / N / Q / G / A / V / L / I / C / R / K / F S152T / N / Q / G / A / V / L / I / C / R / K / F S154T / N / Q / G / A / V / L / I / C / R / K / F, and 2. The method of claim 1, comprising one or more modifications selected from K156S / T / N / Q / G / A / V / L / I / C / R / F. Claim 7: i) the monomer comprises a modification at Q123 and / or Q146; and / or ii) the monomer comprises a modification at K129 and / or E140; and / or iii) The method of claim 1, wherein the monomer comprises a modification at E113 and / or K156.

8. The monomer is - (i) Q123 and / or Q146, and (ii) K129 and / or E140 (i) E113 and / or K156, and (ii) Q123 and / or Q146, or - comprising modifications at (i) E113 and / or K156, and (ii) K129 and / or E140, 2. The method of claim 1, wherein the monomer preferably comprises modifications at (i) E113 and / or K156, (ii) Q123 and / or Q146, and (iii) K129 and / or E140.

9. 2. The method of claim 1, wherein the monomers contain one or more of E113S / N / Y / K / R, Q123S / A / N / M / Y / G / K / R, K129S / N / Y, E140S / N / K / R, Q146S / A / N / M / K / R / G / Y, and K156S / N.

10. 2. The method of claim 1, wherein the monomer is chemically modified, preferably by attachment of a molecule to one or more cysteines, one or more lysines, one or more unnatural amino acids, enzymatic modification of an epitope, or terminal modification.

11. i) the pore is a homo-oligomeric pore comprising a plurality of mutant monomers as defined in any one of claims 1 to 10, the pore preferably being a heptameric pore; or ii) the pore is a hetero-oligomeric pore comprising at least one mutant monomer as defined in any one of claims 1 to 10, the pore preferably being a heptameric pore; or iii) The method of claim 1, wherein the pore comprises a construct comprising two or more covalently linked monomers derived from cytotoxin K, at least one of the monomers being a mutant cytotoxin K monomer as defined in any one of claims 1 to 10.

12. 2. The method of claim 1, wherein the target analyte is a metal ion, an inorganic salt, a polymer, an amino acid, a peptide, a polypeptide, a protein, a nucleotide, an oligonucleotide, a polynucleotide, an oligosaccharide, preferably the target analyte is or comprises a polypeptide or a polynucleotide.

13. A: The target analyte comprises a polynucleotide, and the method comprises: (i) contacting the polynucleotide with a polynucleotide-binding protein capable of controlling movement of the polynucleotide relative to the pore; and (ii) taking one or more measurements characteristic of the polynucleotide as it moves relative to the pore; or B: The method of claim 12, wherein the target analyte comprises a polypeptide, and the method comprises (i) contacting the polypeptide with a polypeptide handling enzyme capable of controlling the movement of the polypeptide relative to the pore, and (ii) taking one or more measurements characteristic of the polypeptide as it moves relative to the pore, preferably wherein the target polypeptide is comprised in a polynucleotide-polypeptide conjugate, and the method comprises (i) contacting the conjugate with a polynucleotide binding protein capable of controlling the movement of the polynucleotide of the conjugate relative to the pore, and (ii) taking one or more measurements characteristic of the polypeptide as it moves relative to the pore.

14. A mutant cytotoxin K monomer comprising a variant of the amino acid sequence of SEQ ID NO: 1, wherein said monomer is capable of forming a pore; A mutant cytotoxin K monomer, wherein said monomer is as defined in any one of claims 1 to 10.

15. i) the pore is a homo-oligomeric pore comprising a plurality of mutant monomers according to claim 14, the pore preferably being a heptameric pore; or ii) the pore is a hetero-oligomeric pore comprising at least one mutant monomer according to claim 14, the pore preferably being a heptameric pore; or iii) A pore comprising a mutant cytotoxin K monomer according to claim 14, wherein the pore comprises at least one construct comprising two or more covalently linked monomers derived from cytotoxin K, at least one of the monomers being a mutant cytotoxin K monomer according to claim 14, preferably wherein the monomers are genetically fused or linked via a linker.