Nanopore with extended channel

By extending the channel of transmembrane protein nanopores with heterologous insertions, the method addresses limitations in analyte characterization, achieving improved sensing regions, volumes, and dwell times for better analyte analysis.

WO2026154144A1PCT designated stage Publication Date: 2026-07-23OXFORD UNIVERSITY INNOVATION LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
OXFORD UNIVERSITY INNOVATION LTD
Filing Date
2026-01-16
Publication Date
2026-07-23

Smart Images

  • Figure EP2026051095_23072026_PF_FP_ABST
    Figure EP2026051095_23072026_PF_FP_ABST
Patent Text Reader

Abstract

Provided herein are methods of characterising analytes using a transmembrane protein nanopore having an extended channel, as well as methods of producing a transmembrane protein nanopore having an increased analyte dwell time.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] NANOPORE WITH EXTENDED CHANNEL

[0002] Cross reference to related applications

[0003] This application claims priority from United Kingdom Patent Application No.

[0004] 2500666.9 filed on 17 January 2025, the contents of which are hereby incorporated by reference.

[0005] Field

[0006] The disclosure relates to a method of characterising an analyte by contacting the analyte with a transmembrane protein nanopore having an extended channel, and taking measurements characteristic of the analyte. The disclosure also relates to a transmembrane protein nanopore having an extended channel, and to methods of producing a transmembrane protein nanopore having an increased analyte dwell time.

[0007] Background

[0008] Nanopores have been applied to measure a range of analyte characteristics at a single molecule level (Bayley, H. & Cremer, P. S. Stochastic sensors inspired by biology. Nature 413, 226-230 (2001)). Nanopore sensing is an approach to analyte detection and characterization that relies on the observation of individual binding or interaction events between the analyte molecules and an ion conducting channel. Nanopore sensors can be created by placing a single pore of nanometre dimensions in an electrically insulating membrane. Electrical and / or optical measurements through the pore can be taken in the presence of analyte molecules. The presence of an analyte inside or near the nanopore alters the measurements obtained, thus allowing the identity of the analyte to be revealed.

[0009] Many approaches have been identified to further improve the characterisation of analytes by nanopore sensors. For example, polynucleotide binding proteins, such as helicases or polymerases, have been used to control the movement of polynucleotide analytes through nanopores, see for example WO 2013 / 057495, WO 2013 / 098562, WO 2014 / 013260, WO 2013 / 098561 and WO 2015 / 055981. In a further approach, nanopores have been modified by amino acid substitution to improve the characterisation of analytes, see for example WO 2012 / 107778, WO 2017 / 184866 and WO 2016 / 034591. In another approach described in WO 2015 / 040423, a highly charged DNA leader sequence is attached to apeptide, polypeptide or protein analyte in order to electrophoretically thread the analyte through a nanopore and improve the characterisation thereof.

[0010] However, there is a need for new and / or improved methods for characterising an analyte using a nanopore system.

[0011] Summary of the invention

[0012] As described in more detail herein, the inventors sought to extend the length of the channel of a transmembrane protein nanopore. The inventors surprisingly found that extension of the channel of the nanopore confers many features beneficial for analyte characterisation, such as longer sensing regions, larger sensing volumes and longer analyte dwell times.

[0013] Accordingly, the disclosure relates to a method of characterising an analyte, the method comprising contacting the analyte with a transmembrane protein nanopore having a first opening, a second opening and a channel therebetween, wherein the nanopore comprises one or more heterologous insertions that extend the length of the channel, and taking one or more measurements characteristic of the analyte as the analyte moves with respect to the nanopore, thereby characterising the analyte. In some embodiments, the analyte is selected from a nucleotide, a polynucleotide, an amino acid, a polypeptide, a saccharide and a polysaccharide.

[0014] In another aspect, the disclosure relates to a transmembrane protein nanopore having a first opening, a second opening and a channel therebetween, wherein the nanopore comprises one or more heterologous insertions that extend the length of the channel.

[0015] In a further aspect, the disclosure relates to a device comprising an array of transmembrane protein nanopores of the disclosure comprised in a membrane.

[0016] In a further aspect, the disclosure relates to a method of increasing the interaction dwell time of an analyte with a transmembrane protein nanopore comprising a first opening, a second opening and a channel therebetween, the method comprising modifying the nanopore by introducing one or more heterologous insertions into the nanopore, thereby extending the length of the channel of the nanopore.

[0017] In some embodiments, the one or more heterologous insertions are one or more heterologous amino acid sequence insertions, wherein each heterologous amino acid sequence insertion comprises at least two natural and / or non-natural amino acids.In some embodiments, the length of the channel is extended by at least about 0.56 nm, by at least about 1 nm, by at least about 2 nm, by at least about 3 nm, by at least about 4 nm or by at least about 5 nm.

[0018] In some embodiments, the nanopore comprises a barrel domain and said one or more heterologous insertions are made in the barrel domain. In some embodiments, the nanopore further comprises a cap domain. In some embodiments, the barrel domain comprises a P-barrel.

[0019] In some embodiments, the nanopore comprises one or more heterologous insertions in a membrane-exposed region of the barrel domain. In some embodiments, the nanopore comprises one or more heterologous insertions in a solvent-exposed region of the barrel domain. In some embodiments, the nanopore comprises one or more heterologous insertions at a position corresponding to between residues 111 and 128 of SEQ ID NO: 1 and / or between residues 130 and 147 of SEQ ID NO: 1. In some embodiments, the nanopore comprises one or more heterologous insertions at a position corresponding to between residues 111 and 117 of SEQ ID NO: 1 and / or between residues 142 and 147 of SEQ ID NO: 1. In some embodiments, the nanopore comprises one or more heterologous insertions at a position corresponding to between residues 117 and 125 of SEQ ID NO: 1 and / or between residues 134 and 142 of SEQ ID NO: 1. In some embodiments, the nanopore comprises one or more heterologous insertions at a position corresponding to between residues 125 and 134 of SEQ ID NO: 1.

[0020] In some embodiments, the nanopore comprises a plurality of heterologous insertions that are substantially coincidental. In some embodiments, the nanopore comprises a plurality of heterologous amino acid sequence insertions which form antiparallel backbone hydrogen bonds in the nanopore.

[0021] In some embodiments, the nanopore comprises two or more constrictions within the channel, wherein the one or more heterologous insertions are positioned between the two or more constrictions. In some embodiments, the one or more heterologous insertions define a constriction within the channel of the nanopore.

[0022] In some embodiments, each heterologous amino acid sequence insertion independently comprises at least 4, at least 6, at least 8, at least 10, at least 12, at least 14, at least 16, at least 18, at least 20, at least 24, at least 28, at least 32, at least 36, at least 40, at least 44, at least 48, at least 52, at least 56, at least 60, at least 70 or at least 80 amino acids. In someembodiments, each heterologous amino acid sequence insertion in said nanopore is the same length.

[0023] In some embodiments, each heterologous amino acid sequence insertion comprises at least two different amino acids. In some embodiments, each heterologous amino acid sequence insertion comprises a sequence motif having a formula (X-Y)nor (Y-X)n; wherein:

[0024] each X is an amino acid independently selected from the group consisting of R / K / N / D / E / Q / H / P / Y / W / S / T / G / A / M / C / F / L / V / I;

[0025] each Y is an amino acid independently selected from the group consisting of K / N / D / E / Q / H / P / S / T / G / A / M; and

[0026] n is at least 1.

[0027] In some embodiments, each X is an amino acid independently selected from the group consisting of K / N / Y / S / T / G / L / V, and / or each Y is an amino acid independently selected from the group consisting of K / N / E / S / T / G / M. In some embodiments, each heterologous insertion each independently comprises a sequence motif selected from GNTN, GNTNEN, KNGNTN, GG, SG, SV, EYMSTLTY, VSIGHTLK, TN, ST, SI, TNGN, KNGN, GNEN, TNKN, ENTN, GNSN, NGNT and NKNGNT, (GNTN)3-GNTIGITN-(GNTN)3, (GNTN)2-GITNGITN-(GNTN)4, (GNTN)3-GITNGITN-(GNTN)3, (GNTN)4-GITNGITN-(GNTN)8, (GNTN)6-GITNGITN-(GNTN)6, (GNTN)6-GITIGITN-(GNTN)6, GNTHGHTN-(GNTN)6, (GNTN)6-GHTHGNTN, (GNTN)2-EN-(GNTN)2, (GNTN)2-KN-(GNTN)2, (GNTNEN)4-GNTIEIGITNEN-(GNTNEN)4, (KNGNTN)4-KNGNTIKIGITN-(KNGNTN)4, TNGK, TNGD, TNGKTNGN, TNGDTNGN, (GNTN)3G, NTN(GNTN)2, GQSQ, GQTQ, NT, ND and NDNT. In some embodiments, each heterologous amino acid sequence insertion independently comprises one or more repeats of one or more sequence motifs such as one or more sequence motifs described herein. In some embodiments, each heterologous amino acid sequence insertion comprises from 1 to 20 repeats. In some embodiments, each heterologous amino acid sequence insertion comprises from 2 to 14 repeats. In some embodiments, each heterologous amino acid sequence insertion comprises from 3 to 8 repeats.

[0028] In some embodiments, the nanopore is an oligomeric nanopore comprising a plurality of monomer subunits, wherein each monomer subunit comprises at least one or more of said heterologous insertions. In some embodiments, the nanopore is a homo-oligomeric nanopore, wherein each monomer subunit comprises the same one or more heterologous insertions. Inother embodiments, the nanopore is a heterooligomeric nanopore, and wherein at least one monomer subunit comprises said one or more heterologous insertions.

[0029] Brief description of the figures

[0030] Figure 1. Design of extended p-barrel proteins scaffolded by the cap and the rim of a-hemolysin. a, The de novo-designed amino acid sequence, GNTN, can be inserted into the P barrel of a-hemolysin above the membrane boundary (between LI 16-T117 and I142-G143) and repeated multiple times to form a stable hydrophilic P barrel of desired length, causing the total length of the transmembrane P barrel to exceed 10.5 nm (protective antigen pore, the longest transmembrane P barrel in nature), b, Design strategies for transmembrane P-barrel structures: Even-numbered amino acids were inserted into the barrel to maintain the alternating pattern in the transmembrane region, while odd-numbered amino acids would flip the pattern, causing the hydrophilic amino acids facing outward to contact the membrane; both inward and outward amino acids were hydrophilic for the insert, as this part will extend into a hydrophilic environment; inward amino acid would prefer smaller side chain compared with the outward one; Glycine kinks were included in the structure to reduce steric clash and relieve strain, c, The insertion of amino acids between L116-T117 and I142-G143 avoids disrupting the side chain hydrogen interactions, membrane boundary anchor and hydrophobic side chains on the transmembrane region, d, Large inward side chains could cause steric clash, e, Glycine residues were included every 4 amino acids to favor P-strand twist. Bioinformatic distribution of 20 amino acids on the transmembrane P barrels (f) or the ‘hydrophilic’ P barrels (g) within 13 P-barrel proteins. Left bar for each amino acid corresponds to outward-facing, right bar corresponds to inward-facing. Threonine and Glycine residues were selected as the inwardfacing amino acid, and Asparagine was selected as the outward-facing amino acid.

[0031] Figure 2. Confirmation of the assembly structure of «HL-(GNTN)nmutants, a, E. coli expressed aHL-(GNTN)nmutants showed both heptameric and monomeric bands with the increasing number of GNTN inserts, b, Representative electrophysiological evaluation of extended oligomers by using single channel recording. The P-barrel of aHL can accommodate the insertion of at least 8 continuous glycine residues on each P-strand, albeit at the cost of reduced stability, as indicated by the noise level in the current trace. The optimized aHL-(GNTN)n (n = 3~5, 8) mutants demonstrated a stable baseline current, although the aHL-(GNTN)s and aHL-(GNTN)s mutants exhibited some sub-conductance, c, aHL-(GNTN)nmutants are considered as series connection of aHL resistor and insert resistor. The insert resistor is proportional to the length of insert, d. The resistance of different extended mutants follows a linear dependence on the number of inserted amino acids. Conditions in b and d: 2 M KC1, 10 mM Tris, 0.1 mM EDTA, pH 8.5, +50 mV (trans), 23 ± 1 °C.

[0032] Figure 3. Structural validation of extended p barrels, a, Single-cysteine mutants based on aHL-(GNTN)3 template. Single cysteine mutagenesis was performed at 14 distinct positions spanning the entirety of the P strand to validate the amino acid orientations at these locations. The representative current traces showed the reaction between 7 inward (b) or outward (c) cysteine residues and DTNB, followed by the reduction by adding DTT. Conditions: 2 M KC1, 10 mM Tris, 0.1 mM EDTA, pH 8.0, 0.2 mM DTNB (trans), 0.8 mM DTT (trans), -50 mV (trans), 23 + 1 °C. d, Structural characterization of aHL-WT nanodiscs and aHL-(GNTN)8-TM nanodiscs by negative stained TEM. Scale bar, 10 nm.

[0033] Figure 4. Extended dwell times of non-covalent adaptor, -Cyclodextrin ( -CD), within extended p barrels, a, Schematic structure of aHL and a molecular adaptor, P-CD (top). Model for the interaction of aHL-(GNTN)3 mutant with P-CD (bottom). The dwell time for P-CD binding was prolonged in aHL-(GNTN)nmutants as the number of inserts increased (b, g).

[0034] Kinetics parameters, the rate constant kon(c), koff(d), and the dissociation constant Kd (e) for the interaction of P-CD with aHL-WT and aHL-(GNTN)nat +50 mV from three or more separate experiments, f, The percentage residual current increases with the addition of more amino acids, h, An enhanced stochastic sensor, aHL-(GNTN)4-P-CD complex for detecting 2-adamantanamine hydrochloride, i, The representative current traces showed the open-pore current (level 1), the binding of P-CD (level 2) and the binding of 2-adamantanamine hydrochloride on an assembled P-CD (level 3). Conditions: 2 M KC1, 10 mM Tris, 0.1 mM EDTA, pH 8.5, +50 mV (trans), 23 ± 1 °C, 40 pM P-CD (trans) (b); 2~80 pM P-CD (trans) (c-g); 55 pM P-CD (trans), 40 pM 2-adamantanamine hydrochloride (i).

[0035] Figure 5. Extended dwell times of analytes within extended p barrels, a, The aHL variants and biopolymer analytes tested, b, scatter plots characterising Ires% and dwell time of 92-nt ssDNA in aHL-WT, aHL-(GNTN)3 and aHL-(GNTN)3-Ml 13R pores, c, representative current traces (left) and scatter plots (right) illustrating the increased dwell time of 10-AA peptide in the extended mutant aHL-(GNTN)3-M113R and aHL-(GNTN)3M113RT115R nanopores compared to aHL-M113R. d, representative current traces (left) and scatter plots(right) illustrating the increased dwell time of a 21-AA peptide in the extended mutant aHL-(GNTN)s-Ml 13R nanopore compared to aHL-Ml 13R.

[0036] Figure 6. Confirmation of the assembly structure of «HL-(GNTN)nmutants.

[0037] Collated dataset expanding on results shown in Figure 2 b,d. a, Electrophysiological evaluation of extended oligomers by using single channel recording. The P-barrel of aHL can accommodate the insertion of at least 8 continuous glycine residues on each P-strand, albeit at the cost of reduced stability, as indicated by the noise level in the current trace. The optimized aHL-(GNTN)n(n = 3~5, 8) mutants demonstrated a stable baseline current, although the aHL-(GNTN)s and aHL-(GNTN)s mutants exhibited some sub-conductance, b. The resistance of different extended mutants follows a linear dependence on the number of inserted amino acids. Conditions: 2 M KC1, 10 mM Tris, 0.1 mM EDTA, pH 8.5, +50 mV (trans), 23 ± 1 °C.

[0038] Figure 7. Extended dwell times of non-covalent adaptor, P-Cyclodextrin (P-CD), within extended p barrels. Collated dataset expanding on results shown in Figure 4. a The percentage residual current increases with the addition of more amino acids in aHL-(GNTN)nmutants. The dwell time for P-CD binding was prolonged in aHL-(GNTN)nmutants as the number of inserts increased (b). Conditions: 2 M KC1, 10 mM Tris, 0.1 mM EDTA, pH 8.5, +50 mV (trans), 23 ± 1 °C, 2~80 pM P-CD (trans).

[0039] Figure 8. Extended dwell times of analytes within extended p barrels, a, The aHL variants and biopolymer analytes tested. Representative current traces (b) and scatter plots (e) illustrating the increased dwell time of 10- AA peptide in the extended mutant aHL-(GNTN)3 compared to aHL-WT. c, Representative current traces showing the increased retention time of 10-AA peptide in the extended mutant aHL-(GNTN)3-M113R compared to aHL-Ml 13R. f, Scatter plots characterising Ires% and dwell time of 10-AA, 21-AA and 31-AA peptides in aHL-Ml 13R and aHL-116 / 142-(GNTN)3-117 / 143-M113R pore. The data points for 21-AA peptide in aHL-Ml 13R pore are highlighted in red to distinguish them from 31-AA peptide. Representative current traces (d) and scatter plots (g) showing the increased retention time of 10-AA peptide event in the extended mutant aHL-(GNTN)3-Ml 13RT115R compared to aHL-Ml 13RT115R. Scatter plots characterising Ires% and dwell time of 21-AA peptide and its PTM variants in aHL-116 / 142-(GNTN)3-117 / 143-M113R pore (h) and aHL-Ml 13R pore (i). j, Scatter plots characterizing translocation events of a 92-nt ssDNA through aHL-WT and aHL-(GNTN)3nanopores. Conditions: 2 M KC1, 10 mM Tris, 0.1 mM EDTA, pH 8.5, 4~20 pM for each peptide (trans), 200 nm or 1 pM 92-nt ssDNA (cis), +100 mV, 23.5 ± 1.5 °C.Detailed description

[0040] The present invention will be described with respect to particular embodiments and with reference to certain drawings but the invention is not limited thereto but only by the claims. Any reference signs in the claims shall not be construed as limiting the scope. Of course, it is to be understood that not necessarily all aspects or advantages may be achieved in accordance with any particular embodiment of the invention. Thus, for example those skilled in the art will recognize that the invention may be embodied or carried out in a manner that achieves or optimizes one advantage or group of advantages as taught herein without necessarily achieving other aspects or advantages as may be taught or suggested herein.

[0041] The invention, both as to organization and method of operation, together with features and advantages thereof, may best be understood by reference to the following detailed description when read in conjunction with the accompanying drawings. The aspects and advantages of the invention will be apparent from and elucidated with reference to the embodiment(s) described hereinafter. Reference throughout this specification to "one embodiment" or "an embodiment" means that a particular feature, structure or characteristic described in connection with the embodiment is included in at least one embodiment of the present invention. Thus, appearances of the phrases "in one embodiment" or "in an embodiment" in various places throughout this specification are not necessarily all referring to the same embodiment, but may. Similarly, it should be appreciated that in the description of exemplary embodiments of the invention, various features of the invention are sometimes grouped together in a single embodiment, figure, or description thereof for the purpose of streamlining the disclosure and aiding in the understanding of one or more of the various inventive aspects. This method of disclosure, however, is not to be interpreted as reflecting an intention that the claimed invention requires more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive aspects lie in less than all features of a single foregoing disclosed embodiment.

[0042] It should be appreciated that “embodiments” of the disclosure can be specifically combined together unless the context indicates otherwise. The specific combinations of all disclosed embodiments (unless implied otherwise by the context) are further disclosed embodiments of the claimed invention.

[0043] In addition as used in this specification and the appended claims, the singular forms “a”, “an”, and “the” include plural referents unless the content clearly dictates otherwise. Thus,for example, reference to “an analyte” includes two or more analytes, reference to “a nanopore” includes two or more such nanopores, reference to “an adduct” includes two or more such adducts, and the like.

[0044] All publications, patents and patent applications cited herein, whether supra or infra, are hereby incorporated by reference in their entirety.

[0045] Definitions

[0046] Where an indefinite or definite article is used when referring to a singular noun e.g. "a" or "an", "the", this includes a plural of that noun unless something else is specifically stated. Where the term "comprising" is used in the present description and claims, it does not exclude other elements or steps. Furthermore, the terms first, second, third and the like in the description and in the claims, are used for distinguishing between similar elements and not necessarily for describing a sequential or chronological order. It is to be understood that the terms so used are interchangeable under appropriate circumstances and that the embodiments of the invention described herein are capable of operation in other sequences than described or illustrated herein. The following terms or definitions are provided solely to aid in the understanding of the invention. Unless specifically defined herein, all terms used herein have the same meaning as they would to one skilled in the art of the present invention. Practitioners are particularly directed to Sambrook et al., Molecular Cloning: A Laboratory Manual, 4thed., Cold Spring Harbor Press, Plainsview, New York (2012); and Ausubel et al., Current Protocols in Molecular Biology (Supplement 114), John Wiley & Sons, New York (2016), for definitions and terms of the art. The definitions provided herein should not be construed to have a scope less than understood by a person of ordinary skill in the art.

[0047] " About" as used herein when referring to a measurable value such as an amount, a temporal duration, and the like, is meant to encompass variations of ± 20 % or ± 10 %, more preferably ± 5 %, even more preferably ± 1 %, and still more preferably ± 0.1 % from the specified value, as such variations are appropriate to perform the disclosed methods.

[0048] “Nucleotide sequence”, “DNA sequence” or “nucleic acid molecule(s)” as used herein refers to a polymeric form of nucleotides of any length, either ribonucleotides or deoxyribonucleotides. This term refers only to the primary structure of the molecule. Thus, this term includes double- and single-stranded DNA, and RNA. The term “nucleic acid” as used herein, is a single or double stranded covalently-linked sequence of nucleotides in whichthe 3' and 5' ends on each nucleotide are joined by phosphodiester bonds. The polynucleotide may be made up of deoxyribonucleotide bases or ribonucleotide bases. Nucleic acids may be manufactured synthetically in vitro or isolated from natural sources. Nucleic acids may further include modified DNA or RNA, for example DNA or RNA that has been methylated, or RNA that has been subject to post-translational modification, for example 5 ’-capping with 7-methylguanosine, 3 ’-processing such as cleavage and polyadenylation, and splicing. Nucleic acids may also include synthetic nucleic acids (XNA), such as hexitol nucleic acid (HNA), cyclohexene nucleic acid (CeNA), threose nucleic acid (TNA), glycerol nucleic acid (GNA), locked nucleic acid (LNA) and peptide nucleic acid (PNA). Sizes of nucleic acids, also referred to herein as “polynucleotides” are typically expressed as the number of base pairs (bp) for double stranded polynucleotides, or in the case of single stranded polynucleotides as the number of nucleotides (nt). One thousand bp or nt equal a kilobase (kb). Polynucleotides of less than around 40 nucleotides in length are typically called “oligonucleotides” and may comprise primers for use in manipulation of DNA such as via polymerase chain reaction (PCR).

[0049] The term “amino acid” in the context of the present disclosure is used in its broadest sense and is meant to include organic compounds containing amine (NH2) and carboxyl (COOH) functional groups, along with a side chain (e.g., a R group) specific to each amino acid. In some embodiments, the amino acids refer to naturally occurring L a-amino acids or residues. The commonly used one and three letter abbreviations for naturally occurring amino acids are used herein: A=Ala; C=Cys; D=Asp; E=Glu; F=Phe; G=Gly; H=His; I=Ile; K=Lys; L=Leu; M=Met; N=Asn; P=Pro; Q=Gln; R=Arg; S=Ser; T=Thr; V=Val; W=Trp; and Y=Tyr (Lehninger, A. L., (1975) Biochemistry, 2d ed., pp. 71-92, Worth Publishers, New York). The general term “amino acid” further includes D-amino acids, retro-inverso amino acids as well as chemically modified amino acids such as amino acid analogues, naturally occurring amino acids that are not usually incorporated into proteins such as norleucine, and chemically synthesised compounds having properties known in the art to be characteristic of an amino acid, such as P-amino acids. For example, analogues or mimetics of phenylalanine or proline, which allow the same conformational restriction of the peptide compounds as do natural Phe or Pro, are included within the definition of amino acid. Such analogues and mimetics are referred to herein as "functional equivalents" of the respective amino acid. Other examples of amino acids are listed by Roberts and Vellaccio, The Peptides: Analysis, Synthesis, Biology,Gross and Meiehofer, eds., Vol. 5 p. 341, Academic Press, Inc., N. Y. 1983, which is incorporated herein by reference.

[0050] The terms “polypeptide”, and “peptide” are interchangeably used herein to refer to a polymer of amino acid residues and to variants and synthetic analogues of the same. Thus, these terms apply to amino acid polymers in which one or more amino acid residues is a synthetic non-naturally occurring amino acid, such as a chemical analogue of a corresponding naturally occurring amino acid, as well as to naturally-occurring amino acid polymers.

[0051] Polypeptides can also undergo maturation or post-translational modification processes that may include, but are not limited to: glycosylation, proteolytic cleavage, lipidization, signal peptide cleavage, propeptide cleavage, phosphorylation, and such like. A peptide can be made using recombinant techniques, e.g., through the expression of a recombinant or synthetic polynucleotide. A recombinantly produced peptide it typically substantially free of culture medium, e.g., culture medium represents less than about 20 %, more preferably less than about 10 %, and most preferably less than about 5 % of the volume of the protein preparation.

[0052] The term “protein” is used to describe a folded polypeptide having a secondary or tertiary structure. The protein may be composed of a single polypeptide, or may comprise multiple polypeptides that are assembled to form a multimer. The multimer may be a homooligomer, or a heterooligmer. The protein may be a naturally occurring, or wild type protein, or a modified, or non-naturally, occurring protein. The protein may, for example, differ from a wild type protein by the addition, substitution or deletion of one or more amino acids.

[0053] A “variant” of a protein encompass peptides, oligopeptides, polypeptides, proteins and enzymes having amino acid substitutions, deletions and / or insertions relative to the unmodified or wild-type protein in question and having similar biological and functional activity as the unmodified protein from which they are derived. The term "amino acid identity" as used herein refers to the extent that sequences are identical on an amino acid-by-amino acid basis over a window of comparison. Thus, a "percentage of sequence identity" is calculated by comparing two optimally aligned sequences over the window of comparison, determining the number of positions at which the identical amino acid residue (e.g., Ala, Pro, Ser, Thr, Gly, Val, Leu, Ile, Phe, Tyr, Trp, Lys, Arg, His, Asp, Glu, Asn, Gln, Cys and Met) occurs in both sequences to yield the number of matched positions, dividing the number ofmatched positions by the total number of positions in the window of comparison (i.e., the window size), and multiplying the result by 100 to yield the percentage of sequence identity.

[0054] For all aspects and embodiments of the present invention, a “variant” has at least 50%, 60%, 70%, 80%, 90%, 95% or 99% complete sequence identity to the amino acid sequence of the corresponding wild-type protein. Sequence identity can also be to a fragment or portion of the full-length polynucleotide or polypeptide. Hence, a sequence may have only 50 % overall sequence identity with a full-length reference sequence, but a sequence of a particular region, domain or subunit could share 80 %, 90 %, or as much as 99 % sequence identity with the reference sequence.

[0055] The term “wild-type” refers to a gene or gene product isolated from a naturally occurring source. A wild-type gene is that which is most frequently observed in a population and is thus arbitrarily designed the “normal” or “wild-type” form of the gene. In contrast, the term “modified”, “mutant” or “variant” refers to a gene or gene product that displays modifications in sequence (e.g., substitutions, truncations, or insertions), post-translational modifications and / or functional properties (e.g., altered characteristics) when compared to the wild-type gene or gene product. It is noted that naturally occurring mutants can be isolated; these are identified by the fact that they have altered characteristics when compared to the wild-type gene or gene product. Methods for introducing or substituting naturally-occurring amino acids are well known in the art. For instance, methionine (M) may be substituted with arginine (R) by replacing the codon for methionine (ATG) with a codon for arginine (CGT) at the relevant position in a polynucleotide encoding the mutant monomer. Methods for introducing or substituting non-naturally-occurring amino acids are also well known in the art. For instance, non-naturally-occurring amino acids may be introduced by including synthetic aminoacyl -tRNAs in the IVTT system used to express the mutant monomer.

[0056] Alternatively, they may be introduced by expressing the mutant monomer in E. coli that are auxotrophic for specific amino acids in the presence of synthetic (i.e. non -naturally-occurring) analogues of those specific amino acids. They may also be produced by naked ligation if the mutant monomer is produced using partial peptide synthesis. Conservative substitutions replace amino acids with other amino acids of similar chemical structure, similar chemical properties or similar side-chain volume. The amino acids introduced may have similar polarity, hydrophilicity, hydrophobicity, basicity, acidity, neutrality or charge to the amino acids they replace. Alternatively, the conservative substitution may introduce another aminoacid that is aromatic or aliphatic in the place of a pre-existing aromatic or aliphatic amino acid. Conservative amino acid changes are well-known in the art and may be selected in accordance with the properties of the 20 main amino acids as defined in Table 1 below. Where amino acids have similar polarity, this can also be determined by reference to the hydropathy scale for amino acid side chains in Table 2.

[0057] Table 1 - Chemical properties of amino acids

[0058] Ala aliphatic, hydrophobic, neutral Met hydrophobic, neutral

[0059] Cys polar, hydrophobic, neutral Asn polar, hydrophilic, neutral Asp polar, hydrophilic, charged (-) Pro hydrophobic, neutral

[0060] Glu polar, hydrophilic, charged (-) Gin polar, hydrophilic, neutral Phe aromatic, hydrophobic, neutral Arg polar, hydrophilic, charged (+) Gly aliphatic, neutral Ser polar, hydrophilic, neutral His aromatic, polar, hydrophilic, Thr polar, hydrophilic, neutral charged (+)

[0061] Ile aliphatic, hydrophobic, neutral Val aliphatic, hydrophobic, neutral Lys polar, hydrophilic, charged(+) Trp aromatic, hydrophobic, neutral

[0062]

[0063] Leu aliphatic, hydrophobic, neutral Tyr aromatic, polar, hydrophobic

[0064] Table 2 - Hydropathy scale

[0065] Side Chain Hydropathy

[0066] Ile 4.5

[0067] Val 4.2

[0068] Leu 3.8

[0069] Phe 2.8

[0070] Cys 2.5

[0071] Met 1.9

[0072] Ala 1.8

[0073] Gly -0.4

[0074] Thr -0.7

[0075] Ser -0.8

[0076] Trp -0.9

[0077] Tyr -1.3

[0078] Pro -1.6

[0079] His -3.2

[0080] Glu -3.5

[0081] Gln -3.5

[0082] Asp -3.5

[0083] Asn -3.5

[0084] Lys -3.9

[0085] Arg -4.5A mutant or modified protein, monomer or peptide can also be chemically modified in any way and at any site. A mutant or modified monomer or peptide is preferably chemically modified by attachment of a molecule to one or more cysteines (cysteine linkage), attachment of a molecule to one or more lysines, attachment of a molecule to one or more non-natural amino acids, enzyme modification of an epitope or modification of a terminus. Suitable methods for carrying out such modifications are well-known in the art. The mutant of modified protein, monomer or peptide may be chemically modified by the attachment of any molecule. For instance, the mutant of modified protein, monomer or peptide may be chemically modified by attachment of a dye or a fluorophore.

[0086] The term “heterologous” in the context of a polynucleotide or a polypeptide refers to a structure, typically a polynucleotide sequence or an amino acid sequence respectively, that is not natively found in the relevant position of the polynucleotide or polypeptide sequence. For example, in some embodiments, a heterologous insertion may be a sequence that is found in a native polynucleotide sequence or a native polypeptide sequence, but is duplicated or transposed such that the sequence is inserted in a different position of the polynucleotide or polypeptide sequence. In some embodiments, a heterologous sequence insertion is a de novo or synthetic sequence insertion, i.e. one that is not derived from a naturally occurring polynucleotide or polypeptide sequence. In some embodiments, a heterologous sequence insertion is a sequence that is derived from a naturally occurring polynucleotide or polypeptide that is different to the polynucleotide or polypeptide comprising the heterologous sequence insertion. A heterologous sequence insertion may be identified by any means known to the skilled person. For example, a heterologous sequence insertion in an a-hemolysin nanopore may be identified by aligning the query sequence with the a-hemolysin nanopore sequence of SEQ ID NO: 1, and identifying any amino acid insertions relative to SEQ ID NO: 1, e.g. in addition to the amino acid positions corresponding to the amino acid positions of SEQ ID NO: 1.

[0087] The term “independently” in the context of a selection from a list refers to the lack of a relationship between any two or more selections from said list. For example, for a feature in which each heterologous amino acid sequence independently comprises a sequence motif selected from GNTN, GNTNEN, KNGNTN, GG, SG, SV, EYMSTLTY, VSIGHTLK, TN, ST, SI, TNGN, KNGN, GNEN, TNKN, ENTN, GNSN, NGNT and NKNGNT, a first heterologous amino acid sequence insertion may be GNTN, whereas a second heterologousamino acid sequence could equally be selected from any of GNTN, GNTNEN, KNGNTN, GG, SG, SV, EYMSTLTY, VSIGHTLK, TN, ST, SI, TNGN, KNGN, GNEN, TNKN, ENTN, GNSN, NGNT and NKNGNT.

[0088] An amino acid position “ corresponding to" a specified position of a reference sequence may be said amino acid position of the reference sequence. For example, an amino acid position corresponding to amino acid position 116 of SEQ ID NO: 1 is position 116 of SEQ ID NO: 1. Additionally, an amino acid position in a query sequence “ corresponding to" a specified position ‘position X’ of a reference sequence (e.g. position ‘116’ of SEQ ID NO: 1) might not be the same position number of the query sequence, but rather the amino acid position that best aligns to position X in the reference sequence. It is within the capabilities of the person skilled in the art to determine which amino acid positions in an alternative amino acid sequence, such as a nanopore sequence, “correspond to" the specified positions in a reference sequence (e.g. SEQ ID NO: 1). For example, the person skilled in the art may perform a sequence alignment of an alternative nanopore amino acid sequence with the reference nanopore sequence (e.g. SEQ ID NO: 1) using a suitable alignment algorithm such as that of Needleman and Wunsch, and determine which region of the alternative nanopore amino acid sequence best aligns to the specified position in the reference nanopore sequence. In some embodiments, the person skilled in the art may perform a structural alignment of an alternative nanopore amino acid sequence with the reference sequence (e.g. SEQ ID NO: 1) using a suitable structural alignment tool, and determine which region of the alternative nanopore sequence best superposes with the specified position of the reference nanopore structure. Structural alignment may be suitable in instances where the reference and alternative sequences have low sequence identity (such as less than 80% sequence identity, less than 50% sequence identity, less than 25% sequence identity, or less than 10% sequence identity) but share a structural feature, such as a P-barrel domain. In some cases, the alignment may be weighted to structural features comprising the position of the reference sequence.

[0089] The term “dwell time" refers to the length of time that an analyte or a part thereof interacts with a reader head or constriction in the channel of a nanopore. The analyte may be transiently immobilised in the channel. The analyte may be moving with respect to, such as into or through the channel. For example, an increased dwell time for an analyte that is transiently immobilised in the channel means that it is immobilised for an increased length oftime. An increased dwell time for an analyte such as a polymer that is moving with respect to, such as into or through the channel means that it is moving at a lower rate, e.g. increasing the amount of time that at least part of the analyte is in the channel. The dwell time of an analyte may result in a combination of the factors discussed above, for example, due to the time that the analyte moves with respect to, such as into or through the channel, and the time the analyte is immobilised in the channel.

[0090] Nanopore

[0091] The disclosure relates to a transmembrane protein nanopore having a first opening, a second opening and a channel therebetween. The channel is solvent-accessible. The nanopore comprises one or more heterologous insertions that extend the length of the channel. Without being bound by theory, the inventors believe that by increasing the length of the channel, the length of the sensing region, the size of the sensing volume and / or the dwell time of an analyte in the channel is increased, thereby enabling improved characterisation of the analyte.

[0092] Typically, the transmembrane protein nanopore is capable of forming the channel when inserted into a membrane as described herein, such as a lipid bilayer (e.g. a 1,2-diphytanoyl-sn-glycero-3 -phosphocholine lipid bilayer). Typically, the transmembrane protein nanopore is capable of forming the channel under conditions as described herein.

[0093] Any suitable transmembrane protein nanopore can be used in the disclosure. A transmembrane protein nanopore is a polypeptide or a collection of polypeptides that permits ions driven by an applied potential to flow from one side of a membrane to the other side of the membrane.

[0094] A transmembrane protein nanopore may be isolated, substantially isolated, purified or substantially purified. A nanopore is isolated or purified if it is completely free of any other components, such as lipids or other nanopores. A nanopore is substantially isolated if it is mixed with carriers or diluents which will not interfere with its intended use. For instance, a nanopore is substantially isolated or substantially purified if it present in a form that comprises less than 10%, less than 5%, less than 2% or less than 1% of other components, such as lipids or other nanopores. The nanopore is typically present in a membrane, for example a lipid bilayer or a synthetic membrane e.g. a block-copolymer membrane.A transmembrane protein nanopore may be a monomer or an oligomer. A transmembrane protein nanopore is often made up of several repeating subunits, such as at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, or at least 16 subunits. The nanopore is typically a hexameric, heptameric, octameric or nonameric nanopore.

[0095] The nanopore may be a homo-oligomer or a hetero-oligomer. A transmembrane protein nanopore may be a heptameric transmembrane protein nanopore.

[0096] The transmembrane protein nanopore comprises a channel through which the ions may flow. In some embodiments, the transmembrane protein nanopore comprises a barrel domain. The barrel domain comprises a first opening, a second opening and a channel therebetween. The barrel domain typically spans the membrane, i.e. is a transmembrane barrel domain. The channel of the transmembrane protein nanopore may comprise or consist of channel of the barrel domain. In other words, the channel of the transmembrane protein nanopore may be formed at least partially or fully by the barrel domain. In some embodiments, the transmembrane protein nanopore further comprises a cap domain. In some embodiments, the cap domain comprises a first opening, a second opening and a channel therebetween. In such embodiments, the channel of the transmembrane protein nanopore is formed by the barrel domain and the cap domain. In some embodiments, the cap domain is required for the folding of the barrel domain into a membrane. In some embodiments, the transmembrane protein nanopore comprises more than one barrel domain, such as two, three or four barrel domains. The two or more barrel domains may each contribute to the channel of the nanopore. Where the transmembrane protein nanopore comprises more than one barrel domain, one of the barrel domains is typically a transmembrane barrel domain and the other barrel domain(s) do not span the membrane.

[0097] If the transmembrane protein nanopore is an oligomeric transmembrane protein nanopore, the subunits of the nanopore typically surround a central axis and contribute strands to a barrel domain, such as a transmembrane barrel domain.

[0098] In some embodiments, the barrel domain is a P-barrel or a a-helix bundle. For example, the barrel domain may be a transmembrane P-barrel or a transmembrane a-helix bundle. Typically, the barrel domain is a P-barrel, such as a transmembrane P-barrel.

[0099] In some embodiments, the nanopore is an oligomeric P-barrel nanopore. In some embodiments the nanopore is an oligomeric P-barrel nanopore comprising a plurality ofmonomer subunits, wherein each monomer subunit contributes one or more P-strands to the P-barrel domain. In some embodiments, each monomer subunit contributes two P-strands to the P-barrel domain. In some embodiments, each monomer subunit contributes no more than two P-strands to the P-barrel domain. In some embodiments, each monomer subunit comprises a first P-strand and a second P-strand that contribute to the P-barrel domain, and the one or more heterologous insertions are in the first P-strand and / or the second P-strand. In some embodiments, each monomer subunit comprises a first P-strand and a second P-strand that contribute to the P-barrel domain, and the one or more heterologous insertions comprise a first heterologous insertion in the first P-strand and a second heterologous insertion in the second P-strand. As described herein, in some embodiments the first heterologous insertion in the first P-strand and the second heterologous insertion in the second P-strand are substantially coincidental.

[0100] The transmembrane protein nanopore for use in the disclosure can be a P-barrel nanopore or an a-helix bundle nanopore. Typically, the transmembrane protein nanopore is a P-barrel nanopore. P-barrel nanopores comprise a barrel domain that is formed from P- strands. Suitable P-barrel nanopores include, but are not limited to, P-toxins, such as a-hemolysin (a-HL / aHL), anthrax toxin and leukocidins, and outer membrane proteins / porins of bacteria, such as Mycobacterium smegmatis porin (Msp), for example MspA, MspB, MspC or MspD, CsgG, outer membrane porin F (OmpF), outer membrane porin G (OmpG), outer membrane phospholipase A, Neisseria autotransporter lipoprotein (NalP), Cytotoxin K (CytK), Vibrio campbellii aHL (VcaHL), Vibrio cholerae Cytolysin (HlyA / VCC) and other nanopores, such as lysenin. In some embodiments, the P-barrel nanopore may be selected from a-HL, such as Staphylococcus aureus a-HL or Vibrio campbellii aHL, aerolysin, epsilon toxin, anthrax toxin, protective antigen pore, a leukocidin, a bacterial outer membrane proteins / porins, such as Mycobacterium smegmatis porin (Msp), for example MspA, MspB, MspC or MspD, Curb specific genes G (CsgG), Clostridium difficile toxin b (CDTb), Cucumaria echinata lectin-III (CEL-III), Clostridium perfringens P-toxin (CPB), Clostridium perfringens iota toxin b, Enterococcus pore-forming toxin 1 (Epxl), Enterococcus poreforming toxin 4 (Epx4), Leukotoxin GH (LukGH), Staphylococcal y-hemolysin, Clostridium perfringens necrotic enteritis B-like toxin (NetB), Pleurotolysin (Ply AB), outer membrane porin F (OmpF), outer membrane porin G (OmpG), outer membrane phospholipase A, Neisseria autotransporter lipoprotein (NalP), Ferrichrome outer membrane transporter(FhuA), Cytotoxin K (CytK), Vibrio cholerae Cytolysin (HlyA / VCC) and lysenin. In some embodiments, the P-barrel nanopore may be selected from a-HL, aerolysin, epsilon toxin, anthrax toxin, protective antigen pore, a leukocidin, a bacterial outer membrane proteins / porins, such as Mycobacterium smegmatis porin (Msp), for example MspA, MspB, MspC or MspD, CsgG, Clostridium difficile toxin b (CDTb), Cucumaria echinata lectin-III (CEL-III), Clostridium perfringens P-toxin (CPB), Clostridium perfringens iota toxin b, Enterococcus pore-forming toxin 1 (Epxl), Enterococcus pore-forming toxin 4 (Epx4), Leukotoxin GH (LukGH), Staphylococcal y-hemolysin, Clostridium perfringens necrotic enteritis B-like toxin (NetB), Pleurotolysin (Ply AB), outer membrane porin F (OmpF), outer membrane porin G (OmpG), outer membrane phospholipase A, Neisseria autotransporter lipoprotein (NalP), Ferrichrome outer membrane transporter (FhuA) and lysenin. a-helix bundle pores comprise a barrel domain that is formed from a-helices. Suitable a-helix bundle nanopores include, but are not limited to, inner membrane proteins and a outer membrane proteins, such as Wza (e.g. see K. R. Mahendran, Nat. Chem. 2016, incorporated by reference) and ClyA toxin. For example, the transmembrane protein nanopore may be derived from or based on Msp, a-hemolysin, lysenin, Phi29, CsgG, CgsF, ClyA, Spl and haemolytic protein fragaceatoxin C (FraC). In some embodiments, the transmembrane protein nanopore may be derived from or based on a-HL such as S. aureus a-HL or V. campbellii aHL, aerolysin, epsilon toxin, anthrax toxin, protective antigen pore, a leukocidin, MspA, MspB, MspC, MspD, CsgF, CsgG, CDTb, CEL-III, CPB, Clostridium perfringens iota toxin b, Epxl, Epx4, LukGH, Staphylococcal y-hemolysin, NetB, Ply AB, OmpF, OmpG, outer membrane phospholipase A, NalP, FhuA, lysenin, ClyA, Wza, SP1, Phi29 portal protein, FraC, CytK and HlyA / VCC. In some embodiments, the transmembrane protein nanopore may be derived from or based on a-HL, aerolysin, epsilon toxin, anthrax toxin, protective antigen pore, a leukocidin, MspA, MspB, MspC, MspD, CsgF, CsgG, CDTb, CEL-III, CPB, Clostridium perfringens iota toxin b, Epxl, Epx4, LukGH, Staphylococcal y-hemolysin, NetB, Ply AB, OmpF, OmpG, outer membrane phospholipase A, NalP, FhuA, lysenin, ClyA, Wza, SP1, Phi29 portal protein and FraC.

[0101] The transmembrane protein nanopore into which the one or more heterologous insertions are inserted may be a wild-type transmembrane protein nanopore (e.g. a wild-type nanopore described herein) or a variant thereof. In some embodiments, the transmembrane protein nanopore into which the one or more heterologous insertions are inserted is a variantof a wild-type transmembrane protein nanopore such as a variant of a transmembrane protein nanopore described herein. In some embodiments, the variant has at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 97%, at least 98% or at least 99% sequence identity to the amino acid sequence of the corresponding wild-type transmembrane protein nanopore. In some embodiments the variant has at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 97%, at least 98% or at least 99% sequence identity to the amino acid sequence of the corresponding wild-type transmembrane protein nanopore, wherein such identity is determined across at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99% or 100% of the amino acid sequence of the wild-type transmembrane protein nanopore.

[0102] In some embodiments, the transmembrane protein nanopore comprising the one or more heterologous insertions further comprises one or more modifications (e.g. mutations) relative to the corresponding wild-type transmembrane protein nanopore. The one or more mutations may be one or more amino acid substitutions, one or more amino acid deletions and / or one or more amino acid insertions other than the one or more heterologous insertions described herein. In some embodiments, the one or more modifications (e.g. mutations) are at a position other than the position of the one or more heterologous insertions. In some embodiments, the transmembrane protein nanopore comprises at least 1, at least 2, at least 3, at least 4, at least 5, at least 10, or more mutations relative to the corresponding wild-type transmembrane protein nanopore, and further comprises the one or more heterologous insertions. In some embodiments, the one or more mutations are in the barrel domain of the nanopore. In some embodiments, the one or more mutations are in a cap domain of the nanopore. In some embodiments, the one or more mutations are at or near a constriction of the nanopore. In some embodiments, the one or more mutations modify one or more properties of the wild-type nanopore, such as modifying the stability, ion selectivity, analyte selectivity, or noise of the nanopore; the size and / or charge of the channel of the nanopore, and / or the interaction of the nanopore with an analyte.

[0103] For example, in some embodiments, the transmembrane protein nanopore into which the one or more heterologous insertions are inserted is a variant of a-hemolysin having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 97%, at least 98% or at least 99% sequence identity to SEQ ID NO: 1. In some embodiments, thetransmembrane protein nanopore into which the one or more heterologous insertions are inserted is a variant of aerolysin having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 97%, at least 98% or at least 99% sequence identity to SEQ ID NO: 72. In some embodiments, the transmembrane protein nanopore into which the one or more heterologous insertions are inserted is a variant of lysenin having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 97%, at least 98% or at least 99% sequence identity to SEQ ID NO: 73. In some embodiments, the transmembrane protein nanopore into which the one or more heterologous insertions are inserted is a variant of CsgG having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 97%, at least 98% or at least 99% sequence identity to SEQ ID NO: 74. In some embodiments, the transmembrane protein nanopore into which the one or more heterologous insertions are inserted is a variant of CytK having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 97%, at least 98% or at least 99% sequence identity to SEQ ID NO: 75. In some embodiments, the transmembrane protein nanopore into which the one or more heterologous insertions are inserted is a variant of Epxl having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 97%, at least 98% or at least 99% sequence identity to SEQ ID NO: 76. In some embodiments, the transmembrane protein nanopore into which the one or more heterologous insertions are inserted is a variant of Epx4 having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 97%, at least 98% or at least 99% sequence identity to SEQ ID NO: 77.

[0104] In some embodiments, the nanopore is derived from a-hemolysin (a-HL). The wild type a-HL pore is formed of seven identical monomers or subunits (i.e. it is heptameric). The sequence of one wild type monomer or subunit of a-hemolysin is shown in SEQ ID NO: 1. Amino acids 1, 7 to 21, 31 to 34, 45 to 51, 63 to 66, 72, 92 to 97, 104 to 111, 124 to 136, 149 to 153, 160 to 164, 173 to 206, 210 to 213, 217, 218, 223 to 228, 236 to 242, 262 to 265, 272 to 274, 287 to 290 and 293 of SEQ ID NO: 1 form loop regions. Residues 111, 113 and 147 of SEQ ID NO: 1 form part of a constriction of the barrel or channel of a-HL. Residues 111 to 128 of SEQ ID NO: 1 and 130 to 147 of SEQ ID NO: 1 form P-strands, said strands from each monomer form the transmembrane P-barrel of a-hemolysin.SEQ ID NO: 1 a-hemolysin ADSDINIKTGTTDIGSNTTVKTGDLVTYDKENGMHKKVFYSFIDDKNHNKKLLVIRT KGTIAGQYRVYSEEGANKSGLAWPSAFKVQLQLPDNEVAQISDYYPRNSIDTKEYMS TLTYGFNGNVTGDDTGKIGGLIGANVSIGHTLKYVQPDFKTILESPTDKKVGWKVIFN NMVNQNWGPYDRDSWNPVYGNQLFMKTRNGSMKAADNFLDPNKASSLLSSGFSPD FATVITMDRKASKQQTNIDVIYERVRDDYQLHWTSTNWKGTNTKDKWTDRSSERYK IDWEKEEMTN

[0105] In some embodiments, the length of the channel of the transmembrane protein nanopore is extended by at least about 0.56 nm. In other words, the transmembrane protein nanopore comprising one or more heterologous insertions as disclosed herein comprises a channel that is at least about 0.56 nm longer than the channel of an otherwise identical transmembrane protein nanopore that does not comprise the heterologous insertion. The otherwise identical transmembrane protein nanopore that does not comprise the heterologous insertion may be a wild-type transmembrane protein nanopore or a variant thereof. The length of the channel is measured as the distance from the first opening to the second opening. In some embodiments, the length of the channel of the transmembrane protein nanopore is extended by at least about 1 nm. In some embodiments, the length of the channel of the transmembrane protein nanopore is extended by at least about 2 nm, by at least about 3 nm, by at least about 4 nm, by at least about 5 nm, by at least about 6 nm, by at least about 7 nm, by at least about 8 nm, by at least about 9 nm or by at least about 10 nm.

[0106] Without wishing to be bound by theory, the inventors believe that there is not necessarily an upper limit to the length of the channel extended by the heterologous insertion, provided that the transmembrane protein nanopore retains the ability to form a solvent-accessible channel having a first opening and a second opening. For example, in the case of transmembrane protein nanopores comprising a P-barrel domain and a cap domain, it is believed (without being bound by theory) that the cap domain drives the initial folding of the P-barrel domain in the membrane, and the folding of the rest of the P-barrel domain propagates from the initial section of the P-barrel. This is exemplified for a-hemolysin in the Examples of the application, in which a heterologous insertion extended the length of the channel of the a-hemolysin pore from 5.2 nm to 27.6 nm whilst retaining the ability to fold.The folding of other P-barrel domain-containing nanopores follows a similar pathway and is expected to similarly lack constraints as to the upper limit of the length of the channel.

[0107] In some embodiments, the length of the channel of the transmembrane protein nanopore is extended by about 0.56 nm to about 50 nm, such as by about 1 nm to about 50 nm, by about 1 nm to about 25 nm, by about 2 nm to about 20 nm, by about 3 nm to about 10 nm. In some embodiments, the length of the nanopore is extended by about 0.56 nm, by about 1 nm, by about 2 nm, by about 3 nm, by about 4 nm, by about 5 nm, by about 6 nm, by about 7 nm, by about 8 nm, by about 9 nm, by about 10 nm, by about 11 nm, by about 12 nm, by about 13 nm, by about 14 nm, by about 15 nm, by about 16 nm, by about 17 nm, by about 18 nm, by about 19 nm or by about 20 nm.

[0108] In some embodiments, the total length of the channel of the transmembrane protein nanopore comprising the heterologous insertion is at least 5 nm, at least 6 nm, at least 7 nm, at least 8 nm, at least 9 nm, at least 10 nm, at least 10.5 nm, at least 10.6 nm, at least 11 nm, at least 12 nm, at least 13 nm, at least 14 nm, at least 15 nm, at least 16 nm, at least 17 nm, at least 18 nm, at least 19 nm or at least 20 nm. In some embodiments, the total length of the channel of the transmembrane protein nanopore comprising the heterologous insertion is at most about 50 nm, such as at most about 40 nm, at most about 30 nm, at most about 25nm, at most about 20 nm or at most about 15 nm. In some embodiments, the total length of the channel of the transmembrane protein nanopore comprising the heterologous insertion is about 5nm to about 50 nm, such as about 5 nm to about 40 nm, about 5 nm to about 30 nm, about 6 nm to about 25 nm, about 7 nm to about 20 nm, about 8 nm to about 15 nm, or about 9 nm to about 12 nm.

[0109] Typically, the one or more heterologous insertions extend the length of the barrel domain, such as a P-barrel domain. The barrel domain comprises a first opening, a second opening and a channel therebetween, and the embodiments disclosed herein related to the extended and / or total length of the channel of the nanopore equally apply to the extended and / or total length of the channel of the barrel domain. For example, the disclosure may relate to a transmembrane protein nanopore comprising a barrel domain, wherein the barrel domain has a first opening, a second opening and a channel therebetween, wherein the nanopore comprises one or more heterologous insertions in the barrel domain that extend the length of the channel, and related methods. In some embodiments, the total length of the channel of the barrel domain comprising the heterologous insertion is at least 5 nm, at least 6nm, at least 7 nm, at least 8 nm, at least 8.5 nm, at least 9 nm, at least 10 nm, at least 10.5 nm, at least 10.6 nm, at least 11 nm, at least 12 nm, at least 13 nm, at least 14 nm, at least 15 nm, at least 16 nm, at least 17 nm, at least 18 nm, at least 19 nm or at least 20 nm. In some embodiments, the total length of the channel of the barrel domain comprising the heterologous insertion is at most about 50 nm, such as at most about 40 nm, at most about 30 nm, at most about 25nm, at most about 20 nm or at most about 15 nm. In some embodiments, the total length of the channel of the barrel domain comprising the heterologous insertion is about 5nm to about 50 nm, such as about 5 nm to about 40 nm, about 5 nm to about 30 nm, about 6 nm to about 25 nm, about 7 nm to about 20 nm, about 8 nm to about 15 nm, or about 9 nm to about 12 nm.

[0110] As described above, and without being bound by theory, the inventors believe that the extended length of the channel increases the analyte interaction dwell time with the nanopore, thereby improving the characterisation of the analyte. In some embodiments, the measurements characteristic of the analyte are improved due to increased analyte interaction dwell time, increased percentage of residual currents, decreased conductance, increased resistance, an increased length of the sensing region of the channel, and / or an increased volume of the channel.

[0111] In some embodiments, the transmembrane protein nanopore has an increased analyte dwell time relative to a corresponding nanopore lacking the heterologous insertion. For example, where the transmembrane protein nanopore is selected from a-HL such as S. aureus a-HL or V. campbellii aHL, aerolysin, epsilon toxin, anthrax toxin, protective antigen pore, a leukocidin, MspA, MspB, MspC, MspD, CsgF, CsgG, CDTb, CEL-III, CPB, Clostridium perfringens iota toxin b, Epxl, Epx4, LukGH, Staphylococcal y-hemolysin, NetB, Ply AB, OmpF, OmpG, outer membrane phospholipase A, NalP, FhuA, ly senin, ClyA, Wza, SP1, Phi29 portal protein, FraC, CytK and HlyA / VCC, and which comprises one or more heterologous insertions, the analyte dwell time is increased relative to a-HL such as S. aureus a-HL or V. campbellii aHL, aerolysin, epsilon toxin, anthrax toxin, protective antigen pore, a leukocidin, MspA, MspB, MspC, MspD, CsgF, CsgG, CDTb, CEL-III, CPB, Clostridium perfringens iota toxin b, Epxl, Epx4, LukGH, Staphylococcal y-hemolysin, NetB, Ply AB, OmpF, OmpG, outer membrane phospholipase A, NalP, FhuA, ly senin, ClyA, Wza, SP1, Phi29 portal protein, FraC, CytK and HlyA / VCC, respectively. In some embodiments, the analyte interaction dwell time is increased by at least 2-fold, as described herein.In some embodiments, the transmembrane protein nanopore has an increased analyte dwell time relative to a corresponding nanopore lacking the heterologous insertion. For example, where the transmembrane protein nanopore is selected from a-HL, aerolysin, epsilon toxin, anthrax toxin, protective antigen pore, a leukocidin, MspA, MspB, MspC, MspD, CsgF, CsgG, CDTb, CEL-III, CPB, Clostridium perfringens iota toxin b, Epxl, Epx4, LukGH, Staphylococcal y-hemolysin, NetB, Ply AB, OmpF, OmpG, outer membrane phospholipase A, NalP, FhuA, lysenin, ClyA, Wza, SP1, Phi29 portal protein and FraC, and which comprises one or more heterologous insertions, the analyte dwell time is increased relative to a-HL, aerolysin, epsilon toxin, anthrax toxin, protective antigen pore, a leukocidin, MspA, MspB, MspC, MspD, CsgF, CsgG, CDTb, CEL-III, CPB, Clostridium perfringens iota toxin b, Epxl, Epx4, LukGH, Staphylococcal y-hemolysin, NetB, Ply AB, OmpF, OmpG, outer membrane phospholipase A, NalP, FhuA, lysenin, ClyA, Wza, SP1, Phi29 portal protein and FraC, respectively. In some embodiments, the analyte interaction dwell time is increased by at least 2-fold, as described herein.

[0112] In some embodiments, the transmembrane protein nanopore has an increased length of the sensing region relative to a corresponding nanopore lacking the heterologous insertion. For example, where the transmembrane protein nanopore is selected from a-HL such as S. aureus a-HL or V. campbellii aHL, aerolysin, epsilon toxin, anthrax toxin, protective antigen pore, a leukocidin, MspA, MspB, MspC, MspD, CsgF, CsgG, CDTb, CEL-III, CPB, Clostridium perfringens iota toxin b, Epxl, Epx4, LukGH, Staphylococcal y-hemolysin, NetB, Ply AB, OmpF, OmpG, outer membrane phospholipase A, NalP, FhuA, lysenin, ClyA, Wza, SP1, Phi29 portal protein, FraC, CytK and HlyA / VCC, and which comprises one or more heterologous insertions, the length of the sensing region is increased relative to a-HL such as S. aureus a-HL or V. campbellii aHL, aerolysin, epsilon toxin, anthrax toxin, protective antigen pore, a leukocidin, MspA, MspB, MspC, MspD, CsgF, CsgG, CDTb, CEL-III, CPB, Clostridium perfringens iota toxin b, Epxl, Epx4, LukGH, Staphylococcal y-hemolysin, NetB, Ply AB, OmpF, OmpG, outer membrane phospholipase A, NalP, FhuA, lysenin, ClyA, Wza, SP1, Phi29 portal protein, FraC, CytK and HlyA / VCC, respectively. For example, the cysteine modification reactions in Figure 3 demonstrate that, along the extended beta barrel, analytes may be covalently sensed based on the perturbations to the current levels. Accordingly, in some embodiments, the one or more heterologous insertions introduce a reactive group, such as a cysteine residue into the channel of the nanopore. In someembodiments, the one or more heterologous insertions introduce a constriction or reader head into the channel of the nanopore.

[0113] In some embodiments, the transmembrane protein nanopore has an increased length of the sensing region relative to a corresponding nanopore lacking the heterologous insertion. For example, where the transmembrane protein nanopore is selected from a-HL, aerolysin, epsilon toxin, anthrax toxin, protective antigen pore, a leukocidin, MspA, MspB, MspC, MspD, CsgF, CsgG, CDTb, CEL-III, CPB, Clostridium perfringens iota toxin b, Epxl, Epx4, LukGH, Staphylococcal y-hemolysin, NetB, Ply AB, OmpF, OmpG, outer membrane phospholipase A, NalP, FhuA, lysenin, ClyA, Wza, SP1, Phi29 portal protein and FraC, and which comprises one or more heterologous insertions, the length of the sensing region is increased relative to a-HL, aerolysin, epsilon toxin, anthrax toxin, protective antigen pore, a leukocidin, MspA, MspB, MspC, MspD, CsgF, CsgG, CDTb, CEL-III, CPB, Clostridium perfringens iota toxin b, Epxl, Epx4, LukGH, Staphylococcal y-hemolysin, NetB, Ply AB, OmpF, OmpG, outer membrane phospholipase A, NalP, FhuA, lysenin, ClyA, Wza, SP1, Phi29 portal protein and FraC, respectively. For example, the cysteine modification reactions in Figure 3 demonstrate that, along the extended beta barrel, analytes may be covalently sensed based on the perturbations to the current levels. Accordingly, in some embodiments, the one or more heterologous insertions introduce a reactive group, such as a cysteine residue into the channel of the nanopore. In some embodiments, the one or more heterologous insertions introduce a constriction or reader head into the channel of the nanopore.

[0114] In some embodiments, the channel of the transmembrane protein nanopore has an increased volume relative to a corresponding nanopore lacking the heterologous insertion. For example, where the transmembrane protein nanopore is selected from a-HL such as S. aureus a-HL or V. campbellii aHL, aerolysin, epsilon toxin, anthrax toxin, protective antigen pore, a leukocidin, MspA, MspB, MspC, MspD, CsgF, CsgG, CDTb, CEL-III, CPB, Clostridium perfringens iota toxin b, Epxl, Epx4, LukGH, Staphylococcal y-hemolysin, NetB, Ply AB, OmpF, OmpG, outer membrane phospholipase A, NalP, FhuA, lysenin, ClyA, Wza, SP1, Phi29 portal protein, FraC, CytK and HlyA / VCC, and which comprises one or more heterologous insertions, the volume of the channel is increased relative to a-HL such as S. aureus a-HL or V. campbellii aHL, aerolysin, epsilon toxin, anthrax toxin, protective antigen pore, a leukocidin, MspA, MspB, MspC, MspD, CsgF, CsgG, CDTb, CEL-III, CPB, Clostridium perfringens iota toxin b, Epxl, Epx4, LukGH, Staphylococcal y-hemolysin,NetB, Ply AB, OmpF, OmpG, outer membrane phospholipase A, NalP, FhuA, lysenin, ClyA, Wza, SP1, Phi29 portal protein, FraC, CytK and HlyA / VCC, respectively.

[0115] In some embodiments, the channel of the transmembrane protein nanopore has an increased volume relative to a corresponding nanopore lacking the heterologous insertion. For example, where the transmembrane protein nanopore is selected from a-HL, aerolysin, epsilon toxin, anthrax toxin, protective antigen pore, a leukocidin, MspA, MspB, MspC, MspD, CsgF, CsgG, CDTb, CEL-III, CPB, Clostridium perfringens iota toxin b, Epxl, Epx4, LukGH, Staphylococcal y-hemolysin, NetB, Ply AB, OmpF, OmpG, outer membrane phospholipase A, NalP, FhuA, lysenin, ClyA, Wza, SP1, Phi29 portal protein and FraC, and which comprises one or more heterologous insertions, the volume of the channel is increased relative to a-HL, aerolysin, epsilon toxin, anthrax toxin, protective antigen pore, a leukocidin, MspA, MspB, MspC, MspD, CsgF, CsgG, CDTb, CEL-III, CPB, Clostridium perfringens iota toxin b, Epxl, Epx4, LukGH, Staphylococcal y-hemolysin, NetB, Ply AB, OmpF, OmpG, outer membrane phospholipase A, NalP, FhuA, lysenin, ClyA, Wza, SP1, Phi29 portal protein and FraC, respectively.

[0116] In some embodiments, the transmembrane protein nanopore comprises a modification that modifies the net charge of the channel of the nanopore. In some embodiments, the one or more heterologous insertions modify the net charge of the channel of the nanopore. The modified net charge of the channel of the nanopore may be relative to the corresponding wildtype nanopore from which the nanopore of the disclosure is derived. Said modifications may be useful to improve the capture and movement of an analyte with respect to the nanopore. In some embodiments, the transmembrane protein nanopore comprises a modification that increases the net negative charge of the nanopore. The net negative charge of the nanopore may be increased, for example, through the substitution of a positively-charged or neutral amino acid with a negatively charged amino acid, the substitution of a positively-charged amino acid with a neutral amino acid, and / or the insertion of a negatively-charged amino acid. In some embodiments, the transmembrane protein nanopore comprises a modification that increases the net positive charge of the nanopore. The net positive charge of the nanopore may be increased, for example, through the substitution of a negatively-charged or neutral amino acid with a positively-charged amino acid, the substitution of a negatively -charged amino acid with a neutral amino acid, and / or the insertion of a positively-charged amino acid. In some embodiments, the transmembrane protein nanopore comprises the substitution of oneor more amino acids at a constriction of the channel to increase or decrease the net charge of the channel. In some embodiments, the transmembrane protein nanopore comprises the substitution of one or more amino acid residues at positions corresponding to residues 113 and / or 115 of SEQ ID NO: 1 with a positively charged amino acid, such as arginine, lysine or histidine. In some embodiments, the transmembrane protein nanopore comprises the substitution of one or more amino acid residues at a position corresponding to residues 113 and / or 115 of SEQ ID NO: 1 with arginine.

[0117] In some embodiments, the transmembrane protein nanopore may comprise a peptide tag such as a poly-aspartate tag and / or a polyhistidine tag, e.g. a D8tag, a H6tag or a D8H6tag. Such tags may be useful in purification of the monomer.

[0118] Heterologous insertion

[0119] The transmembrane protein nanopore disclosed herein comprises one or more heterologous insertions. The one or more heterologous insertions extend the length of the channel of the transmembrane protein nanopore.

[0120] The one or more heterologous insertions may be at the first opening, the second opening, and / or at the channel of the transmembrane protein nanopore comprising the one or more heterologous insertions. A heterologous insertion at an opening of the transmembrane protein nanopore means that the heterologous insertion forms at least part of, or the entirety of the opening. A heterologous insertion at the channel of the transmembrane nanopore means that the heterologous insertion forms at least part of, or the entirety of, the channel that runs through the assembled nanopore. In some embodiments a heterologous insertion at the channel of a transmembrane nanopore may not form at least part of the first or second opening. In some embodiments a heterologous insertion may comprise a portion of an opening of the assembled nanopore and a portion of the channel of the assembled nanopore.

[0121] As described herein, the transmembrane protein nanopore into which the one or more heterologous insertions are inserted may be a wild-type transmembrane protein nanopore or a variant thereof. In embodiments wherein the transmembrane protein nanopore is a variant of a wild-type transmembrane protein nanopore, the one or more heterologous insertions may be at a position in the variant corresponding to the equivalent position in the wild-type transmembrane protein nanopore.In some embodiments, the transmembrane protein nanopore comprises a barrel domain, such as a P-barrel domain, and the one or more heterologous insertions are made in the barrel domain. The transmembrane protein nanopore may comprise one or more further heterologous insertions in a part of the nanopore other than the barrel domain.

[0122] In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 111 and 128 of SEQ ID NO: 1 and / or between residues 130 and 147 of SEQ ID NO: 1. Residues 111 to 128 of SEQ ID NO: 1 form a first P-strand of the a-hemolysin monomer, and residues 130 to 147 of SEQ ID NO: 1 form a second P-strand of the a-hemolysin monomer. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 111 and 128 of SEQ ID NO: 1, and one or more heterologous insertions at a position corresponding to between residues 130 and 147 of SEQ ID NO: 1.

[0123] In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions in a membrane-exposed region of the nanopore. The term

[0124] “ membrane-exposed region" refers to a region of a nanopore, which in the otherwise identical transmembrane protein nanopore that does not comprise a heterologous insertion, comprises outwards-facing residues that are exposed to or contacts a membrane. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions in a membrane-exposed region of a barrel domain. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions in a membrane-exposed region of a P-barrel domain. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 117 and 125 of SEQ ID NO: 1 and / or between residues 134 and 142 of SEQ ID NO: 1. Residues 117 to 125 of SEQ ID NO: 1 form the membrane-exposed region of the first P-strand of the a-hemolysin monomer, and residues 134 to 142 of SEQ ID NO: 1 form the membrane-exposed region of the second P-strand of the a-hemolysin monomer. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 117 and 125 of SEQ ID NO: 1, and one or more heterologous insertions at a position corresponding to between residues 134 and 142 of SEQ ID NO: 1. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to betweenresidues 122 and 123 of SEQ ID NO: 1 and / or between residues 136 and 137 of SEQ ID NO: 1. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 118 and 119 of SEQ ID NO: 1 and / or between residues 139 and 140 of SEQ ID NO: 1.

[0125] In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions in a solvent-exposed region of the nanopore. The term “ solvent-exposed region" refers to a region of a nanopore, which in the otherwise identical transmembrane protein nanopore that does not comprise a heterologous insertion, comprises outward-facing residues that are exposed to or contact solvent, i.e. as opposed to contacting a membrane. Typically, a heterologous insertion in a solvent-exposed region of the nanopore is solvent-exposed itself, i.e. comprises outwards facing residues that are exposed to or contact the solvent. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions in a solvent-exposed region of a barrel domain. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions in a solvent-exposed region of a P-barrel domain. The solvent-exposed region may be either side of the membrane. For example, the solvent-exposed region may be on the cis side of the nanopore, i.e. the side of the nanopore that initially contacts an analyte, and / or on the trans side of the nanopore, i.e. other the side of the membrane to the cis side. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 111 and 117 of SEQ ID NO: 1 and / or between residues 125 and 134 of SEQ ID NO: 1 and / or between residues 142 and 147 of SEQ ID NO: 1. Residues 111 to 117 of SEQ ID NO: 1 form the cis solvent-exposed region of the first P-strand of the a-hemolysin monomer, residues 125 to 134 of SEQ ID NO: 1 form the trans solvent-exposed region of the a-hemolysin monomer, and residues 142 to 147 of SEQ ID NO: 1 form the cis solvent-exposed region of the second P-strand of the a-hemolysin monomer. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 111 and 117 of SEQ ID NO: 1, and / or one or more heterologous insertions at a position corresponding to between residues 142 and 147 of SEQ ID NO: 1. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 111 and 117 of SEQ ID NO: 1, and one or more heterologous insertions at a position corresponding to between residues 142 and 147 of SEQ ID NO: 1. In someembodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 125 and 134 of SEQ ID NO: 1. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 125 and 129 of SEQ ID NO: 1, and one or more heterologous insertions at a position corresponding to between residues 129 and 134 of SEQ ID NO: 1. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 116 and 117 of SEQ ID NO: 1, and / or one or more heterologous insertions at a position corresponding to between residues 142 and 143 of SEQ ID NO: 1. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 127 and 128 of SEQ ID NO: 1, and / or one or more heterologous insertions at a position corresponding to between residues 130 and 131 of SEQ ID NO: 1.

[0126] In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions in a membrane-exposed region of the nanopore and one or more heterologous insertions in a solvent-exposed region of the nanopore.

[0127] In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 215 and 249 of SEQ ID NO: 72 and / or between residues 249 and 282 of SEQ ID NO: 72. Residues 215 to 249 of SEQ ID NO: 72 form a first β-strand of the aerolysin monomer, and residues 249 to 282 of SEQ ID NO: 72 form a second β-strand of the aerolysin monomer. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 215 and 249 of SEQ ID NO: 72, and one or more heterologous insertions at a position corresponding to between residues 249 and 282 of SEQ ID NO: 72.

[0128] In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 215 and 249 of SEQ ID NO: 72 and / or between residues 249 and 281 of SEQ ID NO: 72. Residues 215 to 249 of SEQ ID NO: 72 form a first β-strand of the aerolysin monomer, and residues 249 to 281 of SEQ ID NO: 72 form a second β-strand of the aerolysin monomer. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 215 and 249 of SEQ ID NO: 72, and one or moreheterologous insertions at a position corresponding to between residues 249 and 281 of SEQ ID NO: 72.

[0129] SEQ ID NO: 72 - aerolysin AEPVYPDQLRLFSLGQGVCGDKYRPVNREEAQSVKSNIVGMMGQWQISGLAN GWVIMGPGYNGEIKPGTASNTWCYPTNPVTGEIPTLSALDIPDGDEVDVQWRLVHDS ANFIKPTSYLAHYLGYAWVGGNHSQYVGEDMDVTRDGDGWVIRGNNDGGCDGYR CGDKTAIKVSNFAYNLDPDSFKHGDVTQSDRQLVKTVVGWAVNDSDTPQSGYDVTL RYDTATNWSKTNTYGLSEKVTTKNKFKWPLVGETELSIEIAANQSWASQNGGSTTTS LSQSVRPTVPARSKIPVKIELYKADISYPYEFKADVSYDLTLSGFLRWGGNAWYTHPD NRPNWNHTFVIGPYKDKASSIRYQWDKRYIPGEVKWWDWNWTIQQNGLSTMQNNL ARVLRPVRAGITGDFSAESQFAGNIEIGAPVPLAADSKVRRARSVDGAGQGLRLEIPL DAQELSGLGFNNVSLSVTPAANQ

[0130] In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions in a membrane-exposed region of the nanopore. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 232 and 243 of SEQ ID NO: 72 and / or between residues 255 and 265 of SEQ ID NO: 72. Residues 232 to 243 of SEQ ID NO: 72 form the membrane-exposed region of the first β-strand of the aerolysin monomer, and residues 255 to 265 of SEQ ID NO: 72 form the membrane-exposed region of the second β-strand of the aerolysin monomer. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 232 and 243 of SEQ ID NO: 72, and one or more heterologous insertions at a position corresponding to between residues 232 to 243 of SEQ ID NO: 72.

[0131] In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions in a solvent-exposed region of the nanopore. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 215 and 232 of SEQ ID NO: 72 and / or between residues 243 and 255 of SEQ ID NO: 72 and / or between residues 265 and 282 of SEQ ID NO: 72. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 215 and 232 of SEQID NO: 72 and / or between residues 243 and 255 of SEQ ID NO: 72 and / or between residues 265 and 281 of SEQ ID NO: 72. Residues 215 to 232 of SEQ ID NO: 72 may form the cis solvent-exposed region of the first β-strand of the aerolysin monomer, residues 243 to 255 of SEQ ID NO: 72 may form the trans solvent-exposed region of the aerolysin monomer, and residues 265 to 282 of SEQ ID NO: 72 may form the cis solvent-exposed region of the second β-strand of the aerolysin monomer. Residues 215 to 232 of SEQ ID NO: 72 form the cis solvent-exposed region of the first β-strand of the aerolysin monomer, residues 243 to 255 of SEQ ID NO: 72 form the trans solvent-exposed region of the aerolysin monomer, and residues 265 to 281 of SEQ ID NO: 72 form the cis solvent-exposed region of the second β-strand of the aerolysin monomer. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 215 and 232 of SEQ ID NO: 72, and / or one or more heterologous insertions at a position corresponding to between residues 265 and 282 of SEQ ID NO: 72. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 215 and 232 of SEQ ID NO: 72, and / or one or more heterologous insertions at a position corresponding to between residues 265 and 281 of SEQ ID NO: 72. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 215 and 232 of SEQ ID NO: 72, and one or more heterologous insertions at a position corresponding to between residues 265 and 282 of SEQ ID NO: 72. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 215 and 232 of SEQ ID NO: 72, and one or more heterologous insertions at a position corresponding to between residues 265 and 281 of SEQ ID NO: 72. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 243 and 255 of SEQ ID NO: 72. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 243 and 249 of SEQ ID NO: 72, and one or more heterologous insertions at a position corresponding to between residues 249 and 255 of SEQ ID NO: 72.

[0132] In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 231 and 232 of SEQID NO: 72, and one or more heterologous insertions at a position corresponding to between residues 265 and 266 of SEQ ID NO: 72.

[0133] In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 34 and 70 of SEQ ID NO: 73 and / or between residues 70 and 106 of SEQ ID NO: 73. Residues 34 to 70 of SEQ ID NO: 73 form a first β-strand of the lysenin monomer, and residues 70 to 106 of SEQ ID NO: 73 form a second β-strand of the lysenin monomer. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 34 and 70 of SEQ ID NO: 73, and one or more heterologous insertions at a position corresponding to between residues 70 and 106 of SEQ ID NO: 73.

[0134] SEQ ID NO: 73 - lysenin MSAKAAEGYEQIEVDVVAVWKEGYVYENRGSTSVDQKITITKGMKNVNSETR TVTATHSIGSTISTGDAFEIGSVEVSYSHSHEESQVSMTETEVYESKVIEHTITIPPTSKF TRWQLNADVGGADIEYMYLIDEVTPIGGTQSIPQVITSRAKIIVGRQIILGKTEIRIKHA ERKEYMTVVSRKSWPAATLGHSKLFKFVLYEDWGGFRIKTLNTMYSGYEYAYSSDQ GGIYFDQGTDNPKQRWAINKSLPLRHGDVVTFMNKYFTRSGLCYDDGPATNVYCLD KREDKWILEVVG

[0135] In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions in a membrane-exposed region of the nanopore. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 58 and 65 of SEQ ID NO: 73 and / or between residues 75 and 81 of SEQ ID NO: 73. Residues 58 to 65 of SEQ ID NO: 73 form the membrane-exposed region of the first β-strand of the lysenin monomer, and residues 75 to 81 of SEQ ID NO: 73 form the membrane-exposed region of the second β-strand of the lysenin monomer. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 58 and 65 of SEQ ID NO: 73, and one or more heterologous insertions at a position corresponding to between residues 75 and 81 of SEQ ID NO: 73.In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions in a solvent-exposed region of the nanopore. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 34 and 58 of SEQ ID NO: 73 and / or between residues 65 and 75 of SEQ ID NO: 73 and / or between residues 81 and 106 of SEQ ID NO: 73. Residues 34 to 58 of SEQ ID NO: 73 form the cis solvent-exposed region of the first β-strand of the lysenin monomer, residues 65 to 75 of SEQ ID NO: 73 form the trans solvent-exposed region of the lysenin monomer, and residues 81 to 106 of SEQ ID NO: 73 form the cis solvent-exposed region of the second β-strand of the lysenin monomer. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 34 and 58 of SEQ ID NO: 73, and / or one or more heterologous insertions at a position corresponding to between residues 81 and 106 of SEQ ID NO: 73. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 34 and 58 of SEQ ID NO: 73, and one or more heterologous insertions at a position corresponding to between residues 81 and 106 of SEQ ID NO: 73. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 65 and 75 of SEQ ID NO: 73. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 65 and 70 of SEQ ID NO: 73, and one or more heterologous insertions at a position corresponding to between residues 70 and 75 of SEQ ID NO: 73.

[0136] In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 56 and 57 of SEQ ID NO: 73, and one or more heterologous insertions at a position corresponding to between residues 83 and 84 of SEQ ID NO: 73.

[0137] In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 135 and 144 of SEQ ID NO: 74 and / or between residues 145 and 154 of SEQ ID NO: 74 and / or between residues 182 and 196 of SEQ ID NO: 74 and / or between residues 196 and 210 of SEQ ID NO: 74. These amino acid positions form β-strand 1-4 of the CsgG monomer, respectively. In some embodiments, the transmembrane protein nanopore comprises one or more heterologousinsertions at a position corresponding to between residues 135 and 144 of SEQ ID NO: 74, one or more heterologous insertions at a position corresponding to between residues 145 and 154 of SEQ ID NO: 74, one or more heterologous insertions at a position corresponding to between residues 182 and 196 of SEQ ID NO: 74 and one or more heterologous insertions at a position corresponding to between residues 196 and 210 of SEQ ID NO: 74.

[0138] SEQ ID NO: 74 - CsgG CLTAPPKEAARPTLMPRAQSYKDLTHLPAPTGKIFVSVYNIQDETGQFKPYPAS NFSTAVPQSATAMLVTALKDSRWFIPLERQGLQNLLNERKIIRAAQENGTVAINNRIP LQSLTAANIMVEGSIIGYESNVKSGGVGARYFGIGADTQYQLDQIAVNLRVVNVSTG EILSSVNTSKTILSYEVQAGVFRFIDYQRLLEGEVGYTSNEPVMLCLMSAIETGVIFLIN DGIDRGLWDLQNKAERQNDILVKYRHMSVPPES

[0139] In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions in a membrane-exposed region of the nanopore. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 136 and 142 of SEQ ID NO: 74 and / or between residues 147 and 153 of SEQ ID NO: 74 and / or between residues 183 and 190 of SEQ ID NO: 74 and / or between residues 202 and 209 of SEQ ID NO: 74. These amino acid positions form the membrane-exposed regions of β-strands 1-4 of the CsgG monomer, respectively. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 136 and 142 of SEQ ID NO: 74, one or more heterologous insertions at a position corresponding to between residues 147 and 153 of SEQ ID NO: 74, one or more heterologous insertions at a position corresponding to between residues 183 and 190 of SEQ ID NO: 74 and one or more heterologous insertions at a position corresponding to between residues 202 and 209 of SEQ ID NO: 74.

[0140] In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions in a solvent-exposed region of the nanopore. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 135 and 136 of SEQ ID NO: 74 and / or between residues 153 and 154 of SEQ ID NO: 74 and / or between residues 182 and 183 of SEQ IDNO: 74 and / or between residues 209 and 210 of SEQ ID NO: 74. These amino acid positions form the cis solvent-exposed region of β-strands 1-4 of the CsgG monomer, respectively. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions on each of the cis solvent-exposed region of β-strands 1-4 of the CsgG monomer. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 143 and 144 of SEQ ID NO: 74 and / or between residues 145 and 146 of SEQ ID NO: 74 and / or between residues 191 and 196 of SEQ ID NO: 74 and / or between residues 196 and 201 of SEQ ID NO: 74. These amino acid positions form the trans solvent-exposed region of β-strands 1-4 of the CsgG monomer, respectively. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions on each of the trans solvent-exposed region of β-strands 1-4 of the CsgG monomer.

[0141] In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 137 and 138 of SEQ ID NO: 74 and / or between residues 151 and 152 of SEQ ID NO: 74 and / or between residues 182 and 183 of SEQ ID NO: 74 and / or between residues 209 and 210 of SEQ ID NO: 74.

[0142] In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 143 and 165 of SEQ ID NO: 75 and / or between residues 165 and 187 of SEQ ID NO: 75. Residues 143 to 165 of SEQ ID NO: 75 form a first β-strand of the CytK monomer, and residues 165 to 187 of SEQ ID NO: 75 form a second β-strand of the CytK monomer. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 143 and 165 of SEQ ID NO: 75, and one or more heterologous insertions at a position corresponding to between residues 165 and 187 of SEQ ID NO: 75.

[0143] SEQ ID NO: 75 - CytK MKRSKTYVKCLALSAVLASSALAMHTPVVAAQTTSQVVTDIGQNAKTHTSYN TFNNEQADNMTMSLKVTFIDDPSADKQIAVINTTGSFMKANPTLSDAPVDGYPIPGAS VTLRYPSQYDIAMNLQDNTSRFFHVAPTNAVEETTVTSSVSYQLGGSIKASVTPSGPS GESGATGQVTWSDSVSYKQTSYKTNLIDQTNKHVKWNVFFNGYNNQNWGIYTRDS YHALYGNQLFMYSRTYPHETDARGNLVPMNDLPALTNSGFSPGMIAVVISEKDTEQSSIQVAYTKHADDYTLRPGFTFGTGNWVGNNIKDVDQKTFNKSFVLDWKNKKLVEK K

[0144] (signal peptide underlined, not present in mature pore sequence)

[0145] In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions in a membrane-exposed region of the nanopore. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 152 and 161 of SEQ ID NO: 75 and / or between residues 169 and 178 of SEQ ID NO: 75. Residues 152 to 161 of SEQ ID NO: 75 form the membrane-exposed region of the first β-strand of the CytK monomer, and residues 169 to 178 of SEQ ID NO: 75 form the membrane-exposed region of the second β-strand of the CytK monomer. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 152 and 161 of SEQ ID NO: 75, and one or more heterologous insertions at a position corresponding to between residues 169 and 178 of SEQ ID NO: 75.

[0146] In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions in a solvent-exposed region of the nanopore. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 143 and 151 of SEQ ID NO: 75 and / or between residues 179 and 187 of SEQ ID NO: 75 and / or between residues 162 and 168 of SEQ ID NO: 75. Residues 143 to 151 of SEQ ID NO: 75 form the cis solvent-exposed region of the first β-strand of the CytK monomer, residues 162 to 168 of SEQ ID NO: 75 form the trans solvent-exposed region of the CytK monomer, and residues 179 to 187 of SEQ ID NO: 75 form the cis solvent-exposed region of the second β-strand of the CytK monomer. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 143 and 151 of SEQ ID NO: 75, and / or one or more heterologous insertions at a position corresponding to between residues 179 and 187 of SEQ ID NO: 75. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 143 and 151 of SEQ ID NO: 75, and one or more heterologous insertions at a position corresponding to between residues 179 and 187 of SEQ ID NO: 75. In some embodiments, the transmembrane protein nanopore comprises one or more heterologousinsertions at a position corresponding to between residues 162 to 168 of SEQ ID NO: 75. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 162 and 165 of SEQ ID NO: 75, and one or more heterologous insertions at a position corresponding to between residues 165 and 168 of SEQ ID NO: 75.

[0147] In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 150 and 151 of SEQ ID NO: 75, and one or more heterologous insertions at a position corresponding to between residues 177 and 178 of SEQ ID NO: 75.

[0148] In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 137 and 156 of SEQ ID NO: 76 and / or between residues 157 and 177 of SEQ ID NO: 76. Residues 137 and 156 of SEQ ID NO: 76 form a first β-strand of the Epx1 monomer, and residues 157 and 177 of SEQ ID NO: 76 form a second β-strand of the Epx1 monomer. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 137 and 156 of SEQ ID NO: 76, and one or more heterologous insertions at a position corresponding to between residues 157 and 177 of SEQ ID NO: 76.

[0149] SEQ ID NO: 76 - Epxl DVQVFADNPVLGTEVTFEKDENGRIVKIISKNQYRITIYNSVDAGDTPNDASVSL DVDFIDDKNSGEMGAVASINTFIPSGLRYVEGYKYKGVTNPIYKNLSAGMLWPKKYR VEVVNIPIDQATKIITATPNNNIKEKQVSDTISYGFGGSVSADGKKPGGSIEANLAYTK TTTYDQPDYETSQIKKTTKEAVWDTSFVETRDGYTPNSWNPVYGNQMFMRGRYSN VSPIDNIKKGGEVSSLISGGFSPKMGVVLASPNGTKKSQFVVRVSRMSDMYIMRWSG TEWGGENEINQNVPKEYNALMYEDVKFEIDWEQRTVRTILE

[0150] In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions in a membrane-exposed region of the nanopore. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 143 and 155 of SEQ ID NO: 76 and / or between residues 158 and 169 of SEQ ID NO: 76. Residues 143 to 155 of SEQ IDNO: 76 form the membrane-exposed region of the first β-strand of the Epx1 monomer, and residues 158 to 169 of SEQ ID NO: 76 form the membrane-exposed region of the second β-strand of the Epx1 monomer. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 143 and 155 of SEQ ID NO: 76, and one or more heterologous insertions at a position corresponding to between residues 158 and 169 of SEQ ID NO: 76.

[0151] In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions in a solvent-exposed region of the nanopore. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 137 and 142 of SEQ ID NO: 76 and / or between residues 170 and 177 of SEQ ID NO: 76 and / or between residues 155 and 158 of SEQ ID NO: 73. Residues 137 to 142 of SEQ ID NO: 76 form the cis solvent-exposed region of the first β-strand of the Epx1 monomer, residues 155 to 158 of SEQ ID NO: 76 form the trans solvent-exposed region of the Epx1 monomer, and residues 170 to 177 of SEQ ID NO: 76 form the cis solvent-exposed region of the second β-strand of the Epx1 monomer. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 137 and 142 of SEQ ID NO: 76 and / or one or more heterologous insertions at a position corresponding to between residues 170 and 177 of SEQ ID NO: 76. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 170 and 177 of SEQ ID NO: 76, and one or more heterologous insertions at a position corresponding to between residues 137 and 142 of SEQ ID NO: 76. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 155 and 158 of SEQ ID NO: 76. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 155 and 156 of SEQ ID NO: 76, and one or more heterologous insertions at a position corresponding to between residues 157 and 158 of SEQ ID NO: 76.

[0152] In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 144 and 145 of SEQ ID NO: 76, and one or more heterologous insertions at a position corresponding to between residues 169 and 170 of SEQ ID NO: 76.In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 122 and 141 of SEQ ID NO: 77 and / or between residues 142 and 163 of SEQ ID NO: 77. Residues 122 to 141 of SEQ ID NO: 77 form a first P-strand of the Epx4 monomer, and residues 142 to 163 of SEQ ID NO: 77 form a second P-strand of the Epx4 monomer. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 122 and 141 of SEQ ID NO: 77, and one or more heterologous insertions at a position corresponding to between residues 142 and 163 of SEQ ID NO: 77.

[0153] SEQ ID NO: 77 - Epx4

[0154] SEDNIIGTTTQEIDEHGNVKTIITVKNQQIESYTSTDSGTAKNRSTLTVNANFLND KYSNELTTILSLNGFIPSGRKFIFPKNNTLKGEMLWPQRYSTAVYNIPLDKSVKITNSTP DNTIRSKEVSNSITYGIGGGIKMEGKQPGANLDANAAITKTISYQQPDYETAKTTSTVT GVNWNTNFTETRDGYTRNSWNPVYGNQMFMYGRYTSNIRNNFTPDYQLSSLITSGF SPSYGLVLRAPKDVKKSRIKVVFARRSETYQQNWDGLNWWGRNFYDTKNPDSLSK VTLTFELDWQNHRVTFIELE

[0155] In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions in a membrane-exposed region of the nanopore. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 129 and 140 of SEQ ID NO: 77 and / or between residues 143 and 154 of SEQ ID NO: 77. Residues 129 to 140 of SEQ ID NO: 77 form the membrane-exposed region of the first P-strand of the Epx4 monomer, and residues 143 to 154 of SEQ ID NO: 77 form the membrane-exposed region of the second P-strand of the Epx4 monomer. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 129 and 140 of SEQ ID NO: 77, and one or more heterologous insertions at a position corresponding to between residues 143 and 154 of SEQ ID NO: 77.

[0156] In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions in a solvent-exposed region of the nanopore. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at aposition corresponding to between residues 122 and 128 of SEQ ID NO: 77 and / or between residues 155 and 163 of SEQ ID NO: 77 and / or between residues 140 and 143 of SEQ ID NO: 77. Residues 122 to 128 of SEQ ID NO: 77 form the cis solvent-exposed region of the first P-strand of the Epx4 monomer, residues 140 to 143 of SEQ ID NO: 77 form the trans solvent-exposed region of the Epx4 monomer, and residues 155 to 163 of SEQ ID NO: 77 form the cis solvent-exposed region of the second P-strand of the Epx4 monomer. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 122 and 128 of SEQ ID NO: 77, and / or one or more heterologous insertions at a position corresponding to between residues 155 and 163 of SEQ ID NO: 77. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 122 and 128 of SEQ ID NO: 77, and one or more heterologous insertions at a position corresponding to between residues 155 and 163 of SEQ ID NO: 77. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 140 and 143 of SEQ ID NO: 77. In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 140 and 141 of SEQ ID NO: 77, and one or more heterologous insertions at a position corresponding to between residues 142 and 143 of SEQ ID NO: 77.

[0157] In some embodiments, the transmembrane protein nanopore comprises one or more heterologous insertions at a position corresponding to between residues 129 and 130 of SEQ ID NO: 77, and one or more heterologous insertions at a position corresponding to between residues 154 and 155 of SEQ ID NO: 77.

[0158] In some embodiments, the transmembrane protein nanopore comprises a plurality of heterologous insertions that are substantially coincidental. The term “substantially coincidental” means that each of the heterologous insertions of the plurality form contacts with at least one other heterologous insertion of the plurality. Typically, if the plurality of heterologous insertions are in a barrel domain, such as a P-barrel domain, the plurality of heterologous insertions are substantially planar, substantially parallel with the membrane and / or form a ring around the channel. In some embodiments, the transmembrane protein nanopore comprises more than one plurality of heterologous insertions that are each substantially coincidental, such as two, three, four or five pluralities of heterologousinsertions that are each substantially coincidental. For example, the transmembrane protein nanopore may comprise one plurality of heterologous insertions that are substantially coincidental in a membrane-exposed region of the nanopore and one plurality of heterologous insertions that are substantially coincidental in a solvent-exposed region of the nanopore.

[0159] In some embodiments, the transmembrane protein nanopore comprises a plurality of heterologous insertions that form antiparallel backbone hydrogen bonds in the nanopore. In some embodiments, the plurality of heterologous insertions are each in a P-strand of a P-barrel domain of the nanopore and the heterologous insertions form antiparallel backbone hydrogen bonds in the nanopore. The antiparallel backbone hydrogen bonds may be with an endogenous sequence of the nanopore. The term "endogenous' refers to an otherwise identical nanopore that does not comprise the one or more heterologous insertions. The antiparallel backbone hydrogen bonds may be with another heterologous insertion. As such, the heterologous insertions typically extend the length (e.g. the number of amino acids or monomer units) of a P-strand, and extends the length of the barrel domain as a whole (i.e. the distance from the first opening to the second opening of the barrel domain).

[0160] In some embodiments, the transmembrane protein nanopore comprises a plurality of heterologous insertions that are substantially coincidental, wherein the heterologous insertions form antiparallel backbone hydrogen bonds in the nanopore. The heterologous insertions typically form antiparallel backbone hydrogen bonds with the other heterologous insertions in the nanopore and / or with an endogenous sequence of the nanopore. In some embodiments, the transmembrane protein nanopore comprises two or more, such as three, four or five pluralities of heterologous insertions, wherein the heterologous insertions of each plurality are substantially coincidental and form antiparallel backbone hydrogen bonds in the nanopore.

[0161] The transmembrane protein nanopore may comprise two or more constrictions within the channel. A constriction within the channel may sometimes be referred to as a reader head or a sensing region. In some embodiments, the one or more heterologous insertions are positioned between the two or more constrictions. Without being bound by theory, the inventors believe that the heterologous insertions may be used to control, e.g. increase, the distance between two or more constructions within a channel. An analyte at a constriction is typically capable of being characterised, for example, by measuring the modulation of a current signal. The presence of two constrictions may permit the measurement of twocharacteristics of the analyte. By controlling the distance between the two or more constrictions, the measurement of the characteristics of the analyte may be improved.

[0162] In some embodiments, the one or more heterologous insertions define a constriction within the channel of the nanopore. In other words, the one or more heterologous insertions may form a constriction within the channel of the nanopore.

[0163] In some embodiments, the transmembrane protein nanopore comprises one or more constrictions within the channel and the one or more heterologous insertions are positioned between the one or more constrictions and the cis side of the transmembrane protein nanopore, e.g. the side of the transmembrane protein nanopore that initially contacts the analyte. This may be particularly advantageous for analytes, such as P-cyclodextrin, that may not fully translocate through the nanopore.

[0164] The one or more heterologous insertions may be any structural feature capable of extending the length of the channel of the nanopore.

[0165] In some embodiments, the one or more heterologous insertions may be one or more heterologous amino acid sequence insertions. In some embodiments, the one or more heterologous insertions may be one or more heterologous non-amino acid insertions. The one or more heterologous insertions may be one or more synthetic insertions. The one or more synthetic insertions may comprise or consist of an alkyl chain, a lipid, a polyethylene glycol (PEG), a fatty acid, a saccharide, a polysaccharide, a nucleotide, a polynucleotide, a nonnatural amino acid, a hydroxamate, a phosphonate, a thioacid, or a hydroxyl acid. The one or more synthetic insertions may comprise a polymer, such as a PEG, a polysaccharide, a polynucleotide, a polypeptide comprising a non-natural amino acid, or a depsipeptide

[0166] A transmembrane protein nanopore comprising one or more heterologous amino acids may be produced using any suitable means known to the skilled person. A transmembrane protein nanopore comprising one or more heterologous amino acid sequence insertions may be produced by modifying a polynucleotide encoding a transmembrane protein nanopore such that it encodes a transmembrane protein nanopore comprising the one or more heterologous amino acid sequence insertions. The incorporation of heterologous amino acids may also be achieved using genetic code expansion techniques, wherein non-canonical amino acids are introduced during translation by employing engineered tRNA-synthetase pairs and codons specifically designed for this purpose. Post-translational modifications may be used tointroduce synthetic moieties or non-canonical amino acids, including through chemical or enzymatic derivatization of reactive residues. A transmembrane protein nanopore comprising one or more synthetic insertions may be assembled by protein semi synthesis, e.g. by native chemical ligation or expressed protein ligation.

[0167] In some embodiments, the one or more heterologous insertions comprise a polymer, i.e. comprising a plurality of monomer units. The one or more heterologous insertions may each independently comprise at least 2, at least 6, at least 8, at least 10, at least 12, at least 14, at least 16, at least 18, at least 20, at least 24, at least 28, at least 32, at least 36, at least 40, at least 44, at least 48, at least 52, at least 56, at least 60, at least 70 or at least 80 monomer units. As described above, there is not necessarily a limit to the size of the heterologous insertion. In some embodiments, the one or more heterologous insertions independently comprise about 1 to about 100 monomer units, such as about 2 to about 80, about 4 to about 80, about 6 to about 70, about 8 to about 60, about 12 to about 56, about 12 to about 32, about 16 to about 32, or about 20 to about 32 monomer units. In some embodiments, the one or more heterologous insertions independently comprise about 4, about 6, about 8, about 10, about 12, about 14, about 16, about 18, about 20, about 24, about 28, about 32, about 36, about 40, about 44, about 48, about 52, about 56, about 60, about 70 or about 80 monomer units. In some embodiments, the one or more heterologous insertions independently comprise at most about 80 monomer units, such as at most about 70, at most about 60 or at most about 56 monomer units. In some embodiments, each of the one or more heterologous insertions is the same length. In some embodiments, each of the one or more heterologous insertions comprises the same number of monomer units. In embodiments in which the nanopore comprises two or more pluralities of heterologous insertions, as described herein, the heterologous insertions of each plurality may be the same length as each other heterologous insertion in the plurality.

[0168] The transmembrane protein nanopore may comprise two or more heterologous insertions. The two or more heterologous insertions may be the same, e.g. comprise the same chemical structure. The two or more heterologous insertions may be different.

[0169] The one or more heterologous insertions are typically one or more heterologous amino acid sequence insertions. In other words, the transmembrane protein nanopore as disclosed herein typically comprises one or more heterologous amino acid sequence insertions that extend the length of the channel. In some embodiments, the one or more heterologous aminoacid sequence insertions each comprise at least one amino acid. Typically, the one or more heterologous amino acid sequence insertions each independently comprise at least two amino acids. In some embodiments, the one or more heterologous amino acid sequence insertions independently comprise at least 4, at least 6, at least 8, at least 10, at least 12, at least 14, at least 16, at least 18, at least 20, at least 24, at least 28, at least 32, at least 36, at least 40, at least 44, at least 48, at least 52, at least 56, at least 60, at least 70 or at least 80 amino acids. As described above, there is not necessarily a limit to the size of the heterologous amino acid sequence insertion. In some embodiments, the one or more heterologous amino acid sequence insertions each independently comprise about 1 to about 100 amino acids, such as about 2 to about 80, about 4 to about 80, about 6 to about 70, about 8 to about 60, about 12 to about 56, about 12 to about 32, about 16 to about 32, or about 20 to about 32 amino acids. In some embodiments, the one or more heterologous amino acid sequence insertions each independently comprise about 4, about 6, about 8, about 10, about 12, about 14, about 16, about 18, about 20, about 24, about 28, about 32, about 36, about 40, about 44, about 48, about 52, about 56, about 60, about 70 or about 80 amino acids. In some embodiments, the one or more heterologous amino acid sequence insertions each independently comprise at most about 80 amino acids, such as at most about 70, at most about 60 or at most about 56 amino acids.

[0170] In some embodiments, each of the one or more heterologous amino acid sequence insertions is the same length. In some embodiments, each of the one or more heterologous amino acid sequence insertions each comprise the same number of amino acids. In some embodiments, the transmembrane protein nanopore comprises two or more pluralities of heterologous amino acid sequence insertions, wherein the pluralities are as described herein, and wherein the heterologous amino acid sequence insertions of each plurality may be the same length as each other heterologous insertion in the plurality. For example, each of the one or more heterologous sequence insertions in a first plurality may be 12 amino acids in length, and each of the one or more heterologous sequence insertions in a second plurality may be 20 amino acids in length.

[0171] The one or more heterologous amino acid sequence insertions may comprise natural and / or non-natural amino acids. In some embodiments, the one or more heterologous amino acid sequence insertions may comprise natural amino acids. Typically, the one or more heterologous amino acid sequence insertions each comprise at least two natural and / or non-natural amino acids. In some embodiments, the one or more heterologous amino acid sequence insertions each comprise at least two different natural and / or non-natural amino acids.

[0172] In some embodiments, the one or more heterologous amino acid sequence insertions each comprise a sequence motif having a formula (X-Y)nor (Y-X)n. X and Y may be any amino acid. In some embodiments, each X is an amino acid independently selected from the group consisting of R / K / N / D / E / Q / H / P / Y / W / S / T / G / A / M / C / F / L / V / I. In some embodiments, each X is an amino acid independently selected from the group consisting of K / N / Y / S / T / G / L / V. In some embodiments, each Y is an amino acid independently selected from the group consisting of K / N / D / E / Q / H / P / S / T / G / A / M. In some embodiments, each Y is an amino acid independently selected from the group consisting of K / N / E / S / T / G / M. In some embodiments, each X is an amino acid independently selected from the group consisting of R / K / N / D / E / Q / H / P / Y / W / S / T / G / A / M / C / F / L / V / I and each Y is an amino acid independently selected from the group consisting of K / N / D / E / Q / H / P / S / T / G / A / M. In some embodiments, each X is an amino acid independently selected from the group consisting of K / N / Y / S / T / G / L / V and each Y is an amino acid independently selected from the group consisting of K / N / E / S / T / G / M. n is at least 1. As described above, there is not necessarily a limit to the size of the heterologous amino acid sequence insertion. In some embodiments, n is at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 12, at least 14, at least 16, at least 18, at least 20, at least 22, at least 24, at least 26, at least 28, at least 30, at least 35 or at least 40. In some embodiments, n is about 1 to about 50, such as about 1 to about 40, about 2 to about 40, about 3 to about 35, about 4 to about 30, about 6 to about 14, about 6 to about 16, about 8 to about 16, or about 10 to about 16. In some embodiments, n is about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 12, about 14, about 16, about 18, about 20, about 22, about 24, about 26, about 28, about 30, about 350 or about 40. In some embodiments, n is at most about 40, such as at most about 35, at most about 30 or at most about 28. In embodiments in which n is at least two, each X may be independently selected and / or each Y may be independent selected. For example, the sequence motif (X-Y)2 may be GNTN, and the sequence motif (X-Y)4 may be EYMSTLTY.

[0173] In some embodiments, the one or more heterologous amino acid sequence insertions each independently comprise a sequence motif selected from GNTN, GNTNEN, KNGNTN,GG, SG, SV, EYMSTLTY, VSIGHTLK, TN, ST, SI, TNGN, KNGN, GNEN, TNKN, ENTN, GNSN, NGNT, NKNGNT, (GNTN)3-GNTIGITN-(GNTN)3, (GNTN)2-GITNGITN-(GNTN)4, (GNTN)3-GITNGITN-(GNTN)3, (GNTN)4-GITNGITN-(GNTN)8, (GNTN)6-GITNGITN-(GNTN)6, (GNTN)6-GITIGITN-(GNTN)6, GNTHGHTN-(GNTN)6, (GNTN)6-GHTHGNTN, (GNTN)2-EN-(GNTN)2, (GNTN)2-KN-(GNTN)2, (GNTNEN)4-GNTIEIGITNEN-(GNTNEN)4, (KNGNTN)4-KNGNTIKIGITN-(KNGNTN)4, TNGK, TNGD, TNGKTNGN, TNGDTNGN, (GNTN)3G, NTN(GNTN)2, GQSQ, GQTQ, NT, ND and NDNT. In some embodiments, the heterologous insertion comprises a motif described herein which is 8 amino acids in length or less (such as 6 amino acids in length or less, 4 amino acids in length or less, or 2 amino acids in length) multiple times. In some embodiments, the heterologous insertion consists of a motif described herein which is 10 amino acids in length or more, such as 12 amino acids in length or more, or 16 amino acids in length or more).

[0174] In some embodiments, the one or more heterologous amino acid sequence insertions each independently comprise a sequence motif selected from GNTN, GNTNEN, KNGNTN, GG, SG, SV, EYMSTLTY, VSIGHTLK, TN, ST, SI, TNGN, KNGN, GNEN, TNKN, ENTN, GNSN, NGNT, TNGK, TNGD, TNGKTNGN, TNGDTNGN, GQSQ, GQTQ, NT, ND and NDNT.

[0175] In some embodiments, the one or more heterologous amino acid sequence insertions each independently comprise a sequence motif selected from GNTN, GNTNEN, KNGNTN, GG, SG, SV, EYMSTLTY, VSIGHTLK, TN, ST, SI, TNGN, KNGN, GNEN, TNKN, ENTN, GNSN, NGNT and NKNGNT.

[0176] In some embodiments, the one or more heterologous amino acid sequence insertions each independently comprise one or more repeats of the sequence motifs described herein. In some embodiments, the one or more heterologous amino acid sequence insertions each independently comprise a plurality of repeated sequence motifs. In some embodiments, the one or more heterologous amino acid sequence insertions each independently comprise at least two, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 12, at least 14, at least 16, at least 18, at least 20, at least 22, at least 24, at least 26, at least 28, at least 30, at least 35 or at least 40 repeats of the sequence motifs described herein. In some embodiments, the one or more heterologous amino acid sequence insertions each independently comprise 1 to about 50, such as about 1 to about 40, about 2 to about 40, about3 to about 35, about 4 to about 30, about 6 to about 14, about 2 to about 14, about 3 to about 8, about 6 to about 16, about 8 to about 16, or about 10 to about 16 repeats of the sequence motifs described herein. In some embodiments, the one or more heterologous amino acid sequence insertions each independently comprise about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 12, about 14, about 16, about 18, about 20, about 22, about 24, about 26, about 28, about 30, about 35 or about 40 repeats of the sequence motifs described herein. In some embodiments, n is at most about 40, such as at most about 35, at most about 30 or at most about 28 repeats of the sequence motifs described herein.

[0177] In some embodiments, the one or more heterologous amino acid insertions are each independently selected from GNTN, (GNTN)2, (GNTN)3, (GNTN)4, (GNTN)s, (GNTN)8, (GNTN)14, (GNTN)2O, (GNTN)2-EN-(GNTN)2, (GNTN)2-KN-(GNTN)2, (GNTNEN)2, (KNGNTN)2, (GNTNEN)3, (KNGNTN)3, (GNTNEN)1O, (KNGNTN)io, GG, GGGG, GGGGGG, GGGGGGGG, SG, SGSG, SGSGSG, SGSGSGSG, SV, SVSV, SVSVSV, SVSVSVSV, EYMSTLTY, VSIGHTLK, (EYMSTLTY)2, (VSIGHTLK)2, (EYMSTLTY)s, (VSIGHTLK)5, TNTNTNTNTNTN, STSTSTSTSTST, SISISISISISI, (TNGN)3, (KNGN)3, (GNEN)3, (TNKN)3, (ENTN)3, GNSN, (GNSN)2, (GNSN)3, (GNSN)4, (GNSN)5, (GNSN)8, (GNTN)6, (GNTN)7, (GNTN)9, (GNTN)1O, (GNTN)n, (GNTN)I2, (GNTN)I3, NGNT, NKNGNT, (NKNGNT)2, (GNTN)3-GNTIGITN-(GNTN)3, (GNTN)2-GITNGITN-(GNTN)4, (GNTN)3-GITNGITN-(GNTN)3, (GNTN)4-GITNGITN-(GNTN)8, (GNTN)6-GITNGITN-(GNTN)6, (GNTN)6-GITIGITN-(GNTN)6, (GNTN)8, (GNTN)6-GHTHGNTN, (GNTN)6-GHTHGNTN, GNTHGHTN-(GNTN)6, GNTHGHTN-(GNTN)6, (GNTNEN)4-GNTIEIGITNEN-(GNTNEN)4, (KNGNTN)4-KNGNTIKIGITN-(KNGNTN)4, (TNGK)4(TNGK)8, (TNGK)IO, (TNGK)14, (TNGKTNGN)4, (TNGKTNGN)S, (TNGKTNGN)?, (TNGD)4, (TNGD)8, (TNGD)IO, (TNGD)i4, (TNGDTNGN)4, (TNGDTNGN)s, (TNGDTNGN)7, (GNTN)3G, (GNTN)3G, NTN(GNTN)2, (GQSQ)8, (GQTQ)8, (NGNT)2, (TNGN)4, and (TNGN)s. In some embodiments, the one or more heterologous amino acid insertions are each independently selected from GNTN, (GNTN)2, (GNTN)3, (GNTN)4, (GNTN)s, (GNTN)8, (GNTN)14, (GNTN)20, (GNTN)2-EN-(GNTN)2, (GNTN)2-KN-(GNTN)2, (GNTNEN)2, (KNGNTN)2, (GNTNEN)3, (KNGNTN)3, (GNTNEN)IO, (KNGNTN)io, GG, GGGG, GGGGGG, GGGGGGGG, SG, SGSG, SGSGSG, SGSGSGSG, SV, SVSV, SVSVSV, SVSVSVSV, EYMSTLTY, VSIGHTLK, (EYMSTLTY)2, (VSIGHTLK)2, (EYMSTLTY)5, (VSIGHTLK)5, TNTNTNTNTNTN, STSTSTSTSTST, SISISISISISI,(TNGN)3, (KNGN)3, (GNEN)3, (TNKN)3, (ENTN)3, GNSN, (GNSN)2, (GNSN)3, (GNSN)4, (GNSN)5, (GNSN)8, (GNTN)6, (GNTN)7, (GNTN)9, (GNTN)io, (GNTN)n, (GNTN)n, (GNTN)I3, NGNT, NKNGNT and (NKNGNT)2. In some embodiments, the one or more heterologous amino acid insertions are selected from a row of Table 3. In some embodiments, the one or more heterologous amino acid insertions are selected from a row of Table 4.

[0178] In some embodiments, the one or more heterologous amino acid sequence insertions comprise predominantly hydrophilic amino acids. In some embodiments, the one or more heterologous amino acid sequence insertions comprise at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% hydrophilic amino acids, e.g. Gly, Asn, Thr, Ser, Gin, Lys, Arg, Asp, Glu and His. In some embodiments, the one or more heterologous amino acid sequence insertions do not comprise more than 50%, more than 40%, more than 30%, more than 20%, or more than 10% hydrophobic amino acids, e.g. Vai, He, Leu, Met, Phe, Trp, and Tyr.

[0179] In some embodiments, the transmembrane protein nanopore does not comprise a deletion and / or substitution in the amino acid sequence of the unextended nanopore. In some embodiments, the one or more heterologous amino acid insertions are inserted into the complete amino acid sequence of the unextended transmembrane protein nanopore.

[0180] In some embodiments the transmembrane protein nanopore does not comprise a deletion and / or substitution in the amino acid sequence of the unextended nanopore at the position of the one or more heterologous amino acid insertions. In some embodiments, the transmembrane protein nanopore does not comprise an amino acid deletion within 4 amino acids of the position of the one or more heterologous amino acid insertions. In some embodiments, the transmembrane protein nanopore does not comprise any amino acid deletions within 8 amino acids of the position of the one or more heterologous amino acid insertions. In some embodiments, the transmembrane protein nanopore does not comprise any amino acid deletions. The presence of an amino acid deletion and / or substitution may be determined by comparing the sequence of the extended transmembrane protein nanopore to the sequence of the transmembrane protein nanopore from which it is derived, i.e. prior to modification with the one or more heterologous amino acid insertions.

[0181] In some embodiments, the transmembrane protein nanopore comprises a P-barrel and the one or more heterologous amino acid insertions comprise a sequence from a P-barrel domain. The sequence from the P-barrel domain may be at least 10 amino acids in length,such as at least 12, at least 16, at least 20, or at least 24 amino acids in length. The P-barrel domain may be from the P-barrel domain of the transmembrane protein pore. The P-barrel domain may from the P-barrel domain of a different pore-forming protein, as described herein. Typically, the protein from which the sequence of the P-barrel domain is derived comprises the same number of monomers that form the pore as the unextended transmembrane protein nanopore. In some embodiments, the protein from which the sequence of the P-barrel domain is derived comprises a channel having the about the same diameter as the channel of the transmembrane protein nanopore.

[0182] In some embodiments, the transmembrane protein nanopore is a chimeric transmembrane protein pore. In some embodiments, the transmembrane protein nanopore is not a chimeric transmembrane protein pore. In some embodiments a chimeric transmembrane nanopore comprises a heterologous insertion sequence which is derived from a different transmembrane protein pore. In some embodiments a chimeric transmembrane nanopore comprises a heterologous insertion sequence which is at least 10 amino acids in length, such as at least 12, at least 16, at least 20, or at least 24 amino acids in length, and is derived from a different transmembrane protein nanopore. In some embodiments a heterologous insertion sequence derived from a different transmembrane protein nanopore may have an amino acid sequence that is identical to or substantially identical to an amino acid sequence of the different transmembrane protein nanopore. In some embodiments a chimeric transmembrane nanopore comprises a heterologous insertion sequence which is at least 10 amino acids in length, such as at least 12, at least 16, at least 20, or at least 24 amino acids in length and the heterologous insertion sequence is identical to or substantially identical to a corresponding sequence in a different transmembrane protein nanopore. Thus, a chimeric transmembrane nanopore may comprise a P-barrel domain, and comprise one or more heterologous insertions of at least 10 amino acids in length, such as at least 12, at least 16, at least 20, or at least 24 amino acids in length, of a heterologous P-barrel domain. As above, in some embodiments the transmembrane protein nanopore is not a chimeric transmembrane nanopore as described herein. Thus, in some embodiments wherein the transmembrane protein nanopore comprises a P-barrel domain, in some embodiments the one or more heterologous insertions do not comprise a sequence of at least 10 amino acids in length, such as at least 12, at least 16, at least 20, or at least 24 amino acids in length, of a heterologous P-barrel domain.In some embodiments, the one or more heterologous sequences are one or more synthetic sequences.

[0183] In some embodiments where the one or more heterologous amino acid sequence insertions are in a barrel domain, such as a P-barrel domain, each of the heterologous amino acid sequence insertions are an even number of amino acids in length. This is particularly important for insertions within a native P-strand of a nanopore. An even number of amino acids in the insertion maintains the alternating outward / inward facing pattern of the native P-strand of the nanopore.

[0184] As described above, the transmembrane protein nanopore may be an oligomeric nanopore comprising a plurality of monomer subunits. In some embodiments, each monomer subunit comprises one or more heterologous insertions. The transmembrane protein nanopore typically comprises a barrel domain. The number of heterologous insertions typically corresponds to the number of secondary structure features that each monomer contributes to the barrel domain. For example, if each monomer contributes two a-helices to an a-helical bundle barrel domain, each monomer typically comprises two heterologous insertions or a multiple thereof, e.g. one or more heterologous insertions per a-helix. If each monomer contributes two P-strands to a P-barrel domain, each monomer typically comprises two heterologous insertions or a multiple thereof, e.g. one or more heterologous insertions per P-strand.

[0185] In some embodiments, the nanopore is a homooligomeric nanopore, wherein each monomer subunit comprises the same one or more heterologous insertions. In some embodiments, each monomer subunit comprises one or more heterologous insertions, wherein each heterologous insertion is the same. In some embodiments, each monomer subunit comprises two or more heterologous insertions, wherein at least one of the heterologous insertions is different to another of the heterologous insertions.

[0186] In some embodiments, the nanopore is a heterooligomeric nanopore. At least one monomer of the heterooligomeric nanopore comprises one or more heterologous insertions. In some embodiments, each monomer of the heterooligomeric nanopore comprises one or more heterologous insertions. In some embodiments, each monomer of the heterooligomeric nanopore comprises one or more heterologous insertions, and:at least one monomer comprises a different number of heterologous insertions to the other monomers;

[0187] each monomer comprises the same number of heterologous insertions;

[0188] each monomer comprises one or more heterologous insertions of the same length; each monomer comprises the same number of heterologous insertions, and the heterologous insertions are of the same length;

[0189] at least one monomer comprises a heterologous insertion which is different to the heterologous insertion of the other monomers; and / or

[0190] at least one monomer comprises a heterologous insertion which is in a different position to the other monomers.

[0191] Membrane

[0192] The transmembrane protein nanopore is typically present in a membrane.

[0193] Any suitable membranes may be used and such membranes are well-known in the art. The membrane is typically an amphiphilic layer. An amphiphilic layer is a layer formed from amphiphilic molecules, such as phospholipids, which have both at least one hydrophilic portion and at least one lipophilic or hydrophobic portion. The amphiphilic layer may be a monolayer or a bilayer. The amphiphilic molecules may be synthetic or naturally occurring. Non-naturally occurring amphiphiles and amphiphiles which form a monolayer are known in the art and include, for example, block copolymers (Gonzalez-Perez et al., Langmuir, 2009, 25, 10447-10450).

[0194] In some embodiments the membrane comprises one or more archaebacterial bipolar tetraether lipids or mimics thereof. Such lipids are generally found in extremophiles such as that survive in harsh biological environments, thermophiles, halophiles and acidophiles. Their stability is believed to derive from the fused nature of the final bilayer. It is straightforward to construct block copolymer materials that mimic these biological entities by creating a triblock polymer that has the general motif hydrophilic-hydrophobic-hydrophilic. This material may form monomeric membranes that behave similarly to lipid bilayers and encompass a range of phase behaviours from vesicles through to laminar membranes.

[0195] Block copolymers are polymeric materials in which two or more monomer sub-units polymerized together create a single polymer chain. Block copolymers typically have properties that are contributed by each monomer sub-unit. However, a block copolymer mayhave unique properties that polymers formed from the individual sub-units do not possess. Block copolymers can be engineered such that one of the monomer sub-units is hydrophobic (i.e. lipophilic), whilst the other sub-unit(s) are hydrophilic whilst in aqueous media. In this case, the block copolymer may possess amphiphilic properties and may form a structure that mimics a biological membrane. The block copolymer may be a diblock (consisting of two monomer sub-units), but may also be constructed from more than two monomer sub-units to form more complex arrangements that behave as amphipiles. The copolymer may be a triblock, tetrablock or pentablock copolymer. Typically the copolymer is a triblock copolymer comprising two monomer subunits A and B in an A-B-A pattern; typically the A monomer subunit is hydrophilic and the B subunit is hydrophobic.

[0196] The amphiphilic layer is typically a planar lipid bilayer or a supported bilayer.

[0197] The amphiphilic layer is typically a lipid bilayer. Lipid bilayers are models of cell membranes and serve as excellent platforms for a range of experimental studies. For example, lipid bilayers can be used for in vitro investigation of membrane proteins by singlechannel recording. Alternatively, lipid bilayers can be used as biosensors to detect the presence of a range of substances. The lipid bilayer may be any lipid bilayer. Suitable lipid bilayers include, but are not limited to, a planar lipid bilayer, a supported bilayer or a liposome. The lipid bilayer is usually a planar lipid bilayer. Suitable lipid bilayers are disclosed in WO 2008 / 102121, WO 2009 / 077734 and WO 2006 / 100484).

[0198] Any lipid composition that forms a lipid bilayer may be used. Lipids typically comprise a head group, an interfacial moiety and two hydrophobic tail groups which may be the same or different. Suitable head groups include, but are not limited to, neutral head groups, such as diacylglycerides (DG) and ceramides (CM); zwitterionic head groups, such as phosphatidylcholine (PC), phosphatidylethanolamine (PE) and sphingomyelin (SM); negatively charged head groups, such as phosphatidylglycerol (PG); phosphatidylserine (PS), phosphatidylinositol (PI), phosphatic acid (PA) and cardiolipin (CA); and positively charged headgroups, such as trimethylammonium-Propane (TAP). Suitable interfacial moieties include, but are not limited to, naturally-occurring interfacial moieties, such as glycerol-based or ceramide-based moieties. Suitable hydrophobic tail groups include, but are not limited to, saturated hydrocarbon chains, such as lauric acid (n-Dodecanolic acid), myristic acid (n-Tetradecononic acid), palmitic acid (n-Hexadecanoic acid), stearic acid (n-Octadecanoic) and arachidic (n-Eicosanoic); unsaturated hydrocarbon chains, such as oleic acid (cis-9-Octadecanoic); and branched hydrocarbon chains, such as phytanoyl. The length of the chain and the position and number of the double bonds in the unsaturated hydrocarbon chains can vary. The length of the chains and the position and number of the branches, such as methyl groups, in the branched hydrocarbon chains can vary. The hydrophobic tail groups can be linked to the interfacial moiety as an ether or an ester. The lipids may be mycolic acid.

[0199] The lipids can also be chemically-modified. The head group or the tail group of the lipids may be chemically-modified. Suitable lipids whose head groups have been chemically-modified include, but are not limited to, PEG-modified lipids, such as 1,2-Diacyl-sn-Glycero-3-Phosphoethanolamine-N -[Methoxy(Polyethylene glycol)-2000]; functionalised PEG Lipids, such as 1,2-Distearoyl-sn-Glycero-3-Phosphoethanolamine-N-[Biotinyl(Polyethylene Glycol)2000]; and lipids modified for conjugation, such as 1,2-Dioleoyl-sn-Glycero-3-Phosphoethanolamine-N-(succinyl) and 1,2-Dipalmitoyl-sn-Glycero-3-Phosphoethanolamine-N-(Biotinyl). Suitable lipids whose tail groups have been chemically-modified include, but are not limited to, polymerisable lipids, such as 1,2-bis(10,12-tricosadiynoyl)-sn-Glycero-3-Phosphocholine; fluorinated lipids, such as 1-Palmitoyl-2-(16-Fluoropalmitoyl)-sn-Glycero-3-Phosphocholine; deuterated lipids, such as 1,2-Dipalmitoyl-D62-sn-Glycero-3-Phosphocholine; and ether linked lipids, such as 1,2-Di-O-phytanyl-sn-Glycero-3-Phosphocholine. The lipids may be chemically-modified or functionalised.

[0200] Other components that affect the properties of the amphiphilic layer may be incorporated, such as fatty acids, such as palmitic acid, myristic acid and oleic acid; fatty alcohols, such as palmitic alcohol, myristic alcohol and oleic alcohol; sterols, such as cholesterol, ergosterol, lanosterol, sitosterol and stigmasterol; lysophospholipids, such as 1-Acyl-2-Hydroxy-sn- Glycero-3 -Phosphocholine; and ceramides.

[0201] Methods for forming lipid bilayers are known in the art. Suitable methods are disclosed in the Example. Lipid bilayers are commonly formed by the method of Montal and Mueller (Proc. Natl. Acad. Sci. USA., 1972; 69: 3561-3566), in which a lipid monolayer is carried on aqueous solution / air interface past either side of an aperture which is perpendicular to that interface. The method of Montal & Mueller is popular because it is a cost-effective and relatively straightforward method of forming good quality lipid bilayers that are suitable for protein pore insertion. Other common methods of bilayer formation include tip-dipping, painting bilayers and patch-clamping of liposome bilayers. The lipid bilayer may be formed as described in WO 2009 / 077734. A lipid bilayer may also be a droplet interface bilayerformed between two or more aqueous droplets each comprising a lipid shell such that when the droplets are contacted a lipid bilayer is formed at the interface of the droplets.

[0202] In some embodiments, the nanopore may be present in an amphiphilic membrane or layer within a solid state layer, for instance within a hole, well, gap, channel, trench or slit within the solid state layer. Suitable systems are disclosed in WO 2009 / 020682 and WO 2012 / 005857. Any of the amphiphilic membranes or layers discussed above may be used. A solid-state layer is not of biological origin. In other words, a solid state layer is not derived from or isolated from a biological environment such as an organism or cell, or a synthetically manufactured version of a biologically available structure. Solid state layers can be formed from both organic and inorganic materials including, but not limited to, microelectronic materials, insulating materials such as Si₃N₄, Al₂O₃, and SiO, organic and inorganic polymers such as polyamide, plastics such as Teflon® or elastomers such as two-component additioncure silicone rubber, and glasses. The solid state layer may be formed from monatomic layers, such as graphene, or layers that are only a few atoms thick. Suitable graphene layers are disclosed in WO 2009 / 035647.

[0203] Device

[0204] Also provided herein is a device. The device comprises a transmembrane protein nanopore, as described herein, in a membrane.

[0205] In some embodiments, the device comprises an array of transmembrane protein nanopores, as described herein, in a membrane. In some embodiments, the array comprises about 128, about 256, about 512, about 1024 or more wells, such as about 2000, about 3000, about 4000, about 6000, about 10000, about 12000, about 15000 or more wells.

[0206] Typically, the device is suitable for carrying out the methods disclosed herein. For example, the device may be configured to apply a potential across the nanopore and to take one or more measurements characteristic of the analyte as it moves with respect to the nanopore.

[0207] Characterising an analyte

[0208] Also disclosed herein is a method of characterising an analyte. The method comprises contacting the analyte with a transmembrane protein nanopore as defined herein. The method further comprises taking one or more measurements characteristic of the analyte. The one ormore measurements are typically taken as the analyte moves with respect to, such as through or into, the nanopore, thereby characterising the analyte. In some embodiments, the analyte moves at least partially through the nanopore. For example, the analyte may at least partially translocate the nanopore or may fully translocate the nanopore. In some embodiments, the analyte may move into the nanopore and become transiently or permanently immobilised. The one or more measurements may be taken as the analyte is actively moving with respect to, or through, the nanopore, or once the analyte has moved into the nanopore and is immobilised.

[0209] The method comprises contacting the analyte to be characterised with the nanopore. In some embodiments, the analyte is contacted with the cis side of the nanopore. In some embodiments, the analyte is contacted with the trans side of the nanopore. Typically, the analyte may be introduced into the trans compartment of a device comprising a nanopore as described herein such that it interacts with the trans side of the nanopore.

[0210] In some embodiments the analyte is contacted with the cis side of the nanopore and moves through the nanopore (i.e. translocates the nanopore) in the direction from the cis side of the nanopore to the trans side of the nanopore. In some embodiments the analyte is contacted with the trans side of the nanopore and moves through the nanopore (i.e. translocates the nanopore) in the direction from the trans side of the nanopore to the cis side of the nanopore.

[0211] In some embodiments, the analyte does not fully translocate through the nanopore, i.e. partially translocates through the nanopore or enters into the nanopore. In some embodiments, the analyte is contacted with the cis side of the nanopore and moves into, or at least partially translocates into, the nanopore in the direction from the cis side of the nanopore to the trans side of the nanopore. In some embodiments, the analyte is contacted with the trans side of the nanopore and moves into, or at least partially translocates into, the nanopore in the direction from the trans side of the nanopore to the cis side of the nanopore. One or more measurements characteristic of the analyte may be measured once the analyte moves into, or at least partially translocates into, the nanopore.

[0212] The analyte can be contacted with the nanopore whilst an electrophoretic force is applied across the nanopore. In some embodiments, the analyte is charged and the analyte interacts with the electrophoretic force. In some embodiments, the analyte is uncharged and does not interact with the electrophoretic force. An uncharged analyte may still be caused tomove with respect to the nanopore via electroosmosis and / or diffusion. Such movement may be in a direction from the cis to the trans side of the nanopore or from the trans side to the cis side of the nanopore.

[0213] An analyte may be characterised using the nanopore once or more than once, e.g. at least 2 times, at least 3 times, at least 4 times, at least 5 times, at least 10 times, at least 20 times, at least 50 times, at least 100 times, at least 500 times or more times. Characterising an analyte multiple times using a nanopore as described herein may allow the accuracy of the characterisation of the analyte to be improved.

[0214] The one or more measurements can be any suitable measurements. Typically, the one or more measurements are electrical measurements, e.g. current measurements, and / or one or more optical measurements. Apparatuses for recording suitable measurements, and the information that such measurements can provided, are described in more detail herein.

[0215] In some embodiments, the characterisation involves taking one or more measurements characteristic of the presence or absence or concentration of an analyte. In some embodiments, the concentration of an analyte may be the absolute concentration of an analyte, or the relative concentration of an analyte in a sample comprising two or more analytes.

[0216] In some embodiments, the nanopore is present in a membrane. Membranes suitable for use in accordance with the disclosed methods are described in more detail herein.

[0217] In some embodiments, the disclosed methods comprise applying a potential across the membrane and measuring a signal characteristic of ion flow through the nanopore, wherein the movement of an analyte with respect to the nanopore, such as into the nanopore or through the nanopore, perturbs said signal thereby generating one or more detectable events. In some embodiments taking one or more measurements characteristic of the analyte comprises determining one or more parameters selected from the frequency of the one or more events, the duration of the one or more events, the magnitude of the signal(s) during the one or more events, and / or the noise level of the signal(s) during the one or more events.

[0218] The measurements taken are typically characteristic of one or more characteristics of the analyte. The one or more characteristics may comprise determining the presence or absence of the analyte. The one or more characteristics may comprise the identity of the analyte. Where the analyte is a polynucleotide, the one or more characteristics may be selected from (i) the length of the polynucleotide, (ii) the identity of the polynucleotide, (iii) the sequence of the polynucleotide, (iv) the secondary structure of the polynucleotide, (v) whether or not thepolynucleotide is modified,, and (vi) the number, position(s) and / or location(s) of any modifications on the polynucleotide. Where the analyte is a polypeptide, the one or more characteristics may be selected from (i) the length of the polypeptide, (ii) the identity of the polypeptide, (iii) the sequence of the polypeptide, (iv) the secondary structure of the polypeptide, (v) whether or not the polypeptide is modified, and (vi) the number, position(s) and / or location(s) of any modifications on the polypeptide. The modifications of the polypeptide may be post-translation modifications (PTMs) of a polypeptide.

[0219] Typical post-translational modifications include modification with a hydrophobic group, modification with a cofactor, addition of a chemical group, glycation (the non-enzymatic attachment of a sugar), biotinylation and pegylation. Post-translational modifications can also be non-natural, such that they are chemical modifications (e.g. done in the laboratory) for biotechnological or biomedical purposes. This can allow monitoring the levels of the laboratory made peptide, polypeptide or protein in contrast to the natural counterparts.

[0220] Examples of post-translational modification with a hydrophobic group include myristoylation, attachment of myristate, a C₁₄ saturated acid; palmitoylation, attachment of palmitate, a C₁₆ saturated acid; isoprenylation or prenylation, the attachment of an isoprenoid group; farnesylation, the attachment of a farnesol group; geranylgeranylation, the attachment of a geranylgeraniol group; and glypiation, and glycosylphosphatidylinositol (GPI) anchor formation via an amide bond.

[0221] Examples of post-translational modification with a cofactor include lipoylation, attachment of a lipoate (C₈) functional group; flavination, attachment of a flavin moiety (e.g. flavin mononucleotide (FMN) or flavin adenine dinucleotide (FAD)); attachment of heme C, for instance via a thioether bond with cysteine; phosphopantetheinylation, the attachment of a 4'-phosphopantetheinyl group; and retinylidene Schiff base formation.

[0222] Examples of post-translational modification by addition of a chemical group include acylation, e.g. O-acylation (esters), N-acylation (amides) or S-acylation (thioesters); acetylation, the attachment of an acetyl group for instance to the N-terminus or to lysine; formylation; alkylation, the addition of an alkyl group, such as methyl or ethyl; methylation, the addition of a methyl group for instance to lysine or arginine; amidation; butyrylation; gamma-carboxylation; glycosylation, the enzymatic attachment of a glycosyl group for instance to arginine, asparagine, cysteine, hydroxylysine, serine, threonine, tyrosine or tryptophan; polysialylation, the attachment of polysialic acid; malonylation; hydroxylation;iodination; bromination; citrulination; nucleotide addition, the attachment of any nucleotide such as any of those discussed above, ADP ribosylation; oxidation; phosphorylation, the attachment of a phosphate group for instance to serine, threonine or tyrosine (O-linked) or histidine (N-linked); adenylyl ati on, the attachment of an adenylyl moiety for instance to tyrosine (O-linked) or to histidine or lysine (N-linked); propionylation; pyroglutamate formation; S-glutathionylation; Sumoylation; S-nitrosylation; succinylation, the attachment of a succinyl group for instance to lysine; selenoylation, the incorporation of selenium; and ubiquitinilation, the addition of ubiquitin subunits (N-linked).

[0223] Preferred PTMs for detection by the disclosed methods are phosphorylations, glutathionylations and glycosylations, particularly phosphorylations.

[0224] Analyte

[0225] Any suitable analyte can be characterised using the method disclosed herein. Selection of appropriate analytes for use in accordance with the disclose method is an operational parameter of the disclosed method and is within the capacity of those skilled in the art.

[0226] In some embodiments, the analyte is selected from ions, water, sugars, inorganic salts, lipids, amino acids, peptides, polypeptides, proteins, nucleotides, polynucleotides, carbohydrates, saccharides, polysaccharides, metabolites, neurotransmitters, labels and pharmaceutical agents (i.e. drugs). In some embodiments, the analyte is selected from polypeptides and polynucleotides.

[0227] In some embodiments, the analyte is a nucleotide or a polynucleotide. The polynucleotides may be double-stranded or single-stranded. The nucleotide or polynucleotide may be DNA, RNA, LNA, BNA, PNA and / or morpholino nucleotides. For example, the polynucleotide may be complementary DNA (cDNA) or RNA. The polynucleotides may comprise DNA and / or RNA. Nucleic acids are negatively charged. A polynucleotide is a macromolecule comprising two or more nucleotides. The nucleotides can be naturally occurring or artificial. A nucleotide typically contains a nucleobase, a sugar and at least one phosphate group. The nucleobase is typically heterocyclic. Nucleobases include, but are not limited to, purines and pyrimidines and more specifically adenine, guanine, thymine, uracil and cytosine. The sugar is typically a pentose sugar. Nucleotide sugars include, but are not limited to, ribose and deoxyribose. The nucleotide is typically a ribonucleotide or deoxyribonucleotide. The nucleotide typically contains a monophosphate, diphosphate ortriphosphate. Phosphates may be attached on the 5’ or 3’ side of a nucleotide. Suitable nucleotides include, but are not limited to, adenosine monophosphate (AMP), adenosine diphosphate (ADP), adenosine triphosphate (ATP), guanosine monophosphate (GMP), guanosine diphosphate (GDP), guanosine triphosphate (GTP), thymidine monophosphate (TMP), thymidine diphosphate (TDP), thymidine triphosphate (TTP), uridine monophosphate (UMP), uridine diphosphate (UDP), uridine triphosphate (UTP), cytidine monophosphate (CMP), cytidine diphosphate (CDP), cytidine triphosphate (CTP), cyclic adenosine monophosphate (cAMP), cyclic guanosine monophosphate (cGMP), deoxyadenosine monophosphate (dAMP), deoxyadenosine diphosphate (dADP), deoxyadenosine triphosphate (dATP), deoxyguanosine monophosphate (dGMP), deoxyguanosine diphosphate (dGDP), deoxyguanosine triphosphate (dGTP), deoxythymidine monophosphate (dTMP), deoxythymidine diphosphate (dTDP), deoxythymidine triphosphate (dTTP), deoxyuridine monophosphate (dUMP), deoxyuridine diphosphate (dUDP), deoxyuridine triphosphate (dUTP), deoxycytidine monophosphate (dCMP), deoxycytidine diphosphate (dCDP) and deoxycytidine triphosphate (dCTP). The nucleotides are usually selected from AMP, TMP, GMP, UMP, dAMP, dTMP, dGMP or dCMP. The polynucleotide may be at least about 10, at least about 50, at least about 100, at least about 200, at least about 500, at least about 1000, at least about 5000, at least about 10000, at least about 50000 or at least about 1000000 nucleotides in length. The polynucleotide may comprise from about 1 to about 100000 or more nucleotides, such as from about 10 to about 100000, from about 100 to about 10000, from about 50 to about 500, or from about 500 to about 50000 nucleotides.

[0228] In some embodiments, the analyte may be an amino acid, a peptide, a polypeptide or a protein. The polypeptide may be a protein or a fragment thereof. The polypeptide can be naturally-occurring or non-naturally-occurring. The polypeptide can include within it synthetic or modified amino acids. A number of different types of modification to amino acids are known in the art. The polypeptide may be at least about 2, at least about 10, at least about 50, at least about 100, at least about 500 or at least about 1000 amino acids in length. The peptide may be a polymer of from about 2 to about 50 amino acids or may be a longer polymer of amino acids. Proteins are typically one or more polypeptides that are folded into a functional conformation or form part of a functional complex. For example, the polypeptide may be from about 2 to about 1000 amino acids, such as from about 5 to about 500 amino acids, e.g. from about 10 to about 100 amino acids, such as from about 20 to about 60 aminoacids e.g. from about 30 to about 50 amino acids. Suitable polypeptides include, but are not limited to, proteins such as enzymes, antibodies, hormones, growth factors or growth regulatory proteins, such as cytokines; or fragments of such proteins. The polypeptide may be bacterial, archaeal, fungal, viral or derived from a parasite. The polypeptide may be derived from a plant. The polypeptide is typically mammalian, more usually human. The polypeptide or protein may comprise a post-translational modification as described herein.

[0229] In some embodiments, the analyte is a saccharide or a polysaccharide. A polysaccharide is a polymeric carbohydrate molecule composed of chains of monosaccharide units bound together by glycosidic linkages. A polysaccharide may be linear or branched. A polysaccharide may be homogeneous (comprising only one repeating unit) or heterogeneous (containing modifications of the repeating unit). Polysaccharides include callose or laminarin, chrysolaminarin, xylan, arabinoxylan, mannan, fucoidan and galactomannan. A polysaccharide may be produced by a bacterium such as a pathogenic bacterium. The polysaccharide may be a capsular polysaccharide having a molecular weight of 100-2000 kDa. The polysaccharide may be synthesized from nucleotide-activated precursors (called nucleotide sugars). The polysaccharide may be a lipopolysaccharide. The polysaccharide may be a therapeutic polysaccharide. The polysaccharide may be a toxic polysaccharide. The polysaccharide may be suitable for use as a vaccine. The polysaccharide may be for example bacterial or derived from a plant. The polysaccharide may be useful as an antibiotic, such as streptomycin, neomycins, paromomycine, kanamycin, chalcomycin, erythromycin, magnamycin, spiramycin, oleandomycin, cinerubin and amicetin, or a derivative of any one of the preceding compounds. A polysaccharide, the polysaccharide may comprise from about 2 to about 1000 monosaccharide units, such as from about 5 to about 500 monosaccharide units, e.g. from about 10 to about 100 monosaccharide units, such as from about 20 to about 60 monosaccharide units e.g. from about 30 to about 50 monosaccharide units. In some embodiments, the analyte is a saccharide such as β-cyclodextrin.

[0230] In some embodiments, the analyte is a neurotransmitter. The neurotransmitter may be selected from glutamate, glycine, serotonin, epinephrine, dopamine, opioids, ATP, GBP, nitric oxide, carbon monoxide, y-aminobutyric acid (GABA), acetylcholine, and norepinephrine.

[0231] In some embodiments, the analyte may further comprise large charged substances, such as sulfo-Cy5 and FMN.Those skilled in the art will recognise that the composition of the analyte is not particularly limiting, and any suitable analyte can be characterised in the methods disclosed herein.

[0232] Any number of analytes can be characterised in the disclosed methods. The analyte may be present in a sample comprising a plurality of analytes. For instance, the method may comprise characterising 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 50, 100 or more analytes.

[0233] Conditions

[0234] Any suitable apparatus can be used to enact the methods of the present disclosure. Electrical measurements may be made using standard single channel recording equipment as describe in Stoddart, D. S., et al.. (2009), Proceedings of the National Academy of Sciences of the United States of America 106, p7702-7707, Lieberman KR et al, J Am Chem Soc. 2010;132(50):17961-72, and International Application WO 2000 / 28312, each of which is incorporated by reference in its entirety. Alternatively, electrical measurements may be made using a multi-channel system, for example as described in International Application WO 2009 / 077734 and International Application WO 2011 / 067559, each of which is incorporated by reference in its entirety.

[0235] In some embodiments, the disclosed methods are carried out using an apparatus that is suitable for investigating a membrane / nanopore system in which a nanopore is inserted into a membrane. The disclosed methods may be carried out using any apparatus that is suitable for transmembrane nanopore sensing. For example, the apparatus may comprise a chamber comprising an aqueous solution and a barrier that separates the chamber into two sections. The barrier may have an aperture in which the membrane containing the nanopore is formed. The disclosed methods may also be carried out using droplet interface bilayers (DIBs). Two water droplets may be placed on electrodes and immersed into an oil / phospholipid mixture. The two droplets may be taken in close contact and at the interface a phospholipid membrane may be formed where the nanopores get inserted.

[0236] The disclosed methods may be carried out using the apparatus described in International Application WO 2008 / 102120.

[0237] The disclosed methods typically involve measuring the current flowing through a nanopore. Therefore, the apparatus may also comprise an electrical circuit capable of applying a potential and measuring an electrical signal across a membrane and nanopore. Themethods may be carried out using a patch clamp or a voltage clamp. The methods usually involve the use of a voltage clamp.

[0238] The characterisation methods may comprise optical measurements, for example such as described in WO 2016 / 009180 and WO 2021 / 198695.

[0239] Typically, in the disclosed methods, the nanopore is present in an array of nanopores. For example, in some embodiments the methods may be carried out on a silicon-based array of wells where each array comprises 128, 256, 512, 1024 or more wells, such as 2000, 3000, 4000, 6000, 10000, 12000, 15000 or more wells.

[0240] The methods may be carried out using an array of nanopores as described herein. The use of an array of nanopores may allow the monitoring of the method by monitoring a signal such an electrical or optical signal. The optical detection of analytes using an array of nanopores can be conducted using techniques known in the art, such as those described by Huang et al, Nature Nanotechnology (2015) 10: 986-992

[0241] The methods of the invention may involve the measuring of a current flowing through a nanopore. Suitable conditions for measuring ionic currents through transmembrane nanopores are known in the art and disclosed in the Example. The method is typically carried out with a voltage applied across the membrane and nanopore. The voltage used is typically from +2 V to -2 V, typically -400 mV to +400mV. The voltage used is typically in a range having a lower limit selected from -400 mV, -300 mV, -200 mV, -150 mV, -100 mV, -50 mV, -20mV and 0 mV and an upper limit independently selected from +10 mV, + 20 mV, +50 mV, +100 mV, +150 mV, +200 mV, +300 mV and +400 mV. The voltage used is more often in the range 100 mV to 240m V and most usually in the range of 120 mV to 220 mV.

[0242] The methods may be carried out in the presence of charge carriers, such as metal salts, for example alkali metal salt, halide salts, for example chloride salts, such as alkali metal chloride salt. Charge carriers may include ionic liquids or organic salts, for example tetramethyl ammonium chloride, trimethylphenyl ammonium chloride, phenyltrimethyl ammonium chloride, or l-ethyl-3 -methyl imidazolium chloride. In the exemplary apparatus discussed above, the salt is present in the aqueous solution in the chamber. Potassium chloride (KC1), sodium chloride (NaCl), caesium chloride (CsCl) or a mixture of potassium ferrocyanide and potassium ferricyanide is typically used. KC1, NaCl and a mixture of potassium ferrocyanide and potassium ferricyanide are preferred. The salt concentration may be at saturation. The salt concentration may be 3 M or lower and is typically from 0.1 to 2.5M, from 0.3 to 1.9 M, from 0.5 to 1.8 M, from 0.7 to 1.7 M, from 0.9 to 1.6 M or from 1 M to 1.4 M. The salt concentration is typically from 150 mM to 1 M. The method is usually carried out using a salt concentration of at least 0.3 M, such as at least 0.4 M, at least 0.5 M, at least 0.6 M, at least 0.8 M, at least 1.0 M, at least 1.5 M, at least 2.0 M, at least 2.5 M or at least 3.0 M. The salt concentration used on each side of the membrane may be different, such as 0.1 M at one side and 3 M at the other. The salt and composition used on each side of the membrane may be also different. The use of asymmetric charge conditions can maximise the electroosmotic force through the nanopore.

[0243] In some embodiments, the movement of the analyte through the nanopore is driven by electroosmotic force. The electroosmotic force may be determined, chosen or enhanced according to the requirements of the user using any means known in the art. For example, the electroosmotic force may be increased by reducing the pH. At low pH (e.g. from about pH 2 to about pH 5) basic amino acid side chains in the channel of the nanopore may be protonated and thus have a higher charge. At high pH (e.g. from about pH 8 to about pH 10) acidic amino acid side chains in the channel of the nanopore may be deprotonated and thus have a higher charge. Modifications to increase the charge of the channel through the nanopore may be made in other ways. For example, transmembrane protein nanopores can be modified by mutation to insert charged amino acids into the channel therethrough in order to increase the electroosmotic force through the nanopore. This is described in more detail in

[0244] WO 2024 / 165853, the entire contents of which are incorporated herein by reference.

[0245] In some embodiments the movement of the analyte may be modulated by a physical or chemical force (potential). In some embodiments the physical force is provided by an electrical (e.g. voltage) potential or a temperature gradient, etc. In some embodiments the chemical force is provided by a concentration (e.g. pH) gradient.

[0246] In some embodiments, the movement of the peptide, polypeptide or protein is modulated by mechanically manipulating the analyte thereby moving said analyte with respect to the nanopore.

[0247] In some embodiments the movement of analyte is modulated using a method as described in WO 2020 / 016573, the entire contents of which are incorporated herein by reference.

[0248] The methods are typically carried out in the presence of a buffer. In the exemplary apparatus discussed above, the buffer is present in the aqueous solution in the chamber. Anybuffer may be used in the method of the invention. Typically, the buffer is HEPES. Another suitable buffer is Tris-HCl buffer. The methods are typically carried out at a pH of from 4.0 to 12.0, from 4.5 to 10.0, from 5.0 to 9.0, from 5.5 to 8.8, from 6.0 to 8.7 or from 7.0 to 8.8 or 7.5 to 8.5. The pH used is typically about 7.5. In some embodiments the disclosed methods are conducted between about pH 4 and about pH 10. In some embodiments the disclosed methods are conducted between about pH 5 and about pH 9. In some embodiments the disclosed methods are conducted between about pH 6 and about pH 8. In some embodiments the disclosed methods are conducted about pH 7, such as about pH 7.2.

[0249] A reducing agent such as TCEP (tris(2-carboxyethyl)phosphine) may be present, e.g. at a concentration of from about 1 mM to about 50 mM such as from about 5 mM to about 20 mM, e.g. about 10 mM.

[0250] The methods may be carried out at from 0 °C to 100 °C, from 15 °C to 95 °C, from 16 °C to 90 °C, from 17 °C to 85 °C, from 18 °C to 80 °C, 19 °C to 70 °C, or from 20 °C to 60 °C. The methods are typically carried out at room temperature.

[0251] Other methods

[0252] In a further aspect, the disclosure provides a method of increasing the interaction dwell time of an analyte with a transmembrane protein nanopore as described herein. The method comprises modifying the nanopore by introducing one or more heterologous insertions in to the nanopore, thereby extending the length of the channel of the nanopore.

[0253] In some embodiments, the interaction dwell time of an analyte is increased by at least 2-fold. In some embodiments, the interaction dwell time of an analyte is increased by at least 4-fold, at least 6-fold, at least 8-fold, at least 16-fold, at least 32-fold or at least 80-fold. The interaction dwell time of an analyte is the length of time that an analyte or a part thereof interacts with a constriction in the channel of the nanopore. During the characterisation of an analyte, the interaction between the analyte or part thereof with the constriction in the channel of the nanopore affects the signal being measured. The length of time that the signal is affected by the presence of the analyte in the channel of the nanopore corresponds to the dwell time. In some embodiments, for example wherein the analyte is a polymer such as a polynucleotide, a polypeptide or a polysaccharide, the dwell time can be measured as the dwell time of each monomer unit of the analyte, e.g. as the analyte passes through the nanopore. Dwell time may be measured using any means known to the skilled person. Forexample, as used in the Examples herein, dwell time analysis was performed by using the maximum interval likelihood algorithm of QuB.

[0254] The analyte may be any analyte as described herein. In some embodiments, the analyte is a polynucleotide, a polypeptide or a saccharide. In some embodiments, the dwell time is increased by at least 2-fold. In some embodiments, the analyte is a polynucleotide and the dwell time is increased by at least 2.4-fold. In some embodiments, the analyte is a polypeptide and the dwell time is increased by at least 100-fold. In some embodiments, the analyte is a saccharide, such as β-cyclodextrin, and the dwell time is increased by at least 80-fold.

[0255] The disclosure also provides a method of increasing the distance between two or more constrictions in a transmembrane protein nanopore, wherein the transmembrane protein nanopore comprises a first opening, a second opening and a channel therebetween. The method comprises modifying the nanopore by introducing one or more heterologous insertions between the two or more constrictions, thereby increasing the distance between the two or more constrictions in the nanopore.

[0256] The disclosure also provides a method of increasing the length of a sensing region in a transmembrane protein nanopore, wherein the transmembrane protein nanopore comprises a first opening, a second opening and a channel therebetween. The method comprises modifying the nanopore by introducing one or more heterologous insertions in to the nanopore, thereby increasing the length of the sensing region of the nanopore.

[0257] The disclosure also provides a method of increasing the sensing volume of a transmembrane protein nanopore, wherein the transmembrane protein nanopore comprises a first opening, a second opening and a channel therebetween. The method comprises modifying the nanopore by introducing one or more heterologous insertions in to the nanopore, thereby increasing the sensing volume of the nanopore.

[0258] The transmembrane protein nanopore may be as described herein. The position and / or the identity of the heterologous insertion may be as described herein.

[0259] It is to be understood that although particular embodiments, specific configurations as well as materials and / or molecules, have been discussed herein for methods according to the present invention, various changes or modifications in form and detail may be made without departing from the scope and spirit of this invention. The preceding embodiments andfollowing examples are provided for illustration only, and should not be considered limiting the application. The application is limited only by the claims.

[0260] Examples

[0261] Construction of extended aHL mutant genes

[0262] DNA oligonucleotides were custom-synthesized by Integrated DNA Technologies (IDT). All other reagents were obtained from New England Biolabs (NEB), unless otherwise specified. To construct short a-hemolysin (aHL) mutants, (GNTN)N-D8H6 (N = 1-3), (SG)N-D8H6 (N = 1-4), (SV)N-D8H6 (N = 1-4), (GG)N-D8H6 (N = 1-4), (TN)6-D8H6, (TN)6 / (ST)6-D8H6, (TN)6 / (SI)6-D8H6, (TNGN)3 / (GNTN)3-D8H6, (KNGN)3 / (GNEN)3-D8H6, (TNKN)3 / (ENTN)3-D8H6 and (GNTNEN)n / (KNGNTN)n(n = 2, 3), the inserts were introduced into the gene encoding aHL, which includes a C-terminal D8H6 tag within a pT7 vector, at positions L116-T117 and I142-G143 through two rounds of mutagenic PCR and KLD reactions using Phusion® High-Fidelity PCR Master Mix with GC Buffer and KLD Enzyme Mix (NEB), and amplified in XL10-Gold ultracompetent E. coli cells (Agilent). In addition, for (SG)N-D8H6 (N = 1-4) and (SV)N-D8H6 (N = 1-4) mutants, the inserts were also introduced into the same gene at secondary positions, G122-N123 and I136-G137. For (NGNT)n / (GNTN)n-D8H6 (n = 1-2) mutants, the inserts were introduced into the same gene at third positions, D127-D128 and G130-K131. For (SV)4-D8H6 mutant, the inserts were introduced into the same gene at fourth positions, G126-D127 and I132-G133. To construct long aHL mutants (GNTN)N-D8H6 (N = 4, 5, 8, 14 and 20), (GNTN)2-EN-(GNTN)2 / (GNTN)2-KN-(GNTN)2-D8H6, (GNTNEN)N / (KNGNTN)N-D8H6 (N = 2, 3 and 10), GNTN)3-GNTIGITN-(GNTN)3, (GNTN)2-GITNGITN-(GNTN)4, (GNTN)3-GITNGITN-(GNTN)3, (GNTN)4-GITNGITN-(GNTN)8, (GNTN)6-GITNGITN-(GNTN)6, (GNTN)6-GITIGITN-(GNTN)6, (GNTN)8 / GNTHGHTN-(GNTN)6, (GNTN)6-GHTHGNTN / (GNTN)8, (GNTN)6-GHTHGNTN / GNTHGHTN-(GNTN)6, (GNTNEN)IO / (KNGNTN)IO, (GNTNEN)4-GNTIEIGITNEN-(GNTNEN)4 / (KNGNTN)4-KNGNTIKIGITN-(KNGNTN)4, (TNGK)n / (TNGD)n (n = 4, 8, 10, 14), (TNGKTNGN)n / (TNGDTNGN)n(n = 4, 5, 7), (GNSN)8, (GQSQ)8, and (GQTQ)8, ssDNA fragments for GNTN repeats or relevant repeats were purchased, annealed, and cloned into the same aHL template at the positions L116-T117 and I142-G143 by Gibson Assembly. All the constructs were verified by DNA sequencing.Expression and purification of extended aHL mutants

[0263] Plasmids encoding aHL mutants were transformed into BL21(DE3)pLysS competent cells (Agilent), which were subsequently cultured in 20 mL Luria broth (LB) medium, supplemented with carbenicillin (50 pg / mL) and chloramphenicol (35 pg / mL), at 37 °C with continuous shaking (225 rpm). Protein expression was induced during the exponential growth phase (OD600 = 0.6~0.8) with isopropyl-P-D-l-thiogalactopyranoside (IPTG) (0.4 mM final concentration). After overnight incubation at 18 °C, cells were harvested by centrifugation (10 min, 5,000 × g), resuspended in binding buffer (20 mM sodium phosphate, 300 mM NaCl, 0.05% N-Dodecyl-P-D-maltoside (DDM), pH 7.4), supplemented with a protease inhibitor cocktail (Pierce™, EDTA-free, Thermo), and lysed by three freeze-thaw cycles. The lysate was treated with lysozyme (0.5 mg), nuclease (50 U, Pierce™ Universal Nuclease for Cell Lysis, Thermo) and MgCl₂ (5 mM) on ice for 1 hour. Cell debris was removed by centrifugation at 12,000 x g for 15 min. The supernatant was applied to a NEB Express® Ni Spin Column pre-washed with binding buffer. The column was washed with the wash buffer (20 mM sodium phosphate, 300 mM NaCl, 5 mM Imidazole, 0.05% DDM, pH 7.4) and eluted with the elution buffer (20 mM sodium phosphate, 300 mM NaCl, 500 mM imidazole, 0.05% DDM, pH 7.4). The eluted protein was aliquoted and stored at 4 °C for short-term use, or at -80 °C for long-term storage.

[0264] Single-channel recording

[0265] Planar lipid bilayers of l,2-diphytanoyl-sn-glycero-3 -phosphocholine (Avanti Polar Lipids) were formed by using the Muller-Montal method on a 50 pm-diameter aperture made in a Teflon film (25 pm thick, Goodfellow) separating two 500 pL compartments (cis and trans) of the recording chamber. Each compartment was filled with recording buffer (2 M KC1, 10 mM Tris, 0.1 mM EDTA, pH 8.5). Ionic currents were measured at 23 ± 1 °C by using Ag / AgCl electrodes connected to an Axopatch 200B amplifier. To insert a single pore into the bilayer, the protein sample (1 pL, -4 mg / mL) eluted from NEBExpress® Ni Spin Column was added into the cis compartment. Once a pore had inserted, the solution was replaced with fresh buffer by manual pipetting, to prevent further insertions. 92-nt ssDNA (Merck, PAGE-purified, 5'- AAAAAAAAAAAAAAAAAAAAATTCCCCCCCCCCCCCCCCCCCCCTTAAAAAAAA AATTCCCCCCCCCCTTAAAAAAAAAATTCCCCCCCCCC-3 ) was added to the ciscompartment (200 nm or 1 pM DNA) and record at +100 mV (trans). The 10-AA peptide (H4 peptide (15-24), Cayman, AKRHRKVLRD), 21-AA peptide (H4 peptide (1-21), Cayman, >95% (HPLC), SGRGKGGKGLGKGGAKRHRKV), 21-AA-S1P peptide (H4 peptide (1-21) with phosphorylation at Seri, Chinese peptide company, >95% (HPLC), (SPh)GRGKGGKGLGKGGAKRHRKV), 21-AA-K20Ac peptide (H4 peptide (1-21) with acetylation at Lys20, Chinese peptide company, >95% (HPLC), SGRGKGGKGLGKGGAKRHR(KAc)V), or 31-AA peptide (H4 peptide (1-31), Chinese peptide company, >95% (HPLC), SGRGKGGKGLGKGGAKRHRKVLRDNIQGITK) was added to the trans compartment (4~20 pM) and record at +100 mV (trans). For the data shown in Figure 5, H4 peptide (1-21) of 20 pM (Cayman, >95% (HPLC), SGRGKGGKGLGKGGAKRHRKV) or H4 peptide (15-24) of 8 pM (Cayman, AKRHRKVLRD) was added to the trans compartment and record at +100 mV (trans).

[0266] Chimeric aHL nanopores (Tables 3a, 3b, rows 88-95) were tested under preliminary conditions only; optimisation of experimental conditions is expected to result in successful pore formation. P-cyclodextrin (Merck, Pharmaceutical Secondary Standard, 99.9%) was added into trans compartment (2~80 pM) and record at +50 mV. Signals were low-pass filtered at 10 kHz and sampled at 50 kHz with a Digidata 1440A digitizer (Molecular Devices). Current traces were idealized by using Clampfit 10.3 (Molecular Devices). The idealized data were analyzed with QuB 2.0 software (www.qub.buffalo.edu). Dwell time analysis was performed by using the maximum interval likelihood algorithm of QuB.

[0267] Results

[0268] Transmembrane protein nanopores are widely applied as single-molecule sensors for biomolecule analysis. For example, naturally-occurring transmembrane P barrels vary in the oligomeric state (single-chained or multimers), internal geometry and diameter ( 1 ->25 nm), and length (up to ~10 nm).

[0269] Here, we extended a P-barrel of a nanopore by inserting de-novo-designed peptide sequences. By harnessing the robust pore assembly driven by the cap domain of a multimeric nanopore, a-hemolysin, we constructed transmembrane nanopores containing extended P barrels (Figure 1). In a first aspect, the de novo-designed amino acid sequence, GNTN, was inserted as repeats into the water-accessible segment of the P-barrel of a-hemolysin, above the transmembrane regions (between L116-T117 and I142-G143). We confirmed the assembly ofoligomeric protein pores by SDS-PAGE gel (Figure 2a), and the barrel folding by singlechannel recording and TEM (Figure 2b-d, Figure 3, Figure 6). In single-channel recording, a reduction in conductance was observed with an increase in the number of GNTN inserts in both p strands. The resistance has displayed a direct correlation with the number of GNTN inserts, essentially the length of extended mutants (Figure 2c-d, Figure 6b).

[0270] To prove the folding of extended P barrels with alternating inward- and outward-facing residues as designed, a cysteine residue was introduced at 14 distinct positions spanning the entirety of the P strand on aHL-(GNTN)3 (Figure 3a). The chemical reactivity confirmed that the extended barrels maintain the alternating orientation of amino acids after extension, a crucial characteristic of P barrels (Figure 3b-c). TEM results have demonstrated the construction of a transmembrane P-barrel protein, aHL-(GNTN)8-TM, with a P barrel measuring 14.2 nm, surpassing the length of the longest natural P barrel, the protective antigen pore (10.5 nm) (Figure 3d).

[0271] aHL-(GNTN)nmutants have exhibited novel properties resulting from P-extension, such as prolonged dwell times of binding events with neutral P-cyclodextrin (P-CD) (Figure 4a-g, Figure 7), prolonged translocation time of biopolymers, e.g., ssDNA and peptides (Figure 5, Figure 8). Moreover, PCD can serve as a binding site for channel blockers, providing a platform for detecting various organic molecules using the aHL-pCD complex as a biosensor, such as adamantanamine hydrochloride (Figure 4h-i). The extended barrels exhibited increased dwell times and percentage residual currents of diverse analytes, including P-cyclodextrin, DNA and peptides - key features for high-resolution singlemolecule profiling.

[0272] Additional data for various transmembrane protein nanopores comprising one or more heterologous insertions is shown in Tables 3a and 3b (numbered rows in Table 3b refer to correspondingly numbered rows in Table 3a).

[0273] Table 3a Sequences and positions of insertions in the p barrel of a-hemolysin

[0274] 1. Extended P barrel mutants of aHL, constructed and characterised

[0275] Sequence of inserts Insertion Insertion position

[0276] (N to C) position (Relative to 1ststrand 2ndstrand 1ststrand 2ndstrand transmembrane region (TMR))

[0277]

[0278] GNTN mutantsL116- 1142- 1 GNTN GNTN Above TMR T117 G143

[0279] L116- 1142- 2 (GNTN)2 (GNTN)2 Above TMR T117 G143

[0280] L116- 1142- 3* (GNTN)3 (GNTN)3 Above TMR T117 G143

[0281] L116- 1142- 4 (GNTN)4 (GNTN)4 Above TMR T117 G143

[0282] L116- 1142- 5 (GNTN)5 (GNTN)5 Above TMR T117 G143

[0283] L116- 1142- 6 (GNTN)6 (GNTN)6 Above TMR T117 G143

[0284] L116- 1142- 7 (GNTN)7 (GNTN)7 Above TMR T117 G143

[0285] L116- 1142- 8 (GNTN)8 (GNTN)8 Above TMR T117 G143

[0286] L116- 1142- 9 (GNTN) 14 (GNTN) 14 Above TMR T117 G143

[0287] L116- 1142- 10 (GNTN)20 (GNTN)20 Above TMR T117 G143

[0288] GNTN mutants with outward-facing Isoleucine — mimicking extra hydrophobic ring above membrane found in long transmembrane 0-barrel proteins

[0289] (GNTN)3- (GNTN)3- L116- 1142- 11 GNTIGITN- GNTIGITN- Above TMR T117 G143

[0290] (GNTN)3 (GNTN)3

[0291] (GNTN)2- (GNTN)2- L116- 1142- 12 GITNGITN- GITNGITN- Above TMR T117 G143

[0292] (GNTN)4 (GNTN)4

[0293] (GNTN)3- (GNTN)3- L116- 1142- 13 GITNGITN- GITNGITN- Above TMR T117 G143

[0294] (GNTN)3 (GNTN)3

[0295] (GNTN)4- (GNTN)4- L116- 1142- 14 GITNGITN- GITNGITN- Above TMR T117 G143

[0296] (GNTN)8 (GNTN)8

[0297] (GNTN)6- (GNTN)6- L116- 1142- 15 GITNGITN- GITNGITN- Above TMR T117 G143

[0298] (GNTN)6 (GNTN)6

[0299] (GNTN)6- (GNTN)6- L116- 1142- 16 GITIGITN- GITIGITN- Above TMR T117 G143

[0300] (GNTN)6 (GNTN)6

[0301] GNTN mutants with outward-facing Histidine above membrane boundary GNTHGHTN- L116- 1142- 17 (GNTN)8 Above TMR (GNTN)6 T117 G143

[0302] (GNTN)6- L116- 1142- 18 (GNTN)8 Above TMR GHTHGNTN T117 G143

[0303] (GNTN)6- GNTHGHTN- L116- 1142- 19 Above TMR GHTHGNTN (GNTN)6 T117 G143

[0304] Gm " N mutants with inward-facing salt bridges formed by charged amino acids — forming new constriction sites

[0305] (GNTN)2-EN- (GNTN)2-KN- L116- 1142- 20 Above TMR

[0306]

[0307] (GNTN)2 (GNTN)2 T117 G143L116- 1142- 21 (GNTNEN)2 (KNGNTN)2 Above TMR T117 G143

[0308] L116- 1142- 22 (GNTNEN)3 (KNGNTN)3 Above TMR T117 G143

[0309] L116- 1142- 23 (GNTNEN)IO (KNGNTN)IO Above TMR T117 G143

[0310] (GNTNEN)4- (KNGNTN)4- L116- 1142- 24 GNTIEIGITNEN- KNGNTIKIGITN- Above TMR T117 G143

[0311] (GNTNEN)4 (KNGNTN)4

[0312] L116- 1142- 25 (KNGN)3 (GNEN)3 Above TMR T117 G143

[0313] L116- 1142- 26 (TNKN)3 (ENTN)3 Above TMR T117 G143

[0314] GNTN mutants with outward-facing salt bridges formed by charged amino acids — forming outward staples to stabilise the structure

[0315] L116- 1142- 27 (TNGK)4 (TNGD)4 Above TMR T117 G143

[0316] L116- 1142- 28 (TNGK)8 (TNGD)8 Above TMR T117 G143

[0317] L116- 1142- 29 (TNGK)10 (TNGD)10 Above TMR T117 G143

[0318] L116- 1142- 30 (TNGK)14 (TNGD)14 Above TMR T117 G143

[0319] L116- 1142- 31 (TNGKTNGN)4 (TNGDTNGN)4 Above TMR T117 G143

[0320] L116- 1142- 32 (TNGKTNGN)5 (TNGDTNGN)5 Above TMR T117 G143

[0321] L116- 1142- 33 (TNGKTNGN)7 (TNGDTNGN)7 Above TMR T117 G143

[0322] Unequal insertion into the two strands causes 0-strand mismatch — mis olded structure L116- 1142- 34 (GNTN)3 / Above TMR T117 G143

[0323] L116- 1142- 35 (GNTN)3G / Above TMR T117 G143

[0324] L116- 1142- 36 (GNTN)3G (GNTN)3 Above TMR T117 G143

[0325] L116- 1142- 37 (GNTN)3G NTN(GNTN)2 Above TMR T117 G143

[0326] Glycine mutants — Bland barrel

[0327] L116- 1142- 38 GG GG Above TMR T117 G143

[0328] L116- 1142- 39 GGGG GGGG Above TMR T117 G143

[0329] L116- 1142- 40 GGGGGG GGGGGG Above TMR T117 G143

[0330] L116- 1142- 41 GGGGGGGG GGGGGGGG Above TMR T117 G143

[0331] SG mutants

[0332] L116- 1142- 42 SG SG Above TMR T117 G143

[0333] L116- 1142- 43 SGSG SGSG Above TMR

[0334]

[0335] T117 G143L116- 1142- SGSGSG SGSGSG Above TMR T117 G143

[0336] L116- 1142- SGSGSGSG SGSGSGSG Above TMR T117 G143

[0337] G122- 1136- SG SG Within TMR N123 G137

[0338] G122- 1136- SGSG SGSG Within TMR N123 G137

[0339] G122- 1136- SGSGSG SGSGSG Within TMR N123 G137

[0340] G122- 1136- SGSGSGSG SGSGSGSG Within TMR N123 G137

[0341] SV mutants — Mimicking transmembrane barrel

[0342] L116- 1142- SV SV Above TMR T117 G143

[0343] svsv svsv L116- 1142- Above TMR T117 G143

[0344] svsvsv svsvsv L116- 1142- Above TMR T117 G143

[0345] svsvsvsv svsvsvsv L116- 1142- Above TMR T117 G143

[0346] G122- 1136- SV SV Within TMR N123 G137

[0347] svsv svsv G122- 1136- Within TMR N123 G137

[0348] svsvsv svsvsv G122- 1136- Within TMR N123 G137

[0349] svsvsvsv svsvsvsv G122- 1136- Within TMR N123 G137

[0350] Repeat the sequence of aHL

[0351] Y118- N139- EYMSTLTY VSIGHTLK Within TMR G119 V140

[0352] Y118- N139- (EYMSTLTY)2 (VSIGHTLK)2 Within TMR G119 V140

[0353] Y118- N139- (EYMSTLTY)5 (VSIGHTLK)5 Within TMR G119 V140

[0354] Other sequences

[0355] L116- 1142- TNTNTNTNTNTN TNTNTNTNTNTN Above TMR T117 G143

[0356] L116- 1142- TNTNTNTNTNTN STSTSTSTSTST Above TMR T117 G143

[0357] L116- 1142- TNTNTNTNTNTN SISISISISISI Above TMR T117 G143

[0358] L116- 1142- (TNGN)3 (GNTN)3 Above TMR T117 G143

[0359] L116- 1142- (GNSN)8 (GNSN)8 Above TMR T117 G143

[0360] L116- 1142- (GQSQ)8 (GQSQ)8 Above TMR T117 G143

[0361] L116- 1142- (GQTQ)8 (GQTQ)8 Above TMR T117 G143

[0362]

[0363] Barre extension below the transmembrane regionD127- G130- NGNT GNTN Below TMR D128 K131

[0364] D127- G130- (NGNT)2 (GNTN)2 Below TMR D128 K131

[0365] G126- 1132- SGSGSGSG SGSGSGSG Below TMR D127 G133

[0366] 2. Extended 0 barrel mutants of aHL planned

[0367] L116- 1142- GNSN GNSN Above TMR T117 G143

[0368] L116- 1142- (GNSN)2 (GNSN)2 Above TMR T117 G143

[0369] L116- 1142- (GNSN)3 (GNSN)3 Above TMR T117 G143

[0370] L116- 1142- (GNSN)4 (GNSN)4 Above TMR T117 G143

[0371] L116- 1142- (GNSN)5 (GNSN)5 Above TMR T117 G143

[0372] L116- 1142- (GNTN)3 (TNGN)3 Above TMR T117 G143

[0373] L116- 1142- (GNTN)4 (TNGN)4 Above TMR T117 G143

[0374] L116- 1142- (GNTN)5 (TNGN)5 Above TMR T117 G143

[0375] L116- 1142- (GNTN)9 (GNTN)9 Above TMR T117 G143

[0376] L116- 1142- (GNTN)10 (GNTN) 10 Above TMR T117 G143

[0377] L116- 1142- (GNTN)ll (GNTN) 11 Above TMR T117 G143

[0378] L116- 1142- (GNTN) 12 (GNTN) 12 Above TMR T117 G143

[0379] L116- 1142- (GNTN) 13 (GNTN) 13 Above TMR T117 G143

[0380] D127- G130- GNTN NGNT Below TMR D128 K131

[0381] D127- G130- GNTNEN NKNGNT Below TMR D128 K131

[0382] D127- G130- (GNTNEN)2 (NKNGNT)2 Below TMR D128 K131

[0383] G126- 1132- (GNTN)2 (GNTN)2 Below TMR D127 G133

[0384] 3. Extended β barrel mutants of chimeric αHL, constructed and characterised

[0385] Backbone

[0386] Name Backbone Insert Insert sequence Sequence

[0387] aHL AeL Plan A D222-T274 A1-Y118 Aerolysin

[0388] aHL AeL Plan B S276-D222 (retro) and

[0389] aHL Etx Plan A a-Hemolysin T106-T161 V140- Epsilon

[0390] N163-T106 aHL Etx Plan B N293 Toxin

[0391]

[0392] (retro)92 aHL Lys Plan A T39-V100

[0393] Lysenin E102-T39 93 aHL Lys Plan B

[0394] (retro) 94 aHL PA Long Protective D276-G351 95 aHL PA Short antigen S290-S337

[0395]

[0396] aHL-CHl pore E302-G323 ^Additional mutations introduced with increased peptide dwells: M113R and / or T115R

[0397] **from Gu et al., PNAS, 97(8), pp.3959-3964.

[0398] Table 3b - Characterisation of the extended a-hemolysin variants

[0399] 1. Extended 0 barrel mutants of aHL, constructed and characterised

[0400] 0 barrel

[0401] Characterisation

[0402] length

[0403] Gel confirming SDS- Electrical recording 0-cyclodextrin (0CD) nm resistant oligomer confirming channel binding confirming formation formation extended binding G NTN mutants

[0404] 2-fold increase in 0CD 1 6.3

[0405] residence time 2.5-fold increase in 0CD 2 7.4

[0406] residence time 4-fold increase in 0CD 3* 8.6

[0407] residence time 6-fold increase in 0CD 4 9.7

[0408] residence time 32-fold increase in 0CD 5 10.8

[0409] residence time 6 11.9 - (not characterised) - 7 13.1 - - 80-fold increase in 0CD 8 14.2

[0410] residence time Binding of multiple 0CDs 9 20.9

[0411] observed 10 27.6 - - GNTN mutants with outward-facing Isoleucine — mimicking extra hydrophobic ring above membrane found in long transmembrane 0-barrel proteins

[0412] 60-fold increase in 0CD 11 14.2

[0413] residence time 12 14.2

[0414] 13 14.2

[0415] 14 20.9

[0416] 15 20.9

[0417] 16 20.9

[0418] GNTN mutants with outward-facing Histidine above membrane boundary 17 14.2 - - 18 14.2 - -

[0419]

[0420] 19 14.2 - -GNTN mutants with inward-facing salt bridges formed by charged amino acids — forming new constriction sites

[0421] 1.3-fold increase in PCD residence time 20 10.2 (E / K pairs form salt bridges — likely too narrow for PCD to pass through) 21 8.6 Similar PCD residence time 22 10.2 (E / K pairs form salt bridges — likely too narrow 23 22.0 for PCD to pass through) 24 22.0 - - 25 8.6 - - 26 8.6 - - GNTN mutants with outward-facing salt bridges formed by charged amino acids — forming outward staples to stabilise the structure

[0422] 2.5-fold increase in PCD 27 9.7

[0423] residence time 28 14.2

[0424] 29 16.4

[0425] 30 20.9

[0426] 31 14.2

[0427] 32 16.4

[0428] 33 20.9

[0429] Unequal insertion into the two strands causes 0-strand mismatch — misfolded structure 34 ~5.9 x (misfolded) - 35 ~5.9 x (misfolded) - 36 ~8.6 x (misfolded) - 37 ~8.6 x (misfolded) - Glycine mutants — Bland barrell

[0430] 1.1 -fold increase in PCD 38 5.8

[0431] residence time 39 6.3 - 40 6.9 - 41 7.4 - SG mutants

[0432] 42 5.8 - 43 6.3 - 44 6.9 - 45 7.4

[0433] 46 5.8

[0434] 47 6.3 - 48 6.9 - 2-fold increase in PCD 49 7.4

[0435] residence time SV mutants — Mimicking transmembrane barrel

[0436]

[0437] 50 5.8 - -6.3 - - 6.9 x (not assembled) - - 7.4 X - - 5.8 - - 6.3 X - - 6.9 X - - 7.4 X - - Repeat the sequence of aHL

[0438] 7.4 - - 9.7 X - - 16.4 X - - Other sequences

[0439] 8.6 - 8.6 - - 8.6 X - - 8.6 - - 14.2 - - 14.2 - - 14.2 - - Barrel extension below the transmembrane region

[0440] 6.3 - 7.4 - 7.4 - 2. Extended 0 barrel mutants of aHL planned

[0441] 6.3

[0442] 7.4

[0443] 8.6

[0444] 9.7

[0445] 10.8

[0446] 8.6

[0447] 9.7

[0448] 10.8

[0449] 15.3

[0450] 16.4

[0451] 17.5

[0452] 18.7

[0453] 19.8

[0454] 6.3

[0455] 6.9

[0456] 8.6

[0457] 7.4

[0458] Extended 0 barrel mutants of chimeric aHL, constructed and characterised

[0459]

[0460] CharacterisationP barrel Gel confirming SDS- Electrical recording P-cyclodextrin (PCD) length resistant oligomer confirming channel binding confirming (nm) formation formation extended binding 88 7.7 x (not assembled) - - 89 8.5 X - - 90 9.2 X - - 91 9.6 X - - 92 9.7 X - - 93 10.1 X - - 94 12.5 X - - 95 8.6 X - - 5.2

[0461]

[0462] A dash (-) indicates that the variant has not yet been fully characterised.

[0463] Table 4. Sequences of the chimeric αHL ADSDINIKTGTTDIGSNTTVKTGDLVTYDKENGMHKKVFYSFIDDKNHNKK LLVIRTKGTIAGQYRVYSEEGANKSGLAWPSAFKVQLQLPDNEVAQISDYYP WS / DTXEEA / STZT’EDTATNWSKTNTYGLSEKVTTKNKFKWPLVGET aHL_AeL_

[0464] 88 ELSIEIAANOSWASQNGGSTTVSIGHTLKYVQPDFKTILESPTDKKVGWK Plan A VIFNNMVNQNWGPYDRDSWNPVYGNQLFMKTRNGSMKAADNFLDPNKAS SLLSSGFSPDFATVITMDRKASKQQTNIDVIYERVRDDYQLHWTSTNWKGT NTKDKWTDRSSERYKID WEKEEMTN* ADSDINIKTGTTDIGSNTTVKTGDLVTYDKENGMHKKVFYSFIDDKNHNKK LLVIRTKGTIAGQYRVYSEEGANKSGLAWPSAFKVQLQLPDNEVAQISDYYP RNSIDTKEYMSTLTYSTTTSGGNOSAWSQNAAIEISLETEGVLPWKFK aHL_AeL_

[0465] 89 NKTTVKESLGYTNTKSWNTATDES / GHTZ^WT’DFATTZESPTDA^F Plan B GWKVIFNNMVNQNWGPYDRDSWNPVYGNQLFMKTRNGSMKAADNFLDP NKASSLLSSGFSPDFATVITMDRKASKQQTNIDVIYERVRDDYQLHWTSTNW KGTNTKDKWTDRSSERYKID WEKEEMTN* ADSDINIKTGTTDIGSNTTVKTGDLVTYDKENGMHKKVFYSFIDDKNHNKK LLVIRTKGTIAGQYRVYSEEGANKSGLAWPSAFKVQLQLPDNEVAQISDYYP

[0466] RNSIDTKEYMSTLTYTCKNTDTVTATTTHTV GTSIQATAKFTVPFNET aHL_Etx_

[0467] 90 GNSYWYSASFANYNYNYNSKE1TVSIGHTLKYVQPDFKTILESPTDKKVG Plan A WKVIFNNMVNQNWGPYDRDSWNPVYGNQLFMKTRNGSMKAADNFLDPN KASSLLSSGFSPDFATVITMDRKASKQQTNIDVIYERVRDDYQLHWTSTNWK GTNTKDKWTDRSSERYKIDWEKEEMTN* ADSDINIKTGTTDIGSNTTVKTGDLVTYDKENGMHKKVFYSFIDDKNHNKK LLVIRTKGTIAGQYRVYSEEGANKSGLAWPSAFKVQLQLPDNEVAQISDYYP WS / DTXEFA / STZT’ENHTIEKSNTNTNTNAFSYSTTLSVGTENFPVTFK aHL_Etx_

[0468] 91 ATAOISTGVTHTTTATVTDTNKCTES7GW77. A'ri / G / 7) / 'A77 / . / '.'. S7>77)A'A' Plan B VGWKVIFNNMVNQNWGPYDRDSWNPVYGNQLFMKTRNGSMKAADNFLD PNKASSLLSSGFSPDFATVITMDRKASKQQTNIDVIYERVRDDYQLHWTSTN WKGTNTKDKWTDRSSERYKID WEKEEMTN*

[0469] aHL_Lys_ ADSDINIKTGTTDIGSNTTVKTGDLVTYDKENGMHKKVFYSFIDDKNHNKK 92

[0470]

[0471] Plan A LLVIRTKGTIAGQYRVYSEEGANKSGLAWPSAFKVQLQLPDNEVAQISDYYPRNSIDTKEYMSTLTYTITKGMKNVNSETRTVTATHSIGSTISTGDAFEI GSVEVSYSHSHEESQVSMTETEVYESKVVSIGHTLKYVQPDFKTILESP TDKKVGWKVIFNNMVNQNWGPYDRDSWNPVYGNQLFMKTRNGSMKAAD NFLDPNKASSLLSSGFSPDFATVITMDRKASKQQTNIDVIYERVRDDYQLHW TSTNWKGTNTKDKWTDRSSERYKID WEKEEMTN* ADSDINIKTGTTDIGSNTTVKTGDLVTYDKENGMHKKVFYSFIDDKNHNKK LLVIRTKGTIAGQYRVYSEEGANKSGLAWPSAFKVQLQLPDNEVAQISDYYP RNSIDTKEYMSTLTYEIVKSEYVETETMSVQSEEHSHSYSVEVSGIEFA aHL_Lys_

[0472] 93 DGTSITSGISHTATVTRTESNVNKMGKTITVSIGHTLKYVQPDFKTILES Plan B PTDKKVGWKVIFNNMVNQNWGPYDRDSWNPVYGNQLFMKTRNGSMKAA DNFLDPNKASSLLSSGFSPDFATVITMDRKASKQQTNIDVIYERVRDDYQLH WTSTNWKGTNTKDKWTDRSSERYKID WEKEEMTN* ADSDINIKTGTTDIGSNTTVKTGDLVTYDKENGMHKKVFYSFIDDKNHNKK LLVIRTKGTIAGQYRVYSEEGANKSGLAWPSAFKVQLQLPDNEVAQISDYYP RNSIDTKEYMSTLTYDQSTQNTDSQTRTISKNTSTSRTHTSEVHGNAEV aHL_PA_

[0473] 94 HASFFDIGGSVSAGFSNSNSSTVAIDHSLSLAGERTWAETMGVSIGHT Long

[0474] LKYVQPDFKTILESPTDKKVGWKVIFNNMVNQNWGPYDRDSWNPVYGNQ LFMKTRNGSMKAADNFLDPNKASSLLSSGFSPDFATVITMDRKASKQQTNI DVTYERVRDDYQLHWTSTNWKGTNTKDKWTDRSSERYKIDWEKEEMTN* ADSDINIKTGTTDIGSNTTVKTGDLVTYDKENGMHKKVFYSFIDDKNHNKK LLVIRTKGTIAGQYRVYSEEGANKSGLAWPSAFKVQLQLPDNEVAQISDYYP

[0475] 7? A'S7 / )7,A'FfA7. S77 / 7TSKNTSTSRTHTSEVHGNAEVHASFFDIGGSVSAG aHL_PA_

[0476] 95 FSNSNSSTVAWHSVSIGHTLKYVQPDFKTILESPTDKKVGWKVIFNNMVN Short QNWGPYDRDSWNPVYGNQLFMKTRNGSMKAADNFLDPNKASSLLSSGFS PDFATVITMDRKASKQQTNIDVIYERVRDDYQLHWTSTNWKGTNTKDKWT DRSSERYKID WEKEEMTN* ADSDINIKTGTTDIGSNTTVKTGDLVTYDKENGMHKKVFYSFIDDKNHNKK LLVIRTKGTIAGQYRVYSEEGANKSGLAWPSAFKVQLQLPDNEVAQISDYYP

[0477] 96** RNSIDTKEYMSTLTYEVHASFFDIGGSVSAGVSLGHTLKYVQPDFKTILES * aHL-CH1 PTDKKVGWKVIFNNMVNQNWGPYDRDSWNPVYGNQLFMKTRNGSMKAA DNFLDPNKASSLLSSGFSPDFATVITMDRKASKQQTNIDVIYERVRDDYQLH

[0478]

[0479] WTSTNWKGTNTKDKWTDRSSERYKID WEKEEMTN*

[0480] Italics: aHL Bold: other pore forming toxins; underlined: barrel sequence

[0481] Table 5. Nanopores containing transmembrane barrels that could be potentially extended

[0482] Table 5A. Extended P barrel mutants of transmembrane P-barrel proteins, constructed Sequence of inserts

[0483] Name Insertion position

[0484] (N to C)

[0485] 2nd3rd4th 1st2nd3rd4th 1ststrand

[0486] strand strand strand strand strand strand strand 1 V150- V177- GNTN GNTN / / / /

[0487] S151 T178

[0488] Cytoxin K 2 V150- V177- (GNTN)2 (GNTN)2 / / / / (CytK) S151 T178

[0489] 3 V150- V177- (GNTN)3 (GNTN)3 / / / /

[0490]

[0491] S151 T178Aerolysin 1 N231- W265- GNTN GNTN / / / / (AeL) T232 A266

[0492] 1 A56- H83- Lysenin (Lys) GNTN GNTN / / / /

[0493] T57 E84

[0494] 1 1144- Y169- GNTN GNTN / / / /

[0495] S145 T170 Enterococcus 2 1144- Y169- (GNTN)2 (GNTN)2 / / / / pore-forming S145 T170

[0496] toxin 1 3 1144- Y169- (GNTN)3 (GNTN)3 / / / / (Epxl) S145 T170

[0497] 4 1144- Y169- (GNTN)4 (GNTN)4 / / / /

[0498] S145 T170

[0499] 1 1129- 1154- GNTN GNTN / / / /

[0500] T130 T155 Enterococcus 2 1129- 1154- (GNTN)2 (GNTN)2 / / / / pore-forming T130 T155

[0501] toxin 4 3 1129- 1154- (GNTN)3 (GNTN)3 / / / / (Epx4) T130 T155

[0502] 4 1129- 1154- (GNTN)4 (GNTN)4 / / / /

[0503] T130 T155

[0504] 1 G137- Q151- Y184- T207- GG GG GG GG

[0505] G138 Y156 E185 S208 2 K135- Q153- L182- N209- GG GG GG GG

[0506] S136 L154 S183 E210 3 K135- Q153- L182- N209- Curli specific TN NT TN NT

[0507] S136 L154 S183 E210 genes G

[0508] 4 K135- Q153- L182- N209- (CsgG) TN ND TN NT

[0509] S136 L154 S183 E210 5 K135- Q153- L182- N209- TNTN NDNT TNTN NTNT

[0510] S136 L154 S183 E210 6 K135- Q153- L182- N209- TNGN NDNT TNGN NGNT

[0511]

[0512] S136 L154 S183 E210 Table 5B. Extended P barrel mutants of transmembrane P-barrel proteins, constructed and characterised.

[0513] Insertion P

[0514] Name barrel Characterisation

[0515] position

[0516] length

[0517] Gel Electrical P-cyclodextrin confirming recording (PCD) binding Relative

[0518] nm SDS-resistant confirming confirming to TMR

[0519] oligomer channel extended formation formation binding 1 Above

[0520] 6.9 A / A / A / TMR

[0521] Cytoxin K 2 Above

[0522] 8.0 V V V (CytK) TMR

[0523] 3 Above

[0524] 9.2 V - -

[0525]

[0526] TMR1 Above

[0527] Aerolysin (AeL) 9.6 A / - - TMR

[0528] 1 Above

[0529] Lysenin (Lys) 10.5 - - - TMR

[0530] 1 Above

[0531] 6.1 - - - TMR

[0532] 2 Above

[0533] Enterococcus 7.2 - - - TMR

[0534] pore-forming

[0535] 3 Above

[0536] toxin 1 (Epxl) 8.4 - - - TMR

[0537] 4 Above

[0538] 9.5 - - - TMR

[0539] 1 Above

[0540] 6.1 - - - TMR

[0541] 2 Above

[0542] Enterococcus 7.2 - - - TMR

[0543] pore-forming

[0544] 3 Above

[0545] toxin 4 (Epx4) 8.4 - - - TMR

[0546] 4 Above

[0547] 9.5 - - - TMR

[0548] 1 Within x(not

[0549] 5.3 - - TMR assembled)

[0550] 2 Above

[0551] 5.3 X - - TMR

[0552] 3 Above

[0553] 5.3 X - - Curli specific TMR

[0554] genes G (CsgG) 4 Above

[0555] 5.3 X - - TMR

[0556] 5 Above

[0557] 5.8 X - - TMR

[0558] 6 Above

[0559] 5.8 X - -

[0560]

[0561] TMR

Claims

CLAIMS1. A method of characterising an analyte, the method comprising:contacting the analyte with a transmembrane protein nanopore having a first opening, a second opening and a channel therebetween, wherein the nanopore comprises one or more heterologous insertions that extend the length of the channel; andtaking one or more measurements characteristic of the analyte as the analyte moves with respect to the nanopore, thereby characterising the analyte.

2. A method according to claim 1, wherein the one or more heterologous insertions are one or more heterologous amino acid sequence insertions, wherein each heterologous amino acid sequence insertion comprises at least two natural and / or non-natural amino acids.

3. A method according to claim 1 or 2, wherein the length of the channel is extended by at least about 0.56 nm, by at least about 1 nm, by at least about 2 nm, by at least about 3 nm, by at least about 4 nm or by at least about 5 nm.

4. A method according to any one of claims 1 to 3, wherein the nanopore comprises a barrel domain and optionally a cap domain, and wherein said one or more heterologous insertions are made in the barrel domain.

5. A method according to claim 4, wherein the barrel domain comprises a P-barrel.

6. A method according to claim 5, wherein each P-strand of the P-barrel independently comprises one or more heterologous insertions.

7. A method according to any one of claims 4 to 6, wherein the nanopore comprises one or more heterologous insertions in a membrane-exposed region of the barrel domain.

8. A method according to any one of claims 4 to 7, wherein the nanopore comprises one or more heterologous insertions in a solvent-exposed region of the barrel domain.

9. A method according to any one of the preceding claims, wherein the nanopore comprises one or more heterologous insertions at a position corresponding to between residues 111 and 128 of SEQ ID NO: 1 and / or between residues 130 and 147 of SEQ ID NO: 1, optionally wherein:the nanopore comprises one or more heterologous insertions at a position corresponding to between residues 111 and 117 of SEQ ID NO: 1 and / or between residues 142 and 147 of SEQ ID NO: 1; and / orthe nanopore comprises one or more heterologous insertions at a position corresponding to between residues 117 and 125 of SEQ ID NO: 1 and / or between residues 134 and 142 of SEQ ID NO: 1; and / orthe nanopore comprises one or more heterologous insertions at a position corresponding to between residues 125 and 134 of SEQ ID NO: 1.

10. A method according to any one of the preceding claims, wherein the nanopore comprises a plurality of heterologous insertions that are substantially coincidental.

11. A method according to any one of the preceding claims, wherein the nanopore comprises a plurality of heterologous amino acid sequence insertions which form antiparallel backbone hydrogen bonds in the nanopore.

12. A method according to any one of the preceding claims, wherein the nanopore comprises two or more constrictions within the channel, wherein the one or more heterologous insertions are positioned between the two or more constrictions.

13. A method according to any one of the preceding claims, wherein the one or more heterologous insertions define a constriction within the channel of the nanopore.

14. A method according to any one of claims 2 to 13, wherein each heterologous amino acid sequence insertion independently comprises at least 4, at least 6, at least 8, at least 10, at least 12, at least 14, at least 16, at least 18, at least 20, at least 24, at least 28, at least 32, at least 36, at least 40, at least 44, at least 48, at least 52, at least 56, at least 60, at least 70 or at least 80 amino acids.

15. A method according to any one of claims 2 to 14, wherein each heterologous amino acid sequence insertion in said nanopore is the same length.

16. A method according to any one of claims 2 to 15, wherein each heterologous amino acid sequence insertion comprises at least two different amino acids.

17. A method according to any one of claims 2 to 16, wherein each heterologous amino acid sequence insertion comprises a sequence motif having a formula (X-Y)nor (Y-X)n; wherein:each X is an amino acid independently selected from the group consisting of R / K / N / DZE / Q / H / P / Y / W / S / T / G / A / M / C / F / L / V / I; preferably each X is an amino acid independently selected from the group consisting of K / N / Y / S / T / G / L / V;each Y is an amino acid independently selected from the group consisting of K / N / D / E / Q / H / P / S / T / G / A / M; preferably each Y is an amino acid independently selected from the group consisting of K / N / E / S / T / G / M; andn is at least 1.

18. A method according to any one of claims 2 to 17, wherein each heterologous amino acid sequence insertion each independently comprises a sequence motif selected from GNTN, GNTNEN, KNGNTN, GG, SG, SV, EYMSTLTY, VSIGHTLK, TN, ST, SI, TNGN, KNGN, GNEN, TNKN, ENTN, GNSN, NGNT, NKNGNT, (GNTN)3-GNTIGITN-(GNTN)3, (GNTN)2-GITNGITN-(GNTN)4, (GNTN)3-GITNGITN-(GNTN)3, (GNTN)4-GITNGITN-(GNTN)8, (GNTN)6-GITNGITN-(GNTN)6, (GNTN)6-GITIGITN-(GNTN)6, GNTHGHTN-(GNTN)6, (GNTN)6-GHTHGNTN, (GNTN)2-EN-(GNTN)2, (GNTN)2-KN-(GNTN)2, (GNTNEN)4-GNTIEIGITNEN-(GNTNEN)4, (KNGNTN)4- KNGNTIKIGITN-(KNGNTN)4, TNGK, TNGD, TNGKTNGN, TNGDTNGN, (GNTN)3G, NTN(GNTN)2, GQSQ, GQTQ, NT, ND and NDNT.

19. A method according to claim 17 or 18, wherein each heterologous amino acid sequence insertion independently comprises one or more repeats of said sequence motifs.

20. A method according to claim 19, wherein each heterologous amino acid sequence insertion comprises from 1 to 20 repeats, preferably from 2 to 14 repeats, more preferably from 3 to 8 repeats.

21. A method according to any one of the preceding claims, wherein the nanopore is an oligomeric nanopore comprising a plurality of monomer subunits;wherein each monomer subunit comprises at least one or more of said heterologous insertions.

22. A method according to any one of the preceding claims, wherein the nanopore is a homooligomeric nanopore, wherein each monomer subunit comprises the same one or more heterologous insertions.

23. A method according to any one of claims 1 to 21, wherein the nanopore is a heterooligomeric nanopore, and wherein at least one monomer subunit comprises said one or more heterologous insertions.

24. A method according to any one of the preceding claims, wherein the analyte is selected from a nucleotide, a polynucleotide, an amino acid, a polypeptide, a saccharide and a polysaccharide.

25. A transmembrane protein nanopore having a first opening, a second opening and a channel therebetween, wherein the nanopore comprises one or more heterologous insertions that extend the length of the channel.

26. A transmembrane protein nanopore according to claim 25, wherein the nanopore is as defined in any one of claims 1 to 23.

27. A device comprising an array of transmembrane protein nanopores according to claim 25 or 26 comprised in a membrane.

28. A method of increasing the interaction dwell time of an analyte with a transmembrane protein nanopore comprising a first opening, a second opening and a channel therebetween, the method comprising modifying the nanopore by introducing one or more heterologous insertions into the nanopore, thereby extending the length of the channel of the nanopore.

29. A method according to claim 28, wherein the nanopore, the one or more heterologous insertions and / or the analyte are as defined in any one of claims 1 to 26.