Novel protein pore

By introducing modified CsgF peptides into CsgG pores to form CsgG:CsgF complex, the problem of insufficient differences in current characteristics of CsgG pores is solved, the current characteristics of polynucleotide sequencing is improved, and the analyte sequence recognition ability is enhanced.

CN117106037BActive Publication Date: 2025-07-25VLAAMS INTERUNIVERSITAIR INST VOOR BIOTECHNOLOGIE VZW +2
View PDF 13 Cites 0 Cited by

Patent Information

Application Number
CN202310895485.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2017-06-30
Filing Date
2018-07-02
Publication Date
2025-07-25
Estimated Expiration
2038-07-02

AI Technical Summary

Technical Problem

The sequence dependence of existing CsgG pores in polynucleotide sequencing is poor, affecting the performance of sequencing systems, and it is necessary to improve the nanopore sensing characteristics to improve the current difference between nucleotides.

Method used

By using the modified CsgF peptide to bind to the CsgG pore, additional channel contraction segments are introduced to form a CsgG:CsgF complex, enhancing the current characteristic difference of the nanopore sensing platform.

Benefits of technology

The current characteristic difference of polynucleotide sequencing is improved, the recognition ability of analyte sequences is enhanced, and the performance of the sequencing system is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure GDA0004518353380000141
    Figure GDA0004518353380000141
  • Figure GDA0004518353380000151
    Figure GDA0004518353380000151
  • Figure GDA0004518353380000221
    Figure GDA0004518353380000221
Patent Text Reader

Abstract

The present invention relates to novel protein pores and their use in analyte detection and characterization. In particular, the present invention relates to an isolated pore complex formed by a CsgG-like pore and a modified CsgF peptide, or homologs or mutants thereof, thereby incorporating additional channel constriction segments or reading heads into the nanopore. The present invention also relates to a transmembrane pore complex and a method for producing the pore complex, and its use in molecular sensing and nucleic acid sequencing applications.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This patent application is a divisional application of the patent application with application number 2018800440243, application date July 2, 2018, and invention name “Novel Protein Pore”. Technical Field

[0002] The present invention relates to novel protein pores and their use in analyte detection and characterization. The present invention also relates to a transmembrane pore complex and a method for producing the pore complex, and their use in molecular sensing and nucleic acid sequencing applications. Background Art

[0003] Nanopore sensing is a method for analyte detection and characterization that relies on observing single binding or interaction events between analyte molecules and ion-conducting channels. Nanopore sensors are created by placing a single nanometer-sized pore in an electrically insulating membrane and measuring the voltage-driven ionic current through the pore in the presence of analyte molecules. The presence of an analyte inside or near the nanopore alters the ion flux through the pore, resulting in a modified ion or current measured across the channel. The identity of the analyte is revealed by its unique current signature, particularly the duration and extent of current blockade and the change in current level during the interaction of the analyte with the pore. Analytes can be organic and inorganic small molecules, as well as various biological or synthetic macromolecules and polymers, including polynucleotides, polypeptides, and polysaccharides. Nanopore sensing can reveal the identity of the sensed analyte and perform single-molecule counting of the sensed analyte, but can also provide information on analyte composition, such as nucleotide, amino acid, or glycan sequence, and the presence of base, amino acid, or glycan modifications, such as methylation and acylation, phosphorylation, hydroxylation, oxidation, reduction, glycosylation, decarboxylation, and deamination. Nanopore sensing has the potential to enable rapid and low-cost polynucleotide sequencing, providing single-molecule sequence reads of polynucleotides ranging from tens to tens of thousands of bases in length.

[0004] Two fundamental components of polymer characterization using nanopore sensing are: (1) controlling the movement of the polymer through the pore; and (2) discriminating the building blocks of the polymer as it moves through the pore. During nanopore sensing, the narrowest portion of the pore forms the read head, and the most discriminative portion of the nanopore, in terms of current characteristics, varies with the analyte being passed. CsgG was identified as an ungated, non-selective protein secretion channel from Escherichia coli (Goyal et al., 2014) and has been used as a nanopore for detecting and characterizing analytes. Mutations of the wild-type CsgG pore that improve the properties of the pore in this context have also been disclosed (WO2016 / 034591, WO2017 / 149316, WO2017 / 149317, and WO2017 / 149318, PCT / GB2018 / 051191, all incorporated herein by reference).

[0005] For polynucleotide analytes, nucleotide discrimination is achieved through passage through this mutant pore. However, the current signature has been shown to be sequence-dependent, with multiple nucleotides contributing to the observed current. This means that the height of the channel constriction and the extent of surface interaction with the analyte influence the relationship between the observed current and the polynucleotide sequence. While mutations in the CsgG pore improve the current range for nucleotide discrimination, further improvements in current variability between nucleotides would lead to even higher performance in sequencing systems. Therefore, there is a need to identify novel approaches to improve nanopore sensing properties. Summary of the Invention

[0006] The present disclosure relates to modified CsgF peptides, particularly truncated CsgF fragments, that bind to the CsgG pore and thereby introduce an additional channel or pore constriction within the CsgG pore. Further aspects of the invention relate to isolated transmembrane pore complexes and the use of such CsgG:CsgF complexes and modified CsgF peptides or fragments in nanopore sensing platforms with two sequential read heads.

[0007] A first aspect of the present invention relates to a pore comprising a CsgG pore and a CsgF peptide. In one aspect, the CsgF peptide comprises a CsgG binding region and a region that forms a constriction in the pore. In one aspect, the CsgF peptide is a truncated CsgF peptide lacking the C-terminal head domain of CsgF. In another aspect, the CsgF peptide is a truncated CsgF peptide lacking the C-terminal head domain and a portion of the neck domain of CsgF. In another aspect, the CsgF peptide is a truncated CsgF peptide lacking the C-terminal head domain and a portion of the neck domain of CsgF. Pores are also referred to herein as pore complexes and isolated pore complexes. Isolated pore complexes comprise the CsgG pore, or a homologue or mutant thereof, and a modified CsgF peptide, or a homologue or mutant thereof, particularly a truncated CsgF fragment, or a homologue or mutant thereof. In one embodiment, the modified CsgF peptide, or a homologue or mutant thereof, is located within the lumen of the CsgG pore, or a homologue or mutant thereof. In another embodiment, the isolated pore complex has two or more channel constrictions, one channel constriction positioned or provided by the CsgG pore, formed by its contractile ring, and an additional channel constriction or read head introduced by a modified CsgF peptide, or a homolog or mutant thereof. In one embodiment, the CsgG pore or CsgG-like pore is not a wild-type pore, but rather a mutant CsgG pore, in specific embodiments, for example, wherein a mutation is present in the channel constrictive ring. In another embodiment, the isolated pore complex comprising a modified CsgF peptide, or a homolog or mutant thereof, has a CsgF channel constriction with a diameter ranging from 0.5 nm to 2.0 nm. In one embodiment, a pore complex comprises: (i) a CsgG pore comprising a first opening, a middle section comprising a β-barrel, a second opening, and a lumen extending from the first opening through the middle section to the second opening, wherein the lumenal surface of the middle section defines a CsgG constriction; and (ii) a plurality of modified CsgF peptides, each modified CsgF peptide having a CsgF constriction and a CsgG binding region (also referred to herein as a CsgG binding domain or binding region of CsgF), wherein the modified CsgF peptides form a CsgF constriction within the β-barrel of the CsgG pore, and wherein the CsgG constriction and the CsgF constriction are coaxially spaced apart within the β-barrel of the CsgG pore. The lumenal surface of the CsgG pore may comprise one or more loop regions of a CsgG monomer that define the CsgG constriction. The CsgF constriction and the CsgF binding region typically correspond to the N-terminal portion of the mature CsgF peptide. In one embodiment, the pore complex does not include CsgA, CsgB, and CsgE.

[0008] In a second aspect, the present invention relates to a modified CsgF peptide, or a homolog or mutant thereof, wherein the protein or peptide is modified by truncation or deletion of a portion of the protein to produce a CsgF fragment of SEQ ID NO: 6, or a homolog or mutant thereof. One embodiment relates to a modified or truncated CsgF peptide, or a modified peptide of a CsgF homolog or mutant, comprising SEQ ID NO: 39 or SEQ ID NO: 40, or a homolog or mutant thereof, alternatively comprising SEQ ID NO: 15, or a homolog or mutant thereof, alternatively comprising SEQ ID NO: 54 or SEQ ID NO: 55, or a homolog or mutant thereof. Another embodiment discloses a modified CsgF peptide, wherein one or more positions in the region comprising SEQ ID NO: 15 are modified, and wherein the mutations are required to maintain at least 35% amino acid identity with SEQ ID NO: 15 in the peptide fragment corresponding to the region comprising SEQ ID NO: 15.

[0009] One embodiment relates to a pore comprising a CsgG pore and a modified CsgF peptide, wherein the modified CsgF peptide binds to CsgG and forms a constriction in the pore.

[0010] One embodiment relates to a polynucleotide encoding the modified CsgF peptide according to the second aspect of the invention, or a homologue or mutant thereof. In another embodiment, the isolated pore complex comprising a CsgG pore and a modified CsgF peptide, or a homologue or mutant thereof, is characterized in that the modified CsgF peptide is a peptide provided by a peptide disclosed in the second aspect of the invention.

[0011] Another embodiment relates to an isolated pore complex wherein the modified CsgF peptide is covalently coupled to the CsgG pore or a monomer of the pore, or a homologue or mutant thereof. Even more particularly, the coupling is via a cysteine ​​residue or via a non-native reactive or photoreactive amino acid in the CsgG monomer at a position corresponding to 132, 133, 136, 138, 140, 142, 144, 145, 147, 149, 151, 153, 155, 183, 185, 187, 189, 191, 201, 203, 205, 207 or 209 of SEQ ID NO: 3 or a homologue thereof.

[0012] A preferred embodiment relates to an isolated transmembrane pore complex or membrane composition comprising the isolated pore complex of the present invention and components of a membrane. Specifically, the transmembrane pore complex or membrane composition consists of the isolated pore complex of the present invention and components of a membrane or insulating layer.

[0013] One embodiment relates to a method for producing a pore as disclosed herein, comprising co-expressing one or more CsgG monomers as disclosed herein and a CsgF peptide as disclosed herein in a host cell, thereby allowing formation of a transmembrane pore complex in the cell. The CsgF peptide can be produced by cleavage of a modified CsgF peptide or protein comprising an enzymatic cleavage site at a suitable position in the amino acid sequence.

[0014] One embodiment relates to a method for producing a pore as disclosed herein, comprising contacting one or more purified CsgG monomers with one or more purified modified CsgF peptides, thereby allowing pore formation in vitro. The modified CsgF peptide may be a peptide comprising an enzymatic cleavage site at a suitable position in the amino acid sequence, which is cleaved before or after pore formation.

[0015] A third aspect of the present invention relates to a method for producing the transmembrane pore complex, wherein the pore is an isolated complex formed by a CsgG pore, or a homologue or mutant thereof, and a modified CsgF peptide, or a homologue or mutant thereof, the method comprising the steps of co-expressing CsgG (SEQ ID NO: 2), or a homologue or mutant thereof, and a modified or truncated CsgF (comprising a fragment of SEQ ID NO: 5), or a homologue or mutant thereof, in a suitable host, thereby allowing for in vivo formation of the pore complex. In specific embodiments, the modified CsgF peptide, or a homologue or mutant thereof, comprises SEQ ID NO: 12 or SEQ ID NO: 14, or a homologue or mutant thereof. Alternatively, the method for producing the isolated pore complex comprises the steps of contacting a CsgG monomer of SEQ ID NO: 3, or a homologue or mutant thereof, with a modified CsgF peptide, or a homologue or mutant thereof, to reconstitute the pore complex in vitro. In certain embodiments, the modified CsgF peptide of the method comprises SEQ ID NO: 15 or SEQ ID NO: 16, or a homologue or mutant thereof.

[0016] Another aspect of the present invention relates to a method for determining the presence, absence or one or more characteristics of a target analyte, the method comprising the steps of:

[0017] (i) contacting a target analyte with the isolated pore complex or transmembrane pore complex such that the target analyte moves into the pore channel; and

[0018] (ii) performing one or more measurements as the analyte moves through the pore channel to determine the presence, absence, or one or more characteristics of the analyte.

[0019] In one embodiment, the analyte is a polynucleotide. Specifically, the method using a polynucleotide as an analyte may alternatively comprise determining one or more characteristics selected from the group consisting of: (i) the length of the polynucleotide, (ii) the identity of the polynucleotide, (iii) the sequence of the polynucleotide, (iv) the secondary structure of the polynucleotide, and (v) whether the polynucleotide is modified.

[0020] In another embodiment, the analyte is a protein or peptide, and in other embodiments, the analyte is a polysaccharide or a small organic or inorganic compound, such as but not limited to pharmacologically active compounds, toxic compounds, and pollutants.

[0021] In another embodiment, a method for characterizing a polynucleotide or (poly)peptide using an isolated transmembrane pore complex is described, wherein the pore complex is an isolated complex comprising a CsgG pore, or a homologue or mutant thereof, and a modified CsgF peptide, or a homologue or mutant thereof. Specifically, the CsgG pore, or a homologue or mutant thereof, comprises six to ten CsgG monomers that form the CsgG pore channel.

[0022] Another aspect of the present invention discloses the use of the isolated pore complex or transmembrane pore complex according to the aforementioned aspects of the invention for determining the presence, absence or one or more characteristics of a target analyte. In addition, the present invention also relates to a kit for characterizing a target analyte, comprising (a) the isolated pore complex and (b) a component of a membrane. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The drawings described are only schematic and non-limiting. For illustrative purposes in the drawings, the size of some elements may be exaggerated and not drawn to scale.

[0024] Figure 1: Structure of the CsgG pore and its interface in complex with CsgFSurface (A) and ribbon (B, C) representations show cross-sectional (A), side (B), and top (C) views of a CsgG oligomer (e.g., nonamer) (gold), with a single CsgG precursor in light blue (D) (based on the CsgG X-ray structure PDB entry: 4uv3). The CsgG constrictive loop (CL loop) spans residues 46 to 61 of SEQ ID NO: 3, is shown in dark gray in all figures, and corresponds to the loop provided in the lower left of (E). CsgG residues whose side chains face the lumen of the CsgG β-barrel are shown in medium gray and are labeled in the β-strands of (E) and (D). These residues represent sites available for substitution of natural or unnatural amino acids, e.g., sites suitable for attachment (e.g., covalent cross-linking) of a pore-retaining peptide (including, for example, a modified CsgF peptide or homolog thereof) to the CsgG pore or monomer. In some embodiments, the cross-linking residues include Cys and reactive and photoreactive amino acids such as azidohomoalanine, homopropargylglycine, homoallylglycine, p-acetyl-Phe, p-azido-Phe, p-propargyloxy-Phe, and p-benzoyl-Phe (Wang et al., 2012; Chin et al., 2002), and can be substituted at positions 132, 133, 136, 138, 140, 142, 144, 145, 147, 149, 151, 153, 155, 183, 185, 187, 189, 191, 201, 203, 205, 207, or 209 according to SEQ ID NO: 3. (E) A magnified view of the CL loop and transmembrane β strands of a CsgG monomer is shown. The CsgG constriction loop (dark blue) forms the pore opening or narrowest passage in the CsgG pore (panel A). In some embodiments, three positions 56, 55, and 51 in the CL loop are particularly important for the diameter, chemical, and physical properties of the CsgG channel pore or "read head," according to SED ID NO: 3. These represent preferred positions for modifying the nanopore sensing properties of the CsgG pore and homologs.

[0025] Figure 2: Co-expression of CsgG:CsgF complex proteins and purification of the complex(A) Schematic diagram of the purification protocol for the CsgG:CsgF complex starting from E. coli cultures co-expressing CsgG (SEQ ID NO:2 + C-terminal StrepII tag) and CsgF (SEQ ID NO:4 + C-terminal 6xHis tag). The protocol involves disrupting the resuspended cells and performing a 1% DDM extraction of membrane-bound proteins. The CsgG:CsgF complex and excess CsgF are first enriched by affinity purification on a nickel IMAC column, followed by a second affinity-based enrichment of the CsgG:CsgF complex on a streptavidin column. (B) Coomassie-stained SDS-PAGE of the IMAC (left) and streptavidin (right) purification steps. Protein bands corresponding to CsgG and CsgF are labeled. Notably, the IMAC eluate contained an N-terminally truncated CsgF fragment (marked *), which was not retained in the affinity pull-down using the CsgG-bound Strep tag, indicating that the CsgF N-terminus is required for complex formation with CsgG.

[0026] Figure 3: In vitro Purification of the reconstituted CsgG:CsgF complex protein. (A) Overlay of size exclusion chromatography (SEC) chromatograms (using a BioRad Enrich 650 10 / 300 column) of CsgG (light gray) and CsgG supplemented with excess CsgF (dark gray). The chromatograms show elution peaks corresponding to: CsgG 9-mers (a) and CsgG 18-mers (b) for CsgG chromatography; excess free CsgF (c), as well as 9-mer CsgG:CsgF complexes (d) and 18-mer CsgG:CsgF complexes (e), which elute at higher hydrodynamic radii (molecular weights) due to CsgF incorporation into the complexes. (B) Native PAGE analysis of representative species labeled in panel (A), demonstrating the shift to higher molecular weights due to CsgF incorporation into the CsgG 9-mer and CsgG 18-mer complexes. These experiments demonstrate that the CsgG:CsgF complex can be reconstituted in vitro starting from purified components. (C) Ribbon representation of the CsgG 9amer and CsgG 18amer previously reported in Goyal et al., 2014 (PDB entry 4uv3). The CsgG 18amer is formed from dimers of the CsgG 9amer. SEC and native PAGE analysis shown in panels A and B demonstrate that the CsgG 18amer is suitable for complex formation with CsgF.

[0027] Figure 4: CsgG:CsgF structure determined by cryo-electron microscopy (cryo-EM)(A) Cryo-electron micrographs of CsgG:CsgF complexes show the presence of 9- and 18-mer CsgG:CsgF complexes, with numerous individual particles of the 9- and 18-mer forms highlighted by full and dashed circles, respectively. (B) Two representative class averages of CsgG:CsgF 9-mer complexes viewed from the side. The class averages include 6,020 and 4,159 individual particles, respectively. The class averages reveal the presence of additional density above the CsgG particles, corresponding to oligomeric complexes of CsgF. Three distinct regions are visible in the CsgF oligomer: the "head" and "neck" regions, and a region that resides within the lumen of the CsgG β-barrel and forms a constriction or narrow passage (labeled F) that stacks above the constriction formed by the CsgG CL loop (labeled G). This latter CsgF region is termed the CsgF contractile peptide (FCP).

[0028] Figure 5 Three-dimensional structural model of the CsgG:CsgF complex Cross-sectional view of the 3D cryoEM electron density for the CsgG:CsgF 9-mer complex, calculated from 20,000 particles assigned to 21 class averages. The right panel shows the superposition of the docked CsgG 9-mer X-ray structure (PDB entry: 4uv3) into the cryoEM density. The regions corresponding to CsgG, CsgF, and the CsgF head, neck, and FCP domains are indicated. The cross-section shows that the CsgF FCP domain forms an additional constriction (labeled F) approximately 2 nm above the CsgG contractile ring (labeled G) within the CsgG channel.

[0029] Figure 6: Schematic representation of the CsgG:CsgF pore complex based on the cryo-EM structure. (A) Schematic representation of a CsgG nanopore with a single constriction (labeled (1)) shown in cross-sectional view. The CsgG-based nanopore forms a 3.5–4 nm wide channel with a 0.5–1.5 nm pore opening formed by the CsgG constriction ring (residues 46 to 61 according to SEQ ID NO: 3). Upon complexing with CsgF, a second constriction or pore is introduced into the CsgG channel (labeled (2) / F), and the channel exit is blocked by the CsgF head domain (see Figure 5 When a modified CsgF peptide (e.g., corresponding to a CsgF constricting peptide (FCP) lacking the neck and head regions) is used, a CsgG:CsgF pore complex with two continuous channel constrictions or pores ((1) and (2)) is formed, as shown in the cross-sectional view of the CsgG:CsgF cryo-EM density in (B) and the schematic diagram in (C). Removal of the neck and head regions in the modified CsgF peptide alleviates their blockage of the channel outlet.

[0030] Figure 7: CsgG:CsgF pore complex for nanopore sensing of (bio)polymers (A) or single-molecule analytes (B) A schematic diagram of the application's purpose.When used for polymer sensing, the second channel constriction introduced by the modified CsgF peptide increases the contact area with the analyte and forms a second interaction site and readhead. When used for single-molecule nanopore sensing, the second channel constriction introduced by the modified CsgF peptide creates a second, independent analyte interaction site. (C) Schematic illustration of the theoretical channel conductance curves for a small molecule (represented by hexagons or triangles) passing through and interacting with the continuous CsgG (1) and CsgF (2) constrictions or readhead.

[0031] Figure 8 Multiple sequence alignment of exemplary CsgF homologs The aligned sequences are shown as mature proteins (i.e., lacking their N-terminal signal peptide (SP)). In some embodiments, the boxed sequences indicate CsgF regions where the sequences are conserved (pairwise sequence identities are between 35 and 100% - see Figure 10 ), which corresponds to the CsgF contractile peptide (FCP). The CsgF homologs included in the multiple sequence alignment are Q88H88; A0A143HJA0; Q5E245; Q084E5; F0LZU2; A0A136HQR0; A0A0W1SRL3; B0UH01; Q6NAU5; G8PUY5; A0A0S2ETP7; E3I1Z1; F3Z094; A0A176T7M2; D2QPP8; N2IYT1; W7QHV5; D4ZLW2; D2QT92; A0A167UJA2. The FCP region of E. coli CsgF (SEQ ID NO: 15) and the CsgF homologs shown correspond to SEQ ID NOs: 18-36.

[0032] FIG9 : Experimental evaluation of the E. coli CsgF region that forms the CsgG interacting sequence and the CsgF contractile peptide (FCP).Panel (A) shows the mature sequences (i.e., after removal of the CsgF signal peptide, which corresponds to residues 1-19 of SEQ ID NO:5) of four N-terminal CsgF fragments (SEQ ID NO:8 - CsgF residues 1-27; SEQ ID NO:10; SEQ ID NO:12, and SEQ ID NO:14) co-expressed with E. coli CsgG (SEQ ID NO:2). (B) Anti-Strep (left) and anti-His (right) Western blot analysis of crude cell lysates from CsgG and CsgF co-expression experiments. Anti-Strep analysis demonstrated CsgG expression in all co-expression experiments, while anti-His Western blot analysis showed detectable levels of CsgF fragments only for the truncation mutant CsgF 1-64 (SEQ ID NO:14). A His-tagged Nanobody (Nb) was used as a positive control. (C) Anti-His dot blot analysis of the presence of CsgF fragments in CsgG:CsgF co-expression experiments. The upper row shows whole cell lysates, and the middle and lower rows show the eluate and flowthrough from the Strep affinity pull-down experiment. These data demonstrate that CsgF fragment 1-64, and to a lesser extent CsgF 1-48, are specifically pulled down as complexes with Strep-tagged CsgG. CsgF fragments 1-27 and 1-38 did not produce detectable levels of the corresponding CsgF fragments and showed no signs of complex formation with CsgG.

[0033] Figure 10 : Formation of a multiple sequence alignment of the CsgG interacting sequence and the CsgF region of the CsgF contractile peptide (FCP). This figure shows a multiple sequence alignment and consensus sequence of CsgF peptides and their known homologs in the region corresponding to interaction with CsgG. CsgF homologs are defined by the PFAM domain PF03783. These peptides bind to CsgG and localize to the lumen of the CsgG β-barrel, where they form an additional constriction in the CsgG channel. These peptides and their homologs are examples of CsgF contractile peptides, or FCPs. Pairwise sequence identity among the FCPs shown ranges between 35% and 98%.

[0034] Figure 11: High-resolution cryoEM structure of the CsgG:CsgF complex CsgG is shown in light grey and CsgF is shown in dark grey. A. CsgG:CsgF complex in A. Final electron density map at 100 nm resolution. Side view. B. Top view showing the cryoEM structure of CsgG:CsgF, which contains a 9:9 stoichiometry and has C9 symmetry. C. Internal architecture of the CsgG:CsgF complex. GC, CsgG constriction; FC, CsgF constriction. D. Interactions between the CsgG and CsgF proteins. CsgG and the CsgG constriction are light gray and gray, respectively. CsgF is dark gray. Residues in CsgG and CsgF are labeled in light gray and black, respectively.

[0035] Figure 12 The two read heads of the :CsgG:CsgF complex. CsgG is shown in light gray, and the read head of the CsgG pore is shown in dark gray. CsgF is shown in black, and the read head of CsgF is marked.

[0036] Figure 13 Co-expression of CsgG and CsgFWT in vivo A gene encoding a C-terminally Strep-tagged CsgG polypeptide in the ampicillin-resistant pT7 vector and a gene encoding a C-terminally His-tagged CsgF polypeptide in the kanamycin-resistant pRham vector were co-transformed into Escherichia coli BL21DE3 cells in the presence of both ampicillin and kanamycin. Proteins were expressed overnight at 18°C, 250 rpm, and the CsgG-CsgF complex was purified using a Strep tag followed by a His tag. A. Protein samples before Strep purification (duplicate). B. Protein samples after His purification (three elution fractions). Proteins were electrophoresed on a 4-20% Tris gel.

[0037] Figure 14 In vitro co-expression of CsgG and CsgF and the thermal stability of the CsgG-CsgF complex CsgG and CsgF DNA from different vectors were co-expressed in in vitro transcription and translation reactions. The proteins were radiolabeled with S-35 methionine and exposed to X-ray film. The stability of the complex was assessed by incubating the reaction mixture at different temperatures for 10 minutes.

[0038] FIG15 : Preparation of CsgG:CsgF complex using protease cleavage sites.A. A TEV or C3 or any other protease cleavage site can be incorporated into the CsgF peptide at the desired position (e.g., between 30 and 31, 35 and 36, 40 and 41, 45 and 46 of seq ID No. 6). CsgG is shown in gold and the CsgF domain is shown in red. For clarity, 1-35 of a CsgF subunit are green. 36-45 are shown in purple. The 10-histidine tag is shown in pink and the strep tag on CsgG is shown in blue. B. SDS-PAGE (4-20% TGX) of protease cleavage of the full-length CsgG:CsgF complex, in which a TEV protease cleavage site was inserted between 35-36 of seq ID 6. M: Molecular weight marker. Lane 1: Full-length CsgG:CsgF complex after strep purification. Lane 2: Concentrate after strep purification. Lane 3: After gel filtration. Lane 4: CsgG:CsgF complex cleaved by TEV protease. Lane 5: Flow-through of CsgG:CsgF after strep purification. Lane 6: CsgG:CsgF heated at 60°C for 10 minutes. Lane 7: CsgG:CsgF complex eluted from a streptavidin column. Lane 8: CsgG well as a control. Lane 9: TEV protease as a control.

[0039] Figure 16: Thermal stability of the CsgG:CsgF complex M: molecular weight marker, lane 1: CsgG pore, lane 2: CsgG:CsgF complex at room temperature: lanes 3-9: CsgG:CsgF samples were heated at different temperatures (40, 50, 60, 70, 80, 90, 100°C) for 10 minutes. Lane 1:

[0040] A.Y51A / F56Q / N55V / N91R / K94Q / R97W-del(V105-I107):CsgF-(1-45).

[0041] B.Y51A / F56Q / N55V / N91R / K94Q / R97W-del(V105-I107):CsgF-(1-35).

[0042] C.Y51A / F56Q / N55V / N91R / K94Q / R97W-del(V105-I107):CsgF-(1-30).

[0043] The samples were subjected to SDS-PAGE on 7.5% TGX gels. The CsgG:CsgF complex with both CsgF-(1-45) and CsgF-(1-35) showed a shift from the CsgG pore band in lane 1. Therefore, it is clear that both complexes are thermostable up to 90°C. The complex and the pore dissociated into CsgG monomers at 100°C (lane 9). Although the same thermostability pattern was seen with the CsgG:CsgF complex and CsgF-(1-30), it was difficult to see a shift between the protein bands of the CsgG pore (lane 1) and the CsgG-CsgF complex (lanes 2-8).

[0044] Figure 17 : CsgG:CsgF was formed by in vitro reconstitution using a synthetic CsgF peptide. Native PAGE of CsgG:CsgF reconstitution in vitro using wild-type CsgG or a CsgG mutant with an altered constriction segment, Y51A / F56Q / K94Q / R97W / R192D-del (V105-1107), is shown. Alexa 594-labeled CsgF peptide, corresponding to the first 34 residues of mature CsgF (Seq ID No. 6), was added to purified Strep-tagged CsgG or Y51A / F56Q / K94Q / R97W / R192D-del (V105-1107) in 50 mM Tris, 100 mM NaCl, 1 mM EDTA, and 5 mM LDAO / C8D4 for 15 minutes at room temperature to reconstitute the CsgG-strep. After CsgG-strep was pulled down onto StrepTactin beads, the samples were analyzed on native PAGE. Both WT and Y51A / F56Q / K94Q / R97W / R192D-del(V105-1107) CsgG bound the CsgF N-terminal peptide, as visualized by the fluorescent label.

[0045] FIG18 : Stabilization of CsgG:CsgF or CsgG:FCP complexes. A. Identified amino acid positions that can generate SS bonds in the pairing of CsgG (SEQ ID NO: 3) and CsgF (SEQ ID NO: 6). B. Schematic diagram showing the SS bond between CsgG-Q153C and CsgF-G1C.

[0046] FIG19 : Cysteine ​​cross-linking of the CsgG:CsgF complex.A. Y51A / F56Q / N91R / K94Q / R97W / Q153C-del(V105-I107) and CsgF-G1C proteins were purified individually and incubated together at 4°C for 1 hour or overnight to form a complex and allow SS formation. No oxidizing agent was added to promote SS formation. Control CsgG wells (Y51A / F56Q / N91R / K94Q / R97W / Q153C-del(V105-I107)) and complexes (with and without DTT) were heated at 100°C for 10 minutes to dissociate the complex into CsgG monomers (CsgG m , 30 kDa) and CsgF monomer (CsgF m , 15 kDa). In the absence of reducing agents, CsgG can be seen m and CsgF m Dimers between (CsgG m -CsgF m , 45 kDa), confirming the formation of SS bonds. Overnight incubation resulted in increased dimer formation compared to one-hour incubation. B. Gel-purified CsgG from overnight incubations m -CsgF m The bands were subjected to mass spectrometry analysis. The protein was proteolytically cleaved to generate tryptic peptides. LC-MS / MS sequencing was performed to identify the precursor ion corresponding to the indicated linker peptide. This precursor ion was fragmented to yield the observed fragment ions. These included ions for each peptide, as well as fragments incorporating intact disulfide bonds. This data provides strong evidence for the presence of a disulfide bond between C1 of CsgF and C153 of CsgG.

[0047] Figure 20 : Improve the cysteine ​​cross-linking efficiency of CsgG:CsgF complex. Lane 1: Y51A / F56Q / N91R / K94Q / R97W / N133C-del (V105-I107) and CsgF-T4C proteins were co-expressed and the CsgG:CsgF complex was purified. Lane 2: The complex was heated in the presence of DTT to dissociate the complex into substituent monomers (CsgG m and CsgF m ). If formed, DTT will destroy all SS bonds between CsgG-N133C and CsgF-T4C. Lane 3: The complex is incubated with the oxidant copper-o-phenanthroline to promote SS bond formation. Lane 4: In the absence of DTT, the oxidized sample is heated at 100°C to decompose the complex. m -CsgF m The corresponding new band of 45 kDa confirmed the formation of SS bond.

[0048] Figure 21: Current characteristics of a DNA strand passing through a CsgG:CsgF complex.Complexes were prepared by co-expressing a CsgG pore (Y51A / F56Q / N91R / K94Q / R97W-del (V105-I107)) containing a C-terminal Strep tag with a full-length CsgF protein containing a C-terminal His tag and a TEV protease cleavage site between 35 and 36 of seq ID no. 6. The complex was then purified by TEV protease cleavage to produce the designated CsgG:CsgF complex. Note that TEV cleavage leaves an ENLYFQ sequence at the cleavage site. A. No mutation at position 17 of CsgF. B. N17S mutation in CsgF.

[0049] Figure 22: Current characteristics of a DNA strand passing through a CsgG:CsgF complex. Complexes were prepared by incubating Y51A / N55V / F56Q / N91R / K94Q / R97W-del (V105-I107) wells containing a C-terminal Strep tag with CsgF-(1-35) mutants. A. CsgF-N17S-(1-35). B. CsgF-N17V-(1-35).

[0050] Figure 23: Current characteristics of a DNA strand passing through a CsgG:CsgF complex.

[0051] Complexes were prepared by incubating different CsgG pores containing a C-terminal Strep tag with CsgF-N17S-(1-35). A. The CsgG pore is Y51A / N55V / F56Q / N91R / K94Q / R97W-del (V105-I107). B. The CsgG pore is Y51T / N55V / F56Q / N91R / K94Q / R97W-del (V105-I107). C. The CsgG pore is Y51A / N55I / F56Q / N91R / K94Q / R97W-del (V105-I107). D. The CsgG pore is Y51A / F56A / N91R / K94Q / R97W-del (V105-I107). E. The CsgG pore is Y51A / F56I / N91R / K94Q / R97W-del (V105-I107). F. The CsgG pore is Y51S / N55V / F56Q / N91R / K94Q / R97W-del (V105-I107).

[0052] Figure 24: Current characteristics of a DNA strand passing through a CsgG:CsgF complex.

[0053] Complexes were prepared by incubating E. coli purified Y51A / N55V / F56Q / N91R / K94Q / R97W-del (V105-I107) pores containing a C-terminal strep with three different lengths of CsgF. A. CsgF-(1-29), B. CsgF-(1-35), C. CsgF-(1-45). Arrows indicate the signal range. Surprisingly, the complex with CsgF-(1-29) produced the largest signal range.

[0054] Figure 25 : Signal-to-noise ratio of the current characteristic when a DNA chain passes through the CsgG:CsgF complex. By different CsgG pores (1-Y51A / F56Q / N91R / K94Q / R97W-del(V105-I107)2-Y51A / N55I / F56Q / N91R / K94Q / R97W-del(V105-I107)3-Y51A / N55V / F56Q / N91R / K94Q / R97W-del(V105-I107)4-Y51A / F56A / N91R / K94Q / R97W-del(V105-I107)5-Y51A / F56I / N91R / K94Q / R97W-del(V105-I107) Different CsgG:CsgF complexes were prepared by incubating 6-Y51A / F56V / N91R / K94Q / R97W-del(V105-I107), 7-Y51S / N55A / F56Q / N91R / K94Q / R97W-del(V105-I107), 8-Y51S / N55V / F56Q / N91R / K94Q / R97W-del(V105-I107), and 9-Y51T / N55V / F56Q / N91R / K94Q / R97W-del(V105-I107) with the same CsgF peptide, CsgF-(1-35). Different waveform patterns were observed in DNA translocation experiments, and their signal-to-noise ratios were measured. A higher signal-to-noise ratio indicates higher precision.

[0055] Figure 26: Sequencing errors generated using narrow read heads. Schematic diagram of the interaction between DNA bases and the CsgG pore read head. As the DNA strand translocates through the pore, approximately five bases dominate the current signal at any given time. B. Signal mapping diagram. Event detection signals for multiple reads mapped to simulated signals using a custom HMM for a mixed sequence lacking homopolymer runs and a sequence containing three 10T homopolymer runs.

[0056] Figure 27: Read heads mapping the CsgG:CsgF complex.Read head discrimination plot for the CsgG:CsgF complex. Average change in simulated current when the base at each read head position is changed. To calculate the read head discrimination at position i for a model of length k and alphabet length n, we define the discrimination at read head position i as the median of the standard deviation of the current levels for each of the nk-1 groups of size n where position i is changed while the other positions remain constant. B. Static DNA strands mapping the read heads: A set of polyA DNA strands (SS20 to SS38) were created in which a single base was missing from the DNA backbone (iSpc3). In each strand, the position of iSpc3 was shifted from the 3' end to the 5' end. Based on previous experiments with the CsgG pore, position 7 of the DNA was predicted to be within the CsgG constriction. SS26, corresponding to this DNA, is highlighted. Based on the model from (A), 4-5 bases are predicted to separate the CsgG and CsgF read heads. Therefore, approximately positions 12 and 13 are predicted to be within the CsgF constriction. The SS31 and SS32 DNA strands corresponding to those positions are highlighted. C and D. Two read heads are mapped: the biotin modification at the 3' end of each strand is complexed with monovalent streptavidin, and the current blockade generated by each strand is recorded in a MinION device. When iSpc3 is located above or below the constriction within the pore, no deflection is expected. However, when iSpc3 is located within the constriction, higher current levels are expected to flow through the pore—the additional space created by the missing bases allows more ions to pass. Therefore, by plotting the current flowing through each DNA strand, the positions of the two read heads can be mapped. As expected, the maximum deflection in current is seen when position 7 of the DNA strand is occupied by iSpc3 (C). iSpc3 at positions 6 and 8 also produces deflections higher than the average polyA current level. Therefore, positions 6, 7, and 8 of the DNA strand represent the first read head—the CsgG read head. As expected, another deviation from the baseline polyA is observed when positions 12 and 13 are occupied by iCsp3 (D). This indicates the second read head of the pore - the CsgF read head. The results also confirmed that the two read heads are approximately 4-5 bases apart.

[0057] Figure 28: Read head discrimination and base contributions. The left panel shows the read head difference for each mutant pore: the average change in simulated current when the base at each read head position is changed. To calculate the read head difference at position i for a model of length k and alphabet length n, we define the difference at read head position i as n k-1 The median standard deviation of the current levels for each of the groups where position i is varied while the other positions remain constant. The right panel shows the base contribution plot: the median current for all sequences with base b (A, T, G, or C) at position i of the read head.

[0058] Figure 29: Error curve for dual read head aperture. A. Schematic diagram of the CsgG:CsgF complex and the interactions of DNA bases with the two read heads. Red: strong interaction, orange: weak interaction, gray: no interaction. B. Comparison of deletion errors. Reads from Y51A / F56Q / N91R / K94Q / R97W / R192D-del(V105-I107) and Y51A / N55V / F56Q / N91R / K94Q / R97W-del(V105-I107): Base calls into the CsgF-N17S-(1-35) pore from the same region of E. coli DNA. Reads were aligned to the reference genome using Minimap2 (https: / / arxiv.org / abs / 1708.01492) and the resulting alignments were visualized in the Savant Genome Browser (https: / / www.ncbi.nlm.nih.gov / pubmed / 20562449). Most Y51A / F56Q / N91R / K94Q / R97W / R192D-del (V105-I107) reads contain a single base deletion in the T homopolymer (black box), while most CsgG:CsgF reads do not contain base deletions. C. Comparison of consensus accuracy relative to homopolymer length from raw data generated from Y51A / F56Q / N91R / K94Q / R97W / R192D-del(V105-I107) (blue) and Y51A / N55V / F56Q / N91R / K94Q / R97W-del(V105-I107):CsgF-N17S-(1-35) pores (green).

[0059] Figure 30A Homopolymer recognition of the C:CsgG:CsgF complexDNA with the sequence shown in (A) was translocated through the Y51A / F56Q / N91R / K94Q / R97W / R192D-del(V105-I107) pore (B) and the Y51A / N55V / F56Q / N91R / K94Q / R97W-del(V105-I107):CsgF-N17S-(1-35) pore (C), and their signals were analyzed for the first polyT segment shown in red in (A). When the polyT segment passes through the CsgG pore containing a single read head (the model is based on 5 bases located in the read head), it will produce a flat line in the signal. Therefore, it is difficult to determine the exact number of bases in this region that commonly cause deletion errors. When DNA passes through a CsgG:CsgF complex with two read heads (the model is based on the nine bases located within and between the two read heads), the polyT region displays multiple steps rather than a single flat line. Information from these steps can be used to correctly identify the number of bases in the homopolymer region. This additional information significantly reduces deletion errors and improves overall consensus accuracy.

[0060] Figure 31: CsgG pore Characterization of (Y51A / F56Q / N91R / K94Q / R97W / -del(V105-I107). A. Read head discrimination of the CsgG pore. Average change in simulated current when the base at each read head position is changed. To calculate the read head discrimination at position i for a model of length k and alphabet length n, we define the discrimination at read head position i as n k-1 Figure 2. Median standard deviation of current levels for each of the groups, where position i is varied while other positions remain constant. B. Base contribution plot for the CsgG pore. Median current across all k-mers with base b (A, T, G, or C) at position i of the read head. C. Current signature of a DNA strand passing through a CsgG pore. DETAILED DESCRIPTION

[0061] The present invention will be described with respect to specific embodiments and with reference to certain drawings, but the invention is not limited thereto and is limited only by the claims. Any reference signs in the claims should not be construed as limiting the scope. Of course, it should be understood that not all aspects or advantages may be achieved according to any specific embodiment of the present invention. Thus, for example, those skilled in the art will recognize that the present invention can be embodied or performed in a manner that achieves or optimizes one advantage or group of advantages as taught herein without necessarily achieving other aspects or advantages taught or suggested herein.

[0062] The organization and method of operation of the present invention, as well as its features and advantages, can be best understood by reference to the following detailed description when read in conjunction with the accompanying drawings. Various aspects and advantages of the present invention will become apparent and be illustrated by reference to the embodiments described below. Reference throughout this specification to "one embodiment" or "an embodiment" means that the particular characteristics, structures, or features described in conjunction with that embodiment are included in at least one embodiment of the present invention. Therefore, the phrases "in one embodiment" or "in an embodiment" appearing throughout this specification do not necessarily refer to the same embodiment, but may refer to the same embodiment. Similarly, it should be recognized that in the description of exemplary embodiments of the present invention, various features of the present invention are sometimes grouped together in a single embodiment, figure, or description thereof to simplify the disclosure and aid in understanding one or more aspects of the various inventive aspects. However, this method of disclosure should not be interpreted as reflecting an intention that the claimed invention requires more features than those expressly recited in each claim. On the contrary, as reflected in the appended claims, inventive aspects do not lie in all the features of the aforementioned disclosed single embodiments.

[0063] In addition, as used in the specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the content clearly dictates otherwise. Thus, for example, reference to "a polynucleotide" includes two or more polynucleotides, reference to "a polynucleotide binding protein" includes two or more such proteins, reference to "a helicase" includes two or more helicases, reference to "a monomer" refers to two or more monomers, reference to "a pore" includes two or more pores, and so forth.

[0064] In all discussions herein, the standard single-letter code for amino acids is used. These are as follows: alanine (A), arginine (R), asparagine (N), aspartic acid (D), cysteine ​​(C), glutamic acid (E), glutamine (Q), glycine (G), histidine (H), isoleucine (I), leucine (L), lysine (K), methionine (M), phenylalanine (F), proline (P), serine (S), threonine (T), tryptophan (W), tyrosine (Y), and valine (V). Standard substitution symbols are also used, i.e., Q42R means that the Q at position 42 is replaced by R.

[0065] In the paragraphs herein where the / symbol is used to separate different amino acids at a specific position, the / symbol means "or." For example, Q87R / K means Q87R or Q87K.

[0066] In paragraphs herein where specific positions are separated by a / symbol, the / symbol means "and", such that Y51 / N55 means Y51 and N55.

[0067] Unless otherwise stated, all amino acid substitutions, deletions and / or additions disclosed herein refer to mutant CsgG monomers comprising variants of the sequence set forth in SEQ ID NO: 3.

[0068] Reference to a mutant CsgG monomer comprising a variant of the sequence set forth in SEQ ID NO: 3 encompasses mutant CsgG monomers comprising variants of the sequences set forth in other SEQ ID NOs as disclosed below. CsgG monomers comprising variants of the sequence set forth in SEQ ID NO: 3 may be subjected to amino acid substitutions, deletions, and / or additions equivalent to those disclosed herein for mutant CsgG monomers comprising variants of the sequence set forth in SEQ ID NO: 3.

[0069] All publications, patents and patent applications cited herein, whether supra or infra, are hereby incorporated by reference in their entirety.

[0070] definition

[0071] When referring to a singular noun, an indefinite article or a definite article is used, for example, "a", "a kind of" and "the", unless expressly stated, this includes the plural form of the noun. When the term "comprising" is used in this specification and the claims, other elements or steps are not excluded. In addition, the terms first, second, third, etc. in the specification and the claims are used to distinguish similar elements and are not necessarily used for sequential or chronological order. It should be understood that the terms used in this way are interchangeable where appropriate, and the embodiments of the invention described herein can be operated in other orders than the order described or shown herein. The following terms or definitions are only provided to help understand the present invention. Unless expressly defined herein, all terms used herein have the same meaning for those skilled in the art of the invention. Practitioners are particularly directed to Sambrook et al., Molecular Cloning: A Laboratory Manual, 4th ed., Cold Spring Harbor Press, Plainsview, New York (2012); and Ausubel et al., Current Protocols in Molecular Biology (Supplement 114), John Wiley & Sons, New York (2016) for definitions and terminology in the art. The definitions provided herein should not be construed as being narrower in scope than would be understood by one of ordinary skill in the art.

[0072] As used herein, "about" when referring to a measurable value such as an amount, a time interval, etc., is intended to encompass variations of up to ±20% or ±10%, more preferably ±5%, even more preferably ±1, and still more preferably ±0.1% from the specified value, as such variations are suitable for performing the disclosed methods.

[0073] As used herein, "nucleotide sequence," "DNA sequence," or "nucleic acid molecule" refers to a polymeric form of nucleotides (ribonucleotides or deoxyribonucleotides) of any length. The term refers only to the primary structure of the molecule. Thus, the term includes double-stranded and single-stranded DNA and RNA. As used herein, the term "nucleic acid" is a single-stranded or double-stranded covalently linked sequence of nucleotides in which the 3' and 5' ends of each nucleotide are linked by a phosphodiester bond. A polynucleotide can be composed of deoxyribonucleotide bases or ribonucleotide bases. Nucleic acids can be synthesized and prepared in vitro or isolated from natural sources. Nucleic acids can also include modified DNA or RNA, such as methylated DNA or RNA, or RNA that has been post-translationally modified, such as 5'-capping with 7-methylguanosine, 3'-processing (such as cleavage and polyadenylation), and splicing. Nucleic acids can also include synthetic nucleic acids (XNA), such as hexitol nucleic acids (HNA), cyclohexene nucleic acids (CeNA), threose nucleic acids (TNA), glycerol nucleic acids (GNA), locked nucleic acids (LNA), and peptide nucleic acids (PNA). The size of a nucleic acid (also referred to herein as a "polynucleotide") is typically expressed as the number of base pairs (bp) for double-stranded polynucleotides, or as the number of nucleotides (nt) in the case of single-stranded polynucleotides. One thousand bp or nt equals one thousand bases (kb). Polynucleotides less than about 40 nucleotides in length are often referred to as "oligonucleotides" and may contain primers for DNA manipulation, such as by the polymerase chain reaction (PCR).

[0074] As used herein, "gene" includes the promoter region and coding sequence of a gene. It refers to both the genomic sequence (including possible introns) and the cDNA derived from the spliced ​​message operably linked to the promoter sequence.

[0075] A "coding sequence" is a nucleotide sequence that is transcribed into mRNA and / or translated into a polypeptide when placed under the control of appropriate regulatory sequences. The boundaries of the coding sequence are determined by a translation start codon at the 5' end and a translation stop codon at the 3' end. A coding sequence may include, but is not limited to, mRNA, cDNA, a recombinant nucleotide sequence, or genomic DNA, and in some cases may also contain introns.

[0076] In the context of the present disclosure, the term "amino acid" is used in its broadest sense and is intended to include organic compounds containing amine (NH2) and carboxyl (COOH) functional groups and side chains (e.g., R groups) unique to each amino acid. In some embodiments, amino acids refer to naturally occurring Lα-amino acids or residues. Common single-letter and three-letter abbreviations for naturally occurring amino acids are used herein: A=Ala;C=Cys;D=Asp;E=Glu;F=Phe;G=Gly;H=His;I=Ile;K=Lys;L=Leu;M=Met;N=Asn;P=Pro;Q=Gln;R=Arg;S=Ser;T=Thr;V=Val;W=Trp;And Y=Tyr (Lehninger, AL, (1975) Biochemistry, 2nd ed., pp. 71-92, Worth Publishers, New York). The general term "amino acid" also includes D-amino acids, retro-inverso amino acids, and chemically modified amino acids (such as amino acid analogs), naturally occurring amino acids that are not typically incorporated into proteins (such as norleucine), and chemically synthesized compounds with properties known in the art to be characteristic of amino acids (such as β-amino acids). For example, analogs or mimetics of phenylalanine or proline are included in the definition of an amino acid, which allow the same conformational constraints on peptide compounds as natural Phe or Pro. Such analogs and mimetics are referred to herein as "functional equivalents" of the corresponding amino acids. Roberts and Vellaccio, The Peptides: Analysis, Synthesis, Biology, Gross and Meiehofer, eds., Vol. 5, p. 341, Academic Press, Inc., NY 1983, list other examples of amino acids, which are incorporated herein by reference.

[0077] The terms "protein," "polypeptide," and "peptide" are further used interchangeably herein to refer to polymers of amino acid residues, as well as variants and synthetic analogs of amino acid residues. Thus, these terms apply to amino acid polymers in which one or more amino acid residues is a synthetic, non-naturally occurring amino acid, such as a chemical analog of a corresponding naturally occurring amino acid, as well as to naturally occurring amino acid polymers. Polypeptides may also undergo maturation or post-translational modification processes, which may include, but are not limited to, glycosylation, proteolytic cleavage, lipidation, signal peptide cleavage, propeptide cleavage, phosphorylation, and the like. "Recombinant polypeptide" means a polypeptide produced using recombinant techniques, such as by expression of a recombinant or synthetic polynucleotide. When a chimeric polypeptide or biologically active portion thereof is recombinantly produced, it is also preferably substantially free of culture medium, e.g., culture medium comprises less than about 20%, more preferably less than about 10%, and most preferably less than about 5% of the volume of the protein preparation. "Isolated" means material that is substantially or essentially free from components that normally accompany it in its native state. For example, as used herein, "isolated polypeptide" refers to a polypeptide that has been purified away from molecules that flank the polypeptide in its native state, e.g., a protein complex or CsgF peptide that has been removed from molecules adjacent to the polypeptide in the production host. The isolated CsgF peptide (optionally a truncated CsgF peptide) can be produced by amino acid chemical synthesis or by recombinant production. The isolated complex can be produced by in vitro reconstitution after purification of the components of the complex, such as the CsgG pore and CsgF peptide, or by recombinant co-expression.

[0078] The terms "orthologs" and "paralogs" encompass evolutionary concepts used to describe the ancestral relationships of genes. Paralogs are genes within the same species that originated through duplication of an ancestral gene; orthologs are genes from different organisms that originated through speciation and also derive from a common ancestral gene.

[0079] " Homologs " of proteins encompass peptides, oligopeptides, polypeptides, proteins and enzymes that have amino acid substitutions, deletions and / or insertions relative to the unmodified or wild-type protein in question and have biological and functional activities similar to the unmodified protein from which they are derived. As used herein, the term "amino acid identity" refers to the degree to which the sequences are identical on an amino acid-amino acid basis over a comparison window. Thus, "percentage of sequence identity" is calculated by comparing two optimally aligned sequences over a comparison window, determining the number of positions at which the same amino acid residues (e.g., Ala, Pro, Ser, Thr, Gly, Val, Leu, Ile, Phe, Tyr, Trp, Lys, Arg, His, Asp, Glu, Asn, Gln, Cys and Met) appear in the two sequences to obtain the number of matched positions, dividing the number of matched positions by the total number of positions in the comparison window (i.e., window size), and multiplying the result by 100 to obtain the percentage of sequence identity.

[0080] The term "CsgG pore" defines a pore comprising a plurality of CsgG monomers. Each CsgG monomer can be a wild-type monomer from Escherichia coli (SEQ ID NO: 3), a wild-type homolog of E. coli CsgG, such as a monomer having any of the amino acid sequences set forth in SEQ ID NOs: 68 to 88, or a variant thereof (e.g., a variant of any of SEQ ID NOs: 3 and 68 to 88). Variant CsgG monomers may also be referred to as modified CsgG monomers or mutant CsgG monomers. Modifications or mutations in the variants include, but are not limited to, any one or more modifications or combinations of modifications disclosed herein.

[0081] For all aspects and embodiments of the present invention, a CsgG homologue refers to a polypeptide having at least 50%, 60%, 70%, 80%, 90%, 95% or 99% complete sequence identity to wild-type E. coli CsgG as shown in SEQ ID NO: 3. A CsgG homologue also refers to a polypeptide containing the PFAM domain PF03783, which is unique to CsgG-like proteins. http: / / pfam.xfam.org / / family / PF03783 A list of currently known CsgG homologs and CsgG architecture can be found in . Similarly, a CsgG homologous polynucleotide can comprise a polynucleotide having at least 50%, 60%, 70%, 80%, 90%, 95%, or 99% complete sequence identity to wild-type E. coli CsgG set forth in SEQ ID NO: 1. Examples of CsgG homologs set forth in SEQ ID NO: 3 have the sequences set forth in SEQ ID NOs: 68 to 88.

[0082] The term "modified CsgF peptide" or "CsgF peptide" defines a CsgF peptide that has been truncated from its C-terminus (e.g., as an N-terminal fragment) and / or modified to include a cleavage site. The CsgF peptide may be a fragment of wild-type E. coli CsgF (SEQ ID NO: 5 or SEQ ID NO: 6), or a fragment of a wild-type homolog of E. coli CsgF, such as a peptide comprising any one of the amino acid sequences shown in SEQ ID NOs: 17 to 36, or a variant of any one of these (e.g., a variant modified to include a cleavage site).

[0083] For all aspects and embodiments of the present invention, a CsgF homolog refers to a polypeptide having at least 50%, 60%, 70%, 80%, 90%, 95% or 99% complete sequence identity with wild-type E. coli CsgF as shown in SEQ ID NO: 6. In some embodiments, a CsgG homolog also refers to a polypeptide containing the PFAM domain PF10614 that is unique to CsgF-like proteins. http: / / pfam.xfam.org / / family / PF10614 A list of currently known CsgF homologs and CsgF architecture can be found in . Similarly, a CsgF homologous polynucleotide can comprise a polynucleotide having at least 50%, 60%, 70%, 80%, 90%, 95%, or 99% complete sequence identity to wild-type E. coli CsgG as set forth in SEQ ID NO:4. Examples of truncated regions of the CsgF homolog set forth in SEQ ID NO:6 have the sequences set forth in SEQ ID NOs:17 to 36.

[0084] The term "N-terminal portion of a mature CsgF peptide" refers to a peptide having an amino acid sequence corresponding to the first 60, 50, or 40 amino acid residues from the N-terminus of the mature CsgF peptide (without the signal sequence). The mature CsgF peptide can be wild-type or a mutant (e.g., having one or more mutations).

[0085] Sequence identity can also be to fragments or portions of a full-length polynucleotide or polypeptide. Thus, a sequence may have only 50% overall sequence identity to a full-length reference sequence, but the sequence of a specific region, domain, or subunit may have 80%, 90%, or up to 99% sequence identity to the reference sequence. Nucleic acid sequence homology to SEQ ID NO: 1 for a CsgG homolog or SEQ ID NO: 4 for a CsgF homolog is not limited to sequence identity. Many nucleic acid sequences can exhibit biologically significant homology to one another despite having significantly lower sequence identity. Homologous nucleic acid sequences are considered to be sequences that will hybridize to one another under low stringency conditions (MR. Green, J. Sambrook, 2012, Molecular Cloning: A Laboratory Manual, 4th ed., Books 1-3, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY).

[0086] The term "wild type" refers to a gene or gene product isolated from a naturally occurring source. A wild-type gene is the gene most commonly observed in a population and is therefore arbitrarily designed to be the "normal" or "wild-type" form of the gene. In contrast, the term "modified," "mutation," or "variant" refers to a gene or gene product that exhibits sequence modifications (e.g., substitutions, truncations, or insertions), post-translational modifications, and / or functional properties (e.g., property changes) compared to a wild-type gene or gene product. Note that naturally occurring mutants can be isolated; these mutants are identified by the fact that they have altered characteristics compared to a wild-type gene or gene product. Methods for introducing or replacing naturally occurring amino acids are well known in the art. For example, methionine (M) can be replaced with arginine (R) by replacing the codon for methionine (ATG) with the codon for arginine (CGT) at the relevant position in the polynucleotide encoding the mutant monomer. Methods for introducing or replacing non-naturally occurring amino acids are also well known in the art. For example, non-naturally occurring amino acids can be introduced by including a synthetic aminoacyl-tRNA in the IVTT system for expressing the mutant monomer. Alternatively, non-natural amino acids can be introduced by expressing mutant monomers in E. coli, which are auxotrophic for specific amino acids in the presence of synthetic (i.e. non-natural) analogs of those specific amino acids. If mutant monomers are produced using partial peptide synthesis, they can also be produced by naked connection. Conservative substitutions replace amino acids with other amino acids having similar chemical structures, similar chemical properties or similar side chain volumes. The amino acid introduced can have polarity, hydrophilicity, hydrophobicity, alkalinity, acidity, neutrality or charge similar to the amino acid they replace. Alternatively, conservative substitutions can introduce another aromatic or aliphatic amino acid to replace pre-existing aromatic or aliphatic amino acids. Conservative amino acid changes are well known in the art and can be selected according to the properties of the 20 main amino acids defined in Table 1 below. When amino acid has similar polarity, this can also be determined with reference to the hydrophilicity scale of the amino acid side chain in Table 2.

[0087] Table 1 – Chemical properties of amino acids

[0088]

[0089] Table 2 - Hydrophilicity Scale

[0090]

[0091] Mutation or modified protein, monomer or peptide can also be chemically modified at any site in any way.Mutation or modified monomer or peptide is preferably chemically modified by attachment of molecule to one or more cysteines (cysteine ​​connection), attachment of molecule to one or more lysines, attachment of molecule to one or more non-natural amino acids, enzyme modification of epitope or modification of end. Suitable methods for carrying out such modifications are well known in the art. The mutant of modified protein, monomer or peptide can be chemically modified by attachment of any molecule. For example, the mutant of modified protein, monomer or peptide can be chemically modified by attachment of dye or fluorophore. In some embodiments, the mutant or modified monomer or peptide is chemically modified with a molecular adapter that promotes the interaction between the hole comprising monomer or peptide and the target nucleotide or target polynucleotide sequence. The molecular adapter is preferably a cyclic molecule, cyclodextrin, a substance capable of hybridization, a DNA binder or intercalator, a peptide or peptide analog, a synthetic polymer, an aromatic planar molecule, a positively charged small molecule or a small molecule capable of hydrogen bonding.

[0092] The presence of the adaptor improves the host-guest chemistry between the pore and the nucleotide or polynucleotide sequence, thereby improving the sequencing capability of the pore formed from the mutant monomer. The principles of host-guest chemistry are well known in the art. The adaptor has an effect on the physical or chemical properties of the pore, which improves the interaction between the pore and the nucleotide or polynucleotide sequence. The adaptor can alter the charge of the barrel or channel of the pore, or specifically interact or bind to the nucleotide or polynucleotide sequence, thereby promoting its interaction with the pore. Thus, the modified CsgF peptides provided in the present disclosure can be coupled to enzymes or proteins to provide better accessibility of the protein or enzyme to the pore, which can facilitate certain applications of pore complexes containing the modified CsgF peptide.

[0093] In this context, a protein may also be a fusion protein, in particular referring to a genetic fusion produced, for example, by recombinant DNA technology. As used herein, a protein may also be conjugated or "conjugated to", in particular referring to a chemical and / or enzymatic conjugation resulting in a stable covalent linkage.

[0094] When several polypeptides or protein monomers bind or interact with each other, proteins can form a protein complex. "Binding" means any direct or indirect interaction. Direct interaction implies contact between binding partners, such as contact through covalent bonding or coupling. Indirect interaction means any interaction by which the interacting partners interact in a complex of two or more compounds. The interaction can be completely indirect, with the aid of one or more bridging molecules, or partially indirect, where direct contact still exists between the partners, and this direct contact is stabilized by the additional interaction of one or more compounds. As referred to in this disclosure, a "complex" is defined as a group of two or more associated proteins, which may have different functions. The association between the different polypeptides in a protein complex may be through non-covalent interactions, such as hydrophobic or ionic forces, or may be covalent binding or coupling, such as disulfide bridges or peptide bonds. Covalent "binding" or "coupling" are used interchangeably herein and may also refer to "cysteine ​​coupling" or "reactive or photoreactive amino acid coupling", which refers to biological conjugation between cysteines or (photo) reactive amino acids, respectively, which is a chemical covalent linkage that forms a stable complex. Examples of photoreactive amino acids include azidohomoalanine, homopropargylglycine, homoallylglycine, p-acetyl-Phe, p-azido-Phe, p-propargyloxy-Phe, and p-benzoyl-Phe (Wang et al., 2012, in Protein Engineering, DOI: 10.5772 / 28719; Chin et al., 2002, Proc. Nat. Acad. Sci. USA 99(17); 11020-24).

[0095] A "biological pore" is a transmembrane protein structure that defines a channel or pore that allows molecules and ions to translocate from one side of the membrane to the other. The translocation of ionic species through the pore can be driven by a potential difference applied to either side of the pore. A "nanopore" is a biological pore in which the minimum diameter of the channel through which a molecule or ion passes is on the order of nanometers (10 -9 m). In some embodiments, the biological pore can be a transmembrane protein pore. The transmembrane protein structure of the biological pore can be monomeric or oligomeric in nature. Typically, the pore comprises a plurality of polypeptide subunits arranged around a central axis, thereby forming a protein-lined channel extending substantially perpendicular to the membrane in which the nanopore resides. There is no limit to the number of polypeptide subunits. Typically, the number of subunits is 5 to 30, and suitably the number of subunits is 6 to 10. Alternatively, the number of subunits is not defined as in the case of perfringolysin or related large membrane pores. The protein subunit portion that forms the protein-lined channel in the nanopore typically comprises a secondary structural motif that may include one or more transmembrane β-barrels and / or α-helical portions.

[0096] As used interchangeably herein, the terms "pore," "pore complex," or "complex pore" refer to an oligomeric pore in which, for example, at least a CsgG monomer (including, for example, one or more CsgG monomers, such as two or more CsgG monomers, three or more CsgG monomers) or a CsgG pore (composed of CsgG monomers) and a CsgF peptide (e.g., a modified or truncated CsgF peptide) associate into a complex and together form a pore or nanopore. The pore complex of the present disclosure has the characteristics of a biological pore, i.e., it has a typical transmembrane protein structure. When the pore complex is provided in an environment having a membrane component, a membrane, a cell, or an insulating layer, the pore complex will insert into the membrane or insulating layer, forming a "transmembrane pore complex."

[0097] The pore complexes or transmembrane pore complexes disclosed herein are suitable for analyte characterization. In some embodiments, the pore complexes or transmembrane complexes described herein can be used, for example, to sequence polynucleotide sequences because they can distinguish different nucleotides with high sensitivity. The pore complexes disclosed herein can be substantially isolated, purified, or substantially purified isolated pore complexes. A pore complex disclosed herein is "isolated" or purified if it is completely free of any other components, such as lipids or other pores, or other proteins with which it is normally associated in its native state, such as CsgE, CsgA, or CsgB, or if it is substantially enriched from a membrane compartment. A pore complex is substantially isolated if it is mixed with a carrier or diluent that does not interfere with its intended use. For example, a pore complex is substantially isolated or substantially purified if it is present in a form that contains less than 10%, less than 5%, less than 2%, or less than 1% of any other components, such as triblock copolymers, lipids, or other pores. Alternatively, a pore complex disclosed herein can be a transmembrane pore complex when present in a membrane. The present disclosure provides isolated pore complexes comprising a homo-oligomeric pore derived from CsgG comprising identical mutant monomers, which may also contain mutant forms of CsgG monomers as homologs. Alternatively, isolated pore complexes comprising a hetero-oligomeric CsgG pore are provided, which may be a CsgG pore composed of mutant and wild-type CsgG monomers or different forms of CsgG variants, mutants, or homologs. The isolated pore complexes typically comprise at least 7, at least 8, at least 9, or at least 10 CsgG monomers and one or more (modified) CsgF peptides, such as 2, 3, 4, 5, 6, 7, 8, 9, or 10 CsgF peptides. The pore complexes may comprise any ratio of CsG monomers:CsgF peptides. In one embodiment, the ratio of CsG monomers:CsgF peptides is 1:1.

[0098] As used interchangeably herein, "constriction", "pore opening", "constriction region", "channel constriction" or "constriction site" refers to an aperture defined by the lumen surface of a pore or pore complex that functions to allow ions and target molecules (such as, but not limited to, polynucleotides or single nucleotides) to pass through the pore complex channel, but not other non-target molecules. In some embodiments, the constriction is the narrowest aperture in the pore or pore complex. In this embodiment, the constriction can be used to restrict the passage of molecules through the pore. The size of the constriction is often a key factor in determining the suitability of a nanopore for nucleic acid sequencing applications. If the constriction is too small, the molecules to be sequenced will not be able to pass through. However, in order to have the greatest impact on the flow of ions through the channel, the constriction should not be too large. For example, the constriction should not be wider than the solvent accessible lateral diameter of the target analyte. Ideally, the diameter of any constriction should be as close as possible to the lateral diameter of the analyte passing through. For sequencing of nucleic acids and nucleic acid bases, suitable constriction diameters are in the nanometer range (10 -9 Suitably, the diameter should be in the region of 0.5 to 2.0 nm, and typically, the diameter is in the region of 0.7 to 1.2 nm. The diameter of the constriction in wild-type E. coli CsgG is approximately The diameter of the CsgF constriction formed in the pore complex comprising the CsgG-like pore and the modified CsgF peptide or its homologue or mutant is in the range of 0.5 to 2 nm or in the range of 0.7 to 1.2 nm, and is therefore suitable for nucleic acid sequencing.

[0099] When two or more constrictions are present and spaced apart, each constriction can simultaneously interact with or "read" individual nucleotides within a nucleic acid chain. In this case, the reduction in ion current through the channel will be the result of the combined restriction of flow in all nucleotide-containing constrictions. Thus, in some cases, a dual constriction may result in a composite current signal. In some cases, when two such read heads are present, the current reading from a single constriction, or "read head," may not be determined individually. The constriction of wild-type Escherichia coli CsgG (SEQ ID NO: 3) consists of two annular circles formed by the juxtaposition of a tyrosine residue at position 51 (Tyr 51) and phenylalanine and asparagine residues at positions 56 and 55, respectively (Phe 56 and Asn 55), in two adjacent protein monomers (Figure 1). In most cases, the wild-type pore structure of CsgG is engineered using recombinant genetic techniques to enlarge, alter, or remove one of the two annular circles that comprise the CsgG constriction (referred to herein as the "CsgG channel constriction"), leaving behind a single, well-defined read head. The constriction motif in the CsgG oligomeric pore is located at amino acid residues 38 to 63 in the wild-type monomeric E. coli CsgG polypeptide depicted in SEQ ID NO:3. In considering this region, mutations at any of amino acid residue positions 50 to 53, 54 to 56, and 58 to 59 within the channel of the wild-type CsgG structure, as well as the positioning of the side chains of Tyr51, Asn55, and Phe56, have proven to be beneficial for modifying or altering the characteristics of the readhead. The present disclosure relates to pore complexes comprising a CsgG pore and a modified CsgF peptide, or a homolog or mutant thereof, wherein surprisingly, an additional constriction (referred to herein as the "CsgF channel constriction") is added to the CsgG-containing pore complex, forming a suitable additional second readhead within the pore by complexing with the modified CsgF peptide. The additional CsgF channel constriction or readhead is positioned adjacent to the contractile ring of the CsgG pore or mutant CsgG pore. The additional CsgF channel constriction or readhead is positioned about 10 nm or less, such as 5 nm or less, such as 1, 2, 3, 4, 5, 6, 7, 8, or 9 nm from the constriction ring of the CsgG pore or mutant GcsG pore. The pore complex or transmembrane pore complex of the present disclosure includes a pore complex with two readheads, meaning that the channel constriction is positioned in a manner that provides a suitable separate readhead without interfering with the accuracy of the other constriction channel readheads.The pore complex may thus comprise a CsgG mutant pore (see incorporated references WO 2016 / 034591, WO 2017 / 149316, WO 2017 / 149317, WO 2017 / 149318 and International Patent Application No. PCT / GB2018 / 051191, each of which lists mutations of the wild-type CsgG pore that improve the pore properties) and a wild-type CsgG pore or a homologue thereof, together with a modified CsgF peptide or a homologue or mutant thereof, wherein the CsgF peptide has a further constriction channel that forms the readhead.

[0100] hole

[0101] The present invention relates to a CsgG pore complexed with an extracellularly localized CsgF peptide that surprisingly introduces an additional channel constriction or readhead into the pore complex. Furthermore, the present disclosure provides information about the position of the constriction within the pore complex formed by the CsgF peptide, which is inserted into the lumen of the CsgG pore and has a constriction site in the N-terminal portion of the CsgF protein. Furthermore, the modified or truncated CsgF peptides of the present disclosure are demonstrated to be sufficient to form a pore complex, providing means and methods for biosensing applications. The present disclosure encompasses wild-type and mutant CsgG pores (as disclosed, for example, in WO2016 / 034591, WO2017 / 149316, WO2017 / 149317, WO2017 / 149318, and their international patent application number PCT / GB2018 / 051191), or homologs or mutants thereof, in combination with modified or truncated CsgF peptides and their mutants or homologs, all of which together enhance the ability of the CsgG-like pore complex to interact with analytes (such as polynucleotides). The additional constriction introduced into the CsgG-like nanopore channel by forming a complex with the (modified or truncated) CsgF peptide expands the contact surface with the passing analyte and can serve as a second readhead for analyte detection and characterization. Pores containing mutant CsgG monomers combined with novel mutant or modified forms of CsgF can improve the characterization of analytes, such as polynucleotides, providing a more discriminatory and direct relationship between the currents observed as the polynucleotide moves through the pore. Specifically, by spacing two stacked read heads apart by a defined distance, the CsgG:CsgF pore complex can facilitate the characterization of polynucleotides containing at least one homopolymer stretch, e.g., several consecutive copies of the same nucleotide, that exceeds the interaction length of a single CsgG read head. Additionally, by spacing the two stacked constrictions apart by a defined distance, small molecule analytes (including organic or inorganic drugs and contaminants) passing through the CsgG:CsgF composite pore will pass sequentially through two independent read heads. The chemical properties of either read head can be independently modified, each providing unique interaction properties with the analyte, thereby providing additional discriminatory capabilities during analyte detection.

[0102] In a first aspect, the present invention relates to an isolated pore complex comprising a CsgG pore, or a homologue or mutant thereof, or a CsgG-like pore, and a modified CsgF peptide, or a homologue or mutant thereof. Indeed, the present disclosure relates to a modified CsgG biological pore comprising a modified CsgF peptide (which may be truncated), a mutant, and / or a variant thereof. In one embodiment, the interaction region between the modified CsgF peptide, or a homologue or mutant thereof, is located within the lumen of the CsgG pore, or a homologue or mutant thereof. In another embodiment, the pore complex has two or more constriction sites or readheads, provided by at least one constriction of the CsgG pore and at least one constriction introduced by the CsgF peptide, thereby forming a complex with the CsgG pore. It has been demonstrated that the N-terminal CsgF position, including within the range of amino acid residues 39-64 of SEQ ID NO:5, or more specifically within the range of amino acid residues 49-64 of SEQ ID NO:5, allows for detectable amounts of stable CsgG:CsgF complexes. In one embodiment, the CsgF constriction generated by a modified CsgF peptide (e.g., a CsgF peptide described herein) is adjacent to or head-to-head with the first constriction in the CsgG pore of the pore complex. For CsgG or CsgG-like protein pores, the constriction site has been determined to be formed by a loop region of the β strand (see Figure 1).

[0103] In one embodiment, the modified CsgF peptide is a peptide wherein the modification specifically refers to a truncated CsgF protein or fragment, including an N-terminal CsgF peptide fragment defined by restriction to include a constriction region and bind to a CsgG monomer or a homolog or mutant thereof. The modified CsgF peptide may further comprise a mutation or homologous sequence that may promote certain properties of the pore complex. In a specific embodiment, the modified CsgF peptide comprises a truncation of the CsgF protein compared to the wild-type preprotein (SEQ ID NO: 5) or mature protein (SEQ ID NO: 6) sequence or a homolog thereof. These modified peptides are intended for use as pore complex components, introducing additional constriction sites or read heads into the CsgG-like pore formed by CsgG and the modified or truncated CsgF peptide. Examples of truncated modified peptides are described below.

[0104] Examples of modified homologs of CsgF peptides are identified, for example, in Example 3, and reveal that CsgF-like proteins or CsgF peptides from different bacterial strains containing homologous or similar constriction regions are useful in the use of similar pore complexes. Structural properties and CsgG binding elements are conserved among CsgF peptides derived from various CsgF homologs, allowing the CsgF peptides to be used in combination with different wild-type or mutant CsgG pores. This includes complexes of the CsgG pore with non-homologous CsgFs, meaning that the CsgG pore and the parent CsgF homolog sequence from which the CsgF is derived do not need to originate from the same operon, bacterial species, or strain.

[0105] In alternative embodiments, the CsgG pore in the pore complex is not a wild-type pore, but rather further comprises mutations or modifications to enhance the properties of the pore. The isolated pore complex of the present disclosure, formed from a CsgG pore or homolog thereof and a modified CsgF peptide or homolog thereof, can be formed from a wild-type form of the CsgG pore, or can be further modified in the CsgG pore, for example, by targeted mutagenesis of specific amino acid residues, to further enhance the properties of the CsgG pore required for use in a pore complex. For example, in embodiments of the present invention, mutations are contemplated to alter the number, size, shape, placement, or orientation of constrictions within the channel. Pore complexes comprising modified mutant CsgG pores can be prepared by known genetic engineering techniques that introduce insertions, substitutions, and / or deletions of specific target amino acid residues within the polypeptide sequence. In the case of an oligomeric CsgG pore, mutations can be generated in each monomeric polypeptide subunit, or in any one monomer, or in all monomers. Suitably, in one embodiment of the present invention, the mutations are generated in all monomeric polypeptides within the oligomeric protein structure. A mutant CsgG monomer is a monomer whose sequence differs from that of a wild-type CsgG monomer and which retains the ability to form a pore. Methods for determining the ability of mutant monomers to form a pore are well known in the art. The present disclosure encompasses wild-type and mutant CsgG pores (e.g., as disclosed in WO2016 / 034591, WO2017 / 149316, WO2017 / 149317, WO2017 / 149318, and their International Patent Application No. PCT / GB2018 / 051191), or homologs thereof, in combination with modified or truncated CsgF peptides, mutants, or homologs thereof, all of which together enhance the ability of the CsgG-like pore complex to interact with an analyte, such as a polynucleotide. A mutant CsgG pore may comprise one or more mutant monomers. A CsgG pore may be a homopolymer comprising the same monomer, or a SEQ heteropolymer comprising two or more different monomers. A monomer may have one or more mutations in any combination described below.

[0106] Nanopore complexes comprising modified CsgF peptides differ from the wild-type CsgF protein depicted in SEQ ID NO:6 in that, in certain embodiments, the modified CsgF peptide comprises only an N-terminal fragment or truncation of the wild-type CsgF protein. However, the modified CsgF peptide may additionally or alternatively be a mutant CsgF peptide, in the sense that mutations, such as amino acid substitutions, are generated to allow for a better second constriction site in the pore formed by the complex comprising the CsgF pore and the modified CsgF peptide. When the complex is used for nucleotide sequencing, the mutant peptide may also exhibit improved polynucleotide readout properties, that is, improved polynucleotide capture and nucleotide discrimination, in addition to the improved feature of the complex comprising two read heads. Specifically, pores constructed with the mutant peptides may capture nucleotides and polynucleotides more readily than wild-type pores. Furthermore, pores constructed with the mutant peptides may exhibit an increased current range, which makes it easier to distinguish between different nucleotides, and reduced state changes, which increases the signal-to-noise ratio. Furthermore, the number of nucleotides contributing to the current as a polynucleotide moves through the pore constructed with the mutant peptides may be reduced. This makes it easier to identify a direct relationship between the current observed when a polynucleotide moves through the pore and the polynucleotide sequence. In addition, pores constructed from mutant peptides can show increased throughput, for example, being more likely to interact with analytes such as polynucleotides. This makes it easier to use the pores to characterize analytes. Pores constructed from mutant peptides can be more easily inserted into membranes or can provide an easier way to keep other proteins in the vicinity of the pore complex.

[0107] In an alternative embodiment, the diameter of the CsgF constriction site provided in the pore complex of the invention is in the range of 0.5 nm to 2.0 nm, thereby providing a pore complex suitable for nucleic acid sequencing as described above.

[0108] The pore can be stabilized by covalent attachment of a CsgF peptide to the CsgG pore. The covalent linkage can be, for example, a disulfide bond or click chemistry. The CsgF peptide and the CsgG pore can be covalently linked, for example, via residues at positions corresponding to one or more of the following pairs of positions of SEQ ID NO: 6 and SEQ ID NO: 3, respectively: 1 and 153, 4 and 133, 5 and 136, 8 and 187, 8 and 203, 9 and 203, 11 and 142, 11 and 201, 12 and 149, 12 and 203, 26 and 191, and 29 and 144.

[0109] In the pore, the interaction between the CsgF peptide and the CsgG pore may be stabilized, for example, via hydrophobic interactions or electrostatic interactions at positions corresponding to one or more of the following pairs of positions of SEQ ID NO: 6 and SEQ ID NO: 3, respectively: 1 and 153, 4 and 133, 5 and 136, 8 and 187, 8 and 203, 9 and 203, 11 and 142, 11 and 201, 12 and 149, 12 and 203, 26 and 191, and 29 and 144.

[0110] Residues in CsgF and / or CsgG at one or more of the positions listed above may be modified to enhance the interaction between CsgG and CsgF in the pore.

[0111] In one embodiment, the pores of the present invention can be isolated, substantially isolated, purified, or substantially purified. A pore of the present invention is isolated or purified if it is completely free of any other components, such as lipids or other pores. A pore is substantially isolated if it is mixed with a carrier or diluent that does not interfere with its intended use. For example, a pore is substantially isolated or substantially purified if it is present in a form that contains less than 10%, less than 5%, less than 2%, or less than 1% of other components, such as triblock copolymers, lipids, or other pores. Alternatively, the pores of the present invention can be present in a membrane. Suitable membranes are discussed below.

[0112] The pores of the invention may exist as individual or single pores. Alternatively, the pores of the invention may exist in a homogenous or heterogenous population of two or more pores.

[0113] CsgF peptide

[0114] A second aspect of the present invention relates to novel modified CsgF monomers (peptides), or truncated CsgF proteins, or modified or truncated peptides of CsgF homologs or mutants. These novel modified CsgF peptides can be used in pore complexes to integrate a second or additional read head. The modification or truncation preferably produces a fragment, more preferably an N-terminal fragment, of a wild-type CsgF or mutant or homologous CsgF protein.

[0115] Mature CsgF (as set forth in SEQ ID NO:6) can be divided into three major regions: a "CsgF contractile peptide" (FCP), a "neck" region, and a "head" region (as shown in Figures 4 and 5). The "head" region of the CsgF peptide is distinct from the read head of the pore described herein. The "head" region of the CsgF peptide may also be referred to as the "C-terminal head domain."

[0116] The FCP forms a contact region with the CsgG β-barrel, where it creates an additional constriction. The neck region protrudes from the β-barrel. In the CsgG:CsgF oligomer, it forms a thin-walled hollow tube connecting the FCP to the globular head region.

[0117] Based on multiple sequence alignment ( Figure 8 ), co-purification experiments (Figure 9) and In the cryoEM reconstruction of the CsgG:CsgF complex at 300 nm resolution (Figure 11), the CsgF contractile peptide, neck, and head regions can be defined as three consecutive residue stretches in mature CsgF.

[0118] The FCP spans approximately residues 1 to 35 of mature CsgF (SEQ ID NO: 6). When comparing different CsgF orthologs, the FCP forms the most conserved region of the protein ( Figure 8 、 Figure 10 CryoEM 3D reconstructions revealed that FCP forms a well-defined structure that binds to the interior of the CsgG β-barrel via non-covalent contacts with the CsgG transmembrane hairpins TM1 (residues 134 to 154 of Seq ID NO: 3) and TM2 (residues 184 to 208 of Seq ID NO: 3). Figure 1E ; Figure 11 (TM1 and TM2 are defined in Goyal P et al., 2014). During reconstitution, nine copies of FCP bind to the CsgG oligomer (comprising 9 monomers) and together generate an additional constriction approximately 2 nm above the CsgG constriction, which is formed by a continuous loop spanning residues 46 to 61 of mature CsgG (Seq ID NO: 3; Figure 1E ; Figure 11).

[0119] The cryoEM 3D reconstruction also shows that the CsgF N-terminal residues bind near the base or top of the CsgG β-barrel (depending on orientation) and exit the β-barrel at residue 32. This agrees well with MD simulations showing the average contact times of residue pairs in the CsgG:CsgF binding interface (Table 4). The cryoEM structure and MD simulations reveal that residues 33–34 are located on the exterior of the CsgG β-barrel, where a highly (although not strictly) conserved Pro residue (Pro 35 in Seq ID NO: 6) transitions to the CsgF neck region. The CsgF neck is not resolved in atomic detail in the CsgG:CsgF 3D reconstruction, suggesting its conformational flexibility. Based on multiple sequence alignments and secondary structure predictions, the CsgF neck is predicted to span approximately from residue 36 to residue 50 (SEQ ID NO: 6). The CsgF head region forms the C-terminal portion of CsgF and is predicted to span approximately from residue 51 to the CsgF C-terminus. In the CsgG:CsgF complex, this region oligomerizes to produce a globular structure that appears to cap the CsgG:CsgF channel (Figure 4, Figure 5Multiple sequence alignment of CsgF orthologs revealed that the CsgF neck is the least conserved region, indicating that its length may vary among orthologs ( Figure 8 ).

[0120] The CsgF peptides forming part of the present invention are those that lack the C-terminal head; lack the C-terminal head and a portion of the neck domain of CsgF (e.g., a truncated CsgF peptide may comprise only a portion of the neck domain of CsgF); or lack the C-terminal head and neck domain of CsgF. The CsgF peptide may lack a portion of the neck domain of CsgF, for example, the CsgF peptide may comprise a portion of the neck domain, for example, starting from amino acid residue 36 at the N-terminus of the neck domain (see SEQ ID: NO: 6) (e.g., residues 36-40, 36-41, 36-42, 36-43, 36-45, 36-46 to residues 36-50 or 36-60 of SEQ ID: NO: 6). The CsgF peptide preferably comprises a CsgG binding region and a region that forms a constriction in the pore. The CsgG binding region typically comprises residues 1 to 8 and / or 29 to 32 of the CsgF protein (SEQ ID NO: 6 or a homolog from another species) and may include one or more modifications. The region that forms the constriction in the pore typically comprises residues 9 to 28 of the CsgF protein (SEQ ID NO: 6 or a homolog from another species) and may include one or more modifications. Residues 9 to 17 comprise the conserved motif N9PXFGGXXX 17 And form the turning region. Residues 9 to 28 form an α-helix. 17 (N17 in SEQ ID NO:6) forms the apex of the constriction, corresponding to the narrowest part of the CsgF constriction in the pore. The CsgF constriction also makes stabilizing contacts with the CsgG β-barrel primarily at residues 9, 11, 12, 18, 21, and 22 of SEQ ID NO:6.

[0121] CsgF peptides are typically 28 to 50 amino acids in length, such as 29 to 49, 30 to 45, or 32 to 40 amino acids. Preferably, the CsgF peptide comprises 29 to 35 amino acids or 29 to 45 amino acids. The CsgF peptide comprises all or part of the FCP, corresponding to residues 1 to 35 of SEQ ID NO: 6. In cases where the CsgF peptide is shorter than the FCP, it is preferably truncated at the C-terminus.

[0122] The CsgF fragment of SEQ ID NO: 6, or a homologue or mutant thereof, may be 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, or 55 amino acids in length.

[0123] The CsgF peptide may comprise the amino acid sequence of SEQ ID NO: 6, from residue 1 to any one of residues 25 to 60, such as 27 to 50, for example 28 to 45, of SEQ ID NO: 6, or the corresponding residues from a homologue of SEQ ID NO: 6, or a variant thereof. More specifically, the CsgF peptide may comprise SEQ ID NO: 39 (residues 1 to 29 of SEQ ID NO: 6), or a homologue or variant thereof.

[0124] Examples of such CsgF peptides include, consist essentially of, or consist of the sequence of SEQ ID NO: 15 (residues 1 to 34 of SEQ ID NO: 6), SEQ ID NO: 54 (residues 1 to 30 of SEQ ID NO: 6), SEQ ID NO: 40 (residues 1 to 45 of SEQ ID NO: 6), or SEQ ID NO: 55 (residues 1 to 35 of SEQ ID NO: 6), and homologs or variants of any of these. Other examples of CsgF peptides include, consist essentially of, or consist of the sequence of SEQ ID NO: 7, SEQ ID NO: 8, SEQ ID NO: 9, SEQ ID NO: 10, SEQ ID NO: 11, SEQ ID NO: 12, SEQ ID NO: 13, SEQ ID NO: 14, or SEQ ID NO: 16.

[0125] In the CsgF peptide, for example, one or more residues in SEQ ID NO: 15, SEQ ID NO: 39, SEQ ID NO: 40, SEQ ID NO: 54 or SEQ ID NO: 55 may be modified.

[0126] For example, the CsgF peptide may comprise modifications at positions corresponding to one or more of the following positions in SEQ ID NO: 6: G1, T4, F5, R8, N9, N11, F12, A26, and Q29.

[0127] The CsgF peptide can be modified to introduce cysteine, a hydrophobic amino acid, a charged amino acid, a non-natural reactive amino acid, or a photoreactive amino acid, for example, at positions corresponding to one or more of the following positions in SEQ ID NO: 6: G1, T4, F5, R8, N9, N11, F12, A26, and Q29.

[0128] For example, the CsgF peptide may comprise modifications at positions corresponding to one or more of the following positions in SEQ ID NO: 6: N15, N17, A20, N24, and A28. The CsgF peptide may comprise a modification at a position corresponding to D34 to stabilize the CsgG-CsgF complex. In certain embodiments, the CsgF peptide comprises one or more of the following substitutions: N15S / A / T / Q / G / L / V / I / F / Y / W / R / K / D / C, N17S / A / T / Q / G / L / V / I / F / Y / W / R / K / D / C, A20S / T / Q / N / G / L / V / I / F / Y / W / R / K / D / C, N24S / T / Q / A / G / L / V / I / F / Y / W / R / K / D / C, A28S / T / Q / N / G / L / V / I / F / Y / W / R / K / D / C, and D34F / Y / W / R / K / N / Q / C. The CsgF peptide may, for example, comprise one or more of the following substitutions: G1C, T4C, N17S, and D34Y or D34N.

[0129] CsgF peptides can be produced by enzymatically cleaving longer proteins, such as full-length CsgF. Cleavage at specific sites can be directed by modifying longer proteins (such as full-length CsgF) to include an enzymatic cleavage site at an appropriate position. Examples of CsgF amino acid sequences that have been modified to include such enzymatic cleavage sites are shown in SEQ ID NOs: 56 to 67. After cleavage, all or part of the added enzymatic cleavage site can be present in the CsgF peptide that associates with CsgG to form a pore. Thus, the CsgF peptide can also include all or part of the enzymatic cleavage site at its C-terminus.

[0130] Some examples of suitable CsgF peptides are shown in Table 3 below:

[0131] Table 3: CsgF peptides

[0132]

[0133]

[0134]

[0135] In certain embodiments, the CsgF fragment comprises the amino acid sequence of SEQ ID NO:39, or a mutant or homolog thereof. Specifically, SEQ ID NO:39 comprises the first 29 amino acids of the mature CsgF peptide (SEQ ID NO:6). In another embodiment, the modified CsgF peptide of the present invention is a truncated peptide comprising SEQ ID NO:40. Specifically, SEQ ID NO:40 comprises the first 45 amino acids of the mature CsgF peptide (SEQ ID NO:6). Specifically, the CsgF constriction site and the binding site for CsgG are located within the N-terminal CsgF peptide region, further characterized by amino acids 39 to 64 of SEQ ID NO:5 (present in both SEQ ID NO:39 and SEQ ID NO:40), or in particular, amino acids 49 to 64 of SEQ ID NO:5 (present in SEQ ID NO:40 but not in SEQ ID NO:39; the latter fragment encoded by SEQ ID NO:39 exhibits a weak interaction with CsgG (see Examples)), conferring increased stability to the complex. Thus, the present disclosure provides for modification of CsgF proteins by truncation of the protein to peptides or peptides comprising the N-terminal fragment or constriction site region to allow for in vivo complex formation with the CsgG pore or homologs or mutants thereof. Further limitations are provided in one embodiment regarding the modified CsgF peptide comprising SEQ ID NO: 37 or SEQ ID NO: 38. Finally, CsgF homologous peptides, particularly CsgF homologous peptides aligned within the constriction region (FCP peptides), are identified, and modified CsgF peptide homologs that can form part of the isolated complex are also provided (e.g., see Figure 8 and Figure 10 ).

[0136] Another embodiment relates to a modified or truncated CsgF peptide comprising SEQ ID NO: 15, wherein SEQ ID NO: 15 contains a region of the CsgF protein, including a region from the CsgG binding and / or constriction site, sufficient to reconstitute in vitro a few residues of the pore of a complex comprising CsgG or a homolog thereof and the modified CsgF peptide, to produce an isolated pore complex comprising the constriction of the CsgF channel. Another embodiment describes the modified CsgF peptide comprising SEQ ID NO: 16, which contains an N-terminal fragment of the CsgF protein and two additional amino acids (KD) that increase the solubility and stability of the (synthetic) peptide and also allow for in vitro reconstitution of the complex pore. Further embodiments are provided, wherein the modified CsgF peptide comprises SEQ ID NO: 15, SEQ ID NO: 16, or a homologue or mutant thereof, wherein the modified CsgF peptide is further mutated but still retains at least 35% amino acid identity, e.g., 40%, 50%, 60%, 70%, 80%, 85%, 90% amino acid identity, to SEQ ID NO: 15 or SEQ ID NO: 16, respectively, within the region of the modified CsgF peptide corresponding to SEQ ID NO: 15 or 16. Other embodiments are provided, wherein the modified CsgF peptide comprises SEQ ID NO: 15, SEQ ID NO: 16, or a homolog or mutant thereof, wherein the modified CsgF peptide is further mutated but still retains at least 40%, 45%, 50%, 60%, 70%, 80%, 85%, or 90% amino acid identity to SEQ ID NO: 15 or SEQ ID NO: 16, respectively, within the region of the modified CsgF peptide corresponding to SEQ ID NO: 15 or 16. As discussed above, those mutated regions are intended to alter and / or improve the characteristics of the CsgF constriction site, thereby, for example, enabling more accurate target analysis. Another embodiment discloses a modified CsgF peptide, wherein one or more positions in the region comprising SEQ ID NO: 39, SEQ ID NO: 40, SEQ ID NO: 54 or SEQ ID NO: 55 are modified, and wherein the mutations maintain a minimum of 35% amino acid identity, or 40%, 50%, 60%, 70%, 80%, 85%, 90%, or 95% amino acid identity with SEQ ID NO: 39, SEQ ID NO: 40, SEQ ID NO: 54 or SEQ ID NO: 55 in the peptide fragment corresponding to the region comprising SEQ ID NO: 39, SEQ ID NO: 40, SEQ ID NO: 54 or SEQ ID NO: 55.

[0137] Thus, further embodiments of the present invention relate to an isolated pore complex comprising a CsgG pore, or a homologue or mutant thereof, and a modified CsgF peptide, or a homologue or mutant thereof, wherein the modified CsgF peptide is as defined in the second aspect of the invention.

[0138] Another embodiment relates to an isolated pore complex, wherein the CsgG pore is coupled to a modified CsgF peptide via at least one monomer by covalent binding. In one instance, the covalent linking or binding can be achieved via a cysteine ​​link, wherein the sulfhydryl side group of the cysteine ​​is covalently linked to another amino acid residue or moiety. In a second possibility, the covalent linking is achieved through interactions between non-natural (photo)reactive amino acids. (Photo)reactive amino acids are artificial analogs of natural amino acids that can be used for cross-linking protein complexes and can be incorporated into proteins and peptides in vivo or in vitro. Commonly used photoreactive amino acid analogs are photoreactive diazirine analogs of leucine and methionine, and p-benzoyl-phenyl-alanine, as well as azidohomoalanine, homopropargylglycine, homoallylglycine, p-acetyl-Phe, p-azido-Phe, p-propargyloxy-Phe, and p-benzoyl-Phe (Wang et al., 2012; Chin et al., 2002). Upon exposure to UV light, they become activated and covalently bind to interacting proteins within a few angstroms of the photoreactive amino acid analogs. However, the position within the CsgG monomer where this covalent attachment occurs depends on the exposure to the modified CsgF peptide. As shown in FIG1 , several amino acids are located in positions that provide for covalent attachment, namely positions 132, 133, 136, 138, 140, 142, 144, 145, 147, 149, 151, 153, 155, 183, 185, 187, 189, 191, 201, 203, 205, 207, or 209 of SEQ ID NO: 3 or a homolog thereof.

[0139] Another aspect of the invention relates to constructs comprising the modified CsgF peptide, wherein the peptide is covalently attached. A "construct" comprises two or more covalently attached monomers derived from modified CsgF and / or CsgG or homologs thereof. In other words, a construct may contain more than one monomer. In another aspect, the invention also provides a pore complex comprising at least one construct of the invention. The pore complex contains sufficient constructs and, if necessary, monomers to form the pore. For example, an octameric pore may comprise (a) four constructs each containing two monomers, (b) two constructs each containing four monomers, (c) one construct containing two monomers and six monomers that do not form part of the construct, or (d) one construct containing one or two CsgF monomers and one construct containing six to seven CsgG monomers, or even (e) a construct containing CsgF and CsgG monomers in addition to another construct containing only CsgG monomers. For example, the same and other possibilities are provided for nonameric pores. Other combinations of constructs and monomers can be envisioned by the skilled artisan. One or more constructs of the present invention can be used to form a pore complex for characterizing (such as sequencing) a polynucleotide. The construct can comprise at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10 monomers. The construct preferably comprises two monomers. The two or more monomers can be the same or different and can be CsgF, CsgG, a CsgG / CsgF fusion monomer, or a homolog thereof, or any combination thereof.

[0140] Another embodiment relates to a polynucleotide or nucleic acid molecule encoding the modified CsgF peptide of the present invention or a homologue or mutant thereof, or a polynucleotide encoding a construct as described above.

[0141] Certain embodiments relate to an isolated transmembrane pore complex comprising the isolated pore complex according to the first and second aspects of the invention, and a component of a membrane. The isolated transmembrane pore complex can be directly suitable for molecular sensing, such as nucleic acid sequencing. Alternatively, a membrane composition is provided comprising a modified CsgG / CsgF biological pore as described herein with respect to the isolated pore complex of the invention, and a membrane, a membrane component, or an insulating layer. One embodiment relates to an isolated transmembrane pore complex comprising the isolated pore complex according to the invention and a component of a membrane.

[0142] Although the CsgG:CsgF complex is very stable, when CsgF is truncated, the stability of the CsgG:CsgF complex is reduced compared to the complex containing full-length CsgF. Therefore, disulfide bonds can be generated between CsgG and CsgF to make the complex more stable, for example, after introducing cysteine ​​residues at the positions identified herein. The pore complex can be prepared using any of the methods previously mentioned, and disulfide bond formation can be induced by using an oxidizing agent (e.g., copper-o-phenanthroline). Other interactions (e.g., hydrophobic interactions, charge-charge interactions / electrostatic interactions) can also be used instead of cysteine ​​interactions at those positions.

[0143] In another embodiment, non-natural amino acids can also be incorporated into those positions. In this embodiment, covalent bonds can be formed by click chemistry. For example, non-natural amino acids with azide or alkynes or with dibenzocyclooctyne (DBCO) groups and / or bicyclo [6.1.0] nonyne (BCN) groups can be introduced into one or more of these positions.

[0144] Such stabilizing mutations can be combined with any other modifications to CsgG and / or CsgF, such as those disclosed herein.

[0145] The CsgG pore may comprise at least one, for example 2, 3, 4, 5, 6, 7, 8, 9, or 10, CsgG monomers modified to facilitate attachment to a CsgF peptide. For example, cysteine ​​residues may be introduced at one or more positions corresponding to positions 132, 133, 136, 138, 140, 142, 144, 145, 147, 149, 151, 153, 155, 183, 185, 187, 189, 191, 201, 203, 205, 207, and 209 of SEQ ID NO: 3, and / or at any of the positions identified in Table 4 as predicted to contact CsgF, to facilitate covalent attachment to CsgG. As an alternative to or in addition to covalent attachment via cysteine ​​residues, the pore may be stabilized by hydrophobic or electrostatic interactions. To promote such interactions, non-native reactive or photoreactive amino acids are introduced at positions corresponding to one or more of positions 132, 133, 136, 138, 140, 142, 144, 145, 147, 149, 151, 153, 155, 183, 185, 187, 189, 191, 201, 203, 205, 207, and 209 of SEQ ID NO: 3, and / or at any of the positions identified in Table 4 as predicted to contact CsgF.

[0146] The CsgF peptide can be modified to facilitate attachment to the CsgG pore. For example, cysteine ​​residues can be introduced at one or more positions corresponding to positions 1, 4, 5, 8, 9, 11, 12, 26, or 29 of SEQ ID NO: 6, and / or at any of the positions identified in Table 4 as predicted to contact CsgF, to facilitate covalent attachment to CsgG. As an alternative or in addition to covalent attachment via cysteine ​​residues, the pore can be stabilized by hydrophobic or electrostatic interactions. To facilitate such interactions, a non-naturally reactive or photoreactive amino acid is introduced at one or more positions corresponding to positions 1, 4, 5, 8, 9, 11, 12, 26, or 29 of SEQ ID NO: 6, and / or at any of the positions identified in Table 4 as predicted to contact CsgF.

[0147] Preferred exemplary CsgF peptides include those corresponding to SEQ ID The following mutation of NO:6: N15X1 / N17X2 / A20X3 / N24X4 / A28X5 / D34X6, wherein X1 is N / S / A / T / Q / G / L / V / I / F / Y / W / R / K / D / C, X2 is N / S / A / T / Q / G / L / V / I / F / Y / W / R / K / D / C, X3 is A / S / T / Q / N / G / L / V / I / F / Y / W / R / K / D / C, X4 is N / S / T / Q / A / G / L / V / I / F / Y / W / R / K / D / C, X5 is A / S / T / Q / N / G / L / V / I / F / Y / W / R / K / D / C and X6 is D / F / Y / W / R / K / N / Q / C. Mutations at positions N15, N17, A20, N24, and A28 are contraction mutations, and the mutation at position 34 affects the interaction of CsgF with the bottom of the CsgG pore, stabilizing the interaction.

[0148] CsgG pore

[0149] The CsgG pore may be a homo-oligomeric pore comprising the same mutant monomers of the invention. The CsgG pore may be a hetero-oligomeric pore derived from CsgG, for example comprising at least one mutant monomer as disclosed herein.

[0150] The CsgG pore may contain any number of mutant monomers. The pore typically comprises at least 7, at least 8, at least 9, or at least 10 identical mutant monomers, such as 7, 8, 9, or 10 mutant monomers. The CsgG pore preferably comprises eight or nine identical mutant monomers.

[0151] In a preferred embodiment, all monomers in the hetero-oligomeric CsgG pore (such as 10, 9, 8 or 7 monomers thereof) are mutant monomers as disclosed herein, wherein at least one of them is different from each other. They may all be different from each other.

[0152] The mutant monomers in the CsgG pore are preferably all of approximately the same length or the same length. The barrels of the mutant monomers of the invention in the pore are preferably approximately the same length or the same length. Length can be measured in terms of the number of amino acids and / or length units.

[0153] The mutant monomer may be a variant of SEQ ID NO: 3. Over the entire length of the amino acid sequence of SEQ ID NO: 3, the variant will preferably be at least 50% homologous to the sequence based on amino acid identity. More preferably, over the entire sequence, the variant may be at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, and more preferably at least 95%, 97%, or 99% homologous to the amino acid sequence of SEQ ID NO: 3 based on amino acid identity. Over a stretch of 100 or more (e.g., 125, 150, 175, or 200 or more) consecutive amino acids, there may be at least 80%, e.g., at least 85%, 90%, or 95% amino acid identity ("hard homology").

[0154] The CsgG monomer is highly conserved (as can be readily appreciated from Figures 45 to 47 of WO2017 / 149317). Furthermore, from the knowledge of mutations related to SEQ ID NO: 3, equivalent positions for mutations in CsgG monomers other than SEQ ID NO: 3 can be determined.

[0155] Thus, references to mutant CsgG monomers comprising variants of the sequence set forth in SEQ ID NO:3 and the specific amino acid mutations thereof as set forth in the claims and elsewhere in the specification also encompass mutant CsgG monomers comprising variants of the sequences set forth in SEQ ID NO:68 to 88 and the corresponding amino acid mutations thereof. Similarly, references to constructs, pores, or methods involving the use of pores relating to mutant CsgG monomers comprising variants of the sequence set forth in SEQ ID NO:3 and the specific amino acid mutations thereof as set forth in the claims and elsewhere in the specification also encompass constructs, pores, or methods relating to mutant CsgG monomers comprising variants of the sequences according to the above-disclosed SEQ ID NOs and the corresponding amino acid mutations thereof. It will further be appreciated that the present invention extends to other variant CsgG monomers exhibiting highly conserved regions not explicitly identified in the specification.

[0156] Homology can be determined using standard methods in the art. For example, the UWGCG software package provides the BESTFIT program, which can be used to calculate homology, for example, using (Devereux et al. (1984) Nucleic Acids Research 12, pp. 387-395) with its default settings. PILEUP and BLAST algorithms can be used to calculate homology or correct sequences (such as identifying equivalent residues or corresponding sequences (usually under their default settings), for example, Altschul SF (1993) J Mol Evol 36: 290-300; Altschul, SF et al. (1990) J Mol Biol 215: 403-10. Software for performing BLAST analysis can be publicly available through the National Center for Biotechnology Information (http: / / www.ncbi.nlm.nih.gov / ).

[0157] SEQ ID NO: 3 is a wild-type CsgG monomer from Escherichia coli K-12 substrain MC4100. Variants of SEQ ID NO: 3 may comprise any substitution found in another CsgG homolog. Preferred CsgG homologs are shown in SEQ ID NOs: 68 to 88. Variants may comprise a combination of one or more substitutions found in SEQ ID NOs: 68 to 88 compared to SEQ ID NO: 3. For example, a mutation may be made at any one or more positions in SEQ ID NO: 3 that differ between SEQ ID NO: 3 and any of SEQ ID NOs: 68 to 88. Such a mutation may be a substitution of an amino acid in SEQ ID NO: 3 with an amino acid from the corresponding position in any of SEQ ID NOs: 68 to 88. Alternatively, the mutation at any of these positions may be a substitution with any amino acid, or may be a deletion or insertion mutation, such as a deletion or insertion of 1 to 10 amino acids, such as 2 to 8 or 3 to 6 amino acids. In addition to the mutations disclosed herein, amino acids conserved between SEQ ID NO: 3 and all of SEQ ID NO: 66 to 88 are preferably present in the variants of the invention. However, conservative mutations may be made at any one or more of these positions conserved between SEQ ID NO: 3 and all of SEQ ID NO: 66 to 88.

[0158] The present invention provides pore-forming CsgG mutant monomers comprising any one or more amino acids described herein as substituted into specific positions in SEQ ID NO: 3 at positions corresponding to specific positions in the CsgG monomer structure. Corresponding positions can be determined by standard techniques in the art. For example, the PILEUP and BLAST algorithms mentioned above can be used to align the sequence of the CsgG monomer with SEQ ID NO: 3 to identify corresponding residues.

[0159] The pore-forming mutant monomer generally retains the ability to form the same 3D structure as a wild-type CsgG monomer, such as a CsgG monomer having the sequence of SEQ ID NO: 3. The 3D structure of CsgG is known in the art and is disclosed, for example, in Goyal et al. (2014) Nature 516(7530):250-3. In addition to the mutations described herein, any number of mutations can be made in the wild-type CsgG sequence, provided that the CsgG mutant monomer retains the improved properties conferred upon it by the mutations of the present invention.

[0160] Typically, a CsgG monomer will retain the ability to form a structure comprising three α-helices and five β-sheets. Mutations can be made in at least the CsgG region N-terminal to the first α-helix (starting from S63 in SEQ ID NO: 3), in the second α-helix (from G85 to A99 in SEQ ID NO: 3), in the loop between the second α-helix and the first β-sheet (from Q100 to N120 in SEQ ID NO: 3), in the fourth and fifth β-sheets (S173 to R192 and R198 to T107 in SEQ ID NO: 3, respectively), and in the loop between the fourth and fifth β-sheets (F193 to Q197 in SEQ ID NO: 3) without affecting the ability of the CsgG monomer to form a transmembrane pore capable of translocating a polypeptide. Therefore, it is contemplated that additional mutations can be made in any of these regions of any CsgG monomer without affecting the monomer's ability to form a pore capable of translocating a polynucleotide. It is also contemplated that mutations can be made in other regions, such as in any α-helix (S63 to R76, G85 to A99, or V211 to L236 of SEQ ID NO: 3) or in any β-sheet (I121 to N133, K135 to R142, I146 to R162, S173 to R192, or R198 to T107 of SEQ ID NO: 3), without affecting the ability of the monomer to form a pore that can translocate a polynucleotide. It is also contemplated that deletions of one or more amino acids can be made in any loop region connecting the α-helix and β-sheet and / or in the N-terminal and / or C-terminal region of the CsgG monomer without affecting the ability of the monomer to form a pore that can translocate a polynucleotide.

[0161] In addition to those amino acid substitutions discussed above, amino acid substitutions can also be made to the amino acid sequence of SEQ ID NO: 3, for example, up to 1, 2, 3, 4, 5, 10, 20 or 30 substitutions. Conservative substitutions replace amino acids with other amino acids having similar chemical structures, similar chemical properties or similar side chain volumes. The introduced amino acids can have polarity, hydrophilicity, hydrophobicity, alkalinity, acidity, neutrality or charge similar to the amino acids they replace. Alternatively, conservative substitutions can introduce another aromatic or aliphatic amino acid to replace a pre-existing aromatic or aliphatic amino acid. Conservative amino acid changes are well known in the art and can be selected based on the properties of the 20 major amino acids defined in Table 1 above. In the case where amino acids have similar polarity, this can also be determined with reference to the hydrophilicity scale of the amino acid side chains in Table 2.

[0162] In addition, one or more amino acid residues of the amino acid sequence of SEQ ID NO: 3 may be deleted from the above polypeptides. Up to 1, 2, 3, 4, 5, 10, 20 or 30 or more residues may be deleted.

[0163] Variants may include fragments of SEQ ID NO: 3. Such fragments retain pore-forming activity. Fragments may be at least 50, at least 100, at least 150, at least 200, or at least 250 amino acids in length. Such fragments may be used to generate pores. The fragment preferably comprises the transmembrane domain of SEQ ID NO: 3, i.e., K135-Q153 and S183-S208.

[0164] Alternatively or additionally, one or more amino acids may be added to the above polypeptides. Extensions may be provided at the amino-terminus or carboxyl-terminus of the amino acid sequence of SEQ ID NO:3, or a polypeptide variant or fragment thereof. The extensions may be short, for example, 1 to 10 amino acids in length. Alternatively, the extensions may be longer, for example, up to 50 or 100 amino acids. A carrier protein may be fused to the amino acid sequence according to the present invention. Other fusion proteins are discussed in more detail below.

[0165] The CsgG pore described herein includes the wild-type CsgG pore, or a homologue or mutant / variant thereof. A variant is a polypeptide having an amino acid sequence that differs from the amino acid sequence of SEQ ID NO: 3 and that retains its pore-forming ability. Variants typically contain the region of SEQ ID NO: 3 responsible for pore formation. The pore-forming ability of β-barrel-containing CsgG is provided by a β-sheet in each subunit. Variants of SEQ ID NO: 3 typically contain the β-sheet-forming region of SEQ ID NO: 3, i.e., K134-Q154 and S183-S208. One or more modifications may be made to the β-sheet-forming region of SEQ ID NO: 3, as long as the resulting variant retains its pore-forming ability. Variants of SEQ ID NO: 3 preferably contain one or more modifications, such as substitutions, additions, or deletions, within its α-helices and / or loop regions.

[0166] The mutant CsgG monomer may be a mutant CsgG monomer having a sequence that differs from that of a wild-type CsgG monomer and that retains the ability to form a pore. Mutant monomers may also be referred to herein as variants. Methods for confirming the ability of mutant monomers to form a pore are well known in the art and are discussed in more detail below.

[0167] Specific pore-forming CsgG mutant monomers that may be included in a CsgG pore may comprise one or more of the following modifications:

[0168] - W at the position corresponding to R97 in SEQ ID NO: 3;

[0169] - W at the position corresponding to R93 in SEQ ID NO: 3;

[0170] - Y at the position corresponding to R97 in SEQ ID NO: 3;

[0171] - Y at the position corresponding to R93 in SEQ ID NO: 3;

[0172] - Y at each position corresponding to R93 and R97 in SEQ ID NO: 3;

[0173] - D at the position corresponding to R192 in SEQ ID NO: 3;

[0174] - deletion of the residues at positions corresponding to V105 to I107 in SEQ ID NO: 3;

[0175] - deletion of the residues at one or more positions corresponding to F193 to L199 in SEQ ID NO: 3;

[0176] - deletion of the residues at positions corresponding to F195 to L199 in SEQ ID NO: 3;

[0177] - deletion of the residues at positions corresponding to F193 to L199 in SEQ ID NO: 3;

[0178] - T at the position corresponding to F191 in SEQ ID NO: 3;

[0179] - Q at the position corresponding to K49 in SEQ ID NO: 3;

[0180] - N at the position corresponding to K49 in SEQ ID NO: 3;

[0181] - Q at the position corresponding to K42 in SEQ ID NO: 3;

[0182] - Q at the position corresponding to E44 in SEQ ID NO: 3;

[0183] - N at the position corresponding to E44 in SEQ ID NO: 3;

[0184] - R at the position corresponding to L90 in SEQ ID NO: 3;

[0185] - R at the position corresponding to L91 in SEQ ID NO: 3;

[0186] - R at the position corresponding to I95 in SEQ ID NO: 3;

[0187] - R at the position corresponding to A99 in SEQ ID NO: 3;

[0188] - H at the position corresponding to E101 in SEQ ID NO: 3;

[0189] - K at the position corresponding to E101 in SEQ ID NO: 3;

[0190] - N at the position corresponding to E101 in SEQ ID NO: 3;

[0191] - Q at the position corresponding to E101 in SEQ ID NO: 3;

[0192] - T at the position corresponding to E101 in SEQ ID NO: 3;

[0193] - K at the position corresponding to Q114 in SEQ ID NO: 3.

[0194] The CsgG pore-forming monomer preferably further comprises an A at the position corresponding to Y51 in SEQ ID NO: 3 and / or a Q at the position corresponding to F56 in SEQ ID NO: 3.

[0195] When characterizing (or sequencing) a target polynucleotide, pores constructed from CsgG monomers having an R to W substitution at a position corresponding to position 97 of SEQ ID NO: 3 exhibited improved accuracy compared to otherwise identical pores without the modification at position 97. Improved accuracy was also seen when the CsgG monomers comprised an R to W modification at a position corresponding to position 97 of SEQ ID NO: 3, or an R to Y modification at positions corresponding to positions 93 and 97 of SEQ ID NO: 3, rather than R97W. Thus, a pore can be constructed from one or more mutant CsgG monomers comprising a modification at a position corresponding to R97 or R93 of SEQ ID NO: 3 such that the modification increases the hydrophobicity of the amino acid. For example, such modifications can include amino acid substitutions with any amino acid containing a hydrophobic side chain, including but not limited to W and Y.

[0196] CsgG monomers containing a mutation from R to D, Q, F, S, or T at the position corresponding to position 192 of SEQ ID NO:3 are more easily expressed than monomers without a substitution at position 192, likely due to a reduction in positive charge. Therefore, position 192 can be substituted with an amino acid that reduces positive charge. Monomers containing R192D / Q / F / S / T can also contain other modifications that improve the ability of the mutant pore formed by the monomer to interact with and characterize analytes, such as polynucleotides. However, in one embodiment, it is preferred that the residue at the position corresponding to position 193 of SEQ ID NO:3 is R or K, more preferably R.

[0197] Pores containing CsgG monomers comprising deletions of V105, A106, and I107, deletions of F193, I194, D195, Y196, Q197, R198, and L199, or deletions of D195, Y196, Q197, R198, and L199, and / or F191T, show improved accuracy when characterizing (or sequencing) target polynucleotides. The amino acids at positions 105 to 107 correspond to the cis-loop in the nanopore cap, while the amino acids at positions 193 to 199 correspond to the trans-loop at the other end of the pore. Without wishing to be bound by theory, it is believed that the deletion of the cis-loop improves the interaction of the enzyme with the pore, while the removal of the trans-loop reduces any undesirable interactions between the DNA on the opposite side of the pore.

[0198] Pores comprising a CsgG monomer comprising a K to Q or K to N mutation at position corresponding to K94 of SEQ ID NO: 3 exhibit a reduced number of noisy pores (i.e., those resulting in an increased signal-to-noise ratio) when characterizing (or sequencing) a target polynucleotide compared to the same pore without the mutation at position 94. Position 94 is found within the vestibule of the pore and has been found to be particularly sensitive with respect to noise in the current signal.

[0199] Pores comprising CsgG monomers comprising T104K or T104R, N91R, E101K / N / Q / T / H, E44N / Q, Q114K, A99R, I95R, N91R, L90R, E44Q / N and / or Q42K, or corresponding mutations, all demonstrated improved ability to capture target polynucleotides when used to characterize (or sequence) the target polynucleotide, compared to the same pore without substitutions at these positions.

[0200] In one embodiment, the CsgG pore comprises one or more monomers that are variants of SEQ ID NO: 3, comprising (a) one or more mutations at (i.e., mutations at one or more of) I41, R93, A98, Q100, G103, T104, A106, I107, N108, L113, S115, T117, Y130, K135, E170, S208, D233, D238, and E244 and / or (b) D43S, E44S, F48S / N / Q / Y / W / I / V / H / R / K, Q87N / R / K, N91K / R, K94R / F / Y / W / L / S / N, R97F / Y / W / V / I / K / S / Q / H, E101I / L / A / H, N102K / Q / L / I / V / S / H, R110F / G / N, Q114R / K, R142Q / S, T150Y / A / V / L / S / Q / N, R192D / Q / F / S / T, and D248S / N / Q / K / R. The variant may comprise (a); (b); or (a) and (b). In some embodiments, the variant comprises R97W. In some embodiments, the variant comprises R192D / Q / F / S / T, such as R192D / Q. In (a), the variant can comprise modifications at any number and combination of positions, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, or 19 positions.

[0201] In (a), the variant preferably comprises one or more of I41N, R93F / Y / W / L / I / V / N / Q / S, A98K / R, Q100K / R, G103F / W / S / N / K / R, T104R / K, A106R / K, I107R / K / W / F / Y / L / V, N108R / K, L113K / R, S115R / K, T117R / K, Y130W / F / H / Q / N, K135L / V / N / Q / S, E170S / N / Q / K / R, S208V / I / F / W / Y / L / T, D233S / N / Q / K / R, D238S / N / Q / K / R and E244S / N / Q / K / R.

[0202] In (a), the variant preferably comprises one or more modifications that provide for more consistent movement of the target polynucleotide relative to (such as through) a transmembrane pore comprising a monomer. Specifically, in (a), the variant preferably comprises one or more mutations at the following positions (i.e., mutations at one or more of the following positions): R93, G103, and I107. The variant may comprise R93; G103; I107; R93 and G103; R93 and I107; G103 and I107; or R93, G103, and I107. The variant preferably comprises one or more of R93F / Y / W / L / I / V / N / Q / S, G103F / W / S / N / K / R, and I107R / K / W / F / Y / L / V. These may be present in any combination as shown for positions R93, G103, and I107.

[0203] In (a), the variant preferably comprises one or more modifications that allow the pore constructed from the mutant monomer to preferably capture nucleotides and polynucleotides more easily. Specifically, in (a), the variant preferably comprises one or more mutations at the following positions (i.e., mutations at one or more of the following positions): I41, T104, A106, N108, L113, S115, T117, E170, D233, D238, and E244. The variant may comprise modifications at any number of positions and combinations of positions, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or 11 positions. The variant preferably comprises one or more of I41N, T104R / K, A106R / K, N108R / K, L113K / R, S115R / K, T117R / K, E170S / N / Q / K / R, D233S / N / Q / K / R, D238S / N / Q / K / R and E244S / N / Q / K / R. Additionally or alternatively, the variant may comprise (c) Q42K / R, E44N / Q, L90R / K, N91R / K, I95R / K, A99R / K, E101H / K / N / Q / T and / or Q114K / R.

[0204] In (a), the variant preferably comprises one or more modifications that provide more consistent movement and increase capture. Specifically, in (a), the variant preferably comprises one or more mutations at the following positions (i.e., mutations at one or more of the following positions): (i) A98, (ii) Q100, (iii) G103, and (iv) I107. The variant preferably comprises one or more of the following: (i) A98R / K, (ii) Q100K / R, (iii) G103K / R, and (iv) I107R / K.

[0205] Particularly preferred mutant monomers that increase the capture of analytes, such as polynucleotides, include mutations at one or more of positions Q42, E44, E44, L90, N91, I95, A99, E101, and Q114, which remove negative charge and / or increase positive charge at the mutated position. Specifically, the following mutations can be included in the mutant monomers of the invention to generate CsgG pores with improved ability to capture analytes, preferably polynucleotides: Q42K, E44N, E44Q, L90R, N91R, I95R, A99R, E101H, E101K, E101N, E101Q, E101T, and Q114K. Examples of specific mutant monomers comprising one of these mutations in combination with other beneficial mutations are:

[0206] CsgG-(WT-Y51A / F56Q / K94Q / R97W / R192D-del(V105-I107)-Q42K

[0207] CsgG-(WT-Y51A / F56Q / K94Q / R97W / R192D-del(V105-I107)-E44N

[0208] CsgG-(WT-Y51A / F56Q / K94Q / R97W / R192D-del(V105-I107)-E44Q

[0209] CsgG-(WT-Y51A / F56Q / K94Q / R97W / R192D-del(V105-I107)-L90R

[0210] CsgG-(WT-Y51A / F56Q / K94Q / R97W / R192D-del(V105-I107)-N91R

[0211] CsgG-(WT-Y51A / F56Q / K94Q / R97W / R192D-del(V105-I107)-I95R

[0212] CsgG-(WT-Y51A / F56Q / K94Q / R97W / R192D-del(V105-I107)-A99R

[0213] CsgG-(WT-Y51A / F56Q / K94Q / R97W / R192D-del(V105-I107)-E101H

[0214] CsgG-(WT-Y51A / F56Q / K94Q / R97W / R192D-del(V105-I107)-E101K

[0215] CsgG-(WT-Y51A / F56Q / K94Q / R97W / R192D-del(V105-I107)-E101N

[0216] CsgG-(WT-Y51A / F56Q / K94Q / R97W / R192D-del(V105-I107)-E101Q

[0217] CsgG-(WT-Y51A / F56Q / K94Q / R97W / R192D-del(V105-I107)-E101T

[0218] CsgG-(WT-Y51A / F56Q / K94Q / R97W / R192D-del(V105-I107)-Q114K.

[0219] In (a), variant preferably includes one or more modifications that provide increased characterization accuracy. Specifically, in (a), variant preferably includes one or more mutations (i.e., mutations at one or more positions) at the following positions: Y130, K135, and S208, such as Y130; K135; S208; Y130 and K135; Y130 and S208; K135 and S208; or Y130, K135, and S208. Variant preferably includes one or more of Y130W / F / H / Q / N, K135L / V / N / Q / S, and R142Q / S. These substitutions can exist in any quantity and combination as described for Y130, K135, and S208.

[0220] In (b), the variant may comprise any number and combination of substitutions, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 or 12 substitutions. In (b), the variant preferably comprises one or more modifications that provide for more consistent movement of the target polynucleotide relative to (such as by) a transmembrane pore comprising a monomer. Specifically, in (b), the variant preferably comprises one or more of the following: (i) Q87N / R / K, (ii) K94R / F / Y / W / L / S / N, (iii) R97F / Y / W / V / I / K / S / Q / H, (iv) N102K / Q / L / I / V / S / H and (v) R110F / G / N. More preferably, the variant comprises K94D or K94Q and / or R97W or R97Y. Other preferred variants modified to provide more consistent movement of the target polynucleotide relative to (e.g., through) a transmembrane pore comprising a monomer include (vi) R93W and R93Y. Preferred variants may include R93W and R97W, R93Y and R97W, R93W and R97W, or more preferably R93Y and R97Y.

[0221] In (b), the variant preferably comprises one or more modifications that allow the pore constructed from the mutant monomer to preferably capture nucleotides and polynucleotides more easily. Specifically, in (b), the variant preferably comprises one or more of the following: (i) D43S, (ii) E44S, (iii) N91K / R, (iv) Q114R / K and (v) D248S / N / Q / K / R.

[0222] In (b), the variant preferably includes one or more modifications that provide more consistent motion and increased capture. Specifically, in (b), the variant preferably includes one or more of the following: Q87R / K, E101I / L / A / H, and N102K, such as Q87R / K; E101I / L / A / H; N102K; Q87R / K and E101I / L / A / H; Q87R / K and N102K; E101I / L / A / H and N102K; or Q87R / K, E101I / L / A / H, and N102K.

[0223] In (b), the variant preferably comprises one or more modifications that provide increased characterization accuracy. Specifically, in (a), the variant preferably comprises F48S / N / Q / Y / W / I / V.

[0224] In (b), the variant preferably comprises one or more modifications that provide increased characterization accuracy and increased capture. Specifically, in (a), the variant preferably comprises F48H / R / K.

[0225] Variants may include modifications of (a) and (b) that provide for more consistent movement. Variants may include modifications of (a) and (b) that provide for increased capture.

[0226] The present invention provides variants of SEQ ID NO: 3, wherein the use of pores comprising the variants provides increased throughput for assays for characterizing analytes such as polynucleotides. Such variants may comprise a mutation at K94, preferably K94Q or K94N, more preferably K94Q. Examples of specific mutant monomers comprising K94Q or K94N mutations in combination with other beneficial mutations are:

[0227] CsgG-(WT-Y51A / F56Q / R97W / R192D-StrepII)9-K94N

[0228] CsgG-(WT-Y51A / F56Q / R97W / R192D-StrepII)9-K94Q.

[0229] Using monomers that are variants of SEQ ID NO: 3 to form CsgG pores can provide improved characterization accuracy in assays for characterizing analytes such as polynucleotides. Such variants include variants comprising: a mutation at F191, preferably F191T; a deletion of V105-I107; a deletion of F193-L199 or D195-L199; and / or a mutation at R93 and / or R97, preferably R93Y, R97Y, or more preferably R97W, R93W, or both R97Y and R97Y. Examples of specific mutant monomers comprising one or more of these mutations in combination with other beneficial mutations are:

[0230] CsgG-(WT-Y51A / F56Q / R97W / R192D-StrepII)9-del(D195-L199)

[0231] CsgG-(WT-Y51A / F56Q / R97W / R192D-StrepII)9-del(F193-L199)

[0232] CsgG-(WT-Y51A / F56Q / R97W / R192D-StrepII)9-F191T

[0233] CsgG-(WT-Y51A / F56Q / R97W / R192D-del(V105-I107)-StrepII)9

[0234] CsgG-(WT-Y51A / F56Q / K94Q / R97W / R192D-del(V105-I107)

[0235] CsgG-(WT-Y51A / F56Q / R192D-StrepII)9-R93W

[0236] CsgG-(WT-Y51A / F56Q / R192D-StrepII)9-R93W-del(D195-L199)

[0237] CsgG-(WT-Y51A / F56Q / R192D-StrepII)9-R93Y / R97Y.

[0238] In another embodiment, a variant of SEQ ID NO: 3 comprises (A) a deletion of one or more positions R192, F193, I194, D195, Y196, Q197, R198, L199, L200, and E201 and / or (B) a deletion of one or more of V139 / G140 / D149 / T150 / V186 / Q187 / V204 / G205 (referred to herein as Band 1), G137 / G138 / Q151 / Y152 / Y184 / E185 / Y206 / T207 (referred to herein as Band 2), and A141 / R142 / G147 / A148 / A188 / G189 / G202 / E203 (referred to herein as Band 3).

[0239] In (A), the variant may comprise any number of positions and combinations of positions, such as the deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 positions. In (A), the variant preferably comprises the following deletions:

[0240] -D195, Y196, Q197, R198 and L199;

[0241] -R192, F193, I194, D195, Y196, Q197, R198, L199 and L200;

[0242] -Q197, R198, L199 and L200;

[0243] -I194, D195, Y196, Q197, R198 and L199;

[0244] -D195, Y196, Q197, R198, L199 and L200;

[0245] -Y196, Q197, R198, L199, L200 and E201;

[0246] -Q197, R198, L199, L200 and E201;

[0247] -Q197, R198, L199; or

[0248] -F193, I194, D195, Y196, Q197, R198 and L199.

[0249] More preferably, the variant comprises a deletion of D195, Y196, Q197, R198 and L199 or F193, I194, D195, Y196, Q197, R198 and L199. In (B), any number of bands 1 to 3 and combinations thereof may be deleted, for example, band 1; band 2; band 3; bands 1 and 2; bands 1 and 3; bands 2 and 3; or bands 1, 2 and 3. The variant may comprise a deletion according to (A); (B); or (A) and (B).

[0250] Variants comprising deletions of one or more positions according to (A) and / or (B) above may also comprise any of the modifications or substitutions discussed above and below. If one or more positions appearing after the deletion position in SEQ ID NO: 3 are modified or substituted, the numbering of the modified or substituted position(s) must be adjusted accordingly. For example, if L199 is deleted, E244 becomes E243. Similarly, if band 1 is deleted, R192 becomes R186.

[0251] In another embodiment, the variant of SEQ ID NO: 3 comprises (C) a deletion of one or more positions V105, A106 and 1107. In addition to the deletions according to (A) and / or (B), the deletion according to (C) may also be generated.

[0252] Such deletions generally reduce noise that would be associated with the movement of the target polynucleotide relative to (eg, through) a transmembrane pore comprising a monomer.Thus, the target polynucleotide can be more accurately characterized.

[0253] In the above paragraphs where the / symbol is used to separate different amino acids at a specific position, the / symbol means "or." For example, Q87R / K means Q87R or Q87K.

[0254] Variants of SEQ ID NO:3 that increase capture of analytes such as polynucleotides may include a mutation at T104, preferably T104R or T104K; a mutation at N91, preferably a mutation of N91R; a mutation at E101, preferably E101K / N / Q / T / H; a mutation at position E44, preferably E44N or E44Q, and / or a mutation at position Q42, preferably Q42K.

[0255] Mutations at different positions in SEQ ID NO: 3 may be combined in any possible manner. Specifically, a monomer in a CsgG pore may comprise one or more mutations that improve accuracy, one or more mutations that reduce noise, and / or one or more mutations that enhance analyte capture.

[0256] Variants of SEQ ID NO: 3 preferably comprise one or more of the following: (i) one or more mutations at the following positions (i.e., mutations at one or more of the following positions): N40, D43, E44, S54, S57, Q62, R97, E101, E124, E131, R142, T150, and R192, such as one or more mutations at the following positions (i.e., mutations at one or more of the following positions): N40, D43, E44, S54, S57, Q62, E101, E131, and T150 or N40, D43, E44, E101, and E131; (ii) one or more mutations at S1 / N55, Y51 / F56, N55 / F56, or Y57 / F57. (i) a mutation at position 1 / N55 / F56; (ii) a mutation at position 21; (iii) Q42R or Q42K; (iv) K49R; (v) N102R, N102F, N102Y or N102W; (vi) D149N, D149Q or D149R; (vii) E185N, E185Q or E185R; (viii) D195N, D195Q or D195R; (ix) E201N, E201Q or E201R; (x) E203N, E203Q or E203R; and (xi) a deletion at one or more of the following positions: F48, K49, P50, Y51, P52, A53, S54, N55, F56 and S57. The variant may comprise any combination of (i) to (xi).

[0257] If the variant comprises any of (i) and (iii) to (xi), it may further comprise a mutation at one or more of Y51, N55 and F56, such as at Y51, N55, F56, Y51 / N55, Y51 / F56, N55 / F56 or Y51 / N55 / F56.

[0258] In (i), variants may include mutations in any number and combination of N40, D43, E44, S54, S57, Q62, R97, E101, E124, E131, R142, T150, and R192. In (i), variants preferably include one or more mutations (i.e., mutations at one or more positions) at the following positions: N40, D43, E44, S54, S57, Q62, E101, E131, and T150. In (i), variants preferably include one or more mutations (i.e., mutations at one or more positions) at the following positions: N40, D43, E44, E101, and E131. In (i), variants preferably include mutations at S54 and / or S57. In (i), the variant more preferably comprises a mutation at the following positions: (a) S54 and / or S57 and (b) one or more of Y51, N55 and F56, such as Y51, N55, F56, Y51 / N55, Y51 / F56, N55 / F56 or Y51 / N55 / F56. If S54 and / or S57 are missing in (xi), they cannot be mutated in (i), and vice versa. In (i), the variant preferably comprises a mutation at T150, such as T150I. Alternatively, the variant preferably comprises a mutation at the following positions: (a) T150 and (b) one or more of Y51, N55 and F56, such as Y51, N55, F56, Y51 / N55, Y51 / F56, N55 / F56 or Y51 / N55 / F56. In (i), the variant preferably comprises a mutation at Q62, such as Q62R or Q62K. Alternatively, the variant preferably comprises a mutation at (a) Q62 and (b) one or more of Y51, N55 and F56, such as Y51, N55, F56, Y51 / N55, Y51 / F56, N55 / F56 or Y51 / N55 / F56. The variant may comprise a mutation at D43, E44, Q62 or any combination thereof, such as D43, E44, Q62, D43 / E44, D43 / Q62, E44 / Q62 or D43 / E44 / Q62. Alternatively, the variant preferably comprises a mutation in the following positions: (a) D43, E44, Q62, D43 / E44, D43 / Q62, E44 / Q62 or D43 / E44 / Q62 and (b) one or more of Y51, N55 and F56, such as Y51, N55, F56, Y51 / N55, Y51 / F56, N55 / F56 or Y51 / N55 / F56.

[0259] In (ii) and elsewhere in this application where the / symbol separates different positions, the / symbol means "and", so that Y51 / N55 means Y51 and N55. In (ii), the variant preferably comprises a mutation at Y51 / N55. It has been proposed that the constriction in CsgG is composed of three stacked concentric circles formed by the side chains of residues Y51, N55, and F56 (Goyal et al., 2014, Nature, 516, 250-253). Therefore, mutation of these residues in (ii) can reduce the number of nucleotides that contribute to the current when the polynucleotide moves through the pore, thereby making it easier to identify the direct relationship between the observed current (when the polynucleotide moves through the pore) and the polynucleotide. F56 can be mutated in any of the ways discussed below with reference to the variants and pores that can be used in the methods of the present invention.

[0260] In (v), the variant may comprise N102R, N102F, N102Y or N102W. The variant preferably comprises (a) N102R, N102F, N102Y or N102W and (b) a mutation at one or more of Y51, N55 and F56, such as Y51, N55, F56, Y51 / N55, Y51 / F56, N55 / F56 or Y51 / N55 / F56.

[0261] In (xi), any number and combination of K49, P50, Y51, P52, A53, S54, N55, F56, and S57 may be deleted. Preferably, one or more of K49, P50, Y51, P52, A53, S54, N55, and S57 may be deleted. If Y51, N55, and F56 are deleted in (xi), they cannot be mutated in (ii), and vice versa.

[0262] In (i), the variant preferably comprises one or more of the following substitutions: N40R, N40K, D43N, D43Q, D43R, D43K, E44N, E44Q, E44R, E44K, S54P, S57P, Q62R, Q62K, R97N, R97G, R97L, E101N, E101Q, E101R, E101K, E101F, E101Y, E101W, E124N, E124Q, E124R, E124K, E124F, E124Y, E124W, E131D, R142E, R142N, T150I, R192E, and R192N, such as N40 R, N40K, D43N, D43Q, D43R, D43K, E44N, E44Q, E44R, E44K, S54P, S57P, Q62R, Q62K, E101N, E101Q, E101R, E101K, E101F, E101Y, E101W, E131D, and T150I, or one or more of N40R, N40K, D43N, D43Q, D43R, D43K, E44N, E44Q, E44R, E44K, E101N, E101Q, E101R, E101K, E101F, E101Y, E101W, and E131D. Variants can contain any number of these substitutions and combinations thereof. In (i), the variant preferably comprises S54P and / or S57P. In (i), the variant preferably comprises (a) S54P and / or S57P and (b) a mutation at one or more of Y51, N55, and F56, such as at Y51, N55, F56, Y51 / N55, Y51 / F56, N55 / F56, or Y51 / N55 / F56. The mutation at one or more of Y51, N55, and F56 can be any of the mutations discussed below. In (i), the variant preferably comprises F56A / S57P or S54P / F56A. The variant preferably comprises T150I. Alternatively, the variant preferably comprises mutations in the following positions: (a) T150I and (b) one or more of Y51, N55 and F56, such as Y51, N55, F56, Y51 / N55, Y51 / F56, N55 / F56 or Y51 / N55 / F56.

[0263] In (i), the variant preferably comprises Q62R or Q62K. Alternatively, the variant preferably comprises (a) Q62R or Q62K and (b) a mutation at one or more of Y51, N55 and F56, such as at Y51, N55, F56, Y51 / N55, Y51 / F56, N55 / F56 or Y51 / N55 / F56. The variant may comprise D43N, E44N, Q62R or Q62K or any combination thereof, such as D43N, E44N, Q62R, Q62K, D43N / E44N, D43N / Q62R, D43N / Q62K, E44N / Q62R, E44N / Q62K, D43N / E44N / Q62R or D43N / E44N / Q62K. Alternatively, the variant preferably comprises (a) D43N, E44N, Q62R, Q62K, D43N / E44N, D43N / Q62R, D43N / Q62K, E44N / Q62R, E44N / Q62K, D43N / E44N / Q62R or D43N / E44N / Q62K and (b) a mutation at one or more of Y51, N55 and F56, such as at Y51, N55, F56, Y51 / N55, Y51 / F56, N55 / F56 or Y51 / N55 / F56.

[0264] In (i), the variant preferably comprises D43N.

[0265] In (i), the variant preferably comprises E101R, E101S, E101F or E101N.

[0266] In (i), the variant preferably comprises E124N, E124Q, E124R, E124K, E124F, E124Y, E124W or E124D, such as E124N.

[0267] In (i), the variant preferably comprises R142E and R142N.

[0268] In (i), the variant preferably comprises R97N, R97G or R97L.

[0269] In (i), the variant preferably comprises R192E and R192N.

[0270] In (ii), the variant preferably comprises F56N / N55Q, F56N / N55R, F56N / N55K, F56N / N55S, F56N / N55G, F56N / N55A, F56N / N55T, F56Q / N55Q, F56Q / N55R, F56Q / N55K, F56Q / N55S, F56Q / N55G, F56Q / N55A, F56Q / N55T, F56R / N55Q, F56R / N55R, F56R / N55K, F56R / N55S, F56R / N55G, F56R / N55A, F56R / N55T, F56S / N55Q, F56S / N55R, F56S / N55K, F56S / N55S, F56S / N55G, F56S / N55A, F56S / N55T, F56G / N55Q, F56G / N55R, F56G / N55K, F56G / N55S, F56G / N55G, F56G / N55A, F56G / N55T, F56A / N55Q, F56A / N55R, F56A / N55K, F56A / N55S, F56A / N55G, F56A / N55A, F56A / N55T, F56K / N55Q, F56K / N55R,F56K / N55K, F56K / N55S, F56K / N55G, F56K / N55A, F56K / N55T, F56N / Y51L, F56N / Y51V, F56N / Y51A, F56N / Y51N, F56N / Y51Q, F56N / Y51S, F56N / Y51G, F56Q / Y51L, F56Q / Y51V, F56Q / Y51A, F56Q / Y51N, F56Q / Y51Q, F56Q / Y51S, F56Q / Y51G, F56R / Y51L, F56R / Y51V, F56R / Y51A, F56R / Y51N, F56R / Y51Q, F56R / Y51S, F56R / Y51G, F56S / Y51L, F56S / Y51V, F56S / Y51A, F56S / Y51N, F56S / Y51Q, F56S / Y51S, F56S / Y51G, F56G / Y51L, F56G / Y51V, F56G / Y51A, F56G / Y51N, F56G / Y51Q, F56G / Y51S, F56G / Y51G, F56A / Y51L, F56A / Y51V, F56A / Y51A, F56A / Y51N, F56A / Y51Q, F56A / Y51S, F56A / Y51G, F56K / Y51L, F56K / Y51V, F56K / Y51A, F56K / Y51N, F56K / Y51Q, F56K / Y51S, F56K / Y51G,N55Q / Y51L、N55Q / Y51V、N55Q / Y51A、N55Q / Y51N、N55Q / Y51Q、N55Q / Y51S、N55Q / Y51G、N55R / Y51L、N55R / Y51V、N55R / Y51A、N55R / Y51N、N55R / Y51Q、N55R / Y51S、N55R / Y51G、N55K / Y51L、N55K / Y51V、N55K / Y51A、N55K / Y51N、N55K / Y51Q、N55K / Y51S、N55K / Y51G、N55S / Y51L、N55S / Y51V、N55S / Y51A、N55S / Y51N、N55S / Y51Q、N55S / Y51S、N55S / Y51G、N55G / Y51L、N55G / Y51V、N55G / Y51A、N55G / Y51N、N55G / Y51Q、N55G / Y51S、N55G / Y51G、N55A / Y51L、N55A / Y51V、N55A / Y51A、N55A / Y51N、N55A / Y51Q、N55A / Y51S、N55A / Y51G、N55T / Y51L、N55T / Y51V、N55T / Y51A、N55T / Y51N、N55T / Y51Q、N55T / Y51S、N55T / Y51G、F56N / N55Q / Y51L、F56N / N55Q / Y51V、F56N / N55Q / Y51A、F56N / N55Q / Y51N、F56N / N55Q / Y51Q、F56N / N55Q / Y51S、F56N / N55Q / Y51G、F56N / N55R / Y51L、F56N / N55R / Y51V、F56N / N55R / Y51A、F56N / N55R / Y51N、F56N / N55R / Y51Q、F56N / N55R / Y51S、F56N / N55R / Y51G、F56N / N55K / Y51L、F56N / N55K / Y51V、F56N / N55K / Y51A、F56N / N55K / Y51N、F56N / N55K / Y51Q、F56N / N55K / Y51S、F56N / N55K / Y51G、F56N / N55S / Y51L、F56N / N55S / Y51V、F56N / N55S / Y51A、F56N / N55S / Y51N、F56N / N55S / Y51Q、F56N / N55S / Y51S、F56N / N55S / Y51G、F56N / N55G / Y51L、F56N / N55G / Y51V、F56N / N55G / Y51A、F56N / N55G / Y51N、F56N / N55G / Y51Q、F56N / N55G / Y51S、F56N / N55G / Y51G、F56N / N55A / Y51L、F56N / N55A / Y51V、F56N / N55A / Y51A、F56N / N55A / Y51N、F56N / N55A / Y51Q、F56N / N55A / Y51S、F56N / N55A / Y51G、F56N / N55T / Y51L、F56N / N55T / Y51V、F56N / N55T / Y51A、F56N / N55T / Y51N、F56N / N55T / Y51Q、F56N / N55T / Y51S、F56N / N55T / Y51G、F56Q / N55Q / Y51L、F56Q / N55Q / Y51V、F56Q / N55Q / Y51A、F56Q / N55Q / Y51N、F56Q / N55Q / Y51Q、F56Q / N55Q / Y51S、F56Q / N55Q / Y51G、F56Q / N55R / Y51L、F56Q / N55R / Y51V、F56Q / N55R / Y51A、F56Q / N55R / Y51N、F56Q / N55R / Y51Q、F56Q / N55R / Y51S、F56Q / N55R / Y51G、F56Q / N55K / Y51L、F56Q / N55K / Y51V、F56Q / N55K / Y51A、F56Q / N55K / Y51N、F56Q / N55K / Y51Q、F56Q / N55K / Y51S、F56Q / N55K / Y51G、F56Q / N55S / Y51L、F56Q / N55S / Y51V、F56Q / N55S / Y51A、F56Q / N55S / Y51N、F56Q / N55S / Y51Q、F56Q / N55S / Y51S、F56Q / N55S / Y51G、F56Q / N55G / Y51L、F56Q / N55G / Y51V、F56Q / N55G / Y51A、F56Q / N55G / Y51N、F56Q / N55G / Y51Q、F56Q / N55G / Y51S、F56Q / N55G / Y51G、F56Q / N55A / Y51L、F56Q / N55A / Y51V、F56Q / N55A / Y51A、F56Q / N55A / Y51N、F56Q / N55A / Y51Q、F56Q / N55A / Y51S、F56Q / N55A / Y51G、F56Q / N55T / Y51L、F56Q / N55T / Y51V、F56Q / N55T / Y51A、F56Q / N55T / Y51N、F56Q / N55T / Y51Q、F56Q / N55T / Y51S、F56Q / N55T / Y51G、F56R / N55Q / Y51L、F56R / N55Q / Y51V、F56R / N55Q / Y51A、F56R / N55Q / Y51N、F56R / N55Q / Y51Q、F56R / N55Q / Y51S、F56R / N55Q / Y51G、F56R / N55R / Y51L、F56R / N55R / Y51V、F56R / N55R / Y51A、F56R / N55R / Y51N、F56R / N55R / Y51Q、F56R / N55R / Y51S、F56R / N55R / Y51G、F56R / N55K / Y51L、F56R / N55K / Y51V、F56R / N55K / Y51A、F56R / N55K / Y51N、F56R / N55K / Y51Q、F56R / N55K / Y51S、F56R / N55K / Y51G、F56R / N55S / Y51L、F56R / N55S / Y51V、F56R / N55S / Y51A、F56R / N55S / Y51N、F56R / N55S / Y51Q、F56R / N55S / Y51S、F56R / N55S / Y51G、F56R / N55G / Y51L、F56R / N55G / Y51V、F56R / N55G / Y51A、F56R / N55G / Y51N、F56R / N55G / Y51Q、F56R / N55G / Y51S、F56R / N55G / Y51G、F56R / N55A / Y51L、F56R / N55A / Y51V、F56R / N55A / Y51A、F56R / N55A / Y51N、F56R / N55A / Y51Q、F56R / N55A / Y51S、F56R / N55A / Y51G、F56R / N55T / Y51L、F56R / N55T / Y51V、F56R / N55T / Y51A、F56R / N55T / Y51N、F56R / N55T / Y51Q、F56R / N55T / Y51S、F56R / N55T / Y51G、F56S / N55Q / Y51L、F56S / N55Q / Y51V、F56S / N55Q / Y51A、F56S / N55Q / Y51N、F56S / N55Q / Y51Q、F56S / N55Q / Y51S、F56S / N55Q / Y51G、F56S / N55R / Y51L、F56S / N55R / Y51V、F56S / N55R / Y51A、F56S / N55R / Y51N、F56S / N55R / Y51Q、F56S / N55R / Y51S、F56S / N55R / Y51G、F56S / N55K / Y51L、F56S / N55K / Y51V、F56S / N55K / Y51A、F56S / N55K / Y51N、F56S / N55K / Y51Q、F56S / N55K / Y51S、F56S / N55K / Y51G、F56S / N55S / Y51L、F56S / N55S / Y51V、F56S / N55S / Y51A、F56S / N55S / Y51N、F56S / N55S / Y51Q、F56S / N55S / Y51S、F56S / N55S / Y51G、F56S / N55G / Y51L、F56S / N55G / Y51V、F56S / N55G / Y51A、F56S / N55G / Y51N、F56S / N55G / Y51Q、F56S / N55G / Y51S、F56S / N55G / Y51G、F56S / N55A / Y51L、F56S / N55A / Y51V、F56S / N55A / Y51A、F56S / N55A / Y51N、F56S / N55A / Y51Q、F56S / N55A / Y51S、F56S / N55A / Y51G、F56S / N55T / Y51L、F56S / N55T / Y51V、F56S / N55T / Y51A、F56S / N55T / Y51N、F56S / N55T / Y51Q、F56S / N55T / Y51S、F56S / N55T / Y51G、F56G / N55Q / Y51L、F56G / N55Q / Y51V、F56G / N55Q / Y51A、F56G / N55Q / Y51N、F56G / N55Q / Y51Q、F56G / N55Q / Y51S、F56G / N55Q / Y51G、F56G / N55R / Y51L、F56G / N55R / Y51V、F56G / N55R / Y51A、F56G / N55R / Y51N、F56G / N55R / Y51Q、F56G / N55R / Y51S、F56G / N55R / Y51G、F56G / N55K / Y51L、F56G / N55K / Y51V、F56G / N55K / Y51A、F56G / N55K / Y51N、F56G / N55K / Y51Q、F56G / N55K / Y51S、F56G / N55K / Y51G、F56G / N55S / Y51L、F56G / N55S / Y51V、F56G / N55S / Y51A、F56G / N55S / Y51N、F56G / N55S / Y51Q、F56G / N55S / Y51S、F56G / N55S / Y51G、F56G / N55G / Y51L、F56G / N55G / Y51V、F56G / N55G / Y51A、F56G / N55G / Y51N、F56G / N55G / Y51Q、F56G / N55G / Y51S、F56G / N55G / Y51G、F56G / N55A / Y51L、F56G / N55A / Y51V、F56G / N55A / Y51A、F56G / N55A / Y51N、F56G / N55A / Y51Q、F56G / N55A / Y51S、F56G / N55A / Y51G、F56G / N55T / Y51L、F56G / N55T / Y51V、F56G / N55T / Y51A、F56G / N55T / Y51N、F56G / N55T / Y51Q、F56G / N55T / Y51S、F56G / N55T / Y51G、F56A / N55Q / Y51L、F56A / N55Q / Y51V、F56A / N55Q / Y51A、F56A / N55Q / Y51N、F56A / N55Q / Y51Q、F56A / N55Q / Y51S、F56A / N55Q / Y51G、F56A / N55R / Y51L、F56A / N55R / Y51V、F56A / N55R / Y51A、F56A / N55R / Y51N、F56A / N55R / Y51Q、F56A / N55R / Y51S、F56A / N55R / Y51G、F56A / N55K / Y51L、F56A / N55K / Y51V、F56A / N55K / Y51A、F56A / N55K / Y51N、F56A / N55K / Y51Q、F56A / N55K / Y51S、F56A / N55K / Y51G、F56A / N55S / Y51L、F56A / N55S / Y51V、F56A / N55S / Y51A、F56A / N55S / Y51N、F56A / N55S / Y51Q、F56A / N55S / Y51S、F56A / N55S / Y51G、F56A / N55G / Y51L、F56A / N55G / Y51V、F56A / N55G / Y51A、F56A / N55G / Y51N、F56A / N55G / Y51Q、F56A / N55G / Y51S、F56A / N55G / Y51G、F56A / N55A / Y51L、F56A / N55A / Y51V、F56A / N55A / Y51A、F56A / N55A / Y51N、F56A / N55A / Y51Q、F56A / N55A / Y51S、F56A / N55A / Y51G、F56A / N55T / Y51L、F56A / N55T / Y51V、F56A / N55T / Y51A、F56A / N55T / Y51N、F56A / N55T / Y51Q、F56A / N55T / Y51S、F56A / N55T / Y51G、F56K / N55Q / Y51L、F56K / N55Q / Y51V、F56K / N55Q / Y51A、F56K / N55Q / Y51N、F56K / N55Q / Y51Q, F56K / N55Q / Y51S, F56K / N55Q / Y51G, F56K / N55R / Y51L, F56K / N55R / Y51V, F56K / N55R / Y51A, F56K / N55R / Y51N, F56K / N55R / Y51Q, F56K / N55R / Y51S, F56K / N55R / Y51G, F56K / N55K / Y51L, F56K / N55K / Y51V, F56K / N55K / Y51A, F56K / N55K / Y51N, F56K / N55K / Y51Q, F56K / N55K / Y51S, F56K / N55K / Y51G, F56K / N55S / Y51L, F56K / N55S / Y51V, F56K / N55S / Y51A, F56K / N55S / Y51N, F56K / N55S / Y51Q, F56K / N55S / Y51S, F56K / N55S / Y51G, F56K / N55G / Y51L, F56K / N55G / Y51V, F56K / N55G / Y51A, F56K / N55G / Y51N, F56K / N55G / Y51Q, F56K / N55G / Y51S, F56K / N55G / Y51G, F56K / N55A / Y51L, F56K / N55A / Y51V, F56K / N55A / Y51A, F56K / N55A / Y51N, F56K / N55A / Y51Q, F56K / N55A / Y51S, F56K / N55A / Y51G, F56K / N55T / Y51L, F56K / N55T / Y51V, F56K / N55T / Y51A, F56K / N55T / Y51N, F56K / N55T / Y51Q, F56K / N55T / Y51S, F56K / N55T / Y51G, F56E / N55R, F56E / N55K, F56D / N55R, F56D / N55K, F56R / N55E, F56R / N55D, F56K / N55E or F56K / N55D.,

[0271] In (ii), the variant preferably comprises Y51R / F56Q, Y51N / F56N, Y51M / F56Q, Y51L / F56Q, Y51I / F56Q, Y51V / F56Q, Y51A / F56Q, Y51P / F56Q, Y51G / F56Q, Y51C / F56Q, Y51Q / F56Q, Y51N / F56Q, Y51S / F56Q, Y51E / F56Q, Y51D / F56Q, Y51K / F56Q or Y51H / F56Q.

[0272] In (ii), the variant preferably comprises Y51T / F56Q, Y51Q / F56Q or Y51A / F56Q.

[0273] In (ii), the variant preferably comprises Y51T / F56F, Y51T / F56M, Y51T / F56L, Y51T / F56I, Y51T / F56V, Y51T / F56A, Y51T / F56P, Y51T / F56G, Y51T / F56C, Y51T / F56Q, Y51T / F56N, Y51T / F56T, Y51T / F56S, Y51T / F56E, Y51T / F56D, Y51T / F56K, Y51T / F56H or Y51T / F56R.

[0274] In (ii), the variant preferably comprises Y51T / N55Q, Y51T / N55S or Y51T / N55A.

[0275] In (ii), the variant preferably comprises Y51A / F56F, Y51A / F56L, Y51A / F56I, Y51A / F56V, Y51A / F56A, Y51A / F56P, Y51A / F56G, Y51A / F56C, Y51A / F56Q, Y51A / F56N, Y51A / F56T, Y51A / F56S, Y51A / F56E, Y51A / F56D, Y51A / F56K, Y51A / F56H or Y51A / F56R.

[0276] In (ii), the variant preferably comprises Y51C / F56A, Y51E / F56A, Y51D / F56A, Y51K / F56A, Y51H / F56A, Y51Q / F56A, Y51N / F56A, Y51S / F56A, Y51P / F56A or Y51V / F56A.

[0277] In (xi), the variant preferably comprises deletion and replacement of Y51 / P52, Y51 / P52 / A53, P50 to P52, P50 to A53, K49 to Y51, K49 to A53 with a single proline (P), deletion and replacement of K49 to S54 with a single P, Y51 to A53, Y51 to S54, N55 / F56, N55 to S57, deletion and replacement of N55 / F56 with a single P, deletion and replacement of N55 / F56 with a single glycine (G), deletion and replacement of N55 / F56 with a single alanine (A), deletion of N55 / F56 with a single amino acid residue (G), and replacement of N55 / F56 with a single amino acid residue (A). Deletion and replacement with a single P and Y51N, deletion of N55 / F56 and replacement with a single P and Y51Q, deletion of N55 / F56 and replacement with a single P and Y51S, deletion of N55 / F56 and replacement with a single G and Y51N, deletion of N55 / F56 and replacement with a single G and Y51Q, deletion of N55 / F56 and replacement with a single G and Y51S, deletion of N55 / F56 and replacement with a single A and Y51N, deletion of N55 / F56 and replacement with a single A / Y51Q or deletion of N55 / F56 and replacement with a single A and Y51S.

[0278] More preferably, the variant comprises D195N / E203N, D195Q / E203N, D195N / E203Q, D195Q / E203Q, E201N / E203N, E201Q / E203N, E201N / E203Q, E201Q / E203Q, E185N / E203Q, E185Q / E203Q, E185N / E203N, E185Q / E203N, D195N / E201N / E203N, D195Q / E201N / E203N, D195N / E201Q / E203N, D195Q / E201Q / E203N, E201N / E203Q, D195N / E201Q / E203Q, D195Q / E201Q / E203Q, D149N / E201N, D1 49Q / E201N, D149N / E201Q, D149Q / E201Q, D149N / E201N / D195N, D149Q / E201 N / D195N, D149N / E201Q / D195N, D149N / E201N / D195Q, D149Q / E201Q / D195N, D149Q / E201N / D195Q, D149N / E201Q / D195Q, D149Q / E201Q / D195Q, D149N / E2 03N, D149Q / E203N, D149N / E203Q, D149Q / E203Q, D149N / E185N / E201N, D149 Q / E185N / E201N, D149N / E185Q / E201N, D149N / E185N / E201Q, D149Q / E185Q / E201N, D149Q / E185N / E201Q, D149N / E185Q / E201Q, D149Q / E185Q / E201Q, D1 49N / E185N / E203N, D149Q / E185N / E203N, D149N / E185Q / E203N, D149N / E185 N / E203Q, D149Q / E185Q / E203N, D149Q / E185N / E203Q, D149N / E185Q / E203Q, D149Q / E185Q / E203Q, D149N / E185N / E201N / E203N, D149Q / E185N / E201N / E2 03N, D149N / E185Q / E201N / E203N, D149N / E185N / E201Q / E203N, D149N / E185 N / E201N / E203Q, D149Q / E185Q / E201N / E203N, D149Q / E185N / E201Q / E203N,D149Q / E185N / E201N / E203Q、D149N / E185Q / E201Q / E203N、D149N / E185Q / E201N / E203Q、D149N / E185N / E201Q / E203Q、D149Q / E185Q / E201Q / E203Q、D149Q / E185Q / E201N / E203Q、D149Q / E185N / E201Q / E203Q、D149N / E185Q / E201Q / E203Q、D149Q / E185Q / E201Q / E203N、D149N / E185N / D195N / E201N / E203N、D149Q / E185N / D195N / E201N / E203N、D149N / E185Q / D195N / E201N / E203N、D149N / E185N / D195Q / E201N / E203N、D149N / E185N / D195N / E201Q / E203N、D149N / E185N / D195N / E201N / E203Q、D149Q / E185Q / D195N / E201N / E203N、D149Q / E185N / D195Q / E201N / E203N、D149Q / E185N / D195N / E201Q / E203N、D149Q / E185N / D195N / E201N / E203Q、D149N / E185Q / D195Q / E201N / E203N、D149N / E185Q / D195N / E201Q / E203N、D149N / E185Q / D195N / E201N / E203Q、D149N / E185N / D195Q / E201Q / E203N、D149N / E185N / D195Q / E201N / E203Q、D149N / E185N / D195N / E201Q / E203Q、D149Q / E185Q / D195Q / E201N / E203N、D149Q / E185Q / D195N / E201Q / E203N、D149Q / E185Q / D195N / E201N / E203Q、D149Q / E185N / D195Q / E201Q / E203N、D149Q / E185N / D195Q / E201N / E203Q、D149Q / E185N / D195N / E201Q / E203Q、D149N / E185Q / D195Q / E201Q / E203N、D149N / E185Q / D195Q / E201N / E203Q、D149N / E185Q / D195N / E201Q / E203Q、D149N / E185N / D195Q / E201Q / E203Q、D149Q / E185Q / D195Q / E201Q / E203N、D149Q / E185Q / D195Q / E201N / E203Q、D149Q / E185Q / D195N / E201Q / E203Q、D149Q / E185N / D195Q / E201Q / E203Q、D149N / E185Q / D195Q / E201Q / E203Q、D149Q / E185Q / D195Q / E201Q / E203Q、D149N / E185R / E201N / E203N、D149Q / E185R / E201N / E203N、D149N / E185R / E201Q / E203N、D149N / E185R / E201N / E203Q、D149Q / E185R / E201Q / E203N、D149Q / E185R / E201N / E203Q、D149N / E185R / E201Q / E203Q、D149Q / E185R / E201Q / E203Q、D149R / E185N / E201N / E203N、D149R / E185Q / E201N / E203N、D149R / E185N / E201Q / E203N、D149R / E185N / E201N / E203Q、D149R / E185Q / E201Q / E203N、D149R / E185Q / E201N / E203Q、D149R / E185N / E201Q / E203Q、D149R / E185Q / E201Q / E203Q、D149R / E185N / D195N / E201N / E203N、D149R / E185Q / D195N / E201N / E203N、D149R / E185N / D195Q / E201N / E203N、D149R / E185N / D195N / E201Q / E203N、D149R / E185Q / D195N / E201N / E203Q、D149R / E185Q / D195Q / E201N / E203N、D149R / E185Q / D195N / E201Q / E203N、D149R / E185Q / D195N / E201N / E203Q、D149R / E185N / D195Q / E201Q / E203N、D149R / E185N / D195Q / E201N / E203Q、D149R / E185N / D195N / E201Q / E203Q、D149R / E185Q / D195Q / E201Q / E203N、D149R / E185Q / D195Q / E201N / E203Q、D149R / E185Q / D195N / E201Q / E203Q、D149R / E185N / D195Q / E201Q / E203Q, D149R / E185Q / D195Q / E201Q / E203Q, D 149N / E185R / D195N / E201N / E203N, D149Q / E185R / D195N / E201N / E203N, D1 49N / E185R / D195Q / E201N / E203N, D149N / E185R / D195N / E201Q / E203N, D14 9N / E185R / D195N / E201N / E203Q, D149Q / E185R / D195Q / E201N / E203N, D149Q / E185R / D195N / E201Q / E203N、D149Q / E185R / D195N / E201N / E203Q、D149N / E185R / D195Q / E201Q / E203N, D149N / E185R / D195Q / E201N / E203Q, D149N / E1 85R / D195N / E201Q / E203Q, D149Q / E185R / D195Q / E201Q / E203N, D149Q / E18 5R / D195Q / E201N / E203Q, D149Q / E185R / D195N / E201Q / E203Q, D149N / E185R / D195Q / E201Q / E203Q、D149Q / E185R / D195Q / E201Q / E203Q、D149N / E185R / D195N / E201R / E203N, D149Q / E185R / D195N / E201R / E203N, D149N / E185R / D1 95Q / E201R / E203N, D149N / E185R / D195N / E201R / E203Q, D149Q / E185R / D19 5Q / E201R / E203N, D149Q / E185R / D195N / E201R / E203Q, D149N / E185R / D195Q / E201R / E203Q, D149Q / E185R / D195Q / E201R / E203Q, E131D / K49R, E101N / N 102F, E101N / N102Y, E101N / N102W, E101F / N102F, E101F / N102Y, E101F / N10 2W, E101Y / N102F, E101Y / N102Y, E101Y / N102W, E101W / N102F, E101W / N102 Y, E101W / N102W, E101N / N102R, E101F / N102R, E101Y / N102R or E101W / N102F. ,

[0279] Preferred variants of the invention in which fewer nucleotides contribute to the current as the polynucleotide moves through the pore, in the formed pore, comprise Y51A / F56A, Y51A / F56N, Y51I / F56A, Y51L / F56A, Y51T / F56A, Y51I / F56N, Y51L / F56N or Y51T / F56N or more preferably Y51I / F56A, Y51L / F56A or Y51T / F56A. As discussed above, this makes it easier to identify a direct relationship between the observed current (as the polynucleotide moves through the pore) and the polynucleotide.

[0280] Preferred variants that form pores showing increased range comprise mutations at the following positions:

[0281] Y51, F56, D149, E185, E201, and E203;

[0282] N55 and F56;

[0283] Y51 and F56;

[0284] Y51, N55 and F56; or

[0285] F56 and N102.

[0286] Preferred variants for forming pores showing increased range include:

[0287] Y51N, F56A, D149N, E185R, E201N, and E203N;

[0288] N55S and F56Q;

[0289] Y51A and F56A;

[0290] Y51A and F56N;

[0291] Y51I and F56A;

[0292] Y51L and F56A;

[0293] Y51T and F56A;

[0294] Y51I and F56N;

[0295] Y51L and F56N;

[0296] Y51T and F56N;

[0297] Y51T and F56Q;

[0298] Y51A, N55S, and F56A;

[0299] Y51A, N55S, and F56N;

[0300] Y51T, N55S, and F56Q; or

[0301] F56Q and N102R.

[0302] Preferred variants that form a pore in which fewer nucleotides contribute to the current as the polynucleotide moves through the pore comprise mutations at the following positions:

[0303] N55 and F56, such as N55X and F56Q, wherein X is any amino acid; or

[0304] Y51 and F56, such as Y51X and F56Q, wherein X is any amino acid.

[0305] A particularly preferred variant comprises Y51A and F56Q.

[0306] Preferred variants that form pores showing increased throughput comprise mutations at the following positions:

[0307] D149, E185, and E203;

[0308] D149, E185, E201 and E203; or

[0309] D149, E185, D195, E201 and E203.

[0310] Preferred variations for forming pores that exhibit increased throughput include:

[0311] D149N, E185N, and E203N;

[0312] D149N, E185N, E201N, and E203N;

[0313] D149N, E185R, D195N, E201N and E203N; or

[0314] D149N, E185R, D195N, E201R and E203N.

[0315] Preferred variants that increase capture of polynucleotides in the pore formed comprise the following mutations:

[0316] D43N / Y51T / F56Q;

[0317] E44N / Y51T / F56Q;

[0318] D43N / E44N / Y51T / F56Q;

[0319] Y51T / F56Q / Q62R;

[0320] D43N / Y51T / F56Q / Q62R;

[0321] E44N / Y51T / F56Q / Q62R; or

[0322] D43N / E44N / Y51T / F56Q / Q62R.

[0323] Preferred variants comprise the following mutations:

[0324] D149R / E185R / E201R / E203R or Y51T / F56Q / D149R / E185R / E201R / E203R;

[0325] D149N / E185N / E201N / E203N or Y51T / F56Q / D149N / E185N / E201N / E203N;

[0326] E201R / E203R or Y51T / F56Q / E201R / E203R

[0327] E201N / E203R or Y51T / F56Q / E201N / E203R;

[0328] E203R or Y51T / F56Q / E203R;

[0329] E203N or Y51T / F56Q / E203N;

[0330] E201R or Y51T / F56Q / E201R;

[0331] E201N or Y51T / F56Q / E201N;

[0332] E185R or Y51T / F56Q / E185R;

[0333] E185N or Y51T / F56Q / E185N;

[0334] D149R or Y51T / F56Q / D149R;

[0335] D149N or Y51T / F56Q / D149N;

[0336] R142E or Y51T / F56Q / R142E;

[0337] R142N or Y51T / F56Q / R142N;

[0338] R192E or Y51T / F56Q / R192E; or

[0339] R192N or Y51T / F56Q / R192N.

[0340] Preferred variants comprise the following mutations:

[0341] Y51A / F56Q / E101N / N102R;

[0342] Y51A / F56Q / R97N / N102G;

[0343] Y51A / F56Q / R97N / N102R;

[0344] Y51A / F56Q / R97N;

[0345] Y51A / F56Q / R97G;

[0346] Y51A / F56Q / R97L;

[0347] Y51A / F56Q / N102R;

[0348] Y51A / F56Q / N102F;

[0349] Y51A / F56Q / N102G;

[0350] Y51A / F56Q / E101R;

[0351] Y51A / F56Q / E101F;

[0352] Y51A / F56Q / E101N; or

[0353] Y51A / F56Q / E101G

[0354] The variant preferably further comprises a mutation at T150. Preferred variants that form pores showing increased insertion comprise T150I. Mutations at T150, such as T150I, may be combined with any mutation or combination of mutations discussed above.

[0355] SEQ ID NO:3 preferred variants include (a) R97W and (b) mutations at Y51 and / or F56. SEQ ID NO:3 preferred variants include (a) R97W and (b) Y51R / H / K / D / E / S / T / N / Q / C / G / P / A / V / I / L / M and / or F56R / H / K / D / E / S / T / N / Q / C / G / P / A / V / I / L / M. SEQ ID NO:3 preferred variants include (a) R97W and (b) Y51L / V / A / N / Q / S / G and / or F56A / Q / N. SEQ ID NO:3 preferred variants include (a) R97W and (b) Y51A and / or F56Q. SEQ ID NO:3 preferred variants include R97W, Y51A and F56Q.

[0356] Variants of SEQ ID NO: 3 preferably comprise a mutation at R192. Variants preferably comprise R192D / Q / F / S / T / N / E, R192D / Q / F / S / T or R192D / Q. Preferred variants of SEQ ID NO: 3 comprise (a) R97W, (b) a mutation at Y51 and / or F56 and (c) a mutation at R192, such as R192D / Q / F / S / T / N / E, R192D / Q / F / S / T or R192D / Q. Preferred variants of SEQ ID NO: 3 comprise (a) R97W, (b) Y51 R / H / K / D / E / S / T / N / Q / C / G / P / A / V / I / L / M and / or F56 R / H / K / D / E / S / T / N / Q / C / G / P / A / V / I / L / M and (c) a mutation at R192, such as R192D / Q / F / S / T / N / E, R192D / Q / F / S / T or R192D / Q. Preferred variants of SEQ ID NO: 3 include (a) R97W, (b) Y51L / V / A / N / Q / S / G and / or F56A / Q / N, and (c) a mutation at R192, such as R192D / Q / F / S / T / N / E, R192D / Q / F / S / T or R192D / Q. Preferred variants of SEQ ID NO: 3 include (a) R97W, (b) Y51A and / or F56Q and (c) a mutation at R192, such as R192 D / Q / F / S / T / N / E, R192D / Q / F / S / T or R192D / Q. Preferred variants of SEQ ID NO: 3 include R97W, Y51A, F56Q and R192D / Q / F / S / T or R192D / Q. Preferred variants of SEQ ID NO: 3 include R97W, Y51A, F56Q, and R192D. Preferred variants of SEQ ID NO: 3 include R97W, Y51A, F56Q, and R192Q. In the above paragraphs where the / symbol separates different amino acids at specific positions, the / symbol means "or." For example, R192D / Q means R192D or R192Q.

[0357] Any of the above preferred variants of SEQ ID NO: 3 may further comprise a mutation at R93. Preferred variants of SEQ ID NO: 3 comprise (a) R93W and (b) mutations at Y51 and / or F56, preferably Y51A and F56Q.

[0358] Any of the above preferred variants of SEQ ID NO: 3 may comprise a K94N / Q mutation. Any of the above preferred variants of SEQ ID NO: 3 may comprise a F191T mutation.

[0359] The CsgG monomer can be modified to facilitate attachment to the CsgF peptide. For example, cysteine ​​residues can be introduced at one or more positions corresponding to positions 132, 133, 136, 138, 140, 142, 144, 145, 147, 149, 151, 153, 155, 183, 185, 187, 189, 191, 201, 203, 205, 207, and 209 of SEQ ID NO: 3, and / or at any of the positions identified in Table 4 as predicted to contact CsgF, to facilitate covalent attachment to CsgG. As an alternative to or in addition to covalent attachment via cysteine ​​residues, the pore can be stabilized by hydrophobic or electrostatic interactions. To promote such interactions, non-native reactive or photoreactive amino acids are introduced at positions corresponding to one or more of positions 132, 133, 136, 138, 140, 142, 144, 145, 147, 149, 151, 153, 155, 183, 185, 187, 189, 191, 201, 203, 205, 207, and 209 of SEQ ID NO: 3, and / or at any of the positions identified in Table 4 as predicted to contact CsgF.

[0360] Preferred exemplary pores comprise at least one CsgG monomer having the following mutations relative to SEQ ID NO:3: Y51X1 / N55X2 / F56X3 / N91R / K94Q / R97W / R192D-del (V105-I107), wherein X1 is I / V / S / T, X2 is N / I / V / S / T and / or X3 is Q / I / V / S / T.

[0361] Methods for introducing or replacing naturally occurring amino acids are well known in the art. For example, methionine (M) can be replaced with arginine (R) by replacing the codon for methionine (ATG) with the codon for arginine (CGT) at the relevant position in the polynucleotide encoding the mutant monomer. The polynucleotide can then be expressed as discussed below.

[0362] Double hole

[0363] The CsgG / CsgF pore can be a double pore comprising a first pore and a second pore. At least the first pore is a CsgG / CsgF pore disclosed herein. The second pore can be a CsgG pore or a CsgG / CsgF pore. In one embodiment, both the first pore and the second pore are CsgG / CsgF pores disclosed herein. The first pore and the second pore can be the same or different. In addition to any mutations disclosed herein, the CsgG monomer in the double pore can further comprise one or more additional mutations described below.

[0364] In a dual pore, the first pore can be attached to the second CsgG pore via hydrophobic interactions and / or via one or more disulfide bonds. One or more, such as two, three, four, five, six, eight, nine, or for example all, monomers of the first and / or second pores can be modified to enhance such interactions. This can be achieved in any suitable manner.

[0365] At least one cysteine ​​residue in the amino acid sequence of the first pore at the interface between the first pore and the second pore may be disulfide bonded to at least one cysteine ​​residue in the amino acid sequence of the second pore at the interface between the first pore and the second pore. The cysteine ​​residue in the first pore and / or the cysteine ​​residue in the second pore may be a cysteine ​​residue not present in a wild-type CsgG monomer. Multiple disulfide bonds may be formed between the two pores in the dual pore, such as 2, 3, 4, 5, 6, 7, 8, or 9 to 16, 18, 24, 27, 32, 36, 40, 45, 48, 54, 56, or 63. One or both of the first or second pores may comprise at least one monomer, such as up to 8, 9, or 10 monomers, comprising a cysteine ​​residue at a position corresponding to R97, I107, R110, Q100, E101, N102, and / or L113 of SEQ ID NO: 3 at the interface between the first and second pores.

[0366] At least one monomer in the first pore and / or at least one monomer in the second pore may comprise at least one residue at the interface between the first and second pores that is more hydrophobic than the residue present at the corresponding position in the wild-type CsgG monomer. For example, 2 to 10, such as 3, 4, 5, 6, 7, 8, or 9, residues in the first and / or second pores may be more hydrophobic than the residues at the same position in the corresponding wild-type CsgG monomer. Such hydrophobic residues enhance the interaction between the two pores in the dual pore. The at least one residue at the interface between the first and second pores may be at a position corresponding to R97, I107, R110, Q100, E101, N102, and / or L113 of SEQ ID NO:3. Where the residue at the interface in the wild-type CsgG monomer is R, Q, N, or E, the hydrophobic residue is typically I, L, V, M, F, W, or Y. Where the residue at the interface in the wild-type CsgG monomer is I, the hydrophobic residue is typically L, V, M, F, W, or Y. Where the residue at the interface in the wild-type CsgG monomer is L, the hydrophobic residue is typically I, V, M, F, W or Y.

[0367] The dual pores may comprise one or more monomers comprising one or more cysteine ​​residues at the interface between the pores and one or more monomers comprising one or more introduced hydrophobic residues at the interface between the pores, or may comprise one or more monomers comprising such cysteine ​​residues and such hydrophobic residues. For example, one or more, such as any two, three, or four, of the positions corresponding to R97, I107, R110, Q100, E101, N102, and / or L113 of SEQ ID NO: 3 in the monomer may comprise a cysteine ​​(C) residue and one or more, such as any two, three, or four, of the positions corresponding to R97, I107, R110, Q100, E101, N102, and / or L113 of SEQ ID NO: 3 in the monomer may comprise a hydrophobic residue, such as I, L, V, M, F, W, or Y.

[0368] The double pore may contain bulky residues at one or more (such as 2, 3, 4, 5, 6, or 7) positions in the tail region. These residues are typically located at the interface between the first and second pores and are larger than the residues present at the corresponding positions in the wild-type CsgG monomer. The bulk of these residues prevents the formation of holes in the pore wall at the interface between the first and second pores in the double pore. The at least one bulky residue at the interface between the first and second pores is typically at a position corresponding to A98, A99, T104, V105, L113, Q114, or S115 of SEQ ID NO: 3. In the case where the residue at the interface in the wild-type CsgG monomer is A, the bulky residue is typically I, L, V, M, F, W, Y, N, Q, S, or T. In the case where the residue at the interface in the wild-type CsgG monomer is T, the bulky residue is typically L, M, F, W, Y, N, Q, R, D, or E. When the residue present at the interface in the wild-type CsgG monomer is V, the bulky residue is typically I, L, M, F, W, Y, N, or Q. When the residue present at the interface in the wild-type CsgG monomer is L, the bulky residue is typically M, F, W, Y, N, Q, R, D, or E. When the residue present at the interface in the wild-type CsgG monomer is Q, the bulky residue is typically F, W, or Y. When the residue present at the interface in the wild-type CsgG monomer is S, the bulky residue is typically M, F, W, Y, N, Q, E, or R.

[0369] Particularly where the second pore is located externally of the membrane, the second pore, and optionally the first pore, preferably contain residues in the barrel region of the pore that reduce the negative charge of the barrel interior compared to the charge in the barrel of a wild-type CsgG pore. These mutations render the barrel more hydrophilic. At least one monomer in the first pore and / or at least one monomer in the second pore of the dual pore may contain at least one residue in the barrel region of the pore that has a less negative charge than the residue present at the corresponding position in a wild-type CsgG monomer. The charge inside the barrel is sufficiently neutral or positive that negatively charged analytes (such as polynucleotides) are not repelled from entering the pore due to electrostatic charge. At least one residue, such as two, three, four, or five residues, at positions corresponding to D149, E185, D195, E210, and / or E203 of SEQ ID NO: 3 in the barrel region of the pore may be neutral or positively charged amino acids. At least one residue, such as 2, 3, 4 or 5 residues at positions in the barrel region of the pore corresponding to D149, E185, D195, E210 and / or E203 of SEQ ID NO: 3 is preferably N, Q, R or K.

[0370] Specific examples of mutations that remove charge in SEQ ID NO: 3 include the following: E185N / E203N, D149N / E185R / D195N / E201R / E203N, D149N / E185R / D195N / E201N / E203N, D149R / E185N / D195N / E201N / E203N, D149R / E185N / E201N / E203N, D149N / E185N / D195 / E201N / E203N, D149N / E185N / E203N 1N, D149N / E201N, D195N / E201N, D149N / E203N, D195N / E201N / E203N, E201N / E203N, D195N / E203, E203R, E203N, E201R, E201N, D195R, D195N, E185R, E185N, D149R, and D149N.

[0371] At least one CsgG monomer in the first pore may include at least one residue in the constriction of the barrel region of the first pore that reduces, maintains, or increases the length of the constriction compared to a wild-type CsgG pore, and / or at least one monomer in the second CsgG pore may include at least one residue in the constriction of the barrel region of the second pore that reduces, maintains, or increases the length of the constriction compared to a wild-type CsgG pore. Preferably, the length of the constriction in the first pore and / or the length of the constriction in the second pore is at least as long as in a wild-type pore, more preferably longer.

[0372] The length of the pore can be increased by inserting residues into the region corresponding to the region between positions K49 and F56 of SEQ ID NO: 3. One to five, such as two, three or four, amino acid residues can be inserted at any one or more of the following positions defined with reference to SEQ ID NO: 3: K49 and P50, P50 and Y51, Y51 and P52, P52 and A53, A53 and S54, S54 and N55 and / or N55 and F56. Preferably, a total of 1 to 10, such as 2 to 8, or 3 to 5 amino acid residues are inserted into the monomer sequence. Preferably, all monomers in the first pore and / or all monomers in the second pore have the same number of insertions in this region. The inserted residues can increase the length of the loop between the residues corresponding to Y51 and N55 of SEQ ID NO: 3. The inserted residues may be any combination of A, S, G, or T to maintain flexibility; may be P to add a kink to the loop; and / or may be S, T, N, Q, M, F, W, Y, V, and / or I to contribute to the signal generated when the analyte interacts with the barrel of the pore under an applied potential difference. The inserted amino acids may be any combination of S, G, SG, SGG, SGS, GS, GSS, and / or GSG.

[0373] In a dual pore, the constriction in the barrel of the first and / or second pore may comprise at least one residue, such as 2, 3, 4, or 5 residues, that affects the properties of the pore when used to detect or characterize an analyte, compared to using the first or second pore with a wild-type constriction, wherein the at least one residue in the constriction of the barrel region of the pore is at a position corresponding to Y51, N55, Y51, P52, and / or A53 of SEQ ID NO: 3. The at least one residue may be Q or V at a position corresponding to F56 of SEQ ID NO: 3; A or Q at a position corresponding to Y51 of SEQ ID NO: 3; and / or V at a position corresponding to N55 of SEQ ID NO: 3.

[0374] The dual pore may comprise at least one monomer in the first CsgG pore and / or at least one monomer in the second CsgG pore, the monomers comprising two or more mutations as defined above.

[0375] The CsgG monomer in the dual pore may comprise cysteine ​​residues at positions corresponding to R97, I107, R110, Q100, E101, N102, and L113 of SEQ ID NO:3.

[0376] The CsgG monomer in the dual pore may comprise a residue at a position corresponding to any one or more of R97, Q100, I107, R110, E101, N102, and L113 of SEQ ID NO:3 that is more hydrophobic than the residue present at the corresponding position of SEQ ID NO:3, such as the corresponding position of any one of SEQ ID NOs: 68 to 88, wherein the residue at the position corresponding to R97 and / or I107 is M, the residue at the position corresponding to R110 is I, L, V, M, W, or Y, and / or the residue at the position corresponding to E101 or N102 is V or M. The residue at the position corresponding to Q100 is typically I, L, V, M, F, W, or Y; and / or the residue at the position corresponding to L113 is typically I, V, M, F, W, or Y.

[0377] The specific monomer may have the sequence shown in SEQ ID NO 3, which comprises Y51A, F56Q substitutions and R97I / V / L / M / F / W / Y, I107L / V / M / F / W / Y, R110I / V / L / M / F / W / Y, Q100I / V / L / M / F / W / Y, E101I / V / L / M / F / W / Y, N102I / V / L / M / F / W / Y and L113C1 / V / L / M / F / W / Y, R97I / V / L / M / F / W / Y and N102I / V / L / M / F / W / Y in combination and / or R97I / V / L / M / F / W / Y and E101I / V / L / M / F / W / Y in combination. I107 may have formed a hydrophobic interaction between the two pores.

[0378] The CsgG monomer in at least one of the pores of the dual pore may comprise a residue at a position corresponding to any one or more of A98, A99, T104, V105, L113, Q114, and S115 of SEQ ID NO: 3 that is bulkier than the residue present at the corresponding position of SEQ ID NO: 3, such as the corresponding position of any one of SEQ ID NOs: 68 to 88, wherein the residue at the position corresponding to T104 is L, M, F, W, Y, N, Q, D, or E, the residue at the position corresponding to L113 is M, F, W, Y, N, G, D, or E, and / or the residue at the position corresponding to S115 is M, F, W, Y, N, Q, or E. The residue at the position corresponding to A98 or A99 is typically I, L, V, M, F, W, Y, N, Q, S, or T. The residue at the position corresponding to V105 is I, L, M, F, W, Y, N, or Q. The residue at the position corresponding to Q114 is F, W, or Y. The residue at the position corresponding to E210 is N, Q, R, or K.

[0379] A specific monomer can have the sequence shown in SEQ ID NO 3, which includes Y51A, F56Q substitutions and 1, 2, 3, 4, 5, 6 or all of the following substitutions: A98I / L / V / M / F / W / Y / N / Q / S / T; A99I / L / V / M / F / W / Y / N / Q / S / T; T104N / Q / L / R / D / E / M / F / W / Y; V105I / L / M / F / W / Y / N / Q; L113M / F / W / Y / N / Q / D / E / L / R; Q114Y / F / W; and S115N / Q / M / F / W / Y / E / R.

[0380] The CsgG monomer in at least one of the dual pores may comprise a residue in the barrel region of the pore at positions corresponding to any one or more of D149, E185, D195, E210 and E203 that is less negatively charged than the residue present at the corresponding position of SEQ ID NO: 3, such as any one of SEQ ID NOs: 68 to 88, wherein the residues at positions corresponding to D149, E185, D195 and / or E203 are K.

[0381] The CsgG monomer in at least one of the dual pores may comprise at least one residue in the constriction of the barrel region of the pore that increases the length of the constriction compared to a wild-type CsgG pore. The at least one residue is additional to a residue present in the constriction of the wild-type CsgG pore.

[0382] The length of the pore can be increased by inserting residues into the region corresponding to the region between positions K49 and F56 of SEQ ID NO: 3. One to five, such as two, three, or four, amino acid residues can be inserted at any one or more of the following positions defined with reference to SEQ ID NO: 3: K49 and P50, P50 and Y51, Y51 and P52, P52 and A53, A53 and S54, S54 and N55, and / or N55 and F56. Preferably, a total of 1 to 10, such as 2 to 8, or 3 to 5, amino acid residues are inserted into the monomer sequence. The inserted residues can increase the length of the loop between the residues corresponding to Y51 and N55 of SEQ ID NO: 3. The inserted residues can be any combination of A, S, G, or T to maintain flexibility; can be P to add a kink to the loop; and / or can be S, T, N, Q, M, F, W, Y, V, and / or I to contribute to the signal generated when the analyte interacts with the barrel of the pore under an applied potential difference. The inserted amino acids may be any combination of S, G, SG, SGG, SGS, GS, GSS and / or GSG.

[0383] The CsgG monomer in at least one of the dual pores may comprise at least one residue in the constriction of the barrel region of the pore at a position corresponding to N55, P52 and / or A53 of SEQ ID NO:3 that differs from the residue present in the corresponding wild-type monomer, wherein the residue at the position corresponding to N55 is V.

[0384] Any two or more of the above residues may be present in the same monomer.

[0385] In particular, the monomer may comprise at least one of said cysteine ​​residues, at least one of said hydrophobic residues, at least one of said bulky residues, at least one of said neutral or positively charged residues and / or at least one of said residues that increase the length of the contraction.

[0386] The CsgG monomer in the dual pore may further comprise one or more, such as 2, 3, 4 or 5, residues that affect the properties of the pore when used to detect or characterize an analyte, compared to using a first or second pore having a wild-type constriction, wherein the at least one residue in the constriction of the barrel region of the pore is at a position corresponding to Y51, N55, Y51, P52 and / or A53 of SEQ ID NO: 3. The at least one residue may be Q or V at the position corresponding to F56 of SEQ ID NO: 3; A or Q at the position corresponding to Y51 of SEQ ID NO: 3; and / or V at the position corresponding to N55 of SEQ ID NO: 3.

[0387] Methods for preparing modified proteins

[0388] Methods for introducing or substituting non-naturally occurring amino acids are also well known in the art. For example, non-naturally occurring amino acids can be introduced by including synthetic aminoacyl-tRNAs in the IVTT system used to express mutant monomers. Alternatively, non-naturally occurring amino acids can be introduced by expressing mutant monomers in E. coli, which are auxotrophic for specific amino acids in the presence of synthetic (i.e., non-naturally occurring) analogs of those specific amino acids. If mutant monomers are produced using partial peptide synthesis, they can also be produced by naked ligation.

[0389] Monomers derived from CsgG can be modified to aid their identification or purification, for example by adding a streptavidin tag or by adding a signal sequence to promote their secretion from cells in which the monomers do not naturally contain such sequences. Other suitable tags are discussed in more detail below. The monomers can be labeled with a recognizable marker. The recognizable marker can be any suitable marker that allows for detection of the monomer. Suitable markers are described below.

[0390] CsgG-derived monomers can also be produced using D-amino acids. For example, CsgG-derived monomers can contain a mixture of L-amino acids and D-amino acids. This is routine in the art of producing such proteins or peptides.

[0391] The CsgG-derived monomer contains one or more specific modifications to facilitate nucleotide discrimination. The CsgG-derived monomer may also contain other nonspecific modifications, as long as they do not interfere with pore formation. Many nonspecific side chain modifications are known in the art and can be made to the side chains of the CsgG-derived monomer. Such modifications include, for example, reductive alkylation of amino acids by reaction with an aldehyde followed by reduction with NaBH4, amidation with methyl iminoacetate, or acylation with acetic anhydride.

[0392] Monomers derived from CsgG can be produced using standard methods known in the art. Monomers derived from CsgG can be prepared synthetically or recombinantly. For example, monomers can be synthesized by in vitro translation and transcription (IVTT). Suitable methods for producing pores and monomers are discussed in International Applications WO 2010 / 004273, WO 2010 / 004265, or WO 2010 / 086603. Methods for inserting pores into membranes are known.

[0393] Two or more CsgG monomers in a pore may be covalently attached to each other. For example, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10 monomers may be covalently attached. The covalently attached monomers may be the same or different.

[0394] The monomers may optionally be fused genetically via a linker, or chemically, for example, via a chemical cross-linker. Methods for covalently attaching monomers are disclosed in WO2017 / 149316, WO2017 / 149317 and WO2017 / 149318.

[0395] In some embodiments, the mutant monomer is chemically modified. The mutant monomer can be chemically modified at any site in any way. The mutant monomer is preferably chemically modified by attachment of a molecule to one or more cysteines (cysteine ​​connection), attachment of a molecule to one or more lysines, attachment of a molecule to one or more non-natural amino acids, enzyme modification of an epitope or modification of a terminal. Suitable methods for carrying out such modifications are well known in the art. The mutant monomer can be chemically modified by attachment of any molecule. For example, the mutant monomer can be chemically modified by attachment of a dye or fluorophore.

[0396] In some embodiments, the mutant monomers are chemically modified with molecular adapters that facilitate interaction between the pore comprising the monomer and the target nucleotide or polynucleotide sequence. The presence of the adapter improves the host-guest chemistry between the pore and the nucleotide or polynucleotide sequence, thereby improving the sequencing capacity of the pore formed by the mutant monomer. The principles of host-guest chemistry are well known in the art. The adapter has an effect on the physical or chemical properties of the pore that improves the interaction between the pore and the nucleotide or polynucleotide sequence. The adapter can alter the charge of the barrel or channel of the pore, or specifically interact or bind to the nucleotide or polynucleotide sequence, thereby promoting its interaction with the pore.

[0397] The molecular adaptor is preferably a cyclic molecule, a cyclodextrin, a substance capable of hybridization, a DNA binder or intercalator, a peptide or peptide analog, a synthetic polymer, an aromatic planar molecule, a positively charged small molecule or a small molecule capable of hydrogen bonding.

[0398] The adaptor can be circular. Circular adaptors preferably have the same symmetry as the pore. The adaptor preferably has eight-fold or nine-fold symmetry, as CsgG typically has eight or nine subunits around a central axis. This will be discussed in more detail below.

[0399] Adaptors typically interact with nucleotide or polynucleotide sequences through host-guest chemistry. Adaptors are typically capable of interacting with nucleotide or polynucleotide sequences. Adaptors comprise one or more chemical groups capable of interacting with nucleotide or polynucleotide sequences. The one or more chemical groups preferably interact with nucleotide or polynucleotide sequences through non-covalent interactions, such as hydrophobic interactions, hydrogen bonding, van der Waals forces, π-cation interactions and / or electrostatic forces. The one or more chemical groups capable of interacting with nucleotide or polynucleotide sequences are preferably positively charged. The one or more chemical groups capable of interacting with nucleotide or polynucleotide sequences more preferably comprise amino groups. The amino groups can be attached to primary, secondary or tertiary carbon atoms. Adaptors even more preferably comprise an amino ring, such as a ring of 6, 7 or 8 amino groups. Adaptors most preferably comprise a ring of eight amino groups. The protonated amino ring can interact with negatively charged phosphate groups in the nucleotide or polynucleotide sequence.

[0400] The correct positioning of the adapter in the hole can be promoted by host-guest chemistry between the adapter and the hole containing the mutant monomer. The adapter preferably comprises one or more chemical groups that can interact with one or more amino acids in the hole. The adapter more preferably comprises one or more chemical groups that can interact with one or more amino acids in the hole through non-covalent interactions, such as hydrophobic interactions, hydrogen bonding, van der Waals forces, π-cation interactions and / or electrostatic forces. The chemical group that can interact with one or more amino acids in the hole is typically a hydroxyl or an amine. The hydroxyl group can be attached to a primary carbon, a secondary carbon, or a tertiary carbon atom. The hydroxyl group can form hydrogen bonds with uncharged amino acids in the hole. Any adapter that promotes the interaction between the hole and the nucleotide or polynucleotide sequence can be used.

[0401] Suitable adapters include, but are not limited to, cyclodextrins, cyclic peptides, and cucurbiturils. The adapter is preferably cyclodextrin or a derivative thereof. Cyclodextrin or a derivative thereof may be any of those disclosed in Eliseev, AV and Schneider, HJ. (1994) J. Am. Chem. Soc. 116, 6081-6088. The adapter is more preferably seven-6-amino-β-cyclodextrin (am7-βCD), 6-monodeoxy-6-monoamino-β-cyclodextrin (am1-βCD), or seven-(6-deoxy-6-guanidino)-cyclodextrin (gu7-βCD). The pKa of the guanidino group in gu7-βCD is much higher than that of the primary amine in am7-βCD, and therefore is more positively charged. This gu7-βCD adapter can be used to increase the residence time of the nucleotide in the pore, improve the accuracy of the measured residual current, and improve the base detection rate at high temperature or low data acquisition rate.

[0402] If the succinimidyl 3-(2-pyridyldithio)propionate (SPDP) cross-linker is used as discussed in more detail below, the linker is preferably heptakis(6-deoxy-6-amino)-6-N-mono(2-pyridyl)dithiopropionyl-β-cyclodextrin (am6amPDP1-βCD).

[0403] More suitable adaptors include γ-cyclodextrin, which contains nine sugar units (thus having nine-fold symmetry). γ-cyclodextrin can contain a linker molecule, or can be modified to contain all or more of the modified sugar units used in the β-cyclodextrin example discussed above.

[0404] The molecular adapter can be covalently attached to the mutant monomer. The adapter can be covalently attached to the hole using any method known in the art. The adapter is usually attached by chemical connection. If the molecular adapter is attached by cysteine ​​connection, it is preferably introduced into the mutant by substitution, for example, by introducing into the barrel. The mutant monomer can be chemically modified by the attachment of the molecular adapter to one or more cysteines in the mutant monomer. The one or more cysteines can be naturally occurring, i.e., at position 1 and / or 215 in SEQ ID NO:3. Alternatively, the mutant monomer can be chemically modified by the attachment of one or more cysteines introduced by the molecule to other positions. The cysteine ​​at position 215 can be removed, for example, by substitution, to ensure that the molecular adapter is not attached to this position, but to the cysteine ​​at position 1 or the cysteine ​​introduced at another position.

[0405] The reactivity of cysteine ​​residues can be enhanced by modifying adjacent residues. For example, the basic groups of flanking arginine, histidine, or lysine residues will shift the pKa of the cysteine ​​thiol group to the more reactive S - The pKa of the group. The reactivity of cysteine ​​residues can be protected by thiol protecting groups such as dTNB. They can be reacted with one or more cysteine ​​residues of the mutant monomer before linker attachment.

[0406] The molecule may be attached directly to the mutant monomer. Preferably, the molecule is attached to the mutant monomer using a linker, such as a chemical cross-linker or a peptide linker.

[0407] Suitable chemical cross-linking agents are well known in the art. Preferred cross-linking agents include 2,5-dioxypyrrolidin-1-yl 3-(pyridin-2-yl disulfanyl) propionate, 2,5-dioxypyrrolidin-1-yl 4-(pyridin-2-yl disulfanyl) butyrate, and 2,5-dioxypyrrolidin-1-yl 8-(pyridin-2-yl disulfanyl) octanoate. The most preferred cross-linking agent is succinimidyl-3-(2-pyridyl dithio) propionate (SPDP). Typically, the molecule is covalently attached to a bifunctional cross-linking agent before the molecule / cross-linking agent complex is covalently attached to the mutant monomer, but the bifunctional cross-linking agent can also be covalently attached to the monomer before the bifunctional cross-linking agent / monomer complex is attached to the molecule.

[0408] The linker is preferably resistant to dithiothreitol (DTT). Suitable linkers include, but are not limited to, iodoacetamide-based and maleimide-based linkers.

[0409] In other embodiments, the monomers can be attached to polynucleotide binding proteins. This forms a modular sequencing system that can be used in the sequencing methods of the present invention. Polynucleotide binding proteins are discussed below.

[0410] The polynucleotide binding protein is preferably covalently attached to the mutant monomer. The protein can be covalently attached to the monomer using any method known in the art. The monomer and protein can be chemically fused or genetically fused. If the entire construct is expressed from a single polynucleotide sequence, the monomer and protein are genetically fused. Genetic fusion of monomers to polynucleotide binding proteins is discussed in WO 2010 / 004265.

[0411] If the polynucleotide binding protein is attached via a cysteine ​​linkage, the one or more cysteines are preferably introduced into the mutant by substitution. Preferably, the one or more cysteines are introduced into a loop region that has low conservation in homologs, indicating that mutations or insertions can be tolerated. Therefore, they are suitable for attachment to polynucleotide binding proteins. In such embodiments, the naturally occurring cysteine ​​at position 251 can be removed. As described above, the reactivity of cysteine ​​residues can be enhanced by modification.

[0412] The polynucleotide binding protein can be attached directly to the mutant monomer or through one or more linkers. A hybrid linker as described in WO 2010 / 086602 can be used to attach the molecule to the mutant monomer. Alternatively, a peptide linker can be used. A peptide linker is an amino acid sequence. The length, flexibility and hydrophilicity of the peptide linker are generally designed so that it does not disrupt the function of the monomer and molecule. Preferred flexible peptide linkers are stretches of 2 to 20, such as 4, 6, 8, 10 or 16 serine and / or glycine amino acids. More preferred flexible linkers include (SG)1, (SG)2, (SG)3, (SG)4, (SG)5 and (SG)8, wherein S is serine and G is glycine. Preferred rigid peptide linkers are stretches of 2 to 30, such as 4, 6, 8, 16 or 24 proline amino acids. More preferred rigid linkers include (P) 12 , where P is proline.

[0413] Chemical modification

[0414] Mutant CsgG monomers or CsgF peptides can be chemically modified using molecular adaptors and polynucleotide binding proteins.

[0415] The molecule with which the monomer or peptide is chemically modified may be attached directly to the monomer or peptide, or via a linker, as disclosed in WO 2010 / 004273, WO 2010 / 004265 or WO 2010 / 086603.

[0416] Any of the proteins described herein, such as the CsgG monomer and / or CsgF peptide, can be modified to facilitate their identification or purification, for example by adding a histidine residue (his tag), an aspartic acid residue (asp tag), a streptavidin tag, a flag tag, a SUMO tag, a GST tag, or an MBP tag, or by adding a signal sequence to promote secretion from cells in which the polypeptide does not naturally contain such a sequence. An alternative to introducing a genetic tag is to chemically react the tag with a native or engineered position on the protein. An example of this is reacting a gel shift reagent with an engineered cysteine ​​residue on the outside of the protein. This has been demonstrated to be a method for isolating hemolysin hetero-oligomers (Chem Biol. 1997 Jul;4(7):497-505).

[0417] Any of the proteins described herein, such as CsgG monomers and / or CsgF peptides, may be labeled with a recognizable marker. The recognizable marker may be any suitable marker that allows for detection of the protein. Suitable markers include, but are not limited to, fluorescent molecules, radioisotopes (e.g., 125 I. 35 S), enzymes, antibodies, antigens, polynucleotides, and ligands (such as biotin).

[0418] Any of the proteins described herein, such as CsgG monomers and / or CsgF peptides, can be produced synthetically or recombinantly. For example, the protein can be synthesized by in vitro translation and transcription (IVTT). The amino acid sequence of the protein can be modified to include non-naturally occurring amino acids or to increase the stability of the protein. When the protein is produced synthetically, such amino acids can be introduced during production. The protein can also be altered after synthesis or recombinant production.

[0419] Proteins can also be produced using D-amino acids. For example, a protein can contain a mixture of L-amino acids and D-amino acids. This is conventional in the art of producing such proteins or peptides.

[0420] Proteins may also contain other nonspecific modifications, as long as they do not interfere with the function of the protein. Many nonspecific side chain modifications are known in the art and can produce nonspecific side chain modifications on the side chains of proteins. Such modifications include, for example, reductive alkylation of amino acids by reaction with aldehydes followed by reduction with NaBH4, amidation with imidoacetic acid methyl ester, or acylation with acetic anhydride.

[0421] Any of the proteins described herein, such as CsgG monomers and / or CsgF peptides, can be produced using standard methods known in the art. Polynucleotide sequences encoding proteins can be obtained and replicated using standard methods in the art. Polynucleotide sequences encoding proteins can be expressed in bacterial host cells using standard techniques in the art. Proteins can be produced in cells by expressing the polypeptide in situ from a recombinant expression vector. The expression vector optionally carries an inducible promoter to control expression of the polypeptide. These methods are described in Sambrook, J. and Russell, D. (2001). Molecular Cloning: A Laboratory Manual, 3rd ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY.

[0422] Proteins can be produced on a large scale after purification from the organism that produced the protein or after recombinant expression by any protein liquid chromatography system. Typical protein liquid chromatography systems include FPLC, AKTA systems, Bio-Cad systems, Bio-Rad BioLogic systems, and Gilson HPLC systems.

[0423] Method for generating holes

[0424] In a third aspect, the present invention provides methods for generating CsgG:modified CsgF pore complexes possessing two or more constriction sites in vivo and in vitro. One embodiment provides a method for generating a transmembrane pore complex comprising a CsgG pore, or a homolog or mutant thereof, and a modified CsgF peptide, or a homolog or mutant thereof, by co-expression. The method comprises the steps of expressing a CsgG monomer (expressed as the proprotein provided in SEQ ID NO: 2, or a homolog or mutant thereof) and a modified or truncated CsgF monomer in a suitable host cell, thereby allowing for in vivo formation of the complex pore. The complex comprises the modified CsgF peptide complexed with the CsgG pore to provide an additional read head for the pore. The resulting pore complex generated by the method using the modified CsgF peptide provides a structure sufficient for characterizing a target analyte (such as nucleic acid sequencing) using the pore complex, as it allows passage of the analyte, particularly a polynucleotide chain, and comprises two or more read heads for improved reading of the polynucleotide sequence when used in an appropriate environment for the application.

[0425] More particularly, the modified CsgF peptide expressed in the method comprises the proprotein depicted in SEQ ID NO: 8, 10, 12 or 14, or a homologue thereof. Those sequences limit the method to those CsgF fragments that are capable of introducing a constriction site in the pore complex and binding to the CsgG protein pore to obtain a biological pore.

[0426] Another method for producing isolated pore complexes formed by CsgG and CsgF proteins, etc., involves in vitro reconstitution of the monomers to obtain a functional pore. The method comprises contacting a mature CsgG monomer depicted in SEQ ID NO:3, or a homologue or mutant thereof, with a modified CsgF peptide, or a homologue or mutant thereof, in a suitable system to allow complex formation. The system can be an "in vitro" system, which refers to a system comprising at least the necessary components and environment for performing the method, and utilizing biomolecules, organisms, cells (or portions of cells) outside of their normal, naturally occurring environment, allowing for more detailed, convenient, or efficient analysis than can be performed using whole organisms. The in vitro system can also comprise a suitable buffer composition provided in a test tube to which the protein components forming the complex have been added. Options for providing such systems are known to those skilled in the art. In particular embodiments, the modified CsgF peptide, etc. used in the method for in vitro reconstitution is a peptide comprising SEQ ID NO:15 or SEQ ID NO:16, or a mutant or homologue thereof, which can be produced synthetically or recombinantly. Alternatively, in the method, a modified CsgF peptide comprising SEQ ID NO: 40, 39, 38 or 37, 15, 54, 55, or a homologue or mutant thereof is provided for contacting with a CsgG or CsgG-like pore to generate a pore complex.

[0427] CsgG / CsgF pores can be prepared by any suitable method. Examples of such suitable methods are described.

[0428] In one embodiment, the CsgG / CsgF pore can be produced by co-expression. In this embodiment, at least one gene encoding a CsgG monomer polypeptide (which may be a mutant polypeptide) in one vector and a gene encoding at least one full-length or truncated CsgF polypeptide (which may be a mutant polypeptide) in a second vector can be co-expressed to express the proteins and produce a complex in the transformed cells. This can be done in vivo or in vitro. Alternatively, the two genes encoding the CsgG and CsgF polypeptides can be placed in a single vector under the control of a single promoter or under the control of two separate promoters, which may be the same or different.

[0429] In another embodiment, a CsgG / CsgF pore is produced by separately expressing a CsgG monomer and a CsgF peptide. CsgG monomers or CsgG pores can be purified from cells transformed with a vector encoding at least one CsgG monomer or with more than one vector each expressing a CsgG monomer. CsgF peptides can be purified from cells transformed with a vector encoding at least one CsgF peptide. The purified CsgG monomer / pore can then be incubated with a CsgF peptide to produce a pore complex.

[0430] In another embodiment, CsgG monomers and / or CsgF peptides are produced separately by in vitro translation and transcription (IVTT). The CsgG monomers can then be incubated with the CsgF peptide to prepare the pore complex. Figure 14 The use of this method is described in .

[0431] The above embodiments can be combined, such that, for example, (i) CsgG is produced in vivo and CsgF is produced in vivo; (ii) CsgG is produced in vitro and CsgF is produced in vivo; (iii) CsgG is produced in vivo and CsgF is produced in vitro; (iv) CsgG is produced in vitro and CsgF is produced in vitro.

[0432] One or both of the CsgG monomer and the CsgF peptide can be labeled to facilitate purification. Purification can also be performed when the CsgG monomer and / or the CsgF peptide are unlabeled. Methods known in the art (e.g., ion exchange, gel filtration, hydrophobic interaction column chromatography, etc.) can be used alone or in various combinations to purify pore components.

[0433] Any known tag can be used for either protein. In one embodiment, a dual-tag purification method can be used to purify the CsgG:CsgF complex from both the CsgG pore and CsgF. For example, a Strep tag can be used for CsgG and a His tag for CsgF, or vice versa. Figure 13 This is exemplified in Figure 2. A similar end result is achieved when the two proteins are purified separately and mixed together, followed by another round of Strep and His purification.

[0434] When the full-length CsgF protein forms a complex with CsgG, the neck and head domains of CsgF ( Figure 4B ) (shown in red box) protrudes from the β-barrel of the CsgG pore.

[0435] Therefore, if a pore containing both the CsgG pore and full-length CsgF is used in single-channel recording experiments, the head domains may hinder or prevent the pore from inserting into the membrane. They may also block the path of the analyte in the pore. Therefore, when inserting the pore into the membrane, it is desirable to minimize the number of flexible polypeptides that dangle from the β-barrel. Here, we provide a truncated form of the CsgF protein that mimics the FCP region resolved in the cryo-EM structure of the complex and maintains structural integrity.

[0436] The CsgG / CsgF pore can be prepared before insertion into the membrane or after the CsgG pore is inserted into the membrane. When the pore complex is prepared before insertion into the membrane, a truncation mutant is preferably used. However, the CsgG pore can be inserted into the membrane and then the CsgF peptide is added, allowing the CsgG and CsgF complex to form in situ. For example, in one embodiment of a system in which the reverse side of the membrane is accessible (e.g., in a chip or chamber for electrophysiological measurements), the CsgG pore can be inserted into the membrane and then the CsgF peptide can be added from the reverse side of the membrane, allowing the complex to form in situ. In any embodiment in which the CsgG pore is formed in situ, a larger CsgF peptide can be used. For example, the CsgF peptide can comprise all or part of the neck domain of CsgF (starting approximately at residue 36 of SEQ ID NO: 6). In some embodiments, CsgF can comprise the entire neck domain and part of the head domain (residues 36 to XX of SEQ ID NO: 6).

[0437] CsgG:CsgF and CsgG:FCP complexes can be prepared using different methods depending on the method used to prepare the complex and the stability of the complex with a specific truncation.

[0438] In one embodiment, a truncated form of a CsgF polypeptide of the desired length is used directly.

[0439] Another embodiment uses a full-length CsgF polypeptide or a polypeptide that is longer than the desired truncation (sufficient to maintain complex stability) into which a protease cleavage site (e.g., TEV, HRV 3, or any other protease cleavage site) is inserted, such that a CsgF peptide of the desired length is generated by protease cleavage. In this embodiment, once the CsgG / CsgF complex is formed, a protease is used to cleave CsgF at the desired site. Alternatively, a protease can be used to generate the CsgF peptide prior to complex assembly.

[0440] Some protease sites leave additional tags after cleavage. For example, the TEV protease cleavage sequence is ENLYFQS. TEV protease cleaves the protein between Q and S, leaving ENLYFQ intact at the C-terminus of the CsgF peptide. Figure 15 shows an example of using a modified CsgF containing a TEV cleavage site, followed by cleavage of the modified CsgF using TEV protease after complex formation.

[0441] As another example, the HRV C3 cleavage site is LEVLFQGP, and the enzyme cleaves between the Q and G, leaving LEVLFQ intact at the C-terminus of the CsgF peptide.

[0442] Methods for characterizing analytes

[0443] In another aspect, the present invention provides a method for determining the presence, absence, or one or more characteristics of a target analyte. The method comprises contacting the target analyte with an isolated pore complex or a transmembrane pore (such as a pore of the present invention), allowing the target analyte to move relative to the pore channel, e.g., into or through the pore channel, and performing one or more measurements as the analyte moves relative to the pore, thereby determining the presence, absence, or one or more characteristics of the analyte. The target analyte may also be referred to as a template analyte or a target analyte. The isolated pore complex typically comprises at least 7, at least 8, at least 9, or at least 10 monomers, such as 7, 8, 9, or 10 CsgG monomers. The isolated pore complex preferably comprises eight or nine identical CsgG monomers. Preferably, one or more, such as 2, 3, 4, 5, 6, 7, 8, 9, or 10, CsgG monomers are chemically modified, or the CsgF peptide is chemically modified. The isolated pore complex monomers, such as CsgG monomers, or homologs or mutants thereof, and the modified CsgF monomers, or homologs or mutants thereof, can be derived from any organism. The analyte may first pass through the CsgG constriction and then through the CsgF constriction. In an alternative embodiment, depending on the orientation of the CsgG / CsgF complex in the membrane, the analyte may first pass through the CsgF constriction and then through the CsgG constriction.

[0444] The method is used to determine the presence, absence, or one or more characteristics of a target analyte. The method can be used to determine the presence, absence, or one or more characteristics of at least one analyte. The method can involve determining the presence, absence, or one or more characteristics of two or more analytes. The method can include determining the presence, absence, or one or more characteristics of any number of analytes, such as 2, 5, 10, 15, 20, 30, 40, 50, 100 or more analytes. Any number of characteristics of the one or more analytes can be determined, such as 1, 2, 3, 4, 5, 10 or more characteristics.

[0445] The binding of molecules in the channel of a pore complex or near any of its openings will affect the ion flux through the open channel of the pore, which is the basis of pore channel "molecular sensing". In a manner similar to nucleic acid sequencing applications, changes in the ion flux through the open channel can be measured by changes in current using appropriate measurement techniques (e.g., WO 2000 / 28312 and D. Stoddart et al., Proc. Natl. Acad. Sci., 2010, 106, 7702-7 or WO 2009 / 077734). The degree of reduction in ion flux, as measured by the reduction in current, is related to the size of the obstruction in or near the pore. Therefore, the binding of a target molecule (also called an "analyte") in or near the pore provides a detectable and measurable event, thus forming the basis of a "biosensor". Molecules suitable for nanopore sensing include nucleic acids; proteins; peptides; polysaccharides and small molecules (here refers to low molecular weight (e.g., <900 Da or <500 Da) organic or inorganic compounds) such as drugs, toxins, cytokines and pollutants. Detecting the presence of biomolecules has applications in personalized drug development, medicine, diagnostics, life science research, environmental monitoring, and the security and / or defense industries.

[0446] Alternatively, isolated pore complexes or transmembrane pore complexes containing a wild-type or modified E. coli CsgG nanopore, or a homolog or mutant thereof, and a modified CsgF peptide that provides a channel constriction for the pore in the complex can be used as molecular or biological sensors. In some embodiments, the CsgG nanopore can be derived from or isolated from a bacterial protein (e.g., E. coli, Salmonella typhi). In some embodiments, the CsgG nanopore can be recombinantly produced. Procedures for analyte detection are described in Howorka et al., Nature Biotechnology (2012) Jun 7;30(6):506-7. The analyte molecule to be detected can be bound to either side of the channel or within the lumen of the channel itself. The location of binding can be determined by the size of the molecule to be sensed.

[0447] The target analyte is preferably a metal ion, an inorganic salt, a polymer, an amino acid, a peptide, a polypeptide, a protein, a nucleotide, an oligonucleotide, a polynucleotide, a polysaccharide, a dye, a bleach, a drug, a diagnostic agent, a recreational drug, an explosive material, a toxic compound or an environmental pollutant. The method can involve determining the presence, absence or one or more characteristics of two or more analytes of the same type (such as two or more proteins, two or more nucleotides or two or more drugs). Alternatively, the method can involve determining the presence, absence or one or more characteristics of two or more different types of analytes (such as one or more proteins, one or more nucleotides or one or more drugs).

[0448] The target analyte may be secreted from the cell. Alternatively, the target analyte may be an analyte present inside the cell, such that the analyte must be extracted from the cell before the method can be performed.

[0449] Wild-type pores can act as sensors, but are often modified by recombinant or chemical methods to increase the binding strength, binding location, or binding specificity of the molecule to be sensed. Typical modifications include the addition of a specific binding moiety that is structurally complementary to the molecule to be sensed. Where the analyte molecule comprises a nucleic acid, this binding moiety may comprise a cyclodextrin or oligonucleotide; for small molecules, this may be a known complementary binding region, such as the antigen-binding portion of an antibody or non-antibody molecule, including a single-chain variable fragment (scFv) region or an antigen recognition domain from a T-cell receptor (TCR); or for proteins, it may be a known ligand for the target protein. In this way, wild-type or modified E. coli CsgG nanopores or homologs thereof can be made to function as molecular sensors for detecting the presence of suitable antigens (including epitopes) in a sample. The antigens may include cell surface antigens (including receptors, markers of solid tumors or hematological cancer cells (e.g., lymphomas or leukemias)), viral antigens, bacterial antigens, protozoan antigens, allergens, allergy-related molecules, albumins (e.g., human, rodent, or bovine), fluorescent molecules (including fluorescein), blood group antigens, small molecules, drugs, enzymes, catalytic sites of enzymes or enzyme substrates, and transition state analogs of enzyme substrates. As described above, modifications can be achieved using known genetic engineering and recombinant DNA techniques. The location of any adaptation will depend on the properties of the molecule to be sensed, such as size, three-dimensional structure, and its biochemical properties. The selection of the adapted structure can utilize computational structure design. Techniques such as surface plasmon resonance detection of molecular interactions can be used. (BIAcore, Inc., Piscataway, NJ; see also www.biacore.com) for the determination and optimization of protein-protein or protein-small molecule interactions.

[0450] In one embodiment, the analyte is an amino acid, peptide, polypeptide, or protein. The amino acid, peptide, polypeptide, or protein may be naturally occurring or non-naturally occurring. The polypeptide or protein may contain synthetic or modified amino acids. Several different types of amino acid modifications are known in the art. Suitable amino acids and modifications thereof are described above. It should be understood that the target analyte may be modified by any method available in the art.

[0451] In another embodiment, the analyte is a polynucleotide defined as a macromolecule comprising two or more nucleotides, such as a nucleic acid. Nucleic acids are particularly suitable for nanopore sequencing. Naturally occurring nucleic acid bases in DNA and RNA can be distinguished by their actual size. When a nucleic acid molecule or a single base passes through the channel of a nanopore, the size difference between the bases causes the ion flow through the channel to be directly related to be reduced. Changes in ion flow can be recorded. Electrical measurement techniques suitable for recording changes in ion flow are described in, for example, WO 2000 / 28312 and D. Stoddart et al., Proc. Natl. Acad. Sci., 2010, 106, pp. 7702-7 (single-channel recording equipment); and, for example, in WO 2009 / 077734 (multi-channel recording technology). By suitable calibration, the characteristic reduction of ion flow can be used to identify specific nucleotides and related bases passing through the channel in real time. In typical nanopore nucleic acid sequencing, due to the nucleotide partially blocking the channel, when the single nucleotides of the target nucleic acid sequence pass through the channel of the nanopore in sequence, the open channel ion flow is reduced. This is a reduction in this ion flow measured using the appropriate recording technique described above. The reduction in ion flow can be calibrated to a reduction in ion flow measured for known nucleotides passing through the channel, thereby obtaining a method for determining the nucleotides passing through the channel, and therefore, when performed sequentially, obtaining a way to determine the nucleotide sequence of the nucleic acid passing through the nanopore. In order to accurately determine a single nucleotide, it is generally necessary to directly correlate the reduction in ion flow through the channel with the size of a single nucleotide passing through the contraction (or "reading head"). It should be appreciated that, for example, a complete nucleic acid polymer can be sequenced, and the complete nucleic acid polymer is "passed through" the hole by the action of a related polymerase. Alternatively, the sequence can be determined by passing nucleotide triphosphate bases that have been sequentially removed from the target nucleic acid near the hole (see, for example, WO 2014 / 187924).

[0452] A polynucleotide or nucleic acid may comprise any combination of nucleotides. Nucleotides may be naturally occurring or artificial. One or more nucleotides in a polynucleotide may be oxidized or methylated. One or more nucleotides in a polynucleotide may be damaged. For example, a polynucleotide may comprise a pyrimidine dimer. Such dimers are often associated with UV damage and are a major cause of cutaneous melanoma. One or more nucleotides in a polynucleotide may be modified, for example, with a marker or tag, suitable examples of which are known to those skilled in the art. A polynucleotide may comprise one or more spacers. Nucleotides typically contain a nucleobase, a sugar, and at least one phosphate group. The nucleobase and sugar form a nucleoside. The nucleobase is typically a heterocyclic ring. Nucleobases include, but are not limited to, purines and pyrimidines, and more specifically, adenine (A), guanine (G), thymine (T), uracil (U), and cytosine (C). The sugar is typically a pentose. Nucleotide sugars include, but are not limited to, ribose and deoxyribose. The sugar is preferably deoxyribose. Polynucleotides preferably comprise the following nucleosides: deoxyadenosine (dA), deoxyuridine (dU) and / or deoxythymidine (dT), deoxyguanosine (dG) and deoxycytidine (dC). Nucleotides are typically ribonucleotides or deoxyribonucleotides. Nucleotides typically contain monophosphate, diphosphate or triphosphate. Nucleotides may contain more than three phosphates, for example 4 or 5 phosphates. The phosphates may be attached to the 5' or 3' side of the nucleotide. The nucleotides in a polynucleotide may be attached to each other in any manner. As in nucleic acids, nucleotides are typically attached through sugar and phosphate groups. As in pyrimidine dimers, nucleotides may be linked by their core bases. Polynucleotides may be single-stranded or double-stranded. At least a portion of a polynucleotide is preferably double-stranded. Most preferably, a polynucleotide is ribonucleic acid (RNA) or deoxyribonucleic acid (DNA). In particular, the method using a polynucleotide as an analyte may alternatively comprise determining one or more characteristics selected from the group consisting of: (i) the length of the polynucleotide, (ii) the identity of the polynucleotide, (iii) the sequence of the polynucleotide, (iv) the secondary structure of the polynucleotide and (v) whether the polynucleotide is modified.

[0453] The polynucleotide can be of any length (i). For example, the length of the polynucleotide can be at least 10, at least 50, at least 100, at least 150, at least 200, at least 250, at least 300, at least 400 or at least 500 nucleotides or nucleotide pairs. The length of the polynucleotide can be 1000 or more nucleotides or nucleotide pairs, 5000 or more nucleotides or nucleotide pairs, or a length of 100,000 or more nucleotides or nucleotide pairs. Any number of polynucleotides can be studied. For example, the method can involve characterizing 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 50, 100 or more polynucleotides. If two or more polynucleotides are characterized, they can be different polynucleotides or two examples of the same polynucleotide. The polynucleotide can be naturally occurring or artificial. For example, the method can be used to test the sequence of the oligonucleotide manufactured. The method is usually carried out in vitro.

[0454] The nucleotides may have any identity (ii), including but not limited to adenosine monophosphate (AMP), guanosine monophosphate (GMP), thymidine monophosphate (TMP), uridine monophosphate (UMP), 5-methylcytidine monophosphate, 5-hydroxymethylcytidine monophosphate, cytidine monophosphate (CMP), cyclic adenosine monophosphate (cAMP), cyclic guanosine monophosphate (cGMP), deoxyadenosine monophosphate (dAMP), deoxyguanosine monophosphate (dGMP), deoxythymidine monophosphate (dTMP), deoxyuridine monophosphate (dUMP), deoxycytidine monophosphate (dCMP), and deoxymethylcytidine monophosphate. The nucleotides are preferably selected from AMP, TMP, GMP, CMP, UMP, dAMP, dTMP, dGMP, dCMP, and dUMP. The nucleotides may be abasic (i.e., lack a nucleobase). The nucleotides may also lack a nucleobase and a sugar (i.e., a C3 spacer). The sequence of nucleotides (iii) is determined by the consecutive identity of consecutive nucleotides attached to each other throughout the polynucleotide chain in the 5' to 3' direction of the chain.

[0455] Wells comprising a CsgG pore and a CsgF peptide are particularly useful for analyzing homopolymers. For example, the wells can be used to determine the sequence of a polynucleotide comprising two or more, for example at least 3, 4, 5, 6, 7, 8, 9, or 10, identical contiguous nucleotides. For example, the wells can be used to sequence polynucleotides comprising polyA, polyT, polyG, and / or polyC regions.

[0456] The constriction of the CsgG pore consists of residues at positions 51, 55, and 56 of SEQ ID NO:3. Read heads for CsgG and its constriction mutants are typically very sensitive. When DNA passes through the constriction, at any given time, approximately five bases of DNA interacting with the pore's read head dominate the current signal. While these more sensitive read heads are very good at reading mixed sequence regions of DNA (when A, T, G, and C are mixed), when homopolymer regions (e.g., polyT, polyG, polyA, polyC) are present in the DNA, the signal becomes flat and lacks information. Because these five bases dominate the signal for CsgG and its constriction mutants, it is difficult to distinguish photopolymers longer than five bases without using additional residence time information. However, if the DNA passes through a second read head, more DNA bases interact with the combined read head, increasing the length of homopolymers that can be distinguished. The Examples and Figures demonstrate this improvement in homopolymer sequencing accuracy achieved using a pore comprising the CsgG pore and the CsgF peptide.

[0457] Reagent test kit

[0458] On the other hand, the present invention also provides a kit for characterizing a target polynucleotide. The kit comprises an isolated pore complex according to the present invention, and components of a membrane or insulating layer. The membrane is preferably formed by the components. The isolated pore complex is preferably present in a membrane or insulating layer, together forming a transmembrane pore complex channel. The kit may comprise any type of membrane, such as components of an amphiphilic layer or a triblock copolymer membrane. The kit may also comprise a polynucleotide binding protein. The kit may also contain one or more anchors for coupling the polynucleotide to the membrane. The kit may further comprise one or more other reagents or instruments that enable any of the above-mentioned embodiments to be implemented. Such reagents or instruments include one or more of the following: a suitable buffer (aqueous solution), a device for obtaining a sample from a subject (such as a container or instrument containing a needle), a device for amplifying and / or expressing the polynucleotide, or a voltage clamp or patch clamp device. The reagents may be present in the kit in a dry state so that the liquid sample resuspends the reagents. The kit may also optionally comprise instructions enabling the kit to be used in the method of the present invention or details about which organisms the method can be used for. Finally, the kit may also comprise additional components that can be used for peptide characterization.

[0459] In one embodiment, the isolated pore complex or transmembrane pore complex, as provided herein, is used for nucleic acid sequencing. For such applications, Phi29 DNA polymerase (DNAP) can be used as a molecular motor, with the CsgG:CsgF nanopore complex located within the membrane, to allow controlled movement of oligomeric probe DNA strands through the pore. A voltage can be applied across the pore, and an electric current is generated due to the movement of ions in the saline solution on either side of the nanopore. As the probe DNA moves through the pore, the ion flux through the pore changes relative to the DNA. This information has been shown to be sequence-dependent and allows accurate reading of the probe sequence from current measurements.

[0460] It should be understood that although specific embodiments, specific configurations, and materials and / or molecules have been discussed herein with respect to the engineered cells and methods according to the present invention, various changes or modifications may be made in form and detail without departing from the scope and spirit of the present invention. The following examples are provided to better illustrate specific embodiments and should not be construed as limiting the present application. The present application is limited only by the claims.

[0461] Example

[0462] introduce

[0463] The CsgG pore is part of the type VIII multicomponent secretion system (also known as the curli biosynthesis system), which is responsible for the formation of aggregation fibers called curli in Escherichia coli. Curli are extracellular protein fibers that are primarily involved in bacterial biofilm formation and attachment to abiotic surfaces. Curli biosynthesis is directed by two operons in E. coli, csgBAC and csgDEFG (curli-specific genes) (Hammar et al., 1995). Secretion of the curli subunits CsgA and CsgB depends on CsgG, a specialized lipoprotein found to form oligomeric secretory channels in the outer membrane. For transport, CsgG cooperates with the periplasmic and extracellular accessory proteins CsgE and CsgF. CsgE forms the specificity factor for CsgG-mediated transport, while CsgF appears to couple the secretion of CsgA with its aggregation into extracellular fibers, templated by CsgB.

[0464] The crystal structure of the CsgG secretion channel demonstrates that CsgG forms a and height The nonameric transport complex that transports the OM through the inner diameter The periplasmic domain of the channel is separated from the transmembrane β-barrel by an iris-like septum, forming a 36-stranded β-barrel (Goyal et al., 2014; Figure 1). The septum is formed by a conserved 12-residue "constriction loop" (CL) found in each subunit, which organizes concentrically in the CsgG oligomer to form an orifice with a diameter of approximately 0.6 nm and a height of 1.5 nm, excluding solvent (Figure 1). When acting as a protein secretion channel, this orifice, or constriction, in the CsgG channel forms the primary site of interaction with the translocating polypeptide. When CsgG is used as a nanopore sensing platform, this orifice acts as the primary read head for analytes that reside in or pass through the channel. The diameter of the orifice and its physical and chemical properties can be altered by amino acid substitutions, deletions, or insertions in the protein region corresponding to residues 46 to 61 of SEQ ID NO: 3 ( Figure 1D ). In particular, mutations at positions 51, 55 and 56 according to SEQ ID NO: 3, together or individually, have a beneficial effect on the conductance characteristics of the nanopore and its interaction with analytes, including polynucleotides.

[0465] The assembly factor CsgF represents a component of the curli secretion machinery. The CsgF preprotein reaches the periplasm via the SEC pathway, after which mature CsgF (12.9 kDa) is found as a surface-exposed protein in a CsgG-dependent manner. In the presence of CsgG, CsgF co-isolates with the OM, and co-immunoprecipitation experiments demonstrate direct contact between the two proteins. The available data demonstrate that CsgF is not essential for productive subunit secretion, but rather suggests that this protein forms a coupling factor between CsgA secretion and extracellular polymerization into curli fibers by coordinating or chaperoning the nucleation function of the CsgB subunit.

[0466] Example 1: Production of CsgG:CsgF Complex Protein (Co-expression, in vitro reconstitution, in vitro coupled transcription and translation, and reconstitution of CsgG and CsgF synthetic peptides)

[0467] To produce the CsgG:CsgF complex, the two proteins can be co-expressed in a suitable Gram-negative host, such as Escherichia coli, and extracted and purified as a complex from the outer membrane. In vivo formation of the CsgG pore and the CsgG:CsgF complex requires targeting the proteins to the outer membrane. To this end, CsgG is expressed as a prepro-protein with a lipoprotein signal peptide (Juncker et al. 2003, Protein Sci. 12(8):1652-62) and a Cys residue at the N-terminus of the mature protein (SEQ ID NO:3). An example of such a lipoprotein signal peptide is residues 1-15 of full-length E. coli CsgG as shown in SEQ ID NO:2. Processing of the prepro-CsgG results in cleavage of the signal peptide and lipidation of the mature CsgG, followed by translocation of the mature lipoprotein to the outer membrane, where it is inserted as an oligomeric pore (Goyal et al. 2014, Nature 516(7530):250-3). To form a CsgG:CsgF complex, CsgF can be co-expressed with CsgG and targeted to the periplasm with the aid of a leader sequence, such as the native signal peptide corresponding to residues 1-19 of SEQ ID NO: 5. The CsgG:CsgF combined pore can then be extracted from the outer membrane using detergents and purified to a homogeneous complex by chromatography ( FIG2 ).

[0468] Alternatively, the CsgG:CsgF pore complex can be generated by in vitro reconstitution using the CsgG pore and CsgF—see below and Figure 3.

[0469] For the in vivo formation of the CsgG:CsgF complex in the example shown in FIG2 , E. coli CsgF (SEQ ID NO: 5) and CsgG (SEQ ID NO: 2) were co-expressed using their native signal peptides to ensure periplasmic targeting of both proteins, and the N-terminus of CsgG was lipidated. Additionally, to facilitate purification, CsgF was modified by introducing a C-terminal 6x histidine tag, and the CsgG C-terminus was fused to a Strep-II tag. Co-expression and complex purification were performed as described in the Methods. SDS-PAGE analysis of the His affinity purification eluate revealed enrichment of CsgF-His and co-purification of CsgG-Strep, indicating that the latter is complexed with CsgF ( Figure 2B In addition, SDS-PAGE revealed that a large portion of the eluted CsgF ran at a lower molecular weight due to the loss of the N-terminal fragment of the protein ( Figure 2B , indicated by asterisks). SDS-PAGE analysis of pooled fractions of the His-trap eluate from the second affinity purification revealed the presence of CsgG and CsgF in apparently equimolar concentrations, as well as the loss of the CsgF truncated fragment seen in the His-trap eluate ( Figure 2BThe coelution of CsgF in Strep affinity purification suggests that the protein exists as a non-covalent complex with CsgG. Surprisingly, an N-terminal truncated fragment of CsgF was lost in Strep affinity purification, indicating that the CsgF N-terminus is required for binding to CsgG ( Figure 2B ).

[0470] Figure 13 Another example of the formation of a CsgG:CsgF complex by in vivo co-expression is shown. In this example, the CsgG protein was modified with a C-terminal Strep-II tag, while the full-length CsgF protein was modified with a C-terminal 10X histidine tag. The co-expressed CsgG:CsgF complex was purified from its constituent components by Strep-tag purification followed by histidine tag purification, as described in the Materials and Methods section for analyte characterization. Due to the difference in molecular weight, the CsgG:CsgF complex could be clearly distinguished from the CsgG pore in SDS-PAGE analysis ( Figure 13 A). Figure 13 As shown in Figure .B, two tag purification methods can be successfully applied to purify the CsgG:CsgF complex from its component parts.

[0471] To generate the CsgG:CsgF complex by in vitro reconstitution, CsgG and CsgF were expressed in separate E. coli cultures transformed with pPG1 and pNA101, respectively, and purified, followed by in vitro reconstitution of the CsgG:CsgF complex (see Methods). For comparison, purified CsgG was run similarly to the complex on a Superose 6 column. The CsgG Superose 6 run showed the presence of two pores corresponding to the nonameric CsgG pore ( Figure 3A (a) and 3C) and dimers of the nonameric CsgG pore ( Figure 3A (b) and 3C) as previously described in Goyal et al. (2014). Superose 6 runs of CsgG:CsgF reconstitution revealed the presence of three corresponding to excess CsgF ( Figure 3A (c)), nonameric CsgG:CsgF complex ( Figure 3A (d)) and the dimer of nonameric CsgG:CsgF ( Figure 3A To provide independent confirmation of CsgG:CsgF complex formation, individual Superose 6 elution peaks were analyzed on native PAGE ( Figure 3B ).

[0472] Surprisingly, the CsgG:CsgF complex can also be prepared by an in vitro transcription and translation (IVTT) method, as described in the Materials and Methods section for analyte characterization. The complex can be prepared by expressing CsgG and CsgF proteins in the same IVTT reaction or by reconstituting CsgG and CsgF separately in two different IVTT reactions. Figure 14 In the example shown, a CsgG:CsgF complex was prepared in a reaction mixture using the E. coli T7-S30 circular DNA extraction system (Promega), and the proteins were analyzed on SDS-PAGE. Because protein expression in IVTT does not utilize the native molecular machinery of protein expression, the DNA used to express the protein in IVTT lacks DNA encoding a signal peptide region. When CsgG DNA is expressed in IVTT in the absence of CsgF DNA, only CsgG monomers are produced. Surprisingly, these expressed monomers can be assembled into CsgG oligomeric pores in situ using cell extract membranes present in the IVTT reaction mixture. Figure 14 , lane 1). Although CsgG oligomers are SDS-stabilized, they decompose into their constituent monomers when the sample is heated to 100°C ( Figure 14 , lane 2). When CsgF DNA was expressed in IVTT in the absence of CsgG DNA, only CsgF monomers were seen ( Figure 14 , lane 3). When CsgG and CsgF DNA were mixed at a 1:1 ratio and simultaneously expressed in the same IVTT reaction mixture, the generated CsgF protein efficiently interacted with the assembled CsgG pore to generate a CsgG:CsgF complex ( Figure 14 , lane 5). This SDS-stabilized complex generated in IVTT is thermostable at least up to 70°C ( Figure 14 , lanes 6-12).

[0473] CsgG:CsgF complexes with truncated CsgF can also be prepared by any of the methods described above using DNA encoding a truncated, rather than full-length, form of CsgF. However, when CsgF is truncated below the FCP domain, the stability of the complex may be compromised. Alternatively, once the full-length CsgG:CsgF complex is formed, CsgG:CsgF complexes with truncated CsgF can be prepared by cleaving the full-length CsgF at the appropriate position. Truncations can be performed by modifying the DNA encoding the CsgF protein to incorporate protease cleavage sites at the desired positions ( FIG. 15A ). Seq ID Nos. 56-67 show TEV or HCV C3 protease sites incorporated at various positions in CsgF to generate CsgG:CsgF complexes with truncated CsgF. SDS-PAGE analysis of TEV cleavage of the CsgG:CsgF complex prepared using Seq ID No. 61 is shown ( FIG. 15B ). As described in the Materials and Methods section for analytical characterization, when the CsgG:CsgF complex (with full-length CsgF) was treated with TEV protease, CsgF was truncated at position 35 ( FIG15B , lanes 3 and 4). However, TEV cleavage leaves an additional 6 amino acids C-terminal to the cleavage site. Thus, the remaining CsgF truncated protein complexed with the CsgG pore is 42 amino acids long. The molecular weight difference between this complex and the CsgG pore (without CsgF) was still visible on SDS-PAGE ( FIG15B , lanes 7 and 8).

[0474] Surprisingly, CsgG:CsgF complexes with truncated CsgF can also be prepared by reconstituting the purified CsgG pore (prepared in vivo or in vitro) with a synthetic peptide of appropriate length. Because the reconstitution is performed in vitro, the CsgF signal peptide is not required to prepare the CsgG:CsgF complex. Furthermore, this method does not leave extra amino acids at the C-terminus of CsgF. Mutations and modifications can also be easily incorporated into the synthetic CsgF peptide. Therefore, this method is a very convenient way to reconstitute different CsgG pores, or mutants or homologs thereof, with different CsgF peptides, or mutants or homologs thereof, to generate different CsgG:CsgF complex variants. When CsgF is truncated beyond the FCP domain, the stability of the complex may be compromised. Table 3 shows examples of truncated CsgF and FCP peptides used to generate CsgG:CsgF complex variants. Surprisingly, SDS-PAGE analysis of the thermal stability of CsgG:CsgF complexes prepared using this method with CsgF-(1-45) ( FIG. 16A ), CsgF-(1-35) ( FIG. 16B ), and CsgF-(1-30) ( FIG. 16C ) showed that at least the CsgF-(1-45) and CsgF-(1-35) peptides produced complexes with CsgG that were thermally stable at at least 90°C. Because the CsgG pore breaks down into its component monomers at 90°C, it is difficult to assess the stability of the complexes above 90°C. Because the difference between the bands of the CsgG pore and the CsgG:CsgF-(1-30) complex in SDS-PAGE was minimal, this method was insufficient for analyzing the thermal stability of the CsgG:CsgF-(1-30) complex ( FIG. 16C ). However, CsgG:CsgF complexes were observed in all three cases, and even CsgG:CsgF-(1-29) was observed in the electrophysiological experiments, indicating that even the CsgF-(1-29) peptide generated at least some CsgG:CsgF complexes ( FIG. 24 ).

[0475] Example 2: CsgG:CsgF structural analysis by cryo-EM

[0476] To gain structural insights into the CsgG:CsgF complex, copurified or in vitro reconstituted CsgG:CsgF particles were analyzed by transmission electron microscopy. In preparation for cryo-EM analysis, 500 μL of peak fractions of the dual-affinity purified CsgG:CsgF complex were injected onto a Superose 6 10 / 30 column equilibrated with buffer D (25 mM Tris pH 8, 200 mM NaCl, and 0.03% DDM) and run at 0.5 mL / min. Protein concentration was determined based on calculated absorbance at 280 nm and assuming a 1:1 stoichiometry. Samples for cryo-EM were analyzed as described in the Methods. Figure 4 shows cryo-EM micrographs of the CsgG:CsgF complex and two selected class averages from selected CsgG:CsgF particles. The micrographs reveal the presence of the nonameric pore as well as dimers of the nonameric pore complex. For image reconstruction, nonameric CsgG:CsgF particles were selected and aligned using RELION. Class averages of the CsgG:CsgF complex as a side view and the electron density of the 3D reconstruction showed the presence of additional density corresponding to CsgF, seen as protrusions from the CsgG particle located on the side of the CsgG β-barrel ( Figure 4B 、 Figure 5 The additional density reveals three distinct regions, including a globular head domain, a hollow neck domain, and a domain that interacts with the CsgG β-barrel. The latter CsgF region, termed the CsgF contractile peptide or FCP, inserts into the lumen of the CsgG β-barrel and an additional constriction can be seen forming the CsgG pore (in Figure 4B 、 Figure 5 The constriction is located in the constriction formed by the CsgG constriction ring ( Figure 4B 、 Figure 5 Approximately 2 nm above the surface (marked as G in the figure).

[0477] Example 3: Identification of CsgF-interacting and contractile peptides by truncating CsgF

[0478] Compared to CsgG-only pores, the presence of a second constriction in the CsgG:CsgF pore complex offers opportunities for nanopore sensing applications, providing a second orifice in the nanopore that could serve as a secondary read head or as an extension of the primary read head provided by the CsgG constriction ring (Figures 6, 7). However, when complexed with full-length CsgF, the exit side of the CsgG:CsgF combined pore is blocked by the CsgF neck and head domains. Therefore, we sought to identify the region of CsgF required for interaction with and insertion into the CsgG β-barrel. Our Strep-tactin affinity purification experiments suggested that the N-terminal region of CsgF is required for CsgG interaction, as an N-terminal truncated fragment of CsgF present in the His-trap affinity purification was lost and did not co-purify with CsgG ( Figure 2B CsgF homologues are characterised by the presence of the PFAM domain PF03783. When a multiple sequence alignment (MSA) of CsgG homologues found in Gram-negative bacteria was performed ( Figure 8 (MSA of selected CsgF homologs is shown) revealed a region of sequence conservation corresponding to the first ~30-35 amino acids of mature CsgF (SEQ ID NO:6) (pairwise sequence identities ranging from 35 to 100%). Based on the combined data, it was hypothesized that this N-terminal region of CsgF forms a CsgG-interacting peptide, or FCP. Figure 10 A multiple sequence alignment of FCP among known CsgF homologs is shown in .

[0479] To test the hypothesis that the CsgF N-terminus corresponds to the CsgG binding region and forms a CsgF contractile peptide retained within the lumen of the CsgG β-barrel, Strep-tagged CsgG and His-tagged CsgF truncations were co-overexpressed in Escherichia coli (see Methods). pNA97, pNA98, pNA99, and pNA100 encode N-terminal CsgF fragments corresponding to residues 1-27, 1-38, 1-48, and 1-64 of CsgF (SEQ ID NO:5). These peptides include the CsgF signal peptide corresponding to residues 1-19 of SEQ ID NO:5 and thus will produce peptides that are similar to mature CsgF (SEQ ID NO:6; Figure 9A ), each containing a C-terminal 6x His tag. SDS-PAGE analysis of whole-cell lysates revealed the presence of CsgG in all samples, as well as a CsgF-like protein (SEQ ID NO: 6; Figure 9B) corresponding to the first 45 residues of CsgF fragments. For the shorter N-terminal CsgF fragment, no detectable expression of the peptide was found in whole-cell lysates. After two freeze / thaw cycles, cell pellets of various CsgG:CsgF fragments were further enriched by purification. Whole-cell lysates and eluted fractions from Strep affinity purification were spotted onto nitrocellulose membranes and dot blot analysis was performed using anti-His antibody to detect His-tagged CsgF fragments ( Figure 9C Dot blot analysis showed that the CsgF 20:64 peptide copurified with CsgG, demonstrating that this CsgF fragment is sufficient to form a stable non-covalent complex with CsgG. For the CsgG 20:48 fragment, a small amount of peptide was found to copurify with CsgG, whereas no detectable levels were found for either CsgF 20:27 or CsgF 20:38 in either whole cell lysates or Strep affinity purifications. Figure 9C ), indicating that these latter peptides were not stably expressed in E. coli and / or did not form a stable complex with CsgG.

[0480] Example 4: Characterization of the CsgG:CsgF interaction at atomic resolution.

[0481] To obtain atomic-level details of the CsgG:CsgF interaction, we determined the high-resolution cryoEM structure of the CsgG:CsgF complex. To this end, CsgG and CsgF were recombinantly expressed in Escherichia coli and isolated from the E. coli outer membrane by detergent extraction and purified using tandem affinity purification. Samples were prepared for cryo-electron microscopy by spotting 3 μl of sample onto R2 / 1 porous grids coated with graphene oxide (Quantifoil), and data were collected in counting mode on a 300 kV TITAN Krios with a Gatan K2 direct electron detector. Calculations were performed using 62,000 individual CsgG:CsgF particles. The final electron density map at resolution ( Figure 11A This map allowed the unambiguous docking and local reconstruction of the CsgG crystal structure, as well as the de novo construction of the N-terminal 35 residues of mature CsgF (i.e., residues 20:54 of Seq ID No. 5), which encompass the FCP that binds CsgG and forms the second constriction at the height of the CsgG transmembrane β-barrel ( Figure 11C , D). The cryoEM structure shows that CsgG:CsgF has a 9:9 stoichiometric ratio and C9 symmetry ( Figure 11BThe FCP binds to the interior of the CsgG β-barrel, while the CsgF C-terminus points outward from the CsgG β-barrel and the CsgF N-terminus is located near the CsgG constriction. The structure reveals that P35 in mature CsgF is located outside the CsgG β-barrel and forms a link between the CsgF FCP and the neck region. Due to its flexibility relative to the bulk of the CsgG:CsgF complex, the neck and head regions of CsgF cannot be resolved in high-resolution cryoEM images. Three regions in the CsgG β-barrel stabilize the CsgG:CsgF interaction: (IR1) residues Y130, D155, S183, N209, and T207 in mature CsgG (SEQ ID NO: 3) form an interaction network with the N-terminal amine and residues 1-4 of mature CsgF (SEQ ID NO: 6), including four hydrogen bonds and one electrostatic interaction; (IR2) residues Q187, D149, and E203 in mature CsgG (SEQ ID NO: 3) form an interaction network with R8 and N9 in mature CsgF (SEQ ID NO: 6), encompassing three H-bonds and two electrostatic interactions; and (IR3) residues F144, F191, F193, and L199 in mature CsgG (SEQ ID NO: 3) form a hydrophobic interaction surface with residues F21, L22, and A26 in mature CsgF (SEQ ID NO: 6). The latter is located in the α-helix (helix 1) formed by residues 19-30 of mature CsgF. The conserved sequence NPXFGG (residues 9-14 in SEQ ID NO:6) forms an inward turn connecting the loop region formed by residues 15-19 to the CsgF helix 1. Together, these elements create a constriction in the CsgG:CsgF complex, with residue 17 (mature E. coli CsgF, i.e., N17 in SEQ ID NO:6) forming the narrowest point, resulting in a diameter of The orifice ( Figure 11C The second constriction (F constriction or FC) is located approximately 1 / 2 of the top and 1 / 2 of the bottom constriction (G constriction or GC) formed by CsgG residues 46 to 59, respectively. and Place.

[0482] Example 5: Simulation of Improving the Stability of CcgG-CsgF Complex

[0483] Molecular dynamics simulations were performed to determine which residues in CsgG and CsgF are in close proximity. This information was used to design mutants of CsgG and CsgF that would increase the stability of the complex.

[0484] Simulations were performed using the GROMACS software package, version 4.6.5, with the GROMOS 53a6 force field and the SPC water model. The cryo-EM structure of the CsgG-CsgF complex was used in the simulations. The complex was solvated and then energy minimized using the steepest descent algorithm. Throughout the simulations, the complex backbone was restrained; however, the residue side chains were free to move. The system was simulated in the NPT ensemble for 20 nanoseconds at 300 K using a Berendsen thermostat and Berendsen barostat.

[0485] Contacts between CsgG and CsgF were analyzed using GROMACS analysis software and locally written code. Two residues were defined as making a contact if they were within 3 angstroms of each other. The results are shown in Table 4 below.

[0486] Table 4: Predicted contact frequencies of residue pairs in the CsgG / CsgF complex:

[0487]

[0488]

[0489]

[0490]

[0491] Materials and methods used for structure determination of the CsgG:CsgF complex:

[0492] clone

[0493] To express E. coli CsgG as an outer membrane localized pore, the coding sequence of E. coli CsgG (SEQ ID NO: 1) was cloned into pASK-Iba12 to generate plasmid pPG1 (Goyal et al., 2013).

[0494] To express CsgF tagged with 6x-His at the C-terminus in the E. coli cytoplasm, the coding sequence of mature E. coli CsgF (SEQ ID NO: 6; i.e., CsgF without its signal sequence) was cloned into pET22b via the NdeI and EcoRI sites using PCR products generated with primers "CsgF-His_pET22b_FW" (SEQ ID NO: 46) and "CsgF-His_pET22b_Rev" (SEQ ID NO: 47) to generate the CsgF-His expression plasmid pNA101.

[0495] Based on pGV5403 (integrated with pDEST14 The pTrc99a cassette was replaced with a streptomycin / spectinomycin resistance cassette to generate the pNA62 plasmid, a pTrc99a-based vector expressing csgF-His and csgG-strep. The pGV5403 ampicillin resistance cassette was replaced with a streptomycin / spectinomycin resistance cassette. A PCR fragment encompassing a portion of the E. coli MC4100 csgDEFG operon corresponding to the coding sequences of csgE, csgF, and csgG was generated using primers csgEFG_pDONR221_FW (SEQ ID NO: 48) and csgEFG_pDONR221_Rev (SEQ ID NO: 49) and cloned by BP. The recombinant was inserted into pDONR221 (ThermoFisherScientific). This recombinant csgEFG operon from the pDONR221 donor plasmid was inserted into pGV5403 with a streptomycin / spectinomycin resistance cassette. A 6xHis tag was added to the C-terminus of CsgF by PCR using primers Mut_csgF_His_FW (SEQ ID NO: 50) and Mut_csgF_His_Rev (SEQ ID NO: 51). Finally, csgE was removed by outward PCR using primers DelCsgE_FW (SEQ ID NO: 52) and DelCsgE_Rev (SEQ ID NO: 53) to obtain pNA62.

[0496] A C-terminally His-tagged CsgF fragment corresponding to the putative contractile peptide was generated for periplasmic expression by outward PCR on pNA62, a pTrc99a-based vector expressing CsgF-his and CsgG-strep. Figure 9A ). The primer combinations were as follows: pNa62_CsgF_his tag_Fw (SEQ ID NO: 45) as the forward primer, and CsgF_d27_end (SEQ ID NO: 41), CsgF_d38_end (SEQ ID NO: 42), CsgF_d48_end (SEQ ID NO: 43), or CsgF_d64_end (SEQ ID NO: 44) as the reverse primer to generate pNA97, pNA98, pNA99, and pNA100, respectively.

[0497] In pNA97, csgF was truncated to SEQ ID NO:7, which encodes a CsgF fragment comprising residues 1-27 (SEQ ID NO:8); in pNA98, csgF was truncated to SEQ ID NO:9, which encodes a CsgF fragment comprising residues 1-38 (SEQ ID NO:10); in pNA99, csgF was truncated to SEQ ID NO:11, which encodes a CsgF fragment comprising residues 1-48 (SEQ ID NO:12); and in pNA100, csgF was truncated to SEQ ID NO:13, which encodes a CsgF fragment comprising residues 1-64 (SEQ ID NO:14). Expression of pNA97, pNA98, pNA99, and pNA100 in E. coli indeed resulted in the generation of a CsgG pore in the outer membrane (SEQ ID NO:3), and a CsgF-derived peptide having the following sequence was targeted to the periplasm:

[0498] They are "GTMTFQFRHHHHHH" (SEQ ID NO:37+6xHis), "GTMTFQFRNPNFGGNPNNGHHHHHHH" (SEQ ID NO:38+6xHis), "GTMTFQFRNPNFGGNPNNGAFLLNSAQAQHHHHHH" (SEQ ID NO:39+6xHis) and "GTMTFQFRNPNFGGNPNNGAFLLNSAQAQNSYKDPSYNDDFGIETHHHHHH" (SEQ ID NO:40+6xHis).

[0499] strain

[0500] For all cloning procedures, E. coli Top10 (F - mcrAΔ(mrr - hsdRMS - mcrBC)Φ80lacZΔM15ΔlacX74recA1araD139Δ(araleu)7697galUgalKrpsL(StrR)endA1nupG). For protein production, Escherichia coli C43(DE3) (F – ompT hsdSB(rB - mB - )gal dcm(DE3))and Top10.

[0501] Production of recombinant CsgG:CsgF complexes by co-expression

[0502] To co-express E. coli CsgF (SEQ ID NO: 5) and CsgG (SEQ ID NO: 2), the two recombinant genes (including their native Shine Dalgarno sequences) were placed in a pTrc99a-derived plasmid under the control of the inducible trc promoter to form plasmid pNA62. CsgG and CsgF were overexpressed in E. coli C43 (DE3) cells transformed with plasmid pNA62 and grown in Terrific Broth medium at 37°C. When the cell culture reached an optical density at 600 nm (OD) of 0.7, recombinant protein expression was induced with 0.5 mM IPTG and grown at 28°C for 15 hours before harvesting by centrifugation at 5500 g.

[0503] Generation of recombinant CsgG:CsgF complexes by in vitro recombination

[0504] Full-length E. coli CsgG (SEQ ID NO: 2) modified with a C-terminal StrepII tag was overexpressed in E. coli BL21(DE3) cells transformed with plasmid pPG1 (Goyal et al., 2013). Cells were grown in Terrific Broth medium at 37°C to an OD 600nm of 0.6. Recombinant protein production was induced with 0.0002% anhydrotetracycline (Sigma), and cells were grown for an additional 16 hours at 25°C before harvesting by centrifugation at 5500 g.

[0505] E. coli CsgF (SEQ ID NO: 6; i.e., lacking the CsgF signal sequence) fused to a 6x His tag at the C-terminus was overexpressed in the cytoplasm of E. coli BL21(DE3) cells transformed with plasmid pNA101. Cells were grown at 37°C to an OD of 600 nm, induced with 1 mM IPTG, and allowed to express protein for 15 hours at 37°C before harvesting by centrifugation at 5500 g.

[0506] Purification of CsgG:CsgF complex, recombinant CsgG and CsgF proteins

[0507] Escherichia coli cells transformed with pNA62 and co-expressing CsgG-Strep and CsgF-His were resuspended in 50 mM Tris-HCl (pH 8.0), 200 mM NaCl, 1 mM EDTA, 5 mM MgCl2, 0.4 mM AEBSF, 1 μg / mL leupeptin, 0.5 mg / mL DNase I, and 0.1 mg / mL lysozyme. Cells were disrupted at 20 kPsi using a TS Series cell disruptor (Constant Systems Ltd). The lysed cell suspension was incubated with 1% n-dodecyl-β-d-maltopyranoside (DDM; Inalco) for 30 minutes to further lyse the cells and extract the outer membrane fraction. The remaining cell debris and membranes were then pelleted by ultracentrifugation at 100,000 g for 40 minutes. The supernatant was loaded onto a 5 mL HisTrap column equilibrated in buffer A (25 mM Tris pH 8, 200 mM NaCl, 10 mM imidazole, 10% sucrose, and 0.06% DDM). The column was washed with >10 CV of 5% buffer B (25 mM Tris pH 8, 200 mM NaCl, 500 mM imidazole, 10% sucrose, and 0.06% DDM) and eluted with a gradient of 5-100% buffer B over 60 mL.

[0508] The eluate was diluted 2-fold and then loaded onto a 5 mL Strep-tactin column (IBA GmbH) equilibrated with buffer C (25 mM Tris pH 8, 200 mM NaCl, 10% sucrose, and 0.06% DDM) overnight. The column was washed with >10 CV of buffer C and the protein was eluted by adding 2.5 mM desthiobiotin. Next, 500 μL of the peak fraction of the double affinity purification complex was injected onto a Superose 610 / 30 (GE Healthcare) equilibrated with buffer D (25 mM Tris pH 8, 200 mM NaCl, and 0.03% DDM) and run at 0.5 mL / min to prepare samples for electron microscopy. Protein concentration was determined based on the calculated absorbance at 280 nm and assuming a stoichiometric 1 / 1. Buffer D (25 mM Tris pH 8, 200 mM NaCl, and 0.03% DDM)

[0509] The purification of CsgG-strep for in vitro reconstitution was performed using the same protocol as for CsgG:CsgF, when sucrose was omitted from the buffer and the IMAC and size exclusion steps were bypassed.

[0510] CsgF-His purification for in vitro reconstitution was performed by resuspending the cell pellet in 50 mM Tris-HCl pH 8.0, 200 mM NaCl, 1 mM EDTA, 5 mM MgCl2, 0.4 mM AEBSF, 1 μg / mL leupeptin, 0.5 mg / mL DNase I, and 0.1 mg / mL lysozyme. Cells were disrupted at 20 kPsi using a TS series cell disruptor (Constant Systems Ltd), and the lysed cell suspension was centrifuged at 10.000 g for 30 minutes to remove intact cells and cell debris. The supernatant was added to 5 mL of Ni-IMAC-beads (Workbeads 40IDA, Bio-Works Technologies AB) equilibrated with buffer A (25 mM Tris pH 8, 200 mM NaCl, 10 mM imidazole) and incubated at 4°C for 1 hour. Ni-NTA beads were pooled onto a gravity flow column and washed with 100 mL of 5% buffer B (25 mM Tris pH 8, 200 mM NaCl, 500 mM imidazole diluted in buffer A). Bound proteins were eluted by stepwise addition of buffer B (10% steps of 5 mL each).

[0511] In vitro reconstitution of the CsgG:CsgF complex

[0512] Purified CsgG and CsgF were pooled and used for in vitro reconstitution of the complex. The mixture was then mixed at a molar ratio of 1 CsgG:2 CsgF, resulting in a CsgG barrel filled with CsgF. Next, the reconstitution mixture was injected onto a Superose 6 10 / 30 chromatography column (GE Healthcare) equilibrated with buffer D (25 mM Tris pH 8, 200 mM NaCl, and 0.03% DDM) and run at 0.5 mL / min to prepare samples for electron microscopy ( FIG3 ). Protein concentrations were determined based on calculated absorbance at 280 nm and assuming a 1 / 1 stoichiometry.

[0513] Structural analysis using electron microscopy

[0514] Negative stain electron microscopy was used to probe the sample behavior of the size exclusion fraction. Samples were stained with 1% uranyl formate and imaged using an in-house 120 kV JEM 1400 (JEOL) microscope equipped with LaB6 filaments. Samples for cryo-electron microscopy were prepared by spotting 2 μL of sample onto R2 / 1 continuous carbon (2 nm) coated grids (Quantifoil), manually applied, and inserted into liquid ethane using an in-house insertion device. Sample quality was screened on the in-house JEOL JEM 1400, and data sets were collected on a 200 kV TALOS ARCTICA (FEI) microscope equipped with a Falcon-3 direct electron detection camera. Images were motion corrected using MotionCor2.1 (Zheng et al., 2017), defocus values ​​were determined using ctffind4 (Rohou and Grigorieff, 2015), and data were further analyzed using a combination of RELION (Scheres, 2012) and EMAN2 (Ludtke, 2016). C9 symmetry was imposed on selected 2D class averages during 3D model generation and refinement, characterized by additional density for the head groups.

[0515] For high-resolution cryoEM analysis, CsgG:CsgF samples were prepared for cryo-electron microscopy by spotting 3 μl of sample onto R2 / 1 porous grids (Quantifoil) coated with graphene oxide (Sigma Aldrich), manually applying the sample and inserting it into liquid ethane using a CP3 plunger (Gatan). Sample quality was screened on an in-house JEOL JEM 1400, and data sets were collected on a 300 kV TITAN KRIOS (FEI, Thermo-Scientific) microscope equipped with a K2 Summit direct electron detector (Gatan). The detector was used in counting mode, with 50 frames per second. The cumulative electron dose of the distribution is 56 electrons. 2045 images with a pixel size of Images were motion corrected using MotionCor2.1 (Zheng et al., 2017), and defocus values ​​were determined using ctffind4 (Rohou and Grigorieff, 2015). Particles were automatically picked using Gautomatch (Dr. Kai Zhang), and the data were further analyzed using a combination of RELION2.0 (Kimanius et al., 2016, Elife 5.pii:e18722) and EMAN2 (Ludtke, 2016). During 3D model generation and refinement, C9 symmetry was imposed on the selected 2D class averages, characterized by an additional density for the head group corresponding to CsgF. 62,000 particles were used to The final map was calculated at a resolution of 100 nm. A de novo model of CsgF was constructed using COOT (Brown et al. 2015 Acta Crystallogr D Biol Crystallogr 71(Pt 1):136-53), and an iterative cycle of model construction and refinement of the complete complex was performed using PHENIX (Afonine 2018, Acta Crystallogr D Struct Biol 74(Pt 6):531-544) real-space refinement combined with COOT.

[0516] Protein expression and purification of the CsgG:CsgF fragment

[0517] The CsgF fragment and CsgG were co-expressed, with the CsgF fragment having a His tag at its C-terminus and the CsgG C-terminus fused to a Strep tag. The CsgG:CsgF fragment complex was overexpressed in E. coli Top10 cells transformed with plasmids pNA97, pNA98, pNA99, or pNA100. Plates were incubated at 37°C ON, and colonies were resuspended in LB medium supplemented with streptomycin / spectinomycin. When the cell culture reached an optical density at 600 nm (OD) of 0.7, recombinant protein expression was induced with 0.5 mM IPTG and grown at 28°C for 15 hours before harvesting by centrifugation at 5500 g. The pellet was frozen at -20°C.

[0518] Cell pellets co-expressing various CsgG:CsgF fragments were resuspended in 200 mL of 50 mM Tris-HCl, pH 8.0, 200 mM NaCl, 1 mM EDTA, 5 mM MgCl2, 0.4 mM AEBSF, 1 μg / mL leupeptin, 0.5 mg / mL DNase I, and 0.1 mg / mL lysozyme, sonicated, and incubated with 1% n-dodecyl-β-d-maltopyranoside (DDM; Inalco) to further lyse the cells and extract the outer membrane fraction. Next, the remaining cell debris and membranes were pelleted by centrifugation at 15,000 g for 40 minutes. The supernatant was incubated with 100 μL of Trep-tactin beads at room temperature for 30 minutes. Strep beads were washed with buffer (25 mM Tris pH 8, 200 mM NaCl, and 1% DDM) by centrifugation, and bound protein was eluted by addition of 2.5 mM desthiobiotin in 25 mM Tris pH 8, 200 mM NaCl, 0.01% DDM.

[0519] CsgG:FCP was generated by in vitro reconstitution.

[0520] A synthetic peptide corresponding to the N-terminal 34 residues of mature CsgF (SEQ ID NO: 6) was diluted to 1 mg / ml in a buffer of 0.1 M MES, 0.5 M NaCl, 0.4 mg / ml EDC (1-ethyl-3-(3-dimethylaminopropyl)carbodiimide), and 0.6 mg / ml NHS (N-hydroxysuccinimide) and incubated at room temperature for 15 minutes to allow activation of the peptide carboxyl terminus. Next, a 1 mg / ml solution of Cadaverin-Alexa 594 in PBS was added during a 2-hour incubation to allow covalent coupling at room temperature. The reaction was quenched by exchanging the buffer to 50 mM Tris, NaCl, 1 mM EDTA, and 0.1% DDM using a Zeba Spin filter.

[0521] Labeled peptides were added to strep-affinity purified CsgG in 50 mM Tris, 100 mM NaCl, 1 mM EDTA, and 5 mM LDAO / C8D4 at room temperature for 15 minutes at a 2:1 molar ratio to allow reconstitution of the CsgG:FCP complex. After CsgG-strep was pulled down onto StrepTactin beads, samples were analyzed on native PAGE.

[0522] Example 6: Further stabilization of the CsgG:CsgF complex by covalent cross-linking

[0523] Although the full-length form and some truncated forms of CsgF form stable CsgG:CsgF complexes with the CsgG pore, CsgF can still be forcibly removed from the barrel region of the CsgG pore under certain conditions. Therefore, it is desirable to create a covalent link between the CsgG and CsgF subunits. Based on molecular modeling studies, positions of CsgG and CsgF that are in close proximity have been identified (Example 5 and Table 4). Some of these identified positions have been modified to incorporate cysteines in both CsgG and CsgF. Figure 19 shows an example of a thiol-thiol bond formed between position Q153 of CsgG and position G1 of CsgF. A CsgG pore containing the Q153C mutation was reconstituted with CsgF containing the G1C mutation and incubated for 1 hour to enable SS bond formation. When the complex was heated to 100°C in the absence of DTT, a 45 kDa band corresponding to the dimer between CsgG and CsgF monomers (CsgGm-CsgFm) was visible, indicating the formation of an SS bond between the two monomers (30 kDa for CsgGm and 15 kDa for CsgFm) ( FIG19A ). This band disappeared when heated in the presence of DTT. DTT breaks down the SS bond. When the CsgG:CsgF complex was incubated overnight rather than for 1 hour, the extent of CsgGm-CsgFm dimer formation increased ( FIG19A ). Mass spectrometry was performed to further identify the dimer band. The gel-purified protein was proteolytically cleaved to generate tryptic peptides. LC-MS / MS sequencing was performed to identify the SS bond between position Q153 of CsgG and position G1 of CsgF ( FIG19B ). Oxidizing agents such as copper-o-phenanthroline can be used to enhance SS bond formation. As described in the Methods section, when the CsgG pore containing the N133C modification was reconstituted with the CsgF containing the T4C modification in the presence of copper-o-phenanthroline and then dissociated into its constituent monomers by heating to 100°C in the absence of DTT, a strong dimer band corresponding to CsgGm-CsgFm was observed on SDS-PAGE ( Figure 20 , lanes 3 and 4). When heated in the presence of DTT, the dimer dissociates into its component monomers ( Figure 20 , lanes 1 and 2).

[0524] Example 7: Electrophysiological Characterization of CsgG:CsgF Complex

[0525] When pores were inserted into copolymer membranes and experiments were performed using the Oxford Nanopore Technologies MinION, the signal observed when a DNA strand translocates through CsgG could be well characterized ( FIG31 ). Y51, N55, and F56 of each CsgG subunit form the constriction of the CsgG pore ( Figure 12This sharp constriction acts as the reading head for the CsgG pore ( Figure 31A ) and can accurately distinguish mixed sequences of A, C, G, and T as they pass through the pore. This is because the measured signal contains a characteristic current deflection from which sequence identity can be inferred. However, in homopolymer regions of DNA, the measured signal may not show a current deflection of sufficient magnitude to enable single base identification; making it impossible to accurately determine the length of the homopolymer from the amplitude of the measured signal ( Figure 26B and Figure 26C The decrease in CsgG reader accuracy is correlated with the length of the homopolymer region ( Figure 29C When CsgF interacts with the CsgG pore to form a CsgG:CsgF complex, CsgF introduces a second read head into the CsgG barrel. The second read head consists primarily of position N17 of Seq. ID No. 6. Static chain experiments as described in the Methods section and FIG. 27 were performed to experimentally map the two read heads of the CsgG:CsgF complex, demonstrating the presence of two read heads approximately 5-6 bases apart ( Figure 27B 、 Figure 27C and Figure 27D The reader discrimination map of the CsgG:CsgF complex shows that the second reader introduced by CsgF contributes less to base discrimination than the CsgG reader ( Figure 27A Surprisingly, when a second read head was introduced into the CsgG barrel by CsgF, the previously flat homopolymer region showed a stepping signal ( Figure 30B and Figure 30C These steps contain information that can be used to accurately identify sequences, thereby reducing errors. Compared to the accuracy distribution of the CsgG pore itself, the accuracy of the DNA signal of the CsgG:CsgF complex remains relatively constant over longer homopolymer lengths ( Figure 29C ).

[0526] CsgG:CsgF complexes prepared using any of the methods described in the Methods section can be used to characterize the complexes in DNA sequencing experiments. Figures 21-24 show the signals of a lambda DNA strand passing through various CsgG:CsgF complexes prepared using different methods and composed of different CsgG mutant pores and different CsgF peptides of different lengths. Figures 28 (A-H) show the read head differences and base contribution curves for those pore complexes. Surprisingly, different modifications in the constriction of both the CsgG pore and the CsgF peptide can significantly alter the signal of the CsgG:CsgF pore complex. For example, when CsgG:CsgF complexes were prepared using the same CsgG pore but with two different CsgF peptides of the same length containing either Asn or Ser at position 17 (Seq ID No. 6) (prepared by co-expressing the full-length CsgF protein and then cleaving CsgF between positions 35 and 36 with TEV protease), the signals generated were distinct from each other ( Figure 21 ). Compared to the CsgG:CsgF complex with Asn at position 17, the CsgG:CsgF complex with Ser at position 17 of the CsgF peptide showed lower noise and higher signal-to-noise ratio. Similarly, when the same CsgG pore was reconstituted with two different CsgF peptides of the same length (1-35 of Seq ID No. 6) but with Ser or Val at position 17 to prepare CsgG:CsgF complexes, the complex with Val at position 17 of CsgF showed a noisier signal than the complex with Ser at position 17 of CsgF ( FIG. 22 ). When the same CsgF peptide of the same length was reconstituted with different CsgG pores containing different mutations at positions 51, 55, and 56 of the CsgG read head, the resulting CsgG:CsgF complexes showed very different signals ( Figure 23A -F), with different signal-to-noise ratios ( Figure 25 Surprisingly, when CsgF peptides of different lengths containing the same constriction were reconstituted with the same CsgG pore to prepare CsgG:CsgF complexes, they produced signals of varying ranges ( FIG. 24 ). The CsgG:CsgF complex containing the shortest CsgF peptide (1-29 of Seq ID No. 6) exhibited the greatest range, while the CsgG:CsgF complex containing the longest CsgF peptide (1-45 of Seq ID No. 6) exhibited the smallest range ( FIG. 24 ).

[0527] Materials and methods used for analyte characterization:

[0528] Proteins produced by the methods described below can be used interchangeably with proteins produced by the methods described above for structure determination.

[0529] method

[0530] Expression of CsgG:CsgF or CsgG:FCP complexes by co-expression

[0531] Genes encoding the CsgG protein and its mutants were constructed in the pT7 vector containing the ampicillin resistance gene. Genes encoding the CsgF or FCP protein and its mutants were constructed in the pRham vector containing the kanamycin resistance gene. 1 μL of either plasmid was mixed with 50 μL of Lemo(DE3)ΔCsgEFG on ice for 10 minutes. The sample was then heated at 42°C for 45 seconds and returned to ice for 5 minutes. 150 μL of NEB SOC growth medium was added, and the sample was incubated at 37°C with shaking at 250 rpm for 1 hour. The entire volume was plated onto agar plates containing kanamycin (40 μg / mL), ampicillin (100 μg / mL), and chloramphenicol (34 μg / mL) and incubated overnight at 37°C. A single colony was removed from the plate and inoculated into 100 mL of LB medium containing kanamycin (40 μg / mL), ampicillin (100 μg / mL), and chloramphenicol (34 μg / mL) and incubated overnight at 37°C with shaking at 250 rpm. 25 mL of the starter culture was added to 500 mL of LB medium containing 3 mM ATP, 15 mM MgSO4, kanamycin (40 μg / mL), ampicillin (100 μg / mL), and chloramphenicol (34 μg / ml) and incubated overnight at 37°C. The culture was grown for 7 hours at which time the OD 600 Greater than 3.0. Lactose (final concentration 1.0%), glucose (final concentration 0.2%), and rhamnose (final concentration 2 mM) were added and the temperature was lowered to 18°C ​​while shaking at 250 rpm for 16 hours. The culture was centrifuged at 6000 rpm for 20 minutes at 4°C. The supernatant was discarded and the pellet was retained. The cells were stored at -80°C until purification.

[0532] Expression of CsgG pores with or without a C-terminal Strep tag and CsgF with or without a C-terminal Strep or His tag

[0533] All genes encoding all CsgG proteins and either CsgF or FCP proteins were constructed in a pT7 vector containing an ampicillin resistance gene. The expression procedure was the same as above, except that kanamycin was omitted from all media and buffers.

[0534] Cell lysis (co-expressed complex or individual CsgG / CsgF / FCP proteins)

[0535] Lysis buffer is made up of 50mM Tris (pH 8.0), 150mM NaCl, 0.1% DDM, 1x Bugbuster protein extraction reagent (Merck), 2.5uL Benzonase nuclease (stock solution ≥ 250 units / μL) / 100mL lysis buffer and 1 Sigma protease inhibitor cocktail / 100mL lysis buffer. Use 5X volume of lysis buffer to lyse 1X weight of harvested cells. Resuspend the cells and centrifuge at room temperature for 4 hours until a homogenous lysate is produced. The lysate is centrifuged at 20,000rpm for 35 minutes at 4°C. Carefully extract the supernatant and filter through a 0.2uM Acrodisc syringe filter.

[0536] Strep purification of CsgG or CsgF / FCP proteins or co-expressed complexes when CsgG contains a C-terminal Strep tag and CsgF or FCP contains a C-terminal His tag

[0537] The filtered sample was then loaded onto a 5 mL StrepTrap column using the following parameters: loading rate: 0.8 mL / min, total sample load: 10 mL, washout of unbound: 10 CV (5 mL / min), additional wash: 10 CV (5 mL / min), elution: 3 CV (5 mL / min). Affinity buffer: 50 mL Tris (pH 8.0), 150 mM NaCl, 0.1% DDM; wash buffer: 50 mL Tris (pH 8.0), 2 M NaCl, 0.1% DDM; elution buffer: 50 mL Tris (pH 8.0), 150 mM NaCl, 0.1% DDM, 10 mM desthiobiotin. The eluted sample was collected.

[0538] His purification of CsgG or CsgF / FCP proteins or co-expressed complexes when CsgG contains a C-terminal Strep tag and CsgF or FCP contains a C-terminal His tag

[0539] Filtered samples or pooled elution peaks from Strep purification (in the case of complexes) were loaded onto a 5 mL HisTrap column using the same parameters as above, except using the following buffer: Affinity and Wash Buffer: 50 mL Tris (pH 8.0), 150 mM NaCl, 0.1% DDM, 25 mM Imidazole; Elution: 50 mL Tris (pH 8.0), 150 mM NaCl, 0.1% DDM, 350 mM Imidazole. The peak eluted and concentrated to a volume of 500 uL in a 30 kDa MWCO Merck Milipore centrifugal device.

[0540] Complexes are formed in vitro using components purified in vivo.

[0541] Separately expressed and purified CsgG and CsgF / FCP proteins were mixed at various ratios to identify the correct ratio, always in the presence of excess CsgF. The complex was then incubated overnight at 25°C. To remove excess CsgF and remove DTT from the buffer, the mixture was injected again onto a Superdex Increase 200 10 / 300 column equilibrated in 50 mM Tris (pH 8.0), 150 mM NaCl, and 0.1% DDM. The complex typically eluted from this column in 9 to 10 mL.

[0542] Polishing step of complexes (co-expressed or in vitro prepared) by gel filtration

[0543] If necessary, Strep-purified, His-purified, or His-purified then Strep-purified CsgG:CsgF or CsgG:FCP can be further polished by gel filtration. 500 μL of sample is injected into a 1 mL sample loop and onto a Superdex Increase 200 10 / 300 column equilibrated in 50 mM Tris (pH 8.0), 150 mM NaCl, 0.1% DDM. When run at 1 mL / min, the peak associated with the complex typically elutes in 9 to 10 mL from the column. The sample is heated at 60°C for 15 minutes and centrifuged at 21,000 rcf for 10 minutes. The supernatant is used for analysis. The sample is subjected to SDS-PAGE to confirm and identify the fractions eluting with the complex.

[0544] Cleavage of CsgF or FCP at the TEV protease site

[0545] If CsgF or FCP contains a TEV cleavage site, add TEV protease with a C-terminal histidine tag to the sample containing 2 mM DTT (the amount added is determined by the approximate concentration of the protein complex). Incubate the sample overnight at 4°C on a roller mixer at 25 rpm. The mixture is then run through a 5 mL HisTrap column again, and the flow-through is collected. Any uncleaved material will remain bound to the column, and the cleaved protein will elute. Use the same buffers and parameters as described above for His purification, including a final heating step.

[0546] Purification of CsgG:FCP complexes with in vivo purified CsgG pores and synthetic FCP

[0547] Lyophilized FCP peptides were received from Genscript and Lifetein. Dissolve 1 mg of peptide in 1 mL of nuclease-free ddH2O to obtain a 1 mg / mL sample. Vortex the sample until no peptide is visible. Accurate concentration measurements are difficult due to variable expression levels of CsgG pores and mutants. Protein band intensities against known markers on SDS-PAGE can be used to obtain a rough estimate of sample concentration. CsgG and FCP were then mixed at a molar ratio of approximately 1:50 and incubated overnight at 700 rpm at 25°C. The sample was heated at 60°C for 15 minutes and centrifuged at 21,000 rcf for 10 minutes. The supernatant was used for experiments. If desired, the complex can be purified as described above for co-expression.

[0548] Purification of CsgG:CsgF or CsgG:FCP containing cysteine ​​mutants

[0549] If either or both components contain cysteine, the CsgG:CsgF or CsgG:FCP complexes (I, II, or III below) can be purified using the same procedures as above, except for the composition of the affinity buffer, wash buffer, and elution buffer for His and Strep purifications, as well as the buffer used for gel filtration. For purification of cysteine ​​mutants, all of these buffers should contain 2 mM DTT. 2 mM DTT was also added when synthetic peptides containing cysteine ​​were dissolved in ddH2O.

[0550] I. Co-expression of CsgG and CsgF or FCP

[0551] II. Preparation of CsgG:CsgF or CsgG:FCP Complexes in Vitro Using Individual Components Purified in Vivo III. Preparation of CsgG:CsgF or CsgG:FCP Complexes in Vitro Using Purified CsgG and Synthetic FCP

[0552] Determine cysteine ​​bond formation

[0553] Separate two 50 μL tubes of the final eluate. Add 2 mM DTT to one tube as a reducing agent, and 100 μM Cu(II):1–10 phenanthroline (33 mM:100 mM) to the other tube as an oxidizing agent. Mix the sample 1:1 with Laemmli buffer containing 4% SDS. Half of the sample was heat-treated to 100°C for 10 minutes, while the other half remained untreated and then run on a 4-20% TGX gel (Bio-rad Criterion) in TGS buffer.

[0554] Coupled in vitro transcription and translation (IVTT)

[0555] All proteins were produced by coupled in vitro transcription and translation (IVTT) using the E. coli T7-S30 extraction system (Promega) using circular DNA. A complete 1 mM amino acid mixture minus cysteine ​​and a complete 1 mM amino acid mixture minus methionine were mixed in equal volumes to obtain the working amino acid solution required for high-concentration protein production. Amino acids (10 μL) were mixed with a premix (40 μL), [35S]L-methionine (2 μL, 1175 Ci / mmol, 10 mCi / mL), plasmid DNA (16 μL, 400 ng / μL), T7 S30 extract (30 μL), and rifampicin (2 μL, 20 mg / mL) to produce a 100 μL IVTT protein reaction. Synthesis was performed at 30°C for 4 hours and then incubated overnight at room temperature. If a CsgG:CsgF or CsgG:FCP complex is being prepared for co-expression, equal amounts of plasmid DNA encoding each component are mixed, and a portion of the mixture (16 μL) is used for IVTT. After incubation, the tube is centrifuged at 22,000 g for 10 minutes, and the supernatant is discarded. The resulting pellet is resuspended and washed in MBSA (10 mM MOPS, 1 mg / ml BSA pH 7.4) and centrifuged again under the same conditions. The proteins present in the pellet are resuspended in 1X Laemmli sample buffer and run on a 4-20% TGX gel at 300 V for 25 minutes. The gel is then dried and exposed to The membrane was then processed and the proteins in the gel were visualized.

[0556] Samples used for experiments in MinIONs

[0557] Prior to the assay, all samples were incubated with Brij58 (final concentration 0.1%) for 10 minutes at room temperature and then wells were topped up with the required subsequent well dilutions.

[0558] Method for preparing and operating a static chain

[0559] A group of polyA DNA chains (SS20 to SS38 of Figure 27) were obtained by integrated DNA technology (IDT), in which the DNA backbone (iSpc3) lacked a base. The 3' end of each of these chains also included biotin modification. Static chain was incubated at room temperature for 20 minutes with monovalent streptavidin to cause biotin to bind to streptavidin. Streptavidin-static chain complex was diluted to 500nM (B, Figure 27) and 2uM (C, Figure 27) in 25mM HEPES, 430mM KCl, 30mM ATP, 30mM MgCl2, 2.15mM EDTA (pH 8) (referred to as RBFM). The residual current generated by each static chain was recorded in the MinION device. The MinION flow cell was rinsed according to the standard operation scheme, and then the sequencing program was started with a static flick of 1 minute. Initially, a 10-minute open hole record was produced, and then 150uL of the first streptavidin-static chain complex was added. After 10 minutes, 800 μL of RBFM was flushed through the flow cell before adding the next streptavidin-static chain complex. This process was repeated for all streptavidin-static chains. Once the final streptavidin-static chain complex was incubated on the flow cell, 800 μL of RBFM was flushed through the flow cell and a 10-minute open-pore recording was generated before completing the experiment.

[0560] How to create a difference curve graph

[0561] The read head discrimination curve shows the average change in simulated current when the base at each read head position is changed. To calculate the read head discrimination at position i for a model of length k and alphabet length n, we define the discrimination at read head position i as n k-1 The median of the standard deviations of the current levels for each of the groups where position i is varied while the other positions remain constant.

[0562] Aspects of the Disclosure

[0563] 1. An isolated pore complex comprising a CsgG pore, or a homologue or mutant thereof, and a modified CsgF peptide, or a homologue or mutant thereof.

[0564] 2. The isolated pore complex according to 1, wherein the modified CsgF peptide, or a homologue or mutant thereof, is inserted into the lumen of the CsgG pore, or a homologue or mutant thereof.

[0565] 3. The isolated pore complex of 2, wherein the pore complex has two or more channels comprising a CsgG channel constriction and a CsgF channel constriction.

[0566] 4. The isolated pore complex of any one of 1 to 3, wherein the CsgG pore or homologue or mutant thereof is a mutant CsgG pore.

[0567] 5. The isolated pore complex according to 3 or 4, wherein the diameter of the CsgF channel constriction is in the range of 0.5 nm to 2.0 nm.

[0568] 6. A modified CsgF peptide or a modified peptide of a CsgF homologue or mutant, wherein the modification comprises truncation of the CsgF protein of SEQ ID NO: 6 or a homologue or mutant thereof.

[0569] 7. The modified CsgF peptide or the modified peptide of a CsgF homologue or mutant according to 6, wherein the modified CsgF peptide comprises SEQ ID NO: 39 or SEQ ID NO: 40, or a homologue or mutant thereof.

[0570] 8. The modified CsgF peptide or the modified peptide of a CsgF homologue or mutant according to 6, wherein the modified CsgF peptide comprises SEQ ID NO: 15, or a homologue or mutant thereof.

[0571] 9. The modified CsgF peptide, or the modified peptide of a CsgF homologue or mutant according to 8, comprising one or more position mutations in the region of SEQ ID NO: 15, and having at least 35% amino acid identity with SEQ ID NO: 15.

[0572] 10. A polynucleotide encoding the modified CsgF peptide according to any one of 6 to 9.

[0573] 11. The isolated pore complex according to any one of 1 to 5, wherein the modified CsgF peptide or a homologue or mutant thereof is a peptide according to any one of 6 to 9.

[0574] 12. The isolated pore complex of claim 11, wherein the modified CsgF peptide and the CsgG pore or a homologue or mutant thereof are covalently coupled.

[0575] 13. The isolated pore complex according to 12, wherein the covalent coupling is by means of:

[0576] (i) a cysteine ​​residue at a position corresponding to 132, 133, 136, 138, 140, 142, 144, 145, 147, 149, 151, 153, 155, 183, 185, 187, 189, 191, 201, 203, 205, 207 or 209 of SEQ ID NO: 3 or a homolog thereof;

[0577] (ii) a non-natural reactive or photoreactive amino acid at a position corresponding to 132, 133, 136, 138, 140, 142, 144, 145, 147, 149, 151, 153, 155, 183, 185, 187, 189, 191, 201, 203, 205, 207 or 209 of SEQ ID NO: 3 or a homolog thereof.

[0578] 14. An isolated transmembrane pore complex comprising the isolated pore complex according to any one of 1 to 5 or 11 to 13, and a component of a membrane.

[0579] 15. A method for producing a transmembrane pore complex, wherein the pore complex is formed by a CsgG pore and a modified CsgF peptide, or a homologue or mutant thereof, the method comprising co-expressing CsgG, or a homologue or mutant thereof, as shown in SEQ ID NO: 2, and the modified CsgF peptide, or a homologue or mutant thereof, in a suitable host cell, thereby allowing formation of the transmembrane pore complex in vivo.

[0580] 16. The method according to 15, wherein the modified CsgF peptide or a homologue or mutant thereof comprises SEQ ID NO: 12 or SEQ ID NO: 14 or a homologue or mutant thereof.

[0581] 17. A method for producing an isolated pore complex, wherein the isolated pore is formed by a CsgG pore, or a homologue or mutant thereof, and a modified CsgF peptide, or a homologue or mutant thereof, the method comprising contacting a CsgG monomer of SEQ ID NO: 3, or a homologue or mutant thereof, with the modified CsgF peptide, or a homologue or mutant thereof, thereby allowing reconstitution of the isolated pore complex in vitro.

[0582] 18. The method according to 17, wherein the modified CsgF peptide or a homologue or mutant thereof comprises SEQ ID NO: 15 or SEQ ID NO: 16 or a homologue or mutant thereof.

[0583] 19. A method for determining the presence, absence, or one or more characteristics of a target analyte, comprising the steps of:

[0584] (i) contacting the target analyte with the pore complex according to any one of 1 to 5 or 11 to 13, or with the transmembrane pore complex according to 14, such that the target analyte moves into the pore complex; and

[0585] (ii) performing one or more measurements as the analyte moves through the pore complex to determine the presence, absence, or one or more characteristics of the analyte.

[0586] 20. The method of 19, wherein the analyte is a polynucleotide.

[0587] 21. The method according to 19, wherein the analyte is a (poly)peptide.

[0588] 22. The method of 19, wherein the analyte is a polysaccharide.

[0589] 23. The method according to 19, wherein the analyte is a small organic or inorganic compound, such as a pharmacologically active compound, a toxic compound, and a pollutant.

[0590] 24. The method of claim 20, comprising determining one or more characteristics selected from the group consisting of: (i) the length of the polynucleotide, (ii) the identity of the polynucleotide, (iii) the sequence of the polynucleotide, (iv) the secondary structure of the polynucleotide, and (v) whether the polynucleotide is modified.

[0591] 25. A method for characterising a polynucleotide or (poly)peptide using an isolated transmembrane pore complex, wherein the pore complex comprises a CsgG pore or a homologue or mutant thereof and a modified CsgF peptide or a homologue or mutant thereof.

[0592] 26. The method of 25, wherein the CsgG pore, or a homologue or mutant thereof, comprises six to ten monomers.

[0593] 27. Use of the isolated pore complex according to any one of 1 to 5 or 11 to 13 or the transmembrane pore complex according to 14 for determining the presence, absence or one or more characteristics of a target analyte.

[0594] 28. A kit for characterizing a target analyte comprising (a) the isolated pore complex according to any one of 1 to 5 or 11 to 13 and (b) a component of a membrane.

[0595] sequence

[0596] Sequence Description:

[0597] SEQ ID NO: 1 shows the polynucleotide sequence of wild-type E. coli CsgG from strain K12 including the signal sequence (Gene ID: 945619).

[0598] SEQ ID NO: 2 shows the amino acid sequence of wild-type E. coli CsgG including the signal sequence (Uniprot Accession No. P0AEA2).

[0599] SEQ ID NO: 3 shows the amino acid sequence of wild-type E. coli CsgG as the mature protein (Uniprot Accession No. P0EAEA2).

[0600] SEQ ID NO: 4 shows the polynucleotide sequence of wild-type E. coli CsgF from strain K12 including the signal sequence (Gene ID: 945622).

[0601] SEQ ID NO: 5 shows the amino acid sequence of wild-type E. coli CsgF including the signal sequence (Uniprot Accession No. P0AE98).

[0602] SEQ ID NO: 6 shows the amino acid sequence of wild-type E. coli CsgF as a mature protein (Uniprot Accession No. P0AE98).

[0603] SEQ ID NO: 7 shows the polynucleotide sequence of the wild-type Escherichia coli CsgF fragment encoding amino acids 1 to 27 and a C-terminal 6His tag.

[0604] SEQ ID NO: 8 shows the amino acid sequence of the wild-type E. coli CsgF fragment encompassing amino acids 1 to 27 and a C-terminal 6His tag.

[0605] SEQ ID NO: 9 shows the polynucleotide sequence of the wild-type Escherichia coli CsgF fragment encoding amino acids 1 to 38 and a C-terminal 6His tag.

[0606] SEQ ID NO: 10 shows the amino acid sequence of the wild-type E. coli CsgF fragment encompassing amino acids 1 to 38 and a C-terminal 6His tag.

[0607] SEQ ID NO: 11 shows the polynucleotide sequence of a wild-type Escherichia coli CsgF fragment encoding amino acids 1 to 48 and a C-terminal 6His tag.

[0608] SEQ ID NO: 12 shows the amino acid sequence of the wild-type E. coli CsgF fragment encompassing amino acids 1 to 48 and a C-terminal 6His tag.

[0609] SEQ ID NO: 13 shows the polynucleotide sequence of a wild-type Escherichia coli CsgF fragment encoding amino acids 1 to 64 and a C-terminal 6His tag.

[0610] SEQ ID NO: 14 shows the amino acid sequence of the wild-type E. coli CsgF fragment encompassing amino acids 1 to 64 and a C-terminal 6His tag.

[0611] SEQ ID NO: 15 shows the amino acid sequence of a peptide corresponding to residues 20 to 53 of E. coli CsgF

[0612] SEQ ID NO: 16 shows the amino acid sequence of a peptide corresponding to residues 20 to 42 of E. coli CsgF including a KD at its C-terminus.

[0613] SEQ ID NO: 17 shows the amino acid sequence of a peptide corresponding to residues 23 to 55 of the CsgF homolog Q88H88

[0614] SEQ ID NO: 18 shows the amino acid sequence of a peptide corresponding to residues 25 to 57 of the CsgF homolog A0A143HJA0

[0615] SEQ ID NO: 19 shows the amino acid sequence of a peptide corresponding to residues 21 to 53 of the CsgF homolog Q5E245. SEQ ID NO: 20 shows the amino acid sequence of a peptide corresponding to residues 19 to 51 of the CsgF homolog Q084E5. SEQ ID NO: 21 shows the amino acid sequence of a peptide corresponding to residues 15 to 47 of the CsgF homolog FOLZU2.

[0616] SEQ ID NO: 22 shows the amino acid sequence of a peptide corresponding to residues 26 to 58 of the CsgF homolog A0A136HQR0

[0617] SEQ ID NO: 23 shows the amino acid sequence of a peptide corresponding to residues 21 to 53 of the CsgF homolog A0A0W1SRL3

[0618] SEQ ID NO: 24 shows the amino acid sequence of a peptide corresponding to residues 26 to 59 of the CsgF homolog B0UH01

[0619] SEQ ID NO: 25 shows the amino acid sequence of a peptide corresponding to residues 22 to 53 of the CsgF homolog Q6NAU5

[0620] SEQ ID NO: 26 shows the amino acid sequence of a peptide corresponding to residues 7 to 38 of the CsgF homolog G8PUY5

[0621] SEQ ID NO: 27 shows the amino acid sequence of a peptide corresponding to residues 25 to 57 of the CsgF homolog A0A0S2ETP7

[0622] SEQ ID NO: 28 shows the amino acid sequence of a peptide corresponding to residues 19 to 51 of the CsgF homolog E3I1Z1. SEQ ID NO: 29 shows the amino acid sequence of a peptide corresponding to residues 24 to 55 of the CsgF homolog F3Z094.

[0623] SEQ ID NO: 30 shows the amino acid sequence of a peptide corresponding to residues 21 to 53 of the CsgF homolog A0A176T7M2

[0624] SEQ ID NO: 31 shows the amino acid sequence of a peptide corresponding to residues 14 to 45 of the CsgF homolog D2QPP8. SEQ ID NO: 32 shows the amino acid sequence of a peptide corresponding to residues 28 to 58 of the CsgF homolog N2IYT1.

[0625] SEQ ID NO: 33 shows the amino acid sequence of a peptide corresponding to residues 26 to 58 of the CsgF homolog W7QHV5

[0626] SEQ ID NO: 34 shows the amino acid sequence of a peptide corresponding to residues 23 to 55 of the CsgF homolog D4ZLW2

[0627] SEQ ID NO: 35 shows the amino acid sequence of a peptide corresponding to residues 21 to 53 of the CsgF homolog D2QT92

[0628] SEQ ID NO: 36 shows the amino acid sequence of a peptide corresponding to residues 20 to 51 of the CsgF homolog A0A167UJA2

[0629] SEQ ID NO: 37 shows the amino acid sequence of a wild-type E. coli CsgF fragment encompassing amino acids 20 to 27.

[0630] SEQ ID NO: 38 shows the amino acid sequence of a wild-type E. coli CsgF fragment encompassing amino acids 20 to 38.

[0631] SEQ ID NO: 39 shows the amino acid sequence of a wild-type E. coli CsgF fragment encompassing amino acids 20 to 48.

[0632] SEQ ID NO: 40 shows the amino acid sequence of a wild-type E. coli CsgF fragment encompassing amino acids 20 to 64.

[0633] SEQ ID NO: 41 shows the nucleotide sequence of primer CsgF_d27_terminal

[0634] SEQ ID NO: 42 shows the nucleotide sequence of primer CsgF_d38_terminal

[0635] SEQ ID NO: 43 shows the nucleotide sequence of primer CsgF_d48_terminal

[0636] SEQ ID NO: 44 shows the nucleotide sequence of primer CsgF_d64_terminal

[0637] SEQ ID NO: 45 shows the nucleotide sequence of primer pNa62_CsgF_histag_Fw

[0638] SEQ ID NO: 46 shows the nucleotide sequence of primer CsgF-His_pET22b_FW

[0639] SEQ ID NO: 47 shows the nucleotide sequence of primer CsgF-His_pET22b_Rev

[0640] SEQ ID NO: 48 shows the nucleotide sequence of primer csgEFG_pDONR221_FW

[0641] SEQ ID NO: 49 shows the nucleotide sequence of primer csgEFG_pDONR221_Rev

[0642] SEQ ID NO: 50 shows the nucleotide sequence of primer Mut_csgF_His_FW

[0643] SEQ ID NO: 51 shows the nucleotide sequence of primer Mut_csgF_His_Rev

[0644] SEQ ID NO: 52 shows the nucleotide sequence of primer DelCsgE_Rev

[0645] SEQ ID NO: 53 shows the nucleotide sequence of primer DelCsgE FW

[0646] SEQ ID NO: 54 shows the amino acid sequence of residues 1 to 30 of mature E. coli CsgF

[0647] SEQ ID NO: 55 shows the amino acid sequence of residues 1 to 35 of mature E. coli CsgF

[0648] SEQ ID NO: 56 shows the amino acid sequence of a mutated (T4C / N17S) CsgF sequence having a signal sequence and a TEV protease cleavage site (ENLYFQS) inserted between residues 35 and 36 of the mature protein sequence.

[0649] SEQ ID NO: 57 shows the amino acid sequence of a mutated (N17S-Del) CsgF sequence having a signal sequence and a TEV protease cleavage site (ENLYFQS) inserted between residues 35 and 36 of the mature protein sequence.

[0650] SEQ ID NO: 58 shows the amino acid sequence of a mutated (G1C / N17S) CsgF sequence having a signal sequence and a TEV protease cleavage site (ENLYFQS) inserted between residues 35 and 36 of the mature protein sequence.

[0651] SEQ ID NO: 59 shows the amino acid sequence of a mutant (G1C) CsgF sequence having a signal sequence and a TEV protease cleavage site (ENLYFQS) inserted between residues 35 and 36 of the mature protein sequence.

[0652] SEQ ID NO: 60 shows a protein having a signal sequence, a TEV protease cleavage site (ENLYFQS) inserted between residues 45 and 46 of the mature protein sequence, and a His-terminal 10 The amino acid sequence of the tagged CsgF sequence.

[0653] SEQ ID NO: 61 shows a protein having a signal sequence, a TEV protease cleavage site (ENLYFQS) inserted between residues 35 and 36 of the mature protein sequence, and a His-terminal 10 The amino acid sequence of the tagged CsgF sequence.

[0654] SEQ ID NO: 62 shows a protein having a signal sequence, a TEV protease cleavage site (ENLYFQS) inserted between residues 30 and 31 of the mature protein sequence, and a His-terminal 10 The amino acid sequence of the tagged CsgF sequence.

[0655] SEQ ID NO: 63 shows a protein having a signal sequence, a TEV protease cleavage site (ENLYFQS) inserted between residues 45 and 51 of the mature protein sequence, and a His-terminal 10 The amino acid sequence of the tagged CsgF sequence.

[0656] SEQ ID NO: 64 shows a protein having a signal sequence, a TEV protease cleavage site (ENLYFQS) inserted between residues 30 and 37 of the mature protein sequence, and a His-terminal 10 The amino acid sequence of the tagged CsgF sequence.

[0657] SEQ ID NO: 65 shows a HCVC3 protein having a signal sequence, a cleavage site for HCVC3 protease (LEVLFQGP) inserted between residues 34 and 36 of the mature protein sequence, and a His-terminal 10 The amino acid sequence of the tagged CsgF sequence.

[0658] SEQ ID NO: 66 shows a HCVC3 protease cleavage site (LEVLFQGP) inserted between residues 42 and 43 of the mature protein sequence and a His-terminal 10 The amino acid sequence of the tagged CsgF sequence.

[0659] SEQ ID NO: 67 shows a HCVC3 protein having a signal sequence, a cleavage site for HCVC3 protease (LEVLFQGP) inserted between residues 38 and 47 of the mature protein sequence, and a His-terminal 10 The amino acid sequence of the tagged CsgF sequence.

[0660] SEQ ID NO: 68 shows the amino acid sequence of YP_001453594.1: 1-248 of the hypothetical protein CKO_02032 [Citrobacter koseri ATCC BAA-895], which is 99% identical to SEQ ID NO: 3.

[0661] SEQ ID NO:69 shows the amino acid sequence of WP_001787128.1:16-238 of curli production assembly / transport component CsgG (partial) [Salmonella enterica], which has 98% identity with SEQ ID NO:3.

[0662] SEQ ID NO:70 shows the amino acid sequence of KEY44978.1|:16-277 of the curli production assembly / transport protein CsgG [Citrobacter amalonaticus], which is 98% identical to SEQ ID NO:3.

[0663] SEQ ID NO:71 shows the amino acid sequence of YP_003364699.1:16-277 of the curli production assembly / transport component (partial [Citrobacter rodentium ICC168]), which is 97% identical to SEQ ID NO:3. ...

Claims

1. A truncated CsgF peptide, wherein the truncated CsgF peptide lacks at least a portion of the C-terminal head domain of CsgF and the neck domain of CsgF, and the truncated CsgF peptide comprises a CsgG-binding region and a region that forms a constriction segment in a pore comprising CsgG and the truncated CsgF peptide, and the sequence of the truncated CsgF peptide is: SEQ ID NO:12, SEQ ID NO:14, SEQ ID NO:39, SEQ ID NO:54, SEQ ID NO:40 or SEQ ID NO:55; or SEQ ID NO:55 with an N17S or N17V substitution.

2. A pore comprising a CsgG pore and the truncated CsgF peptide according to claim 1, wherein the truncated CsgF peptide binds to CsgG and forms a constriction segment in the pore.

3. The pore according to claim 2, wherein the truncated CsgF peptide is inserted into the lumen of the CsgG pore.

4. The pore according to claim 2 or 3, wherein the CsgG pore comprises 6 to 10 CsgG monomers.

5. The pore according to any one of claims 2 to 4, wherein the ratio of CsgG monomers to the truncated CsgF peptide in the pore is 1:

1.

6. The pore according to any one of claims 2 to 5, wherein the CsgF peptide and the CsgG pore are covalently coupled, and wherein the covalent coupling is by means of: (i) a cysteine residue at a position corresponding to positions 132, 133, 136, 138, 140, 142, 144, 145, 147, 149, 151, 153, 155, 183, 185, 187, 189, 191, 201, 203, 205, 207 or 209 of SEQ ID NO:3 or its homolog; or (ii) a non-natural reactive or photoreactive amino acid at a position corresponding to positions 132, 133, 136, 138, 140, 142, 144, 145, 147, 149, 151, 153, 155, 183, 185, 187, 189, 191, 201, 203, 205, 207 or 209 of SEQ ID NO:3 or its homolog.

7. The pore according to any one of claims 2 to 6, wherein the CsgG pore comprises at least one monomer, and the at least one monomer comprises one or more modifications corresponding to the following modifications in SEQ ID NO:3: (i) modification at one or more positions Y51, N55 and F56, and optionally at least one substitution selected from Y51A / I / V / S / T, N55A / I / V / S / T and F56 / A / I / V / S / T / Q; (ii) at least one substitution selected from R97W or R97Y and R93W or R93Y; (iii) deletion of V105, A106 and I107 of SEQ ID NO:3; (iv) Deletion of one or more of positions R192, F193, I194, D195, Y196, Q197, R198, L199 and E201 of SEQ ID NO:3, or deletion of D195, Y196, Q197, R198 and L199 of SEQ ID NO:3; (v) At least one substitution selected from K94N / Q / R / F / Y / W / L / S, D43S, E44S, F48S / N / Q / Y / W / I / V / H / R / K, Q87N / R / K, N91K / R, R97F / Y / W / V / I / K / S / Q / H, E101I / L / A / H, N102K / Q / L / I / V / S / H, R110F / G / N, Q114R / K, R142Q / S, T150Y / A / V / L / S / Q / N; (vi) Modification of one or more of positions I41, R93, A98, Q100, G103, T104, A106, I107, N108, L113, S115, T117, Y130, K135, E170, S208, D233, D238, E244, Q42, E44, L90, N91, I95, A99, E101 and Q114; (vii) Deletion of V105, A106 and I107; (viii) At least one substitution selected from Q42K or Q42R; E44N or E44Q; L90R or L90K; N91R or N91K; I95R or I95K; A99R or A99K; E101H, E101K, E101N, E101Q or E101T; and / or Q114K; and / or (ix) Substitution N55V.

8. The pore according to any one of claims 2 to 7, wherein the CsgG pore comprises at least one monomer that comprises R or K at a position corresponding to R192 of SEQ ID NO:

3.

9. A method for generating a pore according to any one of claims 2 to 8, the method comprising co-expressing one or more CsgG monomers and a CsgF peptide in a host cell, thereby allowing formation of a transmembrane pore complex in the cell, or contacting one or more purified CsgG monomers with a modified CsgF peptide, thereby allowing formation of the pore in vitro.

10. The method for generating a pore according to claim 9, wherein the method comprises expressing a modified CsgF peptide comprising an enzyme cleavage site that is positioned such that cleavage of the CsgF peptide produces a truncated CsgF peptide comprising a CsgG binding region and a region that forms a constriction segment in the pore, and wherein the method comprises cleaving the CsgF peptide.

11. A method for determining the presence, absence or one or more characteristics of a target polynucleotide, the method comprising the steps of: (i) contacting the target polynucleotide with a pore according to any one of claims 2 to 8 such that the target polynucleotide moves into the pore complex; and (ii) making one or more measurements as the polynucleotide moves through the pore complex to thereby determine the presence, absence or one or more characteristics of the polynucleotide.

12. A kit for characterizing a target analyte, comprising (a) a pore according to any one of claims 2 - 8 and (b) components of a membrane.

Citation Information

Patent Citations

  • electric kettle

    CN3361190D

  • A miniature support for thin films containing single channels or nanopores and methods for using same

    WO2000028312A1

  • Formation of layers of amphiphilic molecules

    WO2009077734A2

  • Enzyme-PORE constructs

    WO2010004265A1

  • Base-detecting pore

    WO2010004273A1