Novel protein pores
By introducing modified CsgF peptides into CsgG pores, a nanopore sensing platform with two consecutive readout heads is formed, which solves the problem of poor current feature sequence dependence in existing technologies and improves the performance and accuracy of multinucleotide sequencing systems.
Patent Information
- Application Number
- CN202310895585.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2017-06-30
- Filing Date
- 2018-07-02
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2038-07-02
AI Technical Summary
In the prior art, the current characteristics of CsgG wells exhibit poor sequence dependence in multinucleotide sequencing, which affects the performance of the sequencing system.
By using modified CsgF peptides to bind to CsgG pores, a nanopore sensing platform with two consecutive readout heads is formed, which increases the contact area between the analyte and the pore and forms a second interaction site, thereby improving the sequence dependence of current characteristics.
It improves the performance of multinucleotide sequencing systems, enhances the sequence dependence of current characteristics, and improves sequencing accuracy and efficiency.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
[0001] This is a divisional application of patent application no. 2018800440243, filed on 2 July 2018, entitled “Novel protein pore”. TECHNICAL FIELD
[0002] The present invention relates to novel protein pores and their use in analyte detection and characterisation. The invention also relates to a transmembrane pore complex and a method for producing the pore complex, and their use in molecular sensing and nucleic acid sequencing applications. BACKGROUND
[0003] Nanopore sensing is a method of analyte detection and characterisation that relies on observing single binding or interaction events between an analyte molecule and an ion-conducting channel. A nanopore sensor can be created by placing a single nanometre-sized pore in an electrically insulating membrane and measuring the voltage-driven ionic current through the pore in the presence of an analyte molecule. The presence of an analyte inside or near the nanopore changes the flow of ions through the pore, resulting in a changed ionic or electrical current being measured across the channel. The identity of the analyte is revealed by its unique current signature, particularly by the duration and extent of the current block and changes in the current level during the analyte-pore interaction. Analytes can be small organic and inorganic molecules, as well as various biological or synthetic macromolecules and polymers, including polynucleotides, polypeptides and polysaccharides. Nanopore sensing can reveal the identity of the sensed analyte and perform single-molecule counting of the sensed analyte, but can also provide information about the composition of the analyte, such as nucleotide, amino acid or glycan sequence, and the presence of base, amino acid or glycan modifications, such as methylation and acylation, phosphorylation, hydroxylation, oxidation, reduction, glycosylation, decarboxylation, deamination, etc. Nanopore sensing has the potential to enable rapid and low-cost polynucleotide sequencing, providing single-molecule sequence reads of tens to tens of thousands of bases in length.
[0004] Two essential components for polymer characterization using nanopore sensing are: (1) control of the movement of the polymer through the pore; (2) discrimination of the constituent building blocks as the polymer moves through the pore. During nanopore sensing, the narrowest part of the pore forms the read head, and the most discriminating part of the nanopore varies with the analyte passing through in terms of current signatures. CsgG was identified as a gateless, non-selective protein secretion channel from Escherichia coli (Goyal et al., 2014) and has been used as a nanopore to detect and characterise analytes. Mutations to the wild-type CsgG pore that improve the properties of the pore in this context have also been disclosed (WO2016 / 034591, WO2017 / 149316, WO2017 / 149317 and WO2017 / 149318, PCT / GB2018 / 051191, all incorporated herein by reference).
[0005] For analytes that are polynucleotides, discrimination of the nucleotides is achieved via passage through such a mutant pore, but it has been demonstrated that the current signature has sequence dependence and that multiple nucleotides contribute to the observed current, such that the height of the constriction of the channel and the extent of surface-analyte interaction can influence the relationship between the observed current and the polynucleotide sequence. While the mutation through the CsgG pore improves the current range for nucleotide discrimination, if the current difference between nucleotides could be further improved, then the sequencing system would have higher performance. Therefore, there is a need to identify novel methods to improve nanopore sensing properties. SUMMARY
[0006] The present disclosure relates to modified CsgF peptides, in particular truncated CsgF fragments, which bind to CsgG pores and thereby introduce a further additional channel or constriction within the CsgG pore. Further aspects of the invention also relate to isolated transmembrane pore complexes, and the use of the CsgG:CsgF complexes and modified CsgF peptides or fragments in nanopore sensing platforms with two consecutive read heads.
[0007] A first aspect of the invention relates to a pore comprising a CsgG pore and a CsgF peptide. In one aspect, the CsgF peptide comprises a CsgG binding region and a region that forms a constriction in the pore. In one aspect, the CsgF peptide is a truncated CsgF peptide that lacks the C-terminal head domain of CsgF. In another aspect, the CsgF peptide is a truncated CsgF peptide that lacks the C-terminal head domain and a portion of the neck domain of CsgF. In another aspect, the CsgF peptide is a truncated CsgF peptide that lacks the C-terminal head domain and the neck domain of CsgF. The pore is also referred to herein as a pore complex and an isolated pore complex. The isolated pore complex comprises a CsgG pore or a homolog or mutant thereof, and a modified CsgF peptide or a homolog or mutant thereof, in particular a truncated CsgF fragment or a homolog or mutant thereof. In one embodiment, the modified CsgF peptide or a homolog or mutant thereof is located in the lumen of the CsgG pore or a homolog or mutant thereof. In another embodiment, the isolated pore complex has two or more channel constrictions, one channel constriction is located or provided by the CsgG pore, formed by its constriction ring, and the other additional channel constriction or read head is introduced by the modified CsgF peptide or a homolog or mutant thereof. In one embodiment, the CsgG pore or CsgG-like pore is not a wild-type pore, it is a mutated CsgG pore, in particular embodiments, for example, there is a mutation in the channel constriction ring. In another embodiment, the isolated pore complex comprising a modified CsgF peptide or a homolog or mutant thereof has a CsgF channel constriction with a diameter in the range of 0.5 nm to 2.0 nm. In one embodiment, the pore complex comprises: (i) a CsgG pore comprising a first opening, an intermediate segment comprising a beta barrel, a second opening, and a lumen extending from the first opening through the intermediate segment to the second opening, wherein an inner lumen surface of the intermediate segment defines a CsgG constriction; and (ii) a plurality of modified CsgF peptides, each modified CsgF peptide having a CsgF constriction region and a CsgG binding region (also referred to herein as a CsgG binding domain or binding region of CsgF), wherein the modified CsgF peptides form a CsgF constriction within the beta barrel of the CsgG pore, and wherein the CsgG constriction and the CsgF constriction are coaxially spaced apart within the beta barrel of the CsgG pore. The inner lumen surface of the CsgG pore can comprise one or more loop regions of the CsgG monomer that define the CsgG constriction. The CsgF constriction region and the CsgF binding region generally correspond to the N-terminal portion of the CsgF mature peptide. In one embodiment, the pore complex does not include CsgA, CsgB, and CsgE.
[0008] In a second aspect, the present application relates to a modified CsgF peptide or homologue or mutant thereof, wherein the protein or peptide is modified by truncation or deletion of part of the protein, resulting in a CsgF fragment of SEQ ID NO: 6 or a homologue or mutant thereof. One embodiment relates to a modified or truncated CsgF peptide, or a modified peptide of a CsgF homologue or mutant, said modified peptide comprising SEQ ID NO: 39 or SEQ ID NO: 40, or a homologue or mutant thereof, alternatively said modified peptide comprising SEQ ID NO: 15, or a homologue or mutant thereof, alternatively comprising SEQ ID NO: 54, or SEQ ID NO: 55, or a homologue or mutant thereof. Another embodiment discloses a modified CsgF peptide, wherein one or more positions in the region comprising SEQ ID NO: 15 are modified, and wherein the mutation requires maintaining at least 35% amino acid identity to SEQ ID NO: 15 in the peptide fragment corresponding to the region comprising SEQ ID NO: 15.
[0009] One embodiment relates to a pore comprising a CsgG pore and a modified CsgF peptide, wherein the modified CsgF peptide binds to CsgG and forms a constriction in the pore
[0010] One embodiment relates to a polynucleotide encoding the modified CsgF peptide or homologue or mutant thereof according to the second aspect of the application. In another embodiment, an isolated pore complex comprising a CsgG pore and a modified CsgF peptide or homologue or mutant thereof is characterized in that the modified CsgF peptide is the one provided by the peptide disclosed in the second aspect of the application.
[0011] Another embodiment relates to an isolated pore complex, wherein a modified CsgF peptide and a CsgG pore or monomer of the pore, or a homologue or mutant thereof, are covalently coupled. Even more specifically, the coupling is through a cysteine residue or through a non-native reactive or photo-reactive amino acid in the CsgG monomer at a position corresponding to 132, 133, 136, 138, 140, 142, 144, 145, 147, 149, 151, 153, 155, 183, 185, 187, 189, 191, 201, 203, 205, 207, or 209 of SEQ ID NO: 3 or a homologue thereof.
[0012] One preferred embodiment relates to an isolated transmembrane pore complex or membrane composition comprising the isolated pore complex of the application and components of a membrane. In particular, the transmembrane pore complex or membrane composition consists of the isolated pore complex of the application and components of a membrane or insulating layer.
[0013] One embodiment relates to a method for producing a pore as disclosed herein, the method comprising co-expressing one or more CsgG monomers as disclosed herein and a CsgF peptide as disclosed herein in a host cell, thereby allowing formation of a transmembrane pore complex in the cell. The CsgF peptide can be produced by cleavage of a modified CsgF peptide or protein comprising an enzymatic cleavage site at a suitable position in the amino acid sequence.
[0014] One embodiment relates to a method for producing a pore as disclosed herein, the method comprising contacting one or more purified CsgG monomers with one or more purified modified CsgF peptide, thereby allowing formation of a pore in vitro. The modified CsgF peptide can be a peptide comprising an enzymatic cleavage site at a suitable position in the amino acid sequence, which is cleaved before or after pore formation.
[0015] A third aspect of the application relates to a method for producing said transmembrane pore complex, wherein the pore is an isolated complex formed by a CsgG pore or homologue or mutant thereof and a modified CsgF peptide or homologue or mutant thereof, the method comprising the step of co-expressing CsgG SEQ ID NO: 2 or a homologue or mutant thereof, and a modified or truncated CsgF comprising a fragment of SEQ ID NO: 5 or a homologue or mutant thereof in a suitable host, thereby allowing formation of a pore complex in vivo. In a particular embodiment, the modified CsgF peptide or homologue or mutant thereof comprises SEQ ID NO: 12 or SEQ ID NO: 14, or a homologue or mutant thereof. Alternatively, the method for producing an isolated pore complex comprises the step of contacting a CsgG monomer of SEQ ID NO: 3 or a homologue or mutant thereof with a modified CsgF peptide or homologue or mutant thereof, so as to reconstitute a pore complex in vitro. In a particular embodiment, the modified CsgF peptide of the method comprises SEQ ID NO: 15 or SEQ ID NO: 16, or a homologue or mutant thereof.
[0016] A further aspect of the application relates to a method for determining the presence, absence or one or more characteristics of a target analyte, the method comprising the steps of:
[0017] (i) contacting the target analyte with said isolated pore complex or transmembrane pore complex, such that the target analyte moves into the pore channel; and
[0018] (ii) making one or more measurements as the analyte moves through the pore channel, thereby determining the presence, absence or one or more characteristics of said analyte.
[0019] In one embodiment, the analyte is a polynucleotide. In particular, the method using a polynucleotide as an analyte can alternatively comprise determining one or more features selected from (i) the length of the polynucleotide, (ii) the identity of the polynucleotide, (iii) the sequence of the polynucleotide, (iv) the secondary structure of the polynucleotide and (v) whether the polynucleotide is modified.
[0020] In another embodiment, the analyte is a protein or a peptide and in other embodiments, the analyte is a polysaccharide or a small organic or inorganic compound, such as, but not limited to, a pharmacologically active compound, a toxic compound and a pollutant.
[0021] In another embodiment, a method of characterizing a polynucleotide or (poly)peptide using an isolated transmembrane pore complex is described, wherein the pore complex is an isolated complex comprising a CsgG pore or a homolog or mutant thereof and a modified CsgF peptide or a homolog or mutant thereof. In particular, the CsgG pore or a homolog or mutant thereof, comprises six to ten CsgG monomers forming a channel of the CsgG pore.
[0022] Another aspect of the present invention discloses the use of the isolated pore complex or transmembrane pore complex according to the preceding aspects of the present invention for determining the presence, absence or one or more features of a target analyte. Furthermore, the present invention also relates to a kit for characterizing a target analyte comprising (a) the isolated pore complex and (b) components of a membrane. BRIEF DESCRIPTION OF DRAWINGS
[0023] The described drawings are merely schematic and are non-limiting. In the drawings, the size of some of the elements can have been exaggerated and are not drawn on scale for the sake of clarity.
[0024] Figure 1 : Structure of the CsgG pore and interface with CsgF to form a complexFigure 1. Surface (A) and ribbon (B, C) representations of a cross-sectional view (A), side view (B), and top view (C) of a CsgG oligomer (e.g., nonamer) (gold) with individual CsgG precursors in light blue (D) (based on CsgG X-ray structure PDB entry: 4uv3). The CsgG contractile ring (CL ring) spans residues 46 to 61 of SEQ ID NO: 3, is represented in dark gray in all figures, and corresponds to the loops provided in the lower left of (E). CsgG residues with side chains facing the lumen of the CsgG b -barrel are shown in medium gray and are labeled in the b-strands of (E) and (D). These residues represent sites that can be used to substitute natural or non-natural amino acids, e.g., sites suitable for attachment (e.g., covalent cross-linking) of pore-resident peptides (including, e.g., modified CsgF peptides or homologs thereof) to the CsgG pore or monomer. In some embodiments, cross-linking residues include Cys as well as reactive and photoreactive amino acids such as azidohomoalanine, homopropargylglycine, homoallylglycine, p-acetyl-Phe, p-azido-Phe, p- propargyloxy-Phe, and p-benzoyl-Phe (Wang et al., 2012; Chin et al., 2002), and can be substituted into positions 132, 133, 136, 138, 140, 142, 144, 145, 147, 149, 151, 153, 155, 183, 185, 187, 189, 191, 201, 203, 205, 207, or 209 according to SEQ ID NO: 3. (E) shows an enlarged view of the CL ring and the transmembrane b-strands of a CsgG monomer. The CsgG contractile ring (dark blue) forms the orifice or narrowest passage in the CsgG pore (Figure A). In some embodiments, three positions in the CL ring, 56, 55, and 51 according to SED ID NO: 3, are particularly important for the diameter of the CsgG channel orifice or "reading head" as well as chemical and physical properties. These represent preferred positions for altering the nanopore sensing properties of CsgG pores and homologs.
[0025] Figure 2: Co-expression of CsgG:CsgF complex proteins and purification of the complex(A) Schematic of the purification protocol for CsgG:CsgF complexes starting from E. coli cultures co-expressing CsgG (SEQ ID NO: 2 + C-terminal StrepII tag) and CsgF (SEQ ID NO: 4 + C-terminal 6xHis tag). The protocol involves disruption of resuspended cells and 1% DDM extraction of membrane-bound proteins. CsgG:CsgF complexes and excess CsgF are first enriched by affinity purification on a nickel IMAC chromatography column, followed by a second affinity-based CsgG:CsgF complex enrichment on a streptavidin chromatography column. (B) Coomassie-stained SDS-PAGE of the IMAC (left) and streptavidin (right) purification steps. Protein bands corresponding to CsgG and CsgF are labelled. Of note, the IMAC eluate contains an N-terminally truncated CsgF fragment (labelled *) that is not retained in the affinity pull-down using the CsgG-bound Strep tag, indicating that CsgF N-terminus is required for complex formation with CsgG.
[0026] Figure 3: In vitro Purification of reconstituted CsgG:CsgF complex proteins. (A) Superimposed chromatograms of size exclusion chromatography (SEC) runs (using a BioRad Enrich 650 10 / 300 chromatography column) of CsgG (light grey) and CsgG supplemented with excess CsgF (dark grey). The chromatograms show elution peaks corresponding to: CsgG 9-mer (a) and CsgG 18-mer (b) for CsgG chromatographic analysis; and excess free CsgF (c), and 9-mer CsgG:CsgF complex (d) and 18-mer CsgG:CsgF complex (e) eluting at higher hydrodynamic radius (molecular weight) due to incorporation of CsgF into the complex. (B) Native PAGE analysis of representative species labelled in panel (A), confirming the shift to higher molecular weight due to incorporation of CsgF into CsgG 9-mer and CsgG 18-mer complexes. These experiments demonstrate that CsgG:CsgF complexes can be reconstituted in vitro starting from purified components. (C) Ribbon representation of CsgG 9-mer and CsgG 18-mer previously reported in Goyal et al. 2014 (PDB entry 4uv3). CsgG 18-mer is formed from a dimer of CsgG 9-mer. The SEC and native PAGE analysis shown in panels A and B demonstrate that CsgG 18-mer is amenable to complex formation with CsgF.
[0027] Figure 4: CsgG:CsgF structure determined in cryo-EM(A) Cryo-electron micrographs of CsgG:CsgF complexes showing the presence of 9-mer and 18-mer CsgG:CsgF complexes, with many individual particles of the 9-mer and 18-mer forms highlighted with full and dashed circles, respectively. (B) Two representative class averages of CsgG:CsgF 9-mer complexes viewed from the side. The class averages include 6020 and 4159 individual particles, respectively. The class averages reveal the presence of an additional density on top of the CsgG particle, which corresponds to an oligomeric complex of CsgF. Three distinct regions can be seen in the CsgF oligomer: the "head" and "neck" regions, and a region that remains inside the CsgG b-barrel lumen and forms a constriction or narrow passage (labeled F) that stacks on top of the constriction formed by the CsgG CL ring (labeled G). The latter CsgF region is referred to as the CsgF constriction peptide (FCP).
[0028] Figure 5 Three-dimensional structural model of the CsgG:CsgF complex A cross-sectional view of the 3D cryoEM electron density of the CsgG:CsgF 9-mer complex calculated from the 20.000 particles assigned to the 21 class averages. The right panel shows the overlay of the CsgG 9-mer X-ray structure (PDB entry: 4uv3) docked into the cryoEM density. The regions corresponding to CsgG, CsgF, and the CsgF head, neck, and FCP domains are indicated. The cross-section shows that the CsgF FCP region forms an additional constriction (labeled F) in the CsgG channel, about 2 nm above the CsgG constriction ring (labeled G).
[0029] Figure 6: Schematic of the CsgG:CsgF pore complex based on the cryo-EM structure. (A) Schematic representation of the CsgG nanopore with a hidden single constriction (labeled (1)) in a cross-sectional view. The CsgG-based nanopore forms a 3.5-4 nm wide channel that contains a 0.5-1.5 nm aperture formed by the CsgG constriction ring (residues 46 to 61 according to SEQ ID NO: 3). Upon complexation with CsgF, a second constriction or aperture is introduced into the CsgG channel (labeled (2) / F), and the channel exit is blocked by the CsgF head domain (see Figure 5 ). When a modified CsgF peptide is used (e.g., a CsgF constriction peptide (FCP) that lacks the neck and head regions), a CsgG:CsgF pore complex is formed with two consecutive channel constriction segments or apertures ((1) and (2)), as shown in the cross-sectional view of the CsgG:CsgF cryo-EM density in panel (B) and the schematic in panel (C). Removing the neck and head regions in the modified CsgF peptide alleviates their blockage of the channel exit.
[0030] Figure 7: Nanopore sensing of (bio)polymers (A) or single molecule analytes (B) by the CsgG:CsgF pore complex Schematic of the use of the application.When used for polymer sensing, the second channel constriction introduced by the modified CsgF peptide increases the contact area with the analyte and forms a second interaction site and readhead. When used for single molecule nanopore sensing, the second channel constriction introduced by the modified CsgF peptide creates a second, independent analyte interaction site. (C) Schematic of the theoretical channel conductance curve for a small molecule (represented as a hexagon or triangle) passing through and interacting with a consecutive CsgG (1) and CsgF (2) constriction or readhead.
[0031] Figure 8 Multiple sequence alignment of exemplary CsgF homologues The aligned sequences are shown as mature proteins (i.e., lacking their N-terminal signal peptide (SP)). In some embodiments, the boxed sequence indicates a sequence-conserved CsgF region (pairwise sequence identity between 35 and 100% - see Figure 10 ), which corresponds to the CsgF constriction peptide (FCP). The CsgF homologs included in the multiple sequence alignment are Q88H88; A0A143HJA0; Q5E245; Q084E5; F0LZU2; A0A136HQR0; A0A0W1SRL3; B0UH01; Q6NAU5; G8PUY5; A0A0S2ETP7; E3I1Z1; F3Z094; A0A176T7M2; D2QPP8; N2IYT1; W7QHV5; D4ZLW2; D2QT92; A0A167UJA2. The FCP region of E. coli CsgF (SEQ ID NO: 15) and the indicated CsgF homologs correspond to SEQ ID Nos: 18-36.
[0032] Figure 9: Experimental evaluation of the E. coli CsgF region forming the CsgG interaction sequence and the CsgF constriction peptide (FCP).Figure (A) shows the mature sequences (i.e. after removal of the CsgF signal peptide, which corresponds to residues 1-19 of SEQ ID NO: 5) of four N-terminal CsgF fragments (SEQ ID NO: 8_CsgF residues 1-27; SEQ ID NO: 10; SEQ ID NO: 12 and SEQ ID NO: 14) co-expressed with E. coli CsgG (SEQ ID NO: 2). (B) Anti-Strep (left) and anti-His (right) Western blot analysis of SDS-PAGE runs of crude cell lysates of CsgG and CsgF co-expression experiments. The anti-strep analysis demonstrates CsgG expression in all co-expression experiments, while the anti-His Western blot analysis indicates that there are detectable levels of CsgF fragments only for the truncated mutant CsgF 1-64 (SEQ ID NO: 14). His-tagged nanobodies (Nb) were used as a positive control. (C) Anti-His dot blot analysis for the presence of CsgF fragments in CsgG:CsgF co-expression experiments. The top row shows whole cell lysates, the middle and bottom rows show the eluate and flowthrough of Strep affinity pull-down experiments. These data demonstrate that CsgF fragment 1-64, and to a lesser extent CsgF 1-48, are pulled down specifically as a complex with Strep-tagged CsgG. CsgF fragments 1-27 and 1-38 do not produce detectable levels of the corresponding CsgF fragment and show no evidence of complex formation with CsgG.
[0033] Figure 10 Multiple sequence alignment of the CsgF region forming the CsgG interaction sequence and the CsgF contraction peptide (FCP). This figure shows a multiple sequence alignment and consensus sequence of CsgF peptides and their known homologues in the region corresponding to interaction with CsgG. CsgF homologues are defined by the PFAM domain PF03783. These peptides bind to CsgG and localise to the inner lumen of the CsgG β-barrel, where they form an additional constriction in the CsgG channel. These peptides and their homologues are examples of CsgF constriction peptides or FCPs. The range of pairwise sequence identity in the FCPs shown is between 35% and 98%.
[0034] Figure 11 : High resolution cryoEM structure of the CsgG:CsgF complex CsgG is shown in light grey and CsgF in dark grey. A. CsgG:CsgF complex in Final electron density map at 3.5 A resolution. Side view. B. Top view showing the cryoEM structure of CsgG:CsgF, containing a 9:9 stoichiometry with C9 symmetry. C. Internal architecture of the CsgG:CsgF complex. GC, CsgG gate; FC, CsgF gate. D. Interactions between CsgG and CsgF proteins. CsgG and CsgG gate are shown in light and dark grey, respectively. CsgF is shown in dark grey. Residues in CsgG and CsgF are labelled in light grey and black, respectively.
[0035] Figure 12 : two read heads of the CsgG:CsgF complex. CsgG is shown in light grey, and the reading head of the CsgG pore is shown in dark grey. CsgF is shown in black, with the reading head of CsgF labelled.
[0036] Figure 13 Co-expression of CsgG with CsgF WTs in vivo . The gene encoding CsgG polypeptide with C-terminal strep tag in pT7 vector with ampicillin resistance and the gene encoding CsgF polypeptide with C-terminal His tag in pRham vector with kanamycin resistance were transformed together into E. coli BL21 DE3 cells in the presence of both ampicillin and kanamycin. Proteins were expressed overnight at 18 °C, 250 rpm, and the CsgG-CsgF complex was purified using Strep tag purification followed by His tag purification. A. Protein samples before strep purification (in duplicate). B. Protein samples after His purification (three elution fractions). Proteins were run on 4-20% Tris gels.
[0037] Figure 14 In vitro co-expression of CsgG and CsgF and thermal stability of CsgG-CsgF complex . CsgG and CsgF DNA in different vectors were co-expressed in in vitro transcription and translation reactions. Proteins were radiolabelled with S-35 methionine and exposed on X-ray film. The stability of the complex was assessed by incubating the reaction mixture at different temperatures for 10 minutes.
[0038] Figure 15: Preparation of the CsgG:CsgF complex using a protease cleavage site.A. A TEV or C3 or any other protease cleavage site can be incorporated into the desired site of the CsgF peptide (e.g. between 30 and 31, 35 and 36, 40 and 41, 45 and 46 of seq ID No. 6). CsgG is shown in gold and the CsgF domain in red. For clarity, 1-35 of one CsgF subunit is in green. 36-45 is shown in purple. The 10-histidine tag is shown in pink and the strep tag on CsgG in blue. B. SDS-PAGE (4-20% TGX) of protease cleavage of the full-length CsgG:CsgF complex with a TEV protease cleavage site inserted between 35-36 of seq ID 6. M: molecular weight marker, lane 1: CsgG:CsgF full-length complex after strep purification, lane 2: concentrate after strep purification, lane 3: after gel filtration, lane 4: cleavage with TEV protease to generate CsgG:CsgF complex, lane 5: flow through of CsgG:CsgF after strep purification, lane 6: CsgG:CsgF heated at 60°C for 10 min. Lane 7: CsgG:CsgF complex eluted from streptavidin chromatography column, lane 8: CsgG well as control, lane 9: TEV protease as control.
[0039] Figure 16: Thermostability of the CsgG:CsgF complex M: molecular weight marker, lane 1: CsgG well, lane 2: CsgG:CsgF complex at room temperature: lanes 3-9: CsgG:CsgF samples heated at different temperatures (40, 50, 60, 70, 80, 90, 100°C, respectively) for 10 min. Lane 1:
[0040] A. Y51A / F56Q / N55V / N91R / K94Q / R97W-del(V105-I107):CsgF-(1-45).
[0041] B. Y51A / F56Q / N55V / N91R / K94Q / R97W-del(V105-I107):CsgF-(1-35).
[0042] C. Y51A / F56Q / N55V / N91R / K94Q / R97W-del(V105-I107):CsgF-(1-30).
[0043] Samples were run on SDS-PAGE on 7.5% TGX gels. CsgG:CsgF complexes with both CsgF-(1-45) and CsgF-(1-35) showed a shift from the CsgG pore band in lane 1. Thus, it was clear that both complexes were thermostable up to 90°C. The complexes and the pore disintegrated into CsgG monomers at 100°C (lane 9). Although the same thermostability pattern was seen with CsgG:CsgF complexes with CsgF-(1-30), it was difficult to see the shift between the protein bands of the CsgG pore (lane 1) and the CsgG-CsgF complex (lanes 2-8).
[0044] Figure 17 : Formation of CsgG:CsgF by in vitro reconstitution using synthetic CsgF peptides. Native PAGE showing formation of CsgG:CsgF by in vitro reconstitution using wild type CsgG or CsgG mutants with altered constriction Y51A / F56Q / K94Q / R97W / R192D-del(V105-l107). Alexa 594 labeled CsgF peptide corresponding to the first 34 residues of mature CsgF (Seq ID No 6) was added to purified Strep-tagged CsgG or Y51A / F56Q / K94Q / R97W / R192D-del(V105-l107) in 50 mM Tris, 100 mM NaCl, 1 mM EDTA, 5 mM LDAO / C8D4 at a 2:1 molar ratio over 15 minutes at room temperature to allow reconstitution. After pulldown of CsgG-strep onto StrepTactin beads, samples were analyzed on native PAGE. Both WT and Y51A / F56Q / K94Q / R97W / R192D-del(V105-l107) CsgG bound the CsgF N-terminal peptide as visible by the fluorescent tag.
[0045] Figure 18: Stabilization of the CsgG:CsgF or CsgG:FCP complex. A. Amino acid positions in the CsgG (SEQ ID NO: 3) and CsgF (SEQ ID NO: 6) pair that can give rise to S-S bonds. B. Schematic showing S-S bond between CsgG-Q153C and CsgF-G1C.
[0046] Figure 19: Cysteine cross-linking of the CsgG:CsgF complex.A. Y51A / F56Q / N91R / K94Q / R97W / Q153C-del(V105-I107) and CsgF-G1C proteins were purified separately and incubated together at 4°C for 1 hour or overnight to form a complex and allow S-S formation. No oxidizing agent was added to promote S-S formation. Control CsgG pore (Y51A / F56Q / N91R / K94Q / R97W / Q153C-DEL(V105-I107)) and complex (with and without DTT) were heated at 100°C for 10 minutes to dissociate the complex into CsgG monomers (CsgG m , 30 KDa) and CsgF monomers (CsgF m , 15 KDa). In the absence of reducing agent, a dimer between CsgG m and CsgF m (CsgG m -CsgF m , 45 KDa) can be seen, confirming S-S bond formation. Incubation overnight can be seen to increase dimer formation compared to incubation for one hour. B. Mass spectrometry analysis of CsgG m -CsgF m band purified from the gel from overnight incubation. The proteins were proteolytically cleaved to produce tryptic peptides. The LC-MS / MS sequencing method was performed, identifying the precursor ions above corresponding to the linking peptides shown. The precursor ions were fragmented to give the observed fragment ions. These include ions for each peptide, as well as fragments incorporating the intact disulfide bond. This data provides strong evidence for the presence of a disulfide bond between C1 of CsgF and C153 of CsgG.
[0047] Figure 20 : increase the cysteine cross-linking efficiency of the CsgG:CsgF complex. Lane 1 : Y51A / F56Q / N91R / K94Q / R97W / N133C-del(V105-I107) and CsgF-T4C proteins co-expressed, CsgG:CsgF complex purified. Lane 2: Complex heated in the presence of DTT to dissociate the complex into substituent monomers (CsgG m and CsgF m ). If formed, DTT will break all S-S bonds between CsgG-N133C and CsgF-T4C. Lane 3: This complex was incubated with an oxidizing agent, copper-phenanthroline, to promote S-S bond formation. Lane 4: Oxidized sample was heated at 100°C in the absence of DTT to dissociate the complex. A new band of 45 KDa corresponding to CsgG m -CsgF m was seen, confirming S-S bond formation.
[0048] Figure 21 : Current signature of a DNA strand passing through the CsgG:CsgF complex.Complexes were prepared by co-expression of CsgG pore containing a C-terminal strep tag (Y51A / F56Q / N91R / K94Q / R97W-del(V105-I107)) with full length CsgF protein containing a C-terminal His tag and a TEV protease cleavage site between 35 and 36 of seq ID no. 6. The complexes were then purified and cleaved with TEV protease to make the indicated CsgG:CsgF complexes. Note that TEV cleavage leaves the sequence ENLYFQ at the cleavage site. A. Position 17 of CsgF is unmutated. B. N17S mutation in CsgF.
[0049] Figure 22: Current signature of a DNA strand passing through the CsgG:CsgF complex. Complexes were prepared by incubating Y51A / N55V / F56Q / N91R / K94Q / R97W-del(V105-I107) pore containing a C-terminal strep tag with CsgF-(1-35) mutants. A. CsgF-N17S-(1-35). B. CsgF-N17V-(1-35).
[0050] Figure 23: Current signature of a DNA strand passing through the CsgG:CsgF complex.
[0051] Complexes were prepared by incubating different CsgG pores containing a C-terminal strep tag with CsgF-N17S-(1-35). A. CsgG pore is Y51A / N55V / F56Q / N91R / K94Q / R97W-del(V105-I107). B. CsgG pore is Y51T / N55V / F56Q / N91R / K94Q / R97W-del(V105-I107). C. CsgG pore is Y51A / N55I / F56Q / N91R / K94Q / R97W-del(V105-I107). D. CsgG pore is Y51A / F56A / N91R / K94Q / R97W-del(V105-I107). E. CsgG pore is Y51A / F56I / N91R / K94Q / R97W-del(V105-I107). F. CsgG pore is Y51S / N55V / F56Q / N91R / K94Q / R97W-del(V105-I107).
[0052] Figure 24: Current signature of a DNA strand passing through the CsgG:CsgF complex.
[0053] Complexes were prepared by incubating E. coli purified Y51A / N55V / F56Q / N91R / K94Q / R97W-del(V105-I107) pore with three different lengths of CsgF. A. CsgF-(1-29), B. CsgF-(1-35), C. CsgF-(1-45). Arrows indicate the range of signals. Surprisingly, the complex with CsgF-(1-29) produced the largest range of signals.
[0054] Figure 25 : Signal to noise ratio of current signature when DNA strands pass through CsgG:CsgF complex. Different CsgG:CsgF complexes were prepared by incubating different CsgG pores (1-Y51A / F56Q / N91R / K94Q / R97W-del(V105-I107) 2-Y51A / N55I / F56Q / N91R / K94Q / R97W-del(V105-I107) 3-Y51A / N55V / F56Q / N91R / K94Q / R97W-del(V105-I107) 4-Y51A / F56A / N91R / K94Q / R97W-del(V105-I107) 5-Y51A / F56I / N91R / K94Q / R97W-del(V105-I107) 6-Y51A / F56V / N91R / K94Q / R97W-del(V105-I107) 7-Y51S / N55A / F56Q / N91R / K94Q / R97W-del(V105-I107) 8-Y51S / N55V / F56Q / N91R / K94Q / R97W-del(V105-I107) 9-Y51T / N55V / F56Q / N91R / K94Q / R97W-del(V105-I107)) with the same CsgF peptide CsgF-(1-35). Different wave pattern curves were observed in DNA translocation experiments and their signal to noise ratios were measured. The greater the signal to noise ratio, the higher the precision that can be obtained.
[0055] Figure 26: Sequencing errors produced with a narrow read head. Schematic of DNA bases interacting with the CsgG pore reading head. When a DNA strand translocates through the pore, at any given time, about 5 bases dominate the current signal. B. Signal mapping plot. Event detection signals mapped to multiple reads of simulated signals using a custom HMM for mixed sequences run with no homopolymer and sequences containing three 10T homopolymers run.
[0056] Figure 27: Mapping the CsgG:CsgF complex with a read head.A diagram showing the readhead differentiation of the CsgG:CsgF complex. The average change in simulated current when the bases at each readhead position change. To calculate the readhead differentiation at position i for a model of length k and letter length n, we define the differentiation at readhead position i as the median of the standard deviation of the current level for each of nk-1 groups of size n, where position i changes while other positions remain constant. B. Static DNA strands mapping the readheads: A set of polyA DNA strands (SS20 to SS38) were created with one base missing from the DNA backbone (iSpc3). In each strand, the position of iSpc3 was shifted from the 3' end to the 5' end. Based on previous experiments with CsgG wells, position 7 of the DNA was expected to be located within the CsgG contraction. SS26 corresponding to this DNA is highlighted. Based on the model from (A), 4–5 bases are expected to separate the CsgG and CsgF readheads. Therefore, positions approximately 12 and 13 are expected to be within the CsgF contraction. The SS31 and SS32 DNA strands corresponding to those positions are highlighted. C and D. Mapping the two readheads: the biotin modification at the 3' end of each strand is complexed with monovalent streptavidin, and the current blockage generated by each strand is recorded in the MinION apparatus. When the iSpc3 position is located above or below the constriction segment within the well, no deflection is expected. However, when the iSpc3 is located within the constriction segment, a higher current level is expected through the well—the extra space created by the lack of bases allows more ions to pass through. Therefore, the positions of the two readheads can be mapped by plotting the current passing through each DNA strand. As expected, the maximum current deflection is seen when position 7 of the DNA strand is occupied by iSpc3 (C). iSpc3 at positions 6 and 8 also produces a higher deflection than the average polyA current level. Therefore, positions 6, 7, and 8 of the DNA strand represent the first readhead—the CsgG readhead. As expected, another deviation from the baseline polyA is observed when positions 12 and 13 are occupied by iCsp3 (D). This indicates that the second read head of the well is the CsgF read head. The results also confirmed that the two read heads are approximately 4-5 bases apart.
[0057] Figure 28: Read head discrimination and base contribution. The left-hand figure illustrates the readhead difference for each mutation well: the average change in simulated current as the bases at each readhead position change. To calculate the readhead difference at position i for a model of length k and letter length n, we define the difference at readhead position i as an n-fold difference of size n. k-1 The median of the standard deviation of the current level for each of the groups, where position i changes while the others remain constant. The right-hand plot shows the base contribution map: the median current for all sequence cases with base b (A, T, G, or C) at position i of the read head.
[0058] Figure 29: Error profile of a dual read head pore. A. Schematic of the CsgG:CsgF complex and the interaction of DNA bases with both read heads. Red: strong interaction, orange: weak interaction, grey: no interaction. B. Comparison of deletion errors. Reads from Y51A / F56Q / N91R / K94Q / R97W / R192D-del(V105-I107) and Y51A / N55V / F56Q / N91R / K94Q / R97W-del(V105-I107): CsgF-N17S-(1-35) pores: base calls from CsgG:CsgF pores on the same region of E. coli DNA. Reads were aligned to the reference genome using Minimap2 (https: / / arxiv.org / abs / 1708.01492) and the resulting alignments were visualized in the Savant Genome Browser (https: / / www.ncbi.nlm.nih.gov / pubmed / 20562449). Most Y51A / F56Q / N91R / K94Q / R97W / R192D-del(V105-I107) reads contained a single base deletion in the T homopolymer (black box), no base deletions were present in most CsgG:CsgF reads. C. Comparison of consensus accuracy from raw data generated by Y51A / F56Q / N91R / K94Q / R97W / R192D-del(V105-I107) (blue) and Y51A / N55V / F56Q / N91R / K94Q / R97W-del(V105-I107): CsgF-N17S-(1-35) pores (green) versus homopolymer length.
[0059] Figure 30: Homopolymer recognition of the CsgG:CsgF complexDNA with the sequence shown in (A) translocates through Y51A / F56Q / N91R / K94Q / R97W / -del(V105-I107) (B) and Y51A / N55V / F56Q / N91R / K94Q / R97W / -del(V105-I107):CsgF-N17S-(1-35) (C) and their signals are analyzed against the first polyT section shown in red in (A). When the polyT section passes through a CsgG pore containing a single readhead (model based on 5 bases located in the readhead), it will produce a flat line in the signal. Thus, it is difficult to determine the exact number of bases in this region that typically cause deletion errors. When the DNA passes through a CsgG:CsgF complex containing two readheads (model based on 9 bases located inside and between the two readheads), the polyT section shows multiple steps instead of a flat line. The information in these steps can be used to correctly identify the number of bases in the homopolymer region. This additional information can significantly reduce deletion errors and improve overall consensus accuracy.
[0060] Figure 31 : CsgG pore Characterization of (Y51A / F56Q / N91R / K94Q / R97W / -del(V105-I107). A. Readhead discrimination of CsgG pore. The average change in current when the base at each readhead position is changed. To calculate the readhead discrimination of a model of length k and alphabet length n at position i, we define the discrimination at readhead position i as the median of the standard deviations of the current levels of each of the n k-1 groups of size n, where position i is changed while the other positions remain the same. B. Base contribution plot of CsgG pore. The median current on all k-mers with base b (A, T, G, or C) at position i of the readhead. C. Current signature of DNA strand passing through CsgG pore. DETAILED DESCRIPTION
[0061] The application will be described with respect to particular embodiments and with reference to certain drawings but the application is not limited thereto but only by the claims. No reference signs in the claims are to be interpreted as limiting the scope. It will be understood that any particular embodiment disclosed is not necessarily to be construed as being limiting but rather a description of one or more specific embodiments falling within the scope of the application. Thus, for example, it will be appreciated that the application can be embodied in a variety of ways, including but in no way limited to, the ways as taught herein. It will also be appreciated by persons skilled in the art that the present application is not limited to what has been particularly shown and described hereinabove. It will be appreciated that features and / or
[0062] When read in conjunction with the accompanying drawings, the organization and operation of the invention, as well as its features and advantages, can be best understood by referring to the following detailed description. Aspects and advantages of the invention will become apparent and elucidated with reference to the embodiments described below. Throughout this specification, the phrase "an embodiment" or "an embodiment" means that a particular feature, structure, or characteristic described in connection with that embodiment is included in at least one embodiment of the invention. Therefore, the appearance of the phrase "in an embodiment" or "in an embodiment" throughout this specification does not necessarily refer to the same embodiment, but may refer to the same embodiment. Similarly, it should be recognized that in the description of exemplary embodiments of the invention, various features of the invention are sometimes grouped together in a single embodiment, drawing, or description thereof to simplify the disclosure and aid in understanding one or more aspects of the invention. However, this approach of the disclosure should not be construed as reflecting an intention that the claimed invention requires more features than are expressly recited in each claim. Rather, as reflected in the appended claims, inventive aspects do not consist of all the features of the single embodiments disclosed above.
[0063] Furthermore, as used in the specification and appended claims, unless otherwise expressly indicated, the singular forms “a,” “an,” and “the” include plural indicators. Thus, for example, reference to “a polynucleotide” includes two or more polynucleotides, reference to “a polynucleotide-binding protein” includes two or more such proteins, reference to “a helicase” includes two or more helicases, reference to “a monomer” refers to two or more monomers, reference to “a pore” includes two or more pores, and so on.
[0064] Throughout the discussion in this paper, the standard single-letter code for amino acids was used. These are as follows: alanine (A), arginine (R), asparagine (N), aspartic acid (D), cysteine (C), glutamic acid (E), glutamine (Q), glycine (G), histidine (H), isoleucine (I), leucine (L), lysine (K), methionine (M), phenylalanine (F), proline (P), serine (S), threonine (T), tryptophan (W), tyrosine (Y), and valine (V). The standard substitution symbol was also used; Q42R signifies that the Q at position 42 is replaced by R.
[0065] In this article, paragraphs separated by the slash " / " at specific positions mean "or". For example, Q87R / K means either Q87R or Q87K.
[0066] In this article, paragraphs separated by the slash " / " are used to indicate "and", making Y51 / N55 Y51 and N55.
[0067] Unless specified to the contrary, all amino acid substitutions, deletions and / or additions disclosed herein refer to mutant CsgG monomers comprising a variant of the sequence set forth in SEQ ID NO: 3.
[0068] Reference to mutant CsgG monomers comprising a variant of the sequence set forth in SEQ ID NO: 3 encompasses mutant CsgG monomers comprising a variant of the sequence set forth in other SEQ ID NOs disclosed below. Amino acid substitutions, deletions and / or additions can be made to CsgG monomers comprising a variant of a sequence other than that set forth in SEQ ID NO: 3 that are equivalent to those substitutions, deletions and / or additions disclosed herein for mutant CsgG monomers comprising a variant of the sequence set forth in SEQ ID NO: 3.
[0069] All publications, patents and patent applications cited herein, whether supra or infra, are hereby incorporated by reference in their entirety.
[0070] Definitions
[0071] Where a singular term is used in the specification and claims, such term should be read to encompass both the singular and plural forms unless specifically stated otherwise. Where the term comprising is used in the specification and claims, other elements or steps not specifically stated can be included. Furthermore, the terms first, second, third and the like are used merely as labels to distinguish between similar elements and are not necessarily used in a sequential or chronological sense. It is to be understood that the terms so used are interchangeable under appropriate circumstances. The embodiments of the application described herein are intended to be merely illustrative of possible embodiments of the application and do not limit the scope of the application as otherwise expressed by the appended claims. The following terms or definitions are provided solely to aid in the understanding of the application. Unless specifically defined herein, all terms used herein have the same meaning as is commonly understood by one of ordinary skill in the art to which this application belongs. Practitioners are particularly directed to Sambrook et al., Molecular Cloning: A Laboratory Manual, 4th Ed., Cold Spring Harbor Press, Plainsview, New York (2012); and Ausubel et al., Current Protocols in Molecular Biology (Supplement 114), John Wiley & Sons, New York (2016) for definitions and terms of art. None of the definition(s) provided herein should be construed as limiting the scope of the application to only the specifically identified embodiments.
[0072] When referring to measurable values such as amounts, time periods, and the like, "about" as used herein is meant to encompass variations of ±20% or ±10%, more preferably ±5%, even more preferably ±1 and yet more preferably ±0.1% from the specified value, as such variations are appropriate to perform the disclosed methods.
[0073] "Nucleotide sequence," "DNA sequence," or "nucleic acid molecule" as used herein refers to a polymeric form of nucleotides of any length, either ribonucleotides or deoxyribonucleotides. The term refers only to the primary structure of the molecule. Thus, the term includes double-stranded and single-stranded DNA and RNA. The term "nucleic acid" as used herein is a single- or double-stranded covalently linked sequence of nucleotides, in which the 3' and 5' ends of each nucleotide are linked by a phosphodiester bond. A polynucleotide can consist of deoxyribonucleotide bases or ribonucleotide bases. Nucleic acids can be prepared synthetically or isolated from natural sources. Nucleic acids can also include modified DNA or RNA, for example, DNA or RNA that has been methylated, or RNA that has been post-translationally modified, for example, 5'-capping with 7-methylguanosine, 3 '-processing such as cleavage and polyadenylation, and splicing. Nucleic acids can also include synthetic nucleic acids (XNAs) such as hexitol nucleic acid (HNA), cyclohexene nucleic acid (CeNA), threose nucleic acid (TNA), glycerol nucleic acid (GNA), locked nucleic acid (LNA), and peptide nucleic acid (PNA). The size of a nucleic acid, which is also referred to herein as a "polynucleotide," is generally expressed as the number of base pairs (bp) for double-stranded polynucleotides, or in the case of single-stranded polynucleotides, as the number of nucleotides (nt). One thousand bp or nt equals one kilobase (kb). Polynucleotides of less than about 40 nucleotides in length are often referred to as "oligonucleotides," and can comprise primers for DNA manipulation, such as by polymerase chain reaction (PCR).
[0074] As used herein, "gene" includes the promoter region of the gene as well as the coding sequences. It refers both to genomic sequence (including possible introns) and to the cDNA (complementary DNA) derived from messenger RNA that is operably linked to a promoter sequence.
[0075] A "coding sequence" is a nucleotide sequence, which when under the control of appropriate control sequences, is transcribed into mRNA and / or translated into a polypeptide. The boundaries of the coding sequence are determined by a translation start codon at the 5'-terminus and a translation stop codon at the 3'-terminus. A coding sequence can include, but is not limited to, mRNA, cDNA, recombinant nucleotide sequences or genomic DNA, and in certain instances can include introns.
[0076] In the context of the present disclosure, the term "amino acid" is used in its broadest sense and is meant to include organic compounds that contain an amine (NH2) and carboxyl (COOH) functional group, as well as a side chain (e.g., R group) that is characteristic of each amino acid. In some embodiments, amino acid refers to naturally occurring L a-amino acids or residues. The commonly used one- and three-letter abbreviations for naturally occurring amino acids are used herein: A = Ala; C = Cys; D = Asp; E = Glu; F = Phe; G = Gly; H = His; I = lie; K = Lys; L = Leu; M = Met; N = Asn; P = Pro; Q = Gin; R = Arg; S = Ser; T = Thr; V = Val; W = Trp; and Y = Tyr (Lehninger, A. L., (1975) Biochemistry, 2nd Ed., pp. 71-92, Worth Publishers, New York). The general term "amino acid" also includes D-amino acids, retro-inverso amino acids, as well as chemically modified amino acids (such as amino acid analogs), naturally occurring amino acids that are not typically incorporated into proteins (such as norleucine), and chemically synthesized compounds that have properties known in the art to be characteristic of amino acids (such as beta-amino acids). For example, included in the definition of amino acid are analogs or mimetics of phenylalanine or proline that allow the same conformational restrictions on the peptide compound as does natural Phe or Pro. Such analogs and mimetics are referred to herein as "functional equivalents" of the corresponding amino acid. Additional examples of amino acids are listed in Roberts and Vellaccio, The Peptides: Analysis, Synthesis, Biology, Gross and Meiehofer, eds., Vol. 5 pp. 341, Academic Press, Inc., N.Y. 1983, which is incorporated herein by reference.
[0077] The terms "protein," "polypeptide," and "peptide" are used interchangeably herein to refer to polymers of amino acid residues and variants and synthetic analogs of amino acid residues. Thus, these terms apply to amino acid polymers in which one or more amino acid residues are synthetic, non-naturally occurring amino acids, such as chemical analogs of corresponding naturally occurring amino acids, as well as to naturally occurring amino acid polymers. Polypeptides can also undergo maturation or post-translational modification processes, which can include, but are not limited to, glycosylation, proteolytic cleavage, lipidation, signal peptide cleavage, propeptide cleavage, phosphorylation, etc. "Recombinant polypeptide" means a polypeptide that is produced using recombinant technology, e.g., by expression of a recombinant or synthetic polynucleotide. When a chimeric polypeptide or biologically active portion thereof is produced recombinantly, it is also preferably substantially free of culture medium, e.g., less than about 20%, more preferably less than about 10%, most preferably less than about 5% of the volume of the protein preparation is culture medium. "Isolated" means material that is substantially or essentially free from components that normally accompany the material as it exists in its natural state. For example, as used herein, an "isolated polypeptide" refers to a polypeptide that has been purified from the molecules that flank the polypeptide in its natural state, e.g., a protein complex or CsgF peptide that has been removed from molecules adjacent to the polypeptide in the production host. An isolated CsgF peptide (optionally a truncated CsgF peptide) can be produced by amino acid chemical synthesis or can be produced by recombinant production. An isolated complex can be produced by in vitro reconstitution after purification of the components of the complex, e.g., CsgG pore and CsgF peptide, or can be produced by recombinant co-expression.
[0078] "Orthologs" and "paralogs" encompass evolutionary concepts used to describe the ancestral relationships of genes. Paralogs are genes within the same species that originated by duplication of an ancestral gene; orthologs are genes from different organisms that originated by speciation and also derive from a common ancestral gene.
[0079] "homologues" of a protein encompass peptides, oligopeptides, polypeptides, proteins, and enzymes that have amino acid substitutions, deletions, and / or insertions relative to the unmodified or wild-type protein in question and that have similar biological and functional activities to the unmodified protein from which they are derived. As used herein, the term "amino acid identity" refers to the extent that sequences are the same on an amino acid-to-amino acid basis over a comparison window. Thus, the "percentage of sequence identity" is calculated by comparing two optimally aligned sequences over the comparison window, determining the number of positions at which the identical amino acid residue (e.g., Ala, Pro, Ser, Thr, Gly, Val, Leu, He, Phe, Tyr, Trp, Lys, Arg, His, Asp, Glu, Asn, Gin, Cys, and Met) occurs in both sequences to yield the number of matched positions, dividing the number of matched positions by the total number of positions in the comparison window (i.e., the window size), and multiplying the result by 100 to yield the percentage of sequence identity.
[0080] The term "CsgG pore" defines a pore comprising a plurality of CsgG monomers. Each CsgG monomer can be a wild-type monomer from E. coli (SEQ ID NO: 3), a wild-type homologue of E. coli CsgG, such as a monomer having any one of the amino acid sequences set forth in SEQ ID NO: 68 to 88, or a variant of any of the same (e.g., a variant of any one of SEQ ID NO: 3 and 68 to 88). A variant CsgG monomer can also be referred to as a modified CsgG monomer or a mutant CsgG monomer. Modifications or mutations in a variant include, but are not limited to, any one or more of the modifications disclosed herein or a combination of the same.
[0081] For all aspects and embodiments of the application, a CsgG homologue refers to a polypeptide having at least 50%, 60%, 70%, 80%, 90%, 95%, or 99% identity of the entire sequence to the wild-type E. coli CsgG set forth in SEQ ID NO: 3. A CsgG homologue also refers to a polypeptide containing the PFAM domain PF03783 characteristic of CsgG-like proteins. A list of currently known CsgG homologues and CsgG architectures can be found in http: / / pfam.xfam.org / / family / PF03783 A list of currently known CsgG homologues and CsgG architectures can be found in
[0082] The term“modified CsgF peptide” or“CsgF peptide” defines a CsgF peptide that has been truncated from its C-terminus (e.g., to be an N-terminal fragment) and / or modified to include a cleavage site. The CsgF peptide can be a fragment of wild-type E. coli CsgF (SEQ ID NO: 5 or SEQ ID NO: 6), or a fragment of a wild-type homolog of E. coli CsgF, e.g., a peptide comprising any of the amino acid sequences set forth in SEQ ID NOs: 17-36, or a variant of any of these (e.g., a variant modified to include a cleavage site).
[0083] For all aspects and embodiments of the application, a CsgF homolog refers to a polypeptide having at least 50%, 60%, 70%, 80%, 90%, 95%, or 99% identity of the complete sequence to wild-type E. coli CsgF set forth in SEQ ID NO: 6. In some embodiments, a CsgG homolog also refers to a polypeptide containing the PFAM domain PF10614 that is characteristic of CsgF-like proteins. A list of currently known CsgF homologs and CsgF architectures can be found in http: / / pfam.xfam.org / / family / PF10614 The same, a CsgF homolog polynucleotide can comprise a polynucleotide having at least 50%, 60%, 70%, 80%, 90%, 95%, or 99% identity of the complete sequence to wild-type E. coli CsgG set forth in SEQ ID NO: 4. Examples of truncated regions of CsgF homologs set forth in SEQ ID NO: 6 have the sequences set forth in SEQ ID NOs: 17-36.
[0084] The term“N-terminal portion of a CsgF mature peptide” refers to a peptide having an amino acid sequence corresponding to the first 60, 50, or 40 amino acid residues from the N-terminus of a CsgF mature peptide (without signal sequence). The CsgF mature peptide can be wild-type or a mutant (e.g., having one or more mutations).
[0085] Sequence identity can also be a fragment or portion of a full-length polynucleotide or polypeptide. Thus, a sequence can have only 50% overall sequence identity with a full-length reference sequence, but the sequence of a particular region, domain, or subunit can have 80%, 90%, or as much as 99% sequence identity with the reference sequence. Nucleic acid sequence homology of SEQ ID NO: 1 for a CsgG homolog or SEQ ID NO: 4 for a CsgF homolog is not limited to sequence identity. Many nucleic acid sequences, despite having significantly lower sequence identity, can nonetheless exhibit biologically significant homology to one another. Homologous nucleic acid sequences are considered to be sequences that will hybridize to one another under low stringency conditions (M. R. Green, J. Sambrook, 2012, Molecular Cloning: A Laboratory Manual, 4thedition, Books 1-3, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY).
[0086] The term "wild type" refers to a gene or gene product isolated from a naturally occurring source. A wild type gene is that which is most frequently observed in a population and is therefore arbitrarily designed the "normal" or "wild type" form of the gene. In contrast, the term "modified," "mutant," or "variant" refers to a gene or gene product that displays a sequence modification (e.g., substitution, truncation, or insertion), post-translational modification, and / or functional property (e.g., altered characteristic) as compared to a wild type gene or gene product. Note that naturally occurring mutants can be isolated; these are identified by the fact that they have altered characteristics as compared to a wild type gene or gene product. Methods of introducing or substituting naturally occurring amino acids are well known in the art. For example, methionine (M) can be substituted with arginine (R) by replacing the codon for methionine (ATG) with the codon for arginine (CGT) at the relevant position in the polynucleotide encoding the mutant monomer. Methods of introducing or substituting non-naturally occurring amino acids are also well known in the art. For example, non-naturally occurring amino acids can be introduced by including synthetic aminoacyl-tRNAs in the IVTT system used to express the mutant monomer. Alternatively, non-naturally occurring amino acids can be introduced by expressing the mutant monomer in E. coli, which is auxotrophic for a particular amino acid in the presence of a synthetic (i.e., non-naturally occurring) analog of that particular amino acid. If the mutant monomers are produced using a partial peptide synthesis approach, they can also be produced by naked ligation. Conservative substitutions replace an amino acid with another amino acid of similar chemical structure, similar chemical nature, or similar side chain volume. Introduced amino acids can have similar polarity, hydrophilicity, hydrophobicity, basicity, acidity, neutrality, or charge as the amino acids they replace. Alternatively, a conservative substitution can introduce another aromatic or aliphatic amino acid in place of a pre-existing aromatic or aliphatic amino acid. Conservative amino acid changes are well known in the art and can be selected according to the properties of the 20 primary amino acids defined in Table 1 below. Where amino acids have similar polarity, this can also be determined with reference to the scale of hydrophilicity of the side chains of amino acids in Table 2.
[0087] Table 1 - Chemical properties of amino acids
[0088]
[0089]
[0090] Table 2 - Scale of hydrophilicity
[0091]
[0092] A mutant or modified protein, monomer or peptide can also be chemically modified in any way at any site. A mutant or modified monomer or peptide is preferably chemically modified by attachment of a molecule to one or more cysteines (cysteine ligation), attachment of a molecule to one or more lysines, attachment of a molecule to one or more unnatural amino acids, enzymatic modification of an epitope or modification of a terminus. Suitable methods for making such modifications are well known in the art. A mutant of a modified protein, monomer or peptide can be chemically modified by attachment of any molecule. For example, a mutant of a modified protein, monomer or peptide can be chemically modified by attachment of a dye or fluorophore. In some embodiments, a mutant or modified monomer or peptide is chemically modified with a molecular adaptor that facilitates the interaction between a pore comprising a monomer or peptide and a target nucleotide or target polynucleotide sequence. The molecular adaptor is preferably a cyclic molecule, a cyclodextrin, a substance capable of hybridization, a DNA binding agent or intercalator, a peptide or peptide analog, a synthetic polymer, an aromatic planar molecule, a positively charged small molecule or a small molecule capable of hydrogen bonding.
[0093] The presence of an adaptor improves the host-guest chemistry of the pore and the nucleotide or polynucleotide sequence, thereby improving the sequencing ability of a pore formed from a mutant monomer. The principles of host-guest chemistry are well known in the art. The adaptor has an effect on the physical or chemical properties of the pore, which effect improves the interaction of the pore with the nucleotide or polynucleotide sequence. The adaptor can change the charge of the barrel or channel of the pore, or specifically interact or bind with the nucleotide or polynucleotide sequence, thereby facilitating its interaction with the pore. Thus, a modified CsgF peptide as provided in the present disclosure can be coupled to an enzyme or protein, thereby providing better accessibility of the protein or enzyme to the pore, which can facilitate certain applications of a pore complex comprising the modified CsgF peptide.
[0094] In this regard, a protein can also be a fusion protein, especially referring to a genetic fusion produced, for example, by recombinant DNA techniques. A protein can also be conjugated or "conjugated to" as used herein, especially referring to chemical and / or enzymatic conjugations that produce stable covalent linkages.
[0095] When several polypeptide or protein monomers associate or interact with each other, a protein can form a protein complex. By "associate" is meant any direct or indirect interaction. Direct interaction implies contact between binding partners, for example by covalent linkage or coupling. Indirect interaction means any interaction by which interaction partners in a complex of two or more compounds interact. The interaction can be entirely indirect, by means of one or more bridging molecules, or partially indirect, in which there is still direct contact between partners, stabilized by additional interactions of one or more compounds. As referred to in the present disclosure, a "complex" is defined as a group of two or more associated proteins, which can have different functions. Association between different polypeptides of a protein complex can be by non-covalent interactions, such as hydrophobic or ionic forces, or can be covalent binding or coupling, such as disulfide bridges or peptide bonds. Covalent "binding" or "coupling" are used interchangeably herein and can also relate to "cysteine coupling" or "reactive or photoreactive amino acid coupling", referring to bioconjugation between cysteines or between (photo)reactive amino acids, respectively, which is a chemical covalent linkage forming a stable complex. Examples of photoreactive amino acids include azido homoalanine, homoallylglycine, homoallylglycine, p-acetyl-Phe, p-azido-Phe, p-allyloxy-Phe and p-benzoyl-Phe (Wang et al., 2012, in Protein Engineering, DOI: 10.5772 / 28719; Chin et al., 2002, Proc. Nat. Acad. Sci. USA 99(17); 11020-24).
[0096] A "bio-pore" is a transmembrane protein structure that defines a channel or pore that allows the translocation of molecules and ions from one side of the membrane to the other. The translocation of ionic species through the pore can be driven by a potential difference applied to either side of the pore. A "nanopore" is a bio-pore in which the channel through which molecules or ions pass is nanoscale (10 -9 meters) in its smallest dimension. In some embodiments, the bio-pore can be a transmembrane protein pore. The transmembrane protein structure of the bio-pore can be monomeric or oligomeric in nature. Typically, the pore comprises a plurality of polypeptide subunits arranged around a central axis, thereby forming a protein-lined channel extending substantially perpendicular to the membrane in which the nanopore resides. The number of polypeptide subunits is not limited. Typically, the number of subunits is 5 to 30, suitably the number of subunits is 6 to 10. Alternatively, the number of subunits is not defined as in the case of perfringolysin or related large membrane pores. The portion of the protein subunit that forms the protein-lined channel within the nanopore typically comprises secondary structure motifs that can include one or more transmembrane beta-barrels and / or alpha-helical portions.
[0097] As used interchangeably herein, the term "pore," "pore complex," or "complex pore" refers to an oligomer pore in which, for example, at least a CsgG monomer (including, for example, one or more CsgG monomers, such as two or more CsgG monomers, three or more CsgG monomers) or a CsgG pore (comprised of CsgG monomers) and a CsgF peptide (e.g., a modified or truncated CsgF peptide) are associated into a complex and together form a pore or nanopore. The pore complex of the present disclosure has the characteristics of a biological pore, i.e., it has the typical transmembrane protein structure. When the pore complex is provided in an environment having a membrane component, a membrane, a cell, or an insulating layer, the pore complex will insert into the membrane or insulating layer, forming a "transmembrane pore complex."
[0098] The pore complex or transmembrane pore complex of the present disclosure is suitable for analyte characterization. In some embodiments, the pore complex or transmembrane complex described herein can be used, for example, to sequence a polynucleotide sequence, as it can distinguish different nucleotides with high sensitivity. The pore complex of the present disclosure can be an isolated pore complex that is substantially isolated, purified, or substantially purified. The pore complex of the present disclosure is "isolated" or purified if it is completely free of any other components such as lipids or other pores, or other proteins with which it is typically associated in its native state, e.g., CsgE, CsgA CsgB, or if it is sufficiently enriched from a membrane compartment. It is substantially isolated if the pore complex is mixed with a carrier or diluent that will not interfere with its intended use. For example, it is substantially isolated or substantially purified if it is present in a form that contains less than 10%, less than 5%, less than 2%, or less than 1% of other components such as triblock copolymers, lipids, or other pores. Alternatively, the pore complex of the present invention can be a transmembrane pore complex when present in a membrane. The present disclosure provides an isolated pore complex comprising a homo-oligomeric pore derived from CsgG comprising identical mutant monomers, CsgG can also contain a mutant form of CsgG monomer as its homolog. Alternatively, an isolated pore complex comprising a hetero-oligomeric CsgG pore is provided, which can be a CsgG pore comprised of mutant and wild-type CsgG monomers or different forms of CsgG variants, mutants, or homologs. The isolated pore complex typically comprises at least 7, at least 8, at least 9, or at least 10 CsgG monomers and one or more (modified) CsgF peptides, such as 2, 3, 4, 5, 6, 7, 8, 9, 10 CsgF peptides. The pore complex can comprise CsgG monomer:CsgF peptide in any ratio. In one embodiment, the ratio of CsgG monomer:CsgF peptide is 1:1.
[0099] A "constriction", "orifice", "constriction region", "channel constriction", or "constriction site" as used interchangeably herein refers to the aperture defined by the internal lumen surface of a pore or pore complex that functions to allow passage of ions and target molecules (such as, but not limited to, polynucleotides or individual nucleotides) through the pore complex channel, but not other non-target molecules. In some embodiments, the constriction is the narrowest aperture in the pore or pore complex. In this embodiment, the constriction can be used to limit passage of molecules through the pore. The size of the constriction is generally a key factor in determining the suitability of a nanopore for nucleic acid sequencing applications. If the constriction is too small, the molecule to be sequenced will not pass through. However, to have the greatest impact on the ionic current flow through the channel, the constriction should not be too large. For example, the constriction should not be wider than the solvent accessible lateral diameter of the target analyte. Ideally, the diameter of any constriction should be as close as possible to the lateral diameter of the analyte passing through. For sequencing of nucleic acids and nucleic acid bases, a suitable constriction diameter is in the nanometer range (10 -9 Suitably, the diameter should be in the region of 0.5 to 2.0 nm, typically, the diameter is in the region of 0.7 to 1.2 nm. The diameter of the constriction in wild-type E. coli CsgG is approximately (0.9 nm). The diameter of the CsgF constriction formed in pore complexes comprising a CsgG-like pore and a modified CsgF peptide or homolog or mutant thereof is in the range of 0.5 to 2 nm or in the range of 0.7 to 1.2 nm, thus suitable for nucleic acid sequencing.
[0100] When two or more constriction segments are present and spaced apart, each constriction segment can simultaneously interact or "read" a separate nucleotide within the nucleic acid strand. In this case, a reduction in ion flow through the channel would be the result of the combined restriction of flow in all of the nucleotide-containing constriction segments. Thus, in some cases, a double constriction segment can result in a composite current signal. In certain cases, when two such reading heads are present, it can not be possible to separately determine the current reading for one constriction segment or "reading head." The constriction segment of wild-type E. coli CsgG (SEQ ID NO: 3) consists of two annular rings formed by the juxtaposition of a tyrosine residue at position 51 (Tyr 51) and phenylalanine and asparagine residues at positions 56 and 55, respectively (Phe 56 and Asn 55) in two adjacent protein monomers (Figure 1). In most cases, the wild-type pore structure of CsgG is engineered by recombinant genetic techniques to expand, alter, or remove one of the two annular rings that make up the constriction segment of CsgG (referred to herein as the "CsgG channel constriction segment"), leaving a single well-defined reading head. The constriction motif in the CsgG oligomer pore is located at amino acid residues at positions 38 to 63 in the wild-type monomeric E. coli CsgG polypeptide depicted in SEQ ID NO: 3. In considering this region, it is demonstrated that mutations at any of amino acid residue positions 50 to 53, 54 to 56, and 58 to 59 within the channel of the wild-type CsgG structure, as well as the positioning of the Tyr 51, Asn 55, and Phe 56 side chains, are advantageous for altering or changing the characteristics of the reading head. The present disclosure relates to pore complexes comprising a CsgG pore and a modified CsgF peptide or homolog or mutant thereof, surprisingly adding another constriction segment (referred to herein as the "CsgF channel constriction segment") to the CsgG-containing pore complex, forming a suitable additional second reading head in the pore by complexing with the modified CsgF peptide. The additional CsgF channel constriction segment or reading head is positioned adjacent to the constriction ring of the CsgG pore or mutant GcsG pore. The additional CsgF channel constriction segment or reading head is positioned about 10 nm or less, such as 5 nm or less, such as 1, 2, 3, 4, 5, 6, 7, 8, 9 nm from the constriction ring of the CsgG pore or mutant GcsG pore. The pore complexes or transmembrane pore complexes of the present disclosure include pore complexes with two reading heads, meaning that the channel constriction segments are positioned in a manner that provides suitable separate reading heads without interfering with the accuracy of the other constriction channel reading heads.Thus the pore complex can comprise a CsgG mutant pore (see incorporated references WO2016 / 034591, WO2017 / 149316, WO2017 / 149317, WO2017 / 149318 and International Patent Application No. PCT / GB2018 / 051191 each of which lists mutations to the wild type CsgG pore that improve the properties of the pore) as well as a wild type CsgG pore or homologue thereof, together with a modified CsgF peptide or homologue or mutant thereof, wherein the CsgF peptide has a further constriction channel forming a reading head.
[0101] pore
[0102] The present disclosure relates to CsgG pores in complex with extracellularly localised CsgF peptides which surprisingly introduce additional constriction segments or reading heads into the pore complex. Furthermore, the disclosure provides information on the location of the constriction segments formed by the CsgF peptides within the pore complex, which peptide inserts into the lumen of the CsgG pore and the constriction site is in the N-terminal portion of the CsgF protein. In addition, it is demonstrated that modified or truncated CsgF peptides of the disclosure are sufficient to form pore complexes and provide means and methods for biosensing applications. The disclosure comprises wild-type and mutant CsgG pores (as disclosed in, for example, WO2016 / 034591, WO2017 / 149316, WO2017 / 149317, WO2017 / 149318 and International Patent Application Number PCT / GB2018 / 051191) or homologues or mutants thereof in combination with modified or truncated CsgF peptides and mutants or homologues thereof, which all together improve the ability of the CsgG-like pore complex to interact with an analyte, such as a polynucleotide. The additional constriction segments introduced in the CsgG-like nanopore channel by complex formation with the (modified or truncated) CsgF peptide enlarge the contact surface with the passing analyte and can act as a second reading head for analyte detection and characterisation. Pores comprising mutant CsgG monomers in combination with novel CsgF mutant or modified forms can improve the characterisation of an analyte, such as a polynucleotide, providing a more distinctive direct relationship between the current observed as the polynucleotide moves through the pore. In particular, by spacing the two stacked reading heads at a defined distance, the CsgG:CsgF pore complex can facilitate the characterisation of polynucleotides containing at least one homopolymer stretch, for example several consecutive copies of the same nucleotide, which exceeds the interaction length of a single CsgG reading head. In addition, by spacing the two stacked constriction segments at a defined distance, small molecule analytes (including organic or inorganic drugs as well as pollutants) passing through the CsgG:CsgF complex pore will pass through both independent reading heads in succession. The chemical properties of either reading head can be independently modified, each providing unique interaction properties with the analyte, providing additional discrimination power during analyte detection.
[0103] In a first aspect, the present invention relates to an isolated pore complex comprising a CsgG pore or a homolog or mutant thereof, or a CsgG-like pore, and a modified CsgF peptide or a homolog or mutant thereof. Indeed, the present disclosure relates to a modified CsgG biological pore comprising a modified CsgF peptide (which can be truncated), mutants and / or variants thereof. In one embodiment, the interaction zone between the modified CsgF peptide or a homolog or mutant thereof is located in the lumen of the CsgG pore or a homolog or mutant thereof. In another embodiment, the pore complex has two or more constriction sites or reading heads provided by at least one constriction segment of the CsgG pore and by at least one constriction segment introduced by the CsgF peptide, thereby forming a complex with the CsgG pore. It was demonstrated that the N-terminal CsgF positions, including positions ranging from amino acid residues 39-64 of SEQ ID NO: 5, or more particularly positions ranging from amino acid residues 49-64 of SEQ ID NO: 5, allow for a detectable amount of stable CsgG:CsgF complex. In one embodiment, the CsgF constriction segment produced by the modified CsgF peptide (e.g., a CsgF peptide described herein) is adjacent to or head-to-head with the first constriction segment in the CsgG pore of the pore complex. For CsgG or CsgG-like protein pores, it was determined that the constriction sites are formed by the loop regions of the beta strands (see Figure 1).
[0104] In one embodiment, the modified CsgF peptide is a peptide wherein the modification refers, inter alia, to a truncated CsgF protein or fragment, including by way of limitation the N-terminal CsgF peptide fragment comprising the constriction region and binding to the CsgG monomer or a homolog or mutant thereof. The modified CsgF peptide can additionally comprise mutations or homologous sequences that can facilitate certain properties of the pore complex. In a particular embodiment, the modified CsgF peptide comprises a truncation of the CsgF protein as compared to the wild-type proprotein (SEQ ID NO: 5) or mature protein (SEQ ID NO: 6) sequence or homologs thereof. These modified peptides are intended to be used as components of the pore complex, introducing additional constriction sites or reading heads within the CsgG-like pore formed by CsgG and the modified or truncated CsgF peptide. Examples of truncated modified peptides are described below.
[0105] Examples of homologues of the modified CsgF peptide are identified, for example, in Example 3, and reveal CsgF-like proteins or CsgF peptides comprising homologous or similar constriction regions in different bacterial strains, which are useful in the use of similar pore complexes. Structural properties and CsgG binding elements in CsgF peptides derived from various CsgF homologues are conserved, such that CsgF peptides can be used in combination with different wild-type or mutant CsgG pores. This includes complexes of CsgG pores with non-homologous CsgF, meaning that the CsgG pore and the parent CsgF homologue from which the CsgF is derived do not need to be derived from the same operon, bacterial species or strain.
[0106] In alternative embodiments, the CsgG pore in the pore complex is not a wild-type pore, but is also comprises a mutation or modification to increase the properties of the pore. The isolated pore complex of the disclosure formed by a CsgG pore or homologue thereof and a modified CsgF peptide or homologue thereof, can be formed from a wild-type form of the CsgG pore, or further modifications can be made in the CsgG pore, for example by directed mutagenesis of specific amino acid residues, to further enhance the properties required for the CsgG pore to be used in the pore complex. For example, in embodiments of the application, mutations are contemplated to alter the number, size, shape, placement or orientation of constriction segments within the channel. Pore complexes comprising modified mutant CsgG pores can be prepared by known genetic engineering techniques which cause insertion, substitution and / or deletion of specific target amino acid residues in the polypeptide sequence. In the case of oligomeric CsgG pores, the mutation can be made in each monomeric polypeptide subunit, or in any one or all of the monomers. Suitably, in one embodiment of the application, the mutation is made to all of the monomeric polypeptides within the oligomeric protein structure. A mutant CsgG monomer is one whose sequence is different from that of a wild-type CsgG monomer and which retains the ability to form a pore. Methods to confirm the ability of a mutant monomer to form a pore are well known in the art. The disclosure comprises wild-type and mutant CsgG pores (for example, as disclosed in WO2016 / 034591, WO2017 / 149316, WO2017 / 149317, WO2017 / 149318 and international patent application number PCT / GB2018 / 051191) or homologues thereof in combination with modified or truncated CsgF peptides and mutants or homologues thereof, which all together improve the ability of the CsgG-like pore complex to interact with an analyte, such as a polynucleotide. The mutant CsgG pore can comprise one or more mutant monomers. The CsgG pore can be a homopolymer comprising identical monomers, or a SEQ heteropolymer comprising two or more different monomers. The monomers can have one or more mutations of any combination described below.
[0107] The nanopore complex comprising a modified CsgF peptide differs from the wild-type CsgF protein depicted in SEQ ID NO: 6 in that, in certain embodiments, the modified CsgF peptide comprises only an N-terminal fragment or truncation of the wild-type CsgF protein. However, the modified CsgF peptide can additionally or alternatively be a mutant CsgF peptide in the sense that mutations such as amino acid substitutions are made to allow for a better second constriction site in the pore formed by the complex comprising the CsgG pore and the modified CsgF peptide. When the complex is used for nucleotide sequencing, the mutant monomer can likewise have improved polynucleotide reading properties, i.e. in addition to the improved features of the complex comprising two reading heads, it shows improved polynucleotide capture and nucleotide discrimination. In particular, the pore constructed from the mutant peptide can capture nucleotides and polynucleotides more easily than the wild-type. Additionally, the pore constructed from the mutant peptide can show an increased current range, which makes it easier to discriminate between different nucleotides, and a reduced state change, which increases the signal-to-noise ratio. Additionally, the number of nucleotides contributing to the current when a polynucleotide moves through the pore constructed from the mutant can be reduced. This makes it easier to identify a direct relationship between the current observed when a polynucleotide moves through the pore and the polynucleotide sequence. Additionally, the pore constructed from the mutant peptide can show an increased throughput, e.g. it is more likely to interact with an analyte such as a polynucleotide. This makes it easier to use the pore to characterize the analyte. The pore constructed from the mutant peptide can be more easily inserted into a membrane, or can provide an easier way to keep other proteins in the vicinity of the pore complex.
[0108] In an alternative embodiment, the diameter of the CsgF constriction site provided in the pore complex of the application is in the range of 0.5 nm to 2.0 nm, thereby providing a pore complex suitable for nucleic acid sequencing as described above.
[0109] The pore can be stabilized by covalent attachment of the CsgF peptide to the CsgG pore. The covalent linkage can for example be a disulfide bond or click chemistry. The CsgF peptide and the CsgG pore can for example be covalently linked via residues at positions corresponding to one or more of the following pairs of positions of SEQ ID NO: 6 and SEQ ID NO: 3, respectively: 1 and 153, 4 and 133, 5 and 136, 8 and 187, 8 and 203, 9 and 203, 11 and 142, 11 and 201, 12 and 149, 12 and 203, 26 and 191, and 29 and 144.
[0110] The interaction between the CsgF peptide and the CsgG pore in the pore can be stabilized, for example, via hydrophobic or electrostatic interactions at positions corresponding to one or more of the following pairs of positions of SEQ ID NO: 6 and SEQ ID NO: 3, respectively: 1 and 153, 4 and 133, 5 and 136, 8 and 187, 8 and 203, 9 and 203, 11 and 142, 11 and 201, 12 and 149, 12 and 203, 26 and 191, and 29 and 144.
[0111] The residues at one or more of the positions listed above in CsgF and / or CsgG can be modified in order to enhance the interaction between CsgG and CsgF in the pore.
[0112] In one embodiment, the pores of the application can be isolated, substantially isolated, purified, or substantially purified. A pore is isolated or purified if it is completely free of any other components, such as lipids or other pores. A pore is substantially isolated if it is mixed with a carrier or diluent that does not interfere with its intended use. For example, a pore is substantially isolated or substantially purified if it is present in a form that contains less than 10%, less than 5%, less than 2%, or less than 1% of other components, such as triblock copolymers, lipids, or other pores. Alternatively, the pores of the application can be present in a membrane. Suitable membranes are discussed below.
[0113] The pores of the application can exist as individual or single pores. Alternatively, the pores of the application can exist in a homologous or heterologous population of two or more pores.
[0114] CsgF peptide
[0115] A second aspect of the application relates to novel modified CsgF monomers (peptides), or truncated CsgF proteins, or modified or truncated peptides of CsgF homologues or mutants. Those novel modified CsgF peptides can be used in pore complexes to integrate a second or additional readhead. The modification or truncation preferably results in a fragment of the wild-type CsgF or mutant or homologous CsgF protein, more preferably an N-terminal fragment.
[0116] Mature CsgF (as set forth in SEQ ID NO: 6) can be divided into three main regions: a “CsgF constriction peptide” (FCP), a “neck” region, and a “head” region (as shown in Figures 4 and 5). The “head” region of the CsgF peptide is distinct from the readhead of the pores described herein. The “head” region of the CsgF peptide can also be referred to as the “C-terminal head domain”.
[0117] The FCP makes contact with the CsgG b-barrel formation contact region, where an additional constriction segment is created. The neck region protrudes from the b-barrel. In a CsgG:CsgF oligomer, it forms a thin-walled hollow tube that links the FCP to the spherical head region.
[0118] Based on multiple sequence alignment ( Figure 8 ), co-purification experiment (Figure 9) and in CryoEM reconstruction of the CsgG:CsgF complex at resolution (Figure 11). The CsgF contractile peptide, neck, and head regions can be defined as three consecutive residue extensions in mature CsgF.
[0119] The FCP roughly spans residues 1 to 35 of mature CsgF (SEQ ID NO:6). When comparing different CsgF orthologs, the FCP forms the most conserved region of the protein. Figure 8 , Figure 10 CryoEM 3D reconstruction revealed that FCP formed a well-defined structure that was bound to the interior of the CsgG β-barrel via non-covalent contacts with CsgG transmembrane hairpins TM1 (residues 134 to 154 of Seq ID NO:3) and TM2 (residues 184 to 208 of Seq ID NO:3). Figure 1 E (Figure 11) (TM1 and TM2 are defined in Goyal P et al., 2014). During reconstruction, nine copies of FCP bind to the CsgG oligomer (containing nine monomers) and together generate an additional contraction segment approximately 2 nm above the CsgG contraction segment, which is formed by a continuous loop spanning residues 46 to 61 of mature CsgG (Seq ID NO:3). Figure 1 E (Figure 11).
[0120] The cryoEM 3D reconstruction also showed that the N-terminal residues of CsgF bind near the bottom or top of the CsgG β-barrel (depending on orientation) and exit the β-barrel at residue 32. This is in good agreement with the MD simulations that showed the average contact time of residue pairs in the CsgG:CsgF binding interface (Table 4). The cryoEM structure and MD simulations showed that residues 33-34 are located outside the CsgG β-barrel, where the very (albeit not strictly) conserved Pro residue (Pro 35 in Seq ID NO:6) transitions to the CsgF neck region. The CsgF neck was not resolved at atomic detail in the CsgG:CsgF 3D reconstruction, indicating its conformational flexibility. Based on multiple sequence alignment and secondary structure prediction, the CsgF neck is expected to span approximately from residue 36 to residue 50 (SEQ ID NO:6). The CsgF head region forms the C-terminal portion of CsgF and is predicted to span approximately from residue 51 to the C-terminus of CsgF. In the CsgG:CsgF complex, oligomerization in this region produces spherical structures, which appear to cap the CsgG:CsgF channels (Figure 4). Figure 5). A multiple sequence alignment of CsgF orthologues shows that the CsgF neck is the least conserved region, suggesting that its length can vary between orthologues Figure 8
[0121] CsgF peptides forming part of the application are truncated CsgF peptides lacking the C-terminal headpiece; lacking the C-terminal headpiece and part of the neck domain of CsgF (e.g. the truncated CsgF peptide can comprise only part of the neck domain of CsgF); or lacking the C-terminal headpiece and the neck domain of CsgF. The CsgF peptide can lack part of the neck domain of CsgF, e.g. the CsgF peptide can comprise part of the neck domain, e.g. starting at amino acid residue 36 from the N-terminus of the neck domain (see SEQ ID: NO: 6) (e.g. residues 36-40, 36-41, 36-42, 36-43, 36-45, 36-46 up to residues 36-50 or 36-60 of SEQ ID: NO: 6). The CsgF peptide preferably comprises the CsgG binding region and the region that forms the constriction in the pore. The CsgG binding region typically comprises residues 1 to 8 and / or 29 to 32 of the CsgF protein (SEQ ID NO: 6 or a homologue from another species), and can include one or more modifications. The region that forms the constriction in the pore typically comprises residues 9 to 28 of the CsgF protein (SEQ ID NO: 6 or a homologue from another species), and can include one or more modifications. Residues 9 to 17 comprise the conserved motif N9PXFGGXXX 17 and forms the turn region. Residues 9 to 28 form an alpha-helix. X 17 (N17 in SEQ ID NO: 6) forms the apex of the constriction region, corresponding to the narrowest part of the CsgF constriction in the pore. The CsgF constriction region also makes stable contacts with the CsgG β-barrel primarily at residues 9, 11, 12, 18, 21 and 22 of SEQ ID NO: 6.
[0122] The CsgF peptide is 25 to 50 amino acids in length. The CsgF peptide is typically 28 to 50 amino acids in length, such as 29 to 49, 30 to 45 or 32 to 40 amino acids in length. Preferably, the CsgF peptide comprises 29 to 35 amino acids or 29 to 45 amino acids. The CsgF peptide comprises the entire or part of the FCP, corresponding to residues 1 to 35 of SEQ ID NO: 6. Where the CsgF peptide is shorter than the FCP, the truncation is preferably made at the C-terminus.
[0123] A CsgF fragment of SEQ ID NO: 6, or a homolog or mutant thereof, can be 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, or 55 amino acids in length.
[0124] A CsgF peptide can comprise the amino acid sequence of SEQ ID NO: 6, from residue 1 to residue 25 to 60, such as 27 to 50, for example 28 to 45, of any one of SEQ ID NO: 6, or the corresponding residues from a homolog of SEQ ID NO: 6, or a variant of any one thereof. More specifically, the CsgF peptide can comprise SEQ ID NO: 39 (residues 1 to 29 of SEQ ID NO: 6), or a homolog or variant thereof.
[0125] Examples of such CsgF peptides comprise, essentially consist of, or consist of the sequence of SEQ ID NO: 15 (residues 1 to 34 of SEQ ID NO: 6), SEQ ID NO: 54 (residues 1 to 30 of SEQ ID NO: 6), SEQ ID NO: 40 (residues 1 to 45 of SEQ ID NO: 6), or SEQ ID NO: 55 (residues 1 to 35 of SEQ ID NO: 6), and a homolog or variant of any one thereof. Other examples of CsgF peptides comprise, essentially consist of, or consist of the sequence of SEQ ID NO: 7, SEQ ID NO: 8, SEQ ID NO: 9, SEQ ID NO: 10, SEQ ID NO: 11, SEQ ID NO: 12, SEQ ID NO: 13, SEQ ID NO: 14, SEQ ID NO: 16.
[0126] In a CsgF peptide, for example one or more residues in SEQ ID NO: 15, SEQ ID NO: 39, SEQ ID NO: 40, SEQ ID NO: 54, or SEQ ID NO: 55 can be modified.
[0127] For example, a CsgF peptide can comprise a modification at a position corresponding to one or more of the following positions in SEQ ID NO: 6: G1, T4, F5, R8, N9, N11, F12, A26, and Q29.
[0128] The CsgF peptide can be modified to, for example, introduce a cysteine, a hydrophobic amino acid, a charged amino acid, a non-native reactive amino acid, or a photo-reactive amino acid at a position corresponding to one or more of the following positions in SEQ ID NO: 6: G1, T4, F5, R8, N9, N11, F12, A26, and Q29.
[0129] For example, the CsgF peptide can comprise a modification at a position corresponding to one or more of the following positions in SEQ ID NO: 6: N15, N17, A20, N24, and A28. The CsgF peptide can comprise a modification at a position corresponding to D34 to stabilize the CsgG-CsgF complex. In particular embodiments, the CsgF peptide comprises one or more of the following substitutions: N15S / A / T / Q / G / L / V / I / F / Y / W / R / K / D / C, N17S / A / T / Q / G / L / V / I / F / Y / W / R / K / D / C, A20S / T / Q / N / G / L / V / I / F / Y / W / R / K / D / C, N24S / T / Q / A / G / L / V / I / F / Y / W / R / K / D / C, A28S / T / Q / N / G / L / V / I / F / Y / W / R / K / D / C, and D34F / Y / W / R / K / N / Q / C. The CsgF peptide may, for example, comprise one or more of the following substitutions: G1C, T4C, N17S, and D34Y or D34N.
[0130] The CsgF peptide can be produced by cleaving a longer protein, such as full-length CsgF, with an enzyme. Cleavage at a particular site can be directed by modifying the longer protein (such as full-length CsgF) to include an enzyme cleavage site at the appropriate position. Examples of CsgF amino acid sequences that have been modified to include such enzyme cleavage sites are shown in SEQ ID NOs: 56-67. Upon cleavage, the entire or partial added enzyme cleavage site can be present in the CsgF peptide that associates with CsgG to form a pore. Thus, the CsgF peptide can also comprise all or part of the enzyme cleavage site at its C-terminus.
[0131] Some examples of suitable CsgF peptides are shown in Table 3 below:
[0132] Table 3: CsgF peptides
[0133]
[0134]
[0135] In a particular embodiment, the CsgF fragment comprises the amino acid sequence SEQ ID NO: 39, or a mutant or homologue thereof. In particular, SEQ ID NO: 39 comprises the first 29 amino acids of the mature CsgF peptide (SEQ ID NO: 6). In another embodiment, the modified CsgF peptide of the present application is a truncated peptide comprising SEQ ID NO: 40. In particular, SEQ ID NO: 40 comprises the first 45 amino acids of the mature CsgF peptide (SEQ ID NO: 6). In particular, the CsgF retraction site and the binding site for CsgG are located within the N-terminal CsgF peptide region, further characterized by amino acids 39 to 64 of SEQ ID NO: 5 (present in SEQ ID NO: 39 and SEQ ID NO: 40), or especially amino acids 49 to 64 of SEQ ID NO: 5 (present in SEQ ID NO: 40, but not in SEQ ID NO: 39, the latter fragment encoded by SEQ ID NO: 39 showing a weaker interaction with CsgG (see examples)), confer higher stability to the complex. Thus, the present disclosure provides for the modification of the CsgF protein by truncating the protein into said peptide or a peptide comprising said N-terminal fragment or retraction site region to allow the formation of a complex with the CsgG pore or a homologue or mutant thereof in vivo. Further restrictions are provided in one embodiment regarding the modified CsgF peptide comprising SEQ ID NO: 37 or SEQ ID NO: 38. Finally, the identification of CsgF homologous peptides, in particular CsgF homologous peptides (FCP peptides) aligned within the retraction region, also provides for modified CsgF peptide homologues that can form part of said isolated complex (e.g., see Figure 8 and Figure 10 ).
[0136] Another embodiment relates to a modified or truncated CsgF peptide comprising SEQ ID NO: 15, wherein said SEQ ID NO: 15 contains a region of a CsgF protein comprising several residues from the region of the CsgG binding and / or constriction site sufficient to reconstitute a complex pore comprising CsgG or a homolog thereof and the modified CsgF peptide in vitro to generate an isolated pore complex comprising a constriction segment of the CsgF channel. Another embodiment describes said modified CsgF peptide comprising SEQ ID NO: 16, which contains an N-terminal fragment of a CsgF protein and two additional amino acids (KD) that will increase the solubility and stability of the (synthetic) peptide and also allow reconstitution of said complex pore in vitro. Other embodiments are provided, wherein said modified CsgF peptide comprises SEQ ID NO: 15, SEQ ID NO: 16, or a homolog or mutant thereof, wherein said modified CsgF peptide is further mutated, but still maintains a minimum of 35% amino acid identity to SEQ.ID.NO: 15 or SEQ ID NO: 16, respectively, within the region of the modified CsgF peptide corresponding to said SEQ ID NO: 15 or 16, for example 40%, 50%, 60%, 70%, 80%, 85%, 90% amino acid identity. Other embodiments are provided, wherein said modified CsgF peptide comprises SEQ ID NO: 15, SEQ ID NO: 16, or a homolog or mutant thereof, wherein said modified CsgF peptide is further mutated, but still maintains a minimum of 40%, 45%, 50%, 60%, 70%, 80%, 85%, or 90% amino acid identity to SEQ.ID.NO: 15 or SEQ ID NO: 16, respectively, within the region of the modified CsgF peptide corresponding to said SEQ ID NO: 15 or 16. As discussed above, those mutated regions are intended to alter and / or improve the characteristics of the CsgF constriction site, and thus, for example, a more accurate target analysis can be obtained. Another embodiment discloses a modified CsgF peptide, wherein one or more positions in the region comprising SEQ ID NO: 39, SEQ ID NO: 40, SEQ ID NO: 54, or SEQ ID NO: 55 are modified, and wherein said mutations maintain a minimum of 35% amino acid identity to SEQ.ID.NO: 39, SEQ ID NO: 40, SEQ ID NO: 54, or SEQ ID NO: 55, or 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95% amino acid identity, in the peptide fragment corresponding to the region comprising SEQ ID NO: 39, SEQ ID NO: 40, SEQ ID NO: 54, or SEQ ID NO: 55.
[0137] Accordingly, other embodiments of the present application relate to an isolated pore complex comprising a CsgG pore or a homolog or mutant thereof, and a modified CsgF peptide, or a homolog or mutant thereof, wherein the modified CsgF peptide is as defined in the second aspect of the application.
[0138] Further embodiments relate to an isolated pore complex, wherein the CsgG pore is coupled to the modified CsgF peptide via covalent binding of at least one monomer. In one case, the covalent linkage or binding can be achieved by a cysteine linkage, wherein the thiol side group of the cysteine is covalently linked to another amino acid residue or moiety. In a second possibility, the covalent linkage is obtained by interaction between non-natural (photo)reactive amino acids. By (photo)reactive amino acids is meant artificial analogs of natural amino acids, which can be used for cross-linking of protein complexes and can be incorporated into proteins and peptides in vivo or in vitro. Commonly used photo-reactive amino acid analogs are the photo-reactive diazirine analogs of leucine and methionine, and p-benzoyl-phenyl-alanine, and azido homoalanine, homoallylglycine, homoallylglycine, p-acetyl-Phe, p-azido-Phe, p- propargyloxy-Phe and p-benzoyl-Phe (Wang et al., 2012; Chin et al., 2002). Upon exposure to UV light, they are activated and covalently bind to interacting proteins in the range of a few angstroms of the photo-reactive amino acid analog. However, the position in the CsgG monomer where the covalent linkage can occur depends on the exposure of the modified CsgF peptide. As shown in Figure 1, several amino acids are in position to provide a covalent linkage, i.e. positions 132, 133, 136, 138, 140, 142, 144, 145, 147, 149, 151, 153, 155, 183, 185, 187, 189, 191, 201, 203, 205, 207 or 209 of SEQ ID NO: 3 or a homolog thereof.
[0139] Another aspect of the application relates to a construct comprising said modified CsgF peptide, wherein said peptide is covalently attached. A "construct" comprises two or more covalently attached monomers derived from modified CsgF and / or CsgG or homologues thereof. In other words, a construct can contain more than one monomer. In another aspect, the application also provides a pore complex comprising at least one construct of the application. A pore complex contains sufficient constructs, and if necessary, monomers to form said pore. For example, an octameric pore can comprise (a) four constructs each containing two monomers, (b) two constructs each containing four monomers, (c) one construct containing two monomers and six monomers that do not form part of a construct, or (d) one or two CsgF monomers in one construct, and one construct with six to seven CsgG monomers, or even (e) a construct with CsgF and CsgG monomers in addition to another construct containing only CsgG monomers. The same and other possibilities are provided for a nonameric pore, for example. The skilled person can envisage other combinations of constructs and monomers. One or more constructs of the application can be used to form a pore complex for characterizing, such as sequencing, a polynucleotide. A construct can comprise at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9 or at least 10 monomers. A construct preferably comprises two monomers. The two or more monomers can be the same or different, can be CsgF, CsgG, CsgG / CsgF fusion monomers or homologues thereof, or any combination thereof.
[0140] Another embodiment relates to a polynucleotide or nucleic acid molecule encoding said modified CsgF peptide of the application or a homologue or mutant thereof, or a polynucleotide encoding a construct as described above.
[0141] Certain embodiments relate to an isolated transmembrane pore complex comprising an isolated pore complex according to the first and second aspects of the application, and a component of a membrane. The isolated transmembrane pore complex can be directly suitable for molecular sensing, such as nucleic acid sequencing. Alternatively, a membrane composition is provided comprising a modified CsgG / CsgF biological pore as described herein according to the isolated pore complex of the application, and a membrane, a component of a membrane or an insulating layer. One embodiment relates to an isolated transmembrane pore complex consisting of an isolated pore complex according to the application and a component of a membrane.
[0142] Although the CsgG:CsgF complex is very stable, when CsgF is truncated, the stability of the CsgG:CsgF complex is reduced compared to the complex comprising full-length CsgF. Thus, disulfide bonds can be created between CsgG and CsgF to make the complex more stable, for example after introducing cysteine residues at the positions identified herein. The pore complex can be prepared in any of the methods mentioned previously, and the formation of disulfide bonds can be induced by using an oxidizing agent (for example: copper- ortho-phenanthroline). Other interactions (for example: hydrophobic interactions, charge-charge interactions / electrostatic interactions) can also be used in those positions instead of cysteine interactions.
[0143] In another embodiment, non-natural amino acids can also be incorporated in those positions. In this embodiment, covalent bonds can be formed by click chemistry. For example, non-natural amino acids with an azide or alkyne or with a dibenzocyclooctyne (DBCO) group and / or a bicyclo[6.1.0]nonyne (BCN) group can be introduced at one or more of these positions.
[0144] Such stability mutations can be combined with any other modification to CsgG and / or CsgF, for example the modifications disclosed herein.
[0145] The CsgG pore can comprise at least one, for example 2, 3, 4, 5, 6, 7, 8, 9, or 10 CsgG monomers, which are modified to facilitate attachment to the CsgF peptide. For example, a cysteine residue can be introduced at one or more positions corresponding to positions 132, 133, 136, 138, 140, 142, 144, 145, 147, 149, 151, 153, 155, 183, 185, 187, 189, 191, 201, 203, 205, 207, and 209 of SEQ ID NO: 3, and / or at any of the positions identified in Table 4 as expected to be in contact with CsgF, to facilitate covalent attachment to CsgG. In addition to or instead of covalent attachment through a cysteine residue, the pore can be stabilized by hydrophobic interactions or electrostatic interactions. To facilitate such interactions, a non-natural reactive or photoreactive amino acid can be introduced at a position corresponding to one or more of positions 132, 133, 136, 138, 140, 142, 144, 145, 147, 149, 151, 153, 155, 183, 185, 187, 189, 191, 201, 203, 205, 207, and 209 of SEQ ID NO: 3, and / or at any of the positions identified in Table 4 as expected to be in contact with CsgF.
[0146] The CsgF peptide can be modified to facilitate attachment to the CsgG pore. For example, a cysteine residue can be introduced at one or more positions corresponding to positions 1, 4, 5, 8, 9, 11, 12, 26, or 29 of SEQ ID NO: 6, and / or at any one of the positions identified in Table 4 as predicted to be in contact with CsgF, to facilitate covalent attachment to CsgG. Alternatively or in addition to covalent attachment via a cysteine residue, the pore can be stabilized by hydrophobic or electrostatic interactions. To facilitate such interactions, a non-native reactive or photo-reactive amino acid can be introduced at a position corresponding to one or more of positions 1, 4, 5, 8, 9, 11, 12, 26, or 29 of SEQ ID NO: 6, and / or at any one of the positions identified in Table 4 as predicted to be in contact with CsgF.
[0147] Preferred exemplary CsgF peptides include the following mutations relative to SEQ ID NO: 6: N15X1 / N17X2 / A20X3 / N24X4 / A28X5 / D34X6, wherein X1 is N / S / A / T / Q / G / L / V / I / F / Y / W / R / K / D / C, X2 is N / S / A / T / Q / G / L / V / I / F / Y / W / R / K / D / C, X3 is A / S / T / Q / N / G / L / V / I / F / Y / W / R / K / D / C, X4 is N / S / T / Q / A / G / L / V / I / F / Y / W / R / K / D / C, X5 is A / S / T / Q / N / G / L / V / I / F / Y / W / R / K / D / C, and X5 is D / F / Y / W / R / K / N / Q / C. The mutations at positions N15, N17, A20, N24, and A28 are contraction mutations, and the mutation at position 34 affects the interaction of CsgF with the bottom of the CsgG pore to stabilize the interaction.
[0148] CsgG pore
[0149] The CsgG pore can be a homooligomeric pore comprising identical mutant monomers of the application. The CsgG pore can be a heterooligomeric pore derived from CsgG, e.g., comprising at least one mutant monomer as disclosed herein.
[0150] The CsgG pore can contain any number of mutant monomers. The pore typically comprises at least 7, at least 8, at least 9, or at least 10 identical mutant monomers, such as 7, 8, 9, or 10 mutant monomers. The CsgG pore preferably comprises eight or nine identical mutant monomers.
[0151] In a preferred embodiment, all of the monomers in the heterooligomeric CsgG pore, such as 10, 9, 8, or 7 of them, are mutant monomers as disclosed herein, wherein at least one of them is different from one another. They can all be different from one another.
[0152] The mutant monomers in the CsgG pore are preferably all approximately the same length or the same length. The barrels of the mutant monomers of the invention in the pore are preferably approximately the same length or the same length. Length can be measured in number of amino acids and / or length units.
[0153] The mutant monomer can be a variant of SEQ ID NO: 3. The variant will preferably be at least 50% homologous to the sequence over the entire length of the amino acid sequence of SEQ ID NO: 3, based on amino acid identity. More preferably, the variant can be at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90% and more preferably at least 95%, 97% or 99% homologous to the amino acid sequence of SEQ ID NO: 3, based on amino acid identity over the entire sequence. There can be at least 80%, for example at least 85%, 90% or 95% amino acid identity over a stretch of 100 or more (e.g. 125, 150, 175 or 200 or more) contiguous amino acids (“hard homology”).
[0154] CsgG monomers are highly conserved (as can be readily appreciated from Figures 45 to 47 of WO2017 / 149317). Furthermore, from the mutations relating to SEQ ID NO: 3, equivalent positions of mutations of CsgG monomers other than SEQ ID NO: 3 can be determined.
[0155] Thus, reference to a mutant CsgG monomer comprising a variant of the sequence shown as SEQ ID NO: 3 and particular amino acid mutations thereof as set out in the claims and specification also encompasses a mutant CsgG monomer comprising a variant of the sequence shown as SEQ ID NO: 68 to 88 and the corresponding amino acid mutations thereof. Likewise, reference to a construct, pore or method involving the use of a pore relating to a mutant CsgG monomer comprising a variant of the sequence shown as SEQ ID NO: 3 and particular amino acid mutations thereof as set out in the claims and specification also encompasses a construct, pore or method relating to a mutant CsgG monomer comprising a variant of the sequence according to SEQ ID NO above and the corresponding amino acid mutations thereof. It will be further appreciated that the invention extends to other variant CsgG monomers showing regions of high conservation as identified in the specification.
[0156] Homology can be determined using standard methods in the art. For example, the UWGCG package provides the BESTFIT program which performs a best-fit matching between two sequences, e.g., using its default settings (Devereux et al. (1984) Nucleic Acids Research 12, pp. 387-395). The PILEUP and BLAST algorithms can be used to calculate homology or to align sequences, such as identifying equivalent residues or corresponding sequences (typically using their default settings), as described by Altschul S.F. (1993) J Mol Evol 36:290-300; Altschul, S.F. et al. (1990) J Mol Biol 215:403-10. Software for performing BLAST analysis is publicly available through the National Center for Biotechnology Information (http: / / www.ncbi.nlm.nih.gov / ).
[0157] SEQ ID NO: 3 is a wild-type CsgG monomer from E. coli K-12 Subsp. MC4100. A variant of SEQ ID NO: 3 can comprise any substitution present in another CsgG homolog. Preferred CsgG homologs are shown in SEQ ID NOs: 68-88. A variant can comprise a combination of one or more substitutions present in SEQ ID NOs: 68-88 compared to SEQ ID NO: 3. For example, a mutation can be made at any one or more positions in SEQ ID NO: 3 that differ between SEQ ID NO: 3 and any one of SEQ ID NOs: 68-88. Such a mutation can be a substitution of the amino acid in SEQ ID NO: 3 with the amino acid from the corresponding position in any one of SEQ ID NOs: 68-88. Alternatively, the mutation at any one of these positions can be a substitution with any amino acid, or can be a deletion or insertion mutation, such as a deletion or insertion of 1-10 amino acids, such as 2-8 or 3-6 amino acids. In addition to the mutations disclosed herein, amino acids conserved between SEQ ID NO: 3 and all of SEQ ID NOs: 66-88 are preferably present in a variant of the application. However, a conservative mutation can be made at any one or more of these positions that are conserved between SEQ ID NO: 3 and all of SEQ ID NOs: 66-88.
[0158] The present application provides a pore-forming CsgG mutant monomer comprising any one or more of the amino acids described herein substituted into a particular position of SEQ ID NO: 3 at a position in the CsgG monomer structure corresponding to that particular position in SEQ ID NO: 3. The corresponding position can be determined by standard techniques in the art. For example, the PILEUP and BLAST algorithms mentioned above can be used to align the sequence of a CsgG monomer to SEQ ID NO: 3, thereby identifying the corresponding residue.
[0159] The pore-forming mutant monomer will generally retain the ability to form the same 3D structure as the wild-type CsgG monomer, such as the same 3D structure as a CsgG monomer having the sequence of SEQ ID NO: 3. The 3D structure of CsgG is known in the art and is disclosed, for example, in Goyal et al. (2014) Nature 516(7530): 250-3. In addition to the mutations described herein, any number of mutations can be made in the wild-type CsgG sequence, provided that the CsgG mutant monomer retains the improved properties conferred on it by the mutations of the present application.
[0160] Generally, the CsgG monomer will retain the ability to form a structure comprising three alpha-helices and five beta-sheets. Mutations can be made in the CsgG region at least at the N-terminus of the first alpha helix (beginning at S63 in SEQ ID NO: 3), in the second alpha helix (from G85 to A99 of SEQ ID NO: 3), in the loop between the second alpha helix and the first beta sheet (from Q100 to N120 of SEQ ID NO: 3), in the fourth and fifth beta sheets (S173 to R192 and R198 to T107 of SEQ ID NO: 3, respectively), and in the loop between the fourth and fifth beta sheets (F193 to Q197 of SEQ ID NO: 3), without affecting the ability of the CsgG monomer to form a transmembrane pore capable of translocating a polypeptide. Thus, it is contemplated that additional mutations can be made in any of these regions of any CsgG monomer without affecting the ability of the monomer to form a pore that can translocate a polynucleotide. It is also contemplated that mutations can be made in other regions, such as in any alpha helix (S63 to R76, G85 to A99, or V211 to L236 of SEQ ID NO: 3) or in any beta sheet (I121 to N133, K135 to R142, I146 to R162, S173 to R192, or R198 to T107 of SEQ ID NO: 3), without affecting the ability of the monomer to form a pore that can translocate a polynucleotide. It is also contemplated that a deletion of one or more amino acids can be made in any loop region connecting an alpha helix and a beta sheet and / or in the N-terminal and / or C-terminal regions of the CsgG monomer without affecting the ability of the monomer to form a pore that can translocate a polynucleotide.
[0161] In addition to those amino acid substitutions discussed above, amino acid substitutions can be made to the amino acid sequence of SEQ ID NO: 3, for example, up to 1, 2, 3, 4, 5, 10, 20, or 30 substitutions. Conservative substitutions replace an amino acid with another amino acid of similar chemical structure, similar chemical properties, or similar side chain volume. The introduced amino acids can have similar polarity, hydrophilicity, hydrophobicity, basicity, acidity, neutrality, or charge as the amino acids they replace. Alternatively, a conservative substitution can introduce another aromatic or aliphatic amino acid in place of a pre-existing aromatic or aliphatic amino acid. Conservative amino acid changes are well known in the art and can be selected according to the properties of the 20 primary amino acids defined in Table 1 above. In cases where the amino acids have similar polarity, this can also be determined with reference to the scale of hydrophilicity of the side chains of the amino acids in Table 2.
[0162] One or more amino acid residues of the amino acid sequence of SEQ ID NO: 3 can additionally be deleted from the above polypeptides. Up to 1, 2, 3, 4, 5, 10, 20, or 30 or more residues can be deleted.
[0163] Variants can include fragments of SEQ ID NO: 3. Such fragments retain pore-forming activity. Fragments can be at least 50, at least 100, at least 150, at least 200, or at least 250 amino acids in length. Such fragments can be used to generate pores. Fragments preferably comprise the transmembrane domains of SEQ ID NO: 3, i.e., K135-Q153 and S183-S208.
[0164] Alternatively or additionally, one or more amino acids can be added to the above polypeptides. Extensions can be provided at the amino or carboxy terminus of the amino acid sequence of SEQ ID NO: 3 or a polypeptide variant or fragment thereof. Extensions can be short, for example, 1 to 10 amino acids in length. Alternatively, extensions can be longer, for example, up to 50 or 100 amino acids. A carrier protein can be fused to the amino acid sequence according to the application. Other fusion proteins are discussed in more detail below.
[0165] A CsgG pore as described herein includes a wild-type CsgG pore, or a homolog or mutant / variant thereof. A variant is a polypeptide having an amino acid sequence that differs from the amino acid sequence of SEQ ID NO: 3, and which retains its ability to form a pore. A variant typically contains the region of SEQ ID NO: 3 responsible for pore formation. The pore-forming ability of a CsgG containing a beta-barrel is provided by the beta-sheets in each subunit. A variant of SEQ ID NO: 3 typically comprises the regions of SEQ ID NO: 3 that form beta-sheets, i.e. K134-Q154 and S183-S208. One or more modifications can be made to the regions of SEQ ID NO: 3 that form beta-sheets, so long as the resulting variant retains its ability to form a pore. A variant of SEQ ID NO: 3 preferably comprises one or more modifications, such as substitutions, additions or deletions, within its alpha-helices and / or loop regions.
[0166] A mutant CsgG monomer can be a mutant CsgG monomer that is a monomer having a sequence that differs from the sequence of a wild-type CsgG monomer and which retains the ability to form a pore. A mutant monomer can also be referred to herein as a variant. Methods of confirming the ability of a mutant monomer to form a pore are well known in the art and are discussed in more detail below.
[0167] A particular pore-forming CsgG mutant monomer that can be included in a CsgG pore can comprise one or more of the following modifications:
[0168] - W at a position corresponding to R97 in SEQ ID NO: 3;
[0169] - W at a position corresponding to R93 in SEQ ID NO: 3;
[0170] - Y at a position corresponding to R97 in SEQ ID NO: 3;
[0171] - Y at a position corresponding to R93 in SEQ ID NO: 3;
[0172] - Y at each of the positions corresponding to R93 and R97 in SEQ ID NO: 3;
[0173] - D at a position corresponding to R192 in SEQ ID NO: 3;
[0174] - deletion of residues at positions corresponding to V105-I107 in SEQ ID NO: 3;
[0175] - deletion of residues at one or more positions corresponding to F193 to L199 in SEQ ID NO: 3;
[0176] - Residues are missing at positions corresponding to F195 to L199 in SEQ ID NO:3;
[0177] - Residues at positions corresponding to F193 to L199 in SEQ ID NO:3 are missing;
[0178] -T is located at the position corresponding to F191 in SEQ ID NO:3;
[0179] - Q is located at the position corresponding to K49 in SEQ ID NO:3;
[0180] - N is located at the position corresponding to K49 in SEQ ID NO:3;
[0181] - Q is located at the position corresponding to K42 in SEQ ID NO:3;
[0182] - Q is located at the position corresponding to E44 in SEQ ID NO:3;
[0183] -N is located at the position corresponding to E44 in SEQ ID NO:3;
[0184] - R is located at the position corresponding to L90 in SEQ ID NO:3;
[0185] - R is located at the position corresponding to L91 in SEQ ID NO:3;
[0186] - R is located at the position corresponding to I95 in SEQ ID NO:3;
[0187] - R is located at the position corresponding to A99 in SEQ ID NO:3;
[0188] - H is located at the position corresponding to E101 in SEQ ID NO:3;
[0189] -K is located at the position corresponding to E101 in SEQ ID NO:3;
[0190] -N is located at the position corresponding to E101 in SEQ ID NO:3;
[0191] - Q is located at the position corresponding to E101 in SEQ ID NO:3;
[0192] -T is located at the position corresponding to E101 in SEQ ID NO:3;
[0193] - K is located at the position corresponding to Q114 in SEQ ID NO:3.
[0194] The CsgG pore-forming monomer preferably further comprises an A at a position corresponding to Y51 in SEQ ID NO: 3 and / or a Q at a position corresponding to F56 in SEQ ID NO: 3.
[0195] When characterising (or sequencing) a target polynucleotide, a pore constructed from a CsgG monomer having a R to W substitution at a position corresponding to position 97 of SEQ ID NO: 3 shows improved accuracy compared to otherwise identical pores without modification at 97. Improved accuracy is also seen when the CsgG monomer comprises a R to Y modification at positions corresponding to positions 93 and 97 of SEQ ID NO: 3 instead of R97W. Thus, the pore can be constructed from one or more mutant CsgG monomers comprising a modification at a position corresponding to R97 or R93 of SEQ ID NO: 3 such that the modification increases the hydrophobicity of the amino acid. For example, such a modification can comprise an amino acid substitution with any amino acid containing a hydrophobic side chain, including but not limited to W and Y.
[0196] A mutant CsgG monomer comprising a R to D, Q, F, S or T at a position corresponding to position 192 of SEQ ID NO: 3 is more readily expressed than a monomer without a substitution at position 192, which can be due to a reduction in positive charge. Thus, position 192 can be substituted with an amino acid that reduces the positive charge. A monomer comprising R192D / Q / F / S / T can also comprise other modifications that improve the ability of a mutant pore formed from the monomer to interact with and characterise an analyte such as a polynucleotide. However, in one embodiment it is preferred that the residue at a position corresponding to position 193 of SEQ ID NO: 3 is R or K, more preferably R.
[0197] A pore comprising a CsgG monomer comprising a deletion of V105, A106 and I107, a deletion of F193, I194, D195, Y196, Q197, R198 and L199 or a deletion of D195, Y196, Q197, R198 and L199 and / or F191T shows improved accuracy when characterising (or sequencing) a target polynucleotide. The amino acids at positions 105 to 107 correspond to a cis-loop in the nanopore cap, while the amino acids at positions 193 to 199 correspond to a trans-loop on the other end of the pore. Without wishing to be bound by theory, it is thought that deletion of the cis-loop improves the interaction of the enzyme with the pore, while removal of the trans-loop reduces any undesirable interactions between DNA on the opposite face of the pore.
[0198] Pores comprising CsgG monomers comprising a mutation of K to Q or K to N at a position corresponding to K94 of SEQ ID NO: 3 show a decrease in the number of noise holes (i.e. those that result in an increase in signal to noise ratio) when used in the characterization (or sequencing) of a target polynucleotide compared to the same pore with no mutation at 94. Position 94 is found within the vestibule of the pore and this position is found to be particularly sensitive to noise with respect to the current signal.
[0199] Pores comprising CsgG monomers comprising T104K or T104R, N91R, E101K / N / Q / T / H, E44N / Q, Q114K, A99R, I95R, N91R, L90R, E44Q / N, and / or Q42K or corresponding mutations all exhibit an increase in the ability to capture target polynucleotides when used in the characterization (or sequencing) of a target polynucleotide compared to the same pore with no substitutions at these positions.
[0200] In one embodiment, the CsgG pore comprises one or more monomers that are variants of SEQ ID NO: 3 comprising (a) one or more mutations (i.e. mutations at one or more of the following positions) I41, R93, A98, Q100, G103, T104, A106, I107, N108, L113, S115, T117, Y130, K135, E170, S208, D233, D238, and E244 and / or (b) one or more of D43S, E44S, F48S / N / Q / Y / W / I / V / H / R / K, Q87N / R / K, N91K / R, K94R / F / Y / W / L / S / N, R97F / Y / W / V / I / K / S / Q / H, E101I / L / A / H, N102K / Q / L / I / V / S / H, R110F / G / N, Q114R / K, R142Q / S, T150Y / A / V / L / S / Q / N, R192D / Q / F / S / T, and D248S / N / Q / K / R. The variant can comprise (a); (b); or (a) and (b). In some embodiments, the variant comprises R97W. In some embodiments, the variant comprises R192D / Q / F / S / T, such as R192D / Q. In (a), the variant can comprise a modification at any number of positions and combinations of positions, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, or 19 positions.
[0201] In (a), the variant preferably comprises one or more of I41N, R93F / Y / W / L / I / V / N / Q / S, A98K / R, Q100K / R, G103F / W / S / N / K / R, T104R / K, A106R / K, I107R / K / W / F / Y / L / V, N108R / K, L113K / R, S115R / K, T117R / K, Y130W / F / H / Q / N, K135L / V / N / Q / S, E170S / N / Q / K / R, S208V / I / F / W / Y / L / T, D233S / N / Q / K / R, D238S / N / Q / K / R, and E244S / N / Q / K / R.
[0202] In (a), the variant preferably comprises one or more modifications that provide for more consistent movement of a target polynucleotide relative to (such as by) a transmembrane pore comprising monomers. In particular, in (a), the variant preferably comprises one or more mutations (i.e., mutations at one or more of the following positions) at the following positions: R93, G103, and I107. The variant can comprise R93; G103; I107; R93 and G103; R93 and I107; G103 and I107; or R93, G103, and I107. The variant preferably comprises one or more of R93F / Y / W / L / I / V / N / Q / S, G103F / W / S / N / K / R, and I107R / K / W / F / Y / L / V. These can be present in any combination as indicated for positions R93, G103, and I107.
[0203] In (a), the variant preferably comprises one or more modifications that allow for a pore built from the mutant monomers to preferably more readily capture nucleotides and polynucleotides. In particular, in (a), the variant preferably comprises one or more mutations (i.e., mutations at one or more of the following positions) at the following positions: I41, T104, A106, N108, L113, S115, T117, E170, D233, D238, and E244. The variant can comprise modifications at any number of positions and combinations of positions, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or 11 positions. The variant preferably comprises one or more of I41N, T104R / K, A106R / K, N108R / K, L113K / R, S115R / K, T117R / K, E170S / N / Q / K / R, D233S / N / Q / K / R, D238S / N / Q / K / R, and E244S / N / Q / K / R. Additionally or alternatively, the variant can comprise (c) Q42K / R, E44N / Q, L90R / K, N91R / K, I95R / K, A99R / K, E101H / K / N / Q / T, and / or Q114K / R.
[0204] In (a), the variant preferably comprises one or more modifications which provide more consistent movement and increase capture. In particular, in (a), the variant preferably comprises one or more mutations (i.e. mutations at one or more of the following positions) at (i) A98, (ii) Q100, (iii) G103 and (iv) I107. The variant preferably comprises one or more of (i) A98R / K, (ii) Q100K / R, (iii) G103K / R and (iv) I107R / K.
[0205] Particularly preferred mutant monomers which increase capture of an analyte such as a polynucleotide include mutations at one or more of positions Q42, E44, E44, L90, N91, I95, A99, E101 and Q114 which remove the negative charge at the mutated position and / or increase the positive charge at the mutated position. In particular, the following mutations can be included in a mutant monomer of the application to produce a CsgG pore with improved ability to capture an analyte (preferably a polynucleotide): Q42K, E44N, E44Q, L90R, N91R, I95R, A99R, E101H, E101K, E101N, E101Q, E101T and Q114K. Examples of particular mutant monomers comprising one of these mutations in combination with other beneficial mutations are:
[0206] CsgG-(WT-Y51A / F56Q / K94Q / R97W / R192D-del(V105-I107)-Q42K
[0207] CsgG-(WT-Y51A / F56Q / K94Q / R97W / R192D-del(V105-I107)-E44N
[0208] CsgG-(WT-Y51A / F56Q / K94Q / R97W / R192D-del(V105-I107)-E44Q
[0209] CsgG-(WT-Y51A / F56Q / K94Q / R97W / R192D-del(V105-I107)-L90R
[0210] CsgG-(WT-Y51A / F56Q / K94Q / R97W / R192D-del(V105-I107)-N91R
[0211] CsgG-(WT-Y51A / F56Q / K94Q / R97W / R192D-del(V105-I107)-I95R
[0212] CsgG-(WT-Y51A / F56Q / K94Q / R97W / R192D-del(V105-I107)-E101H
[0213] CsgG-(WT-Y51A / F56Q / K94Q / R97W / R192D-del(V105-I107)-E101K
[0214] CsgG-(WT-Y51A / F56Q / K94Q / R97W / R192D-del(V105-I107)-E101N
[0215] CsgG-(WT-Y51A / F56Q / K94Q / R97W / R192D-del(V105-I107)-E101Q
[0216] CsgG-(WT-Y51A / F56Q / K94Q / R97W / R192D-del(V105-I107)-E101T
[0217] CsgG-(WT-Y51A / F56Q / K94Q / R97W / R192D-del(V105-I107)-E101T
[0218] CsgG-(WT-Y51A / F56Q / K94Q / R97W / R192D-del(V105-I107)-Q114K.
[0219] In (a), the variant preferably comprises one or more modifications that provide for increased characterization accuracy. In particular, in (a), the variant preferably comprises one or more mutations (i.e., mutations at one or more of the following positions) at the following positions: Y130, K135, and S208, such as Y130; K135; S208; Y130 and K135; Y130 and S208; K135 and S208; or Y130, K135, and S208. The variant preferably comprises one or more of Y130W / F / H / Q / N, K135L / V / N / Q / S, and R142Q / S. These substitutions can be present in any number and combination as set forth for Y130, K135, and S208.
[0220] In (b), the variant can comprise any number and combination of substitutions, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 substitutions. In (b), the variant preferably comprises one or more modifications that provide more consistent movement of the target polynucleotide relative to (such as by) the transmembrane pore comprising monomers. In particular, in (b), the variant preferably comprises one or more of (i) Q87N / R / K, (ii) K94R / F / Y / W / L / S / N, (iii) R97F / Y / W / V / I / K / S / Q / H, (iv) N102K / Q / L / I / V / S / H, and (v) R110F / G / N. More preferably, the variant comprises K94D or K94Q and / or R97W or R97Y. Other preferred variants modified to provide more consistent movement of the target polynucleotide relative to (for example by) the transmembrane pore comprising monomers include (vi) R93W and R93Y. Preferred variants can comprise R93W and R97W, R93Y and R97W, R93W and R97W, or more preferably R93Y and R97Y.
[0221] In (b), the variant preferably comprises one or more modifications that allow the pore built from the mutant monomers to preferentially more easily capture nucleotides and polynucleotides. In particular, in (b), the variant preferably comprises one or more of (i) D43S, (ii) E44S, (iii) N91K / R, (iv) Q114R / K, and (v) D248S / N / Q / K / R.
[0222] In (b), the variant preferably comprises one or more modifications that provide more consistent movement and increased capture. In particular, in (b), the variant preferably comprises one or more of Q87R / K, E101I / L / A / H, and N102K, such as Q87R / K; E101I / L / A / H; N102K; Q87R / K and E101I / L / A / H; Q87R / K and N102K; E101I / L / A / H and N102K; or Q87R / K, E101I / L / A / H, and N102K.
[0223] In (b), the variant preferably comprises one or more modifications that provide increased characterization accuracy. In particular, in (a), the variant preferably comprises F48S / N / Q / Y / W / I / V.
[0224] In (b), the variant preferably comprises one or more modifications that provide increased characterization accuracy and increased capture. In particular, in (a), the variant preferably comprises F48H / R / K.
[0225] The variants can comprise modifications in (a) and (b) that provide more consistent movement. The variants can comprise modifications in (a) and (b) that provide increased capture.
[0226] The present disclosure provides variants of SEQ ID NO: 3 that use pores comprising the variants provide increased throughput for assays for characterizing an analyte such as a polynucleotide. Such variants can comprise a mutation at K94, preferably K94Q or K94N, more preferably K94Q. Examples of particular mutant monomers comprising a K94Q or K94N mutation in combination with other beneficial mutations are:
[0227] CsgG-(WT-Y51A / F56Q / R97W / R192D-StrepII)9-K94N
[0228] CsgG-(WT-Y51A / F56Q / R97W / R192D-StrepII)9-K94Q.
[0229] Use of monomers that are variants of SEQ ID NO: 3 to form CsgG pores can provide improved characterization accuracy in assays for characterizing an analyte such as a polynucleotide. Such variants include variants comprising a mutation at F191, preferably F191T; a deletion of V105-I107; a deletion of F193-L199 or D195-L199; and / or a mutation at R93 and / or R97, preferably R93Y, R97Y, or more preferably R97W, R93W, or both R97W and R97Y. Examples of particular mutant monomers comprising one or more of these mutations in combination with other beneficial mutations are:
[0230] CsgG-(WT-Y51A / F56Q / R97W / R192D-StrepII)9-del(D195-L199)
[0231] CsgG-(WT-Y51A / F56Q / R97W / R192D-StrepII)9-del(F193-L199)
[0232] CsgG-(WT-Y51A / F56Q / R97W / R192D-StrepII)9-F191T
[0233] CsgG-(WT-Y51A / F56Q / R97W / R192D-del(V105-I107)-StrepII)9
[0234] CsgG-(WT-Y51A / F56Q / K94Q / R97W / R192D-del(V105-I107
[0235] CsgG-(WT-Y51A / F56Q / R192D-StrepII)9-R93W
[0236] CsgG-(WT-Y51A / F56Q / R192D-StrepII)9-R93W-del(D195-L199)
[0237] CsgG-(WT-Y51A / F56Q / R192D-StrepII)9-R93Y / R97Y.
[0238] In another embodiment, the variant of SEQ ID NO: 3 comprises (A) a deletion of one or more of positions R192, F193, I194, D195, Y196, Q197, R198, L199, L200, and E201 and / or (B) a deletion of one or more of V139 / G140 / D149 / T150 / V186 / Q187 / V204 / G205 (referred to herein as Band 1), G137 / G138 / Q151 / Y152 / Y184 / E185 / Y206 / T207 (referred to herein as Band 2), and A141 / R142 / G147 / A148 / A188 / G189 / G202 / E203 (referred to herein as Band 3).
[0239] In (A), the variant can comprise a deletion of any number of positions and combinations of positions, such as a deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 positions. In (A), the variant preferably comprises a deletion of:
[0240] D195, Y196, Q197, R198, and L199;
[0241] R192, F193, I194, D195, Y196, Q197, R198, L199, and L200;
[0242] Q197, R198, L199, and L200;
[0243] I194, D195, Y196, Q197, R198, and L199;
[0244] D195, Y196, Q197, R198, L199, and L200;
[0245] Y196, Q197, R198, L199, L200, and E201;
[0246] Q197, R198, L199, L200, and E201;
[0247] Q197, R198, L199; or
[0248] - F193, I194, D195, Y196, Q197, R198, and L199.
[0249] More preferably, the variant comprises a deletion of D195, Y196, Q197, R198, and L199 or F193, I194, D195, Y196, Q197, R198, and L199. In (B), any number of bands 1-3 and combinations thereof can be deleted, e.g., band 1; band 2; band 3; bands 1 and 2; bands 1 and 3; bands 2 and 3; or bands 1, 2, and 3. The variant can comprise a deletion according to (A); (B); or (A) and (B).
[0250] A variant comprising a deletion according to one or more positions of (A) and / or (B) above can also comprise any modification or substitution discussed above and below. If one or more positions that occur after the deletion position in SEQ ID NO: 3 are modified or substituted, the numbering of the modified or substituted position(s) must be adjusted accordingly. For example, if L199 is deleted, E244 becomes E243. Similarly, if band 1 is deleted, R192 becomes R186.
[0251] In another embodiment, the variant of SEQ ID NO: 3 comprises (C) a deletion of one or more positions V105, A106, and I107. Deletions according to (C) can be made in addition to deletions according to (A) and / or (B).
[0252] The deletions described above generally reduce noise associated with movement of a target polynucleotide relative to (e.g., through) a transmembrane pore comprising monomers. As a result, the target polynucleotide can be more accurately characterized.
[0253] In the paragraphs above that separate different amino acids of a particular position with a / symbol, the / symbol means “or.” For example, Q87R / K means Q87R or Q87K.
[0254] A variant of SEQ ID NO: 3 that increases capture of an analyte, such as a polynucleotide, can comprise a mutation at T104, preferably T104R or T104K; a mutation at N91, preferably N91R; a mutation at E101, preferably E101K / N / Q / T / H; a mutation at position E44, preferably E44N or E44Q, and / or a mutation at position Q42, preferably Q42K.
[0255] Mutations at different positions in SEQ ID NO: 3 can be combined in any possible way. In particular, a monomer in a CsgG pore can comprise one or more mutations that increase accuracy, one or more mutations that reduce noise, and / or one or more mutations that enhance analyte capture.
[0256] A variant of SEQ ID NO: 3 preferably comprises one or more of the following: (i) one or more mutations (i.e. mutations at one or more of the following positions) at positions N40, D43, E44, S54, S57, Q62, R97, E101, E124, E131, R142, T150, and R192, such as one or more mutations (i.e. mutations at one or more of the following positions) at positions N40, D43, E44, S54, S57, Q62, E101, E131, and T150 or N40, D43, E44, E101, and E131; (ii) a mutation at 51 / N55, Y51 / F56, N55 / F56, or Y51 / N55 / F56; (iii) Q42R or Q42K; (iv) K49R; (v) N102R, N102F, N102Y, or N102W; (vi) D149N, D149Q, or D149R; (vii) E185N, E185Q, or E185R; (viii) D195N, D195Q, or D195R; (ix) E201N, E201Q, or E201R; (x) E203N, E203Q, or E203R; and (xi) a deletion at one or more of the following positions: F48, K49, P50, Y51, P52, A53, S54, N55, F56, and S57. A variant can comprise any combination of (i) to (xi).
[0257] If a variant comprises any of (i) and (iii) to (xi), it can also comprise one or more of Y51, N55, and F56, such as a mutation at Y51, N55, F56, Y51 / N55, Y51 / F56, N55 / F56, or Y51 / N55 / F56.
[0258] In (i), the variant can comprise mutations in any number and combination of N40, D43, E44, S54, S57, Q62, R97, E101, E124, E131, R142, T150, and R192. In (i), the variant preferably comprises one or more mutations (i.e., mutations at one or more of the following positions) at positions N40, D43, E44, S54, S57, Q62, E101, E131, and T150. In (i), the variant preferably comprises one or more mutations (i.e., mutations at one or more of the following positions) at positions N40, D43, E44, E101, and E131. In (i), the variant preferably comprises a mutation at S54 and / or S57. In (i), the variant more preferably comprises mutations at (a) S54 and / or S57 and (b) one or more of Y51, N55, and F56, such as Y51, N55, F56, Y51 / N55, Y51 / F56, N55 / F56, or Y51 / N55 / F56. If S54 and / or S57 is / are absent in (xi), it / they cannot be mutated in (i), and vice versa. In (i), the variant preferably comprises a mutation at T150, such as T150I. Alternatively, the variant preferably comprises mutations at (a) T150 and (b) one or more of Y51, N55, and F56, such as Y51, N55, F56, Y51 / N55, Y51 / F56, N55 / F56, or Y51 / N55 / F56. In (i), the variant preferably comprises a mutation at Q62, such as Q62R or Q62K. Alternatively, the variant preferably comprises mutations at (a) Q62 and (b) one or more of Y51, N55, and F56, such as Y51, N55, F56, Y51 / N55, Y51 / F56, N55 / F56, or Y51 / N55 / F56. The variant can comprise mutations at D43, E44, Q62, or any combination thereof, such as D43, E44, Q62, D43 / E44, D43 / Q62, E44 / Q62, or D43 / E44 / Q62. Alternatively, the variant preferably comprises mutations at (a) D43, E44, Q62, D43 / E44, D43 / Q62, E44 / Q62, or D43 / E44 / Q62 and (b) one or more of Y51, N55, and F56, such as Y51, N55, F56, Y51 / N55, Y51 / F56, N55 / F56, or Y51 / N55 / F56.
[0259] In (ii) and elsewhere in this application, the / symbol separating different positions means "and", such that Y51 / N55 is Y51 and N55. In (ii), the variant preferably comprises a mutation at Y51 / N55. It has been suggested that the constriction in CsgG is made up of three stacked concentric circles formed by the side chains of residues Y51, N55 and F56 (Goyal et al., 2014, Nature, 516, 250-253). Therefore, mutation of these residues in (ii) can reduce the number of nucleotides contributing to the current when the polynucleotide is moved through the pore, making it easier to identify a direct relationship between the observed current (when the polynucleotide is moved through the pore) and the polynucleotide. F56 can be mutated in any of the ways discussed below with reference to the variants and pores useful in the methods of the application.
[0260] In (v), the variant can comprise N102R, N102F, N102Y or N102W. The variant preferably comprises (a) N102R, N102F, N102Y or N102W and (b) a mutation at one or more of Y51, N55 and F56, such as at Y51, N55, F56, Y51 / N55, Y51 / F56, N55 / F56 or Y51 / N55 / F56.
[0261] In (xi), any number and combination of K49, P50, Y51, P52, A53, S54, N55, F56 and S57 can be deleted. Preferably, one or more of K49, P50, Y51, P52, A53, S54, N55 and S57 can be deleted. If Y51, N55 and F56 are deleted in (xi), it / they cannot be mutated in (ii) and vice versa.
[0262] In (i), the variant preferably comprises one or more of the following substitutions: N40R, N40K, D43N, D43Q, D43R, D43K, E44N, E44Q, E44R, E44K, S54P, S57P, Q62R, Q62K, R97N, R97G, R97L, E101N, E101Q, E101R, E101K, E101F, E101Y, E101W, E124N, E124Q, E124R, E124K, E124F, E124Y, E124W, E131D, R142E, R142N, T150I, R192E, and R192N, such as one or more of N40R, N40K, D43N, D43Q, D43R, D43K, E44N, E44Q, E44R, E44K, S54P, S57P, Q62R, Q62K, E101N, E101Q, E101R, E101K, E101F, E101Y, E101W, E131D, and T150I, or one or more of N40R, N40K, D43N, D43Q, D43R, D43K, E44N, E44Q, E44R, E44K, E101N, E101Q, E101R, E101K, E101F, E101Y, E101W, and E131D. The variant can comprise any number of these substitutions and combinations thereof. In (i), the variant preferably comprises S54P and / or S57P. In (i), the variant preferably comprises (a) S54P and / or S57P and (b) one or more at Y51, N55, and F56, such as a mutation at Y51, N55, F56, Y51 / N55, Y51 / F56, N55 / F56, or Y51 / N55 / F56. The mutation at one or more of Y51, N55, and F56 can be any of the mutations discussed below. In (i), the variant preferably comprises F56A / S57P or S54P / F56A. The variant preferably comprises T150I. Alternatively, the variant preferably comprises a mutation at (a) T150I and (b) one or more of Y51, N55, and F56, such as Y51, N55, F56, Y51 / N55, Y51 / F56, N55 / F56, or Y51 / N55 / F56.
[0263] In (i), the variant preferably comprises Q62R or Q62K. Alternatively, the variant preferably comprises (a) Q62R or Q62K and (b) a mutation at one or more of Y51, N55, and F56, such as at Y51, N55, F56, Y51 / N55, Y51 / F56, N55 / F56, or Y51 / N55 / F56. The variant can comprise D43N, E44N, Q62R, or Q62K, or any combination thereof, such as D43N, E44N, Q62R, Q62K, D43N / E44N, D43N / Q62R, D43N / Q62K, E44N / Q62R, E44N / Q62K, D43N / E44N / Q62R, or D43N / E44N / Q62K. Alternatively, the variant preferably comprises (a) D43N, E44N, Q62R, Q62K, D43N / E44N, D43N / Q62R, D43N / Q62K, E44N / Q62R, E44N / Q62K, D43N / E44N / Q62R, or D43N / E44N / Q62K and (b) a mutation at one or more of Y51, N55, and F56, such as at Y51, N55, F56, Y51 / N55, Y51 / F56, N55 / F56, or Y51 / N55 / F56.
[0264] In (i), the variant preferably comprises D43N.
[0265] In (i), the variant preferably comprises E101R, E101S, E101F, or E101N.
[0266] In (i), the variant preferably comprises E124N, E124Q, E124R, E124K, E124F, E124Y, E124W, or E124D, such as E124N.
[0267] In (i), the variant preferably comprises R142E and R142N.
[0268] In (i), the variant preferably comprises R97N, R97G, or R97L.
[0269] In (i), the variant preferably comprises R192E and R192N.
[0270] In (ii), the variant preferably comprises F56N / N55Q, F56N / N55R, F56N / N55K, F56N / N55S, F56N / N55G, F56N / N55A, F56N / N55T, F56Q / N55Q, F56Q / N55R, F56Q / N55K, F56Q / N55S, F56Q / N55G, F56Q / N55A, F56Q / N55T, F56R / N55Q, F56R / N55R, F56R / N55K, F56R / N55S, F56R / N55G, F56R / N55A, F56R / N55T, F56S / N55Q, F56S / N55R, F56S / N55K, F56S / N55S, F56S / N55G, F56S / N55A, F56S / N55T, F56G / N55Q, F56G / N55R, F56G / N55K, F56G / N55S, F56G / N55G, F56G / N55A, F56G / N55T, F56A / N55Q, F56A / N55R, F56A / N55K, F56A / N55S, F56A / N55G, F56A / N55A, F56A / N55T, F56K / N55Q, F56K / N55R, F56K / N55K, F56K / N55S, F56K / N55G, F56K / N55A, F56K / N55T, F56N / Y51L, F56N / Y51V, F56N / Y51A, F56N / Y51N, F56N / Y51Q, F56N / Y51S, F56N / Y51G, F56Q / Y51L, F56Q / Y51V, F56Q / Y51A, F56Q / Y51N, F56Q / Y51Q, F56Q / Y51S, F56Q / Y51G, F56R / Y51L, F56R / Y51V, F56R / Y51A, F56R / Y51N, F56R / Y51Q, F56R / Y51S, F56R / Y51G, F56S / Y51L, F56S / Y51V, F56S / Y51A, F56S / Y51N, F56S / Y51Q, F56S / Y51S, F56S / Y51G, F56G / Y51L, F56G / Y51V, F56G / Y51A, F56G / Y51N, F56G / Y51Q, F56G / Y51S, F56G / Y51G, F56A / Y51L, F56A / Y51V, F56A / Y51A, F56A / Y51N, F56A / Y51Q, F56A / Y51S, F56A / Y51G, F56K / Y51L, F56K / Y51V, F56K / Y51A, F56K / Y51N, F56K / Y51Q, F56K / Y51S, F56K / Y51G,N55Q / Y51L, N55Q / Y51V, N55Q / Y51A, N55Q / Y51N, N55Q / Y51Q, N55Q / Y51S, N55Q / Y51G, N55R / Y51L, N55R / Y51V, N55R / Y51A, N55R / Y51N, N55R / Y51Q, N55R / Y51S, N55R / Y51G, N55K / Y51L, N55K / Y51V, N55K / Y51A, N55K / Y51N, N55K / Y51Q, N55K / Y51S, N55K / Y51G, N55S / Y51L, N55S / Y51V, N55S / Y51A, N55S / Y51N, N55S / Y51Q, N55S / Y51S, N55S / Y51G, N55G / Y51L, N55G / Y51V, N55G / Y51A, N55G / Y51N, N55G / Y51Q, N55G / Y51S, N55G / Y51G, N55A / Y51L, N55A / Y51V, N55A / Y51A, N55A / Y51N, N55A / Y51Q, N55A / Y51S, N55A / Y51G, N55T / Y51L, N55T / Y51V, N55T / Y51A, N55T / Y51N, N55T / Y51Q, N55T / Y51S, N55T / Y51G, F56N / N55Q / Y51L, F56N / N55Q / Y51V, F56N / N55Q / Y51A, F56N / N55Q / Y51N, F56N / N55Q / Y51Q, F56N / N55Q / Y51S, F56N / N55Q / Y51G, F56N / N55R / Y51L, F56N / N55R / Y51V, F56N / N55R / Y51A, F56N / N55R / Y51N, F56N / N55R / Y51Q, F56N / N55R / Y51S, F56N / N55R / Y51G, F56N / N55K / Y51L, F56N / N55K / Y51V, F56N / N55K / Y51A, F56N / N55K / Y51N, F56N / N55K / Y51Q, F56N / N55K / Y51S, F56N / N55K / Y51G, F56N / N55S / Y51L, F56N / N55S / Y51V, F56N / N55S / Y51A, F56N / N55S / Y51N, F56N / N55S / Y51Q, F56N / N55S / Y51S, F56N / N55S / Y51G, F56N / N55G / Y51L, F56N / N55G / Y51V, F56N / N55G / Y51A, F56N / N55G / Y51N, F56N / N55G / Y51Q, F56N / N55G / Y51S,F56N / N55G / Y51L, F56N / N55G / Y51V, F56N / N55G / Y51A, F56N / N55G / Y51N, F56N / N55G / Y51Q, F56N / N55G / Y51S, F56N / N55G / Y51G, F56N / N55A / Y51L, F56N / N55A / Y51V, F56N / N55A / Y51A, F56N / N55A / Y51N, F56N / N55A / Y51Q, F56N / N55A / Y51S, F56N / N55A / Y51G, F56N / N55T / Y51L, F56N / N55T / Y51V, F56N / N55T / Y51A, F56N / N55T / Y51N, F56N / N55T / Y51Q, F56N / N55T / Y51S, F56N / N55T / Y51G, F56Q / N55Q / Y51L, F56Q / N55Q / Y51V, F56Q / N55Q / Y51A, F56Q / N55Q / Y51N, F56Q / N55Q / Y51Q, F56Q / N55Q / Y51S, F56Q / N55Q / Y51G, F56Q / N55R / Y51L, F56Q / N55R / Y51V, F56Q / N55R / Y51A, F56Q / N55R / Y51N, F56Q / N55R / Y51Q, F56Q / N55R / Y51S, F56Q / N55R / Y51G, F56Q / N55K / Y51L, F56Q / N55K / Y51V, F56Q / N55K / Y51A, F56Q / N55K / Y51N, F56Q / N55K / Y51Q, F56Q / N55K / Y51S, F56Q / N55K / Y51G, F56Q / N55S / Y51L, F56Q / N55S / Y51V, F56Q / N55S / Y51A, F56Q / N55S / Y51N, F56Q / N55S / Y51Q, F56Q / N55S / Y51S, F56Q / N55S / Y51G, F56Q / N55G / Y51L, F56Q / N55G / Y51V, F56Q / N55G / Y51A, F56Q / N55G / Y51N, F56Q / N55G / Y51Q, F56Q / N55G / Y51S, F56Q / N55G / Y51G, F56Q / N55A / Y51L, F56Q / N55A / Y51V, F56Q / N55A / Y51A, F56Q / N55A / Y51N, F56Q / N55A / Y51Q, F56Q / N55A / Y51S, F56Q / N55A / Y51G, F56Q / N55T / Y51L, F56Q / N55T / Y51V, F56Q / N55T / Y51A, F56Q / N55T / Y51N, F56Q / N55T / Y51Q, F56Q / N55T / Y51S, F56Q / N55T / Y51G, F56R / N55Q / Y51L, F56R / N55Q / Y51V,F56R / N55Q / Y51L, F56R / N55Q / Y51V, F56R / N55Q / Y51A, F56R / N55Q / Y51N, F56R / N55Q / Y51Q, F56R / N55Q / Y51S, F56R / N55R / Y51L, F56R / N55R / Y51V, F56R / N55R / Y51A, F56R / N55R / Y51N, F56R / N55R / Y51Q, F56R / N55R / Y51S, F56R / N55K / Y51L, F56R / N55K / Y51V, F56R / N55K / Y51A, F56R / N55K / Y51N, F56R / N55K / Y51Q, F56R / N55K / Y51S, F56R / N55S / Y51L, F56R / N55S / Y51V, F56R / N55S / Y51A, F56R / N55S / Y51N, F56R / N55S / Y51Q, F56R / N55S / Y51S, F56R / N55S / Y51G, F56R / N55G / Y51L, F56R / N55G / Y51V, F56R / N55G / Y51A, F56R / N55G / Y51N, F56R / N55G / Y51Q, F56R / N55G / Y51S, F56R / N55G / Y51G, F56R / N55A / Y51L, F56R / N55A / Y51V, F56R / N55A / Y51A, F56R / N55A / Y51N, F56R / N55A / Y51Q, F56R / N55A / Y51S, F56R / N55A / Y51G, F56R / N55T / Y51L, F56R / N55T / Y51V, F56R / N55T / Y51A, F56R / N55T / Y51N, F56R / N55T / Y51Q, F56R / N55T / Y51S, F56R / N55T / Y51G, F56S / N55Q / Y51L, F56S / N55Q / Y51V, F56S / N55Q / Y51A, F56S / N55Q / Y51N, F56S / N55Q / Y51Q, F56S / N55Q / Y51S, F56S / N55Q / Y51G, F56S / N55R / Y51L, F56S / N55R / Y51V, F56S / N55R / Y51A, F56S / N55R / Y51N, F56S / N55R / Y51Q, F56S / N55R / Y51S, F56S / N55R / Y51G, F56S / N55K / Y51L, F56S / N55K / Y51V, F56S / N55K / Y51A, F56S / N55K / Y51N, F56S / N55K / Y51Q, F56S / N55K / Y51S, F56S / N55S / Y51L, F56S / N55S / Y51V, F56S / N55S / Y51A, F56S / N55S / Y51N, F56S / N55S / Y51Q, F56S / N55S / Y51S, F56S / N55S / Y51G, F56S / N55G / Y51L, F56S / N55G / Y51V, F56S / N55G / Y51A, F56S / N55G / Y51N, F56S / N55G / Y51Q, F56S / N55G / Y51S, F56S / N55G / Y51G, F56S / N55A / Y51L, F56S / N55A / Y51V, F56S / N55A / Y51A, F56S / N55A / Y51N, F56S / N55A / Y51Q, F56S / N55A / Y51S, F56S / N55A / Y51G, F56S / N55T / Y51L, F56S / N55T / Y51V, F56S / N55T / Y51A, F56S / N55T / Y51N, F56S / N55T / Y51Q, F56S / N55T / Y51S, F56S / N55T / Y51G,F56S / N55K / Y51S, F56S / N55K / Y51G, F56S / N55S / Y51L, F56S / N55S / Y51V, F56S / N55S / Y51A, F56S / N55S / Y51N, F56S / N55S / Y51Q, F56S / N55S / Y51S, F56S / N55S / Y51G, F56S / N55G / Y51L, F56S / N55G / Y51V, F56S / N55G / Y51A, F56S / N55G / Y51N, F56S / N55G / Y51Q, F56S / N55G / Y51S, F56S / N55G / Y51G, F56S / N55A / Y51L, F56S / N55A / Y51V, F56S / N55A / Y51A, F56S / N55A / Y51N, F56S / N55A / Y51Q, F56S / N55A / Y51S, F56S / N55A / Y51G, F56S / N55T / Y51L, F56S / N55T / Y51V, F56S / N55T / Y51A, F56S / N55T / Y51N, F56S / N55T / Y51Q, F56S / N55T / Y51S, F56S / N55T / Y51G, F56G / N55Q / Y51L, F56G / N55Q / Y51V, F56G / N55Q / Y51A, F56G / N55Q / Y51N, F56G / N55Q / Y51Q, F56G / N55Q / Y51S, F56G / N55Q / Y51G, F56G / N55R / Y51L, F56G / N55R / Y51V, F56G / N55R / Y51A, F56G / N55R / Y51N, F56G / N55R / Y51Q, F56G / N55R / Y51S, F56G / N55R / Y51G, F56G / N55K / Y51L, F56G / N55K / Y51V, F56G / N55K / Y51A, F56G / N55K / Y51N, F56G / N55K / Y51Q, F56G / N55K / Y51S, F56G / N55K / Y51G, F56G / N55S / Y51L, F56G / N55S / Y51V, F56G / N55S / Y51A, F56G / N55S / Y51N, F56G / N55S / Y51Q, F56G / N55S / Y51S, F56G / N55S / Y51G, F56G / N55G / Y51L, F56G / N55G / Y51V, F56G / N55G / Y51A, F56G / N55G / Y51N, F56G / N55G / Y51Q, F56G / N55G / Y51S, F56G / N55G / Y51G, F56G / N55A / Y51L,F56G / N55A / Y51L, F56G / N55A / Y51V, F56G / N55A / Y51A, F56G / N55A / Y51N, F56G / N55A / Y51Q, F56G / N55A / Y51S, F56G / N55A / Y51G, F56G / N55T / Y51L, F56G / N55T / Y51V, F56G / N55T / Y51A, F56G / N55T / Y51N, F56G / N55T / Y51Q, F56G / N55T / Y51S, F56G / N55T / Y51G, F56A / N55Q / Y51L, F56A / N55Q / Y51V, F56A / N55Q / Y51A, F56A / N55Q / Y51N, F56A / N55Q / Y51Q, F56A / N55Q / Y51S, F56A / N55Q / Y51G, F56A / N55R / Y51L, F56A / N55R / Y51V, F56A / N55R / Y51A, F56A / N55R / Y51N, F56A / N55R / Y51Q, F56A / N55R / Y51S, F56A / N55R / Y51G, F56A / N55K / Y51L, F56A / N55K / Y51V, F56A / N55K / Y51A, F56A / N55K / Y51N, F56A / N55K / Y51Q, F56A / N55K / Y51S, F56A / N55K / Y51G, F56A / N55S / Y51L, F56A / N55S / Y51V, F56A / N55S / Y51A, F56A / N55S / Y51N, F56A / N55S / Y51Q, F56A / N55S / Y51S, F56A / N55S / Y51G, F56A / N55G / Y51L, F56A / N55G / Y51V, F56A / N55G / Y51A, F56A / N55G / Y51N, F56A / N55G / Y51Q, F56A / N55G / Y51S, F56A / N55G / Y51G, F56A / N55A / Y51L, F56A / N55A / Y51V, F56A / N55A / Y51A, F56A / N55A / Y51N, F56A / N55A / Y51Q, F56A / N55A / Y51S, F56A / N55A / Y51G, F56A / N55T / Y51L, F56A / N55T / Y51V, F56A / N55T / Y51A, F56A / N55T / Y51N, F56A / N55T / Y51Q, F56A / N55T / Y51S, F56A / N55T / Y51G, F56K / N55Q / Y51L, F56K / N55Q / Y51V, F56K / N55Q / Y51A, F56K / N55Q / Y51N,F56K / N55Q / Y51Q, F56K / N55Q / Y51S, F56K / N55Q / Y51G, F56K / N55R / Y51L, F56K / N55R / Y51V, F56K / N55R / Y51A, F56K / N55R / Y51N, F56K / N55R / Y51Q, F56K / N55R / Y51S, F56K / N55R / Y51G, F56K / N55K / Y51L, F56K / N55K / Y51V, F56K / N55K / Y51A, F56K / N55K / Y51N, F56K / N55K / Y51Q, F56K / N55K / Y51S, F56K / N55K / Y51G, F56K / N55S / Y51L, F56K / N55S / Y51V, F56K / N55S / Y51A, F56K / N55S / Y51N, F56K / N55S / Y51Q, F56K / N55S / Y51S, F56K / N55S / Y51G, F56K / N55G / Y51L, F56K / N55G / Y51V, F56K / N55G / Y51A, F56K / N55G / Y51N, F56K / N55G / Y51Q, F56K / N55G / Y51S, F56K / N55G / Y51G, F56K / N55A / Y51L, F56K / N55A / Y51V, F56K / N55A / Y51A, F56K / N55A / Y51N, F56K / N55A / Y51Q, F56K / N55A / Y51S, F56K / N55A / Y51G, F56K / N55T / Y51L, F56K / N55T / Y51V, F56K / N55T / Y51A, F56K / N55T / Y51N, F56K / N55T / Y51Q, F56K / N55T / Y51S, F56K / N55T / Y51G, F56E / N55R, F56E / N55K, F56D / N55R, F56D / N55K, F56R / N55E, F56R / N55D, F56K / N55E, or F56K / N55D.
[0271] In (ii), the variant preferably comprises Y51R / F56Q, Y51N / F56N, Y51M / F56Q, Y51L / F56Q, Y51I / F56Q, Y51V / F56Q, Y51A / F56Q, Y51P / F56Q, Y51G / F56Q, Y51C / F56Q, Y51Q / F56Q, Y51N / F56Q, Y51S / F56Q, Y51E / F56Q, Y51D / F56Q, Y51K / F56Q, or Y51H / F56Q.
[0272] In (ii), the variant preferably comprises Y51T / F56Q, Y51Q / F56Q or Y51A / F56Q.
[0273] In (ii), the variant preferably comprises Y51T / F56F, Y51T / F56M, Y51T / F56L, Y51T / F56I, Y51T / F56V, Y51T / F56A, Y51T / F56P, Y51T / F56G, Y51T / F56C, Y51T / F56Q, Y51T / F56N, Y51T / F56T, Y51T / F56S, Y51T / F56E, Y51T / F56D, Y51T / F56K, Y51T / F56H or Y51T / F56R.
[0274] In (ii), the variant preferably comprises Y51T / N55Q, Y51T / N55S or Y51T / N55A.
[0275] In (ii), the variant preferably comprises Y51A / F56F, Y51A / F56L, Y51A / F56I, Y51A / F56V, Y51A / F56A, Y51A / F56P, Y51A / F56G, Y51A / F56C, Y51A / F56Q, Y51A / F56N, Y51A / F56T, Y51A / F56S, Y51A / F56E, Y51A / F56D, Y51A / F56K, Y51A / F56H or Y51A / F56R.
[0276] In (ii), the variant preferably comprises Y51C / F56A, Y51E / F56A, Y51D / F56A, Y51K / F56A, Y51H / F56A, Y51Q / F56A, Y51N / F56A, Y51S / F56A, Y51P / F56A or Y51V / F56A.
[0277] In (xi), the variant preferably comprises a deletion of Y51 / P52, Y51 / P52 / A53, P50 to P52, P50 to A53, K49 to Y51, K49 to A53, and a substitution with a single proline (P), a deletion of K49 to S54 and a substitution with a single P, a deletion of Y51 to A53, Y51 to S54, N55 / F56, N55 to S57, N55 / F56 and a substitution with a single P, a deletion of N55 / F56 and a substitution with a single glycine (G), a deletion of N55 / F56 and a substitution with a single alanine (A), a deletion of N55 / F56 and a substitution with a single P and Y51N, a deletion of N55 / F56 and a substitution with a single P and Y51Q, a deletion of N55 / F56 and a substitution with a single P and Y51S, a deletion of N55 / F56 and a substitution with a single G and Y51N, a deletion of N55 / F56 and a substitution with a single G and Y51Q, a deletion of N55 / F56 and a substitution with a single G and Y51S, a deletion of N55 / F56 and a substitution with a single A and Y51N, a deletion of N55 / F56 and a substitution with a single A / Y51Q or a deletion of N55 / F56 and a substitution with a single A and Y51S.
[0278] D149N / E201N / D195N / E203N, D149Q / E201N / D195N / E203N, D149N / E201Q / D195N / E203N, D149N / E201N / D195Q / E203N, D149Q / E201Q / D195N / E203N, D149Q / E201N / D195Q / E203N, D149N / E201Q / D195Q / E203N, D149Q / E201Q / D195Q / E203N, D149N / E203N / D195N / E201Q, D149Q / E203N / D195N / E201Q, D149N / E203Q / D195N / E201Q, D149Q / E203Q / D195N / E201Q, D149N / E203N / D195Q / E201Q, D149Q / E203N / D195Q / E201Q, D149N / E203Q / D195Q / E201Q, D149Q / E203Q / D195Q / E201Q, D149N / E185N / D195N / E201N / E203N, D149Q / E185N / D195N / E201N / E203N, D149N / E185Q / D195N / E201N / E203N, D149N / E185N / D195Q / E201N / E203N, D149Q / E185Q / D195N / E201N / E203N, D149N / E185N / D195N / E201Q / E203N, D149Q / E185N / D195N / E201Q / E203N, D149N / E185Q / D195N / E201Q / E203N, D149Q / E185Q / D195N / E201Q / E203N, D149N / E185N / D195Q / E201Q / E203N, D149Q / E185Q / D195Q / E201Q / E203N, D149N / E185N / D195N / E201N / E203Q, D149Q / E185N / D195N / E201N / E203Q, D149N / E185Q / D195N / E201N / E203Q, D149N / E185N / D195Q / E201N / E203Q, D149Q / E185Q / D195Q / E201N / E203Q, D149N / E185N / D195N / E201Q / E203Q, D149Q / E185N / D195N / E201Q / E203Q, D149N / E185Q / D195N / E201Q / E203Q, D149Q / E185Q / D195N / E201Q / E203Q, D149N / E185N / D195Q / E201Q / E203Q, D149Q / E185Q / D195Q / E201Q / E203Q, D149N / E185N / D195N / E201N / E203N / E201Q, D149Q / E185N / D195N / E201N / E203N / E201Q, D149N / E185Q / D195N / E201N / E203N / E201Q, D149N / E185N / D195Q / E201N / E203N / E201Q, D149Q / E185Q / D195N / E201N / E203N / E201Q, D149N / E185N / D195N / E201Q / E203N / E201Q, D149Q / E185N / D195N / E201Q / E203N / E201Q, D149N / E185Q / D195N / E201Q / E203N / E201Q, D149N / E185N / D195Q / E201Q / E203N / E201Q, D149Q / E185Q / D195Q / E201Q / E203N / E201Q, D149N / E185N / D195N / E201N / E203N / E201Q / E203Q, D149Q / E185N / D195N / E201N / E203N / E201Q / E203Q, D149N / E185Q / D195N / E201N / E203N / E201Q / E203Q, D149N / E185N / D195Q / E201N / E203N / E201Q / E203Q, D149Q / E185Q / D195N / E201N / E203N / E201Q / E203Q, D149N / E185N / D195N / E201Q / E203N / E201Q / E203Q, D149Q / E185N / D195N / E201Q / E203N / E201Q / E203Q, D149N / E185Q / D195N / E201Q / E203N / E201Q / E203Q, D149Q / E185Q / D195N / E201Q / E203N / E201Q / E203Q, D149N / E185N / D195Q / E201Q / E203N / E201Q / E203Q, D149Q / E185Q / D195Q / E201Q / E203N / E201Q / E203Q, D149N / E185N / D195N / E201N / E203N / E201Q / E203Q / D149Q, D149N / E185Q / D195N / E201N / E203N / E201Q / E203Q / D149Q, D149N / E185N / D195Q / E201N / E203N / E201Q / E203Q / D149Q, D149Q / E185Q / D195N / E201N / E203N / E201Q / E203Q / D149Q, D149N / E185N / D195N / E201Q / E203N / E201Q / E203Q / D149Q, D149Q / E185N / D195N / E201Q / E203N / E201Q / E203Q / D149Q, D149N / E185Q / D195N / E201Q / E203N / E201Q / E203Q / D149Q, D149Q / E185Q / D195N / E201Q / E203N / E201Q / E203Q / D149Q, D149N / E185N / D195Q / E201Q / E203N / E201Q / E203Q / D149Q, D149Q / E185Q / D195Q / E201Q / E203N / E201Q / E203Q / D149Q, D149N / E185N / D195N / E201N / E203N / E201Q / E203Q / D149Q / E201Q, D149N / E185Q / D195N / E201N / E203N / E201Q / E203Q / D149Q / E201Q, D149N / E185N / D195Q / E201N / E203N / E201Q / E203Q / D149Q / E201Q, D149Q / E185Q / D195N / E201N / E203N / E201Q / E203Q / D149Q / E201Q, D149N / E185N / D195N / E201Q / E203N / E201Q / E203Q / D149Q / E201Q, D149Q / E185N / D195N / E201Q / E203N / E201Q / E203Q / D149Q / E201Q, D149N / E185Q / D195N / E201Q / E203N / E201Q / E203Q / D149Q / E201Q, D149Q / E185Q / D195N / E201Q / E203N / E201Q / E203Q / D149Q / E201Q, D149N / E185N / D195Q / E201Q / E203N / E201Q / E203Q / D149Q / E201Q, D149Q / E185Q / D195Q / E201Q / E203N / E201Q / E203Q / D149Q / E201Q, D149N / E185N / D195N / E201N / E203N / E201Q / E203Q / D149Q / E201Q / E203Q, D149N / E185Q / D195N / E201N / E203N / E201Q / E203Q / D149Q / E201Q / E203Q, D149N / E185N / D195Q / E201N / E203N / E201Q / E203Q / D149Q / E201Q / E203Q, D149Q / E185Q / D195N / E201N / E203N / E201Q / E203Q / D149Q / E201Q / E203Q, D149N / E185N / D195N / E201Q / E203N / E201Q / E203Q / D149Q / E201Q / E203Q, D149Q / E185N / D195N / E201Q / E203N / E201Q / E203Q / D149Q / E201Q / E203Q, D149N / E185Q / D195N / E201Q / E203N / E201Q / E203Q / D149Q / E201Q / E203Q, D149Q / E185Q / D195N / E201Q / E203N / E201Q / E203Q / D149Q / E201Q / E203Q, D149N / E185N / D195Q / E201Q / E203N / E201Q / E203Q / D149Q / E201Q / E203Q, D149Q / E185Q / D195Q / E201Q / E203N / E201Q / E203Q / D149Q / E201Q / E203Q, D149N / E185N / D195N / E201N / E203N / E201Q / E203Q / D149Q / E201Q / E203Q / D149Q, D149N / E185Q / D195N / E201N / E203N / E201Q / E203Q / D149Q / E201Q / E203Q / D149Q, D149N / E185N / D195Q / E201N / E203N / E201Q / E203Q / D149Q / E201Q / E203Q / D149Q, D149Q / E185Q / D195N / E201N / E203N / E201Q / E203Q / D149Q / E201Q / E203Q / D149Q, D149N / E185N / D195N / E201Q / E203N / E201Q / E203Q / D149Q / E201Q / E203Q / D149Q, D149Q / E185N / D195N / E201Q / E203N / E201Q / E203Q / D149Q / E201Q / E203Q / D149Q, D149N / E185Q / D195N / E201QD149Q / E185N / E201N / E203Q, D149N / E185Q / E201Q / E203N, D149N / E185Q / E201N / E203Q, D149N / E185N / E201Q / E203Q, D149Q / E185Q / E201Q / E203Q, D149Q / E185Q / E201N / E203Q, D149Q / E185N / E201Q / E203Q, D149N / E185Q / E201Q / E203Q, D149Q / E185Q / E201Q / E203N, D149N / E185N / D195N / E201N / E203N, D149Q / E185N / D195N / E201N / E203N, D149N / E185Q / D195N / E201N / E203N, D149N / E185N / D195Q / E201N / E203N, D149N / E185N / D195N / E201Q / E203N, D149N / E185N / D195N / E201N / E203Q, D149Q / E185Q / D195N / E201N / E203N, D149Q / E185N / D195Q / E201N / E203N, D149Q / E185N / D195N / E201Q / E203N, D149Q / E185N / D195N / E201N / E203Q, D149N / E185Q / D195Q / E201N / E203N, D149N / E185Q / D195N / E201Q / E203N, D149N / E185Q / D195N / E201N / E203Q, D149N / E185N / D195Q / E201Q / E203N, D149N / E185N / D195Q / E201N / E203Q, D149N / E185N / D195N / E201Q / E203Q, D149Q / E185Q / D195Q / E201N / E203N, D149Q / E185Q / D195N / E201Q / E203N, D149Q / E185Q / D195N / E201N / E203Q, D149Q / E185N / D195Q / E201Q / E203N, D149Q / E185N / D195Q / E201N / E203Q, D149Q / E185N / D195N / E201Q / E203Q, D149N / E185Q / D195Q / E201Q / E203N, D149N / E185Q / D195Q / E201N / E203Q, D149N / E185Q / D195N / E201Q / E203Q, D149N / E185N / D195Q / E201Q / E203Q,D149Q / E185Q / D195Q / E201Q / E203N, D149Q / E185Q / D195Q / E201N / E203Q, D149Q / E185Q / D195N / E201Q / E203Q, D149Q / E185N / D195Q / E201Q / E203Q, D149N / E185Q / D195Q / E201Q / E203Q, D149Q / E185Q / D195Q / E201Q / E203Q, D149N / E185R / E201N / E203N, D149Q / E185R / E201N / E203N, D149N / E185R / E201Q / E203N, D149N / E185R / E201N / E203Q, D149Q / E185R / E201Q / E203N, D149Q / E185R / E201N / E203Q, D149N / E185R / E201Q / E203Q, D149Q / E185R / E201Q / E203Q, D149R / E185N / E201N / E203N, D149R / E185Q / E201N / E203N, D149R / E185N / E201Q / E203N, D149R / E185N / E201N / E203Q, D149R / E185Q / E201Q / E203N, D149R / E185Q / E201N / E203Q, D149R / E185N / E201Q / E203Q, D149R / E185Q / E201Q / E203Q, D149R / E185N / D195N / E201N / E203N, D149R / E185Q / D195N / E201N / E203N, D149R / E185N / D195Q / E201N / E203N, D149R / E185N / D195N / E201Q / E203N, D149R / E185Q / D195N / E201N / E203Q, D149R / E185Q / D195Q / E201N / E203N, D149R / E185Q / D195N / E201Q / E203N, D149R / E185Q / D195N / E201N / E203Q, D149R / E185N / D195Q / E201Q / E203N, D149R / E185N / D195Q / E201N / E203Q, D149R / E185N / D195N / E201Q / E203Q, D149R / E185Q / D195Q / E201Q / E203N, D149R / E185Q / D195Q / E201N / E203Q, D149R / E185Q / D195N / E201Q / E203Q,D149R / E185N / D195Q / E201Q / E203Q, D149R / E185Q / D195Q / E201Q / E203Q, D149N / E185R / D195N / E201N / E203N, D149Q / E185R / D195N / E201N / E203N, D149N / E185R / D195Q / E201N / E203N, D149N / E185R / D195N / E201Q / E203N, D149N / E185R / D195N / E201N / E203Q, D149Q / E185R / D195Q / E201N / E203N, D149Q / E185R / D195N / E201Q / E203N, D149Q / E185R / D195N / E201N / E203Q, D149N / E185R / D195Q / E201Q / E203N, D149N / E185R / D195Q / E201N / E203Q, D149N / E185R / D195N / E201Q / E203Q, D149Q / E185R / D195Q / E201Q / E203N, D149Q / E185R / D195Q / E201N / E203Q, D149Q / E185R / D195N / E201Q / E203Q, D149N / E185R / D195Q / E201Q / E203Q, D149Q / E185R / D195Q / E201Q / E203Q, D149N / E185R / D195N / E201R / E203N, D149Q / E185R / D195N / E201R / E203N, D149N / E185R / D195Q / E201R / E203N, D149N / E185R / D195N / E201R / E203Q, D149Q / E185R / D195Q / E201R / E203N, D149Q / E185R / D195N / E201R / E203Q, D149N / E185R / D195Q / E201R / E203Q, D149Q / E185R / D195Q / E201R / E203Q, E131D / K49R, E101N / N102F, E101N / N102Y, E101N / N102W, E101F / N102F, E101F / N102Y, E101F / N102W, E101Y / N102F, E101Y / N102Y, E101Y / N102W, E101W / N102F, E101W / N102Y, E101W / N102W, E101N / N102R, E101F / N102R, E101Y / N102R, or E101W / N102F.
[0279] Preferred variants of the application that form a pore in which fewer nucleotides contribute to the current as a polynucleotide moves through the pore comprise Y51A / F56A, Y51A / F56N, Y51I / F56A, Y51L / F56A, Y51T / F56A, Y51I / F56N, Y51L / F56N or Y51T / F56N or more preferably Y51I / F56A, Y51L / F56A or Y51T / F56A. As discussed above, this makes it easier to identify a direct relationship between the observed current (as a polynucleotide moves through the pore) and the polynucleotide.
[0280] Preferred variants that form a pore that shows an increased display range comprise mutations in the following positions:
[0281] Y51, F56, D149, E185, E201 and E203;
[0282] N55 and F56;
[0283] Y51 and F56;
[0284] Y51, N55 and F56; or
[0285] F56 and N102.
[0286] Preferred variants that form a pore that shows an increased display range comprise:
[0287] Y51N, F56A, D149N, E185R, E201N and E203N;
[0288] N55S and F56Q;
[0289] Y51A and F56A;
[0290] Y51A and F56N;
[0291] Y51I and F56A;
[0292] Y51L and F56A;
[0293] Y51T and F56A;
[0294] Y51I and F56N;
[0295] Y51L and F56N;
[0296] Y51T and F56N;
[0297] Y51T and F56Q;
[0298] Y51A, N55S and F56A;
[0299] Y51A, N55S and F56N;
[0300] Y51T, N55S and F56Q; or
[0301] F56Q and N102R.
[0302] Preferred variants that form a pore in which fewer nucleotides contribute to the current as a polynucleotide moves through the pore comprise mutations at the following positions:
[0303] N55 and F56, such as N55X and F56Q, where X is any amino acid; or
[0304] Y51 and F56, such as Y51X and F56Q, where X is any amino acid.
[0305] Particularly preferred variants comprise Y51A and F56Q.
[0306] Preferred variants that form a pore that shows increased throughput comprise mutations at the following positions:
[0307] D149, E185 and E203;
[0308] D149, E185, E201 and E203; or
[0309] D149, E185, D195, E201 and E203.
[0310] Preferred variants that form a pore that shows increased throughput comprise:
[0311] D149N, E185N and E203N;
[0312] D149N, E185N, E201N and E203N;
[0313] D149N, E185R, D195N, E201N and E203N; or
[0314] D149N, E185R, D195N, E201R and E203N.
[0315] Preferred variants that form a pore in which capture of a polynucleotide is increased comprise mutations at the following positions:
[0316] D43N / Y51T / F56Q;
[0317] E44N / Y51T / F56Q;
[0318] D43N / E44N / Y51T / F56Q;
[0319] Y51T / F56Q / Q62R;
[0320] D43N / Y51T / F56Q / Q62R;
[0321] E44N / Y51T / F56Q / Q62R; or
[0322] D43N / E44N / Y51T / F56Q / Q62R.
[0323] Preferably the variant comprises the following mutations:
[0324] D149R / E185R / E201R / E203R or Y51T / F56Q / D149R / E185R / E201R / E203R;
[0325] D149N / E185N / E201N / E203N or Y51T / F56Q / D149N / E185N / E201N / E203N;
[0326] E201R / E203R or Y51T / F56Q / E201R / E203R
[0327] E201N / E203R or Y51T / F56Q / E201N / E203R;
[0328] E203R or Y51T / F56Q / E203R;
[0329] E203N or Y51T / F56Q / E203N;
[0330] E201R or Y51T / F56Q / E201R;
[0331] E201N or Y51T / F56Q / E201N;
[0332] E185R or Y51T / F56Q / E185R;
[0333] E185N or Y51T / F56Q / E185N;
[0334] D149R or Y51T / F56Q / D149R;
[0335] D149N or Y51T / F56Q / D149N;
[0336] R142E or Y51T / F56Q / R142E;
[0337] R142N or Y51T / F56Q / R142N;
[0338] R192E or Y51T / F56Q / R192E; or
[0339] R192N or Y51T / F56Q / R192N.
[0340] Preferably the variant comprises the following mutations:
[0341] Y51A / F56Q / E101N / N102R;
[0342] Y51A / F56Q / R97N / N102G;
[0343] Y51A / F56Q / R97N / N102R;
[0344] Y51A / F56Q / R97N;
[0345] Y51A / F56Q / R97G;
[0346] Y51A / F56Q / R97L;
[0347] Y51A / F56Q / N102R;
[0348] Y51A / F56Q / N102F;
[0349] Y51A / F56Q / N102G;
[0350] Y51A / F56Q / E101R;
[0351] Y51A / F56Q / E101F;
[0352] Y51A / F56Q / E101N; or
[0353] Y51A / F56Q / E101G
[0354] Preferably the variant further comprises a mutation at T150. Preferred variants which form a pore showing increased insertion comprise T150I. Mutations at T150, such as T150I, can be combined with any of the mutations or combinations of mutations discussed above.
[0355] Preferred variants of SEQ ID NO: 3 comprise (a) R97W and (b) a mutation at Y51 and / or F56. Preferred variants of SEQ ID NO: 3 include (a) R97W and (b) Y51R / H / K / D / E / S / T / N / Q / C / G / P / A / V / I / L / M and / or F56R / H / K / D / E / S / T / N / Q / C / G / P / A / V / I / L / M. Preferred variants of SEQ ID NO: 3 comprise (a) R97W and (b) Y51L / V / A / N / Q / S / G and / or F56A / Q / N. Preferred variants of SEQ ID NO: 3 comprise (a) R97W and (b) Y51A and / or F56Q. Preferred variants of SEQ ID NO: 3 include R97W, Y51A, and F56Q.
[0356] A variant of SEQ ID NO: 3 preferably comprises a mutation at R192. The variant preferably comprises R192D / Q / F / S / T / N / E, R192D / Q / F / S / T, or R192D / Q. A preferred variant of SEQ ID NO: 3 comprises (a) R97W, (b) a mutation at Y51 and / or F56, and (c) a mutation at R192, such as R192D / Q / F / S / T / N / E, R192D / Q / F / S / T, or R192D / Q. A preferred variant of SEQ ID NO: 3 comprises (a) R97W, (b) Y51 R / H / K / D / E / S / T / N / Q / C / G / P / A / V / I / L / M and / or F56 R / H / K / D / E / S / T / N / Q / C / G / P / A / V / I / L / M, and (c) a mutation at R192, such as R192D / Q / F / S / T / N / E, R192D / Q / F / S / T, or R192D / Q. A preferred variant of SEQ ID NO: 3 comprises (a) R97W, (b) Y51 L / V / A / N / Q / S / G and / or F56 A / Q / N, and (c) a mutation at R192, such as R192D / Q / F / S / T / N / E, R192D / Q / F / S / T, or R192D / Q. A preferred variant of SEQ ID NO: 3 comprises (a) R97W, (b) Y51 A and / or F56 Q, and (c) a mutation at R192, such as R192D / Q / F / S / T / N / E, R192D / Q / F / S / T, or R192D / Q. A preferred variant of SEQ ID NO: 3 comprises R97W, Y51 A, F56 Q, and R192D / Q / F / S / T, or R192D / Q. A preferred variant of SEQ ID NO: 3 comprises R97W, Y51 A, F56 Q, and R192D. A preferred variant of SEQ ID NO: 3 comprises R97W, Y51 A, F56 Q, and R192Q. In the above paragraphs where different amino acids at a particular position are separated by a / symbol, the / symbol means "or". For example, R192D / Q means R192D or R192Q.
[0357] Any of the above preferred variants of SEQ ID NO: 3 described above can further comprise a mutation at R93. A preferred variant of SEQ ID NO: 3 comprises (a) R93W, and (b) a mutation at Y51 and / or F56, preferably Y51 A and F56 Q.
[0358] Any of the above preferred variants of SEQ ID NO: 3 described above can comprise a K94N / Q mutation. Any of the above preferred variants of SEQ ID NO: 3 described above can comprise a F191T mutation.
[0359] The CsgG monomer can be modified to facilitate attachment to the CsgF peptide. For example, a cysteine residue can be introduced at one or more positions corresponding to positions 132, 133, 136, 138, 140, 142, 144, 145, 147, 149, 151, 153, 155, 183, 185, 187, 189, 191, 201, 203, 205, 207, and 209 of SEQ ID NO: 3, and / or at any one of the positions identified in Table 4 as predicted to contact CsgF, to facilitate covalent attachment to CsgG. Alternatively or in addition to covalent attachment via a cysteine residue, the pore can be stabilized by hydrophobic or electrostatic interactions. To facilitate such interactions, a non-native reactive or photo-reactive amino acid can be introduced at a position corresponding to one or more of positions 132, 133, 136, 138, 140, 142, 144, 145, 147, 149, 151, 153, 155, 183, 185, 187, 189, 191, 201, 203, 205, 207, and 209 of SEQ ID NO: 3, and / or at any one of the positions identified in Table 4 as predicted to contact CsgF.
[0360] Preferred exemplary pores include at least one CsgG monomer having the following mutations relative to SEQ ID NO: 3: Y51X1 / N55X2 / F56X3 / N91R / K94Q / R97W / R192D-del(V105-I107), wherein X1 is I / V / S / T, X2 is N / I / V / S / T and / or X3 is Q / I / V / S / T.
[0361] Methods of introducing or substituting naturally occurring amino acids are well known in the art. For example, methionine (M) can be replaced with arginine (R) by replacing the codon for methionine (ATG) with the codon for arginine (CGT) at the relevant position in a polynucleotide encoding the mutant monomer. The polynucleotide can then be expressed as discussed below.
[0362] Dual pore
[0363] The CsgG / CsgF pore can be a dual pore comprising a first pore and a second pore. At least the first pore is a CsgG / CsgF pore disclosed herein. The second pore can be a CsgG pore or a CsgG / CsgF pore. In one embodiment, both the first pore and the second pore are CsgG / CsgF pores disclosed herein. The first pore and the second pore can be the same or different. In addition to any mutations disclosed herein, in a dual pore, the CsgG monomer can further comprise one or more of the additional mutations described below.
[0364] In the dual pore, the first pore can be attached to the second CsgG pore by hydrophobic interactions and / or by one or more disulfide bonds. One or more, such as 2, 3, 4, 5, 6, 8, 9, for example all of the monomers in the first pore and / or in the second pore can be modified to enhance such interactions. This can be achieved in any suitable way.
[0365] At least one cysteine residue in the amino acid sequence of the first pore at the interface between the first pore and the second pore can be disulfide-bonded to at least one cysteine residue in the amino acid sequence of the second pore at the interface between the first pore and the second pore. The cysteine residue in the first pore and / or the cysteine residue in the second pore can be a cysteine residue that is not present in a wild-type CsgG monomer. Between the two pores in the dual pore, a plurality of disulfide bonds can be formed, such as 2, 3, 4, 5, 6, 7, 8, or 9 to 16, 18, 24, 27, 32, 36, 40, 45, 48, 54, 56, or 63. One or both of the first pore or the second pore can comprise at least one monomer, such as up to 8, 9, or 10 monomers, that comprises a cysteine residue at the interface between the first and second pore at a position corresponding to R97, I107, R110, Q100, E101, N102, and / or L113 of SEQ ID NO: 3.
[0366] At least one monomer in the first pore and / or at least one monomer in the second pore can comprise at least one residue at the interface between the first and second pore that is more hydrophobic than the residue present at the corresponding position in a wild-type CsgG monomer. For example, 2 to 10, for example 3, 4, 5, 6, 7, 8, or 9 residues in the first pore and / or the second pore can be more hydrophobic than the residue at the same position in a corresponding wild-type CsgG monomer. Such hydrophobic residues enhance the interaction between the two pores in the dual pore. The at least one residue at the interface between the first and second pore can be at a position corresponding to R97, I107, R110, Q100, E101, N102, and / or L113 of SEQ ID NO: 3. Where the residue at the interface in a wild-type CsgG monomer is R, Q, N, or E, the hydrophobic residue is typically I, L, V, M, F, W, or Y. Where the residue at the interface in a wild-type CsgG monomer is I, the hydrophobic residue is typically L, V, M, F, W, or Y. Where the residue at the interface in a wild-type CsgG monomer is L, the hydrophobic residue is typically I, V, M, F, W, or Y.
[0367] The two-pore can comprise one or more monomers comprising one or more cysteine residues at the interface between the pores and one or more monomers comprising one or more introduced hydrophobic residues at the interface between the pores, or can comprise one or more monomers comprising both such cysteine residues and such hydrophobic residues. For example, one or more, such as any 2, 3, or 4, of the positions in the monomer corresponding to R97, I107, R110, Q100, E101, N102, and / or L113 of SEQ ID NO: 3 can comprise a cysteine (C) residue and one or more, such as any 2, 3, or 4, of the positions in the monomer corresponding to R97, I107, R110, Q100, E101, N102, and / or L113 of SEQ ID NO: 3 can comprise a hydrophobic residue, such as I, L, V, M, F, W, or Y.
[0368] The two-pore can comprise a bulky residue at one or more, such as 2, 3, 4, 5, 6, or 7, positions in the tail region, typically at the interface between the first pore and the second pore and larger than the residue present at the corresponding position in a wild-type CsgG monomer. The bulkiness of these residues prevents the formation of a hole in the pore wall at the interface between the first pore and the second pore in the two-pore. The at least one bulky residue at the interface between the first pore and the second pore is typically at a position corresponding to A98, A99, T104, V105, L113, Q114, or S115 of SEQ ID NO: 3. Where the residue at the interface in a wild-type CsgG monomer is A, the bulky residue is typically I, L, V, M, F, W, Y, N, Q, S, or T. Where the residue present at the interface in a wild-type CsgG monomer is T, the bulky residue is typically L, M, F, W, Y, N, Q, R, D, or E. Where the residue present at the interface in a wild-type CsgG monomer is V, the bulky residue is typically I, L, M, F, W, Y, N, Q. Where the residue present at the interface in a wild-type CsgG monomer is L, the bulky residue is typically M, F, W, Y, N, Q, R, D, or E. Where the residue present at the interface in a wild-type CsgG monomer is Q, the bulky residue is typically F, W, or Y. Where the residue present at the interface in a wild-type CsgG monomer is S, the bulky residue is typically M, F, W, Y, N, Q, E, or R.
[0369] In particular where the second pore is located outside the membrane, the second pore, and optionally the first pore, preferably comprises in the barrel region of the pore residues that reduce the negative charge inside the barrel compared to the charges in the barrel of the wild-type CsgG pore. These mutations make the barrel of the gun more hydrophilic. At least one monomer in the first pore and / or at least one monomer in the second pore of the dual pore can comprise in the barrel region of the pore at least one residue that has less negative charge than the residue present at the corresponding position in the wild-type CsgG monomer. The charge inside the barrel is sufficiently neutral or positive such that negatively charged analytes, such as polynucleotides, are not repelled from entering the pore by electrostatic charge. At least one residue, such as 2, 3, 4, or 5 residues, at positions corresponding to D149, E185, D195, E210, and / or E203 of SEQ ID NO: 3 in the barrel region of the pore can be a neutral or positively charged amino acid. At least one residue, such as 2, 3, 4, or 5 residues, at positions corresponding to D149, E185, D195, E210, and / or E203 of SEQ ID NO: 3 in the barrel region of the pore is preferably N, Q, R, or K.
[0370] Particular examples of charge-removing mutations in SEQ ID NO: 3 include the following: E185N / E203N, D149N / E185R / D195N / E201R / E203N, D149N / E185R / D195N / E201N / E203N, D149R / E185N / D195N / E201N / E203N, D149R / E185N / E201N / E203N, D149N / E185N / D195 / E201N / E203N, D149N / E185N / E201N / E203N, D149N / E185N / E203N, D149N / E185N / E201N, D149N / E203N, D149N / E201N / D195N, D149N / E201N, D195N / E201N / E203N, E201N / E203N, D195N / E203, E203R, E203N, E201R, E201N, D195R, D195N, E185R, E185N, D149R, and D149N.
[0371] At least one CsgG monomer in the first pore can comprise at least one residue in the constriction of the barrel region of the first pore that decreases, maintains or increases the length of the constriction compared to a wild-type CsgG pore, and / or at least one monomer in the second CsgG pore can comprise at least one residue in the constriction of the barrel region of the second pore that decreases, maintains or increases the length of the constriction compared to a wild-type CsgG pore. Preferably, the length of the constriction in the first pore and / or the length of the constriction in the second pore is at least as long as in a wild-type pore, more preferably longer.
[0372] The length of the pore can be increased by inserting residues into a region corresponding to the region between positions K49 and F56 of SEQ ID NO: 3. Between one and five, such as two, three or four amino acid residues can be inserted at any one or more of the following positions defined with reference to SEQ ID NO: 3: K49 and P50, P50 and Y51, Y51 and P52, P52 and A53, A53 and S54, S54 and N55 and / or N55 and F56. Preferably a total of one to ten, such as two to eight, or three to five amino acid residues are inserted into the monomer sequence. Preferably, all monomers in the first pore and / or all monomers in the second pore have the same number of insertions in this region. The inserted residues can increase the length of the loop between residues corresponding to Y51 and N55 of SEQ ID NO: 3. The inserted residues can be any combination of A, S, G or T to maintain flexibility; can be P to add a kink into the loop; and / or can be S, T, N, Q, M, F, W, Y, V and / or I to contribute to the signal generated when an analyte interacts with the barrel of the pore under an applied potential difference. The inserted amino acids can be any combination of S, G, SG, SGG, SGS, GS, GSS and / or GSG.
[0373] In a dual pore, the constriction in the barrel of the first pore and / or the second pore can comprise at least one residue, such as two, three, four or five residues, which when used to detect or characterise an analyte, affects the properties of the pore compared to using a first pore or a second pore having a wild-type constriction, wherein the at least one residue in the constriction of the barrel region of the pore is at a position corresponding to Y51, N55, Y51, P52 and / or A53 of SEQ ID NO: 3. The at least one residue can be Q or V at a position corresponding to F56 of SEQ ID NO: 3; A or Q at a position corresponding to Y51 of SEQ ID NO: 3; and / or V at a position corresponding to N55 of SEQ ID NO: 3.
[0374] The dual pore can comprise at least one monomer in the first CsgG pore and / or at least one monomer in the second CsgG pore, which monomer comprises two or more mutations as defined above.
[0375] The CsgG monomer in the two-pore can comprise a cysteine residue at a position corresponding to R97, I107, R110, Q100, E101, N102 and L113 of SEQ ID NO: 3.
[0376] The CsgG monomer in the two-pore can comprise at a position corresponding to any one or more of R97, Q100, I107, R110, E101, N102 and L113 of SEQ ID NO: 3, a residue that is more hydrophobic than the residue present at the corresponding position of SEQ ID NO: 3, such as the corresponding position of any one of SEQ ID NOs: 68 to 88, wherein the residue at the position corresponding to R97 and / or I107 is M, the residue at the position corresponding to R110 is I, L, V, M, W or Y, and / or the residue at the position corresponding to E101 or N102 is V or M. The residue at the position corresponding to Q100 is typically I, L, V, M, F, W or Y; and / or the residue at the position corresponding to L113 is typically I, V, M, F, W or Y.
[0377] The particular monomer can have the sequence set forth in SEQ ID NO 3 comprising Y51A, F56Q substitutions and R97I / V / L / M / F / W / Y, I107L / V / M / F / W / Y, R110I / V / L / M / F / W / Y, Q100I / V / L / M / F / W / Y, E101I / V / L / M / F / W / Y, N102I / V / L / M / F / W / Y in combination, L113CI / V / L / M / F / W / Y in combination, R97I / V / L / M / F / W / Y and N102I / V / L / M / F / W / Y in combination and / or R97I / V / L / M / F / W / Y and E101I / V / L / M / F / W / Y in combination. I107 can have formed a hydrophobic interaction between the two pores.
[0378] The CsgG monomer in at least one of the two pores can comprise, in the barrel region of the pore, at a position corresponding to any one or more of D149, E185, D195, E210, and E203, a residue that is less negatively charged than the residue present at the corresponding position of SEQ ID NO: 3, such as the corresponding position of any one of SEQ ID NOs: 68-88, wherein the residue at the position corresponding to D149, E185, D195, and / or E203 is K.
[0379] The particular monomer can have the sequence set forth in SEQ ID NO 3 comprising Y51A, F56Q substitution and 1, 2, 3, 4, 5, 6, or all of the following substitutions: A98I / L / V / M / F / W / Y / N / Q / S / T; A99I / L / V / M / F / W / Y / N / Q / S / T; T104N / Q / L / R / D / E / M / F / W / Y; V105I / L / M / F / W / Y / N / Q; L113M / F / W / Y / N / Q / D / E / L / R; Q114Y / F / W; and S115N / Q / M / F / W / Y / E / R.
[0380] The CsgG monomer in at least one of the two pores can comprise, in the barrel region of the pore, at a position corresponding to any one or more of D149, E185, D195, E210, and E203, a residue that is less negatively charged than the residue present at the corresponding position of SEQ ID NO: 3, such as the corresponding position of any one of SEQ ID NOs: 68-88, wherein the residue at the position corresponding to D149, E185, D195, and / or E203 is K.
[0381] The CsgG monomer in at least one of the two pores can comprise, in the constriction of the barrel region of the pore, at least one residue that increases the length of the constriction as compared to a wild-type CsgG pore. The at least one residue is in addition to the residues present in the constriction of a wild-type CsgG pore.
[0382] The length of the pore can be increased by inserting residues into the region corresponding to the region between positions K49 and F56 of SEQ ID NO: 3. One to five, such as two, three or four, amino acid residues can be inserted at any one or more of the following positions defined with reference to SEQ ID NO: 3: K49 and P50, P50 and Y51, Y51 and P52, P52 and A53, A53 and S54, S54 and N55 and / or N55 and F56. Preferably one to ten, such as two to eight, or three to five, amino acid residues in total are inserted into the monomer sequence. The inserted residues can increase the length of the loop between the residues corresponding to Y51 and N55 of SEQ ID NO: 3. The inserted residues can be any combination of A, S, G or T to maintain flexibility; P to add a kink into the loop; and / or S, T, N, Q, M, F, W, Y, V and / or I to contribute to the signal generated when an analyte interacts with the barrel of the pore under an applied potential difference. The inserted amino acids can be any combination of S, G, SG, SGG, SGS, GS, GSS and / or GSG.
[0383] The CsgG monomer in at least one of the pores can comprise at least one residue in the constriction of the barrel region of the pore at a position corresponding to N55, P52 and / or A53 of SEQ ID NO: 3 which is different from the residue present in the corresponding wild-type monomer, wherein the residue at the position corresponding to N55 is V.
[0384] Any two or more of the above-described residues can be present in the same monomer.
[0385] In particular, the monomer can comprise at least one said cysteine residue, at least one said hydrophobic residue, at least one said bulky residue, at least one said neutral or positively charged residue and / or at least one said residue which increases the length of the constriction.
[0386] The CsgG monomer in the dual pore can additionally comprise one or more, such as two, three, four or five, residues which, when used to detect or characterise an analyte, affect the properties of the pore compared to the use of the first or second pore having a wild-type constriction, wherein the at least one residue is in the constriction of the barrel region of the pore at a position corresponding to Y51, N55, Y51, P52 and / or A53 of SEQ ID NO: 3. The at least one residue can be Q or V at a position corresponding to F56 of SEQ ID NO: 3; A or Q at a position corresponding to Y51 of SEQ ID NO: 3; and / or V at a position corresponding to N55 of SEQ ID NO: 3.
[0387] Method of preparing a modified protein
[0388] Methods of introducing or substituting non-naturally occurring amino acids are also well known in the art. For example, non-naturally occurring amino acids can be introduced by including synthetic aminoacyl-tRNAs in the IVTT system used to express the mutant monomer. Alternatively, non-naturally occurring amino acids can be introduced by expressing the mutant monomer in E. coli, which is auxotrophic for the particular amino acid in the presence of synthetic (i.e., non-naturally occurring) analogs of those particular amino acids. If the mutant monomer is produced using a partial peptide synthesis method, they can also be produced by naked ligation.
[0389] The CsgG-derived monomer can be modified to aid in its identification or purification, for example by the addition of a streptavidin tag or by the addition of a signal sequence to facilitate its secretion from a cell in which the monomer does not naturally contain such a sequence. Other suitable tags are discussed in more detail below. The monomer can be labeled with an epitope tag. An epitope tag can be any suitable tag that allows the monomer to be detected. Suitable tags are described below.
[0390] The CsgG-derived monomer can also be produced using D-amino acids. For example, the CsgG-derived monomer can comprise a mixture of L-amino acids and D-amino acids. This is routine in the art of producing such proteins or peptides.
[0391] The CsgG-derived monomer contains one or more specific modifications to facilitate discrimination of nucleotides. The CsgG-derived monomer can also contain other non-specific modifications, so long as they do not interfere with pore formation. Many non-specific side chain modifications are known in the art and can be made to the side chains of the CsgG-derived monomer. Such modifications include, for example, reductive alkylation of amino acids by reaction with an aldehyde followed by reduction with NaBH4, amidation with imidoacetic acid methyl ester or acylation with acetic anhydride.
[0392] The CsgG-derived monomer can be produced using standard methods known in the art. The CsgG-derived monomer can be prepared synthetically or by recombinant means. For example, the monomer can be synthesized by in vitro translation and transcription (IVTT). Suitable methods of producing pores and monomers are discussed in International Applications WO 2010 / 004273, WO 2010 / 004265 or WO 2010 / 086603. Methods for inserting pores into membranes are known.
[0393] Two or more CsgG monomers in the pore can be covalently attached to each other. For example, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9 or at least 10 monomers can be covalently attached. The covalently attached monomers can be the same or different.
[0394] The monomers can optionally be genetically fused by a linker, or chemically fused, for example by a chemical cross-linker. Methods for covalently attaching monomers are disclosed in WO2017 / 149316, WO2017 / 149317 and WO2017 / 149318.
[0395] In some embodiments, the mutant monomers are chemically modified. The mutant monomers can be chemically modified in any way at any site. The mutant monomers are preferably chemically modified by attachment of a molecule to one or more cysteines (cysteine ligation), attachment of a molecule to one or more lysines, attachment of a molecule to one or more unnatural amino acids, enzymatic modification of an epitope or modification of a terminus. Suitable methods for making such modifications are well known in the art. The mutant monomers can be chemically modified by attachment of any molecule. For example, the mutant monomers can be chemically modified by attachment of a dye or fluorophore.
[0396] In some embodiments, the mutant monomers are chemically modified with a molecular adaptor that promotes the interaction between the pore comprising the monomers and the target nucleotide or target polynucleotide sequence. The presence of the adaptor improves the host-guest chemistry of the pore and the nucleotide or polynucleotide sequence, thereby improving the sequencing ability of the pore formed by the mutant monomers. The principles of host-guest chemistry are well known in the art. The adaptor has an influence on the physical or chemical properties of the pore, which influences improves the interaction of the pore with the nucleotide or polynucleotide sequence. The adaptor can change the charge of the barrel or channel of the pore, or specifically interact or bind with the nucleotide or polynucleotide sequence, thereby promoting its interaction with the pore.
[0397] The molecular adaptor is preferably a cyclic molecule, a cyclodextrin, a substance capable of hybridisation, a DNA binding agent or intercalator, a peptide or peptide analogue, a synthetic polymer, an aromatic planar molecule, a small molecule with a positive charge or a small molecule capable of hydrogen bonding.
[0398] The adaptor can be cyclic. The cyclic adaptor preferably has the same symmetry as the pore. The adaptor preferably has eight-fold or nine-fold symmetry, as CsgG typically has eight or nine subunits around the central axis. This is discussed in more detail below.
[0399] Adapters typically interact with nucleotides or polynucleotide sequences through host-guest chemistry. Adapters are typically capable of interacting with nucleotides or polynucleotide sequences. Adapters comprise one or more chemical groups capable of interacting with nucleotides or polynucleotide sequences. The one or more chemical groups preferably interact with nucleotides or polynucleotide sequences through non-covalent interactions, such as hydrophobic interactions, hydrogen bonding, Van der Waal’s forces, pi-cation interactions, and / or electrostatic forces. The one or more chemical groups capable of interacting with nucleotides or polynucleotide sequences are preferably positively charged. The one or more chemical groups capable of interacting with nucleotides or polynucleotide sequences are more preferably comprise an amino group. The amino group can be attached to a primary, secondary, or tertiary carbon atom. The adapter even more preferably comprises an amino ring, such as a ring of 6, 7, or 8 amino groups. The adapter most preferably comprises a ring of eight amino groups. The ring of protonated amino groups can interact with negatively charged phosphate groups in nucleotides or polynucleotide sequences.
[0400] Proper positioning of the adapter in the pore can be facilitated by host-guest chemistry between the adapter and the pore comprising mutant monomers. The adapter preferably comprises one or more chemical groups capable of interacting with one or more amino acids in the pore. The adapter more preferably comprises one or more chemical groups capable of interacting with one or more amino acids in the pore through non-covalent interactions, such as hydrophobic interactions, hydrogen bonding, Van der Waal’s forces, pi-cation interactions, and / or electrostatic forces. The chemical groups capable of interacting with one or more amino acids in the pore are typically hydroxyl groups or amines. The hydroxyl groups can be attached to primary, secondary, or tertiary carbon atoms. The hydroxyl groups can form hydrogen bonds with uncharged amino acids in the pore. Any adapter that facilitates interaction between the pore and nucleotides or polynucleotide sequences can be used.
[0401] Suitable adapters include, but are not limited to, cyclodextrins, cyclic peptides, and cucurbiturils. The adapter is preferably a cyclodextrin or derivative thereof. The cyclodextrin or derivative thereof can be any of those disclosed in Eliseev, A. V. and Schneider, H-J. (1994) J. Am. Chem. Soc. 116, 6081-6088. The adapter is more preferably am7-βCD, am1-βCD, or gu7-βCD. The guanidinium groups in gu7-βCD are much more positively charged than the primary amines in am7-βCD, and thus more positively charged. This gu7-βCD adapter can be used to increase the residence time of nucleotides in the pore, improve the accuracy of the measured residual current, and improve the rate of base detection at high temperatures or low data acquisition rates.
[0402] If a 3-(2-pyridyldithio)propionic acid succinimidyl ester (SPDP) crosslinking reagent is used as discussed in more detail below, the linker is preferably hepta(6-deoxy-6-amino)-6-N- mono(2-pyridyl)dithio propionyl- -cyclodextrin (am6amPDP1-P-CD).
[0403] More suitable linkers include γ-cyclodextrin, which comprises 9 sugar units (and thus has nine-fold symmetry). The γ-cyclodextrin can contain linker molecules, or can be modified to contain all or more of the modified sugar units used in the examples of β-cyclodextrin discussed above.
[0404] The molecular linker can be covalently attached to the mutant monomer. The linker can be covalently attached to the pore using any method known in the art. The linker is typically attached by chemical ligation. If the molecular linker is attached by cysteine ligation, it is preferred that one or more cysteines are introduced into the mutant, for example into the barrel, by substitution. The mutant monomer can be chemically modified by attachment of the molecular linker to one or more cysteines in the mutant monomer. The one or more cysteines can be naturally occurring, i.e. at position 1 and / or 215 in SEQ ID NO: 3. Alternatively, the mutant monomer can be chemically modified by attachment of the molecule to one or more cysteines introduced at other positions. The cysteine at position 215 can be removed, for example by substitution, to ensure that the molecular linker is not attached to this position, but is attached to the cysteine at position 1 or a cysteine introduced at another position.
[0405] The reactivity of a cysteine residue can be enhanced by modification of the adjacent residues. For example, the basic groups of flanking arginine, histidine or lysine residues will shift the pKa of the cysteine thiol group to that of a more reactive S - The reactivity of a cysteine residue can be protected by thiol protecting groups such as dTNB. They can react with one or more cysteine residues of the mutant monomer before attachment of the linker.
[0406] The molecule can be attached directly to the mutant monomer. It is preferred that the molecule is attached to the mutant monomer using a linker, such as a chemical crosslinking reagent or a peptide linker.
[0407] Suitable chemical cross-linking agents are well known in the art. Preferred cross-linking agents include 2,5-dioxopyrrolidin-1-yl 3-(pyridin-2- yldisulfanyl)propanoate, 2,5-dioxopyrrolidin-1-yl 4-(pyridin-2- yldisulfanyl)butanoate, and 2,5-dioxopyrrolidin-1-yl 8-(pyridin-2- yldisulfanyl)octanoate. The most preferred cross-linking agent is 3-(2- pyridyl disulfide)propionic acid succinimidyl (SPDP). Typically, the molecule is covalently attached to the bifunctional cross-linking agent prior to covalently attaching the molecule / cross-linking agent complex to the mutant monomer, but it is also possible to covalently attach the bifunctional cross-linking agent to the monomer prior to attaching the bifunctional cross-linking agent / monomer complex to the molecule.
[0408] The linker is preferably resistant to dithiothreitol (DTT). Suitable linkers include, but are not limited to, iodoacetamide-based and maleimide-based linkers.
[0409] In other embodiments, the monomer can be attached to a polynucleotide binding protein. This creates a modular sequencing system that can be used in the sequencing method of the application. Polynucleotide binding proteins are discussed below.
[0410] The polynucleotide binding protein is preferably covalently attached to the mutant monomer. The protein can be covalently attached to the monomer using any method known in the art. The monomer and protein can be chemically fused or genetically fused. If the entire construct is expressed from a single polynucleotide sequence, the monomer and protein are genetically fused. Genetic fusion of the monomer to the polynucleotide binding protein is discussed in WO 2010 / 004265.
[0411] If the polynucleotide binding protein is attached via a cysteine linkage, the one or more cysteines are preferably introduced into the mutant by substitution. The one or more cysteines are preferably introduced into a loop region that has low conservation in the homolog, indicating that mutations or insertions can be tolerated. As such, they are suitable for attachment of the polynucleotide binding protein. In such embodiments, the naturally occurring cysteine at position 251 can be removed. As discussed above, the reactivity of the cysteine residue can be enhanced by modification.
[0412] The polynucleotide binding protein can be attached directly to the mutant monomer, or attached via one or more linkers. The molecule can be attached to the mutant monomer using a hybridization linker as described in WO 2010 / 086602. Alternatively, a peptide linker can be used. A peptide linker is an amino acid sequence. The length, flexibility and hydrophilicity of the peptide linker are typically designed such that it does not perturb the function of the monomer and the molecule. Preferred flexible peptide linkers are stretches of 2 to 20, such as 4, 6, 8, 10 or 16, serine and / or glycine amino acids. More preferred flexible linkers include (SG)i, (SG)2, (SG)3, (SG)4, (SG)5and (SG)8, where S is serine and G is glycine. Preferred rigid peptide linkers are stretches of 2 to 30, such as 4, 6, 8, 16 or 24, proline amino acids. More preferred rigid linkers include (P) 12 where P is proline.
[0413] Chemical modification
[0414] The mutant CsgG monomer or CsgF peptide can be chemically modified with a molecular adaptor and a polynucleotide binding protein.
[0415] The molecule (with which the monomer or peptide is chemically modified) can be attached directly to the monomer or peptide, or attached via a linker, as disclosed in WO 2010 / 004273, WO 2010 / 004265 or WO 2010 / 086603.
[0416] Any of the proteins described herein, such as CsgG monomers and / or CsgF peptides, can be modified to facilitate their identification or purification, for example by the addition of a histidine residue (his-tag), an aspartic acid residue (asp-tag), a streptavidin tag, a flag tag, a SUMO tag, a GST tag or a MBP tag, or by the addition of a signal sequence to facilitate their secretion from cells in which the polypeptide does not naturally comprise such a sequence. An alternative to introducing a genetic tag is to chemically react a tag with a native or engineered position on the protein. One example of this is to react a gel-migration reagent with an engineered cysteine on the outside of the protein. This has been demonstrated as a method for isolating hemolysin hetero-oligomers (Chem Biol. 1997 Jul;4(7):497-505).
[0417] Any of the proteins described herein, such as CsgG monomers and / or CsgF peptides, can be labelled with an exposable marker. An exposable marker can be any suitable marker that allows detection of the protein. Suitable markers include, but are not limited to, fluorescent molecules, radioisotopes (e.g. 125 I, 35 S), enzymes, antibodies, antigens, polynucleotides and ligands (such as biotin).
[0418] Any of the proteins described herein, such as CsgG monomers and / or CsgF peptides, can be produced synthetically or by recombinant means. For example, the proteins can be synthesized by in vitro translation and transcription (IVTT). The amino acid sequence of the protein can be modified to include non-naturally occurring amino acids or to increase the stability of the protein. Such amino acids can be introduced during production when the protein is produced synthetically. The protein can also be altered after synthetic or recombinant production.
[0419] Proteins can also be produced using D-amino acids. For example, the proteins can comprise a mixture of L-amino acids and D-amino acids. This is routine in the art of producing such proteins or peptides.
[0420] The proteins can also contain other non-specific modifications so long as they do not interfere with the function of the protein. Many non-specific side chain modifications are known in the art and can be made to the side chains of the proteins. Such modifications include, for example, reductive alkylation of amino acids by reaction with an aldehyde followed by reduction with NaBH4, amidation with imidoacetic acid methyl ester, or acylation with acetic anhydride.
[0421] Any of the proteins described herein, such as CsgG monomers and / or CsgF peptides, can be produced using standard methods known in the art. Polynucleotide sequences encoding the proteins can be obtained and copied using standard methods in the art. The polynucleotide sequences encoding the proteins can be expressed in bacterial host cells using standard techniques in the art. The proteins can be produced in the cells by expressing the polypeptides from the recombinant expression vectors in situ. The expression vectors optionally carry an inducible promoter to control expression of the polypeptides. These methods are described in Sambrook, J. and Russell, D. (2001). Molecular Cloning: A Laboratory Manual, 3rd ed. Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY.
[0422] The proteins can be produced on a large scale after purification from the organism producing the protein by any protein liquid chromatography system or after recombinant expression. Typical protein liquid chromatography systems include FPLC, AKTA systems, Bio-Cad systems, Bio-Rad BioLogic systems, and Gilson HPLC systems.
[0423] Methods of producing pores
[0424] In a third aspect, the present invention provides methods of producing a CsgG: modified CsgF pore complex having two or more constriction sites in vivo and in vitro. One embodiment provides a method of producing a transmembrane pore complex comprising a CsgG pore or homolog or mutant form thereof, and a modified CsgF peptide or homolog or mutant thereof, by co-expression. The method comprises the steps of expressing a CsgG monomer (expressed as the proprotein provided in SEQ ID NO: 2, or a homolog or mutant thereof) and expressing a modified or truncated CsgF monomer in a suitable host cell, thereby allowing the formation of a complex pore in vivo. The complex comprises a modified CsgF peptide complexed with a CsgG pore to provide additional reading heads for the pore. The resulting pore complex produced by the method using a modified CsgF peptide provides a structure sufficient to use the pore complex to characterize target analytes, such as nucleic acid sequencing, as it allows the passage of analytes, especially polynucleotide strands, and comprises two or more reading heads for improved reading of the polynucleotide sequence when used in the appropriate environment for said applications.
[0425] More particularly, the modified CsgF peptide expressed in the method comprises the proprotein depicted in SEQ ID NO: 8, 10, 12 or 14, or a homolog thereof. Those sequences limit the method to those CsgF fragments capable of introducing constriction sites in the pore complex and binding to the CsgG protein pore to obtain a biological pore.
[0426] Another method of producing isolated pore complexes formed by CsgG and CsgF proteins, etc. involves in vitro reconstitution of the monomers to obtain functional pores. The method includes the step of contacting the mature CsgG monomer depicted in SEQ ID NO: 3 or a homolog or mutant thereof with a modified CsgF peptide or a homolog or mutant thereof in a suitable system to allow complex formation. The system can be an "in vitro" system, meaning a system that contains at least the necessary components and environment to perform the method, and utilizes biological molecules, organisms, cells (or parts of cells) outside of their normal, naturally occurring environment, permitting analysis that is more detailed, more convenient, or more efficient than can be performed with whole organisms. The in vitro system can also include a suitable buffer composition provided in a test tube, with the protein components that form the complex having been added. One of skill in the art knows the options for providing the system. In particular embodiments, the modified CsgF peptide or the like applied in the method for in vitro reconstitution is a peptide comprising SEQ ID NO: 15 or SEQ ID NO: 16, or a mutant or homolog thereof, which can be produced synthetically or recombinantly. Alternatively, a modified CsgF peptide comprising SEQ ID NO: 40, 39, 38, or 37, 15, 54, 55, or a homolog or mutant thereof is provided in the method for contact with CsgG or CsgG-like pores to produce pore complexes.
[0427] CsgG / CsgF pores can be made by any suitable method. Examples of such suitable methods are described.
[0428] In one embodiment, CsgG / CsgF pores can be produced by co-expression. In this embodiment, at least one gene encoding a CsgG monomer polypeptide (which can be a mutant polypeptide) in one vector and a gene encoding at least one full-length or truncated CsgF polypeptide (which can be a mutant polypeptide) in a second vector can be transformed together to express the proteins and produce the complex in the transformed cells. This can be in vivo or in vitro. Alternatively, the two genes encoding CsgG and CsgF polypeptides can be placed in one vector under the control of a single promoter or under the control of two separate promoters which can be the same or different.
[0429] In another embodiment, CsgG / CsgF pores are produced by expressing CsgG monomers separately from CsgF peptides. CsgG monomers or CsgG pores can be purified from cells transformed with a vector encoding at least one CsgG monomer or with more than one vector each expressing a CsgG monomer. CsgF peptides can be purified from cells transformed with a vector encoding at least one CsgF peptide. Purified CsgG monomers / pores can then be incubated with CsgF peptides to make pore complexes.
[0430] In another embodiment, the CsgG monomer and / or the CsgF peptide are produced separately by in vitro translation and transcription (IVTT). The CsgG monomer can then be incubated with the CsgF peptide to make the pore complex. Figure 14 The use of this method is illustrated in the Examples.
[0431] The above embodiments can be combined, such that, for example, (i) CsgG is produced in vivo and CsgF is produced in vivo; (ii) CsgG is produced in vitro and CsgF is produced in vivo; (iii) CsgG is produced in vivo and CsgF is produced in vitro; (iv) CsgG is produced in vitro and CsgF is produced in vitro.
[0432] One or both of the CsgG monomer and the CsgF peptide can be labeled to facilitate purification. Purification can also be performed when the CsgG monomer and / or the CsgF peptide are not labeled. Methods known in the art (e.g., ion exchange, gel filtration, hydrophobic interaction column chromatography, etc.) can be used, alone or in different combinations, to purify the components of the pore.
[0433] Any known tag can be used in either of the two proteins. In one embodiment, a dual tag purification method can be used to purify the CsgG:CsgF complex from CsgG pore and CsgF. For example, a Strep tag can be used in CsgG and a His tag in CsgF, or vice versa. Figure 13 This is exemplified. A similar end result can be obtained when the two proteins are purified separately and mixed together, followed by another round of Strep and His purification.
[0434] When the full-length CsgF protein forms a complex with CsgG, the neck and head domains of CsgF ( Figure 4B ) protrude from the beta barrel of the CsgG pore.
[0435] Thus, if a pore comprising CsgG pore and full-length CsgF is used in a single channel recording experiment, the head domains can hinder or prevent the pore from inserting into the membrane. They can also block the passage through the pore that the analyte is to pass through. Therefore, it is preferable to reduce the number of flexible polypeptides that hang from the beta-barrel when the pore is inserted into the membrane. Provided herein are truncated forms of the CsgF protein that mimic the FCP region resolved in the cryo EM structure of the complex and maintain structural integrity.
[0436] The CsgG / CsgF pore can be made prior to insertion into the membrane or after insertion of the CsgG pore into the membrane. When the pore complex is made prior to insertion into the membrane, it is preferred to use a truncated mutant. However, the CsgG pore can be inserted into the membrane and then the CsgF peptide can be added, such that the CsgG and CsgF complex can be formed in situ. For example, in one embodiment in a system where the opposite side of the membrane is accessible (e.g. in a chip or chamber for electrophysiological measurements), the CsgG pore can be inserted into the membrane and then the CsgF peptide can be added from the opposite side of the membrane, such that the complex can be formed in situ. In any embodiment where the CsgG pore is formed in situ, a larger CsgF peptide can be used. For example, the CsgF peptide can comprise the entire or partial neck domain of CsgF (approximately starting at residue 36 of SEQ ID NO: 6). In some embodiments, CsgF can comprise the entire neck domain and a portion of the head domain (residues 36 to XX of SEQ ID NO: 6).
[0437] Depending on the method of making the complex and the stability of the complex with a particular truncation, the CsgG:CsgF and CsgG:FCP complexes can be made in different ways.
[0438] In one embodiment, a truncated form of the CsgF polypeptide of the desired length is used directly.
[0439] Another embodiment uses a full-length polypeptide of CsgF or a polypeptide longer than the desired truncation with an inserted protease cleavage site (e.g. TEV, HRV 3, or any other protease cleavage site) such that the CsgF peptide of the desired length is produced by protease cleavage. In this embodiment, once the CsgG / CsgF complex is formed, a protease is used to cleave CsgF at the desired site. Alternatively, the CsgF peptide can be produced using a protease prior to assembly of the complex.
[0440] Some protease sites leave an additional tag after cleavage. For example, the TEV protease cleavage sequence is ENLYFQS. The TEV protease cleaves the protein between Q and S, leaving the entire ENLYFQ at the C-terminus of the CsgF peptide. Figure 15 shows an example of using a modified CsgF comprising a TEV cleavage site, cleaving the modified CsgF using TEV protease after complex formation.
[0441] As another example, the HRV C3 cleavage site is LEVLFQGP, and the enzyme cleaves between Q and G, leaving the entire LEVLFQ at the C-terminus of the CsgF peptide.
[0442] Methods of characterizing an analyte
[0443] In another aspect, the application provides a method of determining the presence, absence, or one or more characteristics of a target analyte. The method comprises contacting the target analyte with an isolated pore complex or transmembrane pore, such as a pore of the application, such that the target analyte moves relative to the pore channel, e.g., into or through the pore channel and making one or more measurements as the analyte moves relative to the pore, thereby determining the presence, absence, or one or more characteristics of the analyte. The target analyte can also be referred to as a template analyte or a target analyte. The isolated pore complex typically comprises at least 7, at least 8, at least 9, or at least 10 monomers, such as 7, 8, 9, or 10 CsgG monomers. The isolated pore complex preferably comprises eight or nine identical CsgG monomers. Preferably one or more, such as 2, 3, 4, 5, 6, 7, 8, 9, or 10 CsgG monomers are chemically modified, or a CsgF peptide is chemically modified. The isolated pore complex monomers, such as CsgG monomers, or homologues or mutants thereof, and the modified CsgF monomers, or homologues or mutants thereof, can be derived from any organism. The analyte can pass through a CsgG constriction segment first, then through a CsgF constriction segment. In an alternative embodiment, the analyte can pass through a CsgF constriction segment first, then through a CsgG constriction segment, depending on the orientation of the CsgG / CsgF complex in the membrane.
[0444] The method is for determining the presence, absence, or one or more characteristics of a target analyte. The method can be for determining the presence, absence, or one or more characteristics of at least one analyte. The method can involve determining the presence, absence, or one or more characteristics of two or more analytes. The method can comprise determining the presence, absence, or one or more characteristics of any number of analytes, such as 2, 5, 10, 15, 20, 30, 40, 50, 100, or more analytes. Any number of characteristics of the one or more analytes can be determined, such as 1, 2, 3, 4, 5, 10, or more characteristics.
[0445] Molecular binding in or near either opening of a channel of a pore complex will have an effect on the open channel ion flow through the pore, which is the basis of pore channel "molecular sensing". Changes in open channel ion flow can be measured by changes in current using suitable measurement techniques in a manner analogous to nucleic acid sequencing applications (e.g. WO 2000 / 28312 and D. Stoddart et al, Proc. Natl. Acad. Sci., 2010, 106, 7702-7 or WO 2009 / 077734). The extent of reduction in ion flow, as measured by a reduction in current, is related to the size of the occlusion in or near the pore. Thus, binding of a target molecule (also referred to as an "analyte") in or near the pore provides a detectable and measurable event, forming the basis of a "biosensor". Molecules suitable for nanopore sensing include nucleic acids; proteins; peptides; polysaccharides and small molecules (here referring to low molecular weight (e.g. <900 Da or <500 Da) organic or inorganic compounds) such as drugs, toxins, cytokines and pollutants. Detecting the presence of biological molecules can be used in personalised medicine development, medicine, diagnostics, life science research, environmental monitoring and in the security and / or defence industries.
[0446] In another aspect, an isolated pore complex or transmembrane pore complex containing a wild-type or modified E. coli CsgG nanopore or homolog or mutant thereof, and a modified CsgF peptide providing a constriction segment of the channel for the pore in the complex, can be used as a molecular or biosensor. In some embodiments, the CsgG nanopore can be derived or isolated from a bacterial protein (e.g. E. coli, Salmonella typhi). In some embodiments, the CsgG nanopore can be produced recombinantly. Procedures for analyte detection are described in Howorka et al. Nature Biotechnology (2012) 7 June; 30(6):506-7. The analyte molecule to be detected can bind on either face of the channel, or in the lumen of the channel itself. The location of binding can be determined by the size of the molecule to be sensed.
[0447] The target analyte is preferably a metal ion, inorganic salt, polymer, amino acid, peptide, polypeptide, protein, nucleotide, oligonucleotide, polynucleotide, polysaccharide, dye, bleach, drug, diagnostic agent, recreational drug, explosive, toxic compound or environmental pollutant. The method can involve determining the presence, absence or one or more characteristics of two or more analytes of the same type, such as two or more proteins, two or more nucleotides or two or more drugs. Alternatively, the method can involve determining the presence, absence or one or more characteristics of two or more analytes of different types, such as one or more proteins, one or more nucleotides or one or more drugs.
[0448] The target analyte can be secreted from a cell. Alternatively, the target analyte can be an analyte that is present inside a cell, such that the analyte must be extracted from the cell before the method can be performed.
[0449] Wild-type pores can act as sensors, but are often modified by recombinant or chemical means to increase the binding strength, binding location, or binding specificity of the molecule to be sensed. Typical modifications include the addition of a specific binding moiety that is complementary to the structure of the molecule to be sensed. In the case of an analyte molecule comprising a nucleic acid, the binding moiety can comprise a cyclodextrin or an oligonucleotide; for small molecules, this can be a known complementary binding region, such as an antibody or an antigen-binding portion of a non-antibody molecule, including a single-chain variable fragment (scFv) region or an antigen recognition domain from a T-cell receptor (TCR); or for proteins, it can be a known ligand for the target protein. In this way, wild-type or modified E. coli CsgG nanopores or homologues thereof can be made to act as molecular sensors for detecting the presence of suitable antigens (including epitopes) in a sample, which can include cell surface antigens (including markers of receptor, solid tumor, or blood cancer cells (e.g., lymphoma or leukemia)), viral antigens, bacterial antigens, protozoal antigens, allergens, allergy-related molecules, albumins (e.g., human, rodent, or bovine), fluorescent molecules (including fluorescein), blood group antigens, small molecules, drugs, enzymes, catalytic sites of enzymes or enzyme substrates, and transition state analogs of enzyme substrates. As noted above, the modifications can be achieved using known genetic engineering and recombinant DNA techniques. The location of any adaptation will depend on the nature of the molecule to be sensed, such as size, three-dimensional structure, and its biochemical properties. The selection of adapted structures can utilize computational structure design. Techniques such as surface plasmon resonance for detecting molecular interactions (BIAcore, Inc., Piscataway, NJ; see also www.biacore.com) can be used to investigate and optimize assays for protein-protein interactions or protein-small molecule interactions. (BIAcore, Inc., Piscataway, NJ; see also www.biacore.com) can be used to investigate and optimize assays for protein-protein interactions or protein-small molecule interactions.
[0450] In one embodiment, the analyte is an amino acid, a peptide, a polypeptide, or a protein. The amino acid, peptide, polypeptide, or protein can be naturally occurring or non-naturally occurring. The polypeptide or protein can comprise synthetic or modified amino acids therein. Several different types of modifications of amino acids are known in the art. Suitable amino acids and modifications thereof are described above. It will be appreciated that the target analyte can be modified by any method available in the art.
[0451] In another embodiment, the analyte is a polynucleotide, such as a nucleic acid, defined as a macromolecule comprising two or more nucleotides. Nucleic acids are particularly suitable for nanopore sequencing. The naturally occurring nucleic acid bases in DNA and RNA can be distinguished by their actual size. When a nucleic acid molecule or individual bases are passed through the channel of a nanopore, the size difference between the bases causes a directly related decrease in the ionic current flow through the channel. Changes in the ionic current flow can be recorded. Electrical measurement techniques suitable for recording changes in ionic current flow are described, for example, in WO 2000 / 28312 and D. Stoddart et al., Proc. Natl. Acad. Sci., 2010, 106, pp. 7702-7 (single channel recording devices); and, for example, in WO 2009 / 077734 (multi-channel recording techniques). With suitable calibration, the characteristic decrease in ionic current flow can be used to identify the particular nucleotide and associated base passing through the channel in real time. In typical nanopore nucleic acid sequencing, as individual nucleotides of a target nucleic acid sequence are passed in turn through the channel of a nanopore, the open channel ionic current flow decreases due to the partial blockage of the channel by the nucleotide. This is the decrease in ionic current flow that is measured using suitable recording techniques described above. The decrease in ionic current flow can be calibrated to the decrease in ionic current flow measured for a known nucleotide passing through the channel, giving a way to determine the nucleotide that is passing through the channel, and hence, when carried out sequentially, a way to determine the nucleotide sequence of a nucleic acid passing through the nanopore. In order to accurately determine individual nucleotides, it is generally necessary to relate the decrease in ionic current flow through the channel directly to the size of the individual nucleotide passing through the constriction (or“read head”). It will be appreciated that, for example, a complete nucleic acid polymer can be sequenced that is“threaded” through the pore by the action of an associated polymerase. Alternatively, the sequence can be determined by passing nucleotide triphosphate bases that have been removed sequentially from a target nucleic acid in close proximity to the pore (see, for example, WO 2014 / 187924).
[0452] A polynucleotide or nucleic acid can comprise any combination of any nucleotides. The nucleotides can be naturally occurring or artificial. One or more nucleotides in a polynucleotide can be oxidized or methylated. One or more nucleotides in a polynucleotide can be damaged. For example, a polynucleotide can comprise pyrimidine dimers. Such dimers are typically associated with UV light damage and are a major cause of skin melanoma. One or more nucleotides in a polynucleotide can be modified, e.g., with a label or tag, suitable examples of which are known to those skilled in the art. A polynucleotide can comprise one or more spacers. A nucleotide typically contains a nucleobase, a sugar, and at least one phosphate group. The nucleobase and sugar form a nucleoside. The nucleobase is typically a heterocycle. Nucleobases include, but are not limited to, purines and pyrimidines, and more specifically, adenine (A), guanine (G), thymine (T), uracil (U), and cytosine (C). The sugar is typically a pentose sugar. Nucleotide sugars include, but are not limited to, ribose and deoxyribose. The sugar is preferably deoxyribose. A polynucleotide preferably comprises the following nucleosides: deoxyadenosine (dA), deoxyuridine (dU) and / or deoxythymidine (dT), deoxyguanosine (dG), and deoxycytidine (dC). The nucleotide is typically a ribonucleotide or a deoxyribonucleotide. The nucleotide typically contains a mono-, di-, or tri-phosphate. The nucleotide can comprise more than three phosphates, e.g., 4 or 5 phosphates. The phosphate can be attached at the 5' or 3' side of the nucleotide. The nucleotides in a polynucleotide can be attached to each other in any manner. As in nucleic acids, the nucleotides are typically attached through the sugar and phosphate groups. As in pyrimidine dimers, the nucleotides can be linked through their nucleobases. A polynucleotide can be single-stranded or double-stranded. At least a portion of a polynucleotide is preferably double-stranded. A polynucleotide is most preferably a ribonucleic acid (RNA) or a deoxyribonucleic acid (DNA). In particular, the method using a polynucleotide as an analyte can alternatively comprise determining one or more features selected from (i) the length of the polynucleotide, (ii) the identity of the polynucleotide, (iii) the sequence of the polynucleotide, (iv) the secondary structure of the polynucleotide, and (v) whether the polynucleotide is modified.
[0453] The polynucleotide can be any length (i). For example, the polynucleotide can be at least 10, at least 50, at least 100, at least 150, at least 200, at least 250, at least 300, at least 400, or at least 500 nucleotides or nucleotide pairs in length. The polynucleotide can be 1000 or more nucleotides or nucleotide pairs, 5000 or more nucleotides or nucleotide pairs, or 100,000 or more nucleotides or nucleotide pairs in length. Any number of polynucleotides can be investigated. For example, the method can involve characterizing 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 50, 100 or more polynucleotides. If two or more polynucleotides are characterized, they can be different polynucleotides or two examples of the same polynucleotide. The polynucleotide can be naturally occurring or artificial. For example, the method can be used to check the sequence of a manufactured oligonucleotide. The method is typically performed in vitro.
[0454] The nucleotides can be of any identity (ii), including but not limited to adenosine monophosphate (AMP), guanosine monophosphate (GMP), thymidine monophosphate (TMP), uridine monophosphate (UMP), 5-methylcytidine monophosphate, 5-hydroxymethylcytidine monophosphate, cytidine monophosphate (CMP), cyclic adenosine monophosphate (cAMP), cyclic guanosine monophosphate (cGMP), deoxyadenosine monophosphate (dAMP), deoxyguanosine monophosphate (dGMP), deoxythymidine monophosphate (dTMP), deoxyuridine monophosphate (dUMP), deoxycytidine monophosphate (dCMP), and deoxymethylcytidine monophosphate. The nucleotides are preferably selected from AMP, TMP, GMP, CMP, UMP, dAMP, dTMP, dGMP, dCMP, and dUMP. The nucleotides can be abasic (i.e. lack a nucleobase). The nucleotides can also lack both a nucleobase and a sugar (i.e. be a C3 spacer). The sequence of the nucleotides (iii) is determined by the consecutive identity of the nucleotides attached to each other in the 5' to 3' direction along the chain throughout the polynucleotide chain.
[0455] The pore comprising a CsgG pore and a CsgF peptide is particularly useful in analysing homopolymers. For example, the pore can be used to determine the sequence of a polynucleotide comprising two or more, for example at least 3, 4, 5, 6, 7, 8, 9, or 10 identical consecutive nucleotides. For example, the pore can be used to sequence a polynucleotide comprising a polyA, polyT, polyG, and / or polyC region.
[0456] The constriction segment of the CsgG pore is composed of residues at positions 51, 55, and 56 of SEQ ID NO: 3. The reading heads of CsgG and its constricted mutants are typically quite sharp. When DNA is threaded through the constriction segment, at any given time, approximately 5 bases of the DNA interact with the reading heads of the pore, dominating the current signal. While these sharper reading heads are very good at reading mixed sequence regions of DNA (when A, T, G, and C are mixed), when homopolymer regions are present in the DNA (e.g., polyT, polyG, polyA, polyC), the signal becomes flat and lacks information. Because 5 bases dominate the signal of CsgG and its constricted mutants, it is difficult to distinguish polymers longer than 5 bases without using other dwell time information. However, if the DNA is threaded through a second reading head, more DNA bases will interact with the combined reading heads, increasing the length of homopolymers that can be distinguished. The examples and figures show that this increase in homopolymer sequencing accuracy is achieved using a pore comprising a CsgG pore and a CsgF peptide.
[0457] Kit
[0458] In another aspect, the present application also provides a kit for characterizing a target polynucleotide. The kit comprises an isolated pore complex according to the present application, and components of a membrane or insulating layer. The membrane is preferably formed from the components. The isolated pore complex is preferably present in the membrane or insulating layer, together forming a transmembrane pore complex channel. The kit can comprise components of any type of membrane, such as an amphiphilic layer or a triblock copolymer membrane. The kit can also comprise a polynucleotide binding protein. The kit can also comprise one or more anchors for coupling a polynucleotide to the membrane. The kit can additionally comprise one or more other reagents or instruments that enable any of the embodiments mentioned above to be implemented. Such reagents or instruments include one or more of the following: a suitable buffer (aqueous solution), a device for obtaining a sample from a subject (such as a container or instrument comprising a needle), a device for amplifying and / or expressing a polynucleotide, or a voltage clamp or patch clamp apparatus. The reagents can be present in the kit in a dry state, such that a liquid sample re-suspends the reagents. The kit can also optionally comprise instructions enabling the kit to be used in the methods of the present application, or details about which organisms the methods can be used on. Finally, the kit can also comprise additional components that can be used in the characterization of the peptide.
[0459] In one embodiment, the isolated pore complex or transmembrane pore complex as provided for by the present application is used for nucleic acid sequencing. For said use, Phi29 DNA polymerase (DNAP) can be used as a molecular motor with the CsgG:CsgF nanopore complex located within the membrane to allow controlled movement of an oligomeric probe DNA strand through the pore. A voltage can be applied across the pore and an electric current is generated due to the movement of ions in the salt solution on either side of the nanopore. As the probe DNA moves through the pore, the flow of ions through the pore changes relative to the DNA. This information has been shown to be sequence dependent and allows the sequence of the probe to be accurately read from the current measurements.
[0460] It is to be understood that even though particular embodiments, specific constructions, and materials have been discussed herein for purposes of the application of the inventive cell and method according to the present application, various modifications, alterations, and improvements can be implemented and incorporated into the application without departing from the scope and spirit of the application. The following examples are provided to better illustrate particular embodiments and should not be considered as limiting the application. The application is limited only by the claims.
[0461] Examples
[0462] Introduction
[0463] The CsgG pore is part of the type VIII multicomponent secretion system, also known as the curli biosynthesis system, which is responsible for the formation of the aggregated fibers called curli in E. coli. Curli are extracellular protein fibers that are mainly involved in the formation of bacterial biofilms and attachment to non-biological surfaces. Curli biosynthesis is directed by two operons in E. coli called csgBAC and csgDEFG (curli-specific genes) (Hammar et al., 1995). Secretion of the curli subunits CsgA and CsgB depends on CsgG, a dedicated lipoprotein that is found to form an oligomeric secretion channel in the outer membrane. For transport, CsgG cooperates with the periplasmic and extracellular helper proteins CsgE and CsgF. CsgE forms a specificity factor for CsgG-mediated transport, while CsgF seems to couple the secretion of CsgA with the template-assembly of extracellular fibers with CsgB.
[0464] The crystal structure of the CsgG secretion channel proved that CsgG forms a nine-mer transport complex with a diameter and a height of 36-strand-barrel (Goyal et al., 2014; Fig. 1) that allows the OM to pass through the inner diameter of the channel. The periplasmic domain of the channel is separated from the transmembrane β-barrel by an iris-like septum, forming large solvent accessible cavity. This membrane is formed by a conserved 12-residue "constriction loop" (CL) found in each subunit, which in the context of the CsgG oligomer forms a ~0.6 nm diameter and ~1.5 nm height pore that does not include solvent (Figure 1). This constriction or constriction segment in the CsgG channel forms the main site of interaction with translocation polypeptides when the protein functions as a protein secretion channel. When CsgG is used as a nanopore sensing platform, this constriction serves as the main readhead for analytes that remain in the channel or pass through the channel. The diameter of the constriction and its physical and chemical properties can be altered by amino acid substitution, deletion, or insertion in the protein region corresponding to residues 46 to 61 of SEQ ID NO: 3 Figure 1 D ). In particular, mutations according to positions 51, 55, and 56 of SEQ ID NO: 3, alone or together, have a beneficial influence on the conductance characteristics of the nanopore and its interaction with analytes, including polynucleotides.
[0465] Assembly factor CsgF represents a component of the curli secretion machinery. The CsgF preprotein reaches the periplasm via the SEC pathway, after which the mature CsgF (12.9 KDa) is found as a surface exposed protein in a CsgG dependent manner. CsgF is isolated with OM in the presence of CsgG and immunoprecipitation experiments indicate a direct contact of these two proteins. Available data demonstrate that CsgF is not essential for productive subunit secretion, rather the protein is suggested to form a coupling factor between CsgA secretion and extracellular polymerization into curli fibers by coordinating or chaperoning the nucleation function of CsgB subunits.
[0466] Example 1 : Production of CsgG:CsgF complex proteins (Co-expression of CsgG with CsgF synthetic peptides, in vitro reconstitution, in vitro coupled transcription and translation, and reconstitution)
[0467] To produce the CsgG:CsgF complex, the two proteins can be co-expressed in a suitable Gram-negative host such as E. coli, and extracted and purified from the outer membrane as a complex. The formation of the CsgG pore and the CsgG:CsgF complex in vivo requires targeting of the proteins to the outer membrane. To this end, CsgG is expressed as a prepro-protein with a lipoprotein signal peptide (Juncker et al. 2003, Protein Sci. 12(8): 1652-62) and a Cys residue at the N-terminus of the mature protein (SEQ ID NO: 3). An example of such a lipoprotein signal peptide is residues 1-15 of the full-length E. coli CsgG shown in SEQ ID NO: 2. Processing of the prepro-CsgG results in cleavage of the signal peptide and lipidation of the mature CsgG, followed by translocation of the mature lipoprotein to the outer membrane, where it inserts as an oligomeric pore (Goyal et al., 2014, Nature 516(7530):250-3). To form the CsgG:CsgF complex, CsgF can be co-expressed with CsgG and targeted to the periplasm by means of a leader sequence, such as the natural signal peptide corresponding to residues 1-19 of SEQ ID NO: 5. The CsgG:CsgF combined pore can then be extracted from the outer membrane using a detergent and purified as a uniform complex by chromatography (Figure 2).
[0468] Alternatively, the CsgG:CsgF pore complex can be produced by in vitro reconstitution using CsgG pores and CsgF - see below and Figure 3.
[0469] For the formation of the CsgG:CsgF complex in vivo in the example shown in Figure 2, E. coli CsgF (SEQ ID NO: 5) and CsgG (SEQ ID NO: 2) were co-expressed using their natural signal peptides to ensure targeting of both proteins to the periplasm, as well as N-terminal lipidation of CsgG. In addition, for ease of purification, CsgF was modified by the introduction of a C-terminal 6x histidine tag, and the CsgG C-terminus was fused to a Strep-II tag. Co-expression and complex purification were performed as described in the Methods. SDS-PAGE analysis of the His-affinity purification eluate revealed enrichment of CsgF-His and co-purification of CsgG-Strep, indicating that the latter is in complex with CsgF Figure 2B ) with an asterisk. SDS-PAGE analysis of pooled fractions of the His-trap eluate of the second affinity purification revealed the presence of apparent equimolar concentrations of CsgG and CsgF, as well as the loss of the CsgF truncation fragment seen in the His-trap eluate Figure 2B Figure 2B ) co-elution of CsgF in Strep affinity purification indicates that the protein exists as a non-covalent complex with CsgG. Surprisingly, N-terminal truncation fragments of CsgF are lost in Strep affinity purification, indicating that the N-terminus of CsgF is required for binding to CsgG Figure 2B
[0470] Figure 13 Another example of CsgG:CsgF complex formation by in vivo co-expression is shown. In this example, the CsgG protein was modified with a C-terminal Strep-II tag, while the CsgF full-length protein was modified with a C-terminal 10X histidine tag. Co-expressed CsgG:CsgF complex was purified from its constituent components by Strep tag purification followed by histidine tag purification as described in the Materials and Methods section for analyte characterization. Due to the difference in molecular weight, CsgG:CsgF complex can be clearly distinguished from CsgG pore in SDS-PAGE analysis Figure 13 A). As shown in B, both tag purification methods can be successfully applied to purify CsgG:CsgF complex from its constituent components. Figure 13
[0471] To generate CsgG:CsgF complex by in vitro reconstitution, CsgG and CsgF were expressed in separate E. coli cultures transformed with pPG1 and pNA101, respectively, and purified, then CsgG:CsgF complex was reconstituted in vitro (see Methods). For comparison, purified CsgG was run on a Superose 6 column similarly to the complex. CsgG Superose 6 run showed the presence of two discrete populations corresponding to the nonamer CsgG pore Figure 3A (a) and 3C) and the dimer of nonamer CsgG pore Figure 3A (b) and 3C) as previously described in Goyal et al. (2014). CsgG:CsgF reconstitution Superose 6 run revealed the presence of three discrete populations corresponding to excess CsgF Figure 3A (c)), nonamer CsgG:CsgF complex Figure 3A (d)), and the dimer of nonamer CsgG:CsgF Figure 3A (e)). To provide independent confirmation of CsgG:CsgF complex formation, individual Superose 6 elution peaks Figure 3B
[0472] Surprisingly, the CsgG:CsgF complex can also be prepared via in vitro transcription and translation (IVTT), as described in the Materials and Methods section for analyte characterization. The complex can be prepared by expressing the CsgG and CsgF proteins in the same IVTT reaction or by reconstructing CsgG and CsgF separately in two different IVTT reactions. Figure 14 In the example shown, a CsgG:CsgF complex was prepared in a reaction mixture using the E. coli T7-S30 circular DNA extraction system (Promega), and the protein was analyzed on SDS-PAGE. Because protein expression in IVTT does not utilize the native molecular mechanisms of protein expression, the DNA used for protein expression in IVTT lacks the DNA encoding the signal peptide region. When CsgG DNA is expressed in IVTT in the absence of CsgF DNA, only CsgG monomers are produced. Surprisingly, these expressed monomers can be in situ assembled into CsgG oligomer pores using cell extract membranes present in the IVTT reaction mixture. Figure 14 Lane 1). Although the oligomers of CsgG are SDS-stable, they decompose into their constituent monomers when the sample is heated to 100°C. Figure 14 Lane 2). When CsgF DNA is expressed in IVTT in the absence of CsgG DNA, only CsgF monomers can be observed. Figure 14 Lane 3). When CsgG and CsgF DNA are mixed in a 1:1 ratio and simultaneously expressed in the same IVTT reaction mixture, the resulting CsgF protein interacts efficiently with the assembled CsgG pores to produce a CsgG:CsgF complex. Figure 14 Lane 5). This SDS-stable complex produced in IVTT exhibits thermal stability at least up to 70°C. Figure 14 (Lane 6-12)
[0473] CsgG:CsgF complexes with truncated CsgF can also be made by using DNA encoding a truncated rather than a full-length form of CsgF by any of the methods shown above. However, upon truncation of CsgF below the FCP domain, the stability of the complex can be compromised. Alternatively, once a full-length CsgG:CsgF complex is formed, a CsgG:CsgF complex with truncated CsgF can be made by cleaving the full-length CsgF at the appropriate position. Truncation can be performed by modifying the DNA encoding the CsgF protein by incorporating a protease cleavage site at the position where cleavage is desired (Figure 15.A). Seq ID Nos. 56-67 show TEV or HCV C3 protease sites incorporated at various positions of CsgF to generate CsgG:CsgF complexes with truncated CsgF. SDS-PAGE analysis of TEV cleavage of a CsgG:CsgF complex made with Seq ID No. 61 is shown (Figure 15.B). Upon treatment of a CsgG:CsgF complex (with full-length CsgF) with TEV protease, CsgF was truncated at position 35 as described in the Materials and Methods section for analytical characterization (Figure 15.B, lanes 3 and 4). However, the TEV cleavage left an additional 6 amino acids C-terminal to the cleavage site. Thus, the remaining CsgF truncated protein complexed with the CsgG pore is 42 amino acids long. The difference in molecular weight of this complex and the CsgG pore (without CsgF) is still visible in SDS-PAGE (Figure 15.B, lanes 7 and 8).
[0474] Surprisingly, CsgG:CsgF complexes with truncated CsgF can also be prepared by reconstituting purified CsgG pores (prepared either in vivo or in vitro) with synthetic peptides of appropriate length. Since reconstitution is performed in vitro, the signal peptide of CsgF is not required for the preparation of CsgG:CsgF complexes. Furthermore, this method does not leave additional amino acids at the C-terminus of CsgF. Mutations and modifications can also be easily incorporated into synthetic CsgF peptides. Thus, this method is a very convenient way to reconstitute different CsgG pores or mutants or homologues thereof with different CsgF peptides or mutants or homologues thereof to generate different CsgG:CsgF complex variants. The stability of the complex can be impaired when the FCP domain is truncated beyond CsgF. Table 3 shows examples of truncated CsgF and FCP peptides used to generate CsgG:CsgF complex variants. Surprisingly, SDS-PAGE analysis of the thermal stability of CsgG:CsgF complexes prepared with this method with CsgF-(1-45) (Figure 16.A), CsgF-(1-35) (Figure 16.B) and CsgF-(1-30) (Figure 16.C) showed that at least CsgF-(1-45) and CsgF-(1-35) peptides generate complexes with CsgG that are thermally stable at least at 90°C. Since CsgG pores disassemble into their constituent monomers at 90°C, it is difficult to assess the stability of the complexes above 90°C. Since the difference between the band of CsgG pores and the band of CsgG:CsgF-(1-30) complexes is minimal in SDS-PAGE, this method is not sufficient to analyze the thermal stability of CsgG:CsgF-(1-30) complexes (Figure 16.C). However, in all three cases CsgG:CsgF complexes were observed, and even CsgG:CsgF-(1-29) was observed in electrophysiological experiments, indicating that even CsgF-(1-29) peptides generate at least some CsgG:CsgF complexes (Figure 24).
[0475] Example 2: Analysis of the CsgG:CsgF structure by cryo-EM
[0476] To gain structural insight into the CsgG:CsgF complex, co-purified or in vitro reconstituted CsgG:CsgF particles were analyzed by transmission electron microscopy. For preparation of cryo-EM analysis, 500 μΙ_ of peak fraction of bi- affinity purified CsgG:CsgF complex was injected onto a Superose 6 10 / 30 column equilibrated with buffer D (25 mM Tris pH 8, 200 mM NaCI and 0.03% DDM) and run at 0.5 mL / min. Protein concentration was determined based on calculated absorbance at 280 nm and assuming a stoichiometry of 1 : 1. Samples for cryo-electron microscopy examination were analyzed as described in the methods. Figure 4 shows cryo-EM micrographs of CsgG:CsgF complex and two selected class averages from picked CsgG:CsgF particles. Micrographs show the presence of both the nonamer pore and the dimer of nonamer pore complexes. For image reconstruction, nonamer CsgG:CsgF particles were picked and aligned using RELION. The class average of the CsgG:CsgF complex as a side view and the 3D reconstructed electron density show the presence of additional density corresponding to CsgF, which is seen as a protrusion from the CsgG particle located at the side of the CsgG β-barrel Figure 4B 、 Figure 5 ). The additional density reveals three distinct regions, including a globular head domain, a hollow neck domain and a domain interacting with the CsgG β-barrel. The latter CsgF region, termed CsgF constriction peptide or FCP, inserts into the lumen of the CsgG β-barrel and can be seen to form an additional constriction segment of the CsgG pore (labeled F in Figure 4B 、 Figure 5 ) that is located about 2 nm above the constriction segment formed by the CsgG constriction ring (labeled G in Figure 4B 、 Figure 5 ).
[0477] Example 3: Identification of CsgF interaction and constriction peptide by truncation of CsgF
[0478] The presence of a second constriction segment in the CsgG:CsgF nanopore complex offers opportunities for nanopore sensing applications, providing a second pore opening in the nanopore that can be used as a second read head or as an extension of the primary read head provided by the CsgG constriction ring (Figures 6, 7) in comparison to the pore of CsgG alone. However, the exit side of the CsgG:CsgF combined pore is blocked by the CsgF neck and head domains when complexed with full-length CsgF. We therefore sought to identify the region of CsgF that is required to interact with the CsgG β-barrel and insert into it. Our Strep-tactin affinity purification experiments suggested that the N-terminal region of CsgF is required for CsgG interaction, as N-terminal truncation fragments of CsgF that were present in the His-trap affinity purification were lost from co-purification with CsgG Figure 2B ). CsgF homologues are characterised by the presence of the PFAM domain PF03783. When a multiple sequence alignment (MSA) of CsgG homologues found in Gram-negative bacteria was performed Figure 8 An MSA of selected CsgF homologues is shown, revealing a region of sequence conservation corresponding to the first ~30-35 amino acids of mature CsgF (SEQ ID NO: 6) (pairwise sequence identity between 35 and 100%). Based on the combined data, it was hypothesised that this N-terminal region of CsgF forms a CsgG interaction peptide or FCP. Figure 10 A multiple sequence alignment of FCPs in known CsgF homologues is shown in Figure
[0479] To test the hypothesis that the CsgF N-terminus corresponds to a CsgG binding region and forms a CsgF constriction peptide that remains within the lumen of the CsgG β-barrel, Strep-tagged CsgG and His-tagged CsgF truncations were co-expressed in E. coli (see Methods). pNA97, pNA98, pNA99 and pNA100 encode N-terminal CsgF fragments corresponding to residues 1-27, 1-38, 1-48 and 1-64 of CsgF (SEQ ID NO: 5). These peptides include the CsgF signal peptide corresponding to residues 1-19 of SEQ ID NO: 5, and so would give rise to periplasmic peptides corresponding to the first 8, 19, 29 and 45 residues of mature CsgF (SEQ ID NO: 6; Figure 9A ) each containing a C-terminal 6x His tag. SDS-PAGE analysis of whole cell lysates revealed that CsgG was present in all samples, and that periplasmic peptides corresponding to the first 8, 19, 29 and 45 residues of mature CsgF (SEQ ID NO: 6; Figure 9Bcorresponding to the first 45 residues of CsgF. For the shorter N-terminal CsgF fragments, no detectable expression of the peptides was found in whole cell lysates. Cell pellets of various CsgG:CsgF fragments were further enriched by purification after two freeze / thaw cycles. Whole cell lysates as well as elution fractions of Strep-affinity purification were spotted onto nitrocellulose membranes and subjected to dot blot analysis using anti-His antibodies to detect His-tagged CsgF fragments Figure 9C ). Dot blotting revealed co-purification of the CsgF 20:64 peptide with CsgG, demonstrating that this CsgF fragment is sufficient to form a stable non-covalent complex with CsgG. For the CsgG 20:48 fragment, a small amount of peptide could be found co-purified with CsgG, while no detectable levels of either CsgF 20:27 or CsgF 20:38 were found in whole cell lysates or Strep-affinity purification Figure 9C ), indicating that the latter peptides are not stably expressed in E. coli and / or do not form stable complexes with CsgG.
[0480] Example 4: Atomic resolution description of the CsgG:CsgF interaction.
[0481] To obtain atomic-level details about the CsgG:CsgF interaction, we determined a high-resolution cryoEM structure of the CsgG:CsgF complex. For this, CsgG and CsgF were recombinantly expressed in E. coli and isolated from the outer membrane of E. coli by detergent extraction and purified using a tandem affinity purification approach. Samples for cryoEM examination were prepared by spotting 3 μΐ of sample on graphene-oxide coated R2 / 1 holey grids (Quantifoil) and data were collected in counting mode on a 300 kV TITAN Krios with a Gatan K2 direct electron detector. The final electron density map at a resolution of 3.0 A was calculated using 62.000 single CsgG:CsgF particles Figure 11 A ). This map allowed for unambiguous docking and local rebuilding of the CsgG crystal structure and de novo building of the N-terminal 35 residues of mature CsgF (i.e. residues 20:54 of Seq ID No. 5), which encompass the FCP Figure 11 C , D) that binds CsgG and forms a second constriction segment at the height of the CsgG transmembrane β-barrel. The cryoEM structure revealed a stoichiometry of 9:9 for CsgG:CsgF with C9 symmetry Figure 11 B). The FCP binds inside the CsgG β-barrel with the C-terminal end of CsgF pointing out of the CsgG β-barrel and the N-terminal end of CsgF located near the constriction of CsgG. This structure shows that P35 in mature CsgF is located outside the CsgG β-barrel and forms a link between the FCP of CsgF and the neck region. Due to the flexibility of the head and neck regions relative to the body of the CsgG:CsgF complex, the neck and head regions of CsgF could not be resolved in the high-resolution cryoEM map. Three regions in the CsgG β-barrel stabilize the CsgG:CsgF interaction: (IR1) residues Y130, D155, S183, N209, and T207 in mature CsgG (SEQ ID NO: 3) form an interaction network with the N-terminal amine and residues 1-4 of mature CsgF (SEQ ID NO: 6) including four hydrogen bonds and one electrostatic interaction; (IR2) residues Q187, D149, and E203 in mature CsgG (SEQ ID NO: 3) form an interaction network with R8 and N9 in mature CsgF (SEQ ID NO 6) encompassing three H bonds and two electrostatic interactions; and (IR3) residues F144, F191, F193, and L199 in mature CsgG (SEQ ID NO: 3) form a hydrophobic interaction surface with residues F21, L22, and A26 in mature CsgF (SEQ ID NO: 6). The latter is located in an alpha-helix (helix 1) formed by residues 19-30 of mature CsgF. The conserved sequence N-P-X-F-G-G (residues 9-14 in SEQ ID NO: 6) forms a loop region connecting residues 15-19 to the inwards turn of helix 1 of CsgF. These elements together create a constriction in the CsgG:CsgF complex with residue 17 (N17 in mature E. coli CsgF, i.e., SEQ ID NO: 6) forming the narrowest point, resulting in a pore opening of Figure 11 C ) approximately and above the top and bottom of the constriction formed by CsgG residues 46 to 59 (G constriction or GC), respectively.
[0482] Example 5: Simulation to improve stability of the CcgG-CsgF complex
[0483] Molecular dynamics simulations were performed to determine which residues in CsgG and CsgF are in close proximity. Using this information, CsgG and CsgF mutants were designed that could increase the stability of the complex.
[0484] Simulations were performed using the GROMACS software package version 4.6.5, with the GROMOS 53a6 force field and SPC water model. The cryo-EM structure of the CsgG-CsgF complex was used in the simulations. The complex was solvated and then the energy was minimized using the steepest descent algorithm. Throughout the simulation, the backbone of the complex was restrained, however the side chains of the residues were free to move. The system was simulated for 20 nanoseconds in the NPT ensemble at 300K using a Berendsen thermostat and Berendsen barostat.
[0485] Contacts between CsgG and CsgF were analyzed using the GROMACS analysis software and locally written code. Two residues were defined to have made contact if they were within 3 angstroms. The results are shown in Table 4 below.
[0486] Table 4: Predicted contact frequencies of residue pairs in the CsgG / CsgF complex:
[0487]
[0488]
[0489]
[0490] Materials and methods for CsgG:CsgF complex structure determination:
[0491] Cloning
[0492] To express E. coli CsgG as an outer membrane-localized pore, the coding sequence for E. coli CsgG (SEQ ID NO: 1) was cloned into pASK-Iba12, resulting in plasmid pPG1 (Goyal et al., 2013).
[0493] To express C-terminal 6x-His tagged CsgF in the E. coli cytoplasm, the coding sequence for mature E. coli CsgF (SEQ ID NO: 6; i.e. CsgF without its signal sequence) was cloned into pET22b via Ndel and EcoRI sites using a PCR product generated with primers "CsgF-His_pET22b_FW" (SEQ ID NO: 46) and "CsgF-His_pET22b_Rev" (SEQ ID NO: 47), resulting in CsgF-His expression plasmid pNA101.
[0494] Based on pGV5403 (integrated pDEST14 The pTrc99a-based cassette was used to generate the pNA62 plasmid, a pTrc99a-based vector expressing csgF-His and csgG-strep. The pGV5403 ampicillin resistance cassette was replaced with a streptomycin / spectinomycin resistance cassette. A PCR fragment covering a portion of the *E. coli* MC4100 csgDEFG operon corresponding to the coding sequences of csgE, csgF, and csgG was generated using primers csgEFG_pDONR221_FW (SEQ ID NO:48) and csgEFG_pDONR221_Rev (SEQ ID NO:49), and was then analyzed via BP. The recombinant was inserted into pDONR221 (ThermoFisherScientific). Next, it was processed via LR... The recombinant csgEFG operon from the pDONR221 donor plasmid was inserted into pGV5403 containing a streptomycin / spectinomycin resistance cassette. A 6xHis tag was added to the C-terminus of the CsgF operon by PCR using primers Mut_csgF_His_FW (SEQ ID NO:50) and Mut_csgF_His_Rev (SEQ ID NO:51). Finally, csgE was removed by outward PCR (primers DelCsgE_FW (SEQ ID NO:52) and DelCsgE_Rev (SEQ ID NO:53)) to obtain pNA62.
[0495] Outward PCR was performed on pNA62 (a pTrc99a-based vector expressing CsgF-his and CsgG-strep) to generate a CsgF fragment with a C-terminal His label corresponding to the putative contractile peptide for periplasmic expression. Figure 9A The constructs were prepared using the following primer combinations: pNa62_CsgF_his tag_Fw (SEQ ID NO:45) as the forward primer, and CsgF_d27_end (SEQ ID NO:41), CsgF_d38_end (SEQ ID NO:42), CsgF_d48_end (SEQ ID NO:43), or CsgF_d64_end (SEQ ID NO:44) as the reverse primers to generate pNA97, pNA98, pNA99, and pNA100, respectively.
[0496] In pNA97, csgF was truncated to SEQ ID NO: 7, which encodes a CsgF fragment comprising residues 1-27 (SEQ ID NO: 8); in pNA98, csgF was truncated to SEQ ID NO: 9, which encodes a CsgF fragment comprising residues 1-38 (SEQ ID NO: 10); in pNA99, csgF was truncated to SEQ ID NO: 11, which encodes a CsgF fragment comprising residues 1-48 (SEQ ID NO: 12); in pNA100, csgF was truncated to SEQ ID NO: 13, which encodes a CsgF fragment comprising residues 1-64 (SEQ ID NO: 14). Expression of pNA97, pNA98, pNA99 and pNA100 in E. coli indeed resulted in the production of CsgG pores (SEQ ID NO: 3) in the outer membrane, and CsgF-derived peptide with the following sequences targeted to the periplasm:
[0497] "GTMTFQFRNPNFGGNPNNGAFLLNSAQAQNSYKDPSYNDDFGIETHHHHHH" (SEQ ID NO: 40 + 6xHis), respectively.
[0498] Strains
[0499] For all cloning procedures, E. coli Top10 (F - mcrA Δ(mrr - hsdRMS - mcrBC) Φ80 lacZ ΔM15 ΔlacX74 recAl araD139 Δ(araleu) 7697 galU galK rpsL (StrR) endAl nupG) was used. For protein production, E. coli C43 (DE3) (F - ompT hsdSB (rB - mB - ) gal dcm (DE3)) and Top10 were used.
[0500] Production of recombinant CsgG:CsgF complex by co-expression
[0501] To co-express E. coli CsgF (SEQ ID NO: 5) and CsgG (SEQ ID NO: 2), the two recombinant genes, including their native Shine-Dalgarno sequences, were placed under the control of the inducible trc promoter in a pTrc99a-derived plasmid to form plasmid pNA62. CsgG and CsgF were overexpressed in E. coli C43(DE3) cells transformed with plasmid pNA62 and grown in Terrific Broth medium at 37°C. When the cell culture reached an optical density (OD) of 0.7 at 600 nm, recombinant protein expression was induced with 0.5 mM IPTG and allowed to grow for 15 hours at 28°C before harvesting by centrifugation at 5500g.
[0502] Production of recombinant CsgG:CsgF complex by in vitro reconstitution
[0503] Full-length E. coli CsgG (SEQ ID NO: 2) modified with a C-terminal StrepII tag was overexpressed in E. coli BL21(DE3) cells transformed with plasmid pPG1 (Goyal et al., 2013). Cells were grown in Terrific Broth medium at 37°C to an OD of 0.6 at 600 nm. Recombinant protein production was induced with 0.0002% anhydrotetracycline (Sigma) and cells were grown for an additional 16 hours at 25°C before harvesting by centrifugation at 5500g.
[0504] E. coli CsgF (SEQ ID NO: 6; i.e., lacking the CsgF signal sequence) fused at the C-terminus to a 6x His tag was overexpressed in the cytoplasm of E. coli BL21(DE3) cells transformed with plasmid pNA101. Cells were grown at 37°C to an OD of 600 nm before induction by 1 mM IPTG and allowed to express protein for 15 hours at 37°C before harvesting by centrifugation at 5500g.
[0505] Recombinant protein purification of CsgG:CsgF complex, CsgG and CsgF
[0506] E. coli cells transformed with pNA62 and co-expressing CsgG-Strep and CsgF-His were resuspended in 50 mM Tris-HCl pH 8.0, 200 mM NaCl, 1 mM EDTA, 5 mM MgCl2, 0.4 mM AEBSF, 1 pg / mL Leupeptin, 0.5 mg / mL DNase I and 0.1 mg / mL Lysozyme. Cells were disrupted using a TS Series Cell Disruptor (Constant Systems Ltd) at 20 kPsi and the lysed cell suspension was incubated with 1% n-dodecyl- -d-maltopyranoside (DDM; Inalco) for 30' to further lyse the cells and extract the outer membrane fraction. Next, the remaining cell debris and membranes were pelleted by ultracentrifugation at 100.000 g for 40'. The supernatant was loaded onto a 5 mL HisTrap chromatography column equilibrated in buffer A (25 mM Tris pH 8, 200 mM NaCl, 10 mM Imidazole, 10% sucrose and 0.06% DDM). The column was washed with >10 CV of 5% buffer B (25 mM Tris pH 8, 200 mM NaCl, 500 mM Imidazole, 10% sucrose and 0.06% DDM) ion buffer A and eluted with 60 mL above a gradient of 5-100% buffer B.
[0507] The eluate was diluted 2-fold and then loaded onto a 5 mL Strep-tactin chromatography column (IBA GmbH) equilibrated with buffer C (25 mM Tris pH 8, 200 mM NaCl, 10% sucrose and 0.06% DDM) overnight. The column was washed with >10 CV of buffer C and the protein was eluted by addition of 2.5 mM desthiobiotin. Next, 500 pL of the peak fraction of the bi- affinity purification complex was injected onto a Superose 610 / 30 (GE Healthcare) equilibrated with buffer D (25 mM Tris pH 8, 200 mM NaCl and 0.03% DDM) and run at 0.5 mL / min to prepare samples for electron microscopy. Protein concentration was determined based on calculated absorbance at 280 nm and assuming a stoichiometry of 1 / 1. Buffer D (25 mM Tris pH 8, 200 mM NaCl and 0.03% DDM)
[0508] CsgG-strep purification for in vitro reconstitution was identical to the protocol for CsgG:CsgF when sucrose was omitted from the buffers and the IMAC and size exclusion steps were bypassed.
[0509] CsgF-His purification for in vitro reconstitution was performed by resuspending cell pellets in 50 mM Tris-HCl pH 8.0, 200 mM NaCl, 1 mM EDTA, 5 mM MgCl2, 0.4 mM AEBSF, 1 pg / mL leupeptin, 0.5 mg / mL DNase I and 0.1 mg / mL lysozyme. Cells were disrupted using a TS Series Cell Disrupter (Constant Systems Ltd) at 20 kPsi and the lysed cell suspension was centrifuged at 10.000 g for 30 min to remove intact cells and cell debris. The supernatant was added to 5 mL Ni-IMAC-beads (Workbeads 40IDA, Bio-Works Technologies AB) equilibrated with buffer A (25 mM Tris pH8, 200 mM NaCl, 10 mM imidazole) and incubated for 1 h at 4°C. The Ni-NTA beads were pooled into a gravity flow column and washed with 100 mL of 5% buffer B (25 mM Tris pH8, 200 mM NaCl, 500 mM imidazole diluted in buffer A). Bound protein was eluted by stepwise increasing buffer B (10% steps of 5 mL each).
[0510] In vitro reconstitution of CsgG:CsgF complex
[0511] Purified CsgG and CsgF were pooled and used to reconstitute the complex in vitro. Therefore a molar ratio of 1 CsgG : 2 CsgF was mixed, filling the CsgG barrel with CsgF. Next, the reconstitution mixture was injected onto a Superose 6 10 / 30 column (GE Healthcare) equilibrated with buffer D (25 mM Tris pH8, 200 mM NaCl and 0.03% DDM) and run at 0.5 mL / min to prepare samples for electron microscopy (Figure 3). Protein concentration was determined based on calculated absorbance at 280 nm and assuming a stoichiometry of 1 / 1.
[0512] Structural analysis using electron microscopy
[0513] Sample behaviour of size exclusion fractions was probed using negative stain electron microscopy. Samples were stained with 1% uranyl formate and imaged using an in-house 120 kV JEM 1400 (JEOL) microscope fitted with a LaB6 filament. Samples for cryo-electron microscopy were prepared by spotting 2 pL of sample onto R2 / 1 continuous carbon (2 nm) coated grids (Quantifoil), hand-stained and plunge-frozen into liquid ethane using an in-house plunge device. Sample quality was screened on an in-house JEOL JEM 1400 before datasets were collected on a 200 kV TALOS ARCTICA (FEI) microscope fitted with a Falcon-3 direct electron detection camera. Images were motion corrected using MotionCor2.1 (Zheng et al., 2017), defocus values were determined using ctffind4 (Rohou and Grigorieff, 2015) and data was further analysed using a combination of RELION (Scheres, 2012) and EMAN2 (Ludtke, 2016). C9 symmetry was imposed on selected 2D class averages during 3D model generation and refinement, which were characterised by additional density for the head group.
[0514] For high resolution cryoEM analysis, CsgG:CsgF samples were prepared for cryo-electron microscopy by spotting 3 pL of sample onto R2 / 1 holey grids (Quantifoil) coated with graphene oxide (Sigma Aldrich), hand-stained and plunge-frozen into liquid ethane using a CP3 plunger (Gatan). Sample quality was screened on an in-house JEOL JEM 1400 before datasets were collected on a 300 kV TITAN KRIOS (FEI, Thermo-Scientific) microscope fitted with a K2 Summit direct electron detector (Gatan). The detector was used in counting mode, collecting 50 frames per second for 20 seconds per image. The cumulative electron dose of the distribution was 56 electrons. 2045 images were collected with a pixel size of 0.65 A and a defocus of 2.5 pm. Images were obtained. Motion correction was performed on the images using MotionCor 2.1 (Zheng et al., 2017), and defocus values were determined using ctffind4 (Rohou and Grigorieff, 2015). Particles were automatically selected using Gautomatch (Dr. Kai Zhang), and the data were further analyzed using a combination of RELION 2.0 (Kimanius et al., 2016, Elife5.pii:e18722) and EMAN2 (Ludtke, 2016). During 3D model generation and refinement, C9 symmetry was applied to the average of selected 2D classes, characterized by an additional density for the head group corresponding to CsgF. 62,000 particles were used to... The final resolution image was calculated. A de novo model of CsgF was constructed using COOT (Brown et al. 2015 Acta Crystallogr D Biol Crystallogr 71(Pt1):136-53), and a complete complex model was constructed and refined iteratively using PHENIX (Afonine 2018, Acta Crystallogr D Struct Biol 74(Pt 6):531-544) for real-space refinement, combined with COOT.
[0515] Protein expression and purification of CsgG:CsgF fragments
[0516] The CsgF and CsgG fragments were co-expressed, with the CsgF fragment having a His tag at its C-terminus and the CsgG fragment fused to a Strep tag at its C-terminus. The CsgG:CsgF fragment complex was overexpressed in *E. coli* Top10 cells transformed with plasmids pNA97, pNA98, pNA99, or pNA100. Plates were incubated at 37°C ON, and colonies were resuspended in LB medium supplemented with streptomycin / spectinomycin. When the cell culture reached an optical density (OD) of 0.7 at 600 nm, recombinant protein expression was induced with 0.5 mM IPTG and incubated at 28°C for 15 h, followed by harvesting by centrifugation at 5500 g.
[0517] Freeze the precipitate at -20°C
[0518] Cell pellets used for co-expression of various CsgG:CsgF fragments were resuspended in 200 mL of 50 mM Tris-HCl pH 8.0, 200 mM NaCl, 1 mM EDTA, 5 mM MgCl2, 0.4 mM AEBSF, 1 pg / mL leupeptin, 0.5 mg / mL DNase I and 0.1 mg / mL lysozyme, sonicated and incubated with 1% n-dodecyl-beta-d-maltopyranoside (DDM; Inalco) to further lyse the cells and extract the outer membrane fraction. Next, the remaining cell debris and membranes were pelleted by centrifugation at 15.000 g for 40'. The supernatant was incubated with 100 pL Strep-tactin beads for 30 min at room temperature. The Strep beads were washed by centrifugation with buffer (25 mM Tris pH 8, 200 mM NaCl and 1% DDM) and the bound proteins were eluted by adding 2.5 mM desthiobiotin in 25 mM Tris pH 8, 200 mM NaCl, 0.01% DDM.
[0519] Production of CsgG:FCP by in vitro reconstitution.
[0520] A synthetic peptide corresponding to the 34 residues of the N-terminus of mature CsgF (SEQ ID NO: 6) was diluted at 1 mg / ml in buffer 0.1 M MES, 0.5 M NaCl, 0.4 mg / ml EDC (1-ethyl-3-(3-dimethylaminopropyl) carbodiimide), 0.6 mg / ml NHS (N-hydroxysuccinimide) and incubated for 15 min at room temperature to allow the activation of the carboxylic terminus of the peptide. Next, 1 mg / ml of a PBS solution of Cadaverin-Alexa594 was added during a 2h incubation to allow the covalent coupling at room temperature. The reaction was quenched by buffer exchange to 50 mM Tris, NaCl, 1 mM EDTA, 0.1% DDM using Zeba Spin filters.
[0521] The labeled peptide was added to the strep-affinity purified CsgG in 50 mM Tris, 100 mM NaCl, 1 mM EDTA, 5 mM LDAO / C8D4 at a 2:1 molar ratio in 15 min at room temperature to allow the reconstitution of the CsgG:FCP complex. After pull-down of CsgG-strep onto StrepTactin beads, the sample was analyzed on native PAGE.
[0522] Example 6: Further stabilization of CsgG:CsgF complex by covalent cross-linking
[0523] Although the full-length and some truncated forms of CsgF form stable CsgG:CsgF complexes with CsgG pores, CsgF can still be forcibly removed from the barrel region of the CsgG pore under certain conditions. Therefore, covalent linkages between the CsgG and CsgF subunits are desirable. Based on molecular simulation studies, positions of CsgG and CsgF that are close to each other have been identified (Example 5 and Table 4). Some of these identified positions have been modified to incorporate cysteine residues in both CsgG and CsgF. Figure 19 shows an example of a thiol-thiol bond formed between the Q153 position of CsgG and the G1 position of CsgF. CsgG pores containing the Q153C mutation were reconstructed with CsgF containing the G1C mutation and incubated for 1 hour to enable the formation of SS bonds. Heating the complex to 100°C in the absence of DTT revealed a 45 kDa band corresponding to the dimer between the CsgG and CsgF monomers (CsgGm-CsgFm), indicating the formation of an SS bond between the two monomers (30 kDa for CsgGm and 15 kDa for CsgFm) (Fig. 19.A). This band disappeared upon heating in the presence of DTT, indicating that DTT breaks down the SS bond. Incubating the CsgG:CsgF complex overnight instead of for 1 hour increased the extent of CsgGm-CsgFm dimer formation (Fig. 19.A). Mass spectrometry was performed to further identify the dimer band. The gel-purified protein was proteolytically cleaved to yield trypsin peptides. LC-MS / MS sequencing was performed to identify the SS bond between the Q153 position of CsgG and the G1 position of CsgF (Fig. 19.B). Oxidizing agents such as copper-1,0-phenanthroline can be used to enhance SS bond formation. As described in the Methods section, in the presence of copper-o-phenanthroline, the pores of CsgG modified with N133C and CsgF modified with T4C were reconstructed. Then, by heating to 100°C in the absence of DTT to decompose them into their constituent monomers, strong dimer bands corresponding to CsgGm-CsgFm could be observed on SDS-PAGE. Figure 20 Lanes 3 and 4). When heated in the presence of DTT, the dimer decomposes into its constituent monomers ( Figure 20 (lanes 1 and 2).
[0524] Example 7: Electrophysiological characterization of CsgG:CsgF complex
[0525] When the pores are inserted into the copolymer membrane and experiments are performed using the MinION from Oxford Nanopore Technologies, the signal observed when the DNA strand translocates through CsgG can be well characterized (Figure 31). The Y51, N55, and F56 of each CsgG subunit form the contractile segment of the CsgG pore (…). Figure 12). This acute constriction serves as a readhead for the CsgG pore Figure 31A ), and is able to accurately distinguish mixed sequences of A, C, G, and T as they pass through the pore. This is because the measured signal contains characteristic current deflections from which the identity of the sequence can be derived. However, in homopolymer regions of DNA, the measured signal can not show a current deflection of sufficient amplitude to allow single base identification; making it impossible to accurately determine the length of the homopolymer from the amplitude of the measured signal Figure 26B and Figure 26C ). The reduction in accuracy of the CsgG readhead is related to the length of the homopolymer region Figure 29C ). When CsgF interacts with the CsgG pore to form a CsgG:CsgF complex, CsgF introduces a second readhead in the CsgG barrel. The second readhead is primarily composed of the N17 position of Seq. ID No. 6. Static strand experiments were performed as described in the methods section and in Figure 27 to experimentally map the two readheads of the CsgG:CsgF complex, and the results indicate that there are two readheads spaced approximately 5-6 bases apart from each other Figure 27B , Figure 27C and Figure 27D ). The readhead discrimination plot for the CsgG:CsgF complex shows that the contribution of the second readhead introduced by CsgF to base discrimination is less than the CsgG readhead Figure 27A ). Surprisingly, when the second readhead is introduced into the CsgG barrel by CsgF, previously flat homopolymer regions show step signals Figure 30B and Figure 30C ). The information contained in these steps can be used to accurately identify the sequence, reducing errors. The accuracy of the DNA signal of the CsgG:CsgF complex remains relatively constant over longer homopolymer lengths compared to the accuracy profile of the CsgG pore itself Figure 29C
[0526] The CsgG:CsgF complexes prepared in any of the methods described in the methods section can be used to characterize the complexes in DNA sequencing experiments. The signals of lambda DNA strands through various CsgG:CsgF complexes, which are prepared by different methods, consisting of different CsgG mutant pores and different CsgF peptides of different lengths, are shown in FIGS. 21-24. The read head discrimination of those pore complexes and their base contribution curves are shown in FIG. 28 (A-H). Surprisingly, different modifications at the constriction segment of both the CsgG pore and the CsgF peptide can significantly change the signal of the CsgG:CsgF pore complex. For example, when the CsgG:CsgF complexes are prepared with the same CsgG pore, but with two different CsgF peptides of the same length containing either Asn or Ser at position 17 (Seq ID No. 6), which are prepared by the same method of co-expressing full-length CsgF protein, then cleaving the CsgF with TEV protease between positions 35 and 36, the resulting signals are different from each other (FIG. 21). The CsgG:CsgF complex with Ser at position 17 of the CsgF peptide shows lower noise and higher signal-to-noise ratio compared to the CsgG:CsgF complex with Asn at position 17 of the CsgF peptide. Similarly, when the same CsgG pore is reconstituted with two different CsgF peptides of the same length (1-35 of Seq ID No. 6), but with Ser or Val at position 17 to prepare CsgG:CsgF complexes, the complex with Val at position 17 of CsgF shows a noisier signal than the complex with Ser at position 17 of CsgF (FIG. 22). When the same length of the same CsgF peptide is reconstituted with different CsgG pores containing different mutations at the CsgG read head (positions 51, 55, and 56), the resulting CsgG:CsgF complexes show very different signals Figure 23A (F) with different signal-to-noise ratios Figure 25 Surprisingly, when different lengths of CsgF peptides containing the same constriction region are reconstituted with the same CsgG pore to prepare CsgG:CsgF complexes, they produce a different range of signals (FIG. 24). The CsgG:CsgF complex containing the shortest CsgF peptide (1-29 of Seq ID No. 6) shows the largest range, and the CsgG:CsgF complex containing the longest CsgF peptide (1-45 of Seq ID No. 6) shows the smallest range (FIG. 24).
[0527] Materials and methods for analyte characterization:
[0528] The proteins produced by the methods described below can be used interchangeably with the proteins produced by the methods described above with respect to structure determination.
[0529] Methods
[0530] Expression of CsgG:CsgF or CsgG:FCP complex by co-expression
[0531] Genes encoding CsgG proteins and mutants thereof were constructed in pT7 vectors containing an ampicillin resistance gene. Genes encoding CsgF or FCP proteins and mutants thereof were constructed in pRham vectors containing a kanamycin resistance gene. 1 uL of both plasmids were mixed with 50 uL Lemo(DE3)ACsgEFG on ice for 10 minutes. The samples were then heated at 42°C for 45 seconds before being placed back on ice for 5 minutes. 150 uL of NEB SOC growth media was added and the samples were incubated at 37°C with 250 rpm shaking for 1 hour. The full volume was plated onto agar plates containing kanamycin (40 ug / mL), ampicillin (100 ug / mL), and chloramphenicol (34 ug / ml) and incubated at 37°C overnight. Individual colonies were removed from the plates and inoculated into 100 mL of LB media containing kanamycin (40 ug / mL), ampicillin (100 ug / mL), and chloramphenicol (34 ug / mL) and incubated at 37°C with 250 rpm shaking overnight. 25 mL of starter culture was added to 500 mL of LB media containing 3 mM ATP, 15 mM MgS04, kanamycin (40 ug / mL), ampicillin (100 ug / mL), and chloramphenicol (34 ug / ml) and incubated at 37°C overnight. The culture was grown for 7 hours at which point the OD 600 Lactose (final concentration of 1.0%), glucose (final concentration of 0.2%), and rhamnose (final concentration of 2 mM) were added and the temperature was reduced to 18°C while maintaining 250 rpm shaking for 16 hours. The culture was centrifuged at 6000 rpm for 20 minutes at 4°C. The supernatant was discarded and the pellet was retained. The cells were stored at -80°C until purification.
[0532] Expression of CsgG pore with or without C-terminal Strep tag and CsgF with or without C-terminal Strep or His tag
[0533] All genes encoding all CsgG proteins and CsgF or FCP proteins were constructed in pT7 vectors containing an ampicillin resistance gene. The expression procedure was the same as above except that kanamycin was omitted from all media and buffers.
[0534] Cell lysis (complexes co-expressed or individual CsgG / CsgF / FCP proteins)
[0535] Lysis buffer was made from 50 mM Tris (pH 8.0), 150 mM NaCl, 0.1% DDM, lx Bugbuster protein extraction reagent (Merck), 2.5 uL Benzonase nuclease (stock > 250 units / uL) per 100 mL lysis buffer, and 1 tablet of Sigma Protease Inhibitor Cocktail per 100 mL lysis buffer. A 5X volume of lysis buffer was used to lyse IX weight of harvested cells. Cells were resuspended and centrifuged for 4 hours at room temperature until a homogenous lysate was produced. The lysate was centrifuged at 20,000 rpm for 35 minutes at 4°C. The supernatant was carefully extracted and filtered through a 0.2 uM Acrodisc syringe filter.
[0536] Strep purification of CsgG or CsgF / FCP proteins or co-expression complexes when CsgG contains a C-terminal Strep tag and CsgF or FCP contains a C-terminal His tag
[0537] The filtered sample was then loaded onto a 5 mL StrepTrap chromatography column with the following parameters: loading speed: 0.8 mL / min, total sample loading volume: 10 mL, wash unbound: 10 CV (5 mL / min), extra wash: 10 CV (5 mL / min), elution: 3 CV (5 mL / min). Affinity buffer: 50 mL Tris (pH 8.0), 150 mM NaCl, 0.1% DDM; wash buffer: 50 mL Tris (pH 8.0), 2 M NaCl, 0.1% DDM; elution buffer: 50 mL Tris (pH 8.0), 150 mM NaCl, 0.1% DDM, 10 mM desthiobiotin. Eluted samples were collected.
[0538] His purification of CsgG or CsgF / FCP proteins or co-expression complexes when CsgG contains a C-terminal Strep tag and CsgF or FCP contains a C-terminal His tag
[0539] The filtered sample from Strep purification (in the case of complexes) or pooled elution peaks were loaded onto a 5 mL HisTrap chromatography column using the same parameters as above, except using the following buffers: affinity and wash buffers: 50 mL Tris (pH 8.0), 150 mM NaCl, 0.1% DDM, 25 mM imidazole; elution: 50 mL Tris (pH 8.0), 150 mM NaCl, 0.1% DDM, 350 mM imidazole. Peaks were eluted and concentrated to a 500 uL volume in a 30 kDa MWCO Merck Milipore centrifugal device.
[0540] Form complexes in vitro with in vivo purified components.
[0541] Mix individually expressed and purified CsgG and CsgF / FCP proteins at different ratios to identify the correct ratio. But always in the presence of excess CsgF. Then incubate the complex overnight at 25°C. To remove excess CsgF and to remove DTT from the buffer, inject the mixture again on a Superdex Increase 200 10 / 300 equilibrated in 50 mM Tris (pH 8.0), 150 mM NaCl, 0.1% DDM. The complex usually elutes at 9 to 10 mL on this column.
[0542] Polish step for complexes (co-expressed or prepared in vitro) by gel filtration
[0543] If necessary, Strep-purified or His-purified or first His-purified then Strep-purified CsgG:CsgF or CsgG:FCP can be further polished by gel filtration. Inject 500 uL of sample into a 1 mL sample loop and inject on a Superdex Increase 200 10 / 300 equilibrated in 50 mM Tris (pH 8.0), 150 mM NaCl, 0.1% DDM. The peak associated with the complex usually elutes at 9 to 10 mL on this column when run at 1 mL / minute. Heat the sample at 60°C for 15 minutes and centrifuge at 21,000 rcf for 10 minutes. Take the supernatant for the assay. Perform SDS-PAGE on the sample to confirm and identify the fractions eluted with the complex.
[0544] Liberate CsgF or FCP at TEV protease site
[0545] If CsgF or FCP contains a TEV cleavage site, add TEV protease with a C-terminal histidine tag to the sample containing 2 mM DTT (the amount added is determined according to the approximate concentration of the protein complex). Incubate the sample on a roller mixer at 25 rpm overnight at 4°C. Then run the mixture again through a 5 mL HisTrap column and collect the flow-through. Any un-cleaved material will remain bound to the column, and the cleaved protein will elute. Use the same buffers and parameters as for the His purification described above and a final heat step.
[0546] Purification of CsgG:FCP complex with in vivo purified CsgG pore and synthetic FCP
[0547] Lyophilized FCP peptides were received from Genscript and Lifetein. 1 mg of peptide was dissolved in 1 mL nuclease-free ddH2O to obtain a 1 mg / mL sample. The sample was vortexed until no peptide was visible. Due to differences in expression levels of CsgG pores and mutants, it is difficult to accurately measure concentrations. Protein band intensities on SDS-PAGE against known markers can be used to obtain a rough estimate of the sample. CsgG and FCP were then mixed in a molar ratio of approximately 1 :50 and incubated overnight at 25 °C with 700 rpm. The sample was heated at 60 °C for 15 minutes and centrifuged at 21,000 ref for 10 minutes. The supernatant was taken for the assay. If needed, the complex can be purified as described above in co-expression.
[0548] Purification of CsgG:CsgF or CsgG:FCP containing cysteine mutants
[0549] If either or both components contain cysteine, CsgG:CsgF or CsgG:FCP complexes can be purified using the same procedure as described above (I or II or III below), except for the composition of the affinity buffers, wash buffers and elution buffers in His and Strep purification and the buffer used for gel filtration. For purification of cysteine mutants, all of these buffers should contain 2 mM DTT. When synthetic peptides containing cysteine are dissolved in ddH2O, 2 mM DTT is also added
[0550] I. Co-expression of CsgG and CsgF or FCP
[0551] II. Preparation of CsgG:CsgF or CsgG:FCP complexes in vitro with individual components purified in vivo III. Preparation of CsgG:CsgF or CsgG:FCP complexes in vitro with CsgG and synthetic FCP purified in vivo
[0552] Determination of cysteine bond formation
[0553] 50 uL of the final eluate was taken from each of the two tubes. In one of the tubes, 2 mM DTT was added as a reducing agent and in the other tube, 100 mM Cu(II): 1-10 o-phenanthroline (33 mM: 100 mM) was added as an oxidizing agent. The samples were mixed 1 : 1 with Laemmli buffer containing 4% SDS. Half of the samples were heat treated for 10 minutes to 100 degrees and half of the samples were not treated, after which they were run on a 4-20% TGX gel (Bio-rad Criterion) in TGS buffer.
[0554] Coupled in vitro transcription and translation (IVTT)
[0555] All proteins were generated by coupled in vitro transcription and translation (IVTT) using the E. coli T7-S30 extract system (Promega) with circular DNA. Equal volumes of complete 1 mM amino acid mix minus cysteine and complete 1 mM amino acid mix minus methionine were mixed to obtain the working amino acid solution required to produce high concentrations of protein. Amino acids (10 uL) were mixed with premix (40 uL), [35S] L-methionine (2 uL, 1175 Ci / mmol, 10 mCi / mL), plasmid DNA (16 uL, 400 ng / uL) and T7 S30 extract (30 uL) and rifampicin (2 uL, 20 mg / mL) to produce 100 uL of IVTT protein reaction. Synthesis was carried out at 30 °C for 4 hours, then incubated overnight at room temperature. If CsgG:CsgF or CsgG:FCP complexes were prepared in co-expression, plasmid DNA encoding each component was mixed in equal amounts and a portion of the mixture (16 uL) was used for IVTT. Following incubation, tubes were centrifuged at 22000 g for 10 minutes and the supernatant was discarded from the tubes. The resulting pellet was resuspended and washed in MBSA (10 mM MOPS, 1 mg / ml BSA pH 7.4) and centrifuged again under the same conditions. The protein present in the pellet was resuspended in IX Laemmli sample buffer and run in a 4-20% TGX gel at 300 V for 25 minutes. The gel was then dried and exposed to MR film overnight. The film was then processed and the proteins in the gel visualised.
[0556] Samples for testing in MinIONs
[0557] Prior to testing, all samples were incubated with Brij58 (final concentration 0.1%) for 10 minutes at room temperature, then made up to the required subsequent well dilutions for the pore insert.
[0558] Method of preparing and running static strands
[0559] A set of polyA DNA strands (SS20 to SS38 of Figure 27) were obtained through Integrated DNA Technologies (IDT) in which the DNA backbone (iSpc3) was missing one base. The 3' end of each of these strands also contained a biotin modification. Incubation of the static strands with monomeric streptavidin for 20 minutes at room temperature resulted in the binding of biotin to streptavidin. The streptavidin-static strand complex was diluted in 25mM HEPES, 430mM KCl, 30mM ATP, 30mM MgCl2, 2.15mM EDTA (pH8) (referred to as RBFM) to 500nM (B, Figure 27) and 2uM (C, Figure 27). The residual current generated by each static strand was recorded in a MinION device. The MinIOn flow cell was flushed according to the standard run protocol and the sequencing procedure was then initiated with a 1 minute static flick. An open pore recording was initially generated for 10 minutes before the addition of 150uL of the first streptavidin-static strand complex. After 10 minutes, 800uL of RBFM was flushed through the flow cell before the addition of the next streptavidin-static strand complex. This process was repeated for all streptavidin-static strand complexes. Once the final streptavidin-static strand complex was incubated on the flow cell, 800uL of RBFM was flushed through the flow cell and an open pore recording was generated for 10 minutes before the completion of the experiment.
[0560] Method of making a discrimination plot
[0561] A read head discrimination plot shows the average change in the simulated current when the base at each read head position is changed. To calculate the read head discrimination of a model of length k and alphabet size n at position i, we define the discrimination at read head position i as the median of the standard deviation of the current levels of each of the n k-1 groups of size n where position i is changed while the other positions remain unchanged.
[0562] Aspects of the disclosure
[0563] 1. An isolated pore complex comprising a CsgG pore or homologue or mutant thereof, and a modified CsgF peptide or homologue or mutant thereof.
[0564] 2. The isolated pore complex of 1, wherein the modified CsgF peptide or homologue or mutant thereof is inserted into the lumen of the CsgG pore or homologue or mutant thereof.
[0565] 3. The isolated pore complex of 2, wherein the pore complex has two or more channels comprising a CsgG channel constriction segment and a CsgF channel constriction segment.
[0566] 4. The isolated pore complex of any one of 1 to 3, wherein the CsgG pore or homolog or mutant thereof is a mutant CsgG pore.
[0567] 5. The isolated pore complex of 3 or 4, wherein the diameter of the CsgF channel constriction is in the range of 0.5 nm to 2.0 nm.
[0568] 6. A modified CsgF peptide or modified peptide of a CsgF homolog or mutant, wherein the modification comprises a truncation of the CsgF protein SEQ ID NO: 6 or a homolog or mutant thereof.
[0569] 7. The modified CsgF peptide or modified peptide of a CsgF homolog or mutant of 6, wherein the modified CsgF peptide comprises SEQ ID NO: 39 or SEQ ID NO: 40, or a homolog or mutant thereof.
[0570] 8. The modified CsgF peptide or modified peptide of a CsgF homolog or mutant of 6, wherein the modified CsgF peptide comprises SEQ ID NO: 15, or a homolog or mutant thereof.
[0571] 9. The modified CsgF peptide of 8, or modified peptide of a CsgF homolog or mutant, wherein one or more positions in the region comprising SEQ ID NO: 15 are mutated, having at least 35% amino acid identity to SEQ ID NO: 15.
[0572] 10. A polynucleotide encoding the modified CsgF peptide of any one of 6 to 9.
[0573] 11. The isolated pore complex of any one of 1 to 5, wherein the modified CsgF peptide or homolog or mutant thereof is the peptide of any one of 6 to 9.
[0574] 12. The isolated pore complex of 11, wherein the modified CsgF peptide and the CsgG pore or homolog or mutant thereof are covalently coupled.
[0575] 13. The isolated pore complex of 12, wherein the covalent coupling is by means of:
[0576] (i) a cysteine residue at a position corresponding to 132, 133, 136, 138, 140, 142, 144, 145, 147, 149, 151, 153, 155, 183, 185, 187, 189, 191, 201, 203, 205, 207, or 209 of SEQ ID NO: 3 or a homolog thereof;
[0577] (ii) a non-native reactive or photoreactive amino acid at a position corresponding to 132, 133, 136, 138, 140, 142, 144, 145, 147, 149, 151, 153, 155, 183, 185, 187, 189, 191, 201, 203, 205, 207, or 209 of SEQ ID NO: 3 or a homolog thereof.
[0578] 14. An isolated transmembrane pore complex comprising the isolated pore complex according to any one of 1 to 5 or 11 to 13, and a component of a membrane.
[0579] 15. A method for producing a transmembrane pore complex, wherein the pore complex is formed by a CsgG pore and a modified CsgF peptide or a homolog or mutant thereof, the method comprising co-expressing CsgG or a homolog or mutant thereof as set forth in SEQ ID NO: 2, and a modified CsgF peptide or a homolog or mutant thereof in a suitable host cell, thereby allowing the formation of a transmembrane pore complex in vivo.
[0580] 16. The method according to 15, wherein the modified CsgF peptide or a homolog or mutant thereof comprises SEQ ID NO: 12 or SEQ ID NO: 14 or a homolog or mutant thereof.
[0581] 17. A method for producing an isolated pore complex, wherein the isolated pore is formed by a CsgG pore or a homolog or mutant thereof, and a modified CsgF peptide or a homolog or mutant thereof, the method comprising contacting a CsgG monomer of SEQ ID NO: 3 or a homolog or mutant thereof with a modified CsgF peptide or a homolog or mutant thereof, thereby allowing the reconstitution of the isolated pore complex in vitro.
[0582] 18. The method according to 17, wherein the modified CsgF peptide or a homolog or mutant thereof comprises SEQ ID NO: 15 or SEQ ID NO: 16 or a homolog or mutant thereof.
[0583] 19. A method for determining the presence, absence or one or more characteristics of a target analyte, comprising the steps of:
[0584] (i) contacting the target analyte with the pore complex according to any one of 1 to 5 or 11 to 13 or with the transmembrane pore complex according to 14, such that the target analyte moves into the pore complex; and
[0585] (ii) making one or more measurements as the analyte moves through the pore complex, thereby determining the presence, absence or one or more characteristics of the analyte.
[0586] 20. The method of 19, wherein the analyte is a polynucleotide.
[0587] 21. The method of 19, wherein the analyte is a (poly)peptide.
[0588] 22. The method of 19, wherein the analyte is a polysaccharide.
[0589] 23. The method of 19, wherein the analyte is a small organic or inorganic compound, such as a pharmacologically active compound, a toxic compound, and a pollutant.
[0590] 24. The method of 20, comprising determining one or more features selected from the group consisting of (i) the length of the polynucleotide, (ii) the identity of the polynucleotide, (iii) the sequence of the polynucleotide, (iv) the secondary structure of the polynucleotide, and (v) whether the polynucleotide is modified.
[0591] 25. A method of characterizing a polynucleotide or a (poly)peptide using an isolated transmembrane pore complex, wherein the pore complex is a complex comprising a CsgG pore or a homolog or mutant thereof and a modified CsgF peptide or a homolog or mutant thereof.
[0592] 26. The method of 25, wherein the CsgG pore or a homolog or mutant thereof comprises six to ten monomers.
[0593] 27. Use of an isolated pore complex according to any one of 1 to 5 or 11 to 13 or a transmembrane pore complex according to 14 for determining the presence, absence, or one or more features of a target analyte.
[0594] 28. A kit for characterizing a target analyte comprising (a) an isolated pore complex according to any one of 1 to 5 or 11 to 13 and (b) components of a membrane.
[0595] Sequences
[0596] SEQUENCE DESCRIPTION:
[0597] SEQ ID NO: 1 shows the polynucleotide sequence of wild-type E. coli CsgG from strain K12 including the signal sequence (Gene ID: 945619).
[0598] SEQ ID NO: 2 shows the amino acid sequence of wild-type E. coli CsgG including the signal sequence (Uniprot accession number P0AEA2).
[0599] SEQ ID NO: 3 shows the amino acid sequence of wild-type E. coli CsgG as a mature protein (Uniprot accession number P0EAEA2).
[0600] SEQ ID NO: 4 shows the polynucleotide sequence of wild-type E. coli CsgF from strain K12 including the signal sequence (Gene ID: 945622).
[0601] SEQ ID NO: 5 shows the amino acid sequence of wild-type E. coli CsgF including the signal sequence (Uniprot accession number P0AE98).
[0602] SEQ ID NO: 6 shows the amino acid sequence of wild-type E. coli CsgF as a mature protein (Uniprot accession number P0AE98).
[0603] SEQ ID NO: 7 shows the polynucleotide sequence of a wild-type E. coli CsgF fragment encoding amino acids 1 to 27 and a C-terminal 6 His tag.
[0604] SEQ ID NO: 8 shows the amino acid sequence of a wild-type E. coli CsgF fragment encompassing amino acids 1 to 27 and a C-terminal 6 His tag.
[0605] SEQ ID NO: 9 shows the polynucleotide sequence of a wild-type E. coli CsgF fragment encoding amino acids 1 to 38 and a C-terminal 6 His tag.
[0606] SEQ ID NO: 10 shows the amino acid sequence of a wild-type E. coli CsgF fragment encompassing amino acids 1 to 38 and a C-terminal 6 His tag.
[0607] SEQ ID NO: 11 shows the polynucleotide sequence of a wild-type E. coli CsgF fragment encoding amino acids 1 to 48 and a C-terminal 6 His tag.
[0608] SEQ ID NO: 12 shows the amino acid sequence of a wild-type E. coli CsgF fragment encompassing amino acids 1 to 48 and a C-terminal 6 His tag.
[0609] SEQ ID NO: 13 shows the polynucleotide sequence of a wild-type E. coli CsgF fragment encoding amino acids 1 to 64 and a C-terminal 6 His tag.
[0610] SEQ ID NO: 14 shows the amino acid sequence of a wild-type E. coli CsgF fragment encompassing amino acids 1 to 64 and a C-terminal 6 His tag.
[0611] SEQ ID NO: 15 shows the amino acid sequence of a peptide corresponding to residues 20 to 53 of E. coli CsgF
[0612] SEQ ID NO: 16 shows the amino acid sequence of a peptide corresponding to residues 20 to 42 of E. coli CsgF including the KD at its C-terminus
[0613] SEQ ID NO: 17 shows the amino acid sequence of a peptide corresponding to residues 23 to 55 of CsgF homolog Q88H88
[0614] SEQ ID NO: 18 shows the amino acid sequence of a peptide corresponding to residues 25 to 57 of CsgF homolog A0A143HJA0
[0615] SEQ ID NO: 19 shows the amino acid sequence of a peptide corresponding to residues 21 to 53 of CsgF homolog Q5E245
[0616] SEQ ID NO: 20 shows the amino acid sequence of a peptide corresponding to residues 19 to 51 of CsgF homolog Q084E5
[0617] SEQ ID NO: 21 shows the amino acid sequence of a peptide corresponding to residues 15 to 47 of CsgF homolog F0LZU2
[0618] SEQ ID NO: 22 shows the amino acid sequence of a peptide corresponding to residues 26 to 58 of CsgF homolog A0A136HQR0
[0619] SEQ ID NO: 23 shows the amino acid sequence of a peptide corresponding to residues 21 to 53 of CsgF homolog A0A0W1SRL3
[0620] SEQ ID NO: 24 shows the amino acid sequence of a peptide corresponding to residues 26 to 59 of CsgF homolog B0UH01
[0621] SEQ ID NO: 25 shows the amino acid sequence of a peptide corresponding to residues 22 to 53 of CsgF homolog Q6NAU5
[0622] SEQ ID NO: 26 shows the amino acid sequence of a peptide corresponding to residues 7 to 38 of CsgF homolog G8PUY5
[0623] SEQ ID NO: 27 shows the amino acid sequence of a peptide corresponding to residues 25 to 57 of CsgF homolog A0A0S2ETP7
[0624] SEQ ID NO:28 shows the amino acid sequence of a peptide corresponding to residues 19 to 51 of CsgF homolog E3I1Z1
[0625] SEQ ID NO:29 shows the amino acid sequence of a peptide corresponding to residues 24 to 55 of CsgF homolog F3Z094
[0626] SEQ ID NO:30 shows the amino acid sequence of a peptide corresponding to residues 21 to 53 of CsgF homolog A0A176T7M2
[0627] SEQ ID NO:31 shows the amino acid sequence of a peptide corresponding to residues 14 to 45 of CsgF homolog D2QPP8
[0628] SEQ ID NO:32 shows the amino acid sequence of a peptide corresponding to residues 28 to 58 of CsgF homolog N2IYT1
[0629] SEQ ID NO:33 shows the amino acid sequence of a peptide corresponding to residues 26 to 58 of CsgF homolog W7QHV5
[0630] SEQ ID NO:34 shows the amino acid sequence of a peptide corresponding to residues 23 to 55 of CsgF homolog D4ZLW2
[0631] SEQ ID NO:35 shows the amino acid sequence of a peptide corresponding to residues 21 to 53 of CsgF homolog D2QT92
[0632] SEQ ID NO:36 shows the amino acid sequence of a peptide corresponding to residues 20 to 51 of CsgF homolog A0A167UJA2
[0633] SEQ ID NO:37 shows the amino acid sequence of a wild-type E. coli CsgF fragment encompassing amino acids 20 to 27.
[0634] SEQ ID NO:38 shows the amino acid sequence of a wild-type E. coli CsgF fragment encompassing amino acids 20 to 38.
[0635] SEQ ID NO:39 shows the amino acid sequence of a wild-type E. coli CsgF fragment encompassing amino acids 20 to 48.
[0636] SEQ ID NO:40 shows the amino acid sequence of a wild-type E. coli CsgF fragment encompassing amino acids 20 to 64.
[0637] SEQ ID NO:41 shows the nucleotide sequence of primer CsgF_d27_end.
[0638] SEQ ID NO: 42 shows the nucleotide sequence of primer CsgF_d38_end
[0639] SEQ ID NO: 43 shows the nucleotide sequence of primer CsgF_d48_end
[0640] SEQ ID NO: 44 shows the nucleotide sequence of primer CsgF_d64_end
[0641] SEQ ID NO: 45 shows the nucleotide sequence of primer pNa62_CsgF_histag_Fw
[0642] SEQ ID NO: 46 shows the nucleotide sequence of primer CsgF-His_pET22b_FW
[0643] SEQ ID NO: 47 shows the nucleotide sequence of primer CsgF-His_pET22b_Rev
[0644] SEQ ID NO: 48 shows the nucleotide sequence of primer csgEFG_pDONR221_FW
[0645] SEQ ID NO: 49 shows the nucleotide sequence of primer csgEFG_pDONR221_Rev
[0646] SEQ ID NO: 50 shows the nucleotide sequence of primer Mut_csgF_His_FW
[0647] SEQ ID NO: 51 shows the nucleotide sequence of primer Mut_csgF_His_Rev
[0648] SEQ ID NO: 52 shows the nucleotide sequence of primer DelCsgE_Rev
[0649] SEQ ID NO: 53 shows the nucleotide sequence of primer DelCsgE FW
[0650] SEQ ID NO: 54 shows the amino acid sequence of residues 1 to 30 of mature E. coli CsgF
[0651] SEQ ID NO: 55 shows the amino acid sequence of residues 1 to 35 of mature E. coli CsgF
[0652] SEQ ID NO:56 shows the amino acid sequence of the mutant (T4C / N17S)CsgF sequence with a signal sequence and a TEV protease cleavage site (ENLYFQS) between residues 35 and 36 of the inserted mature protein sequence.
[0653] SEQ ID NO:57 shows the amino acid sequence of a mutant (N17S-Del)CsgF sequence having a signal sequence and a TEV protease cleavage site (ENLYFQS) between residues 35 and 36 of the inserted mature protein sequ...
Claims
1. A pore consisting of a CsgG pore and a modified CsgF peptide, wherein the modified CsgF peptide binds to CsgG and forms a constriction in the pore, wherein the modified CsgF peptide is: SEQ ID NO: 39, SEQ ID NO: 54, SEQ ID NO: 55, SEQ ID NO: 40, SEQ ID NO: 12, or SEQ ID NO: 14; or SEQ ID NO: 55 with a N17S or N17V modification; wherein the CsgG pore comprises 6 to 10 CsgG monomers which are SEQ ID NO:
3.
2. The pore according to claim 1, wherein the CsgF peptide and the CsgG pore are covalently coupled, and the covalent coupling is by means of: (i) a cysteine residue at a position corresponding to 132, 133, 136, 138, 140, 142, 144, 145, 147, 149, 151, 153, 155, 183, 185, 187, 189, 191, 201, 203, 205, 207, or 209 of SEQ ID NO: 3; (ii) a non-native reactive or photo-reactive amino acid at a position corresponding to 132, 133, 136, 138, 140, 142, 144, 145, 147, 149, 151, 153, 155, 183, 185, 187, 189, 191, 201, 203, 205, 207, or 209 of SEQ ID NO: 3; (iii) residues at positions corresponding to one or more of the following pairs of positions of SEQ ID NO: 6 and SEQ ID NO: 3, respectively: 1 and 153, 4 and 133, 5 and 136, 8 and 187, 8 and 203, 9 and 203, 11 and 142, 11 and 201, 12 and 149, 12 and 203, 26 and 191, and 29 and 144; or (iv) a disulfide bond or click chemistry.
3. The pore according to claim 1, wherein the interaction between the CsgF peptide and the CsgG pore is stabilized by a hydrophobic interaction, or an electrostatic or covalent interaction at positions corresponding to one or more of the following pairs of positions of SEQ ID NO: 6 and SEQ ID NO: 3, respectively: 1 and 153, 4 and 133, 5 and 136, 8 and 187, 8 and 203, 9 and 203, 11 and 142, 11 and 201, 12 and 149, 12 and 203, 26 and 191, and 29 and 144.
4. The pore according to claim 1, wherein the CsgG pore comprises at least one monomer comprising one or more modifications corresponding to the following modifications in SEQ ID NO: 3: (i) a modification at one or more of the positions Y51, N55, and F56; (ii) at least one substitution selected from R97W or R97Y and R93W or R93Y; (iii) a deletion of V105, A106, and I107 of SEQ ID NO: 3; (iv) a deletion of one or more of positions R192, F193, I194, D105, Y196, Q197, R198, L199, and E201 of SEQ ID NO: 3; (v) at least one substitution selected from K94N / Q / R / F / Y / W / L / S, D43S, E44S, F48S / N / Q / Y / W / I / V / H / R / K, Q87N / R / K, N91K / R, R97F / Y / W / V / I / K / S / Q / H, E101I / L / A / H, N102K / Q / L / I / V / S / H, R110F / G / N, Q114R / K, R142Q / S, T150Y / A / V / L / S / Q / N; (vi) a modification at one or more of positions I41, R93, A98, Q100, G103, T104, A106, I107, N108, L113, S115, T117, Y130, K135, E170, S208, D233, D238, E244, Q42, E44, L90, N91, I95, A99, E101, and Q114; (vii) at least one substitution selected from Y51A / I / V / S / T, N55A / I / V / S / T, and F56 / A / I / V / S / T / Q; (viii) substitution R97W; (ix) a deletion of F193, I194, D195, Y196, Q197, R198, and L199 or a deletion of D195, Y196, Q197, R198, and L199; (x) a deletion of V105, A106, and I107; (xi) a substitution selected from K94Q and K94N; (xii) at least one substitution selected from Q42K or Q42R; E44N or E44Q; L90R or L90K; N91R or N91K; I95R or I95K; A99R or A99K; E101H, E101K, E101N, E101Q, or E101T; and / or Q114K; (xiii) substitution N55V; and / or (xiv) R or K at a position corresponding to R192.
5. The pore of claim 1, which is a dual pore comprising two CsgG pores, wherein the modified CsgF peptide is inserted into the lumen of at least one of the CsgG pores.
6. The pore of claim 1, wherein one or more residues of SEQ ID NO: 12, SEQ ID NO: 14, SEQ ID NO: 54, SEQ ID NO: 39, SEQ ID NO: 40, or SEQ ID NO: 55 are modified.
7. The pore of claim 1, wherein the ratio of CsgG monomers to truncated CsgF peptide in the pore is 1: 1.
Citation Information
Patent Citations
electric kettle
CN3361190D
A miniature support for thin films containing single channels or nanopores and methods for using same
WO2000028312A1
Formation of layers of amphiphilic molecules
WO2009077734A2
Enzyme-PORE constructs
WO2010004265A1
Base-detecting pore
WO2010004273A1