Novel protein pores
Modified CsgF peptides form additional constrictions within the CsgG pore, addressing the limitations of current nanopore sensing by improving nucleotide discrimination and sequencing accuracy through enhanced current signatures and interaction sites.
Patent Information
- Application Number
- JP2025097852
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2017-06-30
- Filing Date
- 2025-06-11
- Publication Date
- 2025-10-07
AI Technical Summary
Existing nanopore sensing technologies, particularly those utilizing CsgG pores, face limitations in nucleotide discrimination due to insufficient differences in current signatures, necessitating improvements for enhanced analyte detection and sequencing accuracy.
The introduction of modified CsgF peptides, specifically truncated CsgF fragments, which form additional channel constrictions within the CsgG pore, creating a second interaction site and enhancing the current signature differences for better analyte discrimination.
The modified CsgG:CsgF pore complex significantly improves nucleotide discrimination and sequencing accuracy by introducing a second constriction, increasing the contact area and interaction sites, thereby enhancing the signal-to-noise ratio and accuracy of analyte detection.
Smart Images

Figure 2025148349000023 
Figure 2025148349000024 
Figure 2025148349000025
Abstract
Description
[Technical Field]
[0001] The present invention relates to novel protein pores and their use in analyte detection and characterization. The present invention further relates to transmembrane pore complexes, methods for producing pore complexes, and molecular sensing and for use in nucleic acid sequencing applications. [Background technology]
[0002] Nanopore sensing involves the detection of individual binding events or interactions between analyte molecules and the ion-conducting channel. Nanopore sensors are an approach to analyte detection and characterization that relies on the observation of action events. The researchers placed a single nanometer-sized pore in an electrically insulating membrane, allowing the analyte molecules to pass through. This can be generated by measuring the voltage-driven ionic current through the pore in the presence of The presence of an analyte in or near the pore alters the flow of ions through the pore, resulting in The identity of the analyte determines its unique ionic current or currents measured on the channel. The specific current signature, especially the duration and extent of current block, and its interaction with the pore The difference in current levels during the reaction time is apparent. Analytes are organic and inorganic sub-particles. and various biological or synthetic macromolecules and polynucleotides, polypeptides Nanopore sensing can reveal identity and detect Single molecule counting of the analytes can be performed, but nucleotides, amino acids, etc. Information about the analyte composition, such as amino acid or glycan sequence, as well as methylation and acyl oxidation, phosphorylation, hydroxylation, oxidation, reduction, glycosylation, decarboxylation, deamination, etc. It can also provide information about the presence of bases, amino acids, or glycan modifications. This may enable rapid and inexpensive synthesis of polynucleotides with lengths ranging from tens to tens of thousands of bases. It provides single molecule sequence readouts of polynucleotides.
[0003] Two key components of polymer characterization using nanopore sensing are (1) through the pore; (2) the control of polymer transport and the structure of the polymer as it moves through the pores. During nanopore sensing, the narrowest part of the pore is a function of the analyte passing through it. The reader head is the most differentiated part of the nanopore in terms of current signatures. CsgG acts as an ungated, nonselective protein secretion channel in Escherichia coli (E scherichia coli) (Goyal et al., 2014) and have been used as nanopores for detecting and characterizing analytes. In this context, mutations to the wild-type CsgG pore that improve the properties of the pore have been described. (WO2016 / 034591, WO2017 / 149316, WO2017 / 149 317 and WO2017 / 149318, PCT / GB2018 / 051191, all (which is incorporated herein by reference).
[0004] When the analyte is a polynucleotide, nucleotide discrimination occurs through such a mutant pore. The current signature is achieved by the channel contraction height and the interaction with the analyte. The extent of the interaction surface influences the relationship between the observed current and the polynucleotide sequence. As seen, it is sequence dependent and multiple nucleotides contribute to the observed current. It has been shown that the current range of nucleotide discrimination can be increased through mutations in the CsgG pore. Although improvements have been made, if the current difference between nucleotides can be further improved, the sequencing system Therefore, a new method to improve the function of nanopore sensing is needed. There is a need to identify the law. Summary of the Invention [Problem to be solved by the invention]
[0005] The present disclosure provides a method for binding the CsgG pore, thereby creating additional channels or This invention relates to modified CsgF peptides, particularly truncated CsgF fragments, that induce pore constriction or cleavage. Another aspect of the invention is a method for producing an isolated transmembrane pore complex and a nucleotide sequence encoding two consecutive leader sequences. The CsgG:CsgF complex and its modifications within a nanopore sensing platform with a rRNA Concerning the use of CsgF peptides or fragments. [Means for solving the problem]
[0006] A first aspect of the present invention relates to a CsgG pore and a pore comprising a CsgF peptide. In this case, the CsgF peptide contains a CsgG-binding region and a region that forms the constriction within the pore. In one embodiment, the CsgF peptide is a truncated CsgF peptide lacking the C-terminal head domain of CsgF. In another embodiment, the CsgF peptide is a C-terminal head and a CsgF In another embodiment, the CsgF peptide is a truncated CsgF peptide lacking a portion of the neck domain. The peptide is a truncated CsgF peptide lacking the C-terminal head and neck domains of CsgF. Pores are also referred to herein as pore complexes and as isolated pore complexes. The isolated pore complex may also be a CsgG pore or a homologue or mutant thereof. Variants, as well as modified CsgF peptides or homologs or mutants thereof, particularly truncated CsgF peptides. In one embodiment, the modified Cs comprises a sgF fragment or a homologue or mutant thereof. The gF peptide, or a homologue or mutant thereof, may be a CsgG pore or a homologue thereof. In another embodiment, the isolated pore complex is located within the lumen of the mutant. It has two or more channel constrictions, one of which is positioned or provided by the CsgG pore. , formed by the constriction loop, and another additional channel constriction or The leader head is bound by a modified CsgF peptide or its homologues or mutants. In one embodiment, the CsgG-pore or CsgG-like pore is , a mutant CsgG pore rather than a wild-type pore, and in certain embodiments, e.g., In another embodiment, the modified CsgF peptide has a mutation in the channel contractile loop. Isolated pore complexes containing tid or their homologs or mutants are 0.5 nm In one embodiment, the CsgF channel constriction has a diameter in the range of ∼2.0 nm. The pore complex comprises (i) a first opening, a middle portion containing a β-barrel, and a second opening; a CsgG pore having a lumen extending from the first opening through an intermediate portion to the second opening; (ii) a luminal surface of the intermediate portion defining a CsgG constriction region; and CsgF peptides, including a CsgF contraction domain and a CsgF binding domain (referred to herein as CsgF). the modified CsgF peptide has a CsgF binding domain (also referred to as the CsgF binding domain or region of CsgF), The CsgF constriction region is formed within the β-barrel of the CsgG pore, and the CsgG constriction region and Csg The F constriction is spaced coaxially within the β-barrel of the CsgG pore. The luminal surface of the pore contains one or more loop regions of the CsgG monomer that define the CsgG constriction. The CsgF contraction domain and the CsgF binding domain are typically located at the N of the CsgF mature peptide. In one embodiment, the pore complex comprises CsgA, CsgB, and CsgC. Contains no gE.
[0007] In a second aspect, the present invention provides a modified CsgF peptide or a homologue or mutant thereof. The protein or peptide may be modified by truncation or deletion of a portion of the protein. and the CsgF fragment of SEQ ID NO: 6 or a homologue or mutant thereof. One embodiment is a modified or truncated CsgF peptide, or a CsgF homologue or relates to a mutant modified peptide, said modified peptide being SEQ ID NO: 39, or SEQ ID NO: No. 40 or a homologue or mutant thereof, or alternatively, the modified peptide is SEQ ID NO: 15 or a homologue or mutant thereof, alternatively SEQ ID NO: 54, or SEQ ID NO: 55 or a homologue or mutant thereof. A modified CsgF peptide is disclosed, wherein one or more sequences within a region comprising SEQ ID NO: 15 are The position is modified and the mutant has at least 35% amino acid sequence identity to SEQ ID NO: 15. It is required that the amino acid identity be maintained within the peptide fragment corresponding to the region containing SEQ ID NO: 15. will be done.
[0008] One embodiment relates to a pore comprising a CsgG pore and a modified CsgF peptide, wherein the modified The decorative CsgF peptide binds to CsgG and forms a constriction within the pore.
[0009] One embodiment is the modified CsgF peptide or a homologue thereof according to the second aspect of the invention. or a mutant thereof. An isolated antibody comprising a G pore and a modified CsgF peptide or a homologue or mutant thereof. The modified CsgF peptide is a peptide disclosed in the second aspect of the present invention. It is characterized in that it is a peptide provided by a tide.
[0010] Another embodiment relates to an isolated pore complex, wherein the modified CsgF peptide and Cs a covalently linked IgG pore or a monomer of said pore or a homologue or mutant thereof More specifically, the binding is to a CsgG monomer having SEQ ID NO: 3 or its corresponding sequence. 132, 133, 136, 138, 140, 142, 144, 145, 147, 1 49, 151, 153, 155, 183, 185, 187, 189, 191, 201, 2 via a cysteine residue at position 03, 205, 207 or 209 Alternatively, this can be achieved via unnatural reactive or photoreactive amino acids.
[0011] A preferred embodiment relates to an isolated transmembrane pore complex, or membrane composition, which is The invention includes isolated pore complexes and membrane components, particularly said transmembrane pore complexes. Alternatively, the membrane composition may comprise an isolated pore complex, a membrane component or an insulating layer component of the present invention. It is composed of elements.
[0012] One embodiment relates to a method for generating pores as disclosed herein, the method comprising: One or more CsgG monomers disclosed herein and a CsgF peptide disclosed herein. co-expressing a transmembrane pore complex with the ATPase in a host cell, thereby enabling transmembrane pore complex formation within the cell. The CsgF peptide is prepared by modifying the amino acid sequence to include an enzyme cleavage site at an appropriate position. Decorated CsgF peptides or proteins may be produced intracellularly by cleaving them. .
[0013] One embodiment relates to a method for generating pores as disclosed herein, the method comprising one or more The purified CsgG monomer is contacted with one or more purified modified CsgF peptides. The modified CsgF is then allowed to react with the modified CsgF to form the pores in vitro. The peptide is cleaved by an enzyme at the appropriate position in the amino acid sequence, either before or after the formation of the pore. It may also be a peptide containing a cleavage site.
[0014] A third aspect of the present invention relates to a method for producing the transmembrane pore complex, wherein the The pore comprises a CsgG pore or a homologue or mutant thereof, and a modified CsgF peptide. or a homolog or mutant thereof, and The present invention relates to the production of CsgG SEQ ID NO:2 or a homologue or mutant thereof in a suitable host cell, and and modified or truncated CsgF (including fragments of SEQ ID NO: 5) or homologs thereof. co-expressing the mutants, thereby enabling pore complex formation in vivo. In certain embodiments, the modified CsgF peptide or a homologue or protease thereof is Natural variants include SEQ ID NO: 12 or SEQ ID NO: 14 or homologs or mutants thereof. Alternatively, the method for producing isolated pore complexes may involve in situ analysis of the pore complexes. For in vitro reconstitution, Csg of SEQ ID NO: 3 or its homologues or mutants contacting the G monomer with a modified CsgF peptide or a homologue or mutant thereof In certain embodiments, the modified CsgF peptide of the method comprises the steps of: or SEQ ID NO: 16 or a homologue or mutant thereof.
[0015] Another aspect of the present invention is a method for determining the presence, absence, or one or more characteristics of a target analyte. In relation to the law, this method (i) transferring a target analyte to said isolated pore channel so that the target analyte moves into the pore channel; contacting the membrane with a pore complex or transmembrane pore complex; (ii) taking one or more measurements as the analyte moves through the pore channel; and determining the presence, absence, or one or more characteristics of the analyte by
[0016] In one embodiment, the analyte is a polynucleotide. The method of using the polynucleotide as a target is based on (i) the length of the polynucleotide, (ii) the length of the polynucleotide, (iii) the identity of the polynucleotide; (iv) the secondary structure of the polynucleotide; and (v) one or more characteristics selected from whether the polynucleotide is modified. The method includes determining:
[0017] In another embodiment, the analyte is a protein or a peptide. In some embodiments, the analyte may be a polysaccharide or a compound such as, for example, a pharmacologically active compound, a toxic compound, and Small organic or inorganic compounds such as, but not limited to, pollutants.
[0018] In another embodiment, the isolated transmembrane pore complex is used to detect a polynucleotide or (polynucleotides). (i) A method for characterizing a peptide is described, wherein the pore complex is a CsgG pore. or homologs or mutants thereof, and modified CsgF peptides or homologs thereof. In particular, the CsgG pore or a phase thereof is an isolated complex comprising the CsgG pore or a mutant thereof. The isoform or mutant contains 6–10 CsgG monoclonals that form the CsgG pore channel. Including Ma.
[0019] A further aspect of the invention is a method for determining the presence, absence, or one or more characteristics of a target analyte. Use of said isolated pore complex or transmembrane pore complex according to the above aspect of the invention for The present invention further discloses a method for producing a membrane comprising the steps of: (a) the isolated pore complex; and (b) a membrane comprising: Also related is a kit for characterizing a target analyte that includes the components.
[0020] Text description of the illustration image022.gif. The drawings described are schematic and non-limiting. In the drawings, some elements are shown for illustrative purposes. Sizes may be exaggerated and not drawn to scale. [Brief explanation of the drawings]
[0021] [Figure 1]Structure of the CsgG pore and its interface with CsgF for complex formation. Cross-section (A), side (B), and top (C) views of a CsgG oligomer (e.g., nonamer) (gold color), showing the surface (A) and ribbon (B, C), with a single CsgG protomer colored light blue (D) (based on the CsgG X-ray structure PDB entry: 4uv3). The CsgG contractile loop (CL loop), spanning residues 46-61 according to SEQ ID NO: 3, is shown in dark gray in all panels and corresponds to the loop in the lower left corner of (E). CsgG residues whose side chains face the lumen of the CsgG β-barrel are in medium gray and labeled with β-strands as shown in (E) and (D). These residues represent sites available for substitution with natural or unnatural amino acids, such as those that can be used for attachment (covalent cross-linking) of pore peptides (e.g., including modified CsgF peptides or homologs thereof) to the CsgG pore or monomer. In some embodiments, cross-linking residues include Cys and reactive amino acids, and photoreactive amino acids, azidohomoalanine, homopropargylglycine, homoallelicglycine, p-acetyl-Phe, p-azido-Phe, p-propargyloxy-Phe, and p-benzoyl-Phe (Wang et al. 2012, Chin et al. 2002), and can be substituted at positions 132, 133, 136, 138, 140, 142, 144, 145, 147, 149, 151, 153, 155, 183, 185, 187, 189, 191, 201, 203, 205, 207, or 209 according to SEQ ID NO: 3. (E) A close-up view of the CL loop and transmembrane beta strand of the CsgG monomer is shown. The CsgG contractile loop (colored dark blue) forms the orifice or narrowest passage within the CsgG pore (panel A). In some embodiments, three positions in the CL loop, 56, 55, and 51 according to SEQ ID NO: 3, are particularly important for the diameter and chemical and physical properties of the CsgG channel orifice or "reader head." These represent preferred positions for altering the nanopore-sensing properties of the CsgG pore and homologs. [Figure 2]Coexpression of CsgG:CsgF complex proteins and complex purification. (A) Schematic diagram of the purification protocol for the CsgG:CsgF complex, starting from an E. coli culture coexpressing CsgG (SEQ ID NO:2 + C-terminal Strep II tag) and CsgF (SEQ ID NO:4 + C-terminal 6xHis tag). The protocol includes resuspending cells and 1% DDM extraction of membrane-bound proteins. The CsgG:CsgF complex and excess CsgF undergo a first enrichment by affinity purification on a nickel-IMAC column, followed by a second affinity enrichment of the CsgG:CsgF complex on a streptavidin column. (B) Coomassie-stained SDS-PAGE of the IMAC (left) and streptavidin (right) purification steps. Protein bands corresponding to CsgG and CsgF are labeled. Notably, the IMAC eluate contained an N-terminally truncated CsgF fragment (labeled *) that was not retained in affinity pull-downs using the CsgG-bound Strep tag, indicating that the N-terminus of CsgF is required for complex formation with CsgG. [Figure 3-1]CsgG:CsgF complex protein purification following in vitro reconstitution. (A) Overlaid chromatograms of size-exclusion chromatography (SEC) runs (using a BioRad Enrich 650 10 / 300 column) of CsgG (light gray) and CsgG supplemented with excess CsgF (dark gray). The chromatograms show elution peaks corresponding to the CsgG 9-mer (a) and CsgG 18-mer (b) for CsgG reconstitution, excess free CsgF (c), and the 9-mer CsgG:CsgF complex (d) and the 18-mer CsgG:CsgF complex (e), which elute at higher hydrodynamic radii (molecular weights) due to the incorporation of CsgF into the complex. (B) Native PAGE analysis of representative species labeled in panel (A) confirms the shift to higher molecular weight upon incorporation of CsgF into the CsgG 9-mer and CsgG 18-mer complexes. These experiments demonstrate that the CsgG:CsgF complex can be reconstituted in vitro starting from purified components. (C) Ribbon representation of the previously reported CsgG 9-mer and CsgG 18-mer (Goyal et al. 2014 (PDB entry 4uv3)). The CsgG 18-mer is formed from a dimer of CsgG 9-mers. SEC and native PAGE analysis shown in panels A and B demonstrate the amenability of the CsgG 18-mer to complex with CsgF. [Figure 3-2] (Same as Figure 3-1) [Figure 4-1]CsgG:CsgF structure determined by cryo-EM. (A) Cryo-EM image of the CsgG:CsgF complex, showing the presence of 9- and 18-mer CsgG:CsgF complexes. Numerous single particles of the 9- and 18-mer forms are highlighted by solid and dashed circles, respectively. (B) Two representative class averages of the CsgG:CsgF 9-mer complex from a side view. The class averages contain 6,020 to 4,159 individual particles, respectively. The class averages reveal the presence of additional density at the top of the CsgG particles, corresponding to oligomeric complexes of CsgF. Three distinct regions are observed within the CsgF oligomer: the "head" and "neck" regions, and a region residing within the lumen of the CsgG β-barrel, forming a constriction or narrow passage (labeled F) stacked on top of the constriction formed by the CsgG CL loop (labeled G). This latter CsgF domain is called the CsgF contractile peptide (FCP). [Figure 4-2] (Same as Figure 4-1) [Figure 5] Three-dimensional structural model of the CsgG:CsgF complex. Cross-section of the 3D cryoEM electron density of the CsgG:CsgF nonamer complex, calculated from 20,000 particles assigned to 21 class averages. The right image shows the overlay with the docked X-ray structure of the CsgG nonamer (PDB entry 4uv3) onto the cryoEM density. Regions corresponding to CsgG, CsgF, and the head, neck, and FCP domains of CsgF are shown. The cross-section shows that the CsgF FCP region forms an additional constriction (labeled F) within the CsgG channel approximately 2 nm above the CsgG contractile loop (labeled G). [Figure 6]Schematic of the CsgG:CsgF pore complex based on the cryo-EM structure. (A) Schematic of a CsgG nanopore in cross section, retaining a single constrictor (labeled (1)). The CsgG-based nanopore forms a 3.5-4 nm wide channel containing a 0.5-1.5 nm orifice created by the CsgG contractile loop (residues 46-61 according to SEQ ID NO: 3). Upon complexing with CsgF, a second constrictor or orifice is introduced into the CsgG channel (labeled (2) / F), and the channel exit becomes blocked by the CsgF head domain (see Figure 5). When modified CsgF peptides are used, such as those corresponding to the CsgF contractile peptide (FCP) lacking the neck and head regions, the CsgG:CsgF pore complex forms with two consecutive channel constrictions or orifices ((1) and (2)), as shown in the cross-section of the CsgG:CsgF cryo-EM density in panel (B) and the schematic in panel (C). Removal of the neck and head regions of the modified CsgF peptide relieves blockage of the channel outlet. [Figure 7] Schematic of the use of the CsgG:CsgF pore complex for nanopore sensing applications of (bio)polymers (A) or single-molecule analytes (B). When used in polymer sensing, the second channel constriction introduced by the modified CsgF peptide increases the contact area with the analyte and forms a second interaction site and reader head. When used in single-molecule nanopore sensing, the second channel constriction introduced by the modified CsgF peptide creates a second, independent analyte interaction site. (C) Schematic of the theoretical channel conductance profile (shown as hexagons and triangles) of a small molecule passing through and interacting with consecutive CsgG (1) and CsgF (2) constrictions or reader heads. [Figure 8]Multiple sequence alignment of exemplary CsgF homologs. Aligned sequences are shown as mature proteins (i.e., lacking their N-terminal signal peptides (SPs)). Boxed sequences indicate conserved CsgF regions (between 35-100% pairwise sequence identity—see Figure 10) that, in some embodiments, correspond to the CsgF contractile peptide (FCP). The CsgF homologs included in the multiple sequence alignment are Q88H88, A0A143HJA0, Q5E245, Q084E5, F0LZU2, A0A136HQR0, A0A0W1SRL3, B0UH01, Q6NAU5, G8PUY5, A0A0S2ETP7, E3I1Z1, F3Z094, A0A176T7M2, D2QPP8, N2IYT1, W7QHV5, D4ZLW2, D2QT92, and A0A167UJA2. The FCP regions of E. coli CsgF (SEQ ID NO: 15) and the indicated CsgF homologs correspond to SEQ ID NOs: 18-36. [Figure 9]Experimental evaluation of the E. coli CsgF region that forms the CsgG-interacting sequence and the CsgF contractile peptide (FCP). Panel (A) shows the mature sequences (i.e., after removal of the CsgF signal peptide corresponding to residues 1-19 of SEQ ID NO:5) of four N-terminal CsgF fragments (CsgF residues 1-27 of SEQ ID NO:8, SEQ ID NO:10, SEQ ID NO:12, and SEQ ID NO:14) coexpressed with E. coli CsgG (SEQ ID NO:2). (B) Anti-Strep (left) and anti-His (right) Western blot analysis of SDS-PAGE analysis of crude cell lysates from CsgG and CsgF coexpression experiments. Anti-Strep analysis demonstrates CsgG expression in all coexpression experiments, while anti-His Western blot analysis shows detectable levels of CsgF fragments only for the truncation mutant CsgF 1-64 (SEQ ID NO: 14). His-tagged nanobody (Nb) was used as a positive control. (C) Anti-His dot blot analysis for the presence of CsgF fragments in CsgG:CsgF coexpression experiments. The top row shows whole cell lysates, while the middle and bottom rows show the eluates and flow-through of Strep affinity pull-down experiments. These data demonstrate that CsgF fragment 1-64 and, to a much lesser extent, CsgF 1-48, are specifically pulled down as a complex with Strep-tagged CsgG. CsgF fragments 1-27 and 1-38 do not yield detectable levels of the corresponding CsgF fragments, showing no sign of complex formation with CsgG. [Figure 10-1]Multiple sequence alignment of CsgF regions forming the CsgG-interacting sequence and CsgF contractile peptide (FCP). The figure shows the multiple sequence alignment and consensus sequence of CsgF peptides and their known homologs in the region corresponding to CsgG interaction. CsgF homologs are defined by the PFAM domain PF03783. These peptides bind to CsgG and localize to the lumen of the CsgG β-barrel, where they form an additional constriction within the CsgG channel. These peptides and their homologs are examples of CsgF contractile peptides or FCPs. The pairwise sequence identities of the FCPs shown range from 35 to 98%. [Figure 10-2] (Same as Figure 10-1) [Figure 10-3] (Same as Figure 10-1) [Figure 11] High-resolution cryo-EM structure of the CsgG:CsgF complex. CsgG is shown in light gray, and CsgF is shown in dark gray. A. Final electron density map of the CsgG:CsgF complex at 3.4 Å resolution. Side view. B. Top view of the cryo-EM structure showing that CsgG:CsgF contains 9:9 stoichiometry and has C9 symmetry. C. Internal structure of the CsgG:CsgF complex. GC, CsgG constrictor; FC, CsgF constrictor. D. Interaction between the CsgG and CsgF proteins. CsgG and CsgG constrictor are colored light gray and gray, respectively. CsgF is colored dark gray. CsgG and CsgF residues are labeled light gray and black, respectively. [Figure 12] The two reader heads of the CsgG:CsgF complex. CsgG is shown in light gray, and the reader head of the CsgG pore is shown in dark gray. CsgF is shown in black, and the CsgF reader head is labeled. [Figure 13]Coexpression of CsgG with CsgF WT in vivo. Genes encoding a C-terminally Strep-tagged CsgG polypeptide in a pT7 vector carrying ampicillin resistance and a C-terminally tagged CsgF polypeptide in a pRham vector carrying kanamycin resistance were transformed together into E. coli BL21DE3 cells in the presence of both ampicillin and kanamycin. Proteins were expressed overnight at 18°C and 250 rpm, and the CsgG-CsgF complex was purified using Strep-tag purification followed by His-tag purification. A. Protein sample before Strep purification (duplicate). B. Protein sample after His-purification (triplicate elution fractions). Proteins were run on a 4-20% Tris gel. [Figure 14] In vitro coexpression of CsgG and CsgF and thermostability of the CsgG-CsgF complex. CsgG and CsgF DNA in different vectors were coexpressed in an in vitro transcription and translation reaction. The proteins were radiolabeled with S-35 methionine and exposed to X-ray film. The stability of the complex was assessed by incubating the reaction mixture at different temperatures for 10 minutes. [Figure 15]Preparation of CsgG:CsgF complexes using protease cleavage sites. A. A TEV or C3 or any other protease cleavage site can be introduced into the CsgF peptide at the desired site (e.g., between 30 and 31, 35 and 36, 40 and 41, or 45 and 46 of SEQ ID NO:6). CsgG is shown in gold, and the CsgF domain is shown in red. One CsgF subunit, 1-35, is colored green for clarity. Subunits 36-45 are shown in purple. The ten histidine tag is shown in pink, and the Strep tag of CsgG is shown in blue. B. SDS-PAGE (4-20% TGX) of protease cleavage of the full-length CsgG:CsgF complex, with a TEV protease cleavage site inserted between 35 and 36 of SEQ ID NO:6. M: molecular weight marker; lane 1: strep-purified full-length CsgG:CsgF complex; lane 2: strep-concentrated full-length CsgG:CsgF complex; lane 3: gel filtration; lane 4: TEV protease cleavage to generate the CsgG:CsgF complex; lane 5: strep-purified CsgG:CsgF flow-through; lane 6: CsgG:CsgF heated at 60°C for 10 min; lane 7: CsgG:CsgF complex eluted from the strep column; lane 8: CsgG pore as a control; lane 9: TEV protease as a control. [Figure 16]Thermal stability of the CsgG:CsgF complex. M: molecular weight marker; lane 1: CsgG pore; lane 2: CsgG:CsgF complex at room temperature; lanes 3-9: CsgG:CsgF samples heated at different temperatures (40, 50, 60, 70, 80, 90, and 100 °C, respectively) for 10 min. Lane 1: A. Y51A / F56Q / N55V / N91R / K94Q / R97W-del(V105-I107):CsgF-(1-45). B. Y51A / F56Q / N55V / N91R / K94Q / R97W-del(V105-I107):CsgF-(1-35). C. Y51A / F56Q / N55V / N91R / K94Q / R97W-del(V105-I107):CsgF-(1-30). Samples were subjected to SDS-PAGE on a 7.5% TGX gel. The CsgG:CsgF complex containing both CsgF-(1-45) and CsgF-(1-35) shows a shift from the CsgG pore band in lane 1. Thus, both of these complexes appear to be thermostable up to 90°C. The complex and pore disassemble to CsgG monomers at 100°C (lane 9). The same thermostability pattern is seen for the CsgG:CsgF complex with CsgF-(1-30), although it is difficult to see a shift between the protein bands of the CsgG pore (lane 1) and the CsgG-CsgF complex (lanes 2-8). [Figure 17]Native PAGE showing CsgG:CsgF formation via in vitro reconstitution using synthetic CsgF peptides. The Alexa594-labeled CsgF peptide corresponding to the first 34 residues of mature CsgF (SEQ ID NO: 6) was added to purified Strep-tagged CsgG or Y51A / F56Q / K94Q / R97W / R192D-del(V105-1107) at a 2:1 molar ratio for 15 minutes at room temperature in 50 mM Tris, 100 mM NaCl, 1 mM EDTA, and 5 mM LDAO / C8D4. After CsgG-strep pulldown on StrepTactin beads, samples were analyzed by native PAGE. Both WT and Y51A / F56Q / K94Q / R97W / R192D-del(V105-l107) CsgG bind the CsgF N-terminal peptide, visualized by a fluorescent tag. [Figure 18] Figure 18. Stabilization of the CsgG:CsgF or CsgG:FCP complex. A. Identified amino acid positions in the CsgG (SEQ ID NO: 3) and CsgF (SEQ ID NO: 6) pair that can form disulfide bonds. B. Schematic diagram showing the disulfide bonds between CsgG-Q153C and CsgF-G1C. [Figure 19]Cysteine cross-linking of the CsgG:CsgF complex. A. Y51A / F56Q / N91R / K94Q / R97W / Q153C-del(V105-I107) and CsgF-G1C proteins were separately purified and incubated at 4°C for 1 hour or overnight to allow complex formation and SS formation. No oxidizing agent was added to promote SS formation. Control CsgG pores (Y51A / F56Q / N91R / K94Q / R97W / Q153C-DEL(V105-I107)) and complexes (with or without DTT) were heated at 100°C for 10 minutes to disassemble the complex into CsgG monomers (CsgGm, 30 kDa) and CsgF monomers (CsgFm, 15 kDa). A dimer between CsgGm and CsgFm (CsgGm-CsgFm, 45 Kda) was observed in the absence of a reducing agent, confirming disulfide bond formation. Increased dimer formation was observed in overnight cultures compared to 1-hour cultures. B. Mass spectrometry analysis was performed on the gel-purified m-CsgFm band from the overnight culture. The protein was proteolytically cleaved to generate tryptic peptides. LC-MS / MS sequencing was performed to identify the precursor ion shown above, corresponding to the indicated link peptide. This precursor ion was fragmented, and fragment ions were observed. These included ions for each peptide and fragments incorporating an intact disulfide bond. This data provides strong evidence for the presence of a disulfide bond between C1 of CsgF and C153 of CsgF. [Figure 20]Improving the efficiency of cysteine cross-linking in the CsgG:CsgF complex. Lane 1: Y51A / F56Q / N91R / K94Q / R97W / N133C-del (V105-I107) and CsgF-T4C proteins were coexpressed, and the CsgG:CsgF complex was purified. Lane 2: The complex was heated in the presence of DTT to decompose the complex into substituted monomers (CsgGm and CsgFm). DTT disrupts any disulfide bonds between CsgG-N133C and CsgF-T4C, if any are formed. Lane 3: The complex was incubated with the oxidizing agent copper orthophenanthroline to promote disulfide bond formation. Lane 4: The oxidized sample was heated to 100°C in the absence of DTT to decompose the complex. A new band of 45 kda corresponding to CsgGm-CsgFm appears, confirming disulfide bond formation. [Figure 21] Current signatures as a DNA strand passes through the CsgG:CsgF complex. A CsgG pore containing a C-terminal Strep tag (Y51A / F56Q / N91R / K94Q / R97W-del(V105-I107)) was coexpressed with full-length CsgF protein containing a C-terminal His tag and a TEV protease cleavage site between 35 and 36 of SEQ ID NO: 6 to generate the complex. The purified complex was then cleaved with TEV protease to generate the given CsgG:CsgF complex. Note that TEV cleavage leaves an ENLYFQ sequence at the cleavage site. A. No mutation at position 17 of CsgF. B. N17S mutant in CsgF. [Figure 22] Current signatures as a DNA strand passes through the CsgG:CsgF complex. The complex was generated by incubating the Y51A / N55V / F56Q / N91R / K94Q / R97W-del(V105-I107) pore containing a C-terminal Strep tag with the CsgF-(1-35) mutant. A. CsgF-N17S-(1-35). B. CsgF-N17V-(1-35). [Figure 23-1]Current signatures as a DNA strand passes through the CsgG:CsgF complex. The complexes were generated by incubating different CsgG pores containing a C-terminal Strep tag with CsgF-N17S-(1-35). A. The CsgG pore is Y51A / N55V / F56Q / N91R / K94Q / R97W-del(V105-I107). B. The CsgG pore is Y51T / N55V / F56Q / N91R / K94Q / R97W-del(V105-I107). C. The CsgG pore is Y51A / N55I / F56Q / N91R / K94Q / R97W-del(V105-I107). D. The CsgG pore is Y51A / F56A / N91R / K94Q / R97W-del(V105-I107). E. The CsgG pore is Y51A / F56I / N91R / K94Q / R97W-del(V105-I107). F. The CsgG pore is Y51S / N55V / F56Q / N91R / K94Q / R97W-del(V105-I107). [Figure 23-2] (same as Figure 23-1) [Figure 24] Current signatures as a DNA strand passes through the CsgG:CsgF complex. The complexes were generated by incubating E. coli purified Y51A / N55V / F56Q / N91R / K94Q / R97W-del(V105-I107) pores containing three different lengths of CsgF at the C-terminus: A. CsgF-(1-29), B. CsgF-(1-35), C. CsgF-(1-45). Arrows indicate the range of the signal. Surprisingly, the complex with CsgF-(1-29) produces the signal with the greatest range. [Figure 25]Signal-to-noise ratio of the current signature as a DNA strand passes through the CsgG:CsgF complex. Different CsgG:CsgF complexes are encapsulated in different CsgG pores with the same CsgF peptide, CsgF-(1-35): 1 - Y51A / F56Q / N91R / K94Q / R97W-del(V105-I107) 2 - Y51A / N55I / F56Q / N91R / K94Q / R97W-del(V105-I107) 3 - Y51A / N55V / F56Q / N91R / K94Q / R97W-del(V105-I107) 4 - Y51A / F56A / N91R / K94Q / R97W-del(V105-I107) 5 - Y51 A / F56I / N91R / K94Q / R97W-del(V105-I107)6-Y51A / F56V / N91R / K94Q / R97W-del(V105-I107)7-Y51S / N55A / F56Q / N91R / K94Q / R97W-del(V105-I107)8-Y51S / N55V / F56Q / N91R / K94Q / R97W-del(V105-I107)9-Y51T / N55V / F56Q / N91R / K94Q / R97W-del(V105-I107) was generated by culturing the following DNA translocations: A / F56I / N91R / K94Q / R97W-del(V105-I107)6-Y51A / F56V / N91R / K94Q / R97W-del(V105-I107)7-Y51S / N55A / F56Q / N91R / K94Q / R97W-del(V105-I107)8-Y51S / N55V / F56Q / N91R / K94Q / R97W-del(V105-I107)9. A distinct pattern of short, squiggly lines was observed in DNA translocation experiments, and the signal-to-noise ratio was measured. Greater accuracy can be obtained with a greater signal-to-noise ratio. [Figure 26] A. Sequencing errors with a narrow reader head. B. Representation of DNA base interactions with the reader head of a CsgG pore. Approximately five bases dominate the current signal at any one time as the DNA strand translocates through the pore. C. Mapping plot of signal. Event detection signals for multiple reads mapped to signals modeled using a custom HMM, for a mixed sequence lacking homopolymer runs, and for a sequence containing three homopolymer runs of 10 T. [Figure 27-1]Mapping the reader head of the CsgG:CsgF complex. Reader head discrimination plot for the CsgG:CsgF complex. Average variation in modeled current as the base is varied at each read head position. To calculate the read head discrimination at position i for a model of length k with an alphabet of length n, the discrimination at read head position i was defined as the median of the standard deviations of the current levels for each nk-1 group of size n. Here, position i is varied while the other positions are held constant. B. Resting DNA strands for mapping the reader head: A set of polyA DNA strands (SS20–SS38) was created in which one base was deleted from the DNA backbone (Ispc3). In each strand, the position of ISpc3 shifts from the 3' end to the 5' end. Based on previous experiments with the CsgG pore, the seventh position of the DNA is expected to be located within the CsgG constriction. SS26, which corresponds to this DNA, is highlighted. Based on the model from (A), 4–5 bases are expected to separate CsgG and the CsgG reader head. Therefore, approximately positions 12 and 13 are expected to be within the CsgF constriction. The DNA strands corresponding to these positions, SS31 and SS32, are highlighted. C and D. Mapping of the two reader heads: Biotin modifications at the 3' end of each strand are complexed with monovalent streptavidin, and the current blocks generated from each strand are recorded in the MinION setup. If iSpc3 is located above or below the constriction within the pore, no deflection is expected. However, when iSpc3 is located within the constriction, a higher current level is expected to pass through the pore. This is because the extra space created by the base deletion allows more ions to pass. Therefore, the positions of the two reader heads can be mapped by plotting the current passing through each DNA strand. As expected, the largest deflection of current is observed when position 7 of the DNA strand is occupied by iSpc3 (C). iSpc3 at positions 6 and 8 also produces a higher deflection across the mean polyA current level. Thus, positions 6, 7 and 8 of the DNA strand represent the first leader head (CsgG leader head).As expected, positions 12 and 13 are occupied by iCsp3, and another deviation from the baseline polyA is observed (D). This indicates a second reader head (CsgF reader head) in the pore. The results also confirm that the two reader heads are separated by approximately 4–5 bases. [Figure 27-2] (same as Figure 27-1) [Figure 28] Reader head discrimination and base contribution. The left panel shows the readhead discrimination for each mutant pore, which is the average variation in modeled current when the base at each readhead position is varied. To calculate the readhead discrimination at position i for a model of length k with an alphabet of length n, the discrimination at readhead position i was defined as the median of the standard deviation of the current levels for each nk-1 group of size n, where position i is varied while the other positions are held constant. The right panel shows the base contribution plot. The median current for all sequence contexts with base b (A, T, G, or C) at leader position i. [Figure 29]Error profile of the double reader head pore. A. Schematic of the CsgG:CsgF complex and the interactions of DNA bases with the two reader heads. Red: strong interaction, orange: weak interaction, gray: no interaction. B. Comparison of errors in deletions. Reads from the Y51A / F56Q / N91R / K94Q / R97W / R192D-del(V105-I107) and Y51A / N55V / F56Q / N91R / K94Q / R97W-del(V105-I107):CsgF-N17S-(1-35) pores were base called from the same region of E. coli DNA. Reads were aligned to the reference genome using Minimap2 (https: / / arxiv.org / abs / 1708.01 492), and the resulting alignments were visualized in the Savant Genome Browser (https: / / www.ncbi.nlm.nih.go v / pubmed / 20562449). The majority of Y51A / F56Q / N91R / K94Q / R97W / R192D-del(V105-I107) reads contain a single base deletion (black box) within the T homopolymer, which is absent in the majority of CsgG:CsgF reads. C. Comparison of consensus accuracy from crude data generated from the Y51A / F56Q / N91R / K94Q / R97W / R192D-del(V105-I107) (blue) and Y51A / N55V / F56Q / N91R / K94Q / R97W-del(V105-I107):CsgF-N17S-(1-35) pores (green) relative to the homopolymer length. [Figure 30-1]Homopolymer call of the CsgG:CsgF complex. DNA with the sequence shown in (A) was translocated through the Y51A / F56Q / N91R / K94Q / R97W / R192D-del(V105-I107) pore (B) and the Y51A / N55V / F56Q / N91R / K94Q / R97W-del(V105-I107):CsgF-N17S-(1-35) pore (C), and the signals were analyzed for the first polyT section, shown in red in (A). When the polyT section passes through a CsgG pore containing a single leader head (the model is based on five bases located within the leader head), it generates a flat line in the signal. Therefore, it is difficult to determine the exact number of bases in this region, which often causes deletion errors. When DNA passes through a CsgG:CsgF complex containing two leader heads (the model is based on the nine bases located within and between the two leader heads), the poly-T section shows multiple steps rather than a flat line. Information about these steps can be used to precisely identify the number of bases within a homopolymer region. This additional information significantly reduces deletion errors and improves overall consensus accuracy. [Figure 30-2] (same as Figure 30-1) [Figure 30-3] (same as Figure 30-1) [Figure 30-4] (same as Figure 30-1) [Figure 31] Characterization of the CsgG pore (Y51A / F56Q / N91R / K94Q / R97W / -del(V105-I107). A. CsgG pore reader head discrimination. Average variation in modeled current as the base varies at each reader head position. To calculate reader head discrimination at position i for a model of length k with an alphabet of length n, the discrimination at reader head position i was defined as the median of the standard deviation of the current levels for each k-1 group of size n, where position i is varied while the other positions are held constant. B. Base contribution plot of the CsgG pore. Median current for all k-mers with base b (A, T, G, or C) at reader head position I. C. Current signature as a DNA strand passes through the CsgG pore. DETAILED DESCRIPTION OF THE INVENTION
[0022] While the present invention will be described with respect to particular embodiments and with reference to certain drawings, The invention is not limited thereto, but is limited only by the claims. Of course, it should be understood that the present invention Not necessarily all aspects or advantages may be achieved in accordance with any particular embodiment. It should be understood that the teachings or recommendations of this specification are not intended to be limiting. As taught herein, one or more advantages may be achieved or optimized. It will be appreciated that the invention can be embodied or carried out in any number of ways.
[0023] The present invention, both as to its organization and method of operation, together with its features and advantages, is described in the following detailed description. The present invention can be best understood when read in conjunction with the accompanying drawings. This will be apparent from and elucidated with reference to the embodiments described hereinafter. Throughout this specification, references to "one embodiment" or "an embodiment" refer to an embodiment. A particular feature, structure, or characteristic described in connection therewith may be present in at least one embodiment of the present invention. Thus, in various places throughout this specification, the terms "in one embodiment" and "in one embodiment" are used interchangeably. or "in an embodiment" do not necessarily refer to the same embodiment. Similarly, it should be understood that the examples of the present invention In describing exemplary embodiments, various features of the present invention are presented to streamline the disclosure and provide various inventive features. For the purpose of facilitating understanding of one or more aspects, the present invention may be summarized in a single embodiment, figure, or description. However, this method of the present disclosure may be used in the claims to This is not to reflect an intention that more features are required than are expressly recited in the claims. Rather, as the following claims reflect, An aspect may be less than all features of a single foregoing disclosed embodiment.
[0024] As used in this specification and the appended claims, the singular forms "a," "an," and "an" are used interchangeably. and "the" include plural referents unless the content clearly dictates otherwise. For example, "a polynucleotide" includes two or more polynucleotides, and "a polynucleotide A "helicase" contains two or more such proteins, and a "helicase" contains two "monomer" means two or more "monomers" and "pore" includes two or more pores and the like.
[0025] In all discussions herein, the standard single-letter code for amino acids is used. They are: alanine (A), arginine (R), asparagine (N), Aspartic acid (D), cysteine (C), glutamic acid (E), glutamine (Q), Lysine (G), Histidine (H), Isoleucine (I), Leucine (L), Lysine (K) , methionine (M), phenylalanine (F), proline (P), serine (S), threo threonine (T), tryptophan (W), tyrosine (Y), and valine (V). Standard substitutions The same notation is also used, i.e., Q42R means that Q at position 42 is replaced with R. .
[0026] In paragraphs where different amino acids at a particular position are separated by a / symbol, the / symbol is replaced with " For example, Q87R / K means Q87R or Q87K.
[0027] In paragraphs where different positions are separated by the / symbol, the / symbol means "and", and Y51 / N55 means Y51 and N55.
[0028] All amino acid substitutions, deletions, and / or additions disclosed herein are intended to be construed as being equivalent to the amino acid substitutions, deletions, and / or additions to the amino acid sequences disclosed herein. Unless otherwise indicated, mutant CsgG monomers, including variants of the sequence shown in SEQ ID NO: 3 means...
[0029] Reference to mutant CsgG monomers, including variants of the sequence shown in SEQ ID NO: 3, is , mutant Csg including variants of the sequences set forth in the further SEQ ID NOs. disclosed below Mutant CsgG monomers including variants of the sequence shown in SEQ ID NO: 3 Sequences corresponding to the substitutions, deletions and / or additions disclosed herein with reference to Amino acid substitutions, deletions, and / or additions were made to the CsgG monomer, including the mutants It is possible to do so.
[0030] All publications, patents and patent applications cited herein, whether supra or infra, are incorporated by reference in their entirety. and is incorporated herein by reference in its entirety.
[0031] definition When referring to a singular noun, use an indefinite or definite article (e.g., "a," or "an," or When "the" is used, this includes the plural of that noun unless otherwise specified. When the term "comprises" is used in the present specification and claims, it includes other No element or step is excluded. The terms second, third and similar are used to distinguish between similar elements and are not necessarily used interchangeably. The terms used are not necessarily used to describe sequential or chronological order. are interchangeable under appropriate circumstances, and the embodiments of the invention described herein may be used interchangeably with those described herein. It should be understood that it is possible to operate in other arrangements than those described or illustrated. These terms or definitions are provided solely to aid in the understanding of the present invention. Unless otherwise specified, all terms used herein have the same meaning to those skilled in the art of the present invention. Practitioners are particularly directed to Sambrook et al. for definitions and terminology of the art. al., Molecular Cloning: A Laboratory Man ual, 4 th ed., Cold Spring Harbor Press, Plainsview, New York (2012), and Ausubel et al., Current Protocols in Molecular Biology (Supplement 114), John Wiley & S. ons, New York (2016). The definitions provided herein are , should not be construed to have a range narrower than understood by a person of ordinary skill in the art.
[0032] As used herein, the term "about" refers to an amount that is adequate to perform the disclosed methods. A variation of ±20% or ±10%, more preferably a variation of ±5%, more preferably a variation of ±1% , and more preferably encompasses a variation of ±0.1% and is used herein.
[0033] As used herein, "nucleotide sequence," "DNA sequence," or "nucleic acid molecule" is a polynucleotide of any length, either ribonucleotides or deoxyribonucleotides. This term refers only to the primary structure of the molecule. The term includes double-stranded DNA, and single-stranded DNA, and RNA. When the term "nucleic acid" is used, the 3' and 5' ends of each nucleotide are phosphodiesterases. Single- or double-stranded covalently bound nucleotide sequences linked by ester bonds Polynucleotides are composed of deoxyribonucleotide bases or ribonucleotide salts. Nucleic acids may be synthetically produced in vitro or derived from natural sources. The nucleic acid may be modified DNA or RNA, e.g., methylated DNA or RNA. or RNA, or 5'-capping with 7-methylguaniosine, cleavage and undergoes 3' processing such as polyadenylation and post-translational modifications such as splicing The nucleic acid may further include, for example, hexitol nucleic acid (HNA), cyclohexyl Cetyrene nucleic acid (CEna), threose nucleic acid (TNA), glycerol nucleic acid (GNA), lock It may also include nucleic acids (LNA) and synthetic nucleic acids (XNA) such as peptide nucleic acids (PNA). The size of a nucleic acid referred to herein as a "polynucleotide" is generally determined by the size of a double-stranded polynucleotide. Nucleotides are expressed as the number of base pairs (bp), or single-stranded polynucleotides. In the case of a DNA fragment, it is expressed as the number of nucleotides (nt). 1000 bp or 10 00 nt is equal to 1 kilobase (kb). Polynucleotides less than about 40 nucleotides in length Nucleotides are typically called "oligonucleotides" and are synthesized using the polymerase chain reaction (PCR) The primers may include primers for use in manipulating DNA such as nucleotides R).
[0034] As used herein, a "gene" refers to the promoter region of a gene, as well as the This refers to the genomic sequence (including possible introns). and a spliced messenger operably linked to a promoter sequence. In some cases, the term "cDNA" refers to the cDNA from which it is derived.
[0035] A "coding sequence" is a nucleotide sequence, which is placed under the control of appropriate regulatory sequences. When a sequence is inserted into a cell, it is transcribed into mRNA and / or translated into a polypeptide. The sequence is determined by the translation initiation codon at the 5' end and the translation termination codon at the 3' end. Code sequences include mRNA, cDNA, recombinant nucleotide sequences, or genomic DNA. These may include, but are not limited to, introns, and introns may be present under certain circumstances. .
[0036] The term "amino acid" in the context of this disclosure is used in its broadest sense and refers to an amine (NH2) and carboxyl (COOH) functional groups are attached to the side chains specific to each amino acid (e.g. In some embodiments, the term "amine" refers to an organic compound containing an amine, an amine group ... L α-amino acid refers to a naturally occurring L α-amino acid or residue. The following one-letter and three-letter abbreviations are used herein: A=Ala, C=Cys, D=Asp, E=Glu, F=Phe, G=Gly, H=His, I=Ile, K=Ly s, L=Leu, M=Met, N=Asn, P=Pro, Q=Gln, R=Arg, S= Ser, T=Thr, V=Val, W=Trp, Y=Tyr (Lehninger, A L., (1975) Biochemistry, 2d ed., pp. 7 1-92, Worth Publishers, New York). "Amino Acids" The term further includes D-amino acids, retro-inverso amino acids and amino acid analogs. Chemically modified amino acids, such as norleucine, that are not normally incorporated into proteins Naturally occurring amino acids, as well as those known in the art as characteristic of amino acids such as β-amino acids, For example, peptide compounds may be synthesized from natural Phenylanine or Pro Analogs or mimetics of proline are included in the definition of amino acids. Amino acid mimetics and mimetics are referred to herein as "functional equivalents" of the respective amino acid. Other examples of acids are Roberts and Vellaccio, The Peptide s: Analysis, Synthesis, Biology, Gross a nd Meiehofer, eds., Vol. 5 p. 341, Acade Micronic Press, Inc., NY 1983, incorporated herein by reference. It will be incorporated into the specification.
[0037] The terms "protein," "polypeptide," and "peptide" are further used herein. are used interchangeably to refer to polymers of amino acid residues and variants of amino acid residues and These terms refer to synthetic analogs in which one or more amino acid residues have been replaced by synthetic, non-naturally occurring amino acids. Amino acids (including chemical analogues of the corresponding naturally occurring amino acids) and naturally occurring amino acid polymers Polypeptides also refer to glycosylated, protein-rich polymers of amino acids. These may include proteolytic cleavage, lipidation, signal peptide cleavage, propeptide cleavage, phosphorylation, etc. may undergo maturation or post-translational modification processes, including but not limited to: A "recombinant polypeptide" is a polypeptide that has been produced using recombinant techniques, e.g., recombinantly. Polypeptides produced through expression of recombinant or synthetic polynucleotides. A chimeric polypeptide or a biologically active portion thereof is recombinantly produced. If so, it is also preferably substantially free of culture medium, e.g. The volume of the protein preparation is less than about 20%, more preferably less than about 10%, and most preferably "Isolated" means that the product is free from components that normally accompany it in nature. It means that the material is qualitatively or essentially free of. For example, as used herein, An "isolated polypeptide" is a polypeptide purified from the molecules that naturally surround it. a peptide, e.g., a peptide removed from a molecule present in the production host adjacent to the polypeptide. The term refers to an isolated CsgF peptide (optionally a CsgF protein complex or CsgF peptide). Selectively cleaved CsgF peptides can be produced by amino acid chemical synthesis or recombinantly. The isolated complex can also be generated by chromatin generation. The purified components of the pore and CsgF peptide were then reconstituted in vitro. It can also be produced by recombinant co-expression.
[0038] "Orthologs" and "paralogs" are terms used to describe the ancestral relationships of genes. Paralogs are genes within the same species that originated from the duplication of an ancestral gene. Orthologs are genes from different organisms that originated through speciation and share a common It also comes from ancestral genes.
[0039] A "homologue" or "homologues" of a protein is a homologue of the unmodified or wild-type protein in question. have amino acid substitutions, deletions and / or insertions relative to the unmodified Peptides, oligopeptides, and other compounds with biological and functional activities similar to those of proteins As used herein, "adapted from" includes adducts, polypeptides, proteins and enzymes. The term "amino acid identity" refers to the extent to which sequences are identical amino acid-by-amino acid over a comparison window. Therefore, "percent sequence identity" refers to the degree of identity between two optimally aligned sequences over the comparison window. and compare identical amino acid residues (e.g., Ala, Pro, Ser, Thr, Gly, Val, Leu, Ile, Phe, Tyr, Trp, Lys, Arg, His, Asp, Determine the number of positions where the amino acid sequence (Glu, Asn, Gln, Cys, and Met) occurs in both sequences. The number of matching positions is determined by dividing it by the total number of positions in the comparison frame (i.e., the frame size). The percent sequence identity is calculated by multiplying the result by 100.
[0040] The term "CsgG pore" defines a pore containing multiple CsgG monomers. CsgG monomers were derived from the wild-type monomer from E. coli (SEQ ID NO: 3), E. coli li may be a wild-type homologue of CsgG, for example, as shown in SEQ ID NOs: 68 to 88. Any one of the amino acid sequences, or any variant thereof (e.g., SEQ ID NOs: 3 and 68) Mutant CsgG monomers include those with any one of the following 88 mutants: may also be referred to as modified CsgG monomers or mutant CsgG monomers. The modifications or mutations in may include any one or more of the modifications disclosed herein. or combinations of the above modifications.
[0041] For all aspects and embodiments of the present invention, the CsgG homologue is the CsgG homologue shown in SEQ ID NO: 3. At least 50%, 60%, 70%, 80%, and 9% of live E. coli CsgG CsgG refers to a polypeptide with 0%, 95% or 99% complete sequence identity. The homolog also contains the PFAM domain PF03783, which is characteristic of CsgG-like proteins. Also refers to polypeptides containing CsgG. List of currently known CsgG homologues and CsgG structures Regarding http: / / pfam.xfam.org / / family / PF037 83 Similarly, the CsgG homologous polynucleotide is the wild-type CsgG polynucleotide shown in SEQ ID NO: 1. At least 50%, 60%, 70%, 80%, or 90% against E. coli CsgG , including polynucleotides with 95% or 99% complete sequence identity. Examples of homologues of CsgG shown in SEQ ID NO: 3 have the sequences shown in SEQ ID NOs: 68 to 88.
[0042] The term "modified CsgF peptide" or "CsgF peptide" refers to a peptide having a C-terminus (e.g., CsgF peptides cleaved from the CsgF peptide (e.g., N-terminal fragment) and / or containing the cleavage site. The CsgF peptide is a modified CsgF peptide derived from wild-type E. coli. A fragment of CsgF (SEQ ID NO: 5 or SEQ ID NO: 6), or a fragment of E. coli CsgF It may be a fragment of a wild-type homologue, for example, CsgF (e.g., SEQ ID NOs: 17 to 36). a peptide comprising any one of the amino acid sequences shown), or any variant thereof ( For example, modified to include a cleavage site).
[0043] For all aspects and embodiments of the present invention, the CsgF homologue is the CsgF homologue shown in SEQ ID NO:6. At least 50%, 60%, 70%, 80%, and 9% of the live-type E. coli CsgF It refers to a polypeptide having 0%, 95% or 99% complete sequence identity. In embodiments, the CsgF homologue also contains the PFAM domain, which is characteristic of CsgF-like proteins. The presently known CsgF homologues and polypeptides containing CsgF PF10614 are also included. For a list of CsgF and CsgF structures, see http: / / pfam.xfam.org / / fa mily / PF10614 Similarly, CsgF homologous polynucleotides are described in At least 50%, 60%, and 7% of wild-type E. coli CsgF shown in column 4 Polynucleotides with 0%, 80%, 90%, 95% or 99% complete sequence identity An example of a truncated region of a homologue of CsgG shown in SEQ ID NO: 6 is The array has the arrangement shown in column numbers 17 to 36.
[0044] The term "N-terminal portion of the CsgF mature peptide" refers to the portion of the CsgF mature peptide that is N-terminal to the CsgF mature peptide. The first 60, 50, or 40 amino acid residues (not including the signal sequence) beginning with The CsgF mature peptide refers to a peptide having a corresponding amino acid sequence. It may be mutant (eg, have one or more mutations).
[0045] Sequence identity is to a fragment or portion of a full-length polynucleotide or polypeptide. Thus, a sequence may have only 50% overall sequence identity with a full-length reference sequence. Even in cases where the sequence of a particular region, domain, or subunit is 80%, 90%, or more similar to the reference sequence, or 99% sequence identity, respectively, for CsgG homologs. Homology to the nucleic acid sequence of SEQ ID NO: 1, or SEQ ID NO: 4 for CsgF homologues, is not a single It is not limited purely to sequence identity. Many nucleic acid sequences have clearly low sequence identities. Homologous nucleic acids may exhibit biologically significant homology to one another despite having only a single homolog. The sequences are believed to hybridize with each other under conditions of low stringency (MR Gre en, J. Sambrook, 2012, Molecular Cloning:A Laboratory Manual, Four Edition, Books 1-3 , Cold Spring Harbor Laboratory Press, Col. d Spring Harbor, NY).
[0046] The term "wild-type" refers to a gene or gene product isolated from a naturally occurring source. The wild-type gene is the one most frequently observed in a population and is therefore arbitrarily designed. The "standard" or "wild-type" form of a gene. In contrast, "modified" or "mutated" The terms "variant" or "mutant" refer to a gene or gene product that is a variant of a wild-type gene or gene product. , sequence modifications (e.g., substitutions, truncations, or insertions), post-translational modifications, and / or functional characteristics. refers to a gene or gene product that exhibits a specific trait (e.g., an altered characteristic). can be isolated, which, when compared to the wild-type gene or gene product, It is noteworthy that natural amino acids are distinguished by the fact that their properties have changed. Methods for introducing or substituting are well known in the art. For example, methionine (M) indicates the substitution of methionine at the relevant position in the polynucleotide encoding the mutant monomer. Arginine is obtained by replacing the codon (ATG) with the codon for arginine (CGT). (R). Methods for introducing or substituting unnatural amino acids are also known in the art. For example, unnatural amino acids can be used to express mutant monomers. This was introduced by including synthetic aminoacyl-TrNAs in the IVTT system used. Alternatively, unnatural amino acids can be synthesized by synthesis of those particular amino acids (i.e. In the presence of (unnatural) analogs, E. coli, which is auxotrophic for a specific amino acid, The mutant monomer may be introduced by expressing the mutant monomer in the If the proteins are generated using partial peptide synthesis, they can be synthesized by naked ligation. Conservative substitutions can be made by replacing an amino acid with one that has a similar chemical structure, similar chemical properties, or It is replaced by another amino acid with a similar side chain volume. The introduced amino acid is They may have polarity, hydrophilicity, hydrophobicity, basicity, acidity, neutrality or a charge similar to an acid. Alternatively, a conservative substitution is an aromatic or aliphatic amino acid in place of an existing aromatic or aliphatic amino acid. Other amino acids that are aliphatic may be introduced. Conservative amino acid changes are well known in the art. and may be selected according to the properties of the 20 major amino acids defined in Table 1 below. If the amino acids have similar polarity, this corresponds to the hydropathic scale for amino acid side chains in Table 2. It can also be determined by reference to the [Table 1] [Table 2]
[0047] Mutant or modified proteins, monomers or peptides may be prepared in any manner and in any manner. Mutant or modified monomers or peptides may be chemically modified at one site. Attachment of a molecule to one or more cysteines (cysteine chains), attachment of a molecule to one or more lysines , attachment of the molecule to one or more unnatural amino acids, enzymatic modification of the epitope, or terminal modification. Therefore, chemical modification is preferred. Suitable methods for carrying out such modification include: Mutants of modified proteins, monomers or peptides are well known in the art. They may be chemically modified by the attachment of any molecule, for example, modified proteins, monomers or The peptide mutants can be chemically modified by the attachment of dyes or fluorophores. In some embodiments, the mutant or modified monomer or peptide is a monomer or Interaction between a peptide-containing pore and a target nucleotide or target polynucleotide sequence. The molecule is chemically modified with a molecular adaptor that facilitates its action. The molecular adaptor is a circular molecule, cyclodextrins, hybridizable species, DNA binders or interchelators, Peptides or peptide analogues, synthetic polymers, aromatic planar molecules, positively charged molecules, or Preferably, the hydroxyl group is a small molecule capable of hydrogen bonding.
[0048] The presence of an adaptor allows for the pore and a host of nucleotide or polynucleotide sequences to be Improved guest chemistry and thereby sequenceability of pores formed from mutant monomers The principles of host-guest chemistry are well known in the art. The physical or chemical properties of the pores improve their interaction with the nucleic acid or polynucleotide sequence. The adaptor changes the charge of the pore barrel or channel. or interact specifically with a nucleotide or polynucleotide sequence or may be bound to it, thereby facilitating its interaction with the pore. The modified CsgF peptides provided in the present disclosure are capable of binding to the pore of the protein or enzyme. The modified CsgF peptide may be bound to an enzyme or protein that is in sufficient proximity to the modified CsgF peptide. This may facilitate specific applications of pore complexes containing peptides.
[0049] In this context, proteins also refer to fusion proteins, e.g., proteins produced by recombinant DNA techniques. Proteins may also be conjugated (particularly called gene fusions). They may be bonded or "covalently bonded," which as used herein means stably bonded. By "covalent bond" is meant chemical and / or enzymatic conjugation resulting in a covalent bond.
[0050] Proteins are made up of several polypeptide or protein monomers bound together or interacting with each other. When they interact, they can form a protein complex. "Binding" can be direct or indirect. A direct interaction refers to contact between binding partners, e.g., a covalent bond. Indirect interaction means contact by bonding or bonding. Interactions refer to any interactions that occur within a complex of molecules. This may be entirely indirect through the support of the child, or there may still be direct contact between the parties. It may also be indirect in nature, where the stability is enhanced by the additional interaction of one or more compounds. A "complex" as referred to in this disclosure refers to two or more associated molecules that may have different functions. It is defined as a group of proteins with mutually exclusive functions. The association may be through non-covalent interactions such as hydrophobic or ionic forces, or through disulfides. The bond may be via a covalent bond or bond such as an amino acid bridge or a peptide bond. "Binding" or "coupling" are used interchangeably herein. These are used to refer to bioconjugation between cysteines or between (photo)reactive amino acids, respectively, and are stable. The "cysteine bond" or "reactive or The photoreactive amino acid bond may involve a "photoreactive amino acid bond." Examples of photoreactive amino acids include azidohomoaza. Lanin, homopropargylglycine, homoallergylglycine, p-acetylPhe, p-acetyl These include dihydro-Phe, p-propargyloxy-Phe, and p-benzoyl-Phe. (Wang et al. 2012, in Protein Engineering , DOI: 10.5772 / 28719, Chin et al. 2002, Pr. oc. Nat. Acad. Sci. USA 99(17), 11020-24 ).
[0051] "Biological pores" are membranes that allow the translocation of molecules and ions from one side of a membrane to the other. The pore is a transmembrane protein structure that defines a channel or hole through which ionic species pass. The translocation can be driven by a potential difference applied across the pore. The minimum diameter of the channel through which the ions pass is nanometers (10 -9 meters) In some embodiments, the biological pore is a transmembrane protein pore. The transmembrane protein structure of biological pores is monomeric or oligomeric in nature. Generally, the pore is composed of multiple polypeptide subunits arranged around a central axis. The nanopore is formed by a protein molecule that extends substantially perpendicular to the membrane where the nanopore is located. The number of polypeptide subunits is not limited. Generally, the number of subunits ranges from 5 to a maximum of 30, with the number of subunits being 6-10. Alternatively, the number of subunits may be determined by the number of perfringolysin or related It is not defined in the case of large membrane pores that form protein-lined channels. The portion of the protein subunit within the nanopore generally consists of one or more transmembrane beta-barrels and The amino acid sequence comprises a secondary structure motif which may include an α-helical section and / or an α-helical section.
[0052] The terms "pore," "pore complex," or "complex pore" are used interchangeably herein. The term refers to an oligomeric pore, e.g., a pore containing at least one CsgG monomer (e.g., For example, two or more CsgG monomers, three or more CsgG monomers, etc. CsgG monomer) or CsgG pore (containing CsgG monomer), and CsgF peptide The peptides (e.g., modified or truncated CsgF peptides) associate in a complex and together form a pore or The pore complexes of the present disclosure have the characteristics of biological pores, i.e., It has a typical transmembrane protein structure. The pore complex is a membrane component, a membrane, a cell, or an insulator. When presented in an environment with a rim layer, the pore complexes insert into the membrane or insulating layer, forming "transmembrane pores." A "complex" is formed.
[0053] The pore complexes or transmembrane pore complexes of the present disclosure are suitable for characterizing analytes. In embodiments, the pore complexes or transmembrane complexes described may be, for example, different nucleic acids. It can be used to sequence polynucleotides because it can differentiate between nucleic acids with high sensitivity. The pore complexes of the present disclosure may be substantially isolated, purified, or substantially The pore complexes of the present disclosure may be isolated pore complexes that have been purified in nature. Lipids or other pores associated with the natural state, or other proteins (e.g. CsgE, CsgA, CsgB) or other compounds are not included. A protein is "isolated" or purified when it is sufficiently concentrated from a compartment of a membrane. The conjugates may be substantially For example, the pore complex is less than 10%, less than 5%, less than 2%, or less than 1% other components (such as triblock copolymers, lipids, or other pores) Alternatively, the present invention may be substantially isolated or substantially purified when present in a suitable form. The disclosed pore complexes can be transmembrane pore complexes when present within a membrane. Isolated homo-oligomeric pores derived from CsgG containing a single mutant monomer. We provide a novel pore complex, which also contains mutant forms of the CsgG monomer as its homologue. Alternatively, an isolated pore complex comprising a hetero-oligomeric CsgG pore may be used. provided herein are vectors consisting of mutant and wild-type CsgG monomers or variants thereof. The CsgG pore may be a CsgG pore consisting of a CsgG variant, mutant or homologue in any of the following forms: The isolated pore complexes typically have at least 7, at least 8, at least 9, or at least 10 CsgG monomers, and also 2, 3, 4, 5, 6 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 CsgF peptides, etc. The pore complex may contain any CsG monomer:CsgF peptide ratio. In an embodiment, the CsG monomer:CsgF peptide ratio is 1:1.
[0054] As used interchangeably herein, "constriction," "orifice," "constriction region," and "channel" The "contraction," or "contraction site," is a site that allows ions and target molecules (e.g., polynucleotides or The pore complex allows the passage of nucleotides (including but not limited to individual nucleotides) but not the pore complex. The pore or pore complex has a lumen surface that does not allow the passage of other non-target molecules through the channel. In some embodiments, a constriction refers to an opening defined by a pore or complex of pores. In this embodiment, the constriction is the narrowest opening in the pore that restricts the passage of molecules through the pore. The size of the constriction is typically determined by the size of the nanofibers in nucleic acid sequencing applications. This is the main factor that determines the pore size. If the constriction is too small, the molecules to be sequenced will However, to achieve maximum effect on ion flow through the channel, The constriction should not be too large. For example, the constriction should not be too large, as it may cause the target analyte to be inaccessible to the solvent. Ideally, any constriction should be no wider than the transverse diameter of the analyte. The cross-sectional diameter of the object should be as close as possible to the target. , suitable constriction diameters are in the nanometer range (10-9 meter range). The diameter should be in the range of 0.5-2.0 nm, typically 0.7-1 The constriction of wild-type E. coli CsgG is approximately 9 Å (0.9 The CsgG-like pore and modified CsgF peptide or its homologues have a diameter of 100 nm. The CsgF constriction formed within the pore complex containing the mutant or CsgF constriction was 0.5–2 nanoparticles with diameters in the 0.7-1.2 nm range and therefore suitable for nucleic acid sequencing. are.
[0055] When there are two or more constrictions and the constrictions are spaced apart, each constriction may simultaneously interact with or "read" individual nucleotides within a nucleic acid strand In this situation, the reduction in ion flow through the channel is due to the presence of all nucleotide-containing constricted segments. Therefore, in some instances, the double constriction In certain circumstances, one constriction, or "leading head" The current readout for the "read head" is independent when two such reading heads are present. The contractile region of wild-type E. coli CsgG (SEQ ID NO: 3) may not be determined separately. is a series of tyrosine residues at position 51 (Tyr51) in adjacent protein monomers. and phenylalanine at positions 56 and 55, respectively. The dimers are also formed from phenylalanine and asparagine residues (Phe 56 and Asn 55). The wild-type pore structure of CsgG contains two circular rings (Fig. 1). The two circular G constrictors (referred to herein as "CsgG channel constrictors") Regenerating the ring through recombinant genetic techniques, such as widening, altering, or removing it. engineered to leave a single, distinct leading head. CsgG oligomer contraction within the pore The motif is the wild-type monomeric E. coli CsgG polypeptide shown in SEQ ID NO:3. It is located at amino acid residues 38 to 63 within the domain. Mutants at amino acid residue positions 50–53, 54–56, and 58–59 and wild-type C Positioning of the side chains of Tyr51, Asn55, and Phe56 within the channel of the sgG structure The importance of this and the advantages of modifying or changing the characteristics of the reading head The CsgG pore and modified CsgF peptide or its homologues or mutants were The present disclosure of pore complexes containing variants unexpectedly demonstrates that variants can be used to target CsgG-containing pore complexes. (referred to herein as the "CsgF channel constrictor") to The appropriate additional second reader head is inserted into the pore via complex formation with the F peptide. The additional CsgF channel constriction or reader head forms the CsgG pore. or located adjacent to the contractile loop of the mutated GcsG pore. The additional CsgF channel constriction or reader head acts as a clamping loop for the CsgG pore. 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, Positioned at approximately 10 nm or less, such as 2, 3, 4, 5, 6, 7, 8, and 9 nm The pore complex or transmembrane pore complex of the present disclosure is a pore complex with two leader heads. This includes merging, i.e., channel constrictions positioned in this manner may be combined with other constriction channels. Provides suitable separate reader heads without interfering with the accuracy of the channel reader head Thus, the pore complex is composed of a CsgG mutant pore (incorporated reference WO2 016 / 034591, WO2017 / 149316, WO2017 / 149317, W See International Patent Application No. PCT / GB2017 / 149318 and International Patent Application No. PCT / GB2018 / 051191. Mutations of the wild-type CsgG pore that improve the properties of the pore are listed below. and the wild-type CsgG pore or its homologues, and the modified CsgF peptide or its homologues. The CsgF peptide may comprise a leader head. The stenosis channel has another contraction channel that forms.
[0056] pore The present invention unexpectedly introduces additional channel constrictions or leader heads within the pore complex. The present invention relates to a CsgG pore that forms a complex with an extracellularly located CsgF peptide. Furthermore, the present disclosure provides positional information of the constriction created by the CsgF peptide within the pore complex. The peptide is inserted into the lumen of the CsgG pore, and the contraction site is located in the CsgF protein. Furthermore, the modified or truncated CsgF peptides of the present disclosure are capable of binding to the pore complex. has been shown to be sufficient for the formation and means and methods for biosensing applications - Patent Application 20070122637 The present disclosure provides wild-type and mutant CsgG pores (see, e.g., WO2016 / 034591, WO2017 / 149316, WO2017 / 149317, WO20 17 / 149318 and International Patent Application Publication No. PCT / GB2018 / 051191 or a homolog or mutant thereof, wherein the CsgF peptide is modified or truncated. peptides and their mutants or homologs, collectively, CsgG-like pore These include those that improve the ability of the complex to interact with the analyte (such as a polynucleotide). CsgG-like nanoparticles were synthesized by complex formation with modified or truncated CsgF peptides. Additional constrictions introduced into the pore channel expand the contact surface with the passing analyte, allowing for analysis. It can serve as a second reader head for object detection and characterization. Cells containing mutant CsgG monomers combined with novel mutant or modified forms of F The pores can improve the characterization of analytes such as polynucleotides. This provides a more differentiated and direct relationship between the currents observed when the electrons move through the pores. In particular, having two leader heads stacked with a clear distance between them Thus, the CsgG:CsgF pore complex contains at least one homopolymeric section (e.g., For example, any number of identical nucleotides that exceed the interaction length of a single CsgG leader head. This may facilitate the characterization of polynucleotides containing a sequence of several consecutive copies. The CsgG:CsgF complex has two stacked constrictors at a fixed distance. Small molecule analytes, including organic or inorganic drugs and pollutants, that pass through the pores of the body are classified into two distinct groups: The chemistry of each leader head is independent. can be modified in a number of ways, each with unique interaction properties with the analyte, and therefore This provides additional discrimination during analyte detection.
[0057] In a first aspect, the present invention provides a CsgG pore or a homologue or mutant thereof, or CsgG-like pores and modified CsgF peptides or their homologs or mutants Indeed, the present disclosure relates to an isolated pore complex comprising a modified CsgF peptide (the modified CsgG vectors, including modified CsgG vectors (which may be truncated, mutant and / or variant). In one embodiment, the modified CsgF peptide or homolog thereof The interaction region between the CsgG pore or its homologs or mutants In another embodiment, the pore complex comprises at least one CsgG pore located in a body lumen. and at least one of which is introduced by the CsgF peptide, Contains two or more contractile sites or leader heads that form a complex with the CsgG pore The range of amino acid residues 39 to 64 of SEQ ID NO: 5, or more specifically, the range of amino acid residues 39 to 64 of SEQ ID NO: 5 N-terminal CsgG positions, including those within amino acid residues 49-64, were detected in detectable amounts. It has been shown to enable a stable CsgG:CsgG complex. In one embodiment, CsgF produced by modified CsgF peptides (e.g., those described herein) The constriction is adjacent to or close to the first constriction of the CsgG pore of the pore complex. In the case of sgG or CsgG-like protein pores, the constriction site is bounded by the loop region of the beta strand. It has been determined that the nucleus is formed by the following (see FIG. 1).
[0058] In one embodiment, the modified CsgF peptide is characterized in that the modifications specifically involve the contraction domain and Cs N defined by the constraints that bind the gG monomer or its homologs or mutants Peptide refers to a truncated CsgF protein or fragment, including terminal CsgF peptide fragments. The modified CsgF peptides may further enhance specific properties of the pore complex. The modified CsgF peptide may comprise a mutated or homologous sequence. , the wild-type preprotein (SEQ ID NO: 5) or mature protein (SEQ ID NO: 6) sequence or These modified peptides contain truncated CsgF proteins compared to their homologs. Within the CsgG-like pore formed by G and modified or truncated CsgF peptides , which act as components of the pore complex introducing additional constriction sites or leader heads. Examples of truncated modified peptides are listed below.
[0059] Examples of homologues of modified CsgF peptides are, for example, those determined in Example 3 and shown to be homologous to different bacterial strains. CsgF-like proteins or CsgF peptides containing the same or similar constrictor domains have been identified. This may be useful in the use of similar pore complexes. Various CsgG pores were prepared for use in combination with wild-type or mutant CsgG pores. Conserved structural features and CsgG-binding elements in CsgF peptides derived from gF homologues This includes complexes of the CsgG pore with non-cognate CsgF, which The parent CsgG homologues from which CsgG and CsgF are derived originate from the same operon, bacterial species or strain. It means there is no need.
[0060] In an alternative embodiment, the CsgG pore within the pore complex is not a wild-type pore, but has pore-specific The CsgG pore or its homologues may also contain mutations or modifications to increase the activity of the CsgG pore. and a modified CsgF peptide or a homolog thereof. The pore complex can be formed by the wild-type form of the CsgG pore or by specific amino acids. Further modifications have been made within the CsgG pore, such as by directed mutagenesis of amino acid residues. This further improves the desirable properties of the CsgG pore for use within the pore complex. For example, in embodiments of the present invention, mutations may affect the number, size, shape, placement or structure of constrictions within a channel. Pore complexes containing modified mutant CsgG pores are intended to alter the structure or orientation of the CsgG pore. The method comprises inserting, substituting, and / or deleting specific targeted amino acid residues within a polypeptide sequence. The oligomeric CsgG pore may be prepared by known genetic engineering techniques that result in In this case, mutations may occur within the polypeptide subunit of each monomer, or within any of the monomers. In one embodiment of the present invention, the polymer may be made in any one or all of the monomers described. The mutations made are made to all monomer polypeptides within the oligomeric protein structure. Suitably, the mutant CsgG monomer is prepared by adding a CsgG monomer having a sequence similar to that of the wild-type CsgG monomer. These are monomers that vary from the original sequence and maintain the ability to form pores. Methods for confirming the ability of a mutant monomer to bind to a target molecule are well known in the art. have reported wild-type and mutant CsgG pores (e.g., WO2016 / 034591, W O2017 / 149316, WO2017 / 149317, WO2017 / 149318 and disclosed in International Patent Application Publication No. PCT / GB2018 / 051191) or its derivatives modified or truncated CsgF peptides and mutants or homologs thereof In combination with the nucleotides, the CsgG-like pore complex binds the analyte (e.g., polynucleotide). Mutant CsgG pores include those that improve the ability to interact with one or more protons. The CsgG pore may contain the same monomer or two or more different monomers. The monomer may be a homopolymer, including a heteropolymer containing any of the monomers. The mutant may have one or more of the mutations described below in combination.
[0061] The nanopore complexes comprising modified CsgF peptides are, in certain embodiments, nanopores comprising wild-type CsgF proteins. The wild-type CsgF protein shown in SEQ ID NO: 6 contains only the N-terminal fragment or truncation of the protein. However, the modified CsgF peptides are different from the CsgG pore and modified As an amino acid substitution in the pore formed by a complex containing the decorated CsgF peptide Mutations may additionally or alternatively allow for better second contraction sites. The mutant monomer may be a naturally mutated CsgF peptide. In addition to improving the function of the complex to include a target head, the complex When used for DNA sequencing, i.e., improved polynucleotide capture and nuclease When improved hit discrimination is demonstrated, it is possible to improve polynucleotide reading performance. In particular, pores composed of mutant peptides can more easily accept nucleotides and Furthermore, pores constructed from the mutant peptides capture polynucleotides in the current range. may show an increased range, which facilitates discrimination between different nucleotides and This reduces the signal-to-noise ratio. When a nucleotide translocates through a pore composed of This means that as the polynucleotide moves through the pore and the polynucleotide sequence, This facilitates the identification of a direct relationship between the observed currents. The resulting pores may exhibit increased throughput, e.g., polynucleosides. This allows for the specific characterization of the analyte using the pores. Pores constructed from mutant peptides insert more easily into membranes. or a simpler way to keep additional proteins in the vicinity of the pore complex. It can be provided.
[0062] In another embodiment, the CsgF contraction site provided within the pore complex of the invention is 0.5 n diameters in the range of 1.5 to 2.0 nm, making them suitable for nucleic acid sequencing as described above. A pore composite is provided.
[0063] The pore may be stabilized by covalent attachment of a CsgF peptide to the CsgG pore. The bond may be, for example, a disulfide bond or may be click chemistry. The CsgF peptide and the CsgG pore are, for example, 1 and 153, 4 and 133, and 5 and 1. 36, 8 and 187, 8 and 203, 9 and 203, 11 and 142, 11 and 201, 12 and 14 9, 12 and 203, 26 and 191, and 29 and 144, respectively. may be covalently linked via one or more residues at positions corresponding to pairs of positions in SEQ ID NO:3 good.
[0064] In the pore, the interaction between the CsgF peptide and the CsgG pore is, for example, 1 and 15 3, 4 and 133, 5 and 136, 8 and 187, 8 and 203, 9 and 203, 11 and 142, 1 I took 1 and 201, 12 and 149, 12 and 203, 26 and 191, and 29 and 144. Hydrophobic phases at positions corresponding to one or more pairs of positions in SEQ ID NO: 6 and SEQ ID NO: 3, respectively The cleavage may be stabilized by reciprocal or electrostatic interactions.
[0065] Residues of CsgF and / or CsgG at one or more of the above positions may be involved in the CsgF binding within the pore. It may be modified to strengthen the interaction between G and CsgF.
[0066] In one embodiment, the pores of the present invention may be isolated, substantially isolated, or purified. The pores of the present invention may be purified or substantially purified. It is isolated or purified if it is completely free of any component. It is substantially isolated when mixed with a carrier or diluent that does not interfere with its use. For example, the pores may contain less than 10%, less than 5%, less than 2%, or less than 1% of other components (e.g., When present in a form containing a pore, such as a triblock copolymer, lipid, or other micropores, Alternatively, the pores of the present invention are present in a membrane. Suitable membranes are discussed below.
[0067] The pores of the present invention may be present as individual or single pores. The pores may be present in a homogeneous population or a heterogeneous population of two or more pores.
[0068] CsgF peptide The second aspect of the present invention is a novel modified CsgF monomer (peptide) or CsgF protein. or modified or truncated peptides of CsgF homologs or mutants. These novel modified CsgF peptides incorporate a second or additional leader head. The modification or truncation may be used in a pore complex to inhibit the activity of wild-type CsgF or or a mutant or homologous fragment of the CsgF protein, Terminal fragments are more preferred.
[0069] Mature CsgF (shown in SEQ ID NO: 6) contains the "CsgF contractile peptide" (FCP), "neck " area and "head" area (shown in Figures 4 and 5) The "head" region of the CsgF peptide acts as the leader head of the pore described herein. The "head" region of the CsgF peptide is also called the "C-terminal head domain." You may be called.
[0070] FCP forms a contact area with the CsgG β-barrel, where an additional constriction occurs. The β-barrel protrudes from the β-block region. In the CsgG:CsgF oligomer, this is the FC Form a thin-walled hollow tube connecting P to the spherical head region.
[0071] Multiple sequence alignment (Fig. 8), co-refinement experiments (Fig. 9), and at 3.4 Å resolution (Fig. 11) Based on cryoEM reconstruction of the CsgG:CsgF complex in The neck and head regions can be defined as three consecutive residues in mature CsgF. do.
[0072] The FCP spans approximately residues 1 to 35 of mature CsgF (SEQ ID NO: 6). When comparing CsgF orthologues, these form the most conserved regions of the protein (Figure 1). 8, Figure 10). CryoEM 3D reconstructions show that the CsgG transmembrane hairpin TM1 (SEQ ID NO: 1) 3) and TM2 (residues 134-154 of SEQ ID NO:3) (Fig. 1E, 11) through non-covalent contact with the well-defined CsgG β-barrel interior. The FCP forms are shown in Fig. 1 (TM1 and TM2 are the structures of Goyal P et al., 201 4). In the reconstitution, nine copies of FCP are integrated into the CsgG oligomer (nine monomers). (including the nucleotide sequence 1444-1445) and collectively form residues 46-48 of mature CsgG (SEQ ID NO: 3, Figure 1E, Figure 11). Approximately 2 nm above the top of the CsgG constriction, formed by ~61 continuous loops This results in a high additional shrinkage.
[0073] The cryoEM 3D reconstruction also showed that the CsgF N-terminal residues are also located at the bottom of the CsgG β-barrel. The β-barrel is shown to exit at residue 32, with the β-barrel near the top or bottom (depending on the orientation). This is based on MD simulations showing the average contact times of residue pairs at the CsgG:CsgF binding interface. The cryoEM structure and MD simulations are in good agreement with the previous results (Table 4). Residues 33-34 are the well (but not strictly) conserved Pro residues (SEQ ID NO:6 Pro 35 of CsgF is located outside the CsgG β-barrel and transfers to the CsgF neck region. The CsgF neck has three CsgG:CsgF structures, indicating conformational flexibility. The CsgF neck is not resolved at the atomic level in the D reconstruction. Based on the structure and secondary structure predictions, the sequence spanning from residue 36 to approximately residue 50 (SEQ ID NO: 6) The CsgF head region forms the C-terminal part of CsgF and is predicted to be 51 to the C-terminus of CsgF. This region appears to give rise to a globular structure that appears to occlude the CsgG:CsgF channel. Multiple sequence alignment of CsgF orthologues revealed that C The sgF neck is the least conserved region, and one ortholog and different orthologs, suggesting that their lengths may vary (Fig. 8).
[0074] The CsgF peptides forming part of the present invention may lack a C-terminal head or may comprise a C-terminal head and and part of the neck domain of CsgF (e.g., the truncated CsgF peptide It may contain only a portion of the neck domain, or the C-terminal head and neck domain of CsgF. The CsgF peptide is a truncated CsgF peptide lacking the CsgF neck domain. For example, the CsgF peptide may lack a portion of, for example, the N of the neck domain. From amino acid residue 36 (see SEQ ID NO: 6) at the terminus (e.g., residue 1 of SEQ ID NO: 6) From residues 36-40, 36-41, 36-42, 36-43, 36-45, and 36-46 36-50 or 36-60), and may include a portion of the neck domain. Preferably, the tide comprises a CsgG binding region and a region that forms the constriction within the pore. The CsgG binding region is typically derived from a CsgF protein (SEQ ID NO: 6 or a homologue from another species). The amino acid sequence comprises residues 1-8, and / or residues 29-32 of the amino acid sequence (amino acid sequence), and may contain one or more modifications. The region forming the constriction within the pore is typically a CsgF protein (SEQ ID NO: 6 or another Residues 9-17 of the ribosomal amino acid sequence of the present invention are: Conserved motif N9PXFGGXXX 17 Residues 9-2 form a rotation region. 8 forms an α-helix. X 17 (N17 in SEQ ID NO: 6 is the CsgF contractile site in the pore. The CsgF constriction region also forms the apex of the constriction region, which corresponds to the narrowest part of the CsgF. G β-barrel, primarily contacting residues 9, 11, 12, 18, 21 and 22 of SEQ ID NO:6 Stabilize.
[0075] CsgF peptides typically contain 28-50 amino acids, e.g., 29-49, 30-45, or The CsgF peptide has a length of 29 to 35 amino acids, Alternatively, it preferably contains 29 to 45 amino acids. The CsgF peptide contains all or part of the FCP corresponding to groups 1 to 35. If shorter than P, the truncation is preferably made at the C-terminus.
[0076] The CsgF fragment of SEQ ID NO: 6 or a homologue or mutant thereof may be any of 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 4 0, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53 , 54 or 55 amino acids in length.
[0077] The CsgF peptide can be any one of residues 1 to 25-60 (e.g., SEQ ID NO: 6) from 27 to 50, for example, from 28 to 45) of the amino acid sequence of SEQ ID NO: 6, or 6 homologues, or variants of either thereof. The CsgF peptide is SEQ ID NO: 39 (residues 1-29 of SEQ ID NO: 6) or a homolog thereof. or may comprise a variant.
[0078] Examples of such CsgF peptides include SEQ ID NO: 15 (residues 1-34 of SEQ ID NO: 6), SEQ ID NO: 16 (residues 1-34 of SEQ ID NO: 6), SEQ ID NO: 17 (residues 1-34 of SEQ ID NO: 17), SEQ ID NO: 18 (residues 1-34 of SEQ ID NO: 18), SEQ ID NO: 1 No. 54 (residues 1 to 30 of SEQ ID NO: 6), SEQ ID NO: 40 (residues 1 to 45 of SEQ ID NO: 6), or SEQ ID NO: 55 (residues 1-35 of SEQ ID NO: 6) and any homologues or variants thereof Other examples of CsgF peptides include those consisting essentially of or consisting of variants of the following: SEQ ID NO: No. 7, SEQ ID NO. 8, SEQ ID NO. 9, SEQ ID NO. 10, SEQ ID NO. 11, SEQ ID NO. 12, SEQ ID NO: 13, SEQ ID NO: 14, SEQ ID NO: 16.
[0079] In the CsgF peptide, one or more residues, e.g., SEQ ID NO: 15, SEQ ID NO: 39, SEQ ID NO: SEQ ID NO: 40, SEQ ID NO: 54, or SEQ ID NO: 55 may be modified.
[0080] For example, the CsgF peptide may include modifications at positions corresponding to one or more of the following in SEQ ID NO:6: Also acceptable are: G1, T4, F5, R8, N9, N11, F12, A26 and Q29.
[0081] The CsgF peptide contains cysteines, hydrophobic amino acids, charged amino acids, and non-native reactivity. The amino acid, or photoreactive amino acid, may be, for example, a nucleotide corresponding to one or more of the following of SEQ ID NO: 6: The modifications may be introduced at positions: G1, T4, F5, R8, N9, N11, F1 2, A26 and Q29.
[0082] For example, the CsgF peptide may include modifications at positions corresponding to one or more of the following in SEQ ID NO:6: The CsgF peptide may be: N15, N17, A20, N24, and A28. Modifications may be included at the position corresponding to D34 to stabilize the gG-CsgF complex. In certain embodiments, the CsgF peptide includes one or more of the following substitutions: N15S / A / T / Q / G / L / V / I / F / Y / W / R / K / D / C, N17S / A / T / Q / G / L / V / I / F / Y / W / R / K / D / C, A20S / T / Q / N / G / L / V / I / F / Y / W / R / K / D / C, N24S / T / Q / A / G / L / V / I / F / Y / W / R / K / D / C, A28S / T / Q / N / G / L / V / I / F / Y / W / R / K / D / C and D34F / Y / W / R / K / N / Q / C. The CsgF peptide may be, for example, one or more may include the following substitutions: G1C, T4C, N17S, and D34Y or D34N.
[0083] CsgF peptides are produced by enzymatically cleaving long proteins such as full-length CsgF. Cleavage at specific sites can be achieved by incorporating enzyme cleavage sites at appropriate positions. This can be directed by modifying long proteins such as full-length CsgF. Examples of CsgF amino acid sequences modified to include a cleavage site are shown in SEQ ID NOS: 56-67. After cleavage, all or part of the added enzyme cleavage site is associated with CsgG. The CsgF peptide can be present in the CsgF peptide, which is then split into two to form a pore. , and may further comprise all or part of an enzyme cleavage site at its C-terminus.
[0084] Some examples of suitable CsgF peptides are shown in Table 3 below. [Table 3] JPEG2025148349000004.jpg208162
[0085] In certain embodiments, the CsgF fragment has the amino acid sequence of SEQ ID NO: 39, or a mutant thereof. In particular, SEQ ID NO: 39 is a sequence of the mature CsgF peptide (SEQ ID NO: 6). In another embodiment, the modified CsgF peptide of the invention comprises the first 29 amino acids. In particular, SEQ ID NO: 40 is a truncated peptide containing the mature CsgF peptide ( It contains the first 45 amino acids of SEQ ID NO: 6), in particular the CsgF contraction site and the CsgG The binding site for CsgF is located within the N-terminal CsgF peptide region, and is located between amino acids 39 and 40 of SEQ ID NO:5. 64 (present in SEQ ID NO: 39 and SEQ ID NO: 40), or in particular amino acid 4 of SEQ ID NO: 5 9 to 64 (present in SEQ ID NO: 40 but not in SEQ ID NO: 39, SEQ ID NO: 39 The latter fragment, encoded by , which are further characterized in that they confer greater stability to the complex. forms complexes with the CsgG pore or its homologs or mutants in vivo The peptide, or the N-terminal fragment or the contraction site region, is used to enable synthesis of the peptide. Proteolytic cleavage of the protein to peptides provides modification of the CsgF protein. In one embodiment, a modified CsgF peptide comprising SEQ ID NO: 37 or SEQ ID NO: 38 is provided. Finally, there are further limitations regarding CsgF homologous peptides, especially the contractile domain (FCP). Identification of CsgF homologous peptides located within the isolated complex Also provided are modified CsgF peptide homologs that may form part of the see).
[0086] A further embodiment relates to a modified or truncated CsgF peptide comprising SEQ ID NO: 15, SEQ ID NO: 15 results in an isolated pore complex containing CsgF channel contraction. To achieve this, we investigated the in vivo pore size distribution of a complex containing CsgG or its homologues and a modified CsgF peptide. Some of the regions from the CsgG binding site and / or contraction site sufficient for in vitro reconstitution Another embodiment includes a region of the CsgF protein comprising several residues. The modified CsgF peptide is described, which comprises an N-terminal fragment of the CsgF protein. and two additional amino acids (KD), which contribute to the solubility and This increases the surface area and stability of the complex pore and allows for in vitro reconstitution of the complex pore. The modified CsgF peptide is SEQ ID NO: 15, SEQ ID NO: 16 or a homologue or variant thereof. Further provided are embodiments comprising variants, wherein the modified CsgF peptide further comprises The modified CsgF peptide corresponding to SEQ ID NO: 15 or 16 has further mutations. Within the region of the peptide, for example, 40 for SEQ ID NO: 15 or SEQ ID NO: 16, respectively %, 50%, 60%, 70%, 80%, 85%, 90% amino acid identity, etc. 35% amino acid identity is still maintained. Further provided are embodiments comprising SEQ ID NO: 15, SEQ ID NO: 16 or a homologue or mutant thereof. wherein the modified CsgF peptide is further mutated, Within the region of the modified CsgF peptide corresponding to SEQ ID NO: 15 or 16, respectively, 15, or at least 40%, 45%, 50%, 60%, 70% relative to SEQ ID NO: 16 , 80%, 85%, and 90% amino acid identity are still maintained. The regions are intended to alter and / or improve the properties of the CsgF contraction site, as described above. In another embodiment, this allows for more accurate target analysis, for example. , modified CsgF peptides are disclosed, wherein SEQ ID NO: 39, SEQ ID NO: 40, SEQ ID NO: SEQ ID NO: 54, or one or more positions in the region containing SEQ ID NO: 55 are modified, and the mutation The variant is directed to a region containing SEQ ID NO: 39, SEQ ID NO: 40, SEQ ID NO: 54, or SEQ ID NO: 55. In the corresponding peptide fragment, SEQ ID NO: 39, SEQ ID NO: 40, SEQ ID NO: 54, or the sequence At least 35% amino acid identity with sequence number 55, or 40%, 50%, 60%, 70%, Retain 80%, 85%, 90%, and 95% amino acid identity.
[0087] Therefore, a further embodiment of the present invention provides a CsgG pore or a homologue or mutant thereof. and variants thereof, and single peptides containing modified CsgF peptides or their homologs or mutants. The modified CsgF peptide is related to a second aspect of the invention. is defined as described in
[0088] Additional embodiments relate to isolated pore complexes, wherein at least one monomer The CsgG pore and the modified CsgF peptide are linked via a covalent bond. The covalent bond or bond is one example possible via a cysteine bond, The sulfhydryl side group of the amino acid is covalently bonded to another amino acid residue or moiety. In nature, this is achieved through interactions between covalent, non-natural (photo)reactive amino acids. (Photo)reactive amino acids are artificial derivatives of natural amino acids that can be used to crosslink protein complexes. Refers to analogs of proteins and peptides in vivo or in vitro Photoreactive amino acid analogs in common use are leucine and methyl. Thionine, and para-benzoyl-phenyl-alanine, and azidohomoalanine , homopropargylglycine, homoallergylglycine, p-acetyl-Phe, p-azide -Phe, p-propargyloxy-Phe, and p-benzoyl-Phe It is a reactive diazirine analogue (Wang et al. 2012, Chin et al. al. 2002). When exposed to ultraviolet light, these are activated and form photoreactive amino acids. The analog binds to interacting proteins within a few angstroms of each other, but However, the positions in the CsgG monomer where the covalent bond can occur are determined by the presence of the modified CsgF peptide. As shown in Figure 1, some amino acids provide covalent bonding. positions 132, 133, 136, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 140, 142, 144, 145, 147, 149, 151, 153, 155, 183, 185, 187, 189, 191, 201, 203, 205, 207 or 209 .
[0089] Another aspect of the invention relates to a construct comprising the modified CsgF peptide, wherein the peptide is covalently linked to the CsgF peptide. A "construct" is a modified CsgF and / or CsgG or homologue thereof. In other words, the construct comprises two or more covalently linked monomers derived from In another aspect, the present invention provides a method for preparing a polymeric polymer comprising at least one component of the present invention. The pore composite also includes sufficient components to form pores as needed. For example, an octameric pore may consist of (a) tetrameric pores each containing two monomers; (b) two constructs each containing four monomers; (c) two monomers and six monomers that do not form part of the construct, or (d) one construct containing One or two CsgF monomers and six or seven CsgG monomers in the adult or even (e) another construct containing only CsgG monomers. For example, a nonamer construct may include a construct having CsgF and CsgG monomers in addition to the CsgF and CsgG monomers. The same and additional possibilities are offered for pores. Other combinations of constructs and monomers are possible. Combinations can be envisioned by those skilled in the art. One or more of the constructs of the present invention may be used for sequencing, polymorphism, and other purposes. These can be used to form pore complexes for characterization of oligonucleotides, etc. The body has at least two, at least three, at least four, at least five, at least 6, at least 7, at least 8, at least 9 or at least 10 things Preferably, the construct comprises two monomers. They may be the same or different, and may be CsgF, CsgG, or a CsgG / CsgF fusion monoclonal antibody. The amino acid sequence may be a nucleotide sequence or a homologue thereof, or any combination thereof.
[0090] Another embodiment is a modified CsgF peptide of the invention, or a homologue or mutant thereof. A polynucleotide or nucleic acid molecule encoding the construct, or a polynucleotide encoding the construct described above. Concerning nucleotides.
[0091] Particular embodiments include a method for preparing an isolated pore complex according to the first and second aspects of the invention and a membrane comprising the steps of: The present invention relates to an isolated transmembrane pore complex comprising a component. The combination is directly applicable to molecular sensing uses such as nucleic acid sequencing. Modified CsgG / CsgF biological activity with the isolated pore complex of the invention described herein. A membrane composition is provided that includes a pore and a membrane, membrane component, or insulating layer. The morphology may be an isolated transmembrane pore complex comprising an isolated pore complex according to the present invention, or a membrane pore complex comprising an isolated transmembrane pore complex. and the components of.
[0092] The CsgG:CsgF complex is highly stable, but when CsgF is cleaved, CsgG The stability of the CsgF complex is reduced compared to that of a complex containing full-length CsgF. For example, cysteine residues may be introduced at the positions specified herein, followed by disulfide bond formation. Pore complexes can be created between CsgG and CsgG to make the complex more stable. can be prepared by any of the methods previously described, and disulfide bond formation can be achieved by It can be induced by using an oxidizing agent (eg, copper-orthophenanthroline). Other interactions (e.g., hydrophobic interactions, charge-charge interactions / electrostatic interactions) may also contribute to the synthesis of It can be used in place of stain interactions at those locations.
[0093] In another embodiment, unnatural amino acids may be incorporated at these positions. In the present study, the covalent bond is created via click chemistry. For example, azide or azide unnatural amino acids bearing a dibenzocyclooctyne (DBCO) group and / or Or a bicyclo[6.1.0]nonyne (BCN) group is introduced into one or more of these positions. It is possible.
[0094] Such stabilizing mutations include, for example, CsgG and and / or any other modification to CsgF.
[0095] The CsgG pore has been modified to facilitate attachment to the CsgF peptide. , 6, 7, 8, 9 or 10, etc., at least one CsgG monomer. For example, cysteine residues are located at positions 132, 133, 136, 138, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158 42, 144, 145, 147, 149, 151, 153, 155, 183, 185, 1 One or more corresponding to 87, 189, 191, 201, 203, 205, 207 and 209 and / or contacts CsgF to facilitate covalent binding to CsgG. The amino acid sequence can be introduced into any one of the predicted positions listed in Table 4. Alternatively or in addition to covalent bonds, the pores may be bound by hydrophobic or electrostatic interactions. To promote such interactions, the amino acid sequence at position 132 of SEQ ID NO: 3 may be stabilized by the amino acid sequence at position 132 of SEQ ID NO: 3. 133, 136, 138, 140, 142, 144, 145, 147, 149, 151, 153, 155, 183, 185, 187, 189, 191, 201, 203, 205, at positions corresponding to one or more of 207 and 209 and / or contacting CsgF. Non-natively reactive or photoreactive at any one of the positions listed in Table 4 where Sexual amino acids.
[0096] The CsgF peptide can be modified to facilitate attachment to the CsgG pore. The stain residue is at position 1, 4, 5, 8, 9, 11, 12, 26 or 29 of SEQ ID NO:6. at one or more corresponding positions and / or in contact with CsgF to bind covalently to CsgG. The cysteine may be introduced into any one of the positions listed in Table 4 that are predicted to enhance the Alternatively or additionally to covalent bonding via amino acid residues, the pores may be formed by hydrophobic interactions or It can be stabilized by electrostatic interactions. To promote such interactions, SEQ ID NO:6 at positions corresponding to one or more of positions 1, 4, 5, 8, 9, 11, 12, 26 or 29 of and / or any one of the positions listed in Table 4 that are predicted to contact CsgF , non-native reactive or photoreactive amino acids.
[0097] Preferred exemplary CsgF peptides include the following mutations relative to SEQ ID NO:6: N15X1 / N17X2 / A20X3 / N24X4 / A28XX5 / D34X6, during the ceremony , X1 is N / S / A / T / Q / G / L / V / I / F / Y / W / R / K / D / C, X2 is N / S / A / T / Q / G / L / V / I / F / Y / W / R / K / D / C, X3 is A / S / T / Q / N / G / L / V / I / F / Y / W / R / K / D / C, X4 is N / S / T / Q / A / G / L / V / I / F / Y / W / R / K / D / C, X5 is A / S / T / Q / N / G / L / V / I / F / Y / W / R / K / D / C and X5 are D / F / Y / W / R / K / N / Q / C Mutants at positions N15, N17, A20, N24, and A28 are contraction mutations. The mutant at position 34 inhibits the interaction of pfCsgF with the bottom of the CsgG pore. This affects the interaction and stabilizes it.
[0098] CsgG pore The CsgG pore is a homo-oligomeric pore comprising the same mutant monomer of the invention. The CsgG pore may be, for example, a CsgG pore comprising at least one mutant moiety disclosed herein. The pore may be a hetero-oligomeric pore derived from CsgG, including a CsgG oligomer.
[0099] The CsgG pore can contain any number of mutant monomers. The pore generally contains 7, 8, or 9 mutant monomers. at least 7, at least 8, such as 1, 9, or 10 mutant monomers; Contains at least 9, or at least 10, identical mutant monomers. Preferably, the pore contains eight or nine identical mutant monomers.
[0100] In a preferred embodiment, the hetero-oligomeric CsgG pore (10 of the monomers, 9 All monomers within the range of 1, 8 or 7 are mutant monomers disclosed herein. mers, where at least one of them is different from the others. They can be different from each other.
[0101] The mutant monomers in the CsgG pore are all approximately the same length or The barrels of the mutant monomers in the pores of the present invention are preferably of approximately the same length. Preferably, the length is equal to or greater than the number of amino acids. Or it can be measured in units of length.
[0102] The mutant monomer may be a variant of SEQ ID NO: 3. Amino acid sequence of SEQ ID NO: 3 Over the entire length of the sequence, variants should have at least 50% amino acid identity to the sequence. Preferably, the variant is homologous to the amino acid sequence of SEQ ID NO: 3 throughout the entire sequence. at least 55%, at least 60%, at least 65% based on amino acid identity to the , at least 70%, at least 75%, at least 80%, at least 85%, less Preferably, both sequences have 90% homology, and preferably have at least 95%, 97% or 99% homology. For example, it is more preferable to have a value of 100 or more, for example, 125, 150, 175 or more. or over a stretch of 200 or more contiguous amino acids, at least 80%, e.g., at least may also have 85%, 90% or 95% amino acid identity ("hard homology").
[0103] The CsgG monomer is highly conserved (Figures 45-4 of WO2017 / 149317). 7). Furthermore, from knowledge of the mutations relative to SEQ ID NO:3, It is possible to determine the equivalent positions of CsgG monomer mutations other than no. 3.
[0104] Thus, variants of the sequence shown in SEQ ID NO: 3 and those set forth in the claims and elsewhere herein are Reference is made to mutant CsgG monomers containing the specific amino acid mutations described herein. The variants of the sequences shown in SEQ ID NOs: 68 to 88 and their corresponding amino acid mutations are Also encompassed are mutant CsgG monomers containing the sequence shown in SEQ ID NO: 3. and CsgG monomers containing the specific amino acid mutations described herein. Reference to compositions, pores or methods involving the use of pores in the context of the sequences disclosed above. Numbers and corresponding amino acid mutation variants in mutant CsgG monomers As will be further understood, the present invention encompasses any related compositions, pores, or methods. Other mutant CsgG monomers not explicitly identified by the specificity analysis show conserved regions. - even.
[0105] Standard methods in the art may be used to determine homology. For example, The package can be used to calculate homology, for example, using its default settings. The BESTFIT program (Devereux et al. (1984) Nucleic Acids Research 12, p387-395). P.I. The LEUP and BLAST algorithms search for homology or line up sequences (equivalent residues) or to identify the corresponding sequences (typically with their default settings). For example, see Altschul SF (1993) J Mol E vol 36:290-300, Altschul, SF et al (199 0) J Mol Biol 215:403-10. BLAST analysis was performed. The software used is provided by the National Center for Biotechnology Information. Information(http: / / www.ncbi.nlm.nih It is publicly available at www.gov / .
[0106] SEQ ID NO: 3 is the wild-type CsgG monomer, expressed in E. coli strain K-12 substr. The variant of SEQ ID NO: 3 is derived from MC4100. Preferred CsgG homologues are shown in SEQ ID NOs: 68 to 88. Variants may include substitutions of the sequence Compared to SEQ ID NO: 3, the combination of one or more of the substitutions present in SEQ ID NOs: 68 to 88 is included. For example, a sequence number different from SEQ ID NO: 3 and any one of SEQ ID NOs: 68 to 88 can be used. Mutations may be made at any one or more of positions 3. The amino acid sequence of SEQ ID NO: 3 is replaced with the corresponding amino acid sequence of any one of SEQ ID NOs: 68 to 88. Alternatively, the amino acid at any one of these positions may be substituted with an amino acid at any one of these positions. The mutations may be substitutions with any amino acid, or may be deletion or insertion mutations (1 to 10 amino acid deletions or insertions, 2 to 8 or 3 to 6 amino acid deletions or insertions, etc. In addition to the mutations disclosed herein, SEQ ID NO: 3 and SEQ ID NOs: 66 to 88 may also be used. Preferably, amino acids that are conserved among all of the above are present in the variants of the present invention. However, among these positions that are conserved between SEQ ID NO: 3 and all of SEQ ID NOs: 66 to 88, Conservative mutations may be made at any one or more positions in the
[0107] The present invention provides a method for the preparation of CsgG monomers having the sequence Any one or more of the amino acids described herein as substituted at the specific positions of number 3 Pore-forming CsgG mutant monomers containing acids are provided. Corresponding positions are described in the art. This can be determined by standard techniques, such as the PILEUP algorithm and BL The AST algorithm was used to align the sequence of the CsgG monomer with SEQ ID NO: 3 and The corresponding residue can be identified by
[0108] Pore-forming mutant monomers are generally identical to the CsgG monomer having the sequence of SEQ ID NO:3. retains the ability to form the same 3D structures as wild-type CsgG monomers, including the same 3D structure as C The 3D structure of sgG is known in the art and is described, for example, in Goyal et al. 4) Nature 516(7530):250-3. The improved properties conferred to the CsgG mutant monomer by the mutation are retained. In addition to the mutations described herein, any number of mutations may be introduced into the wild-type CsgG sequence. It can be made.
[0109] Typically, CsgG monomers form a structure containing three α-helices and five β-sheets. The mutant retains the ability to form at least the first alpha helix (S in SEQ ID NO: 3). within the region of CsgG that is N-terminal to the second alpha helix (starting at SEQ ID NO: 63) and No. 3) and the loop between the second alpha helix and the first beta sheet (SEQ ID NO: 3) and the fourth and fifth sheets (S1 73~R192 and R198~T107) and the loop between the fourth and fifth β-sheets (F193 to Q197 of SEQ ID NO: 3), the CsgG monomer is located in the transmembrane pore (here, transmembrane The pore can be manipulated without affecting its ability to form a translocating pore (a membrane pore capable of translocating polypeptides). Thus, to form a pore through which a polynucleotide can be translocated, Further abruption of these regions within any CsgG monomer was not observed without affecting its potency. It is envisioned that mutations can be made. The mutants have pores that can translocate polynucleotides. without affecting the ability of the monomer to form any α-helix (S6 in SEQ ID NO: 3) 3 to R76, G85 to A99, or V211 to L236), or any β-sheet (sequence Row number 3: I121~N133, K135~R142, I146~R162, S173~ It is expected that it may be produced in other regions, such as R192 or R198-T107. Furthermore, deletion of one or more amino acids creates a pore through which polynucleotides can translocate. Any linking of the α-helix and β-sheet without affecting the ability of the monomer to within the loop region and / or within the N- and / or C-terminal regions of the CsgG monomer It is expected that it can also be produced.
[0110] Amino acid substitutions may be made in addition to those mentioned above, for example, up to 1, 2, 3, 4, 5, 10, 20 or more. Conservative substitutions may be made in the amino acid sequence of SEQ ID NO: 3, such as 30 or 30 substitutions. amino acids with other amino acids of similar chemical structure, similar chemical properties, or similar side chain volume. The introduced amino acid has similar polarity, hydrophilicity, hydrophobicity, and It can be basic, acidic, neutral, or charged. Alternatively, conservative substitutions can be made to replace existing aromatic amino groups. In place of an aromatic or aliphatic amino acid, another amino acid, either aromatic or aliphatic, is introduced. Conservative amino acid changes are well known in the art and include the 20 major amino acid changes defined in Table 1 above. The amino acids can be selected according to their properties. If the amino acids have similar polarity, this is shown in Table 2. The hydropathicity scale for the amino acid side chains of .
[0111] One or more amino acid residues of the amino acid sequence of SEQ ID NO: 3 may be additionally substituted from the above polypeptide. up to 1, 2, 3, 4, 5, 10, 20 or 30 or more deletions Residues may be deleted.
[0112] Variants may include fragments of SEQ ID NO: 3. Such fragments retain pore-forming activity. Fragments is at least 50, at least 100, at least 150, at least 200 in length or at least 250 amino acids. Such a fragment may be used to generate a pore. The fragment may be used as a membrane spanning domain of SEQ ID NO: 3, i.e., K135 to Q153. and preferably includes S183 to S208.
[0113] One or more amino acids may alternatively or additionally be added to the above polypeptides The extension may be an amino acid sequence of SEQ ID NO: 3 or a polypeptide variant or fragment thereof. The extension may be provided at the amino or carboxy terminus. Alternatively, the extension may be quite short, for example, up to 50 amino acids or It may be longer, such as 100 amino acids. Other fusion proteins are discussed in detail below.
[0114] The CsgG pore described herein may be a wild-type CsgG pore or a homologue or mutant thereof. Variants include variants / mutants that vary from the amino acid sequence of SEQ ID NO: 3 and do not form pores. A variant is a polypeptide having an amino acid sequence that retains its ability to bind to a target protein. The pore-forming ability of CsgG, which contains the β-barrel, is The β-sheets within each subunit provide the structural identity of the β-sequence. -SEQ ID NO: 3 forming the sheet, i.e. K134 to Q154 and S183 to S208 The one or more modifications may be made such that the resulting mutant retains the ability to form a pore. Mutations in SEQ ID NO: 3 may be made in the region of SEQ ID NO: 3 that forms a β-sheet, as long as they maintain the desired β-sheet structure. The antibody may contain one or more modifications, such as substitutions, additions or deletions, within its alpha helix and / or loop regions. It is preferred to include the above modifications.
[0115] A mutant CsgG monomer has a sequence that is altered from that of the wild-type CsgG monomer, Even mutant CsgG monomers, monomers that retain the ability to form pores, Mutant monomers may also be referred to herein as mutants. Methods for confirming the ability of mutant monomers to This will be considered in more detail.
[0116] Specific pore-forming CsgG mutant monomers that can be included in the CsgG pore include those with the following modifications: It may include any one or more of the following: - W at the position corresponding to R97 in SEQ ID NO: 3, - W at the position corresponding to R93 in SEQ ID NO: 3, - Y at the position corresponding to R97 in SEQ ID NO: 3, - Y at the position corresponding to R93 in SEQ ID NO: 3, - Y at each position corresponding to R93 and R97 in SEQ ID NO: 3, - D at the position corresponding to R192 of SEQ ID NO: 3, - a deletion of residues at positions corresponding to V105-I107 of SEQ ID NO: 3, - a deletion of one or more residues at positions corresponding to F193 to L199 of SEQ ID NO: 3, - a deletion of residues at positions corresponding to 195 to L199 of SEQ ID NO: 3, - a deletion of residues at positions corresponding to 193 to L199 of SEQ ID NO: 3, - T at the position corresponding to R191 of SEQ ID NO: 3, - Q at the position corresponding to K49 in SEQ ID NO: 3, - N at the position corresponding to K49 in SEQ ID NO: 3, - Q at the position corresponding to K42 of SEQ ID NO: 3, - Q at the position corresponding to E44 in SEQ ID NO: 3, - N at the position corresponding to E44 in SEQ ID NO: 3, - R at the position corresponding to L90 of SEQ ID NO: 3, - R at the position corresponding to L91 of SEQ ID NO: 3, - R at the position corresponding to I95 in SEQ ID NO: 3, - R at the position corresponding to A99 of SEQ ID NO: 3, - H at the position corresponding to E101 of SEQ ID NO: 3, - K at the position corresponding to E101 of SEQ ID NO: 3, - N at the position corresponding to E101 of SEQ ID NO: 3, - Q at the position corresponding to E101 of SEQ ID NO: 3, - T at the position corresponding to E101 of SEQ ID NO: 3, - K at the position corresponding to Q114 in SEQ ID NO: 3.
[0117] The CsgG pore-forming monomer contains an A and / or or preferably further comprises Q at a position corresponding to F56 in SEQ ID NO:3.
[0118] Consists of a CsgG monomer containing an R to W substitution at the position corresponding to position 97 of SEQ ID NO:3. The resulting pores are then used to characterize (or sequence) the target polynucleotide. The increase in precision is also due to the R97 Instead of W, the CsgG monomer has an R to W at a position corresponding to position 97 of SEQ ID NO:3. or an R to Y modification at positions corresponding to positions 93 and 97 of SEQ ID NO:3. Therefore, the pores are formed by modifying the amino acids to increase their hydrophobicity. One or more mutant Cs containing a modification at a position corresponding to R97 or R93 in column number 3 For example, such modifications may include, for example, W and Y. This may include, but is not limited to, amino acid substitutions with any amino acid containing a hydrophobic side chain.
[0119] A mutation from R to D, Q, F, S, or T at a position corresponding to position 192 of SEQ ID NO:3 The CsgG monomers containing the nucleotides lack a substitution at position 192 that could result in a reduction of the positive charge. Therefore, position 192 was replaced with an amino acid that reduces the positive charge. Monomers containing R192D / Q / F / S / T may also be used to treat protrusions formed from the monomers. The ability of the mutant pore to interact with and identify analytes such as polynucleotides is improved. However, in one embodiment, the amino acid sequence at position 19 of SEQ ID NO:3 may also be included. The residue at the position corresponding to 3 is preferably R or K, more preferably R. It's nice.
[0120] Deletion of V105, A106 and I107, F193, I194, D195, Y196, Deletion of Q197, R198 and L199, or D195, Y196, Q197, R1 Cells containing CsgG monomers containing deletions of 98 and L199, and / or F191T The pores exhibit increased accuracy when characterizing (or sequencing) target polynucleotides. The amino acids at positions 105–107 correspond to a cis-loop in the nanopore cap. , amino acids at positions 193–199 correspond to a trans-loop at the other end of the pore. Without wishing to be bound by theory, it is possible that the deletion of the cis-loop may impair the enzyme's interaction with the pore. The removal of the trans-loop improves the efficiency of the pore, and the removal of the trans-loop improves the preferential separation of DNA molecules on the trans side of the pore. It is believed that this reduces unwanted interactions.
[0121] Csg containing a K to Q or K to N mutation at the position corresponding to K94 in SEQ ID NO:3 The pore containing the G monomer is used to characterize (or sequence) the target polynucleotide. In contrast, the K94 mutation results in a noisy pore (i.e., signal to noise ratio) compared to the same pore without the mutation. The number of pores (which increases the tone ratio) decreases. Position 94 is found at the entrance to the pores and is the current This has been found to be a particularly sensitive location with respect to signal noise.
[0122] T104K or T104R, N91R, E101K / N / Q / T / H, E44N / Q, Q114K, A99R, I95R, N91R, L90R, E44Q / N and / or Q All pores containing CsgG monomers containing 42K or the corresponding mutations Characterize the target polynucleotide by comparing it to an identical pore without substitutions at the position (sequencing) When the target polynucleotide is captured, the resulting polynucleotide exhibits an increased ability to capture the target polynucleotide.
[0123] In one embodiment, the CsgG pore comprises: (a) I41, R93, A98, Q100, G103 , T104, A106, I107, N108, L113, S115, T117, Y130 , one at K135, E170, S208, D233, D238 and E244 or more mutations (i.e., mutations at one or more of those positions) and / or (b) D43S, E44S, F48S / N / Q / Y / W / I / V / H / R / K, Q87 N / R / K, N91K / R, K94R / F / Y / W / L / S / N, R97F / Y / W / V / I / K / S / Q / H, E101I / L / A / H, N102K / Q / L / I / V / S / H , R110F / G / N, Q114R / K, R142Q / S, T150Y / A / V / L / S / Q / N, R192D / Q / F / S / T and D248S / N / Q / K / R The variants include one or more monomers that are variants of SEQ ID NO: 3, including those described above. In some embodiments, the variant may comprise (a) and (b), or (a) and (b). In some embodiments, the mutation comprises R192D / Q / F / S / T (R192D In (a), the variants include 1, 2, 3, 4, 5, 6, 7, 8, 9, 1 Any number of positions, such as 0, 11, 12, 13, 14, 15, 16, 17, 18 or 19 Modifications in number and combination positions may be included.
[0124] In (a), the mutants are I41N, R93F / Y / W / L / I / V / N / Q / S, A 98K / R, Q100K / R, G103F / W / S / N / K / R, T104R / K, A1 06R / K, I107R / K / W / F / Y / L / V, N108R / K, L113K / R, S115R / K, T117R / K, Y130W / F / H / Q / N, K135L / V / N / Q / S, E170S / N / Q / K / R, S208V / I / F / W / Y / L / T, D233 S / N / Q / K / R, D238S / N / Q / K / R and E244S / N / Q / K / R It is preferable to include one or more of them.
[0125] In (a), the mutant is shown to be occluding (e.g., spanning) a target protein with respect to a transmembrane pore containing a monomer. It is preferred to include one or more modifications that provide for more consistent transfer of oligonucleotides. In (a), the mutant contains one or more mutations at positions R93, G103, and I107. Preferably, the variant contains a mutation (i.e., a mutation at one or more of these positions). R93, G103, I107, R93 and G103, R93 and I107, G103 and I107, or R93, G103 and I107. F / Y / W / L / I / V / N / Q / S, G103F / W / S / N / K / R and I107 It is preferable that the colorant contains one or more of R / K / W / F / Y / L / V. These include R93, G103 and I107 may occur in any combination shown for the positions.
[0126] In (a), the mutant is preferably a nucleic acid molecule, wherein the pore composed of the mutant monomer is The nucleic acid sequence includes one or more modifications that make it easier to capture the nucleic acid and polynucleotide. In particular, in (a), the mutants are I41, T104, A106, N108, At positions L113, S115, T117, E170, D233, D238 and E244 It is preferred to include more than one mutation (i.e., a mutation at one or more of these positions). Mutants can be introduced at any number and combination of positions (1, 2, 3, 4, 5, 6, 7, 8, Mutants may include modifications at positions 141N, 104R / K, 11, 12, 13, 141N, 141T ... A106R / K, N108R / K, L113K / R, S115R / K, T117R / K, E170S / N / Q / K / R, D233S / N / Q / K / R, D238S / N / Q / K / Preferably, the amino acid sequence contains one or more of R and E244S / N / Q / K / R. Alternatively, the variant may be (c) Q42K / R, E44N / Q, L90R / K, N 91R / K, I95R / K, A99R / K, E101H / K / N / Q / T and / or May contain Q114K / R.
[0127] In (a), the mutant contains one or more modifications that provide more consistent migration and increase capture. In particular, in (a), the mutants are preferably (i) A98, (ii) Q100, ( iii) G103, and (iv) one or more mutations at I107 (i.e., Preferably, the mutant contains a mutation at one or more of these positions. K, (ii) Q100K / R, (iii) G103K / R, and (iv) I107R / It is preferred that the aryl group contains one or more of the following:
[0128] Particularly preferred mutant monomers that provide for capture of analytes such as polynucleotides are Place Q42, E44, E44, L90, N91, I95, A99, E101 and Q114 wherein the mutation removes a negative charge at the mutation position; and / or increase the positive charge. In particular, the following mutations can be performed on the analyte, preferably the poly(A)- The mutant models of the invention produce CsgG pores with improved nucleotide trapping capacity. Can be included in models: Q42K, E44N, E44Q, L90R, N91R, I 95R, A99R, E101H, E101K, E101N, E101Q, E101T and and Q114K. One of these mutations in combination with other beneficial mutations Examples of specific mutant monomers, including: CsgG-(WT-Y51A / F56Q / K94Q / R97W / R192D-del( V105-I107)-Q42K CsgG-(WT-Y51A / F56Q / K94Q / R97W / R192D-del( V105-I107)-E44N CsgG-(WT-Y51A / F56Q / K94Q / R97W / R192D-del( V105-I107)-E44Q CsgG-(WT-Y51A / F56Q / K94Q / R97W / R192D-del( V105-I107)-L90R CsgG-(WT-Y51A / F56Q / K94Q / R97W / R192D-del( V105-I107)-N91R CsgG-(WT-Y51A / F56Q / K94Q / R97W / R192D-del( V105-I107)-I95R CsgG-(WT-Y51A / F56Q / K94Q / R97W / R192D-del( V105-I107)-A99R CsgG-(WT-Y51A / F56Q / K94Q / R97W / R192D-del( V105-I107)-E101H CsgG-(WT-Y51A / F56Q / K94Q / R97W / R192D-del( V105-I107)-E101K CsgG-(WT-Y51A / F56Q / K94Q / R97W / R192D-del( V105-I107)-E101N CsgG-(WT-Y51A / F56Q / K94Q / R97W / R192D-del( V105-I107)-E101Q CsgG-(WT-Y51A / F56Q / K94Q / R97W / R192D-del( V105-I107)-E101T CsgG-(WT-Y51A / F56Q / K94Q / R97W / R192D-del( V105-I107)-Q114K.
[0129] In (a), the variant preferably contains one or more modifications that increase the accuracy of characterization. In particular, in (a), the mutants are Y130, K135 and S208 (Y130, etc.), K135, S208, Y130 and K135, Y130 and S208, K135 and and S208, or one or more mutations at positions Y130, K135 and S208 (all Preferably, the mutant contains a mutation at one or more of these positions. One of 0W / F / H / Q / N, K135L / V / N / Q / S and R142Q / S These substitutions preferably include Y130, K135 and S208. They may be present in any number and combination as described above.
[0130] In (b), the mutants are any number and combination of substitutions (1, 2, 3, 4, 5, 6, 7 In (b), the mutant may comprise a monomer More consistent translocation of target polynucleotides with respect to (e.g., through) transmembrane pores, etc. In particular, in (b), the variant preferably comprises one or more modifications that provide: Q87N / R / K, (ii) K94R / F / Y / W / L / S / N, (iii) R97F / Y / W / V / I / K / S / Q / H, (iv)N102K / Q / L / I / V / S / H an Preferably, the mutation comprises one or more of d(v)R110F / G / N. Or more preferably, it contains K94Q and / or R97W or R97Y. More consistent translocation of target polynucleotides into (e.g., through) transmembrane pores containing mers Other preferred variants that are modifications that provide activity include (vi) R93W and R93Y. Preferred variants include R93W and R97W, R93Y and R97W, R93W and R97W, or more preferably R93Y and R97Y.
[0131] In (b), the mutant is preferably a nucleic acid molecule, wherein the pore composed of the mutant monomer is The nucleic acid sequence includes one or more modifications that make it easier to capture the nucleic acid and polynucleotide. In particular, in (b), the mutant is preferably (i) D43S, (ii) E44S, (iii) iii) N91K / R, (iv) Q114R / K and (v) D248S / N / Q / K / It is preferred to include one or more of R.
[0132] In (b), the mutants contain one or more modifications that provide more consistent migration and increase capture. In particular, in (b), the mutation is preferably Q87R / K, E101I / L / A. / H and N102K (such as Q87R / K), E101I / L / A / H, N102K, Q 87R / K and E101I / L / A / H, Q87R / K and N102K, E101I / L / A / H and N102K, or Q87R / K, E101I / L / A / H and N Preferably, it includes one or more of the 102K.
[0133] In (b), the variant preferably contains one or more modifications that increase the accuracy of the characterization. In particular, in (a), the mutation preferably includes F48S / N / Q / Y / W / I / V. It's nice.
[0134] In (b), the variants include one or more modifications that increase the accuracy of characterization and increase capture. In particular, in (a), the mutant may include F48H / R / K. preferable.
[0135] A variant may contain both (a) and (b) modifications that provide more consistent migration. Variants may contain modifications in both (a) and (b) that increase capture.
[0136] The present invention provides a method for characterizing analytes, such as polynucleotides, using pores containing the mutants. Mutants of SEQ ID NO: 3 are provided that increase the throughput of assays for the detection of nucleotides containing ... The mutant is a mutant at K94, preferably at K94Q or K94N, more preferably at K94 The mutation at Q may include the K94Q mutation in combination with other beneficial mutations. Or an example of a specific mutant monomer containing the K94N mutation is: CsgG-(WT-Y51A / F56Q / R97W / R192D-StrepII)9- K94N CsgG-(WT-Y51A / F56Q / R97W / R192D-StrepII)9- K94Q.
[0137] Using a monomer that is a variant of SEQ ID NO: 3 to form the CsgG pore Increased characterization accuracy in assays for characterizing analytes such as oligonucleotides Such mutants may include mutations at F191 (preferably F191T), V10 deletion of 5-I107, deletion of F193-L199 or D195-L199, and / or or R93 and / or R97 (preferably R93Y, R97Y, or more preferably These mutations include R97W, R93W, or both R97Y and R97Y. Certain mutants containing one or more of these mutations in combination with other beneficial mutations Examples of mutant monomers of are: CsgG-(WT-Y51A / F56Q / R97W / R192D-StrepII)9 -del(D195-L199) CsgG-(WT-Y51A / F56Q / R97W / R192D-StrepII)9 -del(F193-L199) CsgG-(WT-Y51A / F56Q / R97W / R192D-StrepII)9 -F191T CsgG-(WT-Y51A / F56Q / R97W / R192D-del(V105- I107)-StrepII)9 CsgG-(WT-Y51A / F56Q / K94Q / R97W / R192D-del( V105-I107) CsgG-(WT-Y51A / F56Q / R192D-StrepII)9-R93W CsgG-(WT-Y51A / F56Q / R192D-StrepII)9-R93W -del(D195-L199) CsgG-(WT-Y51A / F56Q / R192D-StrepII)9-R93Y / R97Y.
[0138] In another embodiment, the variant of SEQ ID NO: 3 comprises: (A) positions R192, F193, I194, One or more of D195, Y196, Q197, R198, L199, L200 and E201 Above deletion and / or (B)V139 / G140 / D149 / T150 / V186 / Q 187 / V204 / G205 (referred to herein as band 1), G137 / G138 / Q151 / Y152 / Y184 / E185 / Y206 / T207 (in this specification, the band 2) and A141 / R142 / G147 / A148 / A188 / G189 / It contains one or more deletions of G202 / E203 (referred to herein as band 3).
[0139] In (A), the mutants are located at any number and combination of positions (1, 2, 3, 4, 5, 6, 7 In (A), the mutant may include a deletion of Prefer: - D195, Y196, Q197, R198, and L199, - R192, F193, I194, D195, Y196, Q197, R198, L1 99, and L200, - Q197, R198, L199 and L200, - I194, D195, Y196, Q197, R198 and L199, - D195, Y196, Q197, R198, L199 and L200, - Y196, Q197, R198, L199, L200 and E201, - Q197, R198, L199, L200 and E201, - Q197, R198, L199, or - F193, I194, D195, Y196, Q197, R198, and L199 .
[0140] The mutants were D195, Y196, Q197, R198 and L199 or F193, I Preferably, the deletions include 194, D195, Y196, Q197, R198, and L199. In (B), band 1, band 2, band 3, bands 1 and 2, and band 1 are more preferable. and 3, bands 2 and 3, or bands 1, 2 and 3, any of bands 1 to 3 Any number and combination of deletions can be made. Mutants can be made by combining (A), (B), or (A) and (B). B) may contain deletions.
[0141] Mutants containing deletions of one or more positions according to (A) and / or (B) above are also described above and and may further include any of the modifications or substitutions discussed below. If a modification or substitution occurs at one or more positions shown after the position, the modification or substitution The numbering of one or more positions must be adjusted accordingly, e.g., L19 If band 9 is deleted, E244 becomes E243. Similarly, if band 1 is deleted, R1 92 becomes R186.
[0142] In another embodiment, the variant of SEQ ID NO: 3 is: (C) positions V105, A106 and I10 7. The deletion by (C) is not related to the deletion by (A) and / or (B). This can be done in addition to the deletions listed above.
[0143] The deletions typically involve (e.g., spanning) the transmembrane pore containing the monomer, This reduces noise associated with the migration of the target polynucleotide. The octides can be characterized more precisely.
[0144] In paragraphs where different amino acids at a particular position are separated by a / symbol, the / symbol is used to indicate "or" For example, Q87R / K means Q87R or Q87K.
[0145] A variant of SEQ ID NO:3 that provides increased capture of analytes such as polynucleotides is T104 (preferably T104R or T104K), mutation at N91 (preferably N91R ), mutation at E101 (preferably E101K / N / Q / T / H), A mutation at position E44 (preferably E44N or E44Q) and / or at position Q42 (preferably Q42K).
[0146] Mutations at different positions in SEQ ID NO: 3 may be combined in any possible way. The monomers in the sgG pore may contain one or more mutations that improve accuracy, one or more that reduce noise, or and / or one or more mutations that enhance analyte capture. .
[0147] Preferably, the variant of SEQ ID NO: 3 comprises one or more of the following: (i) N40, D43 , E44, S54, S57, Q62, R97, E101, E124, E131, R142 , one or more mutations at positions T150 and R192 (i.e., Mutations in one or more of these, e.g., N40, D43, E44, S54, S57, Q6 2, E101, E131 and T150, or N40, D43, E44, E101 and and E131 (i.e., one or more of those positions) (ii) mutations in Y51 / N55, Y51 / F56, N55 / F56, or Y5 mutations at 1 / N55 / F56, (iii) Q42R or Q42K, and (iv) K49 R, (v) N102R, N102F, N102Y or N102W, (vi) D149N , D149Q or D149R, (vii) E185N, E185Q or E185R, (viii) D195N, D195Q or D195R, (ix) E201N, E201 Q or E201R, (x) E203N, E203Q or E203R, and (xi) Positions F48, K49, P50, Y51, P52, A53, S54, N55, F56 and Deletion of one or more of S57. The mutant may include any combination of (i) to (xi). do.
[0148] When the mutant contains any one of (i) and (iii) to (xi), Y51, N5 5, F56, Y51 / N55, Y51 / F56, N55 / F56 or Y51 / N55 / and further comprising a mutation at one or more of Y51, N55, and F56, such as F56. do.
[0149] In (i), the mutants are N40, D43, E44, S54, S57, Q62, R97, E Any number and combination of 101, E124, E131, R142, T150 and R192 In (i), the mutant may contain mutations at N40, D43, E44, S5 4, one or more mutations at positions S57, Q62, E101, E131, and T150 ( That is, it is preferable to include a mutation at one or more of these positions. The variants contained one or more mutations at positions N40, D43, E44, E101, and E131. (i.e., a mutation at one or more of these positions). Preferably, the variant comprises a mutation at S54 and / or S57. The mutants were (a) S54 and / or S57, (b) Y51, N55, F56, and Y51 / N55, Y51 / F56, N55 / F56 or Y51 / N55 / F56), etc. More preferably, it contains a mutation at one or more of S54, S55 and S56. and / or S57 are deleted, then in (xi), it / they are In (i), the mutation cannot be mutated spontaneously, such as T150I. Preferably, the mutant comprises a mutation at T1 50, (b) Y51, N55, F56, Y51 / N55, Y51 / F56, N55 / F5 6, or Y51 / N55 / F56, with one or more of Y51, N55, and F56 In (i), the variant preferably comprises a mutation such as Q62R or Q62K. Alternatively, the variant preferably comprises a mutation at (a) Q62 , (b) Y51, N55, F56, Y51 / N55, Y51 / F56, N55 / F56, Or Y51 / N55 / F56, where one or more of Y51, N55, and F56 are involved. Preferably, the mutant comprises a mutation at D43, E44, Q62, or Any combination thereof, D43, E44, Q62, D43 / E44, D43 / Q62 , E44 / Q62 or D43 / E44 / Q62, etc. Alternatively, mutations The body is (a) D43, E44, Q62, D43 / E44, D43 / Q62, E44 / Q6 2 or D43 / E44 / Q62, (b) Y51, N55, F56, Y51 / N55, Y 51 / F56, N55 / F56, or Y51 / N55 / F56, etc. and F56.
[0150] (ii) and elsewhere in this use, different positions are separated by the / symbol In this case, / means "and", so Y51 / N55 means Y51 and N55. In (ii), the mutant preferably contains a mutation at Y51 / N55. The contraction of gG consists of three stacked structures formed by the side chains of residues Y51, N55, and F56. It has been proposed that the nucleus consists of concentric rings (Goyal et al., 2014, Nat Therefore, mutation of these residues in (ii) The polynucleotide moves through the pore, resulting in the observed current (polynucleotide This allows us to identify a direct relationship between the polynucleotide and the pore (as it translocates through the pore). F56 is a mutation useful in the methods of the present invention. The nuclei may be mutated in any of the ways discussed below with reference to the nuclei and pores.
[0151] In (v), the variant may include N102R, N102F, N102Y or N102W. In (i), the mutant is (a) N102R, N102F, N102Y or N102 W, and (b) at one or more of Y51, N55 and F56, e.g., Y51, N 55, F56, Y51 / N55, Y51 / F56, N55 / F56 or Y51 / N55 It is preferred to include a mutation at / F56.
[0152] In (xi), K49, P50, Y51, P52, A53, S54, N55, F56 and Any number and combination of deletions of K49, P50, Y51, P52, and S57 may be performed. Preferably, one or more of A53, S54, N55 and S57 may be deleted. If any of Y51, N55, and F56 were deleted in (xi), they were i) cannot be mutated, and vice versa.
[0153] In (i), the mutants are the substitutions N40R, N40K, D43N, D43Q, D43R, D4 3K, E44N, E44Q, E44R, E44K, S54P, S57P, Q62R, Q6 2K, R97N, R97G, R97L, E101N, E101Q, E101R, E101 K, E101F, E101Y, E101W, E124N, E124Q, E124R, E1 24K, E124F, E124Y, E124W, E131D, R142E, R142N, one or more of T150I, R192E and R192N, e.g., N40R, N40K; D43N, D43Q, D43R, D43K, E44N, E44Q, E44R, E44K, S54P, S57P, Q62R, Q62K, E101N, E101Q, E101R, E1 One of 01K, E101F, E101Y, E101W, E131D and T150I or above, or N40R, N40K, D43N, D43Q, D43R, D43K, E44 N, E44Q, E44R, E44K, E101N, E101Q, E101R, E101K , E101F, E101Y, E101W and E131D. The variant may contain any number and combination of these substitutions. Preferably, the variant comprises S54P and / or S57P. In (i), the variant is , (a) S54P and / or S57P and (b) Y51, N55 and F56 One or more of these, for example, Y51, N55, F56, Y51 / N55, Y51 / F56, It is preferred to include mutations at N55 / F56 or Y51 / N55 / F56. Mutations at one or more of N51, N55, and F56, as discussed below, may be present. In (i), the mutant may be F56A / S57P or S54P / F5 Preferably, the mutation comprises 6A. Preferably, the mutation comprises T150I. Alternatively, , the mutations include (a) T150I and (b) one or more of Y51, N55, and F56, e.g. For example, Y51, N55, F56, Y51 / N55, Y51 / F56, N55 / F56 or It is preferred to include mutations at Y51 / N55 / F56.
[0154] In (i), the variant preferably comprises Q62R or Q62K. Alternatively, the variant are (a) Q62R or Q62K, and (b) Y51, N55, F56, Y51 / F5 5, Y51, N such as Y51 / F56, N55 / F56, or Y51 / N55 / F56 Preferably, the mutant contains a mutation at one or more of D4, D55 and F56. 3N, E44N, Q62R or Q62K or any combination thereof, e.g. D4 3N, E44N, Q62R, Q62K, D43N / E44N, D43N / Q62R, D4 3N / Q62K, E44N / Q62R, E44N / Q62K, D43N / E44N / Q6 2R or D43N / E44N / Q62K. Alternatively, the mutant may comprise (a) D4 3N, E44N, Q62R, Q62K, D43N / E44N, D43N / Q62R, D4 3N / Q62K, E44N / Q62R, E44N / Q62K, D43N / E44N / Q6 2R or D43N / E44N / Q62K, and (b) Y51, N55, F56, Y5 1 / N55, Y51 / F56, N55 / F56 or Y51 / N55 / F56, etc., Y5 Preferably, the mutations include one or more of N55 and F56.
[0155] In (i), the mutant preferably comprises D43N.
[0156] In (i), the mutant may comprise E101R, E101S, E101F or E101N. It is preferable that:
[0157] In (i), the mutants are E124N, E124Q, E124R, E124K, E124F , E124Y, E124W or E124D (such as E124N).
[0158] In (i), the variant preferably comprises R142E and R142N.
[0159] In (i), the variant preferably comprises R97N, R97G or R97L.
[0160] In (i), the mutation preferably comprises R192E and R192N.
[0161] In (ii), the mutants are F56N / N55Q, F56N / N55R, F56N / N55 K, F56N / N55S, F56N / N55G, F56N / N55A, F56N / N55 T, F56Q / N55Q, F56Q / N55R, F56Q / N55K, F56Q / N55 S, F56Q / N55G, F56Q / N55A, F56Q / N55T, F56R / N55 Q, F56R / N55R, F56R / N55K, F56R / N55S, F56R / N55 G、F56R / N55A、F56R / N55T、F56S / N55Q、F56S / N55 R、F56S / N55K、F56S / N55S、F56S / N55G、F56S / N55 A、F56S / N55T、F56G / N55Q、F56G / N55R、F56G / N55 K、F56G / N55S、F56G / N55G、F56G / N55A、F56G / N55 T、F56A / N55Q、F56A / N55R、F56A / N55K、F56A / N55 S、F56A / N55G、F56A / N55A、F56A / N55T、F56K / N55 Q、F56K / N55R、F56K / N55K、F56K / N55S、F56K / N55 G、F56K / N55A、F56K / N55T、F56N / Y51L、F56N / Y51 V、F56N / Y51A、F56N / Y51N、F56N / Y51Q、F56N / Y51 S、F56N / Y51G、F56Q / Y51L、F56Q / Y51V、F56Q / Y51 A、F56Q / Y51N、F56Q / Y51Q、F56Q / Y51S、F56Q / Y51 G、F56R / Y51L、F56R / Y51V、F56R / Y51A、F56R / Y51 N、F56R / Y51Q、F56R / Y51S、F56R / Y51G、F56S / Y51 L、F56S / Y51V、F56S / Y51A、F56S / Y51N、F56S / Y51 Q、F56S / Y51S、F56S / Y51G、F56G / Y51L、F56G / Y51 V、F56G / Y51A、F56G / Y51N、F56G / Y51Q、F56G / Y51 S、F56G / Y51G、F56A / Y51L、F56A / Y51V、F56A / Y51 A、F56A / Y51N、F56A / Y51Q、F56A / Y51S、F56A / Y51 G、F56K / Y51L、F56K / Y51V、F56K / Y51A、F56K / Y51 N、F56K / Y51Q、F56K / Y51S、F56K / Y51G、N55Q / Y51 L、N55Q / Y51V、N55Q / Y51A、N55Q / Y51N、N55Q / Y51 Q、N55Q / Y51S、N55Q / Y51G、N55R / Y51L、N55R / Y51 V、N55R / Y51A、N55R / Y51N、N55R / Y51Q、N55R / Y51 S、N55R / Y51G、N55K / Y51L、N55K / Y51V、N55K / Y51 A、N55K / Y51N、N55K / Y51Q、N55K / Y51S、N55K / Y51 G、N55S / Y51L、N55S / Y51V、N55S / Y51A、N55S / Y51 N、N55S / Y51Q、N55S / Y51S、N55S / Y51G、N55G / Y51 L、N55G / Y51V、N55G / Y51A、N55G / Y51N、N55G / Y51 Q、N55G / Y51S、N55G / Y51G、N55A / Y51L、N55A / Y51 V、N55A / Y51A、N55A / Y51N、N55A / Y51Q、N55A / Y51 S、N55A / Y51G、N55T / Y51L、N55T / Y51V、N55T / Y51 A、N55T / Y51N、N55T / Y51Q、N55T / Y51S、N55T / Y51 G、F56N / N55Q / Y51L、F56N / N55Q / Y51V、F56N / N55 Q / Y51A、F56N / N55Q / Y51N、F56N / N55Q / Y51Q、F56 N / N55Q / Y51S、F56N / N55Q / Y51G、F56N / N55R / Y51 L、F56N / N55R / Y51V、F56N / N55R / Y51A、F56N / N55 R / Y51N、F56N / N55R / Y51Q、F56N / N55R / Y51S、F56 N / N55R / Y51G、F56N / N55K / Y51L、F56N / N55K / Y51 V、F56N / N55K / Y51A、F56N / N55K / Y51N、F56N / N55 K / Y51Q、F56N / N55K / Y51S、F56N / N55K / Y51G、F56 N / N55S / Y51L、F56N / N55S / Y51V、F56N / N55S / Y51 A、F56N / N55S / Y51N、F56N / N55S / Y51Q、F56N / N55 S / Y51S、F56N / N55S / Y51G、F56N / N55G / Y51L、F56 N / N55G / Y51V、F56N / N55G / Y51A、F56N / N55G / Y51 N、F56N / N55G / Y51Q、F56N / N55G / Y51S、F56N / N55 G / Y51G、F56N / N55A / Y51L、F56N / N55A / Y51V、F56 N / N55A / Y51A、F56N / N55A / Y51N、F56N / N55A / Y51 Q、F56N / N55A / Y51S、F56N / N55A / Y51G、F56N / N55 T / Y51L、F56N / N55T / Y51V、F56N / N55T / Y51A、F56 N / N55T / Y51N、F56N / N55T / Y51Q、F56N / N55T / Y51 S、F56N / N55T / Y51G、F56Q / N55Q / Y51L、F56Q / N55 Q / Y51V、F56Q / N55Q / Y51A、F56Q / N55Q / Y51N、F56 Q / N55Q / Y51Q、F56Q / N55Q / Y51S、F56Q / N55Q / Y51 G、F56Q / N55R / Y51L、F56Q / N55R / Y51V、F56Q / N55 R / Y51A、F56Q / N55R / Y51N、F56Q / N55R / Y51Q、F56 Q / N55R / Y51S、F56Q / N55R / Y51G、F56Q / N55K / Y51 L、F56Q / N55K / Y51V、F56Q / N55K / Y51A、F56Q / N55 K / Y51N、F56Q / N55K / Y51Q、F56Q / N55K / Y51S、F56 Q / N55K / Y51G、F56Q / N55S / Y51L、F56Q / N55S / Y51 V、F56Q / N55S / Y51A、F56Q / N55S / Y51N、F56Q / N55 S / Y51Q、F56Q / N55S / Y51S、F56Q / N55S / Y51G、F56 Q / N55G / Y51L、F56Q / N55G / Y51V、F56Q / N55G / Y51 A、F56Q / N55G / Y51N、F56Q / N55G / Y51Q、F56Q / N55 G / Y51S、F56Q / N55G / Y51G、F56Q / N55A / Y51L、F56 Q / N55A / Y51V、F56Q / N55A / Y51A、F56Q / N55A / Y51 N、F56Q / N55A / Y51Q、F56Q / N55A / Y51S、F56Q / N55 A / Y51G、F56Q / N55T / Y51L、F56Q / N55T / Y51V、F56 Q / N55T / Y51A、F56Q / N55T / Y51N、F56Q / N55T / Y51 Q、F56Q / N55T / Y51S、F56Q / N55T / Y51G、F56R / N55 Q / Y51L、F56R / N55Q / Y51V、F56R / N55Q / Y51A、F56 R / N55Q / Y51N、F56R / N55Q / Y51Q、F56R / N55Q / Y51 S、F56R / N55Q / Y51G、F56R / N55R / Y51L、F56R / N55 R / Y51V、F56R / N55R / Y51A、F56R / N55R / Y51N、F56 R / N55R / Y51Q、F56R / N55R / Y51S、F56R / N55R / Y51 G、F56R / N55K / Y51L、F56R / N55K / Y51V、F56R / N55 K / Y51A、F56R / N55K / Y51N、F56R / N55K / Y51Q、F56 R / N55K / Y51S、F56R / N55K / Y51G、F56R / N55S / Y51 L、F56R / N55S / Y51V、F56R / N55S / Y51A、F56R / N55 S / Y51N、F56R / N55S / Y51Q、F56R / N55S / Y51S、F56 R / N55S / Y51G、F56R / N55G / Y51L、F56R / N55G / Y51 V、F56R / N55G / Y51A、F56R / N55G / Y51N、F56R / N55 G / Y51Q、F56R / N55G / Y51S、F56R / N55G / Y51G、F56 R / N55A / Y51L、F56R / N55A / Y51V、F56R / N55A / Y51 A、F56R / N55A / Y51N、F56R / N55A / Y51Q、F56R / N55 A / Y51S、F56R / N55A / Y51G、F56R / N55T / Y51L、F56 R / N55T / Y51V、F56R / N55T / Y51A、F56R / N55T / Y51 N、F56R / N55T / Y51Q、F56R / N55T / Y51S、F56R / N55 T / Y51G、F56S / N55Q / Y51L、F56S / N55Q / Y51V、F56 S / N55Q / Y51A、F56S / N55Q / Y51N、F56S / N55Q / Y51 Q、F56S / N55Q / Y51S、F56S / N55Q / Y51G、F56S / N55 R / Y51L、F56S / N55R / Y51V、F56S / N55R / Y51A、F56 S / N55R / Y51N、F56S / N55R / Y51Q、F56S / N55R / Y51 S、F56S / N55R / Y51G、F56S / N55K / Y51L、F56S / N55 K / Y51V、F56S / N55K / Y51A、F56S / N55K / Y51N、F56 S / N55K / Y51Q、F56S / N55K / Y51S、F56S / N55K / Y51 G、F56S / N55S / Y51L、F56S / N55S / Y51V、F56S / N55 S / Y51A、F56S / N55S / Y51N、F56S / N55S / Y51Q、F56 S / N55S / Y51S、F56S / N55S / Y51G、F56S / N55G / Y51 L、F56S / N55G / Y51V、F56S / N55G / Y51A、F56S / N55 G / Y51N、F56S / N55G / Y51Q、F56S / N55G / Y51S、F56 S / N55G / Y51G、F56S / N55A / Y51L、F56S / N55A / Y51 V、F56S / N55A / Y51A、F56S / N55A / Y51N、F56S / N55 A / Y51Q、F56S / N55A / Y51S、F56S / N55A / Y51G、F56 S / N55T / Y51L、F56S / N55T / Y51V、F56S / N55T / Y51 A、F56S / N55T / Y51N、F56S / N55T / Y51Q、F56S / N55 T / Y51S、F56S / N55T / Y51G、F56G / N55Q / Y51L、F56 G / N55Q / Y51V、F56G / N55Q / Y51A、F56G / N55Q / Y51 N、F56G / N55Q / Y51Q、F56G / N55Q / Y51S、F56G / N55 Q / Y51G、F56G / N55R / Y51L、F56G / N55R / Y51V、F56 G / N55R / Y51A、F56G / N55R / Y51N、F56G / N55R / Y51 Q、F56G / N55R / Y51S、F56G / N55R / Y51G、F56G / N55 K / Y51L、F56G / N55K / Y51V、F56G / N55K / Y51A、F56 G / N55K / Y51N、F56G / N55K / Y51Q、F56G / N55K / Y51 S、F56G / N55K / Y51G、F56G / N55S / Y51L、F56G / N55 S / Y51V、F56G / N55S / Y51A、F56G / N55S / Y51N、F56 G / N55S / Y51Q、F56G / N55S / Y51S、F56G / N55S / Y51 G、F56G / N55G / Y51L、F56G / N55G / Y51V、F56G / N55 G / Y51A、F56G / N55G / Y51N、F56G / N55G / Y51Q、F56 G / N55G / Y51S、F56G / N55G / Y51G、F56G / N55A / Y51 L、F56G / N55A / Y51V、F56G / N55A / Y51A、F56G / N55 A / Y51N、F56G / N55A / Y51Q、F56G / N55A / Y51S、F56 G / N55A / Y51G、F56G / N55T / Y51L、F56G / N55T / Y51 V、F56G / N55T / Y51A、F56G / N55T / Y51N、F56G / N55 T / Y51Q、F56G / N55T / Y51S、F56G / N55T / Y51G、F56A / N55Q / Y51L、F56A / N55Q / Y51V、F56A / N55Q / Y51A 、F56A / N55Q / Y51N、F56A / N55Q / Y51Q、F56A / N55Q / Y51S、F56A / N55Q / Y51G、F56A / N55R / Y51L、F56A / N55R / Y51V、F56A / N55R / Y51A、F56A / N55R / Y51N 、F56A / N55R / Y51Q、F56A / N55R / Y51S、F56A / N55R / Y51G、F56A / N55K / Y51L、F56A / N55K / Y51V、F56A / N55K / Y51A、F56A / N55K / Y51N、F56A / N55K / Y51Q 、F56A / N55K / Y51S、F56A / N55K / Y51G、F56A / N55S / Y51L、F56A / N55S / Y51V、F56A / N55S / Y51A、F56A / N55S / Y51N、F56A / N55S / Y51Q、F56A / N55S / Y51S 、F56A / N55S / Y51G、F56A / N55G / Y51L、F56A / N55G / Y51V、F56A / N55G / Y51A、F56A / N55G / Y51N、F56A / N55G / Y51Q、F56A / N55G / Y51S、F56A / N55G / Y51G 、F56A / N55A / Y51L、F56A / N55A / Y51V、F56A / N55A / Y51A、F56A / N55A / Y51N、F56A / N55A / Y51Q、F56A / N55A / Y51S、F56A / N55A / Y51G、F56A / N55T / Y51L 、F56A / N55T / Y51V、F56A / N55T / Y51A、F56A / N55T / Y51N、F56A / N55T / Y51Q、F56A / N55T / Y51S、F56A / N55T / Y51G、F56K / N55Q / Y51L、F56K / N55Q / Y51V 、F56K / N55Q / Y51A、F56K / N55Q / Y51N、F56K / N55Q / Y51Q、F56K / N55Q / Y51S、F56K / N55Q / Y51G、F56K / N55R / Y51L、F56K / N55R / Y51V、F56K / N55R / Y51A 、F56K / N55R / Y51N、F56K / N55R / Y51Q、F56K / N55R / Y51S、F56K / N55R / Y51G、F56K / N55K / Y51L、F56K / N55K / Y51V, F56K / N55K / Y51A, F56K / N55K / Y51N , F56K / N55K / Y51Q, F56K / N55K / Y51S, F56K / N55K / Y51G, F56K / N55S / Y51L, F56K / N55S / Y51V, F56K / N55S / Y51A, F56K / N55S / Y51N, F56K / N55S / Y51Q , F56K / N55S / Y51S, F56K / N55S / Y51G, F56K / N55G / Y51L, F56K / N55G / Y51V, F56K / N55G / Y51A, F56K / N55G / Y51N, F56K / N55G / Y51Q, F56K / N55G / Y51S , F56K / N55G / Y51G, F56K / N55A / Y51L, F56K / N55A / Y51V, F56K / N55A / Y51A, F56K / N55A / Y51N, F56K / N55A / Y51Q, F56K / N55A / Y51S, F56K / N55A / Y51G , F56K / N55T / Y51L, F56K / N55T / Y51V, F56K / N55T / Y51A, F56K / N55T / Y51N, F56K / N55T / Y51Q, F56K / N55T / Y51S, F56K / N55T / Y51G, F56E / N55R, F56E / N55K, F56D / N55R, F56D / N55K, F56R / N55E, F56R / N55D, F56K / N55E or F56K / N55D.
[0162] In (ii), the mutants are Y51R / F56Q, Y51N / F56N, Y51M / F56 Q, Y51L / F56Q, Y51I / F56Q, Y51V / F56Q, Y51A / F56 Q, Y51P / F56Q, Y51G / F56Q, Y51C / F56Q, Y51Q / F56 Q, Y51N / F56Q, Y51S / F56Q, Y51E / F56Q, Y51D / F56 It is preferred that the amino acid sequence comprises Q, Y51K / F56Q or Y51H / F56Q.
[0163] In (ii), the mutant is Y51T / F56Q, Y51Q / F56Q or Y51A / F It is preferred to include 56Q.
[0164] In (ii), the mutants are Y51T / F56F, Y51T / F56M, Y51T / F56 L, Y51T / F56I, Y51T / F56V, Y51T / F56A, Y51T / F56 P, Y51T / F56G, Y51T / F56C, Y51T / F56Q, Y51T / F56 N, Y51T / F56T, Y51T / F56S, Y51T / F56E, Y51T / F56 Preferably, the nucleotide sequence includes Y51T / F56K, Y51T / F56H, or Y51T / F56R. I wish.
[0165] In (ii), the mutant is Y51T / N55Q, Y51T / N55S or Y51T / N It is preferred to include 55A.
[0166] In (ii), the mutants are Y51A / F56F, Y51A / F56L, Y51A / F56 I, Y51A / F56V, Y51A / F56A, Y51A / F56P, Y51A / F56 G, Y51A / F56C, Y51A / F56Q, Y51A / F56N, Y51A / F56 T, Y51A / F56S, Y51A / F56E, Y51A / F56D, Y51A / F56 It is preferred that the amino acid sequence contains K, Y51A / F56H or Y51A / F56R.
[0167] In (ii), the mutants are Y51C / F56A, Y51E / F56A, Y51D / F56 A, Y51K / F56A, Y51H / F56A, Y51Q / F56A, Y51N / F56 A, Y51S / F56A, Y51P / F56A or Y51V / F56A. I wish.
[0168] In (xi), the mutants are Y51 / P52, Y51 / P52 / A53, P50-P52, Deletion of P50 to A53, K49 to Y51, K49 to A53, and a single proline (P) substitutions, K49 to S54 and single P substitutions, Y51 to A53, Y51 to S54, N 55 / F56, N55~S57, N55 / F56 and single P substitution, N55 / F5 6 and a single glycine (G) substitution, N55 / F56 and a single alanine (A) substitution substitutions, N55 / F56 and single P and Y51N substitutions, N55 / F56 and Single P and Y51Q substitutions, N55 / F56 and single P and Y51S substitutions substitutions at N55 / F56 and a single G and Y51N; substitutions at N55 / F56 and a single substitutions with G and Y51Q, N55 / F56 and a single G and Y51S substitution, N55 / F56 and a single A and Y51N substitutions, N55 / F56 and a single A / Y51Q substitution, or N55 / F56 and a single A and Y51S substitution It is preferable to do so.
[0169] The mutants were D195N / E203N, D195Q / E203N, and D195N / E203Q. , D195Q / E203Q, E201N / E203N, E201Q / E203N, E20 1N / E203Q, E201Q / E203Q, E185N / E203Q, E185Q / E 203Q, E185N / E203N, E185Q / E203N, D195N / E201N <h2 style=";text-align:left;direction:ltr"> / E203N、D195Q / E201N / E203N、D195N / E201Q / E20<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 3N、D195N / E201N / E203Q、D195Q / E201Q / E203N、D<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 195Q / E201N / E203Q、D195N / E201Q / E203Q、D195Q<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> / E201Q / E203Q、D149N / E201N、D149Q / E201N、D14<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 9N / E201Q、D149Q / E201Q、D149N / E201N / D195N、D<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 149Q / E201N / D195N、D149N / E201Q / D195N、D149N<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> / E201N / D195Q、D149Q / E201Q / D195N、D149Q / E20<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 1N / D195Q、D149N / E201Q / D195Q、D149Q / E201Q / D<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 195Q、D149N / E203N、D149Q / E203N、D149N / E203Q<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 、D149Q / E203Q、D149N / E185N / E201N、D149Q / E18<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 5N / E201N、D149N / E185Q / E201N、D149N / E185N / E<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 201Q、D149Q / E185Q / E201N、D149Q / E185N / E201Q<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 、D149N / E185Q / E201Q、D149Q / E185Q / E201Q、D14<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 9N / E185N / E203N、D149Q / E185N / E203N、D149N / E<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 185Q / E203N、D149N / E185N / E203Q、D149Q / E185Q<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> / E203N、D149Q / E185N / E203Q、D149N / E185Q / E20<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 3Q、D149Q / E185Q / E203Q、D149N / E185N / E201N / E<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 203N、D149Q / E185N / E201N / E203N、D149N / E185Q<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> / E201N / E203N、D149N / E185N / E201Q / E203N、D14 9N / E185N / E201N / E203Q、D149Q / E185Q / E201N / E 203N、D149Q / E185N / E201Q / E203N、D149Q / E185N / E201N / E203Q、D149N / E185Q / E201Q / E203N、D14 9N / E185Q / E201N / E203Q、D149N / E185N / E201Q / E 203Q、D149Q / E185Q / E201Q / E203Q、D149Q / E185Q / E201N / E203Q、D149Q / E185N / E201Q / E203Q、D14 9N / E185Q / E201Q / E203Q、D149Q / E185Q / E201Q / E 203N、D149N / E185N / D195N / E201N / E203N、D149Q / E185N / D195N / E201N / E203N、D149N / E185Q / D19 5N / E201N / E203N、D149N / E185N / D195Q / E201N / E 203N、D149N / E185N / D195N / E201Q / E203N、D149N / E185N / D195N / E201N / E203Q、D149Q / E185Q / D19 5N / E201N / E203N、D149Q / E185N / D195Q / E201N / E 203N、D149Q / E185N / D195N / E201Q / E203N、D149Q / E185N / D195N / E201N / E203Q、D149N / E185Q / D19 5Q / E201N / E203N、D149N / E185Q / D195N / E201Q / E 203N、D149N / E185Q / D195N / E201N / E203Q、D149N / E185N / D195Q / E201Q / E203N、D149N / E185N / D19 <h2 style=";text-align:left;direction:ltr">5Q / E201N / E203Q、D149N / E185N / D195N / E201Q / E<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 203Q、D149Q / E185Q / D195Q / E201N / E203N、D149Q<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> / E185Q / D195N / E201Q / E203N、D149Q / E185Q / D19<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 5N / E201N / E203Q、D149Q / E185N / D195Q / E201Q / E<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 203N、D149Q / E185N / D195Q / E201N / E203Q、D149Q<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> / E185N / D195N / E201Q / E203Q、D149N / E185Q / D19<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 5Q / E201Q / E203N、D149N / E185Q / D195Q / E201N / E<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 203Q、D149N / E185Q / D195N / E201Q / E203Q、D149N<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> / E185N / D195Q / E201Q / E203Q、D149Q / E185Q / D19<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 5Q / E201Q / E203N、D149Q / E185Q / D195Q / E201N / E<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 203Q、D149Q / E185Q / D195N / E201Q / E203Q、D149Q<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> / E185N / D195Q / E201Q / E203Q、D149N / E185Q / D19<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 5Q / E201Q / E203Q、D149Q / E185Q / D195Q / E201Q / E<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 203Q、D149N / E185R / E201N / E203N、D149Q / E185R<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> / E201N / E203N、D149N / E185R / E201Q / E203N、D14<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 9N / E185R / E201N / E203Q、D149Q / E185R / E201Q / E<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 203N、D149Q / E185R / E201N / E203Q、D149N / E185<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> R / E201Q / E203Q、D149Q / E185R / E201Q / E203Q、D1<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 49R / E185N / E201N / E203N、D149R / E185Q / E201N / <h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> E203N、D149R / E185N / E201Q / E203N、D149R / E185<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> N / E201N / E203Q、D149R / E185Q / E201Q / E203N、D1<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 49R / E185Q / E201N / E203Q、D149R / E185N / E201Q / <h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> E203Q、D149R / E185Q / E201Q / E203Q、D149R / E185<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> N / D195N / E201N / E203N、D149R / E185Q / D195N / E2<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 01N / E203N、D149R / E185N / D195Q / E201N / E203N、<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> D149R / E185N / D195N / E201Q / E203N、D149R / E185<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> Q / D195N / E201N / E203Q、D149R / E185Q / D195Q / E2<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 01N / E203N、D149R / E185Q / D195N / E201Q / E203N、<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> D149R / E185Q / D195N / E201N / E203Q、D149R / E185<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> N / D195Q / E201Q / E203N、D149R / E185N / D195Q / E2<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 01N / E203Q、D149R / E185N / D195N / E201Q / E203Q、<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> D149R / E185Q / D195Q / E201Q / E203N、D149R / E185<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> Q / D195Q / E201N / E203Q、D149R / E185Q / D195N / E2<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 01Q / E203Q、D149R / E185N / D195Q / E201Q / E203Q、<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> D149R / E185Q / D195Q / E201Q / E203Q、D149N / E185<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> R / D195N / E201N / E203N、D149Q / E185R / D195N / E2<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 01N / E203N、D149N / E185R / D195Q / E201N / E203N、<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">D149N / E185R / D195N / E201Q / E203N、D149N / E185<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> R / D195N / E201N / E203Q、D149Q / E185R / D195Q / E2<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 01N / E203N、D149Q / E185R / D195N / E201Q / E203N、<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> D149Q / E185R / D195N / E201N / E203Q、D149N / E185<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> R / D195Q / E201Q / E203N、D149N / E185R / D195Q / E2<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 01N / E203Q、D149N / E185R / D195N / E201Q / E203Q、<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> D149Q / E185R / D195Q / E201Q / E203N、D149Q / E185<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> R / D195Q / E201N / E203Q、D149Q / E185R / D195N / E2<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 01Q / E203Q、D149N / E185R / D195Q / E201Q / E203Q、<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> D149Q / E185R / D195Q / E201Q / E203Q、D149N / E185<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> R / D195N / E201R / E203N、D149Q / E185R / D195N / E2<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 01R / E203N、D149N / E185R / D195Q / E201R / E203N、<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> D149N / E185R / D195N / E201R / E203Q、D149Q / E185<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> R / D195Q / E201R / E203N、D149Q / E185R / D195N / E2<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 01R / E203Q、D149N / E185R / D195Q / E201R / E203Q、<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> D149Q / E185R / D195Q / E201R / E203Q、E131D / K49R<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 、E101N / N102F、E101N / N102Y、E101N / N102W、E10<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 1F / N102F、E101F / N102Y、E101F / N102W、E101Y / N<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 102F、E101Y / N102Y、E101Y / N102W、E101W / N102F , E101W / N102Y, E101W / N102W, E101N / N102R, E10 May include 1F / N102R, E101Y / N102R or E101W / N102F More preferable.
[0170] As the polynucleotide translocates through the pore, fewer nucleotides contribute to the current Preferred pore-forming mutants of the invention are Y51A / F56A, Y51A / F56N, Y51I / F56A, Y51L / F56A, Y51T / F56A, Y51I / F56N, Contains Y51L / F56N or Y51T / F56N, or more preferably Y51I / F56A, Y51L / F56A or Y51T / F56A. As mentioned above, this is observed (as the polynucleotide moves through the pore and the polynucleotide) This facilitates the identification of a direct relationship between the currents
[0171] Preferred pore-forming mutants that display increased coverage include mutations at the following positions: nothing: Y51, F56, D149, E185, E201 and E203, N55 and F56, Y51 and F56, Y51, N55 and F56, or F56 and N102.
[0172] Preferred pore-forming variants that display increased coverage are Y51N, F56A, D149N, E185R, E201N and E203N, N55S and F56Q, Y51A and F56A, Y51A and F56N, Y51I and F56A, Y51L and F56A, Y51T and F56A, Y51I and F56N, Y51L and F56N, Y51T and F56N, Y51T and F56Q, Y51A, N55S and F56A, Y51A, N55S, and F56N, Y51T, N55S and F56Q, or F56Q and N102R.
[0173] As the polynucleotide translocates through the pore, the few nucleotides that contribute to the current flow Preferred pore-forming mutants include mutations at the following positions: N55 and F56 (such as N55X and F56Q), where X is any amino acid , Y51 and F56 (such as Y51X and F56Q), where X is any amino acid do.
[0174] Particularly preferred variants include Y51A and F56Q.
[0175] A preferred variant of the pore-forming material that displays increased throughput is the sudden increase in the Mutations include: D149, E185 and E203, D149, E185, E201, and E203, or D149, E185, D195, E201, and E203.
[0176] Preferred pore-forming variants that display increased throughput are D149N, E185N and E203N, D149N, E185N, E201N and E203N, D149N, E185R, D195N, E201N and E203N, or D149N, E185R, D195N, E201R and E203N.
[0177] Preferred mutants that form pores with enhanced polynucleotide capture include the following mutations: include: D43N / Y51T / F56Q, E44N / Y51T / F56Q, D43N / E44N / Y51T / F56Q, Y51T / F56Q / Q62R, D43N / Y51T / F56Q / Q62R, E44N / Y51T / F56Q / Q62R, or D43N / E44N / Y51T / F56Q / Q62R.
[0178] Preferred variants include the following mutations: D149R / E185R / E201R / E203R or Y51T / F56Q / D149 R / E185R / E201R / E203R, D149N / E185N / E201N / E203N or Y51T / F56Q / D149 N / E185N / E201N / E203N, E201R / E203R or Y51T / F56Q / E201R / E203R E201N / E203R or Y51T / F56Q / E201N / E203R, E203R or Y51T / F56Q / E203R, E203N or Y51T / F56Q / E203N, E201R or Y51T / F56Q / E201R, E201N or Y51T / F56Q / E201N, E185R or Y51T / F56Q / E185R, E185N or Y51T / F56Q / E185N, D149R or Y51T / F56Q / D149R, D149N or Y51T / F56Q / D149N, R142E or Y51T / F56Q / R142E, R142N or Y51T / F56Q / R142N, R192E or Y51T / F56Q / R192E, or R192N or Y51T / F56Q / R192N.
[0179] Preferred variants include the following mutations: Y51A / F56Q / E101N / N102R, Y51A / F56Q / R97N / N102G, Y51A / F56Q / R97N / N102R, Y51A / F56Q / R97N, Y51A / F56Q / R97G, Y51A / F56Q / R97L, Y51A / F56Q / N102R, Y51A / F56Q / N102F, Y51A / F56Q / N102G, Y51A / F56Q / E101R, Y51A / F56Q / E101F, Y51A / F56Q / E101N, or Y51A / F56Q / E101G
[0180] Preferably, the mutant further comprises a mutation at T150. A preferred pore-forming mutant includes T150I. The natural mutation may be combined with any of the mutations or combinations of mutations described above.
[0181] Preferred variants of SEQ ID NO: 3 are (a) R97W and (b) Y51 and / or F5 6. Preferred variants of SEQ ID NO: 3 include: (a) R97W, and (b) Y51R / H / K / D / E / S / T / N / Q / C / G / P / A / V / I / L / M and / or F56R / H / K / D / E / S / T / N / Q / C / G / P / A / V / I / L / M Preferred variants of SEQ ID NO: 3 include: (a) R97W, and (b) Y51L / V / A / N / Q / S / G and / or F56A / Q / N. Preferred variants of SEQ ID NO: 3 SEQ ID NO: 3 contains (a) R97W, and (b) Y51A and / or F56Q. Preferred mutants include R97W, Y51A and F56Q.
[0182] Preferably, the variant of SEQ ID NO: 3 comprises a mutation at R192. 2D / Q / F / S / T / N / E, R192D / Q / F / S / T or R192D / Q Preferred variants of SEQ ID NO: 3 include (a) R97W, (b) Y51 and and / or mutations at F56, and (c) mutations at R192, R192D / Q / Includes F / S / T / N / E, R192D / Q / F / S / T or R192D / Q. The preferred variants in row number 3 are (a) R97W, (b) Y51R / H / K / D / E / S / T / N / Q / C / G / P / A / V / I / L / M and / or F56 R / H / K / D / E / S / T / N / Q / C / G / P / A / V / I / L / M, and (c) Sudden in R192 Mutation, R192D / Q / F / S / T / N / E, R192D / Q / F / S / T or R19 2D / Q, etc. Preferred variants of SEQ ID NO: 3 include: (a) R97W, (b) Y51L / V / A / N / Q / S / G and / or F56A / Q / N, and (c) R192 Mutation, R192D / Q / F / S / T / N / E, R192D / Q / F / S / T or R 192D / Q, etc. Preferred variants of SEQ ID NO: 3 include: (a) R97W, (b) Y5 1A and / or F56Q, and (c) mutations at R192, R192 D / Q / Includes F / S / T / N / E, R192D / Q / F / S / T or R192D / Q. Preferred variants in sequence number 3 are R97W, Y51A, F56Q and R192D / Q / F Preferred variants of SEQ ID NO: 3 include R97W, Y5 Preferred variants of SEQ ID NO: 3 include R97W, Y, F56Q and R192D. The amino acids that differ at specific positions are indicated by the / symbol. In paragraphs separated by a / symbol, the / symbol means "or." For example, R192D / Q This means 192D or R192Q.
[0183] Any of the above preferred variants of SEQ ID NO: 3 may further comprise a mutation at R93. Preferred variants of SEQ ID NO: 3 are (a) R93W, and (b) Y51 and / or Mutations at F56 include preferably Y51A and F56Q.
[0184] Any of the above preferred variants of SEQ ID NO: 3 may contain a K94N / Q mutation. Any of the above preferred variants of No. 3 may contain the F191T mutation.
[0185] The CsgG monomer can be modified to facilitate attachment to the CsgF peptide, for example: , cysteine residues at positions 132, 133, 136, 138, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 1 2, 144, 145, 147, 149, 151, 153, 155, 183, 185, 18 One or more corresponding to 7, 189, 191, 201, 203, 205, 207 and 209 and / or contacts CsgF to facilitate covalent binding to CsgG. The amino acid sequence can be introduced into any one of the positions listed in Table 4. Alternatively or in addition to covalent bonds, the pores may be bound by hydrophobic or electrostatic interactions. To promote such interactions, the amino acid sequence at position 132 of SEQ ID NO: 3 may be stabilized by the amino acid sequence at position 132 of SEQ ID NO: 3. 133, 136, 138, 140, 142, 144, 145, 147, 149, 151, 153, 155, 183, 185, 187, 189, 191, 201, 203, 205, at positions corresponding to one or more of 207 and 209 and / or contacting CsgF. Non-natively reactive or photoreactive at any one of the positions listed in Table 4 where Sexual amino acids.
[0186] Preferred exemplary pores include at least one pore having the following mutations relative to SEQ ID NO:3: Contains CsgG monomers: Y51X1 / N55X2 / F56X3 / N91R / K9 4Q / R97W / R192D-del(V105-I107), where X1 is I / V / S / T, X2 is N / I / V / S / T and / or X3 is Q / I / V / S / T.
[0187] Methods for introducing or substituting natural amino acids are well known in the art. For example, A methionine (M) is present at the relevant position in the polynucleotide encoding the mutant monomer. by replacing the methionine codon (ATG) with the arginine codon (CGT) , may be substituted with arginine (R). The polynucleotide is then expressed as discussed below. It can be realized.
[0188] Dual pores The CsgG / CsgF pore may be a dual pore comprising a first pore and a second pore. At least the first pore is a CsgG / CsgF pore as disclosed herein. The second pore may be a CsgG pore or a CsgG / CsgF pore. In the present specification, both the first pore and the second pore are CsgG / C as disclosed herein. The first and second pores may be the same or different. In addition to any of the mutations disclosed herein, in the dual pore, the CsgG monomer may It may contain one or more of the additional mutations described below.
[0189] In a dual pore, the first pore is formed by hydrophobic interactions and / or one or more disulfides. The first pore and / or the second pore may be attached to the second CsgG pore by binding. One or more of the monomers in the pore, e.g., 2, 3, 4, 5, 6, 8, 9, etc., all The polypeptides may be modified to enhance these interactions. This may be achieved in any suitable manner. Ugh.
[0190] At least one amino acid sequence in the first pore at the interface between the first pore and the second pore. Another cysteine residue is located at the interface between the first and second pores. may be disulfide bonded to at least one cysteine residue in the amino acid sequence The cysteine residues in the first pore and / or cysteine residues in the second pore are It may be a cysteine residue that is not present in the CsgG monomer. , 8 or 9 to 16, 18, 24, 27, 32, 36, 40, 45, 48, 54, 56 Multiple disulfide bonds, such as 63 or 64, can form between the two pores within the two pores. One or both of the first or second pores may comprise R97, I107, R110 of SEQ ID NO: 3, The first pore and the second pore are located at positions corresponding to Q100, E101, N102 and / or L113. A small number of monomers, such as up to 8, 9, or 10, containing cysteine residues at the interface between the two pores. The copolymer may contain at least one monomer.
[0191] At least one monomer in the first pore and / or the second pore One monomer may include at least one residue at the interface between the first pore and the second pore. Often, this residue is more hydrophobic than the residue present at the corresponding position in the wild-type CsgG monomer. For example, residues 3, 4, 5, 6, 7, 8, or 9, such as 2 to 10, are the first The residues of the first pore and / or the second pore are at the same positions as in the corresponding wild-type CsgG monomer. In some cases, the residues at the two membranes of the double pore are more hydrophobic. At least one of the first pores and the second pores at the interface between the first pore and the second pore is formed. The three residues are R97, I107, R110, Q100, E101, and N102 of SEQ ID NO: 3. and L113. At the interface within the wild-type CsgG monomer. If the residue is R, Q, N, or E, the hydrophobic residue is generally I, L, V, M, F, W, or or Y. If the residue at the interface in the wild-type CsgG monomer is I, the hydrophobic residue are typically L, V, M, F, W, or Y. Located at the interface within the wild-type CsgG monomer When the residue is L, the hydrophobic residue is generally I, V, M, F, W or Y.
[0192] The dual pores contain one or more monomers containing one or more cysteine residues at the interface between the pores. or may contain one or more hydrophobic residues at the interface between the pores, or It may contain a stain residue and one or more monomers containing such a hydrophobic residue. For example: R97, I107, R110, Q100, E101, N102 and / or is one or more positions (any 2, 3, or 4) in the monomer corresponding to the L113 position. The amino acid sequence of SEQ ID NO: 3 may include R97, I107, R108, R1109, R1110, R1120, R1130, R1140, R1150, R1160, R1170, R1180, R1190, R1200, R1210, R1220, R1230, R1240, R1250, R12 Monomers corresponding to positions 10, Q100, E101, N102 and / or L113 One or more of the positions (e.g., any 2, 3, or 4) in the may contain hydrophobic residues such as Y.
[0193] Thus, the dual pores may be located at one or more positions (e.g., 2, 3, 4, 5, 6, or 7) within the tail region. Although the residues may include bulky residues at the interface between the first and second pores, the residues are generally These residues are bulkier than those at the corresponding positions in the wild-type CsgG monomer. The size of the pore prevents the formation of holes in the pore wall at the interface between the first and second pores of the dual pore. At least one bulky residue at the interface between the first pore and the second pore Generally, A98, A99, T104, V105, L113, Q114 or is the position corresponding to S115. The residue at the interface in the wild-type CsgG monomer is A. When present, the bulky residues are generally I, L, V, M, F, W, Y, N, Q, S, or T. If the residues at the interface in the wild-type CsgG monomer are T, the bulky residues are generally L, M. , F, W, Y, N, Q, R, D or E. Located at the interface within the wild-type CsgG monomer If the residue is V, the bulky residues are generally I, L, M, F, W, Y, N, Q. If the residues at the interface within the CsgG monomer are L, the bulky residues are generally M, F, and W. , Y, N, Q, R, D or E. The residue at the interface in the wild-type CsgG monomer is Q. In the wild-type CsgG monomer, the bulky residues are generally F, W, or Y. If the residues at the face are S, the bulky residues are generally M, F, W, Y, N, Q, E, or R. is.
[0194] If the second pore is on the outside of the membrane, the second pore, and optionally the first pore, is a wild-type Csg G within the barrel region of the pore, reducing the negative charge within the barrel compared to the charge within the barrel of the pore These mutations make the barrel more hydrophilic. At least one monomer in the first pore and / or at least one monomer in the second pore. At least one monomer may comprise at least one residue within the barrel region of the pore, The residues at are less negatively charged than the residues at the corresponding positions in the wild-type CsgG monomer. The charge inside the barrel allows negatively charged analytes, such as polynucleotides, to be attracted by electrostatic charges. Sufficiently neutral or positively charged so as not to enter the pore. 149, E185, D195, E210 and / or E203 At least one residue (e.g., 2, 3, 4, or 5 residues) in the barrel region of the pore is a neuron. D149, E185, D19 of SEQ ID NO: 3 can be either a neutral or positively charged amino acid. 5, at least within the barrel region of the pore at a position corresponding to E210 and / or E203 At least one residue (e.g., 2, 3, 4, or 5 residues) may be N, Q, R, or K. preferable.
[0195] Specific examples of charge-eliminating mutations in SEQ ID NO:3 include: E185N / E2 03N, D149N / E185R / D195N / E201R / E203N, D149N / E185R / D195N / E201N / E203N, D149R / E185N / D195 N / E201N / E203N, D149R / E185N / E201N / E203N, D1 49N / E185N / D195 / E201N / E203N, D149N / E185N / E 201N / E203N, D149N / E185N / E203N, D149N / E185N / E201N, D149N / E203N, D149N / E201N / D195N, D14 9N / E201N, D195N / E201N / E203N, E201N / E203N, D 195N / E203, E203R, E203N, E201R, E201N, D195R, D195N, E185R, E185N, D149R and D149N.
[0196] At least one CsgG monomer in the first pore is located at the constriction of the barrel region of the first pore. and at least one residue in the CsgG pore that is more compact than the wild-type CsgG pore. Reducing, maintaining, or increasing the length of the contraction and / or at least Another CsgG monomer contains at least one residue in the constriction of the barrel region of the second pore. This residue may comprise a maintenance residue that reduces the length of the constriction compared to the wild-type CsgG pore. The length of the constriction in the first pore and / or the constriction in the second pore The length of the portion is at least as long as the wild-type pore, and more preferably longer. It's nice.
[0197] The length of the pore is increased by inserting residues in the region corresponding to the region between positions K49 and F56 in SEQ ID NO:3. The amino acid sequence may be increased by adding 1 to 5 (e.g., 2, 3, or 4) amino acids. An acid residue may be present at any one or more of the following positions as defined by reference to SEQ ID NO:3: The insertion may be performed at the following positions: K49 and P50, P50 and Y51, Y51 and P5 2, P52 and A53, A53 and S54, S54 and N55 and / or N5 5 and F56. A total of 1–10 (e.g., 2–8) or 3–5 amino acid residues are Preferably, all the monomers and / or Preferably, all monomers in the first or second pore have the same number of insertions in this region. The inserted residues are the length of the loop between residues corresponding to Y51 and N55 in SEQ ID NO:3. The inserted residues may be A, S, G, or T to maintain flexibility, or P to add a twist to the P, and / or P to allow the analyte to enter the pore under an applied potential difference. S, T, N, Q, M, F, W, and Y contribute to the signal generated when interacting with the barrel of The inserted amino acids may be any combination of S, G, S, V and / or I. Any combination of G, SGG, SGS, GS, GSS and / or GSG good.
[0198] In the dual pore, the constriction in the barrel of the first pore and / or the second pore may be used for analyte detection. Alternatively, when used for characterization, a first pore or a second pore having a wild-type constrictor is used. At least one residue (2, 3, 4 or more) that affects the properties of the pore compared to when the or 5 residues), where at least one residue in the constriction of the barrel region of the pore One residue corresponds to Y51, N55, Y51, P52 and / or A53 of SEQ ID NO:3. At least one residue is at a position corresponding to Q or F56 in SEQ ID NO:3. or V, A or Q at the position corresponding to Y51 in SEQ ID NO: 3, and / or It can also be a V in the position corresponding to N55.
[0199] The dual pore comprises at least one monomer in the first CsgG pore and / or a second CsgG pore. and at least one monomer within the sgG pore, wherein the monomer is as defined above. It includes two or more mutants defined as follows:
[0200] The CsgG monomer in the dual pore is R97, I107, R110, Q100 of SEQ ID NO: 3 , E101, N102, or L113.
[0201] The CsgG monomer in the dual pore is R97, Q100, I107, R110 of SEQ ID NO: 3 , E101, N102, and L113. This residue may be at the corresponding position in SEQ ID NO: 3 (any one of SEQ ID NOs: 68-88). is more hydrophobic than the residues present at the corresponding positions, etc., where R97 and / or The residue at the position corresponding to I107 is M, and the residue at the position corresponding to R110 is I , L, V, M, W or Y, and / or a position corresponding to E101 or N102 The residue at the position corresponding to Q100 is generally V or M. The residue at the position corresponding to Q100 is generally I, L, or V, M, F, W or Y, and / or the residue at the position corresponding to L113 is Generally I, V, M, F, W or Y.
[0202] The specific monomers are Y51A, F56Q substitutions and R97I / V / L / M / F / W / Y, I107L / V / M / F / W / Y, R110I / V / L / M / F / W / Y, Q100I / V / L / M / F / W / Y, E101I / V / L / M / F / W / Y, N102I / V / L / M / F / W / Y and L113CI / V / L / M / F / W / Y combinations, R97I / Combinations of V / L / M / F / W / Y and N102I / V / L / M / F / W / Y, and and / or R97I / V / L / M / F / W / Y and E101I / V / L / M / F / W / I107 has the sequence shown in SEQ ID NO: 3, which includes a combination of Y and They may already have formed aqueous interactions.
[0203] The CsgG monomer in at least one pore of the dual pore is A98, A99 of SEQ ID NO: 3 9, T104, V105, L113, Q114 and S115 a position corresponding to SEQ ID NO: 3 (e.g., a position corresponding to any one of SEQ ID NOs: 68 to 88) The residue may be bulkier than the residue at the corresponding position of T104. The residue at position L113 is L, M, F, W, Y, N, Q, D, or E. residues at position S115 are M, F, W, Y, N, G, D, or E, and / or at position S115 The residues at the corresponding positions are M, F, W, Y, N, Q, or E. A98 or A99 The residues at the positions corresponding to are generally I, L, V, M, F, W, Y, N, Q, S, or T The residue at the position corresponding to V105 is I, L, M, F, W, Y, N, or Q. The residue at the position corresponding to Q114 is F, W, or Y. The residue at the position corresponding to E210 is F, W, or Y. The residue at position 1 is N, Q, R or K.
[0204] Specific monomers include Y51A, F56Q substitutions and 1, 2, 3, 4, 5, 6 or more of the following: It may have the sequence shown in SEQ ID NO: 3 including all of the substitutions. A98I / L / V / M / F / W / Y / N / Q / S / T, A99I / L / V / M / F / W / Y / N / Q / S / T, T10 4N / Q / L / R / D / E / M / F / W / Y, V105I / L / M / F / W / Y / N / Q , L113M / F / W / Y / N / Q / D / E / L / R, Q114Y / F / W, and S1 15N / Q / M / F / W / Y / E / R.
[0205] At least one CsgG monomer in the dual pore is located at the corresponding position in SEQ ID NO:3. fewer residues than those at positions (such as the corresponding positions in any one of SEQ ID NOS: 68-88) Any one or more of negatively charged D149, E185, D195, E210 and E203 and residues in the barrel region of the pore at positions corresponding to those listed above, where D149, E The residues at positions corresponding to 185, D195 and / or E203 are K.
[0206] At least one CsgG monomer in the dual pore is located in the barrel region of the pore. It may comprise at least one residue within the constriction, which residue is different from the wild-type CsgG pore. at least one residue is located at the constriction of the wild-type CsgG pore This is in addition to the residues present in
[0207] The length of the pore is increased by inserting residues in the region corresponding to the region between positions K49 and F56 in SEQ ID NO:3. The amino acid sequence may be increased by adding 1 to 5 (e.g., 2, 3, or 4) amino acid residues. The group may be positioned at any one or more of the following positions as defined by reference to SEQ ID NO:3: K49 and P50, P50 and Y51, Y51 and P52, P52 and A53, A53 and S54, S54 and N55 and / or N55 and and F56. A total of 1 to 10 (e.g., 2 to 8) or 3 to 5 amino acid residues are The inserted residues are preferably those at positions Y51 and N52 of SEQ ID NO: 3. The length of the loop between residues corresponding to 55 can be increased. The inserted residues are required to maintain flexibility. A, S, G or T to add twist to the loop, P to add twist to the loop, and / or apply contributes to the signal generated when the analyte interacts with the barrel of the pore under a given potential difference. The suffix may be any combination of S, T, N, Q, M, F, W, Y, V and / or I. The inserted amino acids are S, G, SG, SGG, SGS, GS, GSS and / or G Any combination of SGs may be used.
[0208] The CsgG monomers in at least one of the pores within the dual pore are the corresponding wild-type monomers. at N55, P52 and / or A53 of SEQ ID NO: 3 which are different from the residues present in the mer and at least one residue in the constriction of the barrel region of the pore at a corresponding position. , where the residue at the position corresponding to N55 is V.
[0209] Any two or more of the above residues may be present in the same monomer. In particular, the monomer may be At least one cysteine residue, at least one hydrophobic residue, at least one bulky the residues that are neutral or positively charged, at least one of said neutral or positively charged residues, and / or the length of the constriction It may contain at least one of said residues that increases the
[0210] CsgG monomers within the dual pore induce contraction of the barrels of the first and / or second pores. The portion, when used to detect or characterize an analyte, is a first pore or at least one or more of the following may be used to affect the properties of the pores compared to when a second pore is used: The pore may additionally comprise the above residues (e.g., 2, 3, 4 or 5 residues), where At least one residue within the constriction of the allele region is selected from the group consisting of Y51, N55, Y51, N52, Y53, Y54, Y55, Y56, Y57, Y58, Y59, Y60, Y61, Y62, Y63, Y64, Y65, Y66, Y67, Y68, Y69, Y69, Y70, Y71, Y72, Y73, Y7 At least one residue is at a position corresponding to P52 and / or A53. Q or V at the position corresponding to F56 in SEQ ID NO:3, A or It may be Q, and / or V at the position corresponding to N55 of SEQ ID NO:3.
[0211] Methods for producing modified proteins Methods for introducing or substituting non-natural amino acids are also well known in the art. In this case, unnatural amino acids are used in the IVTT system to express mutant monomers. Alternatively, the aminoacyl-TrNA can be introduced by including a synthetic aminoacyl-TrNA in the Unnatural amino acids refer to the existence of synthetic (i.e., non-naturally occurring) analogs of those particular amino acids. In the presence of α-amino acid, mutant monomers were generated in E. coli that were auxotrophic for specific amino acids. The mutant monomer may be introduced by expressing a partial peptide synthesis. When produced using synthesis, they can be produced by naked ligation.
[0212] Monomers derived from CsgG can be synthesized, for example, by the addition of a streptavidin tag, or of a signal sequence that promotes secretion from cells that do not naturally contain such a sequence. Additionally, they may be modified to aid in their identification or purification. Lagging is discussed in more detail below. Monomers can be labeled with distinct labels. The revealing label may be any suitable label that allows the monomer to be detected. The appropriate label is listed below.
[0213] Monomers derived from CsgG may be produced using D-amino acids. Monomers derived from gG may contain a mixture of L- and D-amino acids. It is conventional in the art to produce such proteins or peptides.
[0214] Monomers derived from CsgG contain one or more specific modifications to promote nucleotide discrimination. The monomers derived from CsgG contain other non- Specific modifications may also be included. A number of non-specific side chain modifications are known in the art and may be included in CsgG. Such modifications may be made to the side chains of the derived monomers. Reductive alkylation of amino acids by the reaction, followed by reduction with NaBH4, acetone These include amidation with methyl amidate or acylation with acetic anhydride.
[0215] Monomers derived from CsgG can be produced using standard methods known in the art. Monomers derived from CsgG can be produced synthetically or by recombinant means. For example, the monomer can be prepared by in vitro translation and transcription (IVTT). Suitable methods for producing the pores and monomers are described in International Application WO 2002 / 020944. 10 / 004273, WO 2010 / 004265 or WO 2010 / 08660 3. Methods for inserting pores into membranes are well known.
[0216] Two or more CsgG monomers within the pore may be covalently linked to each other. For example, at least 2 pcs, at least 3 pcs, at least 4 pcs, at least 5 pcs, at least 6 pcs, at least 7 pcs at least 8, at least 9 or at least 10 monomers can be covalently bonded The covalently bonded monomers may be the same or different.
[0217] The monomers can be genetically fused, optionally via a linker, or via, for example, a chemical cross-linker. The method for covalently bonding the monomers is described in WO2017 / 1493 16, as disclosed in WO2017 / 149317 and WO2017 / 149318 .
[0218] In some embodiments, the mutant monomer is chemically modified. The mutant monomer may be chemically modified in any manner and at any site. of cysteine (cysteine bond), attachment of a molecule to one or more lysines, one or more non-natural Chemically modified by attachment of molecules to amino acids, enzymatic modification of epitopes, or terminal modifications Suitable methods for carrying out such modifications are well known in the art. The mutant monomer may be chemically modified by the attachment of any molecule, e.g. , the mutant monomers can be chemically modified by the attachment of dyes or fluorophores.
[0219] In some embodiments, the mutant monomers are mutated by combining the monomer with a target nucleotide sequence or target Chemical modification with a molecular adaptor that promotes interaction between the target polynucleotide sequence and the pore containing the target polynucleotide sequence. The presence of the adaptor is determined by the presence of the pore and the nucleotide or polynucleotide sequence. improve the host-guest chemistry of the mutant monomers, thereby optimizing the alignment of the pores formed from the mutant monomers. Improved sequencing capabilities. The principles of host-guest chemistry are well known in the art. Adapters The physical properties of the pore improve its interaction with the nucleotide or polynucleotide sequence. Adapters have an effect on the electrical or chemical properties of the pore barrel or channel. altering charge or interacting specifically with nucleotide or polynucleotide sequences The pore may be attached to or attached to the pore, thereby facilitating its interaction with the pore.
[0220] Molecular adapters include circular molecules, cyclodextrins, hybridization-competent species, and DNA binding agents or interchelators, peptides or peptide analogues, synthetic polymers, aromatic monolayers It is preferred that the molecule be a planar molecule, a small molecule with a positive charge, or a small molecule capable of hydrogen bonding. It's nice.
[0221] The adaptor may be cyclic. Preferably, the cyclic adaptor has the same symmetry as the pore. The adapter is preferably designed to allow CsgG to be organized into subunits, typically 8 or 9 subunits around a central axis. Since the symmetry is 8-fold or 9-fold, it is preferable to have 8-fold or 9-fold symmetry. This is explained in detail below.
[0222] Adapters are generally synthesized by host-guest chemistry using nucleotide sequences or polynucleotides. Adapters generally interact with nucleotide sequences or polynucleotides. The adaptor has the ability to interact with a nucleotide sequence or polynucleotide. The one or more chemical groups have the ability to interact with a peptide sequence. Aqueous interactions, hydrogen bonding, Van der Waal forces, π-cation interactions and and / or by non-covalent interactions, such as electrostatic forces, to bind nucleotides or polynucleotides. It is preferred that the sequence interacts with the nucleotide or polynucleotide sequence. Preferably, the one or more chemical groups capable of interacting are positively charged. One or more chemical groups capable of interacting with a nucleotide sequence or a polynucleotide sequence are More preferably, the amino group comprises a primary carbon atom, a secondary carbon atom, or an amino group. The adapter can be attached to a tertiary carbon atom. The adapter can be a ring of 6, 7, or 8 amino groups, etc. More preferably, the adaptor comprises a ring of eight amino groups. Most preferably, the ring of protonated amino groups is It may interact with negatively charged phosphate groups within the tide sequence.
[0223] Correct positioning of the adaptor within the pore is achieved by combining the adaptor with the mutant monomer. The adaptor can be attached to one or more amino acids within the pore. Preferably, the adaptor comprises one or more chemical groups capable of interacting with a carboxylic acid. Hydrophobic interactions, hydrogen bonds, Van der Waal forces, π-cation interactions and interact with one or more amino acids within the pore through non-covalent interactions, such as ATP and / or electrostatic forces. More preferably, the pore contains one or more chemical groups capable of acting on the one or more The chemical groups that can interact with amino acids are typically hydroxyls or amines. The hydroxyl group can be attached to a primary, secondary, or tertiary carbon atom. The hydroxyl groups may form hydrogen bonds with uncharged amino acids within the pore. Any molecule that facilitates the interaction between the pore and the nucleotide or polynucleotide sequence. Adapters can be used.
[0224] Suitable adaptors include cyclodextrins, cyclic peptides, and cucurbiturils. The adaptor may be a cyclodextrin or Cyclodextrin or its derivatives are preferably used. Cyclodextrin or its derivatives are disclosed in the following documents: Eliseev, AV, and Schneid er, H. J. (1994) J. Am. Chem. Soc. 116, 6081-6088. The adapter is heptakis-6-amino-β-cyclodextrin (am7-βCD), 6-monodeoxy-6-monoamino-β-cyclodextrin (a m1-CD) or heptakis-(6-deoxy-6-guanidino)-cyclodextrin More preferably, the guanidino group in gu7-βCD is , which has a much higher pKa than the primary amines in am7-βCD and therefore is more positively charged. This gu7-βCD adaptor extends the dwell time of the nucleotide within the pore. This increases the accuracy of the measured residual current and also reduces the salt content at high temperatures or low data acquisition rates. It can be used to increase the rate of radical detection.
[0225] Succinimidyl 3-(2-pyridyldithio)propionate (SPDP) crosslinker is When used as discussed in more detail in, the adapter is heptakis(6-deoxyribonucleic acid). (6-amino)-6-N-mono(2-pyridyl)dithiopropanoyl-β-cyclodecane Preferably, it is stringin (am6amPDP1-βCD).
[0226] More suitable adaptors include gamma-cyclodextrin, which contains nine sugar units (see (Thus, it has 9-fold symmetry.) γ-Cyclodextrin may contain a linker molecule. less than or equal to all of the modified sugar units used in the β-cyclodextrin example above. It may be modified to include more than one.
[0227] The molecular adaptor may be covalently attached to the mutant monomer. The adaptor may be covalently attached to the pore using any method known in the art. When the molecular adaptor is attached via a cysteine bond, One or more cysteines are introduced into the mutant by substitution (e.g., within the barrel) Preferably, the mutant monomer comprises one or more cysteines in the mutant monomer. One or more cysteines can be chemically modified by the attachment of a molecular adaptor to the naturally occurring It may be native, i.e., at positions 1 and / or 215 in SEQ ID NO: 3. In this method, the mutant monomer may have one or more cysteines introduced at other positions. The cysteine at position 215 can be chemically modified by attachment of a molecule to a molecular adaptor. is attached to a cysteine at position 1 or at that position other than a cysteine introduced at another position To ensure that the nucleotide sequence is not modified, it may be removed, for example by substitution.
[0228] The reactivity of cysteine residues can be enhanced by modifying adjacent residues. The basic group of the arginine, histidine, or lysine residue is a more reactive S -The reactivity of cysteine residues is improved by the addition of thiol groups such as dTNB. These can be protected by amine protecting groups. These are used to protect the mutant monolayers before the linker is attached. The cysteine residues of the mer may be reacted with one or more cysteine residues of the mer.
[0229] The molecule may be attached directly to the mutant monomer. The molecule may be attached using a chemical crosslinker or a peptide. Preferably, the mutant monomer is attached to the mutant monomer using a linker such as a linker.
[0230] Suitable chemical cross-linking agents are well known in the art. Preferred cross-linking agents include 2,5-dioxopyrrolidone, Roridin-1-yl 3-(pyridin-2-yldisulfanyl)propanoate, 2,5- dioxopyrrolidin-1-yl 4-(pyridin-2-yldisulfanyl)butanoate, and 2,5-dioxopyrrolidin-1-yl 8-(pyridin-2-yldisulfanyl The most preferred cross-linking agent is succinimidyl 3-(2-pyridinyl) octanoate. Typically, the molecule / crosslinker complex is suddenly Before being covalently linked to the mutant monomer, the molecule is covalently linked to a bifunctional crosslinker; Before the bifunctional crosslinker / monomer complex is bound to a molecule, the bifunctional crosslinker is covalently attached to the monomer. It is also possible to covalently bond the amine to the monomer.
[0231] Preferably, the linker is resistant to dithiothreitol (DTT). Linkers include iodoacetamide-based and maleimide-based linkers. Examples include, but are not limited to:
[0232] In other embodiments, the monomer may be attached to a polynucleotide binding protein. This forms a modular sequencing system that can be used in the sequencing methods of the present invention. Polynucleotide binding proteins are discussed below.
[0233] The polynucleotide binding protein is preferably covalently attached to the mutant monomer. The protein may be covalently attached to the monomer using any method known in the art. The monomer and protein may be chemically fused or genetically fused. When the entire organism is expressed from a single polynucleotide sequence, the monomers and proteins are Genetic fusion of a monomer to a polynucleotide-binding protein is , discussed in WO 2010 / 004265.
[0234] If the polynucleotide binding protein is attached via a cysteine bond, one or more Preferably, the above cysteines are introduced into the mutant by substitution. The tein has low conservation among homologs, indicating that mutations or insertions can be tolerated. Therefore, it is preferable to introduce the polynucleotide binding protein into the loop region. In such an embodiment, the native cysteine at position 251 is The reactivity of cysteine residues can be enhanced by modification as described above. Ugh.
[0235] The polynucleotide binding protein may be a mutant monomer or may be linked via one or more linkers. The molecule can be directly linked to the hydroxyl group as described in WO 2010 / 086602. A hybridization linker may be used to attach the mutant monomer. A peptide linker may be used. A peptide linker is an amino acid sequence. The length, flexibility, and hydrophilicity of the polymer are typically chosen so as not to interfere with the function of the monomer and molecule. Preferred flexible peptide linkers are designed to contain serine amino acids and / or is 2 to 20 amino acids long (e.g., 4, 6, 8, 10, or 16 amino acids), such as glycine amino acids More preferred flexible linkers include (SG)1, (SG)2, (SG)3, ( (SG)4, (SG)5 and (SG)8, where S is serine and G is glycosylated. Preferred rigid linkers are 2 to 30 proline amino acids in length (4, 6, A more preferred rigid linker is (P) 12 Including this where P is proline.
[0236] chemical modification The mutant CsgG monomer or CsgF peptide binds to the molecular adapter and polynucleotide The leukocyte-binding protein may be chemically modified.
[0237] The molecule (monomer or peptide chemically modified) is described in WO 2010 / 004 273, WO 2010 / 004265, or WO 2010 / 086603 It may be attached directly to the monomer or peptide, as described above, or may be attached via a linker. It may be attached via a
[0238] Any of the proteins described herein, such as CsgG monomers and / or CsgF peptides Proteins can be, for example, histidine residues (his tag), aspartic acid residues (asp tag), Add streptavidin tag, Flag tag, SUMO tag, GST tag, and MBP tag or secretion from cells in which the polypeptide does not naturally contain a signal sequence. modified to aid in their identification or purification by the addition of a signal sequence that promotes transcription. An alternative to introducing a genetic tag is to place the tag at a native or engineered position on the protein. One example of this is to chemically react a gel-shift reagent with the outside of a protein. The first step is to react the engineered cysteine with the hemolysin hetero-oligomer. (Chem Biol. 1997 Jul, 4( 7):497-505).
[0239] Any of the proteins described herein, such as CsgG monomers and / or CsgF peptides The protein can be labeled with a distinct label. The distinct label is used to identify the protein by which it can be detected. Suitable labels include fluorescent molecules, radioisotopes, for example 125 I, 35 S, enzymes, antibodies, antigens, polynucleotides, and biotin These include, but are not limited to, ligands.
[0240] Any of the peptides described herein, such as CsgG monomers and / or CsgF peptides Proteins can be made synthetically or by recombinant means. can be synthesized by in vitro translation and transcription (IVTT). The amino acid sequence may contain unnatural amino acids or may be modified to increase protein stability. When proteins are produced by synthetic means, these amino acids may be modified into Proteins can also be introduced after production, either synthetically or recombinantly. may also be changed.
[0241] Proteins may be produced using D-amino acids. For example, proteins may be produced using L-amino acids. The protein or peptide may contain a mixture of D-amino acids and D-amino acids. It is conventional in the art to produce
[0242] The protein may also contain other non-specific modifications, so long as they do not interfere with the function of the protein. A number of non-specific side chain modifications are known in the art and may be made to the side chains of proteins. Such modifications include, for example, reductive alkylation of amino acids by reaction with aldehydes; Subsequent reduction with NaBH4, amidation with methyl acetimidate, or This includes acylation with aqueous acetic acid.
[0243] Any of the peptides described herein, such as CsgG monomers and / or CsgF peptides Proteins can be produced using standard methods well known in the art. Polynucleotide sequences encoding proteins can be derived using standard methods in the art. Polynucleotide sequences encoding proteins may be obtained and replicated. Proteins may be expressed in bacterial host cells using standard techniques. It can be produced intracellularly by expression of the polypeptide from a recombinant expression vector. The present vector optionally carries an inducible promoter to control the expression of the polypeptide. These methods are described in the following references: Sambrook, J. and Rus sell, D. (2001). Molecular Cloning: A La boratory Manual, 3rd Edition. Cold Sprin g Harbor Laboratory Press, Cold Spring H Arbor, NY.
[0244] Proteins can be derived from protein-producing organisms or after recombinant expression. It can be produced on a large scale after purification by a high-performance liquid chromatography system. Typical protein liquid chromatography systems include FPLC, AKTA systems, and B io-Cad system, Bio-Rad BioLogic system, and Gilson HPLC system included.
[0245] Method for generating pores In a third aspect, the present invention provides a method for detecting two or more contractions in vivo and in vitro. Methods for producing a CsgG:modified CsgG pore complex bearing a site are provided. The form comprises a CsgG pore or a homologue or mutant form thereof and a modified CsgF peptide. or a homologue or mutant thereof. The method comprises the step of: expressing the protein or its homologue or mutant form, and modifying or cleaving the protein. The CsgF monomer (both in suitable host cells) was then expressed and subjected to in vivo complexation. and forming a modified CsgG pore. It contains a modified CsgF peptide, providing the pore with an additional reader head. The resulting pore complexes produced by the method using the F peptide are capable of binding to analytes, particularly characterization of target analytes, such as nucleic acid sequencing, to allow passage of polynucleotides and providing sufficient structure for the porous composite to be used in the appropriate setting for said application. two or more reader heads for improved reading of the polynucleotide sequence when Includes:
[0246] More specifically, the modified CsgF peptides expressed in the method are those represented by SEQ ID NOs: 8, 10, 1 2, or 14, or their homologs. These sequences are The method involves introducing a constriction site into the pore complex to bind the CsgG protein to the pore and induce biotinylated proteins. The objective of this study is to limit the CsgF fragments that can acquire biological pores.
[0247] Formed by CsgG and CsgF proteins or similar Another method to generate isolated pore complexes is to reconstitute the monomers in vitro. The method involves obtaining functional pores by using a suitable system that allows complex formation. stem, and the mature CsgG molecule depicted in SEQ ID NO: 3 or a homologue or mutant thereof. contacting the molecule with a modified CsgF peptide or a homologue or mutant thereof. The system may be an "in vitro system", which is an in vitro system It includes at least the elements and circumstances necessary for the execution of the law, and is in line with the normal natural environment. Using external biological molecules, organisms, cells (or parts of cells) and whole organisms allow for more detailed, more convenient, or more efficient analysis than can be performed using An in vitro system also refers to a system that is capable of providing a suitable The composition may contain an appropriate buffer solution, in which the protein components for forming the complex are added. A person skilled in the art is aware of options for providing such a system. In certain embodiments, for in vitro reconstitution, the modifications applied in the method The decorated CsgF peptide or the like may be SEQ ID NO: 15 or SEQ ID NO: 16, or is a peptide, including its mutants or homologues, which may be synthesized or recombinantly Alternatively, SEQ ID NOs: 40, 39, 38, or 37, 15, 54, 55 or The modified CsgF peptide, including its homologues or mutants, can be used in the method to produce Cs and contacting the CsgG- or CsgG-like pore to form a pore complex.
[0248] The CsgG / CsgF pore can be produced by any suitable method. An example method is described.
[0249] In one embodiment, the CsgG / CsgF pore may be produced by co-expression. So, we have a CsgG monomer polypeptide (or mutant polypeptide) in one vector. at least one gene encoding a nucleotide sequence ... A gene encoding a full or truncated CsgF polypeptide (which may be a mutant polypeptide) The gene is transformed together to express the protein, which then forms the complex within the transformed cell. This can be done in vivo or in vitro. Alternatively, two genes encoding the CsgG and CsgF polypeptides can be expressed as , under the control of a single promoter, or under the control of two separate promoters (which are the same The genes can be placed in one vector under the control of a gene encoding the nucleotide sequence (which may be one or different).
[0250] In another embodiment, the CsgG / CsgF pore comprises a CsgG molecule separate from the CsgF peptide. The CsgG monomer or CsgG pore is generated by expressing a small amount of CsgG. vectors encoding at least one CsgG monomer, or It may be purified from cells transformed with one or more vectors expressing the C The sgF peptide was prepared by transforming the cells with a vector encoding at least one CsgF peptide. The purified CsgG monomer / pore may then be purified from the cells. Pore complexes may be produced by incubation with peptides.
[0251] In another embodiment, the CsgG monomer and / or CsgF peptide are synthesized in vitro. The CsgG monomer is then converted to C The pore complex may be generated by incubation with the sgF peptide. Use of this method is shown in Figure 14 is illustrated in.
[0252] The above embodiments may be embodied in, for example, (i) CsgG is expressed in vivo and CsgF is expressed in vivo; (ii) CsgG produced in vitro and CsgF produced in vivo (iii) CsgG produced in vivo and CsgF produced in vitro; or (iv) CsgG is produced in vitro and CsgF is produced in vitro. This may also be done.
[0253] One or both of the CsgG monomer and the CsgF peptide may be tagged to facilitate purification. Purification can be performed using CsgG monomers and / or CsgF peptides tagged with CsgG. This can be done even if the catalyst is not present in the catalyst. The pore components are purified by various methods (e.g., exchange, gel filtration, and hydrophobic interaction column chromatography). They can be used alone or in different combinations to
[0254] Any known tag can be used on either of the two proteins. To purify the CsgG:CsgF complex from the CsgG pore and CsgF, two methods were used. For example, a Strep tag can be used for CsgG. A His tag can be used with CsgF, or vice versa. 13, where two proteins are purified separately and then mixed together. Then, when Strep-purified and His-purified again, similar end results were obtained. It is possible.
[0255] When the full-length CsgF protein forms a complex with CsgG, the network of CsgF (Figure 4B) The q domain and head domain protrude from the β-barrel of the CsgG pore.
[0256] Therefore, CsgG pores and pores containing full-length CsgF were used in single-channel recording experiments. In this case, the head domain may hinder or prevent the insertion of a pore into the membrane. , which may prevent the analyte from passing through the pore. Therefore, when inserting a pore into the membrane, If we could reduce the number of flexible polypeptides hanging from the β-barrel, It mimics the FCP domain resolved in the cryo-EM structure of the complex and is structurally perfect. Provided herein are truncated versions of the CsgF protein that maintain its integrity.
[0257] The CsgG / CsgF pore is generated either before or after the CsgG pore is inserted into the membrane. Truncation mutants can be used when the pore complex is prepared prior to insertion into the membrane. However, it is preferred that CsgG and the CsgG complex be used in situ. Insertion of the CsgG pore into the membrane followed by addition of the CsgF peptide, as formed by For example, in one embodiment, the system allows access to the through-side of the membrane. (e.g., in a chip or chamber for electrophysiology measurements) As shown in Fig. 1, the CsgG pore is inserted into the membrane, and then the CsgF peptide penetrates the membrane. In any embodiment where the CsgG pore is formed in situ, Larger CsgF peptides may also be used. For example, the CsgF peptide may be CsgF( may include all or part of the neck domain of SEQ ID NO:6 (from about residue 36 of SEQ ID NO:6). In some embodiments, CsgF comprises all of the neck domain and part of the head domain. (residues 36 to XX of SEQ ID NO: 6).
[0258] Depending on the method for making the complex and the stability of the complex with a particular cleavage, CsgG:Cs The gF and CsgG:FCP complexes can be generated in different ways.
[0259] In one embodiment, a truncated version of the CsgF polypeptide at the required length is used directly. will be done.
[0260] In another embodiment, the full-length polypeptide of CsgF, or a longer version than the desired truncation, is used. (long enough to keep the complex stable), where it is susceptible to cleavage by proteases. Therefore, a protease cleavage site is required to generate a CsgF peptide of the desired length. Inserted (e.g., TEV, HRV 3 or any other protease cleavage site) In this embodiment, once the CsgG / CsgF complex is formed, a protease is required. Alternatively, a protease can be used to cleave CsGF at a site of interest. It may be used to generate CsgF peptides prior to assembly.
[0261] Some protease sites leave additional tags after cleavage. For example, the TEV protease cleavage site The cleavage sequence is ENLYFQS. TEV protease cleaves proteins between Q and S. This leaves the ENLYFQ intact at the C-terminus of the CsgF peptide. A modified CsgF containing a TEV cleavage site was used, which was cleaved using TEV protease In another example, the HRV C3 cleavage site is LEVLFQGP and the enzyme is It cleaves between Q and G, leaving LEVLFQ intact at the C-terminus of the CsgF peptide.
[0262] Methods for characterizing analytes In a further aspect, the present invention provides a method for determining the presence, absence, or one or more characteristics of a target analyte. The method provides a method for detecting a target analyte in association with (entering or penetrating) a pore channel. a target analyte to pass through the isolated pore complex or transmembrane pore, contacting the analyte with one or more pores (such as the pores of the present invention) and and taking measurements on the sample to thereby determine the presence, absence, or one or more characteristics of the analyte. The target analyte is also called the template analyte or the analyte of interest. Isolated pore complexes generally contain 7, 8, 9, or 10 CsgG monoclonals. At least 7, at least 8, at least 9, or at least 10 The isolated pore complex contains eight or nine identical CsgG monomers. Preferably, the CsgG monomer comprises one or more (2, 3, 4, 5, 6, 7, 8 Preferably, the CsgF peptide is chemically modified, such as by the addition of a nucleotide sequence to the CsgF peptide. The CsgG monomer or its homologues or mutants are chemically modified. The isolated pore complex monomer, such as a monomer or a homolog or mutant thereof, The analyte may be from any organism. After passing through the CsgG constrictor, it passes through the CsgF constrictor. In another embodiment, the analyte may pass through the CsgF constriction and then bind to the CsgG in the membrane. It may pass through the CsgG constriction depending on the orientation of the CsgG / CsgF complex.
[0263] The method is to determine the presence, absence, or one or more characteristics of a target analyte. for determining the presence, absence, or one or more characteristics of at least one analyte, The method may determine the presence, absence, or one or more characteristics of two or more analytes. The method may involve analyzing any number of analytes (2, 5, 10, 15, 20, 30, , 40, 50, 100 or more analytes) to determine the presence, absence, or one or more characteristics of the The method may include determining any number of characteristics (1, 2, 3, 4, 5, 10 or more) of one or more analytes. characteristics of the
[0264] The molecules located within the channel of the pore complex or near one of the openings of the channel The binding of molecules has an effect on the open channel ion flow through the pore, which is This is the essence of "molecular sensing" of open channels in a manner similar to nucleic acid sequencing applications. The change in ion flow can be measured by a change in electrical current using appropriate measurement techniques ( For example, WO 2000 / 28312 and D. Stoddart et al., Proc. Natl. Acad. Sci., 2010, 106, 7702- 7 or WO 2009 / 077734). The decrease in ion flow measured by the decrease in current. The degree of smallness is related to the size of the obstacles in the pores or the proximity of the pores. or nearby binding of a molecule of interest (also called an "analyte") that is detectable and measurable This provides a fundamental event, thereby forming the basis for a "biological sensor." Suitable molecules for this purpose include nucleic acids, proteins, peptides, polysaccharides, and small molecules (e.g., pharmaceuticals). low molecular weight or inorganic compounds (such as drugs, toxins, cytokines, and pollutants) As used herein, low molecular weight (e.g., < 900 Da or < 500 Da) organic or inorganic Detecting the presence of biomolecules is a key step in personalized drug development. development, medical, diagnostics, life science research, environmental monitoring, and security and and / or have applications in the defense industry.
[0265] In another embodiment, an isolated pore complex, or wild-type or modified E. coli Csg G nanopore or its homologues or mutants, and the channel constriction is provided to the pore in the complex. The transmembrane pore complex containing the modified CsgF peptide is a molecular sensor or biological In some embodiments, the CsgG nanopore can act as a sensor for the bacterial protein (e.g., E. coli, Salmonella typhi) derived or isolated from In some embodiments, the CsgG nanopore may be recombinantly produced. The analyte detection procedure is described in Howaka et al. Nature Biotechn ology (2012) Jun 7, 30(6):506-7. The analyte molecules to be detected may be bound to either surface of the channel or to the surface of the channel itself. The location of binding may be determined by the size of the molecule being sensed. .
[0266] The target analytes include metal ions, inorganic salts, polymers, amino acids, peptides, and polypeptides. peptides, proteins, nucleotides, oligonucleotides, polynucleotides, polysaccharides, Dyes, bleaches, pharmaceuticals, diagnostics, recreational drugs, explosives, toxic compounds, and environmental pollutants are preferred. The method comprises the steps of: Determining the presence, absence, or one or more characteristics of two or more analytes of the same type, such as drugs Alternatively, the method may involve detecting one or more proteins, one or more nucleic acids, The presence or absence of two or more analytes of different types, such as methicillin and one or more pharmaceutical agents. , or may involve determining one or more characteristics.
[0267] The target analyte may be secreted from the cell. Alternatively, the target analyte may be secreted from the cell prior to carrying out the method. Analytes may be present intracellularly, so they must be extracted from the cells. .
[0268] The wild-type pore may act as a sensor, but the binding strength and binding position of the molecule to be sensed may vary. modified via recombinant or chemical methods to increase the position or specificity of binding. Typical modifications include the addition of features complementary to the structure of the molecule being sensed. When the analyte molecule comprises a nucleic acid, the binding moiety may be a cyclode. In the case of small molecules, this may include a complementary a suitable binding domain, such as a single chain variable fragment (scFv) domain or a T cell receptor (TCR)-derived domain It may be an antigen-binding portion of an antibody molecule or a non-antibody molecule, including the antigen recognition domain of interest; In the case of proteins, it may be a known ligand of the target protein. The live or modified E. coli CsgG nanopore or its homologues are then injected with the appropriate antigen (E. coli). and the ability to act as a molecular sensor to detect the presence in a sample of a compound (containing a pitope). The antigens can be cell surface antigens (receptors), solid tumor cells, or blood cancer cells (e.g., markers of lymphoma or leukemia), viral antigens, bacterial antigens, protozoan antigens Antigens, allergens, allergy-related molecules, albumin (e.g., human, rodent, or rat) Fluorescent molecules (including fluorescein), blood group antigens, small molecules, drugs, enzymes, enzyme catalysts The catalytic moiety or enzyme substrate, and transition state analogs of the enzyme substrate, may be included. Modifications can be achieved using known genetic engineering and recombinant DNA techniques. The positioning of the fit depends on the properties of the molecule being sensed, e.g., size, three-dimensional structure, and its bioactivity. The selection of the suitable structure may be performed using computational structure design. Determination and optimization of protein-protein or protein-small molecule interactions and BIAcore®, which detects molecular interactions using surface plasmon resonance. Which technologies can be used to investigate (BIAcore, Inc., Piscataway , NJ; see also www.biacore.com).
[0269] In one embodiment, the analyte is an amino acid, peptide, polypeptide, or protein. The amino acid, peptide, polypeptide, or protein may be naturally occurring or non-naturally occurring. Polypeptides or proteins may contain within them synthetic or modified amino acids. Several different types of modifications to amino acids are known in the art. Suitable amino acids and modifications thereof are described above. Of course, the target analyte The modifications may be made by any method available in the art.
[0270] In another embodiment, the analyte is a polynucleotide, such as a nucleic acid, which contains two or more nucleotides. Nucleic acids are particularly suitable for nanopore sequencing. The natural nucleobases in DNA and RNA can be distinguished by their physical size When a nucleic acid molecule or individual bases pass through the channel of a nanopore, the size difference between the bases is called the chirality. This causes a directly correlated decrease in the ion flux through the channel. The fluctuations in ion flux are recorded. Suitable electrical measurement techniques for recording fluctuations in ion currents are described, for example, in WO 200 0 / 28312 and D. Stoddart et al., Proc. tl. Acad. Sci., 2010, 106, pp 7702-7 (single multi-channel recording equipment), and for example WO 2009 / 077734 (multi-channel recording With proper calibration, the characteristic reduction in ion current can be used to measure the real In time, the channel can identify specific nucleotides and related bases. In conventional nanopore nucleic acid sequencing, individual nucleotides of a nucleic acid sequence of interest are identified by nucleotide sequence analysis. When passing through the nanopore channels sequentially, the open channel is partially blocked by the The ion current is reduced. It is this ion current that is measured using the appropriate recording techniques described above. The decrease in ion flux is a measure of the flow of a known nucleotide through the channel. This may be calibrated against the decrease in ion current caused by the nucleotides in the channel. and therefore, when performed sequentially, determine whether the nanopore This provides a method for determining the nucleotide sequence of a nucleic acid. In this case, the size of the individual nucleotides passing through the constriction (or "leading head") It is generally required that the magnitude of the increase in the ion flux through the channel be directly correlated with the decrease in the ion flux through the channel. Of course, sequencing can be achieved by, for example, pore-forming via the action of an associated polymerase. Alternatively, the sequence may be By passing the nucleotide triphosphate groups sequentially removed from the target nucleic acid in close proximity to the pore, (see, for example, WO 2014 / 187924).
[0271] A polynucleotide or nucleic acid can contain any combination of any nucleotides. The nucleotides can be natural or artificial. One or more nucleotides of a polynucleotide can be lost by oxidation or methylation. For example, the polynucleotide may contain pyrimidine dimers. The body is typically associated with UV damage, which is the leading cause of cutaneous melanoma. One or more nucleotides within the nucleotide sequence may be selected from suitable nucleotide sequences known to those skilled in the art, for example. For example, a polynucleotide may be modified with a label or tag. A nucleotide typically comprises a nucleobase, a sugar and at least one spacer. Nucleobases and sugars form nucleosides. Nucleobases are typically are heterocyclic. Nucleobases include purines and pyrimidines and more specifically adenine It contains cytosine (C), uracil (U), guanine (G), thymine (T), and cytosine (C). The sugar is typically a pentose sugar. Nucleotide sugars include, but are not limited to: Sugars include, but are not limited to, ribose and deoxyribose. The polynucleotide preferably contains deoxyadenosine (dA), Deoxyuridine (dU) and / or thymidine (dT), deoxyguaniosine (d Preferably, the nucleotides include deoxycytidine (dG) and deoxycytidine (dC). A nucleotide is typically a ribonucleotide or a deoxyribonucleotide. Contains monophosphate, diphosphate, or triphosphate. Nucleotides contain three or more phosphates. The phosphate may be located at the 5' or 6' end of the nucleotide. The nucleotides in the polynucleotide may be attached to the 3' end in any manner. Nucleotides typically consist of their sugar and phosphorus units in nucleic acids. Nucleotides are attached by their nucleobases in pyrimidine dimers. The polynucleotide may be single-stranded or double-stranded. Preferably, at least a portion of the polynucleotide is double-stranded. Most preferably, it is ribonucleic acid (RNA) or deoxyribonucleic acid (DNA). In addition, the method using a polynucleotide as an analyte includes (i) determining the length of the polynucleotide. (ii) identity of the polynucleotide; (iii) sequence of the polynucleotide; (iv) The secondary structure of the polynucleotide and (v) whether the polynucleotide is modified. The method includes determining one or more characteristics selected from the group consisting of:
[0272] The polynucleotide can be of any length (i). For example, The length should be at least 10, at least 50, at least 100, at least 150, At least 200, at least 250, at least 300, at least 400, or less Each polynucleotide can be up to 500 nucleotides or nucleotide pairs. The length of the nucleotide is 1000 or more nucleotides or nucleotide pairs, 5000 or more nucleotides or nucleotide pairs. nucleotides or nucleotide pairs, or nucleotide pairs greater than or equal to 100,000 in length Any number of polynucleotides can be interrogated. For example, the method can involve 2, 3, 4, 5, 6 , 7, 8, 9, 10, 20, 30, 50, 100 or more polynucleotides. When two or more polynucleotides are characterized, they may be related to different polynucleotides. A polynucleotide can be an example of a polynucleotide or two identical polynucleotides. It can be natural or artificial, for example, an oligonucleotide produced using a method The sequence may be verified, and this method is usually carried out in vitro.
[0273] The nucleotides can have any identity (ii) and can be adenosine monophosphate (AM P), guanosine monophosphate (GMP), thymidine monophosphate (TMP), uridine monophosphate (UMP), 5-methylcytidine monophosphate, 5-hydroxymethylcytosine monophosphate, adenosine monophosphate (CMP), cyclic adenosine monophosphate (cAMP), cyclic guanosine monophosphate (CGMP) cGMP, deoxyadenosine monophosphate (dAMP), deoxyguanosine monophosphate (dAMP) deoxythymidine monophosphate (dGMP), deoxyuridine monophosphate (dTMP), deoxyuridine monophosphate (dUMP), deoxycytidine monophosphate (dCMP) and deoxymethylcytidine monophosphate (dCMP). Nucleotides include, but are not limited to, AMP, TMP, GMP, CM dAMP, dTMP, dGMP, dCMP and dUMP. The nucleotide may be abasic (i.e., lacking a nucleobase). A nucleotide may also lack a nucleobase and a sugar (i.e., a C3 spacer). The sequence of nucleotides (iii) is the sequence of the polynucleotide in the 5' to 3' direction of the strand. The sequence is determined by the identity of the following nucleotides joined together throughout:
[0274] CsgG pores and pores containing CsgF peptides are particularly useful for homopolymer analysis For example, a pore can be used to separate two or more polynucleotides, such as identical consecutive nucleotides. Nucleotides, e.g., at least 3, 4, 5, 6, 7, 8, 9, or 10 polynucleotides The sequence of the nucleotide may be determined. For example, the nucleotide may be a sequence of polyA, polyT, poly may be used to sequence polynucleotides containing G and / or polyC regions. good.
[0275] The CsgG pore constriction is made by residues at positions 51, 55, and 56 of SEQ ID NO:3. The leader head of CsgG and its contraction mutants are generally sharp. As it passes through the constriction, it interacts with the leader head of the pore at any given time, approximately five bases. These sharp reader heads detect DNA (A, T, G and C) readout is very good in mixed sequence regions, but DNA (e.g., p If there is a homopolymer region within the poly(polyT, poly(G), poly(A), poly(C), The signal becomes flat and lacks information. Five bases are missing from the signal of CsgG and its contraction mutants. To control the number of light-transmitting polymers, we discriminated between light-transmitting polymers longer than 5 without using additional dwell time information. However, when the DNA passes through the second reader head, many The DNA bases interact with the combined leader head and can be discriminated against homopolymers. The examples and figures use a CsgG pore and a pore containing a CsgF peptide. We show that this increase in homopolymer sequencing accuracy is achieved using
[0276] kit In a further aspect, the present invention also provides kits for characterizing target polynucleotides. The kit comprises an isolated pore complex according to the present invention and a membrane component or insulating layer. Preferably, the isolated pore complex is formed from constituent elements. which together form a transmembrane pore complex channel. It can include any type of membrane component, such as an amphiphilic layer or a triblock copolymer membrane. The kit may further comprise a polynucleotide binding protein. The kit may further comprise one or more anchors for attaching the protease to the membrane. Adding one or more other reagents or measuring devices that allow for carrying out any of the embodiments Such reagents or equipment may include appropriate buffers (aqueous solutions), obtaining means (such as a container or an instrument including a needle), polynucleotide or voltage or patch The reagents include one or more of the following: a means for amplifying and / or expressing the clamp device; The body sample may be present in the kit in a dry state to resuspend the reagents. Optionally, instructions that enable the kit to be used in the method of the invention, or how the method Finally, the kit may include peptide characterization. It may also include additional components useful for
[0277] In one embodiment, an isolated pore complex or transmembrane pore complex provided by the present invention is used for nucleic acid sequencing. For said use, Phi29 DNA polymerase ( DNAP) with the membrane-located CsgG:CsgF nanopore complex to form the pore The voltage may allow for controlled movement of the oligomer probe DNA strand through the pore. and applied across the nanopore is a current generated from the movement of ions in the salt solution on either side of the nanopore. As the probe DNA moves through the pore, the pore induces an ion flow relative to the DNA. This information has been shown to be sequence dependent, and the sequence of the probe is changed from the current measurement. This allows for accurate reading.
[0278] Certain embodiments, particular constructs and / or molecules are also useful in the engineered cells and Although methods have been described herein, they may be used without departing from the scope and spirit of the present invention. It should be understood that various changes or modifications in form and detail may be made. The following examples are provided to better illustrate certain embodiments, The invention should not be considered to limit the scope of the invention, which is limited only by the scope of the claims. do. [Example]
[0279] Introduction The CsgG pore is part of a multicomponent type VIII secretion system, which is involved in the Also known as the i biosynthetic system, it is expressed as curli in E. coli. Curli inhibits bacterial biofilm formation and the formation of aggregated filaments known as 'curli'. Curli are extracellular protein fibers primarily responsible for attachment to abiotic surfaces. , csgBAC and csgDEFG (curli-specific genes) in E. coli ( Hammar et al., 1995). Secretion of the urli subunits CsgA and CsgB occurs via the oligomeric secretion channels within the outer membrane. It depends on a specialized lipoprotein, CsgG, which is known to form a channel. For transport, CsgG binds periplasmic accessory proteins and extracellular accessory proteins. It acts in conjunction with the proteins CsgE and CsgG. CsgE mediates CsgG-mediated CsgF forms a specificity factor for CsgB-templated aggregation, while CsgB forms a specificity factor for CsgB transport. It appears to bind CsgA secretion to extracellular fibers.
[0280] The crystal structure of the CsgG secretion channel shows that CsgG traverses the OM and is a 36-stranded channel with an inner diameter of 40 Å. Forms a nonameric transport complex with a diameter of 120 Å and a height of 85 Å that spans the β-barrel (Goyal et al., 2014, Figure 1). The periplasmic domain of the channel, separated by an iris-like septum, is 24,000 Å 3 This membrane exists within each subunit and forms a large solvent-accessible cavity. The concentric assembly of CsgG oligomers is formed by a 12-residue “contraction loop” (CL) present in the The fabric forms an orifice with a diameter of about 0.6 nm and a height of about 1.5 nm excluding the solvent ( When acting as a protein secretion channel, this orifice within the CsgG channel The constriction or region forms the main site of interaction with the translocating polypeptide. When used as a pore sensing platform, this orifice is located within the channel. It serves as the primary reader for analytes present in or passing through the flow path. The diameter of the tube and its physical and chemical properties correspond to residues 46-61 of SEQ ID NO:3. The protein may be modified by amino acid substitution, deletion, or insertion in a region of the protein that 1D). In particular, the mutations at positions 51, 55, and 56 according to SEQ ID NO: 3 are taken together or independently, the conductance properties of the nanopore and its interaction with an analyte, including a polynucleotide. It has a beneficial effect on the interaction.
[0281] The assembly factor CsgF represents a component of the curli secretion apparatus. The CsgF preprotein It reaches the periplasm via the SEC pathway and then matures CsgF in a CsgG-dependent manner. (12.9 kDa) was found as a surface-exposed protein in the presence of CsgG. In this study, CsgF was fragmented in the OM, and co-immunoprecipitation experiments demonstrated that the two proteins were in direct contact with each other. Available data suggest that CsgF is involved in productive subunit secretion. Rather than being essential, proteins may regulate or By monitoring the coupling between CsgA secretion and extracellular polymerization, we investigated the mechanism of CsgA secretion and extracellular polymerization within curli fibers. This suggests that it will be formed.
[0282] [Example 1] Production of CsgG:CsgF complex protein (CsgG using CsgF synthetic peptide) Co-expression, in vitro reconstitution, coupled in-vitro transcription and translation, and Cs CsgG reconstitution with gF synthetic peptide) To produce the CsgG:CsgF complex, both proteins must be present in the appropriate Gram-negative host. They can be co-expressed in a host (e.g., E. coli) and extracted as a complex from the outer membrane and purified. The in vivo formation of the CsgG pore and the CsgG:CsgF complex is essential for the transport of CsgG to the outer membrane. Protein targeting is required, and for this, CsgG binds to the lipoprotein signal peptide. protein (Junker et al. 2003, Protein Sci. 12 (8): 1652-62) and Cy at the N-terminal position of the mature protein (SEQ ID NO: 3). These lipoproteins are expressed as preproproteins with a β- and β-residue. An example of a null peptide is residues 1 to 15 of the full-length E. coli CsgG shown in SEQ ID NO:2. Prepro-CsgG processing involves cleavage of the signal peptide and lipidation of mature CsgG. The mature lipoprotein then translocates to the outer membrane, where it forms oligomeric pores. (Goyal et al. 2014, Nature 516(753 0):250-3). To form the CsgG:CsgF complex, CsgF binds to Csg G, and expresses the native signal peptide corresponding to residues 1-19 of SEQ ID NO:5. The CsgG:CsgF leader sequence allows targeting to the periplasm. The combined pores are extracted from the outer membrane using detergent and homogenized by chromatography. The complex can be purified into a complex (Figure 2).
[0283] Alternatively, the CsgG:CsgF pore complex uses the CsgG pore and CsgF. These proteins can be produced by in vitro reconstitution as described below and in Figure 3.
[0284] For the example in vivo CsgG:CsgF complex formation shown in Figure 2, the native The signal peptide was used to encode E. coli CsgF (SEQ ID NO: 5) and CsgG (SEQ ID NO: 2) were co-expressed to allow for periplasmic targeting of both proteins, and CsgG Furthermore, to facilitate purification, CsgF was purified by adding a C-terminal 6x lipid. By introducing a histidine tag and fusing CsgG to a Strep-II tag at the C-terminus, Co-expression and complex purification were performed as described in the methods. SDS-PAGE analysis of the affinity purified eluate demonstrated enrichment of CsgF-His and Csg Co-purification of G-Strep was demonstrated, and the latter was found to be in a complex with CsgF. Furthermore, SDS-PAGE revealed the loss of the N-terminal fragment of the protein. Therefore, we found that a significant fraction of the eluted CsgF was present at low molecular weight ( (Figure 2B, asterisk). Pooled fractions of the His-trap eluate from the second affinity purification. SDS-PAGE analysis indicates the presence of CsgG and CsgF in apparent equimolar concentrates. This revealed the loss of the CsgF cleavage fragment seen in the His-trap eluate and in the His-trap eluate (Fig. 2B The co-elution of CsgF in Strep-affinity purification indicates that the protein is not co-eluted with CsgG. Notably, the N-terminal truncated fragment of CsgF was Fragment is lost during Strep-affinity purification, and the CsgF N-terminus is required to bind CsgG It has been suggested that there is a
[0285] Figure 13 shows another example of CsgG:CsgF complex formation by in vivo co-expression. In this example, the CsgGA protein was modified with a C-terminal Strep-II tag, and CsgG The full-length protein is modified with a C-terminal 10X histidine tag. The complexes were strained as shown in the Materials and Methods section for analyte characterization. The component is isolated from its constituent components by top-tag purification followed by histidine-tag purification. The CsgG:CsgF complex was separated from its constituent components and purified. , which can be clearly distinguished from the CsgG pore in SDS-PAGE analysis (Figure 13 As shown in Figure 13B, two tag purification methods were successfully applied to isolate CsgG:Cs The gF complex can be separated from its component parts and purified.
[0286] To prepare the CsgG:CsgF complex by in vitro reconstitution, and CsgF were transformed with pPG1 and pNA101, respectively. After expression in E. coli culture and purification, the CsgG:CsgF complex was isolated in vivo. For comparison, purified CsgG was reconstituted in vitro (see Methods). The same procedure was followed on a Superose 6 column. CsgG Superose The six runs were performed on the nonameric CsgG pore (Figure 3A(a) and 3C) and the nonameric CsgG pore. We revealed the presence of two discrete populations corresponding to dimers (Figures 3A(b) and 3C). (As previously described in Goyal et al. (2014). CsgG:Csg Superose 6 runs of F reconstitution consisted of excess CsgF (Figure 3A(c)), nonamer C The sgG:CsgF complex ( Figure 3A(d) ) and the nonameric CsgG:CsgF dimer ( This revealed the presence of three distinct groups, corresponding to the CsgG:CsgF complex (Figure 3A(e)). To independently confirm the formation of nucleosomes, various Superose 6 elution peaks were analyzed. and analyzed by native PAGE (Figure 3B).
[0287] Surprisingly, the CsgG:CsgF complex was identified in the Materials and Methods section for analyte characterization. or by the coupled in vitro transcription and translation (IVTT) method described in Section The complex can be generated by combining the CsgG protein and CsgF in the same IVTT reaction. Expressing proteins CsgG and CsgF in two different IVTT reactions In the embodiment shown in FIG. used the E. coli T7-S30 extraction system (Promega) for circulating DNA. The CsgG:CsgF complex was generated in one reaction mixture using the SD Protein expression in IVTT was analyzed by S-PAGE. Therefore, the DNA used to express proteins in IVTT is , lacking the DNA encoding the signal peptide region. When expressed in an IVTT in the absence of Moreover, these expressed monomers were also expressed using cell extract membranes present in the IVTT reaction mixture. CsgG oligomers can be assembled into pores in situ by using (Figure 14, lane 1). CsgG oligomers are SDS-stable, but the sample is 100°C. When heated to 100°C, it decomposes into its component monomers (Figure 14, lane 2). When A was expressed in an IVTT in the absence of CsgG DNA, only CsgF monomers were expressed. The DNA of CsgG and CsgF was 1:1. When mixed in the same ratio and co-expressed in the same IVTT reaction mixture, CsgF proteins Assembled CsgG pore with high efficiency of making CsgG:CsgF complex (Figure 14, lane 5). This SDS-stable complex produced by IVTT interacts with It is thermostable up to at least 70°C (Figure 14, lanes 6 to 12).
[0288] The CsgG:CsgF complex with truncated CsgF also contains the truncated CsgF rather than the full-length version. By using DNA encoding sgF, the vector can be produced by any of the methods described above. However, the stability of the complex is affected by the cleavage of CsgF below the FCP domain. Furthermore, the CsgG:CsgF complex with the truncated CsgF can be After the full-length CsgG:CsgF complex is formed, full-length CsgF is cleaved at the appropriate site Cleavage can be achieved by inserting a protease cleavage site at the desired position. By modifying the DNA encoding the CsgF protein by incorporating SEQ ID NOs: 56 to 67 are CsgF clones with truncated CsgF. TEV or CsgF integrated at various positions to generate G:CsgF complexes indicates the HCV C3 protease site. CsgG:CsgF created with SEQ ID NO:61 SDS-PAGE analysis of TEV cleavage of the complex is shown (Figure 15B). The combined clone (with full-length CsgF) was used in the Materials and Methods section for analyte characterization. When treated with the TEV protease enzyme described in (Fig. 15B, lanes 3 and 4). However, TEV cleavage results in excess C-terminal This leaves six amino acids, thus the remaining CsgF cleavage site in the CsgG pore-containing complex. The protein is 42 amino acids long. The molecular weight difference is still visible on SDS-PAGE (FIG. 15B, lanes 7 and 8).
[0289] Surprisingly, the CsgG:CsgF complex with truncated CsgF also expressed purified CsgG. Pores (produced in vivo or in vitro) are then ligated with synthetic peptides of appropriate length. Reconstitution can be performed in vitro. Therefore, the CsgF signal peptide is required to form the CsgG:CsgF complex. Furthermore, this method does not leave excess amino acids at the C-terminus of CsgF. Mutations and modifications can also be easily incorporated into synthetic CsgF peptides. The method uses different CsgF peptides or mutants or homologs thereof to differentiate different CsgFs. Reconstitution of the sgG pore or mutants or their homologs to differentiate CsgG:CsgF This is a very convenient method for generating complex mutants. The stability of the complex can be determined by the presence of CsgF. Cleavage beyond the FCP domain can impair the CsgG:CsgF complex. Examples of truncated CsgF and FCP peptides used to generate By this method, CsgF-(1-45) (Fig. 16.A) and CsgF-(1-35) (Figure 16.B) and CsgF-(1-30) (Figure 16.C). SDS-PAGE analysis of the thermostability of the G:CsgF complex revealed that at least CsgF-(1-4 5) and CsgF-(1-35) peptide is thermostable up to at least 90°C. The sgG pore was then cooled to 90°C and its components were then formed into a complex with the sgG. It is difficult to assess the stability of the complex above 90°C because it decomposes down to the monomer. The CsgG pore band in SDS-PAGE and the CsgG:CsgF-(1–30) complex Since the difference between the CsgG:CsgF bands is small, this method is suitable for CsgG:CsgF-(1-30 ) complex (Figure 16C). The CsgG:Cs complex was observed in all three cases, and the CsgG:Cs complex was observed in the electrophysiological experiments. Even CsgF-(1–29) peptides were observed, and at least some CsgF-(1–29) peptides were observed. These results demonstrate that the CsgG:CsgF complex is produced by the ATPase inhibitors (Figure 24).
[0290] [Example 2] Cryo-EM analysis of CsgG:CsgF structure To obtain structural insight into the CsgG:CsgF complex, co-purified or in vivo The reconstituted CsgG:CsgF particles were analyzed by transmission electron microscopy. Peak fractions of double affinity purified CsgG:CsgF complexes in preparation for yo-EM analysis 500 μL of the solution was added to Buffer D (25 mM Tris Ph8, 200 mM NaCl, and A Superose 6 10 / 30 column (GE) equilibrated with 0.03% DDM was used. The infusion was performed on a 1000 sieve (1000 sieve) at 0.5 mL / min. The protein concentration was Determined based on calculated absorbance at 280 nm and assuming a 1:1 stoichiometry. Samples for chromophores were analyzed as described in Methods. CsgG:CsgF complex Cryo-EM micrographs of selected CsgG:CsgF particles, as well as selected The two class averages are shown in Figure 4. Micrographs show the two classes of the nonameric pore and the nonameric pore complex. For image reconstruction, select the nonameric CsgG:CsgF particle and Alignment was performed using LION. Class average of the CsgG:CsgF complex as a side view. The 3D reconstructed electron density shows that the CsgG particle is located on the side of the CsgG β-barrel. These show the presence of additional density corresponding to CsgF, seen as protrusions from the nuclei (Fig. 4B, 5). The additional density is due to the globular head domain, the hollow neck domain and the CsgG β-barrel. The CsgF yield reveals three distinct regions that encompass the domains that interact with the CsgF protein. The latter CsgF region, termed the condensed peptide or FCP, is located within the lumen of the CsgG β barrel. The constriction region formed by the CsgG constriction loop inserted into the ribosome (labeled G in Figures 4B and 5) an additional constriction of the CsgG pore (labeled F in Figures 4B and 5), located approximately 2 nm above the It is seen as forming.
[0291] [Example 3] Identification of CsgF-interacting and contractile peptides by cleavage of CsgF The presence of a second constriction in the CsgG:CsgF pore complex is comparable to the CsgG-only pore. In comparison, it offers opportunities for nanopore sensing applications, as a second reader head, or as a Csg The nanofibers can be used as an extension of the main reader head provided by the G contraction loop. It provides a second orifice within the pore (Figs. 6 and 7). When this occurs, the exit side of the CsgG:CsgF combination pore is occupied by the CsgF neck domain and head domain. Therefore, the interaction with the CsgG β-barrel and We aimed to determine the CsgF region required for insertion into Strep-Tactin. In affinity purification experiments, the N-terminal truncated fragment of CsgF present in the His-trap affinity purification The N-terminal region of CsgF was lost and did not co-purify with CsgG, suggesting that the N-terminal region of CsgF is in the CsgG phase. This provided clues that the CsgF homolog is required for the interaction (Fig. 2B). It is characterized by the presence of the FAM domain PF03783. It is found in Gram-negative bacteria. When performing multiple sequence alignment (MSA) of selected CsgG homologues (Figure 8), MSAs of sgF homologs are shown), sequence conserved regions (MSAs of 35–100% pairwise sequence identity). A) corresponds to approximately the first 30-35 amino acids of mature CsgF (SEQ ID NO: 6). Based on the combined data, this N-terminal region of CsgF is related to the CsgG phase. It was hypothesized that they form an interacting peptide or FCP. A multiple sequence alignment of FCPs in is shown in Figure 10.
[0292] The CsgF N-terminus corresponds to the CsgG-binding domain and is located within the lumen of the CsgG β-barrel. To test the hypothesis that CsgF forms a contractile peptide, we used Strep-tagged C sgG and a His-tagged CsgF truncation were co-overexpressed in E. coli ( pNA97, pNA98, pNA99, and pNA100 contain CsgF ( N-terminal Cs corresponding to residues 1-27, 1-38, 1-48, and 1-64 of SEQ ID NO: 5 These peptides encode the Cs gF fragment, which corresponds to residues 1-19 of SEQ ID NO:5. gF signal peptide, and thus the first 8 amino acids of mature CsgF (SEQ ID NO: 6, Figure 9A). , generating periplasmic peptides corresponding to 19, 29, and 45 residues, respectively. Containing a terminal 6x His tag. SDS-PAGE analysis of whole cell lysates revealed that all CsgG in each sample, as well as the first 45 residues of mature CsgF (SEQ ID NO: 6, Figure 9B). The presence of the corresponding CsgF fragment was revealed. No detectable expression of the peptide was observed in the whole cell lysate. After cloning, the cellular mass of the various CsgG:CsgF fragments was further enriched by purification. Whole cell lysates and elution fractions of Strep affinity purification were purified using His-tagged CsgF fragments. on nitrocellulose membranes for dot blot analysis using anti-His antibodies for detection of The dot blot analysis showed that the CsgF20:64 peptide was spotted on the C The CsgG fragment co-purifies with CsgG and forms a stable, non-covalent complex with CsgG. For the CsgG 20:48 fragment, a small amount of the peptide is sufficient to It was found to co-purify with CsgG, whereas CsgF 20:27 or CsgF 20: 38, either for whole cell lysates or Strep affinity purification (Fig. 9C). However, no detectable levels were observed for the latter peptide, which was stably expressed in E. coli. These findings suggest that the ATP-dependent ATPases do not bind to CsgG and / or do not form stable complexes with CsgG.
[0293] [Example 4] Explaining the CsgG:CsgF interaction at atomic resolution To obtain atomic details of the CsgG:CsgF interaction, To this end, we determined the high-resolution cryo-EM structure of the CsgG and gF complexes. CsgF was co-expressed in E. coli, and the CsgG:CsgF complex was isolated by detergent extraction. It was isolated from the E. coli outer membrane and purified using tandem affinity purification. Samples for microscopy were prepared on graphene oxide-coated R2 / 1 Holey grids (Quan The data were prepared by spotting 3 μl samples onto a sieve (tifoil) and analyzed using Gata n 300Kv TITAN Krios using K2 direct electron detector in counting mode 62,000 single CsgG:CsgF particles were used for the final electron density map. The map was calculated at 3.4 Å resolution (Figure 11A). The map provides a clear view of the CsgG crystal structure. The de novo assembly of the N-terminal 35 residues of mature CsgF ( i.e., residues 20:54 of SEQ ID NO: 5), which binds CsgG and It encompasses the FCP, which forms a second constriction at the level of the transmembrane β-barrel (Fig. 11C, D). The cryoEM structure revealed that CsgG:CsgF contains a 9:9 stoichiometry with C9 symmetry. The FCP is located inside the CsgG β-barrel and the C-terminus of CsgF is C The CsgF N-terminus is positioned near the CsgG constrictor. The structure shows that P35 of mature CsgF is located outside the CsgG β-barrel and in the C The CsgG:CsgF complex forms a connection between the CsgF FCP and the neck region. Due to the flexibility of the coalescent body, the CsgF neck and head regions were observed in high-resolution c Three regions within the CsgG β-barrel are not resolved in the ryoEM map. Stabilizes CsgF interactions: i.e., (IR1) of mature CsgG (SEQ ID NO: 3) Residues Y130, D155, S183, N209 and T207 form four H-bonds and a static with the N-terminal amine and residues 1-4 of mature CsgF (SEQ ID NO: 6), including electrical interactions and the (IR2) residue Q18 of mature CsgG (SEQ ID NO: 3) 7, D149 and E203 contain three H-bonds and two electrostatic interactions. It forms an interaction network with R8 and N9 of mature CsgF (SEQ ID NO: 6) and interacts with mature CsgF. (IR3) residues F144, F191, F193 and L199 of sgG (SEQ ID NO: 3) , hydrophobic interaction table with residues F21, L22 and A26 of mature CsgF (SEQ ID NO: 6) The latter is an α-helix formed by residues 19–30 of mature CsgF ( The conserved sequence NPXFGG (SEQ ID NO: 6) is located within helix 1. Residues 9–14) are involved in the loop formed by residues 15–19 in CsgF helix 1. The ATP-binding domain forms an inward turn connecting the CsgF helix 1. These elements generate a constriction within the CsgG:CsgF complex, of which residue 17 ( N17 (SEQ ID NO: 6) in mature E. coli CsgF forms the narrowest point, This results in an orifice with a diameter of 15 Å (FIG. 11C). The constriction formed by residues 46–59 of CsgG (G-constriction or GC) They are located approximately 15 to 30 Å above the upper and lower portions, respectively.
[0294] [Example 5] Simulations to improve the stability of the CcGg-CsgF complex Molecular dynamics simulations were performed to establish which residues in CsgG and CsgG are in close proximity. This information can be used to identify CsgG complexes that can increase their stability. and CsgF mutants were designed.
[0295] Using GROMAS package version 4.6.5, GROMOS 53a6 forced flash Simulations were performed using the field and SPC water models. The cryo-EM structure of the sgF complex was used in the simulation. The energy was minimized using the steepest descent algorithm from the Throughout, restraints were applied to the backbone of the complex, but residue side chains were free to move. The system is equipped with Berendsen thermostats and Berendsen barostats. The results were simulated in the NPT ensemble for 20 ns using a temperature of 300 K.
[0296] Contacts between CsgG and CsgF were described using GROMACS analysis software and in situ. The two residues were within 3 Å of each other. Contact was defined as occurring when the patient was in contact with the skin. The results are shown in Table 4 below.
[0297] [Table 4] JPEG2025148349000006.jpg20984JPEG2025148349000007.jpg15581
[0298] Materials and methods for structural determination of the CsgG:CsgF complex:
[0299] Cloning For expression of E. coli CsgG as an outer membrane-localized pore, E. coli Cs The coding sequence of gG (SEQ ID NO: 1) was cloned into pASK-Iba12, and the plasmid p PG1 was obtained (Goyal et al. 2013).
[0300] For expression of C-terminally 6x-His-tagged CsgF in the E. coli cytoplasm, The codon for mature E. coli CsgF (SEQ ID NO: 6, i.e., CsgF without the signal sequence) The code sequence was determined using primers "CsgF-His_pET22b_FW" (SEQ ID NO: 46) and and "CsgF-His_pET22b_Rev" (SEQ ID NO: 47). Use the PCR product and clone it into pET22b via the NdeI and EcoRI sites. The CsgF-His expression plasmid pNA101 was obtained.
[0301] pTrc99a-based vector expressing csgF-His and CsgG-strep The pNA62 plasmid, which is a vector, was cloned from pGV5403 (pDEST14 Gateway ( It was constructed based on pTrc99a) into which the registered trademark cassette was integrated. The ampicillin resistance cassette is a streptomycin / spectinomycin resistance cassette. The coding sequences for csgE, csgF, and CsgG were replaced with E. A PCR fragment encompassing a portion of the coli MC4100 csgDEFG operon was Timer csgEFG_pDONR221_FW (SEQ ID NO: 48) and csgEFG_p DONR221_Rev (SEQ ID NO: 49) was used to generate pDONR221 (The BP Gateway® integrated into rmoFisher Scientific This recombinant vector was then inserted via recombination from the pDONR221 donor plasmid. The recombinant csgEFG operon was transformed into Insertion into pGV5403 with a streptomycin / spectinomycin resistance cassette The primer Mut_csgF_His_FW (SEQ ID NO: 50) was used via PCR. and Mut_csgF_His_Rev (SEQ ID NO: 51) to obtain a 6xHis tag. was added to the C-terminus of CsgF. Finally, csgE was amplified by outward PCR (primer De lCsgE_FW (SEQ ID NO: 52) and DelCsgE_Rev (SEQ ID NO: 53) This was removed to obtain pNA62.
[0302] Periplasmic reticulum of a C-terminal His-tagged CsgG fragment corresponding to the putative contractile peptide (Figure 9A). The construct for expression of the CsgF-his and CsgG-strep gene was Generated by outward PCR against the Trc99a-based vector pNA62 The primer combinations were as follows: pNa6 as the forward primer; 2_CsgF_histag_Fw (SEQ ID NO: 45) and CsgF_d27_end (SEQ ID NO: Sequence number 41), CsgF_d38_end (sequence number 42), CsgF_d48_end (SEQ ID NO: 43) or CsgF_d64_end (SEQ ID NO: 44) as the reverse primer pNA97, pNA98, pNA99, and pNA100 were created, respectively. Ta.
[0303] In pNA97, csgF is cleaved to SEQ ID NO:7, resulting in residues 1 to 27 (SEQ ID NO:8). In pNA98, csgF is cleaved to SEQ ID NO: 9, and the remaining pNA99 encodes a CsgF fragment containing residues 1 to 38 (SEQ ID NO: 10), and in pNA99, csgF is cleaved to SEQ ID NO: 11 to encode a CsgF fragment containing residues 1 to 48 (SEQ ID NO: 12). In pNA100, csgF is cleaved to SEQ ID NO: 13, and residues 1 to 64 (SEQ ID NO: 13) are inserted. pNA97, pNA Expression of pNA98, pNA99 and pNA100 increased the CsgG pore (SEQ ID NO: 3) in the outer membrane. ) and periplasmic targeting of a CsgF-derived peptide having the sequence :
[0304] "GTMTFQFRHHHHHH" (SEQ ID NO: 37 + 6xHis), "GT MTFQFRNPNFGGNPNNGHHHHHH (SEQ ID NO: 38 + 6xHis) , "GTMTFQFRNPNFGGNPNGAFLLNSAQAQHHHHHH" (distribution Column number 39+ 6xHis), and "GTMTFQFRNPNFGGNPNNGAFL LNSAQAQNSYKDPSYNDDFGIETHHHHHH" (SEQ ID NO: 40+ 6 xHis).
[0305] KK E. coli Top 10(F - mcrA Δ( mrr - hsdRMS - mcrB C) Φ80lacZΔM15 Δ lacX74 recA1 araD139 Δ( araleu) 7697 galU galK rpsL (StrR) endA1 nupG) was used in all cloning steps. 3)(F - ompT hsdSB(rB - mB - )gal dcm(DE3)) and and Top10 were used for protein production.
[0306] Production of recombinant CsgG:CsgF complex by co-expression Co-expression of E. coli CsgF (SEQ ID NO: 5) and CsgG (SEQ ID NO: 2) Both recombinant genes contained their native Shine Dalgarno sequences. The gene is placed under the control of the inducible trc promoter in a pTrc99a-derived plasmid. CsgG and CsgF were expressed in plasmid pNA62. Transformed E. coli grown in Terrific Broth at 37°C The cells were overexpressed in C43(DE3) cells. The cell cultures had an optical density of 0.7 at 600 nm. When the medium reached a density (OD), recombinant protein expression was induced with 0.5 mM IPTG. After incubation at 28°C for 15 hours, the cells were harvested by centrifugation at 5500 g.
[0307] Production of recombinant CsgG:CsgF complex by in vitro reconstitution Full-length E. coli CsgG modified with a C-terminal Strep II-tag (SEQ ID NO: 2) was transformed with the plasmid PpG1 (Goyal et al. 2013) E . The cells were overexpressed in E. coli BL21(DE3) cells at 37°C. The strain was cultured in filtrate broth until the OD at 600 nm reached 0.6. The production of the α-type protein was induced by 0.0002% anhydrotetracycline (Sigma). The cells were incubated for a further 16 hours at 25°C and then harvested by centrifugation at 5500 g. Ta.
[0308] E. coli CsgF (SEQ ID NO: 6) in a C-terminal fusion with a 6x His tag (i.e., lacking the CsgF signal sequence) was transformed with plasmid pNA101 into E. c oli was overexpressed in the cytoplasm of BL21(DE3) cells. The cells were incubated at 37°C for 600 min. The cells were grown to an OD of 100 nm, then induced with 1 mM IPTG and incubated at 37°C for 15 hours. The protein was expressed using the enzyme and then harvested by centrifugation at 5500 g.
[0309] Recombinant protein purification of CsgG:CsgF complex, CsgG, and CsgF Transformed with pNA62 and co-expressing CsgG-Strep and CsgF-His E. coli cells were incubated in 50 mM Tris-HCl pH 8.0, 200 mM NaCl, 1mM EDTA, 5mM MgCl2, 0.4mM AEBSF, 1 μg / mL leupeptin, 0.5 mg / mL DNase I and 0.1 mg / m The cells were resuspended in L-lysozyme. Systems Ltd.) at 20 kPsi, and the lysed cell suspension was % n-dodecyl-β-d-maltopyranoside (DDM, Inalco) for 30 min The remaining cell debris and the The membranes were centrifuged at 100,000 g for 40 minutes in an ultracentrifuge. A (25 mM Tris pH 8, 200 mM NaCl, 10 mM imidazole, 10 5 mL HisTrap cap equilibrated in 0.1% sucrose and 0.06% DDM The column was loaded onto the ram. The column was then filled with >10 CV of 5% Buffer B (25 mM Tris pH 8, 200 mM NaCl, 500 mM imidazole, 10% sucrose, 0.06% DD M) Wash with ionic buffer A and elute with a gradient of 5-100% buffer B over 60 mL. did.
[0310] The eluate was diluted two-fold and purified with buffer C (25 mM Tris pH 8, 200 mM NaCl) , 10% sucrose, and 0.06% DDM) - Loaded onto a Tactin column (IBA GmbH) overnight. The column was loaded with >10 CV buffer. After washing with solution C, the protein was eluted by adding 2.5 mM desthiobiotin. 500 μL of the peak fraction of the two-fold affinity purified complex was added to buffer D (25 mM Tris Supe equilibrated with 1000 mM NaCl (pH 8, 200 mM NaCl, and 0.03% DDM) Rose 6 10 / 30 (GE Healthcare) for electron microscopy. Sample preparation was performed at 0.5 mL / min. Protein concentrations were determined at 280 nm. The absorbance of Buffer D (25 mM) was calculated based on the absorbance of 1 / 1 stoichiometry. Tris pH8, 200mM Nacl, 0.03% DDM)
[0311] CsgG-strep purification for in vitro reconstitution excludes sucrose in the buffer However, when bypassing the IMAC and size exclusion steps, the protocol for CsgG:CsgF is are identical.
[0312] CsgF-His purification for in vitro reconstitution was performed in 50 mM Tris-Hc l pH 8.0, 200mM Nacl, 1mM EDTA, 5mM MgCl2 , 0.4 mM AEBSF, 1 μg / mL leupeptin, 0.5 mg / mL DNA se I, performed by resuspending the cell mass in 0.1 mg / mL lysozyme The cells were then crushed in a TS series cell crusher (Constant Systems Ltd.). The cells were crushed at 20 kPsi using a centrifuge at 10,000 g for 30 minutes to obtain intact cells. The supernatant was collected in buffer A (25 mM Tris pH 8, 200 5 mL Ni-IMAC beads equilibrated with 10 mM NaCl, 10 mM imidazole (Workbis 40 IDA, Bio-Works Technologies A B) and incubated at 4°C for 1 hour. Ni-NTA beads were placed in a gravity flow column. Pool and add 100 mL of 5% Buffer B (25 mM Tris pH 8, 200 mM The resulting mixture was washed with 500 mM NaCl, 500 mM imidazole diluted in buffer A. The proteins were eluted by stepwise increases of buffer B (10% step in each 5 mL). Pu).
[0313] In vitro reconstitution of the CsgG:CsgF complex Purified CsgG and CsgG are pooled and the complex is reconstituted in vitro. Therefore, a molar ratio of 1 CsgG:2 CsgF was mixed to obtain the CsgG barley. The solution was saturated with CsgF. The reconstitution mixture was then added to buffer D (25 mM Tris p Superos equilibrated in HCl (H8, 200 mM NaCl, and 0.03% DDM) e 6 10 / 30 column (GE Healthcare) and analyzed for electron microscopy. The sample was prepared at 0.5 mL / min (Figure 3). The protein concentration was 280 Determined based on calculated absorbance at 100 nm and assuming 1 / 1 stoichiometry.
[0314] Structural analysis using an electron microscope The sample behavior of size exclusion is probed using negative stain electron microscopy. The samples were stained with 1% uranyl formate and measured using an in-house 120 kJ microscope equipped with LaB6 filaments. Imaged using a JEM 1400 (JEOL) microscope. Samples for electronic chromosome testing. 2 μL of sample was applied to an R2 / 1 continuous carbon (2 nm) coated grid (Quantifoil ), manually blotted, and plunged into liquid ethane using an in-house plunge device. The quality of the samples was confirmed by scanning them in-house on a JEOL JEM 1400. After cleaning, a 200 kV Falcon-3 direct electron detection camera was used. The dataset was collected on a TALOS ARCTICA (FEI) microscope. Images are from Mot Motion correction was performed using ionCor2.1 (Zheng et al. 2017). Focus values were determined using ctffind4 (Rohou and Grigo rieff, 2015), data are from RELION (Scheres, 2012) and Further analysis was performed using a combination of EMAN2 and EMAN2 (Ludtke, 2016). C9 symmetry is a selected 2D class average characterized by additional density for the head group. This was imposed during 3D model generation and refinement.
[0315] For high-resolution cryoEM analysis, the CsgG:CsgF sample was prepared using graphene oxide (S R2 / 1 Holey grids (Quanti Aldrich) coated with 3 μl of sample was spotted onto foil, blotted manually, and then CP3 plunger ( prepared for cryo-electron microscopy by plunging into liquid ethane using a Gatan After the sample quality was screened in-house using a JEOL JEM 1400, 300kV TITAN KRIO equipped with Summit electron detector (Gatan) Data sets were collected on a S (FEI, Thermo-Scientific) microscope. The detector measured 56 electrons / Å over 50 frames. 2 Counting module with cumulative electron dose of The image was taken at 1.07 Å pixel size. MotionCor2.1 was used to correct the motion (Zheng et al. 2017). Focus values were determined using ctffind4 (Rohou and Grigo rieff, 2015). The particle is Gautomatch (Dr. Kai Zhan g), and data were automatically captured using RELION 2.0 (Kimanius e t al. 2016, Elife 5. pii: e18722) and EMAN 2 (Ludtke, 2016). C9 symmetry is 3D model generation for selected 2D class means featuring additional density for and refinement. 62,000 particles were used, resulting in a final matrix at 3.4 Å resolution. De novo model building of CsgF was performed using COOT (Brown et al. t al. 2015 Acta Crystallogr D Biol Cryst allogr 71(Pt 1):136-53) and the entire complex The iterative cycle of model building and refinement was performed using PHENIX (Afonine 2018, Acta Crystallogr D Struct Biol 74(Pt 6) :531-544) Real-space refinement was performed using COOT in combination.
[0316] Protein expression and purification of the CsgG:CsgF fragment The CsgF fragment and CsgG were co-expressed, and the CsgF fragment was His-tagged at the C-terminus. , CsgG was fused to a Strep tag at the C-terminus. The CsgG:CsgF fragment complex was E transformed with plasmids pNA97, pNA98, pNA99, or pNA100 The plasmid was overexpressed in E. coli Top10 cells. The plates were incubated at 37°C. Colonies were resuspended in LB medium supplemented with streptomycin / spectomycin. When the cell culture reached an optical density (OD) of 0.7 at 600 nm, the recombinant protein was Expression of the protein was induced with 0.5 mM IPTG, and the mixture was incubated at 28°C for 15 hours. The cells were harvested by centrifugation at 100 xg. The pellet was frozen at -20°C.
[0317] Cell mass of various CsgG:CsgF fragments was diluted in 200 mL of 50 mM Tris-HCl. Cl pH 8.0, 200mM NaCl, 1mM EDTA, 5mM MgCl2 , 0.4 mM AEBSF, 1 μg / mL leupeptin, 0.5 mg / mL DN The cells were resuspended in 0.1 mg / mL lysozyme and sonicated. 1% n-dodecyl-β-d-maltopyranoside was used for cell lysis and extraction of outer membrane components. The remaining cell debris and membranes were then removed by centrifugation. The mixture was centrifuged at 15,000 g for 40 minutes. The supernatant was transferred to 100 μL of Streptomyces tube. The beads were incubated with the Strep beads at room temperature for 30 minutes. Tris pH 8, 200 mM NaCl, and 1% DDM) by centrifugation. Wash with 25 mM Tris pH 8, 200 mM NaCl, 0.01% DD Bound proteins were eluted by adding 2.5 mM desthiobiotin to M.
[0318] Production of CsgG:FCP by in vitro reconstitution A synthetic peptide corresponding to the N-terminal 34 residues of mature CsgF (SEQ ID NO: 6) was diluted in buffer 0.1 M MES, 0.5 M NaCl, 0.4 mg / mL EDC (1-ethyl-3-(3- dimethylaminopropyl)carbodiimide), 0.6 mg / ml NHS (N-hydroxybenzoate) Dilute to 1 mg / ml in HCl (hydroxyaminopropyl carbodiimide) and incubate at room temperature for 15 minutes. The peptide carboxy terminus was activated by incubation with 1 mg / ml Z of cadaverine Alexa594 in PBS was added and allowed to covalently bind at room temperature. Use a eba Spin filter to filter the buffer into 50 mM Tris, NaCl, 1 The reaction was quenched by replacing with 0.5 mM EDTA, 0.1% DDM.
[0319] The labeled peptides were diluted in 50 mM Tris, 100 mM NaCl, 1 mM Strep affinity purified CsgG in 5 mM LDAO / C8D4, EDTA, at room temperature for 1 CsgG:FCP complex was reconstituted by adding Str at a molar ratio of 2:1 for 5 min. After pulling down CsgG-strep on epTactin beads, the samples were The fragments were analyzed by reactive PAGE.
[0320] [Example 6] Further stabilization of the CsgG:CsgF complex by covalent cross-linking Full-length CsgF and several truncated versions of CsgF bind to the CsgG pore and stable C Although the sgG:CsgF complex is formed, under certain conditions, CsgF still binds to the CsgG cell membrane. Therefore, the CsgG and CsgF subunits can be separated from the barrel region of the pore. Based on molecular simulation studies, it is desirable to form covalent bonds between the , the positions of CsgG and CsgF in close proximity to each other were identified (Example 5 and Table 4 Some of these identified positions contain cysteines in both CsgG and CsgF. Figure 19 shows the Q153 position of CsgG and the G1 position of CsgF. An example of thiol-thiol bond formation between Csg containing the Q153C mutation is shown. The G pore was reconstituted with CsgF containing the G1C mutation and incubated for 1 hour to allow for S-S bond formation. When the complex was heated to 100°C in the absence of DTT, the CsgG monomer 45 kDa corresponding to a dimer between CsgGm and CsgF monomers (CsgGm-CsgFm) Bands can be seen, which consist of two monomers (CsgGm at 30 kDa and CsgF This band indicates the formation of an SS bond between the α- and β-amyloids (15 kDa) (Fig. 19A). It disappears when heated in the presence of DTT. DTT breaks the SS bond. CsgG:CsgF The extent of CsgGm-CsgFm dimer formation was reduced when the complex was incubated overnight instead of for 1 h. To further distinguish the dimer band, mass spectrometry was performed. Gel-purified proteins were proteolytically cleaved to generate tryptic peptides. LC-MS / MS sequencing revealed that the Q153 position of CsgG and the G1 position of CsgF The SS bond between the copper and copper sites was identified (Figure 19B). -O-phenanthroline and other oxidizing agents can be used. The CsgG pore was then cleaved in the presence of copper-orthophenanthroline as described in the Methods section. Reconstituted with CsgF containing the T4C modification and then heated to 100°C in the absence of DTT. When decomposed into its component monomers by A distinct dimer band can be observed on SDS-PAGE (Figure 20, lanes 3 and 4). When heating is performed in the presence of DTT, the dimer decomposes into its component monomers (Figure 2). 0, lane 1 and lane 2).
[0321] [Example 7] Electrophysiological characterization of the CsgG:CsgF complex The signal observed when the DNA strand translocated through CsgG was due to the insertion of a pore into the copolymer membrane. The experiment was conducted at Oxford Nanopore Technologies' Micro Each subunit of CsgG was fully characterized when performed using niON (Figure 31). Units Y51, N55, and F56 form the constriction of the CsgG pore (Figure 12). This sharp constriction serves as the leader head of the CsgG pore (Figure 31A). , capable of accurately discriminating between mixed sequences of A, C, G, and T when passing through the pore. This means that the measured signal produces a characteristic current deflection from which the identity of the sequence can be derived. However, in homopolymer regions of DNA, the measured signal is may not exhibit a current deflection large enough to allow single-base discrimination, resulting in Therefore, an accurate determination of the homopolymer length cannot be made from the measured signal magnitude alone. The accuracy of the CsgG leader head was reduced by the homopolymer This correlates with the length of the region (Figure 29C). CsgF interacts with the CsgG pore to form CsgG: When forming the CsgF complex, CsgF inserts a second leader head into the CsgG barrel. This second leader head is mainly composed of the N17 position of SEQ ID NO:6. The static chain experiments described in the "" section and Figure 27 show that the two CsgG:CsgF complexes The results showed that the reader heads of the nucleotide sequences were aligned at approximately 5-6 base pairs. The presence of two leader heads separated from each other is shown (Fig. 27, B, C, and D). The leader-head discrimination plot of the G:CsgF complex shows that the second The contribution of the CsgG leader head to base discrimination is Surprisingly, the second leader head is smaller than the CsgG barcode (Figure 27A). When introduced by CsgF in the nucleus, the previously flat homopolymer region becomes stepped. These steps show the signal structure shown in Figure 30B and C. Contains usable information and reduces errors. DNA of the CsgG:CsgF complex The signal quality was higher for longer homopolymers compared to the quality profile of the CsgG pore itself. It remains relatively constant throughout the ribbon length (Figure 29C).
[0322] CsgG:CsgF complexes produced by either of the methods described in the Methods section , which can be used to characterize the complex in DNA sequencing experiments. Different methods consisting of G mutant pores and different CsgF peptides with different lengths The signals of lambda DNA strands passing through various CsgG:CsgF complexes prepared in Fig. 2 1-24. The pore complexes and their base contribution profiles are shown in the reader's The pore and peptide CsgF were also shown to be highly cleavable. Different modifications at both constrictions significantly alter the signal of the CsgG:CsgF pore complex. For example, even if the CsgG:CsgF complex is produced in the same CsgG pore, the CsgG:CsgF complex at position 17 Two different CsgFs of the same length containing either Asn or Ser in (SEQ ID NO: 6) When peptides are used (after the same co-expression method of the full-length CsgF protein, positions 35 and 36) is generated by TEV protease cleavage of CsgF, and the signal generated is CsgG:Cs, which has Seg at position 17 of the CsgF peptide. The CsgG:CsgF complex has an Asn at position 17 of the CsgF peptide. It shows lower noise and higher signal-to-noise ratio compared to the same Csg The G pore was reconstituted with two different peptides of the same length (1-35 of SEQ ID NO: 6), and Cs When Val or Ser at position 17 was used to generate the gG:CsgF complex Val at position 17 of CsgF is noisier than the complex with Ser at position 17 of CsgF. The same CsgF peptide of the same length shows a significant signal after the CsgG leader header (Figure 22). The pores were reconstituted with different CsgG pores containing different mutations in the pore domain (positions 51, 55, and 56). When CsgG:CsgF complexes were detected, the resulting CsgG:CsgF complexes exhibited different signal-to-noise ratios (Figure 2). 5) (Fig. 23, A–F). Surprisingly, different signals containing the same contraction region were observed. CsgF peptides of different lengths were reconstituted in the same CsgG pore to form CsgG:CsgF complexes. When forming the shortest CsgF peptide ( The CsgG:CsgF complex containing nucleotides 1-29 of SEQ ID NO:6 showed the greatest coverage and The CsgG:CsgF complex containing the longest CsgF peptide (1-45 of SEQ ID NO: 6) The minimum range is shown (Figure 24).
[0323] Materials and methods for analyte characterization: Proteins produced by the methods described below can be analyzed by the methods described above for structure determination. can be used interchangeably with that generated by
[0324] method Coexpression of CsgG:CsgF or CsgG:FCP complexes The gene encoding the CsgG protein and its mutants contains the ampicillin resistance gene. It is constructed with the Pt7 vector containing the gene. The gene encoding it, and its mutants, are expressed in pR 1 uL of both plasmids was incubated on ice for 10 minutes in a 50 uL The sample was then heated to 42°C for 45 seconds. Add 150uL of NEB SOC growth medium and place the sample in 300ml of ice. Incubate at 7°C with shaking at 250 rpm for 1 hour. / mL), ampicillin (100ug / mL) and chloramphenicol (34ug / 1 mL) on an agar plate and incubate overnight at 37°C. The sample was collected and treated with kanamycin (40ug / mL), ampicillin (100ug / mL) and The cells were cultured in 100 ml of LB medium containing chloramphenicol (34 μg / mL) and 37 The 25 mL starter culture was incubated overnight at 3 °C with shaking at 250 rpm. , 15mM MgSO4, kanamycin (40ug / mL), ampicillin (100ug 500 ml of LB medium containing 10 μg / mL of chloramphenicol (34 μg / mL) and incubated overnight at 37°C. The culture was grown for 7 hours, at which point the OD 600 is 3 Lactose (final concentration 1.0%), glucose (final concentration 0.2%) ) and rhamnose (final concentration 2 mM) were added, the temperature was reduced to 18°C, while the shaking was continued for 25 The culture was centrifuged at 6000 rpm for 20 minutes at 4°C. The supernatant was discarded and the pellet was retained. Cells were stored at -80°C until purification.
[0325] CsgG pore with and without C-terminal Strep or His tag, or Expression of CsgF with or without a C-terminal Strep or His tag All CsgG proteins and CsgG proteins or CsgG proteins or All genes encoding FCP proteins were cloned into pT7, which contains the ampicillin resistance gene. The vector is constructed with kanamycin if omitted from all vectors and buffers. Except for this, the expression procedure is the same as above.
[0326] Cell lysis (co-expression complex or individual CsgG / CsgF / FCP proteins) Lysis buffer: 50mM Tris, pH8.0, 150mM Nacl, 0.1% D DM, 1x Bugaster Protein Extraction Reagent (Merck), 2.5uL Benzonase Nuclease lyase (stock ≥ 250 units / µL) / 100 mL of dissolution buffer and 1 tablet of Sigma 5X volume of protease inhibitor cocktail / 100mL lysis buffer. Lyse the harvested cells by 1x weight using lysis buffer. The cells will be resuspended and homogenized. The lysate is incubated at room temperature for 4 hours until the product is formed. Spin at 1,000 rpm. Carefully extract the supernatant and place on a 0.2 μM Acrodisc. Filter through a syringe filter.
[0327] When CsgG contains a C-terminal Strep tag, and when CsgF or FCP contains a C-terminal Hi of CsgG or CsgF / FCP proteins or co-expressed complexes when containing an s tag. Strep Purification The filtered sample was then loaded onto a 5 mL StrepTrap column with the following parameters: Loading speed: 0.8 mL / min, total sample load: 10 mL, unbound wash: 10 CV ( 5 mL / min), additional wash: 10 CV (5 mL / min), elution: 3 CV (5 mL / min). Compatibility buffer: 50mL Tris, pH8.0, 150mM Nacl, 0.1% DD M, wash buffer: 50mL Tris, pH8.0, 2M Nacl, 0.1% DDM , Elution buffer: 50mL Tris, pH8.0, 150mM NaCl, 0.1% D DM, 10 mM desthiobiotin. Eluted samples are collected.
[0328] When CsgG contains a C-terminal Strep tag, and when CsgF or FCP contains a C-terminal Hi of CsgG or CsgF / FCP proteins or co-expressed complexes when containing an s tag. His purification Filtered sample or pooled elution peak from Strep purification (if complex) Prepare a 5 mL HisTrap column using the above parameters except for the following buffer: Loaded: Affinity buffer and wash buffer: 50 mL Tris, pH 8.0, 15 0 mM NaCl, 0.1% DDM, 25 mM imidazole, elution: 50 mL Tris s, pH 8.0, 150 mM NaCl, 0.1% DDM, 350 mM imidazole The peak was eluted and centrifuged in a 30 kDa MWCO Merck Millipore centrifuge. Concentrate to a volume of 500 uL.
[0329] In vitro complex formation using in vivo purified components. Both expressed and purified CsgG and CsgF / FCP proteins were incubated at various ratios. They are mixed separately to determine the correct ratio. However, there is always an excess of CsgG. The complex was incubated overnight at 25°C. Excess CsgF was removed and DTT was removed from the buffer. To do this, the mixture was diluted to 50 mM Tris, pH 8.0, 150 mM NaCl, 0.1 % Superdex Increase 200 10 / 300 equilibrated with DDM The conjugate typically elutes between 9 and 10 mL on this column.
[0330] Gel filtration polishing step for complexes (co-expressed or produced in vitro) Optionally, strep-purify, or His-purify, or His-purify followed by strep The purified CsgG:CsgF or CsgG:FCP was further purified by gel filtration. A 500 μL sample was injected into a 1 mL sample loop and 50 mM Su equilibrated in Tris, pH 8.0, 150 mM NaCl, 0.1% DDM Perdex Increase 200 10 / 300 injected. The peaks typically eluted in 9-10 mL on this column when run at 1 mL / min. The samples were heated to 60°C for 15 minutes and centrifuged at 21,000 rcf for 10 minutes. The supernatant was collected for analysis. The sample was subjected to SDS-PAGE to identify the fractions eluted with the complex. and identified.
[0331] Cleavage of CsgF or FCP at the TEV protease site If CsgF or FCP contains a TEV cleavage site, use a C-terminal histidine tag to cleave the FCP. V-protease is added to the sample along with 2 mM DTT (additional amount is determined by the protein complex). The samples were mixed on a roller mixer at 25 rpm for 4°C. The mixture was then transferred back to the 5 mL HiStrap column and the flow-through was collected. All uncleaved proteins remained bound to the column. The same buffers and parameters were used for His purification described above. A final heating step is used.
[0332] In vivo purified CsgG pore and CsgG:FCP complex with synthetic FCP Purification of Lyophilized FCP peptides from Genscript and Lifetein. g of peptide was dissolved in 1 mL of nuclease-free ddH2O to obtain a 1 mg / mL sample. The sample was vortexed until no peptides were visible. Due to differences in expression levels of the variants, it is difficult to accurately measure the concentration. The intensity of the protein bands on SDS-PAGE relative to the color was used to roughly estimate the quality of the sample. Next, CsgG and FCP were mixed at a molar ratio of approximately 1:50, Incubate overnight at 25°C and 700 rpm. Heat the sample at 60°C for 15 minutes and incubate at 21,000 rpm. The mixture was centrifuged at 47°C for 10 minutes. The supernatant was taken for testing. If necessary, the complex was co-expressed. and purified as described above.
[0333] Purification of CsgG:CsgF or CsgG:FCP containing cysteine mutants His-purified and Strep-purified when one or both components contain cysteine Affinity buffers, wash buffers and elution buffers in the The same procedure as described above was performed to prepare CsgG:CsgF or CsgG:FCP complexes, except for the composition of (I, II, or III below). To purify the variants, all of these buffers should contain 2 mM DTT. 1 mM DTT was also added when cysteine-containing synthetic peptides were dissolved in ddH2O. It was. I. Co-expression of CsgG and CsgF or FCP II. In vivo purified CsgG:CsgF or In vitro production of CsgG:FCP complexes III. In vivo purified CsgG and CsgG:C using synthetic FCP In vitro production of sgF or CsgG:FCP complexes
[0334] Determination of Cys bond formation Two tubes containing 50 μL each were separated from the final elution. 2 mM DTT was added as a reducing agent to the other tube, and 100 μM Cu(II):1 -10-phenanthroline (33 mM:100 mM) was added as an oxidizing agent. Half of the sample was mixed 1:1 with Laemmli buffer containing 4% SDS. After heat treatment for 10 min (denaturing conditions), half of the sample was left untreated, and then 4-2 ml of the sample was diluted in TGS buffer. The analysis was performed on a 0% TGX gel (Bio-Rad Criterion).
[0335] Coupling in vitro transcription and translation (IVTT) All proteins were prepared from E. coli T7-S3 for circular DNA (Promega). By coupled in vitro transcription and translation (IVTT) using a 0-extraction system The complete 1 mM amino acid mixture minus cysteine was used to generate the complete A high-concentration 1 mM amino acid mixture minus methionine was mixed in equal amounts. The working amino acid solution required to produce the desired amount of protein was obtained. Amino acids (10 μL) Premixed solution (40 μL), [35S]L-methionine (2 μL, 1175 Ci / mm ol, 10 mCi / mL), plasmid DNA (16 uL, 400 ng / uL), and T7 S30 extract (30 uL) and rifampicin (2 uL, 20 mg / mL) Mix to create a 100uL reaction of IVTT protein. Synthesis was carried out at 30°C for 4 hours. The cells were then incubated overnight at room temperature. When produced by expression, the plasmid DNA encoding each component is mixed in equal amounts, and the mixture is A portion of the 16 μL was used for IVTT. After incubation, the tubes were heated at 22,000 g for 10 min. The resulting pellet was resuspended in MBSA (10 mM MgCl ). OPS, 1 mg / ml BSA pH 7.4) and centrifuged again under the same conditions. The proteins present in the pellet were resuspended in 1X Laemmli sample buffer and lysed for 3 min. The gel was then dried and analyzed by Carbohydrate. estream® Kodak® BioMax® MR Filter The gel was exposed to HCl overnight, and the film was then processed to visualize the proteins in the gel.
[0336] Samples for testing on MinION All samples were incubated in Brij58 (final concentration 0.1%) ...
Claims
1. A pore comprising a CsgG pore and a modified CsgF peptide, a pore, wherein the CsgG is bound to the pore to form a constriction within the pore.
2. The CsgF peptide comprises a C-terminal head domain of CsgF and a neck domain of CsgF. The pore of claim 1, which is a truncated CsgF peptide lacking at least a portion of the pore.
3. 3. The method of claim 1, wherein the CsgF peptide has a length of 25 to 50 amino acids. pores.
4. The CsgF peptide is any of residues 1 of SEQ ID NO:6 to residues 28-45 of SEQ ID NO:
6. or one of the sequences up to the corresponding residue of a homologue of SEQ ID NO: 6 or a variant of any of the sequences The pore according to any one of claims 1 to 3, comprising an amino acid sequence.
5. The CsgF peptide is SEQ ID NO:39 (residues 1-29 of SEQ ID NO:6) or a homolog thereof. or a variant thereof.
6. The CsgF peptide is selected from the group consisting of SEQ ID NO: 15 (residues 1 to 34 of SEQ ID NO: 6), SEQ ID NO: 40 ( SEQ ID NO:6 (residues 1-45 of SEQ ID NO:6), SEQ ID NO:54 (residues 1-30 of SEQ ID NO:6), or the sequence 55 (residues 1-35 of SEQ ID NO: 6) or a homologue or variant thereof. The pores described in
7. 7. The pore of claim 6, wherein within the CsgF peptide, there is One or more residues of sequence number 39, sequence number 40, sequence number 54, or sequence number 55 The group is modified, the pore.
8. The CsgF peptide is located at the following positions: G1, T4, F5, R8, N9, N11, F12 8. The pore of claim 7, comprising a modification to one or more of A26 and Q29.
9. The modification may be cysteine, a hydrophobic amino acid, a charged amino acid, a non-native reactive amino acid, 9. The pore according to claim 7 or 8, wherein the pore is formed by the introduction of a photoreactive amino acid.
10. The CsgF peptide is located at the following positions: N15, N17, A20, N24, A28 and A pore according to any one of claims 7 to 9, comprising one or more modifications of D34.
11. The CsgF peptide has the following substitutions: N15S / A / T / Q / G / L / V / I / F / Y / W / R / K / D / C, N17S / A / T / Q / G / L / V / I / F / Y / W / R / K / D / C, A20S / T / Q / N / G / L / V / I / F / Y / W / R / K / D / C, N24 S / T / Q / A / G / L / V / I / F / Y / W / R / K / D / C, A28S / T / Q / N / G / L / V / I / F / Y / W / R / K / D / C, and D34F / Y / W / R / K / N The pore of claim 10, comprising one or more of: / Q / C.
12. The CsgF peptide may have the following substitutions: G1C, T4C, N17S, and D34Y; The pore according to any one of claims 7 to 11, comprising one or more of:
13. the CsgF peptide further comprises all or part of an enzyme cleavage site at the C-terminus; The pore according to any one of claims 7 to 12.
14. 14. Any of claims 1 to 13, wherein the truncated CsgF peptide is inserted into the lumen of the CsgG pore. The pore according to any one of claims 1 to 4.
15. 15. Any one of claims 1 to 14, wherein the CsgG pore comprises 6 to 10 CsgG monomers. The pores described in paragraph .
16. 10. The method of claim 1, wherein the ratio of the CsgG monomer to the cleaved CsgF peptide in the pore is 1:
1.
16. The pore according to any one of 1 to 15.
17. 17. The method of claim 1, wherein the CsgF peptide and the CsgG pore are covalently linked. The pore according to any one of claims 1 to 4.
18. 18. The pore of claim 17, wherein the covalent bond is: (i) 132, 133, 136, 138, 140 of SEQ ID NO: 3 or homologues thereof; 142、144、145、147、149、151、153、155、183、185、 Positions corresponding to 187, 189, 191, 201, 203, 205, 207 or 209 The cysteine residues in (ii) 132, 133, 136, 138, 140 of SEQ ID NO: 3 or homologues thereof; 142、144、145、147、149、151、153、155、183、185、 Positions corresponding to 187, 189, 191, 201, 203, 205, 207 or 209 The pore is mediated by non-native reactive or photoreactive amino acids.
19. The CsgF peptide and the CsgG pore are represented by SEQ ID NO: 6 and SEQ ID NO: 3, respectively. 1 and 153, 4 and 133, 5 and 136, 8 and 187, 8 and 203, 9 and 203, 11 and 142, 11 and 201, 12 and 149, 12 and 203, 26 and 191, and 29 and 14 covalently linked via residues at positions corresponding to one or more of the four-position pairs; 19. The pore according to claim 17 or 18.
20. The covalent bond is a disulfide bond or click chemistry.
19. The pore according to any one of claims 19 to 19.
21. The interaction between the CsgF peptide and the CsgG pore is shown in SEQ ID NO: 6 and SEQ ID NO: 7, respectively. and SEQ ID NO: 3, 1 and 153, 4 and 133, 5 and 136, 8 and 187, 8 and 203, 9 and 2 03, 11 and 142, 11 and 201, 12 and 149, 12 and 203, 26 and 191, and and hydrophobic interactions at positions corresponding to one or more of the pairs of positions 29 and 144, or Any of claims 1 to 20, stabilized by electrical or covalent interactions. The pore described in claim 1.
22. The CsgG pore may comprise the following modification of SEQ ID NO: 3: (i) modifications at one or more of positions Y51, N55 and F56; (ii) a group selected from R97W or R97Y and R93W or R93Y; At least one substitution, (iii) deletion of V105, A106, and I107 of SEQ ID NO:3; (iv) positions R192, F193, I194, D105, Y196 of SEQ ID NO: 3; a deletion of one or more of Q197, R198, L199, and E201; (v) K94N / Q / R / F / Y / W / L / S, D43S, E44S, F48S / N / Q / Y / W / I / V / H / R / K, Q87N / R / K, N91K / R, R97F / Y / W / V / I / K / S / Q / H, E101I / L / A / H, N102K / Q / L / I / V / S / H, R110F / G / N, Q114R / K, R142Q / S, T150Y / A / at least one substitution selected from V / L / S / Q / N, and / or (vi) Positions I41, R93, A98, Q100, G103, T104, A106 , I107, N108, L113, S115, T117, Y130, K135, E170 , S208, D233, D238, E244, Q42, E44, L90, N91, I95 , a modification at one or more of A99, E101 and Q114, The pore according to any one of claims 1 to 21, comprising at least one monomer comprising:
23. The CsgG pore may comprise the following modification of SEQ ID NO: 3: (i) Y51A / I / V / S / T, N55A / I / V / S / T and F56 / A At least one substitution selected from: / I / V / S / T / Q; (ii) the substituted R97W; (iii) F193, I194, D195, Y196, Q197, R198 and L deletion of D195, Y196, Q197, R198 and L199, (iv) deletion of V105, A106 and I107; (v) a substitution selected from K94Q and K94N; (vi) Q42K or Q42R, E44N or E44Q, L90R or L9 0K, N91R or N91K, I95R or I95K, A99R or A99K, E 101H, E101K, E101N, E101Q or E101T, and / or Q1 14K, and / or (vii) at least one molecule containing one or more modifications corresponding to the substitution N55V; The pore of claim 22 comprising a polymer.
24. The CsgG pore comprises at least one CsgG pore having an R or K at a position corresponding to R192 of SEQ ID NO:
3. The pore according to any one of claims 1 to 23, wherein both monomers comprise one monomer.
25. A dual pore comprising two CsgG pores, wherein the CsgF peptide is The pore according to any one of claims 1 to 24, inserted into at least one lumen of the 。
26. A pore according to any one of claims 1 to 25, contained within a membrane.
27. The method comprises co-expressing one or more CsgG monomers and a CsgF peptide in a host cell. and thereby allowing transmembrane pore complex formation in the cell.
27. A method for generating pores according to any one of claims 26.
28. The method comprises contacting one or more purified CsgG monomers with a modified CsgF peptide. and thereby allowing the pores to be formed in vitro.
7. A method for generating pores according to any one of claims 6.
29. The modified CsgF peptide is cleaved to generate the pore of claim 27 or 28. method.
30. The method further comprises cleaving the CsgF peptide to form a CsgG binding region and a constriction within the pore. an enzyme cleavage site positioned to generate a truncated CsgF peptide comprising a region wherein the method comprises expressing a modified CsgF peptide comprising:
29. A method for generating a pore according to claim 27 or 28, comprising the step of cutting a cord.
31. the modified CsgF peptide is SEQ ID NO: 15 or SEQ ID NO: 16 or a homologue thereof or The method of any one of claims 27 to 30, comprising a mutant.
32. 1. A method for determining the presence, absence, or one or more characteristics of a target analyte, comprising: (i) a method according to any one of claims 1 to 26, wherein the target analyte is transported into the pore complex; contacting the target analyte with the pore of any one of claims 1 to 4; (ii) taking one or more measurements as the analyte migrates through the pore complex; and determining the presence, absence, or one or more characteristics of said analyte. ,method.
33. The analyte may be a (poly)peptide, a polysaccharide, a small organic compound or an inorganic compound (e.g., a pharmacological 33. The method of claim 32, wherein the active ingredient is a biologically active compound, a toxic compound, a pollutant, or the like.
34. 33. The method of claim 32, wherein the analyte is a polynucleotide.
35. 35. The method of claim 34, wherein the polynucleotide comprises at least one homopolymer region. method.
36. (i) the length of the polynucleotide; (ii) the identity of the polynucleotide; (iii) (iv) the sequence of the polynucleotide; (iv) the secondary structure of the polynucleotide; and (v) determining one or more characteristics of the polynucleotide, including whether it is modified; 36. The method of claim 34 or 35, comprising:
37. The pore according to any one of claims 1 to 26 can be used to detect polynucleotides or (poly) Methods for characterizing peptides.
38. A method according to any one of claims 1 to 26 for determining the presence, absence or one or more properties of a target analyte. Use of the pores according to any one of claims 1 to 4.
39. A target assay comprising (a) a pore according to any one of claims 1 to 26 and (b) a membrane component. A kit for characterizing objects.
40. A modified CsgF peptide containing a CsgG-binding region and a region that forms a constriction within the pore. Hmm, CsgF peptide.
41. 41. The method of claim 40, which is a truncated CsgF peptide lacking the C-terminal head domain of CsgF. CsgF peptide of.
42. 42. The CsgF peptide of claim 41, which lacks at least a portion of the neck domain of CsgF. Petite.
43. a C-terminal head domain of the CsgF and a neck domain of the CsgF, 43. The CsgF peptide of claim 42.
44. The CsgF of any one of claims 40 to 43, having a length of 25 to 50 amino acids. peptide.
45. Residue 1 of SEQ ID NO:6 to residues 28-45 of SEQ ID NO:6 or a homologue or variant thereof The C according to any one of claims 40 to 44, comprising up to one of the amino acid sequences sgF peptide.
46. 39 (residues 1-29 of SEQ ID NO:6) or a homologue or variant thereof. Item 46. The CsgF peptide of any one of Items 40 to 45.
47. SEQ ID NO:15 (residues 1-34 of SEQ ID NO:6), SEQ ID NO:54 (residues 1-30 of SEQ ID NO:6) ), SEQ ID NO:40 (residues 1-45 of SEQ ID NO:6), or SEQ ID NO:55 (residues 1-45 of SEQ ID NO:6) 47. The CsgF peptide of claim 46, comprising a CsgF peptide of any one of groups 1 to 35) or a homologue or variant thereof. Do.
48. 48. The CsgF peptide of any one of claims 45 to 47, comprising SEQ ID NO: 6, SEQ ID NO: 7, SEQ ID NO: 8, SEQ ID NO: 9, SEQ ID NO: 10, SEQ ID NO: 11, SEQ ID NO: 12, SEQ ID NO: 13, SEQ ID NO: 14, SEQ ID NO: 15, SEQ ID NO: 16, SEQ ID NO: 17 Sequence No. 15, residues 1 and 28-4 of SEQ ID NO:39, SEQ ID NO:40, and SEQ ID NO:54 5 or one or more residues in SEQ ID NO:55 (residues 1-35) are modified , CsgF peptide.
49. The following positions: G1, T4, F5, R8, N9, N11, F12, A26 and Q29 49. The CsgF peptide of claim 48, comprising a modification in one or more of:
50. The modification may be cysteine, a hydrophobic amino acid, a charged amino acid, a non-native reactive amino acid, 46. The CsgF peptide of claim 45, wherein the CsgF peptide is a CsgF peptide having a CsgF residue, or a CsgF residue containing a photoreactive amino acid.
51. At one or more of the following positions: N15, N17, A20, N24, A28 and D34 51. The CsgF peptide of any one of claims 48 to 50, comprising a modification in
52. Replacement of the following N15S / A / T / Q / G / L / V / I / F / Y / W / R / K / D / C, N 17S / A / T / Q / G / L / V / I / F / Y / W / R / K / D / C, A20S / T / Q / N / G / L / V / I / F / Y / W / R / K / D / C, N24S / T / Q / A / G / L / V / I / F / Y / W / R / K / D / C, A28S / T / Q / N / G / L / V / I / F / Y / W / R / K / D / C, and one or more of D34F / Y / W / R / K / N / Q / C 52. The CsgF peptide of claim 51, comprising:
53. one or more of the following substitutions: G1C, T4C, N17S, and D34Y or D34N The CsgF peptide of any one of claims 48 to 52, comprising the above.
54. Any of claims 40 to 53, further comprising all or part of an enzyme cleavage site at the C-terminus. A CsgF peptide according to any one of claims 1 to 4.
55. The CsgF peptide of any one of claims 40 to 54, further comprising a signal peptide. Chid.
56. A polynucleotide encoding the CsgF peptide of any one of claims 40 to 55. Do.
57. A pore complex, (i) a first opening, a middle portion including a β-barrel, a second opening, and a lumen extending from the opening through the intermediate portion to the second opening, a CsgG pore, the luminal surface of which defines a CsgG constriction; (ii) a plurality of modified CsgFs, each having a CsgF contraction domain and a CsgG binding domain; a modified CsgF peptide, wherein the modified CsgF peptide is inserted into the CsgG pore and binds the β-valerate. and forming a CsgF constriction region within the chain, the CsgG constriction region and the CsgF constriction region being a pore complex comprising: a pore complex coaxially spaced within the beta barrel of the CsgG pore; 。
58. one or more loop regions of the CsgG monomer whose luminal surface defines the CsgG constriction region 58. The pore composite of claim 57, comprising:
59. The CsgF contraction domain and the CsgG binding domain are located at the N-terminal end of the CsgF mature peptide.
58. The pore complex of claim 57, corresponding to minutes.
60. 58. The pore complex of claim 57, wherein the pore complex does not include CsgA, CsgB, and CsgE. Pore complex.
61. 58. The method of claim 57, wherein the CsgF peptide is covalently attached to the CsgG pore. Pore complex.