Novel protein pore

By incorporating a modified CsgF peptide into the CsgG pore to form a CsgG:CsgF complex with two leader heads, the nanopore sensing system achieves improved nucleotide discrimination, addressing the limitations of current technologies and enhancing sequencing performance.

JP7696948B2Active Publication Date: 2025-06-23VLAAMS INTERUNIVERSITAIR INST VOOR BIOTECHNOLOGIE VZW +2
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2023081191
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2017-06-30
Filing Date
2023-05-17
Publication Date
2025-06-23
Estimated Expiration
2038-07-02

AI Technical Summary

Technical Problem

Current nanopore sensing technologies face limitations in nucleotide discrimination, with the difference in current between nucleotides being insufficient for high-performance sequencing systems.

Method used

A modified CsgF peptide, specifically a cleaved CsgF fragment, is used to bind to the CsgG pore, introducing an additional channel constriction and forming a CsgG:CsgF complex with two consecutive leader heads, enhancing nucleotide discrimination.

Benefits of technology

The modified CsgF peptide improves the nucleotide discrimination capability of nanopore sensing systems by introducing an additional constriction, leading to a more differentiated current signature for polynucleotides, thereby enhancing sequencing performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007696948000023
    Figure 0007696948000023
  • Figure 0007696948000024
    Figure 0007696948000024
  • Figure 0007696948000025
    Figure 0007696948000025
Patent Text Reader

Abstract

To provide a novel method for improving the function of nanopore sensing.SOLUTION: The present invention relates to a novel protein pores and the use thereof in analyte detection and characterization. The present invention is particularly relevant to isolated pore complexes formed by CsgG-like pores and modified CsgF peptides or a homologue or mutant thereof, thereby providing an additional channel constriction or a reader head within the nanopore. The invention further relates to transmembrane pore complexes and methods of making pore complexes and use thereof in molecular sensing and nucleic acid sequencing applications.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to novel protein pores and their use in analyte detection and characterization. The present invention further relates to transmembrane pore complexes and methods for producing pore complexes, and their use in molecular sensing and nucleic acid sequencing applications.

Background Art

[0002] Nanopore sensing is an approach for analyte detection and characterization that relies on the observation of individual binding or interaction events between analyte molecules and an ion-conductive channel. A nanopore sensor can be created by placing a single nanopore of nanometer dimensions in an electrically insulating membrane and measuring the voltage-driven ion current through the pore in the presence of analyte molecules. The presence of an analyte within or near the nanopore changes the flow of ions through the pore, resulting in a change in the ion current or current measured across the channel. The identity of the analyte is revealed through its unique current signature, particularly the duration and extent of current blockades, and differences in current levels during its interaction with the pore. Analytes can be organic and inorganic small molecules, as well as various biological or synthetic macromolecules and polymers including polynucleotides, polypeptides, and polysaccharides. Nanopore sensing can reveal identity and perform single molecule counting of detected analytes, but it can also provide information regarding analyte composition such as nucleotide, amino acid, or glycan sequences, as well as information regarding the presence of modifications of bases, amino acids, or glycans such as methylation and acylation, phosphorylation, hydroxylation, oxidation, reduction, glycosylation, decarboxylation, deamination, and others. Nanopore sensing is capable of changing the flow of ions through the pore, resulting in a change in the ion current or current measured across the channel. The presence of an analyte within or near the nanopore changes the flow of ions through the pore, resulting in a change in the ion current or current measured across the channel. The identity of the analyte is revealed through its unique current signature, particularly the duration and extent of current blockades, and differences in current levels during its interaction with the pore. Analytes can be organic and inorganic small molecules, as well as various biological or synthetic macromolecules and polymers including polynucleotides, polypeptides, and polysaccharides. Nanopore sensing can reveal identity and perform single molecule counting of detected analytes, but it can also provide information regarding analyte composition such as nucleotide, amino acid, or glycan sequences, as well as information regarding the presence of modifications of bases, amino acids, or glycans such as methylation and acylation, phosphorylation, hydroxylation, oxidation, reduction, glycosylation, decarboxylation, deamination, and others. Nanopore sensing is capable of changing the flow of ions through the pore, resulting in a change in the ion current or current measured across the channel. The presence of an analyte within or near the nanopore changes the flow of ions through the pore, resulting in a change in the ion current or current measured across the channel. The identity of the analyte is revealed through its unique current signature, particularly the duration and extent of current blockades, and differences in current levels during its interaction with the pore. Analytes can be organic and inorganic small molecules, as well as various biological or synthetic macromolecules and polymers including polynucleotides, polypeptides, and polysaccharides. Nanopore sensing can reveal identity and perform single molecule counting of detected analytes, but it can also provide information regarding analyte composition such as nucleotide, amino acid, or glycan sequences, as well as information regarding the presence of modifications of bases, amino acids, or glycans such as methylation and acylation, phosphorylation, hydroxylation, oxidation, reduction, glycosylation, decarboxylation, deamination, and others. Nanopore sensing is capable of changing the flow of ions through the pore, resulting in a change in the ion current or current measured across the channel. The presence of an analyte within or near the nanopore changes the flow of ions through the pore, resulting in a change in the ion current or current measured across the channel. The identity of the analyte is revealed through its unique current signature, particularly the duration and extent of current blockades, and differences in current levels during its interaction with the pore. Analytes can be organic and inorganic small molecules, as well as various biological or synthetic macromolecules and polymers including polynucleotides, polypeptides, and polysaccharides. Nanopore sensing can reveal identity and perform single molecule counting of detected analytes, but it can also provide information regarding analyte composition such as nucleotide, amino acid, or glycan sequences, as well as information regarding the presence of modifications of bases, amino acids, or glycans such as methylation and acylation, phosphorylation, hydroxylation, oxidation, reduction, glycosylation, decarboxylation, deamination, and others. Nanopore sensing , which may enable rapid and inexpensive polynucleotides, with lengths of dozens to tens of thousands of bases provides single molecule sequencing of polynucleotides.

[0003] Two important components of polymer characterization using nanopore sensing are: (1) control of polymer movement through the pore, and (2) discrimination of the constituent elements when the polymer moves through the pore. During nanopore sensing, the narrowest part of the pore forms a leader head that is the most differentiated part of the nanopore with respect to the current signature as a function of the analyte passing through. CsgG has been identified from Escherichia coli as a non-selective protein secretion channel without a gate (Goyal et al., 2014) and has been used as a nanopore for detecting and characterizing analytes. Mutations to wild-type CsgG pores that improve pore properties in this context have been disclosed (WO2016 / 034591, WO2017 / 149316, WO2017 / 149317 and WO2017 / 149318, PCT / GB2018 / 051191, all of which are incorporated herein by reference). When the analyte is a polynucleotide, nucleotide discrimination is achieved by the passage through such mutant pores, but the current signature is sequence-dependent such that the height of the channel constriction and the extent of the interaction surface with the analyte affect the relationship between the observed current and the polynucleotide sequence, and it has been shown that multiple nucleotides contribute to the observed current. The current range for nucleotide discrimination is through mutations in the CsgG pore

[0004] Although it has been improved, if the difference in current between nucleotides can be further improved, the sequencing system will have higher performance. Therefore, there is a need to identify new methods for improving the function of nanopore sensing. SUMMARY OF THE INVENTION PROBLEMS TO BE SOLVED BY THE INVENTION

[0005] The present disclosure relates to modified CsgF peptides, particularly cleaved CsgF fragments, that bind to the CsgG pore and thereby introduce another additional channel or pore constriction into the CsgG pore. Other aspects of the present invention also relate to the use of the CsgG:CsgF complex and the modified CsgF peptide or fragment within an isolated transmembrane pore complex and a nanopore sensing platform having two consecutive leader heads. MEANS FOR SOLVING THE PROBLEM

[0006] A first aspect of the present invention relates to a pore comprising a CsgG pore and a CsgF peptide. In one aspect, the CsgF peptide comprises a CsgG binding region and a region that forms a constriction within the pore. In one aspect, the CsgF peptide is a cleaved CsgF peptide lacking the C-terminal head domain of CsgF. In another aspect, the CsgF peptide is a cleaved CsgF peptide lacking the C-terminal head and a portion of the neck domain of CsgF. In another aspect, the CsgF peptide is a cleaved CsgF peptide lacking the C-terminal head domain and the neck domain of CsgF. The pore is also referred to herein as an isolated pore complex, as an isolated pore complex. The isolated pore complex may be a CsgG pore or a homolog or mutant thereof ​​​​​​​​​​​​Variants, as well as modified CsgF peptides or homologues or mutants thereof, particularly truncated CsgF peptides. In one embodiment, the modified Cs comprises a sgF fragment or a homologue or mutant thereof. The gF peptide, or a homolog or mutant thereof, is a CsgG pore or a homolog thereof. In another embodiment, the isolated pore complex is located within the lumen of the mutant. one channel constriction, one positioned or provided by the CsgG pore , the constriction loop is formed by another additional channel constriction or The leader head is bound by a modified CsgF peptide or a homologue or mutant thereof. In one embodiment, the CsgG-pore or CsgG-like pore is , a mutant CsgG pore rather than a wild-type pore, in certain embodiments, e.g. In another embodiment, the modified CsgF peptide has a mutation in the channel contraction loop. Isolated pore complexes containing tid or its homologs or mutants have a pore size of 0.5 nm In one embodiment, the CsgF channel constriction has a diameter in the range of ∼2.0 nm. The pore complex comprises: (i) a first opening, a middle portion including a β-barrel, a second opening, and A CsgG pore having a lumen extending from a first opening through an intermediate portion to a second opening. (ii) a luminal surface of the intermediate portion that defines a CsgG constriction region; and A CsgF peptide comprising a CsgF contraction domain and a CsgF binding domain (referred to herein as CsgF peptide). The modified CsgF peptide has a CsgG binding domain or a region of CsgF, The CsgF constrictor is formed within the β-barrel of the CsgG pore, and the CsgG constrictor and Csg The F constriction is spaced coaxially within the β-barrel of the CsgG pore. The inner cavity surface of the pore may contain one or more loop regions of CsgG monomers that define CsgG constriction. The CsgF constriction region and the CsgF binding region typically correspond to the N-terminal portion of the CsgF mature peptide. In one embodiment, the pore complex does not contain CsgA, CsgB, and CsgE.

[0007] In a second aspect, the present invention relates to a modified CsgF peptide or a homolog or mutant thereof, wherein the protein or peptide is modified through cleavage or deletion of a portion of the protein and comprises a CsgF fragment of SEQ ID NO: 6 or a homolog or mutant thereof. One embodiment relates to a modified or cleaved CsgF peptide, or a modified peptide of a CsgF homolog or mutant, wherein the modified peptide comprises SEQ ID NO: 39, or SEQ ID NO: 40 or a homolog or mutant thereof, or alternatively, the modified peptide comprises SEQ ID NO: 15 or a homolog or mutant thereof, or alternatively SEQ ID NO: 54, or SEQ ID NO: 55 or a homolog or mutant thereof. In another embodiment, a modified CsgF peptide is disclosed, wherein one or more positions within the region containing SEQ ID NO: 15 are modified, and the mutant is required to retain at least 35% amino acid identity to SEQ ID NO: 15 within a peptide fragment corresponding to the region containing SEQ ID NO: 15.

[0008] One embodiment relates to a pore comprising a CsgG pore and a modified CsgF peptide, wherein the modified CsgF peptide binds to CsgG and forms a constriction within the pore.

[0009] One embodiment relates to a polynucleotide encoding the modified CsgF peptide or a homolog or mutant thereof according to the second aspect of the present invention. In another embodiment, the isolated pore complex comprising Csg G pores and the modified CsgF peptide or a homolog or mutant thereof is characterized in that the modified CsgF peptide is the peptide provided by the peptide disclosed in the second aspect of the present invention.

[0010] Another embodiment relates to an isolated pore complex, wherein the modified CsgF peptide and the Cs gG pores or monomers of said pores or homologs or mutants thereof are covalently bonded. More specifically, the bond is made via a cysteine residue at a position corresponding to 132, 133, 136, 138, 140, 142, 144, 145, 147, 1 49, 151, 153, 155, 183, 185, 187, 189, 191, 201, 2 03, 205, 207 or 209 within the CsgG monomer, or via a non-natural reactive or photoreactive amino acid.

[0011] A preferred embodiment relates to an isolated transmembrane pore complex, or a membrane composition, which comprises the isolated pore complex of the present invention and components of the membrane. In particular, the transmembrane pore complex or membrane composition is composed of the isolated pore complex of the present invention, components of the membrane or components of the insulating layer.

[0012] One embodiment relates to a method for generating a pore disclosed herein, the method comprising one or more CsgG monomers disclosed herein and a CsgF peptide disclosed herein ​​​​​​​​​Co-express in the host cell to enable the formation of transmembrane pore complexes intracellularly . The CsgF peptide may be produced intracellularly by cleaving a modified CsgF peptide or protein that contains an enzyme cleavage site at an appropriate position in the amino acid sequence .

[0013] One embodiment relates to a method for generating pores disclosed herein, the method comprising contacting one or more purified CsgG monomers with one or more purified modified CsgF peptides thereby enabling the in vitro formation of pores. The modified CsgF peptide may be a peptide containing an enzyme cleavage site at an appropriate position in the amino acid sequence that is cleaved before or after pore formation .

[0014] A third aspect of the invention relates to a method for generating said transmembrane pore complex, wherein the pore is an isolated complex formed by a CsgG pore or a homolog or mutant thereof, and a modified CsgF peptide or a homolog or mutant thereof, the method comprising co-expressing in a suitable host cell CsgG SEQ ID NO: 2 or a homolog or mutant thereof, and modified or cleaved CsgF (including the fragment of SEQ ID NO: 5) or a homolog or mutant thereof thereby enabling the in vivo formation of pore complexes. In certain embodiments, said modified CsgF peptide or a homolog or mutant thereof comprises SEQ ID NO: 12 or SEQ ID NO: 14 or a homolog or mutant thereof . Alternatively, a method for generating an isolated pore complex is to use SEQ ID NO: 3 or a homolog or mutant thereof for in vitro reconstitution of the pore complex . G or a homolog or mutant thereof Contact the G monomer with a modified CsgF peptide or a homolog or mutant thereof including a step. In certain embodiments, the modified CsgF peptide of the method comprises SEQ ID NO: 15 or SEQ ID NO: 16 or a homolog or mutant thereof.

[0015] Another aspect of the invention relates to a method for determining the presence, absence or one or more characteristics of a target analyte, the method comprising (i) contacting the target analyte with the isolated pore complex or transmembrane pore complex such that the target analyte moves within the pore channel; (ii) taking one or more measurements as the analyte moves through the pore channel and thereby determining the presence, absence or one or more characteristics of the analyte. In one embodiment, the analyte is a polynucleotide. In particular, the method of using a polynucleotide as an analyte comprises determining one or more characteristics selected from (i) the length of the polynucleotide, (ii) the identity of the polynucleotide, (iii) the sequence of the polynucleotide, (iv) the secondary structure of the polynucleotide, and (v) whether the polynucleotide is modified. (ii) taking one or more measurements as the analyte moves through the pore channel and thereby determining the presence, absence or one or more characteristics of the analyte. In one embodiment, the analyte is a polynucleotide. In particular, the method of using a polynucleotide as an analyte comprises determining one or more characteristics selected from (i) the length of the polynucleotide, (ii) the identity of the polynucleotide, (iii) the sequence of the polynucleotide, (iv) the secondary structure of the polynucleotide, and (v) whether the polynucleotide is modified.

[0016] In one embodiment, the analyte is a polynucleotide. In particular, the method of using a polynucleotide as an analyte comprises determining one or more characteristics selected from (i) the length of the polynucleotide, (ii) the identity of the polynucleotide, (iii) the sequence of the polynucleotide, (iv) the secondary structure of the polynucleotide, and (v) whether the polynucleotide is modified. In another embodiment, the analyte is a protein or a peptide, and in a further embodiment, the analyte is a polysaccharide or a small organic or inorganic compound such as, but not limited to, a pharmacologically active compound, a toxic compound, and a contaminant. In another embodiment, the polynucleotide or (poly using an isolated transmembrane pore complex including a step of determining one or more characteristics selected from (i) the length of the polynucleotide, (ii) the identity of the polynucleotide, (iii) the sequence of the polynucleotide, (iv) the secondary structure of the polynucleotide, and (v) whether the polynucleotide is modified.

[0017] In another embodiment, the analyte is a protein or a peptide, and in a further embodiment, the analyte is a polysaccharide or a small organic or inorganic compound such as, but not limited to, a pharmacologically active compound, a toxic compound, and a contaminant. In another embodiment, the polynucleotide or (poly using an isolated transmembrane pore complex

[0018] In another embodiment, a polynucleotide or (poly A method for characterizing a peptide is described, where the pore complex is a CsgG pore or a homolog or mutant thereof, and an isolated complex comprising a modified CsgF peptide or a homolog or mutant thereof. In particular, the CsgG pore or a homolog or mutant thereof comprises 6 - 10 CsgG monomers forming the CsgG pore channel.

[0019] A further aspect of the invention relates to the use of the isolated pore complex or transmembrane pore complex according to the foregoing aspect of the invention for determining the presence, absence, or one or more characteristics of a target analyte. Further, the invention also relates to a kit for characterizing a target analyte comprising (a) the isolated pore complex and (b) components of a membrane.

[0020] Description of the Drawings The drawings described are schematic and non - limiting. In the drawings, the sizes of some elements may be exaggerated for illustrative purposes and may not be drawn to scale. Brief Description of the Drawings

[0021]

Figure 1

Figure 2

Figure 3-1

Figure 3-2

Figure 4-1

Figure 4-2

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10-1

Figure 10-2

Figure 10-3

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23-1

Figure 23-2

Figure 24

Figure 25

Figure 26

Figure 27-1

Figure 27-2

Figure 28

Figure 29

Figure 30-1

Figure 30-2

Figure 30-3

Figure 30-4

Figure 31

Best Mode for Carrying Out the Invention

[0022] The present invention will be described with respect to specific embodiments and with reference to specific drawings, but the invention is not limited thereto and is limited only by the scope of the claims. The reference signs in the claims shall not be construed as limiting the scope. Of course, it goes without saying that not all aspects or advantages will necessarily be achieved according to any particular embodiment of the present invention. Therefore, those skilled in the art will recognize that the present invention can be embodied or performed in a manner that achieves or optimizes one or more of the advantages taught or recommended in this specification, as taught in this specification.

[0023] The present invention, both in terms of its configuration and method of operation, together with its features and advantages, can be best understood when the following detailed description is read in conjunction with the accompanying drawings. The aspects and advantages of the present invention will become apparent and will be elucidated by referring to the embodiments described below. Throughout this specification, references to "one embodiment" or "an embodiment" mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present invention. Thus, the appearances of the phrases "in one embodiment" or "in an embodiment" throughout this specification are not necessarily referring to the same embodiment, although they may be. Similarly, of course, in the illustrative description of the embodiments of the present invention, various features of the present invention may be. However, this method of the present disclosure should not be construed as reflecting an intention that more features are required than those explicitly recited in each claim. Rather, as reflected in the following claims, aspects of the invention are fewer than all of the features of the single embodiment disclosed above.

[0024] As used in this specification and the appended claims, the singular forms "a", "an", and "the" include plural referents unless the content clearly dictates otherwise. Thus, for example, "polynucleotide" includes two or more polynucleotides, "polynucleotide-binding protein" includes two or more such proteins, "helicase" includes two or more helicases, "monomer" means two or more "monomers", and "pore" includes two or more pores and the like.

[0025] Throughout this specification, the standard one-letter codes for amino acids are used. These are as follows: alanine (A), arginine (R), asparagine (N), aspartic acid (D), cysteine (C), glutamic acid (E), glutamine (Q), glycine (G), histidine (H), isoleucine (I), leucine (L), lysine (K), methionine (M), phenylalanine (F), proline (P), serine (S), threonine (T), tryptophan (W), tyrosine (Y), and valine (V). Standard substitution notations are also used, i.e., Q42R means that Q at position 42 is replaced by R.

[0026] In paragraphs where different amino acids at specific positions are separated by the / symbol, the / symbol means "or". For example, Q87R / K means Q87R or Q87K. That is, it means "or". For example, Q87R / K means Q87R or Q87K.

[0027] In paragraphs where different positions are separated by the / symbol, the / symbol means "and", and Y51 / N55 means Y51 and N55.

[0028] All amino acid substitutions, deletions, and / or additions disclosed herein, unless stated to the contrary, mean mutant CsgG monomers containing variants of the sequence shown in SEQ ID NO: 3. That is, it means mutant CsgG monomers containing variants of the sequence shown in SEQ ID NO: 3.

[0029] References to CsgG monomers that are mutants containing variants of the sequence shown in SEQ ID NO: 3 include mutant CsgG monomers containing variants of the sequences described in further SEQ ID NOs. disclosed below. Amino acid substitutions, deletions, and / or additions can be made to CsgG monomers containing variants of the sequences corresponding to the substitutions, deletions, and / or additions disclosed herein with reference to mutant CsgG monomers containing variants of the sequence shown in SEQ ID NO: 3. That is, amino acid substitutions, deletions, and / or additions can be made to CsgG monomers containing variants of the sequences corresponding to the substitutions, deletions, and / or additions disclosed herein with reference to mutant CsgG monomers containing variants of the sequence shown in SEQ ID NO: 3. That is, amino acid substitutions, deletions, and / or additions can be made to CsgG monomers containing variants of the sequences corresponding to the substitutions, deletions, and / or additions disclosed herein with reference to mutant CsgG monomers containing variants of the sequence shown in SEQ ID NO: 3. That is, amino acid substitutions, deletions, and / or additions can be made to CsgG monomers containing variants of the sequences corresponding to the substitutions, deletions, and / or additions disclosed herein with reference to mutant CsgG monomers containing variants of the sequence shown in SEQ ID NO: 3.

[0030] All publications, patents, and patent applications cited herein are hereby incorporated by reference in their entirety, whether supra or infra. That is, all publications, patents, and patent applications cited herein are hereby incorporated by reference in their entirety, whether supra or infra.

[0031] Definition When an indefinite or definite article (e.g., "a", or "an", " the") is used in reference to a singular noun, this includes the plural of that noun, unless otherwise specifically stated. When the term "comprising" is used in this specification and the claims, other That is, when the term "comprising" is used in this specification and the claims, other An element or step is not excluded. Further, the first, second, third and similar terms in the specification and claims are used to distinguish similar elements and are not necessarily used to describe a sequential or temporal order. The terms used are interchangeable under appropriate circumstances, and it should be understood that the embodiments of the invention described herein are capable of operating in other arrangements described or illustrated herein. The following terms or definitions are provided only to assist in understanding the invention. Unless specifically defined herein, all terms used herein have the same meaning to those skilled in the art of the invention. In particular, for implementers, for the definitions and terms of the art, see Sambrook et al., Molecular Cloning: A Laboratory Man ual, 4 ed., Cold Spring Harbor Press, Plainsview, New York (2012), and Ausubel et al., Current Protocols in Molecular Biology (Supplement 114), John Wiley & S th ons, New York (2016). The definitions provided herein should not be construed as having a narrower scope than those understood by those skilled in the art. When used herein, the term "about" includes variations of ±20% or ±10%, more preferably ±5%, more preferably ±1% and more preferably ±0.1% that are appropriate for carrying out the disclosed method and are used herein.

[0032] ​​​​​​

[0033] As used herein, "nucleotide sequence", "DNA sequence", or "nucleic acid molecule" refers to a polymer of nucleotides of either ribonucleotides or deoxyribonucleotides of any length. This term refers only to the primary structure of the molecule. Thus, this term includes double-stranded DNA, as well as single-stranded DNA and RNA. As used herein the term "nucleic acid" refers to a single-stranded or double-stranded covalently linked nucleotide sequence in which the 3' and 5' termini of each nucleotide are joined by phosphodiester bonds. Polynucleotides may be composed of deoxyribonucleotide bases or ribonucleotide bases. Nucleic acids may be produced synthetically in vitro or isolated from natural sources. Nucleic acids may further include modified DNA or RNA, such as methylated DNA or RNA, or RNA that has undergone post-translational modifications such as 5' capping with 7-methylguanosine, 3' processing such as cleavage and polyadenylation, and splicing. Nucleic acids may further include synthetic nucleic acids such as hexitol nucleic acid (HNA), cyclohexene nucleic acid (CeNA), threose nucleic acid (TNA), glycerol nucleic acid (GNA), locked nucleic acid (LNA), and peptide nucleic acid (PNA) (collectively "XNA"). The size of a nucleic acid referred to herein as a "polynucleotide" is generally expressed as the number of base pairs (bp) for double-stranded polynucleotides or as the number of nucleotides (nt) for single-stranded polynucleotides. 1000 bp or 1000 nt is equal to 1 kilobase (kb). Polynucleotides having a length of less Nucleotides are typically referred to as "oligonucleotides" and can include primers for use in DNA manipulations such as polymerase chain reaction (PC R).

[0034] As used herein, "gene" includes both the promoter region of the gene as well as the co ding sequence. This can be the case for genomic sequences (including possible introns), or for cDNA derived from spliced messenger that is operably linked to a promoter sequence. There are both cases where it is the genomic sequence (including possible introns), and cases where it is cDNA derived from spliced messenger that is operably linked to a promoter sequence .

[0035] "Coding sequence" is a nucleotide sequence that is transcribed into mRNA and / or translated into a polypeptide when placed under the control of appropriate regulatory sequences. The boundaries of the coding sequence are determined by a translation start codon at the 5' end and a translation stop codon at the 3' end. The coding sequence can include, but is not limited to, mRNA, cDNA, recombinant nucleotide sequences, or genomic DNA, and may also be present in certain situations with introns . The term "amino acid" in the context of the present disclosure is used in its broadest sense and means an organic compound containing an amine (NH2) and a carboxyl (COOH) functional group together with a side chain (e.g., an R group) specific to each amino acid. In some embodiments, the amino acid refers to a natural L α-amino acid or residue. One-letter and three-letter abbreviations commonly used for natural amino acids are used herein: A = Ala, C = Cys, D = Asp, E = Glu, F = Phe, G = Gly, H = His, I = Ile, K = Lys .

[0036] In some embodiments, the amino acid refers to a natural L α-amino acid or residue. One-letter and three-letter abbreviations commonly used for natural amino acids are used herein: A = Ala, C = Cys, D = Asp, E = Glu, F = Phe, G = Gly, H = His, I = Ile, K = Lys . In some embodiments, the amino acid refers to a natural L α-amino acid or residue. One-letter and three-letter abbreviations commonly used for natural amino acids are used herein: A = Ala, C = Cys, D = Asp, E = Glu, F = Phe, G = Gly, H = His, I = Ile, K = Lys ; D = Asp, E = Glu, F = Phe, G = Gly, H = His, I = Ile, K = Lys ​s, L = Leu, M = Met, N = Asn, P = Pro, Q = Gln, R = Arg, S = Ser, T = Thr, V = Val, W = Trp, Y = Tyr (Lehninger, A . L., (1975) Biochemistry, 2d ed., pp. 7 1 - 92, Worth Publishers, New York). “Amino acid” The term further includes chemically modified amino acids such as D - amino acids, retro - inverso amino acids, and amino acid analogs, natural amino acids such as norleucine which are not normally incorporated into proteins, and chemically synthesized compounds having properties characteristic of amino acids such as β - amino acids as known in the art. For example, analogs or mimetics of phenylalanine or proline that enable a peptide compound to have the same structural constraints as natural Phe or Pro are included in the definition of amino acids. Such analogs and mimetics are herein referred to as “functional equivalents” of the respective amino acids. Other examples of amino acids are described in Roberts and Vellaccio, The Peptide s: Analysis, Synthesis, Biology, Gross a nd Meiehofer, eds., Vol. 5 p. 341, Acade mic Press, Inc., N.Y. 1983 and are incorporated herein by reference. The terms “protein”, “polypeptide”, and “peptide” are used interchangeably herein and refer to polymers of amino acid residues and variants and synthetic analogs of amino acid residues. Thus, these terms include polymers in which one or more amino acid residues are synthetic non - natural amino acids. “Protein”, “polypeptide” and “peptide” are also used interchangeably herein and refer to polymers of amino acid residues, variants and synthetic analogs of amino acid residues. Thus, these terms include polymers in which one or more amino acid residues are synthetic non - natural

[0037] amino acids. The terms “protein”, “polypeptide” and “peptide” are used interchangeably herein and refer to polymers of amino acid residues, variants and synthetic analogs of amino acid residues. Thus, these terms include polymers in which one or more amino acid residues are synthetic non - natural Amino acids (such as chemical analogs of the corresponding natural amino acids), and natural amino acid polymers - are applied to the amino acid polymer. The polypeptide can also undergo maturation modification processes or post-translational modification processes including, but not limited to, glycosylation, proteolytic cleavage, lipidation, signal peptide cleavage, propeptide cleavage, phosphorylation, etc. "Recombinant polypeptide" means a polypeptide produced through the use of recombinant techniques, for example, via the expression of a recombinant polynucleotide or a synthetic polynucleotide. When a chimeric polypeptide or a biologically active portion thereof is produced recombinantly, it is also preferably substantially free of culture medium, for example, the culture medium occupies less than about 20%, more preferably less than about 10%, and most preferably less than about 5% of the volume of the protein preparation. "Isolated" means a material that is substantially or essentially free of the components that are normally associated with it in its natural state. For example, as used herein, "isolated polypeptide" refers to a polypeptide purified from the molecules adjacent to it in its naturally occurring state, for example, a protein complex or a CsgF peptide removed from the molecules present in the production host adjacent to the polypeptide. The isolated CsgF peptide (optionally selectively cleaved CsgF peptide) can be produced by chemical amino acid synthesis or by recombinant production. The isolated complex can be produced by in vitro reconstitution after purification of the components of the complex (e.g., CsgG pores and CsgF peptide), or can be produced by recombinant co-expression. For example, as used herein, "isolated polypeptide" refers to a polypeptide purified from the molecules adjacent to it in its naturally occurring state, for example, a protein complex or a CsgF peptide removed from the molecules present in the production host adjacent to the polypeptide. The isolated CsgF peptide (optionally selectively cleaved CsgF peptide) can be produced by chemical amino acid synthesis or by recombinant production. The isolated complex can be produced by in vitro reconstitution after purification of the components of the complex (e.g., CsgG pores and CsgF peptide), or can be produced by recombinant co-expression. The isolated CsgF peptide (optionally selectively cleaved CsgF peptide) can be generated by chemical amino acid synthesis or by recombinant production. The isolated complex can be generated by in vitro reconstitution after purification of the components of the complex (e.g., CsgG pores and CsgF peptide), or can be generated by recombinant co-expression. - are produced by in vitro reconstitution after purification of the components of the complex (e.g., CsgG pores and CsgF peptide), or can be produced by recombinant co-expression.

[0038] ​"Ortholog" and "paralog" encompass evolutionary concepts used to explain the ancestral relationships of genes. A paralog is a gene within the same species that originated from the duplication of an ancestral gene, while an ortholog is a gene from different organisms that originated through speciation and also derives from a common ancestral gene. A paralog is a gene within the same species that originated from the duplication of an ancestral gene, while an ortholog is a gene from different organisms that originated through speciation and also derives from a common ancestral gene. The "homolog", "homologs" of a protein are peptides, oligopeptides, polypeptides, proteins, and enzymes that have amino acid substitutions, deletions, and / or insertions relative to the unmodified or wild-type protein in question, and have biological and functional activities similar to those of the unmodified protein from which they are derived. As used herein, the term "amino acid identity" refers to the degree to which sequences are identical amino acid by amino acid over a certain comparison frame. Accordingly, the "percentage of sequence identity" is calculated by aligning and comparing two optimally aligned sequences over a comparison frame, determining the number of positions at which identical amino acid residues (e.g., Ala, Pro, Ser, Thr, Gly, Val, Leu, Ile, Phe, Tyr, Trp, Lys, Arg, His, Asp, Glu, Asn, Gln, Cys, and Met) occur in both sequences, determining the number of matching positions, dividing by the total number of positions in the comparison frame (i.e., the frame size), and multiplying the result by 100 to obtain the percentage of identity of the sequences.

[0039] The "homolog", "homologs" of a protein are peptides, oligopeptides, polypeptides, proteins, and enzymes that have amino acid substitutions, deletions, and / or insertions relative to the unmodified or wild-type protein in question, and have biological and functional activities similar to those of the unmodified protein from which they are derived. As used herein, the term "amino acid identity" refers to the degree to which sequences are identical amino acid by amino acid over a certain comparison frame. Accordingly, the "percentage of sequence identity" is calculated by aligning and comparing two optimally aligned sequences over a comparison frame, determining the number of positions at which identical amino acid residues (e.g., Ala, Pro, Ser, Thr, Gly, Val, Leu, Ile, Phe, Tyr, Trp, Lys, Arg, His, Asp, Glu, Asn, Gln, Cys, and Met) occur in both sequences, determining the number of matching positions, dividing by the total number of positions in the comparison frame (i.e., the frame size), and multiplying the result by 100 to obtain the percentage of identity of the sequences. The "CsgG pore" refers to a pore containing multiple CsgG monomers. Each CsgG monomer may be a wild-type monomer from E. coli (SEQ ID NO: 3), a wild-type homolog of E. coli CsgG, such as those shown in SEQ ID NOs: 68 - 88. The "homolog", "homologs" of a protein are peptides, oligopeptides, polypeptides, proteins, and enzymes that have amino acid substitutions, deletions, and / or insertions relative to the unmodified or wild-type protein in question, and have biological and functional activities similar to those of the unmodified protein from which they are derived. As used herein, the term "amino acid identity" refers to the degree to which sequences are identical amino acid by amino acid over a certain comparison frame. Accordingly, the "percentage of sequence identity" is calculated by aligning and comparing two optimally aligned sequences over a comparison frame, determining the number of positions at which identical amino acid residues (e.g., Ala, Pro, Ser, Thr, Gly, Val, Leu, Ile, Phe, Tyr, Trp, Lys, Arg, His, Asp, Glu, Asn, Gln, Cys, and Met) occur in both sequences, determining the number of matching positions, dividing by the total number of positions in the comparison frame (i.e., the frame size), and multiplying the result by 100 to obtain the percentage of identity of the sequences. The "CsgG pore" refers to a pore containing multiple CsgG monomers. Each CsgG monomer may be a wild-type monomer from E. coli (SEQ ID NO: 3), a wild-type homolog of E. coli CsgG, such as those shown in SEQ ID NOs: 68 - 88. The "homolog", "homologs" of a protein are peptides, oligopeptides, polypeptides, proteins, and enzymes that have amino acid substitutions, deletions, and / or insertions relative to the unmodified or wild-type protein in question, and have biological and functional activities similar to those of the unmodified protein from which they are derived. As used herein, the term "amino acid identity" refers to the degree to which sequences are identical amino acid by amino acid over a certain comparison frame. Accordingly, the "percentage of sequence identity" is calculated by aligning and comparing two optimally aligned sequences over a comparison frame, determining the number of positions at which identical amino acid residues (e.g., Ala, Pro, Ser, Thr, Gly, Val, Leu, Ile, Phe, Tyr, Trp, Lys, Arg, His, Asp, Glu, Asn, Gln, Cys, and Met) occur in both sequences, determining the number of matching positions, dividing by the total number of positions in the comparison frame (i.e., the frame size), and multiplying the result by 100 to obtain the percentage of identity of the sequences. The "CsgG pore" refers to a pore containing multiple CsgG monomers. Each CsgG monomer may be a wild-type monomer from E. coli (SEQ ID NO: 3), a wild-type homolog of E. coli CsgG, such as those shown in SEQ ID NOs: 68 - 88. The "CsgG pore" refers to a pore containing multiple CsgG monomers. Each CsgG monomer may be a wild-type monomer from E. coli (SEQ ID NO: 3), a wild-type homolog of E. coli CsgG, such as those shown in SEQ ID NOs: 68 - 88. The "CsgG pore" refers to a pore containing multiple CsgG monomers. Each CsgG monomer may be a wild-type monomer from E. coli (SEQ ID NO: 3), a wild-type homolog of E. coli CsgG, such as those shown in SEQ ID NOs: 68 - 88.

[0040] The "CsgG pore" refers to a pore containing multiple CsgG monomers. Each CsgG monomer may be a wild-type monomer from E. coli (SEQ ID NO: 3), a wild-type homolog of E. coli CsgG, such as those shown in SEQ ID NOs: 68 - 88. The "CsgG pore" refers to a pore containing multiple CsgG monomers. Each CsgG monomer may be a wild-type monomer from E. coli (SEQ ID NO: 3), a wild-type homolog of E. coli CsgG, such as those shown in SEQ ID NOs: 68 - 88. The "CsgG pore" refers to a pore containing multiple CsgG monomers. Each CsgG monomer may be a wild-type monomer from E. coli (SEQ ID NO: 3), a wild-type homolog of E. coli CsgG, such as those shown in SEQ ID NOs: 68 - 88. Any one of the amino acid sequences, or any variant thereof (e.g., any one of SEQ ID NOs: 3 and 68 ~88), such as a monomer having the variant). The variant CsgG monomer may also be referred to as a modified CsgG monomer or a mutant CsgG monomer. The modification or mutation in the variant includes, but is not limited to, any one or more of the modifications disclosed herein , or combinations of the above modifications.

[0041] For all aspects and embodiments of the present invention, the CsgG homolog means a polypeptide having at least 50%, 60%, 70%, 80%, 9 0%, 95% or 99% complete sequence identity to the wild-type E. coli CsgG shown in SEQ ID NO: 3. The CsgG homolog also means a polypeptide containing the PFAM domain PF03783, which is a characteristic of the CsgG-like protein . A list of currently known CsgG homologs and CsgG structures is described in . http: / / pfam.xfam.org / / family / PF037 83 Similarly, the CsgG homologous polynucleotide can include a polynucleotide having at least 50%, 60%, 70%, 80%, 90% to the wild-type E. coli CsgG shown in SEQ ID NO: 1, 95% or 99% complete sequence identity. Examples of CsgG homologs shown in SEQ ID NO: 3 have the sequences shown in SEQ ID NOs: 68 - 88. .

[0042] The term "modified CsgF peptide" or "CsgF peptide" defines a CsgF peptide cleaved from its C-terminus (e.g., the N-terminal fragment), and / or a CsgF peptide modified to include the cleavage site . The CsgF peptide is wild-type E. coli ​A fragment of CsgF (SEQ ID NO: 5 or SEQ ID NO: 6), or a fragment of a wild-type homolog of E. coli CsgF may be, for example, CsgF (e.g., a peptide containing any one of the amino acid sequences shown in SEQ ID NOs: 17 to 36), or any variant thereof ( e.g., one modified to contain a cleavage site), etc. For all aspects and embodiments of the present invention, the CsgF homolog means a polypeptide having at least 50%, 60%, 70%, 80%, 9

[0043] 0%, 95% or 99% complete sequence identity to the wild-type E. coli CsgF shown in SEQ ID NO: 6. In some embodiments, the CsgF homolog also means a polypeptide containing the PFAM domain PF10614, which is a characteristic of CsgF-like proteins. For a list of currently known CsgF homologs and CsgF structures, see http: / / pfam.xfam.org / / fa mily / PF10614 http: / / pfam.xfam.org / / fa mily / PF10614 is described. Similarly, the CsgF homologous polynucleotide can contain a polynucleotide having at least 50%, 60%, 7 0%, 80%, 90%, 95% or 99% complete sequence identity to the wild-type E. coli CsgF shown in SEQ ID NO: 4. An example of the cleaved region of the homolog of CsgG shown in SEQ ID NO: 6 has the sequence shown in SEQ ID NOs: 17 to 36. The term "N-terminal portion of the CsgF mature peptide" refers to a peptide having an amino acid sequence corresponding to the first 60, 50, or 40 amino acid residues starting from the N-terminus of the CsgF mature peptide (excluding the signal sequence). The CsgF mature peptide may be wild-type or a mutant (e.g., having one or more mutants).

[0044] a mutant (e.g., having one or more mutants).

[0045] Sequence identity can also be to a fragment or portion of a full-length polynucleotide or polypeptide. Thus, a sequence can have only 50% overall sequence identity with a full-length reference sequence, but the sequence of a particular region, domain, or subunit can share 80%, 90%, or 99% sequence identity with the reference sequence. For each CsgG homolog, the homology to the nucleic acid sequence of SEQ ID NO: 1, or for a CsgF homolog, SEQ ID NO: 4, is not limited purely to sequence identity. Many nucleic acid sequences can exhibit biologically significant homology to each other even though they clearly have low sequence identity. Homologous nucleic acid sequences are thought to hybridize to each other under low stringency conditions (M.R. Green, J. Sambrook, 2012, Molecular Cloning: A Laboratory Manual, Fourth Edition, Books 1-3 , Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY). The term "wild-type" refers to a gene or gene product isolated from a natural source of origin. A wild-type gene is the one most frequently observed in a population and is thus the "standard" or "wild-type" form of an arbitrarily designed gene. In contrast, the terms "modified," "mutant," or "variant" refer to a gene or gene product that shows a modification of the sequence (e.g., substitution, deletion, or insertion), post-translational modification, and / or functional characteristics (e.g., altered characteristics) when compared to the wild-type gene or gene product. A natural mutant

[0046] The term "wild-type" refers to a gene or gene product isolated from a natural source of origin. Wild-type genes are those most frequently observed in a population and are thus the "standard" or "wild-type" form of an arbitrarily designed gene. In contrast, the terms "modified," "mutant," or "variant" refer to a gene or gene product that, when compared to a wild-type gene or gene product, shows a modification of the sequence (e.g., substitution, deletion, or insertion), post-translational modification, and / or functional characteristics (e.g., altered characteristics). A natural mutant ​​​​​​​can be isolated, and these are identified by the fact that their properties have changed when compared to the wild-type gene or gene product. Methods for introducing or substituting natural amino acids are well known in the art. For example, methionine (M) , can be substituted with arginine (R) by substituting the codon (ATG) of methionine with the codon (CGT) of arginine at the relevant position of the polynucleotide encoding the mutant monomer. Methods for introducing or substituting unnatural amino acids are also well known in the art . For example, unnatural amino acids can be introduced by including synthetic aminoacyl-TrNA in the IVTT system used to express mutant monomers . Alternatively, unnatural amino acids may be introduced by expressing mutant monomers in E. coli that are auxotrophic for specific amino acids in the presence of their synthetic (i.e., unnatural-type) analogs of those specific amino acids . When mutant monomers are generated using solid-phase peptide synthesis, they can be generated by blunt ligation . Conservative substitutions replace an amino acid with another amino acid having a similar chemical structure, similar chemical properties, or similar side-chain volume. The introduced amino acid may have a polarity, hydrophilicity, hydrophobicity, basicity, acidity, neutrality, or charge similar to the amino acid being substituted . Alternatively, conservative substitutions may introduce another aromatic or aliphatic amino acid in place of an existing aromatic or aliphatic amino acid . Conservative amino acid changes are well known in the art and can be selected according to the properties of the 20 major amino acids defined in Table 1 below. When amino acids have similar polarities, this is based on the hydropathy scale for amino acid side chains in Table 2 ​​​​​​​​​​​ It can also be determined by referring to Rule. [Table 1] [Table 2]

[0047] The mutant or modified protein, monomer or peptide may be chemically modified in any way and at any site. The mutant or modified monomer or peptide may be attached to one or more cysteines (cysteine linkage), attached to one or more lysines , attached to one or more non-natural amino acids, enzymatically modified epitopes, or terminal modifications and is preferably chemically modified. Suitable methods for performing such modifications are well known in the art. Mutants of the modified protein, monomer or peptide may be chemically modified by the attachment of any molecule. For example, mutants of the modified protein, monomer or peptide may be chemically modified by the attachment of a dye or fluorophore molecule. In some embodiments, the mutant or modified monomer or peptide is chemically modified by a molecular adapter that facilitates the interaction between the pore containing the monomer or peptide and the target nucleotide or target polynucleotide sequence. The molecular adapter is a cyclic molecule, cyclodextrin, hybridizable species, DNA binder or intercalator, peptide or peptide analog, synthetic polymer, aromatic planar molecule, positively charged molecule, or or a small molecule having the ability to form hydrogen bonds. The presence of the adapter is the host of the pore and the nucleotide sequence or polynucleotide sequence

[0048] ​ Improve guest chemistry and thereby the ability to sequence pores formed from mutant monomers Improve the force. The principles of host-guest chemistry are well known in the art. The adapter is nuc It has an effect on the physical or chemical properties of the pore that improve its interaction with the leotide or polynucleotide sequence. The adapter changes the charge of the pore barrel or channel Or specifically interacts with or binds to a nucleotide sequence or polynucleotide sequence, thereby promoting its interaction with the pore. Thus, the modified CsgF peptide provided in the present disclosure may be bound to an enzyme or protein that is sufficiently proximate to the pore of said protein or enzyme, which may promote a particular use of the pore complex comprising the modified CsgF peptide Or specifically interacts with or binds to a nucleotide sequence or polynucleotide sequence, thereby promoting its interaction with the pore. Thus, the modified CsgF peptide provided in the present disclosure may be bound to an enzyme or protein that is sufficiently proximate to the pore of said protein or enzyme, which may promote a particular use of the pore complex comprising the modified CsgF peptide Or specifically interacts with or binds to a nucleotide sequence or polynucleotide sequence, thereby promoting its interaction with the pore. Thus, the modified CsgF peptide provided in the present disclosure may be bound to an enzyme or protein that is sufficiently proximate to the pore of said protein or enzyme, which may promote a particular use of the pore complex comprising the modified CsgF peptide Or specifically interacts with or binds to a nucleotide sequence or polynucleotide sequence, thereby promoting its interaction with the pore. Thus, the modified CsgF peptide provided in the present disclosure may be bound to an enzyme or protein that is sufficiently proximate to the pore of said protein or enzyme, which may promote a particular use of the pore complex comprising the modified CsgF peptide Or specifically interacts with or binds to a nucleotide sequence or polynucleotide sequence, thereby promoting its interaction with the pore. Thus, the modified CsgF peptide provided in the present disclosure may be bound to an enzyme or protein that is sufficiently proximate to the pore of said protein or enzyme, which may promote a particular use of the pore complex comprising the modified CsgF peptide Or specifically interacts with or binds to a nucleotide sequence or polynucleotide sequence, thereby promoting its interaction with the pore. Thus, the modified CsgF peptide provided in the present disclosure may be bound to an enzyme or protein that is sufficiently proximate to the pore of said protein or enzyme, which may promote a particular use of the pore complex comprising the modified CsgF peptide

[0049] In this context, the protein may also be, for example, a fusion protein (particularly called a gene fusion) produced by recombinant DNA technology The protein may also be conjugated or "conjugated", which, as used herein, means a chemical and / or enzymatic conjugation that results in a stable covalent bond The protein may also be conjugated or "conjugated", which, as used herein, means a chemical and / or enzymatic conjugation that results in a stable covalent bond The protein may also be conjugated or "conjugated", which, as used herein, means a chemical and / or enzymatic conjugation that results in a stable covalent bond

[0050] When several polypeptides or protein monomers bind or interact with each other, a protein complex may be formed. "Binding" refers to any interaction, whether direct or indirect. Direct interaction means contact between the binding partners, for example, contact by covalent bond or binding. Indirect interaction means any interaction in which the interaction partners interact within a complex of two or more compounds. The interaction may be through one or more cross-linking moieties When several polypeptides or protein monomers bind or interact with each other, a protein complex may be formed. "Binding" refers to any interaction, whether direct or indirect. Direct interaction means contact between the binding partners, for example, contact by covalent bond or binding. Indirect interaction means any interaction in which the interaction partners interact within a complex of two or more compounds. The interaction may be through one or more cross-linking moieties When several polypeptides or protein monomers bind or interact with each other, a protein complex may be formed. "Binding" refers to any interaction, whether direct or indirect. Direct interaction means contact between the binding partners, for example, contact by covalent bond or binding. Indirect interaction means any interaction in which the interaction partners interact within a complex of two or more compounds. The interaction may be through one or more cross-linking moieties When several polypeptides or protein monomers bind or interact with each other, a protein complex may be formed. "Binding" refers to any interaction, whether direct or indirect. Direct interaction means contact between the binding partners, for example, contact by covalent bond or binding. Indirect interaction means any interaction in which the interaction partners interact within a complex of two or more compounds. The interaction may be through one or more cross-linking moieties When several polypeptides or protein monomers bind or interact with each other, a protein complex may be formed. "Binding" refers to any interaction, whether direct or indirect. Direct interaction means contact between the binding partners, for example, contact by covalent bond or binding. Indirect interaction means any interaction in which the interaction partners interact within a complex of two or more compounds. The interaction may be through one or more cross-linking moieties Completely indirect, even with the assistance of a child, or parts where there is still direct contact between the parties It may be indirectly indirect in nature, in which case it is stabilized by additional interactions of one or more compounds. The "complex" referred to in this disclosure is defined as a group of two or more related proteins that may have different functions. The relationship between different polypeptides of a protein complex may be through non-covalent interactions such as hydrophobic or ionic forces, or through covalent bonds or linkages such as disulfide bridges or peptide bonds. The covalent "binding" or "coupling" is used interchangeably herein and each means a bioconjugation between cysteines or (photo)reactive amino acids, and a "cysteine bond" or "reactive or photo-reactive amino acid bond", which is a chemical covalent bond that forms a stable complex, may be involved. Examples of photo-reactive amino acids include azidohomoalanine, homopropargylglycine, homoallylglycine, p-acetylPhe, p-azido-Phe, p-propargyloxy-Phe, and p-benzoyl-Phe (Wang et al 2012, in Protein Engineering, DOI: 10.5772 / 28719, Chin et al. 2002, Proc. Nat. Acad. Sci. USA 99(17), 11020-24). A "biological pore" is a transmembrane protein structure that defines a channel or pore that allows the translocation of molecules and ions from one side of the membrane to the other. The translocation of ionic species through the pore can be driven by a potential difference applied across both sides of the pore. A "nanopore" is a molecule through which molecules or ions can pass through a channel or pore formed by a transmembrane protein and is typically on the order of nanometers in size or smaller . (Wang et al 2012, in Protein Engineering, DOI: 10.5772 / 28719, Chin et al. 2002, Proc. Nat. Acad. Sci. USA 99(17), 11020-24 )

[0051] A "biological pore" is a transmembrane protein structure that defines a channel or pore that allows the translocation of molecules and ions from one side of the membrane to the other. The translocation of ionic species through the pore can be driven by a potential difference applied across both sides of the pore. A "nanopore" is a molecule ​​​​​or a biological pore having a minimum diameter of a channel through which ions pass of nanometers (10 -9 meters). In some embodiments, the biological pore is a transmembrane protein pore and can be. The transmembrane protein structure of the biological pore may be essentially monomeric or oligomeric . Generally, the pore includes a plurality of polypeptide subunits arranged around a central axis, thereby forming a channel lined with a protein that extends substantially perpendicular to the membrane in which the nanopore is present . The number of polypeptide subunits is not limited . Generally, the number of subunits ranges from 5 to a maximum of 30, and a number of subunits of 6 to 10 is suitable . Alternatively, the number of subunits is not defined in the case of perfringolysin or related large membrane pores . The portion of the protein subunits within the nanopore that forms the protein-lined channel generally includes secondary structure motifs that can include one or more transmembrane β-barrels and / or α-helical sections . The terms "pore", "pore complex", or "complex pore", used interchangeably herein , refer to oligomeric pores, e.g., at least one CsgG monomer (e.g , two or more CsgG monomers, three or more CsgG monomers, etc., one or more Cs gG monomers) or a CsgG pore (including CsgG monomers), and a CsgF peptide

[0052] (e.g., a modified or cleaved CsgF peptide) are associated in a complex and together form a pore or a nanopore. The pore complexes of the present disclosure have the characteristics of biological pores, i.e , they have a typical transmembrane protein structure. When the pore complex is a membrane component, a membrane, a cell, or an absolute insulator . The pore complexes of the present disclosure have the characteristics of biological pores, i.e., they have a typical transmembrane protein structure. When the pore complex is a membrane component, a membrane, a cell, or an absolute . The pore complexes of the present disclosure have the characteristics of biological pores, i.e., they have a typical transmembrane protein structure. When the pore complex is a membrane component, a membrane, a cell, or an absolute . The pore complex has a typical transmembrane protein structure. When the pore complex is a membrane component, a membrane, a cell, or an absolute When provided in an environment with an edge layer, the pore complex is inserted into a membrane or an insulating layer to form a "membrane-penetrating pore complex".

[0053] The pore complex or membrane-penetrating pore complex of the present disclosure is suitable for characterizing analytes. In some embodiments, the described pore complex or membrane-penetrating complex can, for example, differentiate between different nuc leotides with high sensitivity and can be used for sequencing polynucleotide sequences. The pore complex of the present disclosure can be a substantially isolated, purified, or substantially purified, isolated pore complex. The pore complex of the present disclosure is usually free of lipids or other pores or other proteins (e.g., CsgE, CsgA, CsgB) that are usually associated with its natural state, or is sufficiently concentrated from membrane compartments when it is "isolated" or purified. The pore complex is substantially isolated when mixed with a carrier or diluent that does not interfere with its intended use. For example, the pore complex is present in a form containing less than 10%, less than 5%, less than 2%, or less than 1% of other components (such as triblock copolymers, lipids, or other pores), and is substantially isolated or substantially purified. Alternatively, the pore complex of the present disclosure can be a membrane-penetrating pore complex when present in a membrane. The present disclosure provides an isolated pore complex containing a homo-oligomeric pore derived from CsgG containing the same mutant monomer, which can also include a mutant form of the CsgG monomer as its homolog. Alternatively, an isolated pore complex containing a hetero-oligomeric CsgG pore is provided, which consists of mutant and wild-type CsgG monomers, or different ... ... ... ... ... ... ... It may also be a CsgG pore consisting of a CsgG variant, mutant, or homolog of the described form. The isolated pore complex typically contains at least 7, at least 8, at least 9, or at least 10 CsgG monomers, and also one or more (modified) CsgF pep tides such as 2, 3, 4, 5, 6, 7, 8, 9, 10 CsgF peptides. The pore complex can have any CsgG monomer:CsgF peptide ratio. In one embodiment, the CsgG monomer:CsgF peptide ratio is 1:1.

[0054] As used interchangeably herein, "constriction", "orifice", "constriction region", "channel constriction", or "constriction site" means an opening defined by the inner lumen surface of a pore or pore complex that permits the passage of ions and target molecules (e.g., polynucleotides or individual nucleotides, including but not limited to these), but does not permit the passage of other non-target molecules through the pore complex channel. In some embodiments, the constriction is the narrowest opening within the pore or pore complex. In this embodiment, the constriction can function to restrict the passage of molecules through the pore. The size of the constriction is typically a major factor in determining the suitability of the nanopore for nucleic acid sequencing applications. If the constriction is too small, the molecule to be sequenced cannot pass through. However, the constriction should not be too large in order to achieve the maximum effect on the ion current through the channel. For example, the constriction should not be wider than the transverse diameter accessible to the solvent of the target analyte. Ideally, any constriction should be as close as possible to the transverse diameter of the analyte passing through the analyte. For nucleic acid and nucleobase sequencing, an appropriate constriction diameter is in the nanometer range (10 -9 m)-9 within the meter range). Suitably , the diameter should be in the range of 0.5 - 2.0 nm, and typically, the diameter is 0.7 - 1 .2 nm. The constriction of wild-type E. coli CsgG has a diameter of about 9 Å (0.9 nm). The CsgF constriction formed within a pore complex containing a CsgG-like pore and a modified CsgF peptide or its homolog or mutant has a diameter in the range of 0.5 - 2 nm or in the range of 0.7 - 1.2 nm and is thus suitable for nucleic acid sequencing .

[0055] When two or more constrictions are present and each constriction is spaced apart, each constriction may interact with or "read" individual nucleotides within the nucleic acid strand simultaneously . In this situation, the reduction of ion flow through the channel is the result of the combined restriction of the flow through all constrictions containing nucleotides . Thus, in some examples, a double constriction can result in a complex current signal. In certain situations, the current readout for one constriction, or "reading head" , may not be determinable separately when two such reading heads are present. The constriction of wild-type E. coli CsgG (SEQ ID NO: 3) is formed by the juxtaposition of a series of tyrosine residues at position 51 (Tyr51) within adjacent protein monomers, and also by phenylalanine and asparagine residues (Phe 56 and Asn 55) at positions 56 and 55 respectively, and contains two circular rings (Figure 1). The wild-type pore structure of CsgG mostly consists of two circular rings that constitute the CsgG constriction (referred to herein as the "CsgG channel constriction") . . ​​​​Recombined through recombinant gene technology such as widening, changing, or removing the width of the ring, and then operated to leave a single distinct leading head. The constriction motif within the CsgG oligomeric pore is located at amino acid residues 38 - 63 in the wild-type monomeric E. coli CsgG polypeptide shown in SEQ ID NO: 3. Considering this region, mutants at any amino acid residue positions 50 - 53, 54 - 56, and 58 - 59, and the positioning of the side chains of Tyr51, Asn55, and Phe56 within the channel of the wild-type C sgG structure have been shown to be advantageous for modifying or changing the properties of the leading head. The present disclosure regarding the CsgG pore and the pore complex containing the modified CsgF peptide or its homologs or mutants surprisingly involves adding a separate constriction (referred to herein as the "CsgF channel constriction") to the CsgG-containing pore complex, and forming an appropriate additional second leader head within the pore through complex formation with the modified Csg F peptide. The additional CsgF channel constriction or leader head is positioned adjacent to the constriction loop of the CsgG pore, or the constriction loop of the mutated GcsG pore and is positioned at a location of about 10 nm or less, such as 1, 2, 3, 4, 5, 6, 7, 8, 9 nm, etc., 5 nm or less from the constriction loop of the CsgG pore, or from the constriction loop of the mutated GcsG pore The pore complex or transmembrane pore complex of the present disclosure includes a pore complex with two leader heads, that is, the channel constriction positioned in such a way provides an appropriate separate leader head without interfering with the accuracy of other constriction channel leader heads ​​​​​​​ Thus, the pore complex is a CsgG mutant pore (the incorporated reference WO2 016 / 034591, WO2017 / 149316, WO2017 / 149317, W O2017 / 149318, International Patent Application No. PCT / GB2018 / 051191 are cited, and mutants of wild-type CsgG pores that improve the characteristics of the pores are listed respectively ), and may include a wild-type CsgG pore or its homolog together with a modified CsgF peptide or its homolog or mutant, and the CsgF peptide has another constriction channel that forms a leader head ). The present invention relates, unexpectedly, to a CsgG pore that forms a complex with an extracellularly located CsgF peptide that introduces an additional channel constriction or leader head into the pore complex. Further, the present disclosure provides positional information of the constriction produced by the CsgF peptide within the pore complex, the peptide is inserted into the lumen of the CsgG pore, and the constriction site is within the N-terminal portion of the CsgF protein. Furthermore, the modified or truncated CsgF peptides of the present disclosure have been shown to be sufficient for pore complex formation and provide means and methods for biosensing applications. The present disclosure combines wild-type and mutant CsgG pores (e.g., disclosed in WO2016 / 034591, WO2017 / 149316, WO2017 / 149317, WO20 17 / 149318 and International Patent Application Publication No. PCT / GB2018 / 051191) or their homologs or mutants, together with modified or truncated CsgF

[0056] Pore peptides and their mutants or homologs, collectively, CsgG-like pores ). Furthermore, the present disclosure provides positional information of the constriction produced by the CsgF peptide within the pore complex, the peptide is inserted into the lumen of the CsgG pore, and the constriction site is within the N-terminal portion of the CsgF protein. In addition, the modified or truncated CsgF peptides of the present disclosure have been shown to be sufficient for pore complex formation and provide means and methods for biosensing applications. The present disclosure combines wild-type and mutant CsgG pores (e.g., disclosed in WO2016 / 034591, WO2017 / 149316, WO2017 / 149317, WO20 17 / 149318 and International Patent Application Publication No. PCT / GB2018 / 051191) or their homologs or mutants, together with modified or truncated CsgF peptides and their mutants or homologs, collectively, CsgG-like pores ). The present disclosure relates to wild-type and mutant CsgG pores (e.g., WO2016 / 034591, WO2017 / 149316, WO2017 / 149317, WO20 17 / 149318 and International Patent Application Publication No. PCT / GB2018 / 051191 disclose) or their homologs or mutants, which, in combination with modified or truncated CsgF peptides and their mutants or homologs, collectively, CsgG-like pores ). including those that improve the ability of the complex to interact with an analyte (such as a polynucleotide). CsgG-like nanopores are formed by complex formation using a modified or cleaved CsgF peptide Additional constrictions introduced into the pore channel expand the contact surface with the passing analyte and can serve as a second leader head for analyte detection and characterization. Csg A pore containing a mutant CsgG monomer combined with a novel mutant or modified form of CsgF can improve the characterization of analytes such as polynucleotides, providing a more differentiated direct relationship between the currents observed when the polynucleotide moves through the pore. In particular, with two leader heads stacked at a defined distance with a gap in between, the CsgG:CsgF pore complex can facilitate the characterization of polynucleotides containing at least one homopolymeric region (e.g., several consecutive copies of the same nucleotide that exceed the length of the interaction of a single CsgG leader head). Furthermore, by having two stacked constrictions at a defined distance, small molecule analytes, including organic or inorganic drugs and contaminants, passing through the CsgG:CsgF complex pore pass continuously through two independent leader heads. The chemical properties of either leader head can be independently modified, each given unique interaction characteristics with the analyte, thus providing additional discriminatory power during analyte detection. F, including isolated pore complexes. Indeed, the present disclosure relates to a modified CsgF peptide (the including those that improve the ability of the complex to interact with an analyte (such as a polynucleotide). polynucleotide) moving through the pore. In particular, with two leader heads stacked at a defined distance with a gap in between, the CsgG:CsgF pore complex can facilitate the characterization of polynucleotides containing at least one homopolymeric region (e.g., several consecutive copies of the same nucleotide that exceed the length of the interaction of a single CsgG leader head). Furthermore, by having two stacked constrictions at a defined distance, small molecule analytes, including organic or inorganic drugs and contaminants, passing through the CsgG:CsgF complex pore pass continuously through two independent leader heads. The chemical properties of either leader head can be independently modified, each given unique interaction characteristics with the analyte, thus providing additional discriminatory power during analyte detection. In a first aspect, the invention relates to an isolated pore complex comprising a CsgG pore or a homolog or mutant thereof, or a CsgG-like pore, and a modified CsgF peptide or a homolog or mutant thereof. In fact, the present disclosure relates to a modified CsgF peptide (the In a first aspect, the invention relates to an isolated pore complex comprising a CsgG pore or a homolog or mutant thereof, or a CsgG-like pore, and a modified CsgF peptide or a homolog or mutant thereof. In fact, the present disclosure relates to a modified CsgF peptide (the same nucleotide that exceed the length of the interaction of a single CsgG leader head). Furthermore, by having two stacked constrictions at a defined distance, small molecule analytes, including organic or inorganic drugs and contaminants, passing through the CsgG:CsgF complex pore pass continuously through two independent leader heads. The chemical properties of either leader head can be independently modified, each given unique interaction characteristics with the analyte, thus providing additional discriminatory power during analyte detection. same nucleotide that exceed the length of the interaction of a single CsgG leader head). Furthermore, by having two stacked constrictions at a defined distance, small molecule analytes, including organic or inorganic drugs and contaminants, passing through the CsgG:CsgF complex pore pass continuously through two independent leader heads. The chemical properties of either leader head can be independently modified, each given unique interaction characteristics with the analyte, thus providing additional discriminatory power during analyte detection. In a first aspect, the invention relates to an isolated pore complex comprising a CsgG pore or a homolog or mutant thereof, or a CsgG-like pore, and a modified CsgF peptide or a homolog or mutant thereof. In fact, the present disclosure relates to a modified CsgF peptide (the same nucleotide that exceed the length of the interaction of a single CsgG leader head). Furthermore, by having two stacked constrictions at a defined distance, small molecule analytes, including organic or inorganic drugs and contaminants, passing through the CsgG:CsgF complex pore pass continuously through two independent leader heads. The chemical properties of either leader head can be independently modified, each given unique interaction characteristics with the analyte, thus providing additional discriminatory power during analyte detection. same nucleotide that exceed the length of the interaction of a single CsgG leader head). Furthermore, by having two stacked constrictions at a defined distance, small molecule analytes, including organic or inorganic drugs and contaminants, passing through the CsgG:CsgF complex pore pass continuously through two independent leader heads. The chemical properties of either leader head can be independently modified, each given unique interaction characteristics with the analyte, thus providing additional discriminatory power during analyte detection. In a first aspect, the invention relates to an isolated pore complex comprising a CsgG pore or a homolog or mutant thereof, or a CsgG-like pore, and a modified CsgF peptide or a homolog or mutant thereof. In fact, the present disclosure relates to a modified CsgF peptide (the In a first aspect, the invention relates to an isolated pore complex comprising a CsgG pore or a homolog or mutant thereof, or a CsgG-like pore, and a modified CsgF peptide or a homolog or mutant thereof. In fact, the present disclosure relates to a modified CsgF peptide (the

[0057] In a first aspect, the invention relates to an isolated pore complex comprising a CsgG pore or a homolog or mutant thereof, or a CsgG-like pore, and a modified CsgF peptide or a homolog or mutant thereof. In fact, the present disclosure relates to a modified CsgF peptide (the In fact, the present disclosure relates to a modified CsgF peptide (the Relates to modified CsgG biological pores (which may be truncated, mutants and / or variants). In one embodiment, the interaction region between the modified CsgF peptide or its homolog or mutant is located within the lumen of the CsgG pore or its homolog or mutant. In another embodiment, the pore complex is provided by at least one constriction of the CsgG pore, and at least one is introduced by the CsgF peptide, and has two or more constriction sites or leader heads that form a complex with the CsgG pore. The N-terminal CsgG positions, including the range of amino acid residues 39-64 of SEQ ID NO: 5, or more specifically the range of amino acid residues 49-64 of SEQ ID NO: 5, have been shown to allow for a detectable amount of stable CsgG:CsgG complex. In one embodiment, the CsgF constriction generated by the modified CsgF peptide (e.g., as described herein) is adjacent to or proximal to the first constriction of the CsgG pore of the pore complex. In the case of CsgG or CsgG-like protein pores, the constriction site has been determined to be formed by a loop region of the beta strand (see Figure 1). In one embodiment, the modified CsgF peptide or its homolog or mutant The interaction region between mutants is located within the lumen of the CsgG pore or its homolog or mutant. In another embodiment, the pore complex is provided by at least one constriction of the CsgG pore, and at least one is introduced by the CsgF peptide, and has two or more constriction sites or leader heads that form a complex with the CsgG pore. The N-terminal CsgG positions, including the range of amino acid residues 39-64 of SEQ ID NO: 5, or more specifically the range of amino acid residues 49-64 of SEQ ID NO: 5, have been shown to allow for a detectable amount of stable CsgG:CsgG complex. In one embodiment, the CsgF constriction generated by the modified CsgF peptide (e.g., as described herein) is adjacent to or proximal to the first constriction of the CsgG pore of the pore complex. In the case of CsgG or CsgG-like protein pores, the constriction site has been determined to be formed by a loop region of the beta strand (see Figure 1). In another embodiment, the pore complex is provided by at least one constriction of the CsgG pore, and at least one is introduced by the CsgF peptide, and has two or more constriction sites or leader heads that form a complex with the CsgG pore. The N-terminal CsgG positions, including the range of amino acid residues 39-64 of SEQ ID NO: 5, or more specifically the range of amino acid residues 49-64 of SEQ ID NO: 5, have been shown to allow for a detectable amount of stable CsgG:CsgG complex. In one embodiment, the CsgF constriction generated by the modified CsgF peptide (e.g., as described herein) is adjacent to or proximal to the first constriction of the CsgG pore of the pore complex. In the case of CsgG or CsgG-like protein pores, the constriction site has been determined to be formed by a loop region of the beta strand (see Figure 1). In another embodiment, the pore complex is provided by at least one constriction of the CsgG pore, and at least one is introduced by the CsgF peptide, and has two or more constriction sites or leader heads that form a complex with the CsgG pore. The N-terminal CsgG positions, including the range of amino acid residues 39-64 of SEQ ID NO: 5, or more specifically the range of amino acid residues 49-64 of SEQ ID NO: 5, have been shown to allow for a detectable amount of stable CsgG:CsgG complex. In one embodiment, the CsgF constriction generated by the modified CsgF peptide (e.g., as described herein) is adjacent to or proximal to the first constriction of the CsgG pore of the pore complex. In the case of CsgG or CsgG-like protein pores, the constriction site has been determined to be formed by a loop region of the beta strand (see Figure 1). In another embodiment, the pore complex is provided by at least one constriction of the CsgG pore, and at least one is introduced by the CsgF peptide, and has two or more constriction sites or leader heads that form a complex with the CsgG pore. The N-terminal CsgG positions, including the range of amino acid residues 39-64 of SEQ ID NO: 5, or more specifically the range of amino acid residues 49-64 of SEQ ID NO: 5, have been shown to allow for a detectable amount of stable CsgG:CsgG complex. In one embodiment, the CsgF constriction generated by the modified CsgF peptide (e.g., as described herein) is adjacent to or proximal to the first constriction of the CsgG pore of the pore complex. In the case of CsgG or CsgG-like protein pores, the constriction site has been determined to be formed by a loop region of the beta strand (see Figure 1). The range of amino acid residues 39-64 of SEQ ID NO: 5, or more specifically the range of amino acid residues 49-64 of SEQ ID NO: 5, including positions within the range, have been shown to allow for a detectable amount of stable CsgG:CsgG complex. In one embodiment, the CsgF constriction generated by the modified CsgF peptide (e.g., as described herein) is adjacent to or proximal to the first constriction of the CsgG pore of the pore complex. In the case of CsgG or CsgG-like protein pores, the constriction site has been determined to be formed by a loop region of the beta strand (see Figure 1). The range of amino acid residues 39-64 of SEQ ID NO: 5, or more specifically the range of amino acid residues 49-64 of SEQ ID NO: 5, including positions within the range, have been shown to allow for a detectable amount of stable CsgG:CsgG complex. In one embodiment, the CsgF constriction generated by the modified CsgF peptide (e.g., as described herein) is adjacent to or proximal to the first constriction of the CsgG pore of the pore complex. In the case of CsgG or CsgG-like protein pores, the constriction site has been determined to be formed by a loop region of the beta strand (see Figure 1). The range of amino acid residues 39-64 of SEQ ID NO: 5, or more specifically the range of amino acid residues 49-64 of SEQ ID NO: 5, including positions within the range, have been shown to allow for a detectable amount of stable CsgG:CsgG complex. In one embodiment, the CsgF constriction generated by the modified CsgF peptide (e.g., as described herein) is adjacent to or proximal to the first constriction of the CsgG pore of the pore complex. In the case of CsgG or CsgG-like protein pores, the constriction site has been determined to be formed by a loop region of the beta strand (see Figure 1). In one embodiment, the CsgF constriction generated by the modified CsgF peptide (e.g., as described herein) is adjacent to or proximal to the first constriction of the CsgG pore of the pore complex. In the case of CsgG or CsgG-like protein pores, the constriction site has been determined to be formed by a loop region of the beta strand (see Figure 1). In one embodiment, the CsgF constriction generated by the modified CsgF peptide (e.g., as described herein) is adjacent to or proximal to the first constriction of the CsgG pore of the pore complex. In the case of CsgG or CsgG-like protein pores, the constriction site has been determined to be formed by a loop region of the beta strand (see Figure 1). In the case of CsgG or CsgG-like protein pores, the constriction site has been determined to be formed by a loop region of the beta strand (see Figure 1). In the case of CsgG or CsgG-like protein pores, the constriction site has been determined to be formed by a loop region of the beta strand (see Figure 1).

[0058] In one embodiment, the modified CsgF peptide is a peptide that means a truncated CsgF protein or fragment, including an N-terminal CsgF peptide fragment defined by the modification, which particularly includes a constriction region and is restricted by binding to a CsgG monomer or its homolog or mutant. The modified CsgF peptide may further include a mutation or homologous sequence that can promote specific properties of the pore complex. In certain embodiments, the modified CsgF peptide In one embodiment, the modified CsgF peptide is a peptide that means a truncated CsgF protein or fragment, including an N-terminal CsgF peptide fragment defined by the modification, which particularly includes a constriction region and is restricted by binding to a CsgG monomer or its homolog or mutant. The modified CsgF peptide may further include a mutation or homologous sequence that can promote specific properties of the pore complex. In certain embodiments, the modified CsgF peptide In one embodiment, the modified CsgF peptide is a peptide that means a truncated CsgF protein or fragment, including an N-terminal CsgF peptide fragment defined by the modification, which particularly includes a constriction region and is restricted by binding to a CsgG monomer or its homolog or mutant. The modified CsgF peptide may further include a mutation or homologous sequence that can promote specific properties of the pore complex. In certain embodiments, the modified CsgF peptide In one embodiment, the modified CsgF peptide is a peptide that means a truncated CsgF protein or fragment, including an N-terminal CsgF peptide fragment defined by the modification, which particularly includes a constriction region and is restricted by binding to a CsgG monomer or its homolog or mutant. The modified CsgF peptide may further include a mutation or homologous sequence that can promote specific properties of the pore complex. In certain embodiments, the modified CsgF peptide In one embodiment, the modified CsgF peptide is a peptide that means a truncated CsgF protein or fragment, including an N-terminal CsgF peptide fragment defined by the modification, which particularly includes a constriction region and is restricted by binding to a CsgG monomer or its homolog or mutant. The modified CsgF peptide may further include a mutation or homologous sequence that can promote specific properties of the pore complex. In certain embodiments, the modified CsgF peptide , the wild-type preprotein (SEQ ID NO:5) or mature protein (SEQ ID NO:6) sequence or These modified peptides contain truncated CsgF proteins compared to their homologs. Within the CsgG-like pore formed by G and modified or truncated CsgF peptides , which act as components of the pore complex introducing additional constriction sites or leader heads. Examples of truncated modified peptides are provided below.

[0059] Examples of homologues of modified CsgF peptides are, for example, those determined in Example 3 and shown to be compatible with different bacterial strains. We have identified CsgF-like proteins or CsgF peptides that contain the same or similar constrictor domain. However, this may be useful in the use of similar pore complexes. Various CsgG pores were developed for use in combination with wild-type or mutant CsgG pores. Structural features and CsgG-binding elements are conserved in CsgF peptides derived from CsgG homologues This includes the CsgG pore in complex with the non-cognate CsgF, which The parent CsgG homologues from which CsgG and CsgF are derived originate from the same operon, bacterial species or strain. It means there is no need.

[0060] In an alternative embodiment, the CsgG pore in the pore complex is not a wild-type pore, but has a pore-specific The CsgG pore or its homologues may also contain mutations or modifications to increase the binding affinity of the CsgG pore. and a modified CsgF peptide or a homolog thereof. The pore complex can be formed by the wild-type form of the CsgG pore or by specific amines. Further modifications have been made within the CsgG pore, such as by directed mutagenesis of acid residues. , the desirable properties of the CsgG pore for use within the pore complex are further improved. For example in an embodiment of the present invention, the mutation is intended to change the number, size, shape, arrangement or orientation of constrictions within the channel. A pore complex comprising a modified mutant CsgG pore may be prepared by known genetic engineering techniques that result in the insertion, substitution, and / or deletion of specific target amino acid residues within the polypeptide sequence. In the case of an oligomeric CsgG pore , the mutation may be made within the polypeptide subunit of each monomer, or within any one or all of the monomers. In one embodiment of the present invention, the mutations described are suitably made to all monomer polypeptides within the oligomeric protein structure. A mutant CsgG monomer is a monomer whose sequence is mutated from that of the wild-type CsgG monomer and which maintains the ability to form a pore. Methods for confirming the ability of a mutant monomer to form a pore are well known in the art. The present disclosure relates to wild-type and mutant CsgG pores (e.g., as disclosed in WO2016 / 034591, W O2017 / 149316, WO2017 / 149317, WO2017 / 149318 and International Patent Application Publication No. PCT / GB2018 / 051191) or homologs thereof which, in combination with a modified or truncated CsgF peptide and mutants or homologs thereof , collectively improve the ability of the CsgG-like pore complex to interact with an analyte (such as a polynucleotide ). A mutant CsgG pore may comprise one or more mutant monomers. The CsgG pore may be a homopolymer comprising the same monomers, or a heteropolymer comprising two or more different monomers. The monomers are optional and may be It may have one or more mutants described below in combination.

[0061] In certain embodiments, the nanopore complex comprising the modified CsgF peptide contains only the N-terminal fragment or cleavage of the wild-type CsgF protein, so it is different compared to the wild-type CsgF protein shown in SEQ ID NO: 6. However, the modified CsgF peptide can be additionally or alternatively a mutated CsgF peptide in the sense that it allows a better second constriction site as an amino acid substitution in the pore formed by the CsgG pore and the complex comprising the modified CsgF peptide. Therefore, in addition to the function of the complex being improved such that the mutant monomer contains two leader heads, when the complex is used for nucleotide sequencing, i.e., when improved polynucleotide capture and nucleotide discrimination are shown, it may be possible to improve the polynucleotide reading characteristics. In particular, pores composed of mutant peptides capture nucleotides and polynucleotides more easily than the wild-type. Further, pores composed of mutant peptides may show an increase in the current range, which facilitates discrimination between different nucleotides, reduces the variance of the states, and this increases the signal-to-noise ratio. Further, when the polynucleotide moves through the pore composed of the mutant, the number of nucleotides contributing to the current may decrease. This facilitates the identification of the direct relationship between the observed currents when the polynucleotide moves through the pore and the polynucleotide sequence. Further, pores composed of mutant peptides may show an increase in increased throughput, e.g., polynucleo tide ​​​​​​​​​​​​​It is highly likely to interact with analytes such as chide. This facilitates the characterization of analytes using pores. Pores composed of mutant peptides can be more easily inserted into the membrane or provide an easier way to retain additional proteins in the vicinity of the pore complex.

[0062] In another embodiment, the CsgF constriction site provided within the pore complex of the present invention has a diameter in the range of 0.5 nm to 2.0 nm, thereby providing a pore complex suitable for nucleic acid sequencing as described above.

[0063] The pore can be stabilized by covalent bonding of the CsgF peptide to the CsgG pore. The covalent bond can be, for example, a disulfide bond or click chemistry. The CsgF peptide and the CsgG pore can be covalently bonded through one or more residues at positions corresponding to pairs of positions of SEQ ID NO: 6 and SEQ ID NO: 3, such as 1 and 153, 4 and 133, 5 and 136, 8 and 187, 8 and 203, 9 and 203, 11 and 142, 11 and 201, 12 and 149, 12 and 203, 26 and 191, and 29 and 144.

[0064] In the pore, the interaction between the CsgF peptide and the CsgG pore can be stabilized by hydrophobic or electrostatic interactions at positions corresponding to one or more of the pairs of positions of SEQ ID NO: 6 and SEQ ID NO: 3, such as 1 and 153, 4 and 133, 5 and 136, 8 and 187, 8 and 203, 9 and 203, 11 and 142, 11 and 201, 12 and 149, 12 and 203, 26 and 191, and 29 and 144.

[0065] The residues of CsgF and / or CsgG at one or more of the above positions may be modified to enhance the interaction between Csg G and CsgF within the pore.

[0066] In one embodiment, the pores of the invention can be isolated, substantially isolated, purified , or substantially purified. The pores of the invention are isolated or purified if they contain no other components such as lipids or other pores. The pores are substantially isolated if they are mixed with a carrier or diluent that does not interfere with their intended use. For example, the pores are substantially isolated or substantially purified if they are present in a form that contains less than 10%, less than 5%, less than 2%, or less than 1% of other components (e.g., triblock copolymers, lipids, or other pores). Alternatively, the pores of the invention can be present within a membrane. Suitable membranes are discussed below.

[0067] The pores of the invention can exist individually or as single pores. Alternatively, the pores of the invention can exist in a homogeneous or heterogeneous population of two or more pores.

[0068] CsgF peptide A second aspect of the invention relates to novel modified CsgF monomers (peptides), or CsgF proteins, or modified or truncated peptides of CsgF homologs or mutants. These novel modified CsgF peptides may be used in a pore complex to incorporate a second or additional leader head. The modification or truncation preferably results in a fragment of wild-type CsgF, or a mutant or homolog of the CsgF protein, more preferably an N-terminal fragment.

[0069] Mature CsgF (shown in SEQ ID NO: 6) can be divided into three main regions such as the "CsgF contraction peptide" (FCP), the "neck" region and the "head" region (shown in FIGS. 4 and 5). The "head" region of the CsgF peptide is different from the pore leader head domains described herein. The "head" region of the CsgF peptide may also be referred to as the "C-terminal head domain".

[0070] FCP forms a contact region with the CsgG β-barrel, where an additional constriction occurs. The neck region protrudes from the β-barrel. In the CsgG:CsgF oligomer, this forms a thin-walled hollow tube connecting the FCP to the globular head region.

[0071] Based on the multiple sequence alignment (FIG. 8), co-purification experiments (FIG. 9), and cryoEM reconstruction of the CsgG:CsgF complex at 3.4 Å resolution (FIG. 11), the CsgF contraction peptide, the neck region and the head region can be defined as three consecutive residues within mature CsgF.

[0072] FCP spans approximately residues 1 to 35 of mature CsgF (SEQ ID NO: 6). FCP forms the most conserved region of the protein when comparing different CsgF orthologs (FIGS. 8, FIG. 10). CryoEM 3D reconstruction shows a well-defined structure of the FCP form that binds inside the CsgG β-barrel by non-covalent contacts with the CsgG transmembrane hairpins TM1 (residues 134 - 154 of SEQ ID NO: 3) and TM2 (residues 184 - 208 of SEQ ID NO: 3) (FIGS. 1E, 11) (TM1 and TM2 are defined by Goyal P et al., 201 4). In the reconstruction, 9 copies of FCP are present in the CsgG oligomer (9 monomers ). ​ -containing) are combined and together form a region about 2 nm higher than the upper part of the CsgG constriction formed by the continuous loop spanning residues 46 ~61 of mature CsgG (SEQ ID NO: 3, Figure 1E, Figure 11), resulting in an additional constriction.

[0073] 3D reconstruction by cryoEM also shows that the CsgF N-terminal residues bind near the bottom or top (depending on the orientation) of the CsgG β-barrel and exit the β-barrel at residue 32. This is in good agreement with MD simulations (Table 4) showing the average contact time of residue pairs at the CsgG:CsgF binding interface. The cryoEM structure and MD simulations show that residues 33 - 34 are located outside the CsgG β-barrel where a conserved Pro residue (Pro 35 of SEQ ID NO: 6) migrates to the CsgF neck region. The CsgF neck is not resolved at the atomic level in the 3D reconstruction of the CsgG:CsgF complex. Based on multiple sequence alignment and secondary structure prediction, the CsgF neck is predicted to span from residue 36 to approximately residue 50 (SEQ ID NO: 6). The CsgF head region forms the C-terminal part of CsgF and is predicted to span from approximately residue 51 to the CsgF C-terminus. In the CsgG:CsgF complex, this region forms a globular structure that appears to occlude the CsgG:CsgF channel and oligomerizes (Figures 4, 5). Multiple sequence alignment of CsgF orthologs shows that the CsgF neck is the least conserved region and suggests that the length can vary between different orthologs (Figure 8).

[0074] The CsgF peptide that forms part of the present invention lacks a C-terminal head, or lacks the C-terminal head and also part of the neck domain of CsgF (for example, a cleaved CsgF peptide may contain only a portion of the neck domain of CsgF), or is a cleaved CsgF peptide that lacks the C-terminal head and the neck domain of CsgF. The CsgF peptide may lack part of the CsgF neck domain. For example, the CsgF peptide may consist of, for example, amino acid residue 36 (see SEQ ID NO: 6) at the N-terminus of the neck domain (for example, residues 36-40, 36-41, 36-42, 36-43, 36-45, 36-46 of SEQ ID NO: 6, from residues 36-50 or 36-60), and may include a portion of the neck domain. The CsgF peptide preferably includes a CsgG binding region and a region that forms a constriction within the pore. The CsgG binding region typically includes residues 1-8, and / or residues 29-32, of the CsgF protein (SEQ ID NO: 6 or a homolog from another species), and may include one or more modifications. The region that forms a constriction within the pore typically includes residues 9-28 of the CsgF protein (SEQ ID NO: 6 or a homolog from another species), and may include one or more modifications. Residues 9-17 contain the conserved motif N9PXFGGXXX and form a rotation region. Residues 9-28 form an α helix. X (N17 of SEQ ID NO: 6 forms the apex of the constriction region that corresponds to the narrowest part of the CsgF constriction within the pore. The CsgF constriction region also contacts the CsgG β barrel, mainly residues 9, 11, 12, 18, 21 and 22 of SEQ ID NO: 6 to stabilize it). The region that forms a constriction within the pore typically includes residues 9-28 of the CsgF protein (SEQ ID NO: 6 or a homolog from another species), and may include one or more modifications. Residues 9-17 contain the conserved motif N9PXFGGXXX and form a rotation region. Residues 9-28 form an α helix. X conserved motif N9PXFGGXXX 17 to form a rotation region. Residues 9-2 8 form an α helix. X 17 (N17 of SEQ ID NO: 6 forms the apex of the constriction region of the CsgF constriction within the pore. The CsgF constriction region also contacts the CsgG β barrel, mainly residues 9, 11, 12, 18, 21 and 22 of SEQ ID NO: 6 to stabilize it. The CsgF constriction region also contacts the CsgG β barrel, mainly residues 9, 11, 12, 18, 21 and 22 of SEQ ID NO: 6 to stabilize it. to stabilize it.

[0075] The CsgF peptide typically has a length of 28 to 50 amino acids, such as 29 to 49, 30 to 45 or is 32 to 40 amino acids. The CsgF peptide preferably contains 29 to 35 amino acids, or 29 to 45 amino acids. The CsgF peptide contains all or part of the FCP corresponding to residues 1 to 35 of SEQ ID NO: 6. When the CsgF peptide is shorter than FCP, the cleavage is preferably made at the C-terminus. The CsgF fragment of SEQ ID NO: 6 or its homolog or mutant may have a length of 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 4

[0076] 0, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53 or 54 or 55 amino acids. The CsgF peptide may contain the amino acid sequence of SEQ ID NO: 6 from residue 1 to any one of residues 25 to 60 (e.g., 27 to 50 of SEQ ID NO: 6, such as 28 to 45), or the corresponding residues from a homolog of SEQ ID NO: 6, or any mutant thereof. More specifically the CsgF peptide may contain SEQ ID NO: 39 (residues 1 to 29 of SEQ ID NO: 6) or its homolog

[0077] or mutant. Examples of such CsgF peptides are SEQ ID NO: 15 (residues 1 to 34 of SEQ ID NO: 6), SEQ ID NO: 54 (residues 1 to 30 of SEQ ID NO: 6), SEQ ID NO: 40 (residues 1 to 45 of SEQ ID NO: 6), or SEQ ID NO: 55 (residues 1 to 35 of SEQ ID NO: 6) and any homologs or mutants thereof and are essentially composed of or composed of. Other examples of CsgF peptides are as follows

[0078] The CsgF peptide may contain the amino acid sequence of SEQ ID NO: 6 from residue 1 to any one of residues 25 to 60 (e.g., 27 to 50 of SEQ ID NO: 6, such as 28 to 45), or the corresponding residues from a homolog of SEQ ID NO: 6, or any mutant thereof. More specifically the CsgF peptide may contain SEQ ID NO: 39 (residues 1 to 29 of SEQ ID NO: 6) or its homolog or mutant. Examples of such CsgF peptides are SEQ ID NO: 15 (residues 1 to 34 of SEQ ID NO: 6), SEQ Comprising, consisting essentially of, or consisting of those with an array number: array Number 7, array number 8, array number 9, array number 10, array number 11, array number 12, array number 13, array number 14, array number 16.

[0079] In the CsgF peptide, one or more residues, such as array number 15, array number 39, ar ray number 40, array number 54, or array number 55 may be modified.

[0080] For example, the CsgF peptide may contain modifications at positions corresponding to on e or more of the following of SEQ ID NO: 6: G1, T4, F5, R8, N9, N11, F12, A26, and Q29.

[0081] The CsgF peptide may be modified to introduce cysteine, hydrophobic amino acids, charged amino acids, non-native reactive amino acids, or photoreactive amino acids at positions cor responding to one or more of the following of SEQ ID NO: 6: G1, T4, F5, R8, N9, N11, F1 2, A26, and Q29.

[0082] For example, the CsgF peptide may contain modifications at positions corresponding to on e or more of the following of SEQ ID NO: 6: N15, N17, A20, N24, and A28. The CsgF peptid e may contain modifications at the position corresponding to D34 to stabilize the CsgG-CsgF complex. In certain embodiments, the CsgF peptide contains one or more of the following substitutions: N15S / A / T / Q / G / L / V / I / F / Y / W / R / K / D / C, N17S / A / T / Q / G / L / V / I / F / Y / W / R / K / D / C, A20S / T / Q / N / G / L / V / I / F / Y / W / R / K / D / C, N24S / T / Q / A / G / L / V / I / F / Y / W / R / K / D / C, A28S / T / Q / N / G / L / V / I / F / Y / W / R / K / D / C and D34F / Y / W / R / K / N / Q / C. The CsgF peptide can, for example, include one or more of the following substitutions: G1C, T4C, N17S, and D34Y or D34N.

[0083] The CsgF peptide can be generated by cleaving a long protein such as full-length CsgF using an enzyme. Cleavage at specific sites can be directed by modifying a long protein such as full-length CsgF to include an enzyme cleavage site at an appropriate position. Examples of CsgF amino acid sequences modified to include such enzyme cleavage sites are shown in SEQ ID NOs: 56 - 67. After cleavage, all or part of the added enzyme cleavage site can be present within the CsgF peptide that associates with CsgG to form pores. Thus, the CsgF peptide may further include all or part of the enzyme cleavage site at its C-terminus. by modifying a long protein such as full-length CsgF to include an enzyme cleavage site at an appropriate position. Examples of such enzyme cleavage sites are shown in SEQ ID NOs: 56 - 67. After cleavage, all or part of the added enzyme cleavage site can be present within the CsgF peptide that associates with CsgG to form pores. Thus, the CsgF peptide may further include all or part of the enzyme cleavage site at its C-terminus. cleavage sites are shown in SEQ ID NOs: 56 - 67. After cleavage, all or part of the added enzyme cleavage site can be present within the CsgF peptide that associates with CsgG to form pores. Thus, the CsgF peptide may further include all or part of the enzyme cleavage site at its C-terminus. cleavage sites are shown in SEQ ID NOs: 56 - 67. After cleavage, all or part of the added enzyme cleavage site can be present within the CsgF peptide that associates with CsgG to form pores. Thus, the CsgF peptide may further include all or part of the enzyme cleavage site at its C-terminus. cleavage sites are shown in SEQ ID NOs: 56 - 67. After cleavage, all or part of the added enzyme cleavage site can be present within the CsgF peptide that associates with CsgG to form pores. Thus, the CsgF peptide may further include all or part of the enzyme cleavage site at its C-terminus. cleavage sites are shown in SEQ ID NOs: 56 - 67. After cleavage, all or part of the added enzyme cleavage site can be present within the CsgF peptide that associates with CsgG to form pores. Thus, the CsgF peptide may further include all or part of the enzyme cleavage site at its C-terminus.

[0084] Examples of some appropriate CsgF peptides are shown in Table 3 below.

Table 3

[0085] In certain embodiments, the CsgF fragment includes amino acid sequence number 39, or a mutant or homolog thereof. In particular, SEQ ID NO: 39 includes the first 29 amino acids of the mature CsgF peptide (SEQ ID NO: 6). In another embodiment, the modified CsgF peptide of the invention is a cleavage peptide that includes SEQ ID NO: 40. In particular, SEQ ID NO: 40 is the mature CsgF peptide ( In certain embodiments, the CsgF fragment includes amino acid sequence number 39, or a mutant or homolog thereof. In particular, SEQ ID NO: 39 includes the first 29 amino acids of the mature CsgF peptide (SEQ ID NO: 6). In another embodiment, the modified CsgF peptide of the invention is a cleavage peptide that includes SEQ ID NO: 40. In particular, SEQ ID NO: 40 is the mature CsgF peptide ( In certain embodiments, the CsgF fragment includes amino acid sequence number 39, or a mutant or homolog thereof. In particular, SEQ ID NO: 39 includes the first 29 amino acids of the mature CsgF peptide (SEQ ID NO: 6). In another embodiment, the modified CsgF peptide of the invention is a cleavage peptide that includes SEQ ID NO: 40. In particular, SEQ ID NO: 40 is the mature CsgF peptide ( In certain embodiments, the CsgF fragment includes amino acid sequence number 39, or a mutant or homolog thereof. In particular, SEQ ID NO: 39 includes the first 29 amino acids of the mature CsgF peptide (SEQ ID NO: 6). In another embodiment, the modified CsgF peptide of the invention is a cleavage peptide that includes SEQ ID NO: 40. In particular, SEQ ID NO: 40 is the mature CsgF peptide ( It contains the first 45 amino acids of SEQ ID NO: 6. In particular, the CsgF constriction site and the binding site to CsgG are located within the N-terminal CsgF peptide region, amino acids 39- 64 of SEQ ID NO: 5 (present in SEQ ID NO: 39 and SEQ ID NO: 40), or in particular amino acids 4 9-64 of SEQ ID NO: 5 (present in SEQ ID NO: 40 but not in SEQ ID NO: 39, the latter fragment encoded by SEQ ID NO: 39 shows a weak interaction with CsgG (see examples)) but is further characterized in that it confers higher stability to the complex. Accordingly, the present disclosure provides a modification of the CsgF protein by cleaving the protein with respect to the peptide, or the N-terminal fragment or the constriction site region-containing peptide, to enable in vivo complex formation with the CsgG pore or its homologs or mutants . In one embodiment, there are further restrictions regarding the modified CsgF peptide containing SEQ ID NO: 37 or SEQ ID NO: 38. Finally, the identification of CsgF homologous peptides, particularly those located within the constriction region (FCP peptide), also provides modified CsgF peptide homologs that can form part of the isolated complex (see, for example, FIGS. 8 and 10). In one embodiment, there are further restrictions regarding the modified CsgF peptide containing SEQ ID NO: 37 or SEQ ID NO: 38. Finally, the identification of CsgF homologous peptides, particularly those located within the constriction region (FCP peptide), also provides modified CsgF peptide homologs that can form part of the isolated complex (see, for example, FIGS. 8 and 10). peptide), also provides modified CsgF peptide homologs that can form part of the isolated complex (see, for example, FIGS. 8 and 10). Another embodiment relates to a modified or cleaved CsgF peptide containing SEQ ID NO: 15, where the SEQ ID NO: 15 contains a region of the CsgF protein containing several residues from the CsgG binding site and / or the constriction site region sufficient for in

[0086] vitro reconstitution of a complex pore containing CsgG or its homolog and the modified CsgF peptide to yield an isolated pore complex containing CsgF channel constriction . Another embodiment relates to a modified or cleaved CsgF peptide containing SEQ ID NO: 15, where the SEQ ID NO: 15 contains a region of the CsgF protein containing several residues from the CsgG binding site and / or the constriction site region sufficient for in vitro reconstitution of a complex pore containing CsgG or its homolog and the modified CsgF peptide to yield an isolated pore complex containing CsgF channel constriction . Another embodiment relates to a modified or cleaved CsgF peptide containing SEQ ID NO: 16 describes said modified CsgF peptide, which is an N-terminal fragment of the CsgF protein and contains two additional amino acids (KD), which increases the solubility and stability of said (synthetic) peptide and allows for in vitro reconstitution of said complex pores. Embodiments are further provided wherein said modified CsgF peptide comprises SEQ ID NO: 15, SEQ ID NO: 16 or a homolog or mutant thereof, wherein said modified CsgF peptide is further mutated, but within the region of the modified CsgF peptide corresponding to SEQ ID NO: 15 or 16, there is at least retained at least 35% amino acid identity to SEQ ID NO: 15 or SEQ ID NO: 16, such as 40% %, 50%, 60%, 70%, 80%, 85%, 90% amino acid identity. Embodiments are further provided wherein said modified CsgF peptide comprises SEQ ID NO: 15, SEQ ID NO: 16 or a homolog or mutant thereof, wherein said modified CsgF peptide is further mutated, but within the region of the modified CsgF peptide corresponding to SEQ ID NO: 15 or 16, there is at least retained at least 40%, 45%, 50%, 60%, 70% 80%, 85%, 90% amino acid identity to SEQ ID NO: 15 or SEQ ID NO: 16. These mutated regions are intended to alter and / or improve the properties of the CsgF constriction site as described above, and thus for example more accurate target analysis can be obtained. In another embodiment a modified CsgF peptide is disclosed, wherein one or more positions in the region comprising SEQ ID NO: 39, SEQ ID NO: 40, SEQ ID NO: 54, or SEQ ID NO: 55 are modified, and said mutant is relative to the region comprising SEQ ID NO: 39, SEQ ID NO: 40, SEQ ID NO: 54, or SEQ ID NO: 55 wherein said modified CsgF peptide is further mutated, but within the region of the modified CsgF peptide corresponding to SEQ ID NO: 15 or 16, there is at least retained at least 40%, 45%, 50%, 60%, 70% 80%, 85%, 90% amino acid identity to SEQ ID NO: 15 or SEQ ID NO: 16. These mutated regions are intended to alter and / or improve the properties of the CsgF constriction site as described above, and thus for example more accurate target analysis can be obtained. In another embodiment a modified CsgF peptide is disclosed, wherein one or more positions in the region comprising SEQ ID NO: 39, SEQ ID NO: 40, SEQ ID NO: 54, or SEQ ID NO: 55 are modified, and said mutant is relative to the region comprising SEQ ID NO: 39, SEQ ID NO: 40, SEQ ID NO: 54, or SEQ ID NO: 55 wherein said modified CsgF peptide is further mutated, but within the region of the modified CsgF peptide corresponding to SEQ ID NO: 15 or 16, there is at least retained at least 40%, 45%, 50%, 60%, 70% 80%, 85%, 90% amino acid identity to SEQ ID NO: 15 or SEQ ID NO: 16. These mutated regions are intended to alter and / or improve the properties of the CsgF constriction site as described above, and thus for example more accurate target analysis can be obtained. In another embodiment In the corresponding peptide fragment, at least 35% amino acid identity with SEQ ID NO: 39, SEQ ID NO: 40, SEQ ID NO: 54, or SEQ ID NO: 55, or 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95% amino acid identity is retained.

[0087] Therefore, a further embodiment of the present invention relates to a CsgG pore or its homolog or mutant variant, and an isolated pore complex comprising the denatured CsgF peptide or its homolog or mutant, where the modified CsgF peptide is defined as described in the second aspect of the present invention. As defined as described in

[0088] Additional embodiments relate to an isolated pore complex, where the CsgG pore and the modified CsgF peptide via at least one monomer are covalently bonded. The covalent bond or linkage is one example possible via a cysteine bond, where the sulfhydryl side group of cysteine covalently bonds to another amino acid residue or moiety. In a second possibility, obtained via interactions between covalent, non-natural (photo)reactive amino acids . (Photo)reactive amino acids refer to artificial analogs of natural amino acids that can be used for cross-linking protein complexes and can be incorporated into proteins and peptides in vivo or in vitro. Common photo-reactive amino acid analogs in general use are leucine and methionine thionine, and para-benzoyl-phenyl-alanine, and azidohomoalanine , homopropargylglycine, homoallylglycine, p-acetyl-Phe, p-azido -Phe, p-propargyloxy-Phe, and p-benzoyl-Phe for light reactions. -Phe, and p-propargyloxy-Phe, and p-benzoyl-Phe for light Reactive diazirine analogs (Wang et al. 2012, Chin et al. 2002). When exposed to ultraviolet light, these are activated and bind to interacting proteins within a few angstroms of some of the photoreactive amino acid analogs. However, the positions within the CsgG monomer where such covalent bonding can occur depend on exposure to the modified CsgF peptide fragment. As shown in Figure 1, several amino acids are at positions that provide covalent bonding, namely positions 132, 133, 136, 138, 140, 142, 144, 145, 147, 149, 151, 153, 155, 183, 185, 187, 189, 191, 201, 203, 205, 207 or 209 of SEQ ID NO: 3 or its homologs .< < >

[0089] Another aspect of the present invention relates to a construct comprising the modified CsgF peptide, wherein the peptide is covalently bonded. A "construct" comprises two or more covalently bonded monomers derived from modified CsgF and / or CsgG or their homologs . In other words, a construct can comprise multiple monomers. In another aspect, the present invention also provides a pore complex comprising at least one construct of the present invention . The pore complex comprises a sufficient number of constructs and, optionally, monomers that form pores. For example, an octameric pore can be formed by (a) four constructs each containing two monomers, (b) two constructs each containing four monomers, (c) one construct containing two monomers and six monomers that do not form part of a construct, or (d) one or two CsgF monomers within one construct and six or seven CsgG monomers in one construct, or even (e) another construct containing only CsgG monomers . < > < In addition, it may include constructs having CsgF and CsgG monomers. For example, a nonameric provides identical and additional possibilities for the pores. Other combinations of constructs and monomers can be envisioned by those skilled in the art. One or more constructs of the invention can be used to form pore complexes for characterization such as sequencing, polynucleotides, etc. The constructs can contain at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10 monomers. The construct preferably contains two monomers. Two or more monomers can be the same or different and can be CsgF, CsgG, CsgG / CsgF fusion monomers or homologs thereof, or any combination thereof. Another embodiment relates to a polynucleotide or nucleic acid molecule encoding the modified CsgF peptide or homolog or mutant thereof of the invention, or a polynucleotide encoding the construct described above.

[0090] A particular embodiment relates to an isolated transmembrane pore complex comprising an isolated pore complex according to the first and second aspects of the invention and a membrane constituent. The isolated transmembrane pore complex is directly applicable for use in molecular sensing such as nucleic acid sequencing. Alternatively, there is provided a membrane composition comprising a modified CsgG / CsgF biological pore according to the isolated pore complex of the invention described herein, and a membrane, membrane constituent, or insulating layer. One embodiment

[0091] relates to an isolated transmembrane pore complex consisting of an isolated pore complex according to the invention and a membrane constituent. constituent. The isolated transmembrane pore complex described in the present specification is directly applicable for use in molecular sensing such as nucleic acid sequencing. Alternatively, there is provided a membrane composition comprising a modified CsgG / CsgF biological pore according to the isolated pore complex of the invention described herein, and a membrane, membrane constituent, or insulating layer. One embodiment relates to an isolated transmembrane pore complex consisting of an isolated pore complex according to the invention and a membrane constituent.

[0092] The CsgG:CsgF complex is highly stable, but when CsgF is cleaved, CsgG The stability of the CsgF complex is reduced compared to the complex containing full-length CsgF. For example, cysteine ​​residues can be introduced at positions specified herein followed by formation of disulfide bonds. A pore complex can be created between CsgG and CsgG to make the complex more stable. can be prepared by any of the methods previously described, and disulfide bond formation can be achieved by It can be induced by using an oxidizing agent (eg, copper-orthophenanthroline). Other interactions (e.g., hydrophobic interactions, charge-discharge interactions / electrostatic interactions) may also be involved. It can be used at those locations instead of stain interactions.

[0093] In another embodiment, unnatural amino acids may be incorporated at these positions. In the present study, the covalent bond is created via click chemistry. For example, azide or azide Unnatural amino acids with alkynyl or dibenzocyclooctyne (DBCO) groups and / or Or a bicyclo[6.1.0]nonyne (BCN) group is introduced at one or more of these positions. It is possible.

[0094] Such stabilizing mutations include, for example, CsgG and and / or any other modification to CsgF.

[0095] The CsgG pore has been modified to facilitate attachment to the CsgF peptide. For example, the nucleic acid may contain at least one CsgG monomer. For example, cysteine ​​residues are located at positions 132, 133, 136, 138, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158 42, 144, 145, 147, 149, 151, 153, 155, 183, 185, 1 at one or more positions corresponding to 87, 189, 191, 201, 203, 205, 207, and 209, and / or at any one of the positions predicted to promote covalent bonding to CsgG upon contact with CsgF as described in Table 4 can be introduced. As an alternative or in addition to covalent bonding via cysteine residues, the pore can be stabilized by hydrophobic or electrostatic interactions. To promote such interactions, at positions corresponding to one or more of positions 132, 133, 136, 138, 140, 142, 144, 145, 147, 149, 151, 153, 155, 183, 185, 187, 189, 191, 201, 203, 205, 207, and 209 of SEQ ID NO: 3, and / or at any one of the positions predicted to contact CsgF and described in Table 4, non-native reactive or photoreactive amino acids are. 133, 136, 138, 140, 142, 144, 145, 147, 149, 151, 153, 155, 183, 185, 187, 189, 191, 201, 203, 205, 207, and 209 of SEQ ID NO: 3, and / or at any one of the positions predicted to contact CsgF and described in Table 4, non-native reactive or photoreactive amino acids are. The CsgF peptide can be modified to promote attachment to the CsgG pore. For example, cysteine residues can be introduced at one or more positions corresponding to positions 1, 4, 5, 8, 9, 11, 12, 26, or 29 of SEQ ID NO: 6, and / or at any one of the positions predicted to promote covalent bonding to CsgG upon contact with CsgF as described in Table 4. As an alternative or in addition to covalent bonding via cysteine residues, the pore can be stabilized by hydrophobic or electrostatic interactions. To promote such interactions, at positions corresponding to one or more of positions 1, 4, 5, 8, 9, 11, 12, 26, or 29 of SEQ ID NO: 6,

[0096] The CsgF peptide can be modified to promote attachment to the CsgG pore. For example, cysteine residues can be introduced at one or more positions corresponding to positions 1, 4, 5, 8, 9, 11, 12, 26, or 29 of SEQ ID NO: 6, and / or at any one of the positions predicted to promote covalent bonding to CsgG upon contact with CsgF as described in Table 4. As an alternative or in addition to covalent bonding via cysteine residues, the pore can be stabilized by hydrophobic or electrostatic interactions. To promote such interactions, at positions corresponding to one or more of positions 1, 4, 5, 8, 9, 11, 12, 26, or 29 of SEQ ID NO: 6, and / or at any one of the positions predicted to contact CsgF and described in Table 4. The CsgF peptide can be modified to promote attachment to the CsgG pore. For example, cysteine residues can be introduced at one or more positions corresponding to positions 1, 4, 5, 8, 9, 11, 12, 26, or 29 of SEQ ID NO: 6, and / or at any one of the positions predicted to promote covalent bonding to CsgG upon contact with CsgF as described in Table 4. As an alternative or in addition to covalent bonding via cysteine residues, the pore can be stabilized by hydrophobic or electrostatic interactions. To promote such interactions, at positions corresponding to one or more of positions 1, 4, 5, 8, 9, 11, 12, 26, or 29 of SEQ ID NO: 6, and / or at any one of the positions predicted to promote covalent bonding to CsgG upon contact with CsgF as described in Table 4 can be introduced. As an alternative or in addition to covalent bonding via cysteine residues, the pore can be stabilized by hydrophobic or electrostatic interactions. To promote such interactions, at positions corresponding to one or more of positions 1, 4, 5, 8, 9, 11, 12, 26, or 29 of SEQ ID NO: 6, and / or at any one of the positions predicted to contact CsgF and described in Table 4. The CsgF peptide can be modified to promote attachment to the CsgG pore. For example, cysteine residues can be introduced at one or more positions corresponding to positions 1, 4, 5, 8, 9, 11, 12, 26, or 29 of SEQ ID NO: 6, and / or at any one of the positions predicted to promote covalent bonding to CsgG upon contact with CsgF as described in Table 4. As an alternative or in addition to covalent bonding via cysteine residues, the pore can be stabilized by hydrophobic or electrostatic interactions. To promote such interactions, at positions corresponding to one or more of positions 1, 4, 5, 8, 9, 11, 12, 26, or 29 of SEQ ID NO: 6, and / or at any one of the positions predicted to contact CsgF and described in Table 4. and / or at any one of the positions predicted to contact CsgF and described in Table 4. wherein a non-native reactive or photoreactive amino acid is

[0097] Preferred exemplary CsgF peptides contain the following mutations relative to SEQ ID NO: 6: N15X1 / N17X2 / A20X3 / N24X4 / A28XX5 / D34X6, wherein X1 is N / S / A / T / Q / G / L / V / I / F / Y / W / R / K / D / C, X2 is N / S / A / T / Q / G / L / V / I / F / Y / W / R / K / D / C, X3 is A / S / T / Q / N / G / L / V / I / F / Y / W / R / K / D / C, X4 is N / S / T / Q / A / G / L / V / I / F / Y / W / R / K / D / C, X5 is A / S / T / Q / N / G / L / V / I / F / Y / W / R / K / D / C and X5 is D / F / Y / W / R / K / N / Q / C is. Mutants at positions N15, N17, A20, N24 and A28 are contraction mutants and mutants at position 34 affect the interaction of CsgF with the bottom of the CsgG pore to stabilize the interaction.

[0098] CsgG pore The CsgG pore may be a homo-oligomeric pore containing the same mutant monomers of the present invention. The CsgG pore may be, for example, a hetero-oligomeric pore derived from CsgG containing at least one mutant monomer disclosed herein.

[0099] The CsgG pore can contain any number of mutant monomers. The pore generally contains at least 7, at least 8 at least 9, or at least 10 identical mutant monomers, such as 7, 8 at least 9, or at least 10 identical mutant monomers. The CsgG pore preferably contains 8 or 9 identical mutant monomers.

[0100] In a preferred embodiment, all monomers within the hetero-oligomeric CsgG pore (such as 10, 9, , 8, or 7 of the monomers) are the mutant monomers disclosed herein, where at least one of them is different from the others. These can be different from each other.

[0101] The mutant monomers within the CsgG pore are preferably all of approximately the same length or of the same length. The barrels of the mutant monomers within the pores of the present invention are preferably of approximately the same length or of the same length. The length can be measured in terms of the number of amino acids and / or units of length.

[0102] The mutant monomer can be a variant of SEQ ID NO: 3. Over the entire length of the amino acid sequence of SEQ ID NO: 3, the variant is preferably at least 50% identical to its sequence based on amino acid identity. The variant is preferably at least 55%, at least 60%, at least 65% based on the amino acid identity to the amino acid sequence of SEQ ID NO: 3 over the entire sequence, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, and more preferably at least 95%, 97% or 99% homology can be possessed. For example, over a continuous amino acid range of 100 or more, such as 125, 150, 175 or 200 or more, it can have at least 80%, such as at least also 85%, 90% or 95% amino acid identity ("hard homology").

[0103] The CsgG monomers are highly conserved (Figures 45-4 of WO2017 / 149317 which can be easily understood from 7). Further, from the knowledge of the mutations related to SEQ ID NO: 3, it is possible to determine the equivalent positions of CsgG monomer mutations other than SEQ ID NO: 3. The equivalent positions of CsgG monomer mutations other than SEQ ID NO: 3 can be determined.

[0104] Therefore, references to mutant CsgG monomers comprising variants of the sequences shown in SEQ ID NO: 3 and the specific amino acid mutants thereof recited in the claims and elsewhere in this specification also encompass mutant CsgG monomers comprising variants of the sequences shown in SEQ ID NOs: 68 - 88 and their corresponding amino acid mutations. Similarly, references to constructs, pores or methods involving the use of pores related to CsgG monomers comprising variants of the sequences shown in SEQ ID NO: 3 and the specific amino acid mutations described herein also encompass constructs, pores or methods related to mutant CsgG monomers comprising variants of the sequences and their corresponding amino acid mutations disclosed above. As further understood, the present invention also extends to other mutant CsgG monomers not explicitly specified in the specialization showing highly conserved regions. Homology may be determined using standard methods in the art. For example, the UWGCG package provides the BESTFIT program which can be used to calculate homology, for example, using its default settings (Devereux et al (1984) Nucleic Acids Research 12, p387 - 395). The PILEUP and BLAST algorithms can be used for calculations of homology or alignment of sequences (such as identifying equivalent residues or corresponding sequences, typically using their default settings). which can be easily understood from 7). Further, from the knowledge of the mutations related to SEQ ID NO: 3, it is possible to determine the equivalent positions of CsgG monomer mutations other than SEQ ID NO: 3. not explicitly specified in the specialization showing highly conserved regions. also encompass mutant CsgG monomers comprising variants of the sequences shown in SEQ ID NOs: 68 - 88 and their corresponding amino acid mutations. Similarly, references to constructs, pores or methods involving the use of pores related to CsgG monomers comprising variants of the sequences shown in SEQ ID NO: 3 and the specific amino acid mutations described herein

[0105] Homology may be determined using standard methods in the art. For example, the UWGCG package provides the BESTFIT program which can be used to calculate homology, for example, using its default settings (Devereux et al (1984) Nucleic Acids Research 12, p387 - 395). The PILEUP and BLAST algorithms can be used for calculations of homology or alignment of sequences (such as identifying equivalent residues or corresponding sequences, typically using their default settings). Nucleic Acids Research 12, p387 - 395). The PILEUP and BLAST algorithms can be used for calculations of homology or alignment of sequences (such as identifying equivalent residues or corresponding sequences, typically using their default settings). or corresponding sequences, typically using their default settings). can be used. For example, Altschul S. F. (1993) J Mol E vol 36:290-300, Altschul, S.F et al (199 0) J Mol Biol 215:403-10. The software for performing BLAST analysis is publicly available from the National Center for Biotechnology Information (http: / / www.ncbi.nlm.nih .gov / ).

[0106] SEQ ID NO: 3 is a wild-type CsgG monomer and is derived from Escherichia coli strain K-12 substr. MC4100. Variants of SEQ ID NO: 3 may contain any substitution present in another CsgG homolog. Preferred CsgG homologs are shown in SEQ ID NOs: 68-88. The variant may contain one or more combinations of the substitutions present in SEQ ID NOs: 68-88 as compared to SEQ ID NO: 3. For example, mutations may be made at any one or more of the positions of SEQ ID NO: 3 that differ between SEQ ID NO: 3 and any one of SEQ ID NOs: 68-88. Such mutations may be substitutions of the amino acids in SEQ ID NO: 3 with the corresponding amino acids from any one of SEQ ID NOs: 68-88. Alternatively, the mutation at any one of these positions may be a substitution with any amino acid, or a deletion or insertion mutation (such as a deletion or insertion of 1 to 10 amino acids, a deletion or insertion of 2 to 8 or 3 to 6 amino acids, etc.). In addition to the mutations disclosed herein, the amino acids conserved among all of SEQ ID NO: 3 and SEQ ID NOs: 66-88 are preferably present in the variants of the present invention. However, among these positions conserved among all of SEQ ID NO: 3 and SEQ ID NOs: 66-88 it is preferred that the amino acids present in the mutants of the present invention. However, among these positions conserved among all of SEQ ID NO: 3 and SEQ ID NOs: 66-88 ​​​​​​ Conservative mutations may be made at any one or more positions.

[0107] The present invention provides a pore-forming CsgG mutant monomer containing any one or more amino acids described herein as substituted at a position within the structure of the CsgG monomer corresponding to a specific position of SEQ ID NO: 3. The corresponding position can be determined by standard techniques in the art. For example, using the above-described PILEUP algorithm and the BLAST algorithm, the sequence of the CsgG monomer is aligned with SEQ ID NO: 3, whereby the corresponding residues can be identified.

[0108] The pore-forming mutant monomer generally retains the ability to form the same three-dimensional structure as the wild-type CsgG monomer, such as the same three-dimensional structure as the CsgG monomer having the sequence of SEQ ID NO: 3. The three-dimensional structure of CsgG is well-known in the art and is disclosed, for example, in Goyal et al (2014) Nature 516(7530):250-3. Conditions under which the improved properties imparted to the CsgG mutant monomer by the mutations of the present invention are retained, any number of mutations may be made in the wild-type CsgG sequence in addition to the mutations described herein.

[0109] Typically, the CsgG monomer retains the ability to form a structure containing three alpha helices and five beta sheets. The mutant is at least within the region of CsgG that is the N-terminus relative to the first alpha helix (starting at S63 in SEQ ID NO: 3), within the second alpha helix (G85 - A99 of SEQ ID NO: 3), and within the loop between the second alpha helix and the first beta sheet (SEQ ID NO: ​​​​​​​​​​​​​​within Q100 - N120 of SEQ ID NO: 3, and within the fourth and fifth sheets (SEQ ID NO: 3, S1 within 73 - R192 and R198 - T107, and the loop between the fourth and fifth β - sheets (within F193 - Q197 of SEQ ID NO: 3), a CsgG monomer can be made without affecting its ability to form a transmembrane pore (where the transmembrane pore has the ability to translocate polypeptides). Thus, in order to form a pore capable of translocating polynucleotides, additional mutations can be made in these regions within any CsgG monomer without affecting the monomer's ability. It is assumed that mutants can be made within any α - helix (SEQ ID NO: 3, S6 3 - R76, G85 - A99 or V211 - L236), or within any β - sheet (SEQ ID NO: 3, I121 - N133, K135 - R142, I146 - R162, S173 - R192 or R198 - T107), etc., other regions, without affecting the ability of the monomer to form a pore capable of translocating polynucleotides. It is also expected that one or more amino - acid deletions can be made within any loop region connecting α - helices and β - sheets, and / or within the N - terminal and / or C - terminal regions of the CsgG monomer without affecting the ability of the monomer to form a pore capable of translocating polynucleotides. In addition to those described above, amino - acid substitutions may be made in the amino - acid sequence of SEQ ID NO: 3, for example, up to 1, 2, 3, 4, 5, 10, 20 or 30 substitutions. Conservative substitutions replace an amino - acid with another amino - acid having a similar chemical structure, similar chemical properties, or similar side - chain volume. The introduced amino - acid has a polarity, hydrophilicity, hydrophobicity similar to the amino - acid being replaced. R192 or R198 - T107), etc., other regions, without affecting the ability of the monomer to form a pore capable of translocating polynucleotides. It is also expected that one or more amino - acid deletions can be made within any loop region connecting α - helices and β - sheets, and / or within the N - terminal and / or C - terminal regions of the CsgG monomer without affecting the ability of the monomer to form a pore capable of translocating polynucleotides. It is also expected that mutants can be made within other regions such as within any α - helix (SEQ ID NO: 3, S6 3 - R76, G85 - A99 or V211 - L236), or within any β - sheet (SEQ ID NO: 3, I121 - N133, K135 - R142, I146 - R162, S173 - R192 or R198 - T107), etc., without affecting the ability of the monomer to form a pore capable of translocating polynucleotides. Amino - acid substitutions may be made in the amino - acid sequence of SEQ ID NO: 3 in addition to those described above, for example, up to 1, 2, 3, 4, 5, 10, 20 or 30 substitutions. Conservative substitutions replace an amino - acid with another amino - acid having a similar chemical structure, similar chemical properties, or similar side - chain volume. The introduced amino - acid has a polarity, hydrophilicity, hydrophobicity similar to the amino - acid being replaced. It is also expected that one or more amino - acid deletions can be made within any loop region connecting α - helices and β - sheets, and / or within the N - terminal and / or C - terminal regions of the CsgG monomer without affecting the ability of the monomer to form a pore capable of translocating polynucleotides. Amino - acid substitutions may be made in the amino - acid sequence of SEQ ID NO: 3 in addition to those described above, for example, up to 1, 2, 3, 4, 5, 10, 20 or 30 substitutions. Conservative substitutions replace an amino - acid with another amino - acid having a similar chemical structure, similar chemical properties, or similar side - chain volume. The introduced amino - acid has a polarity, hydrophilicity, hydrophobicity similar to the amino - acid being replaced. It is also expected that mutants can be made within other regions such as within any α - helix (SEQ ID NO: 3, S6

[0110] In addition to those described above, amino - acid substitutions may be made in the amino - acid sequence of SEQ ID NO: 3, for example, up to 1, 2, 3, 4, 5, 10, 20 or 30 substitutions. Conservative substitutions replace an amino - acid with another amino - acid having a similar chemical structure, similar chemical properties, or similar side - chain volume. The introduced amino - acid has a polarity, hydrophilicity, hydrophobicity similar to the amino - acid being replaced. Conservative substitutions replace an amino - acid with another amino - acid having a similar chemical structure, similar chemical properties, or similar side - chain volume. The introduced amino - acid has a polarity, hydrophilicity, hydrophobicity similar to the amino - acid being replaced. Conservative substitutions replace an amino - acid with another amino - acid having a similar chemical structure, similar chemical properties, or similar side - chain volume. The introduced amino - acid has a polarity, hydrophilicity, hydrophobicity similar to the amino - acid being replaced. polarity, hydrophilicity, hydrophobicity similar to the amino - acid being replaced. It may be basic, acidic, neutral or charged. Alternatively, a conservative substitution may introduce another amino acid, which is aromatic or aliphatic, in place of an existing aromatic amino acid or aliphatic amino acid. Changes in conserved amino acids are well known in the art and may be selected according to the properties of the 20 major amino acids defined in Table 1 above. When amino acids have similar polarities, this may also be determined by referring to the hydropathy scale for the amino acid side chains in Table 2 as well.

[0111] One or more amino acid residues of the amino acid sequence of SEQ ID NO: 3 may be additionally deleted from the above-mentioned polypeptide. Up to 1, 2, 3, 4, 5, 10, 20 or 30 or more residues may be deleted.

[0112] The variant may include a fragment of SEQ ID NO: 3. Such fragments retain pore-forming activity. The fragments may be at least 50, at least 100, at least 150, at least 200 or at least 250 amino acids in length. Such fragments may be used to generate pores. The fragments preferably span the membrane of the domain of SEQ ID NO: 3, i.e., K135-Q153 and S183-S208.

[0113] One or more amino acids may alternatively or additionally be added to the above-mentioned polypeptide. The extension may be provided at the amino terminus or carboxy terminus of the amino acid sequence of SEQ ID NO: 3 or the polypeptide variant or a fragment thereof. The extension may be quite short, for example, between 1 and 10 amino acids. Alternatively, the extension may be longer, for example, up to 50 amino acids or 100 amino acids. A carrier protein may be fused to the amino acid sequence according to the present invention ​​​​​ This may be done. Other fusion proteins will be discussed in detail below.

[0114] The CsgG pores described herein include wild-type CsgG pores or homologs or mutants thereof. A mutant is a polypeptide having an amino acid sequence that varies from the amino acid sequence of SEQ ID NO: 3 and that retains its ability to form a pore. Mutants generally include the region of SEQ ID NO: 3 responsible for pore formation. The pore-forming ability of CsgG, which includes a β-barrel, is provided by β-sheets within each subunit. Mutants of SEQ ID NO: 3 typically include a β-sheet, i.e., the region within SEQ ID NO: 3 that forms K134-Q154 and S183-S208. One or more modifications can be made to the region of SEQ ID NO: 3 that forms the β-sheet as long as the resulting mutant retains its ability to form a pore. Mutants of SEQ ID NO: 3 preferably include one or more modifications such as substitutions, additions, or deletions within its α-helix and / or loop regions. -sheet, i.e., the region within SEQ ID NO: 3 that forms K134-Q154 and S183-S208. One or more modifications can be made to the region of SEQ ID NO: 3 that forms the β-sheet as long as the resulting mutant retains its ability to form a pore. Mutants of SEQ ID NO: 3 preferably include one or more modifications such as substitutions, additions, or deletions within its α-helix and / or loop regions. One or more modifications can be made to the region of SEQ ID NO: 3 that forms the β-sheet as long as the resulting mutant retains its ability to form a pore. Mutants of SEQ ID NO: 3 preferably include one or more modifications such as substitutions, additions, or deletions within its α-helix and / or loop regions. A mutant CsgG monomer may be a mutant CsgG monomer whose sequence varies from the sequence of the wild-type CsgG monomer and that retains its ability to form a pore. The mutant monomer may also be referred to herein as a mutant. Methods for confirming the ability of a mutant monomer to form a pore are well known in the art and will be discussed in more detail below.

[0115] A mutant CsgG monomer may be a mutant CsgG monomer whose sequence varies from the sequence of the wild-type CsgG monomer and that retains its ability to form a pore. The mutant monomer may also be referred to herein as a mutant. Methods for confirming the ability of a mutant monomer to form a pore are well known in the art and will be discussed in more detail below. Methods for confirming the ability of a mutant monomer to form a pore are well known in the art and will be discussed in more detail below.

[0116] Specific pore-forming CsgG mutant monomers that may be included in the CsgG pores may include any one or more of the following modifications. - W at the position corresponding to R97 of SEQ ID NO: 3 ​​​​​​​​​- W at the position corresponding to R93 of Sequence No. 3, - Y at the position corresponding to R97 of Sequence No. 3, - Y at the position corresponding to R93 of Sequence No. 3, - Y at each position corresponding to R93 and R97 in Sequence No. 3, - D at the position corresponding to R192 of Sequence No. 3, - Deletion of residues at the position corresponding to V105 - I107 of Sequence No. 3, - Deletion of one or more residues at positions corresponding to F193~L199 of Sequence No. 3, - Deletion of residues at the position corresponding to 195~L199 of Sequence No. 3, - Deletion of residues at the position corresponding to 193~L199 of Sequence No. 3, - T at the position corresponding to R191 of Sequence No. 3, - Q at the position corresponding to K49 of Sequence No. 3, - N at the position corresponding to K49 of Sequence No. 3, - Q at the position corresponding to K42 of Sequence No. 3, - Q at the position corresponding to E44 of Sequence No. 3, - N at the position corresponding to E44 of Sequence No. 3, - R at the position corresponding to L90 of Sequence No. 3, - R at the position corresponding to L91 of Sequence No. 3, - R at the position corresponding to I95 of Sequence No. 3, - R at the position corresponding to A99 of Sequence No. 3, - H at the position corresponding to E101 of Sequence No. 3, - K at the position corresponding to E101 of Sequence No. 3, - N at the position corresponding to E101 of Sequence No. 3, - Q at the position corresponding to E101 of Sequence No. 3, - T at the position corresponding to E101 of Sequence No. 3, - K at the position corresponding to Q114 of Sequence No. 3.

[0117] The CsgG pore-forming monomer preferably further contains A at a position corresponding to Y51 in SEQ ID NO: 3 and / or Q at a position corresponding to F56 in SEQ ID NO: 3.

[0118] Pores composed of CsgG monomers containing a substitution from R to W at a position corresponding to position 97 of SEQ ID NO: 3 show an increase in accuracy when characterizing (or sequencing) a target polynucleotide compared to the same pore without modification at position 97. The increase in accuracy is also seen when the CsgG monomer contains a modification from R to W at a position corresponding to position 97 of SEQ ID NO: 3 or a modification from R to Y at positions corresponding to positions 93 and 97 of SEQ ID NO: 3. Thus, the pore can be composed of one or more mutant CsgG monomers containing a modification at a position corresponding to R97 or R93 of SEQ ID NO: 3 such that the modification increases the hydrophobicity of the amino acid. For example, such modifications can include amino acid substitutions with any amino acid containing a hydrophobic side chain, including but not limited to W and Y.

[0119] CsgG monomers containing a mutation from R to D, Q, F, S or T at a position corresponding to position 192 of SEQ ID NO: 3 are more easily expressed than monomers without the substitution at position 192 which may be due to a decrease in positive charge. Thus, position 192 can be substituted with an amino acid that decreases the positive charge. Monomers containing R192D / Q / F / S / T can also include additional modifications that improve the ability of the mutant pore formed from the monomer to interact with and characterize an analyte such as a polynucleotide. However, in one embodiment, the residue at the position corresponding to position 193 of SEQ ID NO: 3 is preferably R or K, more preferably R. ​​​​​​​​​​​​Yes.

[0120] Deletions of V105, A106, and I107, F193, I194, D195, Y196, Deletions of Q197, R198, and L199, or D195, Y196, Q197, R1 98, and L199, and / or a CsgG monomer containing F191T, the pores show an increase in accuracy when characterizing (or sequencing) a target polynucleotide. The amino acids at positions 105 - 107 correspond to the cis-loop within the cap of the nanopore, and the amino acids at positions 193 - 199 correspond to the trans-loop at the other end of the pore. Without wishing to be bound by theory, deletions in the cis-loop improve the interaction of the enzyme with the pore, and removal of the trans-loop is thought to reduce unfavorable interactions between DNAs on the trans side of the pore.

[0121] Pores containing a CsgG monomer with a mutation from K to Q or K to N at the position corresponding to K94 of SEQ ID NO: 3 show a decrease in the number of pores with more noise (i.e., pores that increase the signal-to-noise ratio) when characterizing (or sequencing) a target polynucleotide, compared to the same pore without the mutation at K94. Position 94 is found at the entrance of the pore and has been found to be a particularly sensitive position in relation to the noise of the current signal.

[0122] All pores containing a CsgG monomer with T104K or T104R, N91R, E101K / N / Q / T / H, E44N / Q, Q114K, A99R, I95R, N91R, L90R, E44Q / N, and / or Q 42K, or corresponding mutations, Characterizing a target polynucleotide (sequencing it) shows an increased ability to capture the target polynucleotide when compared to identical pores without replacement at the position.

[0123] In one embodiment, the CsgG pore is (a) I41, R93, A98, Q100, G103 , T104, A106, I107, N108, L113, S115, T117, Y130 , at one of the positions of K135, E170, S208, D233, D238 and E244 or more mutants (i.e., mutants at one or more of those positions) and / or is (b) one or more monomers that are mutants of SEQ ID NO: 3, including one or more of D43S, E44S, F48S / N / Q / Y / W / I / V / H / R / K, Q87 N / R / K, N91K / R, K94R / F / Y / W / L / S / N, R97F / Y / W / V / I / K / S / Q / H, E101I / L / A / H, N102K / Q / L / I / V / S / H , R110F / G / N, Q114R / K, R142Q / S, T150Y / A / V / L / S / Q / N, R192D / Q / F / S / T and one or more of D248S / N / Q / K / R. The mutant may include (a), ( b), or both (a) and (b). In some embodiments, the mutant includes R97 W. In some embodiments, the mutant includes R192D / Q / F / S / T (R192D / Q, etc.). In (a), the mutant may include modifications at any number and combination of positions, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 1 0, 11, 12, 13, 14, 15, 16, 17, 18 or 19 positions, etc. In (a), the mutant includes I41N, R93F / Y / W / L / I / V / N / Q / S, A

[0124] (a), the mutant includes I41N, R93F / Y / W / L / I / V / N / Q / S, A 98K / R, Q100K / R, G103F / W / S / N / K / R, T104R / K, A1 06R / K, I107R / K / W / F / Y / L / V, N108R / K, L113K / R, S115R / K, T117R / K, Y130W / F / H / Q / N, K135L / V / N / Q / S, E170S / N / Q / K / R, S208V / I / F / W / Y / L / T, D233 S / N / Q / K / R, D238S / N / Q / K / R and E244S / N / Q / K / R of one or more of which are preferably included.

[0125] (a), the variant preferably contains one or more modifications that provide more consistent movement of the target polynucleotide with respect to transmembrane pores containing monomers (such as passing through). In particular (a), the variant preferably contains one or more mutations at the positions of R93, G103, and I107 (i.e., mutations at one or more of these positions). The variant can contain mutations at the positions of R93, G103, and I107. The variant can contain R93 F / Y / W / L / I / V / N / Q / S, G103F / W / S / N / K / R and I107 R / K / W / F / Y / L / V. These can be present in any combination shown for the positions of R93, G103, and I107. Preferably, the variant contains one or more of R93 F / Y / W / L / I / V / N / Q / S, G103F / W / S / N / K / R and I107 R / K / W / F / Y / L / V. These can be present in any combination shown for the positions of R93,

[0126] (a), the variant preferably contains one or more modifications such that the pore composed of mutant monomers preferably captures nucleotides and polynucleotides more easily. In particular (a), the variant preferably contains one or more modifications such that the pore composed of mutant monomers preferably captures nucleotides and polynucleotides more easily. In particular (a), the variant contains I41, T104, A106, N108, at the positions of L113, S115, T117, E170, D233, D238 and E244 preferably includes one or more mutations (i.e., mutations at one or more of these positions). The variant may include modifications at any number and combination of positions (such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or 11, etc.). The variant preferably includes one or more of I41N, T104R / K, A106R / K, N108R / K, L113K / R, S115R / K, T117R / K, E170S / N / Q / K / R, D233S / N / Q / K / R, D238S / N / Q / K / R and E244S / N / Q / K / R. Alternatively or as another method, the variant may include (c) Q42K / R, E44N / Q, L90R / K, N 91R / K, I95R / K, A99R / K, E101H / K / N / Q / T and / or Q114K / R.

[0127] In (a), the variant preferably includes one or more modifications that provide more consistent movement and increase capture. In particular, in (a), the variant preferably includes one or more mutations (i.e., mutations at one or more of these positions) at the positions of (i) A98, (ii) Q100, ( iii) G103, and (iv) I107. The variant preferably includes one or more of (i) A98R / K, (ii) Q100K / R, (iii) G103K / R, and (iv) I107R / K. K.

[0128] Particularly preferred mutant monomers that provide capture of analytes such as polynucleotides have mutations at one or more of the positions Q42, E44, E44, L90, N91, I95, A99, E101 and Q114 wherein the mutation removes a negative charge at the mutation position, ​ and / or increase the positive charge. In particular, the following mutations can be included in the mutant monomers of the present invention that produce CsgG pores with improved ability to capture analytes, preferably polynucleotides: Q42K, E44N, E44Q, L90R, N91R, I 95R, A99R, E101H, E101K, E101N, E101Q, E101T and Q114K. Examples of specific mutant monomers containing one of these mutants in combination with other beneficial mutants are as follows: CsgG-(WT-Y51A / F56Q / K94Q / R97W / R192D-del( V105-I107)-Q42K CsgG-(WT-Y51A / F56Q / K94Q / R97W / R192D-del( V105-I107)-E44N CsgG-(WT-Y51A / F56Q / K94Q / R97W / R192D-del( V105-I107)-E44Q CsgG-(WT-Y51A / F56Q / K94Q / R97W / R192D-del( V105-I107)-L90R CsgG-(WT-Y51A / F56Q / K94Q / R97W / R192D-del( V105-I107)-N91R CsgG-(WT-Y51A / F56Q / K94Q / R97W / R192D-del( V105-I107)-I95R CsgG-(WT-Y51A / F56Q / K94Q / R97W / R192D-del( V105-I107)-A99R CsgG-(WT-Y51A / F56Q / K94Q / R97W / R192D-del( V105-I107)-E101H CsgG-(WT-Y51A / F56Q / K94Q / R97W / R192D-del( V105-I107)-E101K CsgG-(WT-Y51A / F56Q / K94Q / R97W / R192D-del( V105-I107)-E101N V105 - I107)-E101K CsgG-(WT - Y51A / F56Q / K94Q / R97W / R192D - del( V105 - I107)-E101N CsgG-(WT - Y51A / F56Q / K94Q / R97W / R192D - del( V105 - I107)-E101Q CsgG-(WT - Y51A / F56Q / K94Q / R97W / R192D - del( V105 - I107)-E101T CsgG-(WT - Y51A / F56Q / K94Q / R97W / R192D - del( V105 - I107)-Q114K。

[0129] (a) In particular, the variant preferably contains one or more modifications that increase the accuracy of characterization . Specifically in (a), the variant preferably contains one or more mutations at the positions of Y130, K135, and S208 (such as Y130, K135, S208, Y130 and K135, Y130 and S208, K135 and S208, or Y130, K135, and S208 (i.e., mutations at one or more of these positions). The variant preferably contains one or more of Y130W / F / H / Q / N, K135L / V / N / Q / S, and R142Q / S. These substitutions may be present in any number and combination with respect to Y130, K135, and S208.

[0130] (b) In (b), the variant may contain any number and combination of substitutions (such as 1, 2, 3, 4, 5, 6, 7 , 8, 9, 10, 11, or 12, etc.). In (b), the variant shows more consistent movement of the target polynucleotide with respect to transmembrane pores containing monomers (such as passing through). It is preferable to include one or more modifications that provide. In particular, in (b), the variant is (i) Q87N / R / K, (ii) K94R / F / Y / W / L / S / N, (iii) R97F / Y / W / V / I / K / S / Q / H, (iv) N102K / Q / L / I / V / S / H an d (v) preferably includes one or more of R110F / G / N. The variant is K94D or K94Q and / or R97W or R97Y is more preferred. Mono Other preferred variants that provide more consistent movement of the target polynucleotide (such as penetrating) through the transmembrane pore containing the mer are (vi) R93W and R93Y included. Preferred variants may include R93W and R97W, R93Y and R97W, R93W and R97W, or more preferably R93Y and R97Y.

[0131] (b), the variant is preferably one or more modifications such that the pore composed of mutant monomers more easily captures nucleotides and polynucleotides is included. In particular, in (b), the variant is (i) D43S, (ii) E44S, ( iii) N91K / R, (iv) Q114R / K and (v) D248S / N / Q / K / preferably includes one or more of R.

[0132] (b), the variant preferably includes one or more modifications that provide more consistent movement and increase capture. In particular, in (b), the variant is Q87R / K, E101I / L / A / H and N102K (such as Q87R / K), E101I / L / A / H, N102K, Q 87R / K and E101I / L / A / H, Q87R / K and N102K, E101I 87R / K and N102K, E101I / L / A / H and N102K, or Q87R / K, E101I / L / A / H and N Preferably, it contains one or more of them.

[0133] (b) Preferably, the variant contains one or more modifications that increase the accuracy of characterization. In particular, in (a), preferably the variant contains F48S / N / Q / Y / W / I / V.

[0134] (b) Preferably, the variant contains one or more modifications that increase the accuracy of characterization and increase capture. In particular, in (a), preferably the variant contains F48H / R / K.

[0135] The variant can contain modifications of both (a) and (b) that provide more consistent migration. The variant may contain modifications in both (a) and (b) that increase capture.

[0136] The present invention provides a variant of SEQ ID NO: 3 that increases the throughput of an assay for characterizing analytes such as polynucleotides using pores containing the variant. Such variants can contain a mutation at K94, preferably K94Q or K94N, more preferably a mutation at K94Q. Examples of specific mutant monomers containing the K94Q mutation or the K94N mutation in combination with other beneficial mutants are as follows: CsgG-(WT-Y51A / F56Q / R97W / R192D-StrepII)9- K94N CsgG-(WT-Y51A / F56Q / R97W / R192D-StrepII)9- K94Q CsgG-(WT-Y51A / F56Q / R97W / R192D-StrepII)9- K94Q.

[0137] ​​​​​By using the monomer that is a variant of SEQ ID NO: 3 for forming the CsgG pore, the accuracy of characterization in an assay for characterizing analytes such as polynucleotides can be increased. Such variants include mutations at F191 (preferably F191T), deletion of V105-I107, deletion of F193-L199 or D195-L199, and / or mutations at R93 and / or R97 (preferably R93Y, R97Y, or more preferably R97W, R93W or both R97Y and R97Y). Examples of specific mutant monomers that include one or more of these mutants in combination with other beneficial mutants are as follows: in an assay for characterizing analytes such as polynucleotides can be increased Such variants include mutations at F191 (preferably F191T), deletion of V105-I107, deletion of F193-L199 or D195-L199, and / or mutations at R93 and / or R97 (preferably R93Y, R97Y, or more preferably R97W, R93W or both R97Y and R97Y). Other beneficial mutants combinations with one or more of these mutants are as follows: Examples of specific mutant monomers that include one or more of these mutants in combination with other beneficial mutants are as follows: CsgG-(WT-Y51A / F56Q / R97W / R192D-StrepII)9 -del(D195-L199) CsgG-(WT-Y51A / F56Q / R97W / R192D-StrepII)9 -del(F193-L199) CsgG-(WT-Y51A / F56Q / R97W / R192D-StrepII)9 -F191T CsgG-(WT-Y51A / F56Q / R97W / R192D-del(V105- I107)-StrepII)9 CsgG-(WT-Y51A / F56Q / K94Q / R97W / R192D-del( V105-I107) CsgG-(WT-Y51A / F56Q / R192D-StrepII)9-R93W CsgG-(WT-Y51A / F56Q / R192D-StrepII)9-R93W -del(D195-L199) CsgG-(WT-Y51A / F56Q / R192D-StrepII)9-R93Y / R97Y.

[0138] In another embodiment, the variant of SEQ ID NO: 3 has (A) a deletion of one or more of positions R192, F193, I194, D195, Y196, Q197, R198, L199, L200 and E201 and / or (B) a deletion of one or more of V139 / G140 / D149 / T150 / V186 / Q 187 / V204 / G205 (referred to herein as Band 1), G137 / G138 / Q151 / Y152 / Y184 / E185 / Y206 / T207 (referred to herein as Band 2), and A141 / R142 / G147 / A148 / A188 / G189 / G202 / E203 (referred to herein as Band 3). In (A), the variant may include deletions at any number and combination of positions (such as 1, 2, 3, 4, 5, 6, 7

[0139] , 8, 9 or 10). In (A), it is preferred that the variant includes the following deletions: - D195, Y196, Q197, R198, and L199, - R192, F193, I194, D195, Y196, Q197, R198, L1 99, and L200, - Q197, R198, L199 and L200, - I194, D195, Y196, Q197, R198 and L199, - D195, Y196, Q197, R198, L199 and L200, - Y196, Q197, R198, L199, L200 and E201, - Q197, R198, L199, L200 and E201, - Q197, R198, L199, or - F193, I194, D195, Y196, Q197, R198, and L199 .

[0140] ​​ The variant preferably includes a deletion of D195, Y196, Q197, R198, and L199 or F193, I 194, D195, Y196, Q197, R198, and L199. (B) Any number and combination of bands 1-3, such as band 1, band 2, band 3, bands 1 and 2, bands 1 and 3, bands 2 and 3, or bands 1, 2, and 3, may be deleted. The variant may include a deletion by (A), (B), or (A) and ( B). A variant containing a deletion at one or more positions by the above (A) and / or (B) may further include any of the modifications or substitutions considered above and below. When a modification or substitution is made at one or more positions after the deletion position of SEQ ID NO: 3, the numbering of the one or more positions of the modification or substitution must be adjusted accordingly. For example, when L19 9 is deleted, E244 becomes E243. Similarly, when band 1 is deleted, R1

[0141] 92 becomes R186. In another embodiment, the variant of SEQ ID NO: 3 includes a deletion of one or more of (C) position V105, A106, and I10 7. The deletion by (C) can be performed in addition to the deletion by (A) and / or (B).

[0142]

[0143] The above deletions typically reduce noise associated with the movement of the target polynucleotide, such as through transmembrane pores containing monomers (such as passing through). As a result, the target polynucleotide can be more accurately characterized.

[0143]

[0144]

[0144] In paragraphs where different amino acids are separated by a / symbol, the / symbol means "or". For example, Q87R / K means Q87R or Q87K. For example, Q87R / K means Q87R or Q87K.

[0145] Variants of SEQ ID NO: 3 that provide increased capture of analytes such as polynucleotides may include mutations at T104 (preferably T104R or T104K), N91 (preferably N91R ), mutations at E101 (preferably E101K / N / Q / T / H), position mutations at E44 (preferably E44N or E44Q) and / or position Q42 (preferably Q42K).

[0146] Mutations at different positions of SEQ ID NO: 3 can be combined in any possible way. In particular, monomers within the CsG pore can include one or more mutations that improve accuracy, one or more mutations that reduce noise, and / or one or more mutants that enhance capture of analytes . .

[0147] Variants of SEQ ID NO: 3 preferably comprise one or more of the following: (i) one or more mutants at positions N40, D43 , E44, S54, S57, Q62, R97, E101, E124, E131, R142 (i.e., mutations at one or more of those positions), for example, N40, D43, E44, S54, S57, Q6 2, E101, E131 and T150, or one or more mutations at positions such as N40, D43, E44, E101 and E131 (i.e., mutations at one or more of those positions), (ii) Y51 / N55, Y51 / F56, N55 / F56 or Y5 (i.e., mutations at one or more of those positions), (ii) Y51 / N55, Y51 / F56, N55 / F56 or Y5 (i.e., mutations at one or more of those positions), (ii) Y51 / N55, Y51 / F56, N55 / F56 or Y5 Mutation at 1 / N55 / F56, (iii) Q42R or Q42K, (iv) K49 R, (v) N102R, N102F, N102Y or N102W, (vi) D149N , D149Q or D149R, (vii) E185N, E185Q or E185R, (viii) D195N, D195Q or D195R, (ix) E201N, E201 Q or E201R, (x) E203N, E203Q or E203R, and (xi) Deletion of one or more of positions F48, K49, P50, Y51, P52, A53, S54, N55, F56 and S57. The variant may include any combination of (i)-(xi).

[0148] If the variant includes (i) and any one of (iii)-(xi), it may further include a mutation at one or more of Y51, N5 5, F56, Y51 / N55, Y51 / F56, N55 / F56 or Y51 / N55 / F56, etc., i.e., one or more of Y51, N55 and F56.

[0149] (i) In the case of, the variant may include mutations at any number and combination of N40, D43, E44, S54, S57, Q62, R97, E 101, E124, E131, R142, T150 and R192. In (i), the variant preferably includes one or more mutations at positions N40, D43, E44, S5 4, S57, Q62, E101, E131 and T150 (i.e., mutations at one or more of these positions). In (i), the var iant preferably includes one or more mutations at positions N40, D43, E44, E101 and E131 (i.e., mutations at one or more of these positions). In (i), it is preferred that the variant includes one or more mutations at positions N40, D43, E44, E101 and E131 (i.e., mutations at one or more of these positions). In (i), ​​The variant preferably contains a mutation at S54 and / or S57. In (i), the variant contains a mutation at one or more of (a) S54 and / or S57, (b) Y51, N55, F56, Y51 / N55, Y51 / F56, N55 / F56 or Y51 / N55 / F56), etc., among Y5 1, N55 and F56 is more preferred. When S54 and / or S57 is deleted, in (xi), it / they cannot mutate in (i), and vice versa. In (i), the variant preferably contains a mutation at T150 such as T150I . Alternatively, the variant contains a mutation at one or more of (a) T1 50, (b) Y51, N55, F56, Y51 / N55, Y51 / F56, N55 / F5 6, or Y51 / N55 / F56, etc., among Y51, N55 and F56 is preferred. In (i), the variant preferably contains a mutation at Q62 such as Q62R or Q62K . Alternatively, the variant contains a mutation at one or more of (a) Q62 , (b) Y51, N55, F56, Y51 / N55, Y51 / F56, N55 / F56, or Y51 / N55 / F56, etc., among Y51, N55 and F56 is preferred. The variant may contain a mutation at D43, E44, Q62, or any combination thereof, such as D43, E44, Q62, D43 / E44, D43 / Q62 , E44 / Q62 or D43 / E44 / Q62. Alternatively, the variant contains (a) D43, E44, Q62, D43 / E44, D43 / Q62, E44 / Q6 2 or D43 / E44 / Q62, (b) Y51, N55, F56, Y51 / N55, Y 51 / F56, N55 / F56, or Y51 / N55 / F56, etc., among Y51, N55 and O f Y51, N55 and F56 Preferably, it contains mutations at one or more of Y51 and F56.

[0150] (ii) and elsewhere, this use where different positions are separated by the / symbol in this case, / means "and", meaning that Y51 / N55 is Y51 and N55. In (ii), the mutant preferably contains a mutation in Y51 / N55. Cs The constriction of gG has been proposed to consist of three stacked concentric rings formed by the side chains of residues Y51, N55 and F56 (Goyal et al, 2014, Nat ure, 516, 250 - 253). Thus, mutations of these residues in (ii) can reduce the number of nucleotides that can easily identify the direct relationship between the polynucleotide moving through the pore and the current observed (when the polynucleotide moves through the pore) and the polynucleotide. F56 can be mutated in any of the ways discussed below, with reference to the mutants and pores useful in the methods of the present invention.

[0151] (v) In, the mutant may contain N102R, N102F, N102Y or N102W. In (i), the mutant is (a) N102R, N102F, N102Y or N102

[0152] W, and (b) one or more of Y51, N55 and F56, for example Y51, N 55, F56, Y51 / N55, Y51 / F56, N55 / F56 or Y51 / N55 / F56 mutants are preferably included.

[0152] (xi) In, any number and combination of K49, P50, Y51, P52, A53, S54, N55, F56 and S57 can be deleted. K49, P50, Y51, P52 It is preferred that one or more of A53, S54, N55, and S57 can be deleted. If any of Y51, N55, and F56 are deleted in (xi), they cannot be mutated in (i i), and vice versa.

[0153] (i) The mutants include one or more of substitution N40R, N40K, D43N, D43Q, D43R, D4 3K, E44N, E44Q, E44R, E44K, S54P, S57P, Q62R, Q6 2K, R97N, R97G, R97L, E101N, E101Q, E101R, E101 K, E101F, E101Y, E101W, E124N, E124Q, E124R, E1 24K, E124F, E124Y, E124W, E131D, R142E, R142N, T150I, R192E, and R192N, for example, N40R, N40K, D43N, D43Q, D43R, D43K, E44N, E44Q, E44R, E44K, S54P, S57P, Q62R, Q62K, E101N, E101Q, E101R, E1 01K, E101F, E101Y, E101W, E131D, and T150I, or one or more of N40R, N40K, D43N, D43Q, D43R, D43K, E44 N, E44Q, E44R, E44K, E101N, E101Q, E101R, E101K It is preferred to include one or more of E101F, E101Y, E101W, and E131D. The mutants can include any number and combination of these substitutions. In (i), it is preferred that the mutants include S54P and / or S57P. In (i), the mutants are , (a) S54P and / or S57P and (b) Y51, N55, and F56 Preferably, it contains mutants at one or more of them, for example, at Y51, N55, F56, Y51 / N55, Y51 / F56, N55 / F56 or Y51 / N55 / F56. Mutations at one or more of Y 51, N55 and F56 can be any of those considered below. In (i), the mutant preferably contains F56A / S57P or S54P / F5 6A. The mutant preferably contains T150I. Alternatively, the mutant preferably contains (a) T150I and (b) one or more of Y51, N55 and F56, for example Y51, N55, F56, Y51 / N55, Y51 / F56, N55 / F56 or Y51 / N55 / F56.

[0154] In (i), the mutant preferably contains Q62R or Q62K. Alternatively, the mutant contains (a) Q62R or Q62K, and (b) one or more of Y51, N55, F56, Y51 / F5 5, Y51 / F56, N55 / F56, or Y51 / N55 / F56, etc., i.e., one or more of Y51, N 55 and F56. The mutant can contain D4 3N, E44N, Q62R or Q62K or any combination thereof, for example D4 3N, E44N, Q62R, Q62K, D43N / E44N, D43N / Q62R, D4 3N / Q62K, E44N / Q62R, E44N / Q62K, D43N / E44N / Q6 2R or D43N / E44N / Q62K. Alternatively, the mutant contains (a) D4 3N, E44N, Q62R, Q62K, D43N / E44N, D43N / Q62R, D4 3N / Q62K, E44N / Q62R, E44N / Q62K, D43N / E44N / Q6 2R or D43N / E44N / Q62K, and (b) Y51, N55, F56, Y5 1 / N55, Y51 / F56, N55 / F56 or Y51 / N55 / F56, etc., Y5 It preferably contains mutations at one or more of 1, N55 and F56.

[0155] In (i), the variant preferably contains D43N.

[0156] In (i), the variant preferably contains one of E101R, E101S, E101F or E101N and the like.

[0157] In (i), the variant preferably contains one of E124N, E124Q, E124R, E124K, E124F , E124Y, E124W or E124D (such as E124N).

[0158] In (i), the variant preferably contains R142E and R142N.

[0159] In (i), the variant preferably contains one of R97N, R97G or R97L.

[0160] In (i), the variant preferably contains R192E and R192N.

[0161] In (ii), the variant contains F56N / N55Q, F56N / N55R, F56N / N55 K, F56N / N55S, F56N / N55G, F56N / N55A, F56N / N55 T, F56Q / N55Q, F56Q / N55R, F56Q / N55K, F56Q / N55 S, F56Q / N55G, F56Q / N55A, F56Q / N55T, F56R / N55 Q, F56R / N55R, F56R / N55K, F56R / N55S, F56R / N55 G, F56R / N55A, F56R / N55T, F56S / N55Q, F56S / N55 R, F56S / N55K, F56S / N55S, F56S / N55G, F56S / N55 A, F56S / N55T, F56G / N55Q, F56G / N55R, F56G / N55 K, F56G / N55S, F56G / N55G, F56G / N55A, F56G / N55 T, F56A / N55Q, F56A / N55R, F56A / N55K, F56A / N55 S, F56A / N55G, F56A / N55A, F56A / N55T, F56K / N55 Q, F56K / N55R, F56K / N55K, F56K / N55S, F56K / N55 G, F56K / N55A, F56K / N55T, F56N / Y51L, F56N / Y51 V, F56N / Y51A, F56N / Y51N, F56N / Y51Q, F56N / Y51 S, F56N / Y51G, F56Q / Y51L, F56Q / Y51V, F56Q / Y51 A, F56Q / Y51N, F56Q / Y51Q, F56Q / Y51S, F56Q / Y51 G, F56R / Y51L, F56R / Y51V, F56R / Y51A, F56R / Y51 N, F56R / Y51Q, F56R / Y51S, F56R / Y51G, F56S / Y51 L, F56S / Y51V, F56S / Y51A, F56S / Y51N, F56S / Y51 Q, F56S / Y51S, F56S / Y51G, F56G / Y51L, F56G / Y51 V, F56G / Y51A, F56G / Y51N, F56G / Y51Q, F56G / Y51 S, F56G / Y51G, F56A / Y51L, F56A / Y51V, F56A / Y51 A, F56A / Y51N, F56A / Y51Q, F56A / Y51S, F56A / Y51 G, F56K / Y51L, F56K / Y51V, F56K / Y51A, F56K / Y51 N, F56K / Y51Q, F56K / Y51S, F56K / Y51G, N55Q / Y51 L, N55Q / Y51V, N55Q / Y51A, N55Q / Y51N, N55Q / Y51 Q, N55Q / Y51S, N55Q / Y51G, N55R / Y51L, N55R / Y51 V, N55R / Y51A, N55R / Y51N, N55R / Y51Q, N55R / Y51 S, N55R / Y51G, N55K / Y51L, N55K / Y51V, N55K / Y51 A, N55K / Y51N, N55K / Y51Q, N55K / Y51S, N55K / Y51 G, N55S / Y51L, N55S / Y51V, N55S / Y51A, N55S / Y51 N, N55S / Y51Q, N55S / Y51S, N55S / Y51G, N55G / Y51 L, N55G / Y51V, N55G / Y51A, N55G / Y51N, N55G / Y51 Q, N55G / Y51S, N55G / Y51G, N55A / Y51L, N55A / Y51 V, N55A / Y51A, N55A / Y51N, N55A / Y51Q, N55A / Y51 S, N55A / Y51G, N55T / Y51L, N55T / Y51V, N55T / Y51 A, N55T / Y51N, N55T / Y51Q, N55T / Y51S, N55T / Y51 G, F56N / N55Q / Y51L, F56N / N55Q / Y51V, F56N / N55 Q / Y51A, F56N / N55Q / Y51N, F56N / N55Q / Y51Q, F56 N / N55Q / Y51S, F56N / N55Q / Y51G, F56N / N55R / Y51 L, F56N / N55R / Y51V, F56N / N55R / Y51A, F56N / N55 R / Y51N, F56N / N55R / Y51Q, F56N / N55R / Y51S, F56 N / N55R / Y51G、F56N / N55K / Y51L、F56N / N55K / Y51 V、F56N / N55K / Y51A、F56N / N55K / Y51N、F56N / N55 K / Y51Q、F56N / N55K / Y51S、F56N / N55K / Y51G、F56 N / N55S / Y51L、F56N / N55S / Y51V、F56N / N55S / Y51 A、F56N / N55S / Y51N、F56N / N55S / Y51Q、F56N / N55 S / Y51S、F56N / N55S / Y51G、F56N / N55G / Y51L、F56 N / N55G / Y51V、F56N / N55G / Y51A、F56N / N55G / Y51 N、F56N / N55G / Y51Q、F56N / N55G / Y51S、F56N / N55 G / Y51G、F56N / N55A / Y51L、F56N / N55A / Y51V、F56 N / N55A / Y51A、F56N / N55A / Y51N、F56N / N55A / Y51 Q、F56N / N55A / Y51S、F56N / N55A / Y51G、F56N / N55 T / Y51L、F56N / N55T / Y51V、F56N / N55T / Y51A、F56 N / N55T / Y51N、F56N / N55T / Y51Q、F56N / N55T / Y51 S、F56N / N55T / Y51G、F56Q / N55Q / Y51L、F56Q / N55 Q / Y51V、F56Q / N55Q / Y51A、F56Q / N55Q / Y51N、F56 Q / N55Q / Y51Q、F56Q / N55Q / Y51S、F56Q / N55Q / Y51 G、F56Q / N55R / Y51L、F56Q / N55R / Y51V、F56Q / N55 R / Y51A、F56Q / N55R / Y51N、F56Q / N55R / Y51Q、F56 Q / N55R / Y51S、F56Q / N55R / Y51G、F56Q / N55K / Y51 L, F56Q / N55K / Y51V, F56Q / N55K / Y51A, F56Q / N55 K / Y51N, F56Q / N55K / Y51Q, F56Q / N55K / Y51S, F56 Q / N55K / Y51G, F56Q / N55S / Y51L, F56Q / N55S / Y51 V, F56Q / N55S / Y51A, F56Q / N55S / Y51N, F56Q / N55 S / Y51Q, F56Q / N55S / Y51S, F56Q / N55S / Y51G, F56 Q / N55G / Y51L, F56Q / N55G / Y51V, F56Q / N55G / Y51 A, F56Q / N55G / Y51N, F56Q / N55G / Y51Q, F56Q / N55 G / Y51S, F56Q / N55G / Y51G, F56Q / N55A / Y51L, F56 Q / N55A / Y51V, F56Q / N55A / Y51A, F56Q / N55A / Y51 N, F56Q / N55A / Y51Q, F56Q / N55A / Y51S, F56Q / N55 A / Y51G, F56Q / N55T / Y51L, F56Q / N55T / Y51V, F56 Q / N55T / Y51A, F56Q / N55T / Y51N, F56Q / N55T / Y51 Q, F56Q / N55T / Y51S, F56Q / N55T / Y51G, F56R / N55 Q / Y51L, F56R / N55Q / Y51V, F56R / N55Q / Y51A, F56 R / N55Q / Y51N, F56R / N55Q / Y51Q, F56R / N55Q / Y51 S, F56R / N55Q / Y51G, F56R / N55R / Y51L, F56R / N55 R / Y51V, F56R / N55R / Y51A, F56R / N55R / Y51N, F56 R / N55R / Y51Q, F56R / N55R / Y51S, F56R / N55R / Y51 G, F56R / N55K / Y51L, F56R / N55K / Y51V, F56R / N55 K / Y51A, F56R / N55K / Y51N, F56R / N55K / Y51Q, F56 R / N55K / Y51S, F56R / N55K / Y51G, F56R / N55S / Y51 L, F56R / N55S / Y51V, F56R / N55S / Y51A, F56R / N55 S / Y51N, F56R / N55S / Y51Q, F56R / N55S / Y51S, F56 R / N55S / Y51G, F56R / N55G / Y51L, F56R / N55G / Y51 V, F56R / N55G / Y51A, F56R / N55G / Y51N, F56R / N55 G / Y51Q, F56R / N55G / Y51S, F56R / N55G / Y51G, F56 R / N55A / Y51L, F56R / N55A / Y51V, F56R / N55A / Y51 A, F56R / N55A / Y51N, F56R / N55A / Y51Q, F56R / N55 A / Y51S, F56R / N55A / Y51G, F56R / N55T / Y51L, F56 R / N55T / Y51V, F56R / N55T / Y51A, F56R / N55T / Y51 N, F56R / N55T / Y51Q, F56R / N55T / Y51S, F56R / N55 T / Y51G, F56S / N55Q / Y51L, F56S / N55Q / Y51V, F56 S / N55Q / Y51A, F56S / N55Q / Y51N, F56S / N55Q / Y51 Q, F56S / N55Q / Y51S, F56S / N55Q / Y51G, F56S / N55 R / Y51L, F56S / N55R / Y51V, F56S / N55R / Y51A, F56 S / N55R / Y51N, F56S / N55R / Y51Q, F56S / N55R / Y51 S, F56S / N55R / Y51G, F56S / N55K / Y51L, F56S / N55 K / Y51V, F56S / N55K / Y51A, F56S / N55K / Y51N, F56 S / N55K / Y51Q, F56S / N55K / Y51S, F56S / N55K / Y51 G, F56S / N55S / Y51L, F56S / N55S / Y51V, F56S / N55 S / Y51A, F56S / N55S / Y51N, F56S / N55S / Y51Q, F56 S / N55S / Y51S, F56S / N55S / Y51G, F56S / N55G / Y51 L, F56S / N55G / Y51V, F56S / N55G / Y51A, F56S / N55 G / Y51N, F56S / N55G / Y51Q, F56S / N55G / Y51S, F56 S / N55G / Y51G, F56S / N55A / Y51L, F56S / N55A / Y51 V, F56S / N55A / Y51A, F56S / N55A / Y51N, F56S / N55 A / Y51Q, F56S / N55A / Y51S, F56S / N55A / Y51G, F56 S / N55T / Y51L, F56S / N55T / Y51V, F56S / N55T / Y51 A, F56S / N55T / Y51N, F56S / N55T / Y51Q, F56S / N55 T / Y51S, F56S / N55T / Y51G, F56G / N55Q / Y51L, F56 G / N55Q / Y51V, F56G / N55Q / Y51A, F56G / N55Q / Y51 N, F56G / N55Q / Y51Q, F56G / N55Q / Y51S, F56G / N55 Q / Y51G, F56G / N55R / Y51L, F56G / N55R / Y51V, F56 G / N55R / Y51A, F56G / N55R / Y51N, F56G / N55R / Y51 Q, F56G / N55R / Y51S, F56G / N55R / Y51G, F56G / N55 K / Y51L, F56G / N55K / Y51V, F56G / N55K / Y51A, F56 G / N55K / Y51N, F56G / N55K / Y51Q, F56G / N55K / Y51 S, F56G / N55K / Y51G, F56G / N55S / Y51L, F56G / N55 S / Y51V, F56G / N55S / Y51A, F56G / N55S / Y51N, F56 G / N55S / Y51Q, F56G / N55S / Y51S, F56G / N55S / Y51 G, F56G / N55G / Y51L, F56G / N55G / Y51V, F56G / N55 G / Y51A, F56G / N55G / Y51N, F56G / N55G / Y51Q, F56 G / N55G / Y51S, F56G / N55G / Y51G, F56G / N55A / Y51 L, F56G / N55A / Y51V, F56G / N55A / Y51A, F56G / N55 A / Y51N, F56G / N55A / Y51Q, F56G / N55A / Y51S, F56 G / N55A / Y51G, F56G / N55T / Y51L, F56G / N55T / Y51 V, F56G / N55T / Y51A, F56G / N55T / Y51N, F56G / N55 T / Y51Q, F56G / N55T / Y51S, F56G / N55T / Y51G, F56A / N55Q / Y51L, F56A / N55Q / Y51V, F56A / N55Q / Y51A 、F56A / N55Q / Y51N, F56A / N55Q / Y51Q, F56A / N55Q / Y51S, F56A / N55Q / Y51G, F56A / N55R / Y51L, F56A / N55R / Y51V, F56A / N55R / Y51A, F56A / N55R / Y51N 、F56A / N55R / Y51Q, F56A / N55R / Y51S, F56A / N55R / Y51G, F56A / N55K / Y51L, F56A / N55K / Y51V, F56A / N55K / Y51A, F56A / N55K / Y51N, F56A / N55K / Y51Q , F56A / N55K / Y51S, F56A / N55K / Y51G, F56A / N55S / Y51L, F56A / N55S / Y51V, F56A / N55S / Y51A, F56A / N55S / Y51N, F56A / N55S / Y51Q, F56A / N55S / Y51S , F56A / N55S / Y51G, F56A / N55G / Y51L, F56A / N55G / Y51V, F56A / N55G / Y51A, F56A / N55G / Y51N, F56A / N55G / Y51Q, F56A / N55G / Y51S, F56A / N55G / Y51G , F56A / N55A / Y51L, F56A / N55A / Y51V, F56A / N55A / Y51A, F56A / N55A / Y51N, F56A / N55A / Y51Q, F56A / N55A / Y51S, F56A / N55A / Y51G, F56A / N55T / Y51L , F56A / N55T / Y51V, F56A / N55T / Y51A, F56A / N55T / Y51N, F56A / N55T / Y51Q, F56A / N55T / Y51S, F56A / N55T / Y51G, F56K / N55Q / Y51L, F56K / N55Q / Y51V , F56K / N55Q / Y51A, F56K / N55Q / Y51N, F56K / N55Q / Y51Q, F56K / N55Q / Y51S, F56K / N55Q / Y51G, F56K / N55R / Y51L, F56K / N55R / Y51V, F56K / N55R / Y51A , F56K / N55R / Y51N, F56K / N55R / Y51Q, F56K / N55R / Y51S, F56K / N55R / Y51G, F56K / N55K / Y51L, F56K / N55K / Y51V, F56K / N55K / Y51A, F56K / N55K / Y51N , F56K / N55K / Y51Q, F56K / N55K / Y51S, F56K / N55K / Y51G, F56K / N55S / Y51L, F56K / N55S / Y51V, F56K / N55S / Y51A, F56K / N55S / Y51N, F56K / N55S / Y51Q , F56K / N55S / Y51S, F56K / N55S / Y51G, F56K / N55G / Y51L, F56K / N55G / Y51V, F56K / N55G / Y51A, F56K / N55G / Y51N, F56K / N55G / Y51Q, F56K / N55G / Y51S , F56K / N55G / Y51G, F56K / N55A / Y51L, F56K / N55A / Y51V, F56K / N55A / Y51A, F56K / N55A / Y51N, F56K / N55A / Y51Q, F56K / N55A / Y51S, F56K / N55A / Y51G , F56K / N55T / Y51L, F56K / N55T / Y51V, F56K / N55T / Y51A, F56K / N55T / Y51N, F56K / N55T / Y51Q, F56K / N55T / Y51S, F56K / N55T / Y51G, F56E / N55R, F56E / N55K, F56D / N55R, F56D / N55K, F56R / N55E, F56R / N55D, it is preferred to include F56K / N55E or F56K / N55D.

[0162] (ii), the variants are Y51R / F56Q, Y51N / F56N, Y51M / F56 Q, Y51L / F56Q, Y51I / F56Q, Y51V / F56Q, Y51A / F56 Q, Y51P / F56Q, Y51G / F56Q, Y51C / F56Q, Y51Q / F56 Q, Y51N / F56Q, Y51S / F56Q, Y51E / F56Q, Y51D / F56 Preferably, it contains Q, Y51K / F56Q or Y51H / F56Q.

[0163] (ii), the variant preferably contains Y51T / F56Q, Y51Q / F56Q or Y51A / F 56Q.

[0164] (ii), the variant preferably contains Y51T / F56F, Y51T / F56M, Y51T / F56 L, Y51T / F56I, Y51T / F56V, Y51T / F56A, Y51T / F56 P, Y51T / F56G, Y51T / F56C, Y51T / F56Q, Y51T / F56 N, Y51T / F56T, Y51T / F56S, Y51T / F56E, Y51T / F56 D, Y51T / F56K, Y51T / F56H or Y51T / F56R. Preferably.

[0165] (ii), the variant preferably contains Y51T / N55Q, Y51T / N55S or Y51T / N 55A.

[0166] (ii), the variant preferably contains Y51A / F56F, Y51A / F56L, Y51A / F56 I, Y51A / F56V, Y51A / F56A, Y51A / F56P, Y51A / F56 G, Y51A / F56C, Y51A / F56Q, Y51A / F56N, Y51A / F56 T, Y51A / F56S, Y51A / F56E, Y51A / F56D, Y51A / F56 K, Y51A / F56H or Y51A / F56R.

[0167] (ii), the variant preferably contains Y51C / F56A, Y51E / F56A, Y51D / F56 A, Y51K / F56A, Y51H / F56A, Y51Q / F56A, Y51N / F56 Preferably, it contains A, Y51S / F56A, Y51P / F56A or Y51V / F56A. Preferably.

[0168] (xi) In the variant, there are deletions of Y51 / P52, Y51 / P52 / A53, P50 - P52, P50 - A53, K49 - Y51, K49 - A53 and substitution with a single proline (P), substitution with K49 - S54 and a single P, Y51 - A53, Y51 - S54, N 55 / F56, N55 - S57, N55 / F56 and substitution with a single P, N55 / F5 6 and substitution with a single glycine (G), N55 / F56 and substitution with a single alanine (A) substitution, N55 / F56 and substitution with a single P and Y51N, N55 / F56 and substitution with a single P and Y51Q, N55 / F56 and substitution with a single P and Y51S substitution, N55 / F56 and substitution with a single G and Y51N, N55 / F56 and single substitution with G and Y51Q, N55 / F56 and substitution with a single G and Y51S, N55 / F56 and substitution with a single A and Y51N, N55 / F56 and single substitution with A / Y51Q, or N55 / F56 and substitution with a single A and Y51S Preferably, it includes substitution. Preferably.

[0169] The variants are D195N / E203N, D195Q / E203N, D195N / E203Q , D195Q / E203Q, E201N / E203N, E201Q / E203N, E20 1N / E203Q, E201Q / E203Q, E185N / E203Q, E185Q / E 203Q, E185N / E203N, E185Q / E203N, D195N / E201N / E203N, D195Q / E201N / E203N, D195N / E201Q / E20 3N, D195N / E201N / E203Q, D195Q / E201Q / E203N, D 195Q / E201N / E203Q, D195N / E201Q / E203Q, D195Q / E201Q / E203Q, D149N / E201N, D149Q / E201N, D14 9N / E201Q, D149Q / E201Q, D149N / E201N / D195N, D 149Q / E201N / D195N, D149N / E201Q / D195N, D149N / E201N / D195Q, D149Q / E201Q / D195N, D149Q / E20 1N / D195Q, D149N / E201Q / D195Q, D149Q / E201Q / D 195Q, D149N / E203N, D149Q / E203N, D149N / E203Q 、D149Q / E203Q, D149N / E185N / E201N, D149Q / E18 5N / E201N, D149N / E185Q / E201N, D149N / E185N / E 201Q, D149Q / E185Q / E201N, D149Q / E185N / E201Q 、D149N / E185Q / E201Q, D149Q / E185Q / E201Q, D14 9N / E185N / E203N, D149Q / E185N / E203N, D149N / E 185Q / E203N, D149N / E185N / E203Q, D149Q / E185Q / E203N, D149Q / E185N / E203Q, D149N / E185Q / E20 3Q, D149Q / E185Q / E203Q, D149N / E185N / E201N / E 203N, D149Q / E185N / E201N / E203N, D149N / E185Q / E201N / E203N, D149N / E185N / E201Q / E203N, D14 9N / E185N / E201N / E203Q, D149Q / E185Q / E201N / E 203N, D149Q / E185N / E201Q / E203N, D149Q / E185N / E201N / E203Q, D149N / E185Q / E201Q / E203N, D14 9N / E185Q / E201N / E203Q, D149N / E185N / E201Q / E 203Q, D149Q / E185Q / E201Q / E203Q, D149Q / E185Q / E201N / E203Q, D149Q / E185N / E201Q / E203Q, D14 9N / E185Q / E201Q / E203Q, D149Q / E185Q / E201Q / E 203N, D149N / E185N / D195N / E201N / E203N, D149Q / E185N / D195N / E201N / E203N, D149N / E185Q / D19 5N / E201N / E203N, D149N / E185N / D195Q / E201N / E 203N, D149N / E185N / D195N / E201Q / E203N, D149N / E185N / D195N / E201N / E203Q, D149Q / E185Q / D19 5N / E201N / E203N, D149Q / E185N / D195Q / E201N / E 203N, D149Q / E185N / D195N / E201Q / E203N, D149Q / E185N / D195N / E201N / E203Q, D149N / E185Q / D19 5Q / E201N / E203N, D149N / E185Q / D195N / E201Q / E 203N, D149N / E185Q / D195N / E201N / E203Q, D149N / E185N / D195Q / E201Q / E203N, D149N / E185N / D19 5Q / E201N / E203Q, D149N / E185N / D195N / E201Q / E 203Q, D149Q / E185Q / D195Q / E201N / E203N, D149Q / E185Q / D195N / E201Q / E203N, D149Q / E185Q / D19 5N / E201N / E203Q, D149Q / E185N / D195Q / E201Q / E 203N, D149Q / E185N / D195Q / E201N / E203Q, D149Q / E185N / D195N / E201Q / E203Q, D149N / E185Q / D19 5Q / E201Q / E203N, D149N / E185Q / D195Q / E201N / E 203Q, D149N / E185Q / D195N / E201Q / E203Q, D149N / E185N / D195Q / E201Q / E203Q, D149Q / E185Q / D19 5Q / E201Q / E203N, D149Q / E185Q / D195Q / E201N / E 203Q, D149Q / E185Q / D195N / E201Q / E203Q, D149Q / E185N / D195Q / E201Q / E203Q, D149N / E185Q / D19 5Q / E201Q / E203Q, D149Q / E185Q / D195Q / E201Q / E 203Q, D149N / E185R / E201N / E203N, D149Q / E185R / E201N / E203N, D149N / E185R / E201Q / E203N, D14 9N / E185R / E201N / E203Q, D149Q / E185R / E201Q / E 203N, D149Q / E185R / E201N / E203Q, D149N / E185 R / E201Q / E203Q, D149Q / E185R / E201Q / E203Q, D1 49R / E185N / E201N / E203N, D149R / E185Q / E201N / E203N, D149R / E185N / E201Q / E203N, D149R / E185 N / E201N / E203Q, D149R / E185Q / E201Q / E203N, D1 49R / E185Q / E201N / E203Q, D149R / E185N / E201Q / E203Q, D149R / E185Q / E201Q / E203Q, D149R / E185 N / D195N / E201N / E203N, D149R / E185Q / D195N / E2 01N / E203N, D149R / E185N / D195Q / E201N / E203N, D149R / E185N / D195N / E201Q / E203N, D149R / E185 Q / D195N / E201N / E203Q, D149R / E185Q / D195Q / E2 01N / E203N, D149R / E185Q / D195N / E201Q / E203N, D149R / E185Q / D195N / E201N / E203Q, D149R / E185 N / D195Q / E201Q / E203N, D149R / E185N / D195Q / E2 01N / E203Q, D149R / E185N / D195N / E201Q / E203Q, D149R / E185Q / D195Q / E201Q / E203N, D149R / E185 Q / D195Q / E201N / E203Q, D149R / E185Q / D195N / E2 01Q / E203Q, D149R / E185N / D195Q / E201Q / E203Q, D149R / E185Q / D195Q / E201Q / E203Q, D149N / E185 R / D195N / E201N / E203N, D149Q / E185R / D195N / E2 01N / E203N, D149N / E185R / D195Q / E201N / E203N, D149N / E185R / D195N / E201Q / E203N, D149N / E185 R / D195N / E201N / E203Q, D149Q / E185R / D195Q / E2 01N / E203N, D149Q / E185R / D195N / E201Q / E203N, D149Q / E185R / D195N / E201N / E203Q, D149N / E185 R / D195Q / E201Q / E203N, D149N / E185R / D195Q / E2 01N / E203Q, D149N / E185R / D195N / E201Q / E203Q, D149Q / E185R / D195Q / E201Q / E203N, D149Q / E185 R / D195Q / E201N / E203Q, D149Q / E185R / D195N / E2 01Q / E203Q, D149N / E185R / D195Q / E201Q / E203Q, D149Q / E185R / D195Q / E201Q / E203Q, D149N / E185 R / D195N / E201R / E203N, D149Q / E185R / D195N / E2 01R / E203N, D149N / E185R / D195Q / E201R / E203N, D149N / E185R / D195N / E201R / E203Q, D149Q / E185 R / D195Q / E201R / E203N, D149Q / E185R / D195N / E2 01R / E203Q, D149N / E185R / D195Q / E201R / E203Q, D149Q / E185R / D195Q / E201R / E203Q, E131D / K49R , E101N / N102F, E101N / N102Y, E101N / N102W, E10 1F / N102F, E101F / N102Y, E101F / N102W, E101Y / N 102F, E101Y / N102Y, E101Y / N102W, E101W / N102F , E101W / N102Y, E101W / N102W, E101N / N102R, E10 1F / N102R, E101Y / N102R or E101W / N102F is preferably included. More preferably.

[0170] When the polynucleotide moves through the pore, fewer nucleotides contribute to the current Preferred mutants of the present invention that form pores are Y51A / F56A, Y51A / F56N, Y51I / F56A, Y51L / F56A, Y51T / F56A, Y51I / F56N, Y51L / F56N or Y51T / F56N, or more preferably Y51I / F56A, Y51L / F56A or Y51T / F56A. As described above, this facilitates the identification of the direct relationship between the observed currents (when the polynucleotide moves through the pore and the polynucleotide).

[0171] Preferred mutants that form pores displaying an increased range include mutations at the following positions : Y51, F56, D149, E185, E201 and E203, N55 and F56, Y51 and F56, Y51, N55 and F56, or F56 and N102.

[0172] Preferred mutants that form pores displaying an increased range are Y51N, F56A, D149N, E185R, E201N and E203N, N55S and F56Q, Y51A and F56A, Y51A and F56N, Y51I and F56A, Y51L and F56A, Y51T and F56A, Y51I and F56N, Y51L and F56N, Y51T and F56N, Y51T and F56Q, Y51A, N55S and F56A, Y51A, N55S, and F56N, Y51T, N55S and F56Q, or F56Q and N102R.

[0173] When the polynucleotide moves through the pore, a pore that forms with few nucleotides contributing to the current preferred mutants include mutations at the following positions: N55 and F56 (such as N55X and F56Q), where X is any amino acid , Y51 and F56 (such as Y51X and F56Q), where X is any amino acid is.

[0174] Particularly preferred mutants include Y51A and F56Q.

[0175] Preferred variants that form pores displaying increased throughput include mutations at the following positions : D149, E185 and E203, D149, E185, E201, and E203, or D149, E185, D195, E201, and E203.

[0176] Preferred mutants that form pores displaying increased throughput are D149N, E185N and E203N, D149N, E185N, E201N and E203N, D149N, E185R, D195N, E201N and E203N, or D149N, E185R, D195N, E201R and E203N.

[0177] Preferred mutants that form pores that increase the capture of polynucleotides have the following mutations include: D43N / Y51T / F56Q, E44N / Y51T / F56Q, D43N / E44N / Y51T / F56Q, Y51T / F56Q / Q62R, D43N / Y51T / F56Q / Q62R, E44N / Y51T / F56Q / Q62R, or D43N / E44N / Y51T / F56Q / Q62R.

[0178] Preferred mutants include the following mutants: D149R / E185R / E201R / E203R or Y51T / F56Q / D149 R / E185R / E201R / E203R, D149N / E185N / E201N / E203N or Y51T / F56Q / D149 N / E185N / E201N / E203N, E201R / E203R or Y51T / F56Q / E201R / E203R E201N / E203R or Y51T / F56Q / E201N / E203R, E203R or Y51T / F56Q / E203R, E203N or Y51T / F56Q / E203N, E201R or Y51T / F56Q / E201R, E201N or Y51T / F56Q / E201N, E185R or Y51T / F56Q / E185R, E185N or Y51T / F56Q / E185N, D149R or Y51T / F56Q / D149R, D149N or Y51T / F56Q / D149N, R142E or Y51T / F56Q / R142E, R142N or Y51T / F56Q / R142N, R192E or Y51T / F56Q / R192E, or R192N or Y51T / F56Q / R192N.

[0179] Preferred variants include the following mutants: Y51A / F56Q / E101N / N102R, Y51A / F56Q / R97N / N102G, Y51A / F56Q / R97N / N102R, Y51A / F56Q / R97N, Y51A / F56Q / R97G, Y51A / F56Q / R97L, Y51A / F56Q / N102R, Y51A / F56Q / N102F, Y51A / F56Q / N102G, Y51A / F56Q / E101R, Y51A / F56Q / E101F, Y51A / F56Q / E101N, or Y51A / F56Q / E101G

[0180] The variant preferably further comprises a mutation at T150. The bacteria showing increased insertion. Preferred variants forming pores include T150I. Mutations at T150 (such as T150I) may be combined

[0181] with any of the above-described mutations or combinations of mutations. Preferred variants of SEQ ID NO: 3 include (a) R97W and (b) a mutation at Y51 and / or F5 comprises. Preferred variants of SEQ ID NO: 3 are (a) R97W, and (b) Y51L / V / A / N / Q / S / G and / or F56A / Q / N. Preferred variants of SEQ ID NO: 3 comprise (a) R97W, and (b) Y51A and / or F56Q. SEQ ID NO: 3 Preferred variants thereof comprise R97W, Y51A and F56Q.

[0182] Variants of SEQ ID NO: 3 preferably contain a mutation at R192. The variant contains R19 2D / Q / F / S / T / N / E, R192D / Q / F / S / T or R192D / Q preferably. Preferred variants of SEQ ID NO: 3 are (a) R97W, (b) Y51 and or mutations at F56, and (c) mutations at R192, such as R192D / Q / F / S / T / N / E, R192D / Q / F / S / T or R192D / Q. Preferred variants of SEQ ID NO: 3 are (a) R97W, (b) Y51R / H / K / D / E / S / T / N / Q / C / G / P / A / V / I / L / M and / or F56 R / H / K / D / E / S / T / N / Q / C / G / P / A / V / I / L / M, and (c) mutations at R192, such as R192D / Q / F / S / T / N / E, R192D / Q / F / S / T or R19 2D / Q. Preferred variants of SEQ ID NO: 3 are (a) R97W, (b) Y51L / V / A / N / Q / S / G and / or F56A / Q / N, and (c) mutations at R192, such as R192D / Q / F / S / T / N / E, R192D / Q / F / S / T or R 192D / Q. Preferred variants of SEQ ID NO: 3 are (a) R97W, (b) Y5 1A and / or F56Q, and (c) mutations at R192, such as R192 D / Q / F / S / T / N / E, R192D / Q / F / S / T or R192D / Q. Preferred variants of SEQ ID NO: 3 are (a) R97W, (b) Y51R / H / K / D / E / S / T / N / Q / C / G / P / A / V / I / L / M and / or F56 R / H / K / D / E / S / T / N / Q / C / G / P / A / V / I / L / M, and (c) mutations at R192, such as R192D / Q / F / S / T / N / E, R192D / Q / F / S / T or R192D / Q. Preferred variants of SEQ ID NO: 3 are (a) R97W, (b) Y5 1A and / or F56Q, and (c) mutations at R192, such as R192 D / Q / It includes F / S / T / N / E, R192D / Q / F / S / T, or R192D / Q, etc. The preferred variants of SEQ ID NO: 3 include R97W, Y51A, F56Q, and R192D / Q / F / S / T, or R192D / Q. The preferred variants of SEQ ID NO: 3 include R97W, Y5 1A, F56Q, and R192D. The preferred variants of SEQ ID NO: 3 include R97W, Y 51A, F56Q, and R192Q. In paragraphs where different amino acids are separated by the / symbol at a specific position, the / symbol means "or". For example, R192D / Q means R 192D or R192Q.

[0183] Any of the above preferred variants of SEQ ID NO: 3 may further include a mutation at R93 . The preferred variants of SEQ ID NO: 3 include (a) R93W, and (b) mutations at Y51 and / or F56, preferably Y51A and F56Q.

[0184] Any of the above preferred variants of SEQ ID NO: 3 may include the K94N / Q mutation. Any of the above preferred variants of SEQ ID NO: 3 may include the F191T mutation.

[0185] The CsgG monomer can be modified to promote attachment to the CsgF peptide. For example , cysteine residues can be introduced at one or more positions corresponding to positions 132, 133, 136, 138, 140, 14 2, 144, 145, 147, 149, 151, 153, 155, 183, 185, 18 7, 189, 191, 201, 203, 205, 207, and 209 of SEQ ID NO: 3, and / or at any one of the positions described in Table 4 where it is predicted to promote covalent bonding to CsgG upon contact with CsgF. Through cysteine residues As an alternative or in addition to covalent binding, the pores can be stabilized by hydrophobic or electrostatic interactions. To promote such interactions, at a position corresponding to one or more of positions 132, 133, 136, 138, 140, 142, 144, 145, 147, 149, 151, 153, 155, 183, 185, 187, 189, 191, 201, 203, 205, 207 and 209 of SEQ ID NO: 3, and / or at any one of the positions described in Table 4 predicted to contact CsgF, a non-native reactive or photoreactive amino acid is.

[0186] Preferred exemplary pores include at least one CsgG monomer having the following mutations relative to SEQ ID NO: 3: Y51X1 / N55X2 / F56X3 / N91R / K9 4Q / R97W / R192D-del(V105-I107), where X1 is I / V / S / T, X2 is N / I / V / S / T and / or X3 is Q / I / V / S / T.

[0187] Methods for introducing or substituting natural amino acids are well known in the art. For example, methionine (M) can be substituted with arginine (R) by substituting the codon for methionine (ATG) with the codon for arginine (CGT) at the relevant position of the polynucleotide encoding the mutant monomer. Next, the polynucleotide can be found as described below.

[0188] Double pore The CsgG / CsgF pore may be a double pore comprising a first pore and a second pore. At least the first pore is the CsgG / CsgF pore disclosed herein. The second pore may be a CsgG pore or a CsgG / CsgF pore. In one embodiment both the first pore and the second pore are CsgG / C sgF pores as disclosed herein. The first and second pores may be the same or different. In addition to any of the mutations disclosed in this specification, in a double pore, the CsgG monomer may include one or more of the additional mutations described below.

[0189] In a double pore, the first pore may be attached to the second CsgG pore by hydrophobic interactions and / or one or more disulfide bonds. One or more of the monomers within the first pore and / or the second pore, such as 2, 3, 4, 5, 6, 8, 9, etc., for example, all, may be modified to strengthen such interactions. This can be achieved in any suitable way.

[0190] At least one cysteine residue in the amino acid sequence of the first pore at the interface between the first pore and the second pore may be disulfide - bonded to at least one cysteine residue in the amino acid sequence of the second pore at the interface between the first pore and the second pore. The cysteine residues of the first pore and / or cysteine residues within the second pore may be cysteine residues that do not exist in the wild - type CsgG monomer. Multiple disulfide bonds, such as 2, 3, 4, 5, 6, 7 , 8 or 9 - 16, 18, 24, 27, 32, 36, 40, 45, 48, 54, 56 or 63, etc., may be formed between the two pores within the two pores. One or both of the first or second pores are at positions corresponding to R97, I107, R110, Q100, E101, N102 and / or L113 of SEQ ID NO: 3 at the first pore and the second pore​​​​​​​​ At least one monomer, such as a maximum of 8, 9 or 10 monomers, containing cysteine residues at the interface between the two pores, may be included.

[0191] At least one monomer in at least one of the first pore and / or the second pore may contain at least one residue at the interface between the first pore and the second pore. Preferably, this residue is more hydrophobic than the residue present at the corresponding position in the wild-type CsgG monomer. For example, 2 to 10 residues, such as 3, 4, 5, 6, 7, 8, or 9 residues, may be such that the residues in the first pore and / or the second pore are more hydrophobic than the residues at the same position in the corresponding wild-type CsgG monomer. Such hydrophobic residues strengthen the interaction between the two pores of the double pore. At least one residue at the interface between the first pore and the second pore may be at a position corresponding to R97, I107, R110, Q100, E101, N102, and L113 of SEQ ID NO: 3. When the residue at the interface in the wild-type CsgG monomer is R, Q, N, or E, the hydrophobic residue is generally I, L, V, M, F, W, or Y. When the residue at the interface in the wild-type CsgG monomer is I, the hydrophobic residue is generally L, V, M, F, W, or Y. When the residue at the interface in the wild-type CsgG monomer is L, the hydrophobic residue is generally I, V, M, F, W, or Y.

[0192] The double pore may contain one or more monomers containing one or more cysteine residues at the interface between the pores, or may contain one or more hydrophobic residues at the interface between the pores, or may contain one or more monomers containing such cysteine residues and such hydrophobic residues. For example, ​​​​​​​​​​​​R97, I107, R110, Q100, E101, N102 of SEQ ID NO:3 and / or one or more (any 2, 3, or 4, etc.) of the positions within the monomer corresponding to the position of L113 can include a cysteine (C) residue, and one or more (any 2, 3, or 4, etc.) of the positions within the monomer corresponding to the positions of R97, I107, R110, Q100, E101, N102, and / or L113 of SEQ ID NO:3 can include a hydrophobic residue such as I, L, V, M, F, W, or Y.

[0193] Thus, the double pore can include a bulging residue at one or more (such as 2, 3, 4, 5, 6, or 7, etc.) positions within the tail region, and this residue is generally at the interface between the first and second pores and is more bulging than the residue at the corresponding position in the wild-type CsgG monomer. The size of these residues prevents pores from forming in the pore wall at the interface between the first and second pores of the double pore. At least one bulging residue at the interface between the first and second pores is generally at a position corresponding to A98, A99, T104, V105, L113, Q114, or S115 of SEQ ID NO:3. When the residue at the interface in the wild-type CsgG monomer is A the bulging residue is generally I, L, V, M, F, W, Y, N, Q, S, or T. When the residue at the interface in the wild-type CsgG monomer is T, the bulging residue is generally L, M F, W, Y, N, Q, R, D, or E. When the residue at the interface in the wild-type CsgG monomer is V the bulging residue is generally I, L, M, F, W, Y, N, Q. When the residue at the interface in the wild-type CsgG monomer is L the bulging residue is generally M, F, W Y, N, Q, R, D, or E. When the residue at the interface in the wild-type CsgG monomer is Q in the wild-type CsgG monomer, the bulging residue is generally... When this is the case, the bulky residue is generally F, W or Y. When the residue at the interface within the wild-type CsgG monomer is S, the bulky residue is generally M, F, W, Y, N, Q, E or R When this is the case, the bulky residue is generally M, F, W, Y, N, Q, E or R When this is the case, the bulky residue is generally M, F, W, Y, N, Q, E or R

[0194] When the second pore is outside the membrane, the second pore, and optionally the first pore, preferably contain residues within the barrel region of the pore that reduce the negative charge within the barrel compared to the charge within the barrel of the wild-type CsgG pore. These mutations make the barrel more hydrophilic. At least one monomer within the first pore of the double pore and / or at least one monomer within the second pore may contain at least one residue within the barrel region of the pore, but these residues have less negative charge than the residues at the corresponding positions within the wild-type CsgG monomer. The charge within the barrel is sufficiently neutral or positively charged so that a charged analyte such as a polynucleotide does not enter the pore by electrostatic charge. At least one residue (such as 2, 3, 4 or 5 residues) within the barrel region of the pore corresponding to D149, E185, D195, E210 and / or E203 of SEQ ID NO: 3 may be a neutral or positively charged amino acid. At least one residue (such as 2, 3, 4 or 5 residues) within the barrel region of the pore corresponding to D149, E185, D195, E210 and / or E203 of SEQ ID NO: 3 is preferably N, Q, R or K. The charge within the barrel is sufficiently neutral or positively charged so that a charged analyte such as a polynucleotide does not enter the pore by electrostatic charge. At least one residue (such as 2, 3, 4 or 5 residues) within the barrel region of the pore corresponding to D149, E185, D195, E210 and / or E203 of SEQ ID NO: 3 may be a neutral or positively charged amino acid. At least one residue (such as 2, 3, 4 or 5 residues) within the barrel region of the pore corresponding to D149, E185, D195, E210 and / or E203 of SEQ ID NO: 3 is preferably N, Q, R or K. The charge within the barrel is sufficiently neutral or positively charged so that a charged analyte such as a polynucleotide does not enter the pore by electrostatic charge. At least one residue (such as 2, 3, 4 or 5 residues) within the barrel region of the pore corresponding to D149, E185, D195, E210 and / or E203 of SEQ ID NO: 3 may be a neutral or positively charged amino acid. At least one residue (such as 2, 3, 4 or 5 residues) within the barrel region of the pore corresponding to D149, E185, D195, E210 and / or E203 of SEQ ID NO: 3 is preferably N, Q, R or K. The charge within the barrel is sufficiently neutral or positively charged so that a charged analyte such as a polynucleotide does not enter the pore by electrostatic charge. At least one residue (such as 2, 3, 4 or 5 residues) within the barrel region of the pore corresponding to D149, E185, D195, E210 and / or E203 of SEQ ID NO: 3 may be a neutral or positively charged amino acid. At least one residue (such as 2, 3, 4 or 5 residues) within the barrel region of the pore corresponding to D149, E185, D195, E210 and / or E203 of SEQ ID NO: 3 is preferably N, Q, R or K. The charge within the barrel is sufficiently neutral or positively charged so that a charged analyte such as a polynucleotide does not enter the pore by electrostatic charge. At least one residue (such as 2, 3, 4 or 5 residues) within the barrel region of the pore corresponding to D149, E185, D195, E210 and / or E203 of SEQ ID NO: 3 may be a neutral or positively charged amino acid. At least one residue (such as 2, 3, 4 or 5 residues) within the barrel region of the pore corresponding to D149, E185, D195, E210 and / or E203 of SEQ ID NO: 3 is preferably N, Q, R or K. The charge within the barrel is sufficiently neutral or positively charged so that a charged analyte such as a polynucleotide does not enter the pore by electrostatic charge. At least one residue (such as 2, 3, 4 or 5 residues) within the barrel region of the pore corresponding to D149, E185, D195, E210 and / or E203 of SEQ ID NO: 3 may be a neutral or positively charged amino acid. At least one residue (such as 2, 3, 4 or 5 residues) within the barrel region of the pore corresponding to D149, E185, D195, E210 and / or E203 of SEQ ID NO: 3 is preferably N, Q, R or K. The charge within the barrel is sufficiently neutral or positively charged so that a charged analyte such as a polynucleotide does not enter the pore by electrostatic charge. At least one residue (such as 2, 3, 4 or 5 residues) within the barrel region of the pore corresponding to D149, E185, D195, E210 and / or E203 of SEQ ID NO: 3 may be a neutral or positively charged amino acid. At least one residue (such as 2, 3, 4 or 5 residues) within the barrel region of the pore corresponding to D149, E185, D195, E210 and / or E203 of SEQ ID NO: 3 is preferably N, Q, R or K. The charge within the barrel is sufficiently neutral or positively charged so that a charged analyte such as a polynucleotide does not enter the pore by electrostatic charge. At least one residue (such as 2, 3, 4 or 5 residues) within the barrel region of the pore corresponding to D149, E185, D195, E210 and / or E203 of SEQ ID NO: 3 may be a neutral or positively charged amino acid. At least one residue (such as 2, 3, 4 or 5 residues) within the barrel region of the pore corresponding to D149, E185, D195, E210 and / or E203 of SEQ ID NO: 3 is preferably N, Q, R or K. The charge within the barrel is sufficiently neutral or positively charged so that a charged analyte such as a polynucleotide does not enter the pore by electrostatic charge. At least one residue (such as 2, 3, 4 or 5 residues) within the barrel region of the pore corresponding to D149, E185, D195, E210 and / or E203 of SEQ ID NO: 3 may be a neutral or positively charged amino acid. At least one residue (such as 2, 3, 4 or 5 residues) within the barrel region of the pore corresponding to D149, E185, D195, E210 and / or E203 of SEQ ID NO: 3 is preferably N, Q, R or K. The charge within the barrel is sufficiently neutral or positively charged so that a charged analyte such as a polynucleotide does not enter the pore by electrostatic charge. At least one residue (such as 2, 3, 4 or 5 residues) within the barrel region of the pore corresponding to D149, E185, D195, E210 and / or E203 of SEQ ID NO: 3 may be a neutral or positively charged amino acid. At least one residue (such as 2, 3, 4 or 5 residues) within the barrel region of the pore corresponding to D149, E185, D195, E210 and / or E203 of SEQ ID NO: 3 is preferably N, Q, R or K. The charge within the barrel is sufficiently neutral or positively charged so that a charged analyte such as a polynucleotide does not enter the pore by electrostatic charge. At least one residue (such as 2, 3, 4 or 5 residues) within the barrel region of the pore corresponding to D149, E185, D195, E210 and / or E203 of SEQ ID NO: 3 may be a neutral or positively charged amino acid. At least one residue (such as 2, 3, 4 or 5 residues) within the barrel region of the pore corresponding to D149, E185, D195, E210 and / or E203 of SEQ ID NO: 3 is preferably N, Q, R or K. The charge within the barrel is sufficiently neutral or positively charged so that a charged analyte such as a polynucleotide does not enter the pore by electrostatic charge. At least one residue (such as 2, 3, 4 or 5 residues) within the barrel region of the pore corresponding to D149, E185, D195, E210 and / or E203 of SEQ ID NO: 3 may be a neutral or positively charged amino acid. At least one residue (such as 2, 3, 4 or 5 residues) within the barrel region of the pore corresponding to D149, E185, D195, E210 and / or E203 of SEQ ID NO: 3 is preferably N, Q, R or K. The charge within the barrel is sufficiently neutral or positively charged so that a charged analyte such as a polynucleotide does not enter the pore by electrostatic charge. At least one residue (such as 2, 3, 4 or 5 residues) within the barrel region of the pore corresponding to D149, E185, D195, E210 and / or E203 of SEQ ID NO: 3 may be a neutral or positively charged amino acid. At least one residue (such as 2, 3, 4 or 5 residues) within the barrel region of the pore corresponding to D149, E185, D195, E210 and / or E203 of SEQ ID NO: 3 is preferably N, Q, R or K.

[0195] Specific examples of mutations that remove the charge at SEQ ID NO: 3 include: E185N / E203N, D149N / E185R / D195N / E201R / E203N, D149N / E185R / D195N / E201N / E203N, D149R / E185N / D195 N / E201N / E203N, D149R / E185N / E201N / E203N, D1 49N / E185N / D195 / E201N / E203N, D149N / E185N / E 201N / E203N, D149N / E185N / E203N, D149N / E185N / E201N, D149N / E203N, D149N / E201N / D195N, D14 9N / E201N, D195N / E201N / E203N, E201N / E203N, D 195N / E203, E203R, E203N, E201R, E201N, D195R, D195N, E185R, E185N, D149R and D149N.

[0196] At least one CsgG monomer in the first pore is located at the constriction of the barrel region of the first pore. and at least one residue in the CsgG pore that is more compact than the wild-type CsgG pore. Reducing, maintaining, or increasing the length of the constriction and / or at least Another CsgG monomer contains at least one residue in the constriction of the barrel region of the second pore. This residue may comprise a maintenance residue which reduces the length of the constriction compared to the wild-type CsgG pore. maintain or increase the length of the constriction in the first pore and / or the constriction in the second pore The length of the portion is at least as long as the wild-type pore, and more preferably is longer. It is.

[0197] The length of the pore is increased by inserting residues in the region corresponding to the region between positions K49 and F56 of SEQ ID NO:3. The amino acid sequence may be increased by adding 1 to 5 (such as 2, 3, or 4) amino acids. An acid residue may be present at any one or more of the following positions as defined by reference to SEQ ID NO:3: It may be inserted at the position of. K49 and P50, P50 and Y51, Y51 and P5 2, P52 and A53, A53 and S54, S54 and N55 and / or N5 5 and F56. In total, 1 to 10 (such as 2 to 8) or 3 to 5 amino acid residues are preferably inserted into the sequence of the monomer. All monomers in the first pore and / or or all monomers in the second pore preferably have the same number of insertions in this region. The inserted residues can increase the length of the loop between the residues corresponding to Y51 and N55 of SEQ ID NO: 3 . The inserted residues are A, S, G or T for maintaining flexibility, P for adding twist to the loop, and / or S, T, N, Q, M, F, W, Y when the analyte interacts with the barrel of the pore under the applied potential difference. Any combination of V and / or I may be used. The inserted amino acids may be any combination of S, G, S GG, SGS, GS, GSS and / or GSG. In the double pore, the constriction of the barrel of the first pore and / or the second pore, when used for detection or characterization of the analyte, has at least one residue (2, 3, 4 or 5 residues, etc.) that affects the properties of the pore compared to when a first or second pore with a wild - type constriction is used, where at least one residue in the constriction of the barrel region of the pore is at a position corresponding to Y51, N55, Y51, P52 and / or A53 of SEQ ID NO: 3. At least one residue is Q or V at the position corresponding to F56 of SEQ ID NO: 3, A or Q at the position corresponding to Y51 of SEQ ID NO: 3, and / or of SEQ ID NO: 3

[0198] In the double pore, the constriction of the barrel of the first pore and / or the second pore, when used for detection or characterization of the analyte, has at least one residue (2, 3, 4 or 5 residues, etc.) that affects the properties of the pore compared to when a first or second pore with a wild - type constriction is used, where at least one residue in the constriction of the barrel region of the pore is at a position corresponding to Y51, N55, Y51, P52 and / or A53 of SEQ ID NO: 3. At least one residue is Q or V at the position corresponding to F56 of SEQ ID NO: 3, A or Q at the position corresponding to Y51 of SEQ ID NO: 3, and / or of SEQ ID NO: 3 one residue is at a position corresponding to Y51, N55, Y51, P52 and / or A53 of SEQ ID NO: 3. At least one residue is Q or V at the position corresponding to F56 of SEQ ID NO: 3, A or Q at the position corresponding to Y51 of SEQ ID NO: 3, and / or of SEQ ID NO: 3 V at the position corresponding to F56 of SEQ ID NO: 3, A or Q at the position corresponding to Y51 of SEQ ID NO: 3, and / or of SEQ ID NO: 3 It may also be V at the position corresponding to N55.

[0199] The double pore may contain at least one monomer in the first CsgG pore and / or at least one monomer in the second C sgG pore, where the monomer contains two or more mutants defined above.

[0200] The CsgG monomer in the double pore may contain a cysteine residue at a position corresponding to R97, I107, R110, Q100 of SEQ ID NO: 3, E101, N102, or L113.

[0201] The CsgG monomer in the double pore may contain a residue at a position corresponding to any one or more of R97, Q100, I107, R110 of SEQ ID NO: 3, E101, N102, and L113, and this residue is more hydrophobic than the residue present at the corresponding position (such as the position corresponding to any one of SEQ ID NOs: 68 - 88) of SEQ ID NO: 3, where the residue at the position corresponding to R97 and / or I107 is M, the residue at the position corresponding to R110 is I , L, V, M, W, or Y, and / or the residue at the position corresponding to E101 or N102 is V or M. The residue at the position corresponding to Q100 is generally I, L, V, M, F, W, or Y, and / or the residue at the position corresponding to L113 is generally I, V, M, F, W, or Y. A specific monomer is Y51A, F56Q substitution and R97I / V / L / M / F / W / Y, I107L / V / M / F / W / Y, R110I / V / L / M / F / W / Y, Q100I / V / L / M / F / W / Y, E101I / V / L / M / F / W / Y, N102I / V / L /

[0202] I107L / V / M / F / W / Y, R110I / V / L / M / F / W / Y, Q100I / V / L / M / F / W / Y, E101I / V / L / M / F / W / Y, N102I / V / L / ​​​Combinations of M / F / W / Y and L113CI / V / L / M / F / W / Y, R97I / Combinations of V / L / M / F / W / Y and N102I / V / L / M / F / W / Y, and / or Combinations of R97I / V / L / M / F / W / Y and E101I / V / L / M / F / W / Y, and has the sequence shown in SEQ ID NO: 3. I107 may already form a hydrophobic interaction between two pores.

[0203] The CsgG monomer in at least one of the double pores may contain residues that are bulkier than the residues at the corresponding positions in SEQ ID NO: 3, corresponding to any one or more of A98, A9 9, T104, V105, L113, Q114, and S115 (such as any one of the corresponding positions in SEQ ID NOS: 68 - 88). Here, the residue at the position corresponding to T104 is L, M, F, W, Y, N, Q, D, or E, the residue at the position corresponding to L113 is M, F, W, Y, N, G, D, or E, and / or the residue at the position corresponding to S115 is M, F, W, Y, N, Q, or E. The residue at the position corresponding to A98 or A99 is generally I, L, V, M, F, W, Y, N, Q, S, or T The residue at the position corresponding to V105 is I, L, M, F, W, Y, N, or Q The residue at the position corresponding to Q114 is F, W, or Y. The residue at the position corresponding to E210 is N, Q, R, or K. The residue at the position corresponding to A98 or A99 is generally I, L, V, M, F, W, Y, N, Q, S, or T The residue at the position corresponding to V105 is I, L, M, F, W, Y, N, or Q The residue at the position corresponding to Q114 is F, W, or Y. The residue at the position corresponding to E210 is N, Q, R, or K.

[0204] A specific monomer may have the sequence shown in SEQ ID NO: 3, including all of Y51A, F56Q substitution, and 1, 2, 3, 4, 5, 6, or the following substitutions. A98I / L / V / M / F / W / Y / N / Q / S / T, A99I / L / V / M / F / W / Y / N / Q / S / T, T10 4N / Q / L / R / D / E / M / F / W / Y, V105I / L / M / F / W / Y / N / Q , L113M / F / W / Y / N / Q / D / E / L / R, Q114Y / F / W, and S1 15N / Q / M / F / W / Y / E / R.

[0205] At least one CsgG monomer in the pores within the double pores may contain residues at positions corresponding to those in SEQ ID NO: 3 (such as any one of the corresponding positions among SEQ ID NOS: 68 - 88) that have fewer negative charges than any one of D149, E185, D195, E210, and E203 at the corresponding positions, where the residues at positions corresponding to D149, E185, D195, and / or E203 are K. At least one CsgG monomer in the pores within the double pores may contain at least one residue within the constriction of the barrel region of the pore, which increases the length of the constriction compared to the wild - type CsgG pore. The at least one residue is additional to the residues present in the constriction of the wild - type CsgG pore. The length of the pore may be increased by inserting residues into the region corresponding to the region between positions K49 and F56 of SEQ ID NO: 3. One or more residues of 1 - 5 (such as 2, 3, or 4) amino acids may be inserted at any one or more of the following positions defined by reference to SEQ ID NO: 3: K49 and P50, P50 and Y51, Y51 and P52, P52 and A53, A53 and S54, S54 and N55, and / or N55 and D149, E185, D195, and / or E203 are K. At least one CsgG monomer in the pores within the double pores may contain residues at positions corresponding to those in SEQ ID NO: 3 (such as any one of the corresponding positions among SEQ ID NOS: 68 - 88) that have fewer negative charges than any one of D149, E185, D195, E210, and E203 at the corresponding positions, where the residues at positions corresponding to D149, E185, D195, and / or E203 are K.

[0206] At least one CsgG monomer in the pores within the double pores may contain at least one residue within the constriction of the barrel region of the pore, which increases the length of the constriction compared to the wild - type CsgG pore. The at least one residue is additional to the residues present in the constriction of the wild - type CsgG pore. compared to the wild - type CsgG pore. The at least one residue is additional to the residues present in the constriction of the wild - type CsgG pore. At least one residue is additional to the residues present in the constriction of the wild - type CsgG pore. At least one residue is additional to the residues present in the constriction of the wild - type CsgG pore.

[0207] The length of the pore may be increased by inserting residues into the region corresponding to the region between positions K49 and F56 of SEQ ID NO: 3. One or more residues of 1 - 5 (such as 2, 3, or 4) amino acids may be inserted. One or more residues of 1 - 5 (such as 2, 3, or 4) amino acids may be inserted at any one or more of the following positions defined by reference to SEQ ID NO: 3: K49 and P50, P50 and Y51, Y51 and P52, P52 and A53, A53 and S54, S54 and N55, and / or N55 and P52 and A53, A53 and S54, S54 and N55, and / or N55 and P52 and A53, A53 and S54, S54 and N55, and / or N55 and Call F56. In total, 1 to 10 (such as 2 to 8) or 3 to 5 amino acid residues are preferably inserted into the monomer - sequence. The inserted residues can increase the length of the loop between the residues corresponding to Y51 and N 55 of SEQ ID NO: 3. The inserted residues can maintain flexibility A, S, G or T for, P for adding twist to the loop, and / or contribute to the signal generated when the analyte interacts with the pore barrel under the applied potential difference. It may be any combination of S, T, N, Q, M, F, W, Y, V and / or I. The inserted amino acids may be any combination of S, G, SG, SGG, SGS, GS, GSS and / or G SG.

[0208] At least one CsgG monomer in the pores within the double pore may contain at least one residue in the constriction of the pore barrel region at the positions corresponding to N55, P52 and / or A53 of SEQ ID NO: 3 that is different from the residues present in the corresponding wild-type monomer wherein the residue at the position corresponding to N55 is V.

[0209] Any two or more of the above residues may be present in the same monomer. In particular, the monomer may contain at least one cysteine residue, at least one hydrophobic residue, at least one bulky residue, at least one of the neutral or positively charged residues, and / or at least one of the residues that increase the length of the constriction.

[0210] The constriction of the pore barrel of the CsgG monomer in the double pore, when used for the detection or characterization of an analyte, is such that the first pore with a wild-type constriction or ​​​​​​At least one residue (such as 2, 3, 4, or 5 residues) may additionally be included, compared to when the second pore is used, where at least one residue in the constriction portion of the pore region corresponds to the position of Y51, N55, Y51, P52 and / or A53 of SEQ ID NO: 3. At least one residue may be Q or V at the position corresponding to F56 of SEQ ID NO: 3, A or Q at the position corresponding to Y51 of SEQ ID NO: 3, and / or V at the position corresponding to N55 of SEQ ID NO: 3.

[0211] Method for producing a modified protein Methods for introducing or substituting unnatural amino acids are also well known in the art. For example, unnatural amino acids can be introduced by including synthetic aminoacyl - tRNA in an in vitro translation (IVTT) system used to express mutant monomers. As another method, unnatural amino acids can be introduced by expressing mutant monomers in E. coli that are auxotrophic for specific amino acids in the presence of their synthetic (i.e., unnatural - type) analogs of the specific amino acids. When mutant monomers are generated using partial peptide synthesis, they can be generated by blunt ligation.

[0212] Monomers derived from CsgG can be modified to assist in their identification or purification, for example, by the addition of a streptavidin tag or by the addition of a signal sequence that promotes the secretion of monomers from cells that do not naturally contain such sequences. Other suitable tags will be discussed in more detail below. Monomers can be labeled with obvious labels 。An obvious label can be any suitable label that allows the monomer to be detected. Suitable labels are described below.

[0213] The monomers derived from CsgG may be produced using D - amino acids. For example, Cs gG - derived monomers may contain a mixture of L - amino acids and D - amino acids. This is conventional in the art for producing the protein or peptide.

[0214] The monomers derived from CsgG contain one or more specific modifications to facilitate nucleotide discrimination. The monomers derived from CsgG may also contain other non - specific modifications as long as they do not interfere with pore formation. A number of non - specific side - chain modifications are well - known in the art and may be made to the side chains of monomers derived from CsgG. Such modifications include, for example, reductive alkylation of amino acids by reaction with aldehydes, followed by reduction with NaBH4, amidation with methyl acetimidate, or acylation with acetic anhydride.

[0215] The monomers derived from CsgG can be produced using standard methods well - known in the art. The monomers derived from CsgG can be made synthetically or by recombinant means. For example, the monomers can be synthesized by in vitro translation and transcription (IVTT). Suitable methods for producing pores and monomers are described in International Applications WO 20 10 / 004273, WO 2010 / 004265 or WO 2010 / 08660 3. Methods for inserting pores into membranes are well - known.

[0216] Two or more CsgG monomers within the pore may covalently bond to each other. For example, at least 2 Individual, at least 3, at least 4, at least 5, at least 6, at least 7 Individuals, at least 8, at least 9 or at least 10 monomers can be covalently bonded Yes. The covalently bonded monomers may be the same or different.

[0217] The monomers can be genetically fused, optionally via a linker, or chemically fused, for example via a chemical cross-linking agent. The method of covalently bonding monomers is disclosed in WO2017 / 1493 16, WO2017 / 149317 and WO2017 / 149318 .

[0218] In some embodiments, the mutant monomers are chemically modified. The mutant monomers can be chemically modified in any way and at any site. The mutant monomers can be chemically modified by the attachment of a molecule to one or more Of the cysteines (cysteine bonds) of the molecule, the attachment of the molecule to one or more lysines, the attachment of the molecule to one or more non-natural Amino acids, enzymatic modification of epitopes, or terminal modification is preferably chemically modified. Suitable methods for performing such modifications are well known in the art . The mutant monomers may be chemically modified by the attachment of any molecule. For example , the mutant monomers can be chemically modified by the attachment of a dye or a fluorophore molecule.

[0219] In some embodiments, the mutant monomers are chemically modified with a molecular adapter that promotes the interaction between the monomers and the target nucleotide sequence or the target Polynucleotide sequence within the pore. The presence of the adapter improves the host-guest chemistry of the pore and the nucleotide sequence or polynucleotide sequence Thereby improving the pore arrangement formed from the mutant monomers ​​​​Improve column determination ability. The principle of host-guest chemistry is well known in the art. The adapter has an effect on the physical or chemical properties of the pore that improve its interaction with a nucleotide or polynucleotide sequence. The adapter can change the charge of the pore barrel or channel, or specifically interact with or bind to a nucleotide or polynucleotide sequence, thereby promoting its interaction with the pore. The molecular adapter is preferably a cyclic molecule, cyclodextrin, hybridizable species, DNA binding agent or intercalator, peptide or peptide analog, synthetic polymer, aromatic planar molecule, small positively charged molecule, or small molecule with hydrogen bonding ability. The adapter may be cyclic. The cyclic adapter preferably has the same symmetry as the pore. Since CsgG typically has 8 or 9 subunits around the central axis, it preferably has 8-fold or 9-fold symmetry. This will be described in detail below.

[0220]

[0221]

[0222]

[0222] The adapter generally interacts with a nucleotide sequence or polynucleotide sequence by host-guest chemistry. The adapter generally has the ability to interact with a nucleotide sequence or polynucleotide sequence. The adapter contains one or more chemical groups having the ability to interact with a nucleotide sequence or polynucleotide sequence. The one or more chemical groups include hydrophobic interactions, hydrogen bonds, Van der Waal forces, π-cation interactions and The adapter generally interacts with a nucleotide sequence or polynucleotide sequence. The adapter generally has the ability to interact with a nucleotide sequence or polynucleotide sequence. The adapter contains one or more chemical groups having the ability to interact with a nucleotide sequence or polynucleotide sequence. The one or more chemical groups include hydrophobic interactions, hydrogen bonds, Van der Waal forces, π-cation interactions and The adapter generally interacts with a nucleotide sequence or polynucleotide sequence. The adapter generally has the ability to interact with a nucleotide sequence or polynucleotide sequence. The adapter contains one or more chemical groups having the ability to interact with a nucleotide sequence or polynucleotide sequence. The one or more chemical groups include hydrophobic interactions, hydrogen bonds, Van der Waal forces, π-cation interactions and The adapter contains one or more chemical groups having the ability to interact with a nucleotide sequence or polynucleotide sequence. The one or more chemical groups include hydrophobic interactions, hydrogen bonds, Van der Waal forces, π-cation interactions and hydrophobic interactions, hydrogen bonds, Van der Waal forces, π-cation interactions and / or by non-covalent interaction such as electrostatic force with nucleotides or polynucleotides It is preferred to interact with the sequence. One or more chemical groups having the ability to interact with a nucleotide sequence or a polynucleotide sequence preferably carry a positive charge. One or more chemical groups having the ability to interact with a nucleotide sequence or a polynucleotide sequence are more preferably those containing an amino group. The amino group can be attached to the first carbon atom, the second carbon atom or the third carbon atom. The adapter more preferably contains a ring of amino groups such as a ring of 6, 7 or 8 amino groups. Most preferably, the adapter contains a ring of 8 amino groups. The ring of protonated amino groups may interact with the negatively charged phosphate groups within the nucleotide sequence or polynucleotide sequence. The correct positioning of the adapter within the pore can be facilitated by host-guest chemistry between the adapter and the pore containing the mutant monomer. The adapter preferably contains one or more chemical groups having the ability to interact with one or more amino acids within the pore. The adapter more preferably contains one or more chemical groups capable of interacting with one or more amino acids within the pore by non-covalent interactions such as hydrophobic interactions, hydrogen bonds, Van der Waal forces, π-cation interactions and / or electrostatic forces. The chemical group capable of interacting with one or more amino acids within the pore is typically a hydroxyl or an amine. The hydroxyl group can be attached to the first carbon atom, the second carbon atom or the third carbon atom. The hydroxyl group may form a hydrogen bond with an uncharged amino acid within the pore. It is preferred to interact with the sequence. One or more chemical groups having the ability to interact with a nucleotide sequence or a polynucleotide sequence preferably carry a positive charge. One or more chemical groups having the ability to interact with a nucleotide sequence or a polynucleotide sequence are more preferably those containing an amino group. The amino group can be attached to the first carbon atom, the second carbon atom or the third carbon atom. The adapter more preferably contains a ring of amino groups such as a ring of 6, 7 or 8 amino groups. Most preferably, the adapter contains a ring of 8 amino groups. The ring of protonated amino groups may interact with the negatively charged phosphate groups within the nucleotide sequence or polynucleotide sequence. The correct positioning of the adapter within the pore can be facilitated by host-guest chemistry between the adapter and the pore containing the mutant monomer. The adapter preferably contains one or more chemical groups having the ability to interact with one or more amino acids within the pore. The adapter more preferably contains one or more chemical groups capable of interacting with one or more amino acids within the pore by non-covalent interactions such as hydrophobic interactions, hydrogen bonds, Van der Waal forces, π-cation interactions and / or electrostatic forces. The chemical group capable of interacting with one or more amino acids within the pore is typically a hydroxyl or an amine. The hydroxyl group can be attached to the first carbon atom, the second carbon atom or the third carbon atom. The hydroxyl group may form a hydrogen bond with an uncharged amino acid within the pore.

[0223] The correct positioning of the adapter within the pore can be facilitated by host-guest chemistry between the adapter and the pore containing the mutant monomer. The adapter preferably contains one or more chemical groups having the ability to interact with one or more amino acids within the pore. The adapter more preferably contains one or more chemical groups capable of interacting with one or more amino acids within the pore by non-covalent interactions such as hydrophobic interactions, hydrogen bonds, Van der Waal forces, π-cation interactions and / or electrostatic forces. The chemical group capable of interacting with one or more amino acids within the pore is typically a hydroxyl or an amine. The hydroxyl group can be attached to the first carbon atom, the second carbon atom or the third carbon atom. The hydroxyl group may form a hydrogen bond with an uncharged amino acid within the pore. are more preferably those containing an amino group. The amino group can be attached to the first carbon atom, the second carbon atom or the third carbon atom. The adapter more preferably contains a ring of amino groups such as a ring of 6, 7 or 8 amino groups. Most preferably, the adapter contains a ring of 8 amino groups. The ring of protonated amino groups may interact with the negatively charged phosphate groups within the nucleotide sequence or polynucleotide sequence. It is preferred to interact with the sequence. One or more chemical groups having the ability to interact with a nucleotide sequence or a polynucleotide sequence preferably carry a positive charge. One or more chemical groups having the ability to interact with a nucleotide sequence or a polynucleotide sequence are more preferably those containing an amino group. The amino group can be attached to the first carbon atom, the second carbon atom or the third carbon atom. The adapter more preferably contains a ring of amino groups such as a ring of 6, 7 or 8 amino groups. Most preferably, the adapter contains a ring of 8 amino groups. The ring of protonated amino groups may interact with the negatively charged phosphate groups within the nucleotide sequence or polynucleotide sequence. The correct positioning of the adapter within the pore can be facilitated by host-guest chemistry between the adapter and the pore containing the mutant monomer. The adapter preferably contains one or more chemical groups having the ability to interact with one or more amino acids within the pore. The adapter more preferably contains one or more chemical groups capable of interacting with one or more amino acids within the pore by non-covalent interactions such as hydrophobic interactions, hydrogen bonds, Van der Waal forces, π-cation interactions and / or electrostatic forces. The chemical group capable of interacting with one or more amino acids within the pore is typically a hydroxyl or an amine. The hydroxyl group can be attached to the first carbon atom, the second carbon atom or the third carbon atom. The hydroxyl group may form a hydrogen bond with an uncharged amino acid within the pore. are more preferably those containing an amino group. The amino group can be attached to the first carbon atom, the second carbon atom or the third carbon atom. The adapter more preferably contains a ring of amino groups such as a ring of 6, 7 or 8 amino groups. Most preferably, the adapter contains a ring of 8 amino groups. The ring of protonated amino groups may interact with the negatively charged phosphate groups within the nucleotide sequence or polynucleotide sequence. The correct positioning of the adapter within the pore can be facilitated by host-guest chemistry between the adapter and the pore containing the mutant monomer. The adapter preferably contains one or more chemical groups having the ability to interact with one or more amino acids within the pore. The adapter Any that promotes the interaction between the pore and the nucleotide sequence or polynucleotide sequence An adapter can be used.

[0224] Suitable adapters include, but are not limited to, cyclodextrin, cyclic peptide, and cucurbituril Examples include, but are not limited to these. The adapter is preferably cyclodextrin or a derivative thereof. Cyclodextrin or its derivative can be any disclosed in the following documents: Eliseev, A. V., and Schneid er, H-J. (1994) J. Am. Chem. Soc. 116, 6081 - 6088. The adapter is more preferably heptakis-6-amino-β-cyclodextrin (am7-βCD), 6-monodeoxy-6-monoamino-β-cyclodextrin (a m1-CD) or heptakis-(6-deoxy-6-guanidino)-cyclodextrin (gu7-βCD). Since the guanidino group in gu7-βCD has a much higher pKa than the primary amine in am7-βCD, it is more positively charged . This gu7-βCD adapter can be used to extend the dwell time of nucleotides in the pore, increase the accuracy of the measured residual current, and increase the base detection rate at high temperature or low data acquisition speed. ,

[0225] When the succinimidyl 3-(2-pyridyldithio)propionate (SPDP) crosslinking agent is used as discussed in more detail below, the adapter is preferably heptakis(6-deoxy -6-amino)-6-N-mono(2-pyridyl)dithiopropanoyl-β-cyclodextrin (am6amPDP1-βCD). ​

[0226] More suitable adapters include γ-cyclodextrin containing nine sugar units (and thus having nine-fold symmetry). γ-cyclodextrin may contain a linker molecule or may be modified to include all or some of the modified sugar units used in the β-cyclodextrin examples described above. The molecular adapter may be covalently attached to the mutant monomer. The adapter can be covalently attached to the pore using any method well known in the art. The adapter is typically attached via a chemical bond. When the molecular adapter is attached via a cysteine bond, one or more cysteines are preferably introduced into the mutant (e.g., within the barrel) by substitution. The mutant monomer can be chemically modified by attachment of a molecular adapter to one or more cysteines within the mutant monomer. One or more cysteines may be native, i.e., at positions 1 and / or 215 in SEQ ID NO: 3. Alternatively, the mutant monomer can be chemically modified by attachment of a molecule to one or more cysteines introduced at other positions. The cysteine at position 215 may be removed, e.g., by substitution, to ensure that the molecular adapter does not attach at that position other than to the cysteine at position 1 or to a cysteine introduced at another position. The reactivity of the cysteine residue can be enhanced by modification of adjacent residues. For example, the basic groups of adjacent arginine, histidine or lysine residues lower the pKa of the cysteine thiol group to a more reactive S

[0227]

[0228] - ​​​​​​​​​​​​​​​The reactivity of cysteine ​​residues is improved by the addition of thiol groups such as dTNB. These can be protected by aryl protecting groups, which are then added to the mutant monolayer before the linker is attached. The cysteine ​​residues of the mer may be reacted with one or more cysteine ​​residues of the mer.

[0229] The molecule may be attached directly to the mutant monomer. The molecule may be attached using a chemical crosslinker or a peptide crosslinker. It is preferably attached to the mutant monomer using a linker such as a linker.

[0230] Suitable chemical cross-linking agents are well known in the art. Preferred cross-linking agents include 2,5-dioxopyridine. Rolidin-1-yl 3-(pyridin-2-yldisulfanyl)propanoate, 2,5- Dioxopyrrolidin-1-yl 4-(pyridin-2-yldisulfanyl)butanoate, and 2,5-dioxopyrrolidin-1-yl 8-(pyridin-2-yldisulfanyl The most preferred crosslinking agent is succinimidyl 3-(2-pyridinyl) octanoate. Typically, the molecule / crosslinker complex suddenly Before being covalently linked to the mutant monomer, the molecule is covalently linked to a bifunctional crosslinker; Before the bifunctional crosslinker / monomer complex is bound to a molecule, the bifunctional crosslinker is covalently attached to the monomer. It is also possible to covalently bond the amine to the monomer.

[0231] Preferably, the linker is resistant to dithiothreitol (DTT). Linkers include iodoacetamide-based and maleimide-based linkers. Examples include, but are not limited to:

[0232] In other embodiments, the monomer may be attached to a polynucleotide binding protein. This forms a modular sequencing system that can be used in the sequencing method of the present invention. The polynucleotide-binding protein will be discussed below.

[0233] The polynucleotide-binding protein is preferably covalently bound to the mutant monomer. The protein can be covalently bound to the monomer using any method well known in the art. The monomer and the protein can be chemically or genetically fused. When the entire construct is expressed from a single polynucleotide sequence, the monomer and the protein are genetically fused. The gene fusion of the monomer to the polynucleotide-binding protein is discussed in WO 2010 / 004265.

[0234] When the polynucleotide-binding protein is attached via a cysteine bond, it is preferred that more than one cysteine is introduced into the mutant by substitution. More than one cysteine is preferably introduced into a loop region with low conservation among homologs indicating that mutations or insertions are tolerated. Thus, it is suitable for attaching the polynucleotide-binding protein. In such embodiments, the native cysteine at position 251 can be removed. The reactivity of the cysteine residue can be enhanced by modification as described above.

[0235] The polynucleotide-binding protein can bind directly to the mutant monomer or through one or more linkers. The molecule may be attached to the mutant monomer using a hybridizing linker as described in WO 2010 / 086602. Alternatively, a pep A linker may be used. The peptide linker is an amino acid sequence. The length, flexibility, and hydrophilicity of the peptide linker are typically designed so as not to interfere with the function of the monomer and the molecule. A preferred flexible peptide linker is 2 to 20 in length (such as 4, 6, 8, 10, or 16), such as serine amino acids and / or glycine amino acids. More preferred flexible linkers include (SG)1, (SG)2, (SG)3, ( (SG)4, (SG)5, and (SG)8, where S is serine and G is glycine. A preferred rigid linker is 2 to 30 in length (such as 4, 6, 8, 16, or 24) of proline amino acids. A more preferred rigid linker contains (P) where P is proline. 12

[0236] Chemical modification The mutant CsgG monomer or CsgF peptide may be chemically modified with a molecular adapter and a polynucleotide binding protein.

[0237] The molecule (where the monomer or peptide is chemically modified) may be attached directly to the monomer or peptide, or via a linker, as disclosed in WO 2010 / 004 273, WO 2010 / 004265, or WO 2010 / 086603.

[0238] Any protein described herein, such as the CsgG monomer and / or CsgF peptide, may be added with, for example, a histidine residue (his tag), an aspartic acid residue (asp tag) , a streptavidin tag, a flag tag, a SUMO tag, a GST tag, an MBP tag. By doing so, or by the addition of a signal sequence that promotes secretion from cells that do not naturally contain the signal sequence, they can be modified to assist in their identification or purification. An alternative to introducing a genetic tag is to react a tag chemically at a natural or modified position on the protein. One example of this is to react a gel - shift reagent with a cysteine modified on the outside of the protein. This has been demonstrated as a method for separating hemolysin hetero - oligomers (Chem Biol. 1997 Jul, 4( 7):497 - 505).

[0239] Any protein described herein, such as a CsgG monomer and / or a CsgF peptide, can be labeled with an obvious label. An obvious label can be any suitable label that allows the protein to be detected. Suitable labels include, but are not limited to, fluorescent molecules, radioisotopes, for example 125 I, 35 S, enzymes, antibodies, antigens, polynucleotides, and ligands such as biotin.

[0240] Any protein described herein, such as a CsgG monomer and / or a CsgF peptide, can be made synthetically or by recombinant means. For example, the protein can be synthesized by in vitro translation and transcription (IVTT). The amino acid sequence of the protein may contain non - natural amino acids or may be modified to increase the stability of the protein. Such amino acids can be introduced during production. The protein can also be altered after either synthetic or recombinant production.

[0241] The protein may be produced using D - amino acids. For example, the protein may contain a mixture of L - amino acids and D - amino acids. This is conventional in the art for producing such protein or peptide.

[0242] The protein may also include other non - specific modifications, provided that they do not interfere with the function of the protein. A number of non - specific side - chain modifications are well - known in the art and may be made to the side - chains of the protein. Such modifications include, for example, reductive alkylation of amino acids by reaction with an aldehyde, subsequent reduction with NaBH4, amidation with methyl acetimidate, or acylation with acetic anhydride.

[0243] Any protein described herein, such as a CsgG monomer and / or a CsgF peptide, can be manufactured using standard methods well - known in the art. The polynucleotide sequence encoding the protein can be derived and replicated using standard methods in the art. The polynucleotide sequence encoding the protein can be expressed in bacterial host cells using standard techniques in the art. The protein can be produced intracellularly by expression of the polypeptide from an in situ recombinant expression vector. The expression vector optionally carries an inducible promoter to control the expression of the polypeptide. These methods are described in the following references: Sambrook, J. and Russel, D. (2001). Molecular Cloning: A La boratory Manual, 3rd Edition. Cold Sprin boratory Manual, 3rd Edition. Cold Spring g Harbor Laboratory Press, Cold Spring H arbor, NY。

[0244] Proteins can be produced on a large scale after being purified by any protein liquid chromatography system from protein-producing organisms or after recombinant expression. Typical protein liquid chromatography systems include FPLC, AKTA systems, B io-Cad systems, Bio-Rad BioLogic systems, and Gilson HPLC systems. are included.

[0245] Method for generating pores In a third aspect, the present invention provides a method for producing a CsgG:modified CsgG pore complex that retains two or more constriction sites in vivo and in vitro. One embodiment is to co-express a transmembrane pore complex comprising a CsgG pore or a homolog or mutant form thereof and a modified CsgF peptide or a homolog or mutant form thereof. The method includes the steps of expressing a CsgG monomer (expressed as the preprotein provided by SEQ ID NO: 2 or a homolog or mutant form thereof) and expressing a modified or cleaved CsgF monomer (both within a suitable host cell) to form an in vivo complex pore. The complex includes a modified CsgF peptide that forms a complex with the CsgG pore and provides its pore to an additional leader head. The resulting pore complex generated by the method using the modified Csg F peptide allows the passage of analytes (especially polynucleotide chains), for characterizing target analytes such as nucleic acid sequencing including a modified CsgF peptide that forms a complex with the CsgG pore and provides its pore to an additional leader head. The resulting pore complex generated by the method using the modified Csg F peptide allows the passage of analytes (especially polynucleotide chains), for characterizing target analytes such as nucleic acid sequencing and providing sufficient structure for the porous composite to be used in a suitable setting for said application. and when coupled to said polynucleotide sequence, said two or more reader heads for improved reading of said polynucleotide sequence. Includes.

[0246] More specifically, the modified CsgF peptides expressed in the method are those represented by SEQ ID NOs: 8, 10, 1 2, or 14, or their homologs. These sequences are The method further comprises introducing a constriction site into the pore complex to bind the CsgG protein to the pore and induce the binding of the protein to the organism. The present study limited the CsgF fragments to those that could acquire biological pores.

[0247] Formed by the CsgG and CsgF proteins or similar Another method to generate isolated pore complexes is to reconstitute the monomers in vitro. The method involves obtaining a functional pore by using a suitable sieve that allows complex formation. In the stem, a mature CsgG molecule depicted in SEQ ID NO: 3 or a homologue or mutant thereof is contacting the polypeptide with a modified CsgF peptide or a homologue or mutant thereof. The system may be an "in vitro system", which is an in vitro system comprising the method. It contains at least the components and environment necessary to carry out the law, and is consistent with the normal natural environment. Using external biological molecules, organisms, and cells (or parts of cells) and whole organisms Allows for more detailed, more convenient, or more efficient analysis than can be performed using An in vitro system also refers to a system that is capable of providing a suitable The composition may contain an appropriate buffer solution, in which the protein components for forming the complex are contained. A person skilled in the art is aware of the options available for providing said system. For in vitro reconstitution in certain embodiments, the modification CsgF peptide or similar applied in the method is a peptide comprising SEQ ID NO: 15 or SEQ ID NO: 16, or its mutants or homologs, which can be synthesized or recombinantly produced. Alternatively, the modified CsgF peptide comprising SEQ ID NO: 40, 39, 38, or 37, 15, 54, 55, or their homologs or mutants is provided in the method to contact with CsgG or CsgG-like pores to generate pore complexes. For in vitro reconstitution in certain embodiments, the modification CsgF peptide or similar applied in the method is a peptide comprising SEQ ID NO: 15 or SEQ ID NO: 16, or its mutants or homologs, which can be synthesized or recombinantly produced. Alternatively, the modified CsgF peptide comprising SEQ ID NO: 40, 39, 38, or 37, 15, 54, 55, or their homologs or mutants is provided in the method to contact with CsgG or CsgG-like pores to generate pore complexes. For in vitro reconstitution in certain embodiments, the modification CsgF peptide or similar applied in the method is a peptide comprising SEQ ID NO: 15 or SEQ ID NO: 16, or its mutants or homologs, which can be synthesized or recombinantly produced. Alternatively, the modified CsgF peptide comprising SEQ ID NO: 40, 39, 38, or 37, 15, 54, 55, or their homologs or mutants is provided in the method to contact with CsgG or CsgG-like pores to generate pore complexes. For in vitro reconstitution in certain embodiments, the modification CsgF peptide or similar applied in the method is a peptide comprising SEQ ID NO: 15 or SEQ ID NO: 16, or its mutants or homologs, which can be synthesized or recombinantly produced. Alternatively, the modified CsgF peptide comprising SEQ ID NO: 40, 39, 38, or 37, 15, 54, 55, or their homologs or mutants is provided in the method to contact with CsgG or CsgG-like pores to generate pore complexes. For in vitro reconstitution in certain embodiments, the modification CsgF peptide or similar applied in the method is a peptide comprising SEQ ID NO: 15 or SEQ ID NO: 16, or its mutants or homologs, which can be synthesized or recombinantly produced. Alternatively, the modified CsgF peptide comprising SEQ ID NO: 40, 39, 38, or 37, 15, 54, 55, or their homologs or mutants is provided in the method to contact with CsgG or CsgG-like pores to generate pore complexes. For in vitro reconstitution in certain embodiments, the modification CsgF peptide or similar applied in the method is a peptide comprising SEQ ID NO: 15 or SEQ ID NO: 16, or its mutants or homologs, which can be synthesized or recombinantly produced. Alternatively, the modified CsgF peptide comprising SEQ ID NO: 40, 39, 38, or 37, 15, 54, 55, or their homologs or mutants is provided in the method to contact with CsgG or CsgG-like pores to generate pore complexes.

[0248] The CsgG / CsgF pores can be made by any suitable method. Examples of such suitable methods are described. The CsgG / CsgF pores can be made by any suitable method. Examples of such suitable methods are described.

[0249] In one embodiment, the CsgG / CsgF pores can be generated by co-expression. In this embodiment, at least one gene encoding a CsgG monomer polypeptide (which may also be a mutant polypeptide) in one vector and at least one gene encoding a full-length or truncated CsgF polypeptide (which may also be a mutant polypeptide) in a second vector are co-transformed to express the protein and produce the complex in the transformed cells. This can be either in vivo or in vitro. Alternatively, two genes encoding the CsgG and CsgF polypeptides can be placed within one vector under the control of a single promoter or two different separate promoters (which may be the same or different). In one embodiment, the CsgG / CsgF pores can be generated by co-expression. In this embodiment, at least one gene encoding a CsgG monomer polypeptide (which may also be a mutant polypeptide) in one vector and at least one gene encoding a full-length or truncated CsgF polypeptide (which may also be a mutant polypeptide) in a second vector are co-transformed to express the protein and produce the complex in the transformed cells. This can be either in vivo or in vitro. Alternatively, two genes encoding the CsgG and CsgF polypeptides can be placed within one vector under the control of a single promoter or two different separate promoters (which may be the same or different). In one embodiment, the CsgG / CsgF pores can be generated by co-expression. In this embodiment, at least one gene encoding a CsgG monomer polypeptide (which may also be a mutant polypeptide) in one vector and at least one gene encoding a full-length or truncated CsgF polypeptide (which may also be a mutant polypeptide) in a second vector are co-transformed to express the protein and produce the complex in the transformed cells. This can be either in vivo or in vitro. Alternatively, two genes encoding the CsgG and CsgF polypeptides can be placed within one vector under the control of a single promoter or two different separate promoters (which may be the same or different). In one embodiment, the CsgG / CsgF pores can be generated by co-expression. In this embodiment, at least one gene encoding a CsgG monomer polypeptide (which may also be a mutant polypeptide) in one vector and at least one gene encoding a full-length or truncated CsgF polypeptide (which may also be a mutant polypeptide) in a second vector are co-transformed to express the protein and produce the complex in the transformed cells. This can be either in vivo or in vitro. Alternatively, two genes encoding the CsgG and CsgF polypeptides can be placed within one vector under the control of a single promoter or two different separate promoters (which may be the same or different). In one embodiment, the CsgG / CsgF pores can be generated by co-expression. In this embodiment, at least one gene encoding a CsgG monomer polypeptide (which may also be a mutant polypeptide) in one vector and at least one gene encoding a full-length or truncated CsgF polypeptide (which may also be a mutant polypeptide) in a second vector are co-transformed to express the protein and produce the complex in the transformed cells. This can be either in vivo or in vitro. Alternatively, two genes encoding the CsgG and CsgF polypeptides can be placed within one vector under the control of a single promoter or two different separate promoters (which may be the same or different). In one embodiment, the CsgG / CsgF pores can be generated by co-expression. In this embodiment, at least one gene encoding a CsgG monomer polypeptide (which may also be a mutant polypeptide) in one vector and at least one gene encoding a full-length or truncated CsgF polypeptide (which may also be a mutant polypeptide) in a second vector are co-transformed to express the protein and produce the complex in the transformed cells. This can be either in vivo or in vitro. Alternatively, two genes encoding the CsgG and CsgF polypeptides can be placed within one vector under the control of a single promoter or two different separate promoters (which may be the same or different). In one embodiment, the CsgG / CsgF pores can be generated by co-expression. In this embodiment, at least one gene encoding a CsgG monomer polypeptide (which may also be a mutant polypeptide) in one vector and at least one gene encoding a full-length or truncated CsgF polypeptide (which may also be a mutant polypeptide) in a second vector are co-transformed to express the protein and produce the complex in the transformed cells. This can be either in vivo or in vitro. Alternatively, two genes encoding the CsgG and CsgF polypeptides can be placed within one vector under the control of a single promoter or two different separate promoters (which may be the same or different). In one embodiment, the CsgG / CsgF pores can be generated by co-expression. In this embodiment, at least one gene encoding a CsgG monomer polypeptide (which may also be a mutant polypeptide) in one vector and at least one gene encoding a full-length or truncated CsgF polypeptide (which may also be a mutant polypeptide) in a second vector are co-transformed to express the protein and produce the complex in the transformed cells. This can be either in vivo or in vitro. Alternatively, two genes encoding the CsgG and CsgF polypeptides can be placed within one vector under the control of a single promoter or two different separate promoters (which may be the same or different). In one embodiment, the CsgG / CsgF pores can be generated by co-expression. In this embodiment, at least one gene encoding a CsgG monomer polypeptide (which may also be a mutant polypeptide) in one vector and at least one gene encoding a full-length or truncated CsgF polypeptide (which may also be a mutant polypeptide) in a second vector are co-transformed to express the protein and produce the complex in the transformed cells. This can be either in vivo or in vitro. Alternatively, two genes encoding the CsgG and CsgF polypeptides can be placed within one vector under the control of a single promoter or two different separate promoters (which may be the same or different).

[0250] In another embodiment, the CsgG / CsgF pores are provided with CsgG monomers separately from the CsgF peptide. It is generated by expressing the nomer. The CsgG monomer or CsgG pore may be purified from cells transformed with a vector encoding at least one CsgG monomer, or with one or more vectors each expressing the CsgG mo nomer. The CsgF peptide may be purified from cells transformed with a vector encoding at least one CsgF peptide. Next, the purified CsgG monomer / pore may be cultured with the CsgF peptide to produce a pore complex. In another embodiment, the CsgG monomer and / or CsgF peptide are separately produced by in vitro translation and transcription (IVTT). Next, the CsgG monomer may be cultured with the C sgF peptide to produce a pore complex. The use of this method is illustrated in FIG. 14 as shown in the figure.

[0251] In the above embodiments, for example, (i) CsgG is produced in vivo and CsgF is produced in v ivo, (ii) CsgG is produced in vitro and CsgF is produced in vivo production, (iii) CsgG is produced in vivo and CsgF is produced in vitro, or (iv) CsgG is produced in vitro and CsgF is produced in vitro

[0252] It may be produced. One or both of the CsgG monomer and CsgF peptide may be tagged to facilitate purification. Purification may also be carried out when the CsgG monomer and / or CsgF peptide are not tagged. Methods known in the art (e.g., ion exchange

[0253] exchange, etc.) can be used for purification. Purification can also be carried out when the CsgG monomer and / or CsgF peptide are not tagged. Methods known in the art (e.g., ion exchange exchange, etc.) can be used for purification. Exchange, gel filtration, hydrophobic interaction column chromatography) can be used alone or in different combinations to purify the pore components. It can be used to purify the pore components alone or in different combinations.

[0254] Any known tag can be used on either of the two proteins. In one embodiment Two tag purifications can be used to purify the CsgG:CsgF complex from the CsgG pore and CsgF. For example, a strep tag can be used on CsgG and a His tag can be used on CsgF, or vice versa. This is illustrated in FIG. 13. When the two proteins are purified separately and then mixed together and then strep purification and His purification are performed again, similar final results can be obtained. It is possible to obtain.

[0255] When the full-length CsgF protein forms a complex with CsgG, the neck domain and head domain of CsgF (FIG. 4B) protrude from the β-barrel of the CsgG pore.

[0256] Therefore, when pores containing the CsgG pore and full-length CsgF are used in single-channel recording experiments the head domain may interfere with or prevent the insertion of the pore into the membrane. These may also interfere with the analyte passing through the pore. Therefore, when inserting the pore into the membrane it would be better if the number of flexible polypeptides hanging from the β-barrel could be reduced. Mimicking the FCP region resolved within the cryo EM structure of the complex and maintaining structural integrity, a truncated version of the CsgF protein is provided herein. A truncated version of the CsgF protein that mimics the FCP region resolved within the cryo EM structure of the complex and maintains structural integrity is provided herein.

[0257] The CsgG / CsgF pore can be made before or after inserting the CsgG pore into the membrane. It can be produced. When the pore complex is produced before insertion into the membrane, a cleavage mutant is preferably used. However, the CsgG pore can be inserted into the membrane so that the CsgG and CsgG complexes are formed in situ u, and then the CsgF peptide can be added. For example, in one embodiment, in a system accessible from the transmembrane side (e.g., in a chip or chamber for electrophysiological measurements), the CsgG pore is inserted into the membrane so that the complex is formed in situ u, and then the CsgF peptide may be added from the transmembrane side. In any embodiment where the CsgG pore is formed in situ, a larger CsgF peptide may be used. For example, the CsgF peptide may include all or part of the neck domain of CsgF (derived from approximately residue 36 of SEQ ID NO: 6). In some embodiments, CsgF may include all of the neck domain and part of the head domain (residues 36 - XX of SEQ ID NO: 6).

[0258] Depending on the method of making the complex and the stability of the complex with specific cleavage, the CsgG:CsgF and CsgG:FCP complexes can be made in different ways.

[0259] In one embodiment, a cleaved version of the CsgF polypeptide at the required length is used directly.

[0260] In another embodiment, the full-length polypeptide of CsgF, or one longer than the required cleavage (long enough to keep the complex stable), is used, where a protease cleavage site is used so that the CsgF peptide of the required length is generated by cleavage with a protease. is inserted (e.g., TEV, HRV 3 or any other protease cleavage site position). In this embodiment, when the CsgG / CsgF complex is formed, the protease is necessarily used to cleave CsGF at the required site. Alternatively, the protease may be used to generate the CsgF peptide before complex assembly.

[0261] Some protease sites leave additional tags after cleavage. For example, the TEV protease cleavage sequence is ENLYFQS. TEV protease cleaves the protein between Q and S and leaves ENLYFQ intact at the C-terminus of the CsgF peptide. Figure 15 shows an example using a modified CsgF containing the TEV cleavage site cleaved using TEV protease after complex formation. As another example, the HRV C3 cleavage site is LEVLFQGP, and the enzyme cleaves between Q and G, leaving LEVLFQ intact at the C-terminus of the CsgF peptide.

[0262] Method for characterizing an analyte In a further aspect, the present invention provides a method for determining the presence, absence, or one or more characteristics of a target analyte. The method involves contacting the target analyte with an isolated pore complex or a transmembrane pore (such as a pore of the present invention) such that the target analyte moves in relation to the pore channel (such as entering or passing through it), and taking one or more measurements when the analyte moves in relation to the pore, thereby determining the presence, absence, or one or more characteristics of the analyte. The target analyte may also be referred to as a template analyte or a target analyte. The isolated pore complex generally comprises 7, 8, 9, or 10 CsgG mono mers or multimers. ​​​​​​​​​​At least 7, at least 8, at least 9, or at least 10 monomers such as mers are included. The isolated pore complex preferably contains 8 or 9 identical CsgG monomers . One or more (2, 3, 4, 5, 6, 7, 8 , 9, or 10, etc.) of the CsgG monomers are preferably chemically modified, or the CsgF peptide is chemically modified. The isolated pore complex monomers such as CsgG monomers or their homologs or mutants, and modified CsgF monomers or their homologs or mutants can be derived from any organism. After the analyte passes through the CsgG constriction, it may pass through the CsgF constriction . In another embodiment, after the analyte passes through the CsgF constriction, it may pass through the CsgG constriction depending on the orientation of the CsgG / CsgF complex in the membrane .

[0263] The method is to determine the presence, absence, or one or more characteristics of the target analyte. The method is for determining the presence, absence, or one or more characteristics of at least one analyte and may be. The method may be related to determining the presence, absence, or one or more characteristics of two or more analytes . The method may involve determining the presence, absence, or one or more characteristics of any number of analytes (2, 5, 10, 15, 20, 30 , 40, 50, 100 or more analytes, etc.). Any number of characteristics (1, 2, 3, 4, 5, 10 or more characteristics, etc.) of one or more analytes can be determined.

[0264] The binding of molecules either within the channel of the pore complex or near any of the openings at the opening of the channel has an effect on the open channel ion flow through the pore, but this is the pore channel This is the essence of the "molecular sensing" of the channel. In a manner similar to nucleic acid sequencing applications, an open channel Changes in ion flow can be measured using appropriate measurement techniques by changes in current ([ For example, WO 2000 / 28312 and D. Stoddart et al., Proc. Natl. Acad. Sci., 2010, 106, 7702 - 7 or WO 2009 / 077734). The degree of decrease in ion flow measured by the decrease in current is related to the size of the obstacle in the pore or the proximity of the pore. Thus, the binding of a target molecule (also called an "analyte") in or near the pore provides a detectable and measurable event, thereby forming the basis of a "biological sensor". Suitable molecules for nanopore sensing include nucleic acids, proteins, peptides, polysaccharides, and small molecules (e.g., pharmaceuticals, toxins, cytokines, and contaminants, etc.) and other low molecular weight or inorganic compounds (in this specification, it means organic or inorganic compounds with low molecular weight (e.g., < 900 Da or < 500 Da)). Detecting the presence of biomolecules has applications in personalized drug development, medicine, diagnosis, life science research, environmental monitoring, and security and / or defense industries.

[0265] In another aspect, a transmembrane pore complex comprising an isolated pore complex, or a wild - type or modified E. coli Csg G nanopore or its homolog or mutant, and a modified CsgF peptide that provides channel constriction to the pore in the complex can function as a molecular sensor or a biological sensor. In some embodiments, the CsgG nanopore is derived or isolated from bacterial proteins (e.g., E. coli, Salmonella typhi) This may be done. In some embodiments, CsgG nanopores can be recombinantly produced. The procedure for analyte detection is described in Howorka et al. Nature Biotechn ology (2012) Jun 7, 30(6):506-7. The analyte molecules to be detected may bind to either face of the channel or within the lumen of the channel itself. The position of binding can be determined by the size of the molecule being sensed .

[0266] Target analytes include metal ions, inorganic salts, polymers, amino acids, peptides, polypeptides, proteins, nucleotides, oligonucleotides, polynucleotides, polysaccharides, dyes, bleaching agents, pharmaceuticals, diagnostic agents, recreational drugs, explosives, toxic compounds, environmental pollutants are preferred. The method may be related to determining the presence, absence, or one or more characteristics of two or more analytes of the same type, such as two or more proteins, two or more nucleotides, or two or more pharmaceuticals. Alternatively, the method may be related to determining the presence, absence, or one or more characteristics of two or more analytes of different types, such as one or more proteins, one or more nucleotides, and one or more pharmaceuticals.

[0267] Target analytes can be secreted from cells. Alternatively, the target analyte may be an analyte present within the cell such that the analyte must be extracted from the cell before the method is performed. .

[0268] Wild-type pores may act as sensors, but recombinant or chemical methods can be used to modify them to increase the binding strength, binding position, or binding specificity of the molecule being sensed. ​is often decorated. Typical modifications include the addition of specific binding moieties that are complementary to the structure of the molecule being sensed. When the analyte molecule contains nucleic acid, this binding moiety may include cyclodextrin or oligonucleotide, and in the case of small molecules, this may be a complementary binding region, for example, an antigen-binding portion of an antibody molecule or a non-antibody molecule containing a single-chain variable fragment (scFv) region or an antigen recognition domain derived from a T cell receptor (TCR), and in the case of proteins, it may also be a known ligand of the target protein. In this way, wild-type or modified E. coli CsgG nanopores or their homologs can be given the ability to act as molecular sensors for detecting the presence of appropriate antigens (including epitopes) in a sample. The antigens can include cell surface antigens (receptors, solid tumor cells, or markers of blood cancer cells (e.g., lymphoma or leukemia)), viral antigens, bacterial antigens, protozoal antigens, allergens, allergy-related molecules, albumin (e.g., human, rodent, or bovine), fluorescent molecules (including fluorescein), blood group antigens, small molecules, drugs, enzymes, catalytic sites or substrates of enzymes, and analogs of the transition states of enzyme substrates. As described above, the modification can be achieved using known genetic engineering and recombinant DNA techniques. Any suitable positioning depends on the nature of the molecule being sensed, such as size, three-dimensional structure, and its biochemistry. The selection of the adapted structure may also use computational structure design. The determination and optimization of protein-protein interactions or protein-small molecule interactions can be investigated using techniques such as BIAcore (registered trademark) that detect molecular interactions using surface plasmon resonance (BIAcore, Inc., Piscataway ), and the like. The antigens can include cell surface antigens (receptors, solid tumor cells, or markers of blood cancer cells (e.g., lymphoma or leukemia)), viral antigens, bacterial antigens, protozoal antigens, allergens, allergy-related molecules, albumin (e.g., human, rodent, or bovine), fluorescent molecules (including fluorescein), blood ...

Claims

1. (i) a modified CsgG pore comprising one or more deletions within the CsgG constriction loop corresponding to residues 46 to 61 of SEQ ID NO: 3, and (ii) a cleaved CsgF peptide, wherein the CsgF peptide is bound to the modified CsgG pore, forms a constriction within the modified CsgG pore, and the CsgF peptide has a length of 25 to 50 amino acids, the cleaved CsgF peptide comprising an isolated pore complex, wherein the CsgF peptide comprises (i) the amino acid sequence of SEQ ID NO: 6 from residue 1 to any one of residues 28 to 45 of SEQ ID NO: 6, (ii) SEQ ID NO: 39 (residues 1 to 29 of SEQ ID NO: 6), (iii) SEQ ID NO: 15 (residues 1 to 34 of SEQ ID NO: 6), (iv) SEQ ID NO: 40 (residues 1 to 45 of SEQ ID NO: 6), (v) SEQ ID NO: 54 (residues 1 to 30 of SEQ ID NO: 6), or (vi) SEQ ID NO: 55 (residues 1 to 35 of SEQ ID NO: 6), the isolated pore complex.

2. The isolated pore complex according to claim 1, wherein the CsgG pore comprises one or more deletions among positions F48, K49, P50, Y51, P52, A53, S54, N55, F56 and S57.

3. The isolated pore complex according to claim 1 or 2, wherein the CsgF peptide has a length of 25 to 45 amino acids.

4. The isolated pore complex according to any one of claims 1 to 3, wherein the CsgF peptide comprises a modification at one or more of positions G1, T4, F5, R8, N9, N11, F12, A26, Q29, N15, N17, A20, N24, A28 and D34.

5. The isolated pore complex according to any one of claims 1 to 4, wherein the CsgF peptide comprises one or more of the following substitutions: N15S / A / T / Q / G / L / V / I / F / Y / W / R / K / D / C, N17S / A / T / Q / G / L / V / I / F / Y / W / R / K / D / C, A20S / T / Q / N / G / L / V / I / F / Y / W / R / K / D / C, N24S / T / Q / A / G / L / V / I / F / Y / W / R / K / D / C, A28S / T / Q / N / G / L / V / I / F / Y / W / R / K / D / C, D34F / Y / W / R / K / N / Q / C, G1C, T4C, N17S, and D34Y or D34N.

6. The isolated pore complex according to any one of claims 1 to 5, wherein the CsgG pore comprises 6 to 10 CsgG monomers.

7. The CsgF peptide and the CsgG pore are covalently bonded, and the covalent bond is (i) a cysteine residue at a position corresponding to 132, 133, 136, 138, 140, 142, 144, 145, 147, 149, 151, 153, 155, 183, 185, 187, 189, 191, 201, 203, 205, 207, or 209 of SEQ ID NO: 3, (ii) a non-natural reactive or photoreactive amino acid at a position corresponding to 132, 133, 136, 138, 140, 142, 144, 145, 147, 149, 151, 153, 155, 183, 185, 187, 189, 191, 201, 203, 205, 207, or 209 of SEQ ID NO: 3, or (iii) through residues at positions corresponding to one or more of the following pairs of positions of SEQ ID NO: 6 and SEQ ID NO: 3: 1 and 153, 4 and 133, 5 and 136, 8 and 187, 8 and 203, 9 and 203, 11 and 142, 11 and 201, 12 and 149, 12 and 203, 26 and 191, and 29 and 144; the isolated pore complex according to any one of claims 1 to 6. **Claim 8**: The CsgF peptide and the CsgG pore are covalently linked, and the covalent bond is (i) a disulfide bond, or (ii) a covalent bond involving a non-natural amino acid having an azide or alkyne, or a dibenzocyclooctyne (DBCO) group and / or a bicyclo[6.1.0]nonyne (BCN) group. The isolated pore complex according to any one of claims 1 to 7. **Claim 9** The interaction between the CsgF peptide and the CsgG pore is stabilized by hydrophobic interaction, electrostatic interaction or covalent bond interaction at positions corresponding to one or more of the following pairs of positions of SEQ ID NO: 6 and SEQ ID NO: 3: 1 and 153, 4 and 133, 5 and 136, 8 and 187, 8 and 203, 9 and 203, 11 and 142, 11 and 201, 12 and 149, 12 and 203, 26 and 191, and 29 and 144. The isolated pore complex according to any one of claims 1 to 6. **Claim 10** The CsgG pore has the following modifications of SEQ ID NO: 3 (i) Modification at one or more of positions Y51, N55 and F56 (ii) At least one substitution selected from R97W or R97Y and R93W or R93Y (iii) Deletion of V105, A106, and I107 of SEQ ID NO: 3 (iv) Deletion of one or more of positions R192, F193, I194, D105, Y196, Q197, R198, L199, and E201 of SEQ ID NO: 3 (v) At least one substitution selected from K94N / Q / R / F / Y / W / L / S, D43S, E44S, F48S / N / Q / Y / W / I / V / H / R / K, Q87N / R / K, N91K / R, R97F / Y / W / V / I / K / S / Q / H, E101I / L / A / H, N102K / Q / L / I / V / S / H, R110F / G / N, Q114R / K, R142Q / S, T150Y / A / V / L / S / Q / N (vi) Modification at one or more of positions I41, R93, A98, Q100, G103, T104, A106, I107, N108, L113, S115, T117, Y130, K135, E170, S208, D233, D238, E244, Q42, E44, L90, N91, I95, A99, E101 and Q114, (vii) At least one substitution selected from Y51A / I / V / S / T, N55A / I / V / S / T and F56 / A / I / V / S / T / Q, (viii) Substitution R97W, (ix) Deletion of F193, I194, D195, Y196, Q197, R198 and L199, or deletion of D195, Y196, Q197, R198 and L199, (x) Deletion of V105, A106 and I107, (xi) Substitution selected from K94Q and K94N, (xii) At least one substitution selected from Q42K or Q42R, E44N or E44Q, L90R or L90K, N91R or N91K, I95R or I95K, A99R or A99K, E101H, E101K, E101N, E101Q or E101T, and / or Q114K, (xiii) Substitution N55V, and / or (xiv) R or K at the position corresponding to R192 An isolated pore complex according to any one of claims 1 to 9, comprising at least one monomer comprising one or more modifications corresponding to those described above.

11. A double pore comprising two CsgG pores, wherein the CsgF peptide is inserted into the lumen of at least one of the CsgG pores, an isolated pore complex according to any one of claims 1 to 10.

12. An isolated pore complex according to any one of claims 1 to 11, wherein one or more residues of SEQ ID NO: 15, SEQ ID NO: 39, SEQ ID NO: 40, SEQ ID NO: 54, or SEQ ID NO: 55 are modified.

13. The isolated pore complex according to any one of claims 1 to 12, wherein the modification further comprises the introduction of cysteine, a hydrophobic amino acid, a charged amino acid, a non-natural reactive amino acid, or a photoreactive amino acid. **Claim 14** The isolated pore complex according to any one of claims 1 to 13, wherein the ratio of the CsgG monomer:cleaved CsgF peptide in the pores is 1:1.

Citation Information

Patent Citations

  • Mutant pore

    JP2017527284A

  • Mutant pore

    JP2019511217A

  • Mutant pore

    JP2019516344A

  • Mutant pore

    JP2019516346A