Novel Pore Monomers and Pores
Attaching CsgG pore monomers to CsgF peptides at multiple positions in nanopore complexes improves current range and SNR, addressing limitations in analyte discrimination in nanopore sensing.
Patent Information
- Application Number
- JP2025507384
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-08-09
- Filing Date
- 2023-08-09
- Publication Date
- 2025-08-15
AI Technical Summary
Existing nanopore sensing technologies for analyte detection and characterization, particularly for polynucleotides, face challenges in improving the current difference between nucleotides and signal-to-noise ratio, limiting the discrimination capability of sequencing systems.
Formation of pore complexes using CsgG pore monomers attached to CsgF peptides at two or more positions, enhancing the current range and signal-to-noise ratio (SNR) for improved analyte discrimination.
The increased current range and SNR in the pore complexes allow for better discrimination between analytes, enhancing the performance of nanopore sensing systems.
Smart Images

Figure 2025526704000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to novel pore monomer conjugates, pore complexes formed from the conjugates, and their use in the detection and characterization of analytes. [Background technology]
[0002] Nanopore sensing is an approach to analyte detection and characterization that relies on the observation of individual binding or interaction events between analyte molecules and ion-conducting channels. Two of the key elements of analyte characterization using nanopore sensing are (1) the control of analyte movement through the pore and (2) the discrimination of constituent building blocks as the analyte passes through the pore. During nanopore sensing, the narrowest part of the pore forms the most discriminatory part of the nanopore in terms of current signatures as a function of the analyte passing through. CsgG was identified as an ungated, nonselective protein secretion channel from Escherichia coli (Goyal et al., 2014) and has been used as a nanopore for analyte detection and characterization. Mutations to the wild-type CsgG pore that improve the properties of the pore in this context have also been disclosed (WO2016 / 034591, WO2017 / 149316, WO2017 / 149317, WO2017 / 149318, WO2018 / 211241, and WO2019 / 002893, all of which are incorporated by reference in their entirety).
[0003] For polynucleotide analytes, nucleotide discrimination is achieved by measuring the current as the polynucleotide passes through the pore. Because multiple nucleotides contribute to the observed current, the height of the channel constriction and the degree of interaction with the polynucleotide affect the relationship between the observed current and the polynucleotide sequence. Although the current range and signal-to-noise ratio for nucleotide discrimination have been improved through mutations in the CsgG pore, sequencing systems would have higher performance if the current difference between nucleotides could be further improved. Therefore, there is a need to identify novel methods for improving nanopore sensing characteristics. Summary of the Invention
[0004] The present inventors have surprisingly shown that pore complexes formed from pore-monomer conjugates in which the CsgG pore monomer is attached to the CsgF peptide at two or more positions exhibit increased current range and / or increased signal-to-noise ratio (SNR) during analyte characterization compared to conjugates with attachment at only one position. Both the increased current range and increased SNR improve the ability to discriminate between analytes as they pass through the pore. Previous experiments using CsgG and CsgF could not predict either the improved range or the improved SNR. Accordingly, the present invention provides pore-monomer conjugates comprising a CsgG pore monomer attached to a CsgF peptide, wherein the CsgF peptide is attached to the CsgG pore monomer at two or more positions.
[0005] The present invention also provides the following: - a construct comprising two or more covalently attached pore monomer conjugates of the invention, a pore complex comprising at least one pore monomer conjugate of the invention or at least one construct of the invention, wherein said CsgF peptide(s) form a constriction in said pore complex, a pore multimer comprising two or more pores, at least one of said pores being a pore complex of the invention; a membrane comprising a pore complex of the invention or a pore multimer of the invention, - a method for producing a pore monomer conjugate of the invention, the method comprising attaching the CsgF peptide to the CsgG pore at two or more positions, - a method for producing a pore complex of the invention or a pore multimer of the invention, comprising expressing at least one pore monomer conjugate of the invention or a construct of the invention and sufficient pore monomers or constructs to form said pore complex or said pore multimer in a host cell, and allowing said pore complex or said pore multimer to form in said host cell, - a method for producing a pore complex of the invention or a pore multimer of the invention, comprising contacting in vitro at least one pore monomer conjugate of the invention or a construct of the invention with sufficient pore monomer or construct and allowing the formation of said pore complex or said pore multimer, - a method for determining the presence, absence or one or more characteristics of a target analyte, comprising: (i) contacting the target analyte with the pore complex or pore multimer of the present invention such that the target analyte migrates relative to the pore complex or pore multimer of the present invention; (ii) obtaining one or more measurements as the analyte moves relative to the pore complex or pore multimer, thereby determining the presence, absence, or one or more properties of the analyte. - the use of a pore complex of the invention or a pore multimer of the invention in determining the presence, absence or one or more properties of a target analyte, a polynucleotide encoding a pore monomer conjugate of the invention or a construct of the invention, a kit for characterizing a target analyte, comprising (a) a pore complex of the invention or a pore multimer of the invention, and (b) a membrane component; a kit for characterizing a target polynucleotide or a target polypeptide, comprising: (a) a pore complex of the present invention or a pore multimer of the present invention; and (b) a polynucleotide-binding protein; - a device for characterizing a target polynucleotide or a target polypeptide in a sample, the device comprising: (a) a plurality of pore complexes of the invention or a plurality of pore multimers of the invention; and (b) a plurality of polynucleotide binding proteins; an array comprising a plurality of membranes of the invention, a system comprising: (a) a membrane of the invention or an array of the invention; (b) means for applying a potential between the membrane(s); and (c) means for detecting an electrical or optical signal between the membrane(s); a device comprising a pore complex of the invention or a pore multimer of the invention inserted into an in vitro membrane, - A device produced by a method, the method comprising: (i) obtaining a pore complex of the present invention or a pore multimer of the present invention; and (ii) contacting the pore complex or pore multimer with an in vitro membrane such that the pore complex or pore multimer is inserted into the in vitro membrane. [Brief explanation of the drawings]
[0006] [Figure 1]Figure 1 shows the structure and size of the wild-type CsgG pore from E. coli K12 (the databank access code for this structure is 4UV3). Distances shown are measured from backbone to backbone of the amino acids that form the pore structure. The CsgG pore is a tightly interconnected, symmetric nonameric pore resembling a crown. The overall height is 98 Å, and the maximum outer diameter is 120 Å. It defines a central channel and consists of three parts: (A) the cap region, (B) the constriction region, and (C) the transmembrane beta-barrel region. The axial length, or height, of the cap is 39 Å. The inner diameter is 43 Å, and the opening is 66 Å. The beta-barrel has 36 strands, an axial length of 39 Å, and an inner diameter of 55 Å. The transition between the pore cap and the beta-barrel is abrupt, with a constriction located between them at the level of the predicted lipid-aqueous interface. The constriction is approximately 18.5 Å in diameter and 20 Å in length along the axis of the channel. DETAILED DESCRIPTION OF THE INVENTION
[0007] Sequence Listing Description SEQ ID NO: 1 shows the polynucleotide sequence of wild-type E. coli CsgG from strain K12, including the signal sequence (Gene ID: 945619).
[0008] SEQ ID NO: 2 shows the amino acid sequence of wild-type E. coli CsgG, including the signal sequence (Uniprot accession number P0AEA2).
[0009] SEQ ID NO: 3 shows the amino acid sequence of wild-type E. coli CsgG as the mature protein (Uniprot accession number P0AEA2).
[0010] SEQ ID NO: 4 shows the polynucleotide sequence of wild-type E. coli CsgF from strain K12, including the signal sequence (Gene ID: 945622).
[0011] SEQ ID NO: 5 shows the amino acid sequence of wild-type E. coli CsgF, including the signal sequence (Uniprot accession number P0AE98).
[0012] SEQ ID NO: 6 shows the amino acid sequence of wild-type E. coli CsgF as the mature protein (Uniprot accession number P0AE98).
[0013] All publications, patents, and patent applications cited herein, whether supra or infra, are hereby incorporated by reference in their entirety. All publications, patents, and patent applications mentioned herein are hereby incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent that publications and patents or patent applications incorporated by reference conflict with the present invention contained herein, the present specification is intended to supersede and / or supersede any such conflicting material.
[0014] While the present invention will be described with respect to particular embodiments and with reference to certain drawings, the present invention is not limited thereto, but rather only by the claims. Any reference signs in the claims shall not be construed as limiting the scope. Of course, it should be understood that not necessarily all aspects or advantages can be achieved in accordance with any particular embodiment of the invention. Thus, for example, one skilled in the art will recognize that the invention can be embodied or carried out in a manner that achieves or optimizes one advantage or group of advantages taught herein, without necessarily achieving other aspects or advantages that may be taught or suggested herein.
[0015] Additionally, as used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to a "polynucleotide" includes two or more polynucleotides, reference to a "polynucleotide-binding protein" includes two or more such proteins, reference to a "helicase" includes two or more helicases, reference to a "monomer" refers to two or more monomers, reference to a "pore" includes two or more pores, etc.
[0016] In all discussions herein, the standard single-letter codes for amino acids are used, as follows: alanine (A), arginine (R), asparagine (N), aspartic acid (D), cysteine (C), glutamic acid (E), glutamine (Q), glycine (G), histidine (H), isoleucine (I), leucine (L), lysine (K), methionine (M), phenylalanine (F), proline (P), serine (S), threonine (T), tryptophan (W), tyrosine (Y), and valine (V). Standard substitution notation is also used, i.e., Q42R means that Q at position 42 is substituted with R.
[0017] In sections herein where different amino acids at a particular position are separated by a / symbol, the / symbol means "or." For example, Q87R / K means Q87R or Q87K. When different positions are separated by a / symbol, the / symbol means "and," e.g., Y51 / N55 means Y51 and N55.
[0018] The general definitions in WO2019 / 002893 are incorporated herein by reference in their entirety.
[0019] Pore Monomer Conjugate The present invention provides a pore monomer conjugate comprising a CsgG pore monomer attached to a CsgF peptide. The CsgF peptide is attached to the CsgG pore monomer at two or more positions, for example, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, or ten or more positions. The CsgF peptide is preferably covalently attached to the CsgG pore monomer at two or more positions, for example, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, or ten or more positions.
[0020] In the context of the present invention, attachment at two or more positions means that two or more pairs of residues in the CsgF peptide and the CsgG pore monomer are attached to each other, preferably covalently attached. For example, residue 1 in the CsgF peptide may be attached to residue 153 of the CsgG pore monomer (one pair of residues = 1 and 153), and residue 30 in the CsgF peptide may be attached to residue 196 in the CsgG pore monomer (second pair of residues = 30 and 196).
[0021] SEQ ID NO: 3 shows the amino acid sequence of wild-type E. coli CsgG as a mature protein. The two or more positions within the CsgG pore monomer are preferably selected from residues 47-54, 57, 59, 60, 130-134, 136, 137, 138, 140, 142-145, 147, 149, 151, 153, 155, 181, 183, 185, 187, 189, 191, 193, 195-199, 201, 203, 205, 207, 209 and 211-212 within the CsgG pore monomer. The two or more positions within the CsgG pore monomer are preferably selected from residues corresponding to positions 47-54, 57, 59, 60, 130-134, 136, 137, 138, 140, 142-145, 147, 149, 151, 153, 155, 181, 183, 185, 187, 189, 191, 193, 195-199, 201, 203, 205, 207, 209 and 211-212 of SEQ ID NO:3.
[0022] SEQ ID NO: 6 shows the amino acid sequence of wild-type E. coli CsgF as a mature protein. The two or more positions in the CsgF peptide are preferably selected from the N-terminus and residues 1-35 in the CsgF peptide. The two or more positions in the CsgF peptide are preferably selected from the N-terminus and residues corresponding to positions 1-35 of SEQ ID NO: 6. The N-terminus is the amino group of the first residue (i.e., residue 1) of the CsgF peptide. In this context, residue 1 refers to the side chain of residue 1 or the residue corresponding to position 1 of SEQ ID NO: 6. The two or more positions are preferably the following positions / residues in the CsgF peptide or positions / residues in the CsgF peptide corresponding to the following positions of SEQ ID NO: 6: the N-terminus and any one of 1 to 35, 1 and the N-terminus and any one of 2 to 35, 2 and the N-terminus, 1 and any one of 3 to 35, 3 and the N-terminus, 1 to 2 and any one of 4 to 35, 4 and the N-terminus, 1 to 3 and any one of 5 to 35, 5 and the N-terminus, 1 to 4 and any one of 6 to 35, 6 and the N-terminus, 1 to 5 and any one of 7 to 35, 7 and the N-terminus, 1 to 6 and any one of 8 to 35, 8 and the N-terminus, 1 to 7 and any one of 9 to 35, 9 and the N-terminus, 1 to 8 and any one of 10 to 35, 10 and the N-terminus, 1 to 9 and any one of 11 to 35, 11 and the N-terminus, 1 to 10 and any one of 12 to 35, 12 and N-terminus, any of 1 to 11 and 13 to 35, 13 and N-terminus, any of 1 to 12 and 14 to 35, 14 and N-terminus, any of 1 to 13 and 15 to 35, 15 and N-terminus, any of 1 to 14 and 16 to 35, 16 and N-terminus, any of 1 to 15 and 17 to 35, 17 and N-terminus, any of 1 to 16 and 18 to 35, 18 and N-terminus, 1 to 17 and 19 to 35 any of them, 19 and the N-terminus, any of them among 1 to 18 and 20 to 35, 20 and the N-terminus, any of them among 1 to 19 and 21 to 35, 21 and the N-terminus, any of them among 1 to 20 and 22 to 35, 22 and the N-terminus, any of them among 1 to 21 and 23 to 35, 23 and the N-terminus, any of them among 1 to 22 and 24 to 35, 24 and the N-terminus, any of them among 1 to 23 and 25 to 35, 25 and the N-terminus,Any of 1 to 24 and 26 to 35, 26 and the N-terminus, any of 1 to 25 and 27 to 35, 27 and the N-terminus, any of 1 to 26 and 28 to 35, 28 and the N-terminus, any one of 1 to 27 and 29 to 35, 29 and the N-terminus, any of 1 to 28 and 30 to 35, 30 and the N-terminus, any of 1 to 29 and 31 to 35, 31 and the N-terminus, any of 1 to 30 and 32 to 35, 32 and the N-terminus, any of 1 to 31 and 33 to 35, 33 and the N-terminus, any of 1 to 32 and 34 to 35, 34 and the N-terminus, any of 1 to 33 and 35, 35 and the N-terminus, any of 1 to 34.
[0023] Preferred combinations of two or more positions are shown in the table below. Each column indicates the positions / residues in the CsgF peptide and CsgG pore monomer, or the positions in SEQ ID NO:6 and SEQ ID NO:3 to which the positions / residues in the CsgF peptide and CsgG pore monomer correspond. The two or more positions may be in any two or more rows in the table. In each row, the position / residue in the CsgF peptide, or the position / residue corresponding to a position in SEQ ID NO:6, may be attached, preferably covalently attached, to any of the listed positions / residues in the CsgG pore monomer, or to any position / residue corresponding to a listed position in SEQ ID NO:3. For example, in line 3, the residue corresponding to position 2 of the CsgF peptide, or position 2 of SEQ ID NO: 6, is preferably attached, more preferably covalently attached, to a residue corresponding to any of positions 47-54, 57, 59, 60, 130-134, 136, 151, 153, 155, 181, 183, 185, 207, 209, 211-212 in the CsgG pore monomer, or any of positions 47-54, 57, 59, 60, 130-134, 136, 151, 153, 155, 181, 183, 185, 207, 209, 211-212 of SEQ ID NO: 3. [Table 1-1] [Table 1-2]
[0024] One of the two or more attachments is preferably attached to a cysteine residue at position 153 of the CsgG pore monomer, and preferably comprises the N-terminus of a covalently attached CsgF peptide. One of the two or more attachments is preferably attached to a cysteine residue in the CsgG pore monomer corresponding to position 153 of SEQ ID NO:3, and preferably comprises the N-terminus of a covalently attached CsgF peptide.
[0025] One of the two or more attachments is preferably attached to a cysteine residue at position 133 of the CsgG pore monomer, preferably comprising a covalently attached CsgF peptide at position 4. One of the two or more attachments is preferably attached to a cysteine residue in the CsgG pore monomer corresponding to position 133 of SEQ ID NO:3, preferably comprising a position in the CsgF peptide corresponding to a covalently attached position 4 of SEQ ID NO:6.
[0026] One of the two or more attachments is preferably attached to a cysteine residue at position 153 of the CsgG pore monomer, preferably comprising a covalently attached CsgF peptide at position 4. One of the two or more attachments is preferably attached to a cysteine residue in the CsgG pore monomer corresponding to position 153 of SEQ ID NO:3, preferably comprising a position in the CsgF peptide corresponding to a covalently attached position 4 of SEQ ID NO:6.
[0027] One of the two or more attachments is preferably attached to any one of positions 193, 195, 196 and 197 of the CsgG pore monomer, and preferably comprises a covalently attached residue of the CsgF peptide at any one of positions 30, 31, 32 and 33. One of the two or more attachments is preferably attached to a position in the CsgG pore monomer corresponding to any one of positions 193, 195, 196 and 197 of SEQ ID NO:3, and preferably comprises a covalently attached residue of the CsgF peptide corresponding to any one of positions 30, 31, 32 and 33 of SEQ ID NO:6.
[0028] Corresponding positions can be determined by standard techniques in the art, for example, the PILEUP and BLAST algorithms described below can be used to align the sequence of the CsgG pore monomer with SEQ ID NO: 3 and thus identify corresponding residues.
[0029] Attachment at two or more positions preferably comprises one or more reactive groups that react with lysine, cysteine, tyrosine, serine, threonine, proline, tryptophan, arginine, histidine, methionine, or phenylalanine in the CsgG pore monomer. Attachment at two or more positions preferably comprises a reaction between a position, residue, or linker in the CsgF peptide and a lysine, cysteine, tyrosine, serine, threonine, proline, tryptophan, arginine, histidine, methionine, or phenylalanine in the CsgG pore monomer. Attachment at all of the two or more positions preferably comprises one or more reactive groups that react with lysine, cysteine, tyrosine, serine, threonine, proline, tryptophan, arginine, histidine, methionine, or phenylalanine in the CsgG pore monomer. Attachment at all of the two or more positions preferably involves a reaction between a position, residue, or linker in the CsgF peptide and a lysine, cysteine, tyrosine, serine, threonine, proline, tryptophan, arginine, histidine, methionine, or phenylalanine in the CsgG pore monomer.
[0030] The lysine, cysteine, tyrosine, serine, threonine, proline, tryptophan, arginine, histidine, methionine, or phenylalanine may be native to the CsgG pore monomer, or may be introduced into the CsgG pore monomer, preferably by substitution or addition.
[0031] Reactive groups that react with lysine include, but are not limited to, maleimides, activated esters, anhydrides, carbonates, isocyanates, isothiocyanates, a range of other acylating and alkylating agents, oxidatively coupled O-aminophenols, aldehydes, activated carbodiimides, ketenes, sulfonyl halides, fluorosulfates, and sulfonyltriazoles. Positions, residues, or linkers may also be attached to lysines using periodate oxidation, reductive amination, transamination, aniline / arylamine coupling via oxidative coupling, azaelectrocyclization, iminoboronic acid formation, or arenediazonium salt coupling.
[0032] Reactive groups that react with cysteine include, but are not limited to, haloacetamides and other alpha-halocarbonyls, maleimides, acrylates, vinyl sulfones, vinyl pyridines, epoxides, oxanorbornadienes, methylsulfonyl-functional heteroaromatics, allenes, allyl selenosulfates, perfluoroaromatics, thiol-ene and thiol-click chemistry, pyridyldithiols, vinyl sulfones, sulfonyl halides, fluorosulfates, and sulfonyl triazoles. Positions, residues, or linkers may also be attached to cysteine using stress-relief alkylation, nickel(II)-catalyzed oxidative coupling, oxidative coupling with aminophenols, conjugation with allenes (in the presence of gold catalysts or allyl selenosulfates), native chemical ligation, Pd-catalyzed arylation / alkynylation, or allylation followed by cross-metathesis.
[0033] Reactive groups that react with tyrosine include, but are not limited to, sulfonyl halides, fluorosulfates, and sulfonyltriazoles. Positions, residues, or linkers may also be attached to tyrosine using O-alkylation, oxidative coupling of tyrosine, including hydrazone and oxime condensation, electron-deficient alkynes such as alkynones, alkynoates, amides, or esters, addition reactions with cyclic diazodicarboxamides, Pd-catalyzed alkylation, diazonium salts, or aldehydes, Mannich reactions with imines formed from cyclic diazodicarboxamides, and modification with rhodium carbenoids.
[0034] Reactive groups that react with serine or threonine include, but are not limited to, sulfonyl halides, fluorosulfates, and sulfonyltriazoles. Positions, residues, or linkers may also be attached to serine or threonine using periodate oxidation followed by transamination of a ketone / aldehyde with a hydrazide / alkoxyamine. The resulting aldehyde / ketone may also be modified by aldol ligation.
[0035] The position, residue, or linker may be attached to the proline at the N-terminus using oxidative coupling with O-aminophenol.
[0036] Reactive groups that react with tryptophan include, but are not limited to, aldehydes, ketones, and tetrazoles. Positions, residues, or linkers may also be attached to tryptophan using condensation reactions, modification with rhodium carbenoids, conjugation with N / O-centered radicals, and N-terminal Trp modification using the Pictet-Spengler reaction.
[0037] The position, residue, or linker may be attached to the arginine using condensation with an α,β-dicarbonyl compound.
[0038] Reactive groups that react with histidine include, but are not limited to, vinyl sulfone, sulfonyl halides, fluorosulfates, and sulfonyltriazoles. Positions, residues, or linkers may be attached to histidine using C2 alkylation and N3 alkylation / thiophosphorylation.
[0039] A position, residue, or linker may be attached to the methionine using S-alkylation / imidation.
[0040] Positions, residues, or linkers may be attached to the phenylalanine using modification with rhodium carbenoids.
[0041] Attachment at two or more positions preferably includes one or more reactive groups that react with any amino acid within the CsgG pore monomer. Attachment at all of the two or more positions preferably includes one or more reactive groups that react with any amino acid within the CsgG pore monomer. Reactive groups that react with any amino acid include, but are not limited to, activated esters, anhydrides, carbonates, isocyanates, isothiocyanates, and a range of other acylating and alkylating agents, oxidatively coupled O-aminophenols, aldehydes, activated carbodiimides, ketenes, transaminations, and vinylboronic acids.
[0042] Attachment at two or more positions preferably involves reacting a position, residue, or linker in the CsgF peptide with any amino acid in the CsgG pore monomer. Attachment at all two or more positions preferably involves reacting a position, residue, or linker in the CsgF peptide with any amino acid in the CsgG pore monomer. A position, residue, or linker may be attached to any amino acid using periodate oxidation or reductive amination.
[0043] The attachments at two or more positions preferably include one or more reactive groups that undergo click chemistry. The attachments at all of the two or more positions preferably include one or more reactive groups that undergo click chemistry. Suitable click chemistries include, but are not limited to, CuAAC azide / alkyne, Staudinger ligation, stress-promoted azide-alkyne cycloaddition, and the inverse electron demand Diels-Alder reaction between 1,2,4,5-tetrazine and a stressed alkene.
[0044] All of the above descriptions regarding reactive groups and reactions for attaching a CsgF peptide to a residue / amino acid within a CsgG pore monomer equally apply to attaching a CsgG pore monomer to a CsgF peptide. Any reactive group or reaction may be used for attachment within a CsgF peptide. Specific residues within a CsgF peptide may be inherent to the protein. Specific residues may also be introduced into a CsgF peptide, preferably by substitution or addition. Those skilled in the art can attach two proteins at more than one position, preferably by covalent attachment.
[0045] Attachment at two or more positions preferably involves two or more versions of the same or similar reactive group, such as maleimide. Attachment at two or more positions preferably involves two or more versions of the same or similar reaction. Attachment at two or more positions preferably involves two or more maleimide-containing linkers. Attachment at two or more positions preferably involves two or more maleimide reactions. Any of the maleimide groups and linkers discussed above may be used.
[0046] Attachment at two or more positions preferably comprises two or more different reactive groups. Attachment at two or more positions preferably comprises two or more different reactions. The two or more reactive groups or reactions may be any of those discussed above in relation to the CsgF peptide and / or CsgG pore monomer.
[0047] The CsgF peptide is preferably attached to the CsgG pore monomer using two or more linkers, for example, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, or ten or more linkers. The two or more linkers may be the same. The two or more linkers may be different. One skilled in the art can design two or more linkers for use in the present invention. The two or more linkers may be any of the linkers discussed below with respect to the constructs of the present invention.
[0048] The two or more linkers preferably comprise or consist of a linear carbon chain of 2, 3, 4, 5, 6, or more carbon atoms and / or a saturated or unsaturated cyclic group containing 3, 5, or 6 carbon atoms. One or more of the two or more linkers, for example, all, are preferably maleimide-containing linkers. The maleimide group may be used to react with cysteines in the CsgF peptide and / or CsgG pore monomer. The maleimide-containing linker preferably comprises or consists of a maleimide group and a linear carbon chain of 2, 3, 4, 5, 6, or more carbon atoms. The linear carbon chain is typically bonded to a nitrogen atom in the maleimide group. The linear carbon chain also preferably comprises a terminal carboxyl group. This carboxyl group can form an amide bond with an amino acid in the CsgF peptide. The maleimide-containing linker is preferably maleimidoacetic acid, maleimidopropionic acid, maleimidobutyric acid, maleimidopentanoic acid, or maleimidohexanoic acid. Any combination of these linkers may be used in two or more linkers. The maleimide-containing linker is most preferably maleimidopropionic acid.
[0049] The distance between the CsgG pore monomer and the CsgF peptide in the pore monomer conjugate and / or the length of the linker is preferably less than about 3.00 nm, e.g., less than about 2.90 nm, less than about 2.80 nm, less than about 2.70 nm, less than about 2.60 nm, less than about 2.50 nm, less than about 2.40 nm, less than about 2.30 nm, less than about 2.20 nm, less than about 2.10 nm, less than about 2.00 nm. The distance / length is less than about 1.90 nm, less than about 1.80 nm, less than about 1.70 nm, less than about 1.60 nm, less than about 1.50 nm, less than about 1.40 nm, less than about 1.30 nm, less than about 1.20 nm, less than about 1.10 nm, less than about 1.00 nm, less than about 0.90 nm, less than about 0.80 nm, less than about 0.70 nm, less than about 0.60 nm, less than about 0.50 nm, or less than about 0.40 nm. This distance / length can be achieved using any of the specific maleimide-containing linkers discussed above, including maleimide acetic acid, maleimide propionic acid, maleimide butyric acid, maleimide pentanoic acid, or maleimide hexanoic acid. The linker is most preferably maleimide propionic acid.
[0050] The distance and / or linker length between the CsgG pore monomer and the CsgF peptide in the pore monomer conjugate is preferably less than about 1.20 nm. This distance / length can be achieved using maleimidohexanoic acid, as described in more detail above. The distance and / or linker length between the CsgG pore monomer and the CsgF peptide in the pore monomer conjugate is preferably less than about 0.8 nm. This distance / length can be achieved using maleimidopropionic acid, as described above.
[0051] The distance between the CsgG pore monomer and the CsgF peptide in the pore monomer conjugate and / or the length of the linker is preferably about 0.40 nm to about 3.00 nm, for example, about 0.45 nm to about 2.80 nm, about 0.50 nm to about 2.50 nm, about 0.55 nm to about 2.20 nm, about 0.60 nm to about 2.00 nm, about 0.65 nm to about 1.50 nm, about 0.70 nm to about 1.40 nm, about 0.75 nm to about 1.30 nm, about 0.80 nm to about 1.20 nm, about 0.85 nm to about 1.10 nm, and about 0.90 nm to about 1.00 nm. The distance between the CsgG pore monomer and the CsgF peptide in the pore monomer conjugate and / or the length of the linker is preferably about 0.50 nm to about 1.50 nm. The distance and / or linker length between the CsgG pore monomer and the CsgF peptide in the pore monomer conjugate is preferably about 0.60 nm to about 1.20 nm. This distance / length can be achieved using any of the specific maleimide-containing linkers discussed above, including maleimidoacetic acid, maleimidopropionic acid, maleimidobutyric acid, maleimidopentanoic acid, or maleimidohexanoic acid. The linker is most preferably maleimidopropionic acid.
[0052] The pore monomer conjugates of the present invention are capable of forming pores or pore complexes, which can be measured using common methods, including any of the methods described in WO2016 / 034591, WO2017 / 149316, WO2017 / 149317, WO2017 / 149318, WO2018 / 211241 and WO2019 / 002893 (all of which are incorporated herein by reference) and in the Examples.
[0053] CsgG pore monomer A CsgG pore monomer is a monomer capable of forming a CsgG pore. Such monomers are known in the art, particularly from WO2019 / 002893 (incorporated herein by reference in its entirety). The CsgG pore preferably comprises one or more of: (a) a capping region, (b) a constriction region, and (c) a transmembrane beta-barrel region, such as (a), (b), (c), (a) and (b), (a) and (c), (b) and (c), or (a), (b) and (c). The CsgG pore monomer preferably comprises one or more of: (a) a capping region, (b) a constriction region, and (c) a transmembrane beta-barrel region, such as (a), (b), (c), (a) and (b), (a) and (c), (b) and (c), or (a), (b) and (c). The residues of SEQ ID NO: 3 that form these regions are defined below. The CsgG pore formed by the monomers may have any structure, but preferably has or includes the structure of the wild-type CsgG pore (Figure 1). The protein structure of CsgG defines a channel or hole that allows the translocation of molecules and ions from one side of the membrane to the other.
[0054] As used interchangeably herein, the terms "constriction," "opening," "constriction region," "channel constriction," or "constriction site" refer to an opening defined by the luminal surface of a pore or pore complex that acts to allow the passage of ions and target molecules (e.g., but not limited to, polynucleotides or individual nucleotides) but not other non-target molecules through the pore or pore complex channel. The constriction(s) are typically the narrowest opening(s) within the pore or pore complex or within the channel defined by the pore or pore complex. The constriction(s) may serve to limit the passage of molecules through the pore. The size of the constriction is typically an important factor determining the suitability of a pore or pore complex for analyte characterization. If the constriction is too small, the molecules to be characterized will not be able to pass. However, to achieve the greatest effect on ion flow through the channel, the constriction should not be too large. For example, the constriction should not be wider than the solvent-accessible lateral diameter of the target analyte. Ideally, any constriction should be as close as possible to the lateral diameter of the analyte being passed through. The CsgF peptide and CsgG pore monomer typically each provide at least one constriction, such that a pore complex of the invention comprises two or more constrictions.
[0055] The CsgG pore may be of any size, but preferably has the dimensions of a wild-type CsgG pore (Figure 1). The CsgG pore preferably has an outer diameter of about 100 to about 150 Å at its widest point, e.g., about 110 to about 140 Å or about 115 to about 125 Å at its widest point. The CsgG pore preferably has an outer diameter of about 120 Å at its widest point. The CsgG pore preferably has a total length of about 80 to about 120 Å, e.g., about 90 to about 110 Å or about 95 to about 105 Å. The CsgG pore preferably has a total length of about 98 Å. References to "total length" and "length" refer to the length of the pore or pore region when viewed from the side (see, e.g., the side view in Figure 1).
[0056] The cap region preferably has a length of about 20 to about 60 Å, for example, about 30 to about 50 Å or about 35 to about 45 Å. The cap region preferably has a length of about 39 Å. The channel defined by the cap region preferably has an opening with a diameter of about 45 to about 85 Å, for example, a diameter of about 55 to about 75 Å or about 60 to about 70 Å. The channel defined by the cap region preferably has an opening with a diameter of about 66 Å. The channel defined by the cap region preferably has a diameter at its narrowest point of about 30 to about 70 Å, for example, a diameter at its narrowest point of about 35 to about 60 Å or about 40 to about 50 Å. The channel defined by the cap region preferably has a diameter at its narrowest point of about 43 Å.
[0057] The constriction region preferably has a length of about 5 to about 40 Å, for example, about 10 to about 30 Å or about 15 to about 25 Å. The constriction region preferably has a length of about 20 Å. The channel defined by the constriction region preferably has a diameter at its narrowest point of about 2 to about 40 Å, for example, a diameter at its narrowest point of about 5 to about 35 Å, about 8 to about 25 Å, or about 10 to about 20 Å. The channel defined by the constriction region preferably has a diameter of about 9 Å or 12 Å. The channel defined by the constriction region preferably has a diameter of about 18.5 Å. The constriction preferably has a diameter of about 2 to about 40 Å, for example, a diameter of about 5 to about 35 Å, about 8 to about 25 Å, or about 10 to about 20 Å. The constriction preferably has a diameter of about 9 Å or 12 Å. The constriction preferably has a diameter of about 12 Å.
[0058] The transmembrane beta barrel region preferably has a length of about 20 to about 60 Å, for example, about 30 to about 50 Å or about 35 to about 45 Å. The transmembrane beta barrel region preferably has a length of about 39 Å. The channel defined by the transmembrane beta barrel region preferably has a diameter at its narrowest point of about 35 to about 75 Å, for example, a diameter at its narrowest point of about 45 to about 65 Å or about 50 to about 60 Å. The channel defined by the transmembrane beta barrel region preferably has a diameter at its narrowest point of about 55 Å.
[0059] All of the above measurements are based on backbone-to-backbone measurements of the amino acids that make up the different regions (as shown in Figure 1).
[0060] SEQ ID NO: 3 shows the sequence of wild-type E. coli CsgG as a mature protein. Residues 1-41, 64-131, 156-180, and 212-262 of SEQ ID NO: 3 form the cap region. Residues 42-63 of SEQ ID NO: 3 form the constriction region. Residues 132-155 and 181-211 of SEQ ID NO: 3 form the transmembrane beta-barrel region.
[0061] The CsgG pore monomer is preferably a variant of SEQ ID NO: 3. Variant CsgG monomers may also be referred to as modified CsgG pore monomers or mutant CsgG pore monomers. Modifications or mutations in the variants include, but are not limited to, any one or more of the modifications disclosed herein, or combinations of the modifications. The CsgG pore monomer may also be a CsgG homolog monomer. A CsgG homolog pore monomer is a polypeptide that has at least 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, or 99% complete sequence identity to the wild-type E. coli CsgG set forth in SEQ ID NO: 3. A CsgG homolog is also referred to as a polypeptide that contains the PFAM domain PF03783, which is characteristic of CsgG-like proteins. A list of currently known CsgG homologs and CsgG architectures can be found at http: / / pfam.xfam.org / / family / PF03783.
[0062] Across the entire length of the amino acid sequence of SEQ ID NO: 3, a variant is preferably at least 40% homologous to that sequence based on amino acid identity. More preferably, a variant may be at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, more preferably at least 95%, 97%, or 99% homologous to the amino acid sequence of SEQ ID NO: 3 across the entire sequence. Across the entire length of the amino acid sequence of SEQ ID NO: 3, a variant is preferably at least 40% homologous to that sequence. More preferably, a variant may be at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, more preferably at least 95%, 97%, or 99% identical to SEQ ID NO: 3 across the entire sequence.
[0063] Sequence identity can also refer to fragments or portions of the CsgG pore monomer. Thus, a sequence may have less than 40% overall sequence homology / identity with SEQ ID NO:3, but the sequence of a particular region, domain, or subunit can share at least 80%, 90%, or up to 99% sequence homology / identity with the corresponding region of SEQ ID NO:3. There may be at least 80%, e.g., at least 85%, 90%, or 95% amino acid identity ("hard homology") over a stretch of 100 or more, e.g., 125, 150, 175, or 200 or more consecutive amino acids. The CsgG pore monomer is preferably a variant of SEQ ID NO:3 comprising a sequence that is at least 40% homologous to the cap region of SEQ ID NO:3 (residues 1-41, 64-131, 156-180, and 212-262). More preferably, variants may comprise sequences that are at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, more preferably at least 95%, 97% or 99% homologous based on amino acid identity to amino acid residues 1-41, 64-131, 156-180 and 212-262 of SEQ ID NO: 3. Variants preferably comprise sequences that are at least 40% identical to residues 1-41, 64-131, 156-180 and 212-262 of SEQ ID NO: 3. More preferably, variants may comprise sequences that are at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, more preferably at least 95%, 97% or 99% homologous to residues 1-41, 64-131, 156-180 and 212-262 of SEQ ID NO: 3. Homology and / or identity is typically measured over the entire length of the cap region.
[0064] The CsgG pore monomer is preferably a variant of SEQ ID NO: 3 comprising a sequence that is at least 40% homologous to the constriction region (residues 42-63) of SEQ ID NO: 3. More preferably, the variant may comprise a sequence that is at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, more preferably at least 95%, 97% or 99% homologous based on amino acid identity to amino acid residues 42-63 of SEQ ID NO: 3. The variant preferably comprises a sequence that is at least 40% identical to residues 42-63 of SEQ ID NO: 3. More preferably, a variant may comprise a sequence that is at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, and more preferably at least 95%, 97% or 99% identical to residues 42-63 of SEQ ID NO: 3. Homology and / or identity is typically measured over the entire length of the constriction region.
[0065] The CsgG pore monomer is preferably a variant of SEQ ID NO: 3 comprising a sequence that is at least 40% homologous to the transmembrane beta-barrel region (residues 132-155 and 181-211) of SEQ ID NO: 3. More preferably, the variant may comprise a sequence that is at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, more preferably at least 95%, 97%, or 99% homologous based on amino acid identity to amino acid residues 132-155 and 181-211 of SEQ ID NO: 3. The variant preferably comprises a sequence that is at least 40% identical to residues 132-155 and 181-211 of SEQ ID NO: 3. More preferably, variants may comprise sequences that are at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, more preferably at least 95%, 97% or 99% identical to residues 132-155 and 181-211 of SEQ ID NO: 3. Homology and / or identity is typically measured over the entire length of the transmembrane beta-barrel region.
[0066] The CsgG pore monomer is highly conserved (as can be readily seen from Figures 45-47 of WO2017 / 149317). Furthermore, knowledge of the mutations relative to SEQ ID NO:3 allows for the determination of equivalent positions of mutations in the CsgG pore monomer other than those of SEQ ID NO:3.
[0067] Thus, reference to a mutant CsgG pore monomer comprising a variant of the sequence set forth in SEQ ID NO:3 and specific amino acid mutations thereof as described in the claims and elsewhere herein also encompasses mutant CsgG pore monomers comprising variants of any of the sequences set forth in SEQ ID NOs:68-88 of WO2019 / 002893 (incorporated herein by reference in their entirety) and their corresponding amino acid mutations. The CsgG pore monomer may also be any of the sequences set forth in CN113773373A, CN113896776A, CN113912683A, and CN113754743A, or variants thereof. It will further be understood that the present invention extends to other variant CsgG pore monomers not expressly identified herein that exhibit highly conserved regions.
[0068] Standard methods in the art can be used to determine homology. For example, the UWGCG Package provides the BESTFIT program, which can be used to calculate homology, for example, using its default settings (Devereux et al. (1984) Nucleic Acids Research 12, p387-395). For example, the PILEUP and BLAST algorithms can be used to calculate homology or align sequences (identify equivalent residues or corresponding sequences, typically using their default settings), as described in Altschul SF (1993) J Mol Evol 36:290-300, Altschul, SF et al. (1990) J Mol Biol 215:403-10. Software for performing BLAST analysis is publicly available through the National Center for Biotechnology Information (http: / / www.ncbi.nlm.nih.gov / ).
[0069] SEQ ID NO: 3 is the wild-type CsgG pore monomer from E. coli Str. K-12 substrain MC4100. Variants of SEQ ID NO: 3 may include any of the substitutions present in other CsgG homologs. Preferred CsgG homologs are set forth in SEQ ID NOs: 68-88 of WO 2019 / 002893 (incorporated herein by reference in their entireties). Variants may include one or more combinations of the substitutions present in SEQ ID NOs: 68-88 of WO 2019 / 002893 (incorporated herein by reference in their entireties) compared to SEQ ID NO: 3, including one or more substitutions, one or more conservative mutations, one or more deletions, or one or more insertion mutations, for example, a deletion or insertion of 1 to 10 amino acids, such as 2 to 8 or 3 to 6 amino acids.
[0070] The CsgG pore monomer in the pore monomer conjugates of the invention typically retains the ability to form the same 3D structure as a wild-type CsgG pore monomer, for example, a CsgG pore monomer having the sequence of SEQ ID NO: 3. The 3D structure of CsgG is known in the art and is disclosed, for example, in Goyal et al (2014) Nature 516(7530):250-3. In addition to the mutations described herein, any number of mutations may be made in the wild-type CsgG sequence, provided that the CsgG pore monomer retains the improved properties conferred by the mutations of the invention.
[0071] Typically, a CsgG pore monomer retains the ability to form a structure comprising five alpha-helices and five beta strands. It is therefore contemplated that further mutations may be made in any of these regions in any CsgG pore monomer without affecting the ability of the monomer to form a pore capable of translocating a polynucleotide. It is also contemplated that deletion of one or more amino acids may be made in any of the loop regions connecting the alpha helices and beta strands, and / or in the N-terminal and / or C-terminal regions of a CsgG pore monomer, without affecting the ability of the monomer to form a pore capable of translocating a polynucleotide.
[0072] In addition to the above, amino acid substitutions may be made to the amino acid sequence of SEQ ID NO: 3, for example, up to 1, 2, 3, 4, 5, 10, 20, or 30 substitutions. Conservative substitutions replace amino acids with other amino acids of similar chemical structure, similar chemical properties, or similar side chain volume. The introduced amino acids may have similar polarity, hydrophilicity, hydrophobicity, basicity, acidity, neutrality, or charge to the amino acids they replace. Alternatively, conservative substitutions may introduce another amino acid that is aromatic or aliphatic in place of an existing aromatic or aliphatic amino acid. Conservative amino acid changes are well known in the art.
[0073] The CsgG pore monomer may be modified to introduce one or more cysteines, one or more hydrophobic amino acids, one or more charged amino acids, one or more unnatural amino acids, one or more polar amino acids, or one or more photoreactive amino acids. Such introductions may be made in any number and combination. Introduction is preferably made by substitution or addition.
[0074] One or more amino acid residues of the amino acid sequence of SEQ ID NO: 3 may additionally be deleted from the polypeptide. Up to 1, 2, 3, 4, 5, 10, 20, or 30 or more residues may be deleted.
[0075] Variants may include fragments of SEQ ID NO: 3. Such fragments retain pore-forming activity. Fragments may be at least 50, at least 100, at least 150, at least 200, or at least 250 amino acids in length. Such fragments may be used to produce pores. Fragments preferably include the transmembrane beta-barrel region of SEQ ID NO: 3, i.e., residues 132-155 and 181-211, or variants thereof discussed above.
[0076] Alternatively, or in addition, one or more amino acids may be added to the above-described polypeptides. The extension may be provided at the amino or carboxy terminus of the amino acid sequence of SEQ ID NO: 3, or a polypeptide variant or fragment thereof. The extension may be very short, for example, 1 to 10 amino acids in length. Alternatively, the extension may be longer, for example, up to 50 or 100 amino acids. A carrier protein may be fused to the amino acid sequence according to the invention. Other fusion proteins are discussed in more detail below.
[0077] A variant of SEQ ID NO: 3 is a polypeptide having an amino acid sequence different from that of SEQ ID NO: 3 that retains the ability to form a pore. The variant typically contains the region of SEQ ID NO: 3 responsible for pore formation. The pore-forming ability of β-barrel-containing CsgG is provided by the β-strand in the transmembrane beta-barrel of each monomer. A variant of SEQ ID NO: 3 typically contains the regions of SEQ ID NO: 3 that form the β-strand, i.e., residues 132-155 and 181-211, or variants thereof discussed above. One or more modifications can be made to the regions of SEQ ID NO: 3 that form the β-strand, as long as the resulting variant retains the ability to form a pore.
[0078] One or more modifications in a CsgG pore monomer preferably improve the ability of a pore complex comprising the pore monomer to characterize an analyte. For example, modifications / mutations / substitutions are contemplated that alter the number, size, shape, location, or orientation of constrictions in a channel from a pore monomer conjugate of the invention. A CsgG pore monomer or variant of SEQ ID NO: 3 may have any of the specific modifications or substitutions disclosed in WO2016 / 034591, WO2017 / 149316, WO2017 / 149317, WO2017 / 149318, WO2018 / 211241, and WO2019 / 002893 (each of which is incorporated by reference in its entirety).
[0079] Preferred modifications or substitutions in SEQ ID NO: 3 are: (a) a substitution at position Y51, e.g., Y51I, Y51L, Y51A, Y51V, Y51T, Y51S, Y51Q or Y51N; (b) a substitution at position N55, e.g., N55I, N55L, N55A, N55V, N55T, N55S or N55Q; (c) a substitution at position F56, e.g., F56I, F56L, F56A, F56V, F56T, F56S, F56Q or F56N; (d) a substitution at position L90, e.g., L90N, L90D, L90E, L90R, or L90K; (e) a substitution at position N91, e.g., N91D, N91E, N91R, or N91K; (f) a substitution at position K94, e.g., K94R, K94F, K94Y, K94Q, K94W, K94L, K94S, or K94N; (g) substitution at position R192, e.g., R192Q, R192F, R192S R192D, or R192T; (i) substitutions at position C215, e.g., C215T, C215S, C215I, C215L, C215A, C215V, or C215G, including, but not limited to, one or more, e.g., two or more, three or more, four or more, five or more, six or more, seven or more, or all of the following:
[0080] Variants of SEQ ID NO: 3 may further include a deletion of one or more positions, for example a deletion of T104 to N109, a deletion of F193 to L199, or a deletion of F195 to L199.
[0081] Any number of the CsgG pore monomers in a pore or pore complex of the invention may be a variant of SEQ ID NO: 3, such as 6, 7, 8, 9 or 10. Preferably, all 6 to 10 monomers in a pore or pore complex are variants of SEQ ID NO: 3. The variants within a pore complex may be the same or different. The variant is preferably the same in each pore monomer conjugate within a pore complex of the invention.
[0082] CsgF peptide The term "CsgF peptide" preferably defines a CsgF peptide truncated from its C-terminus (i.e., an N-terminal fragment). The CsgF peptide may be a fragment of wild-type E. coli CsgF (SEQ ID NO: 5 or SEQ ID NO: 6) or a fragment of a wild-type homolog of E. coli CsgF, such as a peptide comprising any one of the amino acid sequences set forth in WO2019 / 002893 (incorporated herein by reference in its entirety). A CsgF homolog is a polypeptide having at least 40%, 50%, 60%, 70%, 80%, 90%, 95%, or 99% complete sequence identity to the wild-type E. coli CsgF set forth in SEQ ID NO: 6. A CsgF homolog is also a polypeptide containing the PFAM domain PF10614, characteristic of CsgF-like proteins. A list of currently known CsgF homologs and CsgF architectures can be found at http: / / pfam.xfam.org / / family / PF10614. Mature CsgF (set forth in SEQ ID NO: 6) can be divided into three main regions: the "CsgF constriction peptide" (FCP), the "neck" region, and the "head" region. The "head" region of the CsgF peptide is distinct from the constriction of the pore described herein. The "head" region of the CsgF peptide may also be referred to as the "C-terminal head domain." The structure of CsgF is discussed in detail in WO2019 / 002893, which is incorporated herein by reference in its entirety.
[0083] The CsgF peptides used in the pore monomer conjugates of the present invention preferably lack the C-terminal head, or lack both the C-terminal head and a portion of the CsgF neck domain (e.g., a truncated CsgF peptide may include only a portion of the CsgF neck domain), or are truncated CsgF peptides lacking both the C-terminal head and the CsgF neck domain. The CsgF peptide may lack a portion of the CsgF neck domain; for example, the CsgF peptide may include a portion of the neck domain from amino acid residue 36 at the N-terminus of the neck domain (see SEQ ID NO: 6) (e.g., residues 36-40, 36-41, 36-42, 36-43, 36-45, 36-46, residues 36-50, or 36-60 of SEQ ID NO: 6). The CsgF peptide preferably includes the CsgF binding region and the region that forms the constriction within the pore. The CsgG binding region typically comprises residues 1-11 and / or 29-32 of the CsgF protein (SEQ ID NO: 6 or a homologue from another species) and may include one or more modifications. The region that forms the constriction within the pore typically comprises residues 9-28 of the CsgF protein (SEQ ID NO: 6 or a homologue from another species) and may include one or more modifications. Residues 9-17 contain the conserved motif N9PXFGGXXX 17 and form a turn region. Residues 9 to 28 form an alpha-helix. X 17 (N17 in SEQ ID NO:6) forms the apex of the constriction region, which corresponds to the narrowest part of the CsgF constriction within the pore. The CsgF constriction region also makes stabilizing contacts with the CsgG beta-barrel, primarily at residues 98, 9, 10, 11, 12, 18, 21, 22, 29, and 30 of SEQ ID NO:6.
[0084] The CsgF peptide typically has a length of 28 to 50 amino acids, for example, 29 to 49, 30 to 45, or 32 to 40 amino acids. Preferably, the CsgF peptide contains 29 to 35 amino acids, or 29 to 45 amino acids. The CsgF peptide contains all or part of the FCP corresponding to residues 1 to 35 of SEQ ID NO: 6. When the CsgF peptide is shorter than the FCP, the truncation is preferably at the C-terminus.
[0085] The CsgF peptide may have a length of 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54 or 55 amino acids.
[0086] The CsgF peptide may comprise the amino acid sequence of SEQ ID NO: 6 from residue 1 to any one of residues 25-60, e.g., 27-50, e.g., 28-45 of SEQ ID NO: 6, or the corresponding residues from a homologue of SEQ ID NO: 6, or any variant thereof. More specifically, the CsgF peptide may comprise residues 1-29 of SEQ ID NO: 6, or a homologue or variant thereof.
[0087] The CsgF peptide is preferably a truncated CsgF peptide lacking one or more amino acids from the CsgF shown in SEQ ID NO: 6. The CsgF peptide is preferably a truncated CsgF peptide lacking a stretch of amino acids starting at any one of positions 15 to 35 and ending at position 119 of SEQ ID NO: 6. The CsgF peptide is preferably a truncated CsgF peptide lacking amino acids 15 to 119, 16 to 119, 17 to 119, 18 to 119, 19 to 119, 20 to 119, 21 to 119, 22 to 119, 23 to 119, 24 to 119, 25 to 119, 26 to 119, 27 to 119, 28 to 119, 29 to 119, 30 to 119, 31 to 119, 32 to 119, 33 to 119, 34 to 119, or 35 to 119 from SEQ ID NO:6.
[0088] Examples of such CsgF peptides include, consist essentially of, or consist of residues 1-34 of SEQ ID NO:6, residues 1-30 of SEQ ID NO:6, residues 1-45 of SEQ ID NO:6, or residues 1-35 of SEQ ID NO:6 and any homologs or variants thereof.
[0089] In the CsgF peptide, one or more residues may be modified. For example, the CsgF peptide may include modifications at positions corresponding to one or more of positions G1, T4, F5, R8, N9, N11, F12, N17, A20, N24, A26, Q27, and Q29 in SEQ ID NO:6.
[0090] The CsgF peptide may be modified to introduce one or more cysteines, one or more hydrophobic amino acids, one or more charged amino acids, one or more unnatural amino acids, one or more polar amino acids, or one or more photoreactive amino acids, for example, at positions corresponding to one or more of G1, T4, F5, R8, N9, N11, F12, A26, and Q29 in SEQ ID NO: 6. Such introductions may be made in any number and combination. Introduction is preferably made by substitution or addition.
[0091] For example, the CsgF peptide may contain a modification at a position corresponding to one or more of positions N15, N17, A20, N24, and A28 in SEQ ID NO: 6. The CsgF peptide may contain a modification at a position corresponding to D34 to stabilize the CsgG-CsgF complex. The CsgF peptide may include one or more of the following substituents: N15S / A / T / Q / G / L / V / I / F / Y / W / R / K / D / C / E, N17S / A / T / Q / G / L / V / I / F / Y / W / R / K / D / C / E, A20S / T / Q / N / G / L / V / I / F / Y / W / R / K / D / C / E, N24S / T / Q / A / G / L / V / I / F / Y / W / R / K / D / C / E, A28S / T / Q / N / G / L / V / I / F / Y / W / R / K / D / C / E, and D34F / Y / W / R / K / N / Q / C / E. The CsgF peptide may include, for example, one or more of the following substitutions: G1C, T4C, N17S, and D34Y or D34N.
[0092] CsgF peptides can be produced by cleaving a longer protein, such as full-length CsgF, using an enzyme. Cleavage at a specific site can be directed by modifying a longer protein, such as full-length CsgF, to include an enzyme cleavage site at an appropriate location. Examples of CsgF amino acid sequences modified to include such enzyme cleavage sites are shown in SEQ ID NOS: 56-67 in WO2019 / 002893 (incorporated herein by reference in its entirety). After cleavage, all or part of the added enzyme cleavage site may be present in the CsgF peptide that associates with CsgG to form a pore. Thus, the CsgF peptide may further include all or part of the enzyme cleavage site at its C-terminus.
[0093] Some examples of suitable CsgF peptides are shown in Table 3 of WO2019 / 002893, which is incorporated herein by reference in its entirety.
[0094] The CsgF peptide is preferably a variant of any of the above-described CsgF sequences, including SEQ ID NO: 6, that contains one or more modifications compared to the reference sequence. Over the entire length of the amino acid sequence of SEQ ID NO: 6, the variant is preferably at least 40% homologous to that sequence based on amino acid identity. More preferably, the variant may be at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, more preferably at least 95%, 97%, or 99% homologous to the amino acid sequence of SEQ ID NO: 6 over the entire sequence based on amino acid identity. Over the entire length of the amino acid sequence of SEQ ID NO: 6, the variant is preferably at least 40% homologous to that sequence. More preferably, a variant may be at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, more preferably at least 95%, 97%, or 99% identical to SEQ ID NO: 6 over the entire sequence. There may be at least 80%, e.g., at least 85%, 90%, or 95% amino acid identity ("hard homology") over a stretch of 100 or more, e.g., 125, 150, 175, or 200 or more contiguous amino acids. These homology / identity levels are similar for any of the other CsgF peptides described above.
[0095] Any number of the CsgF peptides in a pore or pore complex of the invention, such as 6, 7, 8, 9, or 10, may contain one or more substitutions compared to SEQ ID NO: 6. All 6 to 10 monomers in a pore or pore complex preferably contain one or more substitutions compared to SEQ ID NO: 6. The CsgF peptides in a pore complex may be the same or different. The CsgF peptide is preferably identical in each pore-monomer conjugate in a pore complex of the invention.
[0096] Stabilizing and other mutations In the pore complex of the invention, the interaction between the CsgF peptide and the CsgG pore may be stabilized by, for example, hydrophobic and / or electrostatic interactions, such as between one or more of the following pairs of positions: 1 and 153, 4 and 133, 5 and 136, 8 and 187, 8 and 203, 9 and 203, 11 and 142, 11 and 201, 12 and 149, 12 and 203, 26 and 191, and 29 and 144 in SEQ ID NO: 6 and SEQ ID NO: 3, respectively.
[0097] Residues of the CsgF peptide and / or CsgG pore monomer at one or more of the positions listed above may be modified to enhance the interaction between CsgG and CsgF in the pore complex. Although the CsgG:CsgF complex is highly stable, truncating CsgF reduces the stability of the CsgG:CsgF complex compared to a complex containing full-length CsgF. Therefore, to make the complex more stable, for example, a disulfide bond can be created between CsgG and CsgF after introducing a cysteine residue at the position identified herein. The pore complex can be prepared by any of the methods described above, and disulfide bond formation can be induced by using an oxidizing agent (e.g., copper-orthophenanthroline). Instead of cysteine interactions, other interactions (e.g., hydrophobic interactions, charge-charge interactions / electrostatic interactions) can also be used at these positions.
[0098] Unnatural amino acids can also be incorporated at these positions. Covalent attachment may be achieved via click chemistry. For example, unnatural amino acids bearing azides or alkynes, or bearing dibenzocyclooctyne (DBCO) and / or bicyclo[6.1.0]nonyne (BCN) groups, can be introduced at one or more of these positions.
[0099] Such stabilizing mutations can be combined with any other modifications to CsgG and / or CsgF, such as those disclosed herein.
[0100] To facilitate such interactions, one or more non-natural or photoreactive amino acids may be included / substituted in the CsgG pore monomer at one or more positions corresponding to one or more of positions 132, 133, 136, 138, 140, 142, 144, 145, 147, 149, 151, 153, 155, 183, 185, 187, 189, 191, 201, 203, 205, 207 and 209 of SEQ ID NO:3.
[0101] To facilitate such interactions, one or more non-naturally reactive or photoreactive amino acids may be included / substituted at one or more positions corresponding to one or more of positions 1, 4, 5, 8, 9, 11, 12, 26 or 29 of SEQ ID NO:6.
[0102] A preferred exemplary CsgF peptide contains the following mutations relative to SEQ ID NO: 6: N15X1 / N17X2 / A20X3 / N24X4 / A28X5 / D34X6, wherein X1 is N / S / A / T / Q / G / L / V / I / F / Y / W / R / K / D / C / E and X2 is N / S / A / T / Q / G / L / V / I / F / Y / W / R / K / D / C / E. X1 is A / S / T / Q / N / G / L / V / I / F / Y / W / R / K / D / C / E, X2 is N / S / T / Q / A / G / L / V / I / F / Y / W / R / K / D / C / E, X3 is A / S / T / Q / N / G / L / V / I / F / Y / W / R / K / D / C / E, X4 is N / S / T / Q / N / G / L / V / I / F / Y / W / R / K / D / C / E, and X6 is D / F / Y / W / R / K / N / Q / C / E. Mutations at N15, N17, A20, N24, and A28 are constriction mutations, and the mutation at position 34 affects and stabilizes the interaction of CsgF with the bottom of the CsgG pore monomer.
[0103] construct The present invention also provides constructs comprising two or more covalently attached pore monomer conjugates of the invention. The constructs may comprise two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, or ten or more pore monomer conjugates of the invention. The constructs may comprise at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten pore monomer conjugates of the invention. The two or more pore monomer conjugates may be the same or different. The two or more pore monomer conjugates may differ based on one or more of: (a) the sequence of the CsgG pore monomer, (b) the sequence of the CsgF peptide, (c) the linker, (d) the attachment position on the CsgG pore monomer, and (e) the attachment position on the CsgF peptide. Pore monomer conjugates are (a), (b), (c), (d), (e), (a) and (b), (a) and (c), (a) and (d), (a) and (e), (b) and (c), (b) and (d), (b) and (e), (c) and (d), (c) and (e), (d) and (e), (a), (b and (c), (a), (b and (d), (a), (b and (e), (a), (c and (d), (a), (c and (e) , (a), (d) and (e), (b), (c) and (d), (b), (c) and (e), (b), (d) and (e), (c), (d) and (e), (a), (b), (c) and (d), (a), (b), (c) and (e), (a), (b), (c), (d) and (e), (a), (b), (c), (d) and (e), (b), (c), (d) and (e), and (a), (b), (c), (d) and (e). Two or more pore monomer conjugates are preferably the same (i.e., identical).
[0104] The construct preferably comprises two pore monomer conjugates. The two or more pore monomer conjugates may be the same or different. The two or more pore monomer conjugates are preferably the same (i.e., identical).
[0105] The pore-monomer conjugates may be genetically fused, optionally via a linker, or chemically fused, for example, via a chemical crosslinker. Methods for covalently attaching the monomers are disclosed in WO2017 / 149316, WO2017 / 149317, and WO2017 / 149318, which are incorporated herein by reference in their entireties.
[0106] The linker is preferably an amino acid sequence and / or a chemical cross-linker. Suitable amino acid linkers, such as peptide linkers, are known in the art. The length, flexibility, and hydrophilicity of the amino acid or peptide linker are typically designed so that the CsgF peptide forms a constriction in the pore complex of the present invention. Preferred flexible peptide linkers are stretches of 2 to 20, e.g., 4, 6, 8, 10, or 16, serine and / or glycine amino acids. More preferred flexible linkers are (SG)1, (SG)2, (SG)3, (SG)4, (SG)5, (SG)8, (SG) 10 , (SG) 15 or (SG) 20 wherein S is serine and G is glycine. A preferred rigid linker is a stretch of 2 to 30, e.g., 4, 6, 8, 16 or 24, proline amino acids. A more preferred rigid linker is (P) 12 wherein P is proline.
[0107] Suitable chemical cross-linking agents are well known in the art, including, but not limited to, those containing the following functional groups: maleimide, active ester, succinimide, azide, alkyne (such as dibenzocyclooctynol (DIBO or DBCO), difluorocycloalkyne, and linear alkyne), phosphine (such as those used in traceless and non-traceless Staudinger ligation), haloacetyl (such as iodoacetamide), phosgene-type reagents, sulfonyl chloride reagents, isothiocyanate, acyl halide, hydrazine, disulfide, vinyl sulfone, aziridine, and photoreactive reagents (such as aryl azide, diaziridine).
[0108] The reaction between the amino acid and the functional group can be spontaneous, such as cysteine / maleimide, or can require an external reagent, such as Cu(I) to link an azide and a linear alkyne.
[0109] Linkers can include any molecule that spans the required distance. Linkers can vary in length from one carbon (phosgene-type linkers) to many angstroms. Examples of linker molecules include, but are not limited to, polyethylene glycol (PEG), polypeptides, polysaccharides, deoxyribonucleic acid (DNA), peptide nucleic acid (PNA), threose nucleic acid (TNA), glycerol nucleic acid (GNA), saturated and unsaturated hydrocarbons, and polyamides. These linkers can be inert or reactive; in particular, they can be chemically cleavable at defined positions or can themselves be modified with fluorophores or ligands. Linkers are preferably resistant to reducing agents such as dithiothreitol (DTT) after covalent attachment of the CsgF peptide to the CsgG pore monomer.
[0110] Preferred crosslinkers include 2,5-dioxopyrrolidin-1-yl 3-(pyridin-2-yldisulfanyl)propanoate, 2,5-dioxopyrrolidin-1-yl 4-(pyridin-2-yldisulfanyl)butanoate, and 2,5-dioxopyrrolidin-1-yl 8-(pyridin-2-yldisulfanyl)octananoate, di-maleimide PEG 1k, di-maleimide PEG 3.4k, di-maleimide PEG 5k, di-maleimide PEG 10k, bis(maleimido)ethane (BMOE), bis-maleimidohexane (BMH), 1,4-bis-maleimidobutane (BMB), 1,4 bis-maleimidyl-2,3-dihydroxybutane (BMDB), BM[PEO]2 (1,8-bis-maleimidodiethylene glycol), BM[PEO]3 (1,11-bis-maleimidotriethylene glycol), tris[2-maleimidoethyl]amine (TMEA), DTME dithiobismaleimidoethane, bis-maleimidoPEG3, bis-maleimidoPEG11, DBCO-maleimide, DBCO-PEG4-maleimide, DBCO-PEG4-NH2, DBCO-PEG4-NHS, DBCO-NHS, DBCO-PEG-DBCO 2.8kDa, DBCO-PEG-DBCO Examples of crosslinkers include 4.0 kDa, DBCO-15 atoms-DBCO, DBCO-26 atoms-DBCO, DBCO-35 atoms-DBCO, DBCO-PEG4-SS-PEG3-biotin, DBCO-SS-PEG3-biotin, DBCO-SS-PEG11-biotin, (succinimidyl 3-(2-pyridyldithio)propionate (SPDP)), and maleimide-PEG(2 kDa)-maleimide (alpha, omega-bis-maleimide poly(ethylene glycol)). The most preferred crosslinker is maleimide-propyl-SRDFWRS-(1,2-diaminoethane)-propyl-maleimide.
[0111] The linker is preferably dithiothreitol (DTT) resistant. Suitable linkers include, but are not limited to, iodoacetamide-based and maleimide-based linkers.
[0112] In particular, pore monomer conjugates may be connected using two or more linkers, each of which comprises a hybridizable region and a group capable of forming a covalent bond. The hybridizable regions within the linkers hybridize and link the CsgG pore monomer and the CsgF peptide. The linked CsgG pore monomer and CsgF peptide are then linked via the formation of a covalent bond between the groups. Any of the specific linkers disclosed in WO2010 / 086602 (incorporated herein by reference in its entirety) may be used in accordance with the present invention.
[0113] The linker may be labeled. Suitable labels include fluorescent molecules (such as Cy3 or AlexaFluor® 555), radioisotopes, e.g. 125 I, 35 S, 32 Labels include, but are not limited to, P, enzymes, antibodies, antigens, polynucleotides, and ligands such as biotin. Such labels allow the amount of linker to be quantified. The label can also be a cleavable purification tag such as biotin, or a specific sequence that is not present in the protein itself but appears in an identification method, such as a peptide released by trypsin digestion.
[0114] A preferred method of connecting pore monomer conjugates is via a cysteine bond, which can be mediated by a bifunctional chemical crosslinker or by an amino acid linker with a terminally presented cysteine residue.
[0115] Another preferred method of attachment is via 4-azidophenylalanine or Faz linkage. This may be mediated by a bifunctional chemical linker or by a polypeptide linker bearing a terminally presented 4-azidophenylalanine or Faz residue. Additional suitable linkers are discussed in more detail below.
[0116] The pore complex of the present invention The terms "pore complex" or "complex pore," as used interchangeably herein, refer to an oligomeric pore complex comprising at least one pore monomer conjugate of the present invention (e.g., including one or more pore monomer conjugates, e.g., two or more pore monomer conjugates, three or more pore monomer conjugates). The pore complex of the present invention has the properties of a biological pore, i.e., it has a typical protein structure and defines a channel. When the pore complex is provided in an environment with a membrane component, a membrane, a cell, or an insulating layer, the pore complex inserts into the membrane or insulating layer to form a "transmembrane pore complex."
[0117] The CsgG portion of a pore complex of the invention (i.e. the portion formed from at least one CsgG pore monomer in at least one conjugate of the invention) preferably has or comprises any of the CsgG pore structures and / or dimensions described above. The CsgG constriction within a pore complex of the invention preferably has or comprises any of the constriction diameters described above.
[0118] At least one CsgF peptide (within at least one pore monomer conjugate or construct) preferably forms a constriction in the pore complex. At least one CsgF peptide is preferably inserted into the lumen of the pore complex. The present invention relates to CsgG pores complexed with a CsgF peptide, which introduce an additional channel constriction into the pore complex, surprisingly increasing the current range and signal-to-noise ratio (SNR). The additional constriction introduced by complexation with the CsgF peptide expands the contact surface with passing analytes and can act as a second constriction for analyte detection and characterization. Pores comprising the pore monomer conjugates of the present invention can improve the characterization of analytes, such as polynucleotides, providing a more discriminatory and direct relationship between the current observed as the polynucleotide translocates through the pore. In particular, by having two stacked constrictions spaced a defined distance apart, the pore complex may facilitate the characterization of at least one homopolymer stretch, e.g., a polynucleotide containing several consecutive copies of the same nucleotide that would otherwise exceed the interaction length of a single CsgG constriction. Additionally, by having two stacked constrictions at a defined distance, small molecule analytes, including organic or inorganic drugs and pollutants, passing through the pore complex pass through the two constrictions sequentially. The chemistry of either constriction can be independently modified, each imparting unique interaction properties with the analyte and thus providing additional discrimination during analyte detection.
[0119] The CsgF constriction formed in the pore complex preferably has a diameter in the range of about 5 to about 20 Å, e.g., about 7 to about 18 Å, about 10 Å to about 15 Å, or about 11 to about 12 Å. The additional CsgF peptide constriction may be about 10 nm or less, e.g., about 5 nm or less, e.g., about 1, 2, 3, 4, 5, 6, 7, 8, 9 nm, from the constriction of the CsgG pore. The distance between the CsgF peptide and the CsgG pore monomer is also discussed above with respect to the pore monomer conjugates of the invention.
[0120] The pore complexes or transmembrane pore complexes of the present invention include pore complexes having two constrictions, i.e., two channel constrictions positioned in such a way that one constriction does not interfere with the precision of the other constriction. The pore complexes may contain any of the mutations, CsgG pore monomers, or CsgF peptides described in WO2016 / 034591, WO2017 / 149316, WO2017 / 149317, WO2019 / 002893, WO2017 / 149318, WO2018 / 211241, and WO2019 / 002893 (the entireties of which are incorporated herein by reference). The pore complexes or transmembrane pore complexes of the present invention include pore complexes having a single constriction. For example, the constriction may be removed from the CsgG pore monomer within the complexes of the present invention so that the pore complexes of the present invention contain only the single constriction provided by the CsgF peptide. The present invention provides a pore composite comprising at least one pore monomer conjugate of the present invention. The pore composite typically comprises at least 6, 7, 8, 9, or 10 pore monomer conjugates of the present invention. The pore composite preferably comprises 8 or 9 pore monomer conjugates of the present invention. The pore monomer conjugates are typically the same (i.e., identical).
[0121] The pore complex is preferably a homo-oligomer comprising 6 to 10, for example 6, 7, 8, 9 or 10, pore monomer conjugates of the invention. The pore monomer conjugates are typically identical. The pore complex preferably comprises 8 or 9 identical pore monomer conjugates of the invention. The pore monomer conjugates may be any one of those discussed above.
[0122] The present invention provides a pore complex comprising at least one construct of the present invention. The pore complex typically comprises at least one, two, three, four, or five constructs of the present invention. The pore complex comprises sufficient CsgG pore monomers to form a pore. For example, an octameric pore may comprise (a) four constructs each comprising two pore monomer conjugates, (b) two constructs each comprising four pore monomer conjugates, (c) one construct comprising two pore monomer conjugates and six pore monomer conjugates that do not form part of the construct, (d) three constructs comprising two pore monomer conjugates and two pore monomer conjugates that do not form part of the construct, and (e) combinations thereof. For example, a nonameric pore provides the same and additional possibilities. Other combinations of constructs and monomers can be envisioned by those skilled in the art. One or more constructs of the present invention can be used to form a pore complex for characterizing polynucleotides, such as sequencing. The pore complex preferably comprises four constructs of the present invention each comprising two pore monomer conjugates. The constructs are typically the same (ie, identical).
[0123] The pore complex is preferably a homo-oligomer comprising one to five, for example, one, two, three, four, or five, constructs of the invention. The constructs are typically the same (i.e., identical). The pore complex preferably comprises four identical constructs of the invention, each comprising two pore monomer conjugates. The constructs may be any of those discussed above.
[0124] The CsgG pore monomers within a CsgG pore are preferably all about the same length, or are the same length. The barrels of the CsgG pore monomers of the invention within a pore are preferably about the same length, or are the same length. Length may be measured in number of amino acids and / or in length units.
[0125] The pore complexes of the present invention may be isolated, substantially isolated, purified, or substantially purified. The pore complexes of the present invention are isolated or purified if they are free of any other components, such as lipids or other pores. The pore complexes are substantially isolated if they are mixed with a carrier or diluent that does not interfere with their intended use. For example, the pore complexes are substantially isolated or substantially purified if they are present in a form that contains less than 10%, less than 5%, less than 2%, or less than 1% of other components, such as block copolymers, lipids, or other pores. Alternatively, the pore complexes of the present invention may be present in a membrane. Suitable membranes are discussed below.
[0126] The pore complexes of the present invention may exist as individual or single pore complexes. Alternatively, the pore complexes of the present invention may exist in two or more pore complexes or homogeneous or heterogeneous populations of pores. Other formats comprising the pore complexes of the present invention are described in more detail below.
[0127] Multimeric pore complexes The present invention also provides pore multimers comprising two or more pores, at least one of which is a pore complex of the present invention. The multimer may comprise any number of pores, for example, 3, 4, 5, 6, 7, or 8 or more pores. Any number of the pores within the multimer, including all of them, may be pore complexes of the present invention.
[0128] The pore multimer may be a dual pore complex comprising a first pore complex of the present invention and a second pore or complex. The second pore or complex is typically derived from CsgG. The second pore complex may be a complex of the present invention. Both the first pore complex and the second pore complex are preferably pore complexes of the present invention. In the dual pore complex, the first pore complex may be attached to the second pore (complex) by hydrophobic interactions and / or one or more disulfide bonds. One or more, for example, 2, 3, 4, 5, 6, 8, 9, or even all, of the monomers in the first pore complex and / or the second pore (complex) may be modified to enhance such interactions. This may be achieved by any suitable method. A specific method for forming a dual pore from a CsgG-derived pore is described in WO2019 / 002893 (incorporated herein by reference in its entirety).
[0129] The pore multimers of the present invention may be isolated, substantially isolated, purified, or substantially purified, as such terms are defined above with respect to the pore complexes of the present invention.
[0130] Membrane embodiments The present invention also provides a pore complex of the present invention or a pore multimer of the present invention contained in a membrane. The present invention also provides a membrane comprising a pore complex of the present invention or a pore multimer of the present invention. These products are directly applicable for use in molecular sensing, such as analyte characterization and polynucleotide sequencing. Suitable membranes are described in more detail below.
[0131] Methods for producing modified proteins Methods for introducing or substituting unnatural amino acids into CsgG pore monomers and CsgF peptides are also well known in the art and are described in WO 2019 / 002893, which is incorporated herein by reference in its entirety. Proteins may also be modified to aid in their identification or purification, for example, by the addition of a streptavidin tag or a signal sequence to facilitate their secretion from cells in which the monomer does not naturally contain such a sequence. Proteins may also be produced using D-amino acids, or a mixture of L- and D-amino acids. This is conventional in the art for producing such proteins or peptides.
[0132] CsgG pore monomers, CsgF peptides, pore monomer conjugates, constructs, pore complexes, or pore multimers (i.e., any protein of the invention) may be chemically modified. Proteins may be chemically modified in any manner and at any site. Proteins are preferably chemically modified by the attachment of a molecule to one or more cysteines (cysteine ligation), one or more lysines, one or more unnatural amino acids, enzymatic modification of an epitope, or terminal modification. Suitable methods for performing such modifications are well known in the art. Proteins may also be chemically modified by the attachment of any molecule, such as a dye or fluorophore.
[0133] The protein may be chemically modified with a molecular adaptor that facilitates interaction between the pore containing the monomer and the target nucleotide or target polynucleotide sequence. Suitable adaptors, including cyclic molecules, cyclodextrins, hybridizable species, DNA binders or interchelators, peptides or peptide analogs, synthetic polymers, aromatic planar molecules, small positively charged molecules, or small molecules capable of hydrogen bonding, are described in WO2019 / 002893 (incorporated herein by reference in its entirety). The molecular adaptor may be attached using any of the methods and linkers described above.
[0134] The protein may be linked to a polynucleotide-binding protein. This forms a modular sequencing system that can be used in the sequencing method of the present invention. Polynucleotide-binding proteins are discussed below. The protein may be covalently linked to the monomer using any method known in the art. The monomer and protein may be chemically fused or genetically fused. Genetic fusion of a monomer to a polynucleotide-binding protein is discussed in WO2010 / 004265 (incorporated herein by reference in its entirety). The polynucleotide-binding protein may be linked via a cysteine bond using any of the methods described above.
[0135] The polynucleotide binding protein may be directly linked to the protein via one or more linkers.The molecule may be linked to the CsgG pore monomer using the hybridization linker described in WO2010 / 086602 (the entirety of which is incorporated herein by reference).Alternatively, a peptide linker may be used.Suitable peptide linkers are described above.
[0136] Any of the proteins can be modified to aid in their identification or purification, for example, by the addition of histidine residues (his tag), aspartic acid residues (asp tag), streptavidin tag, Flag tag, SUMO tag, GST tag, or MBP tag, or by the addition of a signal sequence to facilitate their secretion from cells in which the polypeptide does not naturally contain such a sequence. An alternative to introducing a genetic tag is to chemically react the tag onto a native or engineered position on the protein. One example of this would be to react a gel shift reagent with a cysteine engineered onto the outside of the protein. This has been shown as a method for isolating hemolysin hetero-oligomers (Chem Biol. 1997 July;4(7):497-505).
[0137] Any of the proteins may be labeled with a revealing label. The revealing label may be any suitable label that allows the protein to be detected. Suitable labels include, but are not limited to, fluorescent molecules, radioisotopes, such as 125I, 35S, enzymes, antibodies, antigens, polynucleotides, and ligands such as biotin.
[0138] Proteins may also contain other non-specific modifications as long as they do not interfere with the function of the protein. Many non-specific side chain modifications are known in the art and can be made to the side chains of protein(s). Such modifications include, for example, reductive alkylation of amino acids by reaction with an aldehyde followed by reduction with NaBH4, amidation with methylacetimidate, or acylation with acetic anhydride.
[0139] Any of the proteins can be produced using standard methods known in the art. Polynucleotide sequences encoding proteins can be obtained and replicated using standard methods in the art. Polynucleotide sequences encoding proteins can be expressed in bacterial host cells using standard techniques in the art. Proteins can be produced intracellularly by expressing the polypeptide in situ from a recombinant expression vector. The expression vector optionally carries an inducible promoter to control the expression of the polypeptide. These methods are described in Sambrook, J. and Russell, D. (2001) Molecular Cloning: A Laboratory Manual, 3rd Edition. Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY.
[0140] Proteins can be produced on a large scale from protein-producing organisms following purification by any protein liquid chromatography system, or following recombinant expression. Typical protein liquid chromatography systems include FPLC, AKTA systems, Bio-Cad systems, Bio-Rad BioLogic systems, and Gilson HPLC systems.
[0141] Method for producing porous monomer conjugates The invention provides a method of producing a pore monomer conjugate of the invention, the method comprising attaching, preferably covalently, a CsgF peptide to a CsgG pore monomer at two or more positions.
[0142] The method typically involves modifying a CsgF peptide at two or more positions to include two or more reactive groups that can be attached to two or more positions within the CsgG pore monomer. The two or more reactive groups may be the same. The two or more reactive groups may be different.
[0143] The method preferably comprises contacting a CsgF peptide and a CsgG pore monomer with two or more linkers. The components may be contacted with the two or more linkers in any order, such as first the CsgF peptide and then the CsgG pore monomer, first the CsgG pore monomer and then the CsgF peptide, or both components may be contacted with the linkers simultaneously. One or more linkers may be attached to the CsgF peptide, and one or more linkers may be attached to the CsgG pore monomer, before the two proteins are attached at two or more positions.
[0144] The two or more linkers are preferably first attached to the CsgF peptide or CsgG pore monomer, and then to the other components of the complex. The method preferably comprises covalently attaching the two or more linkers to the CsgF peptide, and then contacting the linkers and CsgF peptide with the CsgG pore monomer under conditions such that the CsgF peptide is attached or covalently attached to the CsgG pore monomer at two or more positions. Such conditions will be familiar to those skilled in the art and are illustrated in the Examples. The method is typically carried out in vitro, as defined below.
[0145] Any of the embodiments discussed above with respect to the microporous monomer conjugates of the present invention apply equally to these methods.
[0146] Method for producing pores The present invention also provides a method for producing a pore complex of the invention or a pore multimer of the invention.
[0147] The method may include expressing a pore complex in a host cell. In particular, the method may include expressing at least one pore monomer conjugate of the invention or a construct of the invention and sufficient pore monomers or constructs to form a pore complex or pore multimer in the host cell, and allowing the pore complex or pore multimer to form in the host cell. The sufficient pore monomers or constructs are preferably sufficient pore monomer conjugates of the invention or sufficient constructs of the invention. The number of CsgG pore monomers, pore monomer conjugates, or constructs required to form a pore complex of the invention or a pore multimer of the invention is discussed above. Suitable host cells and expression systems are known in the art and described in the Examples.
[0148] The method may include forming a pore complex in a non-cellular or in vitro environment. In particular, the method may include contacting at least one pore monomer conjugate of the present invention or a construct of the present invention with sufficient pore monomers or constructs in vitro and allowing the formation of a pore complex or pore multimer. The pore monomer conjugates or constructs may be individually generated by in vitro translation and transcription (IVTT) and then incubated with sufficient pore monomers or constructs. The sufficient pore monomers or constructs are preferably sufficient pore monomer conjugates of the present invention or sufficient constructs of the present invention. The number of CsgG pore monomers, pore monomer conjugates, or constructs required to form a pore complex of the present invention or a pore multimer of the present invention is discussed above. The method may be constructed in an "in vitro system," which refers to a system that includes at least the components and environment necessary to carry out the method, utilizes biomolecules, organisms, cells (or portions of cells) outside their normal naturally occurring environment, and allows for more detailed, convenient, or efficient analysis than can be performed with whole organisms. The in vitro system may also comprise a suitable buffer composition provided in a test tube to which the protein components for complex formation are added. Those skilled in the art are aware of options for providing the system.
[0149] To facilitate purification, some or all of the components of the pore complex or pore multimer may be tagged. Purification can also be performed when the components are not tagged. Methods known in the art (e.g., ion exchange, gel filtration, hydrophobic interaction column chromatography, etc.) can be used alone or in different combinations to purify the components of the pore.
[0150] The pore complexes or pore multimers can be generated prior to insertion into the membrane or after insertion of the components into the membrane.
[0151] Methods for producing the pores and complexes of the present invention and methods for tagging them are disclosed in WO2016 / 034591, WO2017 / 149316, WO2017 / 149317, as well as WO2017 / 149318, WO2018 / 211241, and WO2019 / 002893 (all of which are incorporated herein by reference in their entirety).
[0152] Methods for characterizing an analyte The present invention provides a method for determining the presence, absence, or one or more properties of a target analyte. The method comprises contacting the target analyte with a pore complex of the invention or a pore multimer of the invention such that the target analyte migrates relative to, e.g., to or through, the pore complex or pore multimer, and performing one or more measurements as the analyte migrates relative to the pore complex or pore multimer, thereby determining the presence, absence, or one or more properties of the analyte. The target analyte may also be referred to as a template analyte or an analyte of interest.
[0153] The pore complex of the present invention or the pore multimer of the present invention may be any of those discussed above.
[0154] The method is for determining the presence, absence, or one or more characteristics of a target analyte. The method may be for determining the presence, absence, or one or more characteristics of at least one analyte. The method may relate to determining the presence, absence, or one or more characteristics of two or more analytes. The method may include determining the presence, absence, or one or more characteristics of any number of analytes, for example, 2, 5, 10, 15, 20, 30, 40, 50, 100, or more analytes. Any number of characteristics of one or more analytes may be determined, for example, 1, 2, 3, 4, 5, 10, or more characteristics.
[0155] The binding of a molecule within the channel of a pore complex or pore multimer, or near any opening of the channel, affects the open channel ion flow through the pore complex or pore multimer, which is the essence of "molecular sensing." In a manner similar to nucleic acid sequencing applications, the fluctuation of open channel ion flow can be measured using a suitable measurement technique by changing the current (e.g., WO2000 / 28312 and D. Stoddart et al., Proc. Natl. Acad. Sci., 2010, 106, 7702-7 or WO2009 / 077734, all of which are incorporated herein by reference in their entirety). The degree of decrease in ion flow, measured by the decrease in current, is related to the size of the obstacle within or near the pore. Thus, the binding of a molecule of interest, also referred to as an "analyte," within or near the pore provides a detectable and measurable event, thereby forming the basis of a "biological sensor." Molecules suitable for nanopore sensing include nucleic acids, proteins, peptides, polysaccharides, and small molecules (herein referring to organic or inorganic compounds with low molecular weight (e.g., <900 Da or <500 Da)), such as pharmaceuticals, toxins, cytokines, and pollutants. Detecting the presence of biomolecules finds applications in personalized drug development, medicine, diagnostics, life science research, environmental monitoring, and the security and / or defense industries.
[0156] Pore complexes or pore multimers may function as molecular or biological sensors. Analyte molecules to be detected may bind either to the surface of the channel or within the lumen of the channel itself. The location of binding may be determined by the size of the molecule to be sensed.
[0157] The target analyte is preferably a metal ion, an inorganic salt, a polymer, an amino acid, a peptide, a polypeptide, a protein, a nucleotide, an oligonucleotide, a polynucleotide, a monosaccharide, an oligosaccharide, a polysaccharide, a dye, a bleaching agent, a pharmaceutical, a diagnostic agent, a recreational drug, an explosive, a toxic compound, or an environmental pollutant. The analyte may include two or more different molecules, such as a peptide and a polypeptide. The method may involve determining the presence, absence, or one or more characteristics of two or more analytes of the same type, for example, two or more proteins, two or more nucleotides, or two or more pharmaceuticals. Alternatively, the method may involve determining the presence, absence, or one or more characteristics of two or more different types of analytes, for example, one or more proteins, one or more nucleotides, and one or more pharmaceuticals.
[0158] The target analyte may be secreted from the cell. Alternatively, the target analyte may be an analyte that is present intracellularly and therefore must be extracted from the cell before the method may be performed.
[0159] Pore complexes or pore multimers may be modified via recombinant or chemical methods to increase the strength, location, or specificity of binding of the molecule to be sensed. Typical modifications include the addition of a specific binding moiety complementary to the structure of the molecule to be sensed. If the analyte molecule comprises a nucleic acid, this binding moiety may comprise a cyclodextrin or an oligonucleotide; in the case of a small molecule, it may be a known complementary binding region, including a single-chain variable fragment (scFv) region or an antigen recognition domain from a T-cell receptor (TCR), e.g., the antigen-binding portion of an antibody or non-antibody molecule; or in the case of a protein, it may be a known ligand of the target protein. In this way, the pore complex or pore multimer may be endowed with the ability to act as a molecular sensor for detecting the presence in a sample of suitable antigens (including epitopes), which may include receptors, cell surface antigens including markers of solid tumors or blood cancer cells (e.g., lymphoma or leukemia), viral antigens, bacterial antigens, protozoan antigens, allergens, allergy-related molecules, albumin (e.g., human, rodent, or bovine), fluorescent molecules (including fluorescein), blood group antigens, small molecules, drugs, enzymes, catalytic sites of enzymes or enzyme substrates, and transition-state analogs of enzyme substrates. As noted above, modifications can be achieved using known genetic engineering and recombinant DNA techniques. The location of any adaptations will depend on the properties of the molecule being sensed, such as its size, three-dimensional structure, and its biochemical properties. Selection of the adapted structure may utilize computational structural design. The determination and optimization of protein-protein or protein-small molecule interactions can be investigated using technologies such as BIAcore®, which uses surface plasmon resonance to detect molecular interactions (BIAcore, Inc., Piscataway, NJ, also see www.biacore.com).
[0160] The analyte is preferably an amino acid, peptide, polypeptide, or protein. The amino acid, peptide, polypeptide, or protein may be natural or non-natural. The polypeptide or protein may include synthetic or modified amino acids therein. Several different types of modifications to amino acids are known in the art. Suitable amino acids and their modifications are described above. It should be understood that the target analyte may be modified by any method available in the art.
[0161] The analyte is preferably a polynucleotide, such as a nucleic acid, which is defined as a macromolecule containing two or more nucleotides. Nucleic acids are particularly suitable for nanopore sequencing. Naturally occurring nucleic acid bases in DNA and RNA may be distinguished by their physical size. When a nucleic acid molecule or individual base passes through the nanopore channel, the size difference between the bases causes a directly correlated reduction in ion flow through the channel. The fluctuations in ion flow can be recorded. Suitable electrical measurement techniques for recording the fluctuations in ion flow are discussed above. With suitable calibration, the characteristic reduction in ion flow can be used to identify the specific nucleotide and related base passing through the channel in real time. In typical nanopore nucleic acid sequencing, the open channel ion flow decreases as individual nucleotides of a nucleotide sequence of interest pass sequentially through the nanopore channel due to partial blocking of the channel by the nucleotide. It is this reduction in ion flow that is measured using the suitable recording techniques described above. The reduction in ion flow can be calibrated to the reduction in ion flow measured for known nucleotides passing through the channel, which provides a means for determining which nucleotides pass through the channel, and thus, when performed sequentially, provides a method for determining the nucleotide sequence of the nucleic acid passing through the nanopore.In order to accurately determine individual nucleotides, it is typically necessary that the reduction in ion flow through the channel be directly correlated to the size of each nucleotide passing through the constriction.It will be understood that sequencing can be performed on intact nucleic acid polymers that are "threaded" through the pore, for example, through the action of associated polymerase.Alternatively, sequence can be determined by the passage of nucleotide triphosphate groups that are successively removed from the target nucleic acid adjacent to the pore (see, for example, WO2014 / 187924, the entire contents of which are incorporated herein by reference).
[0162] A polynucleotide or nucleic acid can contain any combination of any nucleotides. Nucleotides can be naturally occurring or artificial. One or more nucleotides in a polynucleotide can be oxidized or methylated. One or more nucleotides in a polynucleotide can be damaged. For example, a polynucleotide can contain pyrimidine dimers. Such dimers are typically associated with UV damage and are a major cause of skin melanoma. One or more nucleotides in a polynucleotide can be modified, for example, with a label or tag, suitable examples of which are known to those skilled in the art. A polynucleotide can contain one or more spacers. A nucleotide typically contains a nucleobase, a sugar, and at least one phosphate group. The nucleobase and sugar form a nucleoside. The nucleobase is typically heterocyclic. Nucleobases include, but are not limited to, purines and pyrimidines, more specifically, adenine (A), guanine (G), thymine (T), uracil (U), and cytosine (C). The sugar is typically a pentose sugar. Nucleotide sugars include, but are not limited to, ribose and deoxyribose. The sugar is preferably deoxyribose. Polynucleotides preferably contain the following nucleosides: deoxyadenosine (dA), deoxyuridine (dU) and / or thymidine (dT), deoxyguanosine (dG), and deoxycytidine (dC). Nucleotides are typically ribonucleotides or deoxyribonucleotides. Nucleotides typically contain monophosphate, diphosphate, or triphosphate. Nucleotides may contain more than three phosphates, for example, four or five phosphates. The phosphates may be attached to the 5' or 3' side of the nucleotide. Nucleotides in a polynucleotide can be attached to each other in any manner. Nucleotides are typically attached via their sugar and phosphate groups, similar to nucleic acids. Nucleotides may be connected via nucleobases, similar to pyrimidine dimers. Polynucleotides can be single-stranded or double-stranded. At least a portion of the polynucleotide is preferably double-stranded. The polynucleotide is most preferably ribonucleic acid (RNA) or deoxyribonucleic acid (DNA).In particular, the method alternatively uses a polynucleotide as an analyte, including determining one or more characteristics selected from (i) the length of the polynucleotide, (ii) the identity of the polynucleotide, (iii) the sequence of the polynucleotide, (iv) the secondary structure of the polynucleotide, and (v) whether the polynucleotide is modified.
[0163] Polynucleotides can be of any length (i). For example, polynucleotides can be at least 10, at least 50, at least 100, at least 150, at least 200, at least 250, at least 300, at least 400, or at least 500 nucleotides or nucleotide pairs in length. Polynucleotides can be 1,000 nucleotides or nucleotide pairs or more, 5,000 nucleotides or nucleotide pairs or more, or 100,000 nucleotides or nucleotide pairs or more in length. Any number of polynucleotides can be investigated. For example, the method can involve characterizing 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 50, 100, or more polynucleotides. When two or more polynucleotides are characterized, they can be different polynucleotides or two instances of the same polynucleotide. Polynucleotides can be naturally occurring or artificial. For example, the method can be used to verify the sequence of a manufactured oligonucleotide. The method is typically performed in vitro.
[0164] The nucleotide may have identity (ii) and may include, but is not limited to, adenosine monophosphate (AMP), guanosine monophosphate (GMP), thymidine monophosphate (TMP), uridine monophosphate (UMP), 5-methylcytidine monophosphate, 5-hydroxymethylcytidine monophosphate, cytidine monophosphate (CMP), cyclic adenosine monophosphate (cAMP), cyclic guanosine monophosphate (cGMP), deoxyadenosine monophosphate (dAMP), deoxyguanosine monophosphate (dGMP), deoxythymidine monophosphate (dTMP), deoxyuridine monophosphate (dUMP), deoxycytidine monophosphate (dCMP), and deoxymethylcytidine monophosphate. The nucleotide is preferably selected from AMP, TMP, GMP, CMP, UMP, dAMP, dTMP, dGMP, dCMP, and dUMP. A nucleotide may be abasic (i.e., lacking a nucleobase). A nucleotide may also lack a nucleobase and a sugar (i.e., a C3 spacer). The sequence (iii) of a nucleotide is determined by the consecutive identities of the following nucleotides, linked together in the 5' to 3' direction of the chain, throughout the polynucleotide strand:
[0165] The pore complexes and pore multimers of the present invention are particularly useful for analyzing homopolymers. For example, they may be used to determine the sequence of polynucleotides that contain two or more identical, e.g., at least 3, 4, 5, 6, 7, 8, 9, or 10 consecutive nucleotides. For example, they may be used to sequence polynucleotides that contain poly-A tracts, poly-T tracts, poly-G tracts, and / or poly-C tracts.
[0166] The CsgG pore constriction is generated by residues 51, 55, and 56 of SEQ ID NO:3. The CsgG constriction and its constriction variants are generally sharp. When DNA passes through the constriction, the interaction of approximately five bases of DNA with the pore constriction at any given time dominates the current signal. These sharper constrictions are excellent at reading mixed-sequence regions of DNA (when A, T, G, and C are mixed), but when there are homopolymeric regions within the DNA (e.g., poly-T, poly-G, poly-A, poly-C), the signal flattens and lacks information. Because five bases dominate the signal for CsgG and its constriction variants, it is difficult to distinguish photopolymers longer than five without using additional dwell time information. However, when DNA passes through the second constriction formed by the CsgF peptide, more DNA bases interact with the combined constriction, increasing the length of homopolymers that can be discriminated.
[0167] The movement of a polynucleotide relative to a pore, such as through a pore, is preferably controlled using a polynucleotide-binding protein. Suitable proteins are described in detail below. The present invention provides a method for determining the presence, absence, or one or more properties of a target polynucleotide, the method comprising: (i) contacting a target polynucleotide with a pore complex of the invention or a pore multimer of the invention and a polynucleotide binding protein such that the polynucleotide binding protein controls the movement of the target analyte relative to, e.g., through, the pore complex or pore multimer; (ii) performing one or more measurements as the polynucleotide translocates relative to, e.g., through, the pore complex or pore multimer, thereby determining the presence, absence, or one or more properties of the polynucleotide.
[0168] In either method, one or more properties of the target analyte are preferably measured by electrical and / or optical measurements. The electrical measurements may be amperometric, impedance, tunneling, or field effect transistor (FET) measurements. The method preferably comprises measuring the current flowing through the pore complex or pore multimer as the analyte moves toward, e.g., through, the pore.
[0169] The general conditions for carrying out the methods of the invention are discussed in more detail below with reference to the kits and systems of the invention.
[0170] Polynucleotides of the Invention The present invention also provides a polynucleotide encoding the pore monomer conjugate of the present invention or the construct of the present invention. The polynucleotide may be any of those discussed above. The present invention also provides an expression vector comprising the polynucleotide of the present invention. The present invention also provides a host cell comprising the polynucleotide of the present invention or the host cell of the present invention. Suitable vectors and host cells are known in the art.
[0171] kit The present invention also provides a kit for characterizing a target analyte. In one embodiment, the kit comprises (a) a pore complex of the present invention or a pore multimer of the present invention, and (b) membrane components. Suitable membranes and components are discussed below.
[0172] In another embodiment, the kit comprises (a) a pore complex or a pore multimer of the present invention and (b) a polynucleotide-binding protein. The kit preferably further comprises a membrane component. The kit may comprise any type of membrane, such as an amphiphilic layer or a triblock copolymer membrane component. Preferred polynucleotide-binding proteins are polymerases, exonucleases, helicases, and topoisomerases, such as gyrases. Suitable enzymes include, but are not limited to, exonuclease I from Escherichia coli, exonuclease III from Escherichia coli, RecJ and bacteriophage lambda exonuclease from Thermus thermophilus, TatD exonuclease, and variants thereof. Three subunits containing the RecJ sequence from Thermus thermophilus or a variant thereof interact to form a trimeric exonuclease. The polymerase may be PyroPhage® 3173 DNA polymerase (commercially available from Lucigen® Corporation), SD polymerase (commercially available from Bioron®), or a variant thereof. The enzyme may be Phi29 DNA polymerase or a variant thereof. The topoisomerase is preferably a member of either of the EC groups 5.99.1.2 and 5.99.1.3.
[0173] The enzyme is most preferably derived from a helicase such as Hel308 Mbu, Hel308 Csy, Hel308 Tga, Hel308 Mhu, TraI Eco, XPD Mbu, or a variant thereof. Any helicase can be used in the present invention. The helicase can be or be derived from Hel308 helicase, RecD helicase, such as TraI helicase or TrwC helicase, XPD helicase, or Dda helicase. The helicase can be any of the helicases, modified helicases, or helicase constructs disclosed in WO2013 / 057495, WO2013 / 098562, WO2013098561, WO2014 / 013260, WO2014 / 013259, WO2014 / 013262, and WO2015 / 055981. all of which are incorporated by reference in their entirety.
[0174] The kit may further include one or more anchors, such as cholesterol, for coupling the target analyte to the membrane. The kit may further include one or more polynucleotide adaptors capable of binding to the target polynucleotide to facilitate characterization of the polynucleotide. The anchor, such as cholesterol, is preferably attached to the polynucleotide adaptor.
[0175] The kit may further include one or more other reagents or equipment that enable any of the above embodiments to be carried out. Such reagents or equipment include one or more of the following: suitable buffer(s) (aqueous solution), means for obtaining a sample from a subject (such as a container or an apparatus including a needle), means for amplifying and / or expressing a polynucleotide, or voltage or patch clamp apparatus. The reagents may be present in the kit in a dry state so that a fluid sample resuspends the reagents. The kit may also optionally include instructions for use to enable the kit to be used in the methods of the invention, or details regarding organisms for which the methods may be used. Finally, the kit may also include additional components useful for characterizing the analyte.
[0176] Device The present invention also provides a device for characterizing a target analyte in a sample, the device comprising: (a) a plurality of pore complexes of the present invention or a plurality of pore multimers of the present invention; and (b) a plurality of polynucleotide binding proteins. The plurality of pore complexes or plurality of pore multimers may be any of those discussed above.
[0177] The present invention also provides a device comprising a pore complex of the present invention or a pore multimer of the present invention inserted into an in vitro membrane.
[0178] The present invention also provides a device produced by a method, the method comprising: (i) obtaining a pore complex of the present invention or a pore multimer of the present invention; and (ii) contacting the pore complex or pore multimer with an in vitro membrane such that the pore complex or pore multimer is inserted into the in vitro membrane.
[0179] Any of the specific embodiments discussed above are equally applicable to the apparatus of the present invention.
[0180] array The present invention also provides an array comprising a plurality of membranes of the present invention. Any of the embodiments discussed above with respect to membranes of the present invention apply equally to arrays of the present invention. The arrays can be configured to perform any of the methods described below.
[0181] In a preferred embodiment, each membrane in the array contains one pore complex or pore multimer. Depending on the manner in which the array is formed, for example, the array may contain one or more membranes that contain no pore complexes or pore multimers and / or one or more membranes that contain two or more pore complexes or pore multimers. The array may contain from about 2 to about 1000 membranes, e.g., from about 10 to about 800, from about 20 to about 600, or from about 30 to about 500 membranes.
[0182] system The invention provides a system comprising: (a) a membrane of the invention, or an array of the invention; (b) means for applying a potential between the membrane(s); and (c) means for detecting an electrical or optical signal between the membrane(s).
[0183] The pores and membranes can be any of those described above and below.
[0184] In one embodiment, the system further comprises a first chamber and a second chamber, the first and second chambers being separated by a membrane(s). When used to characterize a target analyte, the system may further comprise the target analyte, the target analyte being temporarily located within the continuous channel, with one end of the target analyte located in the first chamber and one end of the target analyte located in the second chamber. The target analyte is preferably a target polypeptide or a target polynucleotide.
[0185] In one embodiment, the system further includes a conductive solution in contact with the pore(s), electrodes providing a potential across the membrane(s), and a measurement system for measuring the current through the pore(s). The voltage applied to the membrane and pore is preferably between +5 V and −5 V, e.g., −600 mV to +600 mV or −400 mV to +400 mV. The voltage used is preferably in the range of 100 mV to 240 mV, more preferably in the range of 120 mV to 220 mV. By using an increased applied potential, it is possible to increase the discrimination between different amino acids or nucleotides by the pore. Any suitable conductive solution can be used. For example, the solution can include a charge carrier such as a metal salt, e.g., an alkali metal salt, a halide salt, e.g., a chloride salt, e.g., an alkali metal chloride salt. The charge carrier may include an ionic liquid or an organic salt, such as tetramethylammonium chloride, trimethylphenylammonium chloride, phenyltrimethylammonium chloride, or 1-ethyl-3-methylimidazolium chloride. In an exemplary system, the salt is present in an aqueous solution within the chamber. Potassium chloride (KCl), sodium chloride (NaCl), cesium chloride (CsCl), or a mixture of potassium ferrocyanide and potassium ferricyanide are typically used. KCl, NaCl, and a mixture of potassium ferrocyanide and potassium ferricyanide are preferred. The charge carriers may be asymmetric across the membrane. For example, the type and / or concentration of the charge carriers may be different on each side of the membrane, e.g., within each chamber.
[0186] The salt concentration can be saturating. The salt concentration can be 3 M or less, typically 0.1 to 2.5 M, 0.3 to 1.9 M, 0.5 to 1.8 M, 0.7 to 1.7 M, 0.9 to 1.6 M, or 1 M to 1.4 M. The salt concentration is preferably 150 mM to 1 M. The method is preferably carried out using a salt concentration of at least 0.3 M, e.g., at least 0.4 M, at least 0.5 M, at least 0.6 M, at least 0.8 M, at least 1.0 M, at least 1.5 M, at least 2.0 M, at least 2.5 M, or at least 3.0 M. A high salt concentration provides a high signal-to-noise ratio, allowing identification of currents indicative of the presence of amino acids or nucleotides against a background of normal current fluctuations.
[0187] A buffer solution may be present in the conductive solution. Typically, the buffer solution is a phosphate buffer. Other suitable buffer solutions are HEPES and Tris-HCl buffers. The pH of the conductive solution may be 4.0 to 12.0, 4.5 to 10.0, 5.0 to 9.0, 5.5 to 8.8, 6.0 to 8.7, 7.0 to 8.8, or 7.5 to 8.5. The pH used is preferably about 7.5.
[0188] The system may be included in a device. The device may be any conventional device for analyte analysis, such as an array or a chip. The device is preferably configured to perform the disclosed method. For example, the device may include a chamber containing an aqueous solution and a barrier separating the chamber into two sections. The barrier typically has an opening in which a membrane(s) containing a pore(s) is / are formed. Alternatively, the barrier forms a membrane in which the pores reside.
[0189] The device may also include an electrical circuit capable of applying an electrical potential and measuring the electrical signal between the membrane and the pore.
[0190] The device may be any of those described in WO2008 / 102120, WO2009 / 077734, WO2010 / 122293, WO2011 / 067559, or WO00 / 28312 (all of which are incorporated herein by reference in their entirety).
[0191] film Any suitable membrane can be used in the system. The membrane is preferably an amphiphilic layer. An amphiphilic layer is a layer formed from amphiphilic molecules, such as phospholipids, that have both hydrophilic and lipophilic properties. The amphiphilic molecules can be synthetic or naturally occurring. Non-naturally occurring amphiphiles and amphiphiles that form monolayers are known in the art, including, for example, block copolymers (Gonzalez-Perez et al., Langmuir, 2009, 25, 10447-10450). Block copolymers are polymeric materials in which two or more monomer subunits are polymerized together to create a single polymer chain. Block copolymers typically have properties contributed by each monomer subunit. However, block copolymers can have unique properties not possessed by polymers formed from individual subunits. Block copolymers can be engineered so that, in aqueous media, one of the monomer subunits is hydrophobic (i.e., lipophilic) and the other subunit(s) is / are hydrophilic. In this case, the block copolymer may have amphiphilic properties and form structures that mimic biological membranes. The block copolymer may be diblock (consisting of two monomer subunits), but may also be constructed from more than two monomer subunits to form more complex arrangements that behave as amphiphiles. The copolymer may be a triblock, tetrablock, or pentablock copolymer. The membrane is preferably a triblock copolymer membrane.
[0192] The membrane may comprise one of the membranes disclosed in International Application No. 2014 / 064443 or No. 2014 / 064444.
[0193] The amphiphilic molecules may be chemically modified or functionalized to facilitate coupling of polynucleotides. The amphiphilic layer may be a monolayer or a bilayer. The amphiphilic layer is typically planar. The amphiphilic layer may be curved. The amphiphilic layer may be supported.
[0194] Amphiphilic membranes are typically mobile in nature, with a molecular weight of approximately 10 -8 cms -1 The membrane essentially acts as a two-dimensional fluid with a lipid diffusion rate of 0.05 s, which means that the pore and the coupled polynucleotide can typically move within the amphiphilic membrane.
[0195] The membrane may be a lipid bilayer. Lipid bilayers are models of cell membranes and serve as an excellent platform for a wide range of experimental studies. For example, lipid bilayers can be used for in vitro investigation of membrane proteins by single-channel recording. Alternatively, lipid bilayers can be used as biosensors to detect the presence of a wide range of substances. The lipid bilayer may be any lipid bilayer. Suitable lipid bilayers include, but are not limited to, planar lipid bilayers, supported bilayers, or liposomes. The lipid bilayer is preferably a planar lipid bilayer. Suitable lipid bilayers are disclosed in WO2008 / 102121, WO2009 / 077734, and WO2006 / 100484 (all of which are incorporated herein by reference in their entirety).
[0196] The membrane may include a solid-state layer. The solid-state layer can be formed from both organic and inorganic materials, including, but not limited to, microelectronic materials, insulating materials such as Si3N4, Al2O3, and SiO2, organic and inorganic polymers such as polyamides, plastics such as Teflon®, or elastomers such as two-component addition-cured silicone rubber, and glass. The solid-state layer may also be formed from graphene. Suitable graphene layers are disclosed in WO2009 / 035647, the entire contents of which are incorporated herein by reference. When the membrane includes a solid-state layer, the pores are typically present in an amphiphilic membrane or layer contained within the solid-state layer, for example, within holes, wells, gaps, channels, trenches, or slits within the solid-state layer. Those skilled in the art will be able to prepare suitable solid-state / amphiphilic hybrid systems. Suitable systems are disclosed in WO2009 / 020682 and WO2012 / 005857, both of which are incorporated herein by reference in their entirety. Any of the amphiphilic membranes or layers discussed above may be used.
[0197] The method is typically carried out using (i) an artificial amphiphilic layer containing a pore, (ii) an isolated naturally occurring lipid bilayer containing a pore, or (iii) a cell with a pore inserted therein. The method is typically carried out using an artificial amphiphilic layer such as a diblock copolymer layer or a triblock copolymer layer. In addition to the pore, the layer may contain other transmembrane and / or intramembrane proteins, as well as other molecules. Suitable equipment and conditions are discussed below. The method of the present invention is typically carried out in vitro.
[0198] Sequence Listing SEQ ID NO: 1 (>P0AEA2, coding sequence of WT CsgG from E. coli K12) ATGCAGCGCTTATTTCTTTTGGTTGCCGTCATGTTACTGAGCGGATGCTTAACCGCCCCGCCTAAAGAAGCCGCCAGACCGACATTAATGCCTCGTGCTCAGAGCTACAAAGATTTGACCCATCTGCCAGCGCCGACGGGTAAAATCTTTGTTTCGGTATACAACATTCAGGACGAAACCGGGCAATTTAAACCCTACCCGGCAAGTAACTTCTCCACTGCTGTTCCGCAAAGCGCCACGGCAATGCTGGTCACGGCACTGAAAGATTCTCGCTGGTTTATACCGCTGGAGCGCCAGGGCTTACAAAACCTGCTTAACGAGCGCAAGATTATTCGTGCGGCACAAGAAAACGGCACGGTTGCCATTAATAACCGAATCCCGCTGCAATCTTTAACGGCGGCAAATATCATGGTTGAAGGTTCGATTATCGGTTATGAAAGCAACGTCAAATCTGGCGGGGTTGGGGCAAGATATTTTGGCATCGGTGCCGACACGCAATACCAGCTCGATCAGATTGCCGTGAACCTGCGCGTCGTCAATGTGAGTACCGGCGAGATCCTTTCTTCGGTGAACACCAGTAAGACGATACTTTCCTATGAAGTTCAGGCCGGGGTTTTCCGCTTTATTGACTACCAGCGCTTGCTTGAAGGGGAAGTGGGTTACACCTCGAACGAACCTGTTATGCTGTGCCTGATGTCGGCTATCGAAACAGGGGTCATTTTCCTGATTAATGATGGTATCGACCGTGGTCTGTGGGATTTGCAAAATAAAGCAGAACGGCAGAATGACATTCTGGTGAAATACCGCCATATGTCGGTTCCACCGGAATCCTGA SEQ ID NO: 2 (>P0AEA2(1:277), WT Pro-CsgG from Escherichia coli K12) MQRLFLVAVMLLSGCLTAPPKEAARPTLMPRAQSYKDLTHLPAPTGKIFVSVYNIQDETGQFKPYPASNFSTAVPQSATAMLVTALKDSRWFIPLERQGLQNLLNERKIIRAAQENGTVAINNRIPLQSLTAANIMV EGSIIGYESNVKSGGVGARYFGIGADTQYQLDQIAVNLRVVNVSTGEILSSVNTSKTILSYEVQAGVFRFIDYQRLLEGEVGYTSNEPVMLCLMSAIETGVIFLINDGIDRGLWDLQNKAERQNDILVKYRHMSVPPES SEQ ID NO: 3 (>P0AEA2 (16:277), mature CsgG from E. coli K12) CLTAPPKEAARPTLMPRAQSYKDLTHLPAPTGKIFVSVYNIQDETGQFKPYPASNFSTAVPQSATAMLVTALKDSRWFIPLERQGLQNLLNERKIIRAAQENGTVAINNRIPLQSLTAANIMVEGSIIGYE SNVKSGGVGARYFGIGADTQYQLDQIAVNLRVVNVSTGEILSSVNTSKTILSYEVQAGVFRFIDYQRLLEGEVGYTSNEPVMLCLMSAIETGVIFLINDGIDRGLWDLQNKAERQNDILVKYRHMSVPPES SEQ ID NO: 4 (>P0AE98, coding sequence of WT CsgF from E. coli K12) ATGCGTGTCAAACATGCAGTAGTTCTACTCATGCTTATTTCGCCATTAAGTTGGGCTGGAACCATGACTTTCCAGTTCCGTAATCCAAACTTTGGTGGTAACCCAAATAATGGCGCTTTTTTAAATAGCGCTCAGGCCCAAAACTCTTATAAAAGATCCGAGCTATAACGATGACTTTGGTATTGAAACACCCTCAGCGTTAGAT AACTTTACTCAGGCCATCCAGTCACAAATTTTAGGTGGCTACTGTCGAATATTAATACCGGTAAACCGGGCCGCATGGTGACCAACGATTATATTGTCGATATTGCCAACCGCGATGGTCAATTGCAGTTGAACGTGACAGATCGTAAAACCGGACAAACCTCGACCATCCAGGTTTCGGGTTTACAAAATAACTCAACCGATTTT SEQ ID NO: 5 (>P0AE98 (1:138), WT Pro-CsgF from E. coli K12) MRVKHAVVLLMLISPLSWAGTMTFQFRNPNFGGNPNNNGAFLLNSAQAQNSYKDPSYNDDFGIETPSALDNFTQAIQSQILGGLLSNINTGKPGRMVTNDYIVDIANRDGQLQLNVTDRKTGQTSTIQVSGLQNNSTDF SEQ ID NO: 6 (>P0AE98(20:138), WT mature CsgF from E. coli K12) GTMTFQFRNPNFGGNPNGAFLLNSAQAQNSYKDPSYNDDFGIETPSALDNFTQAIQSQILGGLLSNINTGKPGRMVTNDYIVDIANRDGQLQLNVTDRKTGQTSTIQVSGLQNNSTDF
[0199] The following examples illustrate the present invention. While specific embodiments, specific configurations, and materials and / or molecules have been discussed herein for engineered cells and methods according to the present invention, it should be understood that various changes or modifications in form and detail can be made without departing from the scope and spirit of the present invention. The following examples are provided to better illustrate specific embodiments and should not be construed as limiting the present application, which is limited only by the scope of the claims. [Example]
[0200] Detailed methods for making and testing mutant CsgG pores and mutant CsgG / CsgF complexes are described in WO2016 / 034591, WO2017 / 149316, WO2017 / 149317, WO2017 / 149318, WO2018 / 211241, and WO2019 / 002893 (all of which are incorporated by reference in their entirety).
[0201] Escherichia coli CsgG pore generation A recombinant expression vector encoding a CsgG variant nanopore with a C-terminal Strep affinity tag and ampicillin resistance gene is transformed into chemically competent E. coli cells. Cells are plated on LB agar plates containing the appropriate antibiotic for selection. A single colony from the agar plate is inoculated into LB medium with antibiotic and grown overnight. The culture is diluted with autoinduction medium and the required antibiotic and incubated at 18°C for 68 hours. Cells are harvested by centrifugation before lysis and extraction in 1x Bugbuster extraction reagent (Merck 70921) and 0.1% DDM. The pores are purified from the supernatant using affinity chromatography, heat treatment, and then size exclusion chromatography to select for oligomeric nanopores as judged by SDS-PAGE.
[0202] CsgG / CsgF complex formation protocol CsgG-CsgF complexes are prepared from purified nanopores as described above and chemically synthesized CsgF peptides with or without one or two linkers that can attach to CsgG. The nanopores are buffer-exchanged into a pH 7.0 buffer lacking the reducing agent and incubated at 25°C for 1 hour in an 8-fold molar excess of peptide over CsgG monomer. The reaction is stopped by heating to 60°C for 15 minutes, followed by centrifugation to remove any precipitates and the addition of DTT to 5 mM to prevent any further reaction.
[0203] SDS_PAGE analysis - with heating Add 300 ng of the complex and the CsgG-only pore control to individual 0.5 mL ProteinLoBind Eppendorf tubes (Fisher, 10316752) and bring the volume to 10 μL with reaction buffer. This is brought to a final volume of 20 μL by adding 10 μL of 2x Laemmli buffer. Load the entire sample onto a 4-20% TGX gel (BioRad, 5671093) run in 1x TGS buffer (Sigma, T7777). Run this at 300 V for 21 minutes. To image the gel, use Spyro Ruby (Merk, S4942) stain according to the manufacturer's instructions. Next, image it on a GE Typhoon gel imager using the 450 nm laser.
[0204] SDS-PAGE gel analysis of a CsgG-only pore control and a CsgG / CsgF complex when broken down into their constituent monomeric components upon boiling in the presence of DTT. Lanes in which CsgG is attached to CsgF at one or two positions show a binding shift compared to the CsgG-only control, indicating covalent binding between CsgG and CsgF.
[0205] Irregular curves in DNA (i.e., DNA translocation current traces) Electrical measurements were taken from CsgG-only, single-attached CsgG / CsgF complexes, and double-attached CsgG / CsgF complexes inserted into a MinION flow cell. After inserting the single nanopores into the block copolymer membrane, 1 mL of buffer containing 25 mM potassium phosphate, 150 mM potassium ferrocyanide(II), and 150 mM potassium ferricyanide(III), pH 8.0, was flowed through the system to remove any excess nanopores.
[0206] The Y-adapter is prepared by annealing DNA oligonucleotides as previously described (WO2016 / 034591, incorporated herein in its entirety). The DNA motor is loaded and closed onto the adapter. The subsequent material is purified by HPLC. The Y-adapter contains a 30 C3 leader section to facilitate capture by the nanopore and a side arm for tethering to the membrane.
[0207] The analyte used to assess DNA irregular curvature is a 3.6 kilobase section of DNA from the 3' end of the lambda genome. Analyte preparation, ligation of the analyte to a Y-adapter, SPRI-bead cleanup of the ligated analyte, and application to a minION flow cell are performed using the Oxford Nanopore Technologies Q-SQK-LSK109 protocol.
[0208] Electrical measurements are acquired using a minION Mk1b from Oxford Nanopore Technologies. A standard sequencing script is run at -180 mV for 2-6 h, with static flicking every 5 min to remove elongation nanopore blocks. Raw data is collected into bulk FAST5 files using MinKNOW software (Oxford Nanopore Technologies). A minimum of 150 pores per flow cell are tested per pore type.
[0209] In the absence of the CsgF peptide, the majority of pores are CsgG-only pores. Note that a small number of pores are misclassified as CsgG / CsgF complexes. However, when the pore complex contains CsgF functionalized with a single linker or two linkers that can attach to CsgG, a high proportion of CsgG / CsgF complexes are observed. The thresholds used to classify the type of inserted pore are: CsgG-only pores = pores with an open pore current between 160 pA and 200 pA; CsgG / CsgF complexes = pores with an open pore current between 70 pA and 140 pA. Both classifications also have an open pore noise <18 pA.
[0210] Current traces show ionic current (pA) versus time (s) as single-stranded DNA translocates through CsgG attached to CsgF at one or two positions. Each individual graph corresponds to a single pore inserted into the minION flow cell. The open pore current observed for the CsgG / CsgF complex is approximately 100 pA at an applied voltage of -180 mV. All other channels shown are either CsgG-only channels or empty / blocked channels.
[0211] Box plots show signal metrics for CsgG-based pores. SNR is the signal-to-noise ratio, which is the range of ionic current divided by noise when single-stranded DNA is translocating through the pore. In the presence of a single attachment of CsgF to CsgG, the SNR and / or range are increased. The SNR and / or range of the double-attached CsgG / CsgF complex are also increased compared to the single-attached complex.
Claims
1. A pore monomer conjugate comprising a CsgG pore monomer attached to a CsgF peptide, wherein the CsgF peptide is attached to the CsgG pore monomer at two or more positions.
2. The two or more positions within the CsgG pore monomer are (a) residues 47-54, 57, 59, 60, 130-134, 136, 137, 138, 140, 142-145, 147, 149, 151, 153, 155, 181, 183, 185, 187, 189, 191, 193, 195-199, 201, 203, 205, 207, 209 and 211-212 within the CsgG pore monomer; (b) is selected from residues corresponding to positions 47-54, 57, 59, 60, 130-134, 136, 137, 138, 140, 142-145, 147, 149, 151, 153, 155, 181, 183, 185, 187, 189, 191, 193, 195-199, 201, 203, 205, 207, 209, and 211-212 in SEQ ID NO:
3.
3. 3. The pore monomer conjugate of claim 1 or 2, wherein the two or more positions within the CsgF peptide are selected from the N-terminus and residues 1 to 35 within the CsgF peptide, or from the N-terminus and residues corresponding to positions 1 to 35 within SEQ ID NO:
6.
4. 4. The pore monomer conjugate of claim 1, wherein one of the two or more attachments comprises: (a) the N-terminus of the CsgF peptide attached to a cysteine residue in the CsgG pore monomer corresponding to position 153 of SEQ ID NO:3; (b) the position in the CsgF peptide corresponding to position 4 of SEQ ID NO:6 attached to a cysteine residue in the CsgG pore monomer corresponding to position 133 of SEQ ID NO:3; or (c) the position in the CsgF peptide corresponding to position 4 of SEQ ID NO:6 attached to a cysteine residue in the CsgG pore monomer corresponding to position 153 of SEQ ID NO:
3.
5. 5. The pore monomer conjugate of claim 1, wherein one of the two or more attachments comprises a position in the CsgF peptide corresponding to any one of positions 30, 31, 32 and 33 in SEQ ID NO: 6 attached to a position in the CsgG pore monomer corresponding to any one of positions 193, 195, 196 and 197 in SEQ ID NO:
3.
6. 6. The pore monomer conjugate of any one of claims 1 to 5, wherein the attachment at two or more positions comprises one or more reactive groups that (a) react with lysine, cysteine, tyrosine, serine, threonine, proline, tryptophan, arginine, histidine, methionine, or phenylalanine in the CsgG pore monomer, (b) react with any amino acid in the CsgG pore monomer, and / or (c) undergo click chemistry.
7. 7. The pore monomer conjugate of claim 6, wherein the lysine, cysteine, tyrosine, serine, threonine, proline, tryptophan, arginine, histidine, methionine, or phenylalanine is native to the CsgG pore monomer or is optionally introduced into the CsgG pore monomer by substitution or addition.
8. 8. The pore monomer conjugate of claim 1, wherein the attachment at two or more locations comprises two or more different reactive groups.
9. 9. The pore monomer conjugate of claim 1, wherein the CsgF peptide is attached to the CsgG pore monomer using two or more linkers.
10. 10. The pore monomer conjugate of claim 1, wherein the CsgF peptide is covalently attached to the CsgG pore monomer at two or more positions.
11. 11. The pore monomer conjugate of claim 1, wherein the CsgG pore monomer is a variant of SEQ ID NO: 3 and / or the CsgF peptide is a variant of SEQ ID NO:
6.
12. A construct comprising two or more covalently attached pore monomer conjugates according to any one of claims 1 to 11.
13. The construct of claim 12 , wherein the pore monomer conjugates are genetically fused and / or attached via a linker.
14. A pore complex comprising at least one pore monomer conjugate described in any one of claims 1 to 11 or at least one construct described in claim 12 or 13, wherein the CsgF peptide(s) form a constriction in the pore complex.
15. 15. The pore complex of claim 14, which is a homo-oligomer comprising 6 to 10 pore monomer conjugates according to any one of claims 1 to 11 or 1 to 5 constructs according to claims 12 or 13.
16. 16. The pore complex of claim 14 or 15, wherein the CsgF peptide(s) is inserted into the lumen of the pore complex.
17. A pore multimer comprising two or more pores, at least one of said pores being a pore complex according to any one of claims 14 to 16.
18. A pore complex according to any one of claims 14 to 16 or a pore multimer according to claim 17, which is comprised in a membrane.
19. A membrane comprising a pore complex according to any one of claims 14 to 16 or a pore multimer according to claim 17.
20. 12. A method for producing a pore monomer conjugate according to any one of claims 1 to 11, comprising attaching the CsgF peptide to the CsgG pore monomer at two or more positions.
21. A method for producing a pore complex described in any one of claims 14 to 16 or a pore multimer described in claim 17, comprising expressing at least one pore monomer conjugate described in any one of claims 1 to 11 or a construct described in claim 12 or 13, and sufficient pore monomer or construct to form the pore complex or pore multimer in a host cell, and allowing the pore complex or pore multimer to form in the host cell.
22. A method for producing a pore complex described in any one of claims 14 to 16 or a pore multimer described in claim 17, comprising contacting in vitro at least one pore monomer conjugate described in any one of claims 1 to 11 or a construct described in claim 12 or 13 with sufficient pore monomer or construct, and allowing the pore complex or pore multimer to form.
23. 1. A method for determining the presence, absence, or one or more characteristics of a target analyte, comprising: (i) contacting the target analyte with the pore complex of any one of claims 14 to 16 or the pore multimer of claim 17 such that the target analyte moves relative to the pore complex or the pore multimer; (ii) obtaining one or more measurements as the analyte moves relative to the pore complex or pore multimer, thereby determining the presence, absence, or one or more characteristics of the analyte.
24. 24. The method of claim 23, wherein the analytes are small organic or inorganic compounds such as peptides, polypeptides, monosaccharides, oligosaccharides, polysaccharides, pharmacologically active compounds, toxic compounds, and pollutants.
25. 25. The method of claim 24, wherein the analyte is a polynucleotide.
26. 26. The method of claim 25, wherein the polynucleotide comprises at least one homopolymer region.
27. 27. The method of claim 25 or 26, comprising determining one or more characteristics selected from: (i) the length of the polynucleotide; (ii) the identity of the polynucleotide; (iii) the sequence of the polynucleotide; (iv) the secondary structure of the polynucleotide; and (v) whether the polynucleotide is modified.
28. A method for characterizing a polynucleotide, peptide or polypeptide using a pore complex according to any one of claims 14 to 16 or a pore multimer according to claim 17.
29. 18. Use of a pore complex according to any one of claims 14 to 16 or a pore multimer according to claim 17 for determining the presence, absence or one or more properties of a target analyte.
30. A polynucleotide encoding a pore monomer conjugate according to any one of claims 1 to 11 or a construct according to claim 12 or 13.
31. A kit for characterizing a target analyte, comprising: (a) a pore complex described in any one of claims 14 to 16 or a pore multimer described in claim 17; and (b) a membrane component.
32. A kit for characterizing a target polynucleotide or a target polypeptide, comprising: (a) a pore complex described in any one of claims 14 to 16 or a pore multimer described in claim 17; and (b) a polynucleotide-binding protein or a polypeptide-processing enzyme.
33. An apparatus for characterizing a target polynucleotide or target polypeptide in a sample, comprising: (a) a plurality of pore complexes described in any one of claims 14 to 16 or a plurality of pore multimers described in claim 17; and (b) a plurality of polynucleotide binding proteins or a plurality of polypeptide handling enzymes.
34. 20. An array comprising a plurality of the membrane of claim 19.
35. 35. A system comprising: (a) the membrane of claim 19 or the array of claim 34; (b) means for applying an electric potential across the membrane(s); and (c) means for detecting an electrical or optical signal across the membrane(s).
36. A device comprising a pore complex according to any one of claims 14 to 16 or a pore multimer according to claim 17 inserted into an in vitro membrane.
37. A device manufactured by a method comprising: (i) obtaining a pore complex described in any one of claims 14 to 16 or a pore multimer described in claim 17; and (ii) contacting the pore complex or pore multimer with an in vitro membrane such that the pore complex or pore multimer is inserted into the in vitro membrane.