Novel hole cellular and hole
By using the pore complex formed by functionalized pore monomers and chaperone molecules, combining the functional binding part, the problems of long sample preparation time and low sequencing accuracy in the prior art are solved, and efficient and accurate analyte detection and characterization are achieved.
Patent Information
- Application Number
- CN202380077485.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-11-11
- Filing Date
- 2023-11-10
- Publication Date
- 2025-06-13
AI Technical Summary
Prior art When using nanopore sensing for analyte characterization, sample preparation time is long and sequencing accuracy is low, and hairpin-linked polynucleotide translocation through nanopores will lead to trans strand rehybridization, reducing sequencing accuracy.
By binding to the functional binding moiety, the percentage of double-stranded polynucleotides is increased and the capture of the target polynucleotide is promoted by specific hybridization.
Improves the target detection capability of analytes, reduces sample preparation time, enhances sequencing accuracy, and simplifies the calculation process.
Smart Images

Figure BDA0005389675720000761 
Figure BDA0005389675720000762 
Figure BDA0005389675720000771
Abstract
Description
Technical Field
[0001] The present invention relates to novel pore monomer conjugates comprising a pore monomer and a functionalized partner molecule, pore complexes formed from such conjugates, and their use in analyte detection and characterization. Background Art
[0002] Nanopore sensing is an analyte detection and characterization method that relies on observations of individual binding or interaction events between analyte molecules and ion-conducting channels. Two of the fundamental components in using nanopore sensing for analyte characterization are (1) controlling the movement of the analyte through the pore, and (2) discerning the constituent building blocks as the analyte moves through the pore. During nanopore sensing, the narrowest part of the pore forms the most discerning part of the current signature of the nanopore as a function of the analyte passing through. CsgG has been identified as a non-gated, non-selective protein secretion channel from Escherichia coli (Goyal et al., 2014) and has been used as a nanopore for detecting and characterizing analytes. Mutations of the wild-type CsgG pore that improve the properties of the pore in this context have also been disclosed (WO 2016 / 034591, WO 2017 / 149316, WO 2017 / 149317, WO 2017 / 149318, WO 2018 / 211241, and WO 2019 / 002893, which are incorporated herein by reference in their entirety).
[0003] Methods for sequencing double-stranded polynucleotides have been developed, for example, involving the translocation of both a hairpin-linked template and a complement strand. Measuring both strands in this way is advantageous because information from two complementary-linked strands can be combined and used to provide observations with higher confidence than can be obtained by measuring only the template strand. However, preparing such hairpin-linked polynucleotides increases the sample preparation time and results in loss of valuable analytes. In addition, the translocation of the hairpin-linked template and complement polynucleotide strands through the nanopore causes rehybridization of the strands on the other side (trans) of the nanopore. This alters the translocation rate, resulting in lower sequencing accuracy. Furthermore, due to the difference in current-time data of the template strand and the complement strand, two algorithms are used for calculation, which makes the calculation more complex and intensive. An improved method for characterizing double-stranded polynucleotides with increased accuracy and higher efficiency / throughput is described in WO 2018 / 100370. Summary of the Invention
[0004] The inventors have surprisingly found that chaperone molecules can be used to functionalize pore monomers, and that pore complexes formed from these functionalized pore monomers have an improved ability to determine the presence, absence, or one or more characteristics of a target analyte. For example, in the examples below, the inventors have surprisingly shown that the CsgA chaperone polypeptide can be used to functionalize the CsgG pore with a pore tether and increase the total read percentage of the pore for double-stranded polynucleotides. The pore monomers and pores of the present invention can be used in conjunction with a variety of different applications, as described in more detail below. Accordingly, the present invention provides a pore monomer conjugate comprising a pore monomer, a chaperone molecule, and a functional binding moiety, wherein the chaperone molecule has an affinity for the pore monomer, and wherein the functional binding moiety is linked to the pore monomer via the chaperone molecule.
[0005] The present invention also provides:
[0006] - A construct comprising two or more covalently attached pore monomer conjugates of the present invention;
[0007] - A pore complex comprising at least one pore monomer conjugate of the present invention or at least one construct of the present invention;
[0008] - A pore multimer comprising two or more pores, wherein at least one of the pores is a pore complex of the present invention;
[0009] - A membrane comprising a pore complex of the present invention or a pore multimer of the present invention;
[0010] - A method for determining the presence, absence, or one or more characteristics of a target analyte, the method comprising the steps of:
[0011] (i) contacting the target analyte with a pore complex of the present invention or a pore multimer of the present invention such that the target analyte moves relative to the pore complex or pore multimer; and
[0012] (ii) making one or more measurements as the target analyte moves relative to the pore complex or pore multimer, and thereby determining the presence, absence, or one or more characteristics of the target analyte.
[0013] - Use of a pore complex of the present invention or a pore multimer of the present invention for determining the presence, absence, or one or more characteristics of a target analyte;
[0014] - A kit for characterizing a target analyte, the kit comprising (a) a pore complex of the present invention or a pore multimer of the present invention and (b) components of a membrane;
[0015] - A kit for characterizing a target polynucleotide or a target polypeptide, the kit comprising (a) the pore complex of the present invention or the pore polymer of the present invention, and (b) a polynucleotide-binding protein or a polypeptide-binding protein;
[0016] - An apparatus for characterizing a target polynucleotide or a target polypeptide in a sample, the apparatus comprising (a) a plurality of the pore complexes of the present invention or a plurality of the pore polymers of the present invention, and (b) a plurality of polynucleotide-binding proteins or a plurality of polypeptide-binding proteins;
[0017] - An array comprising a plurality of the membranes of the present invention;
[0018] - A system comprising (a) the membrane of the present invention or the array of the present invention, (b) means for applying an electric potential across the membrane,
[0019] and (c) means for detecting an electrical signal or an optical signal across the membrane;
[0020] - An apparatus comprising the pore complex of the present invention or the pore polymer of the present invention inserted into an ex vivo membrane;
[0021] - An apparatus produced by a method comprising: (i) obtaining the pore complex of the present invention or the pore polymer of the present invention; and (ii) contacting the pore complex or the pore polymer with an ex vivo membrane such that the pore complex or the pore polymer is inserted into the ex vivo membrane;
[0022] - A pore monomer conjugate comprising a CsgG pore monomer, a chaperone molecule, and a functional binding moiety,
[0023] wherein the CsgG pore monomer comprises a sequence that is at least about 40% homologous or identical to the amino acid sequence of SEQ ID NO: 3 over the entire sequence,
[0024] wherein the chaperone molecule comprises SEQ ID NO: 21, 23, 25, 27, 29, 31, 63, 64, 65, 66, 67, or 68,
[0025] wherein the K at position 7 of SEQ ID NO: 21, 23, 25, 27, 29, or 31 is covalently linked to the CsgG pore monomer through a sulfonyl group, or the K at position 8 of SEQ ID NO: 63, 64, 65, 66, 67, or 68 is covalently linked to the CsgG pore monomer through a sulfonyl group,
[0026] wherein the functional binding moiety comprises an oligonucleotide, a polynucleotide, a polynucleotide analogue, or a morpholino that is capable of specifically hybridizing to a target polynucleotide analyte, and
[0027] wherein the functional binding moiety is covalently linked to the C-terminal K of SEQ ID NO: 21, 23, 25, 27, 29, 31, 63, 64, 65, 66, 67 or 68; and
[0028] - A method for determining the presence, absence or one or more characteristics of a target polynucleotide, the method comprising the steps of: (a) contacting a double-stranded polynucleotide comprising a template and a complementary strand with a pore complex or pore polymer comprising at least one pore monomer conjugate of the present invention, wherein the complementary strand comprises a binding region capable of hybridizing to the functional binding moiety such that when the template strand moves through the pore complex or pore polymer, the two strands separate to expose the binding region on the complementary strand, wherein the functional binding moiety hybridizes to the binding region to facilitate capture of the complementary strand through the pore complex or pore polymer; and (b) making one or more measurements as the template strand and the complementary strand move relative to the pore complex or pore polymer, wherein the measurements of both the template strand and the complementary strand are used to determine the presence, absence or one or more characteristics of the target polynucleotide. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 : SDS-PAGE gel analysis of only the pore control of CsgG (CsgG-F56Q) and the CsgG-CsgA-morpholino complex when decomposed into its constituent monomer components after boiling in the presence of DTT. Lanes 2-12 correspond to CsgA polypeptide variants, where the position of the sulfonyl fluoride moves along the CsgA polypeptide. Compared to the sample with only pores in lane 1, lanes 2-12 show a band shift. This means that the CsgG monomers are covalently modified with the CsgA-morpholino strand. Lanes 8-11 show a high reaction efficiency, where most of the CsgG monomers are modified with the CsgA-morpholino strand.
[0030] Figure 2 : SDS-PAGE gel analysis of only the CsgG / CsgF complex control and the CsgG / CsgF-CsgA-N3 complex, where the polypeptide length ranges between 14 and 19 residues, and where each complex is linked with a BCN-morpholino strand. Relative to the sample with only pores in lane 1, lanes 2, 4, 6, 8 and 10 each show a smaller band shift. Lanes 3, 5, 7, 9 and 11 each show a band shift of approximately 40 kDa corresponding to the pore-CsgA-morpholino conjugate, which is absent in the sample with only pores in lane 1. This means that each CsgA polypeptide is covalently linked to the CsgG monomer.
[0031] Figure 3a: Trace of ionic current (pA) vs. time (s) when the DNA strands of a duplex pair translocate through the CsgG-CsgA-morpholino modified pores inserted into the MinION flow cell. The original current trace is shown in black, and the event detection signal is shown in red. Shorter timescale stretches of the template and complement strands of the duplex pair are indicated. Aligned base calls for the pair are shown, demonstrating forward and reverse alignments to the same genomic location in the reference.
[0032] Figure 3b : Period bar graph of sequencing data, where the percentage of DNA bases sequenced during that period is assigned as part of duplex pairs (duplex template in dark grey, duplex complement in black) or 1D strands (light grey). Data from three MinION flow cells are shown, where two flow cells with CsgG-CsgA-morpholino modified pores are compared to unmodified pore controls. The total amount of duplex pair data rises in the modified pores and remains constant during the sequencing experiment.
[0033] Figure 3c : Bar graph showing the percentage of DNA bases sequenced in 3 separate MinION sequencing experiments. The sequenced bases are assigned as part of duplex pairs (duplex template in dark grey, duplex complement in black) or 1D strands (light grey). Data from three MinION flow cells are shown, where two flow cells with CsgG-CsgA-morpholino modified pores are compared to unmodified controls. The total amount of duplex pair data rises in the modified pores.
[0034] Figure 3d : Bar graph showing the total amount of DNA bases sequenced in 8 separate MinION sequencing experiments. The first 3 bar graphs (left to right) show the same data as Figure 3c : Bars 4 - 8 show data from modified pores where the CsgA polypeptides (SEQ ID NOs: 43 - 47 in order from left to right in Table 3) are longer than the CsgA polypeptide (SEQ ID NO: 40 in Table 1) used in Figure 3c : The sequenced bases are assigned as part of duplex pairs (duplex template in dark grey, duplex complement in black) or 1D strands (light grey). The total amount of duplex pair data rises in all tested adapter settings compared to unmodified pore controls.
[0035] Figure 3e: A bar graph showing the total amount of DNA bases sequenced in 8 separate MinION sequencing experiments. The data relates to the CsgA polypeptides shown in Table 5. The first two columns (from left to right) relate to SEQ ID NO:48. The next two columns relate to SEQ ID NO:40. The next two columns relate to SEQ ID NO:49. The last two columns (from left to right) relate to SEQ ID NO:50. The sequenced bases are assigned as part of double-stranded pairs (dark grey double-stranded template, black double-stranded complement) or 1D strands (light grey). The total amount of double-stranded pair data increased in all tested linker environments compared to the unmodified pore control.
[0036] Figure 4 : Native PAGE gel analysis of CsgG pore control and CsgG pores complexed with two linked CsgA polypeptides (SEQ ID NO:69 and 70), where each CsgA polypeptide contains an SO 2 F group. The C-terminus of one CsgA polypeptide (SEQ ID NO:70) is linked to the R group of K at the C-terminus of another CsgA polypeptide (SEQ ID NO:69). This provides a linked construct where one CsgA polypeptide is N-to-C linked to another C-to-N. Then the N-terminus of each CsgA polypeptide can be linked to the CsgG pore. Samples were heated from 62.0 °C to 87.8 °C in lanes 1 - 12. The CsgG single pore is the band shown between 242 kDa and 480 kDa, while the dimer band is between 480 and 720 kDa. For the CsgG pore control, we observed a general slight decrease in the intensity of the single pore band at higher temperatures, which may be due to the decomposition of pore oligomers into individual polypeptide chains. For CsgG pores complexed with linked CsgA polypeptides, no significant differences were observed in the ratio of dimers to single pores within the tested temperature range. We observed a general decrease in band intensity at higher temperatures, which may be due to the decomposition of pore oligomers into individual polypeptide chains. Significantly more dimer pores were observed compared to the control sample, indicating that the CsgA peptide stabilizes the pore dimers.
[0037] Figure 5 : A bar graph showing the total amount of DNA bases sequenced in 3 separate MinION sequencing experiments. The data relates to CsgG - CsgF nanopore complexes variably modified with CsgA - morpholino as indicated. The presence or absence of a competitor for the morpholino sequence is also indicated. The sequenced bases are assigned as part of double-stranded pairs (dark grey double-stranded template, black double-stranded complement) or 1D strands (light grey). When the nanopore is modified and no competitor is present, the total amount of double-stranded pair data increases, and the presence of a competitor restores the total amount of double-stranded pair data to that of the unmodified control pore.
[0038] Figure 6 : Structure and size of the wild-type CsgG pore from Escherichia coli strain K12 (the database accession code for this structure is 4UV3). The distances shown are measured from backbone to backbone of the amino acids forming the pore structure. The CsgG pore is a tightly interconnected symmetric nonameric pore, resembling a crown. The total height is and the maximum outer diameter is It defines a central channel and consists of three parts: (A) the cap region, (B) the constriction region, and (C) the transmembrane β-barrel region. The axial length or height of the cap is Its inner diameter is The opening is The β-barrel has 36 strands, the axial length of and the inner diameter of At the level of the predicted lipid-water interface, the transition between the pore cap and the β-barrel is distinct and is the constriction located between them. The diameter of the constriction is approximately and exhibits along the axis of the channel
[0039] Figure 7 : A higher ratio of polynucleotide-binding protein to pore compared to the pore alone results in an upward shift of the pore band, indicating that the polynucleotide-binding protein and the pore are forming a complex. The white dashed line represents the band of the pore alone, confirming that the polynucleotide-binding protein:pore band is shifted compared to the band of the pore alone. In this figure, the polypeptide-binding protein is labeled "motor".
[0040] Figure 8 : The proportion of polynucleotide-binding protein-driven strand capture events initiated within the first 40 nucleotides of the strand, as determined by mapping the signals to a reference using an algorithm based on an internal hidden Markov model (HMM). When the polynucleotide-binding protein is attached to the nanopore, this proportion is greater than when the polynucleotide-binding protein is unmodified, indicating that the events originate from the polynucleotide-binding protein attached to the nanopore. In this figure, the polypeptide-binding protein is labeled "motor".
[0041] Figure 9 : The total number of polynucleotide-binding protein-driven DNA capture events obtained during the experiment. Experiments using tethered polynucleotide-binding protein showed an increased capture rate of polynucleotide-binding protein-driven DNA strands compared to experiments using control unmodified polynucleotide-binding protein. In summary, these results indicate that by functionalizing the nanopore with morpholino oligonucleotides via CsgA and functionalizing the polynucleotide-binding protein with complementary morpholino oligonucleotides, the polynucleotide-binding protein can be stably complexed with the nanopore and used to generate polynucleotide-binding protein-driven DNA signals through the nanopore. In this figure, the polypeptide-binding protein is labeled "motor".
[0042] Sequence Listing Description
[0043] SEQ ID NO:1 shows the polynucleotide sequence of wild-type Escherichia coli CsgG from strain K12, including the signal sequence (Gene ID: 945619).
[0044] SEQ ID NO:2 shows the amino acid sequence of wild-type Escherichia coli CsgG, including the signal sequence (Uniprot accession number P0AEA2).
[0045] SEQ ID NO:3 shows the amino acid sequence of wild-type Escherichia coli CsgG as a mature protein (Uniprot accession number P0AEA2).
[0046] SEQ ID NO:4 shows the polynucleotide sequence of wild-type Escherichia coli CsgA from strain K12, including the signal sequence.
[0047] SEQ ID NO:5 shows the amino acid sequence of wild-type Escherichia coli CsgA, including the signal sequence.
[0048] SEQ ID NO:6 shows the amino acid sequence of wild-type Escherichia coli CsgA as a mature protein.
[0049] SEQ ID NO:7 shows the sequence of a fragment of CsgA, GVVPQYGGGG.
[0050] SEQ ID NO:8 shows the sequence of the CsgA polypeptide used in the present invention, GVVPQYKGGG.
[0051] SEQ ID NO:9 shows the sequence of a fragment of CsgA, GVVPQYGGGGN.
[0052] SEQ ID NO:10 shows the sequence of the CsgA polypeptide used in the present invention, GVVPQYKGGGN.
[0053] SEQ ID NO:11 shows the sequence of a fragment of CsgA, GVVPQYGGGGNH.
[0054] SEQ ID NO:12 shows the sequence of the CsgA polypeptide used in the present invention, GVVPQYKGGGNH.
[0055] SEQ ID NO:13 shows the sequence of a fragment of CsgA, GVVPQYGGGGNHG.
[0056] SEQ ID NO:14 shows the sequence of the CsgA polypeptide used in the present invention, GVVPQYKGGGNHG.
[0057] SEQ ID NO:15 shows the sequence of a fragment of CsgA, GVVPQYGGGGNHGG.
[0058] SEQ ID NO:16 shows the sequence of the CsgA polypeptide used in the present invention, GVVPQYKGGGNHGG.
[0059] SEQ ID NO:17 shows the sequence of a fragment of CsgA, GVVPQYGGGGNHGGG.
[0060] SEQ ID NO:18 shows the sequence of the CsgA polypeptide used in the present invention, GVVPQYKGGGNHGGG.
[0061] SEQ ID NO:19 shows the sequence of the linker, SGS.
[0062] SEQ ID NO:20 shows the sequence of the CsgA polypeptide used in the present invention, GVVPQYGGGGSGSK.
[0063] SEQ ID NO:21 shows the sequence of the CsgA polypeptide used in Example 1, GVVPQYKGGGSGSK.
[0064] SEQ ID NO:22 shows the sequence of the CsgA polypeptide used in the present invention, GVVPQYGGGGNSGSK.
[0065] SEQ ID NO:23 shows the sequence of the CsgA polypeptide used in Example 1, GVVPQYKGGGNSGSK.
[0066] SEQ ID NO:24 shows the sequence of the CsgA polypeptide used in the present invention, GVVPQYGGGGNHSGSK.
[0067] SEQ ID NO:25 shows the sequence of the CsgA polypeptide used in Example 1, GVVPQYKGGGNHSGSK.
[0068] SEQ ID NO:26 shows the sequence of the CsgA polypeptide used in the present invention, GVVPQYGGGGNHGSGSK.
[0069] SEQ ID NO:27 shows the sequence of the CsgA polypeptide used in Example 1, GVVPQYKGGGNHGSGSK.
[0070] SEQ ID NO:28 shows the sequence of the CsgA polypeptide used in the present invention, GVVPQYGGGGNHGGSGSK.
[0071] SEQ ID NO:29 shows the sequence of the CsgA polypeptide used in Example 1, GVVPQYKGGGNHGGSGSK.
[0072] SEQ ID NO:30 shows the sequence of the CsgA polypeptide used in the present invention, GVVPQYGGGGNHGGGSGSK.
[0073] SEQ ID NO:31 shows the sequence of the CsgA polypeptide used in Example 1, GVVPQYKGGGNHGGGSGSK.
[0074] SEQ ID NO:32 - 42 are the modified CsgA polypeptides in Table 1 of Example 1.
[0075] SEQ ID NO:43 - 47 are the modified CsgA polypeptides in Table 3 of Example 1.
[0076] SEQ ID NO:48 - 50 are the modified CsgA polypeptides in Table 5 of Example 1.
[0077] SEQ ID NO:51 - 56 are the oligonucleotides used in Example 2.
[0078] SEQ ID NO:57 shows the sequence of the CsgA polypeptide used in the present invention, GVVPQYGKGG.
[0079] SEQ ID NO:58 shows the sequence of the CsgA polypeptide used in the present invention, GVVPQYGKGGN.
[0080] SEQ ID NO:59 shows the sequence of the CsgA polypeptide used in the present invention, GVVPQYGKGGNH.
[0081] SEQ ID NO:60 shows the sequence of the CsgA polypeptide used in the present invention, GVVPQYGKGGNHG.
[0082] SEQ ID NO:61 shows the sequence of the CsgA polypeptide used in the present invention, GVVPQYGKGGNHGG.
[0083] SEQ ID NO:62 shows the sequence of the CsgA polypeptide used in the present invention, GVVPQYGKGGNHGGG.
[0084] SEQ ID NO:63 shows the sequence of the CsgA polypeptide used in the present invention, GVVPQYGKGGSGSK.
[0085] SEQ ID NO:64 shows the sequence of the CsgA polypeptide used in the present invention, GVVPQYGKGGNSGSK.
[0086] SEQ ID NO:65 shows the sequence of the CsgA polypeptide used in the present invention, GVVPQYGKGGNHSGSK.
[0087] SEQ ID NO:66 shows the sequence of the CsgA polypeptide used in the present invention, GVVPQYGKGGNHGSGSK.
[0088] SEQ ID NO:67 shows the sequence of the CsgA polypeptide used in the present invention, GVVPQYGKGGNHGGSGSK.
[0089] SEQ ID NO:68 shows the sequence of the CsgA polypeptide used in the present invention, GVVPQYGKGGNHGGGSGSK.
[0090] SEQ ID NO:69 is the sequence used in Example 1, GVVPQY(KSO2F)GPPGGPPK.
[0091] SEQ ID NO:70 is the sequence used in Example 1, GVVPQY(KSO2F)PPGGPPK. Detailed Description
[0092] All publications, patents, and patent applications cited herein, whether supra or infra, are hereby incorporated by reference in their entirety. All publications, patents, and patent applications mentioned in this specification are incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent that the publications and patents or patent applications incorporated by reference conflict with the present invention contained in the specification, the specification is intended to supersede and / or take precedence over any such conflicting material.
[0093] The present invention will be described with respect to specific embodiments and with reference to certain figures, but the invention is not limited thereto and is only limited by the claims. Any reference signs in the claims shall not be construed as limiting the scope. Of course, it should be understood that not all aspects or advantages need to be achieved in accordance with any particular embodiment of the present invention. Thus, for example, those skilled in the art will recognize that the present invention may be practiced or carried out in a manner that achieves or optimizes one or a group of the advantages taught herein without necessarily achieving other aspects or advantages taught or suggested herein.
[0094] In addition, as used in this specification and the appended claims, unless the context clearly indicates otherwise, the singular forms "a / an" and "the" include plural referents. Thus, for example, reference to "a polynucleotide" includes two or more polynucleotides, reference to "a polynucleotide-binding protein" includes two or more such proteins, reference to "a helicase" includes two or more helicases, reference to "a monomer" refers to two or more monomers, reference to "a pore" includes two or more pores, and the like.
[0095] Throughout all discussions herein, the standard single-letter codes for amino acids are used. They are as follows: alanine (A), arginine (R), asparagine (N), aspartic acid (D), cysteine (C), glutamic acid (E), glutamine (Q), glycine (G), histidine (H), isoleucine (I), leucine (L), lysine (K), methionine (M), phenylalanine (F), proline (P), serine (S), threonine (T), tryptophan (W), tyrosine (Y), and valine (V). Standard substitution symbols are also used, i.e., Q42R means that Q at position 42 is replaced by R.
[0096] In paragraphs herein where different amino acids at a specific position are separated by the / symbol, the / symbol means "or". For example, Q87R / K means Q87R or Q87K. In paragraphs herein where different positions are separated by the / symbol, the / symbol means "and", such that Y51 / N55 is Y51 and N55.
[0097] The general definitions in WO 2019 / 002893 are incorporated herein by reference in their entirety.
[0098] Pore monomer conjugate
[0099] The present invention provides pore monomer conjugates, which comprise a pore monomer, a chaperone molecule, and a functional binding moiety. The chaperone molecule has an affinity for the pore monomer. The chaperone molecule typically binds or is linked to the pore monomer. The linkage can include covalent, non-covalent, supramolecular, and / or native interactions. The linkage can include non-covalent, supramolecular, and / or native interactions. The chaperone molecule is preferably covalently linked to the pore monomer. The functional binding moiety is linked, preferably covalently linked, to the chaperone molecule. The functional binding moiety is linked to the pore monomer via the chaperone molecule. The pore monomer conjugate typically has the following structure: pore monomer ∼ chaperone molecule - functional binding moiety, where "∼" = affinity, binding, linkage, or covalent linkage, and "-" = linkage or covalent linkage.
[0100] The pore monomer binds or is linked (preferably covalently linked) to the chaperone molecule, and then the chaperone molecule is linked (preferably covalently linked) to the functional binding moiety. The functional binding moiety is indirectly linked to the pore monomer via the chaperone molecule. The pore monomer is not directly linked to the functional binding moiety. In the pore monomer conjugate of the present invention, the functional binding moiety typically does not serve to link the chaperone molecule to the pore monomer.
[0101] The pore monomer can be from or derived from any pore. The pore is typically a protein pore. Pores suitable for characterizing a target analyte are known in the art. For example, examples are disclosed in WO 2010 / 004265, WO2012 / 107778, WO 2013 / 153359, WO 2015 / 166275, WO 2016 / 055778, WO 2015 / 166276, WO 2016 / 132123, WO 2017 / 174990, WO2018 / 146491, WO 2020 / 095052, WO 2020 / 208357, PCT / GB2022 / 052196, PCT / EP2022 / 077537, WO 2016 / 034591, WO 2017 / 149316, WO 2017 / 149317, WO 2017 / 149318, WO2018 / 211241, and WO 2019 / 002893 (all incorporated herein by reference in their entirety). Examples are also described in patent application numbers GB 2118939.4, US63 / 396,539, GB 2205617.0, GB 2211602.4, GB2216026.1, US 63 / 370,875, and GB 2211607.3 (incorporated herein by reference in their entirety).
[0102] The pore monomer is preferably from or derived from Wza, Iota toxin, anthrax protective antigen, Vibrio cholerae cytolysin, Cytotoxin K (CytK), CELIII, CsgG, Aerolysin, alpha-hemolysin, InvG, GspD, MspA, MspB, MspC, PorARr, PorBRr, PorARc, PilQ, necrotic enteritis B-like toxin (NetB), FraC, gated proteins (including G20c, P23_45, T4, SPP1, P22, and Phi29), gamma-hemolysin, Monalysin, Lysenin, ClyA, and Clostridium perfringens beta-toxin.
[0103] The pore monomer can be composed of two or more different pore monomers from or derived from any of these pores. The pore monomer can be a chimeric pore monomer that contains two or more regions, such as 3, 4, 5, 6, 7 or more regions, where at least two of the two or more regions (such as at least 3, 4, 5, 6 or 7) are from at least two different pores, such as from at least 3, 4, 5, 6 or 7 different pores.
[0104] A pore monomer "is from" or "is derived from" a pore if it shares significant homology / identity with the sequence of a wild-type or naturally occurring pore monomer from the pore. The pore monomer preferably contains a sequence that has at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90% or more preferably at least about 95%, at least about 97%, at least about 98% or at least about 99% homology with the sequence of a wild-type or naturally occurring pore monomer. The pore monomer preferably contains a sequence that has at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90% or more preferably at least about 95%, at least about 97%, at least about 98% or at least about 99% identity with the sequence of a wild-type or naturally occurring pore monomer. Homology and / or identity are typically measured over the entire length of the region. Methods for measuring homology and / or identity will be discussed in more detail below. The sequences of the pores listed above are publicly available through GenBank and various references above and below.
[0105] If a pore monomer shares significant homology / identity with a fragment or portion of the sequence of a wild-type or naturally occurring pore monomer, then the pore monomer can "be from" or "be derived from" the pore. Thus, the sequence can have less than 40% overall sequence homology / identity with the overall sequence of the pore monomer, but the sequence of a particular region, domain, or subunit can share at least about 80%, at least about 90%, or up to at least about 99% sequence homology or identity with the corresponding region of the pore monomer. Within an extension of 100 or more, such as 125, 150, 175, or 200 or more contiguous or adjacent amino acids of the pore monomer, there can be at least about 80%, such as at least about 85%, at least about 90%, or at least about 95% homology or identity ("hard homology").
[0106] The pore is preferably a CsgG pore monomer. Such monomers are discussed in more detail below.
[0107] The chaperone molecule can be any molecule. The chaperone molecule can be a molecule or a modified version of a molecule that interacts with the pore monomer during its natural function. The chaperone molecule can be, for example, a molecule or a modified version of a molecule that forms a complex with the pore monomer in nature.
[0108] The chaperone molecule preferably comprises or consists of: a polymer, an amino acid, a peptide, a polypeptide, a protein, a nucleotide, an oligonucleotide, a polynucleotide, a polynucleotide-polypeptide conjugate, a monosaccharide, an oligosaccharide, or a polysaccharide.
[0109] The chaperone molecule preferably comprises or consists of the following: oligonucleotides or polynucleotides (such as nucleic acids). Oligonucleotides or polynucleotides are defined as macromolecules containing two or more nucleotides. Oligonucleotides, polynucleotides or nucleic acids can contain any combination of any nucleotides. Nucleotides can be naturally occurring or artificial. One or more nucleotides in the oligonucleotide or polynucleotide can be oxidized or methylated. One or more nucleotides in the oligonucleotide or polynucleotide can be damaged. For example, the oligonucleotide or polynucleotide can contain pyrimidine dimers. Such dimers are typically associated with damage caused by ultraviolet light and are a major cause of cutaneous melanoma. One or more nucleotides can be modified, for example, with a label or tag, suitable examples of which are known to those skilled in the art. The oligonucleotide or polynucleotide can contain one or more spacers. Nucleotides generally contain a nucleobase, a sugar and at least one phosphate group. The nucleobase and sugar form a nucleoside. Nucleobases are generally heterocyclic. Nucleobases include but are not limited to purines and pyrimidines, and more specifically include adenine (A), guanine (G), thymine (T), uracil (U) and cytosine (C). Sugars are generally pentoses. Nucleotide sugars include but are not limited to ribose and deoxyribose. The sugar is preferably deoxyribose. The polynucleotide preferably comprises the following nucleosides: deoxyadenosine (dA), deoxyuridine (dU) and / or thymidine (dT), deoxyguanosine (dG) and deoxycytidine (dC). Nucleotides are generally ribonucleotides or deoxyribonucleotides. Nucleotides generally contain mono-, di- or triphosphates. Nucleotides can contain more than three phosphates, such as 4 or 5 phosphates. The phosphate can be attached to the 5' or 3' side of the nucleotide. Nucleotides can be linked to each other in any way. Nucleotides are generally linked through their sugar and phosphate groups, as in nucleic acids. Nucleotides can be linked through their nucleobases, as in pyrimidine dimers. Oligonucleotides or polynucleotides can be single-stranded or double-stranded. At least a portion of the polynucleotide can be double-stranded. Polynucleotides can be ribonucleic acid (RNA) or deoxyribonucleic acid (DNA).
[0110] Oligonucleotides or polynucleotides can be of any length. For example, the length of the chaperone oligonucleotide or polynucleotide can be at least 10, at least 25, at least 30, at least 40, at least 50, at least 100, at least 150, at least 200, at least 250, at least 300, at least 400 or at least 500 nucleotides or nucleotide pairs. The length of the oligonucleotide or polynucleotide can be 1000 or more nucleotides or nucleotide pairs, 5000 or more nucleotides or nucleotide pairs or 100000 or more nucleotides or nucleotide pairs.
[0111] Nucleotides include, but are not limited to, adenosine monophosphate (AMP), guanosine monophosphate (GMP), thymidine monophosphate (TMP), uridine monophosphate (UMP), 5-methylcytidine monophosphate, 5-hydroxymethylcytidine monophosphate, cytidine monophosphate (CMP), cyclic adenosine monophosphate (cAMP), cyclic guanosine monophosphate (cGMP), deoxyadenosine monophosphate (dAMP), deoxyguanosine monophosphate (dGMP), deoxythymidine monophosphate (dTMP), deoxyuridine monophosphate (dUMP), deoxycytidine monophosphate (dCMP), and deoxy-5-methylcytidine monophosphate. The nucleotides are preferably selected from AMP, TMP, GMP, CMP, UMP, dAMP, dTMP, dGMP, dCMP, and dUMP. The nucleotide may be abasic (i.e., lacking a nucleobase). The nucleotide may also lack a nucleobase and a sugar (i.e., be a C3 spacer).
[0112] The chaperone molecule preferably comprises or is an amino acid, peptide, polypeptide, or protein. The amino acid, peptide, polypeptide, or protein may be naturally occurring or non-naturally occurring. The polypeptide or protein may include synthetic or modified amino acids therein. Several different types of modifications to amino acids are known in the art. It should be understood that the chaperone molecule may be modified by any method available in the art.
[0113] The polypeptide may comprise any combination of any amino acid, amino acid analogue, and modified amino acid (i.e., amino acid derivative). The amino acid / derivative / analogue may be naturally occurring or man-made. The polypeptide may comprise any naturally occurring amino acid.
[0114] One or more of the amino acids / derivatives / analogues in the polypeptide may be post-translationally modified. Suitable post-translational modifications are discussed in more detail below with reference to the target analyte.
[0115] The polypeptide may be labeled with a molecular marker. The polypeptide may contain one or more crosslinking segments, e.g., a C-C bridge. The polypeptide may comprise sulfur-containing amino acids and thus have the potential to form disulfide bonds. All of these embodiments are discussed in more detail below with reference to the target analyte. Any polypeptide discussed below with reference to the target analyte may be used as a chaperone molecule.
[0116] The chaperone polypeptide may be of any suitable length. The polypeptide preferably has a length of from about 2 to about 300 peptide units or amino acids. The polypeptide may have a length of from about 2 to about 100 peptide units, e.g., from about 2 to about 50 peptide units, e.g., from about 3 to about 50 peptide units, such as from about 5 to about 25 peptide units, e.g., from about 7 to about 16 peptide units, such as from about 9 to about 13 peptide units. "Peptide unit" may be interchangeable with "amino acid".
[0117] The chaperone polypeptide may be a polypeptide or a modified version of a polypeptide that interacts with the pore monomer during its natural function. The chaperone polypeptide may be from or derived from, for example, a polypeptide that forms a complex with the pore monomer in nature. A chaperone polypeptide is "from" or "derived from" a polypeptide that forms a complex with a pore monomer if the chaperone polypeptide shares significant homology or identity with the sequence of the polypeptide that forms a complex with the pore monomer. The chaperone polypeptide preferably comprises a sequence that has at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, or more preferably at least about 95%, at least about 97%, at least about 98%, or at least about 99% homology to the sequence of the polypeptide that forms a complex with the pore monomer. The pore monomer preferably comprises a sequence that is at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, or more preferably at least about 95%, at least about 97%, at least about 98%, or at least about 99% identical to the polypeptide sequence that forms a complex with the pore monomer. Homology and / or identity is typically measured over the entire length of the region. Methods for measuring homology and / or identity are discussed in more detail below. The sequences of the pores listed above are publicly available through GenBank and various references above and below.
[0118] A companion polypeptide may be "from" or "derived from" the pore if it shares significant homology or identity with a fragment or portion of the sequence of a wild-type or naturally occurring polypeptide that forms a complex with the pore monomer. Thus, a sequence may have less than 40% overall sequence homology or identity with the entire sequence of a polypeptide that forms a complex with the pore monomer, but the sequence of a particular region, domain or subunit may share at least about 80%, at least about 90%, or up to at least about 99% sequence homology or identity with the corresponding region of a polypeptide that forms a complex with the pore monomer. There may be at least about 80%, such as at least about 85%, at least about 90%, or at least about 95% homology or identity over a stretch of 100 or more (e.g. 125, 150, 175 or 200 or more) consecutive amino acids of a polypeptide that forms a complex with the pore monomer ("hard homology").
[0119] The chaperone polypeptide can comprise or consist of the following: a fragment or portion of a polypeptide that forms a complex with a pore monomer. Such fragments or portions typically have 100% homology or identity with the corresponding fragments or portions of the polypeptide that forms a complex with a pore monomer. The chaperone polypeptide can comprise or consist of the following: at least 5 consecutive or contiguous amino acids from a polypeptide that forms a complex with a pore monomer, such as at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 30, at least 40, at least 50, at least 100, or at least 150 consecutive or contiguous amino acids from a polypeptide that forms a complex with a pore monomer.
[0120] The chaperone molecule is preferably a CsgA polypeptide, as discussed in more detail below.
[0121] The chaperone molecule can comprise a polynucleotide and a polypeptide. The chaperone molecule can comprise or can be a polynucleotide-polypeptide conjugate. The conjugate preferably comprises a polynucleotide conjugated to a polypeptide. Polynucleotide-polypeptide conjugates will be discussed in more detail below with reference to the target analyte. Any of these polynucleotide-polypeptide conjugates can be used as a chaperone molecule.
[0122] The chaperone molecule has an affinity for the pore monomer. The chaperone molecule preferably has a high affinity for the pore monomer. If the chaperone molecule binds with a Kd of 1 x 10 -6 M or lower, more preferably 1 x 10 -7 M or lower, 5 x 10 -8 M or lower, more preferably 1 x 10 -8 M or lower, or more preferably 5 x 10 -9 M or lower, then the chaperone molecule has a high affinity for the pore monomer. If a molecule or group binds with a Kd of 1 x 10 -6 M or higher, more preferably 1 x 10 -5 M or higher, more preferably 1 x 10 -4 M or higher, more preferably 1 x 10 -3 M or higher, even more preferably 1 x 10 -2 M or higher, then the molecule or group binds with low affinity. Chaperone molecules, such as chaperone polynucleotides or chaperone polypeptides, preferably have an affinity for the pore monomer without any modification, such as using one or more of the reactive groups discussed below. Chaperone molecules, such as chaperone polynucleotides or chaperone polypeptides, preferably have a natural affinity for the pore monomer, i.e., an affinity based on their wild-type or naturally occurring structure. An example thereof is the CsgA polypeptide that has an affinity for the CsgG pore monomer.
[0123] Chaperone molecules typically bind to pore monomers. The chaperone molecules preferably bind specifically to pore monomers. A chaperone molecule binds specifically to a pore monomer if it binds to the pore monomer with preference or high affinity, but does not bind or binds only with low affinity to other or different molecules, such as other or different pore monomers, other or different polypeptides, and / or polynucleotides. Preferably, the chaperone molecule binds to the pore monomer with an affinity that is at least 10-fold, such as at least 50-fold, at least 100-fold, at least 200-fold, at least 300-fold, at least 400-fold, at least 500-fold, at least 1000-fold, or at least 10,000-fold higher than its affinity for other polynucleotides.
[0124] Affinities can be measured using known binding assays, such as those utilizing fluorescence and radioisotopes. Competitive binding assays are also known in the art. Nanopore force spectroscopy, as described in Hornblower et al., Nature Methods. 4:315 - 317. (2007), or isothermal titration calorimetry (ITC), which is a label-free quantitative technique for studying various biomolecular interactions, can be used to measure the binding strength between a peptide or protein and a polynucleotide. ITC works by directly measuring the heat released or absorbed during a biomolecular binding event.
[0125] Chaperone molecules are typically linked to pore monomers. The chaperone molecules are preferably covalently linked to pore monomers. The chaperone molecule is preferably linked to the pore monomer through one or more reactive groups. The linkage preferably includes one or more reactive groups. The linkage preferably includes a reaction between a position, nucleotide, residue, or linker in the chaperone molecule and the pore monomer. Suitable reactive groups for use in the present invention include, but are not limited to, amine-reactive groups, oxygen-reactive groups, and fluoroacetamide groups. The amine-reactive groups are preferably thioesters, NHS-esters, pentafluorophenyl esters, benzyl halides, sulfonyl fluorides, fluorosulfates, or sulfonyl triazoles. The oxygen-reactive groups are preferably alkyl halides, sulfonyl fluorides, fluorosulfates, or sulfonyl triazoles.
[0126] The chaperone molecule is preferably linked to the pore monomer through a sulfonyl fluoride reaction. When the chaperone molecule is in very close proximity to the pore monomer, a reactive group on the pore monomer, such as the R group of threonine, cysteine, tyrosine, lysine, serine, or histidine, preferably displaces the fluorine group in the sulfonyl fluoride group on the chaperone molecule to covalently link the pore monomer and the chaperone molecule. The pore monomer is preferably covalently linked to the chaperone molecule through a sulfonyl group.
[0127] Additional reactive groups are described in Nature Chemistry, 2021, 13, 1081-1092, Cell Chemical Biology, 2020, 27, 970-985, J. Am. Chem. Soc. 2019, 141, 7, 2782–2799, and Current Opinion in Chemical Biology, 2015, 18-26 (each incorporated herein by reference in its entirety).
[0128] The linkage preferably comprises one or more reactive groups that react with lysine, cysteine, tyrosine, serine, threonine, proline, tryptophan, arginine, histidine, methionine, or phenylalanine in the pore monomer. The linkage preferably comprises a reaction between a position, nucleotide, residue, or linker in the chaperone molecule and lysine, cysteine, tyrosine, serine, threonine, proline, tryptophan, arginine, histidine, methionine, or phenylalanine in the pore monomer.
[0129] Lysine, cysteine, tyrosine, serine, threonine, proline, tryptophan, arginine, histidine, methionine, or phenylalanine may be native to the pore monomer. Lysine, cysteine, tyrosine, serine, threonine, proline, tryptophan, arginine, histidine, methionine, or phenylalanine may preferably be introduced into the pore monomer by substitution or addition.
[0130] Reactive groups that react with lysine include but are not limited to maleimide, activated ester, acid anhydride, carbonate, isocyanate, isothiocyanate, a series of other acylating agents and alkylating agents, oxidative coupling of o-aminophenol, aldehyde, activated carbodiimide, ketene, sulfonyl halide, fluorosulfate, and sulfonyl triazole. A position, residue, or linker may also be attached to lysine using periodate oxidation, reductive amination, transamination, conjugation of aniline / aryl amine via oxidative coupling, azaelectrocyclization, formation of iminoborate, or conjugation of arenediazonium salt.
[0131] Reactive groups that react with cysteine include but are not limited to haloacetamide and other α-halocarbonyls, maleimide, acrylate, vinyl sulfone, vinyl pyridine, epoxide, oxanorbornadiene, mesyl-functionalized heteroaromatics, allene, allyl selenosulfate, perfluoroaromatics, thiol-ene and thiol-yne click chemistry, pyridyl disulfide, vinyl sulfone, sulfonyl halide, fluorosulfate, and sulfonyl triazole. A position, residue, or linker may also be attached to cysteine using strain-release alkylation, nickel(II)-catalyzed oxidative coupling, oxidative coupling with aminophenol, conjugation with allene (in the presence of a gold catalyst or allyl selenosulfate), native chemical ligation, Pd-catalyzed arylation / alkynylation, or cross-metathesis after allylation.
[0132] Reactive groups that react with tyrosine include, but are not limited to, sulfonyl halides, fluorosulfates, and sulfonyltriazoles. A position, residue, or linker can also be attached to tyrosine using oxidative conjugation of tyrosine, including O-alkylation, hydrazone and oxime condensations, addition reactions with electron-deficient alkynes such as alkynones, alkynoates, amides or esters, cyclic diazodicarboxamides, Pd-catalyzed alkylation, diazonium salts or Mannich reactions with imines formed from aldehydes, cyclic diazodicarboxamides, and modification with rhodium carbenes.
[0133] Reactive groups that react with serine or threonine include, but are not limited to, sulfonyl halides, fluorosulfates, and sulfonyltriazoles. A position, residue, or linker can also be attached to serine or threonine using periodate oxidation and subsequent transamination of the resulting ketone / aldehyde with hydrazide / alkoxyamine. The resulting aldehyde / ketone can also be modified by aldol ligation.
[0134] A position, residue, or linker can be attached to proline using oxidative coupling with o-aminophenol at the N-terminus.
[0135] Reactive groups that react with tryptophan include, but are not limited to, aldehydes, ketones, and tetrazoles. A position, residue, or linker can be attached to tryptophan using condensation reactions, modification with rhodium carbenes, conjugation with N / O-centered radicals, and N-terminal Trp modification using the Pictet-Spengler reaction.
[0136] A position, residue, or linker can be attached to arginine using condensation with an α,β-dicarbonyl compound.
[0137] Reactive groups that react with histidine include, but are not limited to, vinyl sulfones, sulfonyl halides, fluorosulfates, and sulfonyltriazoles. A position, residue, or linker can be attached to histidine using C2 alkylation and N3 alkylation / thiophosphorylation.
[0138] A position, residue, or linker can be attached to methionine using S-alkylation / imidation.
[0139] A position, residue, or linker can be attached to phenylalanine using modification with rhodium carbenes.
[0140] Linking, such as covalent linking, preferably includes one or more reactive groups that react with any amino acid in the pore monomer. Reactive groups that react with any amino acid include, but are not limited to, activated esters, acid anhydrides, carbonates, isocyanates, isothiocyanates, and a range of other acylating and alkylating agents, oxidative coupling of o-aminophenol, aldehydes, activated carbodiimides, ketenes, transamination, and vinylboronic acid.
[0141] Linking, such as covalent linking, preferably involves reacting a position, nucleotide, residue or linker in the partner molecule with any amino acid in the pore monomer. The position, residue or linker can be attached to any amino acid using periodate oxidation or reductive amination.
[0142] Linking, such as covalent linking, preferably involves one or more reactive groups that undergo click chemistry reactions. Suitable click chemistries include, but are not limited to, CuAAC azide / alkyne, Staudinger ligation, strain-promoted azide-alkyne cycloaddition, and inverse electron demand Diels-Alder reactions between 1,2,4,5-tetrazines and strained alkenes.
[0143] All of the above discussion regarding reactive groups and reactions for linking a partner molecule to a residue / amino acid in the pore monomer applies equally to linking the pore monomer to the partner molecule, especially if the partner molecule is a partner polypeptide. Any reactive group or reaction can be used for linking in the partner polypeptide. Specific residues in the partner polypeptide can be native. Specific residues can also preferably be introduced into the partner polypeptide by substitution or addition. A person skilled in the art is able to link, preferably covalently link, two proteins or polypeptides.
[0144] The pore monomer is preferably bound to the partner molecule via a linker. The pore monomer is preferably linked (such as covalently linked) to the partner molecule via a linker. The pore monomer is preferably bound to the partner molecule, or preferably linked (such as covalently linked) to the partner molecule via two or more linkers (such as 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more or 10 or more linkers). If two or more linkers are used, they can be the same. If two or more linkers are used, they can be different. A person skilled in the art is able to design one or more linkers for use in the present invention. The linker or one or more linkers can be any of the linkers discussed below. The linker or one or more linkers preferably include any of the reactive groups discussed above for linking the pore monomer to the partner molecule.
[0145] The linker preferably comprises a straight-chain carbon chain of 2, 3, 4, 5, 6 or more carbon atoms and / or a cyclic group containing 3, 5 or 6 carbon atoms or consists thereof.
[0146] The distance between the pore monomer and the partner molecule in the pore monomer conjugate and / or the length of the linker is preferably less than about 2.00 nm, such as less than about 1.90 nm, less than about 1.80 nm, less than about 1.70 nm, less than about 1.60 nm, less than about 1.50 nm, less than about 1.40 nm, less than about 1.30 nm, less than about 1.20 nm, less than about 1.10 nm, less than about 1.00 nm, less than about 0.90 nm, less than about 0.80 nm, less than about 0.70 nm, less than about 0.60 nm, less than about 0.50 nm, or less than about 0.40 nm.
[0147] The distance between the pore monomer and the partner molecule in the pore monomer conjugate and / or the length of the linker is preferably from about 0.40 nm to about 2.0 nm, such as from about 0.45 nm to about 1.90 nm, from about 0.50 nm to about 1.80 nm, from about 0.55 nm to about 1.7 nm, from about 0.60 nm to about 1.6 nm, from about 0.65 nm to about 1.5 nm, from about 0.7 nm to about 1.4 nm, from about 0.75 nm to about 1.3 nm, from about 0.80 nm to about 1.2 nm, from about 0.85 nm to about 1.1 nm, and from about 0.90 nm to about 1.00 nm.
[0148] The pore monomer conjugate can comprise any number of partner molecules, such as 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, or 10 or more partner molecules for each pore monomer. The partner molecules can be the same. The partner molecules can be different. Any of the reactive groups and / or linkers discussed above can be used to attach the partner molecules to the pore monomer.
[0149] Combinations of pore monomers and partner molecules include, but are not limited to, (a) a CsgG pore monomer and a CsgA polypeptide, (b) a CsgG pore monomer and a CsgB polypeptide, (c) a CsgG pore monomer and a CsgC polypeptide, (d) a CsgG pore monomer and a CsgD polypeptide, or (e) a CsgG pore monomer and a CsgE polypeptide. The pore monomer and the partner molecule are preferably selected from (i) a CsgG pore monomer and a CsgA polypeptide, (ii) a CsgG pore monomer and a CsgB polypeptide, (iii) a CsgG pore monomer and a CsgC polypeptide, or (iv) a CsgG pore monomer and a CsgE polypeptide. The pore monomer is preferably a CsgG pore monomer, and the partner molecule is preferably selected from (a) a CsgA polypeptide, (b) a CsgB polypeptide, and (c) a CsgE polypeptide, such as (a), (b), (c), (a) and (b), (a) and (c), (b) and (c), or (a), (b) and (c). The pore monomer is preferably a CsgG pore monomer, and the partner molecule is preferably a CsgA polypeptide. The CsgG pore monomer and the CsgA polypeptide are discussed in more detail below.
[0150] The functional binding moiety has an affinity for another molecule. The functional binding moiety is capable of binding to or linking to another molecule. Affinity and measurement of binding are discussed above.
[0151] The functional binding moiety is capable of binding to or linking to a target analyte. This allows the functionalized pore monomers to be used in methods for characterizing the target analyte. The target analyte preferably comprises or consists of: metal ions, inorganic salts, polymers, amino acids, peptides, polypeptides, proteins, nucleotides, oligonucleotides, polynucleotides, polynucleotide-polypeptide conjugates, monosaccharides, oligosaccharides, polysaccharides, dyes, bleaching agents, drugs, diagnostic agents, recreational drugs, explosives, toxic compounds or environmental pollutants. The target analyte preferably comprises or consists of: polypeptides, proteins, oligonucleotides, polynucleotides, polynucleotide-polypeptide conjugates, oligosaccharides or polysaccharides. These are discussed in more detail below with reference to the methods of the present invention.
[0152] The functional binding moiety comprises a region capable of binding to or linking to a molecule or target analyte. The functional binding moiety preferably comprises oligonucleotides, polynucleotides, polynucleotide analogs, morpholinos, peptide nucleic acids, polypeptides, ligands, cyclodextrins, monosaccharides, oligosaccharides, polysaccharides, boric acid, enzymes, peptides, cyclic peptides, antibodies or fragments thereof or aptamers.
[0153] Oligonucleotides and polynucleotides are defined above and below. Any polynucleotide analog can be used. Polynucleotide analogs typically contain nucleobases linked by a modified backbone. Suitable polynucleotide analogs include, but are not limited to, peptide nucleic acids (PNA), threose nucleic acids (TNA) and glycerol nucleic acids (GNA). Those skilled in the art can determine other polynucleotide analogs suitable for use in the present invention.
[0154] Morpholinos contain nucleobases linked by methylene morpholine rings joined by phosphorodiamide groups. Morpholinos are also known as morpholino oligomers or phosphorodiamide morpholino oligomers (PMO).
[0155] The functional binding moiety can comprise or consist of these oligonucleotides, polynucleotides, polynucleotide analogs or morpholinos. The functional binding moiety preferably comprises an oligonucleotide, polynucleotide, polynucleotide analog or morpholino capable of hybridizing (preferably specifically hybridizing) to a target polynucleotide. An oligonucleotide, polynucleotide, polynucleotide analog or morpholino specifically hybridizes to a target polynucleotide when it hybridizes to the target polynucleotide with preference or high affinity but substantially does not hybridize to other polynucleotides, does not hybridize to other polynucleotides or hybridizes to other polynucleotides only with low affinity. An oligonucleotide, polynucleotide, polynucleotide analog or morpholino specifically hybridizes if it hybridizes to the target polynucleotide with a melting temperature (Tm) that is at least 2 °C higher than its Tm for other sequences, such as at least 3 °C, at least 4 °C, at least 5 °C, at least 6 °C, at least 7 °C, at least 8 °C, at least 9 °C or at least 10 °C higher. More preferably, the oligonucleotide, polynucleotide, polynucleotide analog or morpholino hybridizes to the target polynucleotide with a Tm that is at least 2 °C higher than its Tm for other nucleic acids, such as at least 3 °C, at least 4 °C, at least 5 °C, at least 6 °C, at least 7 °C, at least 8 °C, at least 9 °C, at least 10 °C, at least 20 °C, at least 30 °C or at least 40 °C higher. Preferably, the oligonucleotide, polynucleotide, polynucleotide analog or morpholino hybridizes to the target polynucleotide with a Tm that is at least 2 °C higher than its Tm for a sequence that differs from the target polynucleotide by one or more nucleotides, such as by 1, 2, 3, 4 or 5 or more nucleotides, such as at least 3 °C, at least 4 °C, at least 5 °C, at least 6 °C, at least 7 °C, at least 8 °C, at least 9 °C, at least 10 °C, at least 20 °C, at least 30 °C or at least 40 °C higher. The oligonucleotide, polynucleotide, polynucleotide analog or morpholino typically hybridizes to the target polynucleotide with a Tm of at least 90 °C, such as at least 92 °C or at least 95 °C. The Tm can be experimentally measured using known techniques (including using DNA microarrays) or can be calculated using publicly available Tm calculators (such as those available on the Internet).
[0156] The conditions permitting hybridization are well known in the art (e.g., Sambrook et al., 2001, Molecular Cloning: a laboratory manual, 3rd ed., Cold Spring Harbor Laboratory Press; and Current Protocols in Molecular Biology, Chapter 2, edited by Ausubel et al., Greene Publishing and Wiley-Interscience, New York (1995)). Hybridization can be carried out under low stringency conditions, such as in the presence of a buffer solution of 30% to 35% formamide, 1 M NaCl and 1% SDS (sodium dodecyl sulfate) at 37 °C, followed by washing 20 times in 1-fold (0.1650 M Na+) to 2-fold (0.33 M Na+) SSC (standard sodium citrate) at 50 °C. Hybridization can be carried out under medium stringency conditions, such as in the presence of a buffer solution of 40% to 45% formamide, 1 M NaCl and 1% SDS at 37 °C, followed by washing in 0.5-fold (0.0825 M Na+) to 1-fold (0.1650 M Na+) SSC at 55 °C. Hybridization can be carried out under high stringency conditions, such as in the presence of a buffer solution of 50% formamide, 1 M NaCl, 1% SDS at 37 °C, followed by washing in 0.1-fold (0.0165 M Na+) SSC at 60 °C.
[0157] The oligonucleotide, polynucleotide, polynucleotide analogue or morpholino can be of any length as described above for oligonucleotides and polynucleotides (wherein for polynucleotide analogues and morpholinos, the reference to nucleotides is replaced by nucleobases). The oligonucleotide, polynucleotide, polynucleotide analogue or morpholino preferably comprises a portion or region that is substantially complementary or complementary to a portion or region of the target polynucleotide. This portion or region in the target polynucleotide can also be referred to as the binding region. The length of this portion or region in the oligonucleotide, polynucleotide, polynucleotide analogue or morpholino and / or in the target polynucleotide is typically at least 10 nucleotides or nucleobases, such as at least 15 nucleotides or nucleobases, at least 20 nucleotides or nucleobases, at least 25 nucleotides or nucleobases, at least 30 nucleotides or nucleobases, at least 40 nucleotides or nucleobases or at least 50 nucleotides or nucleobases. Thus, compared to this portion or region in the target polynucleotide, this region or portion of the oligonucleotide, polynucleotide, polynucleotide analogue or morpholino may have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more mismatches across a region of 5, 10, 15, 20, 21, 22, 30, 40 or 50 nucleotides or nucleobases.
[0158] The length of the portion or region is typically 50 nucleotides / nucleobases or fewer, such as 40 nucleotides / nucleobases or fewer, 30 nucleotides / nucleobases or fewer, 20 nucleotides / nucleobases or fewer, 10 nucleotides / nucleobases or fewer, or 5 nucleotides / nucleobases or fewer. The length of the portion or region can be 5 to 50 nucleotides / nucleobases, such as 10 to 40 nucleotides / nucleobases or 20 to 30 nucleotides / nucleobases. The length of the portion or region is preferably 25 nucleotides / nucleobases.
[0159] The term "antibody" includes whole antibodies. Naturally occurring antibodies typically comprise a tetramer, which usually consists of at least two heavy (H) chains and at least two light (L) chains. Each heavy chain consists of a heavy chain variable region (abbreviated herein as VH) and a heavy chain constant region, which typically consists of three domains (CH1, CH2, and CH3). The heavy chain can be any isotype, including IgG (IgG1, IgG2, IgG3, and IgG4 subtypes), IgA (IgA1 and IgA2 subtypes), IgM, and IgE. Each light chain consists of a light chain variable region (abbreviated herein as VL) and a light chain constant region (CL). Light chains include kappa (κ) chains and lambda (λ) chains. The heavy and light chain variable regions typically are responsible for antigen recognition, while the heavy and light chain constant regions can mediate binding of the immunoglobulin to host tissues or factors, including various cells of the immune system (e.g., effector cells) and the first component of the classical complement system (Clq). The VH and VL regions can be further subdivided into hypervariable regions called complementarity determining regions (CDRs), which are interspersed with more conserved regions called framework regions (FRs). Each VH and VL is composed of three CDRs and four FRs, arranged in the following order from the amino terminus to the carboxyl terminus: FR1, CDR1, FR2, CDR2, FR3, CDR3, FR4. The variable regions of the heavy and light chains contain binding domains that interact with antigens.
[0160] The term "functional fragment" refers to a fragment of a full antibody that retains the ability to specifically bind to a given molecule or target analyte. Such fragments include Fab fragments, Fab' fragments, monovalent fragments consisting of the VL, VH, CL, and CH1 domains; F(ab')2 fragments, divalent fragments containing two Fab fragments linked by a disulfide bridge at the hinge region; Fd fragments, which consist of the VH domain and the CH1 domain; Fv fragments, which consist of the VL and VH domains of a single arm of an antibody; dAb fragments (Ward et al., 1989 Nature 341:544-546), which consist of the VH domain; and isolated complementarity-determining regions (CDRs). In addition, although the two domains VL and VH of an Fv fragment are encoded by separate genes, these two domains can be joined using recombinant methods by a synthetic linker that enables the two domains to form a single protein chain in which the VL region and the VH region pair to form a monovalent molecule (referred to as single-chain Fv (scFv); see, for example, Bird et al., 1988 Science 242:423-426; and Huston et al., 1988 Proc. Natl. Acad. Sci. 85:5879-5883). Such single-chain antibodies are also intended to be encompassed within the term "functional fragment" of an antibody. These antibody fragments are obtained using conventional techniques known to those of skill in the art and are screened for utility in the same manner as full antibodies.
[0161] An aptamer is a small molecule that binds to one or more molecules or target analytes. Aptamers can be produced using SELEX (Stoltenburg, R. et al., (2007), Biomolecular Engineering 24, pp. 381-403; Tuerk, C. et al., Science 249, pp. 505-510; Bock, L.C. et al., (1992), Nature 355, pp. 564-566) or NON-SELEX (Berezovski, M. et al. (2006), Journal of the American Chemical Society 128, pp. l410-1411). An aptamer may be able to bind to or link two or more molecules or target analytes. Aptamers that bind to more than one analyte member can be produced using Toggle SELEX (White, R. et al., (2001), Molecular Therapy 4, pp. 567-573).
[0162] The aptamer is preferably a peptide aptamer or an oligonucleotide aptamer. The peptide aptamer can comprise any amino acid. The amino acid can be any of those discussed above. The oligonucleotide aptamer can comprise any nucleotide. The nucleotide can be any of those discussed above. The aptamer can be of any length. The length of the aptamer is typically at least 15 amino acids or nucleotides, such as at least 20, at least 25, at least 30, or at least 35 amino acids or nucleotides. The length of the aptamer is preferably about 15 to about 50, about 20 to about 40, or about 25 to about 30 amino acids or nucleotides.
[0163] The functional binding moiety is preferably capable of binding to or linking to a polynucleotide binding protein, a pore monomer, an aptamer, a cyclic protein, or a DNA origami structure. Linking these structures to a pore monomer or a pore according to the present invention improves its ability to characterize a target analyte, as discussed in more detail below.
[0164] The functional binding moiety is preferably capable of binding to or linking to a polynucleotide binding protein, a polypeptide binding protein, a pore monomer, an aptamer, a cyclic protein, or a DNA origami structure. Linking these structures to a pore monomer or a pore according to the present invention improves its ability to characterize a target analyte, as discussed in more detail below.
[0165] The functional binding moiety may be capable of binding or linking to a polynucleotide binding protein. This forms a modular sequencing system that can be used in the sequencing methods of the present invention. Polynucleotide binding proteins are polymerases, exonucleases, helicases, and topoisomerases, such as gyrase. Suitable polynucleotide binding proteins include, but are not limited to, exonuclease I from Escherichia coli, exonuclease III from Escherichia coli, RecJ from Thermus thermophilus, and bacteriophage λ exonuclease, TatD exonuclease, and variants thereof. Three subunits containing the RecJ sequence from Thermus thermophilus or variants thereof interact to form a trimeric exonuclease. The polymerase can be 3173 DNA polymerase (which is commercially available from Corporation), SD polymerase (commercially available from ), or variants thereof. The polynucleotide binding protein can be Phi29 DNA polymerase or variants thereof. The topoisomerase is preferably a member of any of the partial classification (EC) groups 5.99.1.2 and 5.99.1.3.
[0166] The polynucleotide-binding protein is most preferably derived from a helicase, such as Hel308 Mbu, Hel308 Csy, Hel308 Tga, Hel308 Mhu, TraI Eco, XPD Mbu, or variants thereof. Any helicase can be used in the present invention. The helicase can be or be derived from a Hel308 helicase, a RecD helicase, such as a TraI helicase or a TrwC helicase, an XPD helicase, or a Dda helicase. The helicase can be any one of a helicase, a modified helicase, or a helicase construct, as disclosed in WO 2013 / 057495; WO 2013 / 098562; WO 2013098561; WO 2014 / 013260; WO 2014 / 013259; WO 2014 / 013262, and WO 2015 / 055981. All of these are incorporated herein by reference in their entirety.
[0167] The functional binding portion may be capable of binding or linking to a polypeptide-binding protein. This forms a modular sequencing system that can be used in the sequencing methods of the present invention. Polypeptide-binding proteins are known in the art. The polypeptide-binding protein can be an NTP-driven unfoldase. The unfoldase can be driven by any NTP. Suitable nucleosides that can form the basis of NTP will be discussed below with reference to polynucleotides. The NTP unfoldase can be selected from proteasome ATPases, AAA proteases, AAA+ enzymes, membrane fusion proteins, such as NSF (N-ethylmaleimide-sensitive fusion protein) / Sac18p (N-ethylmaleimide-sensitive fusion protein homolog in yeast) and p97 / VCP / Cdc48p (97-kDa valosin-containing protein), Pex 1p and Pex6p (peroxisome ATPases), Katanin and SKD1 (Vps4p homolog in mice) / Vps4p (vacuolar protein sorting 4 homolog in yeast), Dynein (motor protein), DNA replication proteins, such as ORC (origin recognition complex), Cdc6 (cell division control protein 6), MCM (minichromosome maintenance protein), DnaA, and RFC (replication factor C) / clamp-loader, RuvB (holliday junction ATP-dependent DNA helicase RuvB, EC = 3.6.4.12), and TIP49a / TIP49 and TIP49b / TIP48 (eukaryotic RuvB-like proteins). The NTP unfoldase can be an ATPase (AAA+) enzyme associated with a variety of cellular activities. The AAA+ enzyme can be ClpX, ClpAP, ClpXP, ClpCP, HslVU, or Lon. The polypeptide-binding protein is preferably ClpX. The proteins listed in this paragraph are defined in WO 2013 / 123379, which is incorporated herein by reference in its entirety.
[0168] Aptamers are discussed above.
[0169] Examples of annulins are known in the art. They include, but are not limited to, pentraxins, GroES, SP1, any pore, MspA, aHL, CsgG, lysin, InvG, GspD, leukotoxin, FraC, aerolysin, NetB, and any of those discussed above, as well as hexameric enzymes such as hexameric helicases or hexameric unfoldases.
[0170] The pore monomer can be any of the pore monomers discussed above. The functional binding moiety can bind to the pore monomer in the pore monomer conjugate. This can stabilize the cis - loop in the pore monomer and reduce the signal - to - noise ratio (SNR) of the pore complex formed by the pore monomers. In such embodiments, the pore monomer is not only connected to the partner molecule by the functional binding moiety. In other words, the functional binding moiety is not the only connection between the pore monomer and the partner molecule. In such embodiments, based on the affinity of the partner molecule for the pore monomer, there is another connection or connections between the pore monomer and the partner molecule, preferably a covalent attachment or covalent link, such as one of the reactive groups discussed above and / or one of the linkers discussed above. If the pore monomer is from or derived from CsgG, the partner molecule is preferably not the CsgF peptide.
[0171] The DNA origami structure is preferably a DNA origami pore. Such pores are known in the art, e.g., Langecker et al., Science, 2012; 338:932 - 936).
[0172] The partner molecule is linked (preferably covalently) to the functional binding moiety. The partner molecule can be linked (preferably covalently) to the functional binding moiety using any known method, including using any of the reactive groups discussed above or alternative linking methods discussed below.
[0173] The partner molecule is preferably linked or covalently attached to the functional binding moiety via a linker. The linker can be any of the linkers discussed above or below. The linker is preferably a polypeptide linker or a polynucleotide linker. Polypeptides and polynucleotides are defined above.
[0174] If the chaperone molecule is a polypeptide, such as a CsgA polypeptide, it can be genetically fused to a polynucleotide-binding protein, a pore monomer, a cyclic protein, optionally via a linker. The chaperone polypeptide and the polynucleotide-binding protein, pore monomer, or cyclic protein (with or without a linker) can be expressed as a single polypeptide or protein construct. If the chaperone molecule is a polypeptide, such as a CsgA polypeptide, it can be genetically fused to a polynucleotide-binding protein, a polypeptide-binding protein, a pore monomer, a cyclic protein, optionally via a linker. The chaperone polypeptide and the polynucleotide-binding protein, polypeptide-binding protein, pore monomer, or cyclic protein (with or without a linker) can be expressed as a single polypeptide or protein construct. Any of the linkers discussed above can be used.
[0175] The linker is preferably an amino acid or peptide linker. Suitable amino acid linkers, such as peptide linkers, are known in the art. Flexible peptide linkers are stretches of 2 to 20, such as 4, 6, 8, 10, or 16 serine and / or glycine amino acids. More flexible linkers include (SG) 1 、(SG) 2 、(SG) 3 、(SG) 4 、(SG) 5 、(SG) 8 、(SG) 10 、(SG) 15 or (SG) 20 , where S is serine and G is glycine. Rigid linkers are stretches of 2 to 30, such as 4, 6, 8, 16, or 24 proline amino acids. More rigid linkers include (P) 12 , where P is proline. The linker is preferably SGS (SEQ ID NO:19), where S is serine and G is glycine. The linker is preferably SEQ ID NO:19, further comprising K or K(N3) at its C-terminus.
[0176] The pore monomer conjugates of the present invention are capable of forming pores or pore complexes. This can be measured using conventional methods, including any of those described in WO2016 / 034591, WO 2017 / 149316, WO 2017 / 149317, WO 2017 / 149318, WO 2018 / 211241, and WO2019 / 002893 (all incorporated herein by reference in their entirety) and in the examples.
[0177] CsgG pore monomer
[0178] The pore monomer is preferably a CsgG pore monomer. A CsgG pore monomer is a monomer capable of forming a CsgG pore. Such monomers are known in the art, particularly from WO 2019 / 002893 (incorporated herein by reference in its entirety). The CsgG pore preferably comprises one or more of (a) a cap region, (b) a constriction region, and (c) a transmembrane β-barrel region, such as (a), (b), (c), (a) and (b), (a) and (c), (b) and (c), or (a), (b), and (c). The CsgG pore monomer preferably comprises one or more of (a) a cap-forming region, (b) a constriction-forming region, and (c) a transmembrane β-barrel-forming region, such as (a), (b), (c), (a) and (b), (a) and (c), (b) and (c), or (a), (b), and (c). The residues of SEQ ID NO:3 forming these regions are defined as follows. The CsgG pore formed by the monomer can have any structure, but preferably has or comprises the structure of the wild-type CsgG pore ( Figure 6 ). The protein structure of CsgG defines a channel or pore that allows molecules and ions to translocate from one side of the membrane to the other.
[0179] As used interchangeably herein, "constriction structure", "pore mouth", "constriction structure region", "channel constriction structure", or "constriction structure site" refers to the pore diameter defined by the lumen surface of the pore or pore complex, which functions to allow ions and target analytes (such as but not limited to polynucleotides or individual nucleotides) but not other non-target analytes to pass through the pore or pore complex channel. The constriction is typically the narrowest pore within the pore or pore complex or within the channel defined by the pore or pore complex. The constriction can be used to restrict the passage of molecules through the pore. The size of the constriction is typically a key factor in determining the suitability of the pore or pore complex for analyte characterization. If the constriction is too small, the molecule to be characterized will not be able to pass through. However, in order to achieve the maximum effect on the ion flow through the channel, the constriction structure should not be too large. For example, the constriction structure should not be wider than the solvent-accessible lateral diameter of the target analyte. Ideally, the diameter of any constriction structure should be as close as possible to the lateral diameter of the analyte passing through.
[0180] The CsgG pore can be of any size, but preferably has the dimensions of the wild-type CsgG pore ( Figure 6 ). The CsgG pore preferably has an outer diameter of about to about , such as about to about or about to about The CsgG pore preferably has an outer diameter of about . The CsgG pore preferably has an outer diameter of about 80 to about such as about 90 to about or from about 95 to about total length. The CsgG pore preferably has a length of about total length. References to "total length" and "length" refer to the length of the pore or pore region when viewed from the side (see, e.g., Figure 6 the side view in
[0181] The cap region preferably has a length of from about 20 to about such as from about 30 to about or from about 35 to about The cap region preferably has a length of about The channel defined by the cap region preferably has a diameter of from about 45 to about such as a diameter of from about 55 to about or from about 60 to about opening. The channel defined by the cap region preferably has a diameter of about opening. The diameter of the channel defined by the cap region at its narrowest point is preferably about to about such as at its narrowest point, the diameter is about to about or about to about The diameter of the channel defined by the cap region at its narrowest point is preferably about
[0182] The constriction region preferably has a length of from about 5 to about such as from about 10 to about or from about 15 to about The constriction region preferably has a length of about The diameter of the channel defined by the constriction region at its narrowest point is preferably about to about such as at its narrowest point, the diameter is about to about about to about or about to about The diameter of the channel defined by the constriction region is preferably about or The diameter of the channel defined by the constriction region is preferably about The constricted diameter is preferably about to about such as a diameter of from about to about about to about or about to about The diameter of the constriction is preferably about or The diameter of the constriction is preferably about
[0183] The transmembrane β-barrel region preferably has about 20 to about such as about 30 to about or about 35 to about length. The transmembrane β-barrel preferably has about length. The diameter of the channel defined by the transmembrane β-barrel region at its narrowest point is preferably about to about such as the diameter at its narrowest point is about to about or about to about The diameter of the channel defined by the transmembrane β-barrel region at its narrowest point is preferably about
[0184] All of the above measurements are based on backbone-to-backbone measurements of the amino acids forming the different regions (as Figure 6 shown).
[0185] SEQ ID NO:3 shows the sequence of wild-type Escherichia coli CsgG as a mature protein. Residues 1 to 41, 64 to 131, 156 to 180, and 212 to 262 of SEQ ID NO:3 form the cap region. Residues 42 to 63 of SEQ ID NO:3 form the constriction region. Residues 132 to 155 and 181 to 211 of SEQ ID NO:3 form the transmembrane β-barrel region.
[0186] The CsgG pore monomer is preferably a variant of SEQ ID NO:3. The variant CsgG monomer can also be referred to as a modified CsgG pore monomer or a mutant CsgG pore monomer. Modifications or mutations in the variant include, but are not limited to, any one or more of the modifications disclosed herein or combinations of such modifications. The CsgG pore monomer can be a CsgG homomonomer. A CsgG homomonomer is a polypeptide having at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 99% sequence identity to wild-type Escherichia coli CsgG, as shown in SEQ ID NO:3. CsgG homologs are also referred to as polypeptides containing the PFAM domain PF03783, which is characteristic of CsgG-like proteins. A list of currently known CsgG homologs and CsgG architectures can be found at http: / / pfam.xfam.org / / family / PF03783 .
[0187] Over the entire length of the amino acid sequence of SEQ ID NO:3, the variant will preferably be at least about 40% homologous to the sequence based on amino acid identity. More preferably, based on amino acid identity with the amino acid sequence of SEQ ID NO:3 over the entire sequence, the variant can be at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90% and more preferably at least about 95%, 97% or 99% homologous. Over the entire length of the amino acid sequence of SEQ ID NO:3, the variant will preferably be at least about 40% identical to the sequence. More preferably, the variant can be at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90% and more preferably at least about 95%, 97% or 99% identical to SEQ ID NO:3 over the entire sequence.
[0188] Sequence identity can also relate to fragments or portions of the CsgG pore monomer. Thus, a sequence may have less than 40% overall sequence homology or identity with SEQ ID NO:3, but the sequence of a particular region, domain, or subunit may share at least about 80%, 90%, or up to 99% sequence homology or identity with the corresponding region of SEQ ID NO:3. There may be at least about 80%, such as at least about 85%, 90%, or 95% homology or identity over an extension of 100 or more, such as 125, 150, 175, or 200 or more contiguous or adjacent amino acids. The CsgG pore monomer is preferably a variant of SEQ ID NO:3 that contains a sequence that is at least about 40% homologous to the cap regions (residues 1 to 41, 64 to 131, 156 to 180, and 212 to 262) of SEQ ID NO:3. More preferably, based on amino acid identity with residues 1 to 41, 64 to 131, 156 to 180, and 212 to 262 of SEQ ID NO:3, the variant may contain a sequence that is at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, and more preferably at least about 95%, 97%, or 99% homologous. The variant preferably contains a sequence that is at least about 40% identical to residues 1 to 41, 64 to 131, 156 to 180, and 212 to 262 of SEQ ID NO:3. More preferably, the variant may contain a sequence that is at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, and more preferably at least about 95%, 97%, or 99% identical to the residues 1 to 41, 64 to 131, 156 to 180, and 212 to 262 of SEQ ID NO:3. Homology and / or identity are typically measured over the entire length of the cap region.
[0189] The CsgG pore monomer is preferably a variant of SEQ ID NO:3, which contains a sequence that is at least about 40% homologous to the constriction region (residues 42 to 63) of SEQ ID NO:3. More preferably, based on the amino acid identity with residues 42 to 63 of SEQ ID NO:3, the variant can contain a sequence that is at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90% and more preferably at least about 95%, 97% or 99% homologous. The variant preferably contains a sequence that is at least about 40% identical to residues 42 to 63 of SEQ ID NO:3. More preferably, the variant can contain a sequence that is at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90% and more preferably at least about 95%, 97% or 99% identical to residues 42 to 63 of SEQ ID NO:3. Homology and / or identity are typically measured over the entire length of the constriction region.
[0190] The CsgG pore monomer is preferably a variant of SEQ ID NO:3, which contains a sequence that is at least about 40% homologous to the transmembrane β-barrel regions (residues 132 to 155 and 181 to 211) of SEQ ID NO:3. More preferably, based on the amino acid identity with residues 132 to 155 and 181 to 211 of SEQ ID NO:3, the variant can contain a sequence that is at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90% and more preferably at least about 95%, 97% or 99% homologous. The variant preferably contains a sequence that is at least about 40% identical to residues 132 to 155 and 181 to 211 of SEQ ID NO:3. More preferably, the variant can contain a sequence that is at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90% and more preferably at least about 95%, 97% or 99% identical to residues 132 to 155 and 181 to 211 of SEQ ID NO:3. Homology and / or identity are typically measured over the entire length of the transmembrane β-barrel region.
[0191] The CsgG pore monomer is highly conserved (as can be readily understood from Figures 45 to 47 of WO 2017 / 149317). In addition, based on the knowledge of the mutations associated with SEQ ID NO:3, the equivalent positions of mutations in CsgG pore monomers other than SEQ ID NO:3 can be determined.
[0192] Accordingly, references to mutant CsgG pore monomers that include variants of the sequence shown in SEQ ID NO:3 and their specific amino acid mutations, as set forth in the claims and elsewhere in the specification, also encompass mutant CsgG pore monomers that include variants of any of the sequences shown in SEQ ID NOs:68 to 88 of WO 2019 / 002893 (incorporated herein by reference in its entirety) and their corresponding amino acid mutations. The CsgG pore monomers can also be any of the sequences shown in CN 113773373 A, CN 113896776A, CN113912683 A, and CN 113754743 A or variants thereof. It will be further understood that the invention extends to other variant CsgG pore monomers that exhibit highly conserved regions not explicitly identified in this specification.
[0193] Homology can be determined using standard methods in the art. For example, the UWGCG package provides the BESTFIT program that can be used to calculate homology, such as using it on its default settings (Devereux et al. (1984) Nucleic Acids Research 12, pp. 387-395). The PILEUP and BLAST algorithms can be used to calculate homology or align sequences (such as identifying equivalent residues or corresponding sequences (usually on their default settings)), for example as described in Altschul S.F. (1993) J Mol Evol 36:290-300; Altschul, S.F. et al. (1990) J Mol Biol 215:403-10. Software for performing BLAST analysis is publicly available at the National Center for Biotechnology Information (http: / / www.ncbi.nlm.nih.gov / ).
[0194] SEQ ID NO:3 is the wild-type CsgG pore monomer, which is the transmembrane protein pore of CsgG from Escherichia coli strain K-12 substrain MC4100. Variants of SEQ ID NO:3 can include any substitutions present in another CsgG homolog. CsgG homologs are shown in SEQ ID NOs:68 to 88 of WO 2019 / 002893 (incorporated herein by reference in its entirety). The variants can include a combination of one or more of the substitutions present in SEQ ID NOs:68 to 88 of WO 2019 / 002893 (incorporated herein by reference in its entirety) compared to SEQ ID NO:3, including one or more substitutions, one or more conservative mutations, one or more deletions, or one or more insertion mutations, such as 1 to 10 amino acids, such as deletions or insertions of 2 to 8 or 3 to 6 amino acids.
[0195] The CsgG pore monomers in the pore monomer conjugates of the present invention generally retain the ability to form the same 3D structure as wild-type CsgG pore monomers, such as the same 3D structure as the CsgG pore monomer having the sequence of SEQ ID NO:3. The 3D structure of CsgG is known in the art and is disclosed, for example, in Goyal et al. (2014) Nature 516(7530):250-3. In addition to the mutations described herein, any number of mutations can be made in the wild-type CsgG sequence, provided that the CsgG pore monomer retains the improved properties conferred upon it by the mutations of the present invention.
[0196] Generally, the CsgG pore monomer will retain the ability to form a structure comprising five α-helices and five β-strands. Thus, it is contemplated that further mutations can be made in any of these regions in any CsgG pore monomer without affecting the ability of the monomer to form a pore that can translocate polynucleotides. It is also anticipated that one or more amino acid deletions can be made in any loop region connecting an α-helix and a β-strand and / or in the N-terminal and / or C-terminal regions of the CsgG pore monomer without affecting the ability of the monomer to form a pore that can translocate polynucleotides.
[0197] In addition to the amino acid substitutions discussed above, amino acid substitutions can also be made to the amino acid sequence of SEQ ID NO:3, for example up to 1, 2, 3, 4, 5, 10, 20, or 30 substitutions. Conservative substitutions replace an amino acid with another amino acid having a similar chemical structure, similar chemical properties, or similar side chain volume. The introduced amino acid can have similar polarity, hydrophilicity, hydrophobicity, basicity, acidity, neutrality, or charge to the amino acid it replaces. Alternatively, a conservative substitution can introduce another aromatic or aliphatic amino acid in place of a pre-existing aromatic or aliphatic amino acid. Conservative amino acid changes are well known in the art.
[0198] The CsgG pore monomer can be modified to introduce one or more cysteines, one or more hydrophobic amino acids, one or more charged amino acids, one or more non-natural amino acids, one or more polar amino acids, or one or more photoreactive amino acids. Any number and combination of such introductions can be made. The introduction is preferably by substitution or addition.
[0199] One or more amino acid residues of the amino acid sequence of SEQ ID NO:3 can additionally be deleted from the above polypeptide. Up to 1, 2, 3, 4, 5, 10, 20, or 30 or more residues can be deleted.
[0200] Variants can include fragments of SEQ ID NO:3. Such fragments retain pore-forming activity. The length of the fragment can be at least 50, at least 100, at least 150, at least 200, or at least 250 amino acids. Such fragments can be used to generate pores. The fragment preferably contains the transmembrane β-barrel region of SEQ ID NO:3, i.e., residues 132 to 155 and 181 to 211, or variants thereof as described above.
[0201] One or more amino acids can alternatively or additionally be added to the polypeptides described above. An extension can be provided at the amino terminus or carboxy terminus of the amino acid sequence of SEQ ID NO:3 or its polypeptide variant or fragment. The extension can be very short, e.g., having a length of 1 to 10 amino acids. Alternatively, the extension can be longer, e.g., up to 50 or 100 amino acids. A carrier protein can be fused to the amino acid sequence according to the invention. Other fusion proteins will be discussed in more detail below.
[0202] A variant of SEQ ID NO:3 is a polypeptide having an amino acid sequence different from the amino acid sequence of SEQ ID NO:3 and retaining its ability to form pores. Variants typically contain the pore-forming regions of SEQ ID NO:3. The pore-forming ability of CsgG containing a β-barrel is provided by the β-strands in the transmembrane β-barrel region of each monomer. Variants of SEQ ID NO:3 typically contain the regions in SEQ ID NO:3 that form β-strands, i.e., residues 132 to 155 and 181 to 211, or variants thereof as described above. One or more modifications can be made to the regions of SEQ ID NO:3 that form β-strands, provided that the resulting variant retains its ability to form pores.
[0203] One or more modifications in the CsgG pore monomer preferably improve the ability of the pore complex containing the pore monomer to characterize an analyte. For example, the modification / mutation / substitution is expected to alter the number, size, shape, placement, or orientation of constrictions within the channel of the pore monomer conjugate from the invention. The CsgG pore monomer or a variant of SEQ ID NO:3 can have any of the specific modifications or substitutions disclosed in WO 2016 / 034591, WO 2017 / 149316, WO 2017 / 149317, WO 2017 / 149318, WO 2018 / 211241, and WO 2019 / 002893 (all incorporated herein by reference in their entirety).
[0204] Modifications or substitutions in SEQ ID NO:3 include, but are not limited to, one or more of the following, such as 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, or all:
[0205] (a) Substitutions at position Y51, such as Y51I, Y51L, Y51A, Y51V, Y51T, Y51S, Y51Q or Y51N;
[0206] (b) Substitutions at position N55, such as N55I, N55L, N55A, N55V, N55T, N55S or N55Q;
[0207] (c) Substitutions at position F56, such as F56I, F56L, F56A, F56V, F56T, F56S, F56Q or F56N;
[0208] (d) Substitutions at position L90, such as L90N, L90D, L90E, L90R or L90K;
[0209] (e) Substitutions at position N91, such as N91D, N91E, N91R or N91K;
[0210] (f) Substitutions at position K94, such as K94R, K94F, K94Y, K94Q, K94W, K94L, K94S or K94N;
[0211] (g) Substitutions at position R192, such as R192Q, R192F, R192S, R192D or R192T; and
[0212] (i) Substitutions at position C215, such as C215T, C215S, C215I, C215L, C215A, C215V or C215G.
[0213] Variants of SEQ ID NO:3 preferably contain F56Q. Variants of SEQ ID NO:3 may further contain deletions at one or more positions, such as deletions of T104 - N109, deletions of F193 - L199 or deletions of F195 - L199.
[0214] Any number of CsgG pore monomers in the pore or pore complex of the present invention, such as 6, 7, 8, 9 or 10, may be variants of SEQ ID NO:3. All six to ten monomers in the pore or pore complex are preferably variants of SEQ ID NO:3. The variants in the pore complex may be the same or different. The variants are preferably the same in each pore monomer conjugate in the pore complex of the present invention.
[0215] As discussed above, the chaperone molecule can be covalently linked to the pore monomer. The chaperone molecule is preferably linked to one or more residues in the CsgG pore monomer corresponding to one or more of positions 1 to 9 in SEQ ID NO:3. Any amino acid at these positions can be modified as discussed above, or replaced by an amino acid linked to a reactive group as discussed above, either by addition or substitution.
[0216] The pore monomer conjugate comprising the CsgG pore monomer preferably further comprises a CsgF peptide. Such peptides are described in WO2016 / 034591, WO 2017 / 149316, WO 2017 / 149317, WO 2019 / 002893, WO 2017 / 149318, WO2018 / 211241 and WO 2019 / 002893 (all incorporated herein by reference in their entirety).
[0217] CsgA polypeptide
[0218] The chaperone molecule preferably comprises a CsgA polypeptide. CsgA has a natural affinity for the CsgG pore monomer, and thus it is very simple to design a CsgA polypeptide that binds to any of the CsgG pore monomers discussed above. Wild-type CsgA has a Kd of 23.8 x 10 -6 M (Yan, Z., Yin, M., Chen, J. et al., Assembly and substraterecognition of curlibiogenesis system. Nat Commun 11, 241 (2020)).
[0219] If the pore monomer is a CsgG pore monomer and the chaperone molecule is a CsgA polypeptide, the CsgA polypeptide preferably binds to or is linked to (preferably covalently linked to) the cap region of the CsgG pore monomer. The cap region of CsgG is defined above.
[0220] The CsgA polypeptide is a polypeptide derived from or derived from CsgA. The wild-type Escherichia coli CsgA sequence is shown in SEQ ID NO:5, and the same sequence without the signal peptide is shown in SEQ ID NO:6.
[0221] The CsgA polypeptide can have any length. For example, the CsgA polypeptide can have a length of at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54 or 55 amino acids. The CsgA polypeptide can be 5 to 150 amino acids, such as 7 to 100 amino acids, 10 to 75 amino acids or about 10 to 60 amino acids.
[0222] The CsgA polypeptide is preferably a fragment of SEQ ID NO:5 or SEQ ID NO:6 that has an affinity for, binds to, or attaches to a pore monomer (preferably the CsgG pore monomer). The CsgA polypeptide preferably comprises or consists of at least 5 consecutive or contiguous amino acids from SEQ ID NO:5 or SEQ ID NO:6, such as at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 30, at least 40, at least 50 or at least 100 consecutive or contiguous amino acids from SEQ ID NO:5 or SEQ ID NO:6. The fragment preferably comprises amino acids 21 to 30, 21 to 31, 21 to 32, 21 to 33, 21 to 34 or 21 to 35 from SEQ ID NO:5 or amino acids 1 to 10, 1 to 11, 1 to 12, 1 to 13, 1 to 14 or 1 to 15 from SEQ ID NO:6. The CsgA polypeptide preferably comprises or consists of SEQ ID NO:7, 9, 11, 13, 15 or 17.
[0223] The CsgA polypeptide is preferably a variant of any of the CsgA sequences discussed above, including any fragments and SEQ ID NO:7, 9, 11, 13, 15 or 17. Over the entire length of the amino acid sequence of the CsgA fragment or SEQ ID NO:7, 9, 11, 13, 15 or 17, the variant will preferably be at least about 40% homologous to the sequence based on amino acid identity. More preferably, based on amino acid identity over the entire sequence with the amino acid sequence of the CsgA fragment or SEQ ID NO:7, 9, 11, 13, 15 or 17, the variant can be at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90% and more preferably at least about 95%, 97% or 99% homologous. Over the entire length of the amino acid sequence of the CsgA fragment or SEQ ID NO:7, 9, 11, 13, 15 or 17, the variant will preferably be at least about 40% identical to the sequence. More preferably, the variant can be at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90% and more preferably at least about 95%, 97% or 99% identical over the entire sequence with the CsgA fragment or SEQ ID NO:7, 9, 11, 13, 15 or 17.
[0224] The amino acid at any position in SEQ ID NO:7, 9, 11, 13, 15 or 17 is preferably replaced with K and / or modified to include a reactive group that attaches the CsgA polypeptide to a pore monomer (preferably a CsgG pore monomer). SEQ ID NO:7, 9, 11, 13, 15 or 17 preferably further includes a K at its N-terminus (i.e., before the G at position 1), which is modified to include a reactive group that attaches the CsgA polypeptide to a pore monomer (preferably a CsgG pore monomer). The reactive group in any of these modified versions of SEQ ID NO:7, 9, 11, 13, 15 or 17 can be any of those discussed above. The reactive group is preferably a sulfonyl group.
[0225] The amino acid at any of positions 1 to 10 in SEQ ID NO:7, 9, 11, 13, 15 or 17 (such as G at position 1, V at position 2, V at position 3, P at position 4, Q at position 5, Y at position 6, G at position 7, G at position 8, G at position 9 or G at position 10) is more preferably replaced with K and / or modified to include a reactive group that attaches the CsgA polypeptide to a pore monomer (preferably a CsgG pore monomer).
[0226] The amino acid at any one of positions 1 to 8 of SEQ ID NO: 7, 9, 11, 13, 15 or 17 (such as G at position 1, V at position 2, V at position 3, P at position 4, Q at position 5, Y at position 6, G at position 7 or G at position 8) is more preferably replaced by K and / or modified to include a reactive group that links the CsgA polypeptide to a pore monomer (preferably a CsgG pore monomer).
[0227] The amino acid at positions 6, 7 or 8 of SEQ ID NO: 7, 9, 11, 13, 15 or 17 (such as Y at position 6, G at position 7 or G at position 8) is more preferably replaced by K and / or modified to include a reactive group that links the CsgA polypeptide to a pore monomer (preferably a CsgG pore monomer). The amino acid at position 8 of SEQ ID NO: 7, 9, 11, 13, 15 or 17 is more preferably replaced by K and / or modified to include a reactive group that links the CsgA polypeptide to a pore monomer (preferably a CsgG pore monomer). The reactive group can be any of those discussed above. The reactive group is preferably a sulfonyl group.
[0228] The amino acid at any one of positions 1 to 7 of SEQ ID NO: 7, 9, 11, 13, 15 or 17 (such as G at position 1, V at position 2, V at position 3, P at position 4, Q at position 5, Y at position 6 or G at position 7) is more preferably replaced by K and / or modified to include a reactive group that links the CsgA polypeptide to a pore monomer (preferably a CsgG pore monomer). The amino acid at positions 6 or 7 of SEQ ID NO: 7, 9, 11, 13, 15 or 17 (such as Y at position 6 or G at position 7) is more preferably replaced by K and / or modified to include a reactive group that links the CsgA polypeptide to a pore monomer (preferably a CsgG pore monomer). The amino acid at position 7 of SEQ ID NO: 7, 9, 11, 13, 15 or 17 is more preferably replaced by K and / or modified to include a reactive group that links the CsgA polypeptide to a pore monomer (preferably a CsgG pore monomer). The reactive group can be any of those discussed above. The reactive group is preferably a sulfonyl group.
[0229] The CsgA polypeptide preferably comprises or consists of SEQ ID NO: 7, 9, 11, 13, 15 or 17, and comprises a substitution of G with K at position 7 (G7K). The CsgA polypeptide preferably comprises or consists of SEQ ID NO: 8, 10, 12, 14, 16 or 18. The G at position 7 of SEQ ID NO: 7, 9, 11, 13, 15 or 17 or the K at position 7 of SEQ ID NO: 8, 10, 12, 14, 16 or 18 is preferably modified to comprise a reactive group that attaches the CsgA polypeptide to a pore monomer (preferably a CsgG pore monomer). The reactive group can be any of those discussed above. The reactive group is preferably a sulfonyl group. The CsgA polypeptide comprising or consisting of SEQ ID NO: 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17 or 18 is preferably attached (preferably covalently) to a pore monomer (preferably a CsgG pore monomer) by reaction of a sulfonyl fluoride group at position 7 of SEQ ID NO: 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17 or 18 with the pore monomer.
[0230] The CsgA polypeptide preferably comprises or consists of SEQ ID NO: 7, 9, 11, 13, 15 or 17, and comprises a substitution of G with K at position 8 (G8K). The CsgA polypeptide preferably comprises or consists of SEQ ID NO: 57, 58, 59, 60, 61 or 62. The G at position 8 of SEQ ID NO: 7, 9, 11, 13, 15 or 17 or the K at position 8 of SEQ ID NO: 57, 58, 59, 60, 61 or 62 is preferably modified to comprise a reactive group that attaches the CsgA polypeptide to a pore monomer (preferably a CsgG pore monomer). The reactive group can be any of those discussed above. The reactive group is preferably a sulfonyl group. The CsgA polypeptide comprising or consisting of SEQ ID NO: 7, 9, 11, 13, 15, 17, 57, 58, 59, 60, 61 or 62 is preferably attached (preferably covalently) to a pore monomer (preferably a CsgG pore monomer) by reaction of a sulfonyl fluoride group at position 8 of SEQ ID NO: 7, 9, 11, 13, 15, 17, 57, 58, 59, 60, 61 or 62 with the pore monomer.
[0231] Any of the CsgA polypeptides discussed above preferably further comprises a linker. The linker is preferably attached to the C-terminus of the CsgA polypeptide, such as any of the fragments discussed above or SEQ ID NO: 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17 or 18. The linker is preferably attached to the C-terminus of the CsgA polypeptide, such as SEQ ID NO: 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 57, 58, 59, 60, 61 or 62. The linker is preferably an amino acid or peptide linker. Suitable amino acid linkers, such as peptide linkers, are known in the art and discussed above. The length, flexibility and hydrophilicity of the amino acid or peptide linker are typically designed to control the positioning of the functional binding moiety relative to the CsgA polypeptide. Flexible peptide linkers are stretches of 2 to 20, such as 4, 6, 8, 10 or 16 serine and / or glycine amino acids. More flexible joints include (SG) 1 、(SG) 2 、(SG) 3 、(SG) 4 、(SG) 5 、(SG) 8 、(SG) 10 、(SG) 15 or (SG) 20 , wherein S is serine and G is glycine. A rigid linker is a stretch of 2 to 30, such as 4, 6, 8, 16 or 24 proline amino acids. More rigid linkers include (P) 12 , wherein P is proline. The linker is preferably SGS (SEQ ID NO: 19), wherein S is serine and G is glycine. The linker is preferably SEQ ID NO: 19, further comprising K or K with an azide group, such as K(N3), at its C-terminus. The CsgA polypeptide is preferably linked or covalently linked to the functional binding moiety via a linker. The linker may be any of the linkers discussed above or below.
[0232] The CsgA polypeptide preferably comprises or consists of SEQ ID NO: 20, 22, 24, 26, 28 or 30. These sequences are SEQ ID NO: 7, 9, 11, 13, 15 and 17, wherein a preferred linker is attached to their C-terminus. The amino acid at any position in SEQ ID NO: 20, 22, 24, 26, 28 or 30 is preferably replaced with K and / or modified to include a reactive group that links the CsgA polypeptide to a pore monomer (preferably a CsgG pore monomer). SEQ ID NO: 20, 22, 24, 26, 28 or 30 preferably further comprises K at its N-terminus (i.e., before G at position 1), which K is modified to include a reactive group that links the CsgA polypeptide to a pore monomer (preferably a CsgG pore monomer). The reactive group in any of these modified versions of SEQ ID NO: 20, 22, 24, 26, 28 and 30 can be any of those discussed above. The reactive group is preferably a sulfonyl group.
[0233] The amino acid at any one of positions 1 to 10 in SEQ ID NO: 20, 22, 24, 26, 28 or 30 (such as G at position 1, V at position 2, V at position 3, P at position 4, Q at position 5, Y at position 6, G at position 7, G at position 8, G at position 9 or G at position 10) is more preferably replaced with K and / or modified to include a reactive group that links the CsgA polypeptide to a pore monomer (preferably a CsgG pore monomer).
[0234] The amino acid at any one of positions 1 to 8 in SEQ ID NO: 20, 22, 24, 26, 28 or 30 (such as G at position 1, V at position 2, V at position 3, P at position 4, Q at position 5, Y at position 6, G at position 7 or G at position 8) is more preferably replaced with K and / or modified to include a reactive group that links the CsgA polypeptide to a pore monomer (preferably a CsgG pore monomer).
[0235] The amino acid at position 6, 7, or 8 of SEQ ID NO: 20, 22, 24, 26, 28, or 30 (such as Y at position 6, G at position 7, or G at position 8) is more preferably replaced by K and / or modified to include a reactive group that will link the CsgA polypeptide to a pore monomer (preferably a CsgG pore monomer). The amino acid at position 8 of SEQ ID NO: 20, 22, 24, 26, 28, or 30 is more preferably replaced by K and / or modified to include a reactive group that will link the CsgA polypeptide to a pore monomer (preferably a CsgG pore monomer). The reactive group can be any of those discussed above. The reactive group is preferably a sulfonyl group.
[0236] The amino acid at any one of positions 1 to 7 of SEQ ID NO: 20, 22, 24, 26, 28, or 30 (such as G at position 1, V at position 2, V at position 3, P at position 4, Q at position 5, Y at position 6, or G at position 7) is more preferably replaced by K and / or modified to include a reactive group that will link the CsgA polypeptide to a pore monomer (preferably a CsgG pore monomer). The amino acid at position 6 or 7 of SEQ ID NO: 20, 22, 24, 26, 28, or 30 (such as Y at position 6 or G at position 7) is more preferably replaced by K and / or modified to include a reactive group that will link the CsgA polypeptide to a pore monomer (preferably a CsgG pore monomer). The reactive group can be any of those discussed above. The reactive group is preferably a sulfonyl group.
[0237] A CsgA polypeptide comprising SEQ ID NO: 20, 22, 24, 26, 27, 28, or 30 or any of the modified versions discussed above or consisting thereof is preferably reacted with a pore monomer through a sulfonyl fluoride group to form a sulfonyl group that is linked (preferably covalently linked) to the pore monomer (preferably a CsgG pore monomer). The K at the C-terminus of SEQ ID NO: 20, 22, 24, 26, 27, 28, or 30 or any of the modified versions discussed above is preferably modified with an azide group (such as N3).
[0238] The CsgA polypeptide preferably comprises or consists of SEQ ID NO: 21, 23, 25, 27, 29 or 31. These are SEQ ID NO: 8, 10, 12, 14, 16 or 18, to the C-terminus of which a preferred linker is attached. The K at position 7 in SEQ ID NO: 21, 23, 25, 27, 29 or 31 is preferably modified to comprise a reactive group that attaches the CsgA polypeptide to a pore monomer (preferably a CsgG pore monomer). The reactive group can be any of those discussed above. The reactive group is preferably a sulfonyl group. The CsgA polypeptide comprising or consisting of SEQ ID NO: 21, 23, 25, 27, 29 or 31 is preferably attached (preferably covalently) to a pore monomer (preferably a CsgG pore monomer) by reaction of a sulfonyl fluoride group at position 7 in SEQ ID NO: 21, 23, 25, 27, 29 or 31 with the pore monomer. The K at the C-terminus of SEQ ID NO: 21, 23, 25, 27, 29 or 31 is preferably modified with an azide group such as N3.
[0239] The CsgA polypeptide preferably comprises or consists of SEQ ID NO: 21, 23, 25, 27, 29, 31, 63, 64, 65, 66, 67 or 68. These are SEQ ID NO: 8, 10, 12, 14, 16, 18, 57, 58, 59, 60, 61 or 62, to the C-terminus of which a preferred linker is attached. The K at position 7 in SEQ ID NO: 21, 23, 25, 27, 29 or 31 or the K at position 8 in SEQ ID NO: 63, 64, 65, 66, 67 or 68 is preferably modified to comprise a reactive group that attaches the CsgA polypeptide to a pore monomer (preferably a CsgG pore monomer). The reactive group can be any of those discussed above. The reactive group is preferably a sulfonyl group. The CsgA polypeptide comprising or consisting of SEQ ID NO: 21, 23, 25, 27, 29, 31, 63, 64, 65, 66, 67 or 68 is preferably attached (preferably covalently) to a pore monomer (preferably a CsgG pore monomer) by reaction of a sulfonyl fluoride group at position 7 in SEQ ID NO: 21, 23, 25, 27, 29 or 31 or a sulfonyl fluoride group at position 8 in SEQ ID NO: 63, 64, 65, 66, 67 or 68 with the pore monomer. The K at the C-terminus of SEQ ID NO: 21, 23, 25, 27, 29, 31, 63, 64, 65, 66, 67 or 68 is preferably modified with an azide group such as N3.
[0240] The CsgA polypeptide preferably comprises SEQ ID NO:21, or consists of it. This is SEQ ID NO:8, to the C-terminus of which a preferred linker is attached. The K at position 7 of SEQ ID NO:21 is preferably modified to include a reactive group that attaches the CsgA polypeptide to a pore monomer (preferably a CsgG pore monomer). The reactive group can be any of those discussed above. The reactive group is preferably a sulfonyl group. The CsgA polypeptide comprising or consisting of SEQ ID NO:21 is preferably attached (preferably covalently) to a pore monomer (preferably a CsgG pore monomer) by reaction of the sulfonyl fluoride group at position 7 of SEQ ID NO:21 with the pore monomer. The K at the C-terminus of SEQ ID NO:21 is preferably modified with an azide group such as N3.
[0241] The CsgA polypeptide preferably comprises SEQ ID NO:63, or consists of it. This is SEQ ID NO:57, to the C-terminus of which a preferred linker is attached. The K at position 8 of SEQ ID NO:63 is preferably modified to include a reactive group that attaches the CsgA polypeptide to a pore monomer (preferably a CsgG pore monomer). The reactive group can be any of those discussed above. The reactive group is preferably a sulfonyl group. The CsgA polypeptide comprising or consisting of SEQ ID NO:63 is preferably attached (preferably covalently) to a pore monomer (preferably a CsgG pore monomer) by reaction of the sulfonyl fluoride group at position 8 of SEQ ID NO:63 with the pore monomer. The K at the C-terminus of SEQ ID NO:63 is preferably modified with an azide group such as N3.
[0242] Any of the CsgA polypeptides discussed above (especially those comprising a linker at the C-terminus) can further comprise a modified C-terminal group. For example, the CsgA polypeptide can comprise a -CONH2 group at its C-terminus.
[0243] The CsgA polypeptide preferably comprises, or consists of, SEQ ID NO: 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46 or 47. These sequences are shown in Table 1 or Table 3, where SO2F is a sulfonyl fluoride group and N3 is an azide. The CsgA polypeptide preferably comprises, or consists of, SEQ ID NO: 38, 39, 40, 41, 43, 44, 45, 46 or 47. The CsgA polypeptide comprising, or consisting of, SEQ ID NO: 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46 or 47 is preferably linked (preferably covalently) to a pore monomer (preferably a CsgG pore monomer) by reaction of the sulfonyl fluoride group with the pore monomer to form a sulfonyl group.
[0244] The CsgA polypeptide preferably comprises, or consists of, SEQ ID NO: 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47 or 48. These sequences are shown in Tables 1, 3 and 5, where SO2F is a sulfonyl fluoride group and N3 is an azide. SEQ ID NO: 48 also has a -CONH2 group at its C-terminus. The CsgA polypeptide preferably comprises, or consists of, SEQ ID NO: 38, 39, 40, 41, 43, 44, 45, 46, 47 or 48. The CsgA polypeptide comprising, or consisting of, SEQ ID NO: 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47 or 48 is preferably linked (preferably covalently) to a pore monomer (preferably a CsgG pore monomer) by reaction of the sulfonyl fluoride group with the pore monomer to form a sulfonyl group.
[0245] The CsgA polypeptide preferably comprises or consists of SEQ ID NO: 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49 or 50. These sequences are shown in Tables 1, 3 and 5, where SO2F is a sulfonyl fluoride group and N3 is an azide. SEQ ID NOs: 48 and 49 also have a -CONH2 group at their C-terminus. The CsgA polypeptide preferably comprises or consists of SEQ ID NO: 38, 39, 40, 41, 43, 44, 45, 46, 47, 48, 49 or 50. The CsgA polypeptide comprising or consisting of SEQ ID NO: 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49 or 50 is preferably linked (preferably covalently) to a pore monomer (preferably a CsgG pore monomer) by reaction of the sulfonyl fluoride group with the pore monomer to form a sulfonyl group.
[0246] The CsgA polypeptide preferably comprises or consists of SEQ ID NO: 49 or 50. These sequences are shown in Table 5, where SO2F is a sulfonyl fluoride group and N3 is an azide. SEQ ID NO: 49 also has a -CONH2 group at its C-terminus. The CsgA polypeptide comprising or consisting of SEQ ID NO: 49 or 50 is preferably linked (preferably covalently) to a pore monomer (preferably a CsgG pore monomer) by reaction of the sulfonyl fluoride group with the pore monomer to form a sulfonyl group.
[0247] The CsgA polypeptide preferably comprises or consists of SEQ ID NO: 69 or 70. These sequences are shown in the description of the sequence listing (where SO2F is a sulfonyl fluoride group) and are used in Example 1 and Figure 4 to produce the linked polymer. The CsgA polypeptide comprising and / or consisting of SEQ ID NO: 69 and / or 70 is preferably linked (preferably covalently) to a pore monomer (preferably a CsgG pore monomer) by reaction of the sulfonyl fluoride group with the pore monomer to form a sulfonyl group.
[0248] The CsgA polypeptide is preferably a variant of SEQ ID NO: 8, 10, 12, 14, 16, 18, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46 or 47. Over the entire length of the amino acid sequence of the CsgA fragment or SEQ ID NO: 8, 10, 12, 14, 16, 18, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46 or 47, the variant will preferably be at least about 40% homologous to the sequence based on amino acid identity. More preferably, based on amino acid identity over the entire sequence of the amino acid sequence of the CsgA fragment or SEQ ID NO: 8, 10, 12, 14, 16, 18, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46 or 47, the variant can be at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90% and more preferably at least about 95%, 97% or 99% homologous. Over the entire length of the amino acid sequence of the CsgA fragment or SEQ ID NO: 8, 10, 12, 14, 16, 18, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46 or 47, the variant will preferably be at least about 40% identical to the sequence. More preferably, over the entire sequence, the variant can be at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90% and more preferably at least about 95%, 97% or 99% identical to the CsgA fragment or SEQ ID NO: 8, 10, 12, 14, 16, 18, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46 or 47.
[0249] The CsgA polypeptide is preferably a variant of SEQ ID NO: 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69 or 70. Over the entire length of the amino acid sequence of the CsgA fragment or SEQ ID NO: 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69 or 70, based on amino acid identity, the variant will preferably be at least about 40% homologous to the sequence. More preferably, based on amino acid identity over the entire sequence of the amino acid sequence of the CsgA fragment or SEQ ID NO: 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69 or 70, the variant can be at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, more preferably at least about 95%, 97% or 99% homologous. Over the entire length of the amino acid sequence of the CsgA fragment or SEQ ID NO: 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69 or 70, the variant will preferably be at least about 40% identical to the sequence.More preferably, over the entire sequence, the variant can be at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90% and more preferably at least about 95%, 97% or 99% identical to the CsgA fragment or SEQ ID NO:8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69 or 70.
[0250] In any of the above CsgA polypeptides, including those based on SEQ ID NO:7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46 or 47 or variants thereof, the G at position 1 can be substituted or deleted. In any of the above CsgA polypeptides, including those based on SEQ ID NO:7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 39, 40, 42, 43, 44, 45, 46 or 47 or variants thereof, the Y at position 6 or 7 can be substituted, preferably by W.
[0251] In any of the above CsgA polypeptides, including those based on SEQ ID NO:7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69 or 70 or variants thereof, the G at position 1 can be substituted or deleted. In any of the above CsgA polypeptides, including those based on SEQ ID NO:7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 39, 40, 42, 43, 44, 45, 46, 47, 48, 49, 50, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69 or 70 or variants thereof, the Y at position 6 or 7 can be substituted, preferably by W. In any of the above CsgA polypeptides, including those based on SEQ ID NO:7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 39, 40, 42, 43, 44, 45, 46, 47, 48, 49, 50, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69 or 70 or variants thereof, the Y at position 6 can be substituted, preferably by W.
[0252] The pore monomer is preferably linked or covalently linked to the positions in the CsgA polypeptide discussed above. The pore monomer is typically linked or covalently linked to any one of positions 1 to 10 in the CsgA polypeptide. The pore monomer is typically linked or covalently linked to any one of positions 1 to 8 in the CsgA polypeptide. The pore monomer is typically linked or covalently linked to any one of positions 1 to 7 in the CsgA polypeptide. The pore monomer is typically linked or covalently linked to any one of positions 6, 7 and 8 in the CsgA polypeptide. The pore monomer is typically linked or covalently linked to position 7 or 8 in the CsgA polypeptide. The pore monomer is typically linked or covalently linked to position 7 in the CsgA polypeptide. The pore monomer is typically linked or covalently linked to position 8 in the CsgA polypeptide. The functional binding moiety is typically linked or covalently linked to the C-terminus of the CsgA polypeptide, or to any linker thereon.
[0253] The CsgA polypeptide can be a CsgA multimer comprising two or more linked CsgA polypeptides. The two or more linked CsgA polypeptides can be selected from any of those discussed above, including those from SEQ ID NO: 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69 and 70 and variants thereof. The CsgA multimer can comprise any number of two or more linked CsgA polypeptides, such as 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more or 10 or more CsgA polypeptides. The two or more CsgA polypeptides can be the same CsgA polypeptide. The two or more CsgA polypeptides can be different CsgA polypeptides.
[0254] The two or more CsgA polypeptides can be linked in any manner. The two or more CsgA polypeptides are preferably covalently linked. The two or more CsgA polypeptides are preferably covalently linked via a K at the C-terminus of one or more CsgA polypeptides. The two or more CsgA polypeptides can be directly covalently linked or covalently linked via any of the linkers disclosed herein. Those skilled in the art can link two or more CsgA polypeptides, for example, using a peptide bond, via an R group at any position in the CsgA polypeptide or using the reactive groups discussed above. The two or more CsgA polypeptides can be linked in any orientation. Two or more CsgA polypeptides can be N-terminally linked to C-terminally or C-terminally linked to N-terminally. The two or more CsgA polypeptides can be N-terminally linked to N-terminally. The two or more CsgA polypeptides can be C-terminally linked to C-terminally. One or more CsgA polypeptides can branch from one or more positions within the sequence of one or more CsgA polypeptides.
[0255] The multimer preferably comprises two CsgA polypeptides, wherein the C-terminus of one CsgA polypeptide is linked to the K at the C-terminus of the other CsgA polypeptide. This provides a linked multimer in which one CsgA polypeptide is N-to-C linked to the other C-to-N. Then, the N-terminus of each CsgA polypeptide can be used to link to a pore monomer, such as a CsgG pore monomer. The two CsgA polypeptides can be selected from SEQ ID NO:7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69 and 70 and variants thereof. The two CsgA polypeptides are preferably SEQ ID NO:69 and SEQ ID NO:70. The C-terminus of SEQ ID NO:70 is preferably linked to the K at the C-terminus of SEQ ID NO:69.
[0256] Pore monomer conjugate
[0257] A preferred pore monomer conjugate of the invention is a pore monomer conjugate comprising a CsgG pore monomer, a chaperone molecule, and a functional binding moiety,
[0258] wherein the CsgG pore monomer comprises a sequence that is at least about 40% homologous or identical to the amino acid sequence of SEQ ID NO:3 over the entire sequence,
[0259] wherein the chaperone molecule comprises SEQ ID NO:21, 23, 25, 27, 29, 31, 63, 64, 65, 66, 67 or 68,
[0260] wherein the K at position 7 of SEQ ID NO:21, 23, 25, 27, 29 or 31 is covalently linked to the CsgG pore monomer via a sulfonyl group, or the K at position 8 of SEQ ID NO:63, 64, 65, 66, 67 or 68 is covalently linked to the CsgG pore monomer via a sulfonyl group,
[0261] wherein the functional binding moiety comprises an oligonucleotide, polynucleotide, polynucleotide analogue or morpholino capable of specifically hybridizing to a target polynucleotide analyte, and
[0262] Wherein the functional binding moiety is covalently linked to the C-terminal lysine (K) of SEQ ID NO: 21, 23, 25, 27, 29, 31, 63, 64, 65, 66, 67 or 68. The C-terminal K is at position 14 in SEQ ID NO: 21, position 15 in SEQ ID NO: 23, position 16 in SEQ ID NO: 25, position 17 in SEQ ID NO: 27, position 18 in SEQ ID NO: 29, position 19 in SEQ ID NO: 31, position 14 in SEQ ID NO: 63, position 15 in SEQ ID NO: 64, position 16 in SEQ ID NO: 65, position 17 in SEQ ID NO: 66, position 18 in SEQ ID NO: 67 and position 19 in SEQ ID NO: 68.
[0263] Another preferred pore monomer conjugate of the present invention is a pore monomer conjugate comprising a CsgG pore monomer, a chaperone molecule and a functional binding moiety,
[0264] Wherein the CsgG pore monomer comprises a sequence that is at least about 40% homologous or identical to the amino acid sequence of SEQ ID NO: 3 over the entire sequence,
[0265] Wherein the chaperone molecule comprises SEQ ID NO: 21, 23, 25, 27, 29 or 31,
[0266] Wherein the K at position 7 of SEQ ID NO: 21, 23, 25, 27, 29 or 31 is covalently linked to the CsgG pore monomer via a sulfonyl group,
[0267] Wherein the functional binding moiety comprises an oligonucleotide, polynucleotide, polynucleotide analogue or morpholino capable of specifically hybridizing to a target polynucleotide analyte, and
[0268] Wherein the functional binding moiety is covalently linked to position 14 of SEQ ID NO: 21, 23, 25, 27, 29 or 31.
[0269] The CsgG pore monomer can have any percentage of homology or identity with SEQ ID NO: 3 described above. The CsgG pore monomer sequence preferably contains F56Q. The K at the C-terminus of SEQ ID NO: 21, 23, 25, 27, 29 or 31 is preferably modified with an azide group such as N3. The K at the C-terminus of SEQ ID NO: 21, 23, 25, 27, 29, 31, 63, 64, 65, 66, 67 or 68 is preferably modified with an azide group such as N3. The chaperone molecule preferably comprises or consists of SEQ ID NO: 21 or 63. The chaperone molecule preferably comprises or consists of SEQ ID NO: 21. The chaperone molecule preferably comprises or consists of SEQ ID NO: 63. The chaperone molecule preferably comprises or consists of SEQ ID NO: 40, 43, 44, 45, 46, 47, 48, 49 or 50. The chaperone molecule preferably comprises or consists of SEQ ID NO: 40, 43, 44, 45, 46 or 47. The chaperone molecule preferably comprises or consists of SEQ ID NO: 40. The chaperone molecule preferably contains a -CONH2 group at its C-terminus. The functional binding moiety preferably comprises a morpholino.
[0270] Construct
[0271] The present invention also provides a construct comprising two or more covalently attached pore monomer conjugates of the present invention. The construct may comprise 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, or 10 or more pore monomer conjugates of the present invention. The construct may comprise at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10 pore monomer conjugates of the present invention. Two or more pore monomer conjugates may be the same or different. Two or more pore monomer conjugates may differ based on one or more of the following: (a) the sequence of the CsgG pore monomer, (b) the sequence of the CsgA polypeptide, (c) the linker, (d) the attachment position on the CsgG pore monomer, and (e) the attachment position on the CsgA polypeptide. The pore monomer conjugates may differ based on: (a); (b); (c); (d); (e); (a) and (b); (a) and (c); (a) and (d); (a) and (e); (b) and (c); (b) and (d); (b) and (e); (c) and (d); (c) and (e); (d) and (e); (a), (b) and (c); (a), (b) and (d); (a), (b) and (e); (a), (c) and (d); (a), (c) and (e); (a), (d) and (e); (b), (c) and (d); (b), (c) and (e); (b), (d) and (e); (c), (d) and (e); (a), (b), (c) and (d); (a), (b), (c) and (e); (a), (b), (d) and (e); (a), (c), (d) and (e); (b), (c), (d) and (e); and (a), (b), (c), (d) and (e). Two or more pore monomer conjugates are preferably the same (i.e., identical).
[0272] The construct preferably comprises two pore monomer conjugates. Two or more pore monomer conjugates may be the same or different. Two or more pore monomer conjugates are preferably the same (i.e., identical).
[0273] The pore monomer conjugates may optionally be genetically fused via a linker or chemically fused, for example, via a chemical crosslinking agent. Methods for covalently attaching monomers are disclosed in WO 2017 / 149316, WO 2017 / 149317, and WO 2017 / 149318 (incorporated herein by reference in their entirety).
[0274] The linker is preferably an amino acid sequence and / or a chemical crosslinker. Suitable amino acid linkers, such as peptide linkers, are known in the art. The length, flexibility, and hydrophilicity of the amino acid or peptide linker are typically designed such that the pore monomer conjugates in the construct are in the correct orientation to form a pore complex. Flexible and rigid linkers useful in constructs were discussed above.
[0275] Suitable chemical crosslinkers are well known in the art. Suitable chemical crosslinkers include, but are not limited to, those containing the following functional groups: maleimide, active ester, succinimide, azide, alkyne (such as dibenzocyclooctynol (DIBO or DBCO), difluorocycloalkyne, and linear alkyne), phosphine (such as those for traceless and non-traceless Staudinger ligation), haloacetyl (such as iodoacetamide), phosgene-type reagent, sulfonyl chloride reagent, isothiocyanate, acyl halide, hydrazine, disulfide, vinyl sulfone, aziridine, and photoreactive reagent (such as aryl azide, diazirine).
[0276] The reaction between the amino acid and the functional group may be spontaneous, such as cysteine / maleimide, or may require an external reagent, such as Cu(I) for linking azide and linear alkyne.
[0277] The linker can contain any molecule that extends across the desired distance. The length of the linker can vary from one carbon (phosgene-type linker) to many angstroms. Examples of linker molecules include, but are not limited to, polyethylene glycol (PEG), polypeptide, polysaccharide, deoxyribonucleic acid (DNA), peptide nucleic acid (PNA), threose nucleic acid (TNA), glycerol nucleic acid (GNA), saturated and unsaturated hydrocarbons, polyamide. These linkers can be inert or reactive, especially they can be chemically cleaved at defined positions, or they can be modified with fluorophores or ligands themselves. The linker is preferably resistant to reducing agents (such as dithiothreitol (DTT)) after covalent connection.
[0278] Crosslinkers include 2,5-dioxopyrrolidin-1-yl 3-(pyridin-2-yldisulfanyl)propionate, 2,5-dioxopyrrolidin-1-yl 4-(pyridin-2-yldisulfanyl)butyrate, and 2,5-dioxopyrrolidin-1-yl 8-(pyridin-2-yldisulfanyl)octanoate, dimaleimide PEG 1k, dimaleimide PEG 3.4k, dimaleimide PEG 5k, dimaleimide PEG 10k, bis(maleimide)ethane (BMOE), bismaleimide hexane (BMH), 1,4-bismaleimide butane (BMB), 1,4-bismaleimide-2,3-dihydroxybutane (BMDB), BM[PEO]2 (1,8-bismaleimide diethylene glycol), BM[PEO]3 (1,11-bismaleimide triethylene glycol), tris[2-maleimidoethyl]amine (TMEA), DTME dithiobismaleimide ethane, dimaleimide PEG3, dimaleimide PEG11, DBCO-maleimide, DBCO-PEG4-maleimide, DBCO-PEG4-NH2, DBCO-PEG4-NHS, DBCO-NHS, DBCO-PEG-DBCO 2.8 kDa, DBCO-PEG-DBCO 4.0 kDa, DBCO-15 atoms-DBCO, DBCO-26 atoms-DBCO, DBCO-35 atoms-DBCO, DBCO-PEG4-S-S-PEG3-biotin, DBCO-S-S-PEG3-biotin, DBCO-S-S-PEG11-biotin, (succinimidyl 3-(2-pyridyldithio)propionate (SPDP) and maleimide-PEG(2 kDa)-maleimide (α,ω-bismaleimide poly(ethylene glycol)). One crosslinker is maleimide-propyl-SRDFWRS-(1,2-diaminoethane)-propyl-maleimide.
[0279] The linker is preferably resistant to dithiothreitol (DTT). Suitable linkers include, but are not limited to, iodoacetamide-based and maleimide-based linkers.
[0280] The pore monomer conjugates can be linked using two or more linkers, each linker comprising a hybridizable region and a group capable of forming a covalent bond. The hybridizable regions in the linkers hybridize and connect the pore monomer conjugates. Then, the linked CsgG pore monomer conjugates are coupled by forming a covalent bond between the groups. Any specific linker disclosed in WO 2010 / 086602 (incorporated herein by reference in its entirety) can be used according to the present invention.
[0281] The linker can be labeled. Suitable labels include, but are not limited to, fluorescent molecules (such as Cy3 or 555), radioisotopes (e.g.125 I, 35 S, 32 P), enzymes, antibodies, antigens, polynucleotides, and ligands (such as biotin). Such labels allow quantification of the amount of the linker. The label can also be a cleavable purification tag, such as biotin, or a specific sequence revealed in an identification method, such as a peptide that is not present in the protein itself but is released by trypsin digestion.
[0282] The method of linking the pore monomer conjugate is via a cysteine bond. This can be mediated by a bifunctional chemical crosslinker or by an amino acid linker having a cysteine residue present at the terminus.
[0283] Another method of linking is via 4-azidophenylalanine or a Faz bond. This can be mediated by a bifunctional chemical linker or by a polypeptide linker having a 4-azidophenylalanine or Faz residue present at the terminus. Other suitable linkers are discussed in more detail below.
[0284] Pore complex of the present invention
[0285] The terms "pore complex" or "complex pore" are used interchangeably herein and refer to an oligomeric pore complex comprising at least one pore monomer conjugate of the present invention (including, for example, one or more pore monomer conjugates, such as two or more pore monomer conjugates, three or more pore monomer conjugates, etc.). The pore complex of the present invention has the characteristics of a biological pore, i.e., it has a typical protein structure and defines a channel. When the pore complex is provided in an environment having a membrane component, a membrane, a cell, or an insulating layer, the pore complex will insert into the membrane or insulating layer and form a "transmembrane pore complex".
[0286] If the pore monomer is a CsgG pore monomer, the CsgG portion of the pore complex of the present invention (i.e., the portion formed by at least one CsgG pore monomer in at least one conjugate of the present invention) preferably has or comprises any of the structures and / or dimensions of the CsgG pore discussed above. The CsgG constriction in the pore complex of the present invention preferably has or comprises any of the above constriction diameters.
[0287] Chaperone molecules functionalize the pore complex by attaching or tethering a functional binding moiety to one or more pore monomers in the pore complex. The functionalized pores can be used in any of the applications discussed below with reference to the methods of the present invention.
[0288] The pore complex or transmembrane pore complex of the present invention comprises a pore complex having two constrictions, i.e., two channels are constricted in a manner such that one constriction does not interfere with the accuracy of the other constriction. If the pore monomer is a CsgG pore monomer, the pore complex preferably comprises one or more CsgF peptides. The pore complex preferably comprises CsgF peptides linked (preferably covalently) to each CsgG pore monomer in the pore complex. The pore complex may comprise any of the mutations, CsgG pore monomers or CsgF peptides described in WO 2016 / 034591, WO 2017 / 149316, WO 2017 / 149317, WO 2019 / 002893, WO 2017 / 149318, WO 2018 / 211241 and WO 2019 / 002893 (all incorporated herein by reference in their entirety). The pore complex or transmembrane pore complex of the present invention comprises a pore complex having one constriction.
[0289] The present invention provides a pore complex comprising at least one pore monomer conjugate of the present invention. The pore complex typically comprises at least 6, 7, 8, 9 or 10 pore monomer conjugates of the present invention. The pore complex preferably comprises 8 or 9 pore monomer conjugates of the present invention. The pore monomer conjugates are typically the same (i.e., identical).
[0290] The pore complex is preferably a homo-oligomer comprising 6 to 10, such as 6, 7, 8, 9 or 10 pore monomer conjugates of the present invention. The pore monomer conjugates are typically the same. The pore complex preferably comprises 8 or 9 identical pore monomer conjugates of the present invention. The pore monomer conjugate can be any of those discussed above.
[0291] The present invention provides a pore complex comprising at least one construct of the present invention. The pore complex typically comprises at least 1, 2, 3, 4 or 5 constructs of the present invention. The pore complex comprises sufficient pore monomers to form a pore. For example, an octameric pore can comprise (a) four constructs, each construct comprising two pore monomer conjugates, (b) two constructs, each construct comprising four pore monomer conjugates, (c) one construct comprising two pore monomer conjugates and six pore monomer conjugates that are not part of the construct, (d) three constructs comprising two pore monomer conjugates and two pore monomer conjugates that are not part of the construct, and (e) combinations thereof. For example, the same and additional possibilities are provided for non-polymeric pores. Those skilled in the art can envision other combinations of constructs and monomers. One or more constructs of the present invention can be used to form a pore complex for characterizing (such as sequencing) polynucleotides. The pore complex preferably comprises 4 constructs of the present invention, each construct comprising two pore monomer conjugates. The constructs are typically the same (i.e., identical).
[0292] The pore complex is preferably a homo - oligomer comprising 1 to 5, such as 1, 2, 3, 4, 5 constructs of the present invention. The constructs are typically identical (i.e., the same). The pore complex preferably comprises 4 identical constructs of the present invention, each construct comprising two pore monomer conjugates. The construct can be any of those discussed above.
[0293] The pore monomers in the pore complex preferably all have approximately the same length or have the same length. The barrels of the pore monomers of the present invention in the pore preferably have approximately the same length or have the same length. The length can be measured in terms of the number of amino acids and / or length units.
[0294] The pore complex of the present invention can be isolated, substantially isolated, purified or substantially purified. It is isolated or purified if it is completely free of any other components, such as lipids or other pores. It is substantially isolated if the pore complex is mixed with a carrier or diluent that does not interfere with its intended use. For example, if the pore complex is present in a form comprising less than 10%, less than 5%, less than 2% or less than 1% of other components (such as block copolymers, lipids or other pores), it is substantially isolated or substantially purified. Alternatively, the pore complex of the present invention can be present in a membrane. Suitable membranes are discussed below.
[0295] The pore complex of the present invention can exist as a single or solitary pore complex. Alternatively, the pore complex of the present invention can be present in a homologous or heterologous population of two or more pore complexes or pores. Other forms involving the pore complex of the present invention are discussed in more detail below.
[0296] Multimeric pore complex
[0297] The present invention also provides a pore multimer comprising two or more pores, wherein at least one of the pores is a pore complex of the present invention. The multimer can comprise any number of pores, such as 3, 4, 5, 6, 7 or 8 or more pores. Any number of the pores in the multimer, including all of them, can be the pore complex of the present invention.
[0298] The pore multimer can be a biporous complex comprising a first pore complex and a second pore or complex of the invention. The second pore or complex is typically derived from pores of the same type, such as CsgG. The second pore complex can be a complex of the invention. Both the first pore complex and the second pore complex are preferably pore complexes of the invention. In the biporous complex, the first pore complex can be attached to the second pore (complex) by hydrophobic interactions and / or by one or more disulfide bonds. One or more of the first pore complex and / or the second pore (complex) can be modified, such as 2, 3, 4, 5, 6, 8, 9, for example all monomers, to enhance such interactions. This can be achieved in any suitable manner. A specific method for forming a biporous pore derived from CsgG is described in WO 2019 / 002893 (incorporated herein by reference in its entirety).
[0299] The pore multimers of the invention can be isolated, substantially isolated, purified or substantially purified. These terms were defined above with reference to the pore complexes of the invention.
[0300] Membrane example
[0301] The invention also provides a pore complex or a pore multimer of the invention incorporated in a membrane. The invention also provides a membrane comprising a pore complex or a pore multimer of the invention. These products can be used directly for molecular sensing, such as analyte characterization and polynucleotide sequencing. Suitable membranes are discussed in more detail below.
[0302] Method for preparing a modified protein
[0303] Methods for introducing or substituting non-naturally occurring amino acids in pore monomers and peptides are also well known in the art and described in WO 2019 / 002893 (incorporated herein by reference in its entirety). Proteins can be modified to aid in their identification or purification, for example by adding a streptavidin tag or by adding a signal sequence to facilitate their secretion from cells in which the monomers are not naturally present with such a sequence. D-amino acids or mixtures of L-amino acids and D-amino acids can also be used to produce proteins. This is routine in the field of producing such proteins or peptides.
[0304] A pore monomer, a chaperone polypeptide or protein, a functional binding portion of a polypeptide or protein, a pore monomer conjugate, a construct, a pore complex or a pore multimer (i.e., any protein of the present invention) can be chemically modified. The protein can be chemically modified in any way and at any site. The protein can be chemically modified by attaching a molecule to one or more cysteines (cysteine ligation), attaching a molecule to one or more lysines, attaching a molecule to one or more unnatural amino acids, enzymatic modification of an epitope, or modification of a terminus. Suitable methods for performing such modifications are well known in the art. The protein can be chemically modified by attaching any molecule (e.g., a dye or a fluorophore).
[0305] The protein can be chemically modified with a molecular linker that promotes the interaction between the pore containing the monomer and a target nucleotide or target polynucleotide sequence. Suitable linkers, including cyclic molecules, cyclodextrins, substances capable of hybridization, DNA binders or chelators, peptides or peptide analogs, synthetic polymers, aromatic planar molecules, positively charged small molecules, or small molecules capable of hydrogen bonding, are described in WO 2019 / 002893 (incorporated herein by reference in its entirety). Any of the methods and linkers discussed above can be used to attach the molecular linker.
[0306] The protein can be attached to a polynucleotide-binding protein. This forms a modular sequencing system that can be used in the sequencing methods of the present invention. The polynucleotide-binding protein can be linked using a functional binding portion or by any other method. Polynucleotide-binding proteins are discussed above. Any method known in the art can be used to covalently attach the protein to the monomer. The monomer and the protein can be chemically fused or genetically fused. Gene fusions of monomers with polynucleotide-binding proteins are discussed in WO2010 / 004265 (incorporated herein by reference in its entirety). The polynucleotide-binding protein can be attached via a cysteine bond using any of the methods described above. This also applies to polypeptide-binding proteins.
[0307] The polynucleotide-binding protein can be attached directly to the protein via one or more linkers. The polynucleotide-binding protein or polypeptide-binding protein can be directly linked to the protein via one or more linkers. A hybridization linker as described in WO 2010 / 086602 (incorporated herein by reference in its entirety) can be used to link the molecule to the pore monomer. Alternatively, a peptide linker can be used. Suitable peptide linkers are discussed above.
[0308] Any of the proteins can be modified to aid in their identification or purification, for example by adding histidine residues (His-tag), aspartic acid residues (Asp-tag), streptavidin tag, flag tag, SUMO tag, GST tag or MBP tag, or by adding a signal sequence to facilitate their secretion from the cell in which the polypeptide is not naturally present with such a sequence. An alternative to introducing a genetic tag is to react a tagging chemical to a native or engineered position on the protein. An example of this is reacting a gel transfer reagent to a cysteine engineered on the outside of the protein. This has been shown to be a method for isolating hemolysin hetero-oligomers (Chem Biol. July 1997;4(7):497-505).
[0309] Any of the proteins can be labeled with a display marker. The display marker can be any suitable marker that allows the detection of the protein. Suitable markers include but are not limited to fluorescent molecules, radioisotopes (such as 125I, 35S), enzymes, antibodies, antigens, polynucleotides, and ligands (such as biotin).
[0310] The protein can also contain other non-specific modifications, provided that they do not interfere with the function of the protein. Many non-specific side-chain modifications are known in the art and can be made to the side chains of the protein. Such modifications include, for example, reductive alkylation of amino acids by reaction with an aldehyde followed by reduction with NaBH4, amidation with methyl acetimidate, or acylation with acetic anhydride.
[0311] Any protein can be produced using standard methods known in the art. The polynucleotide sequence encoding the protein can be derived and replicated using standard methods in the art. The polynucleotide sequence encoding the protein can be expressed in a bacterial host cell using standard techniques in the art. The protein can be produced in a cell by in situ expression of the polypeptide from a recombinant expression vector. The expression vector optionally carries an inducible promoter to control the expression of the polypeptide. These methods are described in Sambrook, J. and Russell, D. (2001). Molecular Cloning: A Laboratory Manual, 3rd ed. Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY.
[0312] The protein can be produced on a large scale after purification from the organism producing the protein by any protein liquid chromatography system or after recombinant expression. Typical protein liquid chromatography systems include FPLC, AKTA system, Bio-Cad system, Bio-Rad BioLogic system, and Gilson HPLC system.
[0313] Method for producing a pore monomer conjugate
[0314] The present invention provides methods for generating the pore monomer conjugates of the present invention. The methods include combining or contacting a pore monomer, a chaperone molecule, and a functional binding moiety under conditions for linking (preferably covalently linking) the functional binding moiety to the pore monomer via the chaperone molecule. These conditions are well known to those skilled in the art and are discussed in the examples. The methods are generally carried out in vitro as defined below.
[0315] When the functional binding moiety is combined with the pore monomer, the functional binding moiety is preferably covalently linked to the chaperone molecule. The method preferably includes linking (preferably covalently linking) the pore monomer to the chaperone molecule. Any of the linkers and reactive groups discussed above can be used.
[0316] The method preferably includes linking (preferably covalently linking) the functional binding moiety to the chaperone molecule and then linking (preferably covalently linking) the chaperone molecule to the pore monomer.
[0317] Any of the embodiments discussed above with reference to the pore monomer conjugates of the present invention are equally applicable to these methods.
[0318] Method for functionalizing a pore monomer
[0319] The present invention provides methods for linking a functional binding moiety to a pore monomer. The present invention provides methods for functionalizing a pore monomer. The methods include using a chaperone molecule having an affinity for the pore monomer as a linker. The functional binding moiety is linked (preferably covalently linked) to the pore monomer via the chaperone molecule.
[0320] Any of the embodiments discussed above with reference to the pore monomer conjugates of the present invention are equally applicable to these methods.
[0321] Method for generating pores
[0322] The present invention also provides methods for generating the pore complexes or pore multimers of the present invention.
[0323] The method can involve expressing a pore complex in a host cell. In particular, the method can include expressing at least one pore monomer conjugate of the present invention or a construct of the present invention and a sufficient amount of pore monomers or constructs to form a pore complex or pore multimer in the host cell and allowing the pore complex or pore multimer to form in the host cell. The sufficient amount of pore monomers or constructs is preferably a sufficient amount of pore monomer conjugates of the present invention or a sufficient amount of constructs of the present invention. The number of pore monomers, pore monomer conjugates or constructs required to form the pore complex of the present invention or the pore multimer of the present invention was discussed above. Suitable host cells and expression systems are known in the art and are discussed in the examples.
[0324] The method can involve forming a pore complex in a cell-free or in vitro environment. In particular, the method can include contacting at least one pore monomer conjugate of the present invention or a construct of the present invention with a sufficient amount of pore monomers or constructs in vitro and allowing the formation of a pore complex or pore multimer. The pore monomer conjugate or construct can be produced separately by in vitro transcription and translation (IVTT) and then incubated with a sufficient amount of pore monomers or constructs. The sufficient amount of pore monomers or constructs is preferably a sufficient amount of pore monomer conjugates of the present invention or a sufficient amount of constructs of the present invention. The number of pore monomers, pore monomer conjugates or constructs required to form the pore complex of the present invention or the pore multimer of the present invention was discussed above. The method can be carried out in an "in vitro system", which refers to a system that includes at least the components and environment necessary to carry out the method and utilizes biomolecules, organisms, cells (or parts of cells) outside their normal natural environment, allowing for a more detailed, convenient or efficient analysis than with whole organisms. The in vitro system can also include a suitable buffer composition provided in a test tube, in which the protein components for forming the complex have been added. Those skilled in the art know the options for providing such a system.
[0325] Some or all components of the pore complex or pore multimer can be labeled to facilitate purification. Purification can also be carried out when the components are not labeled. Methods known in the art (e.g., ion exchange, gel filtration, hydrophobic interaction column chromatography, etc.) can be used alone or in different combinations to purify the components of the pore.
[0326] The pore complex or pore multimer can be prepared before insertion into the membrane or after inserting the components into the membrane.
[0327] Methods for preparing the pores and complexes of the present invention and ways of labeling them are disclosed in WO 2016 / 034591, WO 2017 / 149316, WO 2017 / 149317 and WO 2017 / 149318, WO 2018 / 211241 and WO2019 / 002893 (all incorporated herein by reference in their entirety).
[0328] Method for characterizing an analyte
[0329] The present invention provides a method for determining the presence, absence, or one or more characteristics of a target analyte. The method involves contacting the target analyte with a pore complex or a pore polymer of the present invention such that the target analyte moves relative to (such as into or through) the pore complex or the pore polymer, and making one or more measurements as the target analyte moves relative to the pore complex or the pore polymer, and thereby determining the presence, absence, or one or more characteristics of the target analyte. The target analyte may also be referred to as a template analyte or an analyte of interest.
[0330] The pore complex or the pore polymer of the present invention can be any one of those described above.
[0331] The method is used to determine the presence, absence, or one or more characteristics of a target analyte. The method can be used to determine the presence, absence, or one or more characteristics of at least one target analyte. The method can involve determining the presence, absence, or one or more characteristics of two or more target analytes. The method can include determining the presence, absence, or one or more characteristics of any number of target analytes (such as 2, 5, 10, 15, 20, 30, 40, 50, 100 or more analytes). Any number of characteristics of one or more target analytes can be determined, such as 1, 2, 3, 4, 5, 10 or more characteristics.
[0332] Binding of a molecule within the channel of a pore complex or pore polymer, or near either opening of the channel, will affect the ion flow through the open channel of the pore complex or pore polymer, which is the essence of "molecular sensing". In a manner similar to nucleic acid sequencing applications, suitable measurement techniques can be used to measure changes in the open channel ion flow through changes in the current (e.g., WO 2000 / 28312 and D. Stoddart et al., Proc. Natl. Acad. Sci., 2010, 106, 7702-7 or WO2009 / 077734; all incorporated herein by reference in their entirety). The degree of reduction in ion flow measured by a reduction in current is related to the size of the obstacle within or near the pore. Thus, binding of a molecule of interest (also referred to as an "analyte") within or near the pore provides a detectable and measurable event, thus forming the basis of a "biosensor". Suitable molecules for nanopore sensing include nucleic acids; proteins; peptides; polysaccharides and small molecules (herein referring to low molecular weight (e.g., <900 Da or <500 Da) organic or inorganic compounds), such as drugs, toxins, cytokines, and contaminants. Detection of the presence of biomolecules can be applied to personalized drug development, medicine, diagnostics, life science research, environmental monitoring, and the security and / or defense industries.
[0333] A pore complex or pore polymer can be used as a molecular or biosensor. The target analyte molecule to be detected can bind to either face of the channel, or within the lumen of the channel itself. The location of binding can be determined by the size of the molecule to be sensed.
[0334] The target analyte preferably comprises or consists of: metal ions, inorganic salts, polymers, amino acids, peptides, polypeptides, proteins, nucleotides, oligonucleotides, polynucleotides, polynucleotide-polypeptide conjugates, monosaccharides, oligosaccharides, polysaccharides, dyes, bleaches, drugs, diagnostic agents, recreational drugs, explosives, toxic compounds, or environmental contaminants. The target analyte preferably comprises or consists of: polypeptides, proteins, oligonucleotides, polynucleotides, polynucleotide-polypeptide conjugates, oligosaccharides, or polysaccharides. The target analyte can comprise two or more different molecules, such as peptides and polypeptides. The target analyte can be a polynucleotide-polypeptide conjugate. The method can involve determining the presence, absence, or one or more characteristics of two or more target analytes of the same type (such as two or more proteins, two or more nucleotides, or two or more drugs). Alternatively, the method can involve determining the presence, absence, or one or more characteristics of two or more target analytes of different types (such as one or more proteins, one or more nucleotides, and one or more drugs).
[0335] The target analyte can be secreted from cells. Alternatively, the target analyte can be an analyte present inside the cells, such that the target analyte must be extracted from the cells before the method can be carried out. The target analyte can be obtained or extracted from any organism or microorganism. The target analyte can be obtained from a human or an animal, such as from urine, lymph fluid, saliva, mucus, semen, or amniotic fluid, or from whole blood, plasma, or serum. The target analyte can be obtained from plants (e.g., grains, legumes, fruits, or vegetables).
[0336] The pore complex or pore polymer can be modified via recombinant or chemical methods to increase the binding strength, binding location, or binding specificity of the molecule to be sensed. Typical modifications include adding a specific binding moiety that is structurally complementary to the molecule to be sensed. In the case where the analyte molecule contains nucleic acid, the binding moiety can include cyclodextrin or oligonucleotide; for small molecules, this can be a known complementary binding region, such as the antigen-binding portion of an antibody or non-antibody molecule, including a single-chain variable fragment (scFv) region or the antigen recognition domain from a T cell receptor (TCR); or for proteins, it can be a known ligand of the target protein. In this way, the pore complex or pore polymer can act as a molecular sensor for detecting the presence of suitable antigens (including epitopes) in a sample, which may include cell surface antigens, including markers of receptors, solid tumors, or blood cancer cells (such as lymphoma or leukemia), viral antigens, bacterial antigens, protozoan antigens, allergens, allergy-related molecules, albumin (e.g., human, rodent, or bovine), fluorescent molecules (including fluorescein), blood group antigens, small molecules, drugs, enzymes, the catalytic site of an enzyme or enzyme substrate, and transition state analogs of enzyme substrates. As described above, the modification can be achieved using known genetic engineering and recombinant DNA techniques. Any suitable orientation will depend on the nature of the molecule to be sensed, such as size, three-dimensional structure, and its biochemical properties. The selection of a suitable structure can utilize computational structure design. The determination and optimization of protein-protein interactions or protein-small molecule interactions can be studied using techniques such as those that use surface plasmon resonance to detect molecular interactions (BIAcore, Inc., Piscataway, NJ; see also www.biacore.com).
[0337] The target analyte preferably comprises or consists of an amino acid, peptide, polypeptide, or protein. The amino acid, peptide, polypeptide, or protein can be naturally occurring or non-naturally occurring. The polypeptide or protein can include synthetic or modified amino acids therein. Several different types of modifications to amino acids are known in the art. Suitable amino acids and their modifications are as described above. It should be understood that the target analyte can be modified by any method available in the art.
[0338] The target analyte preferably comprises a polypeptide. Any suitable polypeptide can be characterized. The polypeptide can be an unmodified protein or a portion thereof, or a naturally occurring polypeptide or a portion thereof. The target polypeptide can be secreted from a cell. Alternatively, the target polypeptide can be produced intracellularly such that it must be extracted from the cell for characterization. The polypeptide can comprise the cell expression product of a plasmid, e.g., a plasmid for cloning a protein according to the methods described in the following references: Sambrook et al., Molecular Cloning: A Laboratory Manual, 4th Edition, Cold Spring Harbor Press, Plainsview, New York (2012); and Ausubel et al., Current Protocols in Molecular Biology (Supplement 114), John Wiley & Sons, New York (2016).
[0339] The polypeptide can be provided as an impure mixture of one or more polypeptides and one or more impurities. The impurities can comprise truncated forms of the target polypeptide that are different from the "target polypeptide" used for characterization. For example, the target polypeptide can be a full-length protein and the impurities can comprise fractions of the protein. The impurities can also comprise proteins other than the target protein, e.g., they can be co-purified from a cell culture or obtained from a sample.
[0340] The polypeptide can comprise any combination of any amino acids, amino acid analogs, and modified amino acids (i.e., amino acid derivatives). The amino acids (and their derivatives, analogs, etc.) in the polypeptide can be distinguished by their physical size and charge. The amino acid / derivative / analog can be naturally occurring or synthetic. The polypeptide can comprise any naturally occurring amino acids.
[0341] The polypeptide can be modified. The methods of the present invention can be used to modify the polypeptide for detection. The methods can be used to characterize the modifications in the target polypeptide.
[0342] One or more of the amino acids / derivatives / analogs in the polypeptide can be modified. One or more of the amino acids / derivatives / analogs in the polypeptide can be post-translationally modified. Thus, the methods of the present invention can be used to detect the presence, absence, location, and number of post-translational modifications in the polypeptide. The methods can be used to characterize the extent to which the polypeptide has been post-translationally modified.
[0343] Any one or more post-translational modifications can be present in a polypeptide. Typical post-translational modifications include modification with a hydrophobic group, modification with a cofactor, addition of a chemical group, glycation (non-enzymatic attachment of sugars), biotinylation, and polyethylene glycolylation. Post-translational modifications can also be non-natural, such that they are chemical modifications performed in the laboratory for biotechnological or biomedical purposes. This can allow for the monitoring of the levels of laboratory-made peptides, polypeptides, or proteins compared to their natural counterparts.
[0344] Examples of post-translational modification with a hydrophobic group include myristoylation, attachment of myristic acid, a C 14 saturated acid; palmitoylation, attachment of palmitic acid, a C 16 saturated acid; prenylation or isoprenylation, attachment of an isoprenoid group; farnesylation, attachment of a farnesol group; geranylgeranylation, attachment of a geranylgeraniol group; and glycosylation, and formation of a glycosylphosphatidylinositol (GPI) anchor via an amide bond.
[0345] Examples of post-translational modification with a cofactor include lipoylation, attachment of lipoic acid (C 8 ) functional group; flavylation, attachment of a flavin moiety (e.g., flavin mononucleotide (FMN) or flavin adenine dinucleotide (FAD)); attachment of heme C, e.g., via a thioether bond to cysteine; phosphopantetheinylation, attachment of a 4'-phosphopantetheine group; and formation of a retinylidene Schiff base.
[0346] Examples of post-translational modification by addition of a chemical group include acylation, e.g., O-acylation (ester), N-acylation (amide), or S-acylation (thioester); acetylation, e.g., attachment of an acetyl group to the N-terminus or lysine; formylation; alkylation, addition of an alkyl group (such as methyl or ethyl); methylation, e.g., addition of a methyl group to lysine or arginine; amidation; butyrylation; γ-carboxylation; glycosylation, enzymatic attachment of a glycosyl group to, e.g., arginine, asparagine, cysteine, hydroxylysine, serine, threonine, tyrosine, or tryptophan; polysialylation, attachment of polysialic acid; malonylation; hydroxylation; iodination; bromination; citrullination; nucleotide addition, attachment of any nucleotide (such as any of those discussed above), ADP-ribosylation; oxidation; phosphorylation, attachment of a phosphate group to, e.g., serine, threonine, or tyrosine (O-linked) or histidine (N-linked); adenylation, attachment of an adenylyl moiety to, e.g., tyrosine (O-linked) or histidine or lysine (N-linked); propionylation; pyroglutamate formation; S-glutathionylation; sumoylation; S-nitrosylation; succinylation, attachment of a succinyl group to, e.g., lysine; selenylation, incorporation of selenium; and ubiquitination, addition of a ubiquitin subunit (N-linked).
[0347] The polypeptide can be labeled with a molecular marker. The molecular marker can be a modification of the polypeptide that facilitates the detection of the polypeptide in the methods of the present invention. For example, the label can be a modification of the polypeptide that alters the signal obtained when characterizing the conjugate. For example, the label may interfere with the ion flux through the nanopore. In such ways, the label can improve the sensitivity of the method.
[0348] The polypeptide can contain one or more crosslinking segments, e.g., C-C bridges. The polypeptide may not be crosslinked prior to characterization using the method.
[0349] The polypeptide can contain sulfur-containing amino acids and thus have the potential to form disulfide bonds. Typically, in such embodiments, the polypeptide is reduced using a reagent such as DTT (dithiothreitol) or TCEP (tris(2-carboxyethyl)phosphine) prior to characterization using the method.
[0350] The polypeptide can be a full-length protein or a naturally occurring polypeptide. The protein or naturally occurring polypeptide can be fragmented prior to conjugation with the polynucleotide. The protein or polypeptide can be fragmented chemically or enzymatically. The polypeptide or polypeptide fragment can be conjugated to form a longer target polypeptide.
[0351] The polypeptide can have any suitable length. The polypeptide preferably has a length of from about 2 to about 300 peptide units or amino acids. The polypeptide can have a length of from about 2 to about 100 peptide units, e.g., from about 2 to about 50 peptide units, e.g., from about 3 to about 50 peptide units, such as from about 5 to about 25 peptide units, e.g., from about 7 to about 16 peptide units, such as from about 9 to about 12 peptide units. "Peptide unit" can be interchanged with "amino acid".
[0352] One or more characteristics of the polypeptide are preferably selected from (i) the length of the polypeptide, (ii) the identity of the polypeptide, (iii) the sequence of the polypeptide, (iv) the secondary structure of the polypeptide, and (v) whether the polypeptide is modified. One or more characteristics can be the sequence of the polypeptide, or whether the polypeptide is modified, e.g., by one or more post-translational modifications. One or more characteristics are preferably the sequence of the polypeptide.
[0353] The polypeptide can be in a relaxed form. The polypeptide can be kept in a linearized form. Keeping the polypeptide in a linearized form can facilitate the characterization of the polypeptide on a residue-by-residue basis because "aggregation" of the polypeptide within the nanopore is prevented. Any suitable means can be used to keep the polypeptide in a linearized form. For example, if the polypeptide is charged, the polypeptide can be kept in a linearized form by applying a voltage.
[0354] If the polypeptide is uncharged or only weakly charged, the charge can be altered or controlled by adjusting the pH. For example, the polypeptide can be kept in a linearized form by increasing the relative negative charge of the polypeptide using a high pH. Increasing the negative charge of the polypeptide allows it to remain in a linearized form, for example, at a positive voltage. Alternatively, the polypeptide can be kept in a linearized form by increasing the relative positive charge of the polypeptide using a low pH. Increasing the positive charge of the polypeptide allows it to remain in a linearized form, for example, at a negative voltage. In the disclosed methods, polynucleotides are used to process proteins to control the movement of the polynucleotides relative to the nanopore. Since polynucleotides are generally negatively charged, like polynucleotides, it is generally most suitable to increase the linearization of the polypeptide by increasing the pH, thereby making the polypeptide more negatively charged. In this way, the conjugate retains an overall negative charge and can thus be easily moved, for example, under an applied voltage.
[0355] The polypeptide can be kept in a linearized form by using suitable denaturing conditions. Suitable denaturing conditions include, for example, the presence of an appropriate concentration of a denaturing agent, such as guanidine hydrochloride and / or urea. The concentration of such a denaturing agent used in the disclosed methods depends on the target polypeptide to be characterized in the method and can be readily selected by those skilled in the art.
[0356] The polypeptide can be kept in a linearized form by using a suitable detergent. Detergents suitable for use in the disclosed methods include SDS (sodium dodecyl sulfate). The polypeptide can be kept in a linearized form by performing the disclosed method at an elevated temperature. Increasing the temperature overcomes intrastrand bonding and allows the polypeptide to adopt a linearized form.
[0357] The polypeptide can be kept in a linearized form by performing the method under strong electroosmotic forces. Such forces can be provided by using asymmetric salt conditions and / or by providing a suitable charge in the channel of the nanopore. The charge in the channel of the pore can be altered, for example, by mutagenesis. Altering the charge of the pore is well within the capabilities of those skilled in the art. When a voltage potential is applied across the nanopore, altering the charge of the pore generates a strong electroosmotic force due to an imbalanced flow of cations and anions through the nanopore.
[0358] The polypeptide can be kept in a linearized form by passing the polypeptide through a structure such as a nanopillar array, through a nanoslit, or across a nanogap. The physical constraints of such structures can force the polypeptide to adopt a linearized form.
[0359] Preferably, polynucleotide-binding proteins are used to control the movement of the polypeptide relative to the pore, such as through the pore. Suitable proteins are discussed in more detail above. The present invention provides a method for determining the presence, absence, or one or more characteristics of a target polypeptide, the method comprising the following steps:
[0360] (i) contacting a target polypeptide with a pore complex or a pore multimer of the invention and a polypeptide binding protein such that the polypeptide binding protein controls the movement of the target polypeptide relative to (such as through) the pore complex or pore multimer; and
[0361] (ii) making one or more measurements as the polypeptide moves relative to (such as through) the pore complex or pore multimer, and thereby determining the presence, absence or one or more characteristics of the polypeptide.
[0362] The target analyte is preferably a polynucleotide, such as a nucleic acid, which is defined as a macromolecule comprising two or more nucleotides. Nucleic acids are particularly suitable for nanopore sequencing. The naturally occurring nucleic acid bases in DNA and RNA can be distinguished by their physical size. When a nucleic acid molecule or a single base passes through the channel of a nanopore, the size differences between the bases result in a directly related reduction in the ion current passing through the channel. The changes in the ion current can be recorded. Suitable electrical measurement techniques for recording changes in the ion current were discussed above. By appropriate calibration, the characteristic reduction in the ion current can be used to identify in real time the specific nucleotides and associated bases passing through the channel. In typical nanopore nucleic acid sequencing, as each nucleotide of the nucleic acid sequence of interest passes sequentially through the channel of the nanopore, the open-channel ion current decreases due to the partial blockage of the channel by the nucleotide. The above-described suitable recording techniques are used to measure this reduction in the ion current. The reduction in the ion current can be calibrated to the measured ion current reduction of known nucleotides passing through the channel, thereby providing a means for determining which nucleotide is passing through the channel and, thus, when done sequentially, a way to determine the nucleotide sequence of the nucleic acid passing through the nanopore. To accurately determine a single nucleotide, it is generally required that the reduction in the ion current passing through the channel be directly related to the size of the single nucleotide passing through the constriction. It should be understood that an entire nucleic acid polymer that "passes through" the pore via the action of, for example, a relevant polymerase can be sequenced. Alternatively, the sequence can be determined by the passage of nucleoside triphosphate bases that have been sequentially removed from the target nucleic acid near the pore (see, for example, WO 2014 / 187924, which is incorporated herein by reference in its entirety).
[0363] The polynucleotide or nucleic acid can comprise any combination of any nucleotides. The nucleotides can be any of those discussed above with respect to the pore monomer conjugates of the invention. In particular, the method of using a polynucleotide as an analyte alternatively comprises determining one or more characteristics selected from: (i) the length of the polynucleotide, (ii) the identity of the polynucleotide, (iii) the sequence of the polynucleotide, (iv) the secondary structure of the polynucleotide, and (v) whether the polynucleotide is modified.
[0364] A polynucleotide can have any length (i). For example, the length of the polynucleotide can be at least 10, at least 50, at least 100, at least 150, at least 200, at least 250, at least 300, at least 400, or at least 500 nucleotides or nucleotide pairs. The length of the polynucleotide can be 1000 or more nucleotides or nucleotide pairs, 5000 or more nucleotides or nucleotide pairs, or 100000 or more nucleotides or nucleotide pairs. Any number of polynucleotides can be studied. For example, the method can involve characterizing 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 50, 100, or more polynucleotides. If two or more polynucleotides are characterized, they can be different polynucleotides or two instances of the same polynucleotide. The polynucleotide can be naturally occurring or synthetic. For example, the method can be used to verify the sequence of a manufactured oligonucleotide. The method is typically performed in vitro.
[0365] The nucleotide can have any identity (ii). The possible nucleotides were defined above with reference to the pore monomer conjugates of the invention. The sequence of the nucleotides (iii) is determined by the sequential identity of successive nucleotides attached to one another in the 5' to 3' direction along the strand throughout the polynucleotide strain.
[0366] The pore complexes and pore polymers of the invention are particularly useful in analyzing homopolymers. For example, they can be used to determine the sequence of a polynucleotide comprising two or more, such as at least 3, 4, 5, 6, 7, 8, 9, or 10, identical consecutive nucleotides. For example, they can be used to sequence a polynucleotide comprising polyA, polyT, polyG, and / or polyC regions.
[0367] The constriction of CsgG consists of residues at positions 51, 55, and 56 of SEQ ID NO:3. The constriction of CsgG and its constriction mutants is generally clear. When DNA passes through the constriction, the interaction of approximately 5 bases of DNA with the constriction of the pore dominates the current signal at any given time. Although these clearer constrictions are very good for reading mixed sequence regions of DNA (when A, T, G, and C are mixed), when there are homopolymeric regions within the DNA (e.g., polyT, polyG, polyA, polyC), the signal becomes flat and lacks information. Since 5 bases dominate the signal of CsgG and its constriction mutants, it is difficult to distinguish homopolymers longer than 5 bases without using additional dwell time information. However, if the DNA is passing through a second constriction formed by the CsgF peptide, more DNA bases will interact with the combined constrictions, increasing the length of the homopolymer that can be distinguished. This is discussed in WO 2016 / 034591, WO 2017 / 149316, WO 2017 / 149317, WO 2017 / 149318, WO 2018 / 211241, and WO 2019 / 002893 (all incorporated herein by reference in their entirety).
[0368] Preferably, a polynucleotide-binding protein is used to control the movement of the polynucleotide relative to the pore (such as through the pore). Suitable proteins are discussed in more detail above. The present invention provides a method for determining the presence, absence, or one or more characteristics of a target polynucleotide, the method comprising the steps of:
[0369] (i) contacting the target polynucleotide with a pore complex of the present invention or a pore multimer of the present invention and a polynucleotide-binding protein such that the polynucleotide-binding protein controls the movement of the target analyte relative to (such as through) the pore complex or the pore
[0370] multimer; and
[0371] (ii) making one or more measurements while the polynucleotide is moving relative to (such as through) the pore complex or the pore multimer, and thereby determining the presence, absence, or one or more characteristics of the polynucleotide.
[0372] In any of these methods, one or more characteristics of the target analyte are preferably measured by electrical measurement and / or optical measurement. Electrical measurement is current measurement, impedance measurement, tunneling measurement, or field effect transistor (FET) measurement. The method preferably comprises measuring the current flowing through the pore complex or the pore multimer while the target analyte is moving relative to (such as through) the pore.
[0373] The general conditions for performing the methods of the present invention are discussed in more detail below with reference to the kits and systems of the present invention. Specific application of the functional pore complex or multimer of the present invention
[0374] As explained above, the chaperone molecule functionalizes the pore monomers and pore complexes of the present invention by linking functional binding proteins. The functionalized pore complexes and pore polymers of the present invention can be used in a variety of different specific applications of the methods of the present invention.
[0375] The functional binding moiety preferably interacts with or binds to the target analyte and facilitates its movement relative to the pore complex or polymer. In all contexts below, "facilitates" is synonymous with "improves".
[0376] The functional binding moieties capable of binding to any target analyte for use in the methods of the present invention were described above in the context of the pore monomer conjugates of the present invention. Accordingly, the present invention provides a method for determining the presence, absence, or one or more characteristics of a target analyte, the method comprising the steps of: (a) contacting the target analyte with a pore complex of the present invention or a pore polymer of the present invention such that the functional binding moiety interacts with or binds to the target analyte and facilitates its movement relative to (such as into or through) the pore complex or polymer; and (b) making one or more measurements when the target analyte moves relative to the pore complex or pore polymer, and thereby determining the presence, absence, or one or more characteristics of the target analyte.
[0377] The target analyte is preferably a polynucleotide, and the functional binding moiety preferably comprises an oligonucleotide, polynucleotide, polynucleotide analogue, or morpholino capable of hybridizing to the target analyte. Accordingly, the present invention provides a method for determining the presence, absence, or one or more characteristics of a target polynucleotide, the method comprising the steps of: (a) contacting the target polynucleotide with a pore complex of the present invention or a pore polymer of the present invention, wherein the functional binding moiety comprises an oligonucleotide, polynucleotide, polynucleotide analogue, or morpholino capable of hybridizing to the target polynucleotide, such that the functional binding moiety interacts with or binds to the target polynucleotide and facilitates its movement relative to (such as into or through) the pore complex or polymer; and (b) making one or more measurements when the target polynucleotide moves relative to the pore complex or pore polymer, and thereby determining the presence, absence, or one or more characteristics of the target polynucleotide. The functional binding moiety preferably hybridizes to the target polynucleotide. The functional binding moiety preferably facilitates the capture of the target polynucleotide by the pore complex or pore polymer. The functional binding moiety preferably comprises a morpholino.
[0378] The target polynucleotide is preferably double-stranded. A double-stranded polynucleotide typically comprises a template strand and a complement strand. Movement of the target polynucleotide relative to (such as through or into) a pore complex or pore polymer typically separates the two strands. Movement of the template strand relative to (such as through or into) a pore complex or pore polymer typically separates the two strands.
[0379] The movement of the target polynucleotide relative to (such as through or into) the pore complex or pore polymer is preferably controlled by a polynucleotide binding protein (such as a helicase). The polynucleotide binding protein preferably separates the two strands of the double-stranded polynucleotide as the double-stranded polynucleotide or template strand moves relative to (such as through or into) the pore complex or pore polymer.
[0380] The double-stranded polynucleotide (preferably the complement strand) preferably comprises a portion or region on one of the strands to which a functional binding moiety is capable of hybridizing. The double-stranded polynucleotide (preferably the complement strand) preferably comprises a binding region capable of hybridizing to a functional binding moiety.
[0381] The separation of the chains of the double-stranded polynucleotide preferably reveals a portion or region of one of the chains to which the functional binding portion is capable of hybridizing. The separation of the chains of the double-stranded polynucleotide preferably reveals a binding region that is capable of hybridizing with the functional binding portion. The functional binding portion preferably comprises a portion or region that is substantially complementary or complementary to the portion or region or binding region that is revealed when the two chains are separated. When the functional binding portion tethered to the hole hybridizes with this portion or region, it promotes or improves the movement of the chain comprising the portion or region relative to (such as entering or passing through) the hole complex or hole multimer. The functional binding protein promotes or improves the movement of the chain relative to the hole complex or hole multimer by bringing the chain closer to the hole complex or hole multimer or by promoting its capture. When the functional binding portion tethered to the hole hybridizes with this portion or region or binding region, it promotes or improves the capture of the chain (preferably the complement chain) by the hole complex or hole multimer. An example of improved capture according to the present invention is shown in Figure 3a , Figure 3b , Figure 3c and Figure 3d middle.
[0382] Those skilled in the art can use sequencing adapters, such as the adapters described in WO 2016 / 034591 and WO 2018 / 100370 (each incorporated herein by reference in its entirety), to connect suitable portions or regions to double-stranded polynucleotides. These adapters also include suitable binding sites for polynucleotide binding proteins. Those skilled in the art can also design functional binding portions that include portions or regions that can hybridize with the disclosed portions or regions.
[0383] The method preferably comprises: (i) separating the two strands of a double-stranded polynucleotide such that the template strand moves relative to (such as into or through) a pore complex or pore polymer, and such that a portion or region on the complementary strand to which a functional binding moiety can hybridize is exposed; (ii) allowing the functional binding moiety to hybridize to the portion or region on the complementary strand and facilitating the movement of the complementary strand relative to (such as into or through) the pore complex or pore polymer; and (iii) making one or more measurements as the template strand and the complementary strand of the double-stranded target polynucleotide move relative to the pore complex or pore polymer, and thereby determining the presence, absence, or one or more characteristics of the target double-stranded polynucleotide. The separation in step (i) preferably uses a polynucleotide-binding protein. The movement of the complementary strand in step (ii) is also preferably controlled using a polynucleotide-binding protein.
[0384] The method preferably comprises: (i) contacting a double-stranded polynucleotide with a pore complex or pore polymer such that when the template strand moves relative to (such as into or through) the pore complex or pore polymer, the two strands are separated and a binding region on the complementary strand is exposed, wherein a functional binding moiety hybridizes to the binding region to facilitate capture of the complementary strand through the pore complex or pore polymer; and (iii) making one or more measurements as the template strand and the complementary strand move relative to the pore complex or pore polymer, wherein measurements of both the template strand and the complementary strand are used to determine the presence, absence, or one or more characteristics of the target double-stranded polynucleotide. The movement of the template strand and / or the complementary strand is preferably controlled using a polynucleotide-binding protein.
[0385] Accordingly, the present invention provides a method for determining the presence, absence, or one or more characteristics of a target double-stranded polynucleotide, the method comprising the steps of: (a) contacting the target double-stranded polynucleotide with a pore complex of the present invention or a pore polymer of the present invention, wherein the functional binding moiety comprises an oligonucleotide, polynucleotide, polynucleotide analogue, or morpholino capable of hybridizing to a portion or region on the complementary strand of the target double-stranded polynucleotide; (b) separating the two strands of the double-stranded polynucleotide such that the template strand moves relative to (such as into or through) the pore complex or pore polymer and such that a portion or region on the complementary strand is exposed; (c) allowing the functional binding moiety to hybridize to the portion or region on the complementary strand and facilitating the movement of the complementary strand relative to (such as into or through) the pore complex or pore polymer; and (d) making one or more measurements as the template strand and the complementary strand of the double-stranded target polynucleotide move relative to the pore complex or pore polymer, and thereby determining the presence, absence, or one or more characteristics of the target double-stranded polynucleotide. The separation in step (a) preferably uses a polynucleotide-binding protein. The movement of the complementary strand in step (b) is also preferably controlled using a polynucleotide-binding protein. The functional binding moiety preferably comprises a morpholino.
[0386] The method preferably includes modifying the target double-stranded polynucleotide with a double-stranded sequencing adaptor before step (a), the adaptor comprising on one strand a moiety or region to which a functional binding moiety can hybridize. The adaptor is typically linked to the target double-stranded polynucleotide. The moiety or region is typically located on the strand of the adaptor that is linked to the complement strand of the target double-stranded polynucleotide. The adaptor preferably further comprises one or more binding sites and / or leader sequences for polynucleotide-binding proteins.
[0387] The present invention also provides a method for determining the presence, absence or one or more characteristics of a target polynucleotide, the method comprising the steps of: (a) contacting a double-stranded polynucleotide comprising a template strand and a complement strand with a pore complex or pore multimer comprising at least one pore monomer conjugate of the present invention, wherein the complement strand comprises a binding region capable of hybridizing to a functional binding moiety such that when the template strand moves through the pore complex or pore multimer the two strands separate to expose the binding region on the complement strand, wherein the functional binding moiety hybridizes to the binding region to facilitate capture of the complement strand through the pore complex or pore multimer; and (b) making one or more measurements as the template strand and the complement strand move relative to the pore complex or pore multimer, wherein measurements of both the template strand and the complement strand are used to determine the presence, absence or one or more characteristics of the target polynucleotide.
[0388] The method preferably includes modifying the double-stranded polynucleotide with a double-stranded sequencing adaptor before step (a), the adaptor comprising on one strand a binding region capable of hybridizing to a functional binding moiety. The adaptor is typically linked to the target polynucleotide. The moiety or region is typically located on the strand of the adaptor that is linked to the complement strand of the target polynucleotide. The adaptor preferably further comprises one or more binding sites and / or leader sequences for polynucleotide-binding proteins.
[0389] The pore complex or pore multimer can be used in any of the embodiments disclosed in WO 2018 / 100370 (incorporated herein by reference in its entirety). The companion molecule functions as a pore tether in that document.
[0390] In any of these specific embodiments, the most preferred pore monomer conjugate is preferably used in the case where the functional binding moiety hybridizes to the target polynucleotide, moiety or region or binding region.
[0391] The target analyte is preferably a polynucleotide, and the functional binding moiety is preferably linked to a polynucleotide-binding protein that controls the movement of the target polynucleotide relative to the pore complex or pore polymer. Accordingly, the present invention provides a method for determining the presence, absence, or one or more characteristics of a target polynucleotide, the method comprising the steps of: (a) contacting the target polynucleotide with a pore complex of the present invention or a pore polymer of the present invention, wherein the functional binding moiety is linked to a polynucleotide-binding protein that controls the movement of the target polynucleotide relative to (such as into or through) the pore complex or polymer; and (b) making one or more measurements when the target polynucleotide moves relative to the pore complex or pore polymer, and thereby determining the presence, absence, or one or more characteristics of the target polynucleotide. Polynucleotide-binding proteins are discussed in more detail above.
[0392] The target analyte is preferably a ligand of an enzyme, and the functional binding moiety is preferably linked to the enzyme. Accordingly, the present invention provides a method for determining the presence, absence, or one or more characteristics of a target ligand, the method comprising the steps of: (a) contacting the target ligand with a pore complex of the present invention or a pore polymer of the present invention, wherein the functional binding moiety is linked to an enzyme that binds to the target ligand and facilitates its movement relative to (such as into or through) the pore complex or polymer; and (b) making one or more measurements when the target ligand moves relative to the pore complex or pore polymer, and thereby determining the presence, absence, or one or more characteristics of the target ligand. Those skilled in the art can identify suitable ligand and enzyme pairs for use in the present invention.
[0393] The functional binding moiety is preferably an antibody or a functional fragment thereof or an aptamer that binds to the target analyte. Accordingly, the present invention provides a method for determining the presence, absence, or one or more characteristics of a target analyte, the method comprising the steps of: (a) contacting the target analyte with a pore complex of the present invention or a pore polymer of the present invention, wherein the functional binding moiety is an antibody or a functional fragment thereof or an aptamer that binds to the target analyte and facilitates its movement relative to (such as into or through) the pore complex or polymer; and (b) making one or more measurements when the target analyte moves relative to the pore complex or pore polymer, and thereby determining the presence, absence, or one or more characteristics of the target analyte. Antibodies, functional fragments, and aptamers are discussed above. The target analyte is preferably a target polypeptide.
[0394] The pore complexes and pore polymers of the present invention can also be used in a variety of other situations. The functional binding moiety is preferably attached to the pore, the cyclic protein, or the DNA origami structure to increase the distance between the polynucleotide-binding protein and the pore complex or pore polymer, or to provide one or more additional constrictions. Accordingly, the present invention provides a method for determining the presence, absence, or one or more characteristics of a target analyte, the method comprising the steps of: (a) contacting the target analyte with a pore complex of the present invention or a pore polymer of the present invention, wherein the functional binding moiety is attached to the pore, the cyclic protein, or the DNA origami structure to increase the distance between the polynucleotide-binding protein and the pore complex or pore polymer; (b) using the polynucleotide-binding protein to control the movement of the target analyte relative to (such as into or through) the pore complex or polymer; and (c) making one or more measurements as the target analyte moves relative to the pore complex or pore polymer, and thereby determining the presence, absence, or one or more characteristics of the target analyte. The present invention also provides a method for determining the presence, absence, or one or more characteristics of a target analyte, the method comprising the steps of: (a) contacting the target analyte with a pore complex of the present invention or a pore polymer of the present invention, wherein the functional binding moiety is attached to the pore, the cyclic protein, or the DNA origami to provide one or more additional constrictions such that the target analyte moves relative to (such as into or through) the pore complex or polymer; and (b) making one or more measurements as the target analyte moves relative to the pore complex or pore polymer, and thereby determining the presence, absence, or one or more characteristics of the target analyte.
[0395] The functional binding moiety is preferably bound to the pore complex or pore polymer to stabilize the cis-ring and reduce the signal-to-noise ratio of the pore complex or pore polymer. Accordingly, the present invention provides a method for determining the presence, absence, or one or more characteristics of a target analyte, the method comprising the steps of: (a) contacting the target analyte with a pore complex of the present invention or a pore polymer of the present invention, wherein the functional binding moiety is bound to the pore complex or pore polymer to stabilize the cis-ring such that the target analyte moves relative to (such as into or through) the pore complex or polymer; and (b) making one or more measurements as the target analyte moves relative to the pore complex or pore polymer, and thereby determining the presence, absence, or one or more characteristics of the target analyte.
[0396] The functional binding moiety preferably attaches the pore complex or pore polymer to a second pore or a second pore polymer. In Figure 4An example of such a situation is shown. The CsgA polymers of the present invention can be used to link two or more pores or two or more pore polymers. Accordingly, the present invention provides a method for determining the presence, absence, or one or more characteristics of a target analyte, the method comprising the steps of: (a) contacting the target analyte with a pore complex of the present invention or a pore polymer of the present invention, wherein a functional binding moiety links the pore complex or pore polymer to a second pore or a second pore polymer such that the target analyte moves relative to (such as into or through) the pore complex, pore polymer, second pore, or second polymer; and (b) making one or more measurements when the target analyte moves relative to the pore complex, pore polymer, second pore, or second polymer, and thereby determining the presence, absence, or one or more characteristics of the target analyte. The pore and the second pore and / or the pore polymer and the second polymer can be the same or different. The second pore or the second pore polymer can be a pore complex or a pore polymer of the present invention. The second pore can be a pore complex of the present invention. Such embodiments can be represented as follows: A pore monomer conjugate typically has the following structure: pore monomer ∼ chaperone molecule - functional binding moiety - functional binding moiety - chaperone molecule ∼ pore monomer, where "∼" = affinity, binding, linkage, or covalent linkage, and "-" = linkage or covalent linkage. The functional binding moiety - functional binding moiety is preferably a peptide linkage (or gene fusion) between two chaperone molecules, especially when they are chaperone polypeptides. Two chaperone polypeptides can be linked or fused using a linker (such as a polypeptide linker). If the method includes linking two CsgG pores, the C-terminus of one CsgA polypeptide is preferably linked to the R group of K at the C-terminus of another CsgA polypeptide. This provides a linked construct in which one CsgA polypeptide is N-to-C linked to another C-to-N. Then the N-terminus of each CsgA polypeptide can be used to link to a CsgG pore, thereby linking two CsgA pores. In this embodiment, the functional binding moiety - functional binding moiety is a peptide linkage (or gene fusion) between two CsgA polypeptides. The CsgA polypeptides can be any of those discussed above. The CsgA polypeptides can be the same or different. The second pore polymer can be a pore polymer of the present invention.
[0397] The movement of the target polypeptide relative to (such as through or into) the pore complex or pore polymer is preferably controlled by a polypeptide binding protein. Such proteins were discussed above. The polypeptide binding protein preferably unfolds the target polypeptide when the target polypeptide moves relative to (such as through or into) the pore complex or pore polymer. Unfolding enzymes were also discussed above.
[0398] The target analyte is preferably a polypeptide, and the functional binding moiety is preferably linked to a polypeptide-binding protein that controls the movement of the target polypeptide relative to the pore complex or pore multimer. Accordingly, the present invention provides a method for determining the presence, absence, or one or more characteristics of a target polypeptide, the method comprising the steps of: (a) contacting the target polypeptide with a pore complex or pore multimer of the present invention, wherein the functional binding moiety is linked to a polypeptide-binding protein that controls the movement of the target polypeptide relative to (such as into or through) the pore complex or multimer; and (b) making one or more measurements when the target polypeptide moves relative to the pore complex or pore multimer, and thereby determining the presence, absence, or one or more characteristics of the target polypeptide. The polypeptide-binding protein is discussed in more detail above.
[0399] The method preferably includes modifying the target polypeptide with a sequencing adaptor prior to step (a), the adaptor comprising a portion or region to which the functional binding moiety can hybridize. The sequencing adaptor preferably further comprises one or more binding sites and / or leader sequences for the polypeptide-binding protein.
[0400] The present invention also provides a method for determining the presence, absence, or one or more characteristics of a target polypeptide, the method comprising the steps of: (a) modifying the target polypeptide with a sequencing adaptor comprising a portion or region to which the functional binding moiety can hybridize; (b) contacting the modified target polypeptide with a pore complex or pore multimer comprising at least one pore monomer conjugate of the present invention, wherein the functional binding moiety hybridizes to the sequencing adaptor to facilitate capture of the target polypeptide by the pore complex or pore multimer; and (c) making one or more measurements when the target polypeptide moves relative to the pore complex or pore multimer, wherein the measurements are used to determine the presence, absence, or one or more characteristics of the target polypeptide.
[0401] As explained above, the functional binding moiety is preferably an antibody or a functional fragment thereof or an aptamer that binds to the target polypeptide. Accordingly, the present invention provides a method for determining the presence, absence, or one or more characteristics of a target polypeptide, the method comprising the steps of: (a) contacting the target polypeptide with a pore complex or pore multimer of the present invention, wherein the functional binding moiety is an antibody or a functional fragment thereof or an aptamer that binds to the target polypeptide and facilitates its movement relative to (such as into or through) the pore complex or multimer; and (b) making one or more measurements when the target polypeptide moves relative to the pore complex or pore multimer, and thereby determining the presence, absence, or one or more characteristics of the target polypeptide. Antibodies, functional fragments, and aptamers are discussed above.
[0402] In these two targeted embodiments using sequencing adaptors or antibodies, functional fragments or aptamers, the movement of the target polypeptide relative to (such as through or into) the pore complex or pore multimer is preferably controlled by a polypeptide-binding protein.
[0403] Alternative example
[0404] The present invention provides a method for determining the presence, absence or one or more characteristics of a target double-stranded polynucleotide, the method comprising the steps of: (a) contacting the target double-stranded polynucleotide with a pore complex or pore multimer; (b) separating the two strands of the double-stranded polynucleotide such that the template strand moves relative to (such as into or through) the pore complex or pore multimer and such that a portion or region on the complementary strand is exposed; (c) allowing a functional binding moiety to hybridize to the portion or region on the complementary strand and facilitating the movement of the complementary strand relative to (such as into or through) the pore complex or pore multimer; and (d) making one or more measurements as the template strand and the complementary strand of the double-stranded target polynucleotide move relative to the pore complex or pore multimer, and thereby determining the presence, absence or one or more characteristics of the target double-stranded polynucleotide,
[0405] wherein the pore complex or pore multimer comprises at least one pore monomer conjugate comprising a CsgG pore monomer, a chaperone molecule and a functional binding moiety,
[0406] wherein the CsgG pore monomer comprises a sequence that is at least about 40% homologous or identical to the amino acid sequence of SEQ ID NO: 3 over the entire sequence,
[0407] wherein the chaperone molecule comprises SEQ ID NO: 21, 23, 25, 27, 29 or 31,
[0408] wherein the K at position 7 of SEQ ID NO: 21, 23, 25, 27, 29 or 31 is covalently linked to the CsgG pore monomer via a sulfonyl group,
[0409] wherein the functional binding moiety comprises an oligonucleotide, polynucleotide, polynucleotide analogue or morpholino capable of hybridizing to a portion or region on the complementary strand, and
[0410] wherein the functional binding moiety is covalently linked to position 14 of SEQ ID NO: 21, 23, 25, 27, 29 or 31. The CsgG pore monomer sequence preferably comprises F56Q. The K at the C-terminus of SEQ ID NO: 21, 23, 25, 27, 29 or 31 is preferably modified with an azide group such as N3. The chaperone molecule preferably comprises SEQ ID NO: 21, or consists of it. The chaperone molecule preferably comprises SEQ ID NO: 40, 43, 44, 45, 46 or 47, or consists of it. The chaperone molecule preferably comprises SEQ ID NO: 40, or consists of it. The functional binding moiety preferably comprises a morpholino.
[0411] Alternatively, in the method, the pore monomer conjugate comprises a CsgG pore monomer, a chaperone molecule and a functional binding moiety,
[0412] wherein the CsgG pore monomer comprises a sequence that is at least about 40% homologous or identical to the amino acid sequence of SEQ ID NO: 3 over the entire sequence,
[0413] wherein the chaperone molecule comprises SEQ ID NO: 21, 23, 25, 27, 29, 31, 63, 64, 65, 66, 67 or 68,
[0414] wherein the K at position 7 of SEQ ID NO: 21, 23, 25, 27, 29 or 31 is covalently linked to the CsgG pore monomer via a sulfonyl group, or the K at position 8 of SEQ ID NO: 63, 64, 65, 66, 67 or 68 is covalently linked to the CsgG pore monomer via a sulfonyl group,
[0415] wherein the functional binding moiety comprises an oligonucleotide, polynucleotide, polynucleotide analogue or morpholino capable of specifically hybridizing to a target polynucleotide analyte, and
[0416] wherein the functional binding moiety is covalently linked to the C-terminal K of SEQ ID NO: 21, 23, 25, 27, 29, 31, 63, 64, 65, 66, 67 or 68. The CsgG pore monomer can have any percentage homology or identity with SEQ ID NO: 3 as set forth above. The CsgG pore monomer sequence preferably comprises F56Q. The K at the C-terminus of SEQ ID NO: 21, 23, 25, 27, 29, 31, 63, 64, 65, 66, 67 or 68 is preferably modified with an azide group such as N3. The chaperone molecule preferably comprises or consists of SEQ ID NO: 21 or 63. The chaperone molecule preferably comprises or consists of SEQ ID NO: 40, 43, 44, 45, 46, 47, 48, 49 or 50. The chaperone molecule preferably comprises or consists of SEQ ID NO: 40. The chaperone molecule preferably comprises a -CONH2 group at its C-terminus. The functional binding moiety preferably comprises a morpholino.
[0417] The present invention provides a method for determining the presence, absence or one or more characteristics of a target polynucleotide, the method comprising the steps of: (a) contacting a double-stranded polynucleotide comprising a template and a complementary strand with a pore complex or pore multimer, wherein the complementary strand comprises a binding region capable of hybridizing with a functional binding moiety such that when the template strand moves through the pore complex or pore multimer, the two strands separate to expose the binding region on the complementary strand, wherein the functional binding moiety hybridizes with the binding region to facilitate capture of the complementary strand through the pore complex or pore multimer; and (b) making one or more measurements as the template strand and the complementary strand move relative to the pore complex or pore multimer, wherein the measurements of both the template strand and the complementary strand are used to determine the presence, absence or one or more characteristics of the target polynucleotide,
[0418] wherein the pore complex or pore multimer comprises at least one pore monomer conjugate comprising a CsgG pore monomer, a chaperone molecule and a functional binding moiety,
[0419] wherein the CsgG pore monomer comprises a sequence that is at least about 40% homologous or identical to the amino acid sequence of SEQ ID NO: 3 over the entire sequence,
[0420] wherein the chaperone molecule comprises SEQ ID NO: 21, 23, 25, 27, 29 or 31,
[0421] wherein the K at position 7 of SEQ ID NO: 21, 23, 25, 27, 29 or 31 is covalently linked to the CsgG pore monomer via a sulfonyl group,
[0422] wherein the functional binding moiety comprises an oligonucleotide, polynucleotide, polynucleotide analog, or morpholino capable of hybridizing to a portion or region on a complement strand, and
[0423] wherein the functional binding moiety is covalently linked to position 14 of SEQ ID NO: 21, 23, 25, 27, 29, or 31. The CsgG pore monomer sequence preferably comprises F56Q. The K at the C-terminus of SEQ ID NO: 21, 23, 25, 27, 29, or 31 is preferably modified with an azide group such as N3. The chaperone molecule preferably comprises SEQ ID NO: 21, or consists of it. The chaperone molecule preferably comprises SEQ ID NO: 40, 43, 44, 45, 46, or 47, or consists of it. The chaperone molecule preferably comprises SEQ ID NO: 40, or consists of it. The functional binding moiety preferably comprises a morpholino.
[0424] Alternatively, in the method, the pore monomer conjugate comprises a CsgG pore monomer, a chaperone molecule, and a functional binding moiety,
[0425] wherein the CsgG pore monomer comprises a sequence that is at least about 40% homologous or identical to the amino acid sequence of SEQ ID NO: 3 over the entire sequence,
[0426] wherein the chaperone molecule comprises SEQ ID NO: 21, 23, 25, 27, 29, 31, 63, 64, 65, 66, 67, or 68,
[0427] wherein the K at position 7 of SEQ ID NO: 21, 23, 25, 27, 29, or 31 is covalently linked to the CsgG pore monomer via a sulfonyl group, or the K at position 8 of SEQ ID NO: 63, 64, 65, 66, 67, or 68 is covalently linked to the CsgG pore monomer via a sulfonyl group,
[0428] wherein the functional binding moiety comprises an oligonucleotide, polynucleotide, polynucleotide analog, or morpholino capable of specifically hybridizing to a target polynucleotide analyte, and
[0429] Wherein the functional binding moiety is covalently linked to the C-terminal K of SEQ ID NO: 21, 23, 25, 27, 29, 31, 63, 64, 65, 66, 67 or 68. The CsgG pore monomer can have any percentage of homology or identity with SEQ ID NO: 3 as set forth above. The CsgG pore monomer sequence preferably comprises F56Q. The K at the C-terminus of SEQ ID NO: 21, 23, 25, 27, 29, 31, 63, 64, 65, 66, 67 or 68 is preferably modified with an azide group such as N3. The chaperone molecule preferably comprises or consists of SEQ ID NO: 21 or 63. The chaperone molecule preferably comprises or consists of SEQ ID NO: 40, 43, 44, 45, 46, 47, 48, 49 or 50. The chaperone molecule preferably comprises or consists of SEQ ID NO: 40. The chaperone molecule preferably comprises a -CONH2 group at its C-terminus. The functional binding moiety preferably comprises a morpholino.
[0430] These methods are preferably carried out using a pore complex comprising at least one pore monomer conjugate as defined. The pore complex is preferably a homo-oligomer comprising 6 to 10 (such as 6, 7, 8, 9 or 10) pore monomer conjugates. The pore monomer conjugates are typically identical. The pore complex preferably comprises 8 or 9 identical pore monomer conjugates. The CsgG pore monomer can have any percentage of homology or identity with SEQ ID NO: 3 as set forth above.
[0431] The method preferably uses a polynucleotide-binding protein to control the movement of one or both strands of a double-stranded polynucleotide. The movement of the template and / or complement strand is preferably controlled using a polynucleotide-binding protein.
[0432] The method preferably comprises modifying the target double-stranded polynucleotide with a double-stranded sequencing adaptor prior to step (a), the adaptor comprising on one strand a moiety or region to which an oligonucleotide, polynucleotide, polynucleotide analogue or morpholino in the functional binding moiety can hybridize. The method preferably comprises modifying the double-stranded polynucleotide with a double-stranded sequencing adaptor prior to step (a), the adaptor comprising on one strand a binding region capable of hybridizing with the functional binding moiety. The adaptor is typically linked to the target polynucleotide. The moiety or region or binding region is typically located on the strand of the adaptor that is linked to the complement strand of the target polynucleotide. The adaptor preferably further comprises one or more binding sites and / or leader sequences for the polynucleotide-binding protein.
[0433] Polypeptides and polynucleotides of the present invention
[0434] The present invention also provides a polypeptide consisting of the sequences of SEQ ID NO: 7, 9, 11, 13, 15 or 17. The present invention also provides a polypeptide comprising or consisting of SEQ ID NO: 8, 10, 12, 14, 16, 18, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46 or 47.
[0435] The present invention also provides a polypeptide comprising or consisting of SEQ ID NO: 8, 10, 12, 14, 16, 18, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69 or 70.
[0436] The present invention also provides polypeptides comprising a variant of SEQ ID NO:7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46 or 47 or consisting of the same. Over the entire length of the amino acid sequence of the CsgA fragment or of SEQ ID NO:7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46 or 47, the variant will preferably be at least about 40% homologous to the sequence, based on amino acid identity. More preferably, based on amino acid identity over the entire sequence of the amino acid sequence of the CsgA fragment or of SEQ ID NO:7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46 or 47, the variant may be at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90% and more preferably at least about 95%, 97% or 99% homologous. Over the entire length of the amino acid sequence of the CsgA fragment or of SEQ ID NO:7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46 or 47, the variant will preferably be at least about 40% identical to the sequence. More preferably, over the entire sequence, the variant may be at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90% and more preferably at least about 95%, 97% or 99% identical to the CsgA fragment or SEQ ID NO:7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46 or 47.
[0437] The present invention also provides a polypeptide comprising a variant of SEQ ID NO:7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69 or 70 or consisting of the same. Over the entire length of the amino acid sequence of the CsgA fragment or of SEQ ID NO:7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69 or 70, the variant will preferably be at least about 40% homologous to the sequence based on amino acid identity. More preferably, based on the amino acid identity over the entire sequence of the amino acid sequence of the CsgA fragment or of SEQ ID NO:7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69 or 70, the variant may be at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, more preferably at least about 95%, 97% or 99% homologous. Over the entire length of the amino acid sequence of the CsgA fragment or of SEQ ID NO:7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69 or 70, the variant will preferably be at least about 40% identical to the sequence.More preferably, over the entire sequence, the variant can be at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90% and more preferably at least about 95%, 97% or 99% identical to the CsgA fragment or SEQ ID NO:7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69 or 70.
[0438] The amino acid at any one of positions 1 to 7 of SEQ ID NO:7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 or 31 (such as G at position 1, V at position 2, V at position 3, P at position 4, Q at position 5, Y at position 6 or G or K at position 7) is preferably modified to contain a reactive group. The amino acid at positions 6 or 7 of SEQ ID NO:7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 or 31 (such as Y at position 6 or G or K at position 7) is preferably modified to contain a reactive group. The reactive group can be any of those discussed above. The reactive group is preferably sulfonyl fluoride. The K at the C-terminus of SEQ ID NO:20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 or 31 is preferably modified with an azide group (such as N3).
[0439] The amino acid at any one of positions 1 to 10 (such as position 1, position 2, position 3, position 4, position 5, position 6, position 7, position 8, position 9 or position 10) in SEQ ID NO: 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69 or 70 is preferably modified to contain a reactive group. The amino acid at any one of positions 1 to 8 (such as position 1, position 2, position 3, position 4, position 5, position 6, position 7 or position 8) in SEQ ID NO: 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69 or 70 is preferably modified to contain a reactive group. The amino acid at positions 6, 7 or 8 in SEQ ID NO: 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69 or 70 is preferably modified to contain a reactive group. The amino acid at positions 7 or 8 in SEQ ID NO: 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69 or 70 is preferably modified to contain a reactive group.The amino acid at position 7 of SEQ ID NO: 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69 or 70 is preferably modified to contain a reactive group. The amino acid at position 8 of SEQ ID NO: 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69 or 70 is preferably modified to contain a reactive group. The reactive group can be any of those discussed above. The reactive group is preferably sulfonyl fluoride. K at the C-terminus of SEQ ID NO: 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 or 31 is preferably modified with an azide group such as N3.
[0440] G or K at position 7 of SEQ ID NO: 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 or 31 is preferably modified to contain a reactive group. G or K at position 8 of SEQ ID NO: 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69 or 70 is preferably modified to contain a reactive group. The reactive group can be any of those discussed above. The reactive group is preferably sulfonyl fluoride.
[0441] The K at the C-terminus of SEQ ID NO: 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 or 31 is preferably modified with an azide group such as N3. The K at the C-terminus of SEQ ID NO: 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69 or 70 is preferably modified with an azide group such as N3.
[0442] In any of the above CsgA polypeptides (including those based on SEQ ID NO: 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46 or 47 or variants thereof), the G at position 1 can be substituted or deleted. In any of the above CsgA polypeptides (including those based on SEQ ID NO: 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69 or 70 or variants thereof), the G at position 1 can be substituted or deleted.
[0443] Any of the CsgA polypeptides of the present invention may contain a -CONH2 group at its C-terminus.
[0444] The present invention also provides CsgA multimers comprising two or more linked CsgA polypeptides. The two or more linked CsgA polypeptides can be selected from any of those discussed above, including those from SEQ ID NO: 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69 and 70 and variants thereof. The CsgA multimer can comprise any number of two or more linked CsgA polypeptides, such as 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more or 10 or more CsgA polypeptides. The two or more CsgA polypeptides can be the same CsgA polypeptide. The two or more CsgA polypeptides can be different CsgA polypeptides.
[0445] The two or more CsgA polypeptides can be linked in any manner. The two or more CsgA polypeptides are preferably covalently linked. The two or more CsgA polypeptides are preferably covalently linked via a K at the C-terminus of one or more CsgA polypeptides. The two or more CsgA polypeptides can be directly covalently linked or covalently linked via any of the linkers disclosed herein. Those skilled in the art are able to link two or more CsgA polypeptides, for example using peptide bonds, via an R group at any position in the CsgA polypeptide or using the reactive groups discussed above. The two or more CsgA polypeptides can be linked in any orientation. Two or more CsgA polypeptides can be N-terminally linked to C-terminally or C-terminally linked to N-terminally. The two or more CsgA polypeptides can be N-terminally linked to N-terminally. The two or more CsgA polypeptides can be C-terminally linked to C-terminally. One or more CsgA polypeptides can branch from one or more positions within the sequence of one or more CsgA polypeptides.
[0446] The multimer preferably comprises two CsgA polypeptides, wherein the C-terminus of one CsgA polypeptide is linked to the K at the C-terminus of the other CsgA polypeptide. This provides a linked multimer in which one CsgA polypeptide is N-to-C linked to the other C-to-N. Then, the N-terminus of each CsgA polypeptide can be used to link to a pore monomer, such as a CsgG pore monomer. The two CsgA polypeptides can be selected from SEQ ID NO:7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69 and 70 and variants thereof. The two CsgA polypeptides are preferably SEQ ID NO:69 and SEQ ID NO:70. The C-terminus of SEQ ID NO:70 is preferably linked to the K at the C-terminus of SEQ ID NO:69. The invention also provides polynucleotides encoding any of these polypeptides. The polynucleotides can be any of those discussed above. The invention also provides an expression vector comprising the polynucleotide of the invention. The invention also provides a host cell comprising the polynucleotide of the invention or the host cell of the invention. Suitable vectors and host cells are known in the art.
[0447] Kit
[0448] The invention also provides a kit for characterizing a target analyte or for producing a pore monomer conjugate of the invention. The kit comprises (a) a pore monomer and (b) a partner molecule having an affinity for the pore monomer and linked to a functional binding moiety. Any of the embodiments discussed above with reference to the pore monomer conjugates of the invention are equally applicable here.
[0449] The invention also provides a kit for characterizing a target analyte. The kit comprises (a) a pore complex of the invention or a pore multimer of the invention and (b) components of a membrane. Suitable membranes and components are discussed below.
[0450] The present invention also provides a kit for characterizing a target polynucleotide. The kit includes (a) the pore complex of the present invention or the pore polymer of the present invention, and (b) a polynucleotide-binding protein. The present invention also provides a kit for characterizing a target polynucleotide or a target polypeptide. The kit includes (a) the pore complex of the present invention or the pore polymer of the present invention, and (b) a polynucleotide-binding protein or a polypeptide-binding protein. The kit preferably further comprises components of a membrane. The kit can comprise any type of membrane such as an amphiphilic layer, such as components of a triblock copolymer membrane. The polynucleotide-binding protein was discussed above with reference to the pore monomer conjugate of the present invention.
[0451] The kit can further comprise one or more anchors, such as cholesterol, for conjugating the target analyte to the membrane. The kit can further comprise one or more polynucleotide linkers that can be attached to the target polynucleotide to facilitate characterization of the polynucleotide. The anchor, such as cholesterol, is preferably attached to the polynucleotide linker.
[0452] The kit can additionally include one or more other reagents or instruments that enable any of the above embodiments to be carried out. Such reagents or instruments include one or more of the following: a suitable buffer (aqueous solution), a tool for obtaining a sample from a subject (such as a container or instrument containing a needle), a tool for amplifying and / or expressing a polynucleotide, or a voltage or patch clamp device. The reagents can be present in the kit in a dry state such that a fluid sample resuspends the reagents. The kit can also optionally include instructions for using the kit in the methods of the present invention or details regarding which organisms the method can be used for. Finally, the kit can also include additional components useful in analyte characterization.
[0453] Device
[0454] The present invention also provides a device for characterizing a target analyte in a sample, the device comprising (a) a plurality of the pore complexes of the present invention or a plurality of the pore polymers of the present invention, and (b) a plurality of polynucleotide-binding proteins. The present invention also provides a device for characterizing a target analyte in a sample, the device comprising (a) a plurality of the pore complexes of the present invention or a plurality of the pore polymers of the present invention, and (b) a plurality of polynucleotide-binding proteins or a plurality of polypeptide-binding proteins. The present invention also provides a device for characterizing a target polynucleotide or a target polypeptide in a sample, the device comprising (a) a plurality of the pore complexes of the present invention or a plurality of the pore polymers of the present invention and (b) a plurality of polynucleotide-binding proteins or a plurality of polypeptide-binding proteins. The plurality of pore complexes or the plurality of pore polymers can be any of those described above.
[0455] The present invention also provides a device comprising the pore complex of the present invention or the pore polymer of the present invention inserted into an in vitro membrane.
[0456] The present invention also provides a device which is produced by a method comprising the following steps: (i) obtaining a pore complex of the present invention or a pore polymer of the present invention, and (ii) contacting the pore complex or pore polymer with an in vitro membrane such that the pore complex or pore polymer inserts into the in vitro membrane.
[0457] Any specific embodiments discussed above are equally applicable to the device of the present invention.
[0458] Array
[0459] The present invention also provides an array comprising a plurality of membranes of the present invention. Any embodiments discussed above with respect to the membranes of the present invention are equally applicable to the array of the present invention. The array can be arranged to perform any one of the following methods.
[0460] In one embodiment, each membrane in the array comprises a pore complex or a pore polymer. Due to the way the array is formed, for example, the array can comprise one or more layers of membranes that do not contain a pore complex or a pore polymer, and / or one or more layers of membranes that contain two or more pore complexes or polymers. The array can comprise from about 2 to about 1000 layers, such as from about 10 to about 800 layers, from about 20 to about 600 layers or from about 30 to about 500 layers of membranes.
[0461] System
[0462] The present invention provides a system comprising (a) a membrane of the present invention or an array of the present invention, (b) means for applying an electric potential across the membrane, and (c) means for detecting an electrical or optical signal across the membrane. The electrical signal can be a measurement of ions flowing through a nanopore, such as a measurement of current or voltage that varies over time.
[0463] The pores and membranes can be any pores and membranes as described above and below.
[0464] In one embodiment, the system further comprises a first chamber and a second chamber, wherein the first chamber and the second chamber are separated by the membrane. When used to characterize a target analyte, the system can further comprise the target analyte, wherein the target analyte is transiently located within a continuous channel, and wherein one end of the target analyte is located in the first chamber and one end of the target analyte is located in the second chamber. The target analyte is preferably a target polypeptide or a target polynucleotide.
[0465] In one embodiment, the system further includes a conductive solution in contact with the pore, an electrode that provides a voltage potential across the membrane, and a measurement system for measuring the current passing through the pore. The voltage applied across the membrane and the pore is preferably from +5V to -5V, such as from -600mV to +600mV or from -400mV to +400mV. The voltage used is preferably in the range of 100mV to 240mV and more preferably in the range of 120mV to 220mV. By using an increased applied potential, the discrimination of the pore for different amino acids or nucleotides can be increased. Any suitable conductive solution can be used. For example, the solution can contain charge carriers, such as metal salts, for example alkali metal salts, halide salts, for example chloride salts, such as alkali metal chloride salts. The charge carriers can include ionic liquids or organic salts, such as tetramethylammonium chloride, trimethylphenylammonium chloride, phenyltrimethylammonium chloride, or 1-ethyl-3-methylimidazolium chloride. In an exemplary system, the salt is present in an aqueous solution in the chamber. Potassium chloride (KCl), sodium chloride (NaCl), cesium chloride (CsCl), or a mixture of potassium ferrocyanide and potassium ferricyanide is commonly used. KCl, NaCl, and the mixture of potassium ferrocyanide and potassium ferricyanide are preferred. The charge carriers can be transmembrane asymmetric. For example, on each side of the membrane, such as in each chamber, the type and / or concentration of the charge carriers can be different.
[0466] The salt concentration can be at saturation. The salt concentration can be 3M or lower, and is typically 0.1 to 2.5M, 0.3 to 1.9M, 0.5 to 1.8M, 0.7 to 1.7M, 0.9 to 1.6M, or 1M to 1.4M. The salt concentration is preferably 150mM to 1M. The method is preferably performed using a salt concentration of at least 0.3M, such as at least 0.4M, at least 0.5M, at least 0.6M, at least 0.8M, at least 1.0M, at least 1.5M, at least 2.0M, at least 2.5M, or at least 3.0M. The high salt concentration provides a high signal-to-noise ratio and allows the identification of currents indicative of the presence of amino acids or nucleotides against the background of normal current fluctuations.
[0467] A buffer can be present in the conductive solution. Typically, the buffer is a phosphate buffer. Other suitable buffers are HEPES and Tris-HCl buffers. The pH of the conductive solution can be 4.0 to 12.0, 4.5 to 10.0, 5.0 to 9.0, 5.5 to 8.8, 6.0 to 8.7, or 7.0 to 8.8 or 7.5 to 8.5. The pH used is preferably about 7.5.
[0468] The system can be included in a device. The device can be any conventional device for analyte analysis, such as an array or a chip. The device is preferably arranged to perform the disclosed method. For example, the device can include a chamber containing an aqueous solution and a barrier that divides the chamber into two segments. The barrier typically has pores in which a membrane containing pores is formed. Alternatively, the barrier forms a membrane in which pores are present.
[0469] The device can also include circuitry capable of applying an electric potential and measuring an electrical signal across the membrane and the pores.
[0470] The device can be any one of those described in WO 2008 / 102120, WO 2009 / 077734, WO 2010 / 122293, WO 2011 / 067559, WO 2014 / 06442 or WO2020 / 183172 (all incorporated herein by reference in their entirety).
[0471] Measured analysis
[0472] A method for determining the presence, absence, or one or more characteristics of a target polymer analyte can include estimating or determining the sequence of polymer units. The signal measured during the movement of a polymer (such as a polynucleotide) relative to a nanopore may at any time depend on a plurality of polymer units (such as nucleotides). For example, the presence of multiple nucleotides within the lumen of the nanopore and the possible presence of nucleotides outside the nanopore can affect the ion flow and thus the current or voltage signal. The polynucleotide can also contain modified bases that can affect the measured signal, and thus the estimation or determination of the sequence may not be straightforward. Various known mathematical techniques and their variants can be used to determine or estimate the polymer sequence, including probabilistic and machine learning techniques. Such methods are described, for example, in WO2013041878, WO2013121224, WO2018203084 and Zhang et al.: A Guide to Signal Processing Algorithms for Nanopore Sensors, ACS Sens. 2021, 6, 10, 3536–3555, which are all hereby incorporated by reference in their entirety.
[0473] According to aspects of the present invention, the method can include measuring a template and a complementary strand of a target polynucleotide or polynucleotide - polypeptide conjugate, wherein the measurements of both strands can be used to estimate or determine the overall sequence. A variety of known methods can be used. For example, the sequences of the template and the complementary strand can be initially determined from a series of measurements made during the movement of the polynucleotide relative to the nanopore, and the results combined to provide the overall sequence. More preferably, a series of measurements of both the template and the complement can be processed by probabilistic or machine learning techniques into a series of measurements in multiple dimensions, wherein the overall sequence determination is made without initially determining the sequences of the template and the complementary strand. Non - limiting examples of methods suitable for the methods of the present invention are disclosed in WO2015140535 (incorporated herein by reference in its entirety).
[0474] membrane
[0475] Any suitable membrane can be used in the system. The membrane is preferably an amphiphilic layer. An amphiphilic layer is a layer formed from amphiphilic molecules such as phospholipids that have hydrophilic and lipophilic properties. The amphiphilic molecules can be synthetic or naturally occurring. Non - naturally occurring amphiphiles and amphiphiles that form monolayers are known in the art and include, for example, block copolymers (Gonzalez - Perez et al., Langmuir, 2009, 25, 10447 - 10450). A block copolymer is a polymeric material in which two or more monomer subunits are polymerized together to produce a single polymer chain. Block copolymers typically have properties contributed by each monomer subunit. However, block copolymers can have unique properties not possessed by the polymers formed by the individual subunits. Block copolymers can be designed such that one of the monomer sub - units is hydrophobic (i.e., lipophilic) while the other subunits are hydrophilic when in an aqueous medium. In this case, the block copolymer can possess amphiphilic properties and can form structures that mimic biological membranes. Block copolymers can be diblock (consisting of two monomer subunits), but can also be constructed from more than two monomer subunits to form more complex arrangements that behave as amphiphiles. The amphiphilic layer can comprise diblock, triblock, tetrablock, or pentablock copolymers.
[0476] The membrane can include one of the membranes disclosed in International Application No. WO 2014 / 064443 or WO 2014 / 064444 (both incorporated herein by reference in their entirety).
[0477] The amphiphilic molecules can be chemically modified or functionalized to facilitate the coupling of polynucleotides. The amphiphilic layer can be a monolayer or a bilayer. The amphiphilic layer is typically planar. The amphiphilic layer can be curved. The amphiphilic layer can be supported.
[0478] The amphiphilic membrane is typically naturally mobile and essentially acts as having approximately 10 -8 cm s -1A two-dimensional fluid with a lipid diffusion rate. This means that the pores and the coupled polynucleotides can generally move within the amphiphilic membrane.
[0479] The membrane can be a lipid bilayer. The lipid bilayer is a model of the cell membrane and serves as an excellent platform for a series of experimental studies. For example, the lipid bilayer can be used for in vitro studies of membrane proteins by single-channel recording. Alternatively, the lipid bilayer can be used as a biosensor for detecting the presence of a series of substances. The lipid bilayer can be any lipid bilayer. Suitable lipid bilayers include, but are not limited to, planar lipid bilayers, supported bilayers, or liposomes. The lipid bilayer is preferably a planar lipid bilayer. Suitable lipid bilayers are disclosed in WO 2008 / 102121, WO 2009 / 077734, and WO 2006 / 100484 (all incorporated herein by reference in their entirety).
[0480] The membrane can include a solid state layer. The solid state layer can be formed from organic and inorganic materials, including but not limited to microelectronic materials, insulating materials (such as Si 3 N 4 、A1 2 O 3 and SiO), organic and inorganic polymers (such as polyamides), plastics (such as ) or elastomers (such as two-component addition-cured silicone rubber) and glass. The solid state layer can be formed from graphene. Suitable graphene layers are disclosed in WO2009 / 035647 (which is incorporated herein by reference in its entirety). If the membrane includes a solid state layer, then the pores are typically present in the amphiphilic membrane or layer that is contained within the solid state layer, for example, within holes, pores, gaps, channels, grooves, or slits within the solid state layer. Those skilled in the art can prepare suitable solid state / amphiphilic hybrid systems. Suitable systems are disclosed in WO 2009 / 020682 and WO 2012 / 005857 (both incorporated herein by reference in their entirety). Any of the amphiphilic membranes or layers discussed above can be used.
[0481] The method is generally carried out using (i) an artificial amphiphilic layer containing pores, (ii) a separate naturally occurring lipid bilayer containing pores, or (iii) a cell with pores inserted therein. The method is typically carried out using an artificial amphiphilic layer (such as a diblock or triblock copolymer layer). The layer can include other transmembrane and / or intramembrane proteins and other molecules other than pores. Suitable devices and conditions are discussed below. The method of the present invention is generally carried out in vitro.
[0482] Sequence Listing
[0483] SEQ ID NO:1 (>P0AEA2; Coding sequence of WT CsgG from Escherichia coli K12)
[0484] ATGCAGCGCTTATTTCTTTTGGTTGCCGTCATGTTACTGAGCGGATGCTTAACCGCCC
[0485] CGCCTAAAGAAGCCGCCAGACCGACATTAATGCCTCGTGCTCAGAGCTACAAAGATT
[0486] TGACCCATCTGCCAGCGCCGACGGGTAAAATCTTTGTTTCGGTATACAACATTCAGGA
[0487] CGAAACCGGGCAATTTAAACCCTACCCGGCAAGTAACTTCTCCACTGCTGTTCCGCA
[0488] AAGCGCCACGGCAATGCTGGTCACGGCACTGAAAGATTCTCGCTGGTTTATACCGCT
[0489] GGAGCGCCAGGGCTTACAAAACCTGCTTAACGAGCGCAAGATTATTCGTGCGGCAC
[0490] AAGAAAACGGCACGGTTGCCATTAATAACCGAATCCCGCTGCAATCTTTAACGGCGG
[0491] CAAATATCATGGTTGAAGGTTCGATTATCGGTTATGAAAGCAACGTCAAATCTGGCGG
[0492] GGTTGGGGCAAGATATTTTGGCATCGGTGCCGACACGCAATACCAGCTCGATCAGAT
[0493] TGCCGTGAACCTGCGCGTCGTCAATGTGAGTACCGGCGAGATCCTTTCTTCGGTGAA
[0494] CACCAGTAAGACGATACTTTCCTATGAAGTTCAGGCCGGGGTTTTCCGCTTTATTGAC
[0495] TACCAGCGCTTGCTTGAAGGGGAAGTGGGTTACACCTCGAACGAACCTGTTATGCTG
[0496] TGCCTGATGTCGGCTATCGAAACAGGGGTCATTTTCCTGATTAATGATGGTATCGACC
[0497] GTGGTCTGTGGGATTTGCAAAATAAAGCAGAACGGCAGAATGACATTCTGGTGAAAT
[0498] ACCGCCATATGTCGGTTCCACCGGAATCCTGA
[0499] SEQ ID NO:2(>P0AEA2(1:277); WT Pro-CsgG from Escherichia coli K12)
[0500] MQRLFLLVAVMLLSGCLTAPPKEAARPTLMPRAQSYKDLTHLPAPTGKIFVSVYNIQDET
[0501] GQFKPYPASNFSTAVPQSATAMLVTALKDSRWFIPLERQGLQNLLNERKIIRAAQENGTV
[0502] AINNRIPLQSLTAANIMVEGSIIGYESNVKSGGVGARYFGIGADTQYQLDQIAVNLRVVN
[0503] VSTGEILSSVNTSKTILSYEVQAGVFRFIDYQRLLEGEVGYTSNEPVMLCLMSAIETGVIF
[0504] LINDGIDRGLWDLQNKAERQNDILVKYRHMSVPPES
[0505] SEQ ID NO:3(>P0AEA2(16:277); Mature CsgG from Escherichia coli K12)
[0506] CLTAPPKEAARPTLMPRAQSYKDLTHLPAPTGKIFVSVYNIQDETGQFKPYPASNFSTAVP
[0507] QSATAMLVTALKDSRWFIPLERQGLQNLLNERKIIRAAQENGTVAINNRIPLQSLTAANIM
[0508] VEGSIIGYESNVKSGGVGARYFGIGADTQYQLDQIAVNLRVVNVSTGEILSSVNTSKTILS
[0509] YEVQAGVFRFIDYQRLLEGEVGYTSNEPVMLCLMSAIETGVIFLINDGIDRGLWDLQNK
[0510] AERQNDILVKYRHMSVPPES
[0511] SEQ ID NO:4
[0512] ATGAAACTGCTGAAAGTGGCGGCGATTGCGGCGATTGTGTTTAGCGGCAGCGCGCTG
[0513] GCGGGCGTGGTGCCGCAGTATGGCGGCGGCGGCAACCATGGCGGCGGCGGCAACAA
[0514] CAGCGGCCCGAACAGCGAACTGAACATTTATCAGTATGGCGGCGGCAACAGCGCGCT
[0515] GGCGCTGCAGACCGATGCGCGCAACAGCGATCTGACCATTACCCAGCATGGCGGCGG
[0516] CAACGGCGCGGATGTGGGCCAGGGCAGCGATGATAGCAGCATTGATCTGACCCAGCG
[0517] CGGCTTTGGCAACAGCGCGACCCTGGATCAGTGGAACGGCAAAAACAGCGAAATGA
[0518] CCGTGAAACAGTTTGGCGGCGGCAACGGCGCGGCGGTGGATCAGACCGCGAGCAAC
[0519] AGCAGCGTGAACGTGACCCAGGTGGGCTTTGGCAACAACGCGACCGCGCATCAGTAT
[0520] SEQ ID NO:5
[0521] MKLLKVAAIAAIVFSGSALAGVVPQYGGGGNHGGGGNNSGPNSELNIYQYGGGNSAL
[0522] ALQTDARNSDLTITQHGGGNGADVGQGSDDSSIDLTQRGFGNSATLDQWNGKNSEMTV
[0523] KQFGGGNGAAVDQTASNSSVNVTQVGFGNNATAHQY
[0524] SEQ ID NO:6
[0525] GVVPQYGGGGNHGGGGNNSGPNSELNIYQYGGGNSALALQTDARNSDLTITQHGGGN
[0526] GADVGQGSDDSSIDLTQRGFGNSATLDQWNGKNSEMTVKQFGGGNGAAVDQTASNSS
[0527] VNVTQVGFGNNATAHQY
[0528] SEQ ID NO:7-31 are shown in the above description of the sequence listing.
[0529] SEQ ID NO:32-42 are shown in Table 1 below.
[0530] SEQ ID NO:33-47 are shown in Table 3 below.
[0531] SEQ ID NO:48-50 are shown in Table 5 below.
[0532] SEQ ID NO:51-52 are shown in Example 2 below.
[0533] SEQ ID NO:53 ( / 5Phos / = 5'-phosphate)
[0534] / 5Phos / GGTTAAACACCCAAGCAGACGCCTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTT-3’
[0535] SEQ ID NO:54
[0536] 5'-GGCGTCTGCTTGGGTGTTTAACCT-3'
[0537] SEQ ID NO:55
[0538] 5'-GAGGCGAGCGGTCAATTTGGCGTCTGCTTGGGTGTTTAACCT-3'
[0539] SEQ ID NO:56 ( / 3CholTEG / = 3' cholesterol with a triethylene glycol spacer)
[0540] 5'-TTGACCGCTCGCCTC / 3CholTEG /
[0541] SEQ ID NOs: 57 - 70 are shown in the above description of the sequence listing.
[0542] The following examples illustrate the invention. It should be understood that although specific embodiments, specific configurations, and materials and / or molecules have been discussed herein with respect to engineered cells and methods according to the invention, various changes or modifications may be made in form and detail without departing from the scope and spirit of the invention. The following examples are provided to better illustrate specific embodiments and should not be regarded as limiting the present application. The present application is limited only by the claims.
[0543] Examples
[0544] Example 1
[0545] Materials and methods
[0546] Escherichia coli CsgG pore production
[0547] A recombinant expression vector encoding a CsgG variant nanopore with a C-terminal Strep affinity tag and an ampicillin resistance gene was transformed into chemically competent Escherichia coli cells. The cells were plated on LB agar plates containing the appropriate antibiotic for selection. A single colony from the agar plate was inoculated into LB medium containing the antibiotic and grown overnight. The culture was diluted into auto-induction medium supplemented with the necessary antibiotic and incubated at 18 °C for 68 hours. The cells were harvested by centrifugation and then lysed and extracted into 1x Bugbuster extraction reagent (Merck 70921) and 0.1% DDM. The pores were purified from the supernatant using affinity chromatography, heat treatment, and size exclusion chromatography, and oligomeric nanopores were selected as judged by SDS-PAGE.
[0548] CsgG / CsgF complex formation protocol
[0549] Prepare the CsgG-CsgF complex from the purified nanopores and chemically synthesized CsgF peptides as described above. Exchange the nanopore buffer into a pH 7.0 buffer and incubate for 1 hour at 25 °C in a 8-fold molar excess of peptide relative to the CsgG monomer. Terminate the reaction by heating at 60 °C for 15 minutes, followed by centrifugation to remove any precipitate. CsgG / CsgA complex formation protocol
[0550] Prepare the CsgG-CsgA complex from the purified nanopores and chemically synthesized CsgA peptides as described above, with or without sulfonyl fluoride modification. Exchange the nanopore buffer into a pH 7.4 buffer without reducing agent and incubate with 250-fold molar equivalents of peptide relative to the CsgG monomer. To add the morpholino chain to the CsgG-CsgA complex, buffer exchange the complex to remove any unbound CsgA peptide and then add 20-fold molar equivalents of morpholino relative to the CsgG monomer. Each morpholino chain contains 5’ BCN and is click-coupled to the C-terminal azide on the CsgA peptide.
[0551] Figure 1 and Figure 2
[0552] Add 500 ng of the complex ( Figure 1 is CsgG-CsgA, and Figure 2 is CsgG / CsgF-CsgA) and the pore control of CsgG only into individual 0.5 mL ProteinLoBind Eppendorf tubes (Fisher, 10316752) and make up to a 10 μL volume with the reaction buffer. Bring the final volume to 20 μL by adding 10 μL of 2x Laemmli buffer. Load the entire volume of each sample onto a 4%-20% TGX gel (BioRad, 5671093) run with 1x TGS buffer (Sigma, T7777). Run at 300 V for 21 minutes. To image the gel, use the Spyro Ruby (Merk, S4942) stain according to the manufacturer's instructions. Then image it on a GE Typhoon gel imager using a 450 nm laser.
[0553] Figure 3
[0554] CsgA-morpholino-modified CsgG-CsgF nanopore complexes (including CsgG-F56Q) were prepared by first incubating a large excess of a CsgA polypeptide containing sulfonyl fluoride and azide reactive moieties (SEQ ID NO: 40 in Table 1) with the CsgG-CsgF nanopore complex under appropriate buffer conditions. Buffer exchange was used to remove excess reactants. The second reaction was completed by incubating the CsgA-modified pore with a morpholino-DNA strand containing a BCN reactive moiety. The results are shown in Figure 3a , Figure 3b and Figure 3c middle.
[0555] Use ratio Figure 3a , Figure 3b and Figure 3c This process was repeated using the CsgA polypeptide used in (SEQ ID NO: 40 in Table 1) and longer CsgA polypeptides (SEQ ID NO: 43-47 in Table 3). Figure 3d Shown in.
[0556] Further iterations were performed using CsgA polypeptides with reactive SO2F chemistry at the G8 position or with altered C-terminal chemistry (SEQ ID NOs: 48-50 in Table 5). Figure 3e middle.
[0557] Electrical measurements were obtained from unmodified CsgG-CsgF nanopore complexes or CsgG-CsgF-CsgA-morpholino complexes inserted into a MinION flow cell. After inserting a single pore complex into the block copolymer membrane, 1 mL of a buffer containing 25 mM potassium phosphate, 150 mM potassium ferrocyanide (II), 150 mM potassium ferrocyanide (III), pH 8.0 was flowed through the system to remove any excess nanopore.
[0558] Analytes used to evaluate DNA profiles were from a random fragmentation library of the ZymoBIOMICS Microbial Community Standard. Analyte preparation, ligation of analytes to Y-adapters, SPRI bead cleanup of ligated analytes, and addition to the minION flow cell were performed using the Oxford Nanopore Technologies SQK-LSK114 protocol.
[0559] Electrical measurements were obtained using a GridION from Oxford Nanopore Technologies. The standard sequencing script for LSK-114 was run for 20 hours. Raw data were collected in batch FAST5 files using MinKNOW software (Oxford Nanopore Technologies).
[0560] Figure 4
[0561] To 80 ul of the CsgG pore (0.15 mg / ml, 0.4 nmol, in PBS buffer containing 0.1% SDS, pH 7.4), the ligated CsgA polypeptide (SEQ ID NO:69 and 70; see legend; 20 mg / ml in DMSO, 100-fold molar excess) was added. The sample was incubated for 14 h with shaking (25 °C, 600 rpm). Subsequently, 12 x 5 ul aliquots were taken and heated on a PCRmax Alpha Cycler 4 (Stone, UK) for 30 min in a temperature gradient ranging from 62.0 °C to 87.8 °C.
[0562] After heating, each 5 ul sample was mixed with 10 ul of native sample buffer (Bio-Rad, Hercules, USA). The samples were analyzed by blue native PAGE gel (4-20% TGX precast gel, Bio-Rad) in 1x Tris-glycine buffer (Sigma-Aldrich, Burlington, USA). Initially, the gel was run at 180 V for 30 min with 1x NativePage TM Cathode buffer additive (Invitrogen, Waltham, USA) in the cathode chamber. Subsequently, it was replaced with 1x TG buffer and the gel was run at 180 V for an additional 60 min. Unstained NativeMark TM Protein standard (ThermoFisher, Waltham, USA) was present in the leftmost lane. The gel was stained with Quick Coumassie (Generon, Houston, USA) and imaged on a Bio-Rad Gel Doc EZ system.
[0563] Figure 5
[0564] In a further example, a CsgA-CsgG-CsgF variant nanopore complex was prepared with SEQ ID NO:40. Electrical measurements were obtained from unmodified CsgG-CsgF nanopore complexes or CsgG-CsgF-CsgA-morpholino complexes inserted into a MinION flow cell. After inserting a single pore complex into a block copolymer membrane, 1 mL of a buffer containing 25 mM potassium phosphate, 150 mM potassium ferrocyanide (II), 150 mM potassium ferricyanide (III), pH 8.0 was passed through the system to remove any excess nanopores.
[0565] The analytes used to evaluate the DNA curves were from a randomly fragmented library of the ZymoBIOMICS microbial community standard. The analytes were prepared, ligated to Y adapters, purified with SPRI beads for the ligated analytes, and added to the MinION flow cell using the Oxford Nanopore Technologies SQK-LSK114 protocol.
[0566] Before adding the analytes to the MinION flow cell, the standard flow cell preparation wash was adjusted to contain different concentrations of morpholino competitor molecules, and the results are shown in Figure 5 .
[0567] An Oxford Nanopore Technologies GridION was used to obtain electrical measurements. The standard sequencing script for LSK-114 was run and the raw data was collected in batch FAST5 files using MinKNOW software (Oxford Nanopore Technologies).
[0568] Results
[0569] The results are shown in Figures 1 - 4 .
[0570] Figure 1
[0571] Figure 1 The gels in
[0572]
[0573] Figure 1 show the results for the following CsgA polypeptides (Table 1; SO2F is a sulfonyl fluoride group and N3 is an azide group)
[0574]
[0575]
[0576] Figure 2
[0577] Figure 2 The gels in
[0578]
[0579] Figure 2 show the results for the following CsgA polypeptides (Table 3; SO2F is a sulfonyl fluoride group and N3 is an azide group)
[0580]
[0581]
[0582] Figure 3e
[0583] The results of the following CsgA polypeptides are as Figure 3e shown (Table 5).
[0584]
[0585] Figure 4
[0586] Figure 4 key to the gel in
[0587] 1. 62.0 °C
[0588] 2. 63.0 °C
[0589] 3. 65.5 °C
[0590] 4. 68.5 °C
[0591] 5. 71.1 °C
[0592] 6. 73.7 °C
[0593] 7. 76.4 °C
[0594] 8. 79.2 °C
[0595] 9. 81.9 °C
[0596] 10. 84.8 °C
[0597] 11. 86.9 °C
[0598] 12. 87.8 °C
[0599] Example 2
[0600] This example demonstrates the attachment of a polynucleotide-binding protein to a nanopore via a CsgA peptide, which enables polynucleotide-binding protein-regulated DNA translocation through the nanopore.
[0601] Functionalized CsgG nanopores are prepared by reacting a nanopore protein with an excess of modified CsgA peptide. The peptide has two modifications: an internal sulfonyl fluoride group, which reacts with the nanopore to form a covalent bond; and an azide group, which reacts with the BCN-modified morpholino oligomer 5'-GCAATACGTAACTGAACGAAGTACA-3' (SEQ ID NO:51).
[0602] A tethered polynucleotide-binding protein was prepared by functionalizing the N-terminus of the Hel308 helicase with an azide group and then linking a BCN-modified morpholino oligomer 5'-TGTACTTCGTTCAGTTACGTATTGC-3' (SEQ ID NO:52) via a click reaction of BCN with the azide group. The morpholino oligomer has a sequence complementary to the sequence of the morpholino oligomer in the functionalized nanopore.
[0603] It was confirmed by blue native electrophoresis that the tethered polynucleotide-binding protein was able to link to the functionalized nanopore. The tethered polynucleotide-binding protein was mixed with different excess concentrations of the functionalized tethered polynucleotide-binding protein and incubated at room temperature for 15 minutes. The resulting product was loaded into a 4–20% Criterion TM TGX TM precast Midi protein gel (Bio-Rad) after mixing with the native sample buffer for protein gels (Bio-Rad). In the Criterion TM cell (Bio-Rad), at 180 V, using NativePAGE TM cathode buffer additive (20X) diluted to 1X in 1x Tris-glycine buffer (Sigma-Aldrich) as the cathode buffer and 1x Tris-glycine buffer (Sigma-Aldrich) as the anode buffer, the gel was run for 30 minutes. After this period, the cathode buffer was replaced with 1x Tris-glycine buffer (Sigma-Aldrich). Electrophoresis was continued at 180 V for 60 minutes. The gel was stained with Quick Coomassie Dye (Neo-Biotech). The results are shown in Figure 7 it.
[0604] A 3'-overhang adaptor was prepared by annealing DNA oligonucleotides SEQ ID NO:53 and SEQ ID NO:54. A Y-adaptor was prepared by annealing DNA oligonucleotides SEQ ID NO:53 and SEQ ID NO:55. A 3.6-kilobase DNA analyte was obtained from bacteriophage λ via PCR amplification and then end-repaired and dA-tailed using the Ultra II End Repair and dA-Tailing Kit (New England Biolabs) to generate 3'dA overhangs at both ends of each fragment. The samples were ligated to the T overhangs of the 3'-overhang adaptor or the Y-adaptor using LNB (ONT) and T4 DNA ligase. The adaptor-ligated samples were then purified using Agencourt AMPure XP (Beckman Coulter) beads to generate a 3'-overhang-adapted DNA library and a Y-adaptor-adapted DNA library.
[0605] Electrical measurements were obtained using a custom MinION flow cell with functionalized nanopores inserted in its membrane, and data were collected using a GridION X5. An 800 μL volume of SEQ ID NO:56 and FCF(ONT) were flowed through the system at a ratio of 1:40, and then waited for 5 minutes, and then another 200 μL were flowed through the system with the drip port open. The drip port was closed, and a polynucleotide-binding protein wash mixture was prepared by combining 75 μL of SB(ONT), 45 μL of nuclease-free water, and 30 μL of tethered polynucleotide-binding protein (final 800 nM). A control mixture was prepared by combining 75 μL of SB(ONT), 45 μL of nuclease-free water, and 30 μL of unmodified polynucleotide-binding protein (final 800 nM). A total of 150 μL were flowed into each flow cell and incubated for 30 minutes. 500 μL of FCF(ONT) were flowed into each flow cell to displace the polynucleotide-binding protein not paired with the complementary morpholino on the functionalized nanopore. After 5 minutes, an additional 500 μL were flowed into each flow cell, and the drip port was opened. A first sequencing mixture was prepared by combining 37.5 μL of SB(ONT), 11.6 μL of nuclease-free water, 15.0 μL of buffer (consisting of 50 mM HEPES pH 7, 384 mM NaCl, and 5% glycerol), and 10.9 μL of 3'-overhang-adapted DNA library (final 3.56 nM). 75 μL of this sequencing mixture were added to each flow cell through the drip port. MinKNOW(ONT) software was used to obtain current measurements.
[0606] After three hours, a second sequencing mixture was prepared by combining 37.5 μL SB(ONT), 11.8 μL LS(ONT), 15.0 μL buffer (consisting of 50 mM HEPES pH 7, 384 mM NaCl, and 5% glycerol), and 10.7 μL of Y-adapted DNA library (final 3.56 nM). 75 μL of the second sequencing mixture were added to each flow cell through the drip port. After adding the second sequencing mixture, data were collected for another 5 hours. The results are shown in Figure 8 and Figure 9 in.
[0607] Example
[0608] 1. A pore monomer conjugate comprising a pore monomer, a chaperone molecule, and a functional binding moiety, wherein the chaperone molecule has an affinity for the pore monomer, and wherein the functional binding moiety is linked to the pore monomer via the chaperone molecule.
[0609] 2. The pore monomer conjugate according to embodiment 1, wherein the chaperone molecule comprises a chaperone polypeptide or a chaperone protein.
[0610] 3. The pore monomer conjugate according to embodiment 1 or 2, wherein the pore monomer and the chaperone molecule are selected from (a) a CsgG pore monomer and a CsgA polypeptide, (b) a CsgG pore monomer and a CsgB polypeptide, (c) a CsgG pore monomer and a CsgC polypeptide, (d) a CsgG pore monomer and a CsgD polypeptide, or (e) a CsgG pore monomer and a CsgE polypeptide.
[0611] 4. The pore monomer conjugate according to any one of the foregoing embodiments, wherein the pore monomer is a CsgG pore monomer and wherein the chaperone molecule comprises a CsgA polypeptide.
[0612] 5. The pore monomer conjugate according to any one of the foregoing embodiments, wherein the functional binding moiety is covalently linked to the chaperone molecule, optionally via a linker, covalently linked to the chaperone molecule.
[0613] 6. The pore monomer conjugate according to any one of the foregoing embodiments, wherein the functional binding moiety is capable of binding to a target analyte.
[0614] 7. The pore monomer conjugate according to embodiment 6, wherein the target analyte comprises a polypeptide, a protein, an oligonucleotide, a polynucleotide, a polynucleotide - polypeptide conjugate, an oligosaccharide, or a polysaccharide.
[0615] 8. The pore monomer conjugate according to embodiment 7, wherein the functional binding moiety comprises an oligonucleotide, a polynucleotide, a polynucleotide analog, a morpholino, a peptide nucleic acid, a polypeptide, a ligand, a cyclodextrin, a monosaccharide, an oligosaccharide, a polysaccharide, boric acid, an enzyme, a peptide, a cyclic peptide as a functional, an antibody or a fragment thereof, or an aptamer.
[0616] 9. The pore monomer conjugate according to any one of embodiments 1 to 5, wherein the functional binding moiety is capable of binding to a polynucleotide - binding protein, a polypeptide - binding protein, a pore monomer, an aptamer, a cyclic protein, or a DNA origami structure.
[0617] 10. The pore monomer conjugate according to any one of the foregoing embodiments, wherein the chaperone molecule is covalently linked to the pore monomer, optionally via a linker, covalently linked to the pore monomer.
[0618] 11. A construct comprising two or more covalently - linked pore monomer conjugates according to any one of embodiments 1 to 10.
[0619] 12. A pore complex comprising at least one pore monomer conjugate according to any one of embodiments 1 to 10 or at least one construct according to embodiment 11.
[0620] 13. A pore multimer comprising two or more pores, wherein at least one of the pores is a pore complex according to embodiment 12.
[0621] 14. The pore complex according to embodiment 12 or the pore multimer according to embodiment 13, which is comprised in a membrane.
[0622] 15. A membrane comprising the pore complex according to embodiment 12 or the pore multimer according to embodiment 13.
[0623] 16. A method for attaching a functional binding moiety to a pore monomer using a chaperone molecule having an affinity for the pore monomer as a linker.
[0624] 17. A method for determining the presence, absence or one or more characteristics of a target analyte, the method comprising the steps of:
[0625] (i) contacting the target analyte with the pore complex according to embodiment 12 or the pore multimer according to embodiment 13 such that the target analyte moves relative to the pore complex or pore multimer; and
[0626] (ii) making one or more measurements while the analyte moves relative to the pore complex or pore multimer, and thereby determining the presence, absence or one or more characteristics of the target analyte.
[0627] 18. The method according to embodiment 17, wherein the target analyte is as defined in embodiment 7.
[0628] 19. The method according to embodiment 17 or 18, wherein the functional binding moiety interacts with or binds to the target analyte and facilitates its movement relative to the pore complex or multimer.
[0629] 20. The method according to embodiment 19, wherein (a) the target analyte is a polynucleotide and the functional binding moiety comprises an oligonucleotide, polynucleotide, polynucleotide analogue or morpholino capable of hybridizing to the target polynucleotide, (b) the target analyte is a polynucleotide and the functional binding moiety is attached to a polynucleotide-binding protein that controls the movement of the target analyte relative to the pore complex or pore multimer, (c) the target analyte is a ligand of an enzyme and the functional binding moiety is attached to the enzyme, or (d) the functional binding moiety is an antibody or a functional fragment thereof or an aptamer that binds to the target analyte.
[0630] 21. The method according to embodiment 17 or 18, wherein (i) the functional binding moiety is attached to a pore, aptamer, cyclophilin, or DNA origami structure to increase the distance between the polynucleotide-binding protein and the pore complex or pore multimer or to provide one or more additional constrictions, (ii) the functional binding moiety binds to the pore complex or pore multimer to stabilize the cis-loop and reduce the signal-to-noise ratio of the pore complex or pore multimer, or (iii) the functional binding moiety links the pore complex or pore multimer to a second pore or second pore multimer.
[0631] 22. A method of characterizing a target analyte using the pore complex according to embodiment 12 or the pore multimer according to embodiment 13.
[0632] 23. Use of the pore complex according to embodiment 12 or the pore multimer according to embodiment 13 for determining the presence, absence, or one or more characteristics of a target analyte.
[0633] 24. A kit for characterizing a target analyte, the kit comprising (a) the pore complex according to embodiment 12 or the pore multimer according to embodiment 13, and (b) components of a membrane.
[0634] 25. The kit according to embodiment 24, wherein the kit further comprises an enzyme, a pore, a pore multimer, an aptamer, a cyclophilin, or a DNA origami structure.
[0635] 26. A kit for characterizing a target polynucleotide or target polypeptide, the kit comprising (a) the pore complex according to embodiment 12 or the pore multimer according to embodiment 13, and (b) a polynucleotide-binding protein or a polypeptide-binding protein.
[0636] 27. An apparatus for characterizing a target polynucleotide or target polypeptide in a sample, the apparatus comprising (a) a plurality of the pore complexes according to embodiment 12 or a plurality of the pore multimers according to embodiment 13, and (b) a plurality of polynucleotide-binding proteins or a plurality of polypeptide-binding proteins.
[0637] 28. The kit according to embodiment 26 or the apparatus according to embodiment 27, wherein the functional binding moiety is capable of binding to the polynucleotide-binding protein or the polypeptide-binding protein.
[0638] 29. An array comprising a plurality of the membranes according to embodiment 15.
[0639] 30. A system comprising (a) the membrane according to embodiment 15 or the array according to embodiment 29, (b) means for applying an electric potential across the membrane, and (c) means for detecting an electrical or optical signal across the membrane.
[0640] 31. An apparatus comprising a pore complex according to embodiment 12 or a pore polymer according to embodiment 13 inserted into an extracorporeal membrane.
[0641] 32. An apparatus produced by a method comprising: (i) obtaining a pore complex according to embodiment 12 or a pore polymer according to embodiment 13; and (ii) contacting the pore complex or pore polymer with an extracorporeal membrane such that the pore complex or pore polymer is inserted into the extracorporeal membrane.
[0642] 33. A method for determining the presence, absence or one or more characteristics of a target polynucleotide, the method comprising the steps of: (a) contacting a double-stranded polynucleotide comprising a template and a complementary strand with a pore complex or pore polymer comprising at least one pore monomer conjugate according to any one of embodiments 1 to 10, wherein the complementary strand comprises a binding region capable of hybridizing with a functional binding moiety such that when the template strand moves through the pore complex or pore polymer the two strands separate to expose the binding region on the complementary strand, wherein the functional binding moiety hybridizes with the binding region to facilitate capture of the complementary strand through the pore complex or pore polymer; and (b) making one or more measurements as the template strand and the complementary strand move relative to the pore complex or pore polymer, wherein measurements of both the template strand and the complementary strand are used to determine the presence, absence or one or more characteristics of the target polynucleotide.
Claims
1. A pore monomer conjugate comprising a CsgG pore monomer, a chaperone molecule, and a functional binding moiety, wherein the chaperone molecule comprises a polypeptide selected from the group consisting of CsgA polypeptide, CsgB polypeptide, CsgC polypeptide, CsgD polypeptide, and CsgE polypeptide, and wherein the functional binding moiety is linked to the pore monomer via the chaperone molecule.
2. The pore monomer conjugate according to claim 1, wherein the chaperone molecule comprises a polypeptide selected from the group consisting of CsgA polypeptide, CsgB polypeptide, and CsgE polypeptide.
3. The pore monomer conjugate according to claim 1 or 2, wherein the chaperone molecule comprises a CsgA polypeptide.
4. The pore monomer conjugate according to any one of the preceding claims, wherein the functional binding moiety is covalently linked to the chaperone molecule, optionally via a linker.
5. The pore monomer conjugate according to any one of the preceding claims, wherein the functional binding moiety is capable of binding a target analyte.
6. The pore monomer conjugate according to claim 5, wherein the target analyte comprises a polypeptide, a protein, an oligonucleotide, a polynucleotide, a polynucleotide - polypeptide conjugate, an oligosaccharide, or a polysaccharide.
7. The pore monomer conjugate according to claim 6, wherein the functional binding moiety comprises an oligonucleotide, a polynucleotide, a polynucleotide analog, a morpholino, a peptide nucleic acid, a polypeptide, a ligand, a cyclodextrin, a monosaccharide, an oligosaccharide, a polysaccharide, boric acid, an enzyme, a peptide, a cyclic peptide as a functional, an antibody or a fragment thereof, or an aptamer.
8. The pore monomer conjugate according to any one of claims 1 to 4, wherein the functional binding moiety is capable of binding a polynucleotide - binding protein, a polypeptide - binding protein, a pore monomer, an aptamer, a cyclic protein, or a DNA origami structure.
9. The pore monomer conjugate according to any one of the preceding claims, wherein the chaperone molecule is covalently linked to the pore monomer, optionally via a linker.
10. A construct comprising two or more covalently linked pore monomer conjugates according to any one of claims 1 to 9.
11. A pore complex comprising at least one pore monomer conjugate according to any one of claims 1 to 9 or at least one construct according to claim 10.
12. A pore multimer comprising two or more pores, wherein at least one of the pores is a pore complex according to claim 11.
13. The pore complex according to claim 11 or the pore multimer according to claim 12, which is incorporated in a membrane.
14. A membrane comprising the pore complex according to claim 11 or the pore multimer according to claim 12.
15. A method for linking a functional binding moiety to a pore monomer using a chaperone molecule having an affinity for the pore monomer as a linker.
16. A method for determining the presence, absence, or one or more characteristics of a target analyte, the method comprising the steps of: (i) contacting the target analyte with the pore complex according to claim 11 or the pore polymer according to claim 12 such that the target analyte moves relative to the pore complex or the pore polymer; and (ii) making one or more measurements as the analyte moves relative to the pore complex or the pore polymer and thereby determining the presence, absence or one or more characteristics of the target analyte.
17. The method according to claim 16, wherein the target analyte is as defined in claim 6.
18. The method according to claim 16 or 17, wherein the functional binding moiety interacts with or binds to the target analyte and facilitates its movement relative to the pore complex or polymer.
19. The method according to claim 18, wherein (a) the target analyte is a polynucleotide and the functional binding moiety comprises an oligonucleotide, polynucleotide, polynucleotide analogue or morpholino capable of hybridizing to the target polynucleotide, (b) the target analyte is a polynucleotide and the functional binding moiety is linked to a polynucleotide-binding protein that controls the movement of the target analyte relative to the pore complex or pore polymer, (c) the target analyte is a ligand of an enzyme and the functional binding moiety is linked to the enzyme, or (d) the functional binding moiety is an antibody or a functional fragment or aptamer thereof that binds to the target analyte.
20. The method according to claim 16 or 17, wherein (i) the functional binding moiety is linked to a pore, aptamer, cyclophilin or DNA origami structure to increase the distance between the polynucleotide-binding protein and the pore complex or pore polymer or to provide one or more additional constrictions, (ii) the functional binding moiety binds to the pore complex or pore polymer to stabilize the cis-loop and reduce the signal-to-noise ratio of the pore complex or pore polymer, or (iii) the functional binding moiety links the pore complex or pore polymer to a second pore or second pore polymer.
21. A method of characterizing a target analyte using the pore complex according to claim 11 or the pore polymer according to claim 12.
22. Use of the pore complex according to claim 11 or the pore polymer according to claim 12 for determining the presence, absence or one or more characteristics of a target analyte.
23. A kit for characterizing a target analyte, the kit comprising (a) the pore complex according to claim 11 or the pore polymer according to claim 12, and (b) components of a membrane.
24. The kit according to claim 23, wherein the kit further comprises an enzyme, a pore, a pore polymer, an aptamer, a cyclophilin or a DNA origami structure.
25. A kit for characterizing a target polynucleotide or a target polypeptide, the kit comprising (a) a pore complex according to claim 11 or a pore multimer according to claim 12, and (b) a polynucleotide-binding protein or a polypeptide-binding protein.
26. An apparatus for characterizing a target polynucleotide or a target polypeptide in a sample, the apparatus comprising (a) a plurality of pore complexes according to claim 11 or a plurality of pore multimers according to claim 12, and (b) a plurality of polynucleotide-binding proteins or a plurality of polypeptide-binding proteins.
27. The kit according to claim 25 or the apparatus according to claim 26, wherein the functional binding moiety is capable of binding to the polynucleotide-binding protein or the polypeptide-binding protein.
28. An array comprising a plurality of membranes according to claim 14.
29. A system comprising (a) a membrane according to claim 14 or an array according to claim 28, (b) means for applying an electric potential across the membrane, and (c) means for detecting an electrical signal or an optical signal across the membrane.
30. An apparatus comprising a pore complex according to claim 11 or a pore multimer according to claim 12 inserted into an in vitro membrane.
31. An apparatus produced by a method comprising: (i) obtaining a pore complex according to claim 11 or a pore multimer according to claim 12, and (ii) contacting the pore complex or pore multimer with an in vitro membrane such that the pore complex or the pore multimer is inserted into the in vitro membrane.
32. A pore monomer conjugate comprising a CsgG pore monomer, a chaperone molecule, and a functional binding moiety, wherein the CsgG pore monomer comprises a sequence that is at least about 40% homologous or identical to the amino acid sequence of SEQ ID NO:3 over the entire sequence, wherein the chaperone molecule comprises SEQ ID NO:21, 23, 25, 27, 29, 31, 63, 64, 65, 66, 67, or 68, wherein the K at position 7 of SEQ ID NO:21, 23, 25, 27, 29, or 31 is covalently linked to the CsgG pore monomer via a sulfonyl group, or wherein the K at position 8 of SEQ ID NO:63, 64, 65, 66, 67, or 68 is covalently linked to the CsgG pore monomer via a sulfonyl group, wherein the functional binding moiety comprises an oligonucleotide, a polynucleotide, a polynucleotide analogue, or a morpholino that is capable of specifically hybridizing to a target polynucleotide analyte, and wherein the functional binding moiety is covalently linked to the C-terminal K of SEQ ID NO:21, 23, 25, 27, 29, 31, 63, 64, 65, 66, 67, or 68.
33. The pore monomer according to claim 32, wherein the chaperone molecule comprises SEQ ID NO:
21.
34. A method for determining the presence, absence, or one or more characteristics of a target polynucleotide, the method comprising the steps of: (a) contacting a double-stranded polynucleotide comprising a template strand and a complement strand with a pore complex or pore polymer comprising at least one pore monomer conjugate according to any one of claims 1 to 9, 32, and 33, wherein the complement strand comprises a binding region capable of hybridizing with the functional binding moiety such that when the template strand moves through the pore complex or pore polymer, the two strands separate to expose the binding region on the complement strand, and wherein the functional binding moiety hybridizes with the binding region to facilitate capture of the complement strand through the pore complex or pore polymer; and (b) performing one or more measurements as the template strand and the complement strand move relative to the pore complex or pore polymer, wherein the measurements of both the template strand and the complement strand are used to determine the presence, absence, or one or more characteristics of the target polynucleotide.
Citation Information
Patent Citations
Mutant of pore protein monomer, protein pore and application thereof
CN113754743A
Mutant of pore protein monomer, protein pore and application of mutant
CN113773373A
Mutant of pore protein monomer, protein pore and application thereof
CN113896776A
Mutant of pore protein monomer, protein pore, application of mutant and application of protein pore
CN113912683A
Pore
GB202118939D0