Methods for characterizing polypeptides using nanopores - Patents.com

JP2025500398A5Pending Publication Date: 2026-01-06OXFORD NANOPORE TECH LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024537831
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-12-23
Filing Date
2022-12-22
Publication Date
2026-01-06

AI Technical Summary

Technical Problem

Existing methods for characterizing polypeptides, such as mass spectrometry and Edman degradation, are limited by contamination, fragmentation of fragile molecules, and lack of single-molecule analysis, leading to inefficiencies and loss of information about heterogeneity in polypeptide samples.

Method used

A method involving the formation of a polypeptide-polynucleotide conjugate, controlled by a polynucleotide handling protein, which moves through a nanopore detector to enable accurate, single-molecule characterization of polypeptides by repeated binding and rebinding, allowing for detailed measurements of polypeptide properties.

Benefits of technology

This approach allows for rapid, accurate characterization of polypeptides with improved data accuracy and efficiency, overcoming limitations of existing techniques by enabling precise measurement of polypeptide length, identity, sequence, and modifications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Provided herein are methods for characterizing a target polypeptide as it translocates relative to a nanopore, as well as associated kits, systems, and devices for carrying out such methods.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present disclosure relates to methods for characterizing a target polypeptide by forming a conjugate between the target polypeptide and a polynucleotide and controlling the translocation of the conjugate relative to a nanopore using a polynucleotide handling protein. The present disclosure also relates to kits, systems, and devices for carrying out such methods. [Background technology]

[0002] Characterization of biological molecules is becoming increasingly important in biomedical and bioengineering applications. For example, sequencing of nucleic acids allows the study of genomes and the proteins they encode, allowing for associations between, for example, nucleic acid mutations and observable phenomena such as disease manifestations. Nucleic acid sequencing can be used in evolutionary biology to study relationships between organisms. Metagenomics involves identifying the organisms present in a sample, for example, microorganisms in a microbiome, and nucleic acid sequencing allows the identification of such organisms. While techniques to characterize (e.g., sequence) polynucleotides have been widely developed, techniques to characterize polypeptides have been less advanced, despite their enormous biotechnological importance. For example, knowledge of protein sequences allows the establishment of structure-activity relationships, which has influenced rational drug development strategies to develop ligands for specific receptors. Identification of post-translational modifications is also key to understanding the functional properties of many proteins. For example, in eukaryotes, typically 30-50% of protein species are phosphorylated. Some proteins may have multiple phosphorylation sites that serve to activate or inactivate the protein, promote its degradation, or modulate interactions with protein partners. Thus, there is a pressing need for methods for characterizing proteins and other polypeptides.

[0003] Known methods for characterizing polypeptides include mass spectrometry and Edman degradation.

[0004] Protein mass spectrometry involves characterizing whole proteins or fragments thereof in ionized form. Known methods of protein mass spectrometry include electrospray ionization (ESI) and matrix-assisted laser desorption / ionization (MALDI). Although mass spectrometry has several advantages, the results obtained can be affected by the presence of contaminants and fragile molecules can be difficult to process without fragmentation. Furthermore, mass spectrometry is not a single molecule technique and provides only bulk information about the investigated sample. Mass spectrometry is not suitable for characterizing differences within a population of polypeptide samples and is cumbersome when trying to distinguish adjacent residues.

[0005] Edman degradation is an alternative to mass spectrometry that allows for residue-by-residue sequencing of polypeptides. Edman degradation sequences a polypeptide by sequentially cleaving the N-terminal amino acids and then characterizing the individually cleaved residues using chromatography or electrophoresis. However, Edman sequencing is slow, requires the use of expensive reagents, and is not a single-molecule technique like mass spectrometry.

[0006] Thus, there remains a pressing need for new techniques for characterizing polypeptides, especially at the single molecule level. Single molecule techniques for characterizing biomolecules such as polynucleotides have proven particularly attractive due to their high fidelity and avoidance of amplification bias.

[0007] One attractive method for single molecule characterization of biomolecules such as polypeptides is nanopore sensing. Nanopore sensing is an approach to analyte detection and characterization that relies on the observation of individual binding or interaction events between an analyte molecule and an ion-conducting channel. Nanopore sensors can be fabricated by placing a single pore of nanometer dimensions in an insulating membrane and measuring the voltage-driven ionic current through the pore in the presence of an analyte molecule. The presence of an analyte in or near the nanopore will alter the ionic flow through the pore, resulting in a change in the ionic or electrical current measured across the channel. The identity of the analyte is revealed through its unique current signature, particularly the duration and extent of current block, and the variation in current level during the interaction time with the pore. Nanopore sensing has the potential to enable rapid and inexpensive polypeptide characterization.

[0008] Nanopore sensing and characterization of polypeptides have been proposed in the art. For example, WO 2013 / 123379 discloses the use of NTP-driven protein processing unfoldase enzymes to process proteins and translocate them through a nanopore. WO 2021 / 111125 discloses a method for characterizing a target polypeptide using a nanopore. WO 2021 / 133168 discloses a method for fingerprinting and sequencing proteins and peptides using a nanopore. WO 2018 / 064078 discloses the translocation of non-nucleic acid polymers using a polymerase. WO 2021 / 000786 discloses a method for controlling the rate of a polypeptide passing through a nanopore. However, there remains a need for alternative and / or improved methods of characterizing polypeptides.

[0009] In particular, there is also a need for a method to improve the data obtained when characterizing a polypeptide. One problem is that in some cases, it is desirable to improve the accuracy of the characterization data obtained when characterizing a polypeptide. In some known methods, multiple polypeptides are characterized from a sample of polypeptides, and the data obtained are aggregated, thereby improving the overall accuracy. However, this can cause problems. For example, heterogeneity in the sample can mean that when aggregating data obtained from the characterization of multiple polynucleotide strands, useful information about the differences between the strands can be lost. Furthermore, inefficiencies can arise due to the need to capture new strands for characterization after the first strands have been processed. Therefore, there is a need for alternative and / or improved methods of characterizing polypeptides. Summary of the Invention

[0010] The present disclosure relates to a method for characterizing a target polypeptide. The target polypeptide is conjugated to a polynucleotide in a polypeptide-polynucleotide conjugate. The method includes contacting the conjugate with a polynucleotide handling protein. The polynucleotide handling protein can control the movement of the polynucleotide relative to a detector, such as a nanopore. When the conjugate moves in a first direction relative to the detector, one or more measurements characteristic of the polypeptide are performed. The conjugate is unbound from the polynucleotide binding site of the polynucleotide handling protein, and the conjugate moves in a second direction relative to the detector. The conjugate then rebinds to the polynucleotide handling protein. When the conjugate moves in the first direction relative to the detector, one or more measurements characteristic of the polypeptide are performed. In this way, the target polypeptide contained in the conjugate is characterized.

[0011] The conjugate may have a leader attached to the polypeptide and / or the polynucleotide portion of the conjugate may include a blocking moiety, these features are described in more detail herein.

[0012] Thus, provided herein is a method of characterizing a target polypeptide having a first end and a second end, the target polypeptide being included in a conjugate in which the first end of the target polypeptide is conjugated to a first end of a polynucleotide, the method comprising: (i) contacting a detector having a first opening and a second opening, or contained in a structure having a first opening and a second opening, with a conjugate having a polynucleotide handling protein bound thereto, wherein the conjugate is bound to the polynucleotide handling protein at a polynucleotide binding site of the polynucleotide handling protein; (ii) taking one or more measurements characteristic of the target polypeptide when the polynucleotide handling protein controls movement of the conjugate in a first direction relative to a detector; (iii) debinding the conjugate from the polynucleotide binding site of the polynucleotide handling protein such that the conjugate moves in a second direction relative to the detector; (iv) rebinding the conjugate to the polynucleotide binding site of the polynucleotide handling protein and performing one or more measurements characteristic of the target polypeptide as the polynucleotide handling protein controls movement of the conjugate in a first direction relative to a detector; and characterizing the target polypeptide, The second end of the target polypeptide is attached to a leader.

[0013] In some embodiments, prior to step (i), the polynucleotide handling protein is bound to the portion of the conjugate between the polypeptide and the free end of the leader. In some embodiments, the leader is attached to the polypeptide by a polynucleotide linker, and prior to step (i), the polynucleotide handling protein is bound to the linker.

[0014] In some embodiments, the second end of the polynucleotide comprises a blocking moiety to prevent the polynucleotide handling protein from disassociating from the conjugate, hi some embodiments, prior to step (i), the polynucleotide handling protein is bound to a portion of the conjugate between the polypeptide and the blocking moiety.

[0015] Also provided herein is a method of characterizing a target polypeptide having a first end and a second end, the target polypeptide being included in a conjugate, wherein the first end of the target polypeptide is conjugated to a first end of a polynucleotide, the method comprising: (i) contacting a detector having a first opening and a second opening, or contained in a structure having a first opening and a second opening, with a conjugate having a polynucleotide handling protein bound thereto, wherein the conjugate is bound to the polynucleotide handling protein at a polynucleotide binding site of the polynucleotide handling protein; (ii) taking one or more measurements characteristic of the target polypeptide when the polynucleotide handling protein controls movement of the conjugate in a first direction relative to a detector; (iii) debinding the conjugate from the polynucleotide binding site of the polynucleotide handling protein such that the conjugate moves in a second direction relative to the detector; (iv) rebinding the conjugate to the polynucleotide binding site of the polynucleotide handling protein and performing one or more measurements characteristic of the target polypeptide as the polynucleotide handling protein controls movement of the conjugate in a first direction relative to a detector; and characterizing the target polypeptide, The second end of the polynucleotide includes a blocking moiety to prevent disassociation of the polynucleotide handling protein from the conjugate.

[0016] In some embodiments, prior to step (i), the polynucleotide handling protein is bound to part of the conjugate between the polypeptide and the blocking moiety.

[0017] In some embodiments, the second end of the polypeptide is attached to a leader. In some embodiments, prior to step (i), a polynucleotide handling protein is attached to a portion of the conjugate between the polypeptide and the free end of the leader. In some embodiments, the leader is attached to the polypeptide by a polynucleotide linker, and prior to step (i), the polynucleotide handling protein is attached to the linker.

[0018] In some embodiments of the methods provided herein, steps (iii) and (iv) are repeated multiple times, hi some embodiments, steps (iii) and (iv) are repeated at least 2 times, at least 3 times, at least 4 times, at least 5 times, at least 10 times, at least 20 times, at least 50 times, at least 100 times, at least 500 times, at least 1000 times, or more.

[0019] In some embodiments, step (i) comprises contacting the first opening with the leader under conditions such that the leader passes through the first and second openings.

[0020] In some embodiments, in step (iii), the conjugate moves in a direction from the first opening to the second opening, and in steps (ii) and (iv), the polynucleotide handling protein controls movement of the conjugate in a direction from the second opening to the first opening.

[0021] In some embodiments, the detector comprises a transmembrane nanopore spanning a membrane having a cis side and a trans side, a first opening of the nanopore is on the cis side of the membrane and a second opening of the nanopore is on the trans side, a polynucleotide handling protein controls movement of a conjugate through the nanopore in a direction from the trans side to the cis side of the membrane, and when the conjugate unbinds from the polynucleotide binding site of the polynucleotide handling protein, the conjugate moves through the nanopore in a direction from the cis side to the trans side of the membrane.

[0022] In some embodiments, the detector comprises a transmembrane nanopore spanning a membrane having a cis side and a trans side, a first opening of the nanopore is on the trans side of the membrane and a second opening of the nanopore is on the cis side, a polynucleotide handling protein controls movement of a conjugate through the nanopore in a direction from the cis side to the trans side of the membrane, and when the conjugate unbinds from the polynucleotide binding site of the polynucleotide handling protein, the conjugate moves through the nanopore in a direction from the trans side to the cis side of the membrane.

[0023] In some embodiments, the conjugate does not disassociate from the polynucleotide handling protein.

[0024] In some embodiments, the polynucleotide handling protein is modified to prevent the conjugate from disassociating from the polynucleotide handling protein.

[0025] In some embodiments, the polynucleotide handling protein is modified with a closing moiety to topologically close the polynucleotide binding site of the polynucleotide handling protein around the conjugate. In some embodiments, the polynucleotide handling protein is modified to facilitate attachment of a closing moiety to the polynucleotide handling protein. In some embodiments, the polynucleotide handling protein is modified by substituting a cysteine ​​or a non-natural amino acid for at least one amino acid in the polynucleotide handling protein. In some embodiments, the closing moiety comprises a bifunctional crosslinker. In some embodiments, the closing moiety comprises a bond, optionally a disulfide bond.

[0026] In some embodiments, the closing moiety comprises a structure of the formula [ABC], where A and C are each independently reactive functional groups for reacting with an amino acid residue in a polynucleotide handling protein, and B is a linking moiety. In some embodiments, (a) A and C are each independently a cysteine-reactive functional group, and / or (b) linking moiety B comprises a linear or branched, unsubstituted or substituted alkylene, alkenylene, alkynylene, arylene, heteroarylene, carbocyclylene, or heterocyclylene moiety, which is optionally interrupted and / or terminated by one or more atoms or groups selected from O, N(R), S, C(O), C(O)NR, C(O)O, unsubstituted or substituted arylene, arylene-alkylene, heteroarylene, heteroarylene-alkylene, carbocyclylene, carbocyclylene-alkylene, heterocyclylene, and heterocyclylene-alkylene, where R is selected from H, unsubstituted or substituted alkyl, and unsubstituted or substituted aryl. In some embodiments, linking moiety B comprises an alkylene, oxyalkylene, or polyoxyalkylene group, and / or A and C are each a maleimide group.

[0027] In some embodiments, the closing portion has a length of about 1 Å to about 100 Å. In some embodiments, the closing portion has a length of about 5 Å to about 50 Å.

[0028] In some embodiments, the polynucleotide handling protein is a helicase. In some embodiments, the polynucleotide handling protein is a DNA-dependent ATPase (Dda) helicase.

[0029] In some embodiments, the leader comprises a polymer. In some embodiments, the leader is attached to the polypeptide by a linker. In some embodiments, the linker comprises a polynucleotide.

[0030] In some embodiments, a leader comprises one or more of a nucleotide lacking both a nucleobase and a sugar moiety (spacer moiety), a deoxyribonucleotide (DNA), a ribonucleotide (RNA), a peptide nucleotide (PNA), a glycerol nucleotide (GNA), a threose nucleotide (TNA), a locked nucleotide (LNA), a bridged nucleotide (BNA), an abasic nucleotide, or a nucleotide with a modified phosphate linkage.

[0031] In some embodiments, the leader is configured to facilitate unbinding of the polynucleotide binding site of the polynucleotide handling protein from the conjugate.

[0032] In some embodiments, the blocking moiety restricts movement of the conjugate through the polynucleotide binding site of the polynucleotide handling protein, thereby restricting movement of the conjugate in a second direction relative to the detector.

[0033] In some embodiments, the conjugate comprises multiple polypeptide sections and / or multiple polynucleotide sections.

[0034] In some embodiments, the conjugate comprises m -{PN} n or L-{PN} n -P m and wherein the structure comprises one or more structures of the form - L is a leader, L is optionally an N moiety, and the conjugate is LN m -{PN} n and when m is 1, L optionally comprises one or more different types of nucleotides relative to adjacent N moieties; -P is a polypeptide, -N comprises a polynucleotide, - m is 0 or 1, -n is a positive integer, The detector comprises a transmembrane nanopore spanning a membrane having a cis side and a trans side, the method comprising passing a leader (L) through the nanopore, thereby contacting a polypeptide (P) with the nanopore; i) the polynucleotide handling protein is located on the cis side of the nanopore, and the method comprises enabling the polynucleotide handling protein to control the movement of a polynucleotide (N) in a direction from the trans side of the nanopore to the cis side of the nanopore, thereby controlling the movement of a polypeptide (P) through the nanopore; or i) a polynucleotide handling protein is located on the trans side of the nanopore, the method comprising enabling the polynucleotide handling protein to control translocation of a polynucleotide (N) in a direction from the cis side of the nanopore to the trans side of the nanopore, thereby controlling translocation of a polypeptide (P) through the nanopore.

[0035] In some embodiments, the or each polypeptide, independently, has a length of from 2 to about 50 peptide units.

[0036] In some embodiments, the or each polynucleotide, independently, has a length of from about 10 to about 1000 nucleotides.

[0037] In some embodiments, one or more adaptors and / or one or more tethers and / or one or more anchors are attached to the polynucleotide in the conjugate.

[0038] In some embodiments, methods provided include applying a force to a detector, where the polynucleotide handling protein controls movement of the target polynucleotide relative to the detector in a direction opposite to the applied force, hi some embodiments, the force comprises an electric potential applied to the detector.

[0039] In some embodiments, the detector comprises a nanopore, and the nanopore is modified to increase the distance between the polynucleotide handling protein and the constriction region of the nanopore. In some embodiments, the polynucleotide handling protein is modified to increase the distance between the active site of the polynucleotide handling protein and the detector. In some embodiments, the displacer unit separates the polynucleotide handling protein from the detector, thereby increasing the distance between the polynucleotide handling protein and the detector. In some embodiments, the displacer unit comprises one or more proteins.

[0040] In some embodiments, the one or more measurements are characteristic of one or more features of the polypeptide selected from (i) the length of the polypeptide, (ii) the identity of the polypeptide, (iii) the sequence of the polypeptide, (iv) the secondary structure of the polypeptide, and (v) whether the polypeptide is modified.

[0041] Also provided herein is a conjugate comprising a polypeptide conjugated to a polynucleotide, i) a first end of the polypeptide is conjugated to a first end of the polynucleotide; ii) the second end of the polypeptide is attached to a polymer leader, optionally via a linker, which optionally comprises a polynucleotide; iii) if the linker comprises a polynucleotide, the polynucleotide or linker is bound to a polynucleotide binding site of the polynucleotide handling protein, and the polynucleotide binding site is modified with a closing moiety, thereby topologically closing the polynucleotide binding site of the polynucleotide handling protein around the polynucleotide; iv) The polynucleotide comprises a blocking moiety to prevent disassociation of the polynucleotide handling protein from the conjugate.

[0042] In some embodiments, the polypeptides, polynucleotides, leaders, linkers, polynucleotide handling proteins, and blocking moieties are as described in more detail herein.

[0043] Also provided herein is a system comprising a nanopore and a conjugate as defined herein.

[0044] Also provided is a kit comprising: a polynucleotide, wherein a first end of the polynucleotide comprises a reactive functional group for conjugating to a first end of a target polypeptide and a second end of the polynucleotide comprises a blocking moiety; a polynucleotide handling protein, wherein the polynucleotide binding site of the polynucleotide handling protein is modified with a closing moiety for topologically closing the polynucleotide binding site of the polynucleotide handling protein around the polynucleotide; and an adaptor comprising a polymeric leader and a reactive functional group for attachment to a second end of the target polypeptide; A kit comprising:

[0045] In some embodiments, the polynucleotide, polynucleotide handling protein, reader, and nanopore are as described herein. [Brief description of the drawings]

[0046] [Figure 1] Schematic diagram showing a non-limiting example of an embodiment of the disclosed method in which a polynucleotide handling protein on the cis side of a nanopore controls the movement of a conjugate comprising a polynucleotide (DNA2) conjugated to a polypeptide from the trans side of the nanopore to the cis side of the nanopore, thus allowing the polypeptide to be characterized as it moves relative to the nanopore. As shown, an optional leader (DNA1) is attached to the conjugate to facilitate passage of the polypeptide through the nanopore. RED (discussed herein) is shown as the conceptual distance between the constriction in the nanopore and the active site of the polynucleotide handling protein. (A) A substrate can be captured in the nanopore, for example, from the cis side of the membrane, by applying, for example, a positive voltage to the trans side of the membrane. The polynucleotide handling protein moves along the polynucleotide section in the direction indicated by the dotted arrow to pull the conjugate out of the pore and proceed to state (B). As the polynucleotide handling protein moves along the polynucleotide (e.g., in one nucleotide-fueled step), the conjugate is pulled out of the nanopore and the peptide section passes through the nanopore. Debinding of the conjugate from the polynucleotide handling protein allows for a return to state (A), and the process can then be repeated, thereby "flossing" the conjugate against the nanopore. [Diagram 2]Schematic diagram showing a non-limiting example of an embodiment of the disclosed method in which a polynucleotide handling protein on the cis side of a detector, such as a nanopore, controls the movement of a conjugate comprising a polynucleotide (DNA2) conjugated to a polypeptide from the trans side of the nanopore to the cis side of the nanopore, thus allowing the polypeptide to be characterized as it moves relative to the nanopore. As shown, an optional leader is attached to the conjugate to facilitate, for example, the initial passage of the polypeptide through the nanopore. (A) The substrate can be captured in the nanopore, for example, from the cis side of the membrane, by applying, for example, a positive voltage to the trans side of the membrane. The polynucleotide handling protein moves along the polynucleotide section in the direction indicated by the dotted arrow to move the substrate in a first direction relative to the pore, for example, out of the pore, proceeding to state (B). As the polynucleotide handling protein moves along the polynucleotide (e.g., in a one nucleotide fueled step), the protein expels the conjugate from the nanopore. Thus, the polypeptide section of the conjugate passes through the nanopore (state C) and is thus characterized. The conjugate then unbinds from the polynucleotide binding site of the polynucleotide handling protein, and the conjugate moves in a second direction relative to the pore, e.g., into the pore. As depicted, movement of the conjugate results in a return to state (A), but movement of the conjugate relative to the pore may result in a different portion of the conjugate coming into contact with the pore, as discussed herein. The conjugate then rebinds to the same polynucleotide handling protein moiety, which again controls movement of the conjugate in a first direction (e.g., out of the pore), thereby proceeding to states (B) and (C). This cycle can be optionally repeated multiple times, thereby flossing the conjugate against the extractor. [Diagram 3]Schematic diagram showing a non-limiting example of a conjugate for use in the disclosed method. The polynucleotide portion of the conjugate (DNA2) comprises a first strand (DNA2 top) conjugated to a polypeptide at a first end and a second strand (DNA2 bottom) hybridized thereto. An optional double-stranded polynucleotide linker is used to attach an optional leader to the polynucleotide, for example, to assist in threading the conjugate to a detector such as a nanopore. The linker / leader portion may be provided, for example, in the form of a sequencing Y-adapter that includes a leader. An optional blocking moiety (blocker) is included at the second end of the polynucleotide portion of the conjugate. As shown, a polynucleotide handling protein is attached to the conjugate at the portion of the conjugate between the polypeptide and the blocking moiety. [Figure 4] Schematic diagram showing a non-limiting example of a conjugate for use in the disclosed method. The polynucleotide portion of the conjugate (DNA2) comprises a first strand (DNA2 top) conjugated to a polypeptide at a first end and a second strand (DNA2 bottom) hybridized thereto. An optional linker is used to attach an optional leader to the polynucleotide to aid in threading the conjugate to a detector, such as a nanopore. The linker / leader portion may be provided, for example, in the form of a sequencing adaptor that includes a leader. An optional blocking moiety (blocker) is included at the second end of the polynucleotide portion of the conjugate. As shown, a polynucleotide handling protein is attached to the conjugate in the portion of the conjugate between the polypeptide and the free end of the leader. As shown, the polynucleotide handling protein is attached to the linker. [Diagram 5]Schematics showing non-limiting examples of strategies to increase the distance between a nanopore (e.g., a constriction within the nanopore) and the active site of a polynucleotide handling protein used to control the movement of a conjugate relative to the nanopore. FIG. 5A shows: A: Schematic of an unmodified pore showing an unmodified RED. B: The nanopore may be modified to expand the RED. C: A displacer unit can be used to displace the polynucleotide handling protein from the nanopore, thus expanding the RED. FIG. 5B shows how multiple polynucleotide handling proteins can be used to replace the active polynucleotide handling protein that controls the movement of a conjugate from the nanopore relative to the nanopore. These embodiments are described in more detail herein. [Figure 6] Representative current versus time traces for Example 1 with schematics of the corresponding constructs for clarity. States A-D correspond to the following: A - capture of the leader strand by the nanopore, B - translocation of the Y adaptor across the nanopore leader head (RED), C - translocation of the polypeptide across the RED, D - translocation of the polynucleotide tail (DNA2). The first trace represents a partial translocation event of only the Y adaptor (states A and B only), while the second trace shows the translocation of the entire conjugated polynucleotide-polypeptide through the nanopore. Data obtained as described in Example 1 (this data for a polynucleotide-peptide conjugate comprising a peptide of sequence SEQ ID NO: 15). [Figure 7] Current vs. time trace showing high throughput of data collection. In a 3 second period, there are 5 capture events, 4 corresponding to the complete polynucleotide-polypeptide conjugate (event 3 is a partial translocation of only the Y adaptor). Data as described in Example 1 (this data for a polynucleotide-peptide conjugate containing a peptide of sequence SEQ ID NO: 15). [Figure 8]Current traces of translocation of the polynucleotide-peptide conjugate described in Figure 4B and Example 1, corresponding to the peptide sequence GGSGRRSGSG (SEQ ID NO: 16). A: Eleven example traces aligned with respect to the conditions described in Figure 6. B: Overlay of the same eleven traces. C: Stacked plot of the eleven example traces showing normalization of the time axis using a dynamic time warping algorithm that promotes optimal alignment of key trace features. [Figure 9] Current traces of translocation of the polynucleotide-peptide conjugate described in Figure 4B and Example 1 corresponding to the peptide sequence GGSGYYSGSG (SEQ ID NO: 17). A: 12 example traces aligned with respect to the conditions described in Figure 6. B: Overlay of the same 12 traces. C: Stacked plot of the 12 example traces. [Figure 10] Current traces of translocation of the polynucleotide-peptide conjugate described in Figure 4B and in Example 1, corresponding to the peptide sequence GGSGDDSGSG (SEQ ID NO: 15). A: Eleven examples of traces aligned with respect to the conditions described in Figure 6. B: Overlay of the same eleven traces. C: Stacked plot of the eleven example traces. [Figure 11] Schematic structure of the construct obtained using the peptide of SEQ ID NO: 17. A Y adaptor comprising the polynucleotide strands of SEQ ID NOs: 9, 10 and 11; and a polynucleotide tail comprising the polynucleotide strands of SEQ ID NOs: 12 and 14 (described in Example 1). [Figure 12] Representative current versus time traces for Example 2 compared to Example 1. States A-D correspond to: A-capture of the leader strand by the nanopore, B-translocation of the Y-adapter across the nanopore leader head (RED), C-translocation of the polypeptide across the RED, D-translocation of the polynucleotide tail (DNA2). The traces in the top panel were collected according to the protocol of Example 1 (using peptides pre-modified during synthesis). The traces in the bottom panel show translocation of polynucleotide-polypeptide conjugated according to the protocol of Example 2 (using unmodified peptide of SEQ ID NO: 18; i.e., the same sequence as the corresponding traces in Example 1). [Figure 13] Representative current versus time traces for the translocation of polynucleotide-peptide conjugates of a 10 amino acid peptide (top panel, SEQ ID NO: 15) compared to a 21 amino acid peptide (bottom panel, SEQ ID NO: 19). States A-D correspond to the following: A-capture of the leader strand by the nanopore, B-translocation of the Y adaptor across the nanopore leader head (RED), C-translocation of the polypeptide across the RED, D-translocation of the polynucleotide tail. Results are shown in Example 3. [Figure 14-1]Schematic of the design and assembly of the construct followed by a schematic of translocation across the nanopore. (A) A DNA adapter consisting of a polynucleotide handling enzyme loaded and blocked onto the DNA strand between a 5′ blocker (traptavidin bound to biotin at the 5′ of the DNA) and a stall chemistry at the 3′ end. (B) A fully assembled construct including a DNA adapter ligated to a piece of dsDNA, which is then ligated to a linker DNA with a chemistry that allows conjugation to a peptide, followed by a peptide and dsDNA tail with a leader that allows capture by the nanopore and a tether site for localization to the membrane. (C) Schematic of a polynucleotide handling protein on the cis side of the nanopore that controls the translocation of the conjugate (FIG. 14B) in the direction from the trans side of the nanopore to the cis side of the nanopore, thus allowing the polypeptide to be characterized as it translocates relative to the nanopore. As shown, the leader on the dsDNA tail facilitates the initial passage of the polypeptide through the nanopore (i). The substrate is captured into the nanopore from the cis side of the membrane by applying a positive voltage to the trans side of the membrane (i). The construct is pulled into the nanopore stripping dsDNA until the enzyme reaches the pore, where the blocking chemistry behind the enzyme prevents further translocation (ii). When the voltage is stopped, the construct is allowed to flow out of the pore, and over time, the polynucleotide handling enzyme is allowed to diffuse across the stall chemistry of the top strand of the DNA adaptor (iii). As the polynucleotide handling protein moves along the polynucleotide section in the direction indicated by the dotted arrow, the conjugate is displaced out of the pore (iv) until the enzyme stalls at the junction between the DNA and the peptide (v). As the polynucleotide handling protein moves along the polynucleotide (e.g., in a one nucleotide fueled step), the conjugate is displaced from the nanopore, and then voltage is used to pull the conjugate back into the nanopore, resulting in a loop in which the conjugate threads up and down the pore. The polypeptide section of the conjugate can pass through the nanopore multiple times and therefore be reread multiple times. [Figure 14-2] (Continuation of 14-1) [Figure 15] Representative current versus time traces from Example 5 using a conjugate consisting of a DNA adaptor and 800 bp of dsDNA ligated to a GGSGYYSGSG peptide followed by a short dsDNA linker clicked onto the dsDNA tail. (A) Current versus time trace highlighting the section of the trace corresponding to the 800 bp of DNA passing through the nanopore, followed by a section of the polypeptide entering and exiting the nanopore. (B) Zoomed-in views showing different levels of the DNA and peptide current traces, followed by zoomed-out panels focusing on the peptide section moving up and down the pore (read 1 through read 6). [Figure 16] Representative current versus time traces from Example 5 using a conjugate consisting of a DNA adaptor and 3600 bp of dsDNA ligated to a GGSGDDSGSG peptide followed by a short dsDNA linker clicked onto the dsDNA tail. (A) Current versus time trace highlighting the section of the trace corresponding to the DNA passing through the nanopore, followed by a section of the polypeptide entering and exiting the nanopore. (B) Zoomed-in views showing different levels of the DNA and peptide current traces, followed by zoomed-out panels focusing on the peptide section moving up and down the pore (read 1 to read 6). [Figure 17] Representative current versus time traces from Example 5 showing a comparison between three different peptides. Panel (A) shows data for peptide GGSGRRSGSG, panel (B) shows data for peptide GGSGDDSGSG, and (C) corresponds to data for peptide YDDDDY. The top row of traces represents 6 seconds of data, the second row zooms in to 2 seconds, the third row shows 1 second of data, and the bottom row corresponds to 0.5 seconds, with individual peptide reads highlighted with arrows. As the peptides become more negative, the reads show improved reproducibility and better defined characteristics of the peptides. [Figure 18]Schematic structure of the construct obtained using the exemplary peptides of SEQ ID NOs: 39-41, the Y adaptor comprising the polynucleotide strands of SEQ ID NOs: 33, 34 and 35, the clinker comprising the polynucleotide strands of SEQ ID NOs: 36 and 37, and the hairpin back blocker comprising the polynucleotide strand of SEQ ID NO: 38, and the drop back of the bound motor protein after capture of the construct by the nanopore (described in Example 6). [Figure 19] Representative polynucleotide sequences used to generate the constructs shown generally in FIG. [Figure 20] Representative current versus time traces from Example 6 showing a comparison between three different peptides. The top panel shows data for polyD (SEQ ID NO: 39), the middle panel shows data for peptide SG-RR (SEQ ID NO: 40), and the bottom panel corresponds to data for peptide SG-YY (SEQ ID NO: 41). The data demonstrate successful capture and re-reading of peptide chains according to the methods disclosed herein. [Figure 21] Schematic structure of the construct obtained using the exemplary peptides of SEQ ID NOs: 42 to 50. A Y adaptor comprising polynucleotide strands of SEQ ID NOs: 42, 43 and 44, a clinker comprising polynucleotide strands of SEQ ID NOs: 45 and 46, and a DNA tail comprising a polynucleotide strand of SEQ ID NO: 47 (described in Example 7). [Figure 22] Representative polynucleotide sequences used to generate the constructs shown generally in FIG. [Figure 23] Representative current versus time traces from Example 7 showing a comparison between three different peptides. The top panel shows data for peptide PolyD (SEQ ID NO: 48), the middle panel shows data for peptide SG-DD (SEQ ID NO: 49), and the bottom panel corresponds to data for peptide SG-RR (SEQ ID NO: 50). The data demonstrate successful capture and re-reading of peptide chains according to the methods disclosed herein. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0047] The present invention will be described with respect to certain embodiments and with reference to certain drawings, but the present invention is not limited thereto, but only by the claims. Any reference signs in the claims should not be construed as limiting the scope thereof. Of course, it should be understood that not necessarily all aspects or advantages can be achieved in accordance with any particular embodiment of the present invention. Thus, for example, a person skilled in the art will recognize that the present invention can be embodied or performed in a manner that achieves or optimizes one advantage or group of advantages taught herein, without necessarily achieving other aspects or advantages that may be taught or suggested herein.

[0048] The present invention, both as to its organization and method of operation, together with its features and advantages, may best be understood by reference to the following detailed description when read in conjunction with the accompanying drawings. Aspects and advantages of the present invention will be apparent and elucidated with reference to the embodiment(s) described below. Reference throughout this specification to "one embodiment" or "an embodiment" means that a particular feature, structure, or characteristic described in connection with that embodiment is included in at least one embodiment of the present invention. Thus, the appearances of the phrases "in one embodiment" or "in an embodiment" in various places throughout this specification do not necessarily all refer to the same embodiment, although they may. Similarly, in describing exemplary embodiments of the present invention, it should be understood that various features of the invention may be grouped together in a single embodiment, figure, or description thereof in order to simplify the disclosure and aid in understanding one or more of the various inventive aspects. However, this method of disclosure should not be interpreted as reflecting an intention that the claimed invention requires more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive aspects amount to less than all features of a single foregoing disclosed embodiment.

[0049] It is to be understood that "embodiments" of the present disclosure may be specifically combined together, unless the context dictates otherwise. Any specific combination of the disclosed embodiments is a further disclosed embodiment of the claimed invention (unless the context implies otherwise).

[0050] Additionally, as used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural references unless the context clearly indicates otherwise. Thus, for example, reference to a "polynucleotide" includes two or more polynucleotides, reference to a "polynucleotide handling protein" includes two or more such proteins, reference to a "helicase" includes two or more helicases, reference to a "monomer" refers to two or more monomers, and reference to a "pore" includes two or more pores.

[0051] All publications, patents, and patent applications cited herein, whether supra or infra, are hereby incorporated by reference in their entirety.

[0052] definition When an indefinite or definite article is used when referring to a singular noun, such as "a" or "an" or "the," this includes the plural of that noun, unless otherwise specified. When the term "comprises" is used in the present description and claims, it does not exclude other elements or steps. Furthermore, terms such as first, second, third, etc. in the present description and claims are used to distinguish between similar elements and are not necessarily used to describe an order or chronological order that occurred. It is to be understood that the terms used in this manner are interchangeable under appropriate circumstances and that the embodiments of the invention described herein can operate in other sequences than those described or illustrated herein. The following terms or definitions are provided solely to aid in the understanding of the present invention. Unless specifically defined herein, all terms used herein have the same meaning as would be understood by one of ordinary skill in the art of the invention. Skilled artisans should refer to the following publications for definitions and technical terms, particularly those referenced in Sambrook et al., Molecular Cloning: A Laboratory Manual, 4 th ed., Cold Spring Harbor Press, Plainsview, New York (2012), and Ausubel et al., Current Protocols in Molecular Biology (Supplement 114), John Wiley & Sons, New York (2016). The definitions provided herein should not be construed to be narrower than understood by one of ordinary skill in the art.

[0053] As used herein, "about" when referring to a measurable value, such as an amount, temporal duration, and the like, is meant to encompass variations of ±20% or ±10%, more preferably ±5%, even more preferably ±1%, and still more preferably ±0.1% from the specified value, where such variations are appropriate for performing the disclosed methods.

[0054] As used herein, a "nucleotide sequence," "DNA sequence," or "nucleic acid molecule(s)" refers to a polymeric form of nucleotides of any length, either ribonucleotides or deoxyribonucleotides. The term refers only to the primary structure of the molecule. Thus, the term includes double- and single-stranded DNA and RNA. The term "nucleic acid" as used herein is a single- or double-stranded covalently linked nucleotide sequence in which the 3' and 5' ends of each nucleotide are linked by a phosphodiester bond. A polynucleotide may be composed of deoxyribonucleotide or ribonucleotide bases. Nucleic acids may be synthetically produced in vitro or isolated from natural sources. Nucleic acids may further include modified DNA or RNA, e.g., DNA or RNA that is methylated, or RNA that has been subjected to post-translational modifications, e.g., 5'-capping with 7-methylguanosine, 3'-processing such as cleavage and polyadenylation, and splicing. Nucleic acids may also include synthetic nucleic acids (XNAs), such as hexitol nucleic acids (HNAs), cyclohexene nucleic acids (CeNAs), threose nucleic acids (TNAs), glycerol nucleic acids (GNAs), locked nucleic acids (LNAs), and peptide nucleic acids (PNAs). The size of a nucleic acid, also referred to herein as a "polynucleotide", is typically expressed in terms of the number of base pairs (bp) for a double-stranded polynucleotide, or the number of nucleotides (nt) for a single-stranded polynucleotide. 1000 bp or nt is equivalent to a kilobase (kb). Polynucleotides less than about 40 nucleotides in length are typically referred to as "oligonucleotides" and may include primers for use in manipulating DNA, such as via polymerase chain reaction (PCR).

[0055] The term "amino acid" in the context of this disclosure is used in its broadest sense and is intended to include organic compounds containing amine (NH2) and carboxyl (COOH) functional groups, along with a side chain (e.g., R group) specific to each amino acid. In some embodiments, amino acid refers to naturally occurring L α-amino acids or residues. The commonly used one-letter and three-letter abbreviations for naturally occurring amino acids are used herein: A=Ala, C=Cys, D=Asp, E=Glu, F=Phe, G=Gly, H=His, I=Ile, K=Lys, L=Leu, M=Met, N=Asn, P=Pro, Q=Gln, R=Arg, S=Ser, T=Thr, V=Val, W=Trp, and Y=Tyr (Lehninger, AL, (1975) Biochemistry, 2d ed., pp. 71-92, Worth Publishers, New York). The general term "amino acid" further includes D-amino acids, retro-inverso amino acids, and chemically modified amino acids such as amino acid analogs, naturally occurring amino acids that are not normally incorporated into proteins such as norleucine, and chemically synthesized compounds that have properties known in the art that characterize amino acids such as β-amino acids. For example, analogs or mimetics of phenylalanine or proline that allow the same conformational constraints on peptide compounds as natural Phe or Pro are included within the definition of amino acids. Such analogs and mimetics are referred to herein as "functional equivalents" of the respective amino acids. Other examples of amino acids are listed by Roberts and Vellaccio, The Peptides: Analysis, Synthesis, Biology, Gross and Meiehofer, eds., Vol. 5 p. 341, Academic Press, Inc., NY 1983, which is incorporated herein by reference.

[0056] The terms "polypeptide" and "peptide" are used interchangeably herein to refer to a polymer of amino acid residues, as well as variants and synthetic analogs thereof. Thus, these terms apply to amino acid polymers in which one or more amino acid residues are non-naturally occurring synthetic amino acids, such as chemical analogs of the corresponding naturally occurring amino acids, as well as to naturally occurring amino acid polymers. Polypeptides may also undergo maturation or post-translational modification processes, which may include, but are not limited to, glycosylation, proteolytic cleavage, lipidation, signal peptide cleavage, propeptide cleavage, phosphorylation, and the like. Peptides may be produced using recombinant techniques, for example, by expression of recombinant or synthetic polynucleotides. Recombinantly produced peptides are typically substantially free of culture medium, e.g., culture medium represents less than about 20% of the volume of the protein preparation, more preferably less than about 10%, and most preferably less than about 5%.

[0057] The term "protein" is used to describe a folded polypeptide having secondary or tertiary structure. A protein may be composed of a single polypeptide or may contain multiple polypeptides that assemble to form a multimer. A multimer may be a homo- or hetero-oligomer. A protein may be a naturally occurring or wild-type protein, or a modified or non-naturally occurring protein. A protein may differ from a wild-type protein, for example, by the addition, substitution, or deletion of one or more amino acids.

[0058] A "variant" of a protein includes peptides, oligopeptides, polypeptides, proteins, and enzymes that have amino acid substitutions, deletions, and / or insertions compared to the unmodified or wild-type protein in question and have biological and functional activities similar to those of the unmodified protein from which they are derived. As used herein, the term "amino acid identity" refers to the degree to which sequences are identical amino acid-by-amino acid over a comparison window. Thus, "percentage of sequence identity" is calculated by comparing two optimally aligned sequences over a comparison window, determining the number of positions at which identical amino acid residues (e.g., Ala, Pro, Ser, Thr, Gly, Val, Leu, Ile, Phe, Tyr, Trp, Lys, Arg, His, Asp, Glu, Asn, Gln, Cys, and Met) occur in both sequences to calculate the number of matched positions, dividing the number of matched positions by the total number of positions in the comparison window (i.e., the window size), and multiplying the result by 100 to calculate the percentage of sequence identity.

[0059] For all aspects and embodiments of the invention, a "variant" has at least 50%, 60%, 70%, 80%, 90%, 95%, or 99% complete sequence identity to the amino acid sequence of the corresponding wild-type protein. Sequence identity may also be to a fragment or portion of a full-length polynucleotide or polypeptide. Thus, while a sequence may have only 50% sequence identity overall with a full-length reference sequence, the sequence of a particular region, domain, or subunit may share 80%, 90%, or even 99% sequence identity with the reference sequence.

[0060] The term "wild type" refers to a gene or gene product isolated from a naturally occurring source. A wild type gene is that which is most frequently observed in a population and is therefore the arbitrarily designed "normal" or "wild type" form of that gene. In contrast, the terms "modified," "mutant," or "variant" refer to a gene or gene product that exhibits sequence modifications (e.g., substitutions, truncations, or insertions), post-translational modifications, and / or functional properties (e.g., altered characteristics) compared to the wild type gene or gene product. It should be noted that naturally occurring mutants can be isolated, which are identified by the fact that they have altered characteristics compared to the wild type gene or gene product. Methods for introducing or substituting naturally occurring amino acids are well known in the art. For example, methionine (M) can be substituted with arginine (R) by replacing the codon for methionine (ATG) with the codon for arginine (CGT) at the relevant position in the polynucleotide encoding the mutant monomer. Methods for introducing or substituting non-naturally occurring amino acids are also well known in the art. For example, non-naturally occurring amino acids can be introduced by including synthetic aminoacyl-tRNA in the IVTT system used to express the mutant monomers. Alternatively, they can be introduced by expressing mutant monomers in E. coli that are auxotrophic for specific amino acids in the presence of synthetic (i.e., non-naturally occurring) analogs of those specific amino acids. They can also be generated by naked ligation when the mutant monomers are produced using partial peptide synthesis. Conservative substitutions replace amino acids with other amino acids of similar chemical structure, similar chemical properties, or similar side chain volume. The introduced amino acids can have similar polarity, hydrophilicity, hydrophobicity, basic, acidic, neutral, or charged properties as the amino acids they replace. Alternatively, conservative substitutions can introduce another amino acid that is aromatic or aliphatic in place of an existing aromatic or aliphatic amino acid. Conservative amino acid changes are well known in the art and can be selected according to the properties of the 20 main amino acids defined in Table 1 below.If the amino acids have similar polarity, this can also be determined by reference to the hydrophobicity scale of the amino acid side chains in Table 2. [Table 1] [Table 2]

[0061] The mutant or modified protein, monomer, or peptide can also be chemically modified in any manner and at any site. The mutant or modified monomer or peptide is preferably chemically modified by attachment of a molecule to one or more cysteines (cysteine ​​bond), attachment of a molecule to one or more lysines, attachment of a molecule to one or more unnatural amino acids, enzymatic modification of an epitope, or modification of a terminal. Suitable methods for carrying out such modifications are well known in the art. The mutant of the modified protein, monomer, or peptide can be chemically modified by attachment of any molecule. For example, the mutant of the modified protein, monomer, or peptide can be chemically modified by attachment of a dye or fluorophore.

[0062] Disclosed Method The present disclosure relates to methods for characterizing a polypeptide by conjugating it to a polynucleotide and controlling the translocation of the conjugate relative to a nanopore using a polynucleotide handling protein.

[0063] In contrast to methods that attempt to control the translocation of a polypeptide relative to a nanopore using a polypeptide handling enzyme, the methods of the present disclosure enable the use of a polynucleotide handling enzyme to control the translocation of a polypeptide relative to a nanopore.

[0064] The movement of the polynucleotide is controlled by a polynucleotide handling protein, and since the polynucleotide is conjugated to the polypeptide in the conjugate, the movement of the polynucleotide drives the movement of the polypeptide.

[0065] The use of polynucleotide handling proteins to control the movement of polynucleotides, and thus the movement of polypeptides, may be associated with advantages compared to the methods for characterizing polypeptides known in the art.For example, polynucleotide handling proteins can process the handling of polynucleotides with a higher turnover rate compared to polypeptide handling enzymes.This means that characterization data can be obtained more quickly for the polypeptides characterized according to the disclosed methods compared to previously known methods.

[0066] The methods disclosed herein utilize the ability of polynucleotide handling proteins to control the movement of conjugates that do not only contain polynucleotides.In particular, conjugates that contain polypeptides can be moved in a controlled manner using polynucleotide handling proteins as described herein.Polynucleotide handling proteins suitable for use in the disclosed methods are described in more detail herein.

[0067] The methods disclosed herein include taking measurements characteristic of the polypeptide as the conjugate moves in a first direction relative to the detector. The conjugate moves in both a first direction and a second direction relative to the detector in the disclosed methods. Thus, the disclosed methods include "flossing" the conjugate relative to the detector. This is described in more detail herein.

[0068] Although the present disclosure provides a nanopore as an exemplary detector, the methods provided herein are suitable for detectors such as (i) zero mode waveguides, (ii) field effect transistors, optionally nowire field effect transistors, (iii) AFM tips, (iv) nanotubes, optionally carbon nanotubes, and (v) nanopores. The disclosed methods are particularly suitable for methods in which polynucleotides move through a detector or through a structure that contains a detector, such as a well in a detector chip.

[0069] The target polypeptide is included in a conjugate with a polynucleotide. The first end of the target polypeptide is conjugated to the first end of the polynucleotide, thereby forming a conjugate. As described in more detail herein, the polynucleotide handling protein binds to the polynucleotide portion of the conjugate during at least a portion of the method.

[0070] The polynucleotide handling protein is used to control the movement of the conjugate in a first direction relative to the detector. By moving the conjugate in the first direction relative to the detector, the polypeptide portion of the conjugate moves in the first direction relative to the detector. As the conjugate moves in the first direction relative to the detector, a measurement characteristic of the polypeptide is made.

[0071] The conjugate then disassociates from the polynucleotide binding site of the polynucleotide handling protein.The conjugate then moves in a second direction relative to the detector.By moving the conjugate in a second direction relative to the detector, the polypeptide portion of the conjugate moves in a second direction relative to the detector.

[0072] As discussed in more detail herein, the second direction is typically opposite to the first direction. For example, in embodiments in which the detector is a nanopore, the first direction may be "outside" the nanopore (from the "point of view" of the polynucleotide handling protein) and the second direction may be "inside" the nanopore (from the "point of view" of the polynucleotide handling protein). This is described in more detail herein.

[0073] The conjugate then rebinds to the polynucleotide handling protein. The conjugate rebinds to the same polynucleotide handling protein. In other words, the conjugate rebinds to the same molecule of polynucleotide handling protein, not just different molecules of the same type of polynucleotide handling protein. The polynucleotide handling protein is again used to control the movement of the conjugate in a first direction relative to the detector. By moving the conjugate in a first direction relative to the detector, the polypeptide portion of the conjugate moves in a first direction relative to the detector. As the conjugate moves in a first direction relative to the detector, a further measurement characteristic of the polypeptide is made.

[0074] Measurements made as the conjugate migrates in a first direction can, in some embodiments, be combined or compared to improve characterization of the polypeptide.

[0075] In some embodiments, the second end of the polypeptide may be attached to a leader. The second end of the polypeptide may be attached to the leader by a linker, such as a polynucleotide linker. The leader may be a different structure than the linker. This is described in more detail herein.

[0076] Accordingly, provided herein is a method for characterizing a target polypeptide having a first end and a second end, comprising: the target polypeptide is included in a conjugate, wherein a first end of the target polypeptide is conjugated to a first end of a polynucleotide; The method is: (i) contacting a detector having a first opening and a second opening, or contained in a structure having a first opening and a second opening, with a conjugate having a polynucleotide handling protein bound thereto, wherein the conjugate is bound to the polynucleotide handling protein at a polynucleotide binding site of the polynucleotide handling protein; (ii) taking one or more measurements characteristic of the target polypeptide when the polynucleotide handling protein controls movement of the conjugate in a first direction relative to a detector; (iii) debinding the conjugate from the polynucleotide binding site of the polynucleotide handling protein such that the conjugate moves in a second direction relative to the detector; (iv) rebinding the conjugate to the polynucleotide binding site of the polynucleotide handling protein and performing one or more measurements characteristic of the target polypeptide as the polynucleotide handling protein controls movement of the conjugate in a first direction relative to a detector; and characterizing the target polypeptide, The second end of the target polypeptide is attached to a leader.

[0077] The second end of the polynucleotide may include a blocking moiety to prevent the polynucleotide handling protein from disassociating from the conjugate. The polynucleotide handling protein may bind to the portion of the conjugate between the polypeptide and the blocking moiety. For example, the polynucleotide handling protein may bind to the portion of the conjugate between the polypeptide and the blocking moiety prior to step (i) of the disclosed method. This is discussed in more detail below.

[0078] Accordingly, provided herein is a method for characterizing a target polypeptide having a first end and a second end, comprising: the target polypeptide is included in a conjugate in which a first end of the target polypeptide is conjugated to a first end of a polynucleotide; The method is: (i) contacting a detector having a first opening and a second opening, or contained in a structure having a first opening and a second opening, with a conjugate having a polynucleotide handling protein bound thereto, wherein the conjugate is bound to the polynucleotide handling protein at a polynucleotide binding site of the polynucleotide handling protein; (ii) taking one or more measurements characteristic of the target polypeptide when the polynucleotide handling protein controls movement of the conjugate in a first direction relative to a detector; (iii) debinding the conjugate from the polynucleotide binding site of the polynucleotide handling protein such that the conjugate moves in a second direction relative to the detector; (iv) rebinding the conjugate to the polynucleotide binding site of the polynucleotide handling protein and performing one or more measurements characteristic of the target polypeptide as the polynucleotide handling protein controls movement of the conjugate in a first direction relative to a detector; and characterizing the target polypeptide, The second end of the polynucleotide includes a blocking moiety to prevent disassociation of the polynucleotide handling protein from the conjugate.

[0079] In some embodiments, steps (iii) and (iv) of the disclosed method are repeated multiple times by successively binding and rebinding the polynucleotide handling protein to the conjugate. In that way, the conjugate can be oscillated relative to the detector (i.e., "flossed" relative to the first and second openings of the detector). This "flossing" allows repeated characterization of the polypeptide portion of the conjugate. In some embodiments, this can increase the accuracy of the characterization information.

[0080] Any suitable polypeptide can be characterized using the methods disclosed herein. In some embodiments, the target polypeptide is a protein or a naturally occurring polypeptide. In some embodiments, the polypeptide is a synthetic polypeptide. Polypeptides that can be characterized according to the disclosed methods are described in more detail herein.

[0081] Any suitable polynucleotide can be used in forming the conjugate for use in the methods disclosed herein. In some embodiments, the polynucleotide has at least the same length as the portion of the target polypeptide to be characterized. In some embodiments, the polynucleotide has a length longer than the portion of the target polypeptide to be characterized. This ensures that the length of the polypeptide portion that can be characterized is not limited by the amount of polynucleotide that the polynucleotide handling protein uses to control its movement. Suitable polynucleotides for use in the methods disclosed herein are disclosed in more detail herein.

[0082] In the disclosed methods, the target polypeptide can be conjugated to the polynucleotide using any suitable means, several exemplary means are described in more detail herein.

[0083] As discussed herein, the disclosed methods include using a polynucleotide handling protein to control the movement of the conjugate. The polynucleotide handling protein can control the movement of the polynucleotide relative to a detector, such as a nanopore. Exemplary polynucleotide handling proteins are described in more detail herein.

[0084] The polynucleotide handling protein controls the movement of the polynucleotide relative to a detector, such as a nanopore. Thus, the polynucleotide handling protein controls the movement of the conjugate relative to the detector. Any suitable detector can be used in the disclosed methods. Suitable detectors are described in more detail herein. In some embodiments, the detector for use in the disclosed methods is or includes a nanopore. Suitable nanopores for use in the disclosed methods are described in more detail herein.

[0085] In developing the disclosed method, the inventors found that the length of the polypeptide that can be characterized is typically improved when the detector is or includes a nanopore with a longer barrel or channel compared to a nanopore with a shorter barrel or channel. Without being bound by theory, the inventors believe that this may be because a pore with a longer barrel or channel, when used in combination with a polynucleotide handling protein as in the disclosed method, results in a longer distance between the active site of the polynucleotide handling protein and the constriction in the nanopore than a pore with a shorter barrel or channel. This distance may be referred to as the RED (Reader-Enzyme Distance). Those skilled in the art will understand that in such embodiments (as discussed below) the form of the nanopore is not limiting. The nanopore may be a protein nanopore or a solid-state nanopore. If the nanopore does not have a constriction in the channel of the nanopore, the constriction as used herein may be identified, for example, at the opening of the nanopore in one embodiment. Without being bound by theory, it is speculated that the length of the portion of the polypeptide in the conjugate that can be characterized by the nanopore can correspond to or be determined by the RED. In other words, the nanopore can include a read head, one or more measurements are characteristic of the "read portion" of the polypeptide, and the length of the read portion corresponds to or is determined by the distance between the read head and the active site of the polynucleotide handling protein. In some embodiments, the RED can be increased as described herein. Without being bound by theory, this can increase the length of the polypeptide portion of the conjugate that can be characterized according to the disclosed methods. In some embodiments, the RED is increased by modifying the detector used as described herein. In some embodiments, the RED is increased by modifying the polynucleotide handling protein used as described herein. In some embodiments, the RED is increased by using one or more displacer units as described herein.In some embodiments, the RED is increased by modifying the detector, and / or modifying the polynucleotide handling protein, and / or using one or more displacer units described herein.

[0086] The disclosed method includes performing one or more measurements characteristic of the polypeptide as the conjugate translocates relative to the nanopore. The one or more measurements can be any suitable measurements. Typically, the one or more measurements are electrical measurements, such as amperometric measurements, and / or one or more optical measurements. Apparatus for recording suitable measurements, and information such measurements can provide, are described in more detail herein.

[0087] Characterization of target polypeptide This method can be understood by reference to Figures 1 and 2, which show non-limiting examples of the disclosed method. The conjugate can include a polynucleotide and a polypeptide and is bound to a polynucleotide handling protein at its polynucleotide binding site. The conjugate is contacted with a detector, e.g., a nanopore, having a first opening and a second opening. The conjugate is bound to a polynucleotide handling protein described herein at the polynucleotide binding site of the polynucleotide handling protein. The polynucleotide handling protein typically binds to the conjugate at the polynucleotide portion of the conjugate. The polynucleotide handling protein typically binds to a polynucleotide adaptor attached to the polynucleotide portion of the conjugate.

[0088] The conjugate interacts with the detector. For example, the conjugate can pass through the first and second openings of the detector. In the illustrated embodiment, an additional polynucleotide is used to facilitate the passage of the conjugate through the first and second openings. Although such use is within the scope of the disclosed method, this is not required.

[0089] The polynucleotide handling protein processes the polynucleotide conjugated to the polypeptide. When the polynucleotide handling protein processes the polynucleotide, the conjugate moves relative to the detector, and thus the polypeptide moves relative to the detector. As the polypeptide moves relative to the detector (e.g., by passing through the first and second openings of the detector), the polypeptide is characterized.

[0090] In the example shown in Figure 1, the polynucleotide handling protein moves the conjugate in a first direction from the "viewpoint" of the polynucleotide handling protein to "out" of the first opening of the detector. Upon debinding from the polynucleotide handling protein, the conjugate may move in a second direction from the "viewpoint" of the polynucleotide handling protein to "in" of the first opening of the detector.

[0091] For example, as shown, the detector may include a transmembrane nanopore spanning a membrane having a cis side and a trans side. A first opening of the nanopore may be on the cis side of the membrane and a second opening of the nanopore may be on the trans side, such that the polynucleotide handling protein controls the movement of the conjugate through the nanopore in a direction "out" of the nanopore (from the "point of view" of the polynucleotide handling protein). Thus, the polynucleotide handling protein controls the movement of the conjugate through the nanopore in a direction from the trans side to the cis side of the membrane. Thus, in such an embodiment, the movement of the conjugate in a first direction is the movement of the conjugate from the trans side to the cis side of the membrane, i.e., from the trans side to the cis side of the nanopore. Upon debinding from the polynucleotide handling protein, the conjugate moves in a second direction relative to the detector. Thus, movement of the conjugate in the second direction is "into" the nanopore (from the "perspective" of the polynucleotide handling protein), i.e., movement of the conjugate through the nanopore from the cis to the trans side of the membrane, i.e., from the cis to the trans side of the nanopore. Movement of the conjugate controls movement of the polypeptide portion of the conjugate contained therein.

[0092] In some embodiments, the polypeptide portion of the conjugate (or a leader attached thereto, as described in more detail herein) contacts the nanopore in step (i) of the disclosed methods. In some such embodiments, the polypeptide portion or leader contacts the first opening (e.g., the cis opening) of the nanopore. Thus, in some embodiments of step (i) of the disclosed methods, the polypeptide portion of the conjugate is between the first opening (e.g., the cis opening) of the nanopore and the polynucleotide portion of the conjugate, and a polynucleotide handling protein binds at the polynucleotide portion of the conjugate on the cis side of the nanopore.

[0093] Of course, the opposite setup can also be used in the methods disclosed herein. For example, as shown, the detector can include a transmembrane nanopore spanning a membrane having a cis side and a trans side. A first opening of the nanopore can be on the trans side of the membrane and a second opening of the nanopore can be on the cis side, and thus the polynucleotide handling protein controls the movement of the conjugate through the nanopore in a direction "out" of the nanopore (from the "point of view" of the polynucleotide handling protein). Thus, the polynucleotide handling protein controls the movement of the conjugate through the nanopore in a direction from the cis side to the trans side of the membrane. Thus, in such an embodiment, the movement of the conjugate in a first direction is the movement of the conjugate from the cis side to the trans side of the membrane, i.e., from the cis side to the trans side of the nanopore. Upon debinding from the polynucleotide handling protein, the conjugate moves in a second direction relative to the detector. Thus, movement of the conjugate in the second direction is movement of the conjugate "into" the nanopore (from the "perspective" of the polynucleotide handling protein), i.e., from the trans to the cis side of the membrane, i.e., from the trans to the cis side of the nanopore. Movement of the conjugate controls movement of the polypeptide portion of the conjugate contained therein.

[0094] In some embodiments, the polypeptide portion of the conjugate (or a leader attached thereto, as described in more detail herein) contacts the nanopore in step (i) of the disclosed methods. In some such embodiments, the polypeptide portion or leader contacts the first opening of the nanopore (e.g., the trans opening). Thus, in some embodiments of step (i) of the disclosed methods, the polypeptide portion of the conjugate is between the first opening of the nanopore (e.g., the trans opening) and the polynucleotide portion of the conjugate, and a polynucleotide handling protein binds at the polynucleotide portion of the conjugate on the trans side of the nanopore.

[0095] This process can be repeated multiple times by successively binding and rebinding the polynucleotide handling protein to the conjugate. In this manner, the conjugate can be oscillated through the pore (i.e., "flossed" through the nanopore). This "flossing" allows the polypeptide portion of the conjugate to be repeatedly characterized by the nanopore. In some embodiments, this can increase the accuracy of the characterization information.

[0096] As described herein, the conjugate may include a leader. As described above, the leader may facilitate the passage of the conjugate through a detector, such as a nanopore. Any suitable leader may be used, as described in more detail herein. Optionally, the leader may be or may include a polynucleotide. The leader may be the same as or different from the polynucleotide in the conjugate. The leader may be attached to the polypeptide portion of the conjugate. The leader may be linked by a linker as described herein. Any suitable linker may be used. For example, the linker may include or consist of a polynucleotide. When a polynucleotide linker is used, the leader typically comprises or consists of a different type of polynucleotide than the linker. This is described in more detail herein.

[0097] In some embodiments, the conjugate comprises m -{PN} n or L-{PN} n -P m and wherein the structure comprises one or more structures of the form - L is a leader, L is optionally an N moiety, -P is a polypeptide, -N comprises a polynucleotide, - m is 0 or 1, -n is a positive integer.

[0098] In some embodiments, the conjugate comprises m -{PN} n The present invention includes one or more structures of the form:

[0099] The method may include contacting the reader (L) with the detector, for example with a first opening of the detector. The reader (L) may pass through the first opening of the detector. The reader (L) may pass through the first opening and the second opening of the detector. The detector may be a nanopore and the reader (L) may pass through the nanopore.

[0100] Contact between the leader (L) and the detector causes the polypeptide (P) to contact the nanopore. The polypeptide may contact a first opening of the detector. The polypeptide may contact a first opening and a second opening of the detector. In embodiments where the detector is a nanopore, the polypeptide may pass through the nanopore.

[0101] Typically, the conjugate is m -{PN} n where m is 1, L optionally includes one or more different types of nucleotides relative to the adjacent N moieties. This is described in more detail herein.

[0102] In some embodiments, the detector comprises a transmembrane nanopore spanning a membrane having a cis side and a trans side, and the method comprises passing a leader (L) through the nanopore, thereby bringing the polypeptide (P) into contact with the nanopore.

[0103] In some embodiments, the polynucleotide handling protein is located on the cis side of the nanopore, and the method includes enabling the polynucleotide handling protein to control movement of a polynucleotide (N) in a direction from the trans side of the nanopore to the cis side of the nanopore, thereby controlling movement of a polypeptide (P) through the nanopore.

[0104] In other embodiments, the polynucleotide handling protein is located on the trans side of the nanopore, and the method includes enabling the polynucleotide handling protein to control movement of a polynucleotide (N) in a direction from the cis side of the nanopore to the trans side of the nanopore, thereby controlling movement of a polypeptide (P) through the nanopore.

[0105] As described in more detail herein, a conjugate may include one or more adaptors and / or anchors.

[0106] As described in more detail herein, in some embodiments, the conjugate comprises a plurality of polynucleotides and polypeptides. In such embodiments, the polynucleotide handling protein sequentially controls the movement of the polynucleotide relative to the nanopore according to the disclosed methods, thus sequentially moving the polypeptide relative to the nanopore. In this manner, each polypeptide in the conjugate can be sequentially characterized in the disclosed methods.

[0107] For example, the conjugate may be LN m -P1-N-{PN} n or L-P1-N-{PN} n -P m and wherein the structure is one or more of the form: -n is a positive integer, - L is a leader, L is optionally an N moiety, - each P, which may be the same or different, is a polypeptide; - each N, which may be the same or different, comprises a polynucleotide; and -m is 0 or 1.

[0108] In some embodiments, the conjugate comprises m -P1-N-{PN} n The present invention includes one or more structures of the form:

[0109] Typically, in such embodiments, n is from 1 to about 1000, such as from 2 to about 100, for example from about 3 to about 10, such as 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10.

[0110] The method may include contacting the reader (L) with the detector, for example with a first opening of the detector. The leader (L) may pass through the first opening of the detector. The leader (L) may pass through the first and second openings of the detector. The detector may be a nanopore and the leader (L) may pass through the nanopore. Contact between the leader (L) and the detector causes the polypeptide section (P) to contact the nanopore. The polypeptide may contact the first opening of the detector. The polypeptide may contact the first and second openings of the detector. In embodiments where the detector is a nanopore, the polypeptide may pass through the nanopore.

[0111] In some embodiments, the detector comprises a transmembrane nanopore spanning a membrane having a cis side and a trans side, and the method comprises passing a leader (L) through the nanopore, thereby bringing the polypeptide section into contact with the nanopore.

[0112] In some such embodiments, the polynucleotide handling protein is located on the cis side of the nanopore and the method comprises enabling the polynucleotide handling protein to sequentially control the movement of each polynucleotide (N) from the trans side of the nanopore to the cis side of the nanopore, thereby controlling the movement of each polypeptide (P) sequentially through the nanopore. In other such embodiments, the polynucleotide handling protein is located on the trans side of the nanopore and the method comprises enabling the polynucleotide handling protein to sequentially control the movement of each polynucleotide (N) from the cis side of the nanopore to the trans side of the nanopore, thereby controlling the movement of each polypeptide (P) sequentially through the nanopore.

[0113] As described herein, the disclosed methods include "re-reading" the polypeptide portion of the conjugate by "flossing" the polypeptide section against a detector when it moves in a first direction under the control of a polynucleotide handling protein, and when it moves in a second direction when the conjugate decouples from the detector.

[0114] In some embodiments, each polypeptide section of the conjugate is moved sequentially in a first and second direction, i.e., flossed sequentially relative to the detector. For example, a first polypeptide portion of the conjugate may be moved and characterized according to the disclosed methods before subsequent movement and characterization of a second polypeptide portion of the conjugate according to the disclosed methods. The additional polypeptide portions can be further characterized accordingly.

[0115] In other embodiments, multiple polypeptide portions of a conjugate can be characterized simultaneously in the disclosed method. For example, the first and second polypeptide portions can be moved in a first and second direction relative to the detector so that the first and second polypeptide portions are characterized simultaneously. That is, the first and second polypeptide portions can be flossed together relative to the detector. The subsequent polypeptide sections can then be further characterized individually or collectively accordingly.

[0116] One of skill in the art will appreciate that where a conjugate comprises more than one polypeptide, it may be particularly advantageous for the polynucleotide handling protein (as described in more detail herein) to remain bound to the conjugate rather than dissociate when contacting the polypeptide, thereby allowing the polynucleotide handling protein to translocate over successive portions of the polynucleotide, passing over them as it contacts the polypeptide portions in the conjugate, in order to control movement of the conjugate relative to the nanopore.

[0117] In the methods disclosed herein, the polynucleotide handling protein controls the movement of the conjugate in a direction from the second opening of the detector to the first opening of the detector. In the depicted embodiment, as described herein, the polynucleotide handling protein controls the movement of the conjugate in a direction "out" of the pore (from the "viewpoint" of the polynucleotide handling protein). However, one of skill in the art will understand that the methods provided herein also encompass the opposite configuration, in which the polynucleotide handling protein controls the "movement" of the conjugate into the pore (from the "viewpoint" of the polynucleotide handling protein).

[0118] Thus, in some embodiments, the detector may include a transmembrane nanopore spanning a membrane having a cis side and a trans side. A first opening of the nanopore may be on the cis side of the membrane and a second opening of the nanopore may be on the trans side, such that the polynucleotide handling protein controls the movement of the conjugate through the nanopore in a direction "into" the nanopore (from the "point of view" of the polynucleotide handling protein). Thus, the polynucleotide handling protein may control the movement of the conjugate through the nanopore in a direction from the cis side to the trans side of the membrane. Thus, in such embodiments, the movement of the conjugate in a first direction is the movement of the conjugate from the cis side to the trans side of the membrane, i.e., from the cis side to the trans side of the nanopore. Upon debinding from the polynucleotide handling protein, the conjugate moves in a second direction relative to the detector. Thus, movement of the conjugate in the second direction is movement of the conjugate through the nanopore "out" of the nanopore (from the "perspective" of the polynucleotide handling protein), i.e., from the trans to the cis side of the membrane, i.e., from the trans to the cis side of the nanopore. Movement of the conjugate controls movement of the polypeptide portion of the conjugate contained therein.

[0119] Of course, the opposite setup can also be used in the methods disclosed herein. For example, as shown, the detector can include a transmembrane nanopore spanning a membrane having a cis side and a trans side. A first opening of the nanopore can be on the trans side of the membrane and a second opening of the nanopore can be on the cis side, and thus the polynucleotide handling protein controls the movement of the conjugate through the nanopore in a direction "into" the nanopore (from the "point of view" of the polynucleotide handling protein). Thus, the polynucleotide handling protein controls the movement of the conjugate through the nanopore in a direction from the trans side to the cis side of the membrane. Thus, in such an embodiment, the movement of the conjugate in the first direction is the movement of the conjugate from the trans side to the cis side of the membrane, i.e., from the trans side to the cis side of the nanopore. Upon debinding from the polynucleotide handling protein, the conjugate moves in a second direction relative to the detector. Thus, movement of the conjugate in the second direction is movement of the conjugate "out" of the nanopore (from the "perspective" of the polynucleotide handling protein), i.e., from the cis to the trans side of the membrane, i.e., from the cis to the trans side of the nanopore. Movement of the conjugate controls movement of the polypeptide portion of the conjugate contained therein.

[0120] Using similar notation as above, in some embodiments, the conjugate is m -{PN} n or L-{PN} n -P m and wherein the structure comprises one or more structures of the form - L is a leader, L is optionally an N moiety, -P is a polypeptide, -N comprises a polynucleotide, - m is 0 or 1, -n is a positive integer.

[0121] The method may include contacting the reader (L) with the detector, for example with a first opening of the detector. The reader (L) may pass through the first opening of the detector. The reader (L) may pass through the first opening and the second opening of the detector. The detector may be a nanopore and the reader (L) may pass through the nanopore.

[0122] In some such embodiments, the detector comprises a transmembrane nanopore spanning a membrane having a cis side and a trans side, and the method comprises threading a leader (L) through the nanopore, thereby contacting a polypeptide (P) with the nanopore. In some embodiments, the polynucleotide handling protein is located on the cis side of the nanopore, and the method comprises enabling the polynucleotide handling protein to control movement of the polynucleotide (N) in a direction from the cis side of the nanopore to the trans side of the nanopore, thereby controlling movement of the polypeptide (P) through the nanopore. In other embodiments, the polynucleotide handling protein is located on the trans side of the nanopore, and the method comprises enabling the polynucleotide handling protein to control movement of the polynucleotide (N) in a direction from the trans side of the nanopore to the cis side of the nanopore, thereby controlling movement of the polypeptide (P) through the nanopore.

[0123] Polypeptides As explained above, the disclosed methods involve characterizing a target polypeptide within a conjugate as the conjugate migrates relative to a detector.

[0124] Any suitable polypeptide can be characterized by the disclosed methods.

[0125] In some embodiments, the target polypeptide is an unmodified protein or portion thereof, or a naturally occurring polypeptide or portion thereof.

[0126] In some embodiments, the target polypeptide is secreted from the cell. Alternatively, the target polypeptide may be produced intracellularly, such that it must be extracted from the cell for characterization by the disclosed methods. The polypeptide may be contained within a plasmid, e.g., Sambrook et al., Molecular Cloning: A Laboratory Manual, 4 th ed., Cold Spring Harbor Press, Plainsview, New York (2012), and Ausubel et al., Current Protocols in Molecular Biology (Supplement 114), John Wiley & Sons, New York (2016).

[0127] Polypeptides may be obtained or extracted from any organism or microorganism. Polypeptides may be obtained from humans or animals, for example, from urine, lymph, saliva, mucus, semen, or amniotic fluid, or from whole blood, plasma, or serum. Polypeptides may be obtained from plants, for example, cereals, legumes, fruits, or vegetables.

[0128] The target polypeptide may be provided as an impure mixture of one or more polypeptides and one or more impurities. The impurities may include truncated forms of the target polypeptide that are different from the "target polypeptide" for characterization in the disclosed methods. For example, the target polypeptide may be a full-length protein and the impurities may include fractions of the protein. The impurities may also include proteins other than the target protein, which may be, for example, co-purified from a cell culture or obtained from a sample.

[0129] A polypeptide can contain any combination of amino acids, amino acid analogs, and modified amino acids (i.e., amino acid derivatives). The amino acids (and derivatives, analogs, etc.) in a polypeptide can be distinguished by their physical size and charge.

[0130] The amino acids / derivatives / analogs may be naturally occurring or artificial.

[0131] In some embodiments, the polypeptide may comprise any naturally occurring amino acid. Twenty amino acids are encoded by the universal genetic code. These are alanine (A), arginine (R), asparagine (N), aspartic acid (D), cysteine ​​(C), glutamic acid / glutamate (E), glutamine (Q), glycine (G), histidine (H), isoleucine (I), leucine (L), lysine (K), methionine (M), phenylalanine (F), proline (P), serine (S), threonine (T), tryptophan (W), tyrosine (Y), and valine (V). Other naturally occurring amino acids include selenocysteine ​​and pyrrolysine.

[0132] In some embodiments, the polypeptide is modified. In some embodiments, the polypeptide is modified for detection using the disclosed methods. In some embodiments, the disclosed methods are for characterizing modifications in a target polypeptide.

[0133] In some embodiments, one or more of the amino acids / derivatives / analogs in the polypeptide are modified. In some embodiments, one or more of the amino acids / derivatives / analogs in the polypeptide are post-translationally modified. Thus, the methods disclosed herein can be used to detect the presence, absence, and number of positions of post-translational modifications in a polypeptide. The methods disclosed can be used to characterize the degree to which a polypeptide is post-translationally modified.

[0134] Any one or more post-translational modifications may be present on a polypeptide. Typical post-translational modifications include modification with hydrophobic groups, modification with cofactors, addition of chemical groups, glycosylation (non-enzymatic attachment of sugars), biotinylation and pegylation. Post-translational modifications may also be non-natural, such as chemical modifications made in a laboratory for biotechnological or biomedical purposes. This may allow monitoring the levels of peptides, polypeptides or proteins produced in a laboratory, as opposed to their natural counterparts.

[0135] Examples of post-translational modifications by hydrophobic groups include myristoylation, myristic acid, C 14 Attachment of saturated acids; palmitoylation, palmitic acid, C 16 These include attachment of a saturated acid; isoprenylation or prenylation, the attachment of an isoprenoid group; farnesylation, the attachment of a farnesol group; geranylgeranylation, the attachment of a geranylgeraniol group; and glypiation, and the formation of a glycosylphosphatidylinositol (GPI) anchor via an amide bond.

[0136] Examples of post-translational modifications by cofactors include lipoylation, attachment of a lipoate (C8) functional group; flavinylation, attachment of a flavin moiety (e.g., flavin mononucleotide (FMN) or flavin adenine dinucleotide (FAD)); attachment of heme C, e.g., via a thioether bond with cysteine; phosphopantetheinylation, attachment of a 4'-phosphopantetheinyl group; and formation of a retinylidene Schiff base.

[0137] Examples of post-translational modifications by addition of chemical groups include acylation, e.g., O-acylation (ester), N-acylation (amide) or S-acylation (thioester); acetylation, e.g., attachment of an acetyl group to the N-terminus or lysine; formylation; alkylation, addition of an alkyl group such as methyl or ethyl; methylation, e.g., addition of a methyl group to lysine or arginine; amidation; butyration; gamma carboxylation; glycosylation, e.g., enzymatic attachment of a glycosyl group to arginine, asparagine, cysteine, hydroxylysine, serine, threonine, tyrosine or tryptophan; polysialylation, attachment of polysialic acid; malonylation; hydroxylation; iodination; bromination; citrulline. phosphorylation, e.g., the attachment of a phosphate group to serine, threonine or tyrosine (O-linked) or histidine (N-linked); adenylylation, e.g., the attachment of an adenylyl moiety to tyrosine (O-linked) or histidine or lysine (N-linked); propionation; the formation of pyroglutamate; S-glutathionylation; sumoylation; S-nitrosylation; succinylation, e.g., the attachment of a succinyl group to lysine; selenoylation, the incorporation of selenium; and ubiquitination, the addition of a ubiquitin subunit (N-linked).

[0138] It is within the scope of the methods provided herein that the polypeptide is labeled with a molecular label.The molecular label can be a modification to the polypeptide that facilitates the detection of the polypeptide in the methods provided herein.For example, the label can be a modification to the polypeptide that modifies the signal obtained when the conjugate is characterized.For example, the label can interfere with the flow of ions through the nanopore.In this way, the label can improve the sensitivity of the method.

[0139] In some embodiments, the polypeptide comprises one or more cross-linked sections, e.g., CC bridges. In some embodiments, the polypeptide is not cross-linked before being characterized using the disclosed methods.

[0140] In some embodiments, the polypeptide contains sulfide-containing amino acids and therefore has the potential to form disulfide bonds. Typically, in such embodiments, the polypeptide is reduced using a reagent such as DTT (dithiothreitol) or TCEP (tris(2-carboxyethyl)phosphine) before being characterized using the disclosed methods.

[0141] In some embodiments, the polypeptide is a full-length protein or a naturally occurring polypeptide. In some embodiments, the protein or naturally occurring polypeptide is fragmented before being conjugated to the polynucleotide. In some embodiments, the protein or polypeptide is chemically or enzymatically fragmented. In some embodiments, the polypeptide or polypeptide fragment may be conjugated to form a longer target polypeptide.

[0142] The polypeptide can be of any suitable length, hi some embodiments, the polypeptide has a length of about 2 to about 300 peptide units. In some embodiments, the polypeptide has a length of about 2 to about 100 peptide units, for example, about 2 to about 50 peptide units, for example, about 2 to about 40 peptide units, for example, about 2 to about 30 peptide units, for example, about 2 to about 25 peptide units, for example, about 2 to about 20 peptide units; or about 3 to about 50 peptide units, for example, about 3 to about 40 peptide units, for example, about 3 to about 30 peptide units, for example, about 3 to about 25 peptide units, for example, about 3 to about 20 peptide units; or about 5 to about 50 peptide units, for example, about 5 to about 40 peptide units, for example, about 5 to about 30 peptide units, for example, about 5 to about 25 peptide units, for example, about 5 to about 20 peptide units; for example, about 7 to about 16 peptide units, for example, about 9 to about 12 peptide units; or about 16 to about 25 peptide units, for example, about 18 to about 22 peptide units.

[0143] Any number of polypeptides can be characterized with the disclosed methods. For example, the methods can involve the characterization of 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 50, 100, or more polynucleotides. When more than one polypeptide is used, they can be different polypeptides or two or more instances of the same polypeptide.

[0144] It will thus be appreciated that the measurements made in the disclosed methods are typically characteristic of one or more features of the polypeptide selected from (i) the length of the polypeptide, (ii) the identity of the polypeptide, (iii) the sequence of the polypeptide, (iv) the secondary structure of the polypeptide, and (v) whether the polypeptide is modified. In typical embodiments, the measurements are characteristic of the sequence of the polypeptide, or whether the polypeptide is modified, for example, by one or more post-translational modifications. In some embodiments, the measurements are characteristic of the sequence of the polypeptide.

[0145] In some embodiments, the polypeptide is in a relaxed form. In some embodiments, the polypeptide is held in a linearized form. Holding the polypeptide in a linearized form can facilitate characterization of the polypeptide on a residue-by-residue basis, for example, by preventing "bunching" of the polypeptide (within the nanopore, if the detector is or includes a nanopore).

[0146] The polypeptide may be maintained in a linearized form using any suitable means.

[0147] For example, if the polypeptide is charged, applying a voltage can hold the polypeptide in a linearized form.

[0148] If the polypeptide is uncharged or only slightly charged, the charge can be altered or controlled by adjusting the pH. For example, a high pH can be used to increase the relative negative charge of the polypeptide, thereby holding it in a linearized form. Increasing the negative charge of the polypeptide can hold it in a linearized form, for example, under a positive voltage. Alternatively, a low pH can be used to increase the relative positive charge of the polypeptide, thereby holding it in a linearized form. Increasing the positive charge of the polypeptide can hold it in a linearized form, for example, under a negative voltage. In the disclosed method, a polynucleotide handling protein is used to control the movement of the polynucleotide relative to the nanopore. Since polynucleotides are typically negatively charged, it is generally most suitable to increase the linearization of the polypeptide by increasing the pH, as with polynucleotides, thereby making the polypeptide more negatively charged. In this way, the conjugate holds an overall negative charge and can therefore move easily under an applied voltage.

[0149] The polypeptide can be maintained in a linearized form by using suitable denaturing conditions. Suitable denaturing conditions include, for example, the presence of a suitable concentration of a denaturing agent, such as guanidine HCl and / or urea. The concentration of such a denaturing agent used in the disclosed method depends on the target polypeptide to be characterized in the method and can be easily selected by a person skilled in the art.

[0150] The polypeptide can be maintained in a linearized form by using a suitable detergent. Suitable detergents for use in the disclosed methods include SDS (sodium dodecyl sulfate).

[0151] The polypeptide can be maintained in a linearized form by performing the disclosed methods at elevated temperatures: increasing the temperature overcomes the intrachain bonds and allows the polypeptide to assume a linearized form.

[0152] The polypeptide may be held in a linearized form by performing the disclosed method under a strong electroosmotic force. Such a force may be provided by using asymmetric salt conditions and / or by providing a suitable charge to the environment of the polypeptide. For example, if the detector is or includes a nanopore, a charge to linearize the polypeptide may be provided to the channel of the nanopore. The charge of the channel of the protein nanopore may be altered, for example, by mutagenesis. The charge of the channel of the solid-state nanopore may be changed, for example, by chemical modification of the substrate from which the solid-state nanopore is generated. Changing the charge of the nanopore is within the capabilities of one of skill in the art. Changing the charge of the nanopore generates a strong electroosmotic force from the unbalanced flow of cations and anions through the nanopore when a potential is applied across the nanopore.

[0153] A polypeptide can be held in a linearized form by passing through a structure such as an array of nanopillars, a nanoslit, or across a nanogap, in some embodiments, the physical constraints of such structures can force the polypeptide to adopt a linearized form.

[0154] Conjugate formation As described in more detail herein, a conjugate comprises a polynucleotide conjugated to a target polypeptide.

[0155] The target polypeptide can be conjugated to the polynucleotide at any suitable position. For example, the polypeptide can be conjugated to the polynucleotide at the N-terminus or C-terminus of the polypeptide. The polypeptide can be conjugated to the polynucleotide through the side group of a residue (e.g., an amino acid residue) in the polypeptide.

[0156] In some embodiments, a target polypeptide has naturally occurring reactive functional groups that can be used to facilitate conjugation to a polynucleotide, for example, a cysteine ​​residue can be used to form a disulfide bond to a polynucleotide or a modifying group thereon.

[0157] In some embodiments, the target polypeptide is modified to facilitate its conjugation to a polynucleotide. For example, in some embodiments, the polypeptide is modified by attaching a moiety that includes a reactive functional group for attachment to a polynucleotide. For example, in some embodiments, the polypeptide can be extended at the N-terminus or C-terminus by one or more residues (e.g., amino acid residues) that include one or more reactive functional groups for reacting with corresponding reactive functional groups on a polynucleotide. For example, in some embodiments, the polypeptide can be extended at the N-terminus and / or C-terminus by one or more cysteine ​​residues. Such residues can be used for attachment to the polynucleotide portion of the conjugate, for example, by maleimide chemistry (e.g., reaction of cysteine ​​with an azido-maleimide compound such as azido-[Pol]-maleimide, where [Pol] is typically a short-chain polymer such as PEG, e.g., PEG2, PEG3, or PEG4; followed by coupling to an appropriately functionalized polynucleotide, e.g., a polynucleotide bearing a BCN group for reaction with an azide). Such chemistries are described in Example 2. For the avoidance of doubt, where a polypeptide comprises suitable naturally occurring residues at the N-terminus and / or C-terminus (e.g. naturally occurring cysteine ​​residues at the N-terminus and / or C-terminus), such residue(s) may be used for attachment to the polynucleotide.

[0158] In some embodiments, residues in the target polypeptide are modified to facilitate attachment of the target polypeptide to a polynucleotide. In some embodiments, residues (e.g., amino acid residues) in the polypeptide are chemically modified for attachment to a polynucleotide. In some embodiments, residues (e.g., amino acid residues) in the polypeptide are enzymatically modified for attachment to a polynucleotide.

[0159] The conjugation chemistry between the polynucleotide and the polypeptide in the conjugate is not particularly limited. Any suitable combination of reactive functional groups can be used. Many suitable reactive groups and their chemical targets are known in the art. Some exemplary reactive groups and their corresponding targets include arylazides that can react with amines, carbodiimides that can react with amines and carboxyl groups, hydrazides that can react with carbohydrates, hydroxymethylphosphines that can react with amines, imide esters that can react with amines, isocyanates that can react with hydroxyl groups, carbonyls that can react with hydrazines, maleimides that can react with sulfhydryl groups, NHS-esters that can react with amines, PFP-esters that can react with amines, psoralens that can react with thymine, pyridyl disulfides that can react with sulfhydryl groups, vinyl sulfones that can react with sulfhydrylamines and hydroxyl groups, vinyl sulfonamides, and the like.

[0160] Other suitable chemistries for conjugating polypeptides to polynucleotides include click chemistry. Many suitable click chemistry reagents are known in the art. Suitable examples of click chemistry include, but are not limited to: (a) Copper(I)-catalyzed azide-alkyne cycloaddition (azide-alkyne Huisgen cycloaddition), (b) Strain-promoted azide-alkyne cycloaddition; [3+2] cycloaddition of alkenes and azides; inverse demand Diels-Alder reaction of alkenes and tetrazines; and photoclick reaction of alkenes and tetrazoles. (c) Copper-free variants of the 1,3 dipolar cycloaddition reaction in which an azide reacts with a strained alkyne, e.g., in a cyclooctane ring, e.g., in bicyclic [6.1.0]nonyne (BCN); (d) reaction of an oxygen nucleophile on one linker with an epoxide or aziridine reactive moiety on the other; and (e) Staudinger ligation, in which the alkyne moiety can be replaced by an aryl phosphine, resulting in a specific reaction with an azide to give an amide bond.

[0161] Any reactive group can be used to form the conjugate. Some suitable reactive groups include [1,4-bis[3-(2-pyridyldithio)propionamido]butane; 1,1 1-bis-maleimidotriethylene glycol; 3,3'-dithiodipropionic acid di(N-hydroxysuccinimide ester); ethylene glycol-bis(succinic acid N-hydroxysuccinimide ester); 4,4'-diisothiocyanatostilbene-2,2'-disulfonic acid disodium salt; bis[2-(4-azidosalicylamido)ethyl]disulfide; 3-(2-pyridyldithio)propionic acid N-hydroxysuccinimide ester; 4-maleimidobutyric acid N-hydroxysuccinimide ester; iodoacetic acid N-hydroxysuccinimide ester; S-acetylthioglycolic acid N-hydroxysuccinimide ester; azido-PEG-maleimide; and alkyne-PEG-maleimide. The reactive group may be any of those disclosed in WO 2010 / 086602, in particular in Table 3 of that application.

[0162] In some embodiments, the reactive functional group is included in the polynucleotide and the targeting functional group is included in the polypeptide prior to the conjugation step. In other embodiments, the reactive functional group is included in the polypeptide and the targeting functional group is included in the polynucleotide prior to the conjugation step. In some embodiments, the reactive functional group is attached directly to the polypeptide. In some embodiments, the reactive functional group is attached to the polypeptide via a spacer. Any suitable spacer can be used. Suitable spacers include, for example, alkyl diamines such as ethyl diamine, and the like.

[0163] In some embodiments, the polynucleotide is directly conjugated to the polypeptide. In some embodiments, the polynucleotide is ligated to a polynucleotide linker conjugated to the polypeptide. In some embodiments, the polynucleotide is ligated to a polynucleotide having at least 50%, such as at least 60%, such as at least 70%, such as at least 80%, such as at least 90%, such as at least 95%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 29 and / or 30. The use of a linker may be beneficial to facilitate easy conjugation of the polynucleotide to the target polypeptide. A non-limiting example is provided in Example 5.

[0164] As is evident from the above discussion, in some embodiments, the conjugate comprises multiple polypeptide sections and / or multiple polynucleotide sections. For example, the conjugate may comprise a structure of the form ...-PNPNPN..., where P is a polypeptide and N is a polynucleotide. In such embodiments, the polynucleotide handling protein sequentially controls the N portion of the conjugate relative to the nanopore, thus sequentially controlling the movement of the P section relative to the nanopore, thus allowing sequential characterization of the P section. In such embodiments, multiple polynucleotides and polypeptides may be conjugated together by the same or different chemistries.

[0165] leader As described in more detail herein, the disclosed methods involve using a polynucleotide handling protein to control the movement of a polynucleotide-polypeptide conjugate relative to a detector, such as a nanopore.

[0166] As discussed herein, the leader may be included in the conjugate. The leader may be included in the conjugate by binding to the polypeptide. For example, the first end of the polypeptide may be conjugated to the polynucleotide, and the second end of the polypeptide may be attached to the leader. This is shown in FIG. 2. As described in more detail herein, the second end of the polypeptide may be attached to the leader by any suitable means.

[0167] In some embodiments, the leader is attached directly to the second end of the polypeptide, hi some embodiments, the leader is attached to the second end of the polypeptide by a linker.

[0168] Any suitable reader can be used as described herein. In some embodiments, the reader can pass through a first opening of the detector. In some embodiments, the reader can pass through a second opening of the detector. In some embodiments, the reader can pass through a first opening and a second opening of the detector. In some embodiments, the detector is or includes a nanopore and the reader can pass at least a portion of a path through the nanopore. In some embodiments, the detector is or includes a nanopore and the reader can translocate the nanopore.

[0169] In some embodiments, the leader is charged. In some embodiments, the leader is uncharged. In some embodiments, the leader is negatively or positively charged, typically negatively charged. A charged leader can be useful, for example, to pass uncharged or lower charged polypeptides through the first and / or second opening of the detector.

[0170] In some embodiments, the leader is a polymer. In some embodiments, the leader is a charged polymer, e.g., a negatively charged polymer. In some embodiments, the leader comprises a polymer, such as PEG or a polysaccharide. In such embodiments, the leader may be 10-150 monomer units (e.g., ethylene glycol or saccharide units) long, such as 20-120, such as 30-100, such as 40-80, such as 50-70 monomer units (e.g., ethylene glycol or saccharide units) long.

[0171] In some embodiments, the leader is or comprises a polynucleotide. In embodiments where the leader is a polynucleotide, the leader can be the same type of polynucleotide as the polynucleotide used in the conjugate, or the leader can be a different type of polynucleotide. For example, the polynucleotide in the conjugate can be DNA and the leader can be RNA, or vice versa. In some embodiments, the polynucleotide in the conjugate comprises or consists of DNA (e.g., dsDNA) and the leader does not consist of DNA, although in some embodiments the leader can comprise one or more nucleotides in addition to other monomeric units, e.g., the leader can comprise one or more spacers as described herein). In some embodiments, the polynucleotide in the conjugate comprises or consists of ssDNA and the leader does not consist of ssDNA (e.g., the leader can comprise one or more spacers as described herein).

[0172] In some embodiments, the leader comprises one or more spacers as described herein. In some embodiments, the leader may comprise one or more abasic spacers, i.e., one or more spacers in which a base has been removed from one or more nucleotides in the polynucleotide adaptor. In some embodiments, the leader may comprise a peptide nucleic acid (PNA), a glycerol nucleic acid (GNA), a threose nucleic acid (TNA), a locked nucleic acid (LNA), or a synthetic polymer with nucleotide side chains. In some embodiments, the leader comprises one or more nitroindoles, one or more inosines, one or more acridines, one or more 2-aminopurines, one or more 2-6-diaminopurines, one or more 5-bromo-deoxyuridines, one or more inverted thymidines (inverted dT), one or more inverted dideoxy-thymidines (ddT), one or more dideoxy-cytidines (ddC), one or more 5-methylcytidines, one or more 5-hydroxymethylcytidines, one or more 2'-O-methyl RNA bases, one or more isoform ... -deoxycytidine (Iso-dC), one or more iso-deoxyguanosine (Iso-dG), one or more C3(OC3H6OPO3) groups, one or more photocleavable (PC)[OC3H6-C(O)NHCH2-C6H3NO2-CH(CH3)OPO3] groups, one or more hexanediol groups, one or more spacer 9 (iSp9)[(OCH2CH2)3OPO3] groups, or one or more spacer 18 (iSp18)[(OCH2CH2)6OPO3] groups, or one or more thiol linkages. The leader may contain any combination of these groups. Many of these groups are commercially available from IDT® (Integrated DNA Technologies®). For example, C3, iSp9, and iSp18 spacers are all available from IDT®. The leader may contain any number of the above groups. For example, the leader can include about 5 to about 100, such as about 10 to about 50, such as about 20 to about 40, spacers described herein, such as C3, iSp9 and / or iSp18.

[0173] In some embodiments, the leader is attached directly to the polypeptide, i.e., to the second end of the polypeptide in the conjugate. In some embodiments, the leader is attached to the second end of the polypeptide by a chemical bond. In some embodiments, the leader is attached to the second end of the polypeptide by a covalent bond. In some embodiments, the leader is attached to the second end of the polypeptide by a linker. Any suitable linker may be used. In some embodiments, the linker is a polynucleotide as described herein. In some embodiments, the linker is a synthetic polymer, such as PEG. In some embodiments, the linker is the same type of polynucleotide as the polynucleotide portion of the conjugate (e.g., in some embodiments, the polynucleotide portion of the conjugate comprises DNA and the linker comprises DNA), and the leader comprises one or more nucleotides of a different type than those contained in the polynucleotide portion of the conjugate. In some embodiments, the polynucleotide portion of the conjugate comprises DNA, the linker comprises DNA, and the leader comprises one or more non-DNA nucleotides as described herein, e.g., one or more spacers as described herein.

[0174] In some embodiments, the leader is attached to the second end of the polypeptide at the N-terminus or C-terminus of the polypeptide, hi some embodiments, the leader is attached to the second end of a polypeptide side group of a residue (e.g., an amino acid residue) in the polypeptide.

[0175] In some embodiments, a target polypeptide has naturally occurring reactive functional groups that can be used to attach a leader, for example, a cysteine ​​residue can be used to form a disulfide bond to the leader or modifying group thereon.

[0176] In some embodiments, the target polypeptide is modified to facilitate its attachment to a leader. For example, in some embodiments, the polypeptide is modified by attaching a moiety that includes a reactive functional group for attachment to a leader. For example, in some embodiments, the polypeptide may be extended at the N-terminus or C-terminus by one or more residues (e.g., amino acid residues) that include one or more reactive functional groups for reacting with corresponding reactive functional groups on the leader. For example, in some embodiments, the polypeptide may be extended at the N-terminus and / or C-terminus by one or more cysteine ​​residues. Such residues may be used for attachment to the leader, for example, by maleimide chemistry (e.g., reaction of cysteine ​​with an azido-maleimide compound such as azido-[Pol]-maleimide, where [Pol] is typically a short chain polymer such as PEG, e.g., PEG2, PEG3, or PEG4; followed by coupling to an appropriately functionalized leader, e.g., via a BCN group for reaction with an azide). For the avoidance of doubt, where a polypeptide comprises suitable naturally occurring residues at the N-terminus and / or C-terminus (e.g. naturally occurring cysteine ​​residues at the N-terminus and / or C-terminus), such residue(s) may be used for attachment to the leader.

[0177] In some embodiments, residues in the target polypeptide are modified to facilitate attachment of the target polypeptide to a leader. In some embodiments, residues (e.g., amino acid residues) in the polypeptide are chemically modified for attachment to a leader. In some embodiments, residues (e.g., amino acid residues) in the polypeptide are enzymatically modified for attachment to a leader.

[0178] The attachment chemistry between the leader and the polypeptide in the conjugate is not particularly limited. Any suitable combination of reactive functional groups can be used. Many suitable reactive groups and their chemical targets are known in the art. Some exemplary reactive groups and their corresponding targets include arylazides that can react with amines, carbodiimides that can react with amines and carboxyl groups, hydrazides that can react with carbohydrates, hydroxymethylphosphines that can react with amines, imide esters that can react with amines, isocyanates that can react with hydroxyl groups, carbonyls that can react with hydrazines, maleimides that can react with sulfhydryl groups, NHS-esters that can react with amines, PFP-esters that can react with amines, psoralens that can react with thymine, pyridyl disulfides that can react with sulfhydryl groups, vinyl sulfones that can react with sulfhydrylamines and hydroxyl groups, vinyl sulfonamides, and the like.

[0179] Other suitable chemistries for attaching a leader to a polypeptide include click chemistry. Many suitable click chemistry reagents are known in the art. Suitable examples of click chemistry include, but are not limited to: (f) copper(I)-catalyzed azide-alkyne cycloaddition (azide-alkyne Huisgen cycloaddition); (g) Strain-promoted azide-alkyne cycloaddition; [3+2] cycloaddition of alkenes and azides; inverse demand Diels-Alder reaction of alkenes and tetrazines; and photoclick reaction of alkenes and tetrazoles. (h) Copper-free variants of the 1,3 dipolar cycloaddition reaction in which an azide reacts with a strained alkyne, e.g., in a cyclooctane ring, e.g., in bicyclic [6.1.0]nonyne (BCN); (i) reaction of an oxygen nucleophile on one linker with an epoxide or aziridine reactive moiety on the other; and (j) Staudinger ligation, in which the alkyne moiety can be replaced by an aryl phosphine, resulting in a specific reaction with an azide to give an amide bond.

[0180] Any reactive group can be used to attach the leader to the polypeptide. Some suitable reactive groups include [1,4-bis[3-(2-pyridyldithio)propionamido]butane; 1,1 1-bis-maleimidotriethylene glycol; 3,3'-dithiodipropionic acid di(N-hydroxysuccinimide ester); ethylene glycol-bis(succinic acid N-hydroxysuccinimide ester); 4,4'-diisothiocyanatostilbene-2,2'-disulfonic acid disodium salt; bis[2-(4-azidosalicylamido)ethyl]disulfide; 3-(2-pyridyldithio)propionic acid N-hydroxysuccinimide ester; 4-maleimidobutyric acid N-hydroxysuccinimide ester; iodoacetic acid N-hydroxysuccinimide ester; S-acetylthioglycolic acid N-hydroxysuccinimide ester; azido-PEG-maleimide; and alkyne-PEG-maleimide. The reactive group may be any of those disclosed in WO 2010 / 086602, in particular in Table 3 of that application.

[0181] In some embodiments, prior to the attachment step, the reactive functional group is included in the leader and the target functional group is included in the second end of the polypeptide. In other embodiments, prior to the conjugation step, the reactive functional group is included in the second end of the polypeptide and the target functional group is included in the leader. In some embodiments, the reactive functional group is directly attached to the polypeptide. In some embodiments, the reactive functional group is attached to the polypeptide via a spacer. Any suitable spacer can be used. Suitable spacers include, for example, alkyl diamines such as ethyl diamine, and the like.

[0182] In some embodiments, the reader is configured to facilitate unbinding of the polynucleotide binding site of the polynucleotide handling protein from the conjugate, which may be useful to facilitate re-reading of the conjugate by facilitating movement of the conjugate in a second direction relative to the detector when the polynucleotide handling protein contacts the reader.

[0183] In some embodiments, the leader does not include a blocking moiety. In some embodiments, the leader does not include a blocking moiety to prevent translocation of the conjugate through the first and second openings. In some embodiments, the leader does not include a blocking moiety as described herein. The conjugate may include a blocking moiety, for example, at the second end of the polynucleotide portion of the conjugate, but in some embodiments, the leader attached to the second end of the polypeptide does not include a blocking moiety. In some embodiments, the conjugate includes a polypeptide conjugated at a first end to a first end of a polynucleotide and attached at a second end to a leader, the second end of the polynucleotide includes a blocking moiety, and the leader does not include a blocking moiety. For example, a blocking moiety attached to the leader may in some embodiments prevent the leader from passing through the first and / or second openings of the detector, for example, preventing the leader from translocating through the nanopore.

[0184] Polynucleotides As described in more detail herein, the methods provided herein include conjugating a polypeptide to a polynucleotide and using a polynucleotide handling protein to control the movement of the conjugate relative to a detector, such as a nanopore.

[0185] Any suitable polynucleotide can be used in the disclosed methods.

[0186] In some embodiments, the polynucleotides are secreted from the cell. Alternatively, the polynucleotides can be produced intracellularly such that they must be extracted from the cell for use in the disclosed methods.

[0187] The polynucleotide may be provided as an impure mixture of one or more polynucleotides and one or more impurities. The impurities may include a truncated polynucleotide that is different from the polynucleotide used to form the conjugate. For example, the polynucleotide used to form the conjugate may be genomic DNA, and the impurities may include fractions of genomic DNA, plasmids, etc. The target polynucleotide may be a coding region of genomic DNA, and the undesired polynucleotide may include a non-coding region of DNA.

[0188] Examples of polynucleotides include DNA and RNA. The bases in DNA and RNA may be distinguished by their physical size.

[0189] A polynucleotide or nucleic acid may contain any combination of any nucleotides. The nucleotides may be naturally occurring or artificial. One or more nucleotides in a polynucleotide may be oxidized or methylated. One or more nucleotides in a polynucleotide may be damaged. For example, a polynucleotide may contain pyrimidine dimers. Such dimers are typically associated with UV damage and are the main cause of cutaneous melanoma.

[0190] One or more nucleotides in the polynucleotide may be modified, for example, with a label or tag, suitable examples of which are known to those skilled in the art. The polynucleotide may include one or more spacers. An adaptor, such as a sequencing adaptor, may be included in the polynucleotide. Adaptors, tags and spacers are described in more detail herein.

[0191] Exemplary modified bases are disclosed herein and may be incorporated into a polynucleotide by means known in the art, such as by polymerase incorporation of modified nucleotide triphosphates during strand copying (e.g., PCR) or by polymerase fill-in methods. In some embodiments, one or more bases may be modified by chemical means using reagents known in the art.

[0192] A nucleotide typically comprises a nucleobase, a sugar, and at least one phosphate group. The nucleobase and the sugar form a nucleoside. The nucleobase is typically a heterocycle. Nucleobases include, but are not limited to, purines and pyrimidines, more specifically, adenine (A), guanine (G), thymine (T), uracil (U), and cytosine (C). The sugar is typically a pentose sugar. Nucleotide sugars include, but are not limited to, ribose and deoxyribose. The sugar is preferably deoxyribose. Polynucleotides preferably comprise the following nucleosides: deoxyadenosine (dA), deoxyuridine (dU) and / or thymidine (dT), deoxyguanosine (dG), and deoxycytidine (dC). The nucleotide is typically a ribonucleotide or a deoxyribonucleotide. The nucleotide typically comprises a monophosphate, a diphosphate, or a triphosphate. A nucleotide may contain more than three phosphates, for example, four or five phosphates. The phosphate may be attached to the 5' or 3' side of the nucleotide. The nucleotides in a polynucleotide may be attached to each other in any manner. The nucleotides are typically attached by their sugar and phosphate groups, similar to nucleic acids. The nucleotides may be connected through their nucleobases, similar to pyrimidine dimers.

[0193] The polynucleotide may be double-stranded or single-stranded.

[0194] In some embodiments, the polynucleotide is single-stranded DNA. In some embodiments, the polynucleotide is single-stranded RNA. In some embodiments, the polynucleotide is a single-stranded DNA-RNA hybrid. A DNA-RNA hybrid can be prepared by ligating single-stranded DNA to RNA or vice versa. A polynucleotide is most typically a single-stranded deoxyribonucleic acid (DNA) or single-stranded ribonucleic acid (RNA).

[0195] In some embodiments, the polynucleotide is double-stranded DNA. In some embodiments, the polynucleotide is double-stranded RNA. In some embodiments, the polynucleotide is a double-stranded DNA-RNA hybrid. A double-stranded DNA-RNA hybrid can be prepared from single-stranded RNA by reverse transcribing a cDNA complement.

[0196] A polynucleotide can be of any length. For example, a polynucleotide can be at least 10, at least 50, at least 100, at least 150, at least 200, at least 250, at least 300, at least 400, or at least 500 nucleotides or nucleotide pairs in length. A polynucleotide can be 1000 nucleotides or nucleotide pairs or more in length, 5000 nucleotides or nucleotide pairs or more in length, or 100000 nucleotides or nucleotide pairs or more in length.

[0197] More typically, the polynucleotide has a length of about 1 to about 10,000 nucleotides or nucleotide pairs, for example, about 1 to about 1000 nucleotides or nucleotide pairs (e.g., about 10 to about 1000 nucleotides or nucleotide pairs), for example, about 5 to about 500 nucleotides or nucleotide pairs, for example, about 10 to about 100 nucleotides or nucleotide pairs, for example, about 20 to about 80 nucleotides or nucleotide pairs, for example, about 30 to about 50 nucleotides or nucleotide pairs.

[0198] Any number of polynucleotides can be used in the disclosed methods. For example, the methods can include using 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 50, 100, or more polynucleotides. When two or more polynucleotides are used, they can be different polynucleotides or two instances of the same polynucleotide. The polynucleotides can be naturally occurring or artificial.

[0199] The nucleotides may have any identity, including, but not limited to, adenosine monophosphate (AMP), guanosine monophosphate (GMP), thymidine monophosphate (TMP), uridine monophosphate (UMP), 5-methylcytidine monophosphate, 5-hydroxymethylcytidine monophosphate, cytidine monophosphate (CMP), cyclic adenosine monophosphate (cAMP), cyclic guanosine monophosphate (cGMP), deoxyadenosine monophosphate (dAMP), deoxyguanosine monophosphate (dGMP), deoxythymidine monophosphate (dTMP), deoxyuridine monophosphate (dUMP), deoxycytidine monophosphate (dCMP), and deoxymethylcytidine monophosphate. The nucleotides are preferably selected from AMP, TMP, GMP, CMP, UMP, dAMP, dTMP, dGMP, dCMP, and dUMP. A nucleotide can be abasic (i.e., lacking a nucleobase). A nucleotide can also lack a nucleobase and a sugar (i.e., a C3 spacer).

[0200] The polynucleotide may include products of PCR reactions, genomic DNA, products of endonuclease digestion, and / or DNA libraries. The polynucleotide may be obtained or extracted from any organism or microorganism. The polynucleotide may be obtained from humans or animals, for example, from urine, lymph, saliva, mucus, semen, or amniotic fluid, or from whole blood, plasma, or serum. The polynucleotide may be obtained from plants, for example, from cereals, legumes, fruits, or vegetables. The polynucleotide may include genomic DNA. The genomic DNA may be fragmented. The DNA may be fragmented by any suitable method. For example, methods for fragmenting DNA are known in the art, and such methods may use transposases, such as MuA transposase. Genomic DNA is often not fragmented.

[0201] It is within the scope of the methods provided herein that the polynucleotide is labeled with a molecular label. The molecular label can be a modification to the polynucleotide that facilitates the detection of the polynucleotide or conjugate in the methods provided herein. For example, the label can be a modification to the polynucleotide that alters the signal obtained when the conjugate is characterized. For example, the label can interfere with the flow of ions through the nanopore. In this way, the label can improve the sensitivity of the method.

[0202] adapter In some embodiments of the methods provided herein, the polynucleotide included in the conjugate has a polynucleotide adaptor attached thereto. The adaptor typically comprises a polynucleotide strand that can be attached to an end of a polynucleotide.

[0203] In some embodiments, the adaptor is attached to the polynucleotide before the conjugate with the polypeptide is formed, hi some embodiments, the adaptor is attached to the conjugate of the polynucleotide and the polypeptide.

[0204] In some embodiments, the conjugate comprises a leader, and the leader comprises a polynucleotide. In such embodiments, an adaptor may be attached to the leader. In some embodiments, the conjugate comprises a leader, and the leader is attached to the polypeptide of the conjugate by a linker that is a polynucleotide. In such embodiments, the adaptor may be attached to the linker. In some embodiments, the leader is included in an adaptor attached to the conjugate. In such embodiments, the adaptor may be attached to the polypeptide or polynucleotide portion of the conjugate. The adaptor provided with the leader and optionally attached to the conjugate (e.g., attached to the polypeptide portion of the conjugate) may be referred to as a "tail." The tail may comprise a leader as described herein. In some embodiments, the tail comprises a Y adaptor as described herein. In some embodiments, the tail comprises one or more polynucleotides having at least 50%, such as at least 60%, such as at least 70%, such as at least 80%, such as at least 90%, such as at least 95%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NOs: 31 and / or 32.

[0205] In some embodiments, the method includes attaching an adaptor (e.g., an adaptor described herein) to a polynucleotide and forming a conjugate by conjugating the polynucleotide / adapter construct to a target polypeptide. In some embodiments, the conjugate is formed by attaching an adaptor (e.g., an adaptor described herein) to a polynucleotide and forming the conjugate by attaching the adaptor to the target polypeptide.

[0206] In some embodiments, the adaptors may be selected or modified to provide specific sites for conjugation to a polynucleotide.

[0207] An adaptor can be attached to only one end of a polynucleotide or conjugate. A polynucleotide adaptor can be added to both ends of a polynucleotide or conjugate. Alternatively, different adaptors can be added to the two ends of a polynucleotide or conjugate.

[0208] Adapters can be added to both strands of double-stranded polynucleotides. Adapters can be added to single-stranded polynucleotides. Methods for adding adapters to polynucleotides are known in the art. Adapters can be attached to polynucleotides, for example, by ligation, by click chemistry, by tagmentation, by topoisomerase conversion, or by any other suitable method.

[0209] In one embodiment, the or each adapter is synthetic or artificial. Typically, the or each adapter comprises a polymer as described herein. In some embodiments, the or each adapter comprises a spacer as described herein, such as one or more of C3, iSp9 and iSp18 spacer units. In some embodiments, the or each adapter comprises a polynucleotide. The or each polynucleotide adapter may comprise DNA, RNA, modified DNA (e.g., abasic DNA), RNA, PNA, LNA, BNA, and / or PEG. Typically, the or each adapter comprises single-stranded and / or double-stranded DNA or RNA. An adapter may comprise the same type of polynucleotide as the polynucleotide strand to which it is attached. An adapter may comprise a different type of polynucleotide than the polynucleotide strand to which it is attached. In some embodiments, the polynucleotide strand used in the disclosed methods is a single-stranded DNA strand and the adapter comprises DNA or RNA, typically single-stranded DNA. In some embodiments, the polynucleotide is a double-stranded DNA strand and the adapter comprises DNA or RNA, e.g., double-stranded or single-stranded DNA.

[0210] In some embodiments, the adapter is double stranded and both strands are contiguous. In some embodiments, the adapter is double stranded and at least one strand is non-contiguous. In some embodiments, one of the strands is damaged or broken. In some embodiments, the adapter is formed of three strands, two of which hybridize to a third strand. For example, the adapter may include a first strand (e.g., which may be referred to as the "top strand") conjugated to a second strand (e.g., which may be referred to as the "bottom strand 1") at a first portion of the first strand and conjugated to a third strand (e.g., which may be referred to as the "bottom strand 2") at a second portion of the first strand. The first and second portions may abut. The first and second portions may be separated, for example, by a non-hybridized portion of the polynucleotide. The first and second portions may be separated, for example, by at least 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 40, 50, or more nucleotides. A non-hybridized portion of the first strand may occur when the combined length of the second strand and the third strand is shorter than the length of the first strand. A non-hybridized portion of the first strand may occur when the second strand and / or the third strand comprises a sequence that is non-complementary to the first strand. In some embodiments, the adapter comprises one or more polynucleotides having at least 50%, such as at least 60%, such as at least 70%, such as at least 80%, such as at least 90%, such as at least 95%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NOs: 26, 27, and / or 28. A non-limiting example is shown in FIG. 14A, where the hybridized portion formed by bottom strand 1 and bottom strand 2 is separated by a non-hybridized portion of 3 nucleotides in length that arises from a portion of the non-complementary polynucleotides of top strand and bottom strand 1.

[0211] In some embodiments, the adaptor can be a bridging moiety. A bridging moiety can be used to connect two strands of a double-stranded polynucleotide.For example, in some embodiments, a bridging moiety is used to connect the template strand of a double-stranded polynucleotide to the complementary strand of the double-stranded polynucleotide.

[0212] The bridging moiety typically covalently links the two strands of a double-stranded polynucleotide. The bridging moiety can be anything that can link the two strands of a double-stranded polynucleotide, provided that the bridging moiety does not interfere with the translocation of the polynucleotide relative to the nanopore. Suitable bridging moieties include, but are not limited to, polymer linkers, chemical linkers, polynucleotides, or polypeptides. Preferably, the bridging moiety comprises DNA, RNA, modified DNA (e.g., abasic DNA), RNA, PNA, LNA, or PEG. More preferably, the bridging moiety is DNA or RNA.

[0213] In some embodiments, the bridging moiety is a hairpin adaptor. A hairpin adaptor is an adaptor that comprises a single polynucleotide strand, where the ends of the polynucleotide strand can hybridize to each other or are hybridized to each other, forming a loop in the middle of the polynucleotide. Suitable hairpin loop adaptors can be designed using methods known in the art. In some embodiments, the hairpin loop is typically 4-100 nucleotides in length, e.g., 4-50, e.g., 4-20, e.g., 4-8 nucleotides in length. In some embodiments, the bridging moiety (e.g., hairpin adaptor) is attached to one end of the double-stranded polynucleotide. The bridging moiety (e.g., hairpin adaptor) is typically not attached to both ends of the double-stranded polynucleotide.

[0214] In some embodiments, the adaptor is a linear adaptor. The linear adaptor can be attached to either or both ends of a single stranded polynucleotide. If the polynucleotide is a double stranded polynucleotide, the linear adaptor can be attached to either or both ends of either or both strands of the double stranded polynucleotide. The linear adaptor can include a leader as described herein. The linear adaptor can include a moiety for hybridization with a tag (such as a pore tag) as described herein. The linear adaptor can be 10-150 nucleotides in length, such as 20-120, such as 30-100, such as 40-80, such as 50-70 nucleotides in length. The linear adaptor can be single stranded. The linear adaptor can be double stranded.

[0215] In some embodiments, the adaptor can be a Y adaptor. The Y adaptor is typically a polynucleotide adaptor. The Y adaptor is typically double-stranded and includes (a) a region at one end where the two strands are hybridized together, and (b) a region at the other end where the two strands are not complementary. The non-complementary portions of the strands typically form an overhang. The presence of the non-complementary region in the Y adaptor gives the adaptor a Y shape, since unlike the double-stranded portion, the two strands typically do not hybridize to each other. The two single-stranded portions of the Y adaptor can be the same length or different lengths. For example, one single-stranded portion of the Y adaptor can be 10-150 nucleotides in length, e.g., 20-120, e.g., 30-100, e.g., 40-80, e.g., 50-70 nucleotides in length, and the other single-stranded portion of the Y adaptor can independently be 10-150 nucleotides in length, e.g., 20-120, e.g., 30-100, e.g., 40-80, e.g., 50-70 nucleotides in length. The double-stranded "stem" portion of the Y adaptor can be, e.g., 10-150 nucleotides in length, e.g., 20-120, e.g., 30-100, e.g., 40-80, e.g., 50-70 nucleotides in length. In some embodiments, the single-stranded portion of the Y adaptor can include a leader as described herein.

[0216] The adaptor may be linked to the polynucleotide by any suitable means known in the art. The adaptor may be synthesized separately and chemically attached to the target polynucleotide or enzymatically ligated thereto. Alternatively, the adaptor may be generated during processing of the polynucleotide. In some embodiments, the adaptor is linked to the target polynucleotide at or near one end of the polynucleotide. In some embodiments, the adaptor is linked to the polynucleotide within 50, e.g., within 20, e.g., within 10 nucleotides of the end of the polynucleotide. In some embodiments, the adaptor is linked to the polynucleotide at the end of the polynucleotide. When the adaptor is linked to the polynucleotide, the adaptor may contain the same type of nucleotide as the polynucleotide or may contain different nucleotides relative to the polynucleotide.

[0217] Adaptors particularly suitable for use in the disclosed methods can include a linear homopolymer region (e.g., about 5 to about 20 nucleotides, e.g., about 10 to about 30 nucleotides, e.g., thymine or cytidine) and / or modified nucleotides, e.g., spacers or abasic nucleotides, as described herein. Such regions can be useful as leaders, as described herein.

[0218] In some embodiments, the adaptors may include hybridization sites for hybridizing to one or more tethers or anchors (described in more detail herein).

[0219] In some embodiments, the adaptor may include one or more reactive functional groups for binding to the target polypeptide. Click chemistry groups are particularly suitable in this regard. For example, exemplary groups for inclusion in the adaptor include groups that can particulate in copper-free click chemistry, such as groups based on BCN (bicyclo[6.1.0]nonyne) and its derivatives, dibenzocyclooctyne (DBCO) groups, and the like. Reactions of such groups are well known in the art. For example, BCN groups typically react with groups such as azides, tetrazines, and nitrones that can be incorporated into polypeptides. DBCO groups are highly reactive toward azide groups. Other particularly suitable chemical groups include 2-pyridinecarboxaldehyde (2-PCA) groups and their derivatives. For example, 6-(azidomethyl)-2-pyridinecarboxaldehyde can react with the N-terminal amino group of a peptide.

[0220] In some embodiments, the polynucleotide adaptor has a binding site for binding a polynucleotide handling protein described herein, in some embodiments, the binding site comprises a stall moiety for stalling the polynucleotide handling protein on the adaptor.

[0221] In some embodiments, the polynucleotide adaptor comprises a blocking moiety as described herein.

[0222] In some embodiments, the polynucleotide adaptor comprises both a polynucleotide handling protein and a blocking moiety. For example, in some embodiments, the polynucleotide adaptor may comprise a first end comprising an attachment point for attachment to a second end of a polynucleotide portion of a conjugate, and a second end, and the polynucleotide handling protein may be stalled on the polynucleotide adaptor. In some embodiments, the second end of the adaptor may comprise a blocking moiety.

[0223] In some embodiments, the polynucleotide adaptor may comprise a first end comprising an attachment point for attachment to a second end of a polypeptide moiety of a conjugate, and a second end, and the polynucleotide handling protein may be stalled on the adaptor. In some embodiments, the second end of the adaptor may comprise a leader (e.g., attached via a polymer linker as described herein).

[0224] In some embodiments, the conjugate described herein comprises a first adaptor attached to its polynucleotide portion and a second adaptor attached to its polypeptide portion. In some embodiments, the first adaptor comprises a binding site for a polynucleotide handling protein and has a polynucleotide handling protein bound thereto and a blocking moiety to prevent the polynucleotide handling protein from disassociating from the conjugate. In some embodiments, the first adaptor comprises a continuous top strand between the blocking moiety and the attachment to the polypeptide. In some embodiments, the first adaptor comprises a non-continuous bottom strand. In some embodiments, prior to step (i) of the disclosed method, the polynucleotide handling protein is bound to the first adaptor between the blocking moiety and the portion of the top strand that is hybridized to the non-continuous bottom strand. In some embodiments, the second adaptor comprises a polynucleotide Y adaptor having a portion that defines a polynucleotide leader. In some embodiments, the leader comprises a plurality of spacer units as described herein, e.g., a plurality (e.g., about 10 to about 150, e.g., about 20 to about 100, e.g., about 25 to about 50, e.g., about 30 to about 40) of C3, iSp9 and / or iSp18 spacer units.

[0225] In some embodiments, a polynucleotide adaptor comprises a single-stranded polynucleotide region adjacent to one or more double-stranded polynucleotide regions (e.g., a double-stranded polynucleotide may be attached to each end of a single-stranded polynucleotide), where the single-stranded polynucleotide region may bind to (and / or have a motor protein or polynucleotide binding protein that binds, e.g., by stalling thereon), and the region optionally comprises one or more stall moieties as described herein. In some embodiments, the adaptor is attached to a "clinker" region as described herein, where the clinker optionally comprises a double-stranded polynucleotide. In some embodiments, the clinker comprises an attachment point for attachment to a target polypeptide. In some embodiments, the target polypeptide may be attached to a single-stranded polynucleotide "tail" that may comprise a leader as described herein. In some embodiments, the leader and / or adaptor comprises a blocking moiety as described herein.

[0226] Spacer In some embodiments of the methods provided herein, the polynucleotide, conjugate formed by reaction with the polypeptide, or adapter described herein may include a spacer. For example, one or more spacers may be present in the polynucleotide adapter. For example, the polynucleotide adapter may include 1 to about 20 spacers, e.g., 1 to about 10, e.g., 1 to about 5 spacers, e.g., 1, 2, 3, 4, or 5 spacers. The spacer may include any suitable number of spacer units. The spacer may provide an energy barrier that impedes the movement of the polynucleotide handling protein. For example, the spacer may stall the polynucleotide handling protein by reducing the traction force of the polynucleotide handling protein on the polynucleotide. This may be accomplished, for example, by using an abasic spacer, i.e., a spacer in which a base has been removed from one or more nucleotides in the polynucleotide adapter. The spacer may physically block the movement of the polynucleotide handling protein, for example, by introducing a large chemical group that physically impedes the movement of the polynucleotide handling protein.

[0227] In some embodiments, one or more spacers are included in the polynucleotides or conjugates, or in the adaptors, as used in the methods claimed herein, to provide a distinctive signal as they pass through or traverse the nanopore, i.e., as they translocate relative to the nanopore.

[0228] In some embodiments, the spacer may comprise a linear molecule such as a polymer. Typically, such a spacer has a structure different from the polynucleotide used in the conjugate. For example, when the polynucleotide is DNA, the or each spacer typically does not comprise DNA. In particular, when the polynucleotide is deoxyribonucleic acid (DNA) or ribonucleic acid (RNA), the or each spacer preferably comprises peptide nucleic acid (PNA), glycerol nucleic acid (GNA), threose nucleic acid (TNA), locked nucleic acid (LNA), or a synthetic polymer with nucleotide side chains. In some embodiments, the spacer is one or more nitroindoles, one or more inosines, one or more acridines, one or more 2-aminopurines, one or more 2-6-diaminopurines, one or more 5-bromo-deoxyuridines, one or more inverted thymidines (inverted dT), one or more inverted dideoxy-thymidines (ddT), one or more dideoxy-cytidines (ddC), one or more 5-methylcytidines, one or more 5-hydroxymethylcytidines, one or more 2'-O-methyl RNA bases, one or more isopropyl ethers, one or more tert-butyl ... -deoxycytidine (Iso-dC), one or more iso-deoxyguanosine (Iso-dG), one or more C3(OC3H6OPO3) groups, one or more photocleavable (PC)[OC3H6-C(O)NHCH2-C6H3NO2-CH(CH3)OPO3] groups, one or more hexanediol groups, one or more spacer 9 (iSp9)[(OCH2CH2)3OPO3] groups, or one or more spacer 18 (iSp18)[(OCH2CH2)6OPO3] groups, or one or more thiol bonds. The spacer may include any combination of these groups. Many of these groups are commercially available from IDT® (Integrated DNA Technologies®). For example, C3, iSp9, and iSp18 spacers are all available from IDT®. The spacer may include any number of the above groups as spacer units.

[0229] In some embodiments, the spacer may comprise one or more chemical groups that stall the polynucleotide handling protein. In some embodiments, suitable chemical groups are one or more pendant chemical groups. One or more chemical groups may be attached to one or more nucleobases in the polynucleotide, construct or adapter. One or more chemical groups may be attached to the backbone of the polynucleotide adapter. There may be any number of suitable chemical groups, such as 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12 or more. Suitable groups include, but are not limited to, fluorophores, streptavidin and / or biotin, cholesterol, methylene blue, dinitrophenol (DNP), digoxigenin and / or antidigoxigenin, and dibenzylcyclooctyne groups. In some embodiments, the spacer may comprise a polymer. In some embodiments, the spacer may comprise a polymer that is a polypeptide or polyethylene glycol (PEG).

[0230] In some embodiments, the spacer may contain one or more abasic nucleotides (i.e., nucleotides lacking a nucleobase), for example, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12 or more abasic nucleotides. The nucleobase may be replaced by -H (idSp) or -OH in the abasic nucleotide. The abasic spacer may be inserted into the target polynucleotide by removing the nucleobase from one or more adjacent nucleotides. For example, the polynucleotide may be modified to contain 3-methyladenine, 7-methylguanine, 1,N6-ethenoadenine inosine, or hypoxanthine, and the nucleobase may be removed from these nucleotides using human alkyladenine DNA glycosylase (hAAG). Alternatively, the polynucleotide may be modified to contain uracil, and the nucleobase may be removed by uracil DNA glycosylase (UDG). In one embodiment, the one or more spacers do not contain any abasic nucleotides.

[0231] In some embodiments, the polynucleotide handling protein may be stalled with a spacer of the polynucleotide portion of the conjugate prior to the methods disclosed herein, e.g., prior to step (i) and / or step (ii) of the disclosed methods. In some embodiments in which a blocking moiety as described herein is present at a second end of the polynucleotide and a first end of the polynucleotide is conjugated to a first end of a polypeptide in a conjugate, one or more spacers may be included between the polypeptide and the blocking moiety, and the polynucleotide handling protein may be stalled with such a spacer prior to the methods disclosed herein, e.g., prior to step (i) and / or step (ii) of the disclosed methods.

[0232] Methods for stalling polynucleotide handling proteins, such as helicases, on polynucleotide adaptors using spacers are described in International Application No. WO 2014 / 135838, which is incorporated herein by reference in its entirety.

[0233] anchor In some embodiments, the polynucleotide, its conjugate with a polypeptide, or an adaptor attached thereto may include, for example, a membrane anchor or a transmembrane pore anchor attached to the adaptor. In one embodiment, the anchor aids in characterization of the conjugate by the methods disclosed herein. For example, the membrane anchor or the transmembrane pore anchor may facilitate localization of the conjugate around a nanopore in a membrane.

[0234] The anchor can be a polypeptide anchor and / or a hydrophobic anchor that can insert into the membrane. In one embodiment, the hydrophobic anchor is a lipid, a fatty acid, a sterol, a carbon nanotube, a polypeptide, a protein, or an amino acid, such as cholesterol, palmitate, or tocopherol. The anchor can include a thiol, biotin, or a surfactant.

[0235] In some embodiments, the anchor can be biotin (for binding to streptavidin), amylose (for binding to maltose binding protein or fusion proteins), Ni-NTA (for binding to poly-histidine or poly-histidine tagged proteins), or a peptide (such as an antigen).

[0236] In one embodiment, the anchor may include one linker, or two, three, four or more linkers. Preferred linkers include, but are not limited to, polymers, such as polynucleotides, polyethylene glycol (PEG), polysaccharides, and polypeptides. These linkers may be linear, branched, or cyclic. For example, the linker may be a cyclic polynucleotide. The adaptor may hybridize to a complementary sequence on a cyclic polynucleotide linker. One or more anchors or one or more linkers may include a moiety that can be cleaved or degraded, such as a restriction site or a photolabile group. The linker may be functionalized with a maleimide group to attach to a cysteine ​​residue of a protein. Suitable linkers are described in International Application No. 2010 / 086602.

[0237] In one embodiment, the anchor is cholesterol or a fatty acyl chain. Any fatty acyl chain having a length of 6 to 30 carbon atoms, such as, for example, hexadecanoic acid, may be used. Examples of suitable anchors and methods for attaching the anchors to the adaptors are disclosed in WO 2012 / 164270 and WO 2015 / 150786.

[0238] Control of conjugate transfer As described in more detail above, the methods provided herein include contacting a conjugate with a polynucleotide handling protein and performing one or more measurements characteristic of the polypeptide as the conjugate moves relative to a detector, such as a nanopore. The polynucleotide handling protein is used to control the movement of the conjugate as it moves in a first direction relative to first and second openings of the detector. When the conjugate decouples from the polynucleotide handling protein, the conjugate moves in a second direction relative to the detector.

[0239] A polynucleotide handling protein typically controls the movement of a conjugate by controlling the movement of the polynucleotide contained in the conjugate. Thus, a polynucleotide handling protein can typically control the movement of the polynucleotide portion of the conjugate. Since the polynucleotide is conjugated to a polypeptide, controlled movement of the polynucleotide relative to the detector results in controlled movement of the polypeptide, and thus, in controlled movement of the entire conjugate relative to the detector. Thus, reference herein to a polynucleotide handling protein that controls the movement of a conjugate should be understood (unless otherwise required by context) as referring to a polynucleotide handling protein that controls the movement of the polynucleotide portion of the conjugate, thereby controlling the movement of the conjugate. Thus, a polynucleotide handling protein can control the movement of a conjugate by controlling the movement of the polynucleotide portion of the conjugate.

[0240] However, it is not excluded that polynucleotide handling protein can directly control the movement of the polypeptide part of the conjugate.For example, as described herein, polynucleotide handling protein can function as a molecular brake in some embodiments, thereby controlling the movement of the polypeptide part of the conjugate that moves relative to the detector.However, usually, the controlled movement of the conjugate by polynucleotide handling protein is achieved when polynucleotide handling protein controls the movement of the polynucleotide contained in the polypeptide.

[0241] Thus, movement of the conjugate in a first direction is determined by the polynucleotide handling protein, as described in more detail herein, and movement of the conjugate in a second direction relative to the detector can be driven by any suitable means.

[0242] In some embodiments, the movement of the conjugates is driven by a physical or chemical force (electrical potential). In some embodiments, the physical force is provided by an electrical (e.g., voltage) potential or a temperature gradient, etc.

[0243] In some embodiments, the direction of movement in the first direction is opposite to the direction of the force applied to the detector, hi some embodiments, the direction of movement in the second direction is the same as the direction of the force applied to the detector.

[0244] In some embodiments, when a potential is applied to the detector, the conjugate moves relative to the detector. Because the polynucleotide is negatively charged, when a potential is applied to a detector, such as a nanopore, the polynucleotide moves relative to the detector under the influence of the applied potential.

[0245] For example, if the detector is or includes a nanopore and a positive potential is applied to the trans side of the nanopore relative to the cis side of the nanopore, negatively charged analytes are induced to move from the cis side of the nanopore to the trans side of the nanopore. Similarly, if a positive potential is applied to the trans side of the nanopore relative to the cis side of the nanopore, this prevents movement of negatively charged analytes from the trans side of the nanopore to the cis side of the nanopore. The reverse occurs if a negative potential is applied to the trans side of the nanopore relative to the cis side of the nanopore. Apparatus and methods for applying suitable voltages are described in more detail herein.

[0246] In some embodiments, the chemical force is provided by a concentration (eg, pH) gradient.

[0247] In some embodiments, the polynucleotide handling protein controls the movement of the conjugate in the same direction as the physical or chemical force (electrical potential). For example, in some embodiments, the detector is or includes a nanopore, a positive electric potential is applied to the trans side of the nanopore relative to the cis side of the nanopore, and the polynucleotide handling protein controls the movement of the conjugate from the cis side of the nanopore to the trans side of the nanopore. In some embodiments, a positive electric potential is applied to the cis side of the nanopore relative to the trans side of the nanopore, and the polynucleotide handling protein controls the movement of the conjugate from the trans side of the nanopore to the cis side of the nanopore.

[0248] More typically, in the disclosed methods, the polynucleotide handling protein controls the movement of the conjugate in a direction opposite to the physical or chemical force (electrical potential). For example, in some embodiments, the detector is or includes a nanopore, a positive electric potential is applied to the trans side of the nanopore relative to the cis side of the nanopore, and the polynucleotide handling protein controls the movement of the conjugate from the trans side of the nanopore to the cis side of the nanopore. In some embodiments, a positive electric potential is applied to the cis side of the nanopore relative to the trans side of the nanopore, and the polynucleotide handling protein controls the movement of the conjugate from the cis side of the nanopore to the trans side of the nanopore.

[0249] In some embodiments, movement of the conjugate is driven by the polynucleotide handling protein in the absence of an applied potential.

[0250] Polynucleotide Handling Proteins In the disclosed methods, the polynucleotide handling protein can control the movement of the polynucleotide relative to a detector, such as a nanopore. In other words, the polynucleotide handling protein can control the movement of a conjugate. In some embodiments, the polynucleotide handling protein can control the movement of a polynucleotide and a polypeptide.

[0251] Suitable polynucleotide handling proteins are also known as motor proteins or polynucleotide handling enzymes. Suitable polynucleotide handling proteins are known in the art, and some exemplary polynucleotide handling proteins are described in more detail below.

[0252] In one embodiment, the motor protein is or is derived from a polynucleotide handling enzyme. A polynucleotide handling enzyme is a polypeptide that can interact with a polynucleotide and modify at least one property thereof. The enzyme may modify a polynucleotide by cleaving the polynucleotide to form individual nucleotides or shorter chains of nucleotides, such as di- or trinucleotides. The enzyme may modify a polynucleotide by orienting it or moving it to a specific location.

[0253] In some embodiments, the polynucleotide handling protein can be present on the conjugate prior to contact with a detector such as a nanopore. For example, the polynucleotide handling protein can be present on the polynucleotide in the conjugate. In some embodiments, the polynucleotide handling protein can be present on an adapter that comprises part of the conjugate or can be present on a portion of the conjugate.

[0254] In some embodiments, the polynucleotide handling protein can remain bound to the conjugate if the portion of the conjugate that is in contact with the active site of the polynucleotide handling protein comprises a polypeptide. In other words, in some embodiments, the polynucleotide handling protein does not dissociate from the conjugate. In some embodiments, the polynucleotide handling protein does not dissociate from the conjugate when the polynucleotide handling protein contacts the polypeptide portion of the conjugate. In some embodiments, the polynucleotide handling protein is free to move relative to the polypeptide portion until it is contacted by one or more subsequent polynucleotide portions of the conjugate.

[0255] The term "deassociation" as used herein refers to dissociation of the polynucleotide handling protein from the conjugate. Thus, the polynucleotide handling protein may be modified to prevent it from dissociating from the conjugate, e.g., into the reaction medium. It is important to distinguish the potential "deassociation" of the polynucleotide handling protein from, e.g., the "deassociation" of the polynucleotide handling protein from the polynucleotide portion of the conjugate. As used herein, "deassociation" refers to the temporary release of the conjugate (e.g., its polynucleotide portion) from the active site of the polynucleotide handling protein (described in more detail herein), but does not mean deassociation. Thus, for example, the polynucleotide handling protein may be modified to prevent the polynucleotide handling protein from disassociating from the conjugate, but not to prevent the polynucleotide handling protein from deassociating from the conjugate, e.g., from the polynucleotide portion of the conjugate. When deassociated, the polynucleotide handling protein remains associated with the target polynucleotide. For example, a polynucleotide handling protein may maintain association with a conjugate (i.e., may be prevented from disassociating from the conjugate) because it is topologically closed around the conjugate, e.g., around the polynucleotide portion of the conjugate. The polynucleotide binding site may remain free to bind or unbind to the conjugate (e.g., the polynucleotide portion of the conjugate) while the polynucleotide handling protein remains associated with the conjugate, such that the polynucleotide handling protein may bind or unbind to the conjugate. When the polynucleotide handling protein disassociates from the conjugate (e.g., from the polynucleotide portion of the conjugate), it may move over (e.g., along) the conjugate (e.g., the polypeptide and / or polynucleotide portion of the conjugate) under an applied force and may rebind to the conjugate (e.g., the polynucleotide portion of the conjugate).When associated with a conjugate, but is disassociated from the conjugate (e.g., disassociated from the polynucleotide portion of the conjugate), the polynucleotide handling protein is unable to disassociate from the target polynucleotide.

[0256] In some embodiments, a polynucleotide handling protein is modified to prevent disassociation from a conjugate, polynucleotide, or adaptor (other than by passing off the end of the conjugate, polynucleotide, or adaptor, unless a blocking moiety is used, as described in more detail herein) when the polynucleotide handling protein contacts a portion of the conjugate that includes a polypeptide. Such modified polynucleotide handling proteins are particularly suitable for use in the disclosed methods.

[0257] The polynucleotide handling protein can be adapted in any suitable manner. For example, the polynucleotide handling protein can be modified to prevent it from disassociating after it has been loaded onto a polynucleotide, conjugate, or adapter. Alternatively, the polynucleotide handling protein can be modified to prevent the protein from disassociating before loading onto a polynucleotide, conjugate, or adapter. Modification of the polynucleotide handling protein to prevent it from disassociating from a polynucleotide, conjugate, or adapter can be achieved using methods known in the art, for example, the methods discussed in International Application Nos. 2014 / 013260 and 2015 / 110813, each of which is incorporated herein by reference in its entirety, and with particular reference to the passages describing the modification of polynucleotide handling proteins (polynucleotide binding proteins), such as helicases, to prevent the polynucleotide handling protein from disassociating from a polynucleotide chain. For example, the polynucleotide handling protein and / or polynucleotide binding protein can be modified by treatment with tetramethylazodicarboxamide (TMAD). A variety of other blocking moieties are described in more detail herein.

[0258] For example, a polynucleotide handling protein may have a polynucleotide debinding opening, e.g., a cavity, groove, or gap, through which a polynucleotide strand can pass when the polynucleotide handling protein is disassociated from the strand. In some embodiments, the polynucleotide debinding opening of a given motor protein (polynucleotide handling protein) can be determined by reference to its structure, e.g., by reference to its X-ray crystal structure. The X-ray crystal structure can be obtained in the presence and / or absence of a polynucleotide substrate. In some embodiments, the location of the polynucleotide debinding opening in a given polynucleotide handling protein can be predicted or confirmed by molecular modeling using standard packages known in the art. In some embodiments, the polynucleotide debinding opening can be generated transiently by movement of one or more portions of the polynucleotide handling protein, e.g., one or more domains.

[0259] The polynucleotide handling protein (motor protein) may be modified by closing the polynucleotide unbinding opening. Thus, closing the polynucleotide unbinding opening may not only prevent the polynucleotide handling protein from unbinding from the polypeptide portion of the conjugate, but also prevent it from unbinding from the polynucleotide or adaptor. For example, the polynucleotide handling protein may be modified by covalently closing the polynucleotide unbinding opening. In some embodiments, the polynucleotide handling protein for such purposes is a helicase as described herein. Thus, in some embodiments of the disclosed method, the polynucleotide handling protein is modified to fully or partially close the opening present in at least one conformational state of the unmodified protein through which the polynucleotide strand can unbind. However, as described herein, closing the polynucleotide unbinding opening does not necessarily prevent the conjugate from unbinding from the polynucleotide binding site of the polynucleotide handling protein.

[0260] In some embodiments, the polynucleotide handling protein may be modified to prevent the conjugate from disassociating from the polynucleotide handling protein. The polynucleotide handling protein may be modified in any suitable manner.

[0261] Without being bound by theory, the inventors believe that promoting debinding and slowing rebinding may promote re-reading of the polypeptide contained in the conjugate. Without being bound by theory, the inventors believe that this may be because each step that the polynucleotide handling protein takes on the conjugate (e.g., the polynucleotide of the conjugate) is associated with the probability that the polynucleotide handling protein will debind from the conjugate. Such debinding probability may be identified by the so-called off-rate. It is believed that an increase in the off-rate promotes the drop back of the polynucleotide handling protein on the conjugate. Similarly, again without being bound by theory, the inventors believe that the distance that the polynucleotide handling protein can travel along the conjugate before rebinding, once debinding from the target polynucleotide, is associated with the on-rate. Thus, re-reading may be promoted by increasing the off-rate and decreasing the on-rate of the polynucleotide handling protein with respect to the conjugate (e.g., its polynucleotide portion). Tuning the off-rate and on-rate of a polynucleotide handling protein for a given type of conjugate (e.g., a given type of polynucleotide) is within the ability of one of skill in the art in light of the disclosure herein. Thus, a polynucleotide handling protein may be modified to facilitate debinding of a conjugate (e.g., its polynucleotide portion) from the polynucleotide binding site of the polynucleotide handling protein and / or to delay rebinding of a conjugate (e.g., its polynucleotide portion) to the polynucleotide binding site of the polynucleotide handling protein. In some embodiments, a polynucleotide handling protein is modified to facilitate debinding of a conjugate (e.g., its polynucleotide portion) from the polynucleotide binding site of the polynucleotide handling protein and to delay rebinding of a conjugate (e.g., its polynucleotide portion) to the polynucleotide binding site of the polynucleotide handling protein.

[0262] In some embodiments, polynucleotide handling proteins may be modified with a closing moiety to (i) topologically close the polynucleotide binding site around a conjugate (e.g., a polynucleotide portion thereof) and (ii) facilitate debinding of the conjugate (e.g., a polynucleotide portion thereof) from the polynucleotide binding site of the polynucleotide handling protein and / or retard rebinding of the conjugate (e.g., a polynucleotide portion thereof) to the polynucleotide binding site of the polynucleotide handling protein. The polynucleotide handling protein may be modified in any suitable manner to facilitate attachment of such a closing moiety.

[0263] In some embodiments, the closing moiety may comprise a bifunctional crosslinking moiety. The closing moiety may comprise a bifunctional crosslinker. The bifunctional crosslinker may attach at two points on the polynucleotide handling protein and close the polynucleotide debinding opening of the polynucleotide handling protein, thereby preventing disassociation of the conjugate from the polynucleotide handling protein while allowing debinding of the conjugate from the polynucleotide binding site of the polynucleotide handling protein.

[0264] The closing moiety may be attached at any suitable position on the polynucleotide handling protein. For example, the closing moiety may bridge two amino acid residues of the polynucleotide handling protein. Typically, at least one amino acid bridged by the closing moiety is a cysteine ​​or a non-natural amino acid. The cysteine ​​or non-natural amino acid may be introduced into the polynucleotide handling protein by substitution or modification of a naturally occurring amino acid residue of the polynucleotide handling protein. Methods for introducing non-natural amino acids are well known in the art and include, for example, native chemical ligation with a synthetic polypeptide chain that includes such non-natural amino acid. Methods for introducing cysteines into polynucleotide handling proteins are also described, for example, in Sambrook et al., Molecular Cloning: A Laboratory Manual, 4 th ed., Cold Spring Harbor Press, Plainsview, New York (2012), and Ausubel et al., Current Protocols in Molecular Biology (Supplement 114), John Wiley & Sons, New York (2016).

[0265] In some embodiments, the closing moiety has a length of about 1 Å to about 100 Å. The length of the closing moiety can be calculated according to static bond lengths or, more preferably, using molecular dynamics simulations. The length can be, for example, about 2 Å to about 80 Å, such as about 5 Å to about 50 Å, such as about 8 to about 30 Å, such as about 10 to about 25 Å or about 20 Å, such as about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, or 19 Å.

[0266] Without being bound by theory in any way, the inventors believe that, in general, a longer blocking moiety may increase the off-rate of the polynucleotide handling protein from the conjugate, thus facilitating re-reading.

[0267] In some embodiments, the closing moiety comprises a bond. In some embodiments, the closing moiety comprises a disulfide bond. The disulfide bond can be formed by treating the polynucleotide handling protein with any suitable reagent, such as TMAD.

[0268] In some embodiments, the closing moiety comprises a reagent that forms a bond between two click chemistry groups on the polynucleotide handling protein. Examples of click chemistry reagents are provided herein.

[0269] In some embodiments, the closing moiety comprises a protein. For example, a biotin group may be present on the polynucleotide handling protein and the closing moiety may comprise streptavidin. A tag, such as a snoop tag or a spy tag, may be present on the polynucleotide handling protein and the closing moiety may comprise a protein, such as a snoop catcher or a spy catcher, respectively.

[0270] In some embodiments, the closing moiety comprises a structure of the formula [ABC], where A and C are each independently reactive functional groups for reacting with an amino acid residue in a polynucleotide handling protein, and B is a linking moiety. In some embodiments, the closing moiety comprises a bond between a thio group, e.g., a thiol group on a cysteine ​​residue. Thus, in some embodiments, A and C are cysteine-reactive functional groups.

[0271] In some embodiments, the linking moiety B comprises a linear or branched, unsubstituted or substituted alkylene, alkenylene, alkynylene, arylene, heteroarylene, carbocyclylene, or heterocyclylene moiety, which is optionally interrupted or terminated by one or more atoms or groups selected from O, N(R), S, C(O), C(O)NR, C(O)O, unsubstituted or substituted arylene, arylene-alkylene, heteroarylene, heteroarylene-alkylene, carbocyclylene, carbocyclylene-alkylene, heterocyclylene, and heterocyclylene-alkylene, where R is selected from H, unsubstituted or substituted alkyl, and unsubstituted or substituted aryl. Typically, R is H or methyl, more typically H.

[0272] Typically, the alkylene group is 1-20 Typically, the alkenylene group is 2-20 Typically, the alkynylene group is 2-20 An arylene group is typically an alkynylene group. 6-12 Typically, the heteroarylene group is a 5- to 12-membered heteroarylene group. Typically, the carbocyclylene group is 5-12 A carbocyclylene group. Typically, the heterocyclylene group is a 5- to 12-membered heterocyclylene group.

[0273] Typically, the alkylene, alkenylene, or alkynylene moiety may be uninterrupted, interrupted, or terminated by one or more atoms or groups selected from O, N(R), S, C(O), C(O)NR, and C(O)O and unsubstituted or substituted arylene. Usually, the alkylene, alkenylene, or alkynylene moiety may be uninterrupted, interrupted, or terminated by one or more atoms or groups selected from O and N(R) and unsubstituted or substituted arylene. More often, the alkylene, alkenylene, or alkynylene moiety may be uninterrupted, interrupted, or terminated by one or more O atoms.

[0274] For example, the linking moiety is often an unsubstituted or substituted C 1-10 Alkylene, C 2-10 Alkenylene or C 2-10 An alkynylene moiety is uninterrupted or interrupted or terminated with one or more O atoms.

[0275] In some embodiments, linking moiety B comprises an alkylene, oxyalkylene, or polyoxyalkylene group, and / or A and C are each maleimide groups. The alkylene, oxyalkylene, or polyoxyalkylene group can have a length, for example, from about 5 Å to about 50 Å, such as from about 8 to about 30 Å, for example, from about 10 to about 25 Å.

[0276] For example, the linking moiety is (CH2CH2O) x where x is 1 to 10, e.g., 1 to 5, e.g., 1, 2 or 3. Exemplary linking moieties are described in Example 9 and include, for example, BMOE (1,2-bismaleimidoethane), BMOP (1,3-bismaleimidopropane), BMB (1,4-bismaleimidobutane), BM(PEG)2 (1,8-bismaleimido-diethylene glycol) and BM(PEG)3 (1,11-bismaleimido-triethylene glycol).

[0277] Polynucleotide handling proteins suitable for being closed using such closing moieties are discussed in more detail herein. In some preferred embodiments, the polynucleotide handling protein is a helicase, such as the Dda helicase described herein.

[0278] The polynucleotide handling protein can be selected or chosen according to the polynucleotide used in the conjugate characterized in the method disclosed herein. Alternatively, the polynucleotide can be selected or chosen according to the polynucleotide handling protein used to control the movement of the conjugate. For example, when the polynucleotide is DNA, typically, a DNA handling protein can be used. When the polynucleotide is RNA, an RNA handling protein can be used. When the polynucleotide is a hybrid of DNA and RNA, a polynucleotide handling protein that can process both DNA and RNA can be used.

[0279] In one embodiment, the polynucleotide handling protein is derived from any member of Enzyme Classification (EC) groups 3.1.11, 3.1.13, 3.1.14, 3.1.15, 3.1.16, 3.1.21, 3.1.22, 3.1.25, 3.1.26, 3.1.27, 3.1.30, and 3.1.31.

[0280] In some embodiments of the claimed methods, the polynucleotide handling protein is a helicase, polymerase, exonuclease, topoisomerase, or a variant thereof.

[0281] In one embodiment, the polynucleotide handling protein is an exonuclease. Suitable enzymes include, but are not limited to, exonuclease I from E. coli (SEQ ID NO: 1), exonuclease III enzyme from E. coli (SEQ ID NO: 2), RecJ from T. thermophilus (SEQ ID NO: 3) and bacteriophage lambda exonuclease (SEQ ID NO: 4), TatD exonuclease, and variants thereof. Three subunits comprising the sequence shown in SEQ ID NO: 3 or variants thereof interact to form a trimeric exonuclease.

[0282] In one embodiment, the polynucleotide handling protein is a polymerase. The polymerase can be PyroPhage® 3173 DNA polymerase (commercially available from Lucigen® Corporation), SD polymerase (commercially available from Bioron®), Klenow from NEB, or variants thereof. In one embodiment, the enzyme is Phi29 DNA polymerase (SEQ ID NO:5) or variants thereof. Modified versions of Phi29 polymerase that can be used in the disclosed methods are disclosed in U.S. Patent No. 5,576,204.

[0283] In embodiments of the methods provided herein that include controlling the movement of a conjugate by synthesizing a strand complementary to the polynucleotide, the polynucleotide handling protein is typically a polymerase, such as a polymerase described herein.

[0284] In one embodiment, the polynucleotide handling protein is a topoisomerase. In one embodiment, the topoisomerase is a member of any of the subclassification (EC) groups 5.99.1.2 and 5.99.1.3. The topoisomerase can be a reverse transcriptase, an enzyme that can catalyze the formation of cDNA from an RNA template. They are commercially available, for example, from New England Biolabs® and Invitrogen®.

[0285] In one embodiment, the polynucleotide handling protein is a helicase. Any suitable helicase can be used according to the methods provided herein. For example, the or each polynucleotide handling protein used according to the present disclosure can be independently selected from Hel308 helicase, RecD helicase, TraI helicase, TrwC helicase, XPD helicase, and Dda helicase, or variants thereof. Monomeric helicases can include several domains attached together. For example, TraI helicase and TraI subgroup helicase can include two RecD helicase domains, a relaxase domain, and a C-terminal domain. These domains typically form monomeric helicases that can function without forming oligomers. Specific examples of suitable helicases include Hel308, NS3, Dda, UvrD, Rep, PcrA, Pif1, and TraI. These helicases typically act on single-stranded DNA. Examples of helicases that can translocate along both strands of double-stranded DNA include FtfK and hexamer enzyme complexes, or multi-subunit complexes such as RecBCD. In one embodiment, the motor protein is a Dda (DNA-dependent ATPase) helicase.

[0286] Hel308 helicase is described in publications such as International Application No. 2013 / 057495, the entire contents of which are incorporated by reference. RecD helicase is described in publications such as International Application No. 2013 / 098562, the entire contents of which are incorporated by reference. XPD helicase is described in publications such as International Application No. 2013 / 098561, the entire contents of which are incorporated by reference. Dda helicase is described in publications such as International Application No. 2015 / 055981 and International Application No. 2016 / 055777, the respective contents of which are incorporated by reference in their entirety.

[0287] In one embodiment, the helicase comprises the sequence set forth in SEQ ID NO:6 (Trwc Cba) or a variant thereof, the sequence set forth in SEQ ID NO:7 (Hel308 Mbu) or a variant thereof, or the sequence set forth in SEQ ID NO:8 (Dda) or a variant thereof. The variants may differ from the native sequence in any of the ways discussed herein. An exemplary variant of SEQ ID NO:8 includes E94C / A360C. A further exemplary variant of SEQ ID NO:8 includes E94C / A360C, followed by (ΔM1)G1G2 (i.e., deletion of M1, followed by addition of G1 and G2).

[0288] In some embodiments, a polynucleotide handling protein (e.g., a helicase) has at least two active modes of operation (including all components required for the polynucleotide handling protein to facilitate translocation, e.g., ATP and Mg as discussed herein). 2+ The translocation of the conjugate can be controlled in one inactive mode of operation (when the polynucleotide handling protein is equipped with fuel and cofactors such as ribosomal protein) and one inactive mode of operation (when the polynucleotide handling protein is equipped with the components necessary to facilitate translocation).

[0289] When equipped with all the necessary components to facilitate translocation (i.e., in active mode), the polynucleotide handling protein (e.g., helicase) translocates along the polynucleotide in a 5' to 3' or 3' to 5' direction (depending on the polynucleotide handling protein). The polynucleotide handling protein can be used to move the conjugate away (e.g., out) from a detector such as a nanopore (e.g., against an applied force) or to move the conjugate toward (e.g., in) a detector such as a nanopore (e.g., with an applied force). For example, when the end of the conjugate that the polynucleotide handling protein is moving toward is captured by a detector such as a nanopore, the polynucleotide handling protein acts against the direction of the force to pull the threaded conjugate out of the pore (e.g., into the cis chamber). However, when the end of the polynucleotide handling protein moving away is captured by a detector such as a nanopore, the polynucleotide handling protein acts according to the direction of the force and pushes the screw-like conjugate into the pore (e.g., into the trans chamber).

[0290] When the polynucleotide handling protein (e.g., helicase) does not have the necessary components to facilitate translocation (i.e., in the inactive mode), the polynucleotide handling protein can bind to the conjugate and act as a brake to slow the translocation of the construct as it moves relative to the nanopore, for example by being drawn into a detector such as a nanopore by a force. In the inactive mode, it does not matter which end of the conjugate is captured, but rather the applied force that determines the movement of the conjugate relative to the detector is important, and the polynucleotide binding protein acts as a brake. In the inactive mode, the control of the translocation of the conjugate by the polynucleotide binding protein can be described in several ways, including ratcheting, sliding, and braking.

[0291] Polynucleotide handling proteins typically require fuel to handle the processing of polynucleotides. The fuel is typically a free nucleotide or a free nucleotide analog. Free nucleotides include adenosine monophosphate (AMP), adenosine diphosphate (ADP), adenosine triphosphate (ATP), guanosine monophosphate (GMP), guanosine diphosphate (GDP), guanosine triphosphate (GTP), thymidine monophosphate (TMP), thymidine diphosphate (TDP), thymidine triphosphate (TTP), uridine monophosphate (UMP), uridine diphosphate (UDP), uridine triphosphate (UTP), cytidine monophosphate (CMP), cytidine diphosphate (CDP), cytidine triphosphate (CTP), cyclic adenosine monophosphate (cAMP), cyclic guanosine monophosphate (cGMP), deoxyadenosine monophosphate (dAMP), deoxyadeno ...cytidine monophosphate (CMP), cytidine diphosphate (CDP), cytidine triphosphate (CTP), cyclic adenosine monophosphate (cAMP), cyclic guanosine monophosphate (cGMP), deoxyadenosine monophosphate (dAMP), deoxyadenosine triphosphate (GTP), cytidine monophosphate (CAMP), cytidine diphosphate (CGMP), deoxyadenosine monophosphate (DAMP), deoxyadenosine triphosphate (GTP), cytidine monophosphate (CAMP), cytidine diphosphate (CGMP), deoxyadeno The free nucleotide may be, but is not limited to, one or more of adenosine diphosphate (dADP), deoxyadenosine triphosphate (dATP), deoxyguanosine monophosphate (dGMP), deoxyguanosine diphosphate (dGDP), deoxyguanosine triphosphate (dGTP), deoxythymidine monophosphate (dTMP), deoxythymidine diphosphate (dTDP), deoxythymidine triphosphate (dTTP), deoxyuridine monophosphate (dUMP), deoxyuridine diphosphate (dUDP), deoxyuridine triphosphate (dUTP), deoxycytidine monophosphate (dCMP), deoxycytidine diphosphate (dCDP), and deoxycytidine triphosphate (dCTP). The free nucleotide is usually selected from AMP, TMP, GMP, CMP, UMP, dAMP, dTMP, dGMP, or dCMP. The free nucleotide is typically adenosine triphosphate (ATP).

[0292] A cofactor of a polynucleotide handling protein is a factor that enables the polynucleotide handling protein to function. The cofactor is preferably a divalent metal cation. The divalent metal cation is preferably Mg 2+ , Mn 2+ , Ca 2+ , or Co 2+ The cofactor is most preferably Mg2+ It is.

[0293] Blocking part As discussed above, the polynucleotide handling protein can be prevented from disassociating from the conjugate during the methods disclosed herein.

[0294] In some embodiments, the polynucleotide handling protein is prevented from disassociating from the conjugate by a blocking moiety. In some embodiments, the blocking moiety is included at the second end of the polynucleotide included in the conjugate. In some embodiments, the blocking moiety is included in the polynucleotide of the conjugate. In some embodiments, the blocking moiety is included in a polynucleotide adaptor attached to the polynucleotide of the conjugate. In some embodiments, the polynucleotide adaptor, such as the polynucleotide adaptors described herein, includes a blocking moiety.

[0295] In some embodiments, the blocking moiety is to prevent the polynucleotide handling protein from disassociating from the conjugate.

[0296] The blocking moiety is typically too large to pass through the polynucleotide handling protein (e.g., through a polynucleotide binding site of a polynucleotide handling protein that has been modified to topologically close the polynucleotide binding site around the conjugate), so that when movement of the conjugate relative to the detector brings the blocking moiety into contact with the polynucleotide handling protein (e.g., when the conjugate debinds from the polynucleotide binding site of the polynucleotide handling protein), further movement of the conjugate through the polynucleotide handling protein, and thus typically through the detector, is prevented. Thus, in some embodiments, the blocking moiety restricts movement of the conjugate through the polynucleotide binding site of the polynucleotide handling protein, thereby restricting movement of the conjugate in a second direction relative to the detector. At such time, the polynucleotide handling protein may rebind to the polynucleotide binding site of the conjugate. The conjugate may move through the pore in the reverse direction under the control of the polynucleotide handling protein.

[0297] Any suitable blocking moiety can be used in the provided methods. For example, the conjugate can be modified with biotin and the blocking moiety can be, for example, traptavidin, streptavidin, avidin, or neutravidin. The blocking moiety can be a large chemical group such as a dendrimer. The blocking moiety can be a nanoparticle or a bead. The blocking moiety can include one or more of the following: - a polynucleotide secondary structure, preferably a hairpin or G-quadruplex (TBA), - nucleic acid analogues, preferably selected from peptide nucleic acids (PNAs), glycerol nucleic acids (GNAs), threose nucleic acids (TNAs), locked nucleic acids (LNAs), bridged nucleic acids (BNAs) and abasic nucleotides, - fluorophores, avidins such as traptavidin, streptavidin and neutravidin, and / or biotin, cholesterol, methylene blue, dinitrophenol (DNP), digoxigenin and / or antidigoxigenin and dibenzylcyclooctyne groups, and -Polynucleotide binding proteins.

[0298] The blocking moiety may be attached to the polynucleotide in any suitable manner. The blocking moiety may be attached directly to the polynucleotide. The blocking moiety may be attached to the polynucleotide via a linker. Any suitable linker may be used. Any suitable chemical reaction for linking the blocking moiety to the polynucleotide may be used. Any of the attachment methods described herein for attaching a leader to a polypeptide may be used to attach the blocking moiety to the polynucleotide. For the avoidance of doubt, in embodiments where a conjugate includes both a leader attached to a polypeptide as described herein and a blocking moiety attached to a polynucleotide of a conjugate as described herein, the attachment means may be the same or different. In some embodiments, a linker is used to attach the leader to the polypeptide and the blocking moiety is attached directly to the polynucleotide. In some embodiments, the leader is attached directly to the polypeptide and a linker is used to attach the blocking moiety to the polynucleotide. In some embodiments, the leader is attached directly to the polypeptide and the blocking moiety is attached directly to the polynucleotide. In some embodiments, a linker is used to attach the leader to the polypeptide and a linker is used to attach the blocking moiety to the polynucleotide. When linkers are used to attach both the blocking moiety and the leader to the conjugate, the linkers can be the same or different.

[0299] setting In some embodiments, the conjugate comprises a polypeptide conjugated to the 3' end of a polynucleotide. In such embodiments, the first end of the polynucleotide is the 3' end of the polynucleotide. In such embodiments, the second end of the polynucleotide is the 5' end of the polynucleotide. In some embodiments, the C-terminus of the polypeptide may be conjugated to the 3' end of the polynucleotide. In such embodiments, the first end of the polypeptide may be the C-terminus. In such embodiments, the second end of the polypeptide may be the N-terminus. In other embodiments, the N-terminus of the polypeptide may be conjugated to the 3' end of the polynucleotide. In such embodiments, the first end of the polypeptide may be the N-terminus. In such embodiments, the second end of the polypeptide may be the C-terminus.

[0300] In some embodiments, the conjugate comprises a polypeptide conjugated to the 5'-terminus of a polynucleotide. In such embodiments, the first terminus of the polynucleotide is the 5'-terminus of the polynucleotide. In such embodiments, the second terminus of the polynucleotide is the 3'-terminus of the polynucleotide. In some embodiments, the C-terminus of the polypeptide may be conjugated to the 5'-terminus of the polynucleotide. In such embodiments, the first terminus of the polypeptide may be the C-terminus. In such embodiments, the second terminus of the polypeptide may be the N-terminus. In other embodiments, the N-terminus of the polypeptide may be conjugated to the 5'-terminus of the polynucleotide. In such embodiments, the first terminus of the polypeptide may be the N-terminus. In such embodiments, the second terminus of the polypeptide may be the C-terminus.

[0301] In some embodiments, the conjugate comprises a polypeptide conjugated to the 3'-terminus of a polynucleotide, wherein the C-terminus of the polypeptide is conjugated to the 3'-terminus of the polynucleotide and the leader is attached (e.g., directly or via a linker as described herein) to the N-terminus of the polypeptide. In such embodiments, the first terminus of the polynucleotide is the 3'-terminus of the polynucleotide, the second terminus of the polynucleotide is the 5'-terminus of the polynucleotide, the first terminus of the polypeptide is the C-terminus of the polypeptide and the second terminus of the polypeptide is the N-terminus of the polypeptide. In some such embodiments, the blocking moiety is at the 5'-terminus of the polynucleotide.

[0302] In some embodiments, the conjugate comprises a polypeptide conjugated to the 3'-terminus of a polynucleotide, wherein the N-terminus of the polypeptide is conjugated to the 3'-terminus of the polynucleotide and the leader is attached (e.g., directly or via a linker as described herein) to the C-terminus of the polypeptide. In such embodiments, the first terminus of the polynucleotide is the 3'-terminus of the polynucleotide, the second terminus of the polynucleotide is the 5'-terminus of the polynucleotide, the first terminus of the polypeptide is the N-terminus of the polypeptide and the second terminus of the polypeptide is the C-terminus of the polypeptide. In some such embodiments, the blocking moiety is at the 5'-terminus of the polynucleotide.

[0303] In some embodiments, the conjugate comprises a polypeptide conjugated to the 5'-terminus of a polynucleotide, wherein the C-terminus of the polypeptide is conjugated to the 5'-terminus of the polynucleotide and the leader is attached (e.g., directly or via a linker as described herein) to the N-terminus of the polypeptide. In such embodiments, the first terminus of the polynucleotide is the 5'-terminus of the polynucleotide, the second terminus of the polynucleotide is the 3'-terminus of the polynucleotide, the first terminus of the polypeptide is the C-terminus of the polypeptide and the second terminus of the polypeptide is the N-terminus of the polypeptide. In some such embodiments, the blocking moiety is at the 5'-terminus of the polynucleotide.

[0304] In some embodiments, the conjugate comprises a polypeptide conjugated to the 5'-terminus of a polynucleotide, wherein the N-terminus of the polypeptide is conjugated to the 5'-terminus of the polynucleotide and the leader is attached (e.g., directly or via a linker as described herein) to the C-terminus of the polypeptide. In such embodiments, the first terminus of the polynucleotide is the 5'-terminus of the polynucleotide, the second terminus of the polynucleotide is the 3'-terminus of the polynucleotide, the first terminus of the polypeptide is the N-terminus of the polypeptide and the second terminus of the polypeptide is the C-terminus of the polypeptide. In some such embodiments, the blocking moiety is at the 5'-terminus of the polynucleotide.

[0305] In some embodiments, the polynucleotide handling protein binds to the conjugate or the polynucleotide portion of the conjugate. In some embodiments, the polynucleotide handling protein binds to the conjugate or the polynucleotide portion of the conjugate prior to step (i) and / or step (ii) of the disclosed methods. The polynucleotide handling protein may be stalled on the conjugate with a stall moiety, e.g., a spacer as described herein.

[0306] In some embodiments, prior to the disclosed methods (e.g., prior to step (i) and / or step (ii) of the disclosed methods), the polynucleotide handling protein is attached to a portion of the conjugate between the polypeptide and the blocking moiety. This is shown in Figure 3. The polynucleotide handling protein can be stalled on the conjugate with a stall moiety, e.g., a spacer as described herein.

[0307] In some embodiments, the polynucleotide handling protein may bind to the portion of the conjugate between the polypeptide and the leader, for example, between the polypeptide and the "free" end of the leader (i.e. the end of the leader that is not attached to the polypeptide). This is shown in Figure 4.

[0308] In some embodiments, the polynucleotide handling protein binds to a leader included in the conjugate. In some embodiments, the polynucleotide handling protein binds to a leader included in the conjugate prior to step (i) and / or step (ii) of the method.

[0309] In some embodiments, the polynucleotide handling protein is coupled to a linker that attaches a leader to the polypeptide portion of the conjugate. In some embodiments, the linker is or comprises a polynucleotide. In some embodiments, the polynucleotide handling protein is coupled to a linker that attaches a leader to the polypeptide portion of the conjugate prior to step (i) and / or step (ii) of the method.

[0310] In some embodiments, the conjugate comprises a polypeptide conjugated to the 3' end of a polynucleotide, wherein the C-terminus of the polypeptide is conjugated to the 3' end of the polynucleotide, the leader is attached (e.g., directly or via a linker as described herein) to the N-terminus of the polypeptide, and a blocking moiety is present at the 5' end of the polynucleotide, and prior to step (i) of the disclosed methods, a polynucleotide handling protein binds to the portion of the conjugate between the polypeptide and the blocking moiety, or a polynucleotide handling protein binds to the portion of the conjugate between the polypeptide and the free end of the leader.

[0311] In some embodiments, the conjugate comprises a polypeptide conjugated to the 3' end of a polynucleotide, wherein the N-terminus of the polypeptide is conjugated to the 3' end of the polynucleotide, the leader is attached (e.g., directly or via a linker as described herein) to the C-terminus of the polypeptide, and a blocking moiety is present at the 5' end of the polynucleotide, and prior to step (i) of the disclosed methods, a polynucleotide handling protein binds to the portion of the conjugate between the polypeptide and the blocking moiety, or a polynucleotide handling protein binds to the portion of the conjugate between the polypeptide and the free end of the leader.

[0312] In some embodiments, the conjugate comprises a polypeptide conjugated to the 5' end of a polynucleotide, wherein the C-terminus of the polypeptide is conjugated to the 5' end of the polynucleotide, the leader is attached (e.g., directly or via a linker as described herein) to the N-terminus of the polypeptide, and a blocking moiety is present at the 3' end of the polynucleotide, and prior to step (i) of the disclosed methods, a polynucleotide handling protein binds to the portion of the conjugate between the polypeptide and the blocking moiety, or a polynucleotide handling protein binds to the portion of the conjugate between the polypeptide and the free end of the leader.

[0313] In some embodiments, the conjugate comprises a polypeptide conjugated to the 5' end of a polynucleotide, wherein the N-terminus of the polypeptide is conjugated to the 5' end of the polynucleotide, the leader is attached (e.g., directly or via a linker as described herein) to the C-terminus of the polypeptide, and a blocking moiety is present at the 3' end of the polynucleotide, and prior to step (i) of the disclosed methods, a polynucleotide handling protein binds to the portion of the conjugate between the polypeptide and the blocking moiety, or a polynucleotide handling protein binds to the portion of the conjugate between the polypeptide and the free end of the leader.

[0314] In some embodiments in which a polynucleotide handling protein is attached to a leader or to a linker attached to a leader prior to the disclosed methods (e.g., prior to step (i) and / or step (ii) of the disclosed methods), step (i) of the disclosed methods can include contacting a detector with the conjugate under conditions such that the polynucleotide handling protein binds to the polynucleotide portion of the conjugate. For example, step (i) of the disclosed methods can include contacting a detector with the conjugate under conditions such that the polynucleotide handling protein moves along the conjugate (e.g., on the polypeptide portion of the conjugate) and binds to the polynucleotide portion of the conjugate.

[0315] Detector In the methods provided herein, the polynucleotide moves relative to a detector, such as a nanopore. The detector can be selected from (i) a zero mode waveguide, (ii) a field effect transistor, optionally a nowire field effect transistor, (iii) an AFM tip, (iv) a nanotube, optionally a carbon nanotube, and (v) a nanopore. Preferably, the detector is a nanopore.

[0316] The polypeptide portion of the conjugate may be characterized in any suitable manner in the methods provided herein. In one embodiment, the polypeptide and / or conjugate is characterized by detecting ion flow or an optical signal as the conjugate moves relative to the nanopore. This is described in more detail herein. This method is suitable for these and other methods of detecting the polypeptide.

[0317] In another non-limiting example, in one embodiment, the polypeptide and / or conjugate is characterized by detecting the by-product of a processing reaction, such as sequencing by synthesis reaction. Thus, the method may include detecting the product of successive addition of (poly)nucleotides by an enzyme, such as a polymerase, to a nucleic acid strand in the conjugate. The product may be a change in one or more properties of the enzyme, such as the conformation of the enzyme. Thus, such a method may include subjecting an enzyme, such as a polymerase or reverse transcriptase, to a double-stranded polynucleotide under conditions such that the template-dependent incorporation of nucleotide bases into the growing oligonucleotide strand causes a conformational change of the enzyme in response to successively encountered templates, incorporation of stranded nucleic acid bases and / or incorporation of template-specific natural or similar bases (i.e., incorporation events), detection of the conformational change of the enzyme in response to such incorporation events, and thereby detection of the sequence of the template strand. In such a method, a polynucleotide strand may be displaced according to the methods provided herein. Such methods may include detecting and / or measuring uptake events using methods known to those of skill in the art, such as those described in US2017 / 0044605.

[0318] In another embodiment, the by-product may be labeled such that when a nucleotide is added to a synthetic nucleic acid strand complementary to the polynucleotide strand of the conjugate, a phosphate-labeled species is released and the phosphate-labeled species is detected using a detector as described herein. The polynucleotide thus characterized may be translocated according to the methods herein. A suitable label may be an optical label that is detected using a nanopore or a zero-mode waveguide, or by Raman spectroscopy or other detector. A suitable label may be a non-optical label that is detected using a nanopore or other detector.

[0319] In another approach, the nucleoside phosphates (nucleotides) are not labeled, and natural by-product species are detected upon addition of the nucleotide to a synthetic nucleic acid strand complementary to the polynucleotide strand of the conjugate. Suitable detectors can be ion-sensitive field effect transistors, or other detectors.

[0320] These and other detection methods are suitable for use in the methods described herein. A detector can be used to make any suitable measurement as the conjugate moves relative to the detector.

[0321] Nanopore In embodiments of the invention in which the detector is a nanopore, any suitable nanopore can be used, hi one embodiment, the nanopore is a transmembrane pore.

[0322] A transmembrane pore is a structure that traverses a membrane to some extent. It allows hydrated ions driven by an applied electric potential to flow across or within the membrane. A transmembrane pore typically traverses the entire membrane, allowing hydrated ions to flow from one side of the membrane to the other side of the membrane. However, a transmembrane pore does not have to traverse the membrane. It may be closed at one end. For example, a pore may be a well, gap, channel, trench, or slit in a membrane along or into which hydrated ions can flow.

[0323] Any transmembrane pore may be used in the methods provided herein. The pore may be biological or artificial. Suitable pores include, but are not limited to, protein pores, polynucleotide pores, and solid-state pores. In an embodiment, the solid-state pore may comprise a nanochannel. In some embodiments, the solid-state pore is a pore disclosed in International Application No. 2003 / 003446, International Application No. 2009 / 020682, or International Application No. 2016 / 187519, each of which is incorporated by reference in its entirety.

[0324] In some embodiments, the pore may be a DNA origami pore (Langecker et al., Science, 2012;338:932-936). Suitable DNA origami pores are disclosed in International Application No. 2013 / 083983, International Application No. 2018 / 011603, and International Application No. 2020 / 025974, each of which is incorporated by reference in its entirety.

[0325] In one embodiment, the nanopore is a scaffolded polypeptide nanopore. In some embodiments, the pore is a scaffolded polypeptide nanopore as disclosed in International Application No. WO 2020 / 025909 or International Application No. WO 2020 / 074399, each of which is incorporated by reference in their entirety.

[0326] In one embodiment, the nanopore is a transmembrane protein pore. A transmembrane protein pore is a polypeptide or an assembly of polypeptides that allows hydrated ions, such as polynucleotides, to flow from one side of a membrane to the other side of the membrane. In the methods provided herein, the transmembrane protein pore can form a pore that allows hydrated ions driven by an applied potential to flow from one side of a membrane to the other side. The transmembrane protein pore preferably allows polynucleotides to flow from one side of a membrane, such as a triblock copolymer membrane, to the other side. The transmembrane protein pore allows polynucleotides to translocate through the pore.

[0327] In one embodiment, the nanopore is a transmembrane protein pore that is monomeric or oligomeric. The pore is preferably composed of several repeating subunits, such as at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, or at least 16 subunits. The pore is preferably a hexameric, heptameric, octameric, or nanomeric pore. The pore may be a homo-oligomer or a hetero-oligomer.

[0328] In one embodiment, a transmembrane protein pore comprises a barrel or channel through which ions can flow. The subunits of the pore typically surround a central axis and contribute strands to a transmembrane β-barrel or channel or a transmembrane α-helical bundle or channel.

[0329] Typically, the barrel or channel of the transmembrane protein pore comprises amino acids that facilitate interaction with an analyte, such as a target polynucleotide (as described herein). These amino acids are preferably located near the constriction of the barrel or channel. The transmembrane protein pore typically comprises one or more positively charged amino acids, such as arginine, lysine, or histidine, or aromatic amino acids, such as tyrosine or tryptophan. These amino acids typically facilitate interaction between the pore and a nucleotide, polynucleotide, or nucleic acid.

[0330] In one embodiment, the nanopore is a transmembrane protein pore derived from a β-barrel pore or an α-helix bundle pore. A β-barrel pore comprises a barrel or channel formed from β-strands. Suitable β-barrel pores include, but are not limited to, β-toxins such as α-hemolysin, anthrax toxin, and leukocidin, as well as bacterial outer membrane proteins / porins such as Mycobacterium smegmatis porins (Msp), e.g., MspA, MspB, MspC, or MspD, CsgG, outer membrane porin F (OmpF), outer membrane porin G (OmpG), outer membrane phospholipase A, and Neisseria autotransporter lipoprotein (NalP), and other pores such as lysenin. An α-helix bundle pore comprises a barrel or channel formed from α-helices. Suitable α-helix bundle pores include, but are not limited to, inner membrane proteins and α-outer membrane proteins, e.g., WZA and ClyA toxins.

[0331] In one embodiment, the nanopore is a transmembrane pore derived from or based on Msp, α-hemolysin (α-HL), lysenin, CsgG, ClyA, Sp1, or the hemolytic protein Fragaceatoxin C (FraC).

[0332] In one embodiment, the nanopore is a transmembrane protein pore derived from CsgG, for example CsgG from E. coli strain K-12 substrain MC4100. Such pores are oligomeric and typically comprise 7, 8, 9, or 10 monomers derived from CsgG. The pore may be a homo-oligomeric pore derived from CsgG comprising identical monomers. Alternatively, the pore may be a hetero-oligomeric pore derived from CsgG comprising at least one monomer that is different from the others. Examples of suitable pores derived from CsgG are disclosed in International Application No. 2016 / 034591, International Application No. 2017 / 149316, International Application No. 2017 / 149317, International Application No. 2017 / 149318, and International Application No. 2019 / 002893, which are incorporated herein by reference in their entirety.

[0333] In one embodiment, the nanopore is a transmembrane pore derived from lysenin. Examples of suitable pores derived from lysenin are disclosed in International Application No. WO 2013 / 153359, the entire contents of which are incorporated herein by reference.

[0334] In one embodiment, the nanopore is a transmembrane pore derived from or based on α-hemolysin (α-HL). The wild-type α-hemolysin pore is formed from seven identical monomers or subunits (i.e., it is a heptamer). The α-hemolysin pore can be α-hemolysin-NN or a mutant form thereof. The mutant form preferably contains N residues at positions E111 and K147.

[0335] In one embodiment, the nanopore is a transmembrane protein pore derived from an Msp, such as from MspA. An example of a suitable pore derived from MspA is disclosed in WO 2012 / 107778.

[0336] In one embodiment, the nanopore is a transmembrane pore derived from or based on ClyA. Examples of suitable pores derived from ClyA are disclosed in Soskine et al., Nano Letters 2012 12(9), 4895-4900, International Application No. 2014 / 153625, and International Application No. 2017 / 098322, each of which is incorporated herein by reference.

[0337] In one embodiment, the nanopore is a transmembrane pore derived from Phi29. Examples of suitable pores derived from Phi29 are disclosed in Wendell et al., Nature Nanotech 4, 765-772 (2009), International Application No. 2010 / 062697, International Application No. 2019 / 157365 and International Application No. 2019 / 157424, each of which is incorporated herein by reference.

[0338] In some embodiments, the nanopore is selected from M-ring proteins, perforin-2, PlyAB (pleurotolysin), SpoIIIAG, VirB7, type II secretion system protein D, GspD, InvG, PilQ, pentraxins, and portal proteins, including T4, T7, P23_45, G20c, and Phi29 nanopores.

[0339] In one embodiment, the nanopore is a transmembrane pore derived from or based on a bacterium of the Rhodococcus species, such as Rhodococcus corynebacteroides or Rhodococcus ruber, e.g., PorARr, PorBRr or PorARc. Examples of such pores are described in Piselli et al., Eur Biophys J 51, 309-323 (2022).

[0340] In some embodiments, the nanopore is modified to increase the distance between the polynucleotide handling protein and the constriction region of the nanopore. Methods for doing so are disclosed in International Application No. WO 2021 / 111125.

[0341] Increase in RED As noted above, without being bound by theory, the inventors have found that the length of a polypeptide that can be characterized is typically improved when a nanopore with a longer barrel or channel is used compared to a nanopore with a shorter barrel or channel, and as explained above, this is believed to correlate with or be determined by the "RED" distance, as shown diagrammatically in FIG. 5A(A).

[0342] As mentioned above, in some embodiments, the RED can be increased by modifying the detector used. In particular, if the detector is or includes a nanopore, the nanopore can be modified to increase the RED.

[0343] In some embodiments, the nanopore comprises a constriction. A constriction is typically a narrowing of the channel through the nanopore that can determine or control the signal obtained when a conjugate translocates relative to the nanopore. As used herein, both protein and solid state nanopores typically comprise a "constriction."

[0344] In some embodiments, the nanopore is modified to increase the distance between the polynucleotide handling protein and the constriction region of the nanopore. In some embodiments, when the polynucleotide handling protein is being used to control the movement of a conjugate relative to the nanopore, the nanopore is modified to increase the distance between the polynucleotide handling protein and the constriction region of the nanopore. In some embodiments, the nanopore is modified to increase the distance between the polynucleotide handling protein and the constriction region of the nanopore when the polynucleotide handling protein is in contact with the nanopore. The nanopore is typically modified to increase the distance between the polynucleotide handling protein and the constriction region of the nanopore, as determined when the polynucleotide handling protein is being used to control the movement of a conjugate relative to the nanopore. In some embodiments, when used to control the movement of a conjugate relative to the nanopore, the polynucleotide handling protein is in a "seated position" in contact with the nanopore, e.g., in contact with the cis or trans opening of the nanopore. This is described in more detail herein and shown diagrammatically in FIG. 5A(B).

[0345] In some embodiments, the nanopore is modified to increase the distance between the active site of the polynucleotide handling protein and the constriction region of the nanopore, in such embodiments, the distance can be the distance between the active site of the polynucleotide handling protein and the constriction of the nanopore when the polynucleotide handling protein is being used to control the movement of a conjugate relative to the nanopore and / or when the polynucleotide handling protein is in contact with the nanopore.

[0346] The nanopore may be modified in any suitable manner. Modification of nanopores, such as protein nanopores, is within the knowledge of one of ordinary skill in the art. Modification of solid-state nanopores is routine and may be accomplished by controlling the substrate in which the nanopore is formed (e.g., its thickness) or by controlling the components in which the nanopore is formed.

[0347] For example, the nanopore can be modified to increase the length of the channel through the pore.

[0348] Protein nanopores can be modified by introducing additional amino acids into the pore structure. In some embodiments, protein nanopores are modified by introducing one or more loop regions that extend beyond the natural extent of the nanopore. In embodiments where the nanopore comprises multiple subunits, one or more loop regions can be introduced into one or more subunits of the nanopore. The loop regions can, for example, extend beyond the cis entrance of the nanopore.

[0349] Protein nanopores can be modified to extend the length of the barrel or channel through the pore. For example, a beta-barrel pore can be modified by introducing additional amino acids into the protein sequence of the barrel-forming portion, thereby extending the length of the barrel. Rational design of the relevant positions for such modifications can be performed, for example, by reference to the structural (e.g., X-ray) structure of the protein and / or its monomeric subunits.

[0350] The protein nanopore can be modified by the fusion of one or more additional domains to elevate the "seating position" of the polynucleotide handling protein relative to the nanopore.

[0351] In some embodiments, it is possible to modify a protein nanopore by fusing it to another protein nanopore. In this way, a chain of nanopores can be created with a single channel therethrough, extending the distance between the constriction of the channel and the polynucleotide handling protein. In such cases, the multiple nanopores can be the same or different.

[0352] In some embodiments, the RED can be increased by modifying the polynucleotide handling protein used.

[0353] In some embodiments, when a polynucleotide handling protein is used to control the movement of a conjugate relative to a detector (e.g., a nanopore), the polynucleotide handling protein is modified to increase the distance between the polynucleotide handling protein and, for example, the nanopore.

[0354] The polynucleotide handling protein may be modified in any suitable manner. Modification of proteins such as polynucleotide handling proteins is within the knowledge of one of ordinary skill in the art.

[0355] The polynucleotide handling protein may be modified by introducing additional amino acids into the protein structure. In some embodiments, the polynucleotide handling protein is modified by introducing one or more loop regions that extend beyond the natural extent of the protein. In embodiments where the polynucleotide handling protein comprises multiple subunits, one or more loop regions may be introduced into one or more subunits of the polynucleotide handling protein.

[0356] The polynucleotide handling protein may be modified by the fusion of one or more additional domains to displace the nanopore when the polynucleotide handling protein is in a "seated position" relative to the nanopore.

[0357] In some embodiments, the RED can be increased by using one or more displacer units.

[0358] The use of a displacer unit is shown diagrammatically in Figure 5A(C). Thus, in some embodiments, the methods provided herein include providing a displacer unit. In some embodiments, the displacer unit is for separating the polypeptide handling protein from the nanopore, thereby increasing the distance between the polynucleotide handling protein and the nanopore.

[0359] In such embodiments, any suitable displacer unit can be used, for example, the displacer unit can be provided as a protein.

[0360] Any suitable protein can be used as a displacer unit. Exemplary proteins include proteins that adopt a ring-shaped conformation, e.g., a multimer, and thus can be easily positioned at the entrance of the nanopore. Many suitable ring-shaped proteins are known in the art, including the nanopores described herein, helicases (e.g., T7 helicase) and variants thereof. It is not necessary for the displacer to have activity of its own. In some embodiments, the displacer unit does not provide any significant discrimination of either the polynucleotide or peptide in the conjugate.

[0361] In some embodiments, the displacer unit may comprise one or more polynucleotide handling proteins or inactive mutants thereof. This is shown diagrammatically in FIG. 5B. As shown, a polynucleotide handling protein (E1) is used to control the movement of the conjugate relative to the nanopore. Polynucleotide handling proteins E2...E n are initially in contact with the polypeptide portion of the conjugate and therefore do not control the translocation of the conjugate relative to the nanopore, however, they displace the polynucleotide handling protein E1 from the nanopore, increasing the RED.

[0362] In such embodiments, the polynucleotide handling protein used as the displacer unit may be the same or different from the polynucleotide handling protein used to control the movement of the conjugate, hi some embodiments, the polynucleotide handling protein used as the displacer unit is formed from an inactive mutant of the same polynucleotide handling protein used to control the movement of the conjugate.

[0363] One or more displacer units, such as one or more displacer units described herein, can be attached (e.g., covalently or non-covalently) to the nanopore. Alternatively, one or more displacer units can be associated with the nanopore, for example, by controlling their position relative to the nanopore using a polynucleotide.

[0364] In other embodiments, the polynucleotide handling protein is modified to increase the distance from the active site of the polynucleotide handling protein to the nanopore. The polynucleotide handling protein is typically modified to increase the distance between the active site of the polynucleotide handling protein and the nanopore, as determined when the polynucleotide handling protein is used to control the movement of a conjugate relative to the nanopore. This is described in more detail herein.

[0365] tag In some embodiments of the methods provided herein, a tag on the detector (e.g., a nanopore) can be used, for example, to facilitate capture of the conjugate by the detector (e.g., a nanopore).

[0366] The interaction between a tag on a detector, such as a nanopore, and a binding site on a polynucleotide (e.g., a binding site present on the polynucleotide portion of the conjugate, or on an adapter attached to the conjugate, where the binding site can be provided by an anchor or leader sequence of the adapter, or by a capture sequence in the double-stranded stem of the adapter) may be reversible. For example, a polynucleotide can bind to a tag on a detector, such as a nanopore, e.g., via its adapter, and be released at some point during, e.g., characterization of the polynucleotide by the detector / nanopore and / or processing by a polynucleotide handling protein. Strong non-covalent bonds (e.g., biotin / avidin) are still reversible and may be useful in some embodiments of the methods described herein. For example, a pore tag and polynucleotide adapter pair may be designed to provide sufficient interaction between the complement of a double-stranded polynucleotide (or the portion of the adapter attached to its complement) and the nanopore such that the complement is held in the vicinity of the detector / nanopore (without diffusing away from the detector / nanopore) but can be released from the nanopore when processed.

[0367] The tag and polynucleotide adapters can be configured such that the binding strength or affinity of a binding site on the polynucleotide (e.g., a binding site provided by an anchor or leader sequence of the adapter or by a capture sequence within the double-stranded stem of the adapter) to a tag on a detector, such as a nanopore, is sufficient to maintain coupling between the detector / nanopore and the polynucleotide until an applied force is applied to release the bound polynucleotide from the nanopore.

[0368] In some embodiments, the tag or tether is uncharged, which can ensure that the tag or tether is not drawn into a detector, e.g., a nanopore, under the influence of a potential difference, if any.

[0369] One or more molecules that attract or bind to the conjugate, polynucleotide or adapter may be linked to a detector such as a nanopore. Any molecule that hybridizes to the conjugate, adapter and / or polynucleotide may be used. The molecule attached to the detector / pore may be selected from PNA tags, PEG linkers, short oligonucleotides, positively charged amino acids and aptamers. Such molecules are known in the art to be bound to pores. For example, pores with short oligonucleotides attached are disclosed in Howarka et al (2001) Nature Biotech. 19:636-639 and International Application No. 2010 / 086620, and pores with PEG attached within the lumen of the pore are disclosed in Howarka et al (2000) J. Am. Chem. Soc. 122(11):2411-2416.

[0370] Short oligonucleotides attached to a detector, such as a nanopore, that contain a sequence complementary to a sequence in the conjugate (e.g., in a leader sequence in an adapter or another single-stranded sequence) may be used to enhance capture of the conjugate in the methods described herein.

[0371] In some embodiments, the tag or tether may comprise or be an oligonucleotide (e.g., DNA, RNA, LNA, BNA, PNA, or morpholino). The oligonucleotide may have a length of about 10-30 nucleotides or a length of about 10-20 nucleotides. In some embodiments, the oligonucleotide may have at least one end (e.g., the 3'- or 5'-end) modified for conjugation to other modifications or to solid substrate surfaces, including, for example, beads. The end modifier may add a reactive functional group that can be used for conjugation. Examples of functional groups that can be added include, but are not limited to, amino, carboxyl, thiol, maleimide, aminooxy, and any combination thereof. The functional group may be combined with spacers of different lengths (e.g., C3, C9, C12, spacers 9 and 18) to add physical distance of the functional group from the end of the oligonucleotide sequence.

[0372] Examples of 3' and / or 5' end modifications of oligonucleotides include, but are not limited to, 3' affinity tags and functional groups for chemical conjugation (including, for example, 3'-biotin, 3'-primary amine, 3'-disulfide amide, 3'-pyridyldithio, and any combination thereof); 5' end modifications (including, for example, 5'-primary ammine, and / or 5'-dabsyl); modifications for click chemistry (including, for example, 3'-azide, 3'-alkyne, 5'-azide, 5'-alkyne), and any combination thereof.

[0373] In some embodiments, the tag or tether may further comprise a polymer linker, for example, to facilitate coupling to the detector / nanopore. Exemplary polymer linkers include, but are not limited to, polyethylene glycol (PEG). The polymer linker may have a molecular weight of about 500 Da to about 10 kDa, inclusive, or about 1 kDa to about 5 kDa, inclusive. The polymer linker (e.g., PEG) may be functionalized with different functional groups, including, for example, but not limited to, maleimide, NHS ester, dibenzocyclooctyne (DBCO), azide, biotin, amine, alkyne, aldehyde, and any combination thereof.

[0374] Other examples of tags or tethers include, but are not limited to, His tags, biotin or streptavidin, antibodies that bind to the analyte, aptamers that bind to the analyte, analyte binding domains such as DNA binding domains (including, for example, peptide zippers such as leucine zippers, single stranded DNA binding proteins (SSBs)), and any combination thereof.

[0375] The tag or tether may be attached to the exterior surface of the nanopore, e.g., on the cis side of the membrane, using any method known in the art. For example, one or more tags or tethers can be attached to the nanopore via one or more cysteines (cysteine ​​bonds), one or more primary amines such as lysines, one or more unnatural amino acids, one or more histidines (His tags), one or more biotins or streptavidins, one or more antibody-based tags, one or more enzymatic modifications of epitopes (e.g., including acetyltransferases), and any combination thereof. Suitable methods for making such modifications are well known in the art. Suitable unnatural amino acids include, but are not limited to, 4-azido-L-phenylalanine (Faz), and any one of the amino acids numbered 1-71 in Figure 1 of Liu CC and Schultz PG, Annu. Rev. Biochem., 2010, 79, 413-444.

[0376] In some embodiments where one or more tags or tethers are attached to the nanopore via cysteine ​​bond(s), one or more cysteines can be introduced by substitution into one or more of the monomers forming the nanopore. In some embodiments, the nanopore can be chemically modified by attachment of: (i) 4-phenylazomaleinanyl, 1.N-(2-hydroxyethyl)maleimide, N-cyclohexylmaleimide, 1.3-maleimidopropionic acid, 1.1-4-aminophenyl-1H-pyrrole, 2,5,dione, 1.1-4-hydroxyphenyl-1H-pyrrole, 2,5,dione, N-ethylmaleimide, N-methoxycarbonylmaleimide, N-methylmaleimide, N-methyl-2-phenylpropionyl ... Imide, N-tert-butylmaleimide, N-(2-aminoethyl)maleimide, 3-maleimide-proxyl, N-(4-chlorophenyl)maleimide, 1-[4-(dimethylamino)-3,5-dinitrophenyl]-1H-pyrrole-2,5-dione, N-[4-(2-benzimidazolyl)phenyl]maleimide, N-[4-(2-benzoxazolyl)phenyl]maleimide, N-(1-naphthyl)-maleimide, N-(2, 4-xylyl)maleimide, N-(2,4-difluorophenyl)maleimide, N-(3-chloro-para-tolyl)-maleimide, 1-(2-amino-ethyl)-pyrrole-2,5-dione hydrochloride, 1-cyclopentyl-3-methyl-2,5-dihydro-1H-pyrrole-2,5-dione, 1-(3-aminopropyl)-2,5-dihydro-1H-pyrrole-2,5-dione hydrochloride, 3-methyl-1-[2-oxo- 2-(piperazin-1-yl)ethyl]-2,5-dihydro-1H-pyrrole-2,5-dione hydrochloride, 1-benzyl-2,5-dihydro-1H-pyrrole-2,5-dione, 3-methyl-1-(3,3,3-trifluoropropyl)-2,5-dihydro-1H-pyrrole-2,5-dione, 1-[4-(methylamino)cyclohexyl]-2,5-dihydro-1H-pyrrole-2,5-dione trifluoroacetate, SMILES O=C1C=CC(=O)N1CC=2C=CN=CC2, SMILES O=C1C=CC(=O)N1CN2CCNCC2, 1-benzyl-3-methyl-2,5-dihydro-1H-pyrrole-2,(ii) maleimides including diabromomaleimides such as 5-dione, 1-(2-fluorophenyl)-3-methyl-2,5-dihydro-1H-pyrrole-2,5-dione, N-(4-phenoxyphenyl)maleimide, N-(4-nitrophenyl)maleimide, (ii) 3-(2-iodoacetamido)-proxyl, N-(cyclopropylmethyl)-2-iodoacetamide, 2-iodo-N-(2-phenylethyl)acetamide, 2-iodo-N-(2,2,2-trifluoroethyl)acetamide, N-(4-acetylphenyl)-2-iodoacetamide, (iii) iodoacetamides such as N-(4-(aminosulfonyl)phenyl)-2-iodoacetamide, N-(1,3-benzothiazol-2-yl)-2-iodoacetamide, N-(2,6-(diethylphenyl)-2-iodoacetamide, and N-(2-benzoyl-4-chlorophenyl)-2-iodoacetamide; (iv) N-(4-(acetylamino)phenyl)-2-bromoacetamide, N-(2-acetylphenyl)-2-bromoacetamide, 2-bromo-n-(2-cyanophenyl)acetamide, 2-bromo-N-(3-((trifluorophenyl)acetamide, etc.); N-(2-bromophenyl)acetamide, N-(2-benzoylphenyl)-2-bromoacetamide, 2-bromo-N-(4-fluorophenyl)-3-methylbutanamide, N-benzyl-2-bromo-N-phenylpropionamide, N-(2-bromo-butyl)-4-chloro-benzenesulfonamide, 2-bromo-N-methyl-N-phenylacetamide, 2-bromo-N-phenethyl-acetamide, 2-adamantan-1-yl-2-bromo-N-cyclohexyl-acetamide, 2-bromo-N-(2-methylphenyl)butanamide, mono (iv) bromoacetamides such as bromoacetanilide, (iv) disulfides such as aldrithiol-2, aldrithiol-4, isopropyl disulfide, 1-(isobutyldisulfanyl)-2-methylpropane, dibenzyl disulfide, 4-aminophenyl disulfide, 3-(2-pyridyldithio)propionic acid, 3-(2-pyridyldithio)propionic acid hydrazide, 3-(2-pyridyldithio)propionic acid N-succinimidyl ester, am6amPDP1-βCD, and (v) 4-phenylthiazole-2-thiol, perpaldo, 5,Thiols such as 6,7,8-tetrahydro-quinazoline-2-thiol.

[0377] In some embodiments, the tag or tether may be attached to the nanopore directly or via one or more linkers. The tag or tether may be attached to the nanopore using a hybrid linker as described in WO 2010 / 086602. Alternatively, a peptide linker may be used. A peptide linker is an amino acid sequence. The length, flexibility, and hydrophilicity of the peptide linker are typically designed so as not to interfere with the function of the monomer and the pore. Preferred flexible peptide linkers are stretches of 2-20, e.g., 4, 6, 8, 10, or 16 serine and / or glycine amino acids. More preferred flexible linkers include (SG)1, (SG)2, (SG)3, (SG)4, (SG)5, and (SG)8, where S is serine and G is glycine. Preferred rigid linkers are stretches of 2-30, e.g., 4, 6, 8, 16, or 24 proline amino acids. A more preferred rigid linker is one in which P is proline, (P) 12 Includes.

[0378] film Typically, in the disclosed methods, the detector is typically present in the membrane. Any suitable membrane can be used in the system. For example, the detector can be or include a nanopore in the membrane.

[0379] The membrane is preferably an amphiphilic layer. An amphiphilic layer is a layer formed from amphiphilic molecules, such as phospholipids, that have both hydrophilic and lipophilic properties. The amphiphilic molecules can be synthetic or naturally occurring. Non-naturally occurring amphiphiles and amphiphiles that form monolayers are known in the art, including, for example, block copolymers (Gonzalez-Perez et al., Langmuir, 2009, 25, 10447-10450). Block copolymers are polymeric materials in which two or more monomeric subunits are polymerized together to create a single polymer chain. Block copolymers typically have properties contributed by each monomeric subunit. However, block copolymers can have unique properties that polymers formed from individual subunits do not have. Block copolymers can be engineered such that one of the monomeric subunits is hydrophobic (i.e., lipophilic), while the other subunit(s) are hydrophilic in aqueous media. In this case, the block copolymer may have amphiphilic properties and may form structures that mimic biological membranes. The block copolymer may be diblock (consisting of two monomer subunits), but may be constructed from three or more monomer subunits to form more complex arrangements that behave as amphiphiles. The copolymer may be a triblock, tetrablock, or pentablock copolymer. The membrane is preferably a triblock copolymer membrane.

[0380] Archaeal bipolar tetraether lipids are naturally occurring lipids that are constructed such that the lipids form monolayer membranes. These lipids are commonly found in extremophilic, thermophilic, halophilic, and acidophilic bacteria that survive in harsh biological environments. Their stability is believed to derive from the fusogenic nature of the final bilayer. It is straightforward to construct block copolymers that mimic these biological entities by creating triblock polymers with the general motif hydrophilic-hydrophobic-hydrophilic. This material forms monomeric membranes that behave similarly to lipid bilayers and can encompass a wide range of phase behaviors, from vesicles to lamellar membranes. Membranes formed from these triblock copolymers have several advantages over biological lipid membranes. Because the triblock copolymers are synthetic, their exact structure can be carefully controlled to provide the correct chain length and properties required to form membranes and interact with pores and other proteins.

[0381] Block copolymers may also be constructed from subunits that are not classified as lipid submaterials, for example, hydrophobic polymers may be made from siloxanes or other non-hydrocarbon monomers. The hydrophilic subsections of the block copolymers may also have low protein binding properties, allowing for the creation of membranes that are highly resistant when exposed to live biological samples. The head group units may also be derived from non-classified lipid head groups.

[0382] Triblock copolymer membranes also have increased mechanical and environmental stability compared to biological lipid membranes, e.g., much higher operating temperature or pH ranges. The synthetic nature of block copolymers provides a platform for customizing polymer-based membranes for a wide range of applications.

[0383] In some embodiments, the membrane is one of the membranes disclosed in WO 2014 / 064443 or WO 2014 / 064444.

[0384] The amphiphilic molecules may be chemically modified or functionalized to facilitate coupling of polynucleotides. The amphiphilic layer may be a monolayer or a bilayer. The amphiphilic layer is typically planar. The amphiphilic layer may be curved. The amphiphilic layer may be supported.

[0385] Amphiphilic membranes typically have a molecular weight of approximately 10 -8 cm s -1 They are naturally mobile, essentially acting as a two-dimensional fluid with lipid diffusion rates of 0.1 - 0.2 nm, which means that the pore and the coupled polynucleotides can typically move within the amphiphilic membrane.

[0386] The membrane may be a lipid bilayer. Lipid bilayers are models of cell membranes and serve as an excellent platform for a wide range of experimental studies. For example, lipid bilayers can be used for in vitro investigation of membrane proteins by single channel recording. Alternatively, lipid bilayers can be used as biosensors to detect the presence of a wide range of substances. The lipid bilayer may be any lipid bilayer. Suitable lipid bilayers include, but are not limited to, planar lipid bilayers, supported bilayers, or liposomes. The lipid bilayer is preferably a planar lipid bilayer. Suitable lipid bilayers are disclosed in International Application No. 2008 / 102121, International Application No. 2009 / 077734, and International Application No. 2006 / 100484.

[0387] Methods for forming lipid bilayers are known in the art. Lipid bilayers are generally formed by the method of Montal and Mueller (Proc. Natl. Acad. Sci. USA., 1972; 69: 3561-3566), in which a lipid monolayer is supported on the aqueous solution / air interface past both sides of a hole perpendicular to the interface. Lipid is usually added to the surface of an aqueous electrolyte solution by first dissolving it in an organic solvent, and then evaporating a small amount of solvent on the interface of the aqueous solution on both sides of the opening. As the organic solvent evaporates, the solution / air interfaces on both sides of the opening physically move up and down across the opening until a bilayer is formed. Planar lipid bilayers can be formed in a membrane across an opening or in a recess across an opening.

[0388] The Montal & Mueller method is popular because it is a cost-effective and relatively simple method for forming good quality lipid bilayers suitable for protein pore insertion. Other common methods of bilayer formation include tip-dipping, painting bilayers, and patch clamping of liposome bilayers.

[0389] Tip-dipping bilayer formation involves contacting an aperture (e.g., a pipette tip) onto the surface of a test solution carrying a monolayer of lipids. Again, the lipid monolayer is generated at the solution / air interface by first evaporating a small amount of lipid dissolved in an organic solvent at the solution surface. The bilayer is then formed by the Langmuir-Schaefer method, which requires mechanical automation to move the aperture relative to the solution surface.

[0390] For bilayer coating, a small amount of lipid dissolved in an organic solvent is applied directly to an aperture immersed in the test aqueous solution. The lipid solution is spread thinly across the aperture using a paintbrush or equivalent. Thinning the solvent leads to the formation of a lipid bilayer. However, it is difficult to completely remove the solvent from the bilayer, and as a result, bilayers formed by this method are less stable and more prone to noise during electrochemical measurements.

[0391] Patch clamping is commonly used in the study of biological cell membranes. The cell membrane is clamped to the end of a pipette by suction, so that a patch of membrane is attached across the aperture. The method has been adapted to produce lipid bilayers by clamping and then rupturing liposomes to seal the lipid bilayer across the aperture of the pipette. The method requires the production of stable large unilamellar liposomes and a small aperture in the material with a glass surface.

[0392] Liposomes can be formed by sonication, extrusion, or the Mozafari method (Colas et al. (2007) Micron 38:841-847).

[0393] In some embodiments, the lipid bilayer is formed as described in WO 2009 / 077734. Advantageously in this method, the lipid bilayer is formed from dry lipids. In the most preferred embodiment, the lipid bilayer is formed across the aperture as described in WO 2009 / 077734.

[0394] A lipid bilayer is formed from two opposing layers of lipids. These two lipid layers are arranged so that their hydrophobic tail groups face each other to form a hydrophobic interior. The hydrophilic head groups of the lipids face outward toward the aqueous environment on both sides of the bilayer. Bilayers can exist in several lipid phases, including but not limited to liquid disordered phases (fluid lamellar), liquid ordered phases, solid ordered phases (lamellar gel phase, interdigitated gel phase), and planar bilayer crystals (lamellar subgel phase, lamellar crystal phase).

[0395] Any lipid composition that forms a lipid bilayer can be used. The lipid composition is selected to form a lipid bilayer with the required properties, such as surface charge, ability to support membrane proteins, packing density, or mechanical properties. The lipid composition can include one or more different lipids. For example, the lipid composition can include up to 100 lipids. The lipid composition preferably includes 1 to 10 lipids. The lipid composition can include naturally occurring lipids and / or artificial lipids.

[0396] A lipid typically comprises a head group, an interface moiety, and two hydrophobic tail groups, which may be the same or different. Suitable head groups include, but are not limited to, neutral head groups, such as diacylglyceride (DG) and ceramide (CM), zwitterionic head groups, such as phosphatidylcholine (PC), phosphatidylethanolamine (PE), and sphingomyelin (SM), negatively charged head groups, such as phosphatidylglycerol (PG), phosphatidylserine (PS), phosphatidylinositol (PI), phosphatidic acid (PA), and cardiolipin (CA), and positively charged head groups, such as trimethylammonium-propane (TAP). Suitable interface moieties include, but are not limited to, naturally occurring interface moieties, such as glycerol-based moieties or ceramide-based moieties. Suitable hydrophobic tail groups include, but are not limited to, saturated hydrocarbon chains such as lauric acid (n-Dodecanolic acid), myristic acid (n-Tetradecononic acid), palmitic acid (n-Hexadecanoic acid), stearic acid (n-Octadecanoic acid), and arachidic acid (n-Eicosanoic acid), unsaturated hydrocarbon chains such as oleic acid (cis-9-Octadecanoic acid), and branched hydrocarbon chains such as phytanoyl. The length of the chain and the position and number of double bonds in the unsaturated hydrocarbon chains can vary. The length of the chain and the position and number of branches, such as methyl groups in the branched hydrocarbon chains, can vary. The hydrophobic tail group can be linked to the interface moiety as an ether or ester. The lipid can be a mycolic acid.

[0397] Lipids can also be chemically modified. The head group or tail group of lipids can be chemically modified. Suitable lipids with chemically modified head groups include, but are not limited to, PEG-modified lipids, such as 1,2-diacyl-sn-glycero-3-phosphoethanolamine-N-[methoxy(polyethylene glycol)-2000], functionalized PEG lipids, such as 1,2-distearoyl-sn-glycero-3-phosphoethanolamine-N-[biotinyl(polyethylene glycol)2000], and lipids modified for conjugation, such as 1,2-dioleoyl-sn-glycero-3-phosphoethanolamine-N-(succinyl) and 1,2-dipalmitoyl-sn-glycero-3-phosphoethanolamine-N-(biotinyl). Suitable lipids in which the tail group is chemically modified include, but are not limited to, polymerizable lipids such as 1,2-bis(10,12-tricosadiynoyl)-sn-glycero-3-phosphocholine, fluorinated lipids such as 1-palmitoyl-2-(16-fluoropalmitoyl)-sn-glycero-3-phosphocholine, deuterated lipids such as 1,2-dipalmitoyl-D62-sn-glycero-3-phosphocholine, and ether-linked lipids such as 1,2-di-O-phytanyl-sn-glycero-3-phosphocholine. Lipids may be chemically modified or functionalized to facilitate coupling of polynucleotides.

[0398] Amphiphilic layer, for example, lipid composition, typically contains one or more additives that will affect the properties of the layer.Suitable additives include, but are not limited to, fatty acid, for example, palmitic acid, myristic acid, and oleic acid, fatty alcohol, for example, palmitic alcohol, myristic alcohol, and oleic alcohol, sterol, for example, cholesterol, ergosterol, lanosterol, sitosterol, and stigmasterol, lysophospholipid, for example, 1-acyl-2-hydroxy-sn-glycero-3-phosphocholine, and ceramide.

[0399] In another embodiment, the membrane comprises a solid-state layer. The solid-state layer can be formed from both organic and inorganic materials, including but not limited to microelectronic materials, insulating materials such as Si3N4, A12O3, and SiO, organic and inorganic polymers such as polyamides, plastics such as Teflon®, or elastomers such as two-component addition cure silicone rubber, and glass. The solid-state layer can be formed from graphene. A suitable graphene layer is disclosed in WO 2009 / 035647. When the membrane comprises a solid-state layer, the pores are typically present in the amphiphilic membrane or layer contained within the solid-state layer, for example, within holes, wells, gaps, channels, grooves, or slits within the solid-state layer. A person skilled in the art can prepare suitable solid-state / amphiphilic hybrid systems. Suitable systems are disclosed in WO 2009 / 020682 and WO 2012 / 005857. Any of the amphiphilic membranes or layers discussed above can be used.

[0400] The methods disclosed herein are typically carried out using (i) an artificial amphiphilic layer comprising a pore, (ii) an isolated naturally occurring lipid bilayer comprising a pore, or (iii) a cell with a pore inserted therein. The methods are typically carried out using an artificial amphiphilic layer, such as an artificial triblock copolymer layer. The layer may contain other transmembrane and / or intramembrane proteins, as well as other molecules, in addition to the pore. Suitable equipment and conditions are discussed below. The disclosed methods are typically carried out in vitro.

[0401] conditions As explained above, the disclosed methods include characterizing a polypeptide as a conjugate in which the polypeptide is included moves relative to a detector, such as a nanopore.

[0402] In some embodiments, the characterization method may be performed using any device suitable for investigating a membrane / pore system in which a pore is inserted into a membrane. In some embodiments, the characterization method may be performed using any device suitable for transmembrane pore sensing. For example, the device may comprise a chamber containing an aqueous solution and a barrier separating the chamber into two sections. The barrier may have an opening in which a membrane containing a transmembrane pore is formed. The transmembrane pore may be as described herein.

[0403] The characterisation method may be carried out using the apparatus described in WO 2008 / 102120, WO 2010 / 122293 or WO 00 / 28312.

[0404] The characterization method may include measuring the ionic current flow through the pore, typically by measuring the electrical current. Alternatively, the ionic flow through the pore may be measured optically, as disclosed by Heron et al: J.Am.Chem.Soc.9 Vol.131, No.5,2009. Thus, the device may also comprise an electrical circuit capable of applying an electrical potential and measuring an electrical signal across the membrane and the pore. The characterization method may be carried out using a patch clamp or a voltage clamp. The characterization method preferably involves the use of a voltage clamp.

[0405] The characterization method may be performed on silicon-based well arrays, where each array comprises 128, 256, 512, 1024, 2000, 3000, 4000, 6000, 10000, 12000, 15000, or more wells.

[0406] The characterization method may include measuring the current through a detector such as a pore. The method is typically carried out with a voltage applied across the membrane and pore. The voltage used is typically between +2V and -2V, typically between -400mV and +400mV. The voltage used is preferably in a range having a lower limit selected from -400mV, -300mV, -200mV, -150mV, -100mV, -50mV, -20mV, and 0mV, and an upper limit independently selected from +10mV, +20mV, +50mV, +100mV, +150mV, +200mV, +300mV, and +400mV. The voltage used is more preferably in the range of 100mV to 240mV, most preferably in the range of 120mV to 220mV. By using an increased applied potential, it is possible to increase the discrimination between different nucleotides or peptides by the pore.

[0407] The characterization method is typically carried out in the presence of any charge carrier, such as a metal salt, e.g., an alkali metal salt, a halide salt, e.g., a chloride salt, e.g., an alkali metal chloride salt. The charge carrier may include an ionic liquid or an organic salt, e.g., tetramethylammonium chloride, trimethylphenylammonium chloride, phenyltrimethylammonium chloride, or 1-ethyl-3-methylimidazolium chloride. In the exemplary device discussed above, the salt is present in an aqueous solution in the chamber. Potassium chloride (KCl), sodium chloride (NaCl), or cesium chloride (CsCl) are typically used. KCl is preferred. The salt may be an alkaline earth metal salt, such as calcium chloride (CaCl2). The salt concentration may be saturated. The salt concentration may be 3M or less, typically 0.1-2.5M, 0.3-1.9M, 0.5-1.8M, 0.7-1.7M, 0.9-1.6M, or 1M-1.4M. The salt concentration is preferably between 150 mM and 1 M. The characterization method is preferably carried out using a salt concentration of at least 0.3 M, such as at least 0.4 M, at least 0.5 M, at least 0.6 M, at least 0.8 M, at least 1.0 M, at least 1.5 M, at least 2.0 M, at least 2.5 M, or at least 3.0 M. High salt concentrations provide a high signal to noise ratio, allowing identification of currents indicative of binding / no binding against the background of normal current fluctuations.

[0408] The characterization method is typically carried out in the presence of a buffer. In the exemplary device discussed above, the buffer is present in the aqueous solution in the chamber. Any suitable buffer may be used. Typically, the buffer is HEPES. Another suitable buffer is a Tris-HCl buffer. The method is typically carried out at a pH of 4.0-12.0, 4.5-10.0, 5.0-9.0, 5.5-8.8, 6.0-8.7, or 7.0-8.8, or 7.5-8.5. The pH used is preferably about 7.5.

[0409] The characterization method may be carried out at 0° C. to 100° C., 15° C. to 95° C., 16° C. to 90° C., 17° C. to 85° C., 18° C. to 80° C., 19° C. to 70° C., or 20° C. to 60° C. The characterization method is typically carried out at room temperature. The characterization method is optionally carried out at a temperature that supports enzyme function, for example, about 37° C.

[0410] Conjugates Also provided herein are conjugates as described herein. In some embodiments, conjugates are provided herein that include: i) a first end of the polypeptide is conjugated to a first end of the polynucleotide; ii) the second end of the polypeptide is attached to a polymer leader, optionally via a linker, which optionally comprises a polynucleotide; iii) if the linker comprises a polynucleotide, the polynucleotide or linker is bound to a polynucleotide binding site of the polynucleotide handling protein, and the polynucleotide binding site is modified with a closing moiety, thereby topologically closing the polynucleotide binding site of the polynucleotide handling protein around the polynucleotide; iv) The polynucleotide comprises a blocking moiety to prevent disassociation of the polynucleotide handling protein from the conjugate.

[0411] In some embodiments, the polypeptide, polynucleotide, polymer leader, polynucleotide handling protein, blocking moiety, and linker, if present, are as described herein.

[0412] system Also, the system - a nanopore; - a conjugate as described herein.

[0413] In some embodiments, a system comprising: - a nanopore; - a conjugate, comprising: i) a first end of the polypeptide is conjugated to a first end of the polynucleotide; ii) the second end of the polypeptide is attached to a polymer leader, optionally via a linker, which optionally comprises a polynucleotide; iii) if the linker comprises a polynucleotide, the polynucleotide or linker is bound to a polynucleotide binding site of the polynucleotide handling protein, and the polynucleotide binding site is modified with a closing moiety, thereby topologically closing the polynucleotide binding site of the polynucleotide handling protein around the polynucleotide; iv) The polynucleotide comprises a blocking moiety to prevent disassociation of the polynucleotide handling protein from the conjugate.

[0414] In some embodiments, the nanopore, polypeptide, polynucleotide, polymer leader, polynucleotide handling protein, blocking moiety, and linker, if present, are as described herein.

[0415] kit Also, a kit comprising: a polynucleotide, wherein a first end of the polynucleotide comprises a reactive functional group for conjugating to a first end of a target polypeptide and a second end of the polynucleotide comprises a blocking moiety; a polynucleotide handling protein, wherein the polynucleotide binding site of the polynucleotide handling protein is modified with a closing moiety for topologically closing the polynucleotide binding site of the polynucleotide handling protein around the polynucleotide; and an adaptor comprising a polymeric leader and a reactive functional group for attachment to a second end of the target polypeptide; A kit is provided comprising:

[0416] In some embodiments, the polynucleotide, the reactive functional group, the blocking moiety, the polynucleotide handling protein, the adaptor comprising a polymeric leader, and the nanopore are as described herein.

[0417] In some embodiments, the kit further comprises one or more displacer units for increasing the distance between the nanopore and the active site of the polynucleotide handling protein.

[0418] The systems and kits disclosed herein may be configured for use with an algorithm, also provided herein, adapted to run on a computer system. The algorithm may be adapted to detect information characteristic of a polypeptide (e.g., characteristic of the sequence of the polypeptide and / or whether the polypeptide is modified) and selectively process a signal obtained when a conjugate comprising a polypeptide conjugated to a polynucleotide moves relative to a nanopore. Systems are also provided that include computing means configured to detect information characteristic of a polypeptide (e.g., characteristic of the sequence of the polypeptide and / or whether the polypeptide is modified) and selectively process a signal obtained when a conjugate comprising a polypeptide conjugated to a polynucleotide moves relative to a nanopore. In some embodiments, the system includes receiving means for receiving data from the detection of the polypeptide, processing means for processing a signal obtained when the conjugate moves relative to the nanopore, and output means for outputting the characterization information thus obtained.

[0419] Although specific embodiments, specific configurations, and materials and / or molecules are discussed herein for the method according to the invention, it should be understood that various changes or modifications in form and details may be made without departing from the scope and spirit of the invention. The foregoing embodiments and the following examples are provided for illustrative purposes only and should not be considered as limiting the application. The present application is limited only by the claims. EXAMPLES

[0420] Example 1 This example shows the controlled translocation of a conjugate containing two pieces of polynucleotide; a dsDNA Y-adapter (DNA1) and a polypeptide flanked by a dsDNA tail (DNA2). A polynucleotide handling protein on the cis side of the nanopore controls the movement of the conjugate by first unwinding DNA1, translocating it 5'-3' onto the ssDNA, then sliding across the polypeptide section and finally unwinding the DNA2 segment. As this construct moves from the cis to the trans side of the nanopore and passes through the RED, the polypeptide section can be visualized in a current versus time plot, allowing for characterization.

[0421] The Y-adapters were prepared by annealing DNA oligonucleotides (SEQ ID NO:9, SEQ ID NO:10, SEQ ID NO:11). A DNA motor (Dda helicase) was loaded onto the adapters and closed as described in International Application No. 2014 / 013260. The subsequent material was HPLC purified. The Y-adapters contain 30 C3 leader sections to facilitate capture by the nanopore and a side arm for tethering to the membrane. The DNA tails were made by annealing DNA oligonucleotides (SEQ ID NO:12, SEQ ID NO:14).

[0422] In this example, model polypeptide analytes (SEQ ID NOs: 15, 16, 17) were obtained containing presynthesized azide moieties at the N-terminus and immediately after the C-terminus using an ethyldiamine spacer along the peptide backbone. Each analyte was then conjugated to a Y-adapter and DNA tail via a copper-free click chemistry reaction between the azide and a BCN (bicyclo[6.1.0]nonyne) moiety. Samples were purified using Agencourt AMPure XP (Beckman Coulter) beads and washed twice with LNB from the Oxford Nanopore Technologies sequencing kit (SQK-LSK109). Conjugated substrates were eluted in 10 mM Tris-Cl, 50 mM NaCl (pH 8.0).

[0423] Electrical measurements were taken using a MinION Mk1b from Oxford Nanopore Technologies and a custom MinION flow cell with an MspA nanopore. The flow cell was flushed with a tether mix containing 50 nM DNA tether and sequencing buffer lacking ATP. 800 μL of the tether mix was first added for 5 min, and then an additional 200 μL of the mix was flowed through the system with the SpotON port open. DNA-peptide constructs were prepared at a concentration of 0.5 nM in sequencing buffer lacking ATP and LB from the Oxford Nanopore Technologies sequencing kit (SQK-LSK109), resulting in the "sequencing mix". 75 μL of the sequencing mix was added to the MinION flow cell via the SpotON flow cell port. The mixture was incubated on the flow cell for 5–10 min to allow tethering of the constructs and their subsequent capture by the nanopore. In the absence of ATP, the DNA motor remains stalled in the spacer region of the Y-adaptor and the conjugate is captured into the nanopore but does not translocate. After incubation, 200 μL of sequencing buffer containing ATP is added, and in the presence of ATP, the captured DNA-peptide conjugate is translocated across the nanopore by the helicase, resulting in reproducible current footprints.

[0424] Standard sequencing scripts were run at 180 mV for 30 min to 1 h, with static flips every minute to remove extended nanopore blocks. Raw data were collected into bulk FAST5 files using MinKNOW software (Oxford Nanopore Technologies).

[0425] An example of a current versus time trace for one of the model peptides (SEQ ID NO: 15) conjugated to a DNA Y-adaptor and tail can be seen in Figure 6. The Y-adaptor section and dsDNA tail can be separated from the peptide portion of the "snake" (trace) allowing characterization of the peptide.

[0426] Multiple translocation events were observed per second, achieving high throughput. Example current versus time traces showing multiple capture and translocation events are shown in Figure 7 for the same construct used in Figure 6 (i.e., containing the peptide region of SEQ ID NO: 15).

[0427] Characterization of other conjugated polynucleotide-polypeptide constructs was performed as described above. Figures 8-10 show reproducible current versus time traces that allow characterization of constructs incorporating peptide regions containing positively charged amino acids (SEQ ID NO:16, Figure 8); aromatic amino acids (SEQ ID NO:17, Figure 9) and negatively charged amino acids (SEQ ID NO:15, Figure 10).

[0428] For ease of reference, the schematic structure of the construct obtained using peptide SEQ ID NO:17 is shown in FIG.

[0429] Example 2 This example demonstrates the utility of the disclosed methods in characterizing polynucleotide-polypeptide constructs derived from peptides that were not pre-synthesized to contain an attachment group.

[0430] In this example, the Y adapter was the same as in Example 1, and the dsDNA tail was prepared by annealing DNA oligonucleotides (SEQ ID NO: 13, SEQ ID NO: 14). Data collection was performed on a MinION Mk1b from Oxford Nanopore Technologies and a custom MinION flow cell equipped with an MspA nanopore, using the protocol established in Example 1.

[0431] The peptide analyte used in this example was nearly identical to the model peptide used in Example 1 (GGSGDDSGSG, SEQ ID NO: 15 in Example 1; SEQ ID NO: 18 in Example 2), but lacked the pre-synthesized azide molecule for click chemistry conjugation of the polynucleotide adaptor and tail. An additional C-terminal cysteine ​​was included to allow for maleimide chemistry. The N-terminus of the peptide was functionalized with a tetrazine-NHS ester compound (BroadPharm, product code: BP-22946). Unconjugated tetrazines were removed with amino-functionalized magnetic particles (Sigma Aldrich, product code: 53572).

[0432] The peptides were then incubated with DNA tails (SEQ ID NO:13, SEQ ID NO:14) overnight at 4 °C to facilitate the click reaction between tetrazine and TCO (transcyclooctene). After incubation, possible disulfide bonds between the C-terminal peptide cysteines were reduced with 5 mM DTT for 30 min at room temperature and the peptide-DNA conjugates were purified using Agencourt AMPure XP beads (Beckman Coulter) to remove unreacted peptide and DTT. The exposed cysteines were then reacted with azide-PEG3-maleimide (BroadPharm, product code: BP-22468). Excess maleimide linker was removed with Agencourt AMPure XP beads and the constructs were reacted with Y-adapters via click chemistry between BCN and azide overnight at 4 °C. The resulting construct, formed by conjugation between the C-terminus of the peptide and the Y adaptor and the N-terminus and the DNA tail, was purified using Agencourt AMPure XP beads to separate the complete construct from the peptide-DNA tail.

[0433] The final construct was evaluated as described in Example 1 using the exemplary current traces shown in Figure 12. As can be seen, characterization of the peptide was possible without the need for pre-synthesis of the attachment point.

[0434] Example 3 This example compares the disclosed methods in characterizing a 21 amino acid peptide compared to a 10 amino acid peptide.

[0435] Polynucleotide-polypeptide conjugates of 21 amino acid peptides were prepared and analyzed according to the methods described in Example 1. Current versus time traces obtained with the 21 amino acid construct were compared to those obtained with the 10 amino acid construct from Example 2. The peptide sequences used were GDDDGSASGDDDGSASGDDDG (21 aa, SEQ ID NO: 19) and GGSGDDSGSG (10 aa, SEQ ID NO: 15).

[0436] Data showing current versus time traces for the translocation of polynucleotide-peptide conjugates of a 10 amino acid peptide and a 21 amino acid peptide are shown in Figure 13. The two traces, aligned on the same time scale, show that the current section for the 21 amino acid polypeptide is approximately twice as long as that for the 10 amino acid polypeptide.

[0437] Thus, this example confirms that the disclosed methods can be used to characterize polypeptides of various and extended lengths.

[0438] Example 4 This example shows how conjugates can be re-read multiple times using polynucleotide handling proteins with different disulfide-closed linker lengths and different leader chemistries. This example demonstrates, among other things, the controlled movement of the conjugate back and forth to the nanopore detector, and the effect of varying the closure of the polynucleotide binding site of the polynucleotide handling protein and leader.

[0439] Y-adapters with leader arms containing 30 C3 spacer units were prepared by annealing four DNA oligonucleotides with the sequences of SEQ ID NO:20, SEQ ID NO:21, SEQ ID NO:22, and SEQ ID NO:23. A DNA motor (Dda helicase) was loaded onto each adapter and disulfide-closed via reaction with one of the following linkers: diamide (TMAD), BMOE (1,2-bismaleimidoethane), BMOP (1,3-bismaleimidopropane), BMB (1,4-bismaleimidobutane), BM(PEG)2 (1,8-bismaleimido-diethylene glycol), or BM(PEG)3 (1,11-bismaleimido-triethylene glycol).

[0440] E. coli K12 PCR DNA was obtained by extraction from E. coli cells using a Qiagen genomic chip kit, sheared to approximately 10 kb cutoff using a Covaris gTube, end-repaired and dA-tailed using the Ultra II End Repair and dA-Tailing Kit (New England Biolabs), ligated to PCR adapters (PCA; Oxford Nanopore Technologies), and PCR amplified using LongAmp Taq. The resulting double-stranded analytes were end-repaired and dA-tailed by NEBNext End Repair and NEBNext dA Tailing Modules (New England Biolabs (NEB)), generating 3' dA overhangs on both ends of each fragment. Samples were ligated to the T overhangs of the Y adapters using LNB and T4 DNA ligase from the Oxford Nanopore Technologies sequencing kit (LSK-SQK109). Samples were purified using Agencourt AMPure XP (Beckman Coulter) beads and washed twice with LFB from the Oxford Nanopore Technologies sequencing kit (LSK-SQK109). The ligated substrates were eluted in elution buffer (EB) from the same kit to obtain the "DNA library". A DNA library was prepared separately using an adaptor carrying the Dda helicase closed with a disulfide linker, as described above.

[0441] Electrical measurements were taken on a custom MinION flow cell from Oxford Nanopore Technologies with a CsgG nanopore inserted and a MinION Mk1b. To 1170 μL FB (from the Oxford Nanopore Technologies sequencing kit (SQK-LSK109)), 30 μL of FLT was added to obtain the tether mix. 800 μL of tether mix was run through the system, followed by a 5 minute wait and an additional 200 μL of tether mix was run through the system with the SpotON port open. 37.5 μL SQB from the Oxford Nanopore Technologies sequencing kit (SQK-LSK109), 15 μL of DNA library, and 22.5 μL of LB from the Oxford Nanopore Technologies sequencing kit (SQK-LSK109) were mixed to obtain the "sequencing mix". 75 μL of sequencing mix was added to the MinION flow cell via the SpotON flow cell port.

[0442] A custom sequence script was prepared to control the applied potential using the active unblocking circuitry of the MinION. The sequencing voltage was set to 180 mV, and when an enzyme stall level was detected the voltage was switched to zero by disconnecting the channel for 5 seconds to de-stall the motor protein. The classification of stall and chain (sequencing) levels was programmed into the configuration file of the MinKNOW instrument control software to allow detection of stalled species and to apply an unblocking potential that would not cause complete release of the chain. The script worked as follows: if MinKNOW detected that the chain was at the stall level, it applied the unblocking potential for 5 seconds, then returned to a sequencing potential of 180 mV to check for the actively sequenced chain 5 times. If the stall level was still present, the unblocking potential was applied for another 25 seconds and repeated 5 times. A 3 second rest period was built in between each unblocking attempt. If MinKNOW detected an active sequencing chain when it returned the sequencing potential, it stopped the unblocking attempt and only applied the sequencing potential. If no active sequencing strands are generated throughout this process, MinKNOW will turn off the channel. Active deblocking was set to be triggered upon recognition of a block level not related to terminal C3 levels, strands, open pores, or enzyme stall levels. Every 15 min, a "mux scan" was applied to reset the system, globally unblocking all channels of the flow cell and checking for active nanopores at 180 mV. Raw data were collected in bulk FAST5 files using MinKNOW software (Oxford Nanopore Technologies).

[0443] Strand-level events from single-channel data that occurred immediately after the C3 level were recorded as potential re-reads. These re-reads were confirmed by comparing the sequence of base calling and re-reads occurring after the open pore and destalling events with the original read. Events in the same read direction and within the range of the original read were classified as re-reads. Re-read efficiency was quantified in two ways: (i) the percentage of reads that drop back and re-read within 30 seconds of reaching the C3 reader, and (ii) the drop-back distance, which is the length of the re-read, i.e., the distance the enzyme was pushed back from the C3 reader.

[0444] The table below shows the results of this experiment. The results show re-reads with all linkers tested. All linkers tested were useful in this regard, and increasing linker length generally increased the percentage of reads with re-reads within 30 seconds of reaching the C3 reader. [Table 3]

[0445] Additional Y-adapters with leader arms having RNA or C3 leader chemistry were prepared by annealing four DNA oligonucleotides with sequences SEQ ID NO:20, SEQ ID NO:21, and SEQ ID NO:22 and a leader oligonucleotide selected from SEQ ID NO:23, SEQ ID NO:24, and SEQ ID NO:25. A DNA motor (Dda helicase) was loaded onto each adapter and the disulfides were closed via reaction with 1,2-bismaleimidoethane (BMOE).

[0446] Re-read events were scored as above, except when the leader contained RNA, the re-read occurred at the strand level. The table below shows the results of this experiment. The results show re-reads with all leader oligonucleotides tested. [Table 4]

[0447] Example 5 This example shows the controlled translocation of a conjugate containing two pieces of polynucleotide; a dsDNA adaptor (FIG. 14A) and a polypeptide flanked by a dsDNA tail (FIG. 14B-full construct). A polynucleotide handling protein on the cis side of the nanopore controls the movement of the conjugate in the direction from the trans side of the nanopore to the cis side of the nanopore, thus allowing the polypeptide to be characterized as it moves relative to the nanopore. The construct is first captured via the leader section of the dsDNA tail and then pulled across the nanopore using a positive voltage, and the second strand is peeled off as the ssDNA moves through the nanopore. Once the end of the construct reaches the nanopore, a traptavidin-biotin blocker prevents the conjugate from passing through the pore and the voltage is temporarily turned off, allowing the construct to partially diffuse out of the nanopore and the enzyme to move across the stall chemistry weakened by the dsDNA being unwound by the pore. As the enzyme moves along the DNA, it pulls the conjugate out of the nanopore, thus translocating the peptide across the nanopore and allowing it to be characterized. When the enzyme stops at the barrier between the DNA and the peptide, the voltage again pulls the conjugate into the nanopore. The repeated cycle of the enzyme translocating the peptide out of the nanopore, and the voltage pulling it back, allows the same peptide to be read over and over again.

[0448] DNA was prepared by annealing DNA oligonucleotides (SEQ ID NO:26, SEQ ID NO:27, SEQ ID NO:28) for the adaptors, oligonucleotides (SEQ ID NO:31, SEQ ID NO:32) for the dsDNA tails, and oligonucleotides (SEQ ID NO:29, SEQ ID NO:30) for the dsDNA linkers. A DNA motor (Dda helicase) was filled into the adaptors to close them and then incubated with a 3-fold excess of in-house purified traptavidin to prevent the helicase from sliding off the DNA. Subsequent material was then purified using Agencourt AMPure XP (Beckman Coulter) beads and stored in the presence of excess traptavidin to ensure stability of the constructs over time. Asymmetric fragments of dsDNA of variable length (400, 800, 2000, and 3600 bp) were generated by PCR from Lambda DNA and dA-tailed using the NEBNext® Ultra™ II End Repair / dA-Tailing module. A USER digest was performed to reveal the 5'AGGA overhangs. The DNA was then purified using Agencourt AMPure XP (Beckman Coulter) beads.

[0449] In this example, model polypeptide analytes were obtained containing presynthesized azide moieties at the N-terminus and immediately at the C-terminus using an ethyldiamine spacer along the peptide backbone. Each analyte was then conjugated to a dsDNA tail and a dsDNA linker via a copper-free click chemistry reaction between the azide and a BCN (bicyclo[6.1.0]nonyne) moiety. The resulting constructs were purified using Agencourt AMPure XP (Beckman Coulter) beads and washed twice with 28% PEG8000, 2.5 M NaCl and 25 mM Tris pH 8.0. The conjugated substrates were eluted in 20 mM Tris-Cl, 50 mM NaCl (pH 8.0).

[0450] Prior to data collection, the complete construct was assembled via a ligation reaction involving DNA adaptors, asymmetric DNA, dsDNA linkers and peptides clicked onto the dsDNA tail, and T4 ligase. The conjugate was purified using Agencourt AMPure XP (Beckman Coulter) beads.

[0451] Electrical measurements were taken using a custom MinION flow cell equipped with a MinION Mk1b or GridION from Oxford Nanopore Technologies and a CsgG nanopore (CsgG-CsgF pore complex as described in International Application No. 2019 / 002893). The flow cell was flushed with tether mix containing 50 nM DNA tether and SQB buffer. 800 μL of tether mix was first added for 5 min, then another 200 μL of mix was run through the system with the SpotON port open. DNA-peptide constructs were prepared by pre-incubating purified T4 ligation constructs with a 2-fold excess of in-house purified traptavidin to bind to biotin on the DNA adapter. After 5 min incubation, SQB-like buffer from the Oxford Nanopore Technologies sequencing kit (SQK-LSK109) was added along with LB from the Oxford Nanopore Technologies sequencing kit (SQK-LSK109) to obtain the "sequencing mix". 75 μL of the sequencing mix was added to the MinION flow cell via the SpotON flow cell port. The mixture was incubated on the flow cell for 5 min to allow tethering of the construct and subsequent capture by the nanopore.

[0452] Data was collected using a custom script based on the SQK-LSK109 baseline sequencing script provided by Oxford Nanopore Technologies for the R9.4.1 flow cell. The script was configured to recognize and accept open pore, peptide, and chain levels, and active deblocking was configured to trigger upon recognition of terminal biotin-traptavidin blockade levels and other blockades.

[0453] Examples of current versus time traces for two model peptides can be seen in Figures 15 and 16. A comparison of three different peptides is shown in Figure 17.

[0454] Example 6 This example shows the controlled translocation of a conjugate containing two pieces of polynucleotide: a polynucleotide comprising a DNA hairpin containing a 3' Click group, and a "clinker": a polypeptide adjacent to a dsDNA adaptor (see Figures 18 and 19) linked to a peptide via a dsDNA bearing a 5' Click group.

[0455] The polynucleotide handling motor protein, loaded onto the dsDNA adaptor and present on the cis side of the nanopore, controls the movement of the conjugate in the direction from the trans side of the nanopore to the cis side of the nanopore, thus allowing the polypeptide to be characterized as it moves relative to the nanopore. The construct is first captured by the nanopore via the leader section of the dsDNA adaptor, then pulled across the nanopore using an applied positive voltage, and the second strand is peeled off as the ssDNA moves through the nanopore. At the same time, the motor protein is pushed back across the peptide until it reaches a hairpin blocker (blocking moiety). The ends of the hairpin are crosslinked by a psoralen moiety that prevents the hairpin from separating. At this point, the motor protein controls the movement of the conjugate, as it can move along the ssDNA between the hairpin and the peptide. Each time the motor protein reaches a DNA-peptide interface, it stalls briefly, allowing it to be pushed back onto the hairpin bridge by the applied positive voltage acting on the conjugate so that the cycle can start again. This results in repeated reads of the peptide. This scheme allows for a theoretically infinite number of peptide reads until the conjugate is ejected from the nanopore by applying a negative voltage (see FIG. 20, traces 1 and 2: polyD and SG-RR peptides). Note that the hairpin functions as a blocking moiety even in the absence of interstrand bridges within the hairpin. This typically results in a finite number of re-reads until the dsDNA hairpin unwinds and the enzyme falls off the end of the construct (see FIG. 20, trace 3: SG-YY peptide). Three separate peptides were characterized using this scheme, each of which resulted in a distinguishable peptide signal, which is consistently observed across re-reads.

[0456] The DNA components were annealed as follows: DNA oligonucleotides for the adaptor (SEQ ID NO:33, SEQ ID NO:34, SEQ ID NO:35), oligonucleotides for the clinker (SEQ ID NO:36, SEQ ID NO:37), and oligonucleotide for the hairpin (SEQ ID NO:38). A DNA motor protein (modified Dda helicase enzyme) was loaded and closed onto the adaptor using techniques such as those described in International Application No. 2014 / 013260. The subsequent material was then purified using Agencourt AMPure XP (Beckman Coulter) beads. The hairpin psoralen crosslink was formed by irradiating the strands with ultraviolet (UV) light (365 nm) for 5 min.

[0457] In this example, model polypeptide analytes were obtained containing an azide (N3) moiety presynthesized at the N-terminus and a methyl-tetrazine (Tet) moiety synthesized immediately at the C-terminus using an ethyldiamine spacer along the peptide backbone (SEQ ID NO:39, SEQ ID NO:40, SEQ ID NO:41). Each peptide was then conjugated to an irradiated DNA hairpin via a copper-free click chemistry reaction between the tetrazine and TCO (trans-cycloctene) moieties, and then conjugated to a dsDNA clinker via a copper-free click chemistry reaction between the azide and BCN (bicyclo[6.1.0]nonyne) moieties. Finally, an adapter was ligated to the clinker using T4 ligase. The constructs were purified using Agencourt AMPure XP (Beckman Coulter) beads and washed twice with 28% PEG8000, 2.5 M NaCl and 25 mM Tris pH 8.0 after each step of assembly. The final construct was eluted in 50 mM NaCl, 25 mM HEPES, pH 7.5.

[0458] Electrical measurements were taken using Oxford Nanopore Technologies' MinION Mk1b or GridION devices and a custom MinION flow cell with a modified transmembrane nanopore from Rhodococcus inserted. The flow cell was flushed with a tether mix containing 50 nM DNA tether and SQB buffer. First, 800 μL of tether mix was added for 5 min, then an additional 200 μL of mix was run through the system with the SpotON port open. Constructs were prepared for sequencing by adding SQB buffer and water from the Oxford Nanopore Technologies sequencing kit (SQK-LSK109), and 75 μL of this sequencing mix was added to the MinION flow cell via the SpotON flow cell port.

[0459] Data was collected using a custom script based on the SQK-LSK109 baseline sequencing script provided by Oxford Nanopore Technologies for the R9.4.1 flow cell. The script was configured to recognize and accept open nanopore current and alternations ("flicks") from positive to negative voltages every minute (static flicks).

[0460] Examples of current versus time traces showing the readout of three model peptides are shown in FIG. 20: N3-DDDDDDDDDD-Tet (trace 1, "peptide polyD"); N3-GGSGRRSGSG-Tet (trace 2, "peptide SG-RR"), and N3-GGSGYYSGSG-Tet (trace 3, "peptide SY-YY").

[0461] The table below shows the sequences of the polynucleotide / polypeptide chains used in this example. [Table 5]

[0462] Example 7 This example shows the controlled translocation of a conjugate containing a polypeptide flanking two pieces of polynucleotide: a clinker with a 3' click group; a dsDNA adaptor linked to a peptide via dsDNA; and a ssDNA tail containing a leader and a 5' click group (see Figures 21 and 22). (Although this example shows a ssDNA tail, the tail could equally well be comprised of dsDNA.)

[0463] The polynucleotide handling motor protein, loaded onto the dsDNA adaptor and present on the cis side of the nanopore, controls the movement of the conjugate in the direction from the trans side of the nanopore to the cis side of the nanopore, thus allowing the polypeptide to be characterized as it moves relative to the nanopore. In the presence of ATP "fuel" for the motor protein, the protein proceeds along the adaptor and clinker strands to the DNA-peptide interface and stalls as it cannot proceed along the peptide. The construct is first captured by the nanopore via the leader section of the ssDNA tail, then pulled across the nanopore using an applied positive voltage, and the second strand is peeled off as the ssDNA moves through the nanopore. At the same time, the motor protein is pushed against the DNA backblocker (blocking moiety). The motor protein can now move along the ssDNA between the backblocker and the peptide, stalling each time it reaches the DNA-peptide interface, allowing it to be pushed back against the blocker by the applied positive voltage acting on the conjugate so that the cycle can start again. This results in repeated reading of the peptide. Finally, the DNA backblocker can be peeled off by the force of the applied voltage, allowing the motor protein to fall off the end of the construct, thus releasing the strand from the nanopore. Three separate peptides were characterized using this scheme, each of which yielded distinguishable peptide signals that were consistently observed across re-reads (see Figure 23).

[0464] The DNA components were annealed as follows: DNA oligonucleotides for the adaptor (SEQ ID NO: 42, SEQ ID NO: 43, SEQ ID NO: 44), oligonucleotides for the clinker (SEQ ID NO: 45, SEQ ID NO: 46), and oligonucleotides for the ssDNA tail (SEQ ID NO: 47). The DNA motor (modified Dda helicase enzyme) was loaded and closed onto the adaptor using the technique described in the previous example. The subsequent material was then purified using Agencourt AMPure XP (Beckman Coulter) beads.

[0465] In this example, model polypeptide analytes were obtained containing an azide moiety presynthesized at the N-terminus and a methyltetrazine moiety immediately following the C-terminus using an ethyldiamine spacer along the peptide backbone (SEQ ID NO: 48, SEQ ID NO: 49, SEQ ID NO: 50). Each peptide was then conjugated to a dsDNA clinker via a copper-free click chemistry reaction between the tetrazine and TCO (transcycloctene) moieties, and then to a ssDNA tail via a copper-free click chemistry reaction between the azide and BCN (bicyclo[6.1.0]nonyne) moieties. Finally, an adaptor was ligated to the clinker using T4 ligase. The constructs were purified using Agencourt AMPure XP (Beckman Coulter) beads and washed twice with 28% PEG8000, 2.5 M NaCl and 25 mM Tris pH 8.0 after each step of assembly. The final construct was eluted in 50 mM NaCl, 25 mM HEPES, pH 7.5.

[0466] Electrical measurements were taken using Oxford Nanopore Technologies' MinION Mk1b or GridION devices and a custom MinION flow cell with a modified transmembrane nanopore from Rhodococcus inserted. The flow cell was flushed with a tether mix containing 50 nM DNA tether and SQB buffer. First, 800 μL of tether mix was added for 5 min, then an additional 200 μL of mix was run through the system with the SpotON port open. Constructs were prepared for sequencing by adding SQB buffer and water from the Oxford Nanopore Technologies sequencing kit (SQK-LSK109), and 75 μL of this sequencing mix was added to the MinION flow cell via the SpotON flow cell port.

[0467] Data was collected using a custom script based on the SQK-LSK109 baseline sequencing script provided by Oxford Nanopore Technologies for the R9.4.1 flow cell. The script was configured to recognize and accept open nanopore current and alternations ("flicks") from positive to negative voltages every minute (static flicks).

[0468] Examples of current versus time traces showing the readout of three model peptides are shown in FIG. 23: N3-DDDDDDDDDD-Tet (trace 1, "peptide poly D"); N3-GGSGDDSGSG-Tet (trace 2, "peptide SG-DD"), and N3-GGSGRRSGSG-Tet (trace 3, "peptide SG-RR").

[0469] The table below shows the sequences of the polynucleotide / polypeptide chains used in this example. [Table 6]

[0470] Description of sequence listing SEQ ID NO: 1 shows the amino acid sequence of (hexahistidine-tagged) exonuclease I (EcoExo I) from E. coli.

[0471] SEQ ID NO:2 shows the amino acid sequence of the exonuclease III enzyme from E. coli.

[0472] SEQ ID NO: 3 shows the amino acid sequence of the RecJ enzyme from T. thermophilus (TthRecJ-cd).

[0473] SEQ ID NO: 4 shows the amino acid sequence of bacteriophage lambda exonuclease, which is one of three identical subunits that make up the trimer (http: / / www.neb.com / nebecomm / products / productM0262.asp).

[0474] SEQ ID NO:5 shows the amino acid sequence of Phi29 DNA polymerase from Bacillus subtilis.

[0475] SEQ ID NO: 6 shows the amino acid sequence of Trwc Cba (Citromicrobium bathyomarinum) helicase.

[0476] SEQ ID NO: 7 shows the amino acid sequence of Hel308 Mbu (Methanococcoides burtonii) helicase.

[0477] SEQ ID NO: 8 shows the amino acid sequence of Dda helicase 1993 from enterobacteriaceae phage T4.

[0478] SEQ ID NO:9 shows the sequence of the first polynucleotide strand (DNA1-top with C3[(OC3H6OPO3)) leader, 3'BCN click attachment, and enzyme stall chemistry; 8=iSp18[(OCH2CH2)6OPO3]) used to generate the Y adapter as described in Example 1.

[0479] SEQ ID NO: 10 shows the sequence of the second polynucleotide strand (DNA1-back with a side arm for tether) used to generate the Y adapter as described in Example 1.

[0480] SEQ ID NO:11 shows the sequence of the third polynucleotide strand (DNA1-bottom) used to generate the Y adapter as described in Example 1.

[0481] SEQ ID NO: 12 shows the sequence of the first polynucleotide strand (DNA2-top strand, 5'BCN click chemistry) used to generate a dsDNA tail as described in Example 1.

[0482] SEQ ID NO: 13 shows the sequence of the second polynucleotide strand (DNA2-top strand, 5'TCO (orthogonal click chemistry)) used to generate the dsDNA tail as described in Example 1.

[0483] SEQ ID NO: 14 shows the sequence of the polynucleotide strand (DNA2-bottom strand, no side arms) used to generate the dsDNA tail as described in Example 2.

[0484] SEQ ID NO: 15 shows the amino acid sequence of the first peptide fragment used to generate the first polynucleotide-polypeptide construct as described in Example 1.

[0485] SEQ ID NO: 16 shows the amino acid sequence of a second peptide fragment used to generate a second polynucleotide-polypeptide construct as described in Example 1.

[0486] SEQ ID NO:17 shows the amino acid sequence of a third peptide fragment used to generate a third polynucleotide-polypeptide construct as described in Example 1.

[0487] SEQ ID NO: 18 shows the amino acid sequence of the peptide fragment used to generate the polynucleotide-polypeptide construct as described in Example 2.

[0488] SEQ ID NO:19 shows the amino acid sequence of the 21 amino acid peptide fragment used to generate the polynucleotide-polypeptide construct as described in Example 3.

[0489] SEQ ID NOs: 20 to 25 show the polynucleotide sequences of the oligonucleotides described in Example 4.

[0490] SEQ ID NO:26 shows the polynucleotide sequence of the adapter top strand with 5' biotin blocker chemistry as described in Example 5.

[0491] SEQ ID NO:27 shows the polynucleotide sequence of adapter bottom strand 1 as described in Example 5.

[0492] SEQ ID NO:28 shows the polynucleotide sequence of adapter bottom strand 2 as described in Example 5.

[0493] SEQ ID NO: 29 shows the polynucleotide sequence of the linker, top strand, with a click group (BCN) for peptide attachment as described in Example 5.

[0494] SEQ ID NO: 30 shows the polynucleotide sequence of the linker, bottom strand, as described in Example 5.

[0495] SEQ ID NO: 31 shows the polynucleotide sequence of the tail, top strand, with a click group for peptide attachment and a C3 section for capture as described in Example 5.

[0496] SEQ ID NO:32 shows the polynucleotide sequence of the tail, bottom strand, having the tether portion as described in Example 5.

[0497] SEQ ID NOs: 33-50 show the polynucleotide and polypeptide sequences of the chains used in Examples 6 and 7.

[0498] Sequence Listing SEQ ID NO:1 - Exonuclease I from E. coli [ka] SEQ ID NO:2 - Exonuclease III enzyme from E. coli [ka] SEQ ID NO:3 - RecJ enzyme from T. thermophilus [ka] SEQ ID NO:4 - Bacteriophage lambda exonuclease [ka] SEQ ID NO:5 - Phi29 DNA polymerase [ka] SEQ ID NO:6 - Trwc Cba helicase [ka] SEQ ID NO:7 - Hel308 Mbu helicase [ka] SEQ ID NO:8 - Dda helicase [ka] SEQ ID NO:9 33333333333333333333333333333GCAATCGTCGAATGCGTACGTCGTTTTTTTT8AATGTACTTCGTTCAGTTACGTATTGC-BCN SEQ ID NO:10 CGACGTACGCATTCGACGATTGCTTTGAGGCGAGCGGTCAA SEQ ID NO:11 GCAATACGTAACTGAACGAAGT / iBNA-A / / iBNA-meC / / iBNA-A / / iBNAT / / 3BNA-T / SEQ ID NO:12 BCN-AATGTACTTCGTTCAGTTACGTATTGCTGCTTGGGTGTTTAACC SEQ ID NO:13 TCO-AATGTACTTCGTTCAGTTACGTATTGCTGCTTGGGTGTTTAACC SEQ ID NO:14 GGTTAAACACCCAAGCAGCAATACGTAACTGAACGAAGTACATT SEQ ID NO:15 X-GGSGDDSGSG-ed-X (X=azidoacetyl, ed=ethylenediamine) SEQ ID NO:16 X-GGSGRRSGSG-ed-X (X=azidoacetyl, ed=ethylenediamine) SEQ ID NO:17 X-GGSGYYSGSG-ed-X (X=azidoacetyl, ed=ethylenediamine) SEQ ID NO:18 GGSGDDSGSC SEQ ID NO:19 GDDDGSASGDDDGSASGDDDG SEQ ID NO:20 GTTATTCAAGACTTCTTTAATACACTTTTTTTTT / iSp9 / AATGTACTTCGTTCAGTTACGTATTGCTTTGGCGTCTGCTTGGGTGTTTAACCT SEQ ID NO:21 GTGTATTAAAGAAGTCTTGAATAACTTTGAGGCGAGCGGTCAA SEQ ID NO:22 TTTGCAATACGTAACTGAACGAAGT / iBNA-A / / iBNA-MeC / / iBNA-A / / iBNA-T / / 3BNA-T / sequence number 23 / 5Phos / GGTTAAACACCCAAGCAGACGCCTTT / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 C3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / sequence number 24 / 5Phos / GGTTAAACACCCAAGCAGACGCCTTTmUmUmUmUmUmU / iSpC3 / sequence number 25 / 5Phos / GGTTAAACACCCAAGCAGACGCCmUmU / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 C3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / sequence number 26 / 5BiotinTEG / TTTTTTTTTT / iSp18 / AATGTACTTCGTTCAGTTACGTATTGCTTTGGCGTCTGCTTGGGTGTTTAACCT SEQ ID NO:27 TTTGCAATACGTAACTGAACGAAGT / iBNA-A / iBNA-meC / iBNA-A / iBNA-T / iBNA-T SEQ ID NO:28 / 5Phos / GGTTAAACACCCAAGCAGACGCC SEQ ID NO:29 / 5phos / GTTGGAATTTTTTT-BCN SEQ ID NO:30 AAAAAAAATTCCAACTCCT SEQ ID NO:31 BCN-GCAATCGTCGAATGCGTACGTCG333333333333333333333333333333 SEQ ID NO:32 CGACGTACGCATTCGACGATTGCTTTGAGGCGAGCGGTCAA

Claims

1. 1. A method for characterizing a target polypeptide having a first end and a second end, comprising: the target polypeptide is included in a conjugate, wherein the first end of the target polypeptide is conjugated to the first end of a polynucleotide; The method comprises: (i) contacting a detector having a first opening and a second opening, or contained in a structure having a first opening and a second opening, with the conjugate having a polynucleotide handling protein bound thereto, wherein the conjugate is bound to the polynucleotide handling protein at a polynucleotide binding site of the polynucleotide handling protein; (ii) taking one or more measurements characteristic of the target polypeptide when the polynucleotide handling protein controls movement of the conjugate in a first direction relative to the detector; (iii) debinding the conjugate from the polynucleotide binding site of the polynucleotide handling protein such that the conjugate moves in a second direction relative to the detector; (iv) allowing the conjugate to rebind to the polynucleotide binding site of the polynucleotide handling protein and taking one or more measurements characteristic of the target polypeptide as the polynucleotide handling protein controls movement of the conjugate in the first direction relative to the detector; thereby characterizing said target polypeptide, The method, wherein the second end of the target polypeptide is attached to a leader.

2. 1. A method for characterizing a target polypeptide having a first end and a second end, comprising: the target polypeptide is included in a conjugate, wherein the first end of the target polypeptide is conjugated to the first end of a polynucleotide; The method comprises: (i) contacting a detector having a first opening and a second opening, or contained in a structure having a first opening and a second opening, with the conjugate having a polynucleotide handling protein bound thereto, wherein the conjugate is bound to the polynucleotide handling protein at a polynucleotide binding site of the polynucleotide handling protein; (ii) taking one or more measurements characteristic of the target polypeptide when the polynucleotide handling protein controls movement of the conjugate in a first direction relative to the detector; (iii) debinding the conjugate from the polynucleotide binding site of the polynucleotide handling protein such that the conjugate moves in a second direction relative to the detector; (iv) allowing the conjugate to rebind to the polynucleotide binding site of the polynucleotide handling protein and taking one or more measurements characteristic of the target polypeptide as the polynucleotide handling protein controls movement of the conjugate in the first direction relative to the detector; thereby characterizing said target polypeptide, The method, wherein the second end of the polynucleotide comprises a blocking moiety to prevent the polynucleotide handling protein from disassociating from the conjugate.

3. 3. The method of claim 2, wherein prior to step (i), the polynucleotide handling protein is bound to the portion of the conjugate between the polypeptide and the blocking moiety.

4. 2. The method of claim 1, wherein the second end of the polynucleotide comprises a blocking moiety to prevent the polynucleotide handling protein from disassociating from the conjugate.

5. 5. The method of claim 4, wherein prior to step (i), the polynucleotide handling protein is bound to the portion of the conjugate between the polypeptide and the blocking moiety.

6. The method of claim 2 , wherein the second end of the polypeptide is attached to a leader.

7. 10. The method of claim 1 or 6, wherein prior to step (i), the polynucleotide handling protein is attached to the portion of the conjugate between the polypeptide and the free end of the leader.

8. 8. The method of claim 7, wherein the leader is attached to the polypeptide by a polynucleotide linker, and prior to step (i), the polynucleotide handling protein is bound to the linker.

9. 3. The method of claim 1 or 2, comprising repeating steps (iii) and (iv) multiple times.

10. 10. The method of claim 1 or 6, wherein step (i) comprises contacting the first opening with the leader under conditions such that the leader passes through the first and second openings.

11. 3. The method of claim 1 or 2, wherein in step (iii), the conjugate moves in a direction from the first opening to the second opening, and in steps (ii) and (iv), the polynucleotide handling protein controls the movement of the conjugate in a direction from the second opening to the first opening.

12. the detector comprises a transmembrane nanopore spanning a membrane having a cis side and a trans side; (i) the first opening of the nanopore is on the cis side of the membrane and the second opening of the nanopore is on the trans side, and the polynucleotide handling protein controls the movement of the conjugate through the nanopore in a direction from the trans side to the cis side of the membrane, and when the conjugate unbinds from the polynucleotide binding site of the polynucleotide handling protein, the conjugate moves through the nanopore in a direction from the cis side to the trans side of the membrane; or (ii) the first opening of the nanopore is on the trans side of the membrane, the second opening of the nanopore is on the cis side, the polynucleotide handling protein controls the movement of the conjugate through the nanopore in a direction from the cis side to the trans side of the membrane, and when the conjugate unbinds from the polynucleotide binding site of the polynucleotide handling protein, the conjugate moves through the nanopore in a direction from the trans side to the cis side of the membrane.

13. The method of claim 1 or 2, wherein the conjugate does not disassociate from the polynucleotide handling protein.

14. 3. The method of claim 1 or 2, wherein the polynucleotide handling protein is modified to prevent the conjugate from disassociating from the polynucleotide handling protein.

15. 3. The method of claim 1 or 2, wherein the polynucleotide handling protein is modified with a closing moiety to topologically close the polynucleotide binding site of the polynucleotide handling protein around the conjugate.

16. the polynucleotide handling protein is modified to facilitate attachment of the closing moiety to the polynucleotide handling protein; 16. The method of claim 15, wherein optionally, the polynucleotide handling protein is modified by substituting cysteine ​​or an unnatural amino acid for at least one amino acid in the polynucleotide handling protein.

17. i) the closing moiety comprises a bifunctional cross-linker; ii) the closing moiety comprises a bond, optionally a disulfide bond; iii) the closing moiety comprises a structure of the formula [ABC], where A and C are each independently a reactive functional group for reacting with an amino acid residue in the polynucleotide handling protein, and B is a linking moiety; Optionally, wherein (a) A and C are each independently a cysteine-reactive functional group; and / or (b) linking moiety B comprises a linear or branched, unsubstituted or substituted alkylene, alkenylene, alkynylene, arylene, heteroarylene, carbocyclylene, or heterocyclylene moiety, which is optionally interrupted and / or terminated by one or more atoms or groups selected from O, N(R), S, C(O), C(O)NR, C(O)O, unsubstituted or substituted arylene, arylene-alkylene, heteroarylene, heteroarylene-alkylene, carbocyclylene, carbocyclylene-alkylene, heterocyclylene, and heterocyclylene-alkylene, wherein R is selected from H, unsubstituted or substituted alkyl, and unsubstituted or substituted aryl; 16. The method of claim 15, further optionally, wherein linking moiety B comprises an alkylene, oxyalkylene, or polyoxyalkylene group, and / or A and C are each maleimide groups.

18. 16. The method of claim 15, wherein the closing portion has a length of about 1 Å to about 100 Å, optionally about 5 Å to about 50 Å.

19. 3. The method of claim 1 or 2, wherein the polynucleotide handling protein is a helicase, optionally a DNA-dependent ATPase (Dda) helicase.

20. The method of claim 1 or 6, wherein the leader comprises a polymer.

21. 10. The method of claim 1 or 6, wherein the leader is attached to the polypeptide by a linker, and optionally the linker comprises a polynucleotide.

22. 10. The method of claim 1 or 6, wherein the leader comprises one or more of a nucleotide lacking both a nucleobase and a sugar moiety (spacer moiety), a deoxyribonucleotide (DNA), a ribonucleotide (RNA), a peptide nucleotide (PNA), a glycerol nucleotide (GNA), a threose nucleotide (TNA), a locked nucleotide (LNA), a bridged nucleotide (BNA), an abasic nucleotide, or a nucleotide with a modified phosphate linkage.

23. 10. The method of claim 1 or 6, wherein the leader is configured to facilitate unbinding of the polynucleotide binding site of the polynucleotide handling protein from the conjugate.

24. 3. The method of claim 2, wherein the blocking moiety restricts movement of the conjugate through the polynucleotide binding site of the polynucleotide handling protein, thereby restricting movement of the conjugate in the second direction relative to the detector.

25. The method of claim 1 or 2, wherein the conjugate comprises multiple polypeptide sections and / or multiple polynucleotide sections.

26. The conjugate is L-N m -{P-N} n or L-{P-N} n -P m and wherein the structure comprises one or more structures of the form L is a leader, L is optionally an N moiety, The conjugate is L-N m -{P-N} n and when m is 1, L optionally comprises one or more different types of nucleotides relative to the adjacent N moieties; -P is a polypeptide, -N comprises a polynucleotide; -m is 0 or 1, -n is a positive integer, the detector comprises a transmembrane nanopore spanning a membrane having a cis side and a trans side, the method comprising passing the leader (L) through the nanopore, thereby contacting the polypeptide (P) with the nanopore; i) the polynucleotide handling protein is located on the cis side of the nanopore, and the method comprises enabling the polynucleotide handling protein to control the movement of the polynucleotide (N) in a direction from the trans side of the nanopore to the cis side of the nanopore, thereby controlling the movement of the polypeptide (P) through the nanopore; or 3. The method of claim 1 or 2, wherein i) the polynucleotide handling protein is located on the trans side of the nanopore, and the method comprises enabling the polynucleotide handling protein to control movement of the polynucleotide (N) in a direction from the cis side of the nanopore to the trans side of the nanopore, thereby controlling movement of the polypeptide (P) through the nanopore.

27. 3. The method of claim 1 or 2, wherein the or each polypeptide independently has a length of from 2 to about 50 peptide units and / or the or each polynucleotide independently has a length of from about 10 to about 1000 nucleotides.

28. The method of claim 1 or 2, wherein one or more adaptors and / or one or more tethers and / or one or more anchors are attached to the polynucleotide in the conjugate.

29. applying a force to the detector, wherein the polynucleotide handling protein controls movement of the target polynucleotide relative to the detector in a direction opposite to the applied force; The method of claim 1 or 2, wherein the force optionally comprises an electric potential applied to the detector.

30. i) the detector comprises a nanopore, the nanopore modified to increase the distance between the polynucleotide handling protein and the constricted region of the nanopore; and / or ii) the polynucleotide handling protein is modified to increase the distance between the active site of the polynucleotide handling protein and the detector; and / or 3. The method of claim 1 or 2, wherein iii) a displacer unit separates the polynucleotide handling protein from the detector, thereby increasing the distance between the polynucleotide handling protein and the detector, and optionally the displacer unit comprises one or more proteins.

31. 3. The method of claim 1 or 2, wherein the one or more measurements are characteristic of one or more features of the polypeptide selected from (i) the length of the polypeptide, (ii) the identity of the polypeptide, (iii) the sequence of the polypeptide, (iv) the secondary structure of the polypeptide, and (v) whether the polypeptide is modified.

32. 1. A conjugate comprising a polypeptide conjugated to a polynucleotide, i) a first end of the polypeptide is conjugated to a first end of the polynucleotide; ii) the second end of the polypeptide is attached to a polymeric leader, optionally via a linker, and optionally the linker comprises a polynucleotide; iii) if the linker comprises a polynucleotide, the polynucleotide or the linker binds to a polynucleotide binding site of a polynucleotide handling protein, and the polynucleotide binding site is modified with a closing moiety, thereby topologically closing the polynucleotide binding site of the polynucleotide handling protein around the polynucleotide; iv) A conjugate wherein the polynucleotide comprises a blocking moiety to prevent the polynucleotide handling protein from disassociating from the conjugate.

33. 1. A system comprising: - nanopores, - a conjugate as defined in claim 32.

34. A kit comprising: a polynucleotide, wherein a first end of the polynucleotide comprises a reactive functional group for conjugating to a first end of a target polypeptide and a second end of the polynucleotide comprises a blocking moiety; a polynucleotide handling protein, wherein the polynucleotide binding site of said polynucleotide handling protein is modified with a closing moiety to topologically close said polynucleotide binding site of said polynucleotide handling protein around said polynucleotide; an adaptor comprising a polymeric leader and a reactive functional group for attachment to the second end of said target polypeptide; - a nanopore.