method

JP2025500399A5Pending Publication Date: 2026-01-06OXFORD NANOPORE TECH LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024537834
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-12-23
Filing Date
2022-12-22
Publication Date
2026-01-06

AI Technical Summary

Technical Problem

Existing methods for characterizing polynucleotides using nanopore sensors face challenges such as uncontrolled movement of polynucleotides, slippage of motor proteins, and loss of heterogeneity in data aggregation, leading to inaccurate sequence determination and inefficiencies in polynucleotide characterization.

Method used

A method involving a motor protein that controls the movement of a polynucleotide through a detector with a first and second aperture, allowing the polynucleotide to move in a direction opposite to conventional methods, enabling repeated measurements and characterization by binding and rebinding to the motor protein, thereby improving data accuracy and reducing slippage.

Benefits of technology

The method enhances the accuracy and efficiency of polynucleotide characterization by reducing slippage and preserving heterogeneity, allowing for high-precision sequence determination and retention of epigenetic information through repeated readings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Provided herein are methods for characterizing a target polynucleotide as it moves relative to a nanopore using a motor protein. Also provided are polynucleotide adaptors and kits that include such adaptors. The methods, kits, and adaptors find use in characterizing, e.g., sequencing, polynucleotides.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present disclosure provides methods for characterizing a target polynucleotide as it moves relative to a detector, such as a transmembrane nanopore. The present disclosure also provides novel polynucleotide adaptors and kits for use in such methods. The present disclosure also provides methods for re-reading a polynucleotide. [Background technology]

[0002] Nanopore sensing is an approach to analyte detection and characterization that relies on the observation of individual binding or interaction events between analyte molecules and ion-conducting channels. Nanopore sensors can be fabricated by placing a single pore of nanometer dimensions in an insulating membrane and measuring the voltage-driven ionic current through the pore in the presence of an analyte molecule. The presence of an analyte in or near the nanopore will alter the ionic flow through the pore, resulting in a change in the ionic or electrical current measured across the channel. The identity of the analyte is revealed through its unique current signature, particularly the duration and extent of current block, as well as the variation in current level during the interaction time with the pore.

[0003] Polynucleotides are important analytes to sense in this manner. Nanopore sensing of polynucleotide analytes can reveal the identity and perform single molecule counting of sensed analytes, but can also provide information about their composition, such as their nucleotide sequence, as well as the presence of features such as base modifications, oxidation, reduction, decarboxylation, deamination, etc. Nanopore sensing has the potential to enable rapid and inexpensive polynucleotide sequencing, providing single molecule sequence reads for polynucleotides that are tens to tens of thousands of bases long.

[0004] Two of the key components of polymer characterization using nanopore sensing are (1) control of the movement of the polymer through the pore, and (2) identification of the component building blocks as the polymer moves through the pore. During nanopore sensing of an analyte such as a polynucleotide, it is important to control the movement of the polynucleotide relative to the pore. Without control of the movement, accurate characterization of the polynucleotide may be prevented or hindered. For example, if the movement of the polynucleotide relative to the pore is not controlled, it becomes difficult to accurately distinguish each nucleotide in a homopolymer polynucleotide.

[0005] It is known to control the movement of polynucleotides relative to a detector, such as a nanopore, by using motor proteins that control the movement of polynucleotides. Suitable motor proteins include polynucleotide handling enzymes, such as helicases, exonucleases, and topoisomerases. Motor proteins process polynucleotides in a controlled manner. Thus, motor proteins can be used to control the movement of polymers, such as polynucleotides, relative to a detector, such as a nanopore.

[0006] When the detector is a nanopore, the disclosed method typically involves feeding a polynucleotide into the nanopore using a motor protein, the operation of which is described in more detail herein. Methods for feeding polynucleotides into nanopores have been widely developed and have proven to be very useful for characterizing polynucleotides.

[0007] However, there remains a need for additional methods for characterizing polynucleotides. One problem is that it may be desirable to obtain data that differs from that obtained from a method that includes feeding a polynucleotide to a detector, such as a nanopore. For example, the error profile of data resulting from characterizing a polynucleotide in a method that includes feeding a polynucleotide to a detector may not be optimal for accurate characterization of the polynucleotide in some circumstances. Another problem is that when a motor protein is used to feed a polynucleotide to a detector, such as a nanopore, the motor protein may skip forward on the polynucleotide chain in an uncontrolled manner. This phenomenon is also known as slippage. Slippage can be problematic in characterizing a polynucleotide, for example, because one or more nucleotides in the polynucleotide may not be accurately characterized. This is particularly problematic when characterizing a polynucleotide to determine its sequence. Strategies for reducing slippage have traditionally focused on modifying motor proteins to minimize their tendency to slip on the polynucleotide chain. However, other methods of moving a polynucleotide relative to a detector, such as a nanopore, that can reduce slippage, would also be useful.

[0008] There is also a need for a method to improve the data obtained when characterizing polynucleotides. One problem is that in some cases, it is desirable to improve the accuracy of the characterization data obtained when characterizing polynucleotides. In some known methods, multiple polynucleotides are characterized from a sample of polynucleotides, and the data obtained are aggregated, thereby improving the overall accuracy. However, this can cause problems. For example, heterogeneity in the sample can mean that when aggregating data obtained from the characterization of multiple polynucleotide strands, useful information about the differences between strands can be lost. Furthermore, inefficiencies can occur due to the need to capture new strands for characterization after the first strands have been processed. Therefore, there is a need for alternative and / or improved methods of characterizing polynucleotides.

[0009] For these and other reasons, there is a need for new and / or improved methods of translocating polynucleotides relative to detectors, such as nanopores. Summary of the Invention

[0010] The present disclosure relates to a method for characterizing a target polynucleotide by using a motor protein when the target polynucleotide moves relative to a detector having a first opening and a second opening, or a detector contained in a structure having a first opening and a second opening. More specifically, the present disclosure relates to a method in which the motor protein controls the movement of the polynucleotide in a direction from a second opening to a first opening. As described in more detail herein, this direction is typically "outside" the detector from the "point of view" of the motor protein. Thus, the direction of movement of the polynucleotide is opposite to the known way in which the polynucleotide moves into a detector, such as a nanopore. This is described in more detail herein.

[0011] In the disclosed methods, a motor protein first binds to a leader attached to a target polynucleotide. The motor protein may be stalled at a stall moiety on the leader as described herein, and the methods provided herein may include destalling the motor protein such that the motor protein can control the movement of the polynucleotide out of a detector (e.g., a nanopore). Methods for stalling and destalling motor proteins are described in more detail herein.

[0012] Although the present disclosure provides a nanopore as an exemplary detector, the methods provided herein are suitable for detectors such as (i) zero mode waveguides, (ii) field effect transistors, optionally nowire field effect transistors, (iii) AFM tips, (iv) nanotubes, optionally carbon nanotubes, and (v) nanopores. The disclosed methods are particularly suitable for methods in which polynucleotides move through a detector or through a structure that contains a detector, such as a well in a detector chip.

[0013] Thus, provided herein is a method for characterizing a leader-bound target polynucleotide, the method comprising: (i) contacting a detector having a first opening and a second opening, or contained in a structure having a first opening and a second opening, with a reader under conditions such that the first opening is in contact with the reader and the target polynucleotide migrates in a direction from the first opening to the second opening, contacting the leader with a motor protein bound at a polynucleotide binding site of the motor protein; (ii) performing one or more measurements characteristic of the target polynucleotide when the motor protein controls movement of the target polynucleotide in a first direction relative to the detector, the first direction being from the second opening to the first opening; (iii) debinding the target polynucleotide from the polynucleotide binding site of the motor protein such that the target polynucleotide moves in a second direction relative to the detector, the second direction being from the first opening to the second opening; (iv) allowing the target polynucleotide to rebind to the polynucleotide binding site of the motor protein and performing one or more measurements characteristic of the target polynucleotide when the motor protein controls movement of the target polynucleotide in a first direction relative to a detector; and thereby characterizing the target polynucleotide.

[0014] In some embodiments, the target polynucleotide has a first end and a second end, the leader is attached to the first end, and the motor protein is oriented in an orientation to process the target polynucleotide in a direction from the second end toward the first end.

[0015] In some embodiments, the method comprises repeating steps (iii) and (iv) multiple times.

[0016] In some embodiments, prior to step (i), the target polynucleotide is comprised in or consists of a first strand of a double-stranded polynucleotide comprising a first strand and a second strand, in some embodiments, the portion of the first strand between the motor protein and the second end is hybridized to the second strand.

[0017] In some embodiments, the movement of the target polynucleotide in the direction from the first opening to the second opening comprises separation of the first strand from the second strand. In some embodiments, the movement of the target polynucleotide in the direction from the second opening to the first opening comprises annealing of the first strand to the second strand.

[0018] In some embodiments, a first strand of a double-stranded polynucleotide is attached to a second strand of a double-stranded polynucleotide.

[0019] In some embodiments, in step (ii), the motor protein controls movement of a first portion of the target polynucleotide in a first direction relative to the detector, and in step (iv), the motor protein controls movement of a second portion of the target polynucleotide in the first direction relative to the detector, the first portion at least partially overlapping the second portion. In some embodiments, the first portion is the same as the second portion.

[0020] In some embodiments, in step (ii), the motor protein controls movement of a first portion of the target polynucleotide in a first direction relative to the detector, and in step (iv), the motor protein controls movement of a second portion of the target polynucleotide in the first direction relative to the detector, wherein the first portion does not overlap with the second portion.

[0021] In some embodiments, the distance traveled by the target polynucleotide relative to the detector in step (iii) is greater than the distance traveled by the polynucleotide relative to the detector in step (ii) and / or step (iv).

[0022] In some embodiments, (a) in step (iii), the distance traveled by the target polynucleotide relative to the detector is at least 1000 nucleotides in length, and / or (b) in steps (ii) and / or (iv), the distance traveled by the target polynucleotide relative to the detector is each independently at least 100 nucleotides in length.

[0023] In some embodiments, the second end of the target polynucleotide comprises a blocking moiety that prevents the motor protein from disassociating from the polynucleotide, hi some embodiments, the blocking moiety restricts movement of the target polynucleotide through the polynucleotide binding site of the motor protein, thereby restricting movement of the target polynucleotide in a second direction relative to the detector.

[0024] In some embodiments, the detector comprises a transmembrane nanopore spanning a membrane having a cis side and a trans side, a first opening of the nanopore is on the cis side of the membrane and a second opening of the nanopore is on the trans side, a motor protein controls movement of a target polynucleotide through the nanopore from the trans side to the cis side of the membrane, and when the target polynucleotide unbinds from the polynucleotide binding site of the motor protein, the target polynucleotide translocates through the nanopore from the cis side to the trans side of the membrane.

[0025] In some embodiments, the detector comprises a transmembrane nanopore spanning a membrane having a cis side and a trans side, a first opening of the nanopore at the trans side of the membrane and a second opening of the nanopore at the cis side, a motor protein controls movement of a target polynucleotide through the nanopore from the cis side to the trans side of the membrane, and when the target polynucleotide unbinds from the polynucleotide binding site of the motor protein, the target polynucleotide translocates through the nanopore from the trans side to the cis side of the membrane.

[0026] In some embodiments, the target polynucleotide does not disassociate from the motor protein, hi some embodiments, the motor protein is modified to prevent the target polynucleotide from disassociating from the motor protein.

[0027] In some embodiments, the motor protein is modified to facilitate debinding of a target polynucleotide from the polynucleotide binding site of the motor protein and / or to delay rebinding of the target polynucleotide to the polynucleotide binding site of the motor protein.

[0028] In some embodiments, the motor protein is modified with a closing moiety to (i) topologically close the polynucleotide binding site of the motor protein around the target polynucleotide, and / or (ii) facilitate debinding of the target polynucleotide from the polynucleotide binding site of the motor protein and / or retard rebinding of the target polynucleotide to the polynucleotide binding site of the motor protein.

[0029] In some embodiments, the motor protein is modified to facilitate attachment of a closing moiety to the motor protein.

[0030] In some embodiments, the motor protein is modified by substituting cysteine ​​or an unnatural amino acid for at least one amino acid in the motor protein.

[0031] In some embodiments, the closing moiety comprises a bifunctional crosslinker.

[0032] In some embodiments, the closing moiety bridges two amino acid residues of the motor protein, and at least one of the amino acids bridged by the closing moiety is a cysteine ​​or an unnatural amino acid.

[0033] In some embodiments, the closing portion has a length of about 1 Å to about 100 Å. In some embodiments, the closing portion has a length of about 5 Å to about 50 Å.

[0034] In some embodiments, the closing portion comprises a bond, hi some embodiments, the closing portion comprises a disulfide bond.

[0035] In some embodiments, the blocking moiety comprises a structure of the formula [ABC], where A and C are each independently a reactive functional group for reacting with an amino acid residue in a motor protein, and B is a linking moiety. In some embodiments, A and C are each independently a cysteine ​​reactive functional group. In some embodiments, linking moiety B comprises a linear or branched, unsubstituted or substituted alkylene, alkenylene, alkynylene, arylene, heteroarylene, carbocyclylene, or heterocyclylene moiety, which is optionally interrupted or terminated by one or more atoms or groups selected from O, N(R), S, C(O), C(O)NR, C(O)O, unsubstituted or substituted arylene, arylene-alkylene, heteroarylene, heteroarylene-alkylene, carbocyclylene, carbocyclylene-alkylene, heterocyclylene, and heterocyclylene-alkylene, where R is selected from H, unsubstituted or substituted alkyl, and unsubstituted or substituted aryl. In some embodiments, linking moiety B comprises an alkylene, oxyalkylene, or polyoxyalkylene group, and / or A and C are each maleimide groups.

[0036] In some embodiments, the motor protein is a helicase. In some embodiments, the motor protein is a DNA-dependent ATPase (Dda) helicase.

[0037] In some embodiments, prior to step (i), the motor protein is stalled on the leader. In some embodiments, the leader comprises a different type of nucleotide than the target polynucleotide. In some embodiments, the target polynucleotide comprises deoxyribonucleotides (DNA) or ribonucleotides (RNA) and the leader comprises one or more stall units and / or one or more nucleotides lacking both a nucleobase and a sugar moiety (spacer moiety), deoxyribonucleotides (DNA), ribonucleotides (RNA), peptide nucleotides (PNAs), glycerol nucleotides (GNAs), threose nucleotides (TNAs), locked nucleotides (LNAs), bridged nucleotides (BNAs), abasic nucleotides, or nucleotides with modified phosphate linkages.

[0038] In some embodiments, the second strand of the double-stranded polynucleotide comprises a membrane anchor or a transmembrane pore anchor.

[0039] In some embodiments, the method comprises applying a force to the detector, and the motor protein controls movement of the target polynucleotide relative to the detector in a direction opposite to the applied force, hi some embodiments, the force comprises an electric potential applied to the detector.

[0040] Also provided herein is a polynucleotide adaptor having a first end comprising a leader and a second end comprising an attachment point for attachment to a polynucleotide analyte at the first end of the polynucleotide analyte, the polynucleotide adaptor comprising a motor protein stalled thereon in an orientation for processing the adaptor in a direction from the second end to the first end.

[0041] Also provided herein is a kit comprising a first adaptor as defined herein and a second adaptor comprising (i) an attachment point for attachment to a polynucleotide analyte at a second end of the polynucleotide analyte, and (ii) a blocking moiety suitable for preventing a motor protein of the first adaptor from disassociating from the polynucleotide analyte when the first adaptor is attached to the polynucleotide analyte.

[0042] Also provided is a system for characterizing a target polynucleotide, comprising: - one or more polynucleotide adaptors as defined herein; a nanopore for characterizing a target polynucleotide as it translocates relative to the nanopore; - a motor protein for controlling the movement of a target polynucleotide relative to the nanopore.

[0043] In some embodiments the motor protein and / or the blocking moiety is as defined in any one of the preceding claims. [Brief description of the drawings]

[0044] [Figure 1] Schematic diagram showing the distinction between (A) the direction of polynucleotide (PN) movement out of a nanopore under the control of a motor protein according to the methods provided herein and (B) the direction of polynucleotide movement into the pore in a contrasting manner. The open arrows indicate the direction of motor protein (MP) and PN translocation. In both cases, the MP is, for example, a 5'-3' helicase. [Diagram 2]1 is a non-limiting schematic diagram of an embodiment of the method provided herein, in which the target polynucleotide is single-stranded, the target polynucleotide includes a leader (wavy line) at a first end of the target polynucleotide, and the motor protein is bound to the leader, e.g., stalled at the leader. The leader sequence is captured by the nanopore, and the single-stranded polynucleotide moves through the nanopore, pushing the motor protein toward the second end of the target polynucleotide. The motor protein then controls the movement of the polynucleotide "out" of the pore. The motor protein then unbinds from the target polynucleotide, and the target polynucleotide moves back "in" the pore. The motor protein then rebinds to the target polynucleotide and rereads (RR) the target polynucleotide. [Diagram 3] FIG. 1 is a non-limiting schematic diagram of an embodiment of the method provided herein, in which the target polynucleotide is double-stranded, the target polynucleotide (PN) includes a leader (wavy line) located at a first end of the first strand of the target polynucleotide, and the motor protein (MP) is bound to the leader, for example, by stalling at a stall portion included in the leader. The leader is captured by a nanopore, and the first strand of the target polynucleotide moves through the nanopore, pushing the motor protein toward the second end of the first strand of the target polynucleotide. The motor protein then controls the movement of the first strand of the polynucleotide "out" of the pore. The motor protein then unbinds from the first strand of the target polynucleotide, and the first strand of the target polynucleotide moves back "in" the pore. The motor protein then rebinds to the target polynucleotide and rereads (RR) the target polynucleotide. [Figure 4]1 is a non-limiting schematic diagram of an embodiment of the method provided herein, in which the target polynucleotide is double-stranded, the target polynucleotide (PN) includes a leader (wavy line) located at a first end of the first strand of the target polynucleotide, the motor protein (MP) binds to the leader, for example, by stalling at a stall moiety included in the leader, the hairpin adaptor links the second end of the first strand of the target polynucleotide to the first end of the second strand of the target polynucleotide, and the blocking moiety is present at the second end of the second strand. The leader is captured by a nanopore, and the first strand of the target polynucleotide, the hairpin adaptor, and the second strand move through the nanopore, pushing the motor protein toward the second end of the first strand of the target polynucleotide, the hairpin, and the second end of the second strand. The motor protein then controls the movement of the second strand, the hairpin, and the first strand of the polynucleotide "out" of the pore. The motor protein then unbinds from the first strand of the target polynucleotide, and the first strand, the hairpin, and the second strand of the target polynucleotide move back "into" the pore. The blocking moiety prevents the motor protein from unassociating from the target polynucleotide. The motor protein then rebinds to the target polynucleotide and rereads the target polynucleotide (RR). [Diagram 5]FIG. 1 is a non-limiting schematic diagram of an embodiment of the method provided herein, in which the target polynucleotide is double-stranded, the target polynucleotide (PN) includes a leader (wavy line) located at a first end of the first strand of the target polynucleotide, the motor protein (MP) binds to the leader, for example, by stalling at a stall moiety included in the leader, and a blocking moiety is present at the second end of the first strand. As shown, the double-stranded polynucleotide is symmetrical with a second strand identical to the first strand, but this is not required in the disclosed method. The leader is captured by a nanopore, and the first strand of the target polynucleotide moves through the nanopore, pushing the motor protein toward the second end of the first strand of the target polynucleotide. The motor protein then controls the movement of the first strand of the polynucleotide "out" of the pore. The motor protein then unbinds from the first strand of the target polynucleotide, and the first strand of the target polynucleotide moves back "in" the pore. The blocking moiety prevents the motor protein from disassociating from the first strand of the target polynucleotide. The motor protein then rebinds to the target polynucleotide and rereads the target polynucleotide (RR). [Figure 6a](a) Experimental schematic of Example 1 showing capture, "unstalling" and sequencing of both strands of a polynucleotide analyte, with occasional re-reading of the strands. Vs: sequencing potential, Vu: deblocking potential. The polarity of the applied potential is indicated by the arrow. The direction of the applied force is the same as the direction of the arrow. (A) Sequencing potential is applied (120 mV). Open pore, capture of polynucleotide analyte via the 3' leader. Separation of the two strands by the nanopore, template and complementary strands move to the trans compartment. (B) Polynucleotide reaches the enzyme, which is stalled at the spacer portion. The enzyme cannot move over the spacer portion. (C) Deblocking potential is applied (variable, 0 mV to -120 mV) so that the enzyme leaves the nanopore and is free to move over the spacer portion. (D) Sequencing potential is applied (120 mV). The nanopore translocates the polynucleotide until the enzyme reaches the nanopore, after which the enzyme controls the translocation of the polynucleotide from the nanopore. (E) The DNA motor translocates the template section and reaches the hairpin. (F) The DNA motor translocates the complementary section and the template and complementary strands refold in the cis compartment. The motor reaches the leader section and idles in the nanopore. State (F) pushes the enzyme from 3'-5' back to state (E), allowing strand rereading (RR). (G) A deblocking potential is applied, expelling the DNA motor and analyte from the nanopore. [Figure 6b] (b) Representative current-time trace from Example 1 showing an example of double readout of the polynucleotide enzyme via the enzyme being pushed back from the C3 leader under applied potential, zooming in on enzyme control sections (i) and (ii) and also identifying C3 levels. [Figure 6c] (c) Example of six representative re-reads from the experiment described in Example 1. The enzyme control moieties were mapped using an HMM model trained using data from the pore and enzyme combination used. The example reads shown map at least twice to the same strand of a mixture of seven restriction enzyme fragments of bacteriophage lambda DNA. [Figure 7]A non-limiting example of a system for re-reading a polynucleotide strand in a nanopore using a motor protein. A: Adapter for re-reading a polynucleotide strand. The adapter contains three oligonucleotide strands: a top strand (a), a blocker strand (b), and a back blocker strand (c). When hybridized together, these three strands resulted in an adapter with a 5' phosphate and a 3' T overhang that can be ligated to a dA-tailed double-stranded DNA. A motor protein (MP) was loaded onto the adapter. The top strand (a) contained a C3 spacer section (d) complementary to the blocker strand (e), a stall section (f) that stalls the motor protein, and a motor loading site (g). The blocker strand (b) contained a region complementary to (e) with a stall moiety, and a tether complementary arm (h). The back blocker strand (c) contained a region with partial complementarity to the top strand, and an arm containing a biotin-TEG moiety (i). In the control adaptor described in Example 5, this biotin-TEG moiety was absent. B: Motor protein-controlled translocation of a polynucleotide "out" of the nanopore with rereading of the polynucleotide. MP, motor protein; PN, polynucleotide. (a), the adaptor shown in A is ligated to a double-stranded dA-terminated polynucleotide. Adaptors were ligated to both ends of the polynucleotide. (b), the species described in (a) was bound to monovalent traptavidin, referred to as the "analyte molecule". The analyte molecule was added to the cis compartment of the nanopore sequencer and the membrane was biased at 180 mV (trans compartment positive relative to cis). The figure shows the stages of rereading: (i) the analyte molecule was captured at the cis opening of the nanopore via its 3' end. (ii) the blocker strand was removed from the analyte molecule and the analyte was translocated through the nanopore to the stalled motor protein.(ii) to (iii) the motor protein was pushed in the 3'-5' direction along the DNA (or "dropped back" as described in Example X1) until it reached either the distal (5') end where it encountered bound traptavidin or a point 5' from the end where the motor protein was loaded. (iii) to (iv) the motor protein controlled movement of the polynucleotide in the 5'-3' direction out of the nanopore. (iv) to (v) the motor protein continued to move until it reached its original starting point in (i) and (ii). The motor waited in the C3 spacer section of the leader. Then either: (v) to (iii) the motor protein was pushed in the 3'-5' direction according to (ii) to (iii) and the cycle was repeated, or (v) to (vi) the analyte molecule was ejected from the nanopore. The arrows indicate how the motor protein or polynucleotide moved compared to the previous step. [Figure 8] Electrical data related to Example 5 for either (A) a control adaptor ligated to a 3.6 kb polynucleotide analyte or (B) an adaptor containing a 5' biotin moiety. The electrical data show clear current levels for the open pore and the leader. Upon capture of the analyte from the open pore, an increase in current to the leader level was observed, followed by a "drop-back" signal (steps (ii)-(iii) in FIG. 7, B, marked with asterisks) resulting from the enzyme being extruded 3'-5' through the polynucleotide chain. In the example shown in (A), without the 5' biotin moiety, the motor was pushed 3'-5 back to some previous point in the chain, where it rebinded and controlled the translocation of the polynucleotide out of the nanopore, reaching the "leader" moiety, and then "dropped back" again to open the pore. This second "drop-back", beginning approximately 982 seconds later, fully extruded the motor from the 5' end of the polynucleotide chain, resulting in complete translocation of the polynucleotide chain through the nanopore. In (B), in the presence of the 5′ biotin-traptavidin complex, instead of pushing the motor out from the 5′ end upon “drop-back,” the motor controlled translocation out of the nanopore a second time. [Figure 9] Electrical data relevant to Example 6. The data show distinct current levels for the open pore and the reader. When voltage was applied after voltage flick 1, the current level quickly increased to the reader level, indicating that the strand was captured. After a pause, the enzyme was pushed 3'-5' through the polynucleotide strand ("dropped back"), thereby rebinding and controlling the translocation of the polynucleotide out of the nanopore until it reached the leader portion. This process was repeated a total of 14 times, as indicated by the asterisk. After 14 reads, the strand was ejected by reversal of the voltage during voltage flick 2. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0045] The present invention will be described with respect to certain embodiments and with reference to certain drawings, but the present invention is not limited thereto, but only by the claims. Any reference signs in the claims should not be construed as limiting the scope thereof. Of course, it should be understood that not necessarily all aspects or advantages can be achieved in accordance with any particular embodiment of the present invention. Thus, for example, a person skilled in the art will recognize that the present invention can be embodied or performed in a manner that achieves or optimizes one advantage or group of advantages taught herein, without necessarily achieving other aspects or advantages that may be taught or suggested herein.

[0046] The present invention, both as to its organization and method of operation, together with its features and advantages, may best be understood by reference to the following detailed description when read in conjunction with the accompanying drawings. Aspects and advantages of the present invention will be apparent and elucidated with reference to the embodiment(s) described below. Reference throughout this specification to "one embodiment" or "an embodiment" means that a particular feature, structure, or characteristic described in connection with that embodiment is included in at least one embodiment of the invention. Thus, the appearances of the phrases "in one embodiment" or "in an embodiment" in various places throughout this specification do not necessarily all refer to the same embodiment, although they may. Similarly, in describing exemplary embodiments of the present invention, it should be understood that various features of the invention may be grouped together in a single embodiment, figure, or description thereof in order to simplify the disclosure and aid in understanding one or more of the various inventive aspects. However, this method of disclosure should not be interpreted as reflecting an intention that the claimed invention requires more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive aspects have less than all features of a single foregoing disclosed embodiment.

[0047] It is to be understood that "embodiments" of the present disclosure may be specifically combined together, unless the context dictates otherwise. Any specific combination of the disclosed embodiments is a further disclosed embodiment of the claimed invention (unless the context implies otherwise).

[0048] Additionally, as used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural references unless the context clearly indicates otherwise. Thus, for example, reference to a "polynucleotide" includes two or more polynucleotides, reference to a "motor protein" includes two or more such proteins, reference to a "helicase" includes two or more helicases, reference to a "monomer" refers to two or more monomers, and reference to a "pore" includes two or more pores.

[0049] All publications, patents, and patent applications cited herein, whether supra or infra, are hereby incorporated by reference in their entirety.

[0050] definition When an indefinite or definite article is used when referring to a singular noun, such as "a" or "an" or "the," this includes the plural of that noun, unless otherwise specified. When the term "comprises" is used in the present description and claims, it does not exclude other elements or steps. Furthermore, terms such as first, second, third, etc. in the present description and claims are used to distinguish between similar elements and are not necessarily used to describe an order or chronological order that occurred. It is to be understood that the terms used in this manner are interchangeable under appropriate circumstances and that the embodiments of the invention described herein can operate in other sequences than those described or illustrated herein. The following terms or definitions are provided solely to aid in the understanding of the present invention. Unless specifically defined herein, all terms used herein have the same meaning as would be understood by one of ordinary skill in the art of the invention. Skilled artisans should refer to the following publications for definitions and technical terms, particularly those referenced in Sambrook et al., Molecular Cloning: A Laboratory Manual, 4 th ed., Cold Spring Harbor Press, Plainsview, New York (2012), and Ausubel et al., Current Protocols in Molecular Biology (Supplement 114), John Wiley & Sons, New York (2016). The definitions provided herein should not be construed to be narrower than understood by one of ordinary skill in the art.

[0051] As used herein, "about" when referring to a measurable value, such as an amount, temporal duration, and the like, is meant to encompass variations of ±20% or ±10%, more preferably ±5%, even more preferably ±1%, and still more preferably ±0.1% from the specified value, where such variations are appropriate for performing the disclosed methods.

[0052] As used herein, a "nucleotide sequence," "DNA sequence," or "nucleic acid molecule(s)" refers to a polymeric form of nucleotides of any length, either ribonucleotides or deoxyribonucleotides. The term refers only to the primary structure of the molecule. Thus, the term includes double- and single-stranded DNA and RNA. The term "nucleic acid" as used herein is a single- or double-stranded covalently linked nucleotide sequence in which the 3' and 5' ends of each nucleotide are linked by a phosphodiester bond. A polynucleotide may be composed of deoxyribonucleotide or ribonucleotide bases. Nucleic acids may be synthetically produced in vitro or isolated from natural sources. Nucleic acids may further include modified DNA or RNA, e.g., DNA or RNA that is methylated, or RNA that has been subjected to post-translational modifications, e.g., 5'-capping with 7-methylguanosine, 3'-processing such as cleavage and polyadenylation, and splicing. Nucleic acids may also include synthetic nucleic acids (XNAs), such as hexitol nucleic acids (HNAs), cyclohexene nucleic acids (CeNAs), threose nucleic acids (TNAs), glycerol nucleic acids (GNAs), locked nucleic acids (LNAs), and peptide nucleic acids (PNAs). The size of a nucleic acid, also referred to herein as a "polynucleotide", is typically expressed in terms of the number of base pairs (bp) for a double-stranded polynucleotide, or the number of nucleotides (nt) for a single-stranded polynucleotide. 1000 bp or nt is equivalent to a kilobase (kb). Polynucleotides less than about 40 nucleotides in length are typically referred to as "oligonucleotides" and may include primers for use in manipulating DNA, such as via polymerase chain reaction (PCR).

[0053] The term "amino acid" in the context of this disclosure is used in its broadest sense and is intended to include organic compounds containing amine (NH2) and carboxyl (COOH) functional groups, along with a side chain (e.g., R group) specific to each amino acid. In some embodiments, amino acid refers to naturally occurring L α-amino acids or residues. The commonly used one-letter and three-letter abbreviations for naturally occurring amino acids are used herein: A=Ala, C=Cys, D=Asp, E=Glu, F=Phe, G=Gly, H=His, I=Ile, K=Lys, L=Leu, M=Met, N=Asn, P=Pro, Q=Gln, R=Arg, S=Ser, T=Thr, V=Val, W=Trp, and Y=Tyr (Lehninger, AL, (1975) Biochemistry, 2d ed., pp. 71-92, Worth Publishers, New York). The general term "amino acid" further includes D-amino acids, retro-inverso amino acids, and chemically modified amino acids such as amino acid analogs, naturally occurring amino acids that are not normally incorporated into proteins such as norleucine, and chemically synthesized compounds that have properties known in the art that characterize amino acids such as β-amino acids. For example, analogs or mimetics of phenylalanine or proline that allow the same conformational constraints on peptide compounds as natural Phe or Pro are included within the definition of amino acids. Such analogs and mimetics are referred to herein as "functional equivalents" of the respective amino acids. Other examples of amino acids are listed by Roberts and Vellaccio, The Peptides: Analysis, Synthesis, Biology, Gross and Meiehofer, eds., Vol. 5 p. 341, Academic Press, Inc., NY 1983, which is incorporated herein by reference.

[0054] The terms "leader" and "leader sequence" are used interchangeably herein.

[0055] The terms "polypeptide" and "peptide" are used interchangeably herein to refer to a polymer of amino acid residues, as well as variants and synthetic analogs thereof. Thus, these terms apply to amino acid polymers in which one or more amino acid residues are non-naturally occurring synthetic amino acids, such as chemical analogs of the corresponding naturally occurring amino acids, as well as to naturally occurring amino acid polymers. Polypeptides may also undergo maturation or post-translational modification processes, which may include, but are not limited to, glycosylation, proteolytic cleavage, lipidation, signal peptide cleavage, propeptide cleavage, phosphorylation, and the like. Peptides may be produced using recombinant techniques, for example, by expression of recombinant or synthetic polynucleotides. Recombinantly produced peptides are typically substantially free of culture medium, e.g., culture medium represents less than about 20% of the volume of the protein preparation, more preferably less than about 10%, and most preferably less than about 5%.

[0056] The term "protein" is used to describe a folded polypeptide having secondary or tertiary structure. A protein may be composed of a single polypeptide or may include multiple polypeptides that assemble to form a multimer. A multimer may be a homo- or hetero-oligomer. A protein may be a naturally occurring or wild-type protein, or a modified or non-naturally occurring protein. A protein may differ from a wild-type protein, for example, by the addition, substitution, or deletion of one or more amino acids.

[0057] A "variant" of a protein includes peptides, oligopeptides, polypeptides, proteins, and enzymes that have amino acid substitutions, deletions, and / or insertions compared to the unmodified or wild-type protein in question and have biological and functional activities similar to those of the unmodified protein from which they are derived. As used herein, the term "amino acid identity" refers to the degree to which sequences are identical amino acid-by-amino acid over a comparison window. Thus, "percentage of sequence identity" is calculated by comparing two optimally aligned sequences over a comparison window, determining the number of positions at which identical amino acid residues (e.g., Ala, Pro, Ser, Thr, Gly, Val, Leu, Ile, Phe, Tyr, Trp, Lys, Arg, His, Asp, Glu, Asn, Gln, Cys, and Met) occur in both sequences to calculate the number of matched positions, dividing the number of matched positions by the total number of positions in the comparison window (i.e., the window size), and multiplying the result by 100 to calculate the percentage of sequence identity.

[0058] For all aspects and embodiments of the invention, a "variant" has at least 50%, 60%, 70%, 80%, 90%, 95%, or 99% complete sequence identity to the amino acid sequence of the corresponding wild-type protein. Sequence identity may also be to a fragment or portion of a full-length polynucleotide or polypeptide. Thus, while a sequence may have only 50% sequence identity overall with a full-length reference sequence, the sequence of a particular region, domain, or subunit may share 80%, 90%, or even 99% sequence identity with the reference sequence.

[0059] The term "wild type" refers to a gene or gene product isolated from a naturally occurring source. A wild type gene is that which is most frequently observed in a population and is therefore the arbitrarily designed "normal" or "wild type" form of that gene. In contrast, the terms "modified," "mutant," or "variant" refer to a gene or gene product that exhibits sequence modifications (e.g., substitutions, truncations, or insertions), post-translational modifications, and / or functional properties (e.g., altered characteristics) compared to the wild type gene or gene product. It should be noted that naturally occurring mutants can be isolated, which are identified by the fact that they have altered characteristics compared to the wild type gene or gene product. Methods for introducing or substituting naturally occurring amino acids are well known in the art. For example, methionine (M) can be substituted with arginine (R) by replacing the codon for methionine (ATG) with the codon for arginine (CGT) at the relevant position in the polynucleotide encoding the mutant monomer. Methods for introducing or substituting non-naturally occurring amino acids are also well known in the art. For example, non-naturally occurring amino acids can be introduced by including synthetic aminoacyl-tRNA in the IVTT system used to express the mutant monomers. Alternatively, they can be introduced by expressing mutant monomers in E. coli that are auxotrophic for specific amino acids in the presence of synthetic (i.e., non-naturally occurring) analogs of those specific amino acids. They can also be generated by naked ligation when the mutant monomers are produced using partial peptide synthesis. Conservative substitutions replace amino acids with other amino acids of similar chemical structure, similar chemical properties, or similar side chain volume. The introduced amino acids can have similar polarity, hydrophilicity, hydrophobicity, basic, acidic, neutral, or charged properties as the amino acids they replace. Alternatively, conservative substitutions can introduce another amino acid that is aromatic or aliphatic in place of an existing aromatic or aliphatic amino acid. Conservative amino acid changes are well known in the art and can be selected according to the properties of the 20 main amino acids defined in Table 1 below.If the amino acids have similar polarity, this can also be determined by reference to the hydrophobicity scale of the amino acid side chains in Table 2. [Table 1] [Table 2]

[0060] The mutant or modified protein, monomer, or peptide can also be chemically modified in any manner and at any site. The mutant or modified monomer or peptide is preferably chemically modified by attachment of a molecule to one or more cysteines (cysteine ​​bond), attachment of a molecule to one or more lysines, attachment of a molecule to one or more unnatural amino acids, enzymatic modification of an epitope, or modification of a terminal. Suitable methods for carrying out such modifications are well known in the art. The mutant of the modified protein, monomer, or peptide can be chemically modified by attachment of any molecule. For example, the mutant of the modified protein, monomer, or peptide can be chemically modified by attachment of a dye or fluorophore.

[0061] As used herein, an alkylene group is an unsubstituted or substituted, saturated, bidentate moiety obtained by removing two hydrogen atoms, either two from the same carbon atom, or one from each of two different carbon atoms, of a hydrocarbon compound, which may be aliphatic or alicyclic. The hydrocarbon compound may have 1 to 20 carbon atoms, in which case the alkylene group is C 1-20 The alkylene group is C 1-10 When it is alkylene, it can have, for example, 1 to 10 carbon atoms. Typically, C 1-6 Alkylene or C 1-4 Alkylene is, for example, methylene, ethylene, i-propylene, n-propylene, t-butylene, s-butylene or n-butylene.

[0062] An alkenylene group is an unsubstituted or substituted bidentate moiety obtained by removing two hydrogen atoms from the same carbon atom or one hydrogen atom from each of two different carbon atoms of a hydrocarbon compound, which may be aliphatic or alicyclic, and contains one or more carbon-carbon double bonds. The hydrocarbon compound may have from 2 to 20 carbon atoms, in which case the alkenylene group is C 2-20 Alkenylene. The alkenylene group is C 2-10 When it is alkenylene, it may have, for example, 2 to 10 carbon atoms. Typically, it is C 2-6 Alkenylene or C 2-4 It is alkenylene.

[0063] An alkynylene group is an unsubstituted or substituted bidentate moiety obtained by removing two hydrogen atoms from the same carbon atom or one hydrogen atom from each of two different carbon atoms of a hydrocarbon compound, which may be aliphatic or alicyclic, and contains one or more carbon-carbon triple bonds. The hydrocarbon compound may have from 2 to 20 carbon atoms, in which case the alkynylene group is C 2-20 Alkynylene. The alkynylene group is C 2-10 When it is alkynylene, it may have, for example, 2 to 10 carbon atoms. Typically, it is C 2-6 Alkynylene or C 2-4 It is alkynylene.

[0064] An arylene group is an unsubstituted or substituted monocyclic or fused polycyclic bidentate moiety obtained by removing two hydrogen atoms, one from each of two different aromatic ring atoms of an aromatic compound, which moiety has 5 to 14 ring atoms (unless otherwise specified). Typically, each ring has 5 to 7 or 5 to 6 ring atoms. An arylene group may be unsubstituted or substituted.

[0065] Heteroarylene groups are bidentate moieties obtained by removing two hydrogen atoms, one from each of two different ring atoms of a heteroaryl group. Heteroaryl groups are substituted or unsubstituted monocyclic or fused polycyclic (e.g., bicyclic or tricyclic) aromatic groups, typically containing 5 to 14 ring atoms, including at least one heteroatom in the ring portion, e.g., 1, 2 or 3 heteroatoms selected from O, S, N, P, Se and Si, more typically O, S and N. Examples include pyridyl, pyrazinyl, pyrimidinyl, pyridazinyl, furanyl, thienyl, pyrazolidinyl, pyrrolyl, oxadiazolyl, isoxazolyl, thiadiazolyl, thiazolyl, imidazolyl, triazolyl, pyrazolyl, oxazolyl, isothiazolyl, benzofuranyl, isobenzofuranyl, benzothiophenyl, indolyl, indazolyl, carbazolyl, acridinyl, urinyl, cinnolinyl, quinoxalinyl, naphthyridinyl, benzimidazolyl, benzoxazolyl, quinolinyl, quinazolinyl and isoquinolinyl.

[0066] A carbocyclylene group, also known as a cycloalkylene group, is a bidentate moiety obtained by removing two hydrogen atoms, one from each of two carbon atoms of an unsubstituted or substituted cyclic alkyl group. Typically, the moiety contains 3 to 10 ring atoms and has 3 to 10 carbon atoms (unless otherwise specified). Examples include cyclopropane (C3), cyclobutane (C4), cyclopentane (C5), cyclohexane (C6), cycloheptane (C7), methylcyclopropane (C4), dimethylcyclopropane (C5), methylcyclobutane (C5), dimethylcyclobutane (C6), methylcyclopentane (C6), dimethylcyclopentane (C7), methylcyclohexane (C7), dimethylcyclohexane (C8), and menthane (C10).

[0067] A heterocyclylene moiety is a bidentate moiety obtained by removing two hydrogen atoms from two different ring atoms of a heterocyclyl group. Heterocyclyl groups are unsubstituted or substituted cyclic groups, typically containing 5 to 14 atoms, including at least one heteroatom in the ring portion, e.g., 1, 2 or 3 heteroatoms selected from O, S, N, P, Se and Si, more typically O, S and N. Examples include piperazine, piperidine, morpholine, 1,3-oxazinane, pyrrolidine, imidazolidine, oxazolidine, tetrahydropyrazine, tetrahydropyridine, dihydro-1,4-oxazine, tetrahydropyrimidine, dihydro-1,3-oxazine, dihydropyrrole, dihydroimidazole and dihydrooxazole groups.

[0068] An arylene-alkylene group is a group formed by forming a bond between an arylene group and an alkylene group as defined herein. A heteroarylene-alkylene group is a group formed by forming a bond between a heteroarylene group and an alkylene group as defined herein. A carbocyclylene-alkylene group is a group formed by forming a bond between a carbocyclylene group and an alkylene group as defined herein. A heterocyclylene-alkylene group is a group formed by forming a bond between a heterocyclylene group and an alkylene group as defined herein.

[0069] When a group is described as substituted, it is typically substituted by one or more, such as one, two or three, typically one or two, usually one substituent. Suitable substituents may be independently selected from halogen, -OR' and -NR'2, where R' is typically H or unsubstituted C 1-2 alkyl, and unsubstituted C1-C2 alkyl).

[0070] Methods for characterizing an analyte The present disclosure relates to a method for characterizing a target polynucleotide as it moves relative to a detector having a first opening and a second opening, or a detector contained in a structure having a first opening and a second opening, such as a nanopore. Although the present disclosure provides a nanopore as an exemplary detector, the methods provided herein are suitable for detectors such as (i) zero mode waveguides, (ii) field effect transistors, optionally nowire field effect transistors, (iii) AFM tips, (iv) nanotubes, optionally carbon nanotubes, and (v) nanopores. The disclosed methods are particularly suitable for methods in which a polynucleotide moves through a detector or through a structure that includes a detector, such as a well in a detector chip.

[0071] Any suitable polynucleotide can be characterized using the methods disclosed herein. Polynucleotides that can be characterized according to the disclosed methods are described in more detail herein.

[0072] The movement of the target polynucleotide is controlled by using a motor protein. Any suitable motor protein can be used in the methods provided herein. Exemplary motor proteins are described in more detail herein. The method includes re-reading the polynucleotide, such as moving the polynucleotide back and forth relative to the detector. Thus, the disclosed method includes "flossing" the target polynucleotide relative to the detector. This is described in more detail herein.

[0073] In the disclosed methods, the target polynucleotide has a leader attached, which is described in more detail herein.

[0074] First, in the disclosed method, the motor protein binds to the leader at a polynucleotide binding site of the motor protein, which is described in more detail herein. The motor protein can be stalled on the leader at a stall moiety. Suitable stall moieties are described in more detail herein. Stalling the motor protein on the polynucleotide has various advantages. For example, while stalled, the motor protein typically consumes less fuel than when it is not stalled, e.g., when it is moving freely relative to the polynucleotide. This reduction in non-productive fuel usage can be advantageous.

[0075] The methods provided herein typically include destalling the motor protein so that the motor protein can control the movement of the polynucleotide relative to the detector, as described in more detail herein. Methods of destalling the motor protein are described in more detail herein. Controlled destalling of the motor protein has various advantages, including being able to precisely determine the point at which the motor protein starts processing the polynucleotide. This can be useful, for example, to characterize the polynucleotide so that data is not lost as a result of undesired movement of the motor protein on the polynucleotide before the start of data recording.

[0076] The motor protein is used to control movement of the target polynucleotide in a first direction relative to the detector, the first direction being from the second opening of the detector to the first opening of the detector, and the method characteristic of the target polynucleotide occurs as the target polynucleotide moves in the first direction relative to the detector.

[0077] The target polynucleotide then disassociates from the polynucleotide binding site of the motor protein, as described in more detail herein below. Once the target polynucleotide disassociates from the polynucleotide binding site of the motor protein, the target polynucleotide moves in a second direction relative to the detector.

[0078] As discussed herein, the second direction is opposite to the first direction. For example, in embodiments in which the detector is a nanopore, the first direction can be "outside" the nanopore (from the "point of view" of the motor protein) and the second direction can be "inside" the nanopore (from the "point of view" of the motor protein). This is described in more detail herein.

[0079] The target polynucleotide then rebinds to the polynucleotide binding site of the motor protein. The target polynucleotide rebinds to the same motor protein. In other words, the target polynucleotide rebinds to the same molecule of the motor protein, not just to different molecules of the same type of motor protein. The polynucleotide handling protein is again used to control the movement of the target polynucleotide in a first direction relative to the detector. When the conjugate moves in the first direction relative to the detector, a further measurement characteristic of the target polynucleotide is made. The first direction is the same as the first direction described above.

[0080] Thus, disclosed herein is a method for characterizing a leader-bound target polynucleotide, the method comprising: (i) contacting a detector having a first opening and a second opening, or contained in a structure having a first opening and a second opening, with a reader under conditions such that the first opening is in contact with the reader and the target polynucleotide migrates in a direction from the first opening to the second opening, contacting the leader with a motor protein bound at a polynucleotide binding site of the motor protein; (ii) performing one or more measurements characteristic of the target polynucleotide when the motor protein controls movement of the target polynucleotide in a first direction relative to the detector, the first direction being from the second opening to the first opening; (iii) debinding the target polynucleotide from the polynucleotide binding site of the motor protein such that the target polynucleotide moves in a second direction relative to the detector, the second direction being from the first opening to the second opening; (iv) allowing the target polynucleotide to rebind to the polynucleotide binding site of the motor protein and performing one or more measurements characteristic of the target polynucleotide when the motor protein controls movement of the target polynucleotide in a first direction relative to a detector; and characterizing the target polynucleotide. Characterizing the target polynucleotide can include, for example, determining the sequence of the target polynucleotide.

[0081] In some embodiments, steps (iii) and (iv) of the disclosed method are repeated multiple times by continuously binding and rebinding the motor protein to the target polynucleotide. In this way, the target polynucleotide can be oscillated relative to the detector (i.e., "flossed" relative to the first and second openings of the detector). This "flossing" allows the target polynucleotide to be repeatedly characterized. In some embodiments, this can increase the accuracy of the characterization information.

[0082] The disclosed method is based at least in part on the realization that data obtained when a polynucleotide is moved away from a detector, such as a nanopore, may differ from data obtained when the same polynucleotide is moved into a detector (e.g., a nanopore). Data characteristics, including signal profile, noise profile, and error profile, may all be different in some embodiments from a contrasting method in which the same polynucleotide is moved into a detector, such as a nanopore. In some embodiments, data obtained with the disclosed method has advantages compared to data obtained with other known methods. Thus, the disclosed method increases the options available when characterization of a polynucleotide is required. Thus, a user wishing to characterize a polynucleotide can select the method that best suits the particular application at hand. Measurements made when a target polynucleotide is moved in a first direction may, in some embodiments, be combined or compared to improve characterization of the polypeptide.

[0083] The disclosed method has many advantages over previously known methods. For example, each read of the target polynucleotide should be of comparable accuracy since the same strand and the same detection moiety are used. This allows the same base calling model to be used for each read. It also facilitates the combination of data from multiple reads. Furthermore, since the native sequence is re-read multiple times, it is possible to preserve (for example) epigenetic information. The method is also adaptive, and the re-reading can be repeated multiple times until the required accuracy of the data is obtained.

[0084] Characterization of target polynucleotides As discussed above, the target polynucleotide is characterized using a detector.

[0085] The target polynucleotide interacts with the detector. For example, the target polynucleotide can pass through the first and second openings of the detector. A leader is attached to the target polynucleotide. The leader can promote the target polynucleotide to pass through the first and second openings. When the target polynucleotide moves relative to the detector, the target polynucleotide is characterized.

[0086] In some embodiments, the target polynucleotide has a first end and a second end, the leader is attached to the first end, and the motor protein is oriented in an orientation to process the target polynucleotide in a direction from the second end toward the first end. Thus, the motor protein moves relative to the polynucleotide in a direction toward the second end of the polynucleotide, and the polynucleotide moves relative to the motor protein in a direction from the second end toward the first end.

[0087] In the example shown in FIG. 1, the motor protein (MP) moves the target polynucleotide (PN) in a first direction from the "viewpoint" of the motor protein to "out" of the first opening of the detector. When debinding from the motor protein, the target polynucleotide may move in a second direction from the "viewpoint" of the polynucleotide handling protein to "in" of the first opening of the detector. The difference in direction of movement of the polynucleotide out of a detector such as a nanopore in the methods provided herein, as opposed to movement of the polynucleotide into the pore in the contrasting methods, is shown diagrammatically in FIG. 1.

[0088] As will be understood by those skilled in the art, the designation "out" refers to the overall movement of the polynucleotide towards the motor protein. This direction of movement can be contrasted with another mode in which the target polynucleotide is moved "in" the first opening of the detector by the motor protein. The difference between these kinetic schemes is profound. In the methods provided herein in which the polynucleotide is moved "out" of the detector (e.g., out of the nanopore), the direction of movement of the target polynucleotide is from the entrance of the detector that is furthest from the motor protein (i.e., the distal entrance) towards the entrance of the detector that is closest to the motor protein (the proximal entrance). In the contrasting method in which the polynucleotide is moved "in" the detector (e.g., into the nanopore), the direction of movement of the target polynucleotide is from the entrance of the nanopore that is closest to the motor protein (the proximal entrance) towards the entrance of the detector that is furthest from the motor protein (the distal entrance).

[0089] For example, in some embodiments, the detector may be or include a nanopore. When the detector is or includes a nanopore, the first and second openings may be referred to as the cis and trans openings of the nanopore. Often, the first opening is a cis opening and the second opening is a trans opening, but in some embodiments, the first opening is a trans opening and the second opening is a cis opening, respectively. The designation of "cis" and "trans" openings of a nanopore is routine in the art. For example, the cis opening of a nanopore typically faces the cis chamber of a nanopore device, such as the devices described herein having cis and trans chambers, and the trans opening typically faces the trans chamber.

[0090] Thus, in some embodiments of the provided methods, the nanopore spans a membrane having a cis side and a trans side, with a first opening of the nanopore on the cis side of the membrane and a second opening of the nanopore on the trans side. In such embodiments, the motor protein is located on the cis side of the membrane and controls movement of the target polynucleotide through the nanopore in a direction from the trans side to the cis side of the membrane, i.e., from the trans side to the cis side of the nanopore.

[0091] In other embodiments of the provided methods, the nanopore spans a membrane having a cis side and a trans side, with a first opening of the nanopore on the trans side of the membrane and a second opening of the nanopore on the cis side. In such embodiments, the motor protein is located on the trans side of the membrane and controls movement of the target polynucleotide through the nanopore in the cis to trans direction of the membrane, i.e., from the cis to the trans side of the nanopore.

[0092] More specifically, the method includes performing one or more measurements characteristic of the target polynucleotide when the motor protein controls the movement of the target polynucleotide in a first direction relative to the detector. The first direction can be a direction in which the motor protein drives the movement of the polynucleotide. The first direction can be a direction of a force applied to the detector. The first direction can be a direction opposite to the direction of the force applied across the detector.

[0093] In many cases, the detector is included in a structure having a first opening and a second opening, or comprises a transmembrane nanopore having a first opening and a second opening, and step (i) comprises contracting the first opening with a leader attached to the target polynucleotide. Typically, the motor protein controls the movement of the target polynucleotide in a direction from the second opening to the first opening. Typically, when the target polynucleotide debinds from the polynucleotide binding site of the motor protein, the target polynucleotide moves in a direction from the first opening to the second opening.

[0094] Typically, when the detector is or includes a nanopore, the first direction is toward the "out" of the nanopore described herein. Thus, in some embodiments, the movement of the polynucleotide during one or more measurements is toward the out of the nanopore. In some embodiments, the nanopore spans a membrane having a cis side and a trans side, a first opening of the nanopore is on the cis side of the membrane, and a second opening of the nanopore is on the trans side, and the motor protein is located on the cis side of the membrane and controls the movement of the target polynucleotide through the nanopore from the trans side to the cis side of the membrane, i.e., from the trans side to the cis side of the nanopore. In such embodiments, the motor protein controls the movement of the target polynucleotide through the nanopore in a direction toward the "out" of the nanopore (from the "perspective" of the motor protein). Upon debinding from the motor protein, the target polynucleotide moves in a second direction relative to the detector. Thus, movement of the target polynucleotide in the second direction is "into" the nanopore (from the "perspective" of the motor protein), i.e., movement of the polynucleotide through the nanopore from the cis to the trans side of the membrane, i.e., from the cis to the trans side of the nanopore. This is shown in FIG.

[0095] Of course, the opposite setup can also be used in the methods disclosed herein. For example, as shown, the detector can include a transmembrane nanopore spanning a membrane having a cis side and a trans side. In some embodiments, a first opening of the nanopore is on the trans side of the membrane and a second opening of the nanopore is on the cis side, and a motor protein is located on the trans side of the membrane and controls the movement of the target polynucleotide through the nanopore from the cis side to the trans side of the membrane, i.e., from the cis side to the trans side of the nanopore. In such embodiments, the motor protein controls the movement of the target polynucleotide through the nanopore in a direction "out" of the nanopore (from the "perspective" of the motor protein). Upon debinding from the motor protein, the target polynucleotide moves in a second direction relative to the detector. Movement of the target polynucleotide in the second direction is therefore the movement of the polynucleotide through the nanopore in a direction "in" of the nanopore (from the "perspective" of the motor protein), i.e., from the trans side to the cis side of the membrane, i.e., from the trans side to the cis side of the nanopore.

[0096] In other embodiments, the first direction can be a direction "into" a detector, e.g., a nanopore as described herein. Thus, in some embodiments, the detector may include a transmembrane nanopore spanning a membrane having a cis side and a trans side, and the movement of the polynucleotide in the first direction is into the nanopore. In some embodiments, the nanopore spans a membrane having a cis side and a trans side, a first opening of the nanopore is on the cis side of the membrane, and a second opening of the nanopore is on the trans side, and the motor protein is located on the cis side of the membrane and controls the movement of the target polynucleotide through the nanopore from the cis side to the trans side of the membrane, i.e., from the cis side to the trans side of the nanopore. In such embodiments, the motor protein controls the movement of the target polynucleotide through the nanopore in a direction "into" the nanopore (from the "perspective" of the motor protein). Upon debinding from the motor protein, the target polynucleotide moves in a second direction relative to the detector. Thus, movement of the target polynucleotide in the second direction is movement of the polynucleotide through the nanopore in a direction "out" of the nanopore (from the "perspective" of the motor protein), i.e., from the trans side to the cis side of the membrane, i.e., from the trans side to the cis side of the nanopore. In other embodiments, the nanopore spans a membrane having a cis side and a trans side, a first opening of the nanopore is on the trans side of the membrane and a second opening of the nanopore is on the cis side, and the motor protein is located on the trans side of the membrane and controls movement of the target polynucleotide through the nanopore from the trans side to the cis side of the membrane, i.e., from the trans side to the cis side of the nanopore. In such embodiments, the motor protein controls movement of the target polynucleotide through the nanopore in a direction "in" of the nanopore (from the "perspective" of the motor protein). Upon debinding from the motor protein, the target polynucleotide moves in a second direction relative to the detector. Thus, movement of the target polynucleotide in the second direction is movement of the polynucleotide through the nanopore in a direction "out" of the nanopore (from the "perspective" of the motor protein), i.e., from the cis to the trans side of the membrane, i.e., from the cis to the trans side of the nanopore.

[0097] Typically, the first direction of movement of the polynucleotide relative to the first and second openings of the detector is "out" of the first opening, ie, from the second opening to the first opening.

[0098] In some embodiments, the method includes applying a force (e.g., a potential) to a first opening and a second opening of the detector, e.g., across the nanopore, and the motor protein controls movement of the target polynucleotide through the nanopore in a direction opposite the applied force.

[0099] It is important to distinguish the movement of the polynucleotide in the second direction relative to the detector from spontaneous slipping that may occur. For example, slipping of one or two bases is not an example of re-reading as described herein. Typically, in step (iii), the distance that the target polynucleotide moves relative to the detector is at least 10 nucleotides long. In some embodiments, the distance that the target polynucleotide moves relative to the detector is at least 20 nucleotides long, such as at least 30 nucleotides long, such as at least 40 nucleotides long, such as at least 50 nucleotides long, such as at least 100 nucleotides long. Longer distances may be used. In some embodiments, the distance that the target polynucleotide moves relative to the detector in step (iii) is at least 1000 nucleotides (1 kb) long, such as at least 2 kb long, such as at least 5 kb or at least 10 kb long, such as at least 100 kb or at least 1000 kb long.

[0100] Steps (iii) and (iv) of the method can be repeated multiple times to re-read the target polynucleotide multiple times.Steps (iii) and (iv) can be repeated at least once, for example at least twice, for example at least three times, for example at least four times, for example at least five times, for example at least 10 times, for example at least 20 times, for example at least 50 times, for example at least 100 times, for example at least 1000 times, for example at least 10,000 times, for example at least 100,000 times or more.Thus, the method can include "flossing" the polynucleotide back and forth to the detector.

[0101] Thus, if steps (iii) and (iv) are repeated once (and only once), such that the method comprises steps (iii) and (iv) twice and only twice, then the method will comprise steps (i), (ii), (iii), (iv), (iii1), and (iv1), and measurements are made that are characteristic of three portions of the polynucleotide: a first portion in step (ii), a second portion in steps (iii) and (iv), and a third portion in steps (iii1) and (iv1). If steps (iii) and (iv) are repeated twice (and only twice) such that the method comprises steps (iii) and (iv) three times and only three times, then the method comprises steps (i), (ii), (iii), (iv), (iii1), (iv1), (iii2) and (iv2) and measurements characteristic of four portions of the polynucleotide are made, namely the first portion in step (ii), the second portion in steps (iii) and (iv), the third portion in steps (iii1) and (iv1) and the fourth portion in steps (iii2) and (iv2). In other words, if steps (iii) and (iv) are repeated n times, then each repetition will provide measurements characteristic of (n+2) portions of the polynucleotide. Repeating steps (iii) and (iv) multiple times may lead to improved characterization since the portion of the polynucleotide analyzed by the nanopore is sampled multiple times and any random errors that may be recorded in the analysis will not have statistical significance. Thus, the accuracy of the characterization data obtained in this way may be improved. These methods allow very high accuracy levels to be reached, for example at least 99% accuracy, at least 99.9% accuracy, or at least 99.99% accuracy. Thus, in some embodiments, steps (iii) and (iv) are repeated until an accuracy level of at least 99% is reached, such as at least 99.9% accuracy, or at least 99.99% accuracy.

[0102] The portion of the polynucleotide read in step (ii) and the portion of the polynucleotide read in step (iv) of the method typically overlap. In other words, the method includes re-reading at least a portion of the polynucleotide multiple times. Thus, in some embodiments, in step (ii), the motor protein controls the movement of a first portion of the target polynucleotide in a first direction relative to the detector, and in step (iv), the motor protein controls the movement of a second portion of the target polynucleotide in a first direction relative to the detector, the first portion at least partially overlapping with the second portion. In some embodiments, the second portion overlaps with at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% of the first portion. In some embodiments, the first portion is the same as the second portion. Thus, in some embodiments, a portion of the polynucleotide is repeatedly characterized in the provided method. In each iteration, if the second portion of the polynucleotide partially but not completely overlaps with the first portion of the polynucleotide of the previous iteration, the polynucleotide is ratcheted in a zigzag manner relative to the detector. In each iteration, if the second portion of the polynucleotide completely overlaps with the first portion of the polynucleotide of the previous iteration, the same portion of the polynucleotide is flossed back and forth relative to the detector.

[0103] However, in some embodiments, the portion of the polynucleotide read in step (ii) and the portion of the polynucleotide read in step (iv) of the method do not overlap. In other words, the method includes re-reading at least a portion of the polynucleotide multiple times. Thus, in some embodiments, in step (ii), the motor protein controls the movement of a first portion of the target polynucleotide in a first direction relative to the detector, and in step (iv), the motor protein controls the movement of a second portion of the target polynucleotide in a first direction relative to the detector, where the first portion does not overlap with the second portion.

[0104] In some embodiments, the distance that the target polynucleotide first travels relative to the detector in step (i) is at least 1000 nucleotides long. In some embodiments, the distance that the target polynucleotide travels relative to the detector in step (i) is at least 2000 nucleotides long, such as at least 5000 nucleotides long, such as at least 10,000 nucleotides long, such as at least 15000 nucleotides long, such as at least 20,000 nucleotides long, such as at least 25,000 nucleotides long, such as at least 30,000 nucleotides long, such as at least 50,000 nucleotides long, or more, such as at least 100,000 nucleotides long, or even more. This increase in initial "drop back" is a particular advantage of the disclosed method, for example, in some embodiments, it can allow for extended initial reading of the target polynucleotide.

[0105] In some embodiments, the distance the target polynucleotide travels relative to the detector in step (iii) is greater than the distance the polynucleotide travels relative to the detector in step (ii). In some embodiments, the distance the target polynucleotide travels relative to the detector in step (iii) is greater than the distance the polynucleotide travels relative to the detector in step (iv). In some embodiments, the distance the target polynucleotide travels relative to the detector in step (iii) is greater than the distance the polynucleotide travels relative to the detector in step (ii) and / or step (iv).

[0106] In some embodiments, the distance that the target polynucleotide travels relative to the detector in step (iii) is at least 1000 nucleotides in length. In some embodiments, the distance that the target polynucleotide travels relative to the detector in step (iii) is at least 2000 nucleotides in length, such as at least 5000 nucleotides in length, such as at least 10,000 nucleotides in length, such as at least 15000 nucleotides in length, such as at least 20,000 nucleotides in length, such as at least 25,000 nucleotides in length, such as at least 30,000 nucleotides in length, such as at least 50,000 nucleotides in length, or more, such as at least 100,000 nucleotides in length.

[0107] In some embodiments, the distance traveled by the target polynucleotide in steps (ii) and / or (iv) relative to the detector is each independently at least 100 nucleotides in length. In some embodiments, the distance traveled by the target polynucleotide in steps (ii) and / or (iv) relative to the detector is each independently at least 200 nucleotides in length, such as at least 500 nucleotides in length, such as at least 1,000 nucleotides in length, such as at least 2000 nucleotides in length, such as at least 5,000 nucleotides in length, such as at least 10,000 nucleotides in length.

[0108] Forces applied during movement In some embodiments of the disclosed methods, a force may be applied across the detector, e.g., across the nanopore. The force may be controlled to control the method. For example, increasing the force may increase or decrease the movement of the polynucleotide relative to the detector (e.g., nanopore), e.g., to control the rate at which the polynucleotide passes through the pore.

[0109] Any suitable force can be applied in the methods provided herein. The force can be a potential applied across the first and second openings, e.g., if the detector is or includes a nanopore, the force can be a potential applied across the nanopore. In some embodiments, no external force is applied across the first and second openings of the detector (e.g., in some embodiments, if the detector is or includes a nanopore, no external force may be applied across the nanopore). For example, in some embodiments, no potential is applied. Such embodiments are particularly suited in some embodiments to methods in which optical measurements are made as the polynucleotide moves relative to the nanopore.

[0110] In other embodiments, the force can be a voltage force applied across the first and second openings, for example, if the detector is or includes a nanopore, the voltage force can be an electrical potential applied across the nanopore. The voltage can be applied using any suitable device, such as those described herein. Suitable electrical potentials are described in more detail herein.

[0111] In some embodiments, a force is applied across the membrane in which a detector, such as a nanopore, is embedded. The force is typically applied from the cis to the trans side of the membrane, i.e., from the cis to the trans side of the nanopore. The force can be a positive voltage applied to the nanopore or a negative voltage applied to the nanopore.

[0112] Typically, the force is a positive voltage applied across the first and second openings, e.g., if the detector is or includes a nanopore, the force can be a positive voltage across the nanopore such that the trans side of the pore is positive relative to the cis side of the pore. In such embodiments, the force thus attracts the negatively charged polynucleotide to move from the cis side of the pore to the trans side of the pore. In such embodiments, the methods provided herein typically include using a motor protein on the cis side of the pore to control the movement of the polynucleotide from the trans side of the pore to the cis side of the pore against the applied force, i.e., in the opposite direction to the applied force. However, in some embodiments, the methods provided herein include allowing the polynucleotide to move from the cis side of the pore to the trans side of the pore in the same direction as the applied force.

[0113] In other embodiments, the force is a negative voltage applied across the first and second openings, for example, if the detector is or includes a nanopore, the force can be a negative voltage across the nanopore such that the trans side of the pore is negative relative to the cis side of the pore. In such embodiments, the force attracts the negatively charged polynucleotide to move from the trans side of the pore to the cis side of the pore. In such embodiments, the methods provided herein typically include using a motor protein on the trans side of the pore to control the movement of the polynucleotide from the cis side of the pore to the trans side of the pore against the applied force, i.e., in the opposite direction to the applied force. However, in some embodiments, the methods provided herein include allowing the polynucleotide to move from the trans side of the pore to the cis side of the pore in the same direction as the applied force.

[0114] However, as explained below, the methods provided herein do not rely on moving the polynucleotide in the opposite direction to the applied force. In some embodiments, the direction of movement can be the same direction as any applied force, but still outward from a detector, such as a pore. In such embodiments, the motor protein typically controls the movement of the polynucleotide outward from a detector, such as a pore, at a rate that is faster than the rate that would result from the applied force alone.

[0115] Thus, in some embodiments, the force is a positive voltage applied across the first and second openings, e.g., if the detector is or comprises a nanopore, the force can be a positive voltage across the nanopore such that the trans side of the pore is positive relative to the cis side of the pore, and the method can include using a motor protein at the trans side of the pore to control movement of the polynucleotide in a direction from the cis side of the pore to the trans side of the pore with the applied force. In other embodiments, the force is a negative voltage applied across the first and second openings, e.g., if the detector is or comprises a nanopore, the force can be a positive voltage across the nanopore such that the trans side of the pore is negative relative to the cis side of the pore, and the method can include using a motor protein at the cis side of the pore to control movement of the polynucleotide in a direction from the trans side of the pore to the cis side of the pore with the applied force.

[0116] setting In the methods provided, a leader is attached to a target polynucleotide.

[0117] In step (i) of the disclosed method, the reader is contacted with a first opening of the detector. The target polynucleotide then migrates in a direction from the first opening to the second opening.

[0118] The reader can be captured by a detector (e.g., a nanopore) in the methods provided herein.

[0119] In some embodiments, the leader is attached directly to the first end of the polynucleotide. In some embodiments, the leader is attached to the first end of the polynucleotide by a linker. The leader may be included in an adaptor attached to the first end of the polynucleotide. Adaptors are described in more detail herein.

[0120] As described in more detail herein, any suitable reader can be used. In some embodiments, the reader can pass through a first opening of the detector. In some embodiments, the reader can pass through a second opening of the detector. In some embodiments, the reader can pass through a first opening and a second opening of the detector. In some embodiments, the detector is or includes a nanopore and the reader can pass at least a portion of a path through the nanopore. In some embodiments, the detector is or includes a nanopore and the reader can translocate the nanopore.

[0121] Typically, prior to step (i) of the disclosed methods, a leader is provided at a first end of the polynucleotide (e.g., by being included at the first end of the target polynucleotide or by being included in a polynucleotide adaptor attached to the first end of the target polynucleotide) and a motor protein is attached to the leader. For example, a leader may be present at the 3' end of the single stranded polynucleotide and a motor protein may be attached to the leader. Alternatively, a leader sequence may be present at the 5' end of the single stranded polynucleotide and a motor protein may be attached to the leader. The motor protein may be stalled on the leader. The motor protein may be attached to the leader at a stall moiety, as described in more detail herein.

[0122] In some embodiments, the target polynucleotide is single-stranded, the target polynucleotide includes a leader, the leader is located at a first end of the target polynucleotide or is included in an adaptor attached to the first end of the target polynucleotide, and the motor protein binds, for example, to stall at the leader. In such embodiments, the leader is typically captured by a first opening of the detector (e.g., a first opening of the nanopore), and the leader moves through the nanopore until it reaches the stalled motor protein. The motor protein can be "pushed back" onto the target polynucleotide toward the second end of the target polynucleotide. The motor protein controls the movement of the polynucleotide out of the pore. The motor protein then debinds from the target polynucleotide, and the target polynucleotide moves in a second direction relative to the first and second openings of the detector, such as the pore. The second direction can be within the pore. The motor protein then rebinds to the target polynucleotide and controls the movement of the target polynucleotide in the first direction relative to the detector, thus rereading the target polynucleotide. This is illustrated in FIG. 2.

[0123] In some embodiments, the target polynucleotide is double-stranded.

[0124] In some embodiments, the target polynucleotide is double-stranded and comprises a first strand and a second strand, the target polynucleotide comprises a leader, the leader sequence is located at a first end of the polynucleotide and is included in the first strand or attached to the first strand or included in an adaptor attached to the first strand, and the motor protein binds to the leader. The first strand of the double-stranded polynucleotide can be a template strand. The first strand of the double-stranded polynucleotide can be a complementary strand. In some embodiments, the motor protein stalls at the leader attached to the first strand of the target polynucleotide or included in an adaptor attached to the first strand of the target polynucleotide. In some embodiments, the target polynucleotide is double-stranded and comprises a first strand and a second strand, the target polynucleotide comprises a leader sequence, the leader is located at a first end of the polynucleotide and is included in the first strand or attached to the first strand or included in an adaptor attached to the first strand, and the motor protein stalls at the leader.

[0125] For example, a leader sequence can be present at the 3' end of a first strand of a double-stranded polynucleotide and a motor protein can bind, e.g., stall at the leader. Alternatively, a leader sequence can be present at the 5' end of a first strand of a double-stranded polynucleotide and a motor protein can bind, e.g., stall at the leader.

[0126] In such an embodiment, the leader sequence is typically captured by the nanopore, and the first strand moves through the nanopore until it reaches the stalled motor protein. The motor protein can be "pushed back" onto the target polynucleotide towards the second end of the first strand of the target polynucleotide. The motor protein controls the movement of the polynucleotide out of the pore. The motor protein then de-binds from the target polynucleotide, and the target polynucleotide moves in a second direction relative to the first and second openings of a detector, such as a pore. The second direction may be within the pore. The motor protein then re-binds to the target polynucleotide and controls the movement of the first strand of the target polynucleotide in the first direction relative to the detector, thus re-reading the target polynucleotide. This is illustrated in FIG. 3.

[0127] In some embodiments, the first strand and the second strand are attached together by a hairpin adaptor at the second end of the first strand. In some embodiments, the hairpin adaptor is attached at its 5' end to the 3' end of the first strand and at its 3' end to the 5' end of the second strand of the target double-stranded polynucleotide. In some embodiments, the hairpin adaptor is attached at its 3' end to the 5' end of the first strand and at its 5' end to the 3' end of the second strand of the target double-stranded polynucleotide. Thus, the hairpin adaptor links the first strand to the second strand. The hairpin adaptor typically links the second end of the first strand of the double-stranded polynucleotide to the first end of the second strand of the double-stranded polynucleotide.

[0128] In some embodiments, the target polynucleotide is double stranded and comprises a first strand and a second strand, the target polynucleotide comprises a leader, the leader is located at a first end of the polynucleotide and is comprised in the first strand or is comprised in an adaptor attached to the first strand, the first strand and the second strand are attached together at a second end of the first strand by a hairpin adaptor, and the motor protein binds to the hairpin adaptor. In some embodiments, the motor protein is stalled at the leader attached to the first strand of the target polynucleotide or is comprised in an adaptor attached to the first strand of the target polynucleotide. In some embodiments, the target polynucleotide is double stranded and comprises a first strand and a second strand, the target polynucleotide comprises a leader sequence, the leader sequence is located at a first end of the polynucleotide and is either included in the first strand or included in an adaptor attached to the first strand, the first strand and the second strand are attached together by a hairpin adaptor attached to (i) the second end of the first strand and (ii) the first end of the second strand, and the motor protein is stalled at the leader.

[0129] In such an embodiment, the leader sequence is typically captured by the nanopore, and the double-stranded polynucleotide moves through the nanopore until it reaches the stalled motor protein. The motor protein can be "pushed back" on the first strand of the double-stranded polynucleotide, optionally the hairpin adaptor, and optionally the second strand of the double-stranded polynucleotide, toward the second end of the first strand of the target polynucleotide. The motor protein controls the movement of the second strand, and optionally the hairpin adaptor, and also optionally the first strand of the double-stranded polynucleotide, out of the pore. The motor protein then debinds from the target polynucleotide, and the target polynucleotide moves in a second direction relative to the first and second openings of a detector, such as a pore. The second direction may be within the pore. The motor protein then rebinds to the target polynucleotide and controls the movement of the target polynucleotide in the first direction relative to the detector, thus re-reading the target polynucleotide. This is illustrated in FIG. 4.

[0130] The motor protein may be attached to the leader at any point. The motor protein may be attached to the leader at the point where the leader is attached to a first end of the polypeptide. The motor protein may be attached to the free end of the leader. The motor protein may be attached to any point along the leader. The motor protein may be attached to a portion of a polynucleotide, such as a portion of DNA or RNA, contained in a leader, the leader comprising a portion of a polynucleotide and a second type of monomer unit, such as one or more spacer units as described herein.

[0131] A leader may be constructed or designed to promote unbinding of a target polynucleotide from a polynucleotide binding site of a motor protein when the motor protein is in proximity to the leader (e.g., when the motor protein contacts the leader).

[0132] In such embodiments, the motor protein typically has a lower affinity for the leader than for the target polynucleotide, i.e., for the portion of the target polynucleotide being characterized. In some embodiments, the leader has a different structure than the target polynucleotide. In some embodiments, the leader includes a different type of nucleotide than the target polynucleotide.

[0133] In some embodiments, the leader is charged. In some embodiments, the leader is uncharged. In some embodiments, the leader is negatively charged or positively charged, typically negatively charged.

[0134] In some embodiments, the leader is a polymer. In some embodiments, the leader is a linear polymeric group. In some embodiments, the leader is a charged polymer, e.g., a negatively charged polymer. In some embodiments, the leader comprises a polymer such as a polynucleotide, e.g., DNA or RNA, a modified polynucleotide (such as abasic DNA), PNA, LNA, polyethylene glycol (PEG), or a polypeptide. In such embodiments, the leader may be 10-150 monomer units (e.g., ethylene glycol or sugar units) in length, such as 20-120, such as 30-100, such as 40-80, such as 50-70 monomer units (e.g., ethylene glycol or sugar units) in length.

[0135] In some embodiments, a leader may comprise a single-stranded polynucleotide region that does not have significant secondary structure, for example, a leader sequence typically does not form a hairpin or G-quadruplex and is therefore susceptible to capture by a nanopore.

[0136] In some embodiments, the leader is or comprises a polynucleotide. In embodiments where the leader is a polynucleotide, the leader can be the same type of polynucleotide as the target polynucleotide, or the leader can be a different type of polynucleotide. For example, the target polynucleotide can be DNA and the leader can be RNA, or vice versa. In some embodiments, the target polynucleotide comprises or consists of DNA (e.g., dsDNA) and the leader does not consist of DNA, although in some embodiments the leader can comprise one or more nucleotides in addition to other monomeric units, e.g., the leader can comprise one or more spacers as described herein). In some embodiments, the target polynucleotide comprises or consists of ssDNA and the leader does not consist of ssDNA (e.g., the leader can comprise one or more spacers as described herein). In some embodiments, the target polynucleotide comprises or consists of DNA and the leader also consists of DNA.

[0137] In some embodiments, the leader sequence comprises a single stranded DNA, such as a poly dT section. The leader sequence can be of any length, but is typically between 10 and 150 nucleotides in length, e.g., between 20 and 120, 30 and 100, 40 and 80, or 50 and 70 nucleotides in length.

[0138] For example, in some embodiments, the target polynucleotide comprises deoxyribonucleotides (DNA). In such embodiments, the leader may comprise one or more nucleotides lacking both a nucleobase and a sugar moiety (e.g., a spacer moiety). Suitable spacer moieties are described in more detail herein and include C2 spacers, C3 spacers, C6 spacers, iSp9 spacers, iSp18 spacers, and the like. For example, the leader may comprise about 5 to about 100, e.g., about 10 to about 50, e.g., about 20 to about 40, spacers described herein, e.g., C3, iSp9, and / or iSp18. Alternatively or additionally, the leader may comprise ribonucleotides (RNA), peptide nucleotides (PNA), glycerol nucleotides (GNA), threose nucleotides (TNA), locked nucleotides (LNA), bridged nucleotides (BNA), or abasic nucleotides. In some embodiments, the leader may comprise one or more nucleotides having modified phosphate linkages (e.g., including methylphosphonate or phosphothiolate linkages).

[0139] In some other embodiments, the target polynucleotide comprises ribonucleotides (RNA). In such embodiments, the leader may comprise one or more spacers as defined above, deoxyribonucleotides (DNA), peptide nucleotides (PNA), glycerol nucleotides (GNA), threose nucleotides (TNA), locked nucleotides (LNA), bridged nucleotides (BNA), abasic nucleotides, or nucleotides containing modified phosphate linkages.

[0140] Typically, the target polynucleotide comprises deoxyribonucleotides (DNA) and the leader comprises one or more spacer moieties (eg, a C3 spacer) and / or one or more ribonucleotides.

[0141] A leader may only comprise one type of polynucleotide different from the target polynucleotide. For example, if the target polynucleotide is DNA, the leader may comprise a spacer portion or RNA. A leader may comprise more than one type of polynucleotide different from the target polynucleotide. For example, if the target polynucleotide is DNA, the leader may comprise a spacer portion and RNA. A leader may comprise a portion that is the same type of polynucleotide as the target polynucleotide. For example, if the target polynucleotide is DNA, the leader may comprise a portion of DNA in addition to a spacer polynucleotide or RNA. Such a portion may be referred to as a "trap", i.e., a leader based on a spacer (e.g., C3 spacer) and / or RNA (e.g., 2'-methoxyuridine) polynucleotide may comprise one or more DNA traps. A trap typically comprises 1-10 nucleotides, such as 1-6 nucleotides, e.g., 1, 2, 3, 4 or 5 nucleotides, e.g., 1-3 nucleotides. Thus, where the target polynucleotide is DNA, the leader may comprise one or more RNA (e.g., 2'-methoxyuridine) and / or spacer (e.g., C3 spacer) portions and one or more DNA (e.g., thymidine) traps of 1-10 nucleotides in length.

[0142] In some embodiments, the leader comprises one or more stalling units as described herein. In some embodiments, the leader may comprise one or more abasic spacers, i.e., one or more spacers in which a base has been removed from one or more nucleotides in the polynucleotide adaptor. In some embodiments, the leader may comprise a peptide nucleic acid (PNA), a glycerol nucleic acid (GNA), a threose nucleic acid (TNA), a locked nucleic acid (LNA), or a synthetic polymer with nucleotide side chains. In some embodiments, the leader may comprise one or more nitroindoles, one or more inosines, one or more acridines, one or more 2-aminopurines, one or more 2-6-diaminopurines, one or more 5-bromo-deoxyuridines, one or more inverted thymidines (inverted dT), one or more inverted dideoxy-thymidines (ddT), one or more dideoxy-cytidines (ddC), one or more 5-methylcytidines, one or more 5-hydroxymethylcytidines, one or more 2'-O-methyl RNA bases, one or more isoform ... -deoxycytidine (Iso-dC), one or more iso-deoxyguanosine (Iso-dG), one or more C3(OC3H6OPO3) groups, one or more photocleavable (PC)[OC3H6-C(O)NHCH2-C6H3NO2-CH(CH3)OPO3] groups, one or more hexanediol groups, one or more spacer 9 (iSp9)[(OCH2CH2)3OPO3] groups, or one or more spacer 18 (iSp18)[(OCH2CH2)6OPO3] groups, or one or more thiol linkages. The leader may contain any combination of these groups. Many of these groups are commercially available from IDT® (Integrated DNA Technologies®). For example, C3, iSp9, and iSp18 spacers are all available from IDT®. The leader may contain any number of the above groups. For example, the leader can include about 5 to about 100, such as about 10 to about 50, such as about 20 to about 40, spacers described herein, such as C3, iSp9 and / or iSp18.

[0143] In some embodiments, the leader comprises one or more spacers as described herein (e.g., about 20 to about 40 spacers as described herein, e.g., C3, iSp9 and / or iSp18, and / or a region of homopolymeric polynucleotide, e.g., a poly dT section about 20 to about 40 nucleotides in length), a heteropolymeric polynucleotide section, a stall section for stalling a motor protein, which may comprise one or more stall units as described herein (e.g., about 20 to about 40 spacers as described herein, e.g., C3, iSp9 and / or iSp18), and a polynucleotide motor protein binding site.

[0144] One of skill in the art will also understand that when the leader comprises a polynucleotide chain, the sequence of the leader is typically not critical and may be controlled or selected according to other experimental conditions, such as the motor protein and any polynucleotides to be characterized. Exemplary sequences are provided in the Examples, e.g., Example 3, for illustrative purposes only. For example, the leader may comprise a sequence such as one or more of SEQ ID NOs: 53, 54 or 55, 56, 57 or 58, or a polynucleotide sequence having at least 20%, such as at least 30%, such as at least 40%, such as at least 50%, such as at least 60%, such as at least 70%, such as at least 80%, such as at least 90%, such as at least 95% sequence similarity or identity to one or more of SEQ ID NOs: 53, 54 or 55, 56, 57 or 58. The sequence of the leader can typically be altered without adversely affecting the efficacy of the methods provided herein.

[0145] In some embodiments, the leader is attached directly to the target polynucleotide. Typically, the leader is attached to the target polynucleotide using an adapter. That is, in some embodiments, the leader is included in an adapter described herein, and the adapter is attached to the target polynucleotide. Suitable adapters are described in more detail herein. In some embodiments, the leader is attached to a group of adapters that are attached to the target polynucleotide by chemistry, as described herein.

[0146] In some embodiments, the leader is attached to the first end of the polynucleotide by a chemical bond. In some embodiments, the leader is attached to the first end of the polynucleotide by a covalent bond. In some embodiments, the leader is attached to the first end of the polynucleotide by a linker. Any suitable linker may be used. In some embodiments, the linker is a polynucleotide as described herein. In some embodiments, the linker is a synthetic polymer, such as PEG. In some embodiments, the linker is the same type of polynucleotide as the target polynucleotide (e.g., in some embodiments, the target polynucleotide comprises DNA and the linker comprises DNA), and the leader comprises one or more nucleotides of a different type than those contained in the polynucleotide portion of the conjugate. In some embodiments, the target polynucleotide comprises DNA, the linker comprises DNA, and the leader comprises one or more non-DNA nucleotides as described herein, e.g., one or more spacers as described herein.

[0147] In some embodiments, the leader is attached to the 5' end of the target polynucleotide. In some embodiments, the leader is attached to the 3' end of the target polynucleotide.

[0148] In some embodiments, a target polynucleotide has naturally occurring reactive functional groups that can be used to attach a leader.

[0149] The attachment chemistry between the leader and the polynucleotide in the conjugate is not particularly limited. Any suitable combination of reactive functional groups can be used. Many suitable reactive groups and their chemical targets are known in the art. Some exemplary reactive groups and their corresponding targets include arylazides that can react with amines, carbodiimides that can react with amines and carboxyl groups, hydrazides that can react with carbohydrates, hydroxymethylphosphines that can react with amines, imide esters that can react with amines, isocyanates that can react with hydroxyl groups, carbonyls that can react with hydrazines, maleimides that can react with sulfhydryl groups, NHS-esters that can react with amines, PFP-esters that can react with amines, psoralens that can react with thymine, pyridyl disulfides that can react with sulfhydryl groups, vinyl sulfones that can react with sulfhydrylamines and hydroxyl groups, vinyl sulfonamides, and the like.

[0150] Other suitable chemistries for attaching a leader to a polynucleotide include click chemistry. Many suitable click chemistry reagents are known in the art. Suitable examples of click chemistry include, but are not limited to, the following: (a) Copper(I)-catalyzed azide-alkyne cycloaddition (azide-alkyne Huisgen cycloaddition), (b) Strain-promoted azide-alkyne cycloaddition; [3+2] cycloaddition of alkenes and azides; inverse demand Diels-Alder reaction of alkenes and tetrazines; and photoclick reaction of alkenes and tetrazoles. (c) Copper-free variants of the 1,3 dipolar cycloaddition reaction in which an azide reacts with a strained alkyne, e.g., in a cyclooctane ring, e.g., in bicyclic [6.1.0]nonyne (BCN); (d) reaction of an oxygen nucleophile on one linker with an epoxide or aziridine reactive moiety on the other; and (e) Staudinger ligation, in which the alkyne moiety can be replaced by an aryl phosphine, resulting in a specific reaction with an azide to give an amide bond.

[0151] Any reactive group can be used to attach the leader to the polypeptide. Some suitable reactive groups include [1,4-bis[3-(2-pyridyldithio)propionamido]butane; 1,1 1-bis-maleimidotriethylene glycol; 3,3'-dithiodipropionic acid di(N-hydroxysuccinimide ester); ethylene glycol-bis(succinic acid N-hydroxysuccinimide ester); 4,4'-diisothiocyanatostilbene-2,2'-disulfonic acid disodium salt; bis[2-(4-azidosalicylamido)ethyl]disulfide; 3-(2-pyridyldithio)propionic acid N-hydroxysuccinimide ester; 4-maleimidobutyric acid N-hydroxysuccinimide ester; iodoacetic acid N-hydroxysuccinimide ester; S-acetylthioglycolic acid N-hydroxysuccinimide ester; azido-PEG-maleimide; and alkyne-PEG-maleimide. The reactive group may be any of those disclosed in WO2010 / 086602, particularly in Table 3 of that application.

[0152] In some embodiments, prior to the attachment step, the reactive functional group is included in the leader and the target functional group is included in the first end of the polynucleotide. In other embodiments, prior to the conjugation step, the reactive functional group is included in the first end of the polynucleotide and the target functional group is included in the leader. In some embodiments, the reactive functional group is directly attached to the polynucleotide. In some embodiments, the reactive functional group is attached to the polynucleotide via a spacer. Any suitable spacer can be used. Suitable spacers include, for example, alkyl diamines such as ethyl diamine, and the like.

[0153] Motor protein stalling As explained above, in some embodiments, the methods provided herein can include characterization of a target polynucleotide having a motor protein thereon that is stalled at a stall portion, hi some embodiments, the leader comprises a stall portion and the motor protein binds to the leader by stalling on the leader at the stall portion.

[0154] Any suitable stall portion may be used in the methods provided herein. In some embodiments, the stall portion comprises a stall site as described herein. In some embodiments, the stall site comprises one or more stall units.

[0155] Any suitable stall unit can be used. The stall unit typically provides an energy barrier that impedes the movement of the motor protein. For example, the stall unit can stall the motor protein by reducing the pulling force of the motor protein on the polynucleotide. This can be achieved, for example, by using an abasic spacer, i.e., a spacer in which a base has been removed from one or more nucleotides in the polynucleotide adaptor. The spacer can physically block the movement of the polynucleotide handling protein, for example, by introducing a bulky chemical group that physically impedes the movement of the polynucleotide handling protein.

[0156] In some embodiments, the stall unit may comprise a linear molecule such as a polymer. Typically, such stall unit has a structure different from that of the target polynucleotide. For example, when the target polynucleotide is DNA, the or each stall unit typically does not comprise DNA. Specifically, when the target polynucleotide is deoxyribonucleic acid (DNA) or ribonucleic acid (RNA), the or each stall unit preferably comprises peptide nucleic acid (PNA), glycerol nucleic acid (GNA), threose nucleic acid (TNA), locked nucleic acid (LNA), bridged nucleic acid (BNA), or a synthetic polymer with nucleotide side chains. In some embodiments, the stall unit is one or more nitroindoles, one or more inosines, one or more acridines, one or more 2-aminopurines, one or more 2-6-diaminopurines, one or more 5-bromo-deoxyuridines, one or more inverted thymidines (inverted dT), one or more inverted dideoxy-thymidines (ddT), one or more dideoxy-cytidines (ddC), one or more 5-methylcytidines, one or more 5-hydroxymethylcytidines, one or more 2'-O-methyl RNA bases, one or more isoforms. -deoxycytidine (Iso-dC), one or more iso-deoxyguanosine (Iso-dG), one or more C3(OC3H6OPO3) groups, one or more photocleavable (PC)[OC3H6-C(O)NHCH2-C6H3NO2-CH(CH3)OPO3] groups, one or more hexanediol groups, one or more spacer 9 (iSp9)[(OCH2CH2)3OPO3] groups, or one or more spacer 18 (iSp18)[(OCH2CH2)6OPO3] groups, or one or more thiol bonds. The stall site may include any combination of these groups. Many of these groups are commercially available from IDT® (Integrated DNA Technologies®). For example, C3, iSp9, and iSp18 spacers are all available from IDT®. The stall site may include any number of the above groups as stall units. For example, a stall site may include from 1 to about 12 or more (eg, from about 1 to about 6, such as from about 1 to about 8, eg, from 1 to about 4) such stall units.

[0157] In some embodiments, the stall unit may include one or more chemical groups that stall the motor protein. In some embodiments, suitable chemical groups are one or more pendant chemical groups. One or more chemical groups may be attached to one or more nucleobases in the polynucleotide. One or more chemical groups may be attached to the backbone of the polynucleotide. There may be any number of suitable chemical groups, such as 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12 or more. Suitable groups include, but are not limited to, fluorophores, streptavidin and / or biotin, cholesterol, methylene blue, dinitrophenol (DNP), digoxigenin and / or antidigoxigenin, and dibenzylcyclooctyne groups.

[0158] In some embodiments, the stall unit can include a polymer, hi some embodiments, the stall unit can include a polymer that is a polypeptide or polyethylene glycol (PEG).

[0159] In some embodiments, the stall unit may contain one or more abasic nucleotides (i.e., nucleotides lacking a nucleobase), for example, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12 or more abasic nucleotides. The nucleobase may be replaced by -H (idSp) or -OH in the abasic nucleotide. The abasic residue may be inserted into the target polynucleotide by removing the nucleobase from one or more adjacent nucleotides. For example, the polynucleotide may be modified to contain 3-methyladenine, 7-methylguanine, 1,N6-ethenoadenine inosine, or hypoxanthine, and the nucleobase may be removed from these nucleotides using human alkyladenine DNA glycosylase (hAAG). Alternatively, the polynucleotide may be modified to contain uracil, and the nucleobase may be removed by uracil DNA glycosylase (UDG). In one embodiment, one or more stall units do not contain any abasic nucleotides.

[0160] Suitable stall units can be designed or selected depending on the nature of the polynucleotide / polynucleotide adaptor, the motor protein, and the conditions under which the method is to be carried out. For example, many polynucleotide processing proteins process DNA in vivo, and such proteins can typically be stalled using anything but DNA.

[0161] Thus, in some embodiments of the provided methods, the motor protein stalls at a stall site that comprises one or more stall units independently selected from: - a polynucleotide secondary structure, preferably a hairpin or G-quadruplex (TBA), - nucleic acid analogues, preferably selected from peptide nucleic acids (PNAs), glycerol nucleic acids (GNAs), threose nucleic acids (TNAs), locked nucleic acids (LNAs), bridged nucleic acids (BNAs) and abasic nucleotides, - a spacer unit selected from nitroindole, inosine, acridine, 2-aminopurine, 2-6-diaminopurine, 5-bromo-deoxyuridine, inverted thymidine (inverted dTs), inverted dideoxy-thymidine (ddTs), dideoxy-cytidine (ddCs), 5-methylcytidine, 5-hydroxymethylcytidine, 2'-O-methyl RNA bases, isodeoxycytidine (Iso-dCs), isodeoxyguanosine (Iso-dGs), a C3(OC3H6OPO3) group, a photocleavable (PC) [OC3H6-C(O)NHCH2-C6H3NO2-CH(CH3)OPO3] group, a hexanediol group, a spacer 9 (iSp9) [(OCH2CH2)3OPO3] group, a spacer 18 (iSp18) [(OCH2CH2)6OPO3] group, and a thiol linkage, and - fluorophores, avidins such as traptavidin, streptavidin and neutravidin, and / or biotin, cholesterol, methylene blue, dinitrophenol (DNP), digoxigenin and / or anti-digoxigenin and dibenzylcyclooctyne groups.

[0162] The stall moieties described herein can also be used to construct leaders suitable for use in the disclosed rereading methods. As explained above, in some embodiments of such methods, the leader sequences described herein are constructed or designed to promote dissociation of a target polynucleotide from a polynucleotide binding site of a motor protein when the motor protein is in proximity to the leader (e.g., when the motor protein contacts the leader sequence). In some embodiments, the leader sequence can include any of the spacer moieties described above.

[0163] Motor protein destalling In some embodiments, the methods provided herein include contacting a stall moiety described herein with a detector (e.g., a nanopore), thereby stalling a motor protein that may be stalled with a leader bound to a target polynucleotide. Upon destalling, the motor protein can control movement of the polynucleotide in a first direction relative to first and second openings of the detector, as described in more detail herein (e.g., if the detector is or includes a nanopore, the motor protein can control movement of the polynucleotide from the nanopore).

[0164] In its simplest form, the motor protein can be unstalled from the stalled moiety by contacting the stalled moiety with a detector, e.g., a nanopore, however, in some embodiments the method includes actively unstalling the motor protein as described herein.

[0165] In some embodiments, destalling the motor protein comprises applying a destalling force to the polynucleotide, the destalling force being less than and / or in the opposite direction to a read force, the read force being a force applied while the motor protein controls movement of the target polynucleotide and measurements are being taken to determine one or more characteristics of the polynucleotide.

[0166] For example, the read power can be provided as a potential of +2V to -2V, typically -400mV to +400mV. The voltage used is preferably in a range having a lower limit selected from -400mV, -300mV, -200mV, -150mV, -100mV, -50mV, -20mV, and 0mV, and an upper limit independently selected from +10mV, +20mV, +50mV, +100mV, +150mV, +200mV, +300mV, and +400mV. The voltage used is more preferably in the range of 100mV to 240mV, most preferably in the range of 120mV to 220mV. The de-stall power is usually less than the read power. For example, the destall force can be from about −100 mV to +100 mV, such as from about −50 mV to about +50 mV, such as from about −25 mV to about +25 mV.

[0167] For example, in some embodiments, the read power is a potential in the range of +50 mV to +300 mV, more preferably +100 mV to +200 mV, such as +120 mV to +150 mV, and the de-stall power is a potential of -50 to +50 mV, such as -40 mV to +40 mV, or -20 mV to +20 mV, such as 0 mV.

[0168] In some embodiments, the de-stall force is in the opposite direction to the read force. For example, in some embodiments, the read force is applied as a positive potential and the de-stall force is applied as a negative potential. In other embodiments, the read force is applied as a negative potential and the de-stall force is applied as a positive potential. When the de-stall force is in the opposite direction to the read force, it may be the same magnitude as the read force or may be less in magnitude than the read force.

[0169] In some embodiments, the destall force is applied at zero potential. For example, in some embodiments, the read force is applied as a positive potential and the destall force is applied at zero applied potential. In other embodiments, the read force is applied as a negative potential and the destall force is applied at zero applied potential.

[0170] In some embodiments, the destalling force is applied for a time sufficient for the motor protein to destall from the stalled portion, hi some embodiments, the destalling force is applied for a time period of from 1 ms to about 10 s, such as from about 10 ms to about 1 s, e.g., from about 100 ms to about 700 ms, such as from about 300 ms to about 500 ms.

[0171] In some embodiments, destalling the motor protein comprises altering the applied force one or more times between the destall force and the read force. In some embodiments, altering the applied force in this manner comprises stepping or ramping the applied potential between the destall force and the read force. When ramping, any suitable waveform can be used, for example, the ramp can be a linear ramp, an exponential ramp, or a sigmoidal ramp.

[0172] In some embodiments, the applied force varies between a single de-stall force and a read force. In some embodiments, the applied force varies between a series of different de-stall forces and a read force. In some embodiments, the applied force varies stepwise between a series of increasing de-stall forces and a read force. The de-stall force at each step may be any suitable de-stall force, e.g., any of the de-stall forces described herein, and each step may be applied for any suitable period of time, e.g., any of the periods described herein.

[0173] In some embodiments, the de-stall force is the same as the read force, also referred to as the de-stall in a "free running" setting.

[0174] In some embodiments, the motor protein stalls at a stall site that includes one or more stall units and one or more stop moieties, and contacting the one or more stop moieties with a detector, e.g., a nanopore, retards movement of the polynucleotide relative to the detector, thereby causing the motor protein to de-stall from the one or more stall units. Such embodiments are suitable for use in a free-running setting.

[0175] In some embodiments, the stopping moiety provides an energy barrier to the detector that impedes movement of the polynucleotide through, for example, the nanopore. For example, the stopping moiety may impede movement of the polynucleotide through the nanopore by providing a physical block that must be removed before the polynucleotide can pass through the nanopore.

[0176] Without being bound by theory, the inventors believe that the stopping moiety delays translocation of the polynucleotide relative to the first and second openings of the detector (e.g., through the nanopore, if the detector is or includes a nanopore) for a sufficient period of time for the motor protein to overcome the stall unit(s) and stall.

[0177] In some embodiments, the stopping moiety comprises one or more stall units comprising a polynucleotide secondary structure, preferably a hairpin or a G-quadruplex (TBA). Such secondary structures prevent the polynucleotide from passing freely through the nanopore. Upon contacting the stopping moiety with the nanopore, the secondary structure dissociates (e.g., unravels). The time it takes for the secondary structure to dissociate allows the motor protein to unstall from the stall unit(s).

[0178] In some embodiments, the stop moiety comprises one or more stop units comprising a hybridized oligonucleotide. The oligonucleotide may hybridize to a target polynucleotide and prevent the movement of the target polynucleotide through the nanopore. When the stop moiety is contacted with the nanopore, the hybridized oligonucleotide dissociates from the target polynucleotide. The time it takes for the hybridized oligonucleotide to dissociate from the target polynucleotide allows the motor protein to de-stall from the stall unit(s).

[0179] In some embodiments, the stopping moiety comprises one or more stopping units, preferably comprising a nucleic acid analogue selected from peptide nucleic acid (PNA), glycerol nucleic acid (GNA), threose nucleic acid (TNA), locked nucleic acid (LNA), bridged nucleic acid (BNA) and abasic nucleotides. The nucleic acid analogue may be provided with a target polynucleotide or may be hybridized or otherwise bound to the target polynucleotide. When the nucleic acid analogue is provided with the target polynucleotide, contacting the stopping moiety with a nanopore causes the nucleic acid analogue to pass through the nanopore. The time it takes for the nucleic acid analogue to move relative to the detector, e.g., if the detector is or includes a nanopore, to pass through the pore, allows the motor protein to unstall from the stall unit(s). When the nucleic acid analogue hybridizes to the target polynucleotide, contacting the stopping moiety with a detector (e.g., a nanopore) typically causes the nucleic acid analogue to dissociate from the target polynucleotide (e.g., by passing through a nanopore) so that the target polynucleotide can move relative to the detector. The time it takes for the nucleic acid analog to dissociate from the polynucleotide allows the motor protein to unstall from the stall unit(s).

[0180] In some embodiments, the stopping moiety comprises one or more stalling units comprising a fluorophore, an avidin such as traptavidin, streptavidin and neutravidin, and / or a chemical group such as biotin, cholesterol, methylene blue, dinitrophenol (DNP), digoxigenin and / or antidigoxigenin and dibenzylcyclooctyne groups. The chemical group may attach to the target polynucleotide and prevent the target polynucleotide from moving relative to the detector (e.g., through the nanopore). In some embodiments, contacting the stopping moiety with the detector (e.g., nanopore) removes the chemical group from the target polynucleotide. In some embodiments, contacting the stopping moiety with the detector (e.g., nanopore) removes the chemical group from the target polynucleotide. In some embodiments, contacting the stopping moiety with the detector (e.g., nanopore) removes the chemical group from the detector through the first and second openings, e.g., through the nanopore. The time it takes for the chemical group to be removed from the target polynucleotide and / or to move through the first and second openings (e.g., of the nanopore) allows the motor protein to de-stall from the stalling unit(s).

[0181] In some embodiments, the stopping moiety comprises one or more stall units comprising a polynucleotide binding protein. Suitable polynucleotide binding proteins are described in more detail herein. The polynucleotide binding protein can bind to a polynucleotide and prevent the movement of the polynucleotide to a detector (e.g., through a nanopore). When the stopping moiety is contacted with a detector (e.g., a nanopore), the movement of the polynucleotide to the detector (e.g., through a nanopore) is delayed, for example, when the polynucleotide binding protein moves to contact the motor protein. The time taken to run allows the motor protein to de-stall from the stall unit(s).

[0182] Without being bound by theory, the inventors also believe that the stalling moiety often determines the conformation of the polynucleotide in the stalling unit(s). This is particularly true when the stalling moiety contains one or more linear groups such as spacer 18 (iSp18) [(OCH2CH2)6OPO3]. Without being bound by theory, it is believed that when such stalling unit(s) are in contact with a detector such as a nanopore, any applied force across the pore (e.g., an applied voltage field) can extend the stalling moiety in a nearly linear fashion. In this conformation, the motor protein is usually unable to pass through the stalling moiety and destall. However, when the target polynucleotide is stalled at the stalling moiety, the environment of the stalling unit is believed to be similar to that in solution, and the stalling unit may adopt a more compact pseudorandom coil configuration. In this configuration, it may be easier for the motor protein to overcome the stalling unit and destall.

[0183] Thus, in some embodiments, the motor protein stalls at one or more stall units and stall sites comprising one or more stall units independently selected from: - a polynucleotide secondary structure, preferably a hairpin or G-quadruplex (TBA), - nucleic acid analogues, preferably selected from peptide nucleic acids (PNAs), glycerol nucleic acids (GNAs), threose nucleic acids (TNAs), locked nucleic acids (LNAs), bridged nucleic acids (BNAs) and abasic nucleotides, - fluorophores, avidins such as traptavidin, streptavidin and neutravidin, and / or biotin, cholesterol, methylene blue, dinitrophenol (DNP), digoxigenin and / or antidigoxigenin and dibenzylcyclooctyne groups, and - polynucleotide binding proteins, And contacting the one or more stopping moieties with a detector, such as a nanopore, retards translocation of the polynucleotide relative to the detector (e.g., through the nanopore), thereby causing the motor protein to unstall from one or more stall units.

[0184] Motor proteins As one of skill in the art would understand, any suitable motor protein can be used in the methods and products provided herein.

[0185] The motor protein can be any protein that can bind to a polynucleotide and control its movement relative to a detector, e.g., a nanopore, e.g., through a pore.

[0186] More specifically, motor proteins such as helicases typically have at least two modes of active operation (they require all the components necessary to facilitate movement, e.g., ATP and Mg) 2+ The motor proteins can control DNA movement in one active mode (when the motor protein is provided with a cytoplasmic endothelium), and one inactive mode (when the components required to facilitate movement are not provided or when the motor protein is modified to prevent an active mode).

[0187] When all the components necessary to facilitate movement are provided, the motor protein can move along a polynucleotide, such as DNA, in either a 5'-3' or 3'-5' direction. Many motor proteins process polynucleotides, such as DNA, in a 5'-3' direction. Motor proteins that control the movement of polynucleotides in this manner are typically suitable for use in the methods provided herein.

[0188] However, if a motor protein does not have the necessary components to facilitate translocation or is modified to prevent it from actively controlling the movement of a polynucleotide relative to the nanopore, it can still passively control the movement of a polynucleotide relative to the nanopore. For example, a motor protein can bind to a polynucleotide and act as a brake to slow the movement of the polynucleotide when the polynucleotide is drawn into the pore by an applied field (e.g., by a first force in the methods provided herein). In the "inactive" mode, it typically does not matter whether the polynucleotide is captured 3' or 5' (i.e., moves through the nanopore in a 5'-3' or 3'-5' direction), since the applied force provides the driving force to move the polynucleotide through the nanopore. However, in such embodiments, the motor protein can still control the movement of the polynucleotide relative to the nanopore, for example, by acting as a brake. In the inactive mode, the control of the movement of the polynucleotide by the motor protein can be described in several ways, including ratcheting, sliding, and braking. Typically, the methods provided herein do not involve the use of motor proteins operating in a passive mode. However, in embodiments of the methods provided herein that use a polynucleotide binding protein, the polynucleotide binding protein can be a motor protein that operates in a passive mode.

[0189] As explained above, some embodiments of the methods provided herein also include the use of a polynucleotide binding protein as a stop moiety that impedes the movement of a polynucleotide chain through a nanopore. In some embodiments, the polynucleotide binding protein can be a motor protein as described herein. In other embodiments, the polynucleotide binding protein can be a protein that binds to a polynucleotide but does not have polynucleotide processing capabilities, i.e., in some embodiments, it is not a motor protein.

[0190] A polynucleotide handling enzyme is a polypeptide that can interact with a polynucleotide. An enzyme may modify a polynucleotide by cleaving the polynucleotide to form individual nucleotides or shorter chains of nucleotides, such as di- or trinucleotides. An enzyme may modify a polynucleotide by orienting it or moving it to a specific location. A motor protein as used herein may be or be derived from a polynucleotide handling enzyme. A polynucleotide binding protein may be or be derived from a polynucleotide handling enzyme.

[0191] In one embodiment, the motor protein and / or polynucleotide binding protein are independently derived from any member of Enzyme Classification (EC) groups 3.1.11, 3.1.13, 3.1.14, 3.1.15, 3.1.16, 3.1.21, 3.1.22, 3.1.25, 3.1.26, 3.1.27, 3.1.30, and 3.1.31.

[0192] Typically, the motor proteins are each independently a helicase, a polymerase, an exonuclease, a topoisomerase, or a variant thereof.

[0193] In some embodiments, the motor protein and / or polynucleotide binding protein may be modified to prevent the motor protein from disassociating from the polynucleotide, and thus, in some embodiments of such methods, the target polynucleotide does not disassociate from the motor protein.

[0194] The term "disassociation" as used herein refers to the dissociation of the motor protein from the target polynucleotide. Thus, the motor protein may be modified to prevent it from dissociating from the target polynucleotide, for example, into the reaction medium. It is important to distinguish the potential "detachment" of the motor protein from the "disassociation" of the motor protein from the target polynucleotide. "Disassociation" as used herein refers to the temporary release of the motor protein's active site (described in more detail herein) of the target polynucleotide, but does not imply disassociation. Thus, for example, the motor protein may be modified to prevent the motor protein from disassociating from the polynucleotide, but not to prevent the motor protein from disassociating from the polynucleotide. When not bound, the motor protein remains bound to the target polynucleotide. For example, the motor protein may maintain its association with the target polynucleotide (i.e., prevent disassociation from the target polynucleotide), but only because it is topologically closed around the target polynucleotide. The polynucleotide binding site can remain free to bind or unbind to the target polynucleotide while the motor protein remains bound to the target polynucleotide, such that the motor protein can bind or unbind to the target polynucleotide. When the motor protein unbinds from the target polynucleotide, it can move on (e.g., along) the target polynucleotide under an applied force and can rebind to the target polynucleotide. When associated with but unbinds from the target polynucleotide, the motor protein cannot dissociate from the target polynucleotide.

[0195] The motor protein and / or polynucleotide binding protein can be adapted to prevent disassociation in any suitable manner. For example, the motor protein and / or polynucleotide binding protein can be loaded onto a polynucleotide and then modified to prevent it from disassociating from the polynucleotide. Alternatively, the motor protein and / or polynucleotide binding protein can be modified to prevent it from disassociating from the polynucleotide before it is loaded onto the polynucleotide. Modification of the motor protein to prevent it from disassociating from the polynucleotide can be achieved using methods known in the art, for example, the methods discussed in WO2014 / 013260 and WO2015 / 110813, each of which is incorporated herein by reference in its entirety, and with particular reference to the passages describing the modification of motor proteins such as helicases to prevent the motor protein from disassociating from a polynucleotide chain. For example, the motor protein and / or polynucleotide binding protein can be modified by treatment with tetramethylazodicarboxamide (TMAD). Various other closing moieties are described in more detail herein.

[0196] For example, a motor protein and / or polynucleotide binding protein may have a polynucleotide debinding opening, e.g., a cavity, groove, or gap, through which a strand can pass when the motor protein and / or polynucleotide binding protein disassociates from the strand. In some embodiments, the polynucleotide debinding opening is an opening through which a nucleotide can pass when the motor protein and / or polynucleotide binding protein disassociates from the nucleotide. In some embodiments, the polynucleotide debinding opening of a given motor protein and / or polynucleotide binding protein can be determined by reference to its structure, e.g., by reference to its X-ray crystal structure. The X-ray crystal structure can be obtained in the presence and / or absence of a polynucleotide substrate. In some embodiments, the location of the polynucleotide debinding opening in a given motor protein can be estimated or confirmed by molecular modeling using standard packages known in the art. In some embodiments, the polynucleotide debinding opening can be generated transiently by movement of one or more portions of the motor protein, e.g., one or more domains.

[0197] The motor protein and / or polynucleotide binding protein may be modified by closing the polynucleotide debinding opening. The polynucleotide debinding opening may be closed with a closing moiety. Thus, closing the polynucleotide debinding opening may prevent the motor protein and / or polynucleotide binding protein from disassociating from the polynucleotide. For example, the motor protein and / or polynucleotide binding protein may be modified by covalently closing the polynucleotide debinding opening. However, as explained above, closing the polynucleotide debinding opening does not necessarily prevent the target polynucleotide from debinding from the polynucleotide binding site of the motor protein. In some embodiments, a preferred protein for this purpose is a helicase.

[0198] In some embodiments, the motor protein may be modified to prevent the target polynucleotide from disassociating from the target polynucleotide. The motor protein may be modified in any suitable manner.

[0199] Without being bound by theory, the inventors believe that promoting unbinding and slowing rebinding may promote re-reading. Without being bound by theory, the inventors believe that this may be because each step that the motor protein takes on the target polynucleotide is associated with the probability that the motor protein will unbind from the polynucleotide. Such a probability of unbinding may be identified by the so-called off-rate. An increase in the off-rate is believed to promote the drop back of the motor protein on the polynucleotide chain. Similarly, again without being bound by theory, the inventors believe that the distance that the motor protein can move along the target polynucleotide before rebinding, once unbound from the target polynucleotide, is associated with the on-rate. Thus, re-reading may be promoted by increasing the off-rate and decreasing the on-rate of the motor protein with respect to the target polynucleotide. Tuning the off-rate and on-rate of the motor protein for a given type of polynucleotide is within the capabilities of the skilled artisan in view of the disclosure herein. Thus, a motor protein can be modified to facilitate unbinding of a target polynucleotide from the polynucleotide binding site of the motor protein and / or to delay rebinding of a target polynucleotide to the polynucleotide binding site of the motor protein, hi some embodiments, a motor protein is modified to facilitate both unbinding of a target polynucleotide from the polynucleotide binding site of the motor protein and to delay rebinding of a target polynucleotide to the polynucleotide binding site of the motor protein.

[0200] In some embodiments, motor proteins may be modified with a closing moiety to (i) topologically close the polynucleotide binding site of the motor protein around a target polynucleotide and (ii) facilitate debinding of the target polynucleotide from the polynucleotide binding site of the motor protein and / or retard rebinding of the target polynucleotide to the polynucleotide binding site of the motor protein. Motor proteins may be modified in any suitable manner to facilitate attachment of such a closing moiety.

[0201] In some embodiments, the closing moiety may comprise a bifunctional crosslinking moiety. The closing moiety may comprise a bifunctional crosslinker. The bifunctional crosslinker may attach at two points on the motor protein and close the polynucleotide debinding opening of the motor protein, thereby preventing release of the polynucleotide from the motor protein while allowing debinding of the polynucleotide from the polynucleotide binding site of the motor protein.

[0202] The closing moiety may be attached at any suitable position on the motor protein. For example, the closing moiety may bridge two amino acid residues of the motor protein. Typically, at least one amino acid bridged by the closing moiety is a cysteine ​​or a non-natural amino acid. The cysteine ​​or non-natural amino acid may be introduced into the motor protein by substitution or modification of a naturally occurring amino acid residue of the motor protein. Methods for introducing non-natural amino acids are well known in the art and include, for example, native chemical ligation with a synthetic polypeptide chain that includes such a non-natural amino acid. Methods for introducing cysteines into motor proteins are also described, for example, in Sambrook et al., Molecular Cloning: A Laboratory Manual, 4 thed., Cold Spring Harbor Press, Plainsview, New York (2012), and Ausubel et al., Current Protocols in Molecular Biology (Supplement 114), John Wiley & Sons, New York (2016).

[0203] In some embodiments, the closing moiety has a length of about 1 Å to about 100 Å. The length of the closing moiety can be calculated according to static bond lengths or, more preferably, using molecular dynamics simulations. The length can be, for example, about 2 Å to about 80 Å, such as about 5 Å to about 50 Å, such as about 8 to about 30 Å, such as about 10 to about 25 Å or about 20 Å, such as about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, or 19 Å.

[0204] Without being bound by theory in any way, the inventors believe that, in general, a longer closing portion may increase the off-rate of the motor protein from the polynucleotide and thus facilitate re-reading.

[0205] In some embodiments, the closing moiety comprises a bond. In some embodiments, the closing moiety comprises a disulfide bond. The disulfide bond can be formed by treating the motor protein with any suitable reagent, such as TMAD.

[0206] In some embodiments, the closing moiety comprises a reagent that forms a bond between two click chemistry groups on the motor protein. Examples of click chemistry reagents are provided herein.

[0207] In some embodiments, the closing moiety comprises a protein. For example, a biotin group may be present on the motor protein and the closing moiety may comprise streptavidin. A tag, such as a snoop tag or a spy tag, may be present on the motor protein and the closing moiety may comprise a protein, such as a snoop catcher or a spy catcher, respectively.

[0208] In some embodiments, the closing moiety comprises a structure of the formula [ABC], where A and C are each independently reactive functional groups for reacting with an amino acid residue in a motor protein, and B is a linking moiety. In some embodiments, the closing moiety comprises a bond between a thio group, e.g., a thiol group on a cysteine ​​residue. Thus, in some embodiments, A and C are cysteine-reactive functional groups.

[0209] In some embodiments, the linking moiety B comprises a linear or branched, unsubstituted or substituted alkylene, alkenylene, alkynylene, arylene, heteroarylene, carbocyclylene, or heterocyclylene moiety, which is optionally interrupted or terminated by one or more atoms or groups selected from O, N(R), S, C(O), C(O)NR, C(O)O, unsubstituted or substituted arylene, arylene-alkylene, heteroarylene, heteroarylene-alkylene, carbocyclylene, carbocyclylene-alkylene, heterocyclylene, and heterocyclylene-alkylene, where R is selected from H, unsubstituted or substituted alkyl, and unsubstituted or substituted aryl. Typically, R is H or methyl, more typically H.

[0210] Typically, the alkylene group is 1-20 Typically, the alkenylene group is 2-20 Typically, the alkynylene group is 2-20 An arylene group is typically an alkynylene group. 6-12Typically, the heteroarylene group is a 5- to 12-membered heteroarylene group. Typically, the carbocyclylene group is 5-12 A carbocyclylene group. Typically, the heterocyclylene group is a 5- to 12-membered heterocyclylene group.

[0211] Typically, the alkylene, alkenylene, or alkynylene moiety may be uninterrupted, interrupted, or terminated by one or more atoms or groups selected from O, N(R), S, C(O), C(O)NR, and C(O)O and unsubstituted or substituted arylene. Usually, the alkylene, alkenylene, or alkynylene moiety may be uninterrupted, interrupted, or terminated by one or more atoms or groups selected from O and N(R) and unsubstituted or substituted arylene. More often, the alkylene, alkenylene, or alkynylene moiety may be uninterrupted, interrupted, or terminated by one or more O atoms.

[0212] For example, the linking moiety is often an unsubstituted or substituted C 1-10 Alkylene, C 2-10 Alkenylene or C 2-10 An alkynylene moiety is uninterrupted or interrupted or terminated with one or more O atoms.

[0213] In some embodiments, linking moiety B comprises an alkylene, oxyalkylene, or polyoxyalkylene group, and / or A and C are each maleimide groups. The alkylene, oxyalkylene, or polyoxyalkylene group can have a length, for example, from about 5 Å to about 50 Å, such as from about 8 to about 30 Å, for example, from about 10 to about 25 Å.

[0214] For example, the linking moiety is (CH2CH2O) xwhere x is 1 to 10, e.g., 1 to 5, e.g., 1, 2 or 3. Exemplary linking moieties are described in Example 2 and include, for example, BMOE (1,2-bismaleimidoethane), BMOP (1,3-bismaleimidopropane), BMB (1,4-bismaleimidobutane), BM(PEG)2 (1,8-bismaleimido-diethylene glycol) and BM(PEG)3 (1,11-bismaleimido-triethylene glycol).

[0215] Motor proteins suitable for being closed using such closing moieties are discussed in more detail herein, in some preferred embodiments, the motor protein is a helicase, such as the Dda helicase described herein.

[0216] In one embodiment, the motor protein and / or polynucleotide binding protein is an exonuclease or is derived from an exonuclease. Suitable enzymes include, but are not limited to, exonuclease I from E. coli (SEQ ID NO: 1), exonuclease III enzyme from E. coli (SEQ ID NO: 2), RecJ from T. thermophilus (SEQ ID NO: 3) and bacteriophage lambda exonuclease (SEQ ID NO: 4), TatD exonuclease, and variants thereof. Three subunits comprising the sequence shown in SEQ ID NO: 3 or variants thereof interact to form a trimeric exonuclease.

[0217] In one embodiment, the motor protein and / or polynucleotide binding protein is derived from a polymerase. The polymerase can be PyroPhage® 3173 DNA polymerase (commercially available from Lucigen® Corporation), SD polymerase (commercially available from Bioron®), Klenow from NEB, or variants thereof. In one embodiment, the enzyme is Phi29 DNA polymerase (SEQ ID NO: 5) or variants thereof. Modified versions of Phi29 polymerase that can be used in the present invention are disclosed in U.S. Patent No. 5,576,204.

[0218] In one embodiment, the motor protein and / or polynucleotide binding protein is derived from a topoisomerase. In one embodiment, the topoisomerase is a member of any of the subclassification (EC) groups 5.99.1.2 and 5.99.1.3. The topoisomerase can be a reverse transcriptase, an enzyme that can catalyze the formation of cDNA from an RNA template. They are commercially available, for example, from New England Biolabs® and Invitrogen®.

[0219] In one embodiment, the motor protein and / or polynucleotide binding protein are derived from a helicase. Any suitable helicase can be used according to the methods provided herein. For example, the or each enzyme used according to the present disclosure can be independently selected from Hel308 helicase, RecD helicase, TraI helicase, TrwC helicase, XPD helicase, and Dda helicase, or variants thereof. A monomeric helicase can include several domains attached together. For example, TraI helicase and TraI subgroup helicase can include two RecD helicase domains, a relaxase domain, and a C-terminal domain. These domains typically form a monomeric helicase that can function without forming oligomers. Specific examples of suitable helicases include Hel308, NS3, Dda, UvrD, Rep, PcrA, Pif1, and TraI. These helicases typically act on single-stranded DNA. Examples of helicases that can translocate along both strands of double-stranded DNA include FtfK and hexamer enzyme complexes, or multi-subunit complexes such as RecBCD. In one embodiment, the motor protein is a Dda (DNA-dependent ATPase) helicase.

[0220] Hel308 helicase is described in publications such as WO2013 / 057495, the entire contents of which are incorporated by reference. RecD helicase is described in publications such as WO2013 / 098562, the entire contents of which are incorporated by reference. XPD helicase is described in publications such as WO2013 / 098561, the entire contents of which are incorporated by reference. Dda helicase is described in publications such as WO2015 / 055981 and WO2016 / 055777, the entire contents of which are each incorporated by reference.

[0221] In one embodiment, the helicase may comprise a sequence set forth in SEQ ID NO:6 (Trwc Cba) or a variant thereof, a sequence set forth in SEQ ID NO:7 (Hel308 Mbu) or a variant thereof, or a sequence set forth in SEQ ID NO:8 (Dda) or a variant thereof. The variant may differ from the native sequence in any of the ways discussed herein. An exemplary variant of SEQ ID NO:8 includes E94C / A360C. A further exemplary variant of SEQ ID NO:8 includes E94C / A360C, followed by (ΔM1)G1G2 (i.e., deletion of M1, followed by addition of G1 and G2).

[0222] Typically, the motor protein or polynucleotide binding protein may have a fuel binding site. Active unwinding of DNA may be coupled to enhanced hydrolysis, for example, in the motor protein.

[0223] The fuel is typically a free nucleotide or a free nucleotide analog. Free nucleotides include adenosine monophosphate (AMP), adenosine diphosphate (ADP), adenosine triphosphate (ATP), guanosine monophosphate (GMP), guanosine diphosphate (GDP), guanosine triphosphate (GTP), thymidine monophosphate (TMP), thymidine diphosphate (TDP), thymidine triphosphate (TTP), uridine monophosphate (UMP), uridine diphosphate (UDP), uridine triphosphate (UTP), cytidine monophosphate (CMP), cytidine diphosphate (CDP), cytidine triphosphate (CTP), cyclic adenosine monophosphate (cAMP), cyclic guanosine monophosphate (cGMP), deoxyadenosine monophosphate (dAMP), deoxyadeno ...cytidine monophosphate (CMP), cytidine diphosphate (CDP), cytidine triphosphate (CTP), cyclic adenosine monophosphate (cAMP), cyclic guanosine monophosphate (cGMP), deoxyadenosine monophosphate (dAMP), deoxyadenosine triphosphate (GTP), cytidine monophosphate (CAMP), cytidine diphosphate (CGMP), deoxyadenosine monophosphate (DAMP), deoxyadenosine triphosphate (GTP), cytidine monophosphate (CAMP), cytidine diphosphate (CGMP), deoxyadenosine monophosphate (DAMP), deoxyadenosine triphosphate The free nucleotide may be, but is not limited to, one or more of adenosine diphosphate (dADP), deoxyadenosine triphosphate (dATP), deoxyguanosine monophosphate (dGMP), deoxyguanosine diphosphate (dGDP), deoxyguanosine triphosphate (dGTP), deoxythymidine monophosphate (dTMP), deoxythymidine diphosphate (dTDP), deoxythymidine triphosphate (dTTP), deoxyuridine monophosphate (dUMP), deoxyuridine diphosphate (dUDP), deoxyuridine triphosphate (dUTP), deoxycytidine monophosphate (dCMP), deoxycytidine diphosphate (dCDP), and deoxycytidine triphosphate (dCTP). The free nucleotide is usually selected from AMP, TMP, GMP, CMP, UMP, dAMP, dTMP, dGMP, or dCMP. The free nucleotide is typically adenosine triphosphate (ATP).

[0224] A cofactor for a motor protein is a factor that enables the motor protein to function. The cofactor is preferably a divalent metal cation. The divalent metal cation is preferably Mg 2+ , Mn 2+ , Ca 2+ , or Co 2+ The cofactor is most preferably Mg 2+ It is.

[0225] In some embodiments, a polynucleotide binding protein is other than a motor protein as used herein. As used herein, the terms polynucleotide binding protein and polynucleotide binding moiety can be used interchangeably.

[0226] For example, the polynucleotide binding protein or polynucleotide binding moiety may comprise one or more domains independently selected from a helix-hairpin-helix (HhH) domain, a eukaryotic single-stranded binding protein (SSB), a bacterial SSB, an archaeal SSB, a viral SSB, a double-stranded binding protein, a sliding clamp, a processivity factor, a DNA binding loop, a replication initiator protein, a telomere binding protein, a repressor, a zinc finger, and a proliferating cell nuclear antigen (PCNA).

[0227] The helix-hairpin-helix (HhH) domain is a polypeptide motif that binds to DNA in a sequence-nonspecific manner. Suitable domains include domain H (residues 696-751) and domain HI (residues 696-802) from topoisomerase V from Methanopyrus kandleri (SEQ ID NO:37). The polynucleotide binding moiety can be domain HL of SEQ ID NO:37 as set forth in SEQ ID NO:38 or a polynucleotide binding variant thereof. The HhH domain can comprise the sequence set forth in SEQ ID NO:23 or 31 or 32, or a polynucleotide binding variant thereof.

[0228] SSBs bind to single-stranded DNA with high affinity in a sequence-nonspecific manner. SSBs are classified into the following series: Class: All beta proteins, Fold: OB fold, Superfamily: Nucleic acid binding proteins, Family: Single-stranded DNA binding domain, SSBs. SSBs can be derived from eukaryotes such as humans, mice, rats, fungi, protozoa or plants, prokaryotes such as bacteria and archaea, or viruses. Eukaryotic SSBs are also known as replication proteins A (RPA). In most cases, they are heterotrimers formed of units of different sizes. Some of the larger units (e.g., RPA70 in Saccharomyces cerevisiae) are stable and bind to ssDNA in a monomeric form. Bacterial SSBs bind to DNA as stable homotetramers (e.g., E. coli, Mycobacterium smegmatis and Helicobacter pylori) or homodimers (e.g., Deinococcus radiodurans and Thermotoga maritima). Some, such as the SSB encoded by the crenarchaeon Sulfolobus solfataricus, are homotetrameric. Some SSBs from other species have been shown to be monomeric (Methanococcus jannaschii and Methanothermobacter thermoautotrophicum). Still other species of archaea, including Archaeoglobus fulgidus and Methanococcoides burtonii, contain two open reading frames with sequence similarity to RPA. Viral SSBs bind DNA as monomers.

[0229] The SSB is typically selected or modified to have a carboxy-terminal (C-terminal) region that has no or a reduced net negative charge compared to the wild-type protein. Such an SSB typically does not block a transmembrane pore. The C-terminal region of the SSB is typically about the last third, quarter, fifth, or eighth of the SSB at the C-terminus. The C-terminal region is typically about the last 10 amino acids to about the last 60 amino acids of the SSB at the C-terminus, for example about the last 20 amino acids to about the last 40 amino acids of the SSB, such as about the last 30 amino acids of the SSB at the C-terminus.

[0230] Examples of SSBs that contain a C-terminal region that does not have a net negative charge include human mitochondrial SSB (HsmtSSB; SEQ ID NO: 33, human replication protein A 70 kDa subunit, human replication protein A 14 kDa subunit, telomere end, binding protein α subunit from Oxytricha nova, core domain of telomere end binding protein β subunit from Oxytricha nova, protection of telomere protein 1 (Pot1) from Schizosaccharomyces pombe, human Pot1, the OB-fold domain of BRCA2 from mouse or rat, and the p5 protein from phi29 (SEQ ID NO: 34), as well as polynucleotide-binding variants thereof. Examples of SSBs whose C-terminal regions can be modified to reduce the net negative charge include SSB from E. coli (EcoSSB; SEQ ID NO: 35, SSB from Mycobacterium tuberculosis, SSB from Deinococcus radiodurans, SSB from Thermus thermophiles, Sulfolobus solfataricus SSB, human replication protein A32kDa subunit (RPA32) fragment, Saccharomyces cerevisiae CDC13SSB, E. coli Primosomal replication protein N (PriB), Arabidopsis thaliana PriB, hypothetical protein At4g28440, T4 SSB (gp32; SEQ ID NO: 36), RB69 SSB (gp32; SEQ ID NO: 24), T7 SSB (gp2.5; SEQ ID NO: 25), and polynucleotide-binding variants thereof. Suitable modifications to reduce the net negative charge are disclosed in WO2014 / 013259.

[0231] Double-stranded binding proteins bind to double-stranded DNA with high affinity. Suitable double-stranded binding proteins include Mutator S (MutS, NCBI Reference Sequence: NP_417213.1, SEQ ID NO: 39), Sso7d (Sufolobus solfataricus P2, NCBI Reference Sequence: NP_343889.1, SEQ ID NO: 40, Nucleic Acids Research, 2004, Vol. 32, No. 3, 1197-1207), Sso10b1 (NCBI Reference Sequence: NP_342446.1, SEQ ID NO: 41), Sso10b2 (NCBI Reference Sequence: NP_342448.1, SEQ ID NO: 42), tryptophan repressor (Trp repressor, NCBI Reference Sequence: NP_291006.1, SEQ ID NO: 43), lambda repressor (NCBI Reference Sequence: NP_040628.1, SEQ ID NO: 44), These include, but are not limited to, Cren7 (NCBI Reference Sequence: NP_342459.1, SEQ ID NO: 45), major histone classes H1 / H5, H2A, H2B, H3 and H4 (NCBI Reference Sequence: NP_066403.2, SEQ ID NO: 46), dsbA (NCBI Reference Sequence: NP_049858.1, SEQ ID NO: 47), Rad51 (NCBI Reference Sequence: NP_002866.2, SEQ ID NO: 48), sliding clamp and topoisomerase V Mka (SEQ ID NO: 37), or polynucleotide-binding variants of any of these proteins.

[0232] Other polynucleotide-binding proteins include sliding clamps. Sliding clamps are usually multimeric proteins (homodimers or homotrimers) that surround dsDNA. Sliding clamps usually require accessory proteins (clamp loaders) to assemble them around the DNA helix in an ATP-dependent process. They also do not directly contact DNA, but function as topological tethers. Associated with DNA sliding clamps are processivity factors, which are viral proteins that anchor the cognate polymerase to DNA, dramatically increasing the length of the fragments generated. They can be monomeric (as in the case of UL42 from herpes simplex virus type 1) or multimeric (UL44 from cytomegalovirus is a dimer). UL42 typically comprises the sequence shown in SEQ ID NO:26 or SEQ ID NO:30 or a polynucleotide-binding variant thereof.

[0233] Another polynucleotide binding protein is the thioredoxin binding domain (TBD) (residues 258-333) of bacteriophage T7 DNA polymerase. Binding of the TBD to thioredoxin (e.g., from E. coli) causes a conformational change of the polypeptide to one that binds to DNA. Other polynucleotide binding proteins include the accessory protein cisA from phage Φx174 and the geneII protein from phage M13. These proteins have unique DNA binding capabilities, and some of them recognize specific DNA sequences. Other polynucleotide binding proteins include telomere binding proteins.

[0234] Small DNA-binding motifs (such as helix-turn-helix) recognize specific DNA sequences. In the case of the bacteriophage 434 repressor, a 62-residue fragment was engineered and shown to retain DNA-binding ability and specificity. Zinc fingers consist of about 30 amino acids that bind to DNA in a specific manner. Typically, each zinc finger recognizes only three DNA bases, but multiple fingers can be linked to recognize longer sequences.

[0235] Proliferating cell nuclear antigen (PCNA) forms a very tight clamp that slides up and down dsDNA or ssDNA. PCNA from Crenarchaea is a heterotrimer of SEQ ID NOs: 27, 28 and 29. Thus, the polynucleotide binding protein can be a trimer comprising the sequences shown in SEQ ID NOs: 27, 28 and 29 or a polynucleotide binding variant thereof. Another PCNA sliding clamp (NCBI Reference Sequence: ZP_06863050.1; SEQ ID NO: 49) forms a dimer. Thus, the polynucleotide binding protein can be a dimer comprising SEQ ID NO: 49 or a polynucleotide binding variant thereof.

[0236] The polynucleotide binding motif may be selected from any of the following: [Table 3-1] [Table 3-2] [Table 3-3]

[0237] Polynucleotides The methods of the invention involve characterization of a target polynucleotide as it moves relative to a detector, such as a nanopore.

[0238] Polynucleotides, such as nucleic acids, are polymers that contain two or more nucleotides.Polynucleotides can be single-stranded or double-stranded.Double-stranded polynucleotides are made of two single-stranded polynucleotides that hybridize together.Target polynucleotides can be single-stranded or double-stranded polynucleotides, which are described in more detail herein.

[0239] A polynucleotide can contain any combination of any nucleotides, whether naturally occurring or artificial.

[0240] A nucleotide typically comprises a nucleobase, a sugar, and at least one phosphate group. The nucleobase and sugar form a nucleoside.

[0241] Nucleobases are typically heterocyclic. Nucleobases include, but are not limited to, purines and pyrimidines, more specifically, adenine (A), guanine (G), thymine (T), uracil (U), and cytosine (C).

[0242] The sugar is typically a pentose. Nucleotide sugars include, but are not limited to, ribose and deoxyribose. The sugar is preferably deoxyribose. The polynucleotide preferably includes the following nucleosides: deoxyadenosine (dA), deoxyuridine (dU) and / or thymidine (dT), deoxyguanosine (dG), and deoxycytidine (dC).

[0243] The nucleotide is typically a ribonucleotide or a deoxyribonucleotide. The nucleotide typically contains a monophosphate, a diphosphate, or a triphosphate. The nucleotide may contain more than three phosphates, for example, four or five phosphates. The phosphate may be attached to the 5' or 3' side of the nucleotide. Nucleotides include, but are not limited to, adenosine monophosphate (AMP), guanosine monophosphate (GMP), thymidine monophosphate (TMP), uridine monophosphate (UMP), 5-methylcytidine monophosphate, 5-hydroxymethylcytidine monophosphate, cytidine monophosphate (CMP), cyclic adenosine monophosphate (cAMP), cyclic guanosine monophosphate (cGMP), deoxyadenosine monophosphate (dAMP), deoxyguanosine monophosphate (dGMP), deoxythymidine monophosphate (dTMP), deoxyuridine monophosphate (dUMP), deoxycytidine monophosphate (dCMP), and deoxymethylcytidine monophosphate. The nucleotides are preferably selected from AMP, TMP, GMP, CMP, UMP, dAMP, dTMP, dGMP, dCMP, and dUMP.

[0244] A nucleotide can be abasic (i.e., lacking a nucleobase). A nucleotide can also lack a nucleobase and a sugar (i.e., a C3 spacer).

[0245] The nucleotides in a polynucleotide can be attached to each other in any manner. The nucleotides are typically attached by their sugar and phosphate groups, as in nucleic acids. The nucleotides can be connected through their nucleobases, as in pyrimidine dimers.

[0246] A polynucleotide can be a nucleic acid, such as deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). A polynucleotide can include a single strand of RNA hybridized to a single strand of DNA. A polynucleotide can be any synthetic nucleic acid known in the art, such as peptide nucleic acid (PNA), glycerol nucleic acid (GNA), threose nucleic acid (TNA), bridged nucleotides (BNA), locked nucleic acid (LNA), or other synthetic polymers with nucleotide side chains. The PNA backbone is composed of repeating N-(2-aminoethyl)-glycine units linked by peptide bonds. The GNA backbone is composed of repeating glycol units linked by phosphodiester bonds. The TNA backbone is composed of repeating threose sugars linked together by phosphodiester bonds. LNA is formed from ribonucleotides with an extra bridge connecting the 2' oxygen and the 4' carbon in the ribose moiety, as discussed above.

[0247] The polynucleotide is preferably DNA, RNA, or a DNA or RNA hybrid, most preferably DNA. A DNA / RNA hybrid may contain DNA and RNA on the same strand. Preferably, a DNA / RNA hybrid contains one DNA strand hybridized to an RNA strand.

[0248] The backbone of a polynucleotide can be altered to reduce the likelihood of strand breaks. For example, DNA is known to be more stable than RNA under many conditions. The backbone of a polynucleotide chain can be modified to avoid damage caused by extreme chemicals such as free radicals.

[0249] DNA or RNA containing unnatural or modified bases can be produced by amplifying a natural DNA or RNA polynucleotide in the presence of modified NTPs using an appropriate polymerase.

[0250] The nucleotides in the polynucleotide may be modified. The nucleotides may be oxidized or methylated. One or more nucleotides in the polynucleotide may be damaged. For example, the polynucleotide may contain pyrimidine dimers. Such dimers are typically associated with UV damage and are the main cause of skin melanoma. One or more nucleotides in the polynucleotide may be modified, for example, by a label or tag.

[0251] Single-stranded polynucleotides may contain regions with strong secondary structures, such as hairpins, quadruplexes, or triplex DNA. These types of structures can be used to control the movement of the polynucleotide relative to the nanopore. For example, secondary structures can be used to stop the movement of the polynucleotide through the nanopore, as described in more detail herein. Each successive secondary structure along the strand stops the movement of the strand relative to the nanopore. The polynucleotide may reform the secondary structure after it has moved through the nanopore. Such secondary structures can be used to prevent the polynucleotide from moving back through the nanopore when a low or no negative voltage is applied (applied to the trans side of the nanopore), thus helping to control the movement of the polynucleotide so that it only occurs in a controlled manner in the relevant steps of the methods provided herein.

[0252] As used herein, a double-chain polypeptide can contain single-stranded regions as well as regions having other structures, such as hairpin loops, triplexes, and / or quadruplexes. Such secondary structures can be useful as described above in the context of single-stranded polynucleotides.

[0253] In some embodiments, the target polynucleotide is a double-stranded polynucleotide. In some embodiments, prior to step (i), the target polynucleotide is comprised in or consists of a first strand of a double-stranded polynucleotide comprising said first strand and a second strand.

[0254] In some embodiments, the target polynucleotide is a double-stranded polynucleotide, and the first strand is hybridized to the second strand. In some embodiments, the portion of the first strand between the motor protein and the second end is hybridized to the second strand. In some embodiments, the motor protein binds to a leader at the first end of the target polynucleotide, and the portion of the first strand between the motor protein and the second end is hybridized to the second strand. In some embodiments (e.g., during a method of controlling movement of a target polynucleotide using a motor protein), the motor protein may bind to the target polynucleotide, and the portion of the first strand between the motor protein and the second end is hybridized to the second strand.

[0255] Thus, in some embodiments, the portion of the target polynucleotide "above" the motor protein (i.e., between the motor protein and the second end of the polynucleotide) is hybridized (e.g., as a polynucleotide duplex) and the portion of the target polynucleotide "below" the motor protein (i.e., between the motor protein and the leader) is not hybridized.

[0256] In some embodiments, the movement of the double-stranded polynucleotide changes the portions of the double-stranded polynucleotide that are hybridized. Thus, the motor protein and / or detector (e.g., nanopore) can act to determine the portions of the polynucleotide that are hybridized as double stranded and the portions that are dehybridized.

[0257] In some embodiments, the movement of the target polynucleotide in the direction from the first opening of the detector to the second opening of the detector comprises separation of the first strand from the second strand. The separation of the first strand from the second strand may comprise dehybridization of the first strand from the second strand. In some embodiments, the first strand dehybridizes from the second strand as the first strand moves through the polynucleotide binding site of the motor protein. Thus, the motor protein acts to unfold the double-stranded polynucleotide as the first strand moves through the polynucleotide binding site of the motor protein in the direction from the first opening of the detector to the second opening.

[0258] In some embodiments, the movement of the target polynucleotide in the direction from the second opening of the detector to the first opening of the detector comprises annealing of the first strand to the second strand. The annealing of the first strand to the second strand may comprise hybridization of the first strand to the second strand. In some embodiments, the first strand hybridizes to the second strand as the first strand moves through the polynucleotide binding site of the motor protein in the direction from the second opening of the detector to the first opening. Thus, the separated (dehybridized) strands of the polynucleotide are reattached to each other.

[0259] The two strands of a double-stranded molecule can be attached to each other. For example, in an embodiment where the target polynucleotide comprises a first strand and a second strand, the first strand can be attached to the second strand. The two strands can be covalently linked at the ends of the molecule, for example, by linking the 5' end of one strand to the 3' end of the other strand in a hairpin structure.

[0260] The target polynucleotide can be of any length. For example, the target polynucleotide can be at least 10, at least 50, at least 100, at least 150, at least 200, at least 250, at least 300, at least 400, or at least 500 nucleotides or nucleotide pairs in length. The target polynucleotide can be 1000 or more nucleotides or nucleotide pairs in length, or 5000 or more nucleotides or nucleotide pairs in length, or 100000 or more nucleotides or nucleotide pairs in length, or 500,000 or more nucleotides or nucleotide pairs in length, or 1,000,000 or more nucleotides or nucleotide pairs in length, or 10,000,000 or more nucleotides or nucleotide pairs in length, or 100,000,000 or more nucleotides or nucleotide pairs in length, or 200,000,000 or more nucleotides or nucleotide pairs in length, or the entire length of a chromosome.

[0261] The target polynucleotide may be an oligonucleotide. An oligonucleotide is a short nucleotide polymer, usually having 50 or less nucleotides, such as 40 or less, 30 or less, 20 or less, 10 or less, or 5 or less nucleotides. The target oligonucleotide is preferably about 15 to about 30 nucleotides in length, such as about 20 to about 25 nucleotides in length. For example, the oligonucleotide may be about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, or about 30 nucleotides in length.

[0262] The target polynucleotide may be a fragment of a longer polynucleotide. In this embodiment, the longer polynucleotide is typically fragmented into multiple fragments, such as two or more shorter polynucleotides.

[0263] The target polynucleotide may comprise the products of a PCR reaction, genomic DNA, the products of endonuclease digestion, and / or a DNA library.

[0264] The target polynucleotide may be naturally occurring. The target polynucleotide may be secreted from the cell. Alternatively, the target analyte may be an analyte that is present intracellularly and thus the analyte must be extracted from the cell before the method can be performed.

[0265] Target polynucleotides may be derived from common organisms such as viruses, bacteria, archaea, plants or animals. Such organisms may be selected or modified to tailor the sequence of the target polynucleotide, for example, by adjusting base composition, removing undesirable sequence elements, etc. The selection and modification of organisms to arrive at desired polynucleotide properties is routine for those of skill in the art.

[0266] The source organism of the target polynucleotide can be selected based on the desired characteristics of the sequence. Desirable characteristics include the ratio of single-stranded polynucleotides to double-stranded polynucleotides produced by the organism, the sequence complexity of the polynucleotides produced by the organism, the composition (such as GC composition) of the polynucleotides produced by the organism, or the length of the continuous polynucleotide chain produced by the organism. For example, if a continuous polynucleotide chain of about 50 kb is required, lambda phage DNA can be used. If a longer continuous chain is required, other organisms can be used to produce the polynucleotides, for example, E. coli produces about 4.5 Mb of continuous dsDNA.

[0267] Target polynucleotides are often obtained from humans or animals, for example, from urine, lymph, saliva, mucus, semen, or amniotic fluid, or from whole blood, plasma, or serum. Target polynucleotides can be obtained from plants, for example, cereals, legumes, fruits, or vegetables. Target polynucleotides can include genomic DNA. Genomic DNA can be fragmented. DNA can be fragmented by any suitable method. For example, methods of fragmenting DNA are known in the art, and such methods can use transposases, such as MuA transposase. Genomic DNA is often not fragmented.

[0268] In some embodiments, the polynucleotides are synthetic or semi-synthetic. For example, DNA or RNA can be purely synthetic, synthesized by conventional DNA synthesis methods, such as phosphoramidite-based chemistry. Synthetic polynucleotide subunits can be linked together by known means, such as ligation or chemical bonding, to generate longer chains. In some embodiments, internal self-forming structures (e.g., hairpins, quadruplexes) can be engineered into the substrate, for example, by ligating appropriate sequences. Synthetic polynucleotides can be replicated and scaled up for production by means known in the art, including PCR, incorporation into bacterial factories, and the like.

[0269] In some embodiments, the polynucleotide may have a simplified nucleotide composition. In some embodiments, the polynucleotide has a repeating pattern of the same subunits. For example, the repeating unit may be (AmGn)q, where m, n, and q are positive integers. For example, m is often 1 to 20, such as 1 to 10, such as 1 to 5, such as 1, 2, 3, 4, or 5. n is often 1 to 20, such as 1 to 10, such as 1 to 5, such as 1, 2, 3, 4, or 5. m and n may be the same or different. In many cases, q is 1 to about 100,000. A typical repeating unit may be, for example, (AAAAAAGGGGGG)q. Repetitive polynucleotides can be made by many means known in the art, for example, by linking together synthetic subunits with sticky ends that allow for ligation. In some embodiments, the polynucleotide may be a linked polynucleotide. Methods for linking polynucleotides are described in PCT / GB2017 / 051493.

[0270] In some embodiments, the polynucleotide may include a base that includes a reactive side chain. Any suitable reactive functional group may be incorporated into the side chain as desired. Suitable examples of reactive functional groups include click chemistry reagents. Suitable examples of click chemistry include, but are not limited to: (f) Copper-free variants of the 1,3 dipolar cycloaddition reaction in which an azide reacts with an alkyne that is strained, e.g., to the cyclooctane ring; (g) reaction of an oxygen nucleophile on one linker with an epoxide or aziridine reactive moiety on the other; and (h) Staudinger ligation, in which the alkyne moiety can be replaced by an aryl phosphine, resulting in a specific reaction with an azide to give an amide bond.

[0271] Polynucleotide Adapters In some embodiments, the leader to which the motor protein first binds is included in a polynucleotide adaptor. WO2015 / 110813 describes the loading of motor proteins onto target polynucleotides such as adaptors and is incorporated herein by reference in its entirety.

[0272] An adaptor typically comprises a polynucleotide strand capable of binding to an end of a target polynucleotide, the target polynucleotide typically intended for characterization by the methods disclosed herein.

[0273] Polynucleotide adaptors can be added to both ends of target polynucleotide. Alternatively, different adaptors can be added to those two ends of target polynucleotide. Adaptors can be added to only one end of target polynucleotide. Methods of adding adaptors to polynucleotides are known in the art. Adaptors can be attached to polynucleotides, for example, by ligation, by click chemistry, by tagmentation, by topoisomerase conversion, or by any other suitable method.

[0274] The adaptor may be synthetic or artificial. Typically, the adaptor comprises a polymer as described herein. In some embodiments, the adaptor comprises a polynucleotide. In some embodiments, the adaptor may comprise a single-stranded polynucleotide strand. In some embodiments, the adaptor may comprise a double-stranded polynucleotide. The polynucleotide adaptor may comprise DNA, RNA, modified DNA (such as basic DNA), RNA, PNA, LNA, BNA, and / or PEG. Typically, the adaptor comprises single-stranded and / or double-stranded DNA or RNA.

[0275] The adaptor may comprise a stall moiety as described herein. The adaptor may comprise a loading site for a motor protein or a polynucleotide binding protein. The adaptor may comprise a tag.

[0276] The adaptor can be a Y adaptor. A Y adaptor is typically double-stranded and includes (a) a region at one end where the two strands are hybridized together and (b) a region at the other end where the two strands are not complementary. The non-complementary portions of the strands form an overhang. The hybridized stem of the adaptor typically attaches to the 5' end of the first strand of the double-stranded polynucleotide and the 3' end of the second strand of the double-stranded polynucleotide, or to the 3' end of the first strand of the double-stranded polynucleotide and the 5' end of the second strand of the double-stranded polynucleotide. The presence of the non-complementary region in the Y adaptor gives the adaptor a Y shape because, unlike the double-stranded portion, the two strands typically do not hybridize to each other. A motor protein or polynucleotide can bind to the overhang of an adaptor such as a Y adaptor. In another embodiment, a motor protein or polynucleotide binding protein can bind to the double-stranded region. In other embodiments, the motor protein or polynucleotide binding protein may bind to a single-stranded region and / or a double-stranded region of the adaptor. In other embodiments, a first motor protein or polynucleotide binding protein may bind to a single-stranded region of such an adaptor and a second motor protein or polynucleotide binding protein may bind to the double-stranded region of that adaptor. In some embodiments, one of the non-complementary strands of a polynucleotide adaptor, such as a Y adaptor, may provide a leader as described herein that can thread through the nanopore when contacted with the transmembrane pore.

[0277] In one embodiment, the adaptor comprises a membrane anchor or a pore anchor, hi some embodiments, the anchor is complementary to the overhang to which the motor protein or polynucleotide binding protein binds, and thus can bind to a polynucleotide hybridized to it.

[0278] In one embodiment, the polynucleotide adaptor is a hairpin loop adaptor, which is described in more detail herein. A hairpin loop adaptor is an adaptor that comprises a single polynucleotide strand, and the ends of the polynucleotide strand can hybridize to each other or are hybridized to each other, and the central section of the polynucleotide forms a loop. A suitable hairpin loop adaptor can be designed using methods known in the art. Typically, the 3' end of the hairpin loop adaptor is attached to the 5' end of the first strand of a double-stranded polynucleotide, and the 5' end of the hairpin loop adaptor is attached to the 3' end of the second strand of the double-stranded polynucleotide, or the 5' end of the hairpin loop adaptor is attached to the 3' end of the first strand of a double-stranded polynucleotide, and the 3' end of the hairpin loop adaptor is attached to the 5' end of the second strand of the double-stranded polynucleotide.

[0279] One of skill in the art will also understand that when the adaptor comprises a polynucleotide strand, the sequence of the adaptor is typically not critical and may be controlled or selected according to other experimental conditions, such as the motor protein and any polynucleotides to be characterized. Exemplary sequences are provided in the Examples for illustrative purposes only. For example, the adaptor may comprise a sequence such as one or more of SEQ ID NOs: 10-15 or 17-22, or a polynucleotide sequence having at least 20%, such as at least 30%, such as at least 40%, such as at least 50%, such as at least 60%, such as at least 70%, such as at least 80%, such as at least 90%, such as at least 95% sequence similarity or identity with one or more of SEQ ID NOs: 10-15 or 17-22. The sequence of the adaptor may typically be altered without adversely affecting the efficacy of the methods provided herein.

[0280] In some embodiments, the polynucleotide adaptor may include a loading site for loading a motor protein and / or a polynucleotide binding protein. The loading site may be, for example, a single-stranded region that can be targeted by a motor protein or a polynucleotide binding protein. The loading site may be a region of the polynucleotide adaptor to which an exogenous polynucleotide strand containing a motor protein or a polynucleotide binding protein can bind in order to translocate the motor protein or the polynucleotide binding protein to the polynucleotide being evaluated in the methods provided herein.

[0281] In some embodiments, a polynucleotide adaptor comprises a double-stranded polynucleotide region for binding to an end of a target polynucleotide (e.g., a target double-stranded polynucleotide), a single-stranded polynucleotide region capable of binding to (and / or to which) a motor protein or polynucleotide binding protein has bound, e.g., by stalling thereon, optionally comprising one or more stall moieties as described herein, an optional further polynucleotide region that optionally attaches (e.g., hybridizes) to a further polynucleotide strand that provides a tether to a membrane anchor as described herein, and a leader. In some embodiments, the tether to the membrane anchor is not complementary to the leader, such that the adaptor defines a Y adaptor.

[0282] Blocking part In some embodiments, a blocking moiety can be used to prevent the motor protein from disassociating from the target polynucleotide.

[0283] In some embodiments, the blocking moiety is comprised in a target polynucleotide. In some embodiments, the blocking moiety is comprised in a polynucleotide adaptor attached to the target polynucleotide. In some embodiments, a polynucleotide adaptor, such as a polynucleotide adaptor described herein, comprises a blocking moiety. In some embodiments, the blocking moiety is to prevent the motor protein from disassociating from the polynucleotide.

[0284] A blocking moiety can be used to prevent the motor protein from disassociating from the target polynucleotide. The blocking moiety is typically present at the opposite end of the target polynucleotide from the leader. For example, if the leader is attached to the 3' end of the polynucleotide strand in the target polynucleotide, the blocking moiety is typically located between the motor protein and the 5' end of the strand. If the leader is present at the 5' end of the polynucleotide strand of the target polynucleotide, the blocking moiety is typically located between the motor protein and the 3' end of the strand.

[0285] For example, in some embodiments, a target polynucleotide has a leader attached to it at a first end, and a second end of the target polynucleotide includes a blocking moiety to prevent the motor protein from disassociating from the polynucleotide.

[0286] In some embodiments, the target polynucleotide comprises a leader sequence at a first end of the target polynucleotide, the motor protein binds to (e.g., stalls at) the leader, and the blocking moiety is positioned between the motor protein and the second end of the polynucleotide (i.e., the end of the polynucleotide that is at the second end of the polynucleotide), thereby preventing the motor protein from disassociating from the target polynucleotide at the second end of the target polynucleotide.

[0287] For example, in some embodiments, the target polynucleotide comprises a leader at the 5' end of the first strand, the motor protein is stalled on the leader, and a blocking moiety is located between the motor protein and the 3' end of the first strand of the polynucleotide, thereby preventing the motor protein from disassociating from the target polynucleotide at the 3' end of the first strand of the target polynucleotide. In other embodiments, the target polynucleotide comprises a leader sequence at the 3' end of the first strand, the motor protein is stalled on the leader, and a blocking moiety is located between the motor protein and the 5' end of the first strand of the polynucleotide, thereby preventing the motor protein from disassociating from the target polynucleotide at the 5' end of the first strand of the target polynucleotide. The leader may be included in an adaptor as described herein.

[0288] When the target polynucleotide is a double-stranded polynucleotide, the blocking moiety is typically located on the same strand as the motor protein. When a motor protein is present on each strand of a double-stranded polynucleotide (e.g., when the double-stranded polynucleotide has rotational symmetry), a blocking moiety is typically present on each strand of the polynucleotide. An example of this is shown in Figure 5.

[0289] However, in some embodiments, the target polynucleotide comprises a double stranded polynucleotide, e.g., a first and second strand are joined together, e.g., with a hairpin adaptor as described herein. In such embodiments, the leader may be at a first end of the first strand, the second end of the first strand may be attached to a first end of the second strand, and the blocking moiety may be at a second end of the second strand. An example of this is shown in FIG.

[0290] The blocking moiety is typically too large to pass through the motor protein (e.g., through a polynucleotide binding site of a motor protein that has been modified to topologically close the polynucleotide binding site around the target polynucleotide), so that when movement of the polynucleotide relative to the detector brings the blocking moiety into contact with the motor protein (e.g., when the target polynucleotide unbinds from the polynucleotide binding site of the motor protein), further movement of the polynucleotide through the motor protein, and thus typically through the detector, is prevented. Thus, in some embodiments, the blocking moiety restricts movement of the polynucleotide through the polynucleotide binding site of the motor protein, thereby restricting movement of the target polynucleotide in a second direction relative to the detector. At such point, the motor protein may rebind to the polynucleotide. The polynucleotide may then move through the pore in the reverse direction under the control of the motor protein.

[0291] Any suitable blocking moiety can be used in the methods provided. Suitable blocking moieties include many of the same groups that can be used as terminating moieties described herein. For example, blocking moieties can include one or more of the following: - a polynucleotide secondary structure, preferably a hairpin or G-quadruplex (TBA), - nucleic acid analogues, preferably selected from peptide nucleic acids (PNAs), glycerol nucleic acids (GNAs), threose nucleic acids (TNAs), locked nucleic acids (LNAs), bridged nucleic acids (BNAs) and abasic nucleotides, - fluorophores, avidins such as traptavidin, streptavidin and neutravidin, and / or biotin, cholesterol, methylene blue, dinitrophenol (DNP), digoxigenin and / or antidigoxigenin and dibenzylcyclooctyne groups, and -Polynucleotide binding proteins.

[0292] The blocking moiety may be attached to the polynucleotide in any suitable manner. The blocking moiety may be attached directly to the polynucleotide. The blocking moiety may be attached to the polynucleotide via a linker. Any suitable linker may be used. Any suitable chemical reaction for linking the blocking moiety to the polynucleotide may be used. Any of the attachment methods described herein for attaching the leader to the polynucleotide may be used to attach the blocking moiety to the polynucleotide. The attachment means may be the same or different. In some embodiments, a linker is used to attach the leader to the polynucleotide and the blocking moiety is attached directly to the polynucleotide. In some embodiments, the leader is attached directly to the polynucleotide and a linker is used to attach the blocking moiety to the polynucleotide. In some embodiments, the leader is attached directly to the polynucleotide and the blocking moiety is attached directly to the polynucleotide. In some embodiments, a linker is used to attach the leader to the polynucleotide and a linker is used to attach the blocking moiety to the polynucleotide. When linkers are used to attach both the blocking moiety and the leader to the polynucleotide, the linkers may be the same or different.

[0293] Spacer In some embodiments, the polynucleotide or polynucleotide adaptor may include one or more spacers, e.g., 1 to about 10 spacers, e.g., 1 to about 5 spacers, e.g., 1, 2, 3, 4, or 5 spacers. The spacer may include any suitable number of spacer units. The spacer typically provides an energy barrier that impedes the movement of the polynucleotide binding protein. For example, the spacer may impede the movement of the motor protein or polynucleotide binding protein by reducing the traction force of the protein, e.g., using an abasic spacer. The spacer may physically block the movement of the polynucleotide binding protein, e.g., by introducing bulky chemical groups to physically impede the movement of the protein.

[0294] In some embodiments, one or more spacers are included in the polynucleotide or polynucleotide adaptor to provide a distinctive signal when they pass through the nanopore. One or more spacers may be used to define or separate one or more regions of the polynucleotide, for example, to separate an adaptor from a target polynucleotide.

[0295] In some embodiments, the spacer may comprise a linear molecule, such as a polymer, for example a polypeptide or polyethylene glycol (PEG). Typically, such a spacer has a structure different from the target polynucleotide. For example, if the target polynucleotide is DNA, the or each spacer typically does not comprise DNA. In particular, if the target polynucleotide is deoxyribonucleic acid (DNA) or ribonucleic acid (RNA), the or each spacer preferably comprises peptide nucleic acid (PNA), glycerol nucleic acid (GNA), threose nucleic acid (TNA), locked nucleic acid (LNA), or a synthetic polymer with nucleotide side chains. In some embodiments, the spacer is one or more nitroindoles, one or more inosines, one or more acridines, one or more 2-aminopurines, one or more 2-6-diaminopurines, one or more 5-bromo-deoxyuridines, one or more inverted thymidines (inverted dT), one or more inverted dideoxy-thymidines (ddT), one or more dideoxy-cytidines (ddC), one or more 5-methylcytidines, one or more 5-hydroxymethylcytidines, one or more 2'-O-methyl RNA bases, one or more isopropyl ethers, one or more tert-butyl ... -deoxycytidine (Iso-dC), one or more iso-deoxyguanosine (Iso-dG), one or more C3(OC3H6OPO3) groups, one or more photocleavable (PC)[OC3H6-C(O)NHCH2-C6H3NO2-CH(CH3)OPO3] groups, one or more hexanediol groups, one or more spacer 9 (iSp9)[(OCH2CH2)3OPO3] groups, or one or more spacer 18 (iSp18)[(OCH2CH2)6OPO3] groups, or one or more thiol bonds. The spacer may include any combination of these groups. Many of these groups are commercially available from IDT® (Integrated DNA Technologies®). For example, C3, iSp9, and iSp18 spacers are all available from IDT®. The spacer may include any number of the above groups as spacer units.

[0296] In some embodiments, the spacer may include one or more chemical groups, for example, one or more pendant chemical groups. One or more chemical groups may be attached to one or more nucleobases in the polynucleotide adaptor. One or more chemical groups may be attached to the backbone of the polynucleotide adaptor. There may be any number of suitable chemical groups, such as 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12 or more. Suitable groups include, but are not limited to, fluorophores, streptavidin and / or biotin, cholesterol, methylene blue, dinitrophenol (DNP), digoxigenin and / or antidigoxigenin, and dibenzylcyclooctyne groups.

[0297] In some embodiments, the spacer may contain one or more abasic nucleotides (i.e., nucleotides lacking a nucleobase), for example, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12 or more abasic nucleotides. The nucleobase may be replaced by -H (idSp) or -OH in the abasic nucleotide. The abasic spacer may be inserted into the target polynucleotide by removing the nucleobase from one or more adjacent nucleotides. For example, the polynucleotide may be modified to contain 3-methyladenine, 7-methylguanine, 1,N6-ethenoadenine inosine, or hypoxanthine, and the nucleobase may be removed from these nucleotides using human alkyladenine DNA glycosylase (hAAG). Alternatively, the polynucleotide may be modified to contain uracil, and the nucleobase may be removed by uracil DNA glycosylase (UDG). In one embodiment, the one or more spacers do not contain any abasic nucleotides.

[0298] Suitable spacers can be designed or selected depending on the nature of the polynucleotide or polynucleotide adaptor, the motor protein, and the conditions under which the method is to be performed.

[0299] tag In some embodiments, the polynucleotide or polynucleotide adaptor may include a tag or tether. For example, the polynucleotide may be attached to a tag on the nanopore, e.g., via its adaptor, and released at some point during characterization of the polynucleotide, e.g., by the nanopore. Strong non-covalent bonds (e.g., biotin / avidin) are still reversible and may be useful in some embodiments of the methods described herein.

[0300] A pore tag and polynucleotide adaptor pair can be configured such that the binding strength or affinity of a binding site on the polynucleotide (e.g., a binding site provided by an anchor or leader sequence of the adaptor or by a capture sequence within the double-stranded stem of the adaptor) to the tag on the nanopore is sufficient to maintain coupling between the nanopore and the polynucleotide until an applied force is applied, releasing the bound polynucleotide from the nanopore.

[0301] In some embodiments, the tag or tether is uncharged, which can ensure that the tag or tether is not drawn into the nanopore under the influence of a potential difference.

[0302] One or more molecules that attract or bind to the polynucleotide or adapter may be linked to the pore. Any molecule that hybridizes to the adapter and / or target polynucleotide may be used. The molecule attached to the pore may be selected from PNA tags, PEG linkers, short oligonucleotides, positively charged amino acids and aptamers. Such molecules are known in the art to be bound to pores. For example, pores with short oligonucleotides attached are disclosed in Howarka et al (2001) Nature Biotech. 19:636-639 and WO2010 / 086620, and pores with PEG attached within the lumen of the pore are disclosed in Howarka et al (2000) J. Am. Chem. Soc. 122(11):2411-2416.

[0303] Short oligonucleotides attached to a detector (e.g., a transmembrane pore) that contain a sequence complementary to the sequence of the leader sequence or another single-stranded sequence of the adapter may be used to enhance capture of target polynucleotides in the methods described herein.

[0304] In some embodiments, the tag or tether may include or be an oligonucleotide (e.g., DNA, RNA, LNA, BNA, PNA, or morpholino). The oligonucleotide (e.g., DNA, RNA, LNA, BNA, PNA, or morpholino) may have a length of about 10-30 nucleotides or about 10-20 nucleotides. In some embodiments, the oligonucleotide (e.g., DNA, RNA, LNA, BNA, PNA, or morpholino) for use in the tag or tether may have at least one end (e.g., the 3'- or 5'-end) modified for attachment to other sites of modification or to a solid substrate surface, including, for example, a bead. The end modifier may add a reactive functional group that can be used for conjugation. Examples of functional groups that can be added include, but are not limited to, amino, carboxyl, thiol, maleimide, aminooxy, and any combination thereof. The functional groups can be combined with spacers of different lengths (eg, C3, C9, C12, spacers 9 and 18) to add physical distance of the functional group from the end of the oligonucleotide sequence.

[0305] In some embodiments, the tag or tether may comprise or be a morpholino oligonucleotide. The morpholino oligonucleotide may have a length of about 10-30 nucleotides or about 10-20 nucleotides. The morpholino oligonucleotide may be modified or unmodified. For example, in some embodiments, the morpholino oligonucleotide may be modified at the 3' and / or 5' end of the oligonucleotide. Examples of 3' and / or 5' end modifications of the morpholino oligonucleotide include, but are not limited to, 3' affinity tags and functional groups for chemical conjugation (e.g., including 3'-biotin, 3'-primary amine, 3'-disulfide amide, 3'-pyridyldithio, and any combination thereof); 5' end modifications (e.g., including 5'-primary ammine, and / or 5'-dabsyl); modifications for click chemistry (e.g., including 3'-azide, 3'-alkyne, 5'-azide, 5'-alkyne), and any combination thereof.

[0306] In some embodiments, the tag or tether may further comprise a polymer linker, for example, to facilitate attachment to a detector, for example, a nanopore. Exemplary polymer linkers include, but are not limited to, polyethylene glycol (PEG). The polymer linker may have a molecular weight of about 500 Da to about 10 kDa, inclusive, or about 1 kDa to about 5 kDa, inclusive. The polymer linker (e.g., PEG) may be functionalized with different functional groups, including, for example, but not limited to, maleimide, NHS ester, dibenzocyclooctyne (DBCO), azide, biotin, amine, alkyne, aldehyde, and any combination thereof. In some embodiments, the tag or tether may also comprise a 1 kDa PEG having a 5'-maleimide group and a 3'-DBCO group. In some embodiments, the tag or tether may further comprise a 2 kDa PEG having a 5'-maleimide group and a 3'-DBCO group. In some embodiments, the tag or tether may further comprise a 3 kDa PEG having a 5'-maleimide group and a 3'-DBCO group. In some embodiments, the tag or tether may further comprise a 5 kDa PEG having a 5'-maleimide group and a 3'-DBCO group.

[0307] Other examples of tags or tethers include, but are not limited to, His tags, biotin or streptavidin, antibodies that bind to the analyte, aptamers that bind to the analyte, analyte binding domains such as DNA binding domains (including, for example, peptide zippers such as leucine zippers, single stranded DNA binding proteins (SSBs)), and any combination thereof.

[0308] The tag or tether may be attached to the exterior surface of the nanopore, e.g., on the cis side of the membrane, using any method known in the art. For example, one or more tags or tethers can be attached to the nanopore via one or more cysteines (cysteine ​​bonds), one or more primary amines such as lysines, one or more unnatural amino acids, one or more histidines (His tags), one or more biotins or streptavidins, one or more antibody-based tags, one or more enzymatic modifications of epitopes (e.g., including acetyltransferases), and any combination thereof. Suitable methods for making such modifications are well known in the art. Suitable unnatural amino acids include, but are not limited to, 4-azido-L-phenylalanine (Faz), and any one of the amino acids numbered 1-71 in Figure 1 of Liu CC and Schultz PG, Annu. Rev. Biochem., 2010, 79, 413-444.

[0309] In some embodiments where one or more tags or tethers are attached to the nanopore via cysteine ​​bond(s), one or more cysteines can be introduced by substitution into one or more of the monomers forming the nanopore. In some embodiments, the nanopore can be chemically modified by attachment of: (i) 4-phenylazomaleinanyl, 1.N-(2-hydroxyethyl)maleimide, N-cyclohexylmaleimide, 1.3-maleimidopropionic acid, 1.1-4-aminophenyl-1H-pyrrole, 2,5,dione, 1.1-4-hydroxyphenyl-1H-pyrrole, 2,5,dione, N-ethylmaleimide, N-methoxycarbonylmaleimide, N-methylmaleimide, N-methyl-2-phenylpropionyl ... Imide, N-tert-butylmaleimide, N-(2-aminoethyl)maleimide, 3-maleimide-proxyl, N-(4-chlorophenyl)maleimide, 1-[4-(dimethylamino)-3,5-dinitrophenyl]-1H-pyrrole-2,5-dione, N-[4-(2-benzimidazolyl)phenyl]maleimide, N-[4-(2-benzoxazolyl)phenyl]maleimide, N-(1-naphthyl)-maleimide, N-(2, 4-xylyl)maleimide, N-(2,4-difluorophenyl)maleimide, N-(3-chloro-para-tolyl)-maleimide, 1-(2-amino-ethyl)-pyrrole-2,5-dione hydrochloride, 1-cyclopentyl-3-methyl-2,5-dihydro-1H-pyrrole-2,5-dione, 1-(3-aminopropyl)-2,5-dihydro-1H-pyrrole-2,5-dione hydrochloride, 3-methyl-1-[2-oxo- 2-(piperazin-1-yl)ethyl]-2,5-dihydro-1H-pyrrole-2,5-dione hydrochloride, 1-benzyl-2,5-dihydro-1H-pyrrole-2,5-dione, 3-methyl-1-(3,3,3-trifluoropropyl)-2,5-dihydro-1H-pyrrole-2,5-dione, 1-[4-(methylamino)cyclohexyl]-2,5-dihydro-1H-pyrrole-2,5-dione trifluoroacetate, SMILES O=C1C=CC(=O)N1CC=2C=CN=CC2, SMILES O=C1C=CC(=O)N1CN2CCNCC2, 1-benzyl-3-methyl-2,5-dihydro-1H-pyrrole-2,(ii) maleimides including diabromomaleimides such as 5-dione, 1-(2-fluorophenyl)-3-methyl-2,5-dihydro-1H-pyrrole-2,5-dione, N-(4-phenoxyphenyl)maleimide, N-(4-nitrophenyl)maleimide, (ii) 3-(2-iodoacetamido)-proxyl, N-(cyclopropylmethyl)-2-iodoacetamide, 2-iodo-N-(2-phenylethyl)acetamide, 2-iodo-N-(2,2,2-trifluoroethyl)acetamide, N-(4-acetylphenyl)-2-iodoacetamide, (iii) iodoacetamides such as N-(4-(aminosulfonyl)phenyl)-2-iodoacetamide, N-(1,3-benzothiazol-2-yl)-2-iodoacetamide, N-(2,6-(diethylphenyl)-2-iodoacetamide, and N-(2-benzoyl-4-chlorophenyl)-2-iodoacetamide; (iv) N-(4-(acetylamino)phenyl)-2-bromoacetamide, N-(2-acetylphenyl)-2-bromoacetamide, 2-bromo-n-(2-cyanophenyl)acetamide, 2-bromo-N-(3-((trifluorophenyl)acetamide, etc.); N-(2-bromophenyl)acetamide, N-(2-benzoylphenyl)-2-bromoacetamide, 2-bromo-N-(4-fluorophenyl)-3-methylbutanamide, N-benzyl-2-bromo-N-phenylpropionamide, N-(2-bromo-butyl)-4-chloro-benzenesulfonamide, 2-bromo-N-methyl-N-phenylacetamide, 2-bromo-N-phenethyl-acetamide, 2-adamantan-1-yl-2-bromo-N-cyclohexyl-acetamide, 2-bromo-N-(2-methylphenyl)butanamide, mono (iv) bromoacetamides such as bromoacetanilide, (iv) disulfides such as aldrithiol-2, aldrithiol-4, isopropyl disulfide, 1-(isobutyldisulfanyl)-2-methylpropane, dibenzyl disulfide, 4-aminophenyl disulfide, 3-(2-pyridyldithio)propionic acid, 3-(2-pyridyldithio)propionic acid hydrazide, 3-(2-pyridyldithio)propionic acid N-succinimidyl ester, am6amPDP1-βCD, and (v) 4-phenylthiazole-2-thiol, perpaldo, 5,Thiols such as 6,7,8-tetrahydro-quinazoline-2-thiol.

[0310] In some embodiments, the tag or tether may be attached to the nanopore directly or via one or more linkers. The tag or tether may be attached to the nanopore using a hybrid linker as described in WO2010 / 086602. Alternatively, a peptide linker may be used. A peptide linker is an amino acid sequence. The length, flexibility, and hydrophilicity of the peptide linker are typically designed so as not to interfere with the function of the monomer and the pore. A preferred flexible peptide linker is a stretch of 2-20, e.g., 4, 6, 8, 10, or 16, serine, and / or glycine amino acids. More preferred flexible linkers include (SG)1, (SG)2, (SG)3, (SG)4, (SG)5, and (SG)8, where S is serine and G is glycine. A preferred rigid linker is a stretch of 2-30, e.g., 4, 6, 8, 16, or 24, proline amino acids. A more preferred rigid linker is one in which P is proline, (P) 12 Includes.

[0311] anchor In one embodiment, the polynucleotide or polynucleotide adaptor may include a membrane anchor or a transmembrane pore anchor. In one embodiment, the anchor assists in characterizing the target polynucleotide according to the methods disclosed herein. For example, the membrane anchor or transmembrane pore anchor may facilitate localization of the selected polynucleotide around the nanopore.

[0312] The anchor can be a polypeptide anchor and / or a hydrophobic anchor that can insert into the membrane. In one embodiment, the hydrophobic anchor is a lipid, a fatty acid, a sterol, a carbon nanotube, a polypeptide, a protein, or an amino acid, such as cholesterol, palmitate, or tocopherol. The anchor can include a thiol, biotin, or a surfactant.

[0313] In one aspect, the anchor can be biotin (for binding to streptavidin), amylose (for binding to maltose binding protein or fusion proteins), Ni-NTA (for binding to poly-histidine or poly-histidine tagged proteins), or a peptide (such as an antigen).

[0314] In one embodiment, the anchor may include one linker, or two, three, four or more linkers. Preferred linkers include, but are not limited to, polymers, such as polynucleotides, polyethylene glycol (PEG), polysaccharides, and polypeptides. These linkers may be linear, branched, or cyclic. For example, the linker may be a cyclic polynucleotide. The adaptor may hybridize to a complementary sequence on the cyclic polynucleotide linker. One or more anchors or one or more linkers may include a moiety that can be cleaved or degraded, such as a restriction site or a photolabile group. The linker is functionalized with a maleimide group to attach to a cysteine ​​residue of the protein. Suitable linkers are described in WO2010 / 086602.

[0315] In one embodiment, the anchor is cholesterol or a fatty acyl chain. Any fatty acyl chain having a length of 6 to 30 carbon atoms, such as, for example, hexadecanoic acid, can be used. Examples of suitable anchors and methods for attaching the anchors to the adaptors are disclosed in WO2012 / 164270 and WO2015 / 150786.

[0316] In another embodiment, the anchor may consist of or include a hydrophobic modification to the polynucleotide or polynucleotide adaptor. The hydrophobic modification may include a modified phosphate group contained within the polynucleotide or polynucleotide anchor. The hydrophobic modification may include, for example, phosphorothioates, such as charge-neutralized alkyl phosphorothioates (PPTs), which are described in Jones et al, J. Am. Chem. Soc. 2021, 143, 22, 8305, the entire contents of which are incorporated herein by reference. Suitable alkyl groups include, for example, C1-C alkyl groups, such as C2-C6 alkyl groups. 10 Alkyl groups include, for example, methyl, ethyl, propyl, butyl, pentyl, and hexyl groups. Incorporation of charge-neutralized alkyl-phosphorothioates into a polynucleotide allows the polynucleotide to engage hydrophobic regions such as lipid bilayers.

[0317] Detector In the methods provided herein, the polynucleotide moves relative to a detector, such as a nanopore. The detector can be selected from (i) a zero mode waveguide, (ii) a field effect transistor, optionally a nowire field effect transistor, (iii) an AFM tip, (iv) a nanotube, optionally a carbon nanotube, and (v) a nanopore. Preferably, the detector is a nanopore.

[0318] The polynucleotide may be characterized in any suitable manner in the methods provided herein. In one embodiment, the polynucleotide is characterized by detecting an ion flow or optical signal as the polynucleotide translocates relative to the nanopore, as described in more detail herein. The method is suitable for these and other methods of detecting polynucleotides.

[0319] In another non-limiting example, in one embodiment, polynucleotides are characterized by detecting by-products of polynucleotide processing reactions, such as sequencing by synthesis reactions. Thus, the method may include detecting products of sequential addition of (poly)nucleotides by an enzyme, such as a polymerase, to a nucleic acid strand. The product may be a change in one or more properties of the enzyme, such as the conformation of the enzyme. Thus, such a method may include subjecting an enzyme, such as a polymerase or reverse transcriptase, to a double-stranded polynucleotide under conditions such that template-dependent incorporation of nucleotide bases into the growing oligonucleotide strand causes conformational changes in the enzyme in response to successively encountered templates, incorporation of stranded nucleic acid bases and / or incorporation of template-specific natural or similar bases (i.e., incorporation events), detection of conformational changes in the enzyme in response to such incorporation events, and thereby detection of the sequence of the template strand. In such a method, the polynucleotide strand may be displaced according to the methods provided herein. Such a method may include detecting and / or measuring incorporation events using methods known to those skilled in the art, such as those described in US2017 / 0044605.

[0320] In another embodiment, the by-products may be labeled such that upon addition of a nucleotide to a synthetic nucleic acid strand complementary to a template strand, a phosphate-labeled species is released and the phosphate-labeled species is detected using a detector as described herein. Polynucleotides characterized in this manner may be translocated according to the methods herein. Suitable labels may be optical labels that are detected using a nanopore or zero-mode waveguide, or by Raman spectroscopy or other detectors. Suitable labels may be non-optical labels that are detected using a nanopore or other detectors.

[0321] In another approach, the nucleoside phosphates (nucleotides) are not labeled, and natural by-product species are detected upon addition of the nucleotide to a synthetic nucleic acid strand complementary to a template strand. Suitable detectors can be ion-sensitive field effect transistors, or other detectors.

[0322] These and other detection methods are suitable for use in the methods described herein. Any suitable measurement can be made using the detector as the polynucleotides migrate relative to the detector.

[0323] Nanopore In embodiments of the invention in which the detector is a nanopore, any suitable nanopore can be used, hi one embodiment, the nanopore is a transmembrane pore.

[0324] A transmembrane pore is a structure that traverses a membrane to some extent. It allows hydrated ions driven by an applied electric potential to flow across or within the membrane. A transmembrane pore typically traverses the entire membrane, allowing hydrated ions to flow from one side of the membrane to the other side of the membrane. However, a transmembrane pore does not have to traverse the membrane. It may be closed at one end. For example, a pore may be a well, gap, channel, trench, or slit in a membrane along or into which hydrated ions can flow.

[0325] In the methods provided herein, the nanopore typically has a first opening and a second opening. The first opening is typically a cis opening and the second opening is typically a trans opening. However, in some embodiments, the first opening is a trans opening and the second opening is a cis opening. The motor protein used in the methods provided herein is typically provided at the first opening of the nanopore, thus controlling the movement of the target polynucleotide in a direction from the second opening of the nanopore toward the first opening of the nanopore.

[0326] Any transmembrane pore may be used in the methods provided herein. The pore may be biological or artificial. Suitable pores include, but are not limited to, protein pores, polynucleotide pores, and solid-state pores.

[0327] In one embodiment, the solid-state pore may comprise a nanochannel, hi some embodiments, the solid-state pore is a pore disclosed in WO2003 / 003446, WO2009 / 020682 or WO2016 / 187519, each of which is incorporated by reference in their entirety.

[0328] In one embodiment, the pore may be a DNA origami pore (Langecker et al., Science, 2012;338:932-936). Suitable DNA origami pores are disclosed in WO2013 / 083983, WO2018 / 011603, and WO2020 / 025974, each of which is incorporated by reference in its entirety.

[0329] In one embodiment, the nanopore is a scaffolded polypeptide nanopore. In some embodiments, the pore is a scaffolded polypeptide nanopore as disclosed in WO2020 / 025909 or WO2020 / 074399, each of which is incorporated by reference in their entirety.

[0330] In one embodiment, the nanopore is a transmembrane protein pore. A transmembrane protein pore is a polypeptide or an assembly of polypeptides that allows hydrated ions, such as polynucleotides, to flow from one side of a membrane to the other side of the membrane. In the methods provided herein, the transmembrane protein pore can form a pore that allows hydrated ions driven by an applied potential to flow from one side of a membrane to the other side. The transmembrane protein pore preferably allows polynucleotides to flow from one side of a membrane, such as a triblock copolymer membrane, to the other side. The transmembrane protein pore allows polynucleotides to translocate through the pore.

[0331] In one embodiment, the nanopore is a transmembrane protein pore that is monomeric or oligomeric. The pore is preferably composed of several repeating subunits, such as at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, or at least 16 subunits. The pore is preferably a hexameric, heptameric, octameric, or nanomeric pore. The pore may be a homo-oligomer or a hetero-oligomer.

[0332] In one embodiment, a transmembrane protein pore comprises a barrel or channel through which ions can flow. The subunits of the pore typically surround a central axis and contribute strands to a transmembrane β-barrel or channel or a transmembrane α-helical bundle or channel.

[0333] Typically, the barrel or channel of the transmembrane protein pore comprises amino acids that facilitate interaction with an analyte, such as a target polynucleotide (as described herein). These amino acids are preferably located near the constriction of the barrel or channel. The transmembrane protein pore typically comprises one or more positively charged amino acids, such as arginine, lysine, or histidine, or aromatic amino acids, such as tyrosine or tryptophan. These amino acids typically facilitate interaction between the pore and a nucleotide, polynucleotide, or nucleic acid.

[0334] In one embodiment, the nanopore is a transmembrane protein pore derived from a β-barrel pore or an α-helix bundle pore. A β-barrel pore comprises a barrel or channel formed from β-strands. Suitable β-barrel pores include, but are not limited to, β-toxins such as α-hemolysin, anthrax toxin, and leukocidin, as well as bacterial outer membrane proteins / porins such as Mycobacterium smegmatis porins (Msp), e.g., MspA, MspB, MspC, or MspD, CsgG, outer membrane porin F (OmpF), outer membrane porin G (OmpG), outer membrane phospholipase A, and Neisseria autotransporter lipoprotein (NalP), and other pores such as lysenin. An α-helix bundle pore comprises a barrel or channel formed from α-helices. Suitable α-helix bundle pores include, but are not limited to, inner membrane proteins and α-outer membrane proteins, e.g., WZA and ClyA toxins.

[0335] In one embodiment, the nanopore is a transmembrane pore derived from or based on Msp, α-hemolysin (α-HL), lysenin, CsgG, ClyA, Sp1, or the hemolytic protein Fragaceatoxin C (FraC).

[0336] In one embodiment, the nanopore is a transmembrane protein pore derived from CsgG, for example CsgG from E. coli strain K-12 substrain MC4100. Such pores are oligomeric and typically comprise 7, 8, 9, or 10 monomers derived from CsgG. The pore may be a homo-oligomeric pore derived from CsgG comprising identical monomers. Alternatively, the pore may be a hetero-oligomeric pore derived from CsgG comprising at least one monomer that is different from the others. Examples of suitable pores derived from CsgG are disclosed in WO2016 / 034591, WO2017 / 149316, WO2017 / 149317, WO2017 / 149318, and WO2019 / 002893, which are incorporated herein by reference in their entirety.

[0337] In one embodiment, the nanopore is a transmembrane pore derived from lysenin. Examples of suitable pores derived from lysenin are disclosed in WO2013 / 153359, which is incorporated herein by reference in its entirety.

[0338] In one embodiment, the nanopore is a transmembrane pore derived from or based on α-hemolysin (α-HL). The wild-type α-hemolysin pore is formed from seven identical monomers or subunits (i.e., it is a heptamer). The α-hemolysin pore can be α-hemolysin-NN or a mutant form thereof. The mutant form preferably contains N residues at positions E111 and K147.

[0339] In one embodiment, the nanopore is a transmembrane protein pore derived from an Msp, such as from MspA. An example of a suitable pore derived from MspA is disclosed in WO2012 / 107778.

[0340] In one embodiment, the nanopore is a transmembrane pore derived from or based on ClyA. Examples of suitable pores derived from ClyA are disclosed in Soskine et al., Nano Letters 2012 12(9), 4895-4900, WO2014 / 153625, and WO2017 / 098322, each of which is incorporated herein by reference.

[0341] In one embodiment, the nanopore is a transmembrane pore derived from Phi29. Examples of suitable pores derived from Phi29 are disclosed in Wendell et al., Nature Nanotech 4, 765-772 (2009), WO2010 / 062697, WO2019 / 157365 and WO2019 / 157424, each of which is incorporated herein by reference.

[0342] In some embodiments, the nanopore is selected from M-ring proteins, perforin-2, PlyAB (pleurotolysin), SpoIIIAG, VirB7, type II secretion system protein D, GspD, InvG, PilQ, pentraxins, and portal proteins, including T4, T7, P23_45, G20c, and Phi29 nanopores.

[0343] In one embodiment, the nanopore is a transmembrane pore derived from or based on a bacterium of the Rhodococcus species, such as Rhodococcus corynebacteroides or Rhodococcus ruber, e.g., PorARr, PorBRr or PorARc. Examples of such pores are described in Piselli et al., Eur Biophys J 51, 309-323 (2022).

[0344] film In the disclosed methods, the detector is typically a nanopore present in the membrane. Any suitable membrane can be used.

[0345] The membrane is preferably an amphiphilic layer. An amphiphilic layer is a layer formed from amphiphilic molecules, such as phospholipids, that have both hydrophilic and lipophilic properties. The amphiphilic molecules can be synthetic or naturally occurring. Non-naturally occurring amphiphiles and amphiphiles that form monolayers are known in the art, including, for example, block copolymers (Gonzalez-Perez et al., Langmuir, 2009, 25, 10447-10450). Block copolymers are polymeric materials in which two or more monomeric subunits are polymerized together to create a single polymer chain. Block copolymers typically have properties contributed by each monomeric subunit. However, block copolymers can have unique properties that polymers formed from individual subunits do not have. Block copolymers can be engineered such that one of the monomeric subunits is hydrophobic (i.e., lipophilic), while the other subunit(s) are hydrophilic in aqueous media. In this case, the block copolymer may have amphiphilic properties and may form structures that mimic biological membranes. The block copolymer may be diblock (consisting of two monomer subunits), but may be constructed from three or more monomer subunits to form more complex arrangements that behave as amphiphiles. The copolymer may be a triblock, tetrablock, or pentablock copolymer. The membrane is preferably a triblock copolymer membrane.

[0346] Archaeal bipolar tetraether lipids are naturally occurring lipids that are constructed such that the lipids form monolayer membranes. These lipids are commonly found in extremophilic, thermophilic, halophilic, and acidophilic bacteria that survive in harsh biological environments. Their stability is believed to derive from the fusogenic nature of the final bilayer. It is straightforward to construct block copolymers that mimic these biological entities by creating triblock polymers with the general motif hydrophilic-hydrophobic-hydrophilic. This material forms monomeric membranes that behave similarly to lipid bilayers and can encompass a wide range of phase behaviors, from vesicles to lamellar membranes. Membranes formed from these triblock copolymers have several advantages over biological lipid membranes. Because the triblock copolymers are synthetic, their exact structure can be carefully controlled to provide the correct chain length and properties required to form membranes and interact with pores and other proteins.

[0347] Block copolymers may also be constructed from subunits that are not classified as lipid submaterials, for example, hydrophobic polymers may be made from siloxanes or other non-hydrocarbon monomers. The hydrophilic subsections of the block copolymers may also have low protein binding properties, allowing for the creation of membranes that are highly resistant when exposed to live biological samples. The head group units may also be derived from non-classified lipid head groups.

[0348] Triblock copolymer membranes also have increased mechanical and environmental stability compared to biological lipid membranes, e.g., much higher operating temperature or pH ranges. The synthetic nature of block copolymers provides a platform for customizing polymer-based membranes for a wide range of applications.

[0349] In some embodiments, the membrane is one of the membranes disclosed in WO 2014 / 064443 or WO 2014 / 064444.

[0350] The amphiphilic molecules may be chemically modified or functionalized to facilitate coupling of polynucleotides. The amphiphilic layer may be a monolayer or a bilayer. The amphiphilic layer is typically planar. The amphiphilic layer may be curved. The amphiphilic layer may be supported.

[0351] Amphiphilic membranes typically have a molecular weight of approximately 10 -8 cm s -1 They are naturally mobile, essentially acting as a two-dimensional fluid with lipid diffusion rates of 0.1 - 0.2 nm, which means that the pore and the coupled polynucleotide can typically move within the amphiphilic membrane.

[0352] The membrane can be a lipid bilayer. Lipid bilayers are models of cell membranes and serve as an excellent platform for a wide range of experimental studies. For example, lipid bilayers can be used for in vitro investigation of membrane proteins by single channel recording. Alternatively, lipid bilayers can be used as biosensors to detect the presence of a wide range of substances. The lipid bilayer can be any lipid bilayer. Suitable lipid bilayers include, but are not limited to, planar lipid bilayers, supported bilayers, or liposomes. The lipid bilayer is preferably a planar lipid bilayer. Suitable lipid bilayers are disclosed in WO2008 / 102121, WO2009 / 077734, and WO2006 / 100484.

[0353] Methods for forming lipid bilayers are known in the art. Lipid bilayers are generally formed by the method of Montal and Mueller (Proc. Natl. Acad. Sci. USA., 1972; 69: 3561-3566), in which a lipid monolayer is supported on the aqueous solution / air interface past both sides of a hole perpendicular to the interface. Lipid is usually added to the surface of an aqueous electrolyte solution by first dissolving it in an organic solvent, and then evaporating a small amount of solvent on the interface of the aqueous solution on both sides of the opening. As the organic solvent evaporates, the solution / air interfaces on both sides of the opening physically move up and down across the opening until a bilayer is formed. Planar lipid bilayers can be formed in a membrane across an opening or in a recess across an opening.

[0354] The Montal & Mueller method is popular because it is a cost-effective and relatively simple method for forming good quality lipid bilayers suitable for protein pore insertion. Other common methods of bilayer formation include tip-dipping, painting bilayers, and patch clamping of liposome bilayers.

[0355] Tip-dipping bilayer formation involves contacting an aperture (e.g., a pipette tip) onto the surface of a test solution carrying a monolayer of lipids. Again, the lipid monolayer is generated at the solution / air interface by first evaporating a small amount of lipid dissolved in an organic solvent at the solution surface. The bilayer is then formed by the Langmuir-Schaefer method, which requires mechanical automation to move the aperture relative to the solution surface.

[0356] For bilayer coating, a small amount of lipid dissolved in an organic solvent is applied directly to an aperture immersed in the test aqueous solution. The lipid solution is spread thinly across the aperture using a paintbrush or equivalent. Thinning the solvent leads to the formation of a lipid bilayer. However, it is difficult to completely remove the solvent from the bilayer, and as a result, bilayers formed by this method are less stable and more prone to noise during electrochemical measurements.

[0357] Patch clamping is commonly used in the study of biological cell membranes. The cell membrane is clamped to the end of a pipette by suction, so that a patch of membrane is attached across the aperture. The method has been adapted to produce lipid bilayers by clamping and then rupturing liposomes to seal the lipid bilayer across the aperture of the pipette. The method requires the production of stable large unilamellar liposomes and a small aperture in the material with a glass surface.

[0358] Liposomes can be formed by sonication, extrusion, or the Mozafari method (Colas et al. (2007) Micron 38:841-847).

[0359] In some embodiments, the lipid bilayer is formed as described in WO 2009 / 077734. Advantageously, in this method, the lipid bilayer is formed from dry lipids. In the most preferred embodiment, the lipid bilayer is formed across the aperture as described in WO 2009 / 077734.

[0360] A lipid bilayer is formed from two opposing layers of lipids. These two lipid layers are arranged so that their hydrophobic tail groups face each other to form a hydrophobic interior. The hydrophilic head groups of the lipids face outward toward the aqueous environment on both sides of the bilayer. Bilayers can exist in several lipid phases, including but not limited to liquid disordered phases (fluid lamellar), liquid ordered phases, solid ordered phases (lamellar gel phase, interdigitated gel phase), and planar bilayer crystals (lamellar subgel phase, lamellar crystal phase).

[0361] Any lipid composition that forms a lipid bilayer can be used. The lipid composition is selected to form a lipid bilayer with the required properties, such as surface charge, ability to support membrane proteins, packing density, or mechanical properties. The lipid composition can include one or more different lipids. For example, the lipid composition can include up to 100 lipids. The lipid composition preferably includes 1 to 10 lipids. The lipid composition can include naturally occurring lipids and / or artificial lipids.

[0362] A lipid typically comprises a head group, an interface moiety, and two hydrophobic tail groups, which may be the same or different. Suitable head groups include, but are not limited to, neutral head groups, such as diacylglyceride (DG) and ceramide (CM), zwitterionic head groups, such as phosphatidylcholine (PC), phosphatidylethanolamine (PE), and sphingomyelin (SM), negatively charged head groups, such as phosphatidylglycerol (PG), phosphatidylserine (PS), phosphatidylinositol (PI), phosphatidic acid (PA), and cardiolipin (CA), and positively charged head groups, such as trimethylammonium-propane (TAP). Suitable interface moieties include, but are not limited to, naturally occurring interface moieties, such as glycerol-based moieties or ceramide-based moieties. Suitable hydrophobic tail groups include, but are not limited to, saturated hydrocarbon chains such as lauric acid (n-Dodecanolic acid), myristic acid (n-Tetradecononic acid), palmitic acid (n-Hexadecanoic acid), stearic acid (n-Octadecanoic acid), and arachidic acid (n-Eicosanoic acid), unsaturated hydrocarbon chains such as oleic acid (cis-9-Octadecanoic acid), and branched hydrocarbon chains such as phytanoyl. The length of the chain and the position and number of double bonds in the unsaturated hydrocarbon chains can vary. The length of the chain and the position and number of branches, such as methyl groups in the branched hydrocarbon chains, can vary. The hydrophobic tail group can be linked to the interface moiety as an ether or ester. The lipid can be a mycolic acid.

[0363] Lipids can also be chemically modified. The head group or tail group of lipids can be chemically modified. Suitable lipids with chemically modified head groups include, but are not limited to, PEG-modified lipids, such as 1,2-diacyl-sn-glycero-3-phosphoethanolamine-N-[methoxy(polyethylene glycol)-2000], functionalized PEG lipids, such as 1,2-distearoyl-sn-glycero-3-phosphoethanolamine-N-[biotinyl(polyethylene glycol)2000], and lipids modified for conjugation, such as 1,2-dioleoyl-sn-glycero-3-phosphoethanolamine-N-(succinyl) and 1,2-dipalmitoyl-sn-glycero-3-phosphoethanolamine-N-(biotinyl). Suitable lipids in which the tail group is chemically modified include, but are not limited to, polymerizable lipids such as 1,2-bis(10,12-tricosadiynoyl)-sn-glycero-3-phosphocholine, fluorinated lipids such as 1-palmitoyl-2-(16-fluoropalmitoyl)-sn-glycero-3-phosphocholine, deuterated lipids such as 1,2-dipalmitoyl-D62-sn-glycero-3-phosphocholine, and ether-linked lipids such as 1,2-di-O-phytanyl-sn-glycero-3-phosphocholine. Lipids may be chemically modified or functionalized to facilitate coupling of polynucleotides.

[0364] Amphiphilic layer, for example, lipid composition, typically contains one or more additives that will affect the properties of the layer.Suitable additives include, but are not limited to, fatty acid, for example, palmitic acid, myristic acid, and oleic acid, fatty alcohol, for example, palmitic alcohol, myristic alcohol, and oleic alcohol, sterol, for example, cholesterol, ergosterol, lanosterol, sitosterol, and stigmasterol, lysophospholipid, for example, 1-acyl-2-hydroxy-sn-glycero-3-phosphocholine, and ceramide.

[0365] In another embodiment, the membrane comprises a solid-state layer. The solid-state layer can be formed from both organic and inorganic materials, including but not limited to microelectronic materials, insulating materials such as Si3N4, A12O3, and SiO, organic and inorganic polymers such as polyamides, plastics such as Teflon®, or elastomers such as two-component addition cure silicone rubber, and glass. The solid-state layer can be formed from graphene. A suitable graphene layer is disclosed in WO2009 / 035647. When the membrane comprises a solid-state layer, the pores are typically present in the amphiphilic membrane or layer contained within the solid-state layer, for example, within holes, wells, gaps, channels, grooves, or slits within the solid-state layer. Those skilled in the art can prepare suitable solid / amphiphilic hybrid systems. Suitable systems are disclosed in WO2009 / 020682 and WO2012 / 005857. Any of the amphiphilic membranes or layers discussed above can be used.

[0366] The methods disclosed herein are typically carried out using (i) an artificial amphiphilic layer comprising a pore, (ii) an isolated naturally occurring lipid bilayer comprising a pore, or (iii) a cell with a pore inserted therein. The methods are typically carried out using an artificial amphiphilic layer, such as an artificial triblock copolymer layer. The layer may contain other transmembrane and / or intramembrane proteins, as well as other molecules, in addition to the pore. Suitable equipment and conditions are discussed below. The methods of the invention are typically carried out in vitro.

[0367] General method As noted above, the methods provided herein can be operated using any suitable detector, and therefore any suitable device for detecting polynucleotides.

[0368] The method may be carried out in some embodiments using any device suitable for sensing transmembrane pores. For example, the device may include a chamber containing an aqueous solution and a barrier separating the chamber into two sections. The barrier may have an opening in which a membrane containing a transmembrane pore is formed. The transmembrane pore is described herein.

[0369] The method may be carried out using the devices described in WO2008 / 102120, WO2010 / 122293, or WO00 / 28312. Briefly, the binding of a molecule (e.g., a target polynucleotide) within the channel of the pore affects the open channel ion flow through the pore, which is the essence of "molecular sensing" of the pore channel. The variation in the open channel ion flow can be measured using a suitable measurement technique by changing the current. The degree of reduction in the ion flow, measured by the reduction in the current, is related to the size of the obstacle in or near the pore. Thus, the binding of a molecule of interest (e.g., a target polynucleotide) within or near the pore provides a detectable and measurable event, thereby forming the basis of a "biological sensor". Detecting the presence of a biomolecule finds applications in personalized drug development, medicine, diagnostics, life science research, environmental monitoring, and the security and / or defense industries.

[0370] When used to characterize polynucleotides, the presence, absence, or one or more characteristics of a target polynucleotide are determined. The method can be for determining the presence, absence, or one or more characteristics of at least one target polynucleotide. The method can relate to determining the presence, absence, or one or more characteristics of two or more target polynucleotides. The method can include determining the presence, absence, or one or more characteristics of any number of target polynucleotides, for example, 2, 5, 10, 15, 20, 30, 40, 50, 100 or more target polynucleotides. Any number of characteristics of one or more target polynucleotides can be determined, for example, 1, 2, 3, 4, 5, 10 or more characteristics. Characteristics that can be detected by the methods provided herein include the identity or sequence of the polynucleotide, the length of the polynucleotide, whether the polynucleotide is modified, etc. In some embodiments, the methods provided herein are methods for sequencing a target polynucleotide. In some embodiments, the polynucleotide sequence can be determined in real time by aligning real-time signals or base calling to a known reference. Exemplary methods for determining polynucleotide sequences are described in WO2016 / 059427, which is incorporated herein by reference.

[0371] When used to characterize polynucleotides, the method may include measuring the flow of ionic current through the pore, typically by measuring electrical current. Alternatively, the flow of ions through the pore may be measured optically, as disclosed by Heron et al: J.Am.Chem.Soc.9 Vol.131, No.5,2009. Thus, the device may also comprise an electrical circuit capable of applying an electrical potential and measuring an electrical signal across the membrane and the pore. The characterization method may be carried out using a patch clamp or a voltage clamp. The characterization method preferably involves the use of a voltage clamp.

[0372] The method can include measuring an optical signal, as described in Chen et al., Nature Communications (2018) 9:1733, the entire contents of which are incorporated herein by reference. For example, nanopores, such as optically engineered nanopore structures (such as plasmonic nanoslits), can be used to locally enable single molecule surface-enhanced Raman spectroscopy (SERS) and characterize polynucleotides by direct Raman spectroscopic detection.

[0373] The method may be performed on silicon-based well arrays, each array comprising 128, 256, 512, 1024, 2000, 3000, 4000, 6000, 10000, 12000, 15000 or more wells.

[0374] The method may include measuring the current through the pore. The method is typically carried out with a voltage applied across the membrane and pore. The voltage used is typically between +2V and -2V, typically between -400mV and +400mV. The voltage used is preferably in a range having a lower limit selected from -400mV, -300mV, -200mV, -150mV, -100mV, -50mV, -20mV, and 0mV, and an upper limit independently selected from +10mV, +20mV, +50mV, +100mV, +150mV, +200mV, +300mV, and +400mV. The voltage used is more preferably in the range of 100mV to 240mV, most preferably in the range of 120mV to 220mV. By using an increased applied potential, it is possible to increase the discrimination between different nucleotides per pore.

[0375] In some embodiments, the methods include providing conditions to promote debinding of a target polynucleotide from a polynucleotide binding site of a motor protein and / or conditions to retard rebinding of a target polynucleotide to a polynucleotide binding site of a motor protein.

[0376] The method is typically carried out in the presence of any charge carrier, such as a metal salt, e.g., an alkali metal salt, a halogen salt, e.g., a chloride salt, e.g., an alkali metal chloride salt. The charge carrier may include an ionic liquid or an organic salt, e.g., tetramethylammonium chloride, trimethylphenylammonium chloride, phenyltrimethylammonium chloride, or 1-ethyl-3-methylimidazolium chloride. In the exemplary device discussed above, the salt is present in an aqueous solution in the chamber. Potassium chloride (KCl), sodium chloride (NaCl), or cesium chloride (CsCl) are typically used. KCl is preferred. The salt may be an alkaline earth metal salt, such as calcium chloride (CaCl2). The salt concentration may be saturated. The salt concentration may be 3M or less, typically 0.1-2.5M, 0.3-1.9M, 0.5-1.8M, 0.7-1.7M, 0.9-1.6M, or 1M-1.4M. The salt concentration is preferably between 150 mM and 1 M. The characterization method is preferably carried out using a salt concentration of at least 0.3 M, such as at least 0.4 M, at least 0.5 M, at least 0.6 M, at least 0.8 M, at least 1.0 M, at least 1.5 M, at least 2.0 M, at least 2.5 M, or at least 3.0 M. High salt concentrations provide a high signal to noise ratio, allowing identification of currents indicative of binding / no binding against the background of normal current fluctuations.

[0377] In some embodiments, providing the conditions comprises providing a salt concentration to increase the rate of dissociation of the target polynucleotide from the polynucleotide binding site of the motor protein. In some embodiments, providing the conditions comprises providing a salt concentration to decrease the rate of rebinding of the target polynucleotide to the polynucleotide binding site of the motor protein. It is within the skill of the art in the art in view of the disclosure herein to determine suitable salt concentrations to promote debinding and / or delay rebinding of the target polynucleotide from the polynucleotide binding site of the motor protein.

[0378] In some embodiments, providing the conditions comprises providing an osmolality to increase the rate of dissociation of the target polynucleotide from the polynucleotide binding site of the motor protein. In some embodiments, providing the conditions comprises providing an osmolality to decrease the rate of rebinding of the target polynucleotide to the polynucleotide binding site of the motor protein. It is within the ability of one of skill in the art in view of the disclosure herein to determine a suitable osmolality to promote debinding and / or delay rebinding of the target polynucleotide from the polynucleotide binding site of the motor protein.

[0379] The method is usually carried out in the presence of a buffer. In the exemplary device discussed above, the buffer is present in the aqueous solution in the chamber. Any suitable buffer can be used. Typically, the buffer is HEPES. Another suitable buffer is Tris-HCl buffer. The method is typically carried out at a pH of 4.0-12.0, 4.5-10.0, 5.0-9.0, 5.5-8.8, 6.0-8.7, or 7.0-8.8, or 7.5-8.5. The pH used is preferably about 7.5.

[0380] The method may be carried out at 0° C. to 100° C., 15° C. to 95° C., 16° C. to 90° C., 17° C. to 85° C., 18° C. to 80° C., 19° C. to 70° C., or 20° C. to 60° C. The method may typically be carried out at room temperature. The method is optionally carried out at a temperature that supports enzyme function, for example, about 37° C.

[0381] In some embodiments, providing the conditions comprises increasing the temperature to increase the rate of dissociation of the target polynucleotide from the polynucleotide binding site of the motor protein. In some embodiments, providing the conditions comprises increasing the temperature to decrease the rate of rebinding of the target polynucleotide to the polynucleotide binding site of the motor protein. Without being bound by theory in any way, the inventors believe that increasing the temperature may facilitate re-reading, for example by increasing the rate of dissociation of the motor protein from the polynucleotide. It is within the skill of the art in the art in view of the disclosure herein to determine a suitable temperature for facilitating debinding and / or delaying rebinding of the target polynucleotide from the polynucleotide binding site of the motor protein.

[0382] Examples of providing conditions to promote re-reading by providing a temperature to promote re-reading are provided herein, see, e.g., Example 4. In some embodiments, providing conditions to promote debinding of a target polynucleotide from a polynucleotide binding site of a motor protein and / or to retard rebinding of a target polynucleotide to a polynucleotide binding site of a motor protein can include providing a temperature of about 20°C to about 50°C, e.g., about 30°C to about 45°C, e.g., about 34°C to about 40°C, e.g., about 31, 32, 33, 34, 35, 36, 37, 38, or 39°C.

[0383] Polynucleotide Adapters Also provided is a polynucleotide adapter comprising a motor protein. It will be understood that any of the polynucleotide adapters disclosed herein can be applied to the method embodiments discussed herein and above.

[0384] In one embodiment, provided herein is a polynucleotide adaptor having a first end comprising a leader and a second end comprising an attachment point for attachment to a polynucleotide analyte at the first end of the polynucleotide analyte, the polynucleotide adaptor comprising a motor protein stalled thereon in an orientation for processing the adaptor in a direction from the second end to the first end.

[0385] In one embodiment, the polynucleotide adaptor is a polynucleotide adaptor described in more detail herein. In one embodiment, the motor protein is a motor protein described herein.

[0386] The motor protein is oriented to process the polynucleotide adaptor in a direction away from the attachment point on the adaptor for binding to the double-stranded polynucleotide, i.e., toward the leader. The motor protein can be oriented on the polynucleotide adaptor to control the movement of the target polynucleotide in a trans to cis direction. The motor protein can be oriented on the polynucleotide adaptor to control the movement of the target polynucleotide toward a detector, such as a nanopore, in a direction toward the motor protein, i.e., out of the detector, e.g., out of the nanopore, as described in more detail herein.

[0387] In some embodiments, the polynucleotide adaptor comprises a stall moiety as described herein. In some embodiments, the polynucleotide adaptor comprises a stop moiety as described herein.

[0388] kit Kits are also provided that include the polynucleotide adaptor and the motor protein. It will be understood that any of the polynucleotide adaptors disclosed herein can be applied to the kit embodiments discussed herein and above.

[0389] In some embodiments, provided herein is a kit comprising a first adaptor having a first end comprising a leader and a second end comprising an attachment point for attachment to a polynucleotide analyte at the first end of the polynucleotide analyte, the first adaptor comprising a motor protein stalled thereon in an orientation for processing the adaptor from the second end towards the first end; and a second adaptor comprising (i) an attachment point for attachment to a polynucleotide analyte at the second end of the polynucleotide analyte and (ii) a blocking moiety suitable for preventing the motor protein of the first adaptor from disassociating from the polynucleotide analyte when the first adaptor is attached to the polynucleotide analyte.

[0390] system Also provided is a system comprising a polynucleotide adaptor, a motor protein, and a nanopore. It will be understood that any of the polynucleotide adaptors disclosed herein can be applied to the system embodiments discussed herein and above.

[0391] In one embodiment, a system for characterizing a target double-stranded polynucleotide is provided, the system comprising: - a polynucleotide adaptor having a first end comprising a leader and a second end comprising an attachment point for attachment to a polynucleotide analyte at the first end of the polynucleotide analyte, the polynucleotide adaptor comprising a motor protein stalled thereon in an orientation for processing the adaptor in a direction from the second end to the first end; a nanopore for characterizing a target polynucleotide as it translocates relative to the nanopore; and a motor protein for controlling the movement of the target polynucleotide relative to the nanopore.

[0392] In one embodiment, the polynucleotide adaptor is a polynucleotide adaptor as described in more detail herein. In one embodiment, the motor protein is a motor protein as described herein. In one embodiment, the nanopore is a nanopore as described herein. The system may further include a membrane, a controller, etc.

[0393] The systems and kits disclosed herein may be configured for use with algorithms, also provided herein, adapted to run on a computer system. The algorithm may be adapted to detect information characteristic of a polynucleotide (e.g., characteristic of a sequence of the polynucleotide) and selectively process signals obtained as the polynucleotide moves relative to the nanopore. Systems are also provided that include computing means configured to detect information characteristic of a polynucleotide (e.g., characteristic of a sequence of the polynucleotide) and selectively process signals obtained as the polynucleotide moves relative to the nanopore. In some embodiments, the system includes receiving means for receiving data from the detection of the polynucleotide, processing means for processing signals obtained as the polynucleotide moves relative to the nanopore, and output means for outputting the characterization information thus obtained.

[0394] Although specific embodiments, specific configurations, and materials and / or molecules are discussed herein for the method according to the invention, it should be understood that various changes or modifications in form and details may be made without departing from the scope and spirit of the invention. The following examples are provided to better illustrate specific embodiments and should not be considered as limiting the present application. The present application is limited only by the scope of the claims. EXAMPLES

[0395] Example 1 This example demonstrates the controlled translocation of a DNA polynucleotide strand through a nanopore using a DNA motor that unwinds dsDNA while translocating 5'-3' of ssDNA. The DNA motor was first stalled on a Y adaptor linked to the polynucleotide. The Y adaptor contained an oligonucleotide with a leader with 30 3'-terminal C3 spacer residues. The polynucleotide translocated through the nanopore in distinct stages: (1) an enzyme-free stage, in which the 3' end of the polynucleotide was captured by the nanopore, which translocated and separated the two strands under a positive applied potential until it reached the DNA motor, which was stalled at the 5' end; (2) an "de-stalled" stage, in which the DNA motor initially failed to overcome the stall under the positive bias but was activated ("de-stalled") by applying a reverse potential; (3) a DNA motor-controlled stage, in which the motor began to translocate the DNA 5'-3' out of the nanopore against the applied potential; (4) upon reaching the end of the polynucleotide, a certain level of blockage was observed, which could be removed by reversing the potential to eject the strand, as well as occasionally (5) under the force of the applied sequencing potential, the enzyme spontaneously slipped backwards and re-bound to the upstream DNA, repeating from step (3).

[0396] The Y adaptors were prepared by annealing DNA oligonucleotides (SEQ ID NO: 17, SEQ ID NO: 22, SEQ ID NO: 19 and SEQ ID NO: 21). A DNA motor (Dda helicase) was loaded onto the adaptors. The oligonucleotide SEQ ID NO: 22 contained the C3 spacer residue described above.

[0397] A seven-fragment DNA library was obtained by digestion of bacteriophage lambda DNA with SnaBI and BamHI restriction enzymes, and end-repaired and dA-tailed by the NEBNext End Repair and NEBNext dA Tailing Modules (New England Biolabs (NEB)), generating 3′ dA overhangs on both ends of each fragment.

[0398] The seven fragment DNA library was ligated to the dA ends of the Y adaptors using LNB and T4 DNA ligase (NEB) from the Oxford Nanopore Technologies sequencing kit (LSK-SQK109). Samples were purified using Agencourt AMPure XP (Beckman Coulter) beads and washed twice with LFB from the Oxford Nanopore Technologies sequencing kit (LSK-SQK109). The ligated substrate was eluted in 10 mM Tris-Cl, 50 mM NaCl (pH 8.0) to obtain the "DNA library".

[0399] Electrical measurements were taken with a FLO-MIN106 MinION flow cell from Oxford Nanopore Technologies and a MinION Mk1b. To 1170 μL FB (from the Oxford Nanopore Technologies sequencing kit (SQK-LSK109)), 30 μL of FLT was added to obtain the Tether Mix. 800 μL of Tether Mix was run through the system, followed by a 5 minute wait and an additional 200 μL of Tether Mix with the SpotON port open. 37.5 μL of SQB from the Oxford Nanopore Technologies sequencing kit (SQK-LSK109), 15 μL of DNA library, and 22.5 μL of LB from the Oxford Nanopore Technologies sequencing kit (SQK-LSK109) were mixed to obtain the "Sequencing Mix". 75 μL of the Sequencing Mix was added to the MinION flow cell via the SpotON flow cell port.

[0400] DNA libraries were run using a custom sequencing script that controlled the applied potentials as follows: 55 s capture phase (+120 mV), 5 s destall phase (-20 mV), 55 s sequencing (+120 mV), ejection phase (0 mV, 1 s; -120 mV, 3 s). This sequence of applied potentials was repeated multiple times. Raw data was collected into bulk FAST5 files using MinKNOW software (Oxford Nanopore Technologies).

[0401] Figure 6a shows a schematic of the experiment. The experiment includes a "reread" step (RR) where the enzyme unbinds and returns to its previous position on the DNA strand from the 3'C3 (non-DNA) leader (E) and translocates once more from 5' to 3', resulting in multiple reads of the same DNA strand. No open pore levels were observed during the reread, i.e. it is unlikely that the molecule was ejected from the nanopore. Figure 6b shows an example current-time trace of a molecule that was read twice (i and ii). A hidden Markov model was trained to map the enzyme control moieties against the references for each restriction digested fragment (Figure 6c). The data shows that the reads mapped to the same fragment in the reference, and instances where some mapped twice or three times were recorded, confirming that a strand was read multiple times.

[0402] Example 2 This example demonstrates how motor proteins with different linker lengths for disulfide closure can be used to reread native DNA analytes multiple times.

[0403] Y-adapters with leader arms containing 30 C3 spacer units were prepared by annealing four DNA oligonucleotides with the sequences of SEQ ID NO:50, SEQ ID NO:51, SEQ ID NO:52, and SEQ ID NO:53. A DNA motor (Dda helicase) was loaded onto each adapter and disulfide-closed via reaction with one of the following linkers: diamide (TMAD), BMOE (1,2-bismaleimidoethane), BMOP (1,3-bismaleimidopropane), BMB (1,4-bismaleimidobutane), BM(PEG)2 (1,8-bismaleimido-diethylene glycol), or BM(PEG)3 (1,11-bismaleimido-triethylene glycol).

[0404] E. coli K12 PCR DNA was obtained by extraction from E. coli cells using a Qiagen genomic chip kit, sheared to approximately 10 kb cutoff using a Covaris gTube, end-repaired and dA-tailed using the Ultra II End Repair and dA-Tailing Kit (New England Biolabs), ligated to PCR adapters (PCA; Oxford Nanopore Technologies), and PCR amplified using LongAmp Taq. The resulting double-stranded analytes were end-repaired and dA-tailed by NEBNext End Repair and NEBNext dA Tailing Modules (New England Biolabs (NEB)), generating 3' dA overhangs on both ends of each fragment. Samples were ligated to the T overhangs of the Y adapters using LNB and T4 DNA ligase from the Oxford Nanopore Technologies sequencing kit (LSK-SQK109). Samples were purified using Agencourt AMPure XP (Beckman Coulter) beads and washed twice with LFB from the Oxford Nanopore Technologies sequencing kit (LSK-SQK109). The ligated substrates were eluted in elution buffer (EB) from the same kit to obtain the "DNA library". A DNA library was prepared separately using an adaptor carrying the Dda helicase closed with a disulfide linker, as described above.

[0405] Electrical measurements were taken on a custom MinION flow cell from Oxford Nanopore Technologies with a CsgG nanopore inserted and a MinION Mk1b. To 1170 μL FB (from the Oxford Nanopore Technologies Sequencing Kit (SQK-LSK109)), 30 μL of FLT was added to obtain the tether mix. 800 μL of tether mix was run through the system, followed by a 5 minute wait and an additional 200 μL of tether mix was run through the system with the SpotON port open. 37.5 μL of SQB from the Oxford Nanopore Technologies Sequencing Kit (SQK-LSK109), 15 μL of DNA library, and 22.5 μL of LB from the Oxford Nanopore Technologies Sequencing Kit (SQK-LSK109) were mixed to obtain the "sequencing mix". 75 μL of the sequencing mix was added to the MinION flow cell via the SpotON flow cell port.

[0406] A custom sequence script was prepared to control the applied potential using the active unblocking circuitry of the MinION. The sequencing voltage was set to 180 mV, and the motor protein was de-stalled by switching the voltage to zero by disconnecting the channel for 5 seconds when an enzyme stall level was detected. The classification of stall and strand (sequencing) levels was programmed into the configuration file of the MinKNOW instrument control software to allow detection of stalled species and to apply an unblocking potential that did not cause complete release of the strand. The script worked as follows: if MinKNOW detected that the strand was at the stall level, it applied the unblocking potential for 5 seconds, then returned to a sequencing potential of 180 mV to check for the actively sequenced strand 5 times. If the stall level was still present, it applied the unblocking potential for another 25 seconds and repeated 5 times. A 3 second rest period was built in between each unblocking attempt. If MinKNOW detected an active sequencing strand when it returned the sequencing potential, it stopped the unblocking attempt and applied only the sequencing potential. If no active sequencing strands are generated throughout this process, MinKNOW will turn off the channel. Active deblocking was set to be triggered upon recognition of a block level not related to terminal C3 levels, strands, open pores, or enzyme stall levels. Every 15 min, a "mux scan" was applied to reset the system, globally unblocking all channels of the flow cell and checking for active nanopores at 180 mV. Raw data were collected in bulk FAST5 files using MinKNOW software (Oxford Nanopore Technologies).

[0407] Strand-level events from single-channel data that occurred immediately after the C3 level (marked "C3" in Figure 6b) were recorded as potential re-reads (e.g., marked "ii" in Figure 6b). These re-reads were confirmed by comparing the sequence of base calling and re-reads that occurred after the open pore and destalling events with the original read (e.g., marked "i" in Figure 6b) as described in Example 1. Events in the same read direction and within the range of the original read were classified as re-reads. Re-read efficiency was quantified in two ways: (i) the percentage of reads that drop back and re-read within 30 seconds of reaching the C3 reader, and (ii) the drop-back distance, which is the length of the re-read, i.e., the distance the enzyme was pushed back from the C3 reader.

[0408] The table below shows the results of this experiment, which show re-reads for all linkers tested, and that increasing linker length increases the percentage of reads with re-reads within 30 seconds of reaching the C3 reader. [Table 4]

[0409] Example 3 This example demonstrates how a native DNA analyte can be read multiple times using adapters with different sequences in the leader that are encountered by the Dda helicase at the 3' end of the sequenced strand.

[0410] Y-adapters with leader arms having RNA or C3 leader chemistry were prepared by annealing four DNA oligonucleotides with sequences of SEQ ID NO:50, SEQ ID NO:51, and SEQ ID NO:52 and a leader oligonucleotide selected from SEQ ID NO:53, SEQ ID NO:54, and SEQ ID NO:55. A DNA motor (Dda helicase) was loaded onto each adapter and the disulfides were closed via reaction with 1,2-bismaleimidoethane (BMOE).

[0411] A DNA library was prepared by ligating the above Y adaptors to E. coli DNA prepared as described in Example 2.

[0412] Electrical measurements were taken on a custom MinION flow cell from Oxford Nanopore Technologies with a CsgG nanopore inserted and a MinION Mk1b. To 1170 μL FB (from the Oxford Nanopore Technologies Sequencing Kit (SQK-LSK109)), 30 μL of FLT was added to obtain the tether mix. 800 μL of tether mix was run through the system, followed by a 5 minute wait and an additional 200 μL of tether mix was run through the system with the SpotON port open. 37.5 μL of SQB from the Oxford Nanopore Technologies Sequencing Kit (SQK-LSK109), 15 μL of DNA library, and 22.5 μL of LB from the Oxford Nanopore Technologies Sequencing Kit (SQK-LSK109) were mixed to obtain the "sequencing mix". 75 μL of the sequencing mix was added to the MinION flow cell via the SpotON flow cell port.

[0413] Data were collected using a custom script similar to that described in Example 2, but with re-reads generated from the strand level when the reader contained RNA.

[0414] The following table shows the results of this experiment, which demonstrate rereading with all leader oligonucleotides tested, with the best rereading efficiency achieved using leader oligonucleotide SEQ ID NO:55, as judged by the reduction in median time between rereads. [Table 5]

[0415] Example 4 This example demonstrates how a native DNA analyte can be re-read multiple times at different sequencing run temperatures.

[0416] A Y-adapter with a leader arm with C3 leader chemistry was prepared by annealing four DNA oligonucleotides with sequences SEQ ID NO:50, SEQ ID NO:51, SEQ ID NO:52, and SEQ ID NO:53. A DNA motor (Dda helicase) was loaded onto the adapter and the disulfide was closed via reaction with 1,2-bismaleimidoethane (BMOE).

[0417] A DNA library was prepared by ligating the above Y adaptors to E. coli DNA prepared as described in Example 2.

[0418] Electrical measurements were taken on a custom MinION flow cell from Oxford Nanopore Technologies with a CsgG nanopore inserted and a MinION Mk1b. To 1170 μL FB (from the Oxford Nanopore Technologies Sequencing Kit (SQK-LSK109)), 30 μL of FLT was added to obtain the tether mix. 800 μL of tether mix was run through the system, followed by a 5 minute wait and an additional 200 μL of tether mix was run through the system with the SpotON port open. 37.5 μL of SQB from the Oxford Nanopore Technologies Sequencing Kit (SQK-LSK109), 15 μL of DNA library, and 22.5 μL of LB from the Oxford Nanopore Technologies Sequencing Kit (SQK-LSK109) were mixed to obtain the "sequencing mix". 75 μL of the sequencing mix was added to the MinION flow cell via the SpotON flow cell port.

[0419] Data was collected using custom scripts similar to those described in Example 2.

[0420] The table below shows the results of this experiment. The results show re-reads at all temperatures tested, and that re-read efficiency improved with increasing temperature, as judged by the increasing percentage of reads that were re-read within 30 seconds of reaching the C3 reader, and the median drop back distance. [Table 6]

[0421] Example 5 This example demonstrates how a DNA analyte can be reread multiple times by a nanopore and the translocation of the DNA through the nanopore is controlled by a motor protein.

[0422] The Y-adapter was prepared by annealing DNA oligonucleotides 56, 57, and 58. The control Y-adapter was prepared by annealing DNA oligonucleotides 59, 58, and 60. A DNA motor protein (Dda helicase) was loaded onto the adapter and closed by incubation with trimethylthiodiamide (TMAD).

[0423] Components FB, FLT, SFB, LNB, LB, EB and SQB were obtained from the Ligation Sequencing Kit SQK-LSK109 from Oxford Nanopore Technologies plc.

[0424] A 3.6 kilobase DNA analyte was obtained via PCR amplification from bacteriophage lambda, then end-repaired and dA-tailed using the Ultra II End Repair and dA Tailing Kit (New England Biolabs) to generate 3' dA overhangs on both ends of each fragment. Using LNB and T4 DNA ligase, the samples were ligated to the T overhangs of either the Y-adapter or the control Y-adapter to obtain two substrates, which were then purified using Agencourt AMPure XP (Beckman Coulter) beads. The beads were washed twice using the SFB from the Oxford Nanopore Technologies sequencing kit (SQK-LSK109). The ligated substrates were eluted in the elution buffer (EB) of the same kit. Monovalent traptavidin was added to a final concentration of 500 nM to obtain two "DNA libraries".

[0425] Electrical measurements were taken using a custom MinION flow cell containing the CsgG nanopore, and data was collected using a GridION X5 (Oxford Nanopore Technologies plc). To 1170 μL of FB, 30 μL of FLT was added to obtain the Tether Mix. 800 μL of Tether Mix was run through the system, followed by a 5-minute wait and an additional 200 μL of Tether Mix running through the system with the SpotON port open. 37.5 μL of SQB, 15 μL of each DNA library, and 22.5 μL of LB were mixed to obtain the "sequencing mix." 75 μL of the sequencing mix was added to the MinION flow cell via the SpotON flow cell port.

[0426] Data was collected using a custom script based on the SQK-LSK109 baseline sequencing script provided by Oxford Nanopore Technologies plc on an R9.4.1 flow cell. The script was configured to recognize and accept open pore, C3 spacer and chain levels, and active deblocking was configured to trigger upon recognition of terminal biotin-traptavidin blockade levels and other blockades.

[0427] A schematic of the adaptor is shown in FIG. 7, A, and a schematic of the rereading is shown in FIG. 7, B. Exemplary electrical data are shown in FIG. 8. These data show capture of DNA library molecules via the 3' end, which resulted in a current slightly higher than the open pore current ("leader level"). After a brief pause at this level, motor-controlled movement was observed at the "strand" level, followed by a return to the "leader level". This pattern ("leader level", "strand level") was repeated several times. Base calling of the strand-level events confirmed that the same DNA strand was read multiple times without dissociating from the nanopore, and that controlled movement of the DNA strand from the pore occurred up to the leader section, after which the motor "retreated" to an earlier point on the strand (3' to 5' movement), resulting in controlled movement up to the 3' end (5' to 3' movement). In the absence of a biotin moiety acting as a stopper (in the control adaptor), the motor protein was more frequently pushed away ("push off") from the 5' end of the polynucleotide strand.

[0428] Example 6 This example demonstrates how a DNA analyte can be reread multiple times by a nanopore and the translocation of the DNA through the nanopore is controlled by a motor protein.

[0429] Y-adapters were prepared by annealing DNA oligonucleotides SEQ ID NO: 61, 62 and 63. A DNA motor protein (Dda helicase) was loaded onto the adapters and closed by incubation with trimethylthiodiamide (TMAD).

[0430] Components FB, FLT, SFB, LNB, LB, EB and SQB were obtained from the Ligation Sequencing Kit SQK-LSK109 from Oxford Nanopore Technologies plc.

[0431] NA12878 DNA was fragmented to 180 bp using Bioruptor Pico (diagenode) according to the manufacturer's protocol, then end-repaired and dA-tailed using Ultra II End Repair and dA Tailing Kit (New England Biolabs) to generate 3' dA overhangs on both ends of each fragment. Using LNB and T4 DNA ligase, samples were ligated to the T overhangs of Y adapters and then purified using Agencourt AMPure XP (Beckman Coulter) beads. Beads were washed twice using SFB from Oxford Nanopore Technologies' Sequencing Kit (SQK-LSK109). Ligated products were eluted in the elution buffer (EB) of the same kit. Monovalent traptavidin was added to a final concentration of 500 nM to obtain the "DNA library".

[0432] Electrical measurements were taken using a custom MinION flow cell containing the CsgG nanopore, and data was collected using a GridION X5 (Oxford Nanopore Technologies plc). To 1170 μL of FB, 30 μL of FLT was added to obtain the Tether Mix. 800 μL of Tether Mix was run through the system, followed by a 5-minute wait and an additional 200 μL of Tether Mix running through the system with the SpotON port open. 37.5 μL of SQB, 15 μL of each DNA library, and 22.5 μL of LB were mixed to obtain the "sequencing mix." 75 μL of the sequencing mix was added to the MinION flow cell via the SpotON flow cell port.

[0433] Data was collected using a custom script based on the SQK-LSK109 baseline sequencing script provided by Oxford Nanopore Technologies plc on an R9.4.1 flow cell. The script was configured to recognize and accept open pore, C3 spacer and chain levels, and active deblocking was configured to trigger upon recognition of terminal biotin-traptavidin blockade levels and other blockades.

[0434] A schematic of the adapter is shown in FIG. 7, A, and a schematic of the rereading is shown in FIG. 7, B. Exemplary electrical data are shown in FIG. 9. These data show capture of DNA library molecules via the 3' end, which resulted in a current slightly higher than the open pore current ("reader level"). After a brief pause at this level, motor-controlled translocation was observed at the "strand" level, followed by a return to the "reader level". This pattern ("reader level", "strand level") was repeated several times. Base calling of strand-level events confirmed that the same DNA strand was read multiple times without dissociating from the nanopore, and that controlled translocation of the DNA strand from the pore occurred up to the leader section, after which the motor "retreated" to an earlier point on the strand (3' to 5' translocation), resulting in controlled translocation to the 3' end (5' to 3' translocation).

[0435] Description of sequence listing SEQ ID NO: 1 shows the amino acid sequence of (hexahistidine-tagged) exonuclease I (EcoExo I) from E. coli.

[0436] SEQ ID NO:2 shows the amino acid sequence of the exonuclease III enzyme from E. coli.

[0437] SEQ ID NO: 3 shows the amino acid sequence of the RecJ enzyme from T. thermophilus (TthRecJ-cd).

[0438] SEQ ID NO: 4 shows the amino acid sequence of bacteriophage lambda exonuclease, which is one of three identical subunits that make up the trimer (http: / / www.neb.com / nebecomm / products / productM0262.asp).

[0439] SEQ ID NO:5 shows the amino acid sequence of Phi29 DNA polymerase from Bacillus subtilis phage Phi29.

[0440] SEQ ID NO: 6 shows the amino acid sequence of Trwc Cba (Citromicrobium bathyomarinum) helicase.

[0441] SEQ ID NO: 7 shows the amino acid sequence of Hel308 Mbu (Methanococcoides burtonii) helicase.

[0442] SEQ ID NO: 8 shows the amino acid sequence of Dda helicase 1993 from enterobacteriaceae phage T4.

[0443] SEQ ID NOs: 9 to 22 show the nucleotide sequences of the DNA strands discussed in the Examples.

[0444] SEQ ID NO: 23 shows the amino acid sequence of a preferred HhH domain.

[0445] SEQ ID NO: 24 shows the amino acid sequence of ssb from bacteriophage RB69 encoded by the gp32 gene.

[0446] SEQ ID NO: 25 shows the amino acid sequence of ssb from bacteriophage T7 encoded by the gp2.5 gene.

[0447] SEQ ID NO: 26 shows the amino acid sequence of the UL42 processivity factor from herpesvirus type 1.

[0448] SEQ ID NO: 27 shows the amino acid sequence of subunit 1 of PCNA.

[0449] SEQ ID NO: 28 shows the amino acid sequence of subunit 2 of PCNA.

[0450] SEQ ID NO: 29 shows the amino acid sequence of subunit 3 of PCNA.

[0451] SEQ ID NO: 30 shows the amino acid sequence (1 to 319) of the UL42 processivity factor derived from herpesvirus type 1.

[0452] SEQ ID NO: 31 shows the amino acid sequence of the (HhH)2 domain.

[0453] SEQ ID NO: 32 shows the amino acid sequence of the (HhH)2-(HhH)2 domain.

[0454] SEQ ID NO: 33 shows the amino acid sequence of human mitochondrial SSB (HsmtSSB).

[0455] SEQ ID NO: 34 shows the amino acid sequence of the p5 protein from Phi29 DNA polymerase.

[0456] SEQ ID NO: 35 shows the amino acid sequence of wild-type SSB from E. coli.

[0457] SEQ ID NO: 36 shows the amino acid sequence of ssb from bacteriophage T4 encoded by the gp32 gene.

[0458] SEQ ID NO: 37 shows the amino acid sequence of topoisomerase V Mka (Methanopyrus Kandleri).

[0459] SEQ ID NO: 38 shows the amino acid sequence of domain HL of topoisomerase V Mka (Methanopyrus Kandleri).

[0460] SEQ ID NO: 39 shows the amino acid sequence of Mutant S (Escherichia coli).

[0461] SEQ ID NO: 40 shows the amino acid sequence of Sso7d (Sufolobus solfataricus).

[0462] SEQ ID NO: 41 shows the amino acid sequence of Sso10b1 (Sulfolobus solfataricus P2).

[0463] SEQ ID NO: 42 shows the amino acid sequence of Sso10b2 (Sulfolobus solfataricus P2).

[0464] SEQ ID NO: 43 shows the amino acid sequence of tryptophan repressor (Escherichia coli).

[0465] SEQ ID NO: 44 shows the amino acid sequence of the lambda repressor (Enterobacteriaceae phage lambda).

[0466] SEQ ID NO: 45 shows the amino acid sequence of Cren7 (Histone crenarchaea Cren7 Sso).

[0467] SEQ ID NO: 46 shows the amino acid sequence of human histone (Homo sapiens).

[0468] SEQ ID NO: 47 shows the amino acid sequence of dsbA (Enterobacteriaceae phage T4).

[0469] SEQ ID NO: 48 shows the amino acid sequence of Rad51 (Homo sapiens).

[0470] SEQ ID NO: 49 shows the amino acid sequence of PCNA sliding clamp (Citromicrobium bathyomarinum JL354).

[0471] SEQ ID NOs:50 to 55 show the polynucleotide sequences of the oligonucleotides described in Examples 2 to 4.

[0472] SEQ ID NOs: 56-60 show the polynucleotide sequences of the oligonucleotides described in Example 5, where / 5Phos / =5' phosphate, 3=C3 spacer, 8=spacer 18, underlined bases=locked nucleic acid (LNA), and / 5BioTEG / =5' biotin TEG.

[0473] SEQ ID NOs: 61-63 show the polynucleotide sequences of the oligonucleotides described in Example 6, where / 5Phos / =5' phosphate, 3=C3 spacer, 8=spacer 18, underlined bases=locked nucleic acid (LNA), and / 5BioTEG / =5' biotin TEG.

[0474] Sequence Listing SEQ ID NO:1 - Exonuclease I from E. coli [ka] SEQ ID NO:2 - Exonuclease III enzyme from E. coli [ka] SEQ ID NO:3 - RecJ enzyme from T. thermophilus [ka] SEQ ID NO:4 - Bacteriophage lambda exonuclease [ka] SEQ ID NO:5 - Phi29 DNA polymerase [ka] SEQ ID NO:6 - Trwc Cba helicase [ka] SEQ ID NO:7 - Hel308 Mbu helicase [ka] SEQ ID NO:8 - Dda helicase [ka] SEQ ID NO:9 SEQ ID NO:10 / 5BiotinTEG / TTTTTTTTTT / iSp18 / AATGTACTTCGTTCAGTTACGTATTGCT SEQ ID NO:11 / 5Phos / GCAATACGTAACTGAACGAAGT / iBNA-A / / iBNA-MeC / / iBNA-A / / iBNA-T / / iBNA-T / TTTGAGGCGAGCGGTCAATTTTTTTTTTTTTTTTTTTT SEQ ID NO: 12 / 5Phos / TGCAATACGTAACTGAACGAAGTACATTAATGTACTTCGTTCAGTTACGTATTGCATCCT SEQ ID NO:13 / 5Phos / TGCAATACGTAACTGAACGAAGTACATTTTTTTGAAGATAGAGCGATTTTTTTTTTTTTTTTGTACTTCGTTCAGTTACGTATTGCATCCT SEQ ID NO:14 / 5Phos / TGCAATACGTAACTGAACGAAGTACATTTTTTTGAAGATAGAGCGATTTTTTTTTTTTTTTTGTACTTCGTTCAGTTACGTATTGCAT SEQ ID NO:15 / 5Phos / TGCAATACGTAACTGAACGAAGTACATTTTTTTGAAGATAGAGCGATTTTT / iFluorT / / iFluorT / / iFluorT / TTTTTTTTGTACTTCGTTCAGTTACGTATTGCATCCT SEQ ID NO:16 / 5BNA-T / / iBNA-MeC / / iBNA-G / / iBNA-MeC / / iBNA-T / CTATCTTC SEQ ID NO:17 GTTATTCAAGACTTCTTTAATACACTTTTTTTTT / iSp18 / AATGTACTTCGTTCAGTTACGTATTGCTTTGGCGTCTGCTTGGGTGTTTAACCT SEQ ID NO:18 / 5Phos / GGTTAAACACCCAAGCAGACGCCTTTGAGGCGAGCGGTCAATTTTTTTTTTTTTTTTTT SEQ ID NO:19 GCAATACGTAACTGAACGAAGT / iBNA-A / / iBNA-MeC / / iBNA-A / / iBNA-T / / 3BNA-T / SEQ ID NO:20 / 5BNA-G / / iBNA-T / / iBNA-G / / iBNA-T / / iBNA-A / TTAAAGAAGTCTTGAAT / iBNA-A / / iBNA-A / / 3BNA-meC / SEQ ID NO:21 GTGTATTAAAGAAGTCTTGAATAAC SEQ ID NO:22 / 5Phos / GGTTAAACACCCAAGCAGACGCCTTTGAGGCGAGCGGTCAA / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSp C3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / 3SpC3 / SEQ ID NO:23 GTGSGAWKEWLERKVGEGRARRLIEYFGSAGEVGKLVENAEVSKLLEVPGIGDEAVARLV PGGSS SEQ ID NO:24 [ka] SEQ ID NO:25 [ka] SEQ ID NO:26 [ka] SEQ ID NO:27 [ka] SEQ ID NO:28 [ka] SEQ ID NO:29 [ka] SEQ ID NO:30 [ka] SEQ ID NO:31 WKEWLERKVGEGRARRLIEYFGSAGEVGKLVENAEVSKLLEVPGIGDEAVARLVP SEQ ID NO:32 [ka] SEQ ID NO:33 [ka] SEQ ID NO:34 [ka] SEQ ID NO:35 [ka] SEQ ID NO:36 [ka] SEQ ID NO:37 [ka] SEQ ID NO:38 [ka] SEQ ID NO:39 [ka] SEQ ID NO:40 [ka] SEQ ID NO:41 [ka] SEQ ID NO:42 [ka] SEQ ID NO:43 [ka] SEQ ID NO:44 [ka] SEQ ID NO:45 MSSGKKPVKVKTPAGKEAELVPEKVWALAPKGRKGVKIGLFKDPETGKYFRHKLPDDYPI SEQ ID NO:46 [ka] SEQ ID NO:47 [ka] SEQ ID NO:48 [ka] SEQ ID NO:49 [ka] SEQ ID NO:50 GTTATTCAAGACTTCTTTAATACACTTTTTTTTT / iSp9 / AATGTACTTCGTTCAGTTACGTATTGCTTTGGCGTCTGCTTGGGTGTTTAACCT SEQ ID NO:51 GTGTATTAAAGAAGTCTTGAATAACTTTGAGGCGAGCGGTCAA SEQ ID NO:52 TTTGCAATACGTAACTGAACGAAGT / iBNA-A / / iBNA-MeC / / iBNA-A / / iBNA-T / / 3BNA-T / sequence number 53 / 5Phos / GGTTAAACACCCAAGCAGACGCCTTT / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 C3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / sequence number 54 / 5Phos / GGTTAAACACCCAAGCAGACGCCTTTmUmUmUmUmUmU / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / sequence number 55 / 5Phos / GGTTAAACACCCAAGCAGACGCCmUmU / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 C3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / iSpC3 / / sequence number 56 / 5Phos / GGCGTCTGCTTGGGTGTTTAACCTTTTTTTTT33333333333333333333AATGTACTTCGTTCAGTTACGTATTGCTTTTTTTTTTTTTTTTTT sequence number 57 / 5BioTEG / TTTGGTTAAACACCCAAGCAGACGCCT SEQ ID NO:58 [ka] SEQ ID NO:59 / 5Phos / GGACTATACAAGGCGTCTGCTTGGGTGTTTAACCTTTTTTTTTT33333333333333333333AATGTACTTCGTTCAGTTACGTATTGCTTTTTTTTTTTTTTTTTT SEQ ID NO:60 GGTTAAACACCCAAGCAGACGCCTTGTATAGTCCT SEQ ID NO:61 / 5Phos / GGCGTCTGCTTGGGTGTTTAACCTTTTTTTTTT333333333333333333333AATGTACTTCGTTCAGTTACGTATTGCTTTTTTTTTTTTTTTTTT SEQ ID NO:62 / 5BioTEG / TTTGGTTAAACACCCAAGCAGACGCCT SEQ ID NO:63 [ka]

Claims

1. 1. A method for characterizing a leader-attached target polynucleotide, comprising: The method comprises: (i) contacting a detector having a first opening and a second opening, or contained in a structure having a first opening and a second opening, with the reader under conditions such that the first opening contacts the reader and the target polynucleotide migrates in a direction from the first opening to the second opening; contacting the leader with a motor protein bound at a polynucleotide binding site of the motor protein; (ii) performing one or more measurements characteristic of the target polynucleotide when the motor protein controls movement of the target polynucleotide in a first direction relative to the detector, the first direction being from the second opening to the first opening; (iii) debinding the target polynucleotide from the polynucleotide binding site of the motor protein such that the target polynucleotide moves in a second direction relative to the detector, the second direction being from the first opening to the second opening; (iv) allowing the target polynucleotide to rebind to the polynucleotide-binding site of the motor protein and performing one or more measurements characteristic of the target polynucleotide when the motor protein controls movement of the target polynucleotide in the first direction relative to the detector; thereby characterizing the target polynucleotide.

2. 2. The method of claim 1, wherein the target polynucleotide has a first end and a second end, the leader is attached to the first end, and the motor protein is oriented in an orientation to process the target polynucleotide in a direction from the second end toward the first end.

3. 10. The method of claim 1, comprising repeating steps (iii) and (iv) multiple times.

4. 2. The method of claim 1, wherein prior to step (i), the target polynucleotide is comprised in or consists of the first strand of a double-stranded polynucleotide comprising a first strand and a second strand.

5. 5. The method of claim 4, wherein the portion of the first strand between the motor protein and the second end is hybridized to the second strand.

6. 5. The method of claim 4, wherein movement of the target polynucleotide in the direction from the first opening to the second opening comprises separation of the first strand from the second strand.

7. 5. The method of claim 4, wherein movement of the target polynucleotide in the direction from the second opening to the first opening comprises annealing of the first strand to the second strand.

8. 5. The method of claim 4, wherein the first strand of the double-stranded polynucleotide is attached to the second strand of the double-stranded polynucleotide.

9. 2. The method of claim 1, wherein in step (ii), the motor protein controls movement of a first portion of the target polynucleotide in the first direction relative to the detector, and wherein in step (iv), the motor protein controls movement of a second portion of the target polynucleotide in the first direction relative to the detector, wherein the first portion at least partially overlaps the second portion.

10. The method of claim 1 , wherein the first portion is the same as the second portion.

11. 2. The method of claim 1, wherein in step (ii), the motor protein controls movement of a first portion of the target polynucleotide in the first direction relative to the detector, and in step (iv), the motor protein controls movement of a second portion of the target polynucleotide in the first direction relative to the detector, wherein the first portion does not overlap with the second portion.

12. 2. The method of claim 1, wherein in step (iii), the distance traveled by the target polynucleotide relative to the detector is greater than the distance traveled by the polynucleotide relative to the detector in step (ii) and / or step (iv).

13. 2. The method of claim 1, wherein: (a) in step (iii), the distance traveled by the target polynucleotide relative to the detector is at least 1000 nucleotides in length; and / or (b) in steps (ii) and / or (iv), the distance traveled by the target polynucleotide relative to the detector is each independently at least 100 nucleotides in length.

14. 2. The method of claim 1, wherein the second end of the target polynucleotide comprises a blocking moiety to prevent the motor protein from disassociating from the polynucleotide.

15. 15. The method of Claim 14, wherein the blocking moiety restricts movement of the target polynucleotide through the polynucleotide binding site of the motor protein, thereby restricting movement of the target polynucleotide in the second direction relative to the detector.

16. the detector comprises a transmembrane nanopore spanning a membrane having a cis side and a trans side; (i) the first opening of the nanopore is on the cis side of the membrane and the second opening of the nanopore is on the trans side, and the motor protein controls the movement of the target polynucleotide through the nanopore from the trans side to the cis side of the membrane, and when the target polynucleotide unbinds from the polynucleotide binding site of the motor protein, the target polynucleotide moves through the nanopore from the cis side to the trans side of the membrane; or 2. The method of claim 1, wherein (ii) the first opening of the nanopore is on the trans-side of the membrane and the second opening of the nanopore is on the cis-side, the motor protein controls the movement of the target polynucleotide through the nanopore from the cis-side to the trans-side of the membrane, and when the target polynucleotide unbinds from the polynucleotide-binding site of the motor protein, the target polynucleotide moves through the nanopore from the trans-side to the cis-side of the membrane.

17. 10. The method of claim 1, wherein the target polynucleotide does not disassociate from the motor protein.

18. 10. The method of claim 1, wherein the motor protein is modified to prevent the target polynucleotide from disassociating from the motor protein.

19. 10. The method of claim 1, wherein the motor protein is modified to facilitate debinding of the target polynucleotide from the polynucleotide binding site of the motor protein and / or to delay rebinding of the target polynucleotide to the polynucleotide binding site of the motor protein.

20. 10. The method of claim 1, wherein the motor protein is modified with a closing moiety to (i) topologically close the polynucleotide binding site of the motor protein around the target polynucleotide, and / or (ii) facilitate debinding of the target polynucleotide from the polynucleotide binding site of the motor protein and / or delay rebinding of the target polynucleotide to the polynucleotide binding site of the motor protein.

21. the motor protein is modified to facilitate attachment of the closing moiety to the motor protein; 21. The method of claim 20, optionally wherein the motor protein is modified by substituting cysteine ​​or an unnatural amino acid for at least one amino acid in the motor protein.

22. i) the closing moiety comprises a bifunctional cross-linker; ii) the closing moiety comprises a bond, optionally a disulfide bond; iii) the closing moiety comprises a structure of the formula [ABC], where A and C are each independently a reactive functional group for reacting with an amino acid residue in the motor protein, and B is a linking moiety; Optionally, wherein A and C are each independently a cysteine-reactive functional group, and / or linking moiety B comprises a linear or branched, unsubstituted or substituted alkylene, alkenylene, alkynylene, arylene, heteroarylene, carbocyclylene, or heterocyclylene moiety, which moiety is optionally interrupted and / or terminated by one or more atoms or groups selected from O, N(R), S, C(O), C(O)NR, C(O)O, unsubstituted or substituted arylene, arylene-alkylene, heteroarylene, heteroarylene-alkylene, carbocyclylene, carbocyclylene-alkylene, heterocyclylene, and heterocyclylene-alkylene, wherein R is selected from H, unsubstituted or substituted alkyl, and unsubstituted or substituted aryl; 21. The method of claim 20, further optionally, wherein linking moiety B comprises an alkylene, oxyalkylene, or polyoxyalkylene group, and / or A and C are each maleimide groups.

23. 21. The method of claim 20, wherein the closing moiety bridges two amino acid residues of the motor protein, and at least one amino acid bridged by the closing moiety is cysteine ​​or an unnatural amino acid.

24. 21. The method of claim 20, wherein the closing portion has a length of about 1 Å to about 100 Å, optionally about 5 Å to about 50 Å.

25. The method of claim 1 , wherein the motor protein is a helicase.

26. 2. The method of claim 1, wherein the motor protein is a DNA-dependent ATPase (Dda) helicase.

27. 2. The method of claim 1, wherein prior to step (i), the motor protein is stalled on the leader.

28. 10. The method of claim 1, wherein the leader comprises a different type of nucleotide than the target polynucleotide.

29. 2. The method of claim 1, wherein the target polynucleotide comprises deoxyribonucleotides (DNA) or ribonucleotides (RNA), and the leader comprises one or more stall units and / or one or more nucleotides lacking both a nucleobase and a sugar moiety (spacer moiety), deoxyribonucleotides (DNA), ribonucleotides (RNA), peptide nucleotides (PNAs), glycerol nucleotides (GNAs), threose nucleotides (TNAs), locked nucleotides (LNAs), bridged nucleotides (BNAs), abasic nucleotides, or nucleotides with modified phosphate linkages.

30. 5. The method of claim 4, wherein the second strand of the double-stranded polynucleotide comprises a membrane anchor or a transmembrane pore anchor.

31. applying a force to the detector, wherein the motor protein controls movement of the target polynucleotide relative to the detector in a direction opposite to the applied force; Optionally, the force comprises an electrical potential applied to the detector.

32. 1. A polynucleotide adaptor having a first end comprising a leader and a second end comprising an attachment point for attachment to a polynucleotide analyte at the first end of the polynucleotide analyte, the polynucleotide adaptor comprising a stalled motor protein thereon in an orientation for processing the adaptor in a direction from the second end to the first end.

33. 33. A kit comprising: a first adaptor of claim 32; and a second adaptor comprising (i) an attachment point for attachment to a polynucleotide analyte at a second end of the polynucleotide analyte, and (ii) a blocking moiety suitable for preventing the motor protein of the first adaptor from disassociating from the polynucleotide analyte when the first adaptor is attached to the polynucleotide analyte.

34. 1. A system for characterizing a target polynucleotide, comprising: - one or more polynucleotide adapters as defined in claim 32, - a nanopore for characterizing the target polynucleotide as it translocates relative to the nanopore; - a motor protein for controlling the movement of the target polynucleotide relative to the nanopore.

35. 33. The polynucleotide adaptor, kit, or system of claim 32, wherein the polynucleotide adaptor, the motor protein, and / or the blocking moiety are as defined in claim 1.