A method for characterizing polynucleotides that travel through nanopores.
The method of stalling and destalling motor proteins to control polynucleotide movement through nanopores enhances characterization accuracy by reducing slippage and heterogeneity, facilitating efficient strand-specific data collection.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- OXFORD NANOPORE TECH LTD
- Filing Date
- 2021-06-18
- Publication Date
- 2026-05-15
AI Technical Summary
Existing methods for characterizing polynucleotides using nanopores face challenges such as uncontrolled movement of polynucleotides, slippage by motor proteins, and aggregation of data from multiple strands leading to loss of strand-specific information and inefficiencies.
A method involving a motor protein that stalls and destalls to control the movement of polynucleotides through a detector, allowing for accurate characterization by measuring characteristics in both directions relative to the detector, and using adapters with blocking moieties to prevent unassociation.
Improves the accuracy of polynucleotide characterization by reducing slippage and heterogeneity, enabling efficient data collection and strand-specific analysis.
Smart Images

Figure 0007860000000009 
Figure 0007860000000010 
Figure 0007860000000011
Abstract
Description
[Technical Field]
[0001] This disclosure provides a method for characterizing target polynucleotides as they move toward a detector such as a transmembrane nanopore. The disclosure also provides novel polynucleotide adapters and kits for use in such a method. The disclosure also provides a method for rereading polynucleotides. [Background technology]
[0002] Nanopore sensing is an approach to analytes detection and characterization that relies on observing individual binding or interaction events between analyte molecules and ion conduction channels. Nanopore sensors can be fabricated by placing a single nanometer-sized pore within an insulating film and measuring the voltage-driven ion current through the pore in the presence of analyte molecules. If analyte is present inside or near the nanopore, the ion flow through the pore will change, resulting in an ion current or current change measured across the channel. The identity of the analyte is revealed through its unique current signature, particularly the duration and extent of current blocks, as well as fluctuations in current levels during interaction with the pore.
[0003] Polynucleotides are important analytes for sensing in this manner. Nanopore sensing of polynucleotide analytes can reveal their identity and perform single-molecule counting of the sensed analytes, but can also provide information about their composition, such as their nucleotide sequences, as well as the presence of features such as base modifications, oxidation, reduction, decarboxylation, and deamination. Nanopore sensing has the potential to enable rapid and inexpensive polynucleotide sequencing, providing single-molecule sequence reads of polynucleotides ranging from tens to tens of thousands of base pairs in length.
[0004] Two key components of polymer characterization using nanopore sensing are (1) controlling the movement of the polymer through the pore, and (2) identifying the component building blocks as the polymer moves through the pore. During nanopore sensing of analytes such as polynucleotides, controlling the movement of the polynucleotide relative to the pore is crucial. Failure to control the movement can hinder or impede accurate characterization of the polynucleotide. For example, if the movement of polynucleotide relative to the pore is uncontrolled, accurately distinguishing each nucleotide in homopolymer polynucleotides becomes difficult.
[0005] It is known that the movement of polynucleotides to detectors such as nanopores can be controlled by using motor proteins that regulate the movement of polynucleotides. Suitable motor proteins include polynucleotide handling enzymes such as helicases, exonucleases, and topoisomerases. Motor proteins process polynucleotides in a controlled manner. Therefore, motor proteins can be used to control the movement of polymers such as polynucleotides to detectors such as nanopores.
[0006] When the detector is a nanopore, the disclosed method typically involves using a motor protein to feed polynucleotides into the nanopore. This operation is described in more detail herein. Methods for feeding polynucleotides into nanopores have been widely developed and have proven to be very useful for characterizing polynucleotides.
[0007] However, there remains a need for further methods for characterizing polypeptides. One problem is that it is sometimes desirable to obtain data different from that obtained from methods involving feeding polynucleotides to detectors such as nanopores. For example, the error profile of data resulting from polynucleotide characterization in methods involving feeding polynucleotides to detectors may, in some situations, not be optimal for accurate characterization of polynucleotides. Another problem is that when using motor proteins to feed polynucleotides to detectors such as nanopores, the motor proteins may skip forward on the polynucleotide chain in an uncontrolled manner. This phenomenon is also known as slippage. Slippage can be problematic when characterizing polynucleotides, for example, because one or more nucleotides within the polynucleotide may not be accurately characterized. This is particularly problematic when characterizing a polynucleotide determines its sequence. As a strategy to reduce slippage, the conventional focus has been on modifying motor proteins to minimize their tendency to slip on the polynucleotide chain. However, other methods for transporting polynucleotides to detectors such as nanopores that can reduce slippage would also be useful.
[0008] There is also a need for methods to improve the data obtained when characterizing polynucleotides. One problem is that, in some cases, it is desirable to improve the accuracy of the characterization data obtained when characterizing polynucleotides. In some known methods, multiple polynucleotides are characterized from a sample of polynucleotides, and the resulting data is aggregated, thereby improving the overall accuracy. However, this can lead to problems. For example, heterogeneity in the sample may mean that useful information about differences between strands may be lost when aggregating data obtained from the characterization of multiple polynucleotide strands. Furthermore, inefficiencies can arise because a new strand needs to be captured for characterization after the first strand has been processed. Therefore, alternative and / or improved methods for characterizing polynucleotides are needed.
[0009] For these and other reasons, there is a need for novel and / or improved methods for transferring polynucleotides to detectors such as nanopores. [Overview of the project]
[0010] This disclosure relates to a method for characterizing a target polynucleotide that moves toward a detector using a motor protein. More specifically, this disclosure relates to a method by which a motor protein moves a polynucleotide toward the outside of a detector. Thus, the direction of movement of the polynucleotide is opposite to the known methods by which polynucleotides move toward nanopores. This is described in more detail herein.
[0011] In the disclosed method, the motor protein is initially stalled at a stalling portion on the polynucleotide, and the method provided herein also includes destalling the motor protein so that it can control the movement of the polynucleotide exiting a detector (e.g., a nanopore). Methods for stalling and destalling the motor protein are described in more detail herein.
[0012] While this disclosure provides nanopores as exemplary detectors, the methods provided herein are suitable for detectors such as (i) zero-mode waveguides, (ii) field-effect transistors, optionally Noy field-effect transistors, (iii) AFM chips, (iv) nanotubes, optionally carbon nanotubes, and (v) nanopores. The disclosed methods are particularly suitable for methods of moving polynucleotides through a detector or through a structure containing a detector, such as wells in a detector chip.
[0013] Therefore, this specification provides a method for characterizing a target polypeptide, and this method is (i) A detector having a first opening and a second opening, or (ii) a structure including a detector, wherein the structure having a first opening and a second opening is brought into contact with a target polynucleotide, such that the target polynucleotide has a stalled motor protein on it, and the motor protein is stalled at the stalled portion. (ii) Bringing the stalled portion into contact with the nanopore, thereby causing the motor protein to destall, (iii) Measuring one or more measurements characteristic of a target polynucleotide as it controls the movement of a motor protein through a detector or structure from a second opening to a first opening, thereby characterizing the target polynucleotide.
[0014] Furthermore, this specification provides a method for characterizing a target polynucleotide, and this method is (i) Contacting the detector with a target polynucleotide to which a motor protein is bound, wherein the target polynucleotide is bound to the motor protein at the polynucleotide binding site of the motor protein, (ii) obtaining one or more measurements characteristic of the target polynucleotide when the motor protein controls the movement of the target polynucleotide in a first direction relative to the detector; (iii) dissociating the target polynucleotide from the polynucleotide binding site of the motor protein so that the target polynucleotide moves in a second direction relative to the detector; (iv) re-associating the target polynucleotide with the polynucleotide binding site of the motor protein and obtaining one or more measurements characteristic of the target polynucleotide when the motor protein controls the movement of the target polynucleotide in a first direction relative to the detector, thereby characterizing the target polynucleotide.
[0015] Also provided herein is a method of characterizing a target polynucleotide, the method comprising: (i) contacting a first opening of a transmembrane nanopore having a first opening and a second opening with the target polynucleotide, wherein the target polynucleotide has a stalled motor protein thereon and the motor protein is stalled at the stalled portion; (ii) contacting the stalled portion with the nanopore, thereby de-stalling the motor protein; (iii) obtaining one or more measurements characteristic of the target polynucleotide when the motor protein controls the movement of the target polynucleotide through the nanopore in a direction from the second opening of the nanopore to the first opening of the nanopore, thereby characterizing the target polynucleotide.
[0016] In some embodiments, the nanopore spans a membrane having a cis side and a trans side, with a first opening of the nanopore on the cis side of the membrane and a second opening on the trans side, and a motor protein controls the movement of a target polynucleotide through the nanopore from the trans side to the cis side of the membrane.
[0017] In some embodiments, the method involves applying a force across a nanopore, where a motor protein controls the movement of a target polynucleotide through the nanopore in the opposite direction to the applied force, and the force preferably includes an electrical potential applied across the nanopore.
[0018] In some embodiments, the motor protein is a helicase. In some embodiments, the motor protein is a DNA-dependent ATPase (Dda) helicase.
[0019] In some embodiments, the adapter is attached to one or both ends of the target polynucleotide. In some embodiments, the motor protein is stalled on the adapter.
[0020] In some embodiments, the nanopore captures the leader sequence at the first end of the target polynucleotide, and the motor protein stalls at the second end of the target polynucleotide, or on an adapter attached to the second end of the target polynucleotide.
[0021] In some embodiments, - The target polynucleotide is single-stranded, - The target polynucleotide includes a leader sequence, the leader sequence is located at the first end of the target polynucleotide or is contained in an adapter attached to the first end of the target polynucleotide, and - The motor protein is either stalled at the second end of the target polynucleotide or stalled on the adapter at the second end of the target polynucleotide.
[0022] In some embodiments, the target polynucleotide is double-stranded.
[0023] In some embodiments, - The target polynucleotide is double-stranded and comprises a first strand and a second strand. - The target polynucleotide includes a leader sequence, the leader sequence is located at the first end of the polynucleotide and is included in the first strand or in an adapter attached to the first strand, and - The motor protein stalls at the second end of the target polynucleotide.
[0024] In some embodiments, the motor protein stalls at the second end of the first strand of the target polynucleotide, or stalls on an adapter at the second end of the first strand of the target polynucleotide. In some embodiments, the first and second strands are attached together by a hairpin adapter at the second end of the first strand, and the motor protein stalls at the hairpin adapter. In some embodiments, the first and second strands are attached together by hairpin adapters attached (i) to the second end of the first strand and (ii) to the first end of the second strand, and the motor protein stalls at the second end of the second strand of the double-stranded polynucleotide, or stalls on the adapter at the second end of the second strand.
[0025] In some embodiments, the target polynucleotide includes a portion complementary to the tag sequence. In some embodiments, the target polynucleotide includes a portion having an oligonucleotide hybridized thereto, the oligonucleotide comprising (a) a hybridizing portion for hybridizing to the target polynucleotide, and (b) (i) a portion complementary to the tag sequence, or (ii) an affinity molecule capable of binding to the tag. In some embodiments, the target polynucleotide is double-stranded, the portion complementary to the tag sequence is the first chain portion of the polynucleotide, and / or the portion having an oligonucleotide hybridized thereto is the first chain portion of the polynucleotide.
[0026] In some embodiments, the motor protein stalls at a stall site containing one or more stall units independently selected from the following: -Polynucleotide secondary structure, preferably hairpin or G-quadrivalent (TBA), - Preferably, a nucleic acid analog selected from peptide nucleic acid (PNA), glycerol nucleic acid (GNA), threose nucleic acid (TNA), locked nucleic acid (LNA), cross-linked nucleic acid (BNA), and debasic nucleotide. -Spacer units selected from nitroindole, inosine, acridine, 2-aminopurine, 2-6-diaminopurine, 5-bromodoxyuridine, inverted thymidine (inverted dTs), inverted dideoxythymidine (ddTs), dideoxycytidine (ddCs), 5-methylcytidine, 5-hydroxymethylcytidine, 2'-O-methylRNA base, isodeoxycytidine (Iso-dCs), isodeoxyguanosine (Iso-dGs), C3(OC3H6OPO3) group, photocleavable (PC)[OC3H6-C(O)NHCH2-C6H3NO2-CH(CH3)OPO3] group, hexanediol group, spacer 9 (iSp9)[(OCH2CH2)3OPO3] group, spacer 18 (iSp18)[(OCH2CH2)6OPO3] group, and thiol linkage, and -Avidins such as fluorophores, traptabidine, streptavidin and neutraavidin, and / or biotin, cholesterol, methylene blue, dinitrophenol (DNP), digoxigenin and / or anti-digoxigenin and dibenzylcyclooctin groups.
[0027] In some embodiments, stalling a motor protein involves applying a stalling force to a polynucleotide, the stalling force being smaller than and / or in the opposite direction to a read force, the read force being the force applied while the motor protein controls the movement of the target polynucleotide and measurements are being taken to determine one or more features of the polynucleotide. In some embodiments, stalling a motor protein involves stepping the applied force one or more times between the stalling force and the read force.
[0028] In some embodiments, the motor protein stalls at a stall site comprising one or more stall units and one or more pausing portions, and contacting one or more pausing portions with a nanopore delays the movement of polynucleotides through the nanopore, thereby causing the motor protein to destall from one or more stall units. In some embodiments, the pausing portion comprises one or more pausing units independently selected from the following: -Polynucleotide secondary structure, preferably hairpin or G-quadrivalent (TBA), - Preferably, a nucleic acid analog selected from peptide nucleic acid (PNA), glycerol nucleic acid (GNA), threose nucleic acid (TNA), locked nucleic acid (LNA), cross-linked nucleic acid (BNA), and debasic nucleotide. -Avidins such as fluorophores, traptabidine, streptavidin and neutraavidin, and / or biotin, cholesterol, methylene blue, dinitrophenol (DNP), digoxigenin and / or anti-digoxigenin and dibenzylcyclooctin groups, and -Polynucleotide-binding protein.
[0029] In some embodiments, the target polynucleotide includes a blocking moiety that prevents the motor protein from unassociating with the polynucleotide. In some embodiments, the target polynucleotide includes a leader sequence at its first end, the motor protein stalls at the second end of the target polynucleotide, or on an adapter attached to the second end of the target polynucleotide, and the blocking moiety is located between the motor protein and the second end of the polynucleotide, thereby preventing the motor protein from unassociating with the target polynucleotide at its second end.
[0030] A polynucleotide adapter is also provided having a first end and a second end containing an attachment site for attaching to a double-stranded polynucleotide analyte, the polynucleotide adapter comprising (i) a motor protein stalled thereon in an orientation for processing the adapter in the direction of the attachment site, and (ii) a blocking portion located between the motor protein and the second end of the adapter.
[0031] A kit is also provided which includes the first adapter described herein and a second adapter which includes a single-stranded leader sequence at the first end and an attachment site at the second end for attachment to a double-stranded polynucleotide analyte.
[0032] In some embodiments of the polynucleotide adapter or kit provided herein, the polynucleotide adapter, the motor protein, and / or the blocking portion are as defined herein. [Brief explanation of the drawing]
[0033] [Figure 1]A schematic diagram illustrating the distinction between (A) the direction of polynucleotide (PN) transfer from a nanopore under the control of a motor protein by the method provided herein, and (B) the direction of polynucleotide transfer to a pore by a contrasting method. The open arrows indicate the direction of transfer of the motor protein (MP) and PN. In both cases, the MP is, for example, a 5'-3' helicase. [Figure 2] This is a schematic diagram of one embodiment of the method provided herein, in which the target polynucleotide is single-stranded, the target polynucleotide comprising a leader sequence located at the first end of the target polynucleotide, and a motor protein stalled at the second end of the target polynucleotide by a stall region (x). The leader sequence is captured by a nanopore, and the single-stranded polynucleotide moves through the nanopore until it reaches the stalled motor protein. Once destalled, the motor protein controls the movement of the polynucleotide from the pore. [Figure 3] This is a schematic diagram of one embodiment of the method provided herein, in which the target polynucleotide is double-stranded, the target polynucleotide comprising a leader sequence (wavy line) located at the first end of the first strand of the target polynucleotide, and a motor protein stalled at a stalled region (x) at the second end of the first strand of the target polynucleotide. The leader sequence is captured by a nanopore, and the first strand of the target polynucleotide moves through the nanopore until it reaches the stalled motor protein. After destallation, the motor protein (MP) controls the movement of the first strand of the target polynucleotide (PN) from the pore. [Figure 4]This is a schematic diagram of one embodiment of the method provided herein, wherein the target polynucleotide is double-stranded, and the target polynucleotide includes a leader sequence (wavy line) located at the first end of the first strand of the target polynucleotide, and a motor protein (MP) is stalled at a stall portion (x) of a hairpin adapter connecting the second end of the first strand of the target polynucleotide to the first end of the second strand of the target polynucleotide. The leader sequence is captured by a nanopore, and the first strand of the target polynucleotide moves through the nanopore until it reaches the stalled motor protein. After destallation, the motor protein controls the movement of the first strand of the target polynucleotide (PN) from the pore. [Figure 5] This is a schematic diagram of one embodiment of the method provided herein, wherein the target polynucleotide is double-stranded, the target polynucleotide comprising a leader sequence located at the first end of the first chain of the target polynucleotide, and a hairpin adapter ligating the second end of the first chain of the target polynucleotide to the first end of the second chain of the target polynucleotide. A motor protein (MP) stalls at a stall region (x) at the second end of the second chain of the target polynucleotide. The leader sequence (wavy line) is captured by a nanopore, and the first chain of the target polynucleotide, the hairpin adapter, and the second chain of the target polynucleotide pass through the nanopore until they reach the stalled motor protein. After destallation, the motor protein controls the movement of the second chain of the target polynucleotide (PN), the hairpin adapter, and the first chain as they exit the pore. [Figure 6]This is a nanopore sequencing adapter with a DNA helicase that moves from 5' to 3', where the 3' strand is preferentially captured in the nanopore. The adapter contains two oligonucleotides known as the top strand (A) and the bottom strand (B). The top strand contains a 5' biotin moiety (C) complexed with monovalent traptabidine (D), where the DNA motor (direction 5'-3') is loaded into a closed poly(dT) binding site (E), stalled by an internal spacer 18 moiety (F), and the 3' dT base is offered for ligation to a double strand (G) with a dA tail. The bottom strand contains a 5' phosphate moiety (H), a double-stranded region (I) containing a BNA base as a stalling chemical group, 20 consecutive 3' terminal thymidine bases as a leader (wavy line, J), and a site (K) for hybridizing a hydrophobic tether. See Example 1. [Figure 7] Figure 6 is a conceptual diagram showing how a sequencing adapter (A) is ligated to both ends of a double-stranded DNA polynucleotide (B) with a dA tail to generate a continuous double helix. [Figure 8] This is a schematic diagram of the experiment in Example 1, showing the capture, destallation, and sequencing of polynucleotide analytes. Vs: sequencing potential, Vu: deblocking potential. The polarity of the applied potential is indicated by the arrow. The direction of the applied force is the same as the direction of the arrow. (A) Applying the sequencing potential (120mV). Capture of polynucleotide analytes via the open pore and 3' leader (from Figure 7). Separation of the double helix by the nanopore, and the complementary strand is removed. (B) The polynucleotide reaches the enzyme, which is stalled in the spacer region. The enzyme cannot move over the spacer region. (C) Applying the deblocking potential (0mV) so that the enzyme moves away from the nanopore and moves freely over the spacer region. (D) Applying the sequencing potential (120mV). The polynucleotide moves through the nanopore until the enzyme reaches the nanopore, and then the enzyme controls the movement of the polynucleotide from the nanopore. (E) The DNA motor reaches the leader and becomes idle. (F) The deblocking potential is applied, and the DNA motor and analyte are expelled from the nanopore. Repeat the cycle from (A). [Figure 9]Top: Representative current-time traces from Example 1. States A-F correspond to the states described in Figure 8. Bottom: Magnified view (1 second) of the area enclosed by the rectangle in the trace above, showing the controlled movement of polynucleotides from the nanopore. The applied potentials are as follows: A and B: 120 mV, C: 0 mV, D, E and F: 120 mV, and the cycle is repeated. [Figure 10] The components of the experiment described in Example 2, in which both chains of the polynucleotide analyte are first rearranged through the nanopore without enzyme, then the enzyme “de-stalls”, and then the enzyme controls the movement of both chains of the polynucleotide analyte from the nanopore. A. Adapter including a hairpin portion and a 3'-TCCT overhang that specifically binds to one end of the polynucleotide analyte. B. Sequencing adapter identical to that described in Example 1 and Figure 6. C. Polynucleotide analyte with asymmetric ends, one having a 3'dA tail and the other having a 3'-AGGA overhang. The template chain and complementary chain are shown by dashed and solid lines, respectively. Ligation of DA, B, and C yields library molecule D. [Figure 11]This is a schematic diagram of the experiment in Example 2, showing the capture of both strands of a polynucleotide analyte, "de-stalling," and sequencing. Vs: sequencing potential, Vu: deblocking potential. The polarity of the applied potential (if not zero) is indicated by the arrow. The direction of the applied force is the same as the direction of the arrow. (A) Applying the sequencing potential (120mV). Capture of the polynucleotide analyte via the open pore and 3' leader (from Figure 7). Separation of the double strand by the nanopore, the template and complementary strands move into the transcompartment. (B) The polynucleotide reaches the enzyme, which is stalled in the spacer region. The enzyme cannot move over the spacer region. (C) Applying the deblocking potential (variable, 0mV to -120mV) so that the enzyme leaves the nanopore and moves freely over the spacer region. (D) Applying the sequencing potential (120mV). The polynucleotide moves through the nanopore until the enzyme reaches the nanopore, and then the enzyme controls the movement of the polynucleotide from the nanopore. (E) The DNA motor moves through the template portion and reaches the hairpin. (F) The DNA motor moves through the complementary portion, and the template and complementary strands refold in the cis-compartment. The motor reaches the leader section and idles in the nanopore. A deblocking potential is applied, and the DNA motor and analyte are ejected from the nanopore. [Figure 12] (a) A typical current-time trace of the data from Example 2, where the destamping voltage varies between 0 and -120mV. When the discharge potential increases above -60mV, no event is observed, suggesting that the hairpin formed in the transformer provides resistance to chain discharge up to this voltage. The portion where movement is controlled is shown as a box enclosed by a dashed line. (b) A typical current-time trace of the event described in Example 2. States A to G correspond to those described in Figure 11. [Figure 13]This is a representative current-time trace from Example 3, showing the capture of polynucleotide analytes into the nanopore and controlled movement from the nanopore. The DNA motor was “de-stalled” using the “active de-stalling” process described in Example 3. Asterisks indicate where the active stall potential was applied, with a 3-second pause between de-stalling attempts, initially up to 5 times at 5 seconds, then up to 5 times at 25 seconds. After the first attempt at 5 seconds, the enzyme de-stalled, as in Examples 1 and 2, controlling the movement of polynucleotides out of the nanopore with the template (Temp.) and complementary (Comp.) sections, followed by the leader state. A: Current-time trace showing the behavior of a “1D DNA library” similar to that described in Example 1, de-stalled after the first attempt. B: Current-time trace showing the behavior of a bound template-complementary polynucleotide (“2D DNA library”) de-stalled after 4 attempts, similar to that described in Example 2, bound by a hairpin portion. [Figure 14] The hairpin portion of the experiment described in Example 4 is used, in which both strands of the polynucleotide analyte are first rearranged through the nanopore without enzyme, then the enzyme is “de-stalled”, and then the enzyme controls the movement of both strands of the polynucleotide analyte from the nanopore. Additional portions of the hairpin introduce an additional signal during the initial enzyme-free capture step. These portions are illustrated as follows: (A) No portion in the hairpin as a control. (B) Hairpin with oligonucleotide i hybridized into the hairpin loop. (C) Three consecutive fluorescein dT bases ii in the hairpin loop, indicated by the asterisk. (D) As in (C), but using oligonucleotide hybridized into the hairpin loop. [Figure 15]This schematic diagram illustrates the capture and enzyme-free rearrangement of a double-stranded polynucleotide analyte with a hairpin portion, where the hairpin portion has an optionally large fluorophore and optionally an oligonucleotide that hybridizes into a hairpin loop. The schematic diagram shows two additional detectable intermediates, A1 and A2, which correspond to the oligonucleotide hybridized into the hairpin loop at the top of the nanopore using the fluorophore in the lumen of the nanopore, and the fluorophore in the lumen of the nanopore only. An additional state D1 corresponds to the fluorophore in the lumen of the nanopore and an enzyme moving along the fluorophore. [Figure 16(a)](a) Data showing the identification of the aenzymatic transfer of polynucleotides in which the template and complementary strands are linked via a hairpin portion. The polynucleotides are guided through the nanopore via an applied potential before the enzyme-controlled transfer step. The schematic diagram of the experiment is similar to that described in Example 2 and Figure 11. The hairpin is shown in Figure 14A. (i) A sequencing adapter containing only DNA and a polynucleotide library linked to the hairpin adapter. (ii) Representative current-time traces of the molecules shown in (i). The assignment of components A-G is based on states A-G described in Figure 11. (iii) A magnified view of the area enclosed by a rectangle shown in (ii), showing the identification of open pore level A and stall level B. The area enclosed by an asterisk has a different shape and noise than B, and also differs in relation to other representative molecules described in this example, and is presumed to arise from the aenzymatic rearrangement portion. (b) Data showing the identification of aenzymatic transfer of polynucleotides, where the template and complementary strands are linked via a hairpin portion, and the oligonucleotide hybridizes to the hairpin. The polynucleotide is guided through the nanopore via an applied potential before an enzyme-controlled transfer step. The schematic diagram of the experiment is similar to that described in Example 2 and Figure 11. The hairpin is shown in Figure 14B. (i) A polynucleotide library linked to a sequencing adapter and a hairpin adapter containing DNA hybridized with oligonucleotides (ON). (ii) Representative current-time traces of the molecules shown in (i). The assignment of components A-G is based on states A-G described in Figure 11. (iii) A magnified view of the area enclosed by a rectangle shown in (ii), showing the identification of open pore level A and stall level B. Compared to the example shown in Figure 16a, an additional level A2 (described in Figure 15) arises from the hybridized oligonucleotide. Thus, the asterisked region corresponds to aenzymatic rearrangement. (c) This data shows the identification of the non-enzymatic transfer of polynucleotides in which the template and complementary strands are linked via a hairpin region, and the hairpin has three bulky groups (three consecutive fluorescein dT bases, FAM).The polynucleotides are guided through the nanopore via an applied potential before the enzyme-controlled migration step. The schematic diagram of the experiment is similar to that described in Example 2 and Figure 11. The hairpin is shown in Figure 14C. (i) Polynucleotide library linked to a sequencing adapter containing fluorescein bases and a hairpin adapter. (ii) Representative current-time trace of the molecule shown in (i). The assignment of components A-G is based on states A-G described in Figure 11. The additional level D1 is presumed to result from the slow migration of the enzyme through the bulky FAM region. (State F is not seen in this example because the complementary region E is reduced for the efflux step G). (iii) Enlarged view of the region enclosed in the rectangle shown in (ii), showing the identification of open pore level A and stall level B. Compared to the example shown in Figure 16a, an additional downtick current level A1 of approximately 20 pA (see Figure 15) is generated from the FAM group. Thus, the asterisked region corresponds to an aenzymatic rearrangement. (d) Data showing the identification of non-enzymatic transfer of polynucleotides in which the template and complementary strands are linked via a hairpin region, with three bulky groups (three consecutive fluorescein-dT bases, FAM) present on the hairpin, where oligonucleotides (ON) hybridize. The polynucleotides are guided through the nanopore via an applied potential before an enzyme-controlled transfer step. The schematic diagram of the experiment is similar to that described in Example 2 and Figure 11. The hairpin is shown in Figure 14D. (i) Oligonucleotides (ON) hybridize to a polynucleotide library ligated to a sequencing adapter containing fluorescein bases (FAM) and a hairpin adapter. (ii) Representative current-time trace of the molecule shown in (i). The assignment of components A-G is based on states A-G described in Figure 11. An additional level D1 with a downtick in the current level is presumed to result from the slow movement of the enzyme in the bulky FAM region. (iii)(ii) is an enlarged view of the area enclosed by the rectangle, showing the distinction between open pore level A and stall level B.Compared to the examples shown in Figures 16a and 16c, an additional downtick current level A1 (see Figure 15) of approximately 20 pA is generated from the FAM group. Compared with Figure 16b, an additional level A2 due to hybridized ON is also observed. Therefore, the asterisked region corresponds to an enzymatic rearrangement. (e) Measurement of the duration of enzymatic rearrangement in the E. coli test library. (i) Four representative examples from the random E. coli test library described in Example 4, where double-stranded polynucleotides are ligated to a sequencing adapter at one end and to a hairpin portion at the other end. Oligonucleotides are hybridized to the hairpin portion. Therefore, the resulting polynucleotides are similar to those in Figure 16b, except that the polynucleotides are of random length. The four examples shown are event-fitted current-time traces to simplify the raw data. Level A2 and the enzymatic portion (indicated by an asterisk) are shown in each example. A threshold of 60 pA (dotted line) was used to distinguish the enzyme-free portion A2. Therefore, the periods marked with an asterisk were measured as the time during which the threshold of 60 pA between the pore level A (where the current opened) and the oligonucleotide level A2 exceeded. (ii) The relationship between the enzyme-controlled chain duration (measured as the sum of periods D and E shown in Figure 16b, ii) and the unenzymatic capture duration (measured as described in Part i of this figure) was measured for 30 examples and is shown as a scatter plot. A linear regression line with an R2 value of 0.414 is shown, indicating a positive correlation. [Figure 16(b)] (As stated above.) [Figure 16(c)] (As stated above.) [Figure 16(d)] (As stated above.) [Figure 16(e)] (As stated above.) [Figure 17(a)](a) A nanopore sequencing adapter having a DNA helicase that moves from 5' to 3', where the 3' strand is preferentially captured by the nanopore. The enzyme stalls via another blocker strand containing a BNA region, and the helicase stalls via a spacer portion on the strand loaded with it. The adapter contains oligonucleotides known as the top strand (A), bottom strand (B), blocker strand (C), and back blocker (D). Both the blocker strand and the back blocker hybridize to the double-stranded top strand forming region. The DNA motor (direction 5'-3') is loaded into a closed poly(dT) binding site (E) in the single-stranded region between C and D and stalls via an internal spacer 18 portion (F). The top strand has a 3'dT base for ligation into a dA-tailed double strand. The bottom chain includes a 5' phosphate moiety (circled P), 20 consecutive thymidine bases as a leader (wavy line G), and a site for hybridizing the hydrophobic tether (H). (b) A schematic diagram showing the sequencing adapter (A) described in Figure 17a, ligated at both ends to a double-stranded polynucleotide analyte (B). (c) A schematic diagram of the experiment in Example 5 showing the capture, destallation, and sequencing of the polynucleotide analyte. Vs: sequencing potential, Vu: deblocking potential. The polarity of the applied potential is indicated by the arrow. The direction of the applied force is the same as the direction of the arrow. (A) Applying the sequencing potential (120mV). Capture of the polynucleotide analyte via the open pore and 3' leader (from Figure 7). Separation of the double strand by the nanopore, and the complementary strand is removed. (B) The nanopore is temporarily stalled at the blocker strand portion. (C) The polynucleotide reaches the enzyme stalled at the spacer portion. The enzyme cannot move over the spacer portion. (D) A deblocking potential is applied (0mV) to allow the enzyme to move away from the nanopore and move freely over the spacer portion. (E) A sequencing potential is applied (120mV). The nanopore moves polynucleotides until the enzyme reaches the nanopore, after which the enzyme controls the movement of polynucleotides from the nanopore. (F) The DNA motor reaches the leader and becomes idle. (G) A deblocking potential is applied, and the DNA motor and analyte are ejected from the nanopore. Repeat the cycle from (A).(d)i, a representative current-time trace of Example 5, showing the capture of polynucleotide analytes into and controlled movement from the nanopore using an adapter in which the biotin-traptabidine backblocker is replaced with another backblocker oligonucleotide, as described in Figures 17a and 17b. The DNA motor was “destalled” using the “activation destallation” process described in Example 5 and the earlier Example 3. Levels A to G (described in Figure 17c) are assigned in relation to the previous example. The squared regions ii (enzyme-free rearrangement) and iii (enzyme-controlled rearrangement) are shown in magnified view. [Figure 17(b)] (As stated above.) [Figure 17(c)] (As stated above.) [Figure 17(d)] (As stated above.) [Figure 18(a)](a) Schematic diagram of Example 6, showing the capture, destallation, and sequencing of both strands of the polynucleotide analyte, with occasional rereading of the strands. Vs: sequencing potential, Vu: deblocking potential. The polarity of the applied potential is indicated by the arrow. The direction of the applied force is the same as the direction of the arrow. (A) Applying the sequencing potential (120mV). Capture of the polynucleotide analyte via the open pore and 3' reader (from Figure 7). Separation of the double strand by the nanopore, the template and complementary strands move into the transcompartment. (B) The polynucleotide reaches the enzyme stalled in the spacer region. The enzyme cannot move over the spacer region. (C) Applying the deblocking potential (variable, 0mV to -120mV) so that the enzyme detaches from the nanopore and moves freely over the spacer region. (D) Applying the sequencing potential (120mV). The nanopore moves the polynucleotide until the enzyme reaches the nanopore, and then the enzyme controls the movement of the polynucleotide from the nanopore. (E) The DNA motor moves the template portion and reaches the hairpin. (F) The DNA motor moves the complementary portion, and the template and complementary strands refold in the cis-compartment. The motor reaches the leader section and idles in the nanopore. State (F) is pushed back to state (E) by the enzyme from 3'-5', enabling a strand reread (RR). (G) A deblocking potential is applied, and the DNA motor and analyte are ejected from the nanopore. (b) A representative current-time trace from Example 6, showing an example where the polynucleotide enzyme is read twice via the enzyme being pushed backward from the C3 reader under the applied potential. Enlarged view of enzyme-controlled portions (i) and (ii) and the C3 level is also identified. (c) Six representative reread examples from the experiment described in Example 6. The enzyme-regulated regions were mapped using an HMM model trained with data from the pore and enzyme combinations used. The example readings indicate that the same strand of the bacteriophage lambda DNA mixture of seven restriction enzyme fragments is mapped at least twice. [Figure 18(b)] (As stated above.) [Figure 18(c)] (As stated above.) [Figure 19(a)] (a) An example of a typical HMM mapping of the data described in Example 7, using data collected at a sequencing potential of 120 mV. (b) An example of a typical HMM mapping of the data described in Example 7, using data collected at a sequencing potential of 140 mV. (c) An example of a typical HMM mapping of the data described in Example 7, using data collected at a sequencing potential of 160 mV. (d) Histograms of single-molecule enzyme rates extracted from the data in Figures 19a, 19b, and 19c. The number of molecules in each population is shown. The medians for each population are as follows: 120 mV, 319 bp / sec, 140 mV, 259 bp / sec, 160 mV, 196 bp / sec. [Figure 19(b)] (As stated above.) [Figure 19(c)] (As stated above.) [Figure 19(d)] (As stated above.) [Figure 20(a)] (a) This is an experimental diagram, identical to Figure 18a / Example 6. In addition, the “entry” stage used to measure the enzyme-free rearrangement (between steps A and C) is marked with an asterisk. (b) These are representative current-time traces of the three library examples shown in Example 8, with a 10kb PCR fragment (top), bacteriophage lambda DNA (center), and T4 DNA (bottom). Full-length readings of T4 DNA are not recorded; examples of partial fragments are shown. In each example, the “entry” stage is marked with an asterisk, and the enzyme-controlled stage is marked with an E. The duration of each portion is measured manually and marked on the trace. A magnified view of the entry stage of the T4 example is shown. The portion marked with B in Figure 20a (blocker oligonucleotide at the top of the pore) cannot be reliably detected. (c) This is a log-log scatter plot of measured capture times measured from the traces of the 31 examples described in Example 8. Markers are colored grayscale according to the library of origin. [Figure 20(b)] (As stated above.) [Figure 20(c)] (As stated above.) [Modes for carrying out the invention]
[0034] The present invention is described with respect to specific embodiments and with reference to certain drawings, but the present invention is not limited thereto and is limited only by the claims. None of the reference numerals in the claims should be construed as limiting the scope. Needless to say, it should be understood that not all aspects or advantages can necessarily be achieved according to any particular embodiment of the present invention. Accordingly, for example, a person skilled in the art will recognize that the present invention can be embodied or performed in a manner that achieves or optimizes one or a group of advantages taught herein without necessarily achieving other aspects or advantages that can be taught or suggested herein.
[0035] The present invention, with regard to both its organization and method of operation, along with its features and advantages, is best understood by referring to the embodiments for carrying out the invention as described below, when read in conjunction with the accompanying drawings. The aspects and advantages of the present invention will become apparent and clarified by referring to the embodiments described below. Throughout this specification, any reference to “one embodiment” or “a certain embodiment” means that a particular feature, structure, or characteristic described in relation to that embodiment is included in at least one embodiment of the present invention. Thus, the appearance of the phrase “in one embodiment” or “in a certain embodiment” in various places throughout this specification does not necessarily all refer to the same embodiment, although it may. Similarly, in the description of exemplary embodiments of the present invention, it should be understood that various features of the present invention may be summarized in a single embodiment, figure, or description thereof in order to simplify this disclosure and to aid in the understanding of one or more of the various embodiments of the invention. However, the method of this disclosure should not be interpreted as reflecting an intention that the claimed invention requires more features than those explicitly enumerated in each claim. Rather, as reflected in the following claims, the embodiments of the invention are less than all the features of a single, aforementioned disclosed embodiment.
[0036] Unless otherwise indicated by the context, it should be understood that the “embodiments” of this disclosure may be specifically combined together. Any specific combination of all disclosed embodiments (unless otherwise implied by the context) constitutes a further disclosed embodiment of the claimed invention.
[0037] In addition, as used herein and in the appended claims, the singular forms "a," "an," and "the" include multiple referents unless the context otherwise explicitly indicates. Thus, for example, a reference to "polynucleotide" includes two or more polynucleotides, a reference to "motor protein" includes two or more such proteins, a reference to "helicase" includes two or more helicases, a reference to "monomer" refers to two or more monomers, and a reference to "pore" includes two or more pores.
[0038] All publications, patents, and patent applications cited above or below in this Spec. are incorporated herein by reference in their entirety.
[0039] definition When an indefinite or definite article, such as "a" or "an" or "the," is used to refer to a singular noun, unless otherwise specified, this includes the plural form of that noun. When the term "includes" is used in this description and claims, it does not exclude other elements or steps. Furthermore, terms such as first, second, third, etc., in this description and claims are used to distinguish similar elements and are not necessarily used to describe the order in which they occurred or the order in which they occurred. It should be understood that such terms are interchangeable under appropriate circumstances and that embodiments of the invention described herein may operate in an order other than that described or illustrated herein. The following terms or definitions are provided solely to aid in the understanding of the invention. Unless specifically defined herein, all terms used herein have the same meaning as understood by those skilled in the art. Those skilled in the art should refer to the definitions and technical terms, in particular Sambrook et al., Molecular Cloning: A Laboratory Manual, 4 th See ed., Cold Spring Harbor Press, Plainsview, New York (2012), and Ausubel et al., Current Protocols in Molecular Biology (Supplement 114), John Wiley & Sons, New York (2016). The definitions provided herein should not be construed as narrower than those understood by those skilled in the art.
[0040] As used herein, "about" when referring to measurable values such as quantity or temporal duration means that it includes a variation of ±20% or ±10%, more preferably ±5%, even more preferably ±1%, and still more preferably ±0.1% from the specified value, such variation being appropriate for carrying out the disclosed method.
[0041] As used herein, “nucleotide sequence,” “DNA sequence,” or “nucleic acid molecule” refers to a polymeric form of nucleotides of any length, either ribonucleotides or deoxyribonucleotides. This term refers only to the primary structure of the molecule. Therefore, this term includes double-stranded and single-stranded DNA and RNA. As used herein, the term “nucleic acid” is a single-stranded or double-stranded covalent nucleotide sequence in which the 3' and 5' ends of each nucleotide are linked by phosphodiester bonds. Polynucleotides may consist of deoxyribonucleotide bases or ribonucleotide bases. Nucleic acids may be produced synthetically in vitro or isolated from natural sources. Nucleic acids may further include modified DNA or RNA, e.g., methylated DNA or RNA, or RNA subjected to post-translational modifications, e.g., 3'-processing such as 5'-capping, cleavage, and polyadenylation with 7-methylguanosine, and splicing. Nucleic acids may also include synthetic nucleic acids (XNAs) such as hexitol nucleic acids (HNAs), cyclohexene nucleic acids (CeNAs), threose nucleic acids (TNAs), glycerol nucleic acids (GNAs), locked nucleic acids (LNAs), and peptide nucleic acids (PNAs). The size of nucleic acids, also referred to herein as “polynucleotides,” is typically expressed in terms of the number of base pairs (bp) for double-stranded polynucleotides, or the number of nucleotides (nts) for single-stranded polynucleotides. 1000 bp or nts corresponds to kilobases (kb). Polynucleotides with a length of less than approximately 40 nucleotides are typically called “oligonucleotides” and may include primers used for manipulating DNA, such as through polymerase chain reaction (PCR).
[0042] In the context of this disclosure, the term “amino acid” is used in its broadest sense and is intended to include organic compounds containing amine (NH2) and carboxyl (COOH) functional groups, along with a side chain specific to each amino acid (e.g., an R group). In some embodiments, an amino acid refers to a naturally occurring L α-amino acid or residue. The one- and three-letter abbreviations commonly used for naturally occurring amino acids are: A=Ala, C=Cys, D=Asp, E=Glu, F=Phe, G=Gly, H=His, I=Ile, K=Lys, L=Leu, M=Met, N=Asn, P=Pro, Q=Gln, R=Arg, S=Ser, T=Thr, V=Val, W=Trp, and Y=Tyr as used herein (Lehninger, AL, (1975) Biochemistry, 2nd ed., pp. 71-92, Worth Publishers, New York). The general term “amino acid” further includes chemically modified amino acids such as D-amino acids, retro-inversoamino acids, and amino acid analogs, naturally occurring amino acids not typically incorporated into proteins such as norleucine, and chemically synthesized compounds that exhibit amino acid characteristics known in the art, such as β-amino acids. For example, analogs or mimics of phenylalanine or proline that allow for the same stereochemical constraints on peptide compounds as natural Phe or Pro are included within the definition of an amino acid. Such analogs and mimics are referred to herein as “functional equivalents” of their respective amino acids. Other examples of amino acids are listed by reference in Roberts and Vellaccio, The Peptides: Analysis, Synthesis, Biology, Gross and Meiehofer, eds., Vol. 5 p. 341, Academic Press, Inc., NY 1983.
[0043] The terms “polypeptide” and “peptide” are used interchangeably herein to refer to polymers of amino acid residues, as well as their variants and synthetic analogs. Therefore, these terms apply to amino acid polymers, where one or more amino acid residues are synthetic amino acids that do not exist naturally, such as chemical analogs of corresponding naturally occurring amino acids, as well as naturally occurring amino acid polymers. Polypeptides may undergo maturation or post-translational modification processes, including but not limited to glycosylation, proteolytic cleavage, lipidation, signal peptide cleavage, propeptide cleavage, and phosphorylation. Peptides can be prepared using recombination techniques, for example, by the expression of recombinant or synthetic polynucleotides. Peptides produced by recombination typically contain substantially no culture medium; for example, the culture medium constitutes less than about 20%, more preferably less than about 10%, and most preferably less than about 5% of the volume of the protein preparation.
[0044] The term "protein" is used to describe folded polypeptides that have a secondary or tertiary structure. A protein may consist of a single polypeptide or may contain multiple polypeptides that assemble to form a multimer. A multimer may be a homooligomer or a heterooligomer. A protein may be a naturally occurring protein or a wild-type protein, or it may be a modified protein or a protein that does not exist in nature. A protein may differ from a wild-type protein, for example, by the addition, substitution, or deletion of one or more amino acids.
[0045] Protein "mutants" include peptides, oligopeptides, polypeptides, proteins, and enzymes that have amino acid substitutions, deletions, and / or insertions compared to the unmodified or wild-type protein in question, and that have similar biological and functional activity to the unmodified protein from which they are derived. As used herein, the term "amino acid identity" refers to the degree to which sequences are identical amino acid-wise across a comparison window. Thus, the "percentage of sequence identity" is calculated by comparing two optimally aligned sequences across a comparison window, determining the number of positions in which identical amino acid residues (e.g., Ala, Pro, Ser, Thr, Gly, Val, Leu, Ile, Phe, Tyr, Trp, Lys, Arg, His, Asp, Glu, Asn, Gln, Cys, and Met) occur in both sequences, calculating the number of matching positions, dividing the number of matching positions by the total number of positions in the comparison window (i.e., the window size), and multiplying the result by 100 to calculate the percentage of sequence identity.
[0046] In all aspects and embodiments of the present invention, the "mutant" has at least 50%, 60%, 70%, 80%, 90%, 95%, or 99% complete sequence identity with respect to the amino acid sequence of the corresponding wild-type protein. Sequence identity may also be to a full-length polynucleotide or a fragment or portion of a polypeptide. Thus, a sequence may have only 50% sequence identity overall with a full-length reference sequence, while the sequence of a particular region, domain, or subunit may share as much as 80%, 90%, or 99% sequence identity with the reference sequence.
[0047] The term "wild-type" refers to a gene or gene product isolated from a naturally occurring source. A wild-type gene is the most frequently observed and therefore arbitrarily designed "normal" or "wild-type" form of that gene in a given population. In contrast, the terms "modified," "mutant," or "variant" refer to a gene or gene product that exhibits sequence modifications (e.g., substitutions, cleavage, or insertions), post-translational modifications, and / or functional characteristics (e.g., altered features) compared to a wild-type gene or gene product. It should be noted that naturally occurring mutants can be isolated, and these are identified by the fact that they have altered features compared to a wild-type gene or gene product. Methods for introducing or substituting naturally occurring amino acids are well known in the art. For example, methionine (M) can be substituted with arginine (R) by replacing the methionine codon (ATG) with the arginine codon (CGT) at the relevant position in the polynucleotide encoding the mutant monomer. Methods for introducing or substituting amino acids that do not exist in nature are also well known in the art. For example, amino acids that do not exist in nature can be introduced by including synthetic aminoacyl-tRNA in the IVTT system used to express mutant monomers. Alternatively, they can be introduced by expressing mutant monomers in E. coli that are nutrientally required for those particular amino acids, in the presence of synthetic (i.e., non-naturally occurring) analogs of those particular amino acids. They may also be produced by naked ligation when the mutant monomers are produced using partial peptide synthesis. Conservative substitutions replace an amino acid with another amino acid having a similar chemical structure, similar chemical properties, or similar side-chain volume. The introduced amino acids may have similar polarity, hydrophilicity, hydrophobicity, basicity, acidity, neutrality, or charge to the amino acid they replace. Alternatively, conservative substitutions can introduce another amino acid that is aromatic or aliphatic in place of an existing aromatic or aliphatic amino acid. Conservative amino acid changes are well known in the art and can be selected according to the properties of 20 major amino acids defined in Table 1 below.If amino acids have similar polarity, this can also be determined by referring to the hydrophobicity scale of the amino acid side chains in Table 2. [Table 1] [Table 2]
[0048] Mutant or modified proteins, monomers, or peptides can also be chemically modified in any manner and at any site. Preferably, mutant or modified monomers or peptides are chemically modified by attachment of molecules to one or more cysteine molecules (cysteine bonding), attachment of molecules to one or more lysine molecules, attachment of molecules to one or more non-natural amino acids, enzymatic modification of epitopes, or terminal modification. Suitable methods for carrying out such modifications are well known in the art. Mutants of modified proteins, monomers, or peptides can be chemically modified by attachment of any molecule. For example, mutants of modified proteins, monomers, or peptides can be chemically modified by attachment of dyes or fluorophores.
[0049] As used herein, an alkylene group is an unsubstituted or substituted, saturated bidentate moiety obtained by removing two hydrogen atoms from the same carbon atom, or any two hydrogen atoms from each of two different carbon atoms, from a hydrocarbon compound that may be aliphatic or alicyclic. The hydrocarbon compound may have 1 to 20 carbon atoms, in which case the alkylene group is C 1-20 It is an alkylene. The alkylene group is C 1-10 If it is an alkylene, it may have, for example, 1 to 10 carbon atoms. Typically, C 1-6 Alkylene, or C 1-4 Alkylenes include, for example, methylene, ethylene, i-propylene, n-propylene, t-butylene, s-butylene, or n-butylene.
[0050] An alkenylene group is an unsubstituted or substituted bidentate moiety obtained by removing two hydrogen atoms from the same carbon atom or one hydrogen atom from each of two different carbon atoms of a hydrocarbon compound which may be aliphatic or alicyclic, and contains one or more carbon-carbon double bonds. The hydrocarbon compound may have 2 to 20 carbon atoms, in which case the alkenylene group is C 2-20 alkenylene. When the alkenylene group is C 2-10 alkenylene, it may have, for example, 2 to 10 carbon atoms. Usually, it is C 2-6 alkenylene, or C 2-4 alkenylene.
[0051] An alkynylene group is an unsubstituted or substituted bidentate moiety obtained by removing two hydrogen atoms from the same carbon atom or one hydrogen atom from each of two different carbon atoms of a hydrocarbon compound which may be aliphatic or alicyclic, and contains one or more carbon-carbon triple bonds. The hydrocarbon compound may have 2 to 20 carbon atoms, in which case the alkynylene group is C 2-20 alkynylene. When the alkynylene group is C 2-10 alkynylene, it may have, for example, 2 to 10 carbon atoms. Usually, it is C 2-6 alkynylene, or C 2-4 alkynylene.
[0052] An arylene group is an unsubstituted or substituted monocyclic or fused polycyclic bidentate moiety obtained by removing one hydrogen atom each from two different aromatic ring atoms of an aromatic compound, and the moiety has (unless otherwise specified) 5 to 14 ring atoms. Typically, each ring has 5 to 7 or 5 to 6 ring atoms. The arylene group may be unsubstituted or substituted.
[0053] A heteroarylene group is a bidentate moiety obtained by removing two hydrogen atoms, one from each of two different ring atoms of a heteroaryl group. Heteroaryl groups are substituted or unsubstituted monocyclic or fused polycyclic (e.g., bicyclic or tricyclic) aromatic groups that typically contain 5 to 14 ring atoms, with at least one heteroatom in the ring portion, e.g., 1, 2, or 3 heteroatoms selected from O, S, N, P, Se, and Si, more typically O, S, and N. Examples include pyridyl, pyrazinyl, pyrimidinyl, pyridadinyl, furanyl, thienyl, pyrazolidinyl, pyrrolyl, oxadiazolyl, isoxazolyl, thiadiazolyl, thiazolyl, imidazolyl, triazolyl, pyrazolyl, oxazolyl, isothiazolyl, benzofuranyl, isobenzofuranyl, benzothiophenyl, indolyl, indazolyl, carbazolyl, acridinyl, urinyl, sinnolinyl, quinoxalinyl, naphthylidinyl, benzimidazolyl, benzoxazolyl, quinolinyl, quinazolinyl, and isoquinolinyl.
[0054] Carbocyclylene groups, also known as cycloalkylene groups, are bidentate moieties obtained by removing two hydrogen atoms, one from each of the two carbon atoms of an unsubstituted or substituted cyclic alkyl group. Typically, these moieties contain 3 to 10 ring atoms and 3 to 10 carbon atoms (unless otherwise specified). Examples include cyclopropane (C3), cyclobutane (C4), cyclopentane (C5), cyclohexane (C6), cycloheptane (C7), methylcyclopropane (C4), dimethylcyclopropane (C5), methylcyclobutane (C5), dimethylcyclobutane (C6), methylcyclopentane (C6), dimethylcyclopentane (C7), methylcyclohexane (C7), dimethylcyclohexane (C8), and menthane (C10).
[0055] The heterocyclylene moiety is a bidentate moiety obtained by removing two hydrogen atoms from two different ring atoms of a heterocyclyl group. Heterocyclyl groups are unsubstituted or substituted cyclic groups and typically contain 5 to 14 atoms, with at least one heteroatom in the ring portion, e.g., one, two, or three heteroatoms selected from O, S, N, P, Se, and Si, more typically from O, S, and N. Examples include piperazine, piperidine, morpholin, 1,3-oxazinane, pyrrolidine, imidazolidine, oxazolidine, tetrahydropyrazine, tetrahydropyridine, dihydro-1,4-oxazine, tetrahydropyrimidine, dihydro-1,3-oxazine, dihydropyrrole, dihydroimidazole, and dihydrooxazole groups.
[0056] An arylene-alkylene group is a group formed by forming a bond between an arylene group and an alkylene group as defined herein. A heteroarylene-alkylene group is a group formed by forming a bond between a heteroarylene group and an alkylene group as defined herein. A carbocyclylene-alkylene group is a group formed by forming a bond between a carbocyclylene group and an alkylene group as defined herein. A heterocyclylene-alkylene group is a group formed by forming a bond between a heterocyclylene group and an alkylene group as defined herein.
[0057] When a group is described as substituted, it is typically substituted by one or more substituents, such as one, two, or three, usually one or two, and most commonly one substituent. Preferred substituents can be independently selected from halogens, -OR', and -NR'2 (wherein R' is typically H or unsubstituted C). 1-2 (Alkyl and unsubstituted C1-C2 alkyl groups).
[0058] Methods for characterizing analytes This disclosure relates to a method for characterizing target polynucleotides that move toward a detector such as a nanopore by using a motor protein. Any suitable motor protein may be used in the manner provided herein. Exemplary motor proteins are described in more detail herein.
[0059] This disclosure also relates to a method for characterizing a target polynucleotide, which includes contacting the polynucleotide with a detector and rereading the polynucleotide, such as by moving the polynucleotide back and forth relative to the detector. This is described in more detail herein.
[0060] More specifically, in some embodiments, the disclosure relates to a method by which a motor protein moves polynucleotides out of a detector (e.g., out of a nanopore). Thus, the direction of polynucleotide movement in such embodiments is opposite to known methods by which polynucleotides move inward of a nanopore. This is described in more detail herein.
[0061] While this disclosure provides nanopores as exemplary detectors, the methods provided herein are suitable for detectors such as (i) zero-mode waveguides, (ii) field-effect transistors, optionally Noy field-effect transistors, (iii) AFM chips, (iv) nanotubes, optionally carbon nanotubes, and (v) nanopores. The disclosed methods are particularly suitable for methods of moving polynucleotides through a detector or through a structure containing a detector, such as wells in a detector chip.
[0062] In the disclosed method, the motor protein is typically first stalled on a polynucleotide at a stall region. Preferred stall regions are described in more detail herein. Stalling a motor protein on a polynucleotide offers several advantages. For example, while stalled, the motor protein typically consumes less fuel than when not stalled, for example, when freely moving with respect to the polynucleotide. This reduction in unproductive fuel consumption can be advantageous.
[0063] The methods provided herein typically involve destalling a motor protein so that the motor protein can control the movement of polynucleotides from a detector (e.g., a nanopore). Methods for destalling a motor protein are described in more detail herein. Controlled destalling of a motor protein has several advantages, including the ability to precisely determine the point at which the motor protein begins processing polynucleotides. This can be useful, for example, for characterizing polynucleotides so that data is not lost as a result of undesirable movement of the motor protein over the polynucleotide before data recording begins.
[0064] The disclosed method is at least partially based on the recognition that data obtained when polynucleotides are moved from a detector such as a nanopore may differ from data obtained when the same polynucleotides are moved into a detector (e.g., a nanopore). Data characteristics, including signal profiles, noise profiles, and error profiles, may, in some embodiments, differ from those obtained by contrasting methods in which the same polynucleotides are moved into a detector such as a nanopore. In some embodiments, data obtained by the disclosed method have advantages compared to data obtained by other known methods. Therefore, the disclosed method increases the available options when polynucleotide characterization is required. Thus, users who wish to characterize polynucleotides can select the method best suited to their specific application.
[0065] As described above, the disclosed method relates in some embodiments to transferring a target polynucleotide from a detector such as a nanopore. While nanopores are considered as exemplary detectors in this specification, the method is not limited thereto.
[0066] Nanopores typically have two openings, namely a first opening and a second opening. These openings are often referred to as the cis and trans openings of the nanopore. In many cases, the first opening is a cis opening and the second opening is a trans opening, but in some embodiments, the first opening is a trans opening and the second opening is a cis opening. The notation of "cis" and "trans" openings of a nanopore is routine in the art. For example, the cis opening of a nanopore typically faces the cis chamber of a nanopore device, such as the apparatus described herein which has cis and trans chambers, and the trans opening typically faces the trans chamber.
[0067] In a particular method provided herein, a first opening of a nanopore contacts a polynucleotide having a stalled motor protein on it. The method involves using the motor protein to control the movement of a target polynucleotide through the nanopore in the direction from a second opening of the nanopore to the first opening of the nanopore.
[0068] Therefore, from the motor protein's perspective, the target polynucleotide moves out of the nanopore. The notation "outward" refers to the overall movement of the polynucleotide toward the motor protein. This direction of movement may be in contrast to another mode in which the target polynucleotide moves "inward" of the nanopore by the motor protein.
[0069] The differences in these kinetic schemes are significant. In the method provided herein, where polynucleotides move "out" of the pore, the direction of movement is from the nanopore entrance furthest from the motor protein (i.e., the distal entrance) to the nanopore entrance closest to the motor protein (the proximal entrance). In contrast, in the method where polynucleotides move "in" of the pore, the direction of movement is from the nanopore entrance closest to the motor protein (the proximal entrance) to the nanopore entrance furthest from the motor protein (the distal entrance).
[0070] Accordingly, in some embodiments of the provided method, the nanopore spans a membrane having a cis side and a trans side, with the first opening of the nanopore on the cis side of the membrane and the second opening of the nanopore on the trans side. In such embodiments, a motor protein is located on the cis side of the membrane and controls the movement of a target polynucleotide through the nanopore from the trans side to the cis side of the membrane.
[0071] In another embodiment of the provided method, the nanopore spans a membrane having a cis side and a trans side, with the first opening of the nanopore on the trans side of the membrane and the second opening of the nanopore on the cis side. In such an embodiment, the motor protein is located on the trans side of the membrane and controls the movement of a target polynucleotide through the nanopore from the cis side to the trans side of the membrane.
[0072] In contrast to the movement of polynucleotides into the pore in the contrasting method, Figure 1 schematically illustrates the difference in the direction of movement of polynucleotides from the pore in the method provided herein.
[0073] Reread In some embodiments, the methods provided herein include rereading a polynucleotide to characterize it. Rereading a polynucleotide involves obtaining one or more measurements characteristic of the polynucleotide as it moves back and forth relative to a detector.
[0074] In one embodiment, this specification provides a method for characterizing a target polypeptide, which is (i) Contacting the detector with a target polynucleotide to which a motor protein is bound, wherein the target polynucleotide is bound to the motor protein at the polynucleotide binding site of the motor protein, (ii) Obtaining one or more characteristic measurements of the target polynucleotide when the motor protein controls the movement of the target polynucleotide in a first direction relative to the detector, (iii) Debinding the target polynucleotide from the polynucleotide binding site of the motor protein so that the target polynucleotide moves in a second direction relative to the detector, (iv) When the target polynucleotide is re-bound to the polynucleotide binding site of the motor protein and the motor protein controls the movement of the target polynucleotide in a first direction relative to the detector, one or more characteristic measurements of the target polynucleotide are obtained. This includes characterizing the target polynucleotide.
[0075] In related embodiments, this specification provides a method for characterizing a target polypeptide, which is: (i) Contacting the detector with a target polynucleotide to which a motor protein is bound, wherein the target polynucleotide is bound to the motor protein at the polynucleotide binding site of the motor protein, (ii) Obtaining one or more characteristic measurements of the target polynucleotide when the motor protein controls the movement of the target polynucleotide in a first direction relative to the detector, (iii) Allowing the target polynucleotide to dissociate from the polynucleotide binding site of the motor protein so that the target polynucleotide moves in a second direction relative to the detector, (iv) When the target polynucleotide is re-bound to the polynucleotide binding site of the motor protein and the motor protein controls the movement of the target polynucleotide in a first direction relative to the detector, one or more characteristic measurements of the target polynucleotide are obtained. This includes characterizing the target polynucleotide.
[0076] The disclosed method offers several advantages over conventionally known methods. For example, each read of a target polynucleotide should have equivalent precision because the same chain and detection region are used. This allows the same base calling model to be used for each read. It also facilitates the joining of data from multiple reads. Furthermore, because the native sequence is reread multiple times, it is possible to preserve (e.g.) epigenetic information. This method is also adaptive, allowing for multiple rereads until data of the required precision is obtained.
[0077] More specifically, the method may involve obtaining one or more measurements characteristic of a target polynucleotide when a motor protein controls the movement of the target polynucleotide in a first direction relative to the detector. The first direction may be the direction in which the motor protein drives the movement of the polynucleotide. The first direction may be the direction of a force applied to the detector. The first direction may be the direction opposite to the direction of a force applied across the detector.
[0078] In many cases, the detector comprises a structure having a first opening and a second opening, or a transmembrane nanopore having a first opening and a second opening, and step (i) comprises contracting the first opening with a target polynucleotide. Typically, a motor protein controls the movement of the target polynucleotide in the direction from the second opening to the first opening. Typically, when the target polynucleotide detaches from the polynucleotide binding site of the motor protein, the target polynucleotide moves in the direction from the first opening to the second opening.
[0079] Therefore, if the detector is a nanopore or contains a nanopore, the first direction may be “into” the nanopore as described herein. Thus, in some embodiments, the movement of polynucleotides while one or more measurements are being performed is made outward from the nanopore. In some embodiments, the nanopore spans a membrane having a cis side and a trans side, the first opening of the nanopore is on the cis side of the membrane, the second opening of the nanopore is on the trans side, and the motor protein is located on the cis side of the membrane and controls the movement of the target polynucleotide through the nanopore from the cis side to the trans side of the membrane. In other embodiments, the nanopore spans a membrane having a cis side and a trans side, the first opening of the nanopore is on the cis side of the membrane, the second opening of the nanopore is on the trans side, and the motor protein is located on the trans side of the membrane and controls the movement of the target polynucleotide through the nanopore from the trans side to the cis side of the membrane.
[0080] More often, if the detector is a nanopore or contains a nanopore, the first direction is the direction "outward" of the nanopore as described herein. Thus, in some embodiments, the movement of polynucleotides while one or more measurements are being performed is made outward of the nanopore. In some embodiments, the nanopore spans a membrane having a cis side and a trans side, the first opening of the nanopore is on the cis side of the membrane, the second opening of the nanopore is on the trans side, and the motor protein is located on the cis side of the membrane and controls the movement of the target polynucleotide through the nanopore from the trans side to the cis side of the membrane. In other embodiments, the nanopore spans a membrane having a cis side and a trans side, the first opening of the nanopore is on the cis side of the membrane, the second opening of the nanopore is on the trans side, and the motor protein is located on the trans side of the membrane and controls the movement of the target polynucleotide through the nanopore from the cis side to the trans side of the membrane.
[0081] The provided method may include debinding a target polynucleotide from a polynucleotide binding site of a motor protein, which is described in detail below herein. When the target polynucleotide is debound from the polynucleotide binding site of the motor protein, the target polynucleotide moves in a second direction relative to the detector. The second direction is typically opposite to the first direction.
[0082] Therefore, in some embodiments of the detector being a nanopore or a method comprising a nanopore, the first direction in which the target polynucleotide moves relative to the detector is into the nanopore, and the second direction in which the target polynucleotide moves relative to the detector is out of the nanopore. In other embodiments, the first direction in which the target polynucleotide moves relative to the detector is outward through the nanopore, and the second direction in which the target polynucleotide moves relative to the detector is inward through the nanopore.
[0083] The provided method may then include rejoining the target polynucleotide to a polynucleotide binding site of a motor protein. The motor protein then controls the movement of the target polynucleotide in a first direction when one or more measurements characteristic of the polynucleotide are performed. The first direction is the same as the first direction described above.
[0084] Therefore, in one embodiment, this specification provides a method for characterizing a target polynucleotide, and this method is (i) Contacting the first opening of a transmembrane nanopore having a first opening and a second opening with a target polynucleotide having a motor protein bound thereto, wherein the target polynucleotide is bound to the motor protein at the polynucleotide binding site of the motor protein, (ii) Obtaining one or more characteristic measurements of the target polynucleotide when the motor protein controls the movement of the target polynucleotide from the first opening of the nanopore to the second opening of the nanopore, (iii) Debinding the target polynucleotide from the polynucleotide binding site of the motor protein so that the target polynucleotide moves from the second opening of the nanopore to the first opening of the nanopore, (iv) Re-binding the target polynucleotide to the polynucleotide binding site of the motor protein, and obtaining one or more characteristic measurements of the target polynucleotide as the motor protein controls the movement of the target polynucleotide in the direction from the first opening of the nanopore to the second opening of the nanopore, This includes characterizing the target polynucleotide. Characterizing the target polynucleotide may include, for example, determining the sequence of the target polynucleotide.
[0085] For example, in some embodiments, the nanopore spans a membrane having a cis side and a trans side, with the first opening of the nanopore on the cis side of the membrane and the second opening on the trans side, and a motor protein controls the movement of the target polynucleotide through the nanopore from the cis side to the trans side of the membrane. In other embodiments, the first opening of the nanopore is on the trans side of the membrane and the second opening is on the cis side, and a motor protein controls the movement of the target polynucleotide through the nanopore from the trans side to the cis side of the membrane. In some embodiments, the method involves applying a force (e.g., an electric potential) across the nanopore, and the motor protein controls the movement of the target polynucleotide through the nanopore in the same direction as the applied force.
[0086] In another embodiment, this specification provides a method for characterizing a target polypeptide, which is (i) Contacting the first opening of a transmembrane nanopore having a first opening and a second opening with a target polynucleotide having a motor protein bound thereto, wherein the target polynucleotide is bound to the motor protein at the polynucleotide binding site of the motor protein, (ii) Obtaining one or more measurements characteristic of the target polynucleotide when the motor protein controls the movement of the target polynucleotide from the second opening of the nanopore to the first opening of the nanopore, (iii) Debinding the target polynucleotide from the polynucleotide binding site of the motor protein so that the target polynucleotide moves from the second opening of the nanopore to the first opening of the nanopore, (iv) Re-binding the target polynucleotide to the polynucleotide binding site of the motor protein, and obtaining one or more characteristic measurements of the target polynucleotide as the motor protein controls the movement of the target polynucleotide in the direction from the second opening of the nanopore to the first opening of the nanopore, This includes characterizing the target polynucleotide. Characterizing the target polynucleotide may include, for example, determining the sequence of the target polynucleotide.
[0087] For example, in some embodiments, the nanopore spans a membrane having a cis side and a trans side, with the first opening of the nanopore on the cis side of the membrane and the second opening on the trans side, and a motor protein controls the movement of the target polynucleotide through the nanopore from the trans side to the cis side of the membrane. In other embodiments, the first opening of the nanopore is on the trans side of the membrane and the second opening is on the cis side, and a motor protein controls the movement of the target polynucleotide through the nanopore from the cis side to the trans side of the membrane. In some embodiments, the method involves applying a force (e.g., an electric potential) across the nanopore, and the motor protein controls the movement of the target polynucleotide through the nanopore in the direction opposite to the applied force.
[0088] It is important to distinguish the movement of the polynucleotide in a second direction relative to the detector from any spontaneous slipping that may occur. For example, a slip of one or two bases is not an example of rereading as described herein. Typically, in step (iii), the distance the target polynucleotide travels relative to the detector is at least 10 nucleotides long. In some embodiments, the distance the target polynucleotide travels relative to the detector is at least 20 nucleotides long, e.g., at least 30 nucleotides long, e.g., at least 40 nucleotides long, e.g., at least 50 nucleotides long, e.g., at least 100 nucleotides long. Longer distances may be used. In some embodiments, the distance the target polynucleotide travels relative to the detector in step (iii) is at least 1000 nucleotides (1 kb) long, e.g., at least 2 kb long, e.g., at least 5 kb or at least 10 kb long, e.g., at least 100 kb or at least 1000 kb long.
[0089] Steps (iii) and (iv) of this method may be repeated multiple times to reread the target polynucleotide multiple times. Steps (iii) and (iv) may be repeated at least once, for example, at least twice, for example, at least three times, for example, at least four times, for example, at least five times, for example, at least ten times, for example, at least twenty times, for example, at least fifty times, for example, at least 100 times, for example, at least 1,000 times, for example, at least 10,000 times, for example, at least 100,000 times or more. Thus, this method may include "flooring" the polynucleotide back and forth against the detector.
[0090] Therefore, if steps (iii) and (iv) are repeated once (and only once), and as a result the method includes steps (iii) and (iv) twice and only twice, the method will include steps (i), (ii), (iii), (iv), (iii1), and (iv1), and characteristic measurements will be performed on three parts of the polynucleotide, namely the first part in step (ii), the second part in steps (iii) and (iv), and the third part in steps (iii1) and (iv1). If steps (iii) and (iv) are repeated twice (and only twice) so that the method includes steps (iii) and (iv) three times and only three times, the method will include steps (i), (ii), (iii), (iv), (iii1), (iv1), (iii2), and (iv2), and characteristic measurements will be taken for four parts of the polynucleotide: the first part in step (ii), the second part in steps (iii) and (iv), the third part in steps (iii1) and (iv1), and the fourth part in steps (iii2) and (iv2). In other words, if steps (iii) and (iv) are repeated n times, each repeat will yield characteristic measurements for the (n+2) part of the polynucleotide. Repeating steps (iii) and (iv) multiple times can lead to improved characterization because the part of the polynucleotide analyzed by the nanopore is sampled multiple times, and any stochastic errors that may be recorded in the analysis will lose their statistical significance. Therefore, the accuracy of the characteristic data obtained in this way can be improved. These methods make it possible to achieve very high levels of accuracy, such as at least 99%, at least 99.9%, or at least 99.99%. Therefore, in some embodiments, steps (iii) and (iv) are repeated until an accuracy level of at least 99% is achieved, such as at least 99.9% or at least 99.99%.
[0091] The portion of the polynucleotide read in step (ii) and the portion of the polynucleotide read in step (iv) of this method typically overlap. In other words, this method involves rereading at least a portion of the polynucleotide multiple times. Thus, in some embodiments, in step (ii), a motor protein controls the movement of a first portion of the target polynucleotide in a first direction relative to the detector, and in step (iv), a motor protein controls the movement of a second portion of the target polynucleotide in a first direction relative to the detector, with the first portion overlapping at least partially with the second portion. In some embodiments, the second portion overlaps with at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% of the first portion. In some embodiments, the first portion is identical to the second portion. Thus, in some embodiments, a portion of the polynucleotide is repeatedly characterized in the provided manner. If, in each repeat, the second portion of the polynucleotide partially but not completely overlaps with the first portion of the polynucleotide from the previous repeat, the polynucleotide is ratcheted to the detector in a zigzag pattern. If, in each repeat, the second portion of the polynucleotide completely overlaps with the first portion of the polynucleotide from the previous repeat, the same portion of the polynucleotide is flossed back and forth to the detector.
[0092] Forces applied during movement In some embodiments of the disclosed method, a force can be applied to the entire detector, for example, the entire nanopore. To control the method, the force can be controlled. For example, by increasing the force, the movement of polynucleotides through the detector (e.g., the nanopore) can be increased or decreased, and the rate at which polynucleotides pass through the pore can be controlled.
[0093] In the methods provided herein, any suitable force can be applied. The force may be an electrical potential applied to the entire detector, for example, the entire nanopore. In some embodiments, no external force is applied to the entire nanopore. For example, in some embodiments, no electrical potential is applied. Such embodiments are particularly suitable in some embodiments for methods in which optical measurements are performed as polynucleotides move relative to the nanopore.
[0094] In other embodiments, the force may be a voltage force applied to the entire nanopore. The voltage may be applied using any suitable apparatus, such as the apparatus described herein. Suitable potentials are described in more detail herein.
[0095] In some embodiments, a force is applied across the film in which the nanopores are embedded. The force is typically applied from the cis side to the transformer side of the film, i.e., from the cis side to the transformer side of the nanopores. The force can be a positive voltage applied to the nanopores, or a negative voltage applied to the nanopores.
[0096] Typically, the force is a positive voltage applied across the nanopore such that the trans side of the pore is positive relative to the cis side of the pore. In such embodiments, the force thus attracts negatively charged polynucleotides and moves them from the cis side to the trans side of the pore. In such embodiments, the methods provided herein typically involve using a motor protein on the cis side of the pore to control the movement of the polynucleotides against the applied force, in the direction from the trans side to the cis side of the pore, i.e., in the direction opposite to the applied force. However, in some embodiments, the methods provided herein (e.g., a method for rereading polynucleotides) may involve using a motor protein on the cis side of the pore to control the movement of the polynucleotides from the cis side to the trans side of the pore, in the same direction as the applied force.
[0097] In other embodiments, the force is a negative voltage applied across the nanopore such that the trans side of the pore is negative relative to the cis side of the pore. In such embodiments, the force attracts negatively charged polynucleotides and moves from the trans side to the cis side of the pore. In such embodiments, the methods provided herein typically involve using a motor protein on the trans side of the pore to control the movement of the polynucleotide against the applied force, in the direction from the cis side to the trans side of the pore, i.e., in the direction opposite to the applied force. However, in some embodiments, the methods provided herein (e.g., a method for rereading polynucleotides) may involve using a motor protein on the trans side of the pore to control the movement of the polynucleotide from the trans side to the cis side of the pore, in the same direction as the applied force.
[0098] However, as will be explained below, the methods provided herein do not rely on moving the polynucleotide in the opposite direction to the applied force. In some embodiments, the direction of movement may be the same as the applied force, but still in the direction away from the pore. In such embodiments, the motor protein controls the movement of the polynucleotide from the pore at a rate typically faster than the rate resulting from the applied force alone.
[0099] Therefore, in some embodiments, the force is a positive voltage applied across the nanopore such that the trans side of the pore is positive relative to the cis side of the pore, and the method may include using a motor protein on the trans side of the pore to control the movement of polynucleotides in the direction from the cis side of the pore to the trans side of the pore by the applied force. In other embodiments, the force is a negative voltage applied across the nanopore such that the trans side of the pore is negative relative to the cis side of the pore, and the method may include using a motor protein on the cis side of the pore to control the movement of polynucleotides in the direction from the trans side of the pore to the cis side of the pore by the applied force.
[0100] setting In some embodiments of the methods provided, the leader sequence is contained in or attached to the target polynucleotide. The leader sequence can be captured by a detector (e.g., a nanopore) in the methods provided herein.
[0101] Leader sequences are described in more detail herein. Typically, leader sequences are single-stranded polynucleotide regions that do not have significant secondary structures. For example, leader sequences typically do not form hairpins or G-quadrilaterals and are therefore readily captured by nanopores.
[0102] The leader sequence is typically provided at the first end of a polynucleotide or contained in an adapter attached to the first end of a polynucleotide. Adapters are described in more detail herein.
[0103] Typically, the leader sequence is provided at the first end of the polynucleotide (e.g., by being contained within the first end of the target polynucleotide, or by being contained within a polynucleotide adapter attached to the first end of the target polynucleotide), and the motor protein stalls at the second end of the target polynucleotide, or on an adapter attached to the second end of the target polynucleotide. For example, the leader sequence may be located at the 3' end of a single-stranded polynucleotide, and the motor protein may be located at the 5' end of the single-stranded polynucleotide. Alternatively, the leader sequence may be located at the 5' end of a single-stranded polynucleotide, and the motor protein may be located at the 3' end of the single-stranded polynucleotide. This configuration allows the first end of the polynucleotide to be captured by the nanopore and to pass through the nanopore, for example, from the first end to the second end. The motor protein at the second end of the polynucleotide typically prevents the polynucleotide from moving completely through the nanopore. In the methods provided herein, a motor protein at the second end of a polynucleotide can control the movement of the polynucleotide from the nanopore toward the motor protein, typically by processing the polynucleotide from the second end toward the first end.
[0104] In some embodiments, the target polynucleotide is single-stranded and includes a leader sequence, which is located at the first end of the target polynucleotide or contained in an adapter bound to the first end of the target polynucleotide, and the motor protein stalls at the second end of the target polynucleotide or stalls on the adapter at the second end of the target polynucleotide. In such embodiments, the leader sequence is typically captured by a nanopore, and the single-stranded polynucleotide moves through the nanopore until it reaches the stalled motor protein. Once destalled, the motor protein controls the movement of the polynucleotide from the pore, as illustrated in Figure 2.
[0105] In some embodiments, the target polynucleotide is double-stranded.
[0106] In some embodiments, the target polynucleotide is double-stranded and comprises a first and a second strand, the target polynucleotide includes a leader sequence, the leader sequence is located at the first end of the polynucleotide and is contained within the first strand or contained within an adapter attached to the first strand, and a motor protein stalls at the second end of the target polynucleotide. This configuration allows the first end of the first strand of the double-stranded polynucleotide to be captured by the nanopore and to pass through the nanopore from the first end to the second end. The motor protein at the second end of the polynucleotide typically prevents the polynucleotide from moving completely through the nanopore. The first strand of the double-stranded polynucleotide may be a template strand. The first strand of the double-stranded polynucleotide may be a complementary strand.
[0107] In some embodiments, the motor protein stalls at the second end of the first strand of the target polynucleotide or on an adapter at the second end of the first strand of the target polynucleotide. In some embodiments, the target polynucleotide is double-stranded and comprises a first strand and a second strand, the target polynucleotide comprises a leader sequence, the leader sequence is located at the first end of the polynucleotide and is contained in the first strand or contained in an adapter attached to the first strand, and the motor protein stalls at the second end of the first strand of the target polynucleotide or on an adapter at the second end of the first strand of the target polynucleotide. For example, the leader sequence may be located at the 3' end of the first strand of the double-stranded polynucleotide, and the motor protein may be located at the 5' end of the first strand of the double-stranded polynucleotide. Alternatively, the leader sequence may be located at the 5' end of the first strand of the double-stranded polynucleotide, and the motor protein may be located at the 3' end of the first strand of the double-stranded polynucleotide. In such embodiments, the leader sequence is typically captured by a nanopore, and the single-stranded polynucleotide moves through the nanopore until it reaches a stall motor protein. After destallation, the motor protein controls the movement of the first strand of the polynucleotide from the pore, as shown in Figure 3.
[0108] In some embodiments, the first and second strands are joined together by a hairpin adapter at the second end of the first strand, and the motor protein is stalled at the hairpin adapter. In some embodiments, the hairpin adapter binds to the 3' end of the first strand at its 5' end and attaches to the 5' end of the second strand of the target double-stranded polynucleotide at its 3' end. In some embodiments, the hairpin adapter binds to the 5' end of the first strand at its 3' end and attaches to the 3' end of the second strand of the target double-stranded polynucleotide at its 5' end. Thus, the hairpin adapter ligates the first strand to the second strand. Typically, the hairpin adapter ligates the second end of the first strand of the double-stranded polynucleotide to the first end of the second strand of the double-stranded polynucleotide.
[0109] In some embodiments, the target polynucleotide is double-stranded and comprises a first and a second strand, the target polynucleotide includes a leader sequence, the leader sequence located at the first end of the polynucleotide and contained in the first strand or contained in an adapter attached to the first strand, the first and second strands being attached together by a hairpin adapter at the second end of the first strand, and the motor protein stalling at the hairpin adapter. In such embodiments, the leader sequence is typically captured by a nanopore, and the first strand of the double-stranded polynucleotide moves through the nanopore until it reaches the stalled motor protein. After destallation, the motor protein controls the movement of the first strand of the double-stranded polynucleotide from the pore, as shown in Figure 4.
[0110] In some embodiments, the first and second strands are attached together by a hairpin adapter bonded to (i) the second end of the first strand and (ii) the first end of the second strand, and the motor protein stalls at the second end of the second strand of the double-stranded polynucleotide or at the second end of the second strand on the adapter. In some embodiments, the hairpin adapter is attached at its 5' end to the 3' end of the first strand of the target double-stranded polynucleotide and at its 3' end to the 5' end of the second strand of the target double-stranded polynucleotide, and the motor protein stalls at the 3' end of the second strand. In some embodiments, the hairpin adapter is attached at its 3' end to the 5' end of the first strand of the target double-stranded polynucleotide and at its 5' end to the 3' end of the second strand of the target double-stranded polynucleotide, and the motor protein stalls at the 5' end of the second strand. Therefore, the hairpin adapter connects the first chain to the second chain.
[0111] In some embodiments, the target polynucleotide is double-stranded and comprises a first and a second chain, the target polynucleotide comprising a leader sequence, the leader sequence located at the first end of the polynucleotide and contained in the first chain or contained in an adapter attached to the first chain, the first and second chains being attached together by a hairpin adapter attached to (i) the second end of the first chain and (ii) the first end of the second chain, the motor protein stalls at the second end of the second chain of the double-stranded polynucleotide or stalls at the second end of the second chain on the adapter. In such embodiments, the leader sequence is typically captured by a nanopore, and the first chain of the double-stranded polynucleotide, the hairpin adapter, and the second chain of the double-stranded polynucleotide move through the nanopore until they reach the stalled motor protein. After destallation, the motor protein controls the movement of the second chain, and possibly the hairpin adapter, and possibly the first chain of the double-stranded polynucleotide, from leaving the pore. This is illustrated in Figure 5.
[0112] It will be apparent that motor proteins can stall along the polynucleotide rather than at its terminus. In the context of this specification, the motor protein stalls at the terminus of the portion of the polynucleotide characterized in the method provided herein. Those skilled in the art will understand that in the method provided herein, the portion of the polynucleotide characterized is a parameter that can be determined by the arrangement of the motor protein on the polynucleotide and can be controlled by the user of the method.
[0113] In embodiments of the disclosed method, which includes rereading a target polynucleotide (for example, a method comprising: obtaining one or more characteristic measurements of the target polynucleotide when a motor protein controls the movement of the target polynucleotide in a first direction relative to a detector; debinding the target polynucleotide from the polynucleotide binding site of the motor protein so that the target polynucleotide moves in a second direction relative to the detector; rebinding the target polynucleotide to the polynucleotide binding site of the motor protein; and obtaining one or more characteristic measurements of the target polynucleotide when a motor protein controls the movement of the target polynucleotide in a first direction relative to a detector), the leader sequence can be constructed or designed to facilitate the debinding of the target polynucleotide from the polynucleotide binding site of the motor protein when the motor protein is near the leader sequence (when the motor protein is in contact with the leader sequence).
[0114] In such embodiments, the motor protein typically has a lower affinity for the leader sequence than for the target polynucleotide, i.e., a lower affinity for the characterized portion of the target polynucleotide. In some embodiments, the leader has a different structure from the target polynucleotide. In some embodiments, the leader contains a different type of nucleotide from the target polynucleotide.
[0115] For example, in some embodiments, the target polynucleotide includes deoxyribonucleotide (DNA). In such embodiments, the leader may include one or more nucleotides lacking both nucleic acid bases and sugar moieties (e.g., spacer moieties). Preferred spacer moieties are described in detail herein and include C2 spacers, C3 spacers, C6 spacers, iSp9 spacers, iSp18 spacers, and the like. Alternatively or additionally, the leader may include ribonucleotides (RNA), peptide nucleotides (PNA), glycerol nucleotides (GNA), threose nucleotides (TNA), locked nucleotides (LNA), cross-linked nucleotides (BNA), or baseless nucleotides. In some embodiments, the leader may include one or more nucleotides having modified phosphate bonds (e.g., methylphosphonate or phosphothiolate bonds).
[0116] In some other embodiments, the target polynucleotide includes ribonucleotide (RNA). In such embodiments, the leader may include one or more spacers as defined above, deoxyribonucleotide (DNA), peptide nucleotide (PNA), glycerol nucleotide (GNA), threose nucleotide (TNA), locked nucleotide (LNA), cross-linked nucleotide (BNA), debasalized nucleotide, or nucleotides containing a modified phosphate bond.
[0117] Typically, the target polynucleotide comprises a deoxyribonucleotide (DNA), and the leader comprises one or more spacer segments (e.g., C3 spacers) and / or one or more ribonucleotides.
[0118] A leader may contain only one type of polynucleotide different from the target polynucleotide. For example, if the target polynucleotide is DNA, the leader may contain a spacer portion or RNA. A leader may contain multiple types of polynucleotides different from the target polynucleotide. For example, if the target polynucleotide is DNA, the leader may contain a spacer portion and RNA. A leader may contain a portion that is the same type of polynucleotide as the target polynucleotide. For example, if the target polynucleotide is DNA, the leader may contain a portion of DNA in addition to a spacer polynucleotide or RNA. Such portions may be called “traps,” i.e., a leader based on a spacer (e.g., a C3 spacer) and / or RNA (e.g., 2'-methoxyuridine) polynucleotide may contain one or more DNA traps. A trap typically contains 1 to 10 nucleotides, such as 1 to 6 nucleotides, e.g., 1, 2, 3, 4, or 5 nucleotides, e.g., 1 to 3 nucleotides. Therefore, if the target polynucleotide is DNA, the reader may include one or more RNA (e.g., 2'-methoxyuridine) and / or spacer (e.g., C3 spacer) portions and one or more DNA (e.g., thymidine) traps of 1 to 10 nucleotides in length.
[0119] Those skilled in the art will understand that, if the leader comprises a polynucleotide chain, the leader sequence is typically not deterministic and can be controlled or selected according to other experimental conditions, such as the motor protein and any polynucleotide to be characterized. Exemplary sequences are provided for illustrative purposes only in the examples, particularly in Example 10. For example, the leader may comprise a sequence such as one or more of SEQ ID NOs: 70, 71, or 72, or a polynucleotide sequence having at least 20%, e.g., at least 30%, e.g., at least 40%, e.g., at least 50%, e.g., at least 60%, e.g., at least 70%, e.g., at least 80%, e.g., at least 90%, e.g., at least 95% sequence similarity or identity with one or more of SEQ ID NOs: 70, 71, or 72. The leader sequence can typically be modified without adversely affecting the effectiveness of the method provided herein.
[0120] Stalling of motor proteins As described above, the method provided herein involves characterizing a target polynucleotide having a motor protein stalled at a stalled region.
[0121] Any suitable stall portion may be used in the manner provided herein. In some embodiments, the stall portion includes the stall section described herein. In some embodiments, the stall portion includes one or more stall units.
[0122] Any suitable stall unit can be used. Stall units typically provide an energy barrier that hinders the movement of motor proteins. For example, a stall unit can stall a motor protein by reducing its traction on the polynucleotide. This can be achieved, for example, by using a debase spacer, i.e., a spacer from which a base has been removed from one or more nucleotides in the polynucleotide adapter. Spacers can physically block the movement of polynucleotide handling proteins, for example, by introducing a large chemical group that physically hinders the movement of the polynucleotide handling protein.
[0123] In some embodiments, the stall unit may include a linear molecule such as a polymer. Typically, such stall units have a structure different from the target polynucleotide. For example, when the target polynucleotide is DNA, the stall unit or each stall unit typically does not contain DNA. Specifically, when the target polynucleotide is deoxyribonucleic acid (DNA) or ribonucleic acid (RNA), the stall unit or each stall unit preferably includes peptide nucleic acid (PNA), glycerol nucleic acid (GNA), threose nucleic acid (TNA), locked nucleic acid (LNA), cross-linked nucleic acid (BNA), or a synthetic polymer having nucleotide side chains. In some embodiments, the stall unit is one or more nitroindoles, one or more inosines, one or more acridines, one or more 2-aminopurines, one or more 2-6-diaminopurines, one or more 5-bromodeoxyuridines, one or more inverted thymidines (inverted dTs), one or more inverted dideoxythymidines (ddTs), one or more dideoxycytidines (ddCs), one or more 5-methylcytidines, one or more 5-hydroxymethylcytidines, one or more 2'-O-methylRNA bases, one or more iso -The compound may contain deoxycytidine (Iso-dC), one or more iso-deoxyguanosine (Iso-dG), one or more C3 (OC3H6OPO3) groups, one or more photocleavable (PC)[OC3H6-C(O)NHCH2-C6H3NO2-CH(CH3)OPO3] groups, one or more hexanediol groups, one or more spacer 9 (iSp9)[(OCH2CH2)3OPO3] groups, or one or more spacer 18 (iSp18)[(OCH2CH2)6OPO3] groups, or one or more thiol bonds. The stall site may contain any combination of these groups. Many of these groups are commercially available from IDT® (Integrated DNA Technologies®). For example, C3, iSp9, and iSp18 spacers are all available from IDT®. The stall site may contain any number of the above groups as stall units. For example, a stalled section may contain 1 to approximately 12 or more such stall units (e.g., approximately 1 to approximately 8 units, or 1 to approximately 4 units, etc., or 1 to approximately 6 units).
[0124] In some embodiments, the stall unit may comprise one or more chemical groups that stall the motor protein. In some embodiments, preferred chemical groups are one or more pendant chemical groups. One or more chemical groups may be bound to one or more nucleic acid bases in the polynucleotide. One or more chemical groups may be attached to the polynucleotide backbone. Any number of suitable chemical groups may exist, such as 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12 or more. Preferred groups include, but are not limited to, fluorophores, streptavidin and / or biotin, cholesterol, methylene blue, dinitrophenol (DNP), digoxigenin and / or anti-digoxigenin, and dibenzylcyclooctin groups.
[0125] In some embodiments, the stall unit may include a polymer. In some embodiments, the stall unit may include a polymer that is a polypeptide or polyethylene glycol (PEG).
[0126] In some embodiments, the stall unit may consist of one or more debasalized nucleotides (i.e., nucleotides lacking a nucleic acid base), for example, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12 or more debasalized nucleotides. The nucleic acid base may be replaced by -H(idSp) or -OH in the debasalized nucleotide. The debasalized residue may be inserted into the target polynucleotide by removing a nucleic acid base from one or more adjacent nucleotides. For example, the polynucleotide may be modified to contain 3-methyladenine, 7-methylguanine, 1,N6-ethenoadenine inosine, or hypoxanthine, and the nucleic acid base may be removed from these nucleotides using human alkyladenine DNA glycosylase (hAAG). Alternatively, the polynucleotide may be modified to contain uracil, and the nucleic acid base may be removed by uracil DNA glycosylase (UDG). In one embodiment, the stall unit does not contain any debasalized nucleotides.
[0127] A suitable stall unit can be designed or selected depending on the properties of the polynucleotide / polynucleotide adapter, the motor protein, and the conditions under which the method is to be performed. For example, many polynucleotide processing proteins process DNA in vivo, and such proteins can typically stall using something other than DNA.
[0128] Therefore, in some embodiments of the provided method, the motor protein stalls at a stall site containing one or more stall units independently selected from the following: -Polynucleotide secondary structure, preferably hairpin or G-quadrivalent (TBA), - Preferably, a nucleic acid analog selected from peptide nucleic acid (PNA), glycerol nucleic acid (GNA), threose nucleic acid (TNA), locked nucleic acid (LNA), cross-linked nucleic acid (BNA), and debasic nucleotide. -Spacer units selected from nitroindole, inosine, acridine, 2-aminopurine, 2-6-diaminopurine, 5-bromodoxyuridine, inverted thymidine (inverted dTs), inverted dideoxythymidine (ddTs), dideoxycytidine (ddCs), 5-methylcytidine, 5-hydroxymethylcytidine, 2'-O-methylRNA base, isodeoxycytidine (Iso-dCs), isodeoxyguanosine (Iso-dGs), C3(OC3H6OPO3) group, photocleavable (PC)[OC3H6-C(O)NHCH2-C6H3NO2-CH(CH3)OPO3] group, hexanediol group, spacer 9 (iSp9)[(OCH2CH2)3OPO3] group, spacer 18 (iSp18)[(OCH2CH2)6OPO3] group, and thiol linkage, and -Avidins such as fluorophores, traptabidine, streptavidin and neutraavidin, and / or biotin, cholesterol, methylene blue, dinitrophenol (DNP), digoxigenin and / or anti-digoxigenin and dibenzylcyclooctin groups.
[0129] The stall portions described herein can also be used to construct a leader suitable for use in the disclosed rereading method. As described above, in some embodiments of such a method, the leader sequence described herein is constructed or designed to facilitate the dissociation of a target polynucleotide from the polynucleotide binding site of a motor protein when the motor protein is near the leader (e.g., when the motor protein is in contact with the leader sequence). In some embodiments, the leader sequence may include any of the spacer portions described above.
[0130] Motor protein destallation In some embodiments, the methods provided herein include bringing the stalled portion into contact with a detector (e.g., a nanopore) to destall the motor protein. After destallation, the motor protein can control the movement of polynucleotides exiting the detector (e.g., the nanopore), as described in more detail herein.
[0131] In its simplest form, the motor protein can be destalled from the stalled portion by bringing it into contact with a detector, such as a nanopore. However, in some embodiments, this method involves actively destalling the motor protein as described herein.
[0132] In some embodiments, stalling a motor protein involves applying a stalling force to a polynucleotide, the stalling force being smaller than and / or in the opposite direction to a read force, the read force being the force applied while the motor protein controls the movement of the target polynucleotide and measurements are being taken to determine one or more features of the polynucleotide.
[0133] For example, the reading force can typically be provided as a potential of +2V to -2V, usually -400mV to +400mV. The voltage used is preferably within a range having a lower limit selected from -400mV, -300mV, -200mV, -150mV, -100mV, -50mV, -20mV, and 0mV, and an upper limit independently selected from +10mV, +20mV, +50mV, +100mV, +150mV, +200mV, +300mV, and +400mV. The voltage used is more preferably in the range of 100mV to 240mV, most preferably in the range of 120mV to 220mV. Typically, the destabilization force is smaller than the reading force. For example, the stall force can range from approximately -100mV to +100mV, for instance, from approximately -50mV to approximately +50mV, or for instance, from approximately -25mV to approximately +25mV.
[0134] For example, in some embodiments, the reading force is a potential in the range of +100mV to +200mV, such as +50mV to +300mV, more preferably +120mV to +150mV, and the destabilization force is a potential in the range of -20mV to +20mV, such as -40mV to +40mV, such as -50mV to +50mV, or 0mV.
[0135] In some embodiments, the destall force is in the opposite direction to the reading force. For example, in some embodiments, the reading force is applied as a positive potential and the destall force is applied as a negative potential. In other embodiments, the reading force is applied as a negative potential and the destall force is applied as a positive potential. When the destall force is in the opposite direction to the reading force, it may be the same magnitude as the reading force or it may be smaller in magnitude than the reading force.
[0136] In some embodiments, the destall force is applied at zero potential. For example, in some embodiments, the reading force is applied as a positive potential and the destall force is applied at zero application potential. In other embodiments, the reading force is applied as a negative potential and the destall force is applied at zero application potential.
[0137] In some embodiments, the destalling force is applied for a sufficient amount of time for the motor protein to destall from the stalled portion. In some embodiments, the destalling force is applied for a period of about 1 ms to about 10 s, such as about 10 ms to about 1 s, or about 100 ms to about 700 ms, such as about 300 ms to about 500 ms.
[0138] In some embodiments, stalling a motor protein involves changing the force applied between the stalling force and the read force one or more times. In some embodiments, changing the applied force in this manner involves stepping or tilting the applied potential between the stalling force and the read force. When tilting, any suitable waveform can be used, for example, the tilt may be a linear, exponential, or sigmoid tilt.
[0139] In some embodiments, the applied force varies between a single destall force and a reading force. In some embodiments, the applied force varies between a series of different destall forces and reading forces. In some embodiments, the applied force varies stepwise between a series of increasing destall forces and reading forces. The destall force in each step may be any suitable destall force, e.g., any of the destall forces described herein, and each step may be applied over any suitable period, e.g., any period described herein.
[0140] In some embodiments, the destall force is the same as the reading force. This is also referred to as destall in a "free-running" setting.
[0141] In some embodiments, the motor protein stalls at a stall site comprising one or more stall units and one or more stopping portions, and when one or more stopping portions come into contact with a nanopore, the movement of polynucleotides through the nanopore is delayed, thereby causing the motor protein to destall from one or more stall units. Such embodiments are suitable for use in self-propelled settings.
[0142] In some embodiments, the stopping portion provides an energy barrier that prevents the movement of polynucleotides through the nanopore. For example, the stopping portion may prevent the movement of polynucleotides through the nanopore by providing a physical block that must be removed before the polynucleotides can pass through the nanopore.
[0143] While not bound by theory, the inventors believe that the stopping portion delays the movement of polynucleotides through the nanopore for a sufficient amount of time for the motor protein to overcome the stall unit and stall.
[0144] In some embodiments, the stopping portion comprises one or more stopping units, each containing a polynucleotide secondary structure, preferably a hairpin or a G quadruple chain (TBA). Such secondary structures prevent the polynucleotide from freely passing through the nanopore. When the stopping portion comes into contact with the nanopore, the secondary structure dissociates (e.g., unravels). The time it takes for the secondary structure to dissociate allows the motor protein to destall from the stall unit.
[0145] In some embodiments, the stop portion includes one or more stop units containing hybridized oligonucleotides. The oligonucleotides may hybridize to a target polynucleotide, preventing the target polynucleotide from moving through the nanopore. When the stop portion is brought into contact with the nanopore, the hybridized oligonucleotides dissociate from the target polynucleotide. The time it takes for the hybridized oligonucleotides to dissociate from the target polynucleotide allows the motor protein to destall from the stall unit.
[0146] In some embodiments, the stop portion preferably comprises one or more stop units including nucleic acid analogs selected from peptide nucleic acids (PNA), glycerol nucleic acids (GNA), threose nucleic acids (TNA), locked nucleic acids (LNA), crosslinked nucleic acids (BNA), and debasalized nucleotides. The nucleic acid analogs may be supplied with the target polynucleotide, or they may hybridize to the target polynucleotide or otherwise bind to it. When the nucleic acid analogs are supplied with the target polynucleotide, contact with the nanopore allows the nucleic acid analogs to pass through the nanopore. The time it takes for the nucleic acid analogs to pass through the pore allows the motor protein to destall from the stall unit. When the nucleic acid analogs hybridize to the target polynucleotide, contact with the nanopore typically allows the nucleic acid analogs to dissociate from the target polynucleotide, enabling the target polynucleotide to pass through the nanopore. The time it takes for the nucleic acid analogs to dissociate from the polynucleotide allows the motor protein to destall from the stall unit.
[0147] In some embodiments, the stop portion comprises one or more stop units containing chemical groups such as fluorophores, avidins such as traptaavidin, streptavidin, and neutraavidin, and / or biotin, cholesterol, methylene blue, dinitrophenol (DNP), digoxigenin, and / or anti-digoxigenin and dibenzylcyclooctin groups. The chemical groups may attach to the target polynucleotide and prevent the target polynucleotide from passing through the nanopore. In some embodiments, contact of the stop portion with the nanopore removes the chemical groups from the target polynucleotide. In some embodiments, contact of the stop portion with the nanopore allows the chemical groups to pass through the nanopore. The time it takes for the chemical groups to be removed from the target polynucleotide and / or to pass through the nanopore allows the motor protein to destall from the stall unit.
[0148] In some embodiments, the stopping portion comprises one or more stopping units, each containing a polynucleotide-binding protein. Suitable polynucleotide-binding proteins are described in more detail herein. The polynucleotide-binding protein can bind to a polynucleotide and prevent its movement through the nanopore. When the stopping portion is brought into contact with the nanopore, the movement of the polynucleotide through the nanopore is delayed, for example, when the polynucleotide-binding protein moves and comes into contact with the motor protein. The time required for execution allows the motor protein to destall from the stall unit.
[0149] While not bound by theory, the inventors also believe that the stalling region often determines the three-dimensional structure of the polynucleotide at the stalling unit. This is especially true when the stalling region contains one or more linear groups such as spacer 18(iSp18)[(OCH2CH2)6OPO3]. While not bound by theory, if such a stalling unit is in contact with a nanopore, it is thought that a force applied across the pore (e.g., an applied voltage field) can extend the stalling region almost linearly. In this three-dimensional structure, the motor protein is usually unable to pass through the stalling region and destall. However, if the target polynucleotide is stalled at the stalling region, the environment of the stalling unit is thought to be similar to that in solution, and the stalling unit may adopt a more compact pseudo-random coil configuration. In this configuration, it may be easier for the motor protein to overcome the stalling unit and destall.
[0150] Therefore, in some embodiments, the motor protein stalls at a stall site comprising one or more stall units and one or more stop units, independently selected from the following: -Polynucleotide secondary structure, preferably hairpin or G-quadrivalent (TBA), - Preferably, a nucleic acid analog selected from peptide nucleic acid (PNA), glycerol nucleic acid (GNA), threose nucleic acid (TNA), locked nucleic acid (LNA), cross-linked nucleic acid (BNA), and debasic nucleotide. -Avidins such as fluorophores, traptabidine, streptavidin and neutraavidin, and / or biotin, cholesterol, methylene blue, dinitrophenol (DNP), digoxigenin and / or anti-digoxigenin and dibenzylcyclooctin groups, and -Polynucleotide-binding proteins, Furthermore, when one or more stalled regions come into contact with the nanopore, the movement of polynucleotides through the nanopore is delayed, thereby causing the motor protein to destall from one or more stalled units.
[0151] Motor protein As those skilled in the art will understand, any suitable motor protein can be used in the methods and products provided herein.
[0152] A motor protein can be any protein that can bind to a polynucleotide and control its movement toward a detector, such as a nanopore, for example, through a pore.
[0153] More specifically, motor proteins such as helicases typically have at least two modes of operation (all the components necessary to facilitate movement, e.g., ATP and Mg) 2+ DNA movement can be controlled in two modes: (when necessary components are provided) and in one inactive mode (when the motor protein is modified to prevent the active mode, or when the components necessary to facilitate movement are not provided).
[0154] Once all the necessary components to facilitate movement are provided, the motor protein can move along a polynucleotide, such as DNA, in either the 5'-3' or 3'-5' direction. Many motor proteins process polynucleotides, such as DNA, in the 5'-3' direction. Motor proteins that thus control the movement of polynucleotides are typically well-suited for use in the methods provided herein.
[0155] However, if the motor protein lacks the components necessary to facilitate movement, or is modified to prevent active control of the movement of the polynucleotide relative to the nanopore, it can still passively control the movement of the polynucleotide relative to the nanopore. For example, a motor protein can act as a brake, slowing down the movement of the polynucleotide once it has bound to the polynucleotide and is drawn into the pore by the site to which the polynucleotide is applied (e.g., by the first force in the method provided herein). In “inactive” mode, it usually doesn’t matter whether the DNA is captured at 3’ or 5’ (i.e., whether it moves through the nanopore in the 5’-3’ direction or the 3’-5’ direction), since the applied force provides the propulsion force that moves the polynucleotide through the nanopore. However, in such embodiments, the motor protein can still control the movement of the polynucleotide relative to the nanopore, for example, by acting as a brake. In inactive mode, the control of polynucleotide movement by the motor protein can be described in several ways, including ratcheting, sliding, and braking. Typically, the methods provided herein do not involve the use of motor proteins operating in passive mode. However, in embodiments of the methods provided herein that use polynucleotide-binding proteins, the polynucleotide-binding protein may be a motor protein that operates in a passive mode.
[0156] As described above, some embodiments of the methods provided herein also include the use of a polynucleotide-binding protein as a stopping portion that prevents the movement of polynucleotide chains through a nanopore. In some embodiments, the polynucleotide-binding protein may be a motor protein as described herein. In other embodiments, the polynucleotide-binding protein may be a protein that binds to polynucleotides but does not have the ability to process polynucleotides; that is, in some embodiments, it is not a motor protein.
[0157] A polynucleotide handling enzyme is a polypeptide capable of interacting with polynucleotides. The enzyme may modify a polynucleotide by cleaving it to form individual nucleotides or shorter chains of nucleotides, such as di- or trinucleotides. The enzyme may also modify a polynucleotide by orienting it or moving it to a specific position. The motor proteins used herein may be polynucleotide handling enzymes or derived therefrom. The polynucleotide-binding proteins may be polynucleotide handling enzymes or derived therefrom.
[0158] In one embodiment, the motor protein is more preferably derived from any member of enzyme classification (EC) groups 3.1.11, 3.1.13, 3.1.14, 3.1.15, 3.1.16, 3.1.21, 3.1.22, 3.1.25, 3.1.26, 3.1.27, 3.1.30, and 3.1.31.
[0159] Typically, motor proteins are helicases, polymerases, exonucleases, topoisomerases, or variants thereof.
[0160] In some embodiments, motor proteins and / or polynucleotide-binding proteins can be modified to prevent the motor protein from unassociating with the polynucleotide. This is particularly useful in methods disclosed herein that involve rereading a target polynucleotide. Thus, in some embodiments of such methods, the target polynucleotide does not unassociate with the motor protein.
[0161] As used herein, the term “disassociation” refers to the dissociation of a motor protein from a target polynucleotide. Therefore, a motor protein can be modified to prevent it from dissociating from the target polynucleotide, for example, to a reaction medium. It is important to distinguish the potential “detachment” of a motor protein from the “debinding” of the motor protein from the target polynucleotide. As used herein, “debinding” refers to the transient release of the motor protein’s active site (described in more detail herein) from the target polynucleotide, but does not imply detachment. Therefore, for example, a motor protein can be modified to prevent it from disassociating from the polynucleotide, but not to prevent it from debinding from the polynucleotide. If not bound, the motor protein remains bound to the target polynucleotide. For example, a motor protein can maintain its binding to a target polynucleotide (i.e., prevent disassociation from the target polynucleotide) because it is topologically closed around the target polynucleotide. The polynucleotide binding site may remain freely bound to or detached from the target polynucleotide while the motor protein remains bound to the target polynucleotide, allowing the motor protein to bind to or detach from the target polynucleotide. When the motor protein detaches from the target polynucleotide, it can move along (e.g., along) the target polynucleotide under the applied force and re-bind to the target polynucleotide. When associated with but detached from the target polynucleotide, the motor protein cannot dissociate from the target polynucleotide.
[0162] Motor proteins and / or polynucleotide-binding proteins can be adapted to prevent detachment in any preferred manner. For example, a motor protein and / or polynucleotide-binding protein may be modified to prevent it from being loaded onto a polynucleotide and then from detaching from the polynucleotide. Alternatively, a motor protein and / or polynucleotide-binding protein may be modified to prevent it from detaching from the polynucleotide before it is loaded onto the polynucleotide. Modification of a motor protein to prevent it from detaching from a polynucleotide can be achieved using methods known in the art, e.g., the methods discussed in WO2014 / 013260, which is incorporated entirely herein by reference, and with particular reference to the section describing the modification of a motor protein, such as a helicase, to prevent it from detaching from a polynucleotide chain. For example, a motor protein and / or polynucleotide-binding protein can be modified by treatment with tetramethylazodicarboxamide (TMAD). Various other closed parts are described in more detail herein.
[0163] For example, motor proteins and / or polynucleotide-binding proteins may have polynucleotide debinding openings, e.g., cavities, grooves, or voids, through which a chain can pass when the motor protein and / or polynucleotide-binding protein deassociates from a chain. In some embodiments, a polynucleotide debinding opening is an opening through which a nucleotide can pass when the motor protein and / or polynucleotide-binding protein deassociates from a nucleotide. In some embodiments, the polynucleotide debinding opening of a given motor protein and / or polynucleotide-binding protein can be determined by reference to its structure, e.g., by reference to its X-ray crystal structure. The X-ray crystal structure can be obtained in the presence and / or absence of a polynucleotide substrate. In some embodiments, the location of the polynucleotide debinding opening in a given motor protein and / or polynucleotide-binding protein can be estimated or confirmed by molecular modeling using a standard package known in the art. In some embodiments, the polynucleotide debinding opening may be transiently generated by the movement of one or more parts of the motor protein, e.g., one or more domains.
[0164] Motor proteins and / or polynucleotide-binding proteins can be modified by closing polynucleotide debinding openings. These openings can be closed by a closure. Therefore, closing the polynucleotide debinding openings can prevent motor proteins and / or polynucleotide-binding proteins from deassociating with polynucleotides. For example, motor proteins and / or polynucleotide-binding proteins can be modified by covalently closing polynucleotide debinding openings. However, as described above, closing the polynucleotide debinding openings does not necessarily prevent the target polynucleotide from debinding from the polynucleotide-binding site of the motor protein. In some embodiments, a preferred protein for this purpose is a helicase.
[0165] In some embodiments, particularly in embodiments of the disclosed method involving rereading a target polynucleotide, the motor protein may be modified to prevent the deassociation of the target polynucleotide from the target polynucleotide. The motor protein may be modified in any preferred manner.
[0166] While not bound by theory, the inventors believe that promoting detachment and delaying re-binding may facilitate re-reading. Again, while not bound by theory, the inventors believe this is because each step the motor protein takes with respect to a target polynucleotide is related to the probability that the motor protein will detach from the polynucleotide. Such a probability of detachment can be identified by a so-called off-velocity. An increase in the off-velocity is thought to promote the dropback of the motor protein relative to the polynucleotide chain. Similarly, again, while not bound by theory, the inventors believe that once detached from the target polynucleotide, the distance the motor protein can travel along the target polynucleotide before re-binding is related to the on-velocity. Therefore, re-reading may be facilitated by increasing the off-velocity and decreasing the on-velocity of the motor protein with respect to the target polynucleotide. Adjusting the off-velocity and on-velocity of the motor protein for a given type of polynucleotide is within the capabilities of those skilled in the art, given the disclosures herein. Therefore, motor proteins can be modified to promote the debinding of a target polynucleotide from the polynucleotide binding site of the motor protein, and / or to delay the rebinding of the target polynucleotide to the polynucleotide binding site of the motor protein. In some embodiments, motor proteins are modified to promote both the debinding of a target polynucleotide from the polynucleotide binding site of the motor protein and the rebinding of the target polynucleotide to the polynucleotide binding site of the motor protein.
[0167] In some embodiments, the motor protein may be modified with a closure moiety to (i) topologically close the polynucleotide binding site of the motor protein around the target polynucleotide, and (ii) promote the debinding of the target polynucleotide from the polynucleotide binding site of the motor protein, and / or delay the rebinding of the target polynucleotide to the polynucleotide binding site of the motor protein. The motor protein may be modified in any preferred manner to promote the attachment of such closure moieties.
[0168] In some embodiments, the closure portion may include a bifunctional crosslinking portion. The closure portion may include a bifunctional crosslinking agent. The bifunctional crosslinking agent attaches to the motor protein at two points, closing the polynucleotide debinding opening of the motor protein, thereby preventing the detachment of polynucleotides from the motor protein, while allowing the debinding of polynucleotides from the polynucleotide binding site of the motor protein.
[0169] The cloning portion can attach to any suitable position on the motor protein. For example, the cloning portion can crosslink two amino acid residues of the motor protein. Typically, at least one amino acid crosslinked by the cloning portion is cysteine or a non-natural amino acid. Cysteine or a non-natural amino acid can be introduced into the motor protein by substitution or modification of naturally occurring amino acid residues of the motor protein. Methods for introducing non-natural amino acids are well known in the art and include, for example, natural chemical ligation with a synthetic polypeptide chain containing such a non-natural amino acid. A method for introducing cysteine into a motor protein is also described, for example, in Sambrook et al., Molecular Cloning: A Laboratory Manual, 4. thThe techniques described in references such as, ed., Cold Spring Harbor Press, Plainsview, New York (2012), and Ausubel et al., Current Protocols in Molecular Biology (Supplement 114), John Wiley & Sons, New York (2016), are also within the scope of the skill of a person skilled in the art.
[0170] In some embodiments, the closure portion has a length of about 1 Å to about 100 Å. The length of the closure portion can be calculated according to the static bond length, or more preferably using molecular dynamics simulations. The length may be, for example, about 2 Å to about 80 Å, for example about 5 Å to about 50 Å, for example about 8 to about 30 Å, for example about 10 to about 25 Å or about 20 Å, for example about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, or 19 Å.
[0171] While not bound by theory, the inventors generally believe that longer closure regions can increase the off-rate of motor proteins from polynucleotides and thus facilitate re-reading.
[0172] In some embodiments, the closure portion includes a bond. In some embodiments, the closure portion includes a disulfide bond. The disulfide bond can be formed by treating the motor protein with any suitable reagent such as TMAD.
[0173] In some embodiments, the closure portion includes a reagent that forms a bond between two click chemistry groups on a motor protein. Examples of click chemistry reagents are provided herein.
[0174] In some embodiments, the closure portion contains a protein. For example, a biotin group may be present on the motor protein, and the closure portion may contain streptavidin. Tags such as a snoop tag or spy tag may be present on the motor protein, and the closure portion may contain proteins such as a snoop catcher or spy catcher, respectively.
[0175] In some embodiments, the closed portion comprises the structure of formula [ABC], where A and C are independent reactive functional groups for reacting with amino acid residues in the motor protein, and B is the linking portion. In some embodiments, the closed portion comprises a bond between thio groups, such as thiol groups on a cysteine residue. Thus, in some embodiments, A and C are cysteine reactive functional groups. In some embodiments, linking portion B includes a linear or branched unsubstituted or substituted alkylene, alkenylene, alkylylene, arylene, heteroarylene, carbocyclylene, or heterocyclene portion, which is optionally interrupted or terminated by one or more atoms or groups selected from O, N(R), S, C(O), C(O)NR, C(O)O, unsubstituted or substituted arylene, arylene-alkylene, heteroarylene, heteroarylene-alkylene, carbocyclylene, carbocyclylene-alkylene, heterocyclylene, and heterocyclylene-alkylene, where R is selected from H, unsubstituted or substituted alkyl, and unsubstituted or substituted aryl. Typically, R is H or methyl, and more typically H.
[0176] Typically, the alkylene group is C 1-20 It is an alkylene group. Typically, an alkenylene group is C 2-20 It is an alkenylene group. Typically, an alkenylene group is C 2-20 It is an alkynylene group. Typically, the allerene group is C 6-12 It is an arylene group. Typically, a heteroarylene group is a heteroarylene group with 5 to 12 members. Typically, a carbocyclylene group is C 5-12It is a carbocyclylene group. Typically, a heterocyclylene group is a heterocyclylene group with 5 to 12 members.
[0177] Typically, alkylene, alkenylene, or alkynylene moieties may be interrupted or terminated by atoms or groups selected from O, N(R), S, C(O), C(O)NR, and C(O)O and unsubstituted or substituted arylenes. Usually, alkylene, alkenylene, or alkynylene moieties are not interrupted, are interrupted, or terminated by one or more atoms or groups selected from O and N(R) and unsubstituted or substituted arylenes. More often, alkylene, alkenylene, or alkynylene moieties are not interrupted, are interrupted, or terminated by one or more oxygen atoms.
[0178] For example, the bonding portion is often unsubstituted or substituted C. 1-10 Alkylene, C 2-10 Alkenylene, or C 2-10 This is an alkynylene moiety, which is either not interrupted by one or more oxygen atoms, is interrupted by one or more oxygen atoms, or terminates by an oxygen atom.
[0179] In some embodiments, linking portion B comprises an alkylene, oxyalkylene, or polyoxyalkylene group, and / or A and C are each maleimide groups. The alkylene, oxyalkylene, or polyoxyalkylene group may have a length of about 8 to about 30 Å, for example, about 5 Å to about 50 Å, or for example, about 10 to about 25 Å.
[0180] For example, the connecting part is (CH2CH2O) xThe formula may include PEG portions such as, where x is 1 to 10, for example 1 to 5, for example 1, 2, or 3. Exemplary linked portions are described in Example 9 and include, for example, BMOE (1,2-bismaleimide ethane), BMOP (1,3-bismaleimide propane), BMB (1,4-bismaleimide butane), BM(PEG)2 (1,8-bismaleimide diethylene glycol), and BM(PEG)3 (1,11-bismaleimide triethylene glycol).
[0181] Motor proteins suitable for closure using the closure portion described above are discussed in more detail herein. In some preferred embodiments, the motor protein is a helicase, such as the Dda helicase described herein.
[0182] In one embodiment, the motor protein and / or polynucleotide-binding protein is or derived from an exonuclease. Preferred enzymes include, but are not limited to, exonuclease I (SEQ ID NO: 1) from E. coli, exonuclease III (SEQ ID NO: 2) from E. coli, RecJ (SEQ ID NO: 3) and bacteriophage lambda exonuclease (SEQ ID NO: 4) from T. thermophilus, TatD exonuclease, and their variants. Three subunits, including the sequence shown in SEQ ID NO: 3 or its variants, interact to form a trimer exonuclease.
[0183] In one embodiment, the motor protein and / or polynucleotide-binding protein is derived from a polymerase. The polymerase may be PyroPhage® 3173 DNA polymerase (commercially available from Lucigen® Corporation), SD polymerase (commercially available from Bioron®), Klenow from NEB, or a variant thereof. In one embodiment, the enzyme is Phi29 DNA polymerase (SEQ ID NO: 5) or a variant thereof. Modified versions of Phi29 polymerase that may be used in the present invention are disclosed in U.S. Patent No. 5,576,204.
[0184] In one embodiment, the motor protein and / or polynucleotide-binding protein is derived from a topoisomerase. In one embodiment, the topoisomerase is preferably a member of subgroup (EC) groups 5.99.1.2 and 5.99.1.3. The topoisomerase may be a reverse transcriptase, which is an enzyme capable of catalyzing the formation of cDNA from an RNA template. These are commercially available, for example, from New England Biolabs® and Invitrogen®.
[0185] In one embodiment, the motor protein and / or polynucleotide-binding protein is derived from a helicase. Any suitable helicase can be used according to the methods provided herein. For example, the enzymes used according to this disclosure may be independently selected from Hel308 helicase, RecD helicase, TraI helicase, TrwC helicase, XPD helicase, and Dda helicase, or variants thereof. Monomer helicases may comprise several domains attached together. For example, TraI helicase and TraI subgroup helicases may comprise two RecD helicase domains, a relaxase domain, and a C-terminal domain. These domains typically form a monomer helicase that can function without forming an oligomer. Specific examples of suitable helicases include Hel308, NS3, Dda, UvrD, Rep, PcrA, Pif1, and TraI. These helicases typically act on single-stranded DNA. Examples of helicases that can move along both strands of double-stranded DNA include the FtfK and hexamer enzyme complex, or multi-subunit complexes such as RecBCD. In one embodiment, the motor protein is a Dda (DNA-dependent ATPase) helicase.
[0186] The Hel308 helicase is described in its entirety in publications such as WO2013 / 057495, which is incorporated by reference. The RecD helicase is described in its entirety in publications such as WO2013 / 098562, which is incorporated by reference. The XPD helicase is described in its entirety in publications such as WO2013 / 098561, which is incorporated by reference. The Dda helicases are described in their respective entireties in publications such as WO2015 / 055981 and WO2016 / 055777, which are incorporated by reference.
[0187] In one embodiment, the helicase may include the sequence shown in SEQ ID NO: 6 (TrwcCba) or a variant thereof, the sequence shown in SEQ ID NO: 7 (Hel308Mbu) or a variant thereof, or the sequence shown in SEQ ID NO: 8 (Dda) or a variant thereof. The variant may differ from the natural sequence in any of the ways discussed herein. An exemplary variant of SEQ ID NO: 8 includes E94C / A360C. Further exemplary variants of SEQ ID NO: 8 include E94C / A360C followed by (ΔM1)G1G2 (i.e., deletion of M1 followed by addition of G1 and G2).
[0188] Typically, motor proteins or polynucleotide-binding proteins may have fuel-binding sites. Active unwinding of DNA can be coupled, for example, to the promotion of hydrolysis in motor proteins.
[0189] The fuel is typically free nucleotides or free nucleotide analogs. Free nucleotides include adenosine monophosphate (AMP), adenosine diphosphate (ADP), adenosine triphosphate (ATP), guanosine monophosphate (GMP), guanosine diphosphate (GDP), guanosine triphosphate (GTP), thymidine monophosphate (TMP), thymidine diphosphate (TDP), thymidine triphosphate (TTP), uridine monophosphate (UMP), uridine diphosphate (UDP), uridine triphosphate (UTP), cytidine monophosphate (CMP), cytidine diphosphate (CDP), cytidine triphosphate (CTP), cyclic adenosine monophosphate (cAMP), cyclic guanosine monophosphate (cGMP), deoxyadenosine monophosphate (dAMP), and deoxyadenosine The free nucleotide may be one or more of the following: adenosine triphosphate (dADP), deoxyadenosine triphosphate (dATP), deoxyguanosine monophosphate (dGMP), deoxyguanosine diphosphate (dGDP), deoxyguanosine triphosphate (dGTP), deoxythymidine monophosphate (dTMP), deoxythymidine diphosphate (dTDP), deoxythymidine triphosphate (dTTP), deoxyuridine monophosphate (dUMP), deoxyuridine diphosphate (dUDP), deoxyuridine triphosphate (dUTP), deoxycytidine monophosphate (dCMP), deoxycytidine diphosphate (dCDP), and deoxycytidine triphosphate (dCTP). The free nucleotide is usually selected from AMP, TMP, GMP, CMP, UMP, dAMP, dTMP, dGMP, or dCMP. The free nucleotide is typically adenosine triphosphate (ATP).
[0190] A motor protein cofactor is a factor that enables the motor protein to function. The cofactor is preferably a divalent metal cation. The divalent metal cation is preferably Mg 2+ Mn 2+ Ca 2+ , or Cole 2+ The cofactor is most preferably Mg 2+ That is the case.
[0191] In some embodiments, the polynucleotide-binding protein is other than the motor protein as used herein. As used herein, the terms polynucleotide-binding protein and polynucleotide-binding moiety are interchangeable.
[0192] For example, a polynucleotide-binding protein or polynucleotide-binding moiety may contain one or more domains independently selected from helix-hairpin-helix (HhH) domains, eukaryotic single-strand binding proteins (SSBs), bacterial SSBs, archaeal SSBs, viral SSBs, double-strand binding proteins, sliding clamps, progression factors, DNA-binding loops, replication initiation proteins, telomere-binding proteins, repressors, zinc fingers, and proliferating cell nuclear antigens (PCNAs).
[0193] The helix-hairpin-helix (HhH) domain is a polypeptide motif that binds to DNA in a sequence-nonspecific manner. Preferred domains include domain H (residues 696-751) and domain H1 (residues 696-802) (SEQ ID NO: 54) derived from topoisomerase V from Metanopyrus candrelli. The polynucleotide binding portion may be domain HL of SEQ ID NO: 54 as shown in SEQ ID NO: 55, or its polynucleotide-binding variant. The HhH domain may include the sequences shown in SEQ ID NO: 40, 48, or 49, or their polynucleotide-binding variants.
[0194] SSBs bind to single-stranded DNA with high affinity in a sequence-nonspecific manner. SSBs are classified into the following series: Class: all beta proteins, Fold: OB fold, Superfamily: nucleoside-binding proteins, Family: single-stranded DNA-binding domain, SSB. SSBs can originate from eukaryotes such as humans, mice, rats, fungi, protists, or plants, prokaryotes such as bacteria and archaea, or viruses. Eukaryotic SSBs are also known as replication protein A (RPA). In most cases, they are heterotrimers formed from units of different sizes. Some of the larger units (e.g., RPA70 of Saccharomyces cerevisiae) are stable and bind to ssDNA in monomeric form. Bacterial SSBs bind to DNA as stable homotetramers (e.g., E. coli, Mycobacterium smegmatis, and Helicobacter pylori) or homodimers (e.g., Deinococcus radiodurans and Thermotoga maritima). Some SSBs, such as those encoded by the clenarchaeon Sulfolobus solfataricus, are homotetramers. Several SSBs from other species have been shown to be monomeric (Methanococcus jannaschii and Methanothermobacter thermoautotrophicum). Even more archaeal species, including Archaeoglobus fulgidus and Methanococcoides burtonii, contain two open reading frames that are sequence-similar to RPAs. Viral SSBs bind to DNA as monomers.
[0195] SSBs are typically selected or modified to have a carboxy-terminal (C-terminal) region with no net negative charge or a reduced net negative charge compared to wild-type proteins. Such SSBs usually do not block transmembrane pores. The C-terminal region of an SSB is typically the last approximately one-third, one-quarter, one-fifth, or one-eighth of the SSB at the C-terminus. The C-terminal region is typically the last approximately 20 to 40 amino acids of the SSB, such as the last approximately 10 to 60 amino acids of the C-terminus of the SSB, or the last approximately 30 amino acids of the C-terminus of the SSB.
[0196] Examples of SSBs containing a C-terminal region without a net negative charge include human mitochondrial SSB (HsmtSSB; SEQ ID NO: 50), human replication protein A70kDa subunit, human replication protein A14kDa subunit, telomere terminus, Oxytrichanova-derived binding protein α subunit, Oxytrichanova-derived telomere terminus binding protein β subunit core domain, Schizosaccharomyces pombe-derived telomere protein 1 protection (Pot1), human Pot1, mouse or rat-derived BRCA2 OB folded domain, phi29-derived p5 protein (SEQ ID NO: 51), and their polynucleotide-binding mutants. Examples of SSBs whose C-terminal region can be modified to reduce the net negative charge include E. coli SSB (EcoSSB; SEQ ID NO: 52), Mycobacterium tuberculosis SSB, Deinococcus radiodurans SSB, Thermus thermophiles-derived SSB, and Sulfolobus Examples include SSBs from *Solfatricus*, human replication protein A32kDa subunit (RPA32) fragments, CDC13SSB from *Saccharomyces cerevisiae*, Primosomal replication protein N (PriB) from *E. coli*, PriB from *Arabidopsis thaliana*, virtual protein At4g28440, SSB from T4 (gp32; SEQ ID NO: 53), SSB from RB69 (gp32; SEQ ID NO: 41), SSB from T7 (gp2.5; SEQ ID NO: 42), and polynucleotide links and their variants. Preferred modifications for reducing the net negative charge are disclosed in WO2014 / 013259.
[0197] Double-strand binding proteins bind to double-stranded DNA with high affinity. Suitable double-strand binding proteins include mutator S (MutS; NCBI reference sequence: NP_417213.1; SEQ ID NO: 56) and Sso7d (Sufolobus solfataricus P2; NCBI reference sequence: NP_343889.1; SEQ ID NO: 57; Nucleic Acids Research, 2004, Vol. 32, No.3, 1197-1207), Sso10b1 (NCBI reference sequence: NP_342446.1; SEQ ID NO: 58), Sso10b2 (NCBI reference sequence: NP_342448.1; SEQ ID NO: 59), Tryptophan repressor (Trp repressor; NCBI reference sequence: NP_291006.1; SEQ ID NO: 60), Lambda repressor (NCBI reference sequence: NP_040628.1; SEQ ID NO: 61), Cren7 (NCBI reference sequence: NP_342459.1; SEQ ID NO: 59), 62) Examples include, but are not limited to, major histone classes H1 / H5, H2A, H2B, H3, and H4 (NCBI reference sequence: NP_066403.2, SEQ ID NO: 63), dsbA (NCBI reference sequence: NP_049858.1; SEQ ID NO: 64), Rad51 (NCBI reference sequence: NP_002866.2; SEQ ID NO: 65), sliding clamp and topoisomerase VMka (SEQ ID NO: 54), or polynucleotide-linked mutants of any of these proteins.
[0198] Other polynucleotide-binding proteins include sliding clamps. Sliding clamps are typically multimer proteins (homodimers or homotrimers) that surround dsDNA. Sliding clamps usually require accessory proteins (clamp loaders) to assemble them around the DNA helix in an ATP-dependent process. They also function as topology tethers without direct contact with the DNA. Related to DNA sliding clamps are processivity factors, which are viral proteins that fix homogeneous polymerases to DNA, dramatically increasing the length of the resulting fragments. They can be monomers (in the case of UL42 from herpes simplex virus 1) or (UL44 from cytomegalovirus is a multimer dimer). UL42 typically contains the sequence shown in SEQ ID NO: 43 or SEQ ID NO: 47 or its polynucleotide-binding variants.
[0199] Another polynucleotide-binding protein is the thioredoxin-binding domain (TBD) (residues 258-333) of bacteriophage T7 DNA polymerase. Binding of the TBD to thioredoxin (e.g., from E. coli) causes a conformational change in the polypeptide that binds to DNA. Other polynucleotide-binding proteins include the accessory protein cisA from phage Φx174 and the geneII protein from phage M13. These proteins possess intrinsic DNA-binding capabilities, and some recognize specific DNA sequences. Other polynucleotide-binding proteins include telomere-binding proteins.
[0200] Small DNA-binding motifs (such as helix-turn-helix) recognize specific DNA sequences. In the case of bacteriophage 434 repressor, a 62-residue fragment was manipulated and shown to retain DNA-binding ability and specificity. Zinc fingers consist of approximately 30 amino acids that bind to DNA in a specific manner. Typically, each zinc finger recognizes only three DNA bases, but multiple fingers can be linked together to recognize longer sequences.
[0201] The proliferating cell nuclear antigen (PCNA) forms a very tight clamp that slides up and down dsDNA or ssDNA. PCNAs from Clenearchaea are heterotrimers of SEQ ID NOs. 44, 45, and 46. Therefore, the polynucleotide-binding protein may be a trimer containing the sequences shown in SEQ ID NOs. 44, 45, and 46, or a polynucleotide-binding variant thereof. Another PCNA sliding clamp (NCBI reference sequence: ZP_06863050.1; SEQ ID NO: 66) forms a dimer. Therefore, the polynucleotide-binding protein may be a dimer containing SEQ ID NO: 66, or a polynucleotide-binding variant thereof.
[0202] The polynucleotide binding motif can be selected from the following: [Table 3-1] [Table 3-2] [Table 3-3]
[0203] Polynucleotides The method of the present invention includes characterizing a target polynucleotide as it moves toward a detector such as a nanopore.
[0204] Polynucleotides, such as nucleic acids, are macromolecules containing two or more nucleotides. Polynucleotides can be single-stranded or double-stranded. Double-stranded polynucleotides are created from two single-stranded polynucleotides that have hybridized together. Target polynucleotides can be single-stranded or double-stranded polynucleotides.
[0205] Polynucleotides can contain any combination of any nucleotides. Nucleotides may be naturally occurring or artificially created.
[0206] Nucleotides typically consist of a nucleic acid base, a sugar, and at least one phosphate group. The nucleic acid base and sugar form a nucleoside.
[0207] Nucleic acid bases are typically heterocyclic. Nucleic acid bases include, but are not limited to, purines and pyrimidines, more specifically, adenine (A), guanine (G), thymine (T), uracil (U), and cytosine (C).
[0208] The sugar is typically a pentose. Nucleotide sugars include, but are not limited to, ribose and deoxyribose. The sugar is preferably deoxyribose. The polynucleotide preferably comprises the following nucleosides: deoxyadenosine (dA), deoxyuridine (dU) and / or thymidine (dT), deoxyguanosine (dG), and deoxycytidine (dC).
[0209] Nucleotides are typically ribonucleotides or deoxyribonucleotides. Nucleotides typically contain monophosphate, diphosphate, or triphosphate. Nucleotides may contain more than three phosphates, for example, four or five. Phosphates may be attached to the 5' or 3' end of the nucleotide. Nucleotides include, but are not limited to, adenosine monophosphate (AMP), guanosine monophosphate (GMP), thymidine monophosphate (TMP), uridine monophosphate (UMP), 5-methylcytidine monophosphate, 5-hydroxymethylcytidine monophosphate, cytidine monophosphate (CMP), cyclic adenosine monophosphate (cAMP), cyclic guanosine monophosphate (cGMP), deoxyadenosine monophosphate (dAMP), deoxyguanosine monophosphate (dGMP), deoxythymidine monophosphate (dTMP), deoxyuridine monophosphate (dUMP), deoxycytidine monophosphate (dCMP), and deoxymethylcytidine monophosphate. The nucleotides are preferably selected from AMP, TMP, GMP, CMP, UMP, dAMP, dTMP, dGMP, dCMP, and dUMP.
[0210] Nucleotides can be debased (i.e., lacking a nucleic acid base). Nucleotides can also lack both a nucleic acid base and a sugar (i.e., they are C3 spacers).
[0211] Nucleotides in a polynucleotide can attach to each other in any manner. Typically, nucleotides attach via their sugar and phosphate groups, similar to nucleic acids. Nucleotides can also be linked via their nucleic acid bases, similar to pyrimidine dimers.
[0212] Polynucleotides can be nucleic acids such as deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). A polynucleotide may include a strand of RNA hybridized to a strand of DNA. A polynucleotide can be any synthetic nucleic acid known in the art, such as peptide nucleic acid (PNA), glycerol nucleic acid (GNA), threose nucleic acid (TNA), cross-linked nucleotide (BNA), locked nucleic acid (LNA), or other synthetic polymers having nucleotide side chains. The PNA backbone consists of repeating N-(2-aminoethyl)-glycine units linked by peptide bonds. The GNA backbone consists of repeating glycol units linked by phosphodiester bonds. The TNA backbone consists of repeating threose sugars linked together by phosphodiester bonds. LNA is formed from ribonucleotides having extra crosslinks connecting the 2' oxygen and 4' carbon in the ribose moiety, as discussed above.
[0213] The polynucleotide is preferably DNA, RNA, or a DNA or RNA hybrid, most preferably DNA. The DNA / RNA hybrid may contain both DNA and RNA on the same strand. Preferably, the DNA / RNA hybrid contains a single DNA strand hybridized to an RNA strand.
[0214] The main chain of a polynucleotide can be modified to reduce the likelihood of breaks. For example, DNA is known to be more stable than RNA under many conditions. The main chain of a polynucleotide can be modified to avoid damage caused by radical chemicals such as free radicals.
[0215] DNA or RNA containing non-natural or modified bases can be produced by amplifying natural DNA or RNA polynucleotides in the presence of modified NTPs using a suitable polymerase.
[0216] Nucleotides in a polynucleotide can be modified. Nucleotides can be oxidized or methylated. One or more nucleotides in a polynucleotide may be damaged. For example, a polynucleotide may contain pyrimidine dimers. Such dimers are typically associated with UV damage and are a major cause of cutaneous melanoma. One or more nucleotides in a polynucleotide can be modified, for example, by labels or tags.
[0217] Single-stranded polynucleotides may contain regions with strong secondary structures, such as hairpin, quadruple-stranded, or triple-stranded DNA. These types of structures can be used to control the movement of polynucleotides toward a nanopore. For example, secondary structures can be used to stop the movement of polynucleotides through a nanopore, as will be described in more detail herein. Each successive secondary structure along the strand stops the movement of the strand toward the nanopore. After moving through the nanopore, the polynucleotide may reform the secondary structure. Using such secondary structures, it is possible to prevent the polynucleotide from returning through the nanopore when the negative voltage is low or not applied (applied to the trans side of the nanopore), and thus it is useful in controlling the movement of polynucleotides so that in the relevant steps of the method provided herein, it occurs only in a controlled manner.
[0218] When used herein, a double-stranded polypeptide may include a single-stranded region as well as regions having other structures, such as hairpin loops, triple-stranded, and / or quadruple-stranded regions. Such secondary structures may be useful in the context of single-stranded polynucleotides as described above.
[0219] The two chains of a double-stranded molecule can be covalently bonded, for example, by connecting the 5' end of one chain to the 3' end of the other chain in a hairpin structure at the end of the molecule.
[0220] The target polynucleotide can be of any length. For example, the target polynucleotide may be at least 10, at least 50, at least 100, at least 150, at least 200, at least 250, at least 300, at least 400, or at least 500 nucleotides or nucleotide pairs in length. The target polynucleotide may be at least 1,000 nucleotides or nucleotide pairs, or at least 5,000 nucleotides or nucleotide pairs in length, or at least 100,000 nucleotides or nucleotide pairs in length, or at least 500,000 nucleotides or nucleotide pairs in length, or at least 1,000,000 nucleotides or nucleotide pairs in length, or at least 10,000,000 nucleotides or nucleotide pairs in length, or at least 200,000,000 nucleotides or nucleotide pairs in length, or the entire length of the chromosome.
[0221] The target polynucleotide may be an oligonucleotide. An oligonucleotide is typically a short nucleotide polymer having 50 or fewer nucleotides, for example, 40 or fewer, 30 or fewer, 20 or fewer, 10 or fewer, or 5 or fewer nucleotides. The target oligonucleotide is preferably about 15 to about 30 nucleotides long, for example, about 20 to about 25 nucleotides long. For example, the oligonucleotide may be about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, or about 30 nucleotides long.
[0222] The target polynucleotide can be a fragment of a long polynucleotide. In this embodiment, the long target polynucleotide is typically fragmented into multiple shorter target polynucleotides, for example.
[0223] The target polynucleotide may include the product of a PCR reaction, genomic DNA, the product of endonuclease digestion, and / or a DNA library.
[0224] The target polynucleotide may occur naturally. The target polynucleotide may be secreted from cells. Alternatively, the target analyte may be an analyte present within a cell, and therefore the analyte must be extracted from the cell before the method can be performed.
[0225] The target polynucleotide may be derived from common organisms such as viruses, bacteria, archaea, plants, or animals. Such organisms may be selected or modified to adjust the sequence of the target polynucleotide, for example, by adjusting the base composition or removing undesirable sequence elements. The selection and modification of organisms to achieve desired polynucleotide properties is routine for those skilled in the art.
[0226] The source organism for the target polynucleotide can be selected based on the desired characteristics of the sequence. Desired characteristics include the ratio of single-stranded to double-stranded polynucleotides produced by the organism, the complexity of the sequence of the polynucleotides produced by the organism, the composition of the polynucleotides produced by the organism (such as the GC composition), or the length of the continuous polynucleotide chain produced by the organism. For example, if a continuous polynucleotide chain of about 50 kb is required, lambda phage DNA can be used. If a longer continuous chain is required, other organisms can be used to produce polynucleotides; for example, E. coli produces a continuous dsDNA of about 4.5 Mb.
[0227] Target polynucleotides are often obtained from humans or animals, for example, from urine, lymph, saliva, mucus, semen, or amniotic fluid, or from whole blood, plasma, or serum. Target polynucleotides can also be obtained from plants, for example, cereals, legumes, fruits, or vegetables. Target polynucleotides may include genomic DNA. Genomic DNA can be fragmented. DNA can be fragmented by any suitable method. For example, methods for fragmenting DNA are known in the art, and such methods can use transposases such as MuA transposase. Genomic DNA is often not fragmented.
[0228] In some embodiments, polynucleotides are synthetic or semi-synthetic. For example, DNA or RNA may be pure synthetics synthesized by conventional DNA synthesis methods, such as phosphoramidite-based chemical reactions. Synthetic polynucleotide subunits can be linked together by known means, such as ligation or chemical bonding, to produce longer chains. In some embodiments, internally self-forming structures (e.g., hairpins, quadruples) can be designed within the substrate, for example, by ligating a suitable sequence. Synthetic polynucleotides can be replicated and scaled up for production by means known in the art, including PCR and integration into bacterial factories.
[0229] In some embodiments, polynucleotides may have a simplified nucleotide composition. In some embodiments, polynucleotides have a repeating pattern of the same subunit. For example, the repeating unit may be (AmGn)q, where m, n, and q are positive integers. For example, m is often 1 to 20, e.g., 1 to 10, e.g., 1 to 5, e.g., 1, 2, 3, 4, or 5. n is often 1 to 20, e.g., 1 to 10, e.g., 1 to 5, e.g., 1, 2, 3, 4, or 5. m and n may be the same or different. Often, q is 1 to about 100,000. A typical repeating unit may be, for example, (AAAAAAGGGGGG)q. Repeating polynucleotides can be prepared by many means known in the art, for example, by linking together synthetic subunits having sticky ends that allow ligation. In some embodiments, polynucleotides may be linked polynucleotides. Methods for linking polynucleotides are described in PCT / GB2017 / 051493.
[0230] In some embodiments, the polynucleotide can include a base comprising a reactive side chain. Optionally, any suitable reactive functional group can be incorporated into the side chain. Suitable examples of reactive functional groups include click chemistry reagents. Suitable examples of click chemistry include, but are not limited to, the following. (a) Copper-free variant of the 1,3-dipolar cycloaddition reaction, where an azide reacts with a strained alkyne, for example, in a cyclooctane ring (b) Reaction of an oxygen nucleophile on one linker with an epoxide or aziridine reactive moiety on the other, and (c) Staudinger ligation, where an alkyne moiety is replaced with an arylphosphine to provide a specific reaction with an azide to afford an amide bond.
[0231] Polynucleotide adapter In some embodiments, the motor protein and / or polynucleotide-binding protein, if present, can be provided on the polynucleotide adapter. WO2015 / 110813 describes the loading of a target polynucleotide, such as an adapter for a motor protein, which is incorporated herein by reference in its entirety.
[0232] The adapter typically includes a polynucleotide chain capable of binding to the end of the target polynucleotide. The target polynucleotide is typically intended for characterization by the methods disclosed herein.
[0233] A polynucleotide adapter can be added to both ends of a target polynucleotide. Alternatively, different adapters can be added to those two ends of the target polynucleotide. The adapter can be added to only one end of the target polynucleotide. Methods for adding an adapter to a polynucleotide are known in the art. The adapter can be attached to the polynucleotide, for example, by ligation, by click chemistry, by tagmentation, by topoisomerase conversion, or by any other suitable method.
[0234] The adapter can be a composition or an artifact. Typically, the adapter includes the polymers described herein. In some embodiments, the adapter includes a polynucleotide. In some embodiments, the adapter can include a single-stranded polynucleotide chain. In some embodiments, the adapter can include a double-stranded polynucleotide. The polynucleotide adapter can include DNA, RNA, modified DNA (such as alkaline DNA), RNA, PNA, LNA, BNA, and / or PEG. Usually, the adapter includes single-stranded and / or double-stranded DNA or RNA.
[0235] The adapter can include a stall portion as described herein. The adapter can include a loading site for a motor protein or a polynucleotide-binding protein. The adapter can include a tag.
[0236] The adapter may be a Y-adapter. A Y-adapter is typically double-stranded and includes (a) a region at one end where the two strands hybridize together, and (b) a region at the other end where the two strands are not complementary. The non-complementary portion of the strands forms an overhang. The hybridized stem of the adapter typically attaches to the 5' end of the first strand of the double-stranded polynucleotide and the 3' end of the second strand of the double-stranded polynucleotide, or to the 3' end of the first strand of the double-stranded polynucleotide and the 5' end of the second strand of the double-stranded polynucleotide. The presence of the non-complementary region in the Y-adapter gives the adapter a Y shape because, unlike the double-stranded portion, the two strands do not typically hybridize to each other. A motor protein or polynucleotide may bind to the overhang of an adapter such as a Y-adapter. In another embodiment, a motor protein or polynucleotide-binding protein may bind to the double-stranded region. In other embodiments, a motor protein or polynucleotide-binding protein may bind to the single-stranded and / or double-stranded regions of the adapter. In other embodiments, a first motor protein or polynucleotide-binding protein may bind to the single-stranded region of such adapter, and a second motor protein or polynucleotide-binding protein may bind to the double-stranded region of the adapter.
[0237] In one embodiment, the adapter includes a membrane anchor or a pore anchor. In some embodiments, the anchor is complementary to an overhang to which a motor protein or polynucleotide-binding protein binds, and can therefore bind to a polynucleotide to which it hybridizes.
[0238] In some embodiments, one of the non-complementary strands of a polynucleotide adapter, such as a Y adapter, may include a leader sequence that can pass through a nanopore when in contact with a transmembrane pore.
[0239] The leader sequence typically comprises polynucleotides, such as DNA or RNA, modified polynucleotides (e.g., debasalized DNA), PNA, LNA, polyethylene glycol (PEG), or polymers such as polypeptides. In some embodiments, the leader sequence comprises single-stranded DNA, such as a polydT section. The leader sequence can be of any length, but is typically 10 to 150 nucleotides long, for example, 20 to 120, 30 to 100, 40 to 80, or 50 to 70 nucleotides long.
[0240] In one embodiment, the polynucleotide adapter is a hairpin loop adapter. The hairpin loop adapter is an adapter containing a single polynucleotide chain, the ends of which can or are hybridized to each other so that the central sections of the polynucleotide form a loop. A suitable hairpin loop adapter can be designed using methods known in the art. Typically, the 3' end of the hairpin loop adapter is bound to the 5' end of the first strand of a double-stranded polynucleotide, and the 5' end of the hairpin loop adapter is bound to the 3' end of the second strand of the double-stranded polynucleotide, or the 5' end of the hairpin loop adapter is bound to the 3' end of the first strand of a double-stranded polynucleotide, and the 3' end of the hairpin loop adapter is bound to the 5' end of the second strand of the double-stranded polynucleotide. As will be described in more detail below, the polynucleotide adapter can be bound to a target polynucleotide in order to characterize the target polynucleotide.
[0241] Those skilled in the art will understand that, if the adapter comprises a polynucleotide chain, the sequence of the adapter is typically not deterministic and can be controlled or selected according to other experimental conditions, such as the motor protein and any polynucleotide to be characterized. Exemplary sequences are provided in the examples for illustrative purposes only. For example, the adapter may comprise one or more sequences of SEQ ID NOs: 21-26 or 28-33, or a polynucleotide sequence having at least 20%, e.g., at least 30%, e.g., at least 40%, e.g., at least 50%, e.g., at least 60%, e.g., at least 70%, e.g., at least 80%, e.g., at least 90%, e.g., at least 95% sequence similarity or identity with one or more of SEQ ID NOs: 21-26 or 28-33. The sequence of the adapter can typically be modified without adversely affecting the effectiveness of the method provided herein.
[0242] In some embodiments, the polynucleotide adapter may include a loading site for loading a motor protein and / or a polynucleotide-binding protein. The loading site may be, for example, a single-stranded region that can be targeted by the motor protein or polynucleotide-binding protein. The loading site may be a region of the polynucleotide adapter to which an exogenous polynucleotide chain containing a motor protein or polynucleotide-binding protein can be bound in order to transfer the motor protein or polynucleotide-binding protein to the polynucleotide evaluated by the method provided herein.
[0243] Therefore, the motor protein used in the methods provided herein may stall on the polynucleotide adapter. In other embodiments, the motor protein stalls on the target polynucleotide but not on the polynucleotide adapter.
[0244] Blocking section In some embodiments, a blocking moiety can be used to prevent the motor protein from unassociating with the target polynucleotide.
[0245] In some embodiments, the blocking portion is included in the target polynucleotide. In some embodiments, the blocking portion is included in a polynucleotide adapter attached to the target polynucleotide. In some embodiments, the polynucleotide adapter, for example, the polynucleotide adapter described herein, includes the blocking portion.
[0246] Blocking regions can be used to prevent motor proteins from unassociating with the target polynucleotide. For example, if the motor protein is located at the 3' end of the polynucleotide chain in the target polynucleotide or polynucleotide adapter, the blocking region is typically located between the motor protein and the 3' end of the chain. If the motor protein is located at the 5' end of the polynucleotide chain in the target polynucleotide or polynucleotide adapter, the blocking region is typically located between the motor protein and the 5' end of the chain.
[0247] For example, in some embodiments, the polynucleotide adapter may include a first end containing an attachment site for attaching to a target polynucleotide analyte, and a second end, and the motor protein may stall on the polynucleotide adapter in an orientation for processing the adapter in the direction of the attachment site. In such embodiments, a blocking portion may be placed between the motor protein and the second end of the adapter to prevent the motor protein from unassociating from the second end of the polynucleotide adapter.
[0248] For example, in some embodiments, the polynucleotide adapter may include a 3' end containing an attachment site for binding to the 5' end of a target polynucleotide analyte, and a 5' end, and the motor protein may stall on the polynucleotide adapter in an orientation for processing the adapter in the direction of the 3' end, i.e., 5'→3' direction. In such embodiments, a blocking portion may be placed between the motor protein and the 5' end of the adapter to prevent the motor protein from unassociating with the 5' end of the polynucleotide adapter. In other embodiments, the polynucleotide adapter may include a 5' end containing an attachment site for binding to the 3' end of a target polynucleotide analyte, and a 3' end, and the motor protein may stall on the polynucleotide adapter in an orientation for processing the adapter in the direction of the 5' end, i.e., 3'→5' direction. In such embodiments, a blocking portion may be placed between the motor protein and the 3' end of the adapter to prevent the motor protein from unassociating with the 3' end of the polynucleotide adapter.
[0249] In some embodiments, the target polynucleotide includes a leader sequence at its first end, the motor protein stalls on the second end of the target polynucleotide or on an adapter bound to the second end of the target polynucleotide, and the blocking portion is located between the motor protein and the second end of the polynucleotide (i.e., the end of the polynucleotide at the second end of the polynucleotide), thereby preventing the motor protein from unassociating with the target polynucleotide at the second end of the target polynucleotide.
[0250] For example, in some embodiments, the target polynucleotide includes a leader sequence at the 5' end of the first strand, and the motor protein stalls at the 3' end of the first strand on an adapter attached to the 3' end of the first strand of the target polynucleotide, with a blocking portion located between the motor protein and the 3' end of the first strand of the polynucleotide, thereby preventing the motor protein from unassociating with the target polynucleotide at the 3' end of the first strand. In other embodiments, the target polynucleotide includes a leader sequence at the 3' end of the first strand, and the motor protein stalls at the 5' end of the first strand on an adapter attached to the 5' end of the first strand of the target polynucleotide, with a blocking portion located between the motor protein and the 5' end of the first strand of the polynucleotide, thereby preventing the motor protein from unassociating with the target polynucleotide at the 5' end of the first strand. Of course, the polynucleotide adapter can be attached to a double-stranded polynucleotide or a single-stranded polynucleotide. When the target polynucleotide is a double-stranded polynucleotide, the blocking region is typically located on the same chain as the motor protein. When motor proteins are present on each chain of a double-stranded polynucleotide (for example, when the double-stranded polynucleotide is rotationally symmetric), the blocking region is typically present on each chain of the polynucleotide.
[0251] Any suitable blocking portion can be used in the manner provided. Suitable blocking portions include many of the same groups that can be used as stopping portions as described herein. For example, a blocking portion may include one or more of the following: -Polynucleotide secondary structure, preferably hairpin or G-quadrivalent (TBA), - Preferably, a nucleic acid analog selected from peptide nucleic acid (PNA), glycerol nucleic acid (GNA), threose nucleic acid (TNA), locked nucleic acid (LNA), cross-linked nucleic acid (BNA), and debasic nucleotide. - Fluorophores, traptavidin, streptavidin and avidin such as neutravidin, and / or biotin, cholesterol, methylene blue, dinitrophenol (DNP), digoxigenin and / or anti-digoxigenin and dibenzylcyclooctyne groups, and - Polynucleotide-binding proteins. These elements are described in more detail herein in the context of the stop portion.
[0252] Spacer In some embodiments, the polynucleotide or polynucleotide adapter can include, for example, from 1 to about 10 spacers, for example, from 1 to about 5 spacers, for example, 1, 2, 3, 4, or 5 spacers. The spacer can include any suitable number of spacer units. The spacer typically provides an energy barrier that impedes the movement of the polynucleotide-binding protein. For example, the spacer can impede the movement of a motor protein or polynucleotide-binding protein by reducing the traction of the protein, for example, using a abasic spacer. The spacer can physically block the movement of the polynucleotide-binding protein, for example, by introducing a bulky chemical group to physically impede the movement of the protein.
[0253] In some embodiments, one or more spacers are included in the polynucleotide or polynucleotide adapter to provide a unique signal when they pass through the nanopore. One or more spacers can be used to define or separate one or more regions of the polynucleotide, for example, separating the adapter from the target polynucleotide.
[0254] In some embodiments, the spacer may include a linear molecule such as a polymer, for example, a polypeptide or polyethylene glycol (PEG). Typically, such a spacer has a structure different from the target polynucleotide. For example, when the target polynucleotide is DNA, the spacer or each spacer typically does not contain DNA. In particular, when the target polynucleotide is deoxyribonucleic acid (DNA) or ribonucleic acid (RNA), the spacer or each spacer preferably includes peptide nucleic acid (PNA), glycerol nucleic acid (GNA), threose nucleic acid (TNA), locked nucleic acid (LNA), or a synthetic polymer having nucleotide side chains. In some embodiments, the spacer is one or more nitroindoles, one or more inosines, one or more acridines, one or more 2-aminopurines, one or more 2-6-diaminopurines, one or more 5-bromodeoxyuridines, one or more inverted thymidines (inverted dT), one or more inverted dideoxythymidines (ddT), one or more dideoxycytidines (ddC), one or more 5-methylcytidines, one or more 5-hydroxymethylcytidines, one or more 2'-O-methylRNA bases, one or more iso - May contain deoxycytidine (Iso-dC), one or more iso-deoxyguanosine (Iso-dG), one or more C3 (OC3H6OPO3) groups, one or more photocleavable (PC)[OC3H6-C(O)NHCH2-C6H3NO2-CH(CH3)OPO3] groups, one or more hexanediol groups, one or more spacer 9 (iSp9)[(OCH2CH2)3OPO3] groups, or one or more spacer 18 (iSp18)[(OCH2CH2)6OPO3] groups, or one or more thiol bonds. Spacers may contain any combination of these groups. Many of these groups are commercially available from IDT® (Integrated DNA Technologies®). For example, C3, iSp9, and iSp18 spacers are all available from IDT®. Spacers may contain any number of the above groups as spacer units.
[0255] In some embodiments, the spacer may include one or more chemical groups, for example, one or more pendant chemical groups. One or more chemical groups may be attached to one or more nucleic acid bases in the polynucleotide adapter. One or more chemical groups may be attached to the backbone of the polynucleotide adapter. Any number of suitable chemical groups may be present, such as 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12 or more. Suitable groups include, but are not limited to, fluorophores, streptavidin and / or biotin, cholesterol, methylene blue, dinitrophenol (DNP), digoxigenin and / or anti-digoxigenin, and dibenzylcyclooctin groups.
[0256] In some embodiments, the spacer may contain one or more debased nucleotides (i.e., nucleotides lacking a nucleic acid base), for example, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12 or more debased nucleotides. The nucleic acid base may be replaced by -H (idSp) or -OH in the debased nucleotide. The debased spacer may be inserted into the target polynucleotide by removing a nucleic acid base from one or more adjacent nucleotides. For example, the polynucleotide may be modified to contain 3-methyladenine, 7-methylguanine, 1,N6-ethenoadenine inosine, or hypoxanthine, and the nucleic acid base may be removed from these nucleotides using human alkyladenine DNA glycosylase (hAAG). Alternatively, the polynucleotide may be modified to contain uracil, and the nucleic acid base may be removed by uracil DNA glycosylase (UDG). In one embodiment, the one or more spacers do not contain any debased nucleotides.
[0257] A suitable spacer can be designed or selected depending on the properties of the polynucleotide or polynucleotide adapter, the motor protein, and the conditions under which the method is to be performed.
[0258] tag In some embodiments, the polynucleotide or polynucleotide adapter may include a tag or tether. For example, the polynucleotide can be bound to a tag on a nanopore, for example, via its adapter, and released at some point during, for example, the characterization of the polynucleotide by the nanopore. Strong non-covalent bonds (e.g., biotin / avidin) are still reversible and are useful in some embodiments of the methods described herein.
[0259] A pair of pore tags and polynucleotide adapters may be configured such that the binding strength or affinity of the binding site on the polynucleotide to the tag on the nanopore (e.g., provided by an anchor or leader sequence on the adapter, or by a capture sequence in the double-stranded stem of the adapter) is sufficient to maintain the coupling between the nanopore and the polynucleotide until an applied force is applied and the bound polynucleotide is released from the nanopore.
[0260] In some embodiments, the tag or tether is uncharged. This ensures that the tag or tether is not drawn into the nanopore under the influence of a potential difference.
[0261] One or more molecules that attract or bind to a polynucleotide or adapter may be ligated to the pore. Any molecule that hybridizes to the adapter and / or target polynucleotide may be used. The molecules attached to the pore may be selected from PNA tags, PEG linkers, short oligonucleotides, positively charged amino acids, and aptamers. The binding of such molecules to pores is known in the art. For example, pores with attached short oligonucleotides are disclosed in Howarka et al (2001) Nature Biotech. 19:636-639 and WO2010 / 086620, and pores containing PEG attached within the lumen of the pore are disclosed in Howarka et al (2000) J.Am.Chem.Soc. 122(11):2411-2416.
[0262] The capture of the target polynucleotide in the method herein may be enhanced by using a short oligonucleotide attached to a detector (e.g., a transmembrane pore) that contains a sequence complementary to the leader sequence or another single-stranded sequence of the adapter.
[0263] In some embodiments, the tag or tether may include, or be composed of, an oligonucleotide (e.g., DNA, RNA, LNA, BNA, PNA, or morpholino). The oligonucleotide (e.g., DNA, RNA, LNA, BNA, PNA, or morpholino) may have a length of about 10 to 30 nucleotides or about 10 to 20 nucleotides. In some embodiments, the oligonucleotide (e.g., DNA, RNA, LNA, BNA, PNA, or morpholino) for use in the tag or tether may have at least one modified end (e.g., a 3'- or 5'-end) for bonding to other modification sites or to the surface of a solid substrate, such as beads. The end modifier may have reactive functional groups that can be used for bonding. Examples of functional groups that can be added include, but are not limited to, amino, carboxyl, thiol, maleimide, aminooxy, and any combination thereof. Functional groups can be combined with spacers of different lengths (e.g., C3, C9, C12, spacers 9 and 18) to add to the physical distance of the functional group from the ends of the oligonucleotide sequence.
[0264] In some embodiments, the tag or tether may contain or be a morpholinooligonucleotide. The morpholinooligonucleotide may have a length of about 10 to 30 nucleotides or about 10 to 20 nucleotides. The morpholinooligonucleotide may be modified or unmodified. For example, in some embodiments, the morpholinooligonucleotide may be modified at the 3' and / or 5' ends of the oligonucleotide. Examples of 3' and / or 5' end modifications of morpholinooligonucleotides include, but are not limited to, 3' affinity tags and functional groups for chemical bonding (e.g., 3'-biotin, 3'-primary amine, 3'-disulfideamide, 3'-pyridyldithio, and any combination thereof); 5' end modifications (e.g., 5'-primary ammine, and / or 5'-dabsyl); modifications for click chemistry (e.g., 3'-azide, 3'-alkyne, 5'-azide, 5'-alkyne), and any combination thereof.
[0265] In some embodiments, the tag or tether may further include a polymer linker to facilitate bonding to a detector, such as a nanopore. Exemplary polymer linkers include, but are not limited to, polyethylene glycol (PEG). The polymer linker may have a molecular weight of about 500 Da to about 10 kDa (inclusive of both ends), or about 1 kDa to about 5 kDa (inclusive of both ends). The polymer linker (e.g., PEG) can be functionalized with different functional groups, including, but not limited to, maleimide, NHS esters, dibenzocyclooctin (DBCO), azides, biotin, amines, alkynes, aldehydes, and any combination thereof. In some embodiments, the tag or tether may also include 1 kDa PEG having a 5'-maleimide group and a 3'-DBCO group. In some embodiments, the tag or tether may further include 2 kDa PEG having a 5'-maleimide group and a 3'-DBCO group. In some embodiments, the tag or tether may further comprise a 3 kDa PEG having a 5'-maleimide group and a 3'-DBCO group.
[0266] Other examples of tags or tethers include, but are not limited to, His tags, biotin or streptavidin, antibodies that bind to the sample, aptamers that bind to the sample, sample-binding domains such as DNA-binding domains (including peptide zippers like leucine zippers, single-stranded DNA-binding proteins (SSBs)), and any combination thereof.
[0267] Tags or tethers may be attached to the outer surface of a nanopore, for example, on the cis side of a membrane, using any method known in the art. For example, one or more tags or tethers can be attached to a nanopore via one or more cysteines (cysteine bonds), one or more primary amines such as lysine, one or more non-natural amino acids, one or more histidines (His tags), one or more biotins or streptavidins, one or more antibody-based tags, one or more enzymatic modifications of an epitope (including, for example, acetyltransferases), and any combination thereof. Preferred methods for carrying out such modifications are well known in the art. Preferred non-natural amino acids include, but are not limited to, 4-azido-L-phenylalanine (Faz) and any of the amino acids numbered 1 to 71 in Figure 1 of Liu CC and Schultz PG, Annu. Rev. Biochem., 2010, 79, 413-444.
[0268] In some embodiments, where one or more tags or tethers are attached to the nanopore via cysteine bonds, one or more cysteines can be introduced by substitution into one or more monomers forming the nanopore. In some embodiments, the nanopore may be chemically modified by the attachment of: (i) 4-phenylazomareinanyl, 1.N-(2-hydroxyethyl)maleimide, N-cyclohexylmaleimide, 1.3-maleimidopropionic acid, 1.1-4-aminophenyl-1H-pyrrole,2,5,dione, 1.1-4-hydroxyphenyl-1H-pyrrole,2,5,dione, N-ethylmaleimide, N-methoxycarbonylmale Imide, N-tert-butylmaleimide, N-(2-aminoethyl)maleimide, 3-maleimide-proxyl, N-(4-chlorophenyl)maleimide, 1-[4-(dimethylamino)-3,5-dinitrophenyl]-1H-pyrrole-2,5-dione, N-[4-(2-benzimidazolyl)phenyl]maleimide, N-[4-(2-benzoxazolyl)phenyl]maleimide, N-(1-naphthyl)-maleimide, N-(2, 4-Xylyl)maleimide, N-(2,4-difluorophenyl)maleimide, N-(3-chloro-para-tolyl)-maleimide, 1-(2-amino-ethyl)-pyrrole-2,5-dione hydrochloride, 1-cyclopentyl-3-methyl-2,5-dihydro-1H-pyrrole-2,5-dione, 1-(3-aminopropyl)-2,5-dihydro-1H-pyrrole-2,5-dione hydrochloride, 3-methyl-1-[2-oxo- 2-(piperazine-1-yl)ethyl]-2,5-dihydro-1H-pyrrole-2,5-dione hydrochloride, 1-benzyl-2,5-dihydro-1H-pyrrole-2,5-dione, 3-methyl-1-(3,3,3-trifluoropropyl)-2,5-dihydro-1H-pyrrole-2,5-dione, 1-[4-(methylamino)cyclohexyl]-2,5-dihydro-1H-pyrrole-2,5-dione trifluoroacetic acid, SMILES O=C1C=CC(=O)N1CC=2C=CN=CC2, SMILES O=C1C=CC(=O)N1CN2CCNCC2, 1-benzyl-3-methyl-2,5-dihydro-1H-pyrrole-2,5-dione, 1-(2-fluorophenyl)-3-methyl-2,Maleimides including diabromomaleimides such as 5-dihydro1H-pyrrole-2,5-dione, N-(4-phenoxyphenyl)maleimide, N-(4-nitrophenyl)maleimide, (ii)3-(2-iodoacetamide)-proxyl, N-(cyclopropylmethyl)-2-iodoacetamide, 2-iodo-N-(2-phenylethyl)acetamide, 2-iodo-N-(2,2,2-trifluoroethyl)acetamide, N-(4-acetylphenyl)-2-iodoacetamide, N-(4-(aminosulfonyl)phenyl)-2- iodoacetamides such as iodoacetamide, N-(1,3-benzothiazole-2-yl)-2-iodoacetamide, N-(2,6-(diethylphenyl)-2-iodoacetamide, N-(2-benzoyl-4-chlorophenyl)-2-iodoacetamide, (iii)N-(4-(acetylamino)phenyl)-2-bromoacetamide, N-(2-acetylphenyl)-2-bromoacetamide, 2-bromo-n-(2-cyanophenyl)acetamide, and 2-bromo-N-(3-((trifluoromethyl)phenyl)acetamide) N-(2-benzoylphenyl)-2-bromoacetamide, 2-bromo-N-(4-fluorophenyl)-3-methylbutanamide, N-benzyl-2-bromo-N-phenylpropionamide, N-(2-bromo-butyl)-4-chlorobenzenesulfonamide, 2-bromo-N-methyl-N-phenylacetamide, 2-bromo-N-phenethyl-acetamide, 2-adamantan-1-yl-2-bromo-N-cyclohexyl-acetamide, 2-bromo-N-(2-methylphenyl)butanamide, monobromoacetanilide Disulfides such as (iv) bromoacetamide, (iv) aldrithiol-2, aldrithiol-4, isopropyl disulfide, 1-(isobutyldisulfanyl)-2-methylpropane, dibenzyl disulfide, 4-aminophenyl disulfide, 3-(2-pyridyldithio)propionic acid, 3-(2-pyridyldithio)propionic acid hydrazide, 3-(2-pyridyldithio)propionic acid N-succinimidyl ester, am6amPDP1-βCD, and (v) 4-phenylthiazole-2-thiol, perpaldo, 5,6,7,Thiols such as 8-tetrahydroquinazoline-2-thiol.
[0269] In some embodiments, the tag or tether may be attached directly to the nanopore or via one or more linkers. The tag or tether may be attached to the nanopore using a hybrid linker as described in WO2010 / 086602. Alternatively, a peptide linker may be used. The peptide linker is an amino acid sequence. The length, flexibility, and hydrophilicity of the peptide linker are typically designed so as not to interfere with the function of the monomer and pore. A preferred flexible peptide linker is a stretch of 2 to 20, e.g., 4, 6, 8, 10, or 16 serine and / or glycine amino acids. A more preferred flexible linker comprises (SG)1, (SG)2, (SG)3, (SG)4, (SG)5, and (SG)8, where S is serine and G is glycine. A preferred rigid linker is a stretch of 2 to 30, e.g., 4, 6, 8, 16, or 24 proline amino acids. A more preferred rigid linker is one in which P is proline, (P) 12 Includes.
[0270] anchor In one embodiment, the polynucleotide or polynucleotide adapter may include a membrane anchor or a transmembrane pore anchor. In one embodiment, the anchor assists in the characterization of a target polynucleotide according to the method disclosed herein. For example, the membrane anchor or transmembrane pore anchor may facilitate the localization of selected polynucleotides around nanopores.
[0271] The anchor may be a polypeptide anchor and / or a hydrophobic anchor that can be inserted into the membrane. In one embodiment, the hydrophobic anchor is a lipid, fatty acid, sterol, carbon nanotube, polypeptide, protein, or amino acid, for example, cholesterol, palmitate, or tocopherol. The anchor may include thiols, biotin, or surfactants. In one embodiment, the anchor may be biotin (for binding to streptavidin), amylose (for binding to maltose-binding proteins or fusion proteins), Ni-NTA (for binding to polyhistidine or polyhistidine-tagged proteins), or a peptide (such as an antigen).
[0272] In one embodiment, the anchor may comprise one linker, or two, three, four or more linkers. Preferred linkers include, but are not limited to, polymers such as polynucleotides, polyethylene glycol (PEG), polysaccharides, and polypeptides. These linkers may be linear, branched, or cyclic. For example, the linker may be a cyclic polynucleotide. The adapter may hybridize to a complementary sequence on the cyclic polynucleotide linker. One or more anchors or one or more linkers may comprise components that can be cleaved or degraded, such as limiting sites or photodissociable groups. The linkers are functionalized with maleimide groups to attach to cysteine residues of proteins. Preferred linkers are described in WO2010 / 086602.
[0273] In one embodiment, the anchor is cholesterol or a fatty acyl chain. For example, any fatty acyl chain having a carbon atom length of 6 to 30, such as hexadecanoic acid, can be used. Examples of suitable anchors and methods for attaching the anchor to the adapter are disclosed in WO2012 / 164270 and WO2015 / 150786.
[0274] In another embodiment, the anchor may consist of or include a hydrophobic modification of a polynucleotide or polynucleotide adapter. The hydrophobic modification may include a modified phosphate group contained within the polynucleotide or polynucleotide anchor. The hydrophobic modification may include phosphorothioates such as charge-neutralized alkyl phosphorothioates (PPTs), which are described in Jones et al, J.Am.Chem.Soc.2021,143,22,8305, the full contents of which are incorporated herein by reference. Suitable alkyl groups include, for example, C1-C2 alkyl groups such as C2-C6 alkyl groups. 10 Alkyl groups include, for example, methyl, ethyl, propyl, butyl, pentyl, and hexyl groups. Incorporation of charge-neutralized alkyl-phosphorothioates into polynucleotides allows the polynucleotides to engage with hydrophobic regions such as lipid bilayers.
[0275] detector In the method provided herein, polynucleotides are moved toward a detector such as a nanopore. The detector may be selected from (i) a zero-mode waveguide, (ii) a field-effect transistor, optionally a Noyer field-effect transistor, (iii) an AFM chip, (iv) a nanotube, optionally a carbon nanotube, and (v) a nanopore. Preferably, the detector is a nanopore.
[0276] Polynucleotides can be characterized in any preferred manner in the methods provided herein. In one embodiment, the polynucleotide is characterized by detecting an ionic flow or optical signal as the polynucleotide moves toward a nanopore. This is described in more detail herein. This method is suitable for these and other methods of detecting polynucleotides.
[0277] In another non-limiting example, in one embodiment, a polynucleotide is characterized by detecting a byproduct of a polynucleotide processing reaction, such as sequencing by a synthetic reaction. Thus, this method may include detecting the product of the sequential addition of (poly)nucleotides to a nucleic acid chain by an enzyme such as polymerase. The product may be a change in one or more properties of the enzyme, such as the conformation of the enzyme. Thus, such a method may include subjecting an enzyme such as polymerase or reverse transcriptase to a double-stranded polynucleotide under conditions such that the template-dependent incorporation of nucleotide bases into an elongating oligonucleotide chain causes a conformational change of the enzyme in response to the templates encountered sequentially; the incorporation of strand nucleic acid bases and / or template-specific native or analogous bases (i.e., an incorporation event); detection of the conformational change of the enzyme in response to such incorporation event; and detection of the sequence of the template chain as a result. In such a method, the polynucleotide chain may be moved according to a method provided herein. Such a method may include detecting and / or measuring the incorporation event using a method known to those skilled in the art, such as the method described in US2017 / 0044605.
[0278] In another embodiment, when a nucleotide is added to a synthetic nucleic acid chain complementary to the template chain, a phosphate-labeled species is released, and the byproduct can be labeled so that the phosphate-labeled species can be detected using a detector described herein. The polynucleotide thus characterized can be moved according to the method described herein. Suitable labels may be optical labels detected using a nanopore or zero-mode waveguide, or by Raman spectroscopy or other detectors. Suitable labels may be non-optical labels detected using a nanopore or other detector.
[0279] In an alternative approach, the nucleoside phosphate (nucleotide) is not labeled, and when the nucleotide is added to a synthetic nucleic acid chain complementary to the template chain, the natural byproduct species is detected. Suitable detectors may be ion-sensitive field-effect transistors or other detectors.
[0280] These and other detection methods are suitable for use in the methods described herein. Any suitable measurement can be performed using the detector as the polynucleotides move toward the detector.
[0281] nanopore In embodiments of the present invention where the detector is a nanopore, any suitable nanopore can be used. In one embodiment, the nanopore is a transmembrane pore.
[0282] A transmembrane pore is a structure that extends to some extent across a membrane. This allows hydrated ions, driven by an applied potential, to flow across or within the membrane. Typically, transmembrane pores traverse the entire membrane, allowing hydrated ions to flow from one side of the membrane to the other. However, transmembrane pores do not necessarily have to traverse the membrane; one end may be closed. For example, a pore can be a well, gap, channel, groove, or slit in the membrane, through which hydrated ions can flow.
[0283] In the methods provided herein, the nanopore typically has a first opening and a second opening. The first opening is typically a cis opening, and the second opening is typically a trans opening. However, in some embodiments, the first opening is a trans opening and the second opening is a cis opening. The motor protein used in the methods provided herein is typically supplied to the first opening of the nanopore and thus controls the movement of a target polynucleotide in the direction from the second opening of the nanopore toward the first opening of the nanopore.
[0284] Any transmembrane pore may be used in the manner provided herein. The pore may be biological or artificial. Suitable pores include, but are not limited to, protein pores, polynucleotide pores, and solid pores. The pore may be a DNA origami pore (Langecker et al., Science, 2012;338:932-936). Suitable DNA origami pores are disclosed in WO2013 / 083983.
[0285] In one embodiment, the nanopore is a transmembrane protein pore. A transmembrane protein pore is a polypeptide or a polypeptide aggregate that allows hydrated ions, such as polynucleotides, to flow from one side of a membrane to the other. In the methods provided herein, the transmembrane protein pore can form a pore that allows hydrated ions, driven by an applied potential, to flow from one side of a membrane to the other. Preferably, the transmembrane protein pore allows polynucleotides to flow from one side of a membrane to the other, such as a triblock copolymer membrane. The transmembrane protein pore allows polynucleotides to move through the pore.
[0286] In one embodiment, the nanopore is a transmembrane protein pore that is a monomer or oligomer. The pore is preferably composed of several repeating subunits, such as at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, or at least 16 subunits. The pore is preferably a hexamer, heptamer, octamaer, or nanomeric pore. The pore may be a homo-oligomer or a hetero-oligomer.
[0287] In one embodiment, a transmembrane protein pore includes a barrel or channel through which ions can flow. Subunits of the pore typically surround a central axis and contribute chains to a transmembrane β-barrel or channel or a transmembrane α-helix bundle or channel.
[0288] Typically, the barrel or channel of a transmembrane protein pore contains amino acids that facilitate interaction with an analyte, such as a target polynucleotide (as described herein). These amino acids are preferably located near the constriction of the barrel or channel. The transmembrane protein pore typically contains one or more positively charged amino acids, such as arginine, lysine, or histidine, or aromatic amino acids, such as tyrosine or tryptophan. These amino acids typically facilitate interaction between the pore and nucleotides, polynucleotides, or nucleic acids.
[0289] In one embodiment, the nanopore is a transmembrane protein pore derived from a β-barrel pore or an α-helix bundle pore. The β-barrel pore includes a barrel or channel formed from a β-chain. Suitable β-barrel pores include, but are not limited to, β-toxins, e.g., α-hemolytic toxins, anthrax toxins, and leucosidines, as well as bacterial outer membrane proteins / porins, e.g., Mycobacterium smegmatis porins (Msp), e.g., MspA, MspB, MspC, or MspD, CsgG, outer membrane porin F (OmpF), outer membrane porin G (OmpG), outer membrane phospholipase A, and Neisseria autotransporter lipoprotein (NalP), and other pores, e.g., lysenin. The α-helix bundle pore includes a barrel or channel formed from an α-helix. Suitable α-helix bundle pores include, but are not limited to, inner membrane proteins and α-outer membrane proteins, such as WZA and ClyA toxin.
[0290] In one embodiment, the nanopore is derived from or based on Msp, α-hemolysin (α-HL), lysenin, CsgG, ClyA, Sp1, or the hemolytic protein fragaceatoxin C (FraC).
[0291] In one embodiment, the nanopore is derived from CsgG, for example, CsgG from E. coli strain K-12 sub-strain MC4100. Such a pore is an oligomer and typically contains 7, 8, 9, or 10 monomers derived from CsgG. The pore may be a homooligomeric pore derived from CsgG containing the same monomers. Alternatively, the pore may be a heterooligomeric pore derived from CsgG containing at least one monomer different from the others. Examples of suitable pores derived from CsgG are disclosed in WO2016 / 034591.
[0292] In one embodiment, the nanopore is a transmembrane pore derived from lysenine. Examples of suitable pores derived from lysenine are disclosed in WO2013 / 153359.
[0293] In one embodiment, the nanopore is a transmembrane pore derived from or based on α-hemolysin (α-HL). The wild-type α-hemolysin pore is formed from seven identical monomers or subunits (i.e., it is a nanomer). The α-hemolysin pore may be α-hemolysin-NN or a variant thereof. The variant preferably contains N residues at positions E111 and K147.
[0294] In one embodiment, the nanopore is a transmembrane protein pore derived from Msp, for example, MspA. Examples of suitable pores derived from MspA are disclosed in WO2012 / 107778.
[0295] In one embodiment, the nanopore is a transmembrane pore derived from or based on ClyA.
[0296] film In the disclosed method, the detector is typically a nanopore present in a film. Any suitable film may be used.
[0297] The membrane is preferably an amphiphilic layer. The amphiphilic layer is a layer formed from amphiphilic molecules such as phospholipids that have both hydrophilic and lipophilic properties. The amphiphilic molecules may be synthetic or naturally occurring. Amphiphilic substances that do not exist naturally and amphiphilic substances that form monolayers are known in the art and include, for example, block copolymers (Gonzalez-Perez et al., Langmuir, 2009, 25, 10447-10450). Block copolymers are polymer materials in which two or more monomer subunits are polymerized together to produce a single polymer chain. Block copolymers typically have properties contributed by each monomer subunit. However, block copolymers may have unique properties that polymers formed from individual subunits do not have. Block copolymers can be manipulated so that one of the monomer subunits is hydrophobic (i.e., lipophilic) while the other subunits are hydrophilic in an aqueous medium. In this case, the block copolymer may have amphiphilic properties and can form structures that mimic biological membranes. The block copolymer may be a diblock (consisting of two monomer subunits), but it may be constructed from three or more monomer subunits to form more complex arrangements that behave as amphiphilic materials. The copolymer may be a triblock, tetrablock, or pentablock copolymer. The membrane is preferably a triblock copolymer membrane.
[0298] Archaeal bipolar tetraether lipids are naturally occurring lipids that are constructed to form monolayer membranes. These lipids are commonly found in extremophilic, thermophilic, halophilic, and acidophilic bacteria that survive in harsh biological environments. Their stability is thought to derive from the fusion properties of the final bilayer. It is straightforward to construct block copolymers that mimic these biological entities by creating triblock polymers with a common hydrophilic-hydrophobic-hydrophilic motif. This material forms monomer membranes that behave similarly to lipid bilayers and can encompass a wide range of phase behaviors, from vesicles to layered membranes. Membranes formed from these triblock copolymers have several advantages over biological lipid membranes. Because triblock copolymers are synthetic, their precise structure can be carefully controlled to provide the correct chain length and properties necessary to form membranes and interact with pores and other proteins.
[0299] Block copolymers may be constructed from subunits not classified as lipid submaterials; for example, hydrophobic polymers can be made from siloxanes or other non-hydrocarbon monomers. The hydrophilic subsections of block copolymers may also possess low protein-binding properties, which allows for the creation of membranes that are highly resistant when exposed to raw biological samples. These head group units may also originate from unclassified lipid head groups.
[0300] Triblock copolymer membranes also possess increased mechanical and environmental stability compared to biolipid membranes, such as a much higher operating temperature or pH range. The synthetic properties of block copolymers provide a basis for customizing polymer-based membranes for a wide range of applications.
[0301] In some embodiments, the film is one of the films disclosed in International Application No. WO2014 / 064443 or WO2014 / 064444.
[0302] The amphiphilic molecules may be chemically modified or functionalized to facilitate the coupling of polynucleotides. The amphiphilic layer may be a monolayer or a bilayer. The amphiphilic layer is typically planar. The amphiphilic layer may be curved. The amphiphilic layer may be supported.
[0303] Amphiphilic membranes are typically about 10 -8 cm s -1 It is naturally mobile, essentially acting as a two-dimensional fluid with a lipid diffusion rate. This means that pores and coupled polynucleotides can typically move within amphiphilic membranes.
[0304] The membrane may be a lipid bilayer. Lipid bilayers are a model of cell membranes and serve as an excellent base for a wide range of experimental studies. For example, lipid bilayers can be used for in vitro investigations of membrane proteins by single-channel recording. Alternatively, lipid bilayers can be used as biosensors for detecting the presence of a wide range of substances. The lipid bilayer can be any lipid bilayer. Preferred lipid bilayers include, but are not limited to, planar lipid bilayers, supported bilayers, or liposomes. The lipid bilayer is preferably a planar lipid bilayer. Preferred lipid bilayers are disclosed in WO2008 / 102121, WO2009 / 077734, and WO2006 / 100484.
[0305] Methods for forming lipid bilayers are known in the art. Lipid bilayers are generally formed by the method of Montal and Mueller (Proc. Natl. Acad. Sci. USA., 1972; 69: 3561-3566), in which a lipid monolayer is supported on an aqueous solution / air interface, extending across an opening perpendicular to the interface. The lipid is usually added to the surface of the aqueous electrolyte by first dissolving it in an organic solvent and then evaporating a small amount of solvent over the aqueous solution interface on both sides of the opening. As the organic solvent evaporates, the solution / air interfaces on both sides of the opening physically move up and down across the opening until the bilayer is formed. Planar lipid bilayers can be formed across the opening in a membrane or across the opening in a recess.
[0306] The Montal & Mueller method is popular because it is a cost-effective, relatively simple method for forming high-quality lipid bilayers suitable for protein pore insertion. Other common methods for bilayer formation include tip-dipping, painting bilayers, and liposome bilayer patch clamping.
[0307] Tip-immersion bilayer formation involves bringing an opening (e.g., a pipette tip) into contact with the surface of a test solution supporting a lipid monolayer. In this case, too, the lipid monolayer is formed at the solution / air interface by evaporating a small amount of lipid initially dissolved in an organic solvent at the solution surface. The bilayer is then formed by the Langmuir-Shafer method, requiring mechanical automation to move the opening relative to the solution surface.
[0308] In the case of bilayer coating, a small amount of lipid dissolved in an organic solvent is applied directly to the opening immersed in the test aqueous solution. The lipid solution is spread thinly across the opening using a paint brush or equivalent. Diluting the solvent leads to the formation of a lipid bilayer. However, it is difficult to completely remove the solvent from the bilayer, and as a result, the bilayer formed by this method is less stable and prone to generating noise during electrochemical measurements.
[0309] Patch clamping is commonly used in the study of biological cell membranes. The cell membrane is secured to the end of the pipette by aspiration, allowing a patch of membrane to adhere across the opening. This method has been adapted to produce lipid bilayers by securing liposomes and then rupturing them to seal the lipid bilayer across the pipette opening. This method requires the creation of stable, large monolayer liposomes and smaller openings in materials with glass surfaces.
[0310] Liposomes can be formed by sonication, extrusion, or the Mozafari method (Colas et al. (2007) Micron 38:841-847).
[0311] In some embodiments, the lipid bilayer is formed as described in International Application WO2009 / 077734. Advantageously, this method forms the lipid bilayer from dry lipids. In the most preferred embodiment, the lipid bilayer is formed across the opening as described in WO2009 / 077734.
[0312] A lipid bilayer is formed from two opposing layers of lipids. These two lipid layers are arranged such that their hydrophobic tail groups face each other, forming a hydrophobic interior. The hydrophilic head groups of the lipids face outward toward the aqueous environment on both sides of the bilayer. Bilayers can exist in several lipid phases, including, but not limited to, liquid disordered phases (fluid layered), liquid ordered phases, solid ordered phases (layered gel phases, comb-shaped gel phases), and planar bilayer crystals (layered subgel phases, lamellar crystal phases).
[0313] Any lipid composition that forms a lipid bilayer can be used. The lipid composition is selected so as to form a lipid bilayer having the desired properties, such as surface charge, ability to support membrane proteins, packing density, or mechanical properties. The lipid composition may contain one or more different lipids. For example, the lipid composition may contain up to 100 lipids. The lipid composition preferably contains 1 to 10 lipids. The lipid composition may contain naturally occurring lipids and / or artificial lipids.
[0314] Lipids typically comprise a head group, an interfacial moiety, and two hydrophobic tail groups, which may be the same or different. Suitable head groups include, but are not limited to, neutral head groups, e.g., diacylglycerides (DG) and ceramides (CM); zwitterionic head groups, e.g., phosphatidylcholine (PC), phosphatidylethanolamine (PE), and sphingomyelin (SM); negatively charged head groups, e.g., phosphatidylglycerol (PG), phosphatidylserine (PS), phosphatidylinositol (PI), phosphatidic acid (PA), and cardiolipin (CA); and positively charged head groups, e.g., trimethylammonium-propane (TAP). Suitable interfacial moieties include, but are not limited to, naturally occurring interfacial moieties, e.g., glycerol-based moieties or ceramide-based moieties. Suitable hydrophobic tail groups include, but are not limited to, saturated hydrocarbon chains such as lauric acid (n-dodecanolic acid), myristic acid (n-tetradecononic acid), palmitic acid (n-hexadecanoic acid), stearic acid (n-octadecanoic acid), and arachidic acid (n-eicosanoic acid), unsaturated hydrocarbon chains such as oleic acid (cis-9-octadecanoic acid), and branched hydrocarbon chains such as phytanoyl. The length of the chain and the position and number of double bonds in the unsaturated hydrocarbon chains may vary. The length of the chain and the position and number of branches, such as methyl groups in the branched hydrocarbon chains, may vary. The hydrophobic tail groups may be linked to the interface as ethers or esters. The lipids may be mycolic acids.
[0315] Lipids can also be chemically modified. The head or tail groups of lipids can be chemically modified. Suitable lipids with chemically modified head groups include, but are not limited to, PEG-modified lipids, e.g., 1,2-diacyl-sn-glycero-3-phosphoethanolamine-N-[methoxy(polyethylene glycol)-2000], functionalized PEG lipids, e.g., 1,2-distearoyl-sn-glycero-3-phosphoethanolamine-N-[biotinyl(polyethylene glycol)2000], and lipids modified for conjugation, e.g., 1,2-dioleoyl-sn-glycero-3-phosphoethanolamine-N-(succinyl) and 1,2-dipalmitoyl-sn-glycero-3-phosphoethanolamine-N-(biotinyl). Suitable lipids with chemically modified tail groups include, but are not limited to, polymerizable lipids such as 1,2-bis(10,12-tricosadiynoyl)-sn-glycero-3-phosphocholine, fluorinated lipids such as 1-palmitoyl-2-(16-fluoropalmitoyl)-sn-glycero-3-phosphocholine, deuterated lipids such as 1,2-dipalmitoyl-D62-sn-glycero-3-phosphocholine, and ether-binding lipids such as 1,2-di-O-phytanyl-sn-glycero-3-phosphocholine. Lipids may be chemically modified or functionalized to facilitate polynucleotide coupling.
[0316] The amphiphilic layer, for example, a lipid composition, typically contains one or more additives that will affect the properties of the layer. Suitable additives include, but are not limited to, fatty acids, such as palmitic acid, myristic acid, and oleic acid; fatty alcohols, such as palmitic alcohol, myristic alcohol, and oleic alcohol; sterols, such as cholesterol, ergosterol, lanosterol, sitosterol, and stigmasterol; lysophospholipids, such as 1-acyl-2-hydroxy-sn-glycero-3-phosphocholine; and ceramides.
[0317] In another embodiment, the film includes a solid layer. The solid layer may include, but is not limited to, microelectronic materials, insulating materials such as Si3N4, A12O3, and SiO, organic and inorganic polymers such as polyamides, plastics such as Teflon®, or elastomers such as two-component addition-cured silicone rubber, and glass, and may be formed from both organic and inorganic materials. The solid layer may be formed from graphene. A suitable graphene layer is disclosed in WO2009 / 035647. When the film includes a solid layer, pores are typically located within the solid layer, e.g., within holes, wells, gaps, channels, grooves, or slits in the amphiphilic film or layer contained within the solid layer. Those skilled in the art can prepare suitable solid / amphiphilic hybrid systems. Suitable systems are disclosed in WO2009 / 020682 and WO2012 / 005857. Any of the amphiphilic films or layers considered above may be used.
[0318] The methods disclosed herein are typically carried out using (i) an artificial amphiphilic layer containing a pore, (ii) an isolated, naturally occurring lipid bilayer containing a pore, or (iii) a cell into which a pore has been inserted. The method is typically carried out using an artificial amphiphilic layer, such as an artificial triblock copolymer layer. In addition to the pore, the layer may contain other transmembrane proteins and / or intramembrane proteins, as well as other molecules. Preferred apparatus and conditions are discussed below. The methods of the present invention are typically carried out in vitro.
[0319] General method As described above, the methods provided herein may be operated using any suitable detector, and therefore any suitable apparatus for detecting polynucleotides may be used.
[0320] In some embodiments, this method can be carried out using any apparatus suitable for sensing transmembrane pores. For example, the apparatus may comprise a chamber containing an aqueous solution and a barrier separating the chamber into two sections. The barrier may have an opening on which a membrane containing a transmembrane is formed. Transmembrane pores are described herein.
[0321] This method can be carried out using the apparatus described in WO2008 / 102120, WO2010 / 122293, or WO00 / 28312. In short, the binding of a molecule (e.g., a target polynucleotide) within a pore channel affects the open-channel ion flow through the pore, which is the essence of "molecular sensing" of pore channels. Fluctuations in the open-channel ion flow can be measured using suitable measurement techniques based on changes in current. The degree of decrease in ion flow, measured by a decrease in current, is related to the size of obstacles within or near the pore. Thus, the binding of a target molecule (e.g., a target polynucleotide) within or near the pore provides a detectable and measurable event, thereby forming the basis for a "biological sensor." By detecting the presence of biomolecules, applications can be found in personalized drug development, medicine, diagnostics, life science research, environmental monitoring, and the security and / or defense industries.
[0322] When used to characterize a polynucleotide, the presence or absence of a target polynucleotide, or one or more features, is determined. The method may be for determining the presence or absence of at least one target polynucleotide, or one or more features. The method may relate to determining the presence or absence of two or more target polynucleotides, or one or more features. The method may include determining the presence or absence of any number of target polynucleotides, e.g., 2, 5, 10, 15, 20, 30, 40, 50, 100 or more target polynucleotides, or one or more features. Any number of features of one or more target polynucleotides, e.g., 1, 2, 3, 4, 5, 10 or more features, can be determined. Features that can be detected by the methods provided herein include the identity or sequence of a polynucleotide, the length of a polynucleotide, and whether or not a polynucleotide is modified. In some embodiments, the methods provided herein are methods for sequencing target polynucleotides. In some embodiments, the polynucleotide sequence can be determined in real time by aligning a real-time signal or base calling against a known reference. An exemplary method for determining polynucleotide sequences is described in WO2016 / 059427, which is incorporated herein by reference.
[0323] When used for characterizing polynucleotides, the method may typically involve measuring the flow of ionic current through the pore by measuring electric current. Alternatively, the flow of ions through the pore may be measured optically, as disclosed by Heron et al: J.Am.Chem.Soc.9 Vol.131, No.5, 2009. Thus, the apparatus may also include an electrical circuit capable of applying a potential and measuring electrical signals across the membrane and pore. The characterization method may be performed using patch clamp or voltage clamp. The characterization method preferably involves the use of voltage clamp.
[0324] This method may include measuring optical signals, as described by Chen et al., Nature Communications (2018) 9:1733, the full details of which are incorporated herein by reference. For example, nanopores such as optically designed nanopore structures (e.g., plasmonic nanoslits) can be used to locally enable single-molecule surface-enhanced Raman spectroscopy (SERS) and characterize polynucleotides by direct Raman spectroscopic detection.
[0325] This method can be implemented with silicon-based well arrays, each having 128, 256, 512, 1024, 2000, 3000, 4000, 6000, 10000, 12000, or 15000 or more wells.
[0326] This method may include measuring the current flowing through the pore. This method is typically performed with a voltage applied across the membrane and the pore. The voltage used is typically +2V to -2V, and typically -400mV to +400mV. The voltage used is preferably within a range having a lower limit selected from -400mV, -300mV, -200mV, -150mV, -100mV, -50mV, -20mV, and 0mV, and an upper limit independently selected from +10mV, +20mV, +50mV, +100mV, +150mV, +200mV, +300mV, and +400mV. The voltage used is more preferably in the range of 100mV to 240mV, and most preferably in the range of 120mV to 220mV. By using an increased applied potential, it is possible to increase the discrimination between different nucleotides in each pore.
[0327] In some embodiments of the methods disclosed, particularly those involving the re-reading of a target polynucleotide as described herein, the methods include providing conditions for promoting the debinding of the target polynucleotide from the polynucleotide binding site of a motor protein, and / or delaying the rebinding of the target polynucleotide to the polynucleotide binding site of a motor protein.
[0328] This method is typically carried out in the presence of any charge carrier, such as metal salts, alkali metal salts, halogen salts, chloride salts, or alkali metal chloride salts. The charge carrier may include ionic liquids or organic salts, such as tetramethylammonium chloride, trimethylphenylammonium chloride, phenyltrimethylammonium chloride, or 1-ethyl-3-methylimidazolium chloride. In the exemplary apparatus considered above, the salt is present in an aqueous solution in the chamber. Potassium chloride (KCl), sodium chloride (NaCl), or cesium chloride (CsCl) are typically used. KCl is preferred. The salt may be an alkaline earth metal salt such as calcium chloride (CaCl2). The salt concentration may be at saturation. The salt concentration may be 3 M or less, and is typically 0.1–2.5 M, 0.3–1.9 M, 0.5–1.8 M, 0.7–1.7 M, 0.9–1.6 M, or 1 M–1.4 M. The salt concentration is preferably 150 mM to 1 M. The characterization method is preferably performed using salt concentrations of at least 0.3 M, for example, at least 0.4 M, at least 0.5 M, at least 0.6 M, at least 0.8 M, at least 1.0 M, at least 1.5 M, at least 2.0 M, at least 2.5 M, or at least 3.0 M. Higher salt concentrations provide a high signal-to-noise ratio and allow for the identification of currents that show coupling / uncoupling to the background of normal current fluctuations.
[0329] In some embodiments, providing the condition involves providing a salt concentration to increase the rate at which a target polynucleotide dissociates from the polynucleotide binding site of a motor protein. In some embodiments, providing the condition involves providing a salt concentration to decrease the rate at which a target polynucleotide re-binds to the polynucleotide binding site of a motor protein. Determining a suitable salt concentration to promote the debinding of a target polynucleotide from the polynucleotide binding site of a motor protein and / or to delay its re-binding is within the scope of the skill of those skilled in the art, given the disclosures herein.
[0330] In some embodiments, providing such conditions includes providing an osmotic pressure to increase the rate at which a target polynucleotide dissociates from the polynucleotide binding site of a motor protein. In some embodiments, providing such conditions includes providing an osmotic pressure to decrease the rate at which a target polynucleotide rebinds to the polynucleotide binding site of a motor protein. Determining a suitable osmotic pressure to promote the debinding of a target polynucleotide from the polynucleotide binding site of a motor protein and / or to delay its rebinding is within the capabilities of those skilled in the art, given the disclosures herein.
[0331] The method is typically carried out in the presence of a buffer. In the exemplary apparatus considered above, the buffer is present in an aqueous solution in the chamber. Any suitable buffer can be used. Typically, the buffer is HEPES. Another suitable buffer is Tris-HCl buffer. The method is typically carried out at pH values of 4.0–12.0, 4.5–10.0, 5.0–9.0, 5.5–8.8, 6.0–8.7, or 7.0–8.8, or 7.5–8.5. The pH used is preferably about 7.5.
[0332] The method can be carried out at temperatures of 0°C–100°C, 15°C–95°C, 16°C–90°C, 17°C–85°C, 18°C–80°C, 19°C–70°C, or 20°C–60°C. The method is typically carried out at room temperature. This method may optionally be carried out at a temperature that supports enzyme function, for example, around 37°C.
[0333] In some embodiments, providing the conditions involves increasing the temperature to increase the rate at which the target polynucleotide dissociates from the polynucleotide binding site of the motor protein. In some embodiments, providing the conditions involves increasing the temperature to decrease the rate at which the target polynucleotide re-binds to the polynucleotide binding site of the motor protein. While not bound by theory, the inventors believe that increasing the temperature may facilitate re-reading, for example, by increasing the rate at which the motor protein dissociates from the polynucleotide. Determining a suitable temperature to facilitate the de-binding of the target polynucleotide from the polynucleotide binding site of the motor protein and / or to delay re-binding is within the scope of the skill of those skilled in the art, given the disclosures herein.
[0334] Examples of providing conditions for promoting rereading by providing a temperature to promote rereading are provided herein, see, for example, Example 11. In some embodiments, providing conditions for promoting the debinding of a target polynucleotide from the polynucleotide binding site of a motor protein and / or delaying the rebinding of a target polynucleotide to the polynucleotide binding site of a motor protein may include providing a temperature of about 20°C to about 50°C, e.g., about 30°C to about 45°C, e.g., about 34°C to about 40°C, e.g., about 31, 32, 33, 34, 35, 36, 37, 38, or 39°C.
[0335] Further aspects of the methods to be disclosed The following are further aspects of the disclosed method: 1. A method for characterizing a target polynucleotide, (i) Contacting the detector with a target polynucleotide to which a motor protein is bound, wherein the target polynucleotide is bound to the motor protein at the polynucleotide binding site of the motor protein, (ii) Obtaining one or more characteristic measurements of the target polynucleotide when the motor protein controls the movement of the target polynucleotide in a first direction relative to the detector, (iii) Debinding the target polynucleotide from the polynucleotide binding site of the motor protein so that the target polynucleotide moves in a second direction relative to the detector, (iv) When the target polynucleotide is re-bound to the polynucleotide binding site of the motor protein and the motor protein controls the movement of the target polynucleotide in a first direction relative to the detector, one or more characteristic measurements of the target polynucleotide are obtained. A method comprising characterizing a target polynucleotide thereby.
[0336] 2. The method according to embodiment 1, comprising repeating steps (iii) and (iv) multiple times.
[0337] 3. The method according to embodiment 1 or 2, wherein in step (ii), a motor protein controls the movement of a first portion of the target polynucleotide in a first direction relative to the detector, and in step (iv), a motor protein controls the movement of a second portion of the target polynucleotide in a first direction relative to the detector, such that the first portion at least partially overlaps with the second portion.
[0338] 4. The method according to any one of the prior art, wherein the first part is the same as the second part.
[0339] 5. The method according to any one of the prior embodiments, wherein in step (iii), the distance the target polynucleotide travels relative to the detector is at least 100 nucleotides long.
[0340] 6. The method according to any one of the prior art, wherein the detector comprises a structure having a first aperture and a second aperture, or comprises a transmembrane nanopore having a first aperture and a second aperture, and step (i) comprises shrinking the first aperture with a target polynucleotide.
[0341] 7. The method according to embodiment 6, wherein (i) a motor protein controls the movement of a target polynucleotide from a second opening to a first opening, and (ii) when the target polynucleotide detaches from the polynucleotide binding site of the motor protein, the target polynucleotide moves from the first opening to the second opening.
[0342] 8. The method according to any one of the preceding embodiments, comprising applying a force to a detector, wherein a motor protein controls the movement of a target polynucleotide relative to the detector in the opposite direction to the applied force.
[0343] 9. The detector includes a transmembrane nanopore spanning a film having a cis side and a trans side. (i) The first opening of the nanopore is on the cis side of the membrane, the second opening of the nanopore is on the trans side, the motor protein controls the movement of the target polynucleotide through the nanopore from the trans side to the cis side of the membrane, and when the target polynucleotide detaches from the polynucleotide binding site of the motor protein, the target polynucleotide moves through the nanopore from the cis side to the trans side of the membrane, and (ii) The method according to any one of the prior embodiments, wherein the first opening of the nanopore is on the trans side of the membrane, the second opening of the nanopore is on the cis side, a motor protein controls the movement of a target polynucleotide through the nanopore from the cis side to the trans side of the membrane, and when the target polynucleotide detaches from the polynucleotide binding site of the motor protein, the target polynucleotide moves through the nanopore from the trans side to the cis side of the membrane.
[0344] 10. The method according to any one of the prior embodiments, wherein the target polynucleotide is attached to a leader configured to promote the detachment of the polynucleotide binding site of a motor protein from the target polynucleotide in the vicinity of the leader, or comprises a leader.
[0345] 11. The method according to embodiment 10, wherein the target polynucleotide is detached from the polynucleotide binding site of the motor protein when the motor protein comes into contact with the leader.
[0346] 12. The method according to embodiment 10 or embodiment 11, wherein the motor protein has a lower affinity for the leader than for the nucleotides of the target polynucleotide.
[0347] 13. The method according to any one of embodiments 10 to 12, wherein the leader comprises a nucleotide of a different type from the target polynucleotide.
[0348] 14. (i) The method according to any one of embodiments 10 to 13, wherein the target polynucleotide comprises deoxyribonucleotide (DNA) and the leader comprises one or more nucleotides lacking both nucleic acid bases and sugar moieties (spacer moieties), ribonucleotide (RNA), peptide nucleotide (PNA), glycerol nucleotide (GNA), threose nucleotide (TNA), locked nucleotide (LNA), cross-linked nucleotide (BNA), debased nucleotide, or nucleotide having a modified phosphate bond, or (ii) The method according to any one of embodiments 10 to 13, wherein the target polynucleotide comprises ribonucleotide (RNA) and the leader comprises nucleotides lacking both nucleic acid bases and sugar moieties (spacer moieties), deoxyribonucleotide (DNA), peptide nucleotide (PNA), glycerol nucleotide (GNA), threose nucleotide (TNA), locked nucleotide (LNA), cross-linked nucleotide (BNA), debased nucleotide, or nucleotide having a modified phosphate bond.
[0349] 15. The method according to any one of embodiments 10 to 14, wherein the target polynucleotide comprises a deoxyribonucleotide (DNA), and the leader comprises one or more spacer portions and / or one or more ribonucleotides.
[0350] 16. The method according to any one of the prior embodiments, wherein the target polynucleotide does not disassociate from the motor protein.
[0351] 17. The method according to any one of the prior embodiments, wherein a motor protein is modified to prevent a target polynucleotide from unassociating with the target polynucleotide.
[0352] 18. The method according to any one of the prior embodiments, wherein a motor protein is modified to promote the debinding of a target polynucleotide from the polynucleotide binding site of the motor protein, and / or delay the rebinding of the target polynucleotide to the polynucleotide binding site of the motor protein.
[0353] 19. The method according to any one of the preceding embodiments, wherein the motor protein is modified with a closure portion for (i) topologically closing the polynucleotide binding site of the motor protein around the target polynucleotide, and (ii) promoting the debinding of the target polynucleotide from the polynucleotide binding site of the motor protein and / or delaying the rebinding of the target polynucleotide to the polynucleotide binding site of the motor protein.
[0354] 20. The method according to embodiment 19, wherein the motor protein is modified to facilitate the binding of the closure portion to the motor protein.
[0355] 21. The method according to embodiment 20, wherein the motor protein is modified by substituting at least one amino acid in the motor protein with cysteine or a non-natural amino acid.
[0356] 22. The method according to any one of embodiments 19 to 21, wherein the closed portion contains a bifunctional crosslinking agent.
[0357] The method according to any one of embodiments 19 to 22, wherein the closed portion cross-links two amino acid residues of the motor protein, and at least one of the amino acids cross-linked by the closed portion is cysteine or a non-natural amino acid.
[0358] 24. The method according to any one of embodiments 19 to 23, wherein the closed portion has a length of about 1 Å to about 100 Å.
[0359] 25. The method according to any one of embodiments 19 to 21, wherein the closed portion includes a bond, preferably a disulfide bond.
[0360] 26. The method according to any one of embodiments 19 to 24, wherein the closed portion comprises the structure of formula [ABC], where A and C are independent reactive functional groups for reacting with amino acid residues in a motor protein, and B is a linking portion.
[0361] 27. The method according to embodiment 26, wherein A and C are each independently cysteine-reactive functional groups.
[0362] 28. The method according to embodiment 26 or 27, wherein the linking portion B includes a linear or branched unsubstituted or substituted alkylene, alkenylene, alkylylene, arylene, heteroarylene, carbocyclylene, or heterocyclylene portion, the portion being optionally interrupted or terminated by one or more atoms or groups selected from O, N(R), S, C(O), C(O)NR, C(O)O, unsubstituted or substituted arylene, arylene-alkylene, heteroarylene, heteroarylene-alkylene, carbocyclylene, carbocyclylene-alkylene, heterocyclylene, and heterocyclylene-alkylene, wherein R is selected from H, unsubstituted or substituted alkyl, and unsubstituted or substituted aryl.
[0363] 29. The method according to any one of embodiments 26 to 28, wherein the connecting portion B comprises an alkylene, oxyalkylene, or polyoxyalkylene group, and / or A and C are each maleimide groups.
[0364] 30. The method according to any one of embodiments 19-25 or 26-29, wherein the closed portion has a length of about 5 Å to about 50 Å.
[0365] 31. The method according to any one of the prior art, comprising providing conditions for promoting the debinding of a target polynucleotide from a polynucleotide binding site of a motor protein, and / or conditions for delaying the rebinding of a target polynucleotide to a polynucleotide binding site of a motor protein.
[0366] 32. The method according to embodiment 31, wherein providing the conditions includes increasing the temperature to increase the rate at which a target polynucleotide dissociates from the polynucleotide binding site of the motor protein.
[0367] 33. The method according to embodiment 31 or 32, wherein providing the conditions includes increasing the temperature to reduce the rate of rebinding of the target polynucleotide to the polynucleotide binding site of the motor protein.
[0368] 34. The method according to any one of the prior embodiments, wherein the motor protein is a helicase.
[0369] These embodiments relate to features that will be described in more detail herein.
[0370] Polypolymer adapter Polynucleotide adapters containing motor proteins are also provided. It will be understood that any of the polynucleotide adapters disclosed herein can be applied to embodiments of the methods discussed herein and above.
[0371] In one embodiment, a polynucleotide adapter is also provided having a first end and a second end including an attachment site for attaching to a double-stranded polynucleotide analyte, the polynucleotide adapter comprising (i) a motor protein stalled thereon in an orientation for processing the adapter in the direction of the attachment site, and (ii) a blocking portion located between the motor protein and the second end of the adapter.
[0372] In one embodiment, the polynucleotide adapter is a polynucleotide adapter as described in detail herein. In one embodiment, the motor protein is a motor protein as described herein. In one embodiment, the blocking portion is a blocking portion as described herein.
[0373] Motor proteins are oriented to process polynucleotide adapters toward attachment points on the adapter for binding to double-stranded polynucleotides. Motor proteins may also be oriented on polynucleotide adapters to control the movement of target polynucleotides from trans to cis.
[0374] The motor protein is oriented on a polynucleotide adapter to control the movement of the target polynucleotide toward a detector such as a nanopore, in the direction toward the motor protein, i.e., away from the detector, e.g., away from the nanopore, as described in more detail herein.
[0375] In some embodiments, the polynucleotide adapter includes a stall portion as described herein. In some embodiments, the polynucleotide adapter includes a stop portion as described herein.
[0376] kit Kits containing polynucleotide adapters and motor proteins are also provided. It will be understood that any of the polynucleotide adapters disclosed herein can be applied to embodiments of the kits discussed herein and above.
[0377] In one embodiment, a kit for modifying a target polynucleotide is provided, the kit comprising a first adapter provided in the present invention and a second adapter comprising a single-stranded leader sequence at a first end and an attachment site at a second end for attachment to a double-stranded polynucleotide analyte.
[0378] In one embodiment, the second adapter is an adapter described in detail herein.
[0379] system Systems comprising polynucleotide adapters, motor proteins, and nanopores are also provided. It will be understood that any of the polynucleotide adapters disclosed herein can be applied to embodiments of the systems discussed herein and above.
[0380] In one embodiment, a system for characterizing a target double-stranded polynucleotide is provided, the system is, - A stop portion and a polynucleotide adapter that optionally includes a stop portion, - Nanopores for characterizing the target polynucleotide when the target polynucleotide moves relative to the nanopore, -Includes a motor protein for moving double-stranded polynucleotides in a first direction relative to the nanopore.
[0381] In one embodiment, the polynucleotide adapter is a polynucleotide adapter as described in detail herein. In one embodiment, the motor protein is a motor protein as described herein. In one embodiment, the nanopore is a nanopore as described herein. The system may further include a membrane, a control device, and the like.
[0382] While specific embodiments, configurations, and materials and / or molecules are discussed herein in relation to the methods of the present invention, it should be understood that various changes or modifications of form and detail may be made without departing from the scope and spirit of the invention. The following examples are provided to better illustrate specific embodiments and should not be considered as limiting this application. This application is limited solely by the claims. [Examples]
[0383] Example 1 This embodiment demonstrates controlled movement of a DNA polynucleotide strand through a nanopore using a DNA motor that unwinds dsDNA while transferring the 5'-3' of ssDNA. The DNA motor initially stalled on a Y-adapter ligated to the polynucleotide. The polynucleotide moved through the nanopore in the following distinct stages: (1) an unenzymatic stage where the 3' end of the polynucleotide was captured by the nanopore, separating the double helix under a positive applied potential until the nanopore shifted and reached the DNA motor stalled at the 5' end; (2) a "de-stalled" stage where the DNA motor was initially unable to overcome the stall under a positive bias but was activated ("de-stalled") by applying a reverse potential; (3) a DNA motor-controlled stage where the motor began to move the DNA 5'-3' away from the nanopore against the applied potential; and (4) a certain block level was observed where, after reaching the end of the polynucleotide, the strand could be removed by reversing the potential to eject it.
[0384] Asymmetric 3.6 kilobase double-stranded DNA analytes (bacteriophage lambda DNA fragments; SEQ ID NO: 20) were obtained by PCR, and a 3'dA overhang was generated at one end and a 3'AGGA overhang at the other end using the NEBNext end repair and NEBNext dA-tailing module (New England Biolabs (NEB)) and USER digest.
[0385] The Y adapter was prepared by annealing DNA oligonucleotides (SEQ ID NOs. 21 and 22). A DNA motor (Dda helicase) was loaded into the adapter. Monomer traptabidine was added to the adapter and bound to the 5' biotin portion as a blocker to (1) prevent the DNA motor from diffusing in the reverse direction from the 5' end, and (2) prevent the 5' end of the library from being unintentionally captured by nanopores.
[0386] Double-stranded DNA analytes were ligated to the dA end of a Y adapter using LNB and T4 DNA ligase (NEB) from the Oxford Nanopore Technologies sequencing kit SKQ-LSK109 (also referred to herein as LSK-SQK109; see https: / / community.nanoporetech.com / protocols / gDNA-sqk-lsk109 / v / gde_9063_v109_revt_14aug2019 for details). Samples were purified using Agencourt AMPure XP (Beckman Coulter) beads and washed twice with LFB from the Oxford Nanopore Technologies sequencing kit (LSK-SQK109). The ligated substrates were eluted in 10 mM Tris-Cl, 50 mM NaCl (pH 8.0) to obtain a "DNA library".
[0387] Electrical measurements were obtained using Oxford Nanopore Technologies' FLO-MIN106 MinION flow cell and MinION Mk1b. A tether mix was obtained by adding a 50 nM DNA tether to 1200 μL of FB (from Oxford Nanopore Technologies sequencing kit (SQK-LSK109)). After running 800 μL of the tether mix through the system, a 5-minute wait was allowed, and then another 200 μL of the tether mix was run through the system with the SpotON port open. A "sequencing mix" was obtained by mixing 37.5 μL of SQB from Oxford Nanopore Technologies sequencing kit (SQK-LSK109), 15 μL of DNA library, 0.7 μL of excess monomer traptabidine (approximately 100 nM tetramer), and 22.5 μL of LB from Oxford Nanopore Technologies sequencing kit (SQK-LSK109). 75 μL of sequencing mix was added to the MinION flow cell via the SpotON flow cell port.
[0388] The custom sequencing script was prepared to control the applied potentials as follows: a 0.5-second destall phase (0mV), an 85.5-second sequencing phase (+120mV), and an efflux phase (0mV to -120mV, varying between 1 second and -120mV, for 3 seconds). This sequence of applied potentials was repeated multiple times.
[0389] Raw data was collected into bulk FAST5 files using MinKNOW software (Oxford Nanopore Technologies).
[0390] Figure 6 shows the adapter used in this embodiment. Figure 7 shows the adapter attached to a double-stranded polynucleotide analyte. Figure 8 shows a schematic diagram of the experiment in this embodiment, illustrating the pattern of applied potentials required to capture, desorb, and characterize the polynucleotide analyte. Figure 9 shows an example of current-vs-time traces in this embodiment. The data show that after the polynucleotide analyte is captured by the nanopore, followed by "destallation" of the DNA by lowering the applied potential to 0 to -120 mV, the DNA is controlled and moved stepwise from the nanopore. Few enzyme-mediated events were recorded above the destallation potential of -40 mV, suggesting that single-stranded DNA is retained in the nanopore during the destallation phase between 0 and -40 mV.
[0391] Example 2 This embodiment demonstrates the controlled movement of both strands of a DNA polynucleotide double helix through a nanopore using a DNA motor that unwinds the dsDNA while transferring the 5'-3' of the ssDNA. The DNA motor initially stalled on a Y-adapter ligated to the polynucleotide. The template and complementary strands were linked via a hairpin portion. The polynucleotides moved through the nanopore in the following distinct stages: (1) an enzyme-free stage where the 3' end of the polynucleotide is captured by the nanopore, the nanopore rearranges, and separates the double helix under a positive applied potential until it reaches a DNA motor stalled at the 5' end, with the complementary strand passing first and then the template strand; (2) a "de-stalled" stage where the DNA motor was initially unable to overcome the stall under a positive bias but was activated ("de-stalled") by applying a reverse potential; (3) a DNA motor-controlled stage where the motor begins to move the DNA 5'-3' away from the nanopore against the applied potential, with the DNA motor moving first along the template strand, passing over a hairpin, and then along the complementary strand; and (4) a certain block level was observed where, after reaching the end of the polynucleotide, it could be removed by reversing the potential to eject the strand.
[0392] Asymmetric 3.6 kilobase double-stranded DNA analytes (bacteriophage lambda DNA fragments; SEQ ID NO: 20) were obtained by PCR, and a 3'dA overhang was generated at one end and a 3'AGGA overhang at the other end using the NEBNext end repair and NEBNext dA-tailing module (New England Biolabs (NEB)) and USER digest.
[0393] The Y adapter was prepared by annealing DNA oligonucleotides (SEQ ID NOs: 21, SEQ ID NOs: 22). A DNA motor (Dda helicase) was loaded into the adapter. Monomer traptabidine was added to the adapter and bound to the 5' biotin portion as a blocker to (1) prevent the DNA motor from diffusing in the reverse direction from the 5' end, and (2) prevent the 5' end of the library from being unintentionally captured by nanopores.
[0394] Hairpins with a 3'-TCCT overhang were prepared by heating a DNA oligonucleotide (SEQ ID NO: 23) in double-stranded annealing buffer (Integrated DNA Technologies, Inc.) at 1 μM and 95°C for 2 minutes, followed by cooling with wet ice.
[0395] Double-stranded DNA analytes and hairpins were ligated to a Y-adapter using LNB and T4 DNA ligase (NEB) from the Oxford Nanopore Technologies sequencing kit (LSK-SQK109). Samples were purified using Agentcourt AMPure XP (Beckman Coulter) beads and washed twice with LFB from the Oxford Nanopore Technologies sequencing kit (LSK-SQK109). Ligated substrates were eluted in 10 mM Tris-Cl, 50 mM NaCl (pH 8.0) to obtain a "DNA library".
[0396] Electrical measurements were obtained using Oxford Nanopore Technologies' FLO-MIN106 MinION flow cell and MinION Mk1b. A tether mix was obtained by adding a 50 nM DNA tether to 1200 μL of FB (from Oxford Nanopore Technologies sequencing kit (SQK-LSK109)). After running 800 μL of the tether mix through the system, a 5-minute wait was allowed, and then another 200 μL of the tether mix was run through the system with the SpotON port open. A "sequencing mix" was obtained by mixing 37.5 μL of SQB from Oxford Nanopore Technologies sequencing kit (SQK-LSK109), 15 μL of DNA library, 0.7 μL of excess monomer traptabidine (approximately 100 nM tetramer), and 22.5 μL of LB from Oxford Nanopore Technologies sequencing kit (SQK-LSK109). 75 μL of sequencing mix was added to the MinION flow cell via the SpotON flow cell port.
[0397] The custom sequencing script was prepared to control the applied potential as follows: a 0.5-second destall phase (variable experimentally within the range of 0mV to -120mV), an 85.5-second sequencing phase (+120mV), and an efflux phase (0mV, 1 second, -120mV, 3 seconds). This series of applied potentials was repeated multiple times. Raw data was collected into bulk FAST5 files using MinKNOW software (Oxford Nanopore Technologies).
[0398] Figure 10 shows the components used in this embodiment: hairpin (A), adapter (B), and polynucleotide analyte (C), (D) all components ligated together. Figure 11 shows a schematic diagram of the experiment in this embodiment, showing the pattern of applied potentials required to capture, destallize, and characterize the hairpin-derivativeized polynucleotide analyte. Figure 12a shows some examples of current-vs-time traces in this embodiment. The data show the capture of the polynucleotide analyte by the nanopore and the subsequent controlled stepwise movement of the DNA from the nanopore after "destallization". The destallization potential was varied between 0mV and -120mV, but no enzyme-mediated events were observed above -60mV, which suggests that the hairpin folded in the transcompartment encounters resistance to efflux during destallization up to a potential of -60mV and additional resistance compared to single-stranded DNA alone (from Example 1). Figure 12b assigns states A-G to the current trace examples in Figure 11. Compared to Example 1, an additional state E is observed in Figure 12b, which may be due to the enzymatic transfer of the complementary portion of the polynucleotide from the nanopore immediately following the template portion D.
[0399] Example 3 This embodiment demonstrates controlled destallation of a DNA motor via an "active destallation" process. One or both strands of a double helix of DNA polynucleotides were passed through a nanopore using a DNA motor that unwinds dsDNA. The DNA motor first stalled on a Y-adapter ligated to the polynucleotide. Optionally, the template and complementary strands were linked at the distal end of the polynucleotide via a hairpin portion; otherwise, the template and complement strands were not linked, omitting the hairpin. The polynucleotides moved through the nanopore in the following distinct stages: (1) an enzyme-free stage where the 3' end of the polynucleotide was captured by the nanopore, separating the double helix until the nanopore rearranged and reached a DNA motor stalled at the 5' end; (2) an activated "de-stalled" stage where the DNA motor was initially unable to overcome the stall under a positive bias but was activated ("de-stalled") by repeatedly applying an efflux potential followed by a return to the sequencing potential; (3) a DNA motor-controlled stage where the motor began moving the DNA 5'-3' away from the nanopore against the applied potential; and (4) a certain block level was observed after reaching the end of the polynucleotide, which could be removed by reversing the potential that effluxes the strand.
[0400] Asymmetric 3.6 kilobase double-stranded DNA analytes (bacteriophage lambda DNA fragments; SEQ ID NO: 20) were obtained by PCR, and a 3'dA overhang was generated at one end and a 3'AGGA overhang at the other end using the NEBNext end repair and NEBNext dA-tailing module (New England Biolabs (NEB)) and USER digest.
[0401] Symmetrical 3.6 kilobase double-stranded DNA analytes (bacteriophage lambda DNA fragments; SEQ ID NO: 20) were obtained by PCR, and 3'dA overhangs were generated at both ends using the NEBNext end repair and NEBNext dA-tailing module (New England Biolabs (NEB)).
[0402] The Y adapter was prepared by annealing DNA oligonucleotides (SEQ ID NOs. 21 and 22). A DNA motor (Dda helicase) was loaded into the adapter. Monomer traptabidine was added to the adapter and bound to the 5' biotin portion as a blocker to (1) prevent the DNA motor from diffusing in the reverse direction from the 5' end, and (2) prevent the 5' end of the library from being unintentionally captured by nanopores.
[0403] Hairpins with a 3'-TCCT overhang were prepared by heating a DNA oligonucleotide (SEQ ID NO: 23) in double-stranded annealing buffer (Integrated DNA Technologies, Inc.) at 1 μM and 95°C for 2 minutes, followed by cooling with wet ice.
[0404] Double-stranded DNA analytes were ligated to a Y-adapter using LNB and T4 DNA ligase (NEB) from the Oxford Nanopore Technologies sequencing kit (LSK-SQK109). Samples were purified using Agentcourt AMPure XP (Beckman Coulter) beads and washed twice with LFB from the Oxford Nanopore Technologies sequencing kit (LSK-SQK109). Ligated substrates were eluted in 10 mM Tris-Cl, 50 mM NaCl (pH 8.0) to obtain a "1D DNA library".
[0405] Asymmetric double-stranded DNA analytes were ligated to Y-adapters and hairpins using the Oxford Nanopore Technologies sequencing kit (LSK-SQK109) and T4 DNA ligase (NEB) LNB. The samples were purified using Agentcourt AMPure XP (Beckman Coulter) beads and washed twice with Oxford Nanopore Technologies sequencing kit (LSK-SQK109) LFB. The ligated substrates were eluted in 10 mM Tris-Cl, 50 mM NaCl (pH 8.0) to obtain a "2D DNA library".
[0406] Electrical measurements were obtained using Oxford Nanopore Technologies' FLO-MIN106 MinION flow cell and MinION Mk1b. A tether mix was obtained by adding a 50 nM DNA tether to 1200 μL of FB (from Oxford Nanopore Technologies sequencing kit (SQK-LSK109)). After running 800 μL of the tether mix through the system, a 5-minute wait was allowed, and then another 200 μL of the tether mix was run through the system with the SpotON port open. A "sequencing mix" was obtained by mixing 37.5 μL of SQB from Oxford Nanopore Technologies sequencing kit (SQK-LSK109), 15 μL of a 1D or 2D DNA library, 0.7 μL of excess monomer traptabidine (approximately 100 nM tetramer), and 22.5 μL of LB from Oxford Nanopore Technologies sequencing kit (SQK-LSK109). 75 μL of sequencing mix was added to the MinION flow cell via the SpotON flow cell port.
[0407] A custom sequence script was prepared to control the applied potential using MinION's active unblocking circuit. The sequencing voltage was set to 120mV, and the active unblocking potential ("active destabilization" stage, stage (2) above) was set to -12mV for 1D libraries and -48mV for 2D libraries. The classification of stalled levels and chain (sequencing) levels was programmed into the configuration file of the MinKNOW instrument control software to enable detection of stalled species, and the knowledge of static unblocking potentials from Examples 1 and 2 was used to apply an unblocking potential that would not cause complete chain release. The script was made to function as follows: If MinKNOW detects that a chain is at a stalled level, it first applies the unblocking potential for 5 seconds, then returns to the 120mV sequencing potential to actively check the sequencing chain 5 times. If a stalled level is still present, it applies the unblocking potential for another 25 seconds and repeats 5 times. A 3-second rest period was incorporated between each unblocking attempt. If MinKNOW detects an active sequencing strand when it returns a sequencing potential, it stops the unblocking attempt and applies only the sequencing potential. If no active sequencing strand is generated throughout this process, MinKNOW turns off the channel. Every 15 minutes, a "mux scan" is applied to reset the system, unblocking all channels in the flow cell globally and checking for active nanopores at 120mV.
[0408] Raw data was collected into bulk FAST5 files using MinKNOW software (Oxford Nanopore Technologies).
[0409] Figures 7 and 10, D show the polynucleotide analytes used in this embodiment. Their preparation is described in Examples 1 and 2. Figure 13 shows examples of current traces of a 1D DNA library (A) and a 2D DNA library (B). Regions where destallization was attempted are marked with asterisks. The data demonstrate that both 1D and 2D libraries can be destallized using these methods, and that several attempts can be repeated to destallize the enzyme and confirm the enzymatically controlled transfer of polynucleotides from the nanopore.
[0410] Example 4 This example demonstrates a method for estimating the size of a double-stranded DNA molecule when the template and complementary strands are linked by a hairpin before the terminal 5'-3' DNA motor actively moves the DNA strands away from the nanopore in opposite directions, using the signal duration from the initial enzyme-free portion (3'-5') of the DNA transposition via the nanopore. In addition, this example demonstrates a method for distinguishing signals using markers added to the hairpin.
[0411] The DNA motor initially stalled on the Y-adapter linked to the polynucleotide. The template and complementary strands were linked together via a hairpin portion according to Example 2. Optionally, the hairpin portion contained bulky fluorophore groups or debasing groups and / or hybridized additional oligonucleotides to the hairpin.
[0412] An asymmetric 3.6 kilobase double-stranded DNA analyte (bacteriophage lambda DNA fragment; SEQ ID NO: 20) was obtained by primer-assisted PCR, one of which contained multiple dUTP bases. End repair and dA tailing were performed using the NEBNext end repair and NEBNext dA tailing module (New England Biolabs (NEB)), followed by NEBUSER digest, generating a 3' dA overhang at one end and a 3' AGGA overhang at the other end.
[0413] Random libraries of Escherichia coli double-stranded DNA were generated by ligating generic adapters to E. coli SCS110 DNA, which had been sheared to approximately 20kb in size using Covaris gTubes, and amplifying by PCR. The fragments were then repaired and tailed using the NEBNext End Repair and NEBNext dA Tailing Module (New England Biolabs (NEB)) to generate 3'dA overhangs at both ends.
[0414] The Y adapter was prepared by annealing DNA oligonucleotides (SEQ ID NOs. 21 and 22). A DNA motor (Dda helicase) was loaded into the adapter. Monomer traptabidine was added to the adapter and bound to the 5' biotin portion as a blocker to (1) prevent the DNA motor from diffusing in the reverse direction from the 5' end, and (2) prevent the 5' end of the library from being unintentionally captured by nanopores.
[0415] Hairpins with 3'-TCCT or 3'-T overhangs were prepared by heating 1 μM of DNA from SEQ ID NO: 24, SEQ ID NO: 25, or SEQ ID NO: 26 in duplex annealing buffer (Integrated DNA Technologies, Inc.) at 95°C for 2 minutes, followed by quenching on wet ice.
[0416] Asymmetric 3.6 kb double-stranded DNA analytes and hairpins (SEQ ID NO: 24 or 26) were ligated to a Y-adapter using LNB and T4 DNA ligase (NEB) from the Oxford Nanopore Technologies sequencing kit (LSK-SQK109). Samples were purified using Agentcourt AMPure XP (Beckman Coulter) beads and washed twice with LFB from the Oxford Nanopore Technologies sequencing kit (LSK-SQK109). The ligated substrates were eluted in 10 mM Tris-Cl, 50 mM NaCl (pH 8.0) to obtain a "3.6 kb DNA library".
[0417] Using LNB and T4 DNA ligase (NEB) from the Oxford Nanopore Technologies sequencing kit (LSK-SQK109), Escherichia coli double-stranded DNA and hairpin (SEQ ID NO: 25) were ligated to a Y-adapter. The samples were purified using Agencourt AMPure XP (Beckman Coulter) beads and washed twice with LFB from the Oxford Nanopore Technologies sequencing kit (LSK-SQK109). The ligated substrates were eluted in 10 mM Tris-Cl, 50 mM NaCl (pH 8.0) to obtain a "random E. coli test library".
[0418] Electrical measurements were obtained using Oxford Nanopore Technologies' FLO-MIN106 MinION flow cell and MinION Mk1b. A tether mix was obtained by adding a 50 nM DNA tether to 1200 μL of FB (from Oxford Nanopore Technologies sequencing kit (SQK-LSK109)). After running 800 μL of the tether mix through the system, a 5-minute wait was allowed, and then another 200 μL of the tether mix was run through the system with the SpotON port open. A "sequencing mix" was obtained by mixing 37.5 μL of SQB from Oxford Nanopore Technologies sequencing kit (SQK-LSK109), 15 μL of either a 3.6 kb library or a random E. coli test library, 0.7 μL of excess monomer traptabidine (approximately 100 nM tetramer), and 22.5 μL of LB from Oxford Nanopore Technologies sequencing kit (SQK-LSK109). A portion of the reaction solution was also mixed with the oligonucleotide of SEQ ID NO: 27 at a concentration of 50 nM. 75 μL of the sequencing mix was added to the MinION flow cell via the SpotON flow cell port.
[0419] Two libraries were tested with different execution scripts. The 3.6kb library was run using a custom sequencing script to control the applied potentials as follows: 0.5-second destall phase (-40mV), 85.5-second sequencing phase (+120mV), and efflux phase (0mV, 1 second, -120mV, 3 seconds). This sequence of applied potentials was repeated multiple times. The random E. coli test library was run using the custom active destall phase script described in Example 3, with a capture / sequencing voltage of 120mV and an efflux voltage of -48mV.
[0420] Raw data was collected into bulk FAST5 files using MinKNOW software (Oxford Nanopore Technologies).
[0421] Figure 14 shows the hairpin-oligonucleotide combinations used in this embodiment. A 3.6kb DNA library was used to initially characterize the capture step signal. Figure 15 shows a schematic diagram of the intermediates expected to be detected by electrical measurements of enzyme-free and enzyme-mediated transfer. Compared to Figure 11, two additional states A1 and A2 (see Figure 15), corresponding to the bulky group in the nanopore and the blocker oligonucleotide on the nanopore, respectively, are expected during the initial enzyme-free capture. Additional state D1 corresponds to the enzyme moving via the bulky group of the hairpin portion and is expected between the template (D) and complementary chain (E) steps of enzyme-mediated transfer. Figures 16a–16d show examples of traces for each hairpin-oligonucleotide combination. The hairpin-only portion (Figure 16a) showed a relatively flat but detectable capture step (marked with an asterisk). The addition of a hybridized oligonucleotide to the hairpin introduced an additional uptick intermediate (indicated as A2 in Figure 16b), and three bulky fluorescein-dT bases introduced a downtick (indicated as A1 in Figure 16c). The combination of the hybridized oligonucleotide and fluorescein-dT base to the hairpin introduced both types of signals (shown in Figure 16d). The introduction of the additional signals made it possible to measure the duration of the enzyme-free capture / entry phase of the polynucleotide (indicated by asterisks in Figures 16a-d).
[0422] The enzyme-free capture stage of a random E. coli test library was measured using an example with the scheme shown in Figure 16b (hairpin and hybridized oligonucleotide) (Figure 16e). Figure 16e,i shows the simplified (event-fitted) raw data for four examples. The enzyme-free capture period between states A and A2, indicated by asterisks, was measured using a threshold of 60 pA. Figure 16e,ii shows the duration of enzyme-mediated transposition plotted against the capture period of 30 molecules. Linear regression analysis shows that the enzyme-free capture period correlates with the duration of enzyme-mediated chain, confirming that this method can be used to estimate chain size before sequencing.
[0423] Example 5 This embodiment demonstrates the controlled movement of a DNA polynucleotide chain through a nanopore using a DNA motor that unwinds dsDNA while transferring the 5'-3' of ssDNA. This embodiment describes an adapter configuration that replaces that described in a previous embodiment. The DNA motor initially stalled on a Y adapter ligated to the polynucleotide. The polynucleotide moved through the nanopore in the following distinct stages: (1) an enzyme-free stage where the 3' end of the polynucleotide was captured by the nanopore, the nanopore rearranged, and the double helix was separated under a positive applied potential until it reached the DNA motor stalled at the 5' end; (2) a "de-stalled" stage where the DNA motor was initially unable to overcome the stall under a positive bias but was activated ("de-stalled") by applying a reverse potential; (3) a DNA motor-controlled stage where the motor began to move the DNA 5'-3' away from the nanopore against the applied potential; and (4) a certain block level was observed where, after reaching the end of the polynucleotide, the chain could be removed by reversing the potential to eject it.
[0424] The Y adapter was prepared by annealing DNA oligonucleotides (SEQ ID NOs. 28, 29, 30, and 31). A DNA motor (Dda helicase) was loaded into the adapter. Compared to previous examples, oligonucleotide SEQ ID NO 31 replaced the function of the biotin-streptavidin complex; that is, the oligonucleotide formed a double-stranded region behind the enzyme, preventing the enzyme from diffusing in the reverse direction from the 5' end of the enzyme-loaded strand and preventing the 5' end strand from being trapped by nanopores. Oligonucleotide SEQ ID NO 30 acted as a forward blocker, stalling the enzyme in solution. A schematic diagram of this adapter is shown in Figure 17a.
[0425] Symmetrical 3.6 kilobase double-stranded DNA analytes (bacteriophage lambda DNA fragments; SEQ ID NO: 20) were obtained by PCR, and 3'dA overhangs were generated at both ends using the NEBNext end repair and NEBNext dA-tailing module (New England Biolabs (NEB)).
[0426] Double-stranded DNA analytes were ligated to the dA end of a Y-adapter using LNB and T4 DNA Ligase (NEB) from the Oxford Nanopore Technologies sequencing kit (LSK-SQK109). The samples were purified using Agencourt AMPure XP (Beckman Coulter) beads and washed twice with LFB from the Oxford Nanopore Technologies sequencing kit (LSK-SQK109). The ligated substrates were eluted in 10 mM Tris-Cl, 50 mM NaCl (pH 8.0) to obtain a "DNA library". A schematic diagram of the library is shown in Figure 17b.
[0427] Electrical measurements were obtained using the Oxford Nanopore Technologies FLO-MIN106 MinION flow cell and MinION Mk1b. 1170 μL of FB was mixed with 30 μL of FLT (from the Oxford Nanopore Technologies sequencing kit (SQK-LSK109)) to obtain a tether mix. After flowing 800 μL of the tether mix through the system, a 5-minute wait was allowed, and then another 200 μL of the tether mix was flowed through the system with the SpotON port open. A "sequencing mix" was obtained by mixing 37.5 μL SQB, 15 μL of DNA library, and 22.5 μL of LB from the Oxford Nanopore Technologies sequencing kit (SQK-LSK109). 75 μL of the sequencing mix was added to the MinION flow cell via the SpotON flow cell port.
[0428] Enzyme destallation was controlled using the custom sequencing script described in Example 3, with a sequencing voltage of 120 mV and an efflux voltage of 12 mV. Raw data were collected in bulk FAST5 files using MinKNOW software (Oxford Nanopore Technologies).
[0429] Figure 17c shows a schematic diagram of the intermediate steps expected to occur between polynucleotide capture and enzyme destallation. Compared to Example 1, an additional intermediate (blocker chain on the nanopore, followed by removal of the blocker by the nanopore; state B in Figure 17c) is expected. Figure 17d shows a representative current-time trace (i). The area enclosed by the square (ii) corresponds to the capture / entry stage. In the example shown, the enzyme was destalled in a second 5-second destallation attempt (D), and the enzyme controlled the movement of the polynucleotide from the nanopore during E (magnified in iii). The data show that (a) the function of the biotin-traptoavidin binding described in Example 1 can be replaced with an oligonucleotide "backblocker", and (b) the enzyme blocker oligonucleotide can exist as a separate fragment removed by the nanopore.
[0430] Example 6 This embodiment demonstrates controlled movement of a DNA polynucleotide chain through a nanopore using a DNA motor that unwinds the dsDNA while transferring the 5'-3' of the ssDNA. The DNA motor initially stalled on a Y-adapter ligated to the polynucleotide. Compared to the previous embodiment, the Y-adapter contained an oligonucleotide with a leader having 30 3' terminal C3 spacer residues. The polynucleotide moved through the nanopore in the following distinct stages: (1) an enzyme-free stage where the 3' end of the polynucleotide was captured by the nanopore, separating the double helix under a positive applied potential until the nanopore rearranged and reached a DNA motor stalled at the 5' end; (2) a "de-stalled" stage where the DNA motor was initially unable to overcome the stall under a positive bias but was activated ("de-stalled") by applying a reverse potential; (3) a DNA motor-controlled stage where the motor began to move the DNA 5'-3' away from the nanopore against the applied potential; (4) upon reaching the end of the polynucleotide, a distinct block level was observed that was clearly different from the poly(dT) level in the previous example, which could be removed by reversing the possibility of strand ejection, as well as occasionally, (5) under the force of the applied sequencing potential, the enzyme spontaneously slipped backward and rejoined the upstream DNA, repeating from step (3).
[0431] The Y adapter was prepared by annealing DNA oligonucleotides (SEQ ID NO: 28, SEQ ID NO: 33, SEQ ID NO: 30, and SEQ ID NO: 32). A DNA motor (Dda helicase) was loaded into the adapter. The oligonucleotide of SEQ ID NO: 33 contained the C3 spacer residue described above.
[0432] A DNA library of seven fragments was obtained by digesting bacteriophage lambda DNA using SnaBI and BamHI restriction enzymes. End repair and dA tailing were performed using the NEBNext end repair and NEBNextdA tailing module (New England Biolabs (NEB)), generating 3'dA overhangs at both ends of each fragment.
[0433] A DNA library of seven fragments was ligated to the dA end of a Y-adapter using the LNB and T4 DNA ligase (NEB) from the Oxford Nanopore Technologies sequencing kit (LSK-SQK109). The samples were purified using Agentcourt AMPure XP (Beckman Coulter) beads and washed twice with the LFB from the Oxford Nanopore Technologies sequencing kit (LSK-SQK109). The ligated substrates were eluted in 10 mM Tris-Cl, 50 mM NaCl (pH 8.0) to obtain the "DNA library".
[0434] Electrical measurements were obtained using the Oxford Nanopore Technologies FLO-MIN106 MinION flow cell and MinION Mk1b. 1170 μL of FB (from the Oxford Nanopore Technologies sequencing kit (SQK-LSK109)) was mixed with 30 μL of FLT to obtain a tether mix. After flowing 800 μL of the tether mix through the system, a 5-minute wait was followed by flowing another 200 μL of tether mix through the system with the SpotON port open. A "sequencing mix" was obtained by mixing 37.5 μL SQB, 15 μL of DNA library, and 22.5 μL of LB from the Oxford Nanopore Technologies sequencing kit (SQK-LSK109). 75 μL of the sequencing mix was added to the MinION flow cell via the SpotON flow cell port.
[0435] The DNA library was run using a custom sequencing script that controlled the applied potential as follows: a 5-second destall phase (-20mV), a 55-second sequencing phase (+120mV), and an efflux phase (0mV, 1 second, -120mV, 3 seconds). This sequence of applied potentials was repeated multiple times. Raw data were collected into bulk FAST5 files using MinKNOW software (Oxford Nanopore Technologies).
[0436] Figure 18a shows a schematic diagram of the experiment. Compared to Figure 17c, this experiment introduces an additional “re-read” step (RR), where the enzyme debonds and returns to its previous position on the DNA strand from the 3’C3 (non-DNA) reader (E), then rearranges again from 5’ to 3’, resulting in multiple reads of the same DNA strand. No open pore levels were observed during the re-reads, meaning it is unlikely that molecules were ejected from the nanopore. Figure 18b shows an example of current-time traces of molecules read twice (i and ii). A hidden Markov model was trained to map the enzyme-controlled region to the reference for each restriction enzyme-treated fragment (Figure 18c). The data showed that reads mapped to the same fragment in the reference, with some examples recorded where some were mapped two or three times, confirming that the strand was read multiple times.
[0437] Example 7 This embodiment illustrates how the rate of dislocation of a DNA polynucleotide chain passing through a nanopore can be controlled using an applied voltage, by using a DNA motor that unwinds dsDNA while dislocating it in the 5'-3' direction on ssDNA in the opposite direction to the force applied to the DNA by the electric field.
[0438] The Y adapter was prepared by annealing DNA oligonucleotides (SEQ ID NOs. 28, 33, 30, and 32). A DNA motor (Dda helicase) was loaded into the adapter.
[0439] A library of seven bacteriophage lambda fragments was prepared according to Example 6. The libraries were ligated to the dA end of a Y adapter using the LNB and T4 DNA ligase (NEB) of the Oxford Nanopore Technologies sequencing kit (LSK-SQK109). The samples were purified using Agencourt AMPure XP (Beckman Coulter) beads and washed twice with the LFB of the Oxford Nanopore Technologies sequencing kit (LSK-SQK109). The ligated substrates were eluted in 10 mM Tris-Cl, 50 mM NaCl (pH 8.0) to obtain a "DNA library".
[0440] Electrical measurements were obtained using the Oxford Nanopore Technologies FLO-MIN106 MinION flow cell and MinION Mk1b. 1170 μL of FB (from the Oxford Nanopore Technologies sequencing kit (SQK-LSK109)) was mixed with 30 μL of FLT to obtain a tether mix. After flowing 800 μL of the tether mix through the system, a 5-minute wait was followed by flowing another 200 μL of tether mix through the system with the SpotON port open. A "sequencing mix" was obtained by mixing 37.5 μL SQB, 15 μL of DNA library, and 22.5 μL of LB from the Oxford Nanopore Technologies sequencing kit (SQK-LSK109). 75 μL of the sequencing mix was added to the MinION flow cell via the SpotON flow cell port.
[0441] The DNA library was run using a custom sequencing script that controlled the applied potential as follows: a 55-second capture phase (+120 to +200 mV), a 5-second discharging phase (-20 mV), a 55-second sequencing phase (+120 mV), and an efflux phase (0 mV, 1 second, -120 mV, 3 seconds). This series of applied potentials was repeated multiple times. Raw data were collected into bulk FAST5 files using MinKNOW software (Oxford Nanopore Technologies).
[0442] The experimental scheme is shown in Figure 17c, but in this example, the capture / sequencing voltage was varied between 120mV and 200mV. The data were mapped using the HMM model described in Example 6. Figures 19a–d show the HMM mapping for 16 example reads of data collected at 120mV, 140mV, and 160mV. The mapping was used to estimate the enzyme rate during the enzyme-controlled rearrangement step. At 120mV, the median of the enzyme rate was 319 bp / s, at 140mV it was 259 bp / s, and at 160mV it was 196 bp / s. The data show that, theoretically, the enzyme rate can be reduced to zero by increasing the applied potential.
[0443] Example 8 This embodiment demonstrates a method for estimating the size of one strand of a double-stranded DNA molecule before fully characterizing it, based solely on the capture / entry phase, using the duration of the signal from the initial, enzyme-free portion (3'-5') of the DNA transposition through a nanopore.
[0444] The Y adapter was prepared by annealing DNA oligonucleotides (SEQ ID NOs. 28, 33, 30, and 32). A DNA motor (Dda helicase) was loaded into the adapter.
[0445] A 10kb fragment was obtained from bacteriophage lambda by PCR. Commercially available bacteriophage lambda DNA (approximately 48kb) and T4 DNA (approximately 169kb) were obtained. These double-stranded analytes were repaired and dA-tailed using the NEBNext end repair and NEBNext dA tailing module (New England Biolabs (NEB)) to generate 3'dA overhangs at both ends of each fragment. Each sample was ligated (separately) to the dA end of a Y-adapter using LNB and T4 DNA ligase (NEB) from the Oxford Nanopore Technologies sequencing kit (LSK-SQK109). The samples were purified using Agencourt AMPure XP (Beckman Coulter) beads and washed twice with LFB from the Oxford Nanopore Technologies sequencing kit (LSK-SQK109). The ligated substrates were eluted with 10 mM Tris-Cl and 50 mM NaCl (pH 8.0) to obtain a "10kb library," a "lambda library," and a "T4 library."
[0446] Electrical measurements were obtained using the Oxford Nanopore Technologies FLO-MIN106 MinION flow cell and MinION Mk1b. 1170 μL of FB (from the Oxford Nanopore Technologies sequencing kit (SQK-LSK109)) was mixed with 30 μL of FLT to obtain a tether mix. After flowing 800 μL of the tether mix through the system, a 5-minute wait was followed by flowing another 200 μL of tether mix through the system with the SpotON port open. A "sequencing mix" was obtained by mixing 37.5 μL SQB, 15 μL of DNA library, and 22.5 μL of LB from the Oxford Nanopore Technologies sequencing kit (SQK-LSK109). 75 μL of the sequencing mix was added to the MinION flow cell via the SpotON flow cell port.
[0447] Data were collected at a capture / arrangement voltage of 120 mV using a custom script similar to that described in Example 3.
[0448] Figure 20a shows an experimental diagram similar to that of Example 5 (Figure 17c) above. The enzyme-free capture phase was manually measured as the asterisked period between the pore level (A) and the stall level (C), and is shown in detail in the lower panel of Figure 20b. The capture phase can be identified by its distinct noise and median current level characteristics. The enzyme-mediated rearrangement time (E) was also measured. Figure 20b shows representative current-time traces obtained in separate flow cells for each of the three libraries described above. For example, the 10kb library had an enzyme-free capture time of 1.6 seconds and an enzyme-mediated rearrangement time of 35.3 seconds. The T4 library yielded longer captures, but no examples of full length were recorded, which was thought to be likely due to an increased possibility of nicks entering the chain. Figure 20c shows a logarithmic plot of the capture periods (A-C) versus the enzyme-mediated rearrangement periods. Linear correlations (R) were observed from 31 examples. 2 A value of 0.74 was obtained, confirming that it is possible to estimate the chain size using this method before decoding the sequence.
[0449] Example 9 This example demonstrates a method for rereading a native DNA analyte multiple times using motor proteins with different disulfide closure linker lengths.
[0450] Y adapters with leader arms containing 30 C3 spacer units were prepared by annealing four DNA oligonucleotides having the sequences of SEQ ID NOs. 67, 68, 69, and 70. A DNA motor (Dda helicase) was loaded into each adapter, and the disulfide bonds were closed by reaction with one of the following linkers: diamide (TMAD), BMOE (1,2-bismaleimide ethane), BMOP (1,3-bismaleimide propane), BMB (1,4-bismaleimide butane), BM(PEG)2 (1,8-bismaleimide-diethylene glycol), or BM(PEG)3 (1,11-bismaleimide-triethylene glycol).
[0451] E. coli K12 PCR DNA was obtained by extraction from E. coli cells using the Qiagen Genome Chip Kit, sheared to approximately 10kb cutoffs using Covaris gTube, repaired the ends and added dA tails using the Ultra II End Repair and dA-Taling Kit (New England Biolabs), ligated to a PCR adapter (PCA; Oxford Nanopore Technologies), and PCR amplified using LongAmp Taq. The resulting double-stranded analytes were repaired and dA-tailed using the NEBNext End Repair and NEBNext dA-Taling Module (New England Biolabs (NEB)) to generate 3'dA overhangs at both ends of each fragment. The samples were ligated to the T overhangs of the Y adapter using LNB and T4 DNA ligase from the Oxford Nanopore Technologies sequencing kit (LSK-SQK109). The samples were purified using Agencourt AMPure XP (Beckman Coulter) beads and washed twice with LFB from the Oxford Nanopore Technologies sequencing kit (LSK-SQK109). The ligated substrates were eluted with elution buffer (EB) from the same kit to obtain a "DNA library". The DNA libraries were prepared separately using an adapter carrying Dda helicase closed with a disulfide linker, as described above.
[0452] Electrodynamic measurements were obtained using a custom MinION flow cell from Oxford Nanopore Technologies with inserted CsgG nanopores and a MinION Mk1b. 30 μL of FLT was added to 1170 μL of FB (from the Oxford Nanopore Technologies sequencing kit (SQK-LSK109)) to obtain a tether mix. After flowing 800 μL of the tether mix through the system, a 5-minute wait was allowed, and then another 200 μL of the tether mix was flowed through the system with the SpotON port open. A "sequencing mix" was obtained by mixing 37.5 μL SQB, 15 μL of DNA library, and 22.5 μL of LB from the Oxford Nanopore Technologies sequencing kit (SQK-LSK109). 75 μL of the sequencing mix was added to the MinION flow cell via the SpotON flow cell port.
[0453] Using a custom script similar to that described in Example 3, data was collected at a capture / sequencing voltage of 180 mV, however, the motor protein was destalled by switching the voltage to zero by cleaving the channel for 5 seconds when an enzyme stall level was detected. Activated deblocking was configured to be triggered when a block level unrelated to the terminal C3 level, chain, open pore, or enzyme stall level was recognized.
[0454] Chain-level events from single-channel data occurring immediately after the C3 level (marked "C3" in Figure 18b) were recorded as potential rereads (e.g., labeled "ii" in Figure 18b). These rereads were identified by comparing the base calling and reread sequences occurring after open pore and destallation events with the original read (e.g., labeled "i" in Figure 18b), as described in Example 6. Events in the same reading direction and within the range of the original read were classified as rereads. Reread efficiency was quantified in two ways: (i) the percentage of reads that drop back and are reread within 30 seconds of reaching the C3 reader, and (ii) the dropback distance, which is the length of the reread, i.e., the distance the enzyme is pushed back from the C3 reader.
[0455] The table below shows the results of this experiment. The results show rereads in all linkers tested, indicating that increasing the linker length increased the percentage of reads with rereads within 30 seconds of reaching the C3 reader. [Table 4]
[0456] Example 10 This embodiment demonstrates a method for reading a native DNA analyte multiple times using an adapter with different reader sequences that the Dda helicase encounters at the 3' end of the sequenced strand.
[0457] Y adapters having leader arms with RNA or C3 leader chemistry were prepared by annealing four DNA oligonucleotides having the sequences of SEQ ID NOs. 67, 68, and 69, and a leader oligonucleotide selected from SEQ ID NOs. 70, 71, and 72. A DNA motor (Dda helicase) was loaded into each adapter, and the disulfide was closed via reaction with 1,2-bismaleimide ethane (BMOE).
[0458] The DNA library was prepared by ligating the above-mentioned Y adapter to E. coli DNA prepared as described in Example 9.
[0459] Electrodynamic measurements were obtained using a custom MinION flow cell from Oxford Nanopore Technologies with inserted CsgG nanopores and a MinION Mk1b. 30 μL of FLT was added to 1170 μL of FB (from the Oxford Nanopore Technologies sequencing kit (SQK-LSK109)) to obtain a tether mix. After flowing 800 μL of the tether mix through the system, a 5-minute wait was allowed, and then another 200 μL of the tether mix was flowed through the system with the SpotON port open. A "sequencing mix" was obtained by mixing 37.5 μL SQB, 15 μL of DNA library, and 22.5 μL of LB from the Oxford Nanopore Technologies sequencing kit (SQK-LSK109). 75 μL of the sequencing mix was added to the MinION flow cell via the SpotON flow cell port.
[0460] Using a custom script similar to that described in Example 3, data was collected at a capture / sequencing voltage of 180 mV, however, the motor protein was destalled by switching the voltage to zero by cleaving the channel for 5 seconds when an enzyme stall level was detected. Activated deblocking was configured to be triggered when a block level unrelated to the terminal C3 level, chain, open pore, or enzyme stall level was recognized.
[0461] Reread events were scored according to Example 9, with a few exceptions, and if the leader contained RNA, rereads occurred at the strand level.
[0462] The table below shows the results of this experiment. The results show rereading with all the reader oligonucleotides tested, and indicate that optimal rereading efficiency was obtained when using reader oligonucleotide SEQ ID NO: 72, as determined by the reduction in median time between rereadings. [Table 5]
[0463] Example 11 This embodiment demonstrates a method in which a natural DNA analyte is reread multiple times at various sequencing temperatures.
[0464] A Y-adapter with a leader arm having C3 leader chemistry was prepared by annealing four DNA oligonucleotides having the sequences of SEQ ID NOs. 67, 68, 69, and 70. A DNA motor (Dda helicase) was loaded into the adapter, and the disulfide was closed via a reaction with 1,2-bismaleimide ethane (BMOE).
[0465] The DNA library was prepared by ligating the above-mentioned Y adapter to E. coli DNA prepared as described in Example 9.
[0466] Electrodynamic measurements were obtained using a custom MinION flow cell from Oxford Nanopore Technologies with inserted CsgG nanopores and a MinION Mk1b. 30 μL of FLT was added to 1170 μL of FB (from the Oxford Nanopore Technologies sequencing kit (SQK-LSK109)) to obtain a tether mix. After flowing 800 μL of the tether mix through the system, a 5-minute wait was allowed, and then another 200 μL of the tether mix was flowed through the system with the SpotON port open. A "sequencing mix" was obtained by mixing 37.5 μL SQB, 15 μL of DNA library, and 22.5 μL of LB from the Oxford Nanopore Technologies sequencing kit (SQK-LSK109). 75 μL of the sequencing mix was added to the MinION flow cell via the SpotON flow cell port.
[0467] Data was collected at a capture / sequencing voltage of 180 mV using a custom script similar to that described in Example 3, however, the motor protein was destalled by switching the voltage to zero by cleaving the channel for 5 seconds when an enzyme stall level was detected. Activated deblocking was configured to be triggered when a block level unrelated to terminal C3 levels, chains, open pores, or enzyme stall levels was recognized. Reread events were scored according to Example 9.
[0468] The table below shows the results of this experiment. The results show rereading at all temperatures tested, and indicate that rereading efficiency improved as the temperature increased, as judged by the increase in the percentage of reads reread within 30 seconds of reaching the C3 reader and the median dropback distance. [Table 6]
[0469] Explanation of the sequence list Sequence ID 1 shows the amino acid sequence of E. coli-derived (hexahistidine-tagged) exonuclease I (EcoExo I).
[0470] Sequence ID 2 shows the amino acid sequence of the exonuclease III enzyme derived from E. coli.
[0471] Sequence ID 3 shows the amino acid sequence of the RecJ enzyme (TthRecJ-cd) derived from T. thermophilus.
[0472] Sequence ID 4 shows the amino acid sequence of bacteriophage lambda exonuclease. This sequence is one of three identical subunits that make up the trimer. (http: / / www.neb.com / nebecomm / products / productM0262.asp)
[0473] Sequence ID 5 shows the amino acid sequence of Phi29 DNA polymerase derived from the Bacillus subtilis phage Phi29.
[0474] Sequence ID 6 shows the amino acid sequence of Trwc Cba (Citromicrobium bathyomarinum) helicase.
[0475] Sequence ID 7 shows the amino acid sequence of Hel308 Mbu (Methanococcoides burtonii) helicase.
[0476] Sequence ID 8 shows the amino acid sequence of Dda helicase 1993 derived from the intestinal bacterium phage T4.
[0477] Sequence IDs 20-33 show the nucleotide sequences of the polynucleotide chain discussed in Example 1.
[0478] Sequence ID 40 shows the amino acid sequence of the preferred HhH domain.
[0479] Sequence ID 41 shows the amino acid sequence of ssb derived from bacteriophage RB69 encoded by the gp32 gene.
[0480] Sequence ID 42 shows the amino acid sequence of ssb derived from bacteriophage T7 encoded by the gp2.5 gene.
[0481] Sequence ID 43 shows the amino acid sequence of UL42, which is derived from the herpesvirus.
[0482] Sequence ID 44 shows the amino acid sequence of subunit 1 of PCNA.
[0483] Sequence ID 45 shows the amino acid sequence of subunit 2 of PCNA.
[0484] Sequence ID 46 shows the amino acid sequence of subunit 3 of PCNA.
[0485] Sequence ID 47 shows the amino acid sequence (from 1 to 319) of the UL42 processivity factor from herpesvirus type 1.
[0486] Sequence ID 48 shows the amino acid sequence of the (HhH)2 domain.
[0487] Sequence ID 49 shows the amino acid sequence of the (HhH)2-(HhH)2 domain.
[0488] Sequence ID 50 shows the amino acid sequence of human mitochondrial SSB (HsmtSSB).
[0489] Sequence ID 51 shows the amino acid sequence of the p5 protein from Phi29 DNA p...
Claims
1. A method for characterizing a target polynucleotide, (i) Contacting the first opening of a transmembrane nanopore having a first opening and a second opening with the target polynucleotide, wherein the target polynucleotide has a stalled motor protein thereon, and the motor protein is stalled at the stalled portion. (ii) Bringing the stalled portion into contact with the nanopore, thereby causing the motor protein to destall, (iii) A method comprising obtaining one or more measurements characteristic of the target polynucleotide when the motor protein controls the movement of the target polynucleotide through the nanopore in the direction from the second opening of the nanopore to the first opening of the nanopore, thereby characterizing the target polynucleotide.
2. The method according to claim 1, wherein the nanopore extends across a membrane having a cis side and a trans side, the first opening of the nanopore is on the cis side of the membrane, the second opening of the nanopore is on the trans side, and the motor protein controls the movement of the target polynucleotide through the nanopore from the trans side to the cis side of the membrane.
3. The method according to claim 1, wherein the nanopore extends across a membrane having a cis side and a trans side, the first opening of the nanopore is on the trans side of the membrane, the second opening of the nanopore is on the cis side, and the motor protein controls the movement of the target polynucleotide through the nanopore from the cis side to the trans side of the membrane.
4. The method according to claim 1, comprising applying a force across the nanopore, wherein the motor protein controls the movement of the target polynucleotide through the nanopore in the direction opposite to the applied force.
5. The method according to claim 4, wherein the force includes a potential applied across the nanopore.
6. The method according to claim 1, wherein the motor protein is a helicase.
7. The method according to claim 1, wherein the motor protein is a DNA-dependent ATPase (Dda) helicase.
8. The method according to claim 1, wherein the adapter is attached to one or both ends of the target polynucleotide.
9. The method according to claim 8, wherein the motor protein is stalled on the adapter.
10. The method according to claim 1, wherein the nanopore captures a leader sequence at the first end of the target polynucleotide, and the motor protein is stalled at the second end of the target polynucleotide or on an adapter attached to the second end of the target polynucleotide.
11. - The target polynucleotide is single-stranded, - The target polynucleotide includes a leader sequence, the leader sequence is located at the first end of the target polynucleotide, or is included in an adapter attached to the first end of the target polynucleotide, - The method according to claim 1, wherein the motor protein is stalled at the second end of the target polynucleotide or on the adapter at the second end of the target polynucleotide.
12. - The target polynucleotide is double-stranded and comprises a first strand and a second strand, - The target polynucleotide includes a leader sequence, the leader sequence is located at the first end of the polynucleotide and is included in the first chain or included in an adapter attached to the first chain, - The method according to claim 11, wherein the motor protein is stalled at the second end of the target polynucleotide.
13. The method according to claim 12, wherein the motor protein is stalled at the second end of the first chain of the target polynucleotide, or on the adapter at the second end of the first chain of the target polynucleotide.
14. The method according to claim 12 or 13, wherein the first chain and the second chain are attached together by a hairpin adapter at the second end of the first chain, and the motor protein is stalled at the hairpin adapter.
15. The method according to claim 12, wherein the first strand and the second strand are attached together by hairpin adapters attached to (i) the second end of the first strand and (ii) the first end of the second strand, and the motor protein is stalled at the second end of the second strand of the double-stranded polynucleotide or on the adapter at the second end of the second strand.
16. The motor protein comprises one or more stall units independently selected from the following: - Polynucleotide secondary structure, - Nucleic acid analogs selected from peptide nucleic acids (PNA), glycerol nucleic acids (GNA), threose nucleic acids (TNA), locked nucleic acids (LNA), cross-linked nucleic acids (BNA), and debasalized nucleotides. - Nitroindole, inosine, acridine, 2 - aminopurine, 2 - 6 - diamino - purine, 5 - bromo - deoxyuridine, inverted thymidine (inverted dTs), inverted dideoxy - thymidine (ddTs), dideoxy - cytidine (ddCs), 5 - methylcytidine, 5 - hydroxymethylcytidine, 2’ - O - methyl RNA base, isodeoxycytidine (Iso - dCs), isodeoxyguanosine (Iso - dGs), C3(OC 3 H 6 OPO 3 ), photocleavable (PC) [OC 3 H 6 - C(O)NHCH 2 - C 6 H 3 NO 2 - CH(CH 3 )OPO 3 , hexanediol group, spacer 9 (iSp9) [(OCH 2 CH 2 ) 3 OPO 3 , spacer 18 (iSp18) [(OCH 2 CH 2 ) 6 OPO 3 , and spacer units selected from thiol linkages, and The method according to claim 1, wherein the vehicle is stalled at stall sites containing avidins such as fluorophores, traptabidine, streptavidin, and neutraavidin, and / or biotin, cholesterol, methylene blue, dinitrophenol (DNP), digoxigenin, and / or anti-digoxigenin and dibenzylcyclooctin groups.
17. The method according to claim 1, wherein destalling the motor protein includes applying a destalling force to the polynucleotide, the destalling force being less than and / or in the opposite direction to a read force, the read force being a force applied while the motor protein controls the movement of the target polynucleotide and measurements are being taken to determine one or more features of the polynucleotide.
18. The method according to claim 17, wherein the destalling of the motor protein includes stepping the applied force one or more times between the destalling force and the reading force.
19. The method according to claim 1, wherein the motor protein stalls at a stall site comprising one or more stall units and one or more stopping portions, and contacting the one or more stopping portions with the nanopore delays the movement of the polynucleotide through the nanopore, thereby causing the motor protein to destall from the one or more stall units.
20. The aforementioned stopping portion is one or more stopping portions independently selected from the following: - Polynucleotide secondary structure, - Nucleic acid analogs selected from peptide nucleic acids (PNA), glycerol nucleic acids (GNA), threose nucleic acids (TNA), locked nucleic acids (LNA), cross-linked nucleic acids (BNA), and debasalized nucleotides. - Fluorophores, avidins such as traptabidine, streptavidin and neutraavidin, and / or biotin, cholesterol, methylene blue, dinitrophenol (DNP), digoxigenin and / or anti-digoxigenin and dibenzylcyclooctin groups, and - The method according to claim 19, comprising a polynucleotide-binding protein.
21. The target polynucleotide includes a blocking portion to prevent the motor protein from unassociating with the polynucleotide. The target polynucleotide includes a leader sequence at its first end, the motor protein stalls on the second end of the target polynucleotide or on an adapter attached to the second end of the target polynucleotide, and the blocking portion is located between the motor protein and the second end of the polynucleotide, thereby preventing the motor protein from unassociating from the target polynucleotide at its second end. The method according to claim 1.