Polynucleotide sequencing using nanopores

JP2024534436A5Pending Publication Date: 2025-09-30ILLUMINA INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024516872
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-09-22
Filing Date
2022-09-19
Publication Date
2025-09-30

AI Technical Summary

Technical Problem

Existing methods for polynucleotide sequencing using nanopores are not sufficiently robust, reproducible, or cost-effective for practical applications, particularly in clinical and other settings requiring precise genome sequencing.

Method used

A method and system for sequencing polynucleotides using nanopores that involves positioning the polynucleotide through the nanopore with a 3' end on one side and a 5' end on the other, forming a duplex, applying forces to inhibit translocation, and measuring electrical properties such as current, resistance, or voltage drop to identify nucleotides, including the use of modified bases and nucleotide analogs to enhance stability and accuracy.

Benefits of technology

This approach provides accurate and reproducible sequencing with improved throughput and control over translocation, resolving issues with homopolymer detection and reducing errors, enabling sequencing of long reads up to 10,000 bases without the need for optical components.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Provided herein is sequencing of a polynucleotide using a nanopore. A polynucleotide is positioned through the opening of the nanopore such that its 3' end is on a first side of the nanopore and its 5' end is on a second side of the nanopore. A duplex is formed on the first side of the nanopore with the polynucleotide including the 3' end. The duplex is extended on the first side of the nanopore by adding a nucleotide to the 3' end of the duplex. A first force is applied that positions the 3' end of the duplex within the opening, and the nanopore inhibits translocation of the 3' end of the duplex to the second side of the nanopore. Values ​​of electrical properties of the 3' end of the duplex and the single-stranded portion of the polynucleotide are measured. The measurements are used to identify the nucleotide at the 3' end of the duplex.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] (CROSS REFERENCE TO RELATED APPLICATIONS) This application claims the benefit of U.S. Provisional Patent Application No. 63 / 247,155, entitled "Sequencing Polynucleotides Using Nanopores," filed September 22, 2021, the entire contents of which are incorporated by reference herein.

[0002] (Sequence Listing) The material in the attached Sequence Listing is incorporated by reference into this application. The attached Sequence Listing XML file (named "G1094_IP_2048_PCT.xml") was created on September 16, 2022 and is 8kB in size.

[0003] FIELD OF THEINVENTION This application relates generally to polynucleotide sequencing using nanopores. [Background technology]

[0004] A huge amount of time and energy has been spent by academics and companies on using nanopores for polynucleotide sequencing. For example, residence times have been measured for complexes of DNA with Klenow fragment (KF) of DNA polymerase I on nanopores in an applied electric field. Or, for example, current or flux measuring sensors have been used in experiments on DNA trapped in an α-hemolysin nanopore. Or, for example, KF-DNA complexes have been differentiated based on their properties when trapped in an electric field above an α-hemolysin nanopore. In yet another example, polynucleotide sequencing is performed using a single polymerase enzyme complex that includes a polymerase enzyme and a template nucleic acid attached proximal to a nanopore, and nucleotide analogs in solution. The nucleotide analogs include a charge blocking label attached to the polyphosphate portion of the nucleotide analog such that the charge blocking label is cleaved when the nucleotide analog is incorporated into the polynucleotide being synthesized. The charge blocking label is detected by the nanopore to determine the presence and identity of the incorporated nucleotide, thereby determining the sequence of the template polynucleotide. In yet another example, the construct comprises a transmembrane protein nanopore subunit and a nucleic acid handling enzyme.

[0005] Olasagasti et al., "Replication of individual DNA molecules under electronic control using a protein nanopore," Nature Nanotechnology 5(11):798-806 (2010), discloses positioning a DNA template through a nanopore. A DNA template-polymerase complex is formed on a first side of the α-hemolysin nanopore and includes a DNA duplex and a polymerase. The DNA template initially includes an abasic reporter nucleotide that is located on a second side of the α-hemolysin nanopore. The ionic current through the nanopore is measured (IEBS, where EBS refers to the enzyme-bound state) while the polymerase is used to add nucleotides to the duplex based on the sequence of the DNA template. As these nucleotides are added, the abasic reporter nucleotide is attracted toward and subsequently passes through the α-hemolysin, which is the I EBS causes changes in

[0006] However, such previously known compositions, systems, and methods are not always sufficiently robust, reproducible, or sensitive, and may not have a sufficiently high throughput for practical implementation, requiring commercial applications such as, for example, genomic sequencing in clinical and other settings that require cost-effective and highly accurate operation. Thus, there is a need for improved compositions, systems, and methods for sequencing polynucleotides. Summary of the Invention [Means for solving the problem]

[0007] Provided herein is sequencing of polynucleotides using nanopores.

[0008] Some examples herein provide a method of sequencing a polynucleotide using a nanopore that includes a first side, a second side, and an opening extending through the first and second sides. The method may include (a) positioning a polynucleotide through the opening of the nanopore such that a 3' end of the polynucleotide is on the first side of the nanopore and a 5' end of the polynucleotide is on the second side of the nanopore. The method may include (b) forming a duplex with the polynucleotide on the first side of the nanopore, the duplex including a 3' end. The method may include (c) extending the duplex on the first side of the nanopore by adding a first nucleotide to the 3' end of the duplex. The method may include (d) applying a first force that positions the 3' end of the extended duplex within the opening. The method may include using the nanopore to inhibit translocation of the 3' end of the extended duplex to a second side of the nanopore while the first force is applied, and measuring values ​​of an electrical property of the 3' end of the extended duplex and the single-stranded portion of the polynucleotide. The method may include (e) identifying the first nucleotide using the value measured in act (d).

[0009] In some examples, the value measured in operation (d) includes a current, an ionic current, an electrical resistance, or a voltage drop across the nanopore. In some examples, the value measured in operation (d) includes noise in the current, the ionic current, the electrical resistance, or the voltage drop across the nanopore. In some examples, the value measured in operation (d) includes a standard deviation of the noise.

[0010] In some examples, the value measured in act (d) is based on at least the M nucleotides of the single stranded portion of the polynucleotide and the D pairs of hybridized nucleotides of the extended duplex. In some examples, M is 2 or more and D is 1 or more. In some examples, M is 3 or more. In some examples, D is 2 or more. In some examples, at least one of the M nucleotides of the single stranded portion comprises a modified base, and the method includes identifying the modified base using the value measured in act (d). In some examples, the modified base comprises a methylated base.

[0011] In some examples, the method further includes inhibiting addition of another nucleotide to the 3' end of the extended duplex while the first force is applied in operation (d). In some examples, the nanopore inhibits addition of another nucleotide.

[0012] In some examples, the nanopore is oriented such that a first side of the nanopore comprises a majority of the opening, hi some examples, the nanopore is oriented such that a second side of the nanopore comprises a majority of the opening.

[0013] In some examples, the method further includes (f) applying a modified first force that repositions the 3' end of the extended duplex within the opening. The method may further include using the nanopore to inhibit translocation of the 3' end of the extended duplex to a second side of the nanopore while applying the modified first force, and measuring values ​​of an electrical property of the 3' end of the extended duplex and the single-stranded portion of the polynucleotide. The method may further include (g) identifying the first nucleotide using the value measured in (f).

[0014] In some examples, the first nucleotide is added using a polymerase in contact with the 3' end of the duplex. In some examples, the method may further include reversibly inhibiting the polymerase from adding a second nucleotide to the 3' end of the extended duplex. In some examples, the blocking moiety reversibly inhibits the polymerase from adding a second nucleotide to the 3' end of the extended duplex. In some examples, the first nucleotide is bound to the blocking moiety. In some examples, the blocking moiety comprises a 3'-blocking group. In some examples, the blocking moiety is reversibly associated with the extended duplex. In some examples, the method further includes detecting association of the blocking moiety with the extended duplex. In some examples, the method further includes detecting absence of the blocking moiety from the extended duplex. In some examples, the method further includes removing the blocking moiety to allow the polymerase to add a second nucleotide to the 3' end of the extended duplex.

[0015] In some examples, the first force applied in (d) removes the polymerase from contact with the 3' end of the extended duplex. In some examples, the method includes applying a second force to remove the polymerase from contact with the 3' end of the extended duplex, where the second force is greater than the first force.

[0016] In some examples, the method includes (f) applying a third force that places the polymerase in contact with the 3' end of the duplex and within or adjacent to the opening of the first side of the nanopore. The method may further include inhibiting movement of the polymerase into or further into the opening using the nanopore while applying the third force, and measuring a value of an electrical property of the polymerase. The method may include (g) using the value measured in (f) to identify contact of the polymerase with the 3' end of the duplex. In some examples, the third force is less than the first force. In some examples, operation (f) is performed after operation (c) and before operation (d). In some examples, the first nucleotide is associated with a blocking moiety. In some examples, operation (g) further includes using the value measured in operation (f) to confirm the presence of a blocking moiety associated with the first nucleotide.

[0017] In some examples, the polymerase includes a DNA polymerase. In some examples, the polymerase includes an RNA polymerase. In some examples, the polymerase includes a reverse transcriptase.

[0018] In some examples, the first nucleotide is associated with a blocking moiety, and operation (e) further includes using the value measured in operation (d) to confirm the presence of the blocking moiety associated with the nucleotide. In some examples, the method includes removing the blocking moiety from the first nucleotide after operation (d). In some examples, the method includes, after removing the blocking moiety, (f) again applying the first force to position the 3' end of the extended duplex within the opening and position the single-stranded portion of the polynucleotide within the opening. The method may include using the nanopore to inhibit translocation of the 3' end of the extended duplex to a second side of the nanopore while again applying the first force, and measuring values ​​of the electrical properties of the 3' end of the extended duplex and the single-stranded portion of the polynucleotide. The method may include (g) again identifying the first nucleotide using the value measured in operation (f). In some examples, the method includes removing the blocking moiety from the first nucleotide before operation (d).

[0019] In some examples, the extended duplex comprises one or more nucleotide analogs. In some examples, the one or more nucleotide analogs enhance the stability of the extended duplex compared to a natural nucleotide. In some examples, the one or more nucleotide analogs comprise one or more locked nucleic acids (LNAs). In some examples, the one or more nucleotide analogs comprise one or more 2'-methoxy (2'-OMe) nucleotides. In some examples, the one or more nucleotide analogs comprise one or more 2'-fluorinated (2'-F) nucleotides. In some examples, the one or more nucleotide analogs alter the value of an electrical property compared to a natural nucleotide. In some examples, the first nucleotide comprises one of the one or more nucleotide analogs. In some examples, the one or more nucleotide analogs comprise a 2' modification. In some examples, the one or more nucleotide analogs comprise a base modification.

[0020] In some instances, the first force is insufficiently strong to cause dissociation of the extended duplex.

[0021] In some examples, the first force includes a first voltage.

[0022] In some examples, operations (b) and (c) are performed in the absence of the first force.

[0023] In some examples, operation (c) is performed in the presence of a fourth force that opposes the first force.

[0024] In some examples, the first locking structure is attached to a 3' end of the polynucleotide on a first side of the nanopore. The first locking structure can inhibit translocation of the 3' end of the polynucleotide through the opening to a second side of the nanopore. In some examples, the first locking structure is removable.

[0025] In some examples, the second locking structure is attached to a 5' end of the polynucleotide on the second side of the nanopore. The second locking structure can inhibit translocation of the 5' end of the polynucleotide through the opening to the first side of the nanopore. In some examples, the second locking structure is removable.

[0026] In some examples, the method further includes, after operation (d), dissociating the extended duplex from the polynucleotide and forming a new duplex with the polynucleotide on the first side of the nanopore, the new duplex including a new 3' end.

[0027] In some examples, operation (a) includes contacting the nanopore with a polynucleotide hybridized to a substantially complementary polynucleotide and applying a sixth force that dehybridizes the substantially complementary polynucleotide from the polynucleotide.

[0028] In some examples, the nanopore comprises a solid-state nanopore. In some examples, the nanopore comprises a biological nanopore. In some examples, the biological nanopore comprises MspA.

[0029] In some instances, the polynucleotide comprises RNA. In some instances, the polynucleotide comprises DNA.

[0030] In some instances, the extended duplex comprises a primer hybridized to the polynucleotide.

[0031] Some examples herein provide a sequencing system. The sequencing system may include a nanopore including a first side, a second side, and an opening extending through the first and second sides. The sequencing system may include a polynucleotide disposed through the opening of the nanopore such that a 3' end of the polynucleotide is on the first side of the nanopore and a 5' end of the polynucleotide is on the second side of the nanopore. The sequencing system may include a duplex having a polynucleotide disposed on the first side of the nanopore, the duplex including a 3' end at which a first nucleotide is disposed. The sequencing system may include circuitry configured to apply a first force that positions the 3' end of the duplex within the opening. The circuitry may also be configured to measure values ​​of an electrical property of the 3' end of the duplex and the single-stranded portion of the polynucleotide while applying the first force. The circuitry may also be configured to identify the first nucleotide using the measurements. The nanopore may inhibit translocation of the 3' end of the duplex to a second side of the nanopore while the first force is applied.

[0032] In some examples, the value measured by the circuitry includes a current, an ionic current, an electrical resistance, or a voltage drop across the nanopore. In some examples, the value measured by the circuitry includes noise in the current, an ionic current, an electrical resistance, or a voltage drop across the nanopore. In some examples, the value measured by the circuitry includes a standard deviation of the noise.

[0033] In some examples, the value measured by the circuit is based on at least M nucleotides of the single stranded portion of the polynucleotide and D pairs of hybridized nucleotides of the double strand. M may be 2 or more and D may be 1 or more. In some examples, M is 3 or more. In some examples, D is 2 or more. In some examples, at least one of the M nucleotides of the single stranded portion comprises a modified base, and the circuit is configured to identify the modified base using the value measured by the circuit. In some examples, the modified base comprises a methylated base.

[0034] In some examples, while the first force is applied, addition of another nucleotide to the 3' end of the duplex is inhibited. In some examples, the nanopore inhibits addition of another nucleotide.

[0035] In some examples, the nanopore is oriented such that a first side of the nanopore comprises a majority of the opening, hi some examples, the nanopore is oriented such that a second side of the nanopore comprises a majority of the opening.

[0036] In some examples, the circuitry is further configured to apply a modified first force that repositions the 3' end of the duplex within the opening. The circuitry may be further configured to measure values ​​of an electrical property of the 3' end of the duplex and the single-stranded portion of the polynucleotide while applying the modified first force. The circuitry may be further configured to identify the first nucleotide using the measurements. The nanopore may inhibit translocation of the 3' end of the duplex to a second side of the nanopore while the modified first force is applied.

[0037] In some examples, the system further comprises a polymerase configured to contact the 3' end of the duplex and add a first nucleotide. In some examples, the polymerase is reversibly inhibited from adding a second nucleotide to the 3' end of the duplex. In some examples, the system further comprises a blocking moiety that reversibly inhibits the polymerase from adding a second nucleotide to the 3' end of the duplex. In some examples, the first nucleotide is bound to the blocking moiety. In some examples, the blocking moiety comprises a 3'-blocking group. In some examples, the blocking moiety is reversibly associated with the duplex. In some examples, the circuit is further configured to detect association of the blocking moiety with the duplex. In some examples, the circuit is further configured to detect absence of the blocking moiety from the duplex. In some examples, the blocking moiety is removable to allow the polymerase to add a second nucleotide to the 3' end of the duplex.

[0038] In some examples, the first force removes the polymerase from contact with the 3' end of the duplex. In some examples, the circuit is further configured to apply a second force to remove the polymerase from contact with the 3' end of the duplex, the second force being greater than the first force. In some examples, the circuit is further configured to apply a third force that places the polymerase in contact with the 3' end of the duplex within or adjacent to the opening of the first side of the nanopore. The circuit may be configured to measure a value of an electrical property of the polymerase while applying the third force. The circuit may be configured to use the measured value to identify contact of the polymerase with the 3' end of the duplex. The nanopore may inhibit movement of the polymerase into or further into the opening. In some examples, the third force is less than the first force. In some examples, the circuit is configured to apply the third force prior to applying the first force. In some examples, the first nucleotide is associated with a blocking moiety. In some examples, the circuitry is configured to use the measured value to confirm the presence of a blocking moiety associated with the first nucleotide.

[0039] In some examples, the polymerase includes a DNA polymerase. In some examples, the polymerase includes an RNA polymerase. In some examples, the polymerase includes a reverse transcriptase.

[0040] In some examples, the first nucleotide is associated with a blocking moiety, and the circuit is further configured to use the measured value to confirm the presence of the blocking moiety associated with the nucleotide. In some examples, the blocking moiety is removed from the first nucleotide after applying the first force. In some examples, the circuit is configured to place the 3' end of the duplex within the opening and reapply the first force to place the single-stranded portion of the polynucleotide within the opening after the blocking moiety is removed. The circuit may be further configured to measure values ​​of the electrical properties of the 3' end of the duplex and the single-stranded portion of the polynucleotide while reapplying the first force. The circuit may be further configured to again identify the first nucleotide using the measured value. The nanopore may inhibit translocation of the 3' end of the duplex to a second side of the nanopore. In some examples, the blocking moiety is removed from the first nucleotide before the first force is applied.

[0041] In some examples, the duplex comprises one or more nucleotide analogs. In some examples, the one or more nucleotide analogs enhance the stability of the duplex compared to a natural nucleotide. In some examples, the one or more nucleotide analogs comprise one or more locked nucleic acids (LNAs). In some examples, the one or more nucleotide analogs comprise one or more 2'-methoxy (2'-OMe) nucleotides. In some examples, the one or more nucleotide analogs comprise one or more 2'-fluorinated (2'-F) nucleotides. In some examples, the one or more nucleotide analogs alter the value of an electrical property compared to a natural nucleotide. In some examples, the first nucleotide comprises one of the one or more nucleotide analogs. In some examples, the one or more nucleotide analogs comprise a 2' modification. In some examples, the one or more nucleotide analogs comprise a base modification.

[0042] In some examples, the first force is insufficiently strong to cause dissociation of the duplex. In some examples, the first force comprises a first voltage. In some examples, in the absence of the first force, the polynucleotide is disposed through the opening and the duplex is disposed on a first side of the nanopore. In some examples, the circuit is configured to apply a fourth force that counteracts the first force.

[0043] In some examples, the first locking structure is attached to a 3' end of a polynucleotide on a first side of the nanopore and inhibits translocation of the 3' end of the polynucleotide through the opening to a second side of the nanopore. In some examples, the first locking structure is removable.

[0044] In some examples, a second locking structure is attached to a 5' end of the polynucleotide on the second side of the nanopore, and the second locking structure inhibits translocation of the 5' end of the polynucleotide through the opening to the first side of the nanopore. In some examples, the second locking structure is removable.

[0045] In some examples, the circuit is configured to dissociate the double strand from the polynucleotide after applying the first force.

[0046] In some examples, the circuitry is configured to apply a sixth force that dehybridizes the substantially complementary polynucleotide from the polynucleotide to position the polynucleotide through the opening of the nanopore.

[0047] In some examples, the nanopore comprises a solid-state nanopore. In some examples, the nanopore comprises a biological nanopore. In some examples, the biological nanopore comprises MspA.

[0048] In some instances, the polynucleotide comprises RNA. In some instances, the polynucleotide comprises DNA.

[0049] In some instances, the duplex comprises a primer hybridized to the polynucleotide.

[0050] Some examples herein provide a method of sequencing an unknown polynucleotide. The method may include providing a plurality of measurements of electrical properties of a single-stranded portion of the unknown polynucleotide and a 3' end of a duplex having the unknown polynucleotide within the opening of a nanopore as input to a nucleotide identification module. The method may include using the nucleotide identification module to compare the plurality of measurements to values ​​in a data structure, the data structure correlating different measurements with different combinations of nucleotides within the single-stranded portion of the known polynucleotide and the 3' end of a known duplex comprising the known polynucleotide within the opening of the nanopore. The method may include using the nucleotide identification module to determine a sequence of nucleotides in the sequence of the unknown polynucleotide using the comparison. The method may include receiving a representation of the determined sequence of nucleotides as output from the nucleotide identification module.

[0051] In some examples, the nucleotide identification module comprises a trained machine learning algorithm. In some examples, the nucleotide identification module comprises a trained deep learning algorithm. In some examples, the data structure comprises neurons of a trained machine learning algorithm.

[0052] In some examples, the data structure includes a read map, hi some examples, the read map includes a lookup table that stores different measurements and representations of different combinations of nucleotides within the 3' end of the known double stranded and single stranded portion of the known nucleotides.

[0053] In some examples, the method further includes using, by the computer, the measurement module to generate a plurality of measurements using the opening of the nanopore.

[0054] In some examples, the method further includes computerized use of the nucleotide addition module, the measurement module, and the nucleotide identification module to generate a data structure using the opening of the nanopore.

[0055] Some examples herein provide a system for sequencing an unknown polynucleotide. The system may include a processor and at least one computer-readable medium. The computer-readable medium may store a plurality of measurements of electrical properties of a single-stranded portion of an unknown polynucleotide and a 3' end of a duplex that includes the unknown polynucleotide within the opening of a nanopore. The computer-readable medium may store a data structure that correlates different measurements with different combinations of nucleotides within the single-stranded portion of a known polynucleotide and the 3' end of a known duplex within the opening of a nanopore. The computer-readable medium may store instructions for causing the processor to perform operations. The operations may include comparing the plurality of measurements to values ​​in the data structure. The operations may include using the comparison to determine a sequence of nucleotides in the sequence of the unknown polynucleotide. The operations may include outputting a representation of the determined sequence of nucleotides.

[0056] In some examples, the nucleotide identification module comprises a trained machine learning algorithm. In some examples, the nucleotide identification module comprises a trained deep learning algorithm. In some examples, the data structure comprises neurons of a trained machine learning algorithm.

[0057] In some examples, the data structure includes a read map, hi some examples, the read map includes a lookup table that stores different measurements and representations of different combinations of nucleotides within the 3' end of the known double stranded and single stranded portion of the known nucleotides.

[0058] In some examples, the instructions are further for causing the processor to generate a plurality of measurements using the opening of the nanopore.

[0059] In some examples, the instructions are further for causing the processor to generate a data structure using the opening of the nanopore.

[0060] Some examples herein provide a method of locking a polynucleotide into a nanopore that includes a first side, a second side, and an opening extending through the first and second sides. The method may include (a) attaching a first locking group to a 3' end of the polynucleotide. The method may include (b) positioning the polynucleotide through the opening of the nanopore such that the 3' end of the polynucleotide and the first locking group are on the first side of the nanopore and the 5' end of the polynucleotide is on the second side of the nanopore. The method may include (c) attaching a second locking group to the 5' end of the polynucleotide on the second side of the nanopore.

[0061] In some instances, the first locking group comprises a locked nucleic acid (LNA) or a peptide nucleic acid (PNA).

[0062] In some instances, the second locking group comprises a locked nucleic acid (LNA) or a peptide nucleic acid (PNA).

[0063] In some examples, the polynucleotide is hybridized to a complementary polynucleotide prior to operation (a), and the method further comprises dehybridizing the complementary polynucleotide between operations (b) and (c).

[0064] It should be understood that any respective feature / example of each of the aspects of the present disclosure described herein may be implemented together in any suitable combination, and any feature / example from any one or more of these aspects may be implemented together in any suitable combination with any of the features of the other aspects described herein, in order to achieve the benefits described herein. [Brief description of the drawings]

[0065] [Figure 1A] 1A-1C generally illustrate the use of an exemplary sequencing system and exemplary compositions and operations for sequencing a polynucleotide using a nanopore. [Figure 1B] 1A-1C generally illustrate the use of an exemplary sequencing system and exemplary compositions and operations for sequencing a polynucleotide using a nanopore. [Figure 1C] 1A-1C generally illustrate the use of an exemplary sequencing system and exemplary compositions and operations for sequencing a polynucleotide using a nanopore. [Figure 1D] 1A-1C generally illustrate the use of an exemplary sequencing system and exemplary compositions and operations for sequencing a polynucleotide using a nanopore. [Figure 1E] 1A-1C generally illustrate the use of an exemplary sequencing system and exemplary compositions and operations for sequencing a polynucleotide using a nanopore. [Figure 1F] 1A-1C generally illustrate the use of an exemplary sequencing system and exemplary compositions and operations for sequencing a polynucleotide using a nanopore. [Figure 1G] 1A-1C generally illustrate the use of an exemplary sequencing system and exemplary compositions and operations for sequencing a polynucleotide using a nanopore. [Figure 1H] 1A-1C generally illustrate the use of an exemplary sequencing system and exemplary compositions and operations for sequencing a polynucleotide using a nanopore. [Figure 2A] 1 generally illustrates the use of an exemplary sequencing system, as well as additional exemplary compositions and operations, for sequencing polynucleotides using nanopores. [Figure 2B] 1 generally illustrates the use of an exemplary sequencing system, as well as additional exemplary compositions and operations, for sequencing polynucleotides using nanopores. [Figure 2C] 1 generally illustrates the use of an exemplary sequencing system, as well as additional exemplary compositions and operations, for sequencing polynucleotides using nanopores. [Figure 2D]1 generally illustrates the use of an exemplary sequencing system, as well as additional exemplary compositions and operations, for sequencing polynucleotides using nanopores. [Figure 2E] 1 generally illustrates the use of an exemplary sequencing system, as well as additional exemplary compositions and operations, for sequencing polynucleotides using nanopores. [Diagram 3] 1 illustrates the flow of operations in an exemplary method for sequencing a polynucleotide. [Figure 4A] (SEQ ID NO:3, SEQ ID NO:4) Schematically illustrates the use of the sequencing system of Figures 1A-1H to resequence the same polynucleotide. [Figure 4B] (SEQ ID NO:3, SEQ ID NO:4) Schematically illustrates the use of the sequencing system of Figures 1A-1H to resequence the same polynucleotide. [Figure 4C] (SEQ ID NO:3, SEQ ID NO:4) Schematically illustrates the use of the sequencing system of Figures 1A-1H to resequence the same polynucleotide. [Figure 5A] (SEQ ID NO:3, SEQ ID NO:4) Schematically illustrates the use of the sequencing system of Figures 1A-1H to prepare a nanopore for sequencing different polynucleotides. [Figure 5B] (SEQ ID NO:3, SEQ ID NO:4) Schematically illustrates the use of the sequencing system of Figures 1A-1H to prepare a nanopore for sequencing different polynucleotides. [Figure 6] (SEQ ID NO:5, SEQ ID NO:6) Schematically illustrates the use of the sequencing system of Figures 1A-1H to generate and use polynucleotides for sequencing or polynucleotide synthesis. [Figure 7] 1A-1H are schematic illustrations of alternative configurations of the sequencing system of FIG. [Figure 8A] 1A-1H illustrate example values ​​of electrical characteristics that may be measured using the systems of FIGS. [Figure 8B]1A-1H illustrate example values ​​of electrical characteristics that may be measured using the systems of FIGS. [Figure 8C] 1A-1H illustrate example values ​​of electrical characteristics that may be measured using the systems of FIGS. [Figure 8D] 1A-1H illustrate example values ​​of electrical characteristics that may be measured using the systems of FIGS. [Figure 8E] 1A-1H illustrate example values ​​of electrical characteristics that may be measured using the systems of FIGS. [Figure 9] Illustrates an exemplary N-dimensional read map of electrical features that can be used to identify nucleotides using the system of FIGS. 1A-1H. [Figure 10] 1A-1H are schematic diagrams illustrating exemplary circuits that may be used in the systems of FIGS. [Figure 11] 1 illustrates a plot of values ​​measured as a function of time during sequencing of an exemplary polynucleotide. [Figure 12] 1 illustrates a plot of values ​​measured during resequencing of an exemplary polynucleotide under a set of measurement conditions. [Figure 13A] 1 illustrates a plot of values ​​measured during resequencing of an exemplary polynucleotide under different sets of measurement conditions. [Figure 13B] 1 illustrates a plot of values ​​measured during resequencing of an exemplary polynucleotide under different sets of measurement conditions. [Figure 13C] 1 illustrates a plot of values ​​measured during resequencing of an exemplary polynucleotide under different sets of measurement conditions. [Figure 14] 8 illustrates plots of values ​​measured during resequencing of an exemplary polynucleotide under different sets of measurement conditions, where the sequencing system has an alternative configuration as described with reference to FIG. 7. [Figure 15A] (SEQ ID NO:2) Illustrates a plot of values ​​measured during sequencing of a polynucleotide containing modified bases. [Figure 15B](SEQ ID NO:2) Illustrates a plot of values ​​measured during sequencing of a polynucleotide containing modified bases. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0066] Provided herein is sequencing of polynucleotides using nanopores.

[0067] More specifically, in a manner as described in more detail below, a single-stranded target polynucleotide to be sequenced is placed through the opening of the nanopore. A duplex is formed with a portion of the target polynucleotide at a first side of the nanopore. A force is then applied to pin the duplex within the opening of the nanopore, and an electrical measurement is made. The specific value of the measurement is based on the specific complementary bases located at the 3' end of the duplex and the specific sequence of bases in the single-stranded portion of the target polynucleotide within the opening of the nanopore. The specific measurement thus provides information from which the sequence of bases in the target polynucleotide can be determined. A force is then applied to remove the duplex from within the opening of the nanopore so that the duplex can be extended by nucleotides, and the measurement is repeated. The repeated measurements provide further information from which the sequence of bases in the target polynucleotide can be determined.

[0068] It will be appreciated that the subject matter can be used to sequence polynucleotides such as DNA or RNA using relatively few reagents and without the need for optical components that may otherwise add cost, weight, and complexity. The subject matter can be used to discriminate modified bases, such as methylated bases, without the need to chemically or enzymatically modify the modified bases. The subject matter accommodates relatively long reads, for example, up to about 1,000 bases, or up to about 2,000 bases, or up to about 5,000 bases, or even up to 10,000 or more bases. The subject matter overcomes the homopolymer problem traditionally associated with strand-based nanopore sequencing to improve accuracy of sequencing regions that contain repetitive nucleotides. The subject matter provides for controllable translocation of polynucleotides through the nanopore, thus inhibiting or preventing translocation events that are too fast to be detected and may result in deletion errors or other types of errors that adversely affect accuracy. These and other problems are solved by the present systems, compositions, and methods, as will be appreciated from the present disclosure.

[0069] We first provide a brief description of some of the terms used herein, then we describe some exemplary systems, compositions, and methods for sequencing polynucleotides using nanopores.

[0070] term Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art. The use of the term "including" and other forms such as "include", "includes" and "included" is not limiting. The use of the term "having" and other forms such as "have", "has" and "had" is not limiting. As used herein, whether in a transitional phrase or in the body of a claim, the terms "comprise" and "comprising" should be interpreted as having an open-ended meaning. That is, the above terms should be interpreted as synonymous with the phrase "having at least" or "comprising at least". For example, when used in the context of a process, the term "comprising" means that the process includes at least the recited steps, but may include additional steps. When used in the context of a compound, composition, or system, the term "comprising" means that the compound, composition, or system includes at least the recited features or components, but may include additional features or components.

[0071] As used herein, the singular forms "a," "an," and "the" include plural referents unless the content clearly dictates otherwise.

[0072] As used throughout this specification, the terms "substantially," "approximately," and "about" are used to describe and take into account small variations due to processing variations, etc. For example, they can refer to ±10% or less, such as ±5% or less, such as ±2% or less, such as ±1% or less, such as ±0.5% or less, such as ±0.2% or less, such as ±0.1% or less, such as ±0.05% or less.

[0073] As used herein, the term "nucleotide" is intended to mean a molecule that includes a sugar and at least one phosphate group, and in some instances also includes a nucleobase. A nucleotide that lacks a nucleobase may be referred to as "abasic." Nucleotides include deoxyribonucleotides, modified deoxyribonucleotides, ribonucleotides, modified ribonucleotides, peptide nucleotides, modified peptide nucleotides, modified phosphate sugar backbone nucleotides, and mixtures thereof. Examples of nucleotides include adenosine monophosphate (AMP), adenosine diphosphate (ADP), adenosine triphosphate (ATP), thymidine monophosphate (TMP), thymidine diphosphate (TDP), thymidine triphosphate (TTP), cytidine monophosphate (CMP), cytidine diphosphate (CDP), cytidine triphosphate (CTP), guanosine monophosphate (GMP), guanosine diphosphate (GDP), guanosine triphosphate (GTP), uridine monophosphate (UMP), uridine diphosphate (UDP), uridine triphosphate (GT ... deoxyadenosine monophosphate (UTP), deoxyadenosine monophosphate (dAMP), deoxyadenosine diphosphate (dADP), deoxyadenosine triphosphate (dATP), and deoxythymidine monophosphate (DTMP).These include deoxythymidine diphosphate (dTMP), deoxythymidine diphosphate (dTDP), deoxythymidine triphosphate (dTTP), deoxycytidine diphosphate (dCDP), deoxycytidine triphosphate (dCTP), deoxyguanosine monophosphate (dGMP), deoxyguanosine diphosphate (dGDP), deoxyguanosine triphosphate (dGTP), deoxyuridine monophosphate (dUMP), deoxyuridine diphosphate (dUDP), and deoxyuridine triphosphate (dUTP).

[0074] As used herein, the term "nucleotide" is also intended to encompass any nucleotide analog, which is a type of nucleotide that contains a modified nucleobase, sugar, backbone and / or phosphate moiety compared to naturally occurring nucleotides. Nucleotide analogs may also be referred to as "modified nucleic acids." Exemplary modified nucleobases include inosine, xanthate, hypoxanthate, isocytosine, isoguanine, 2-aminopurine, 5-methylcytosine, 5-hydroxymethylcytosine, 2-aminoadenine, 6-methyladenine, 6-methylguanine, 2-propylguanine, 2-propyladenine, 2-thiouracil, 2-thiothymine, 2-thiocytosine, 15-halouracil, 15-halocytosine, 5-propynyluracil, 5-propynylcytosine, 6-azouracil ... These include cytosine, 6-azothymine, 5-uracil, 4-thiouracil, 8-halo adenine or guanine, 8-amino adenine or guanine, 8-thiol adenine or guanine, 8-thioalkyl adenine or guanine, 8-hydroxyl adenine or guanine, 5-halo substituted uracil or cytosine, 7-methylguanine, 7-methyladenine, 8-azaguanine, 8-azaadenine, 7-deazaguanine, 7-deazaadenine, 3-deazaguanine, 3-deazaadenine, and the like. As known in the art, certain nucleotide analogs cannot become incorporated into polynucleotides, such as nucleotide analogs such as adenosine 5'-phosphosulfate. A nucleotide can include any suitable number of phosphates, such as 3, 4, 5, 6, or more than 6 phosphates. Nucleotide analogs also include locked nucleic acids (LNA), peptide nucleic acids (PNA), and 5-hydroxybutynyl-2'-deoxyuridine ("Super T").

[0075] As used herein, the term "polynucleotide" refers to a molecule that comprises a sequence of nucleotides linked together. A polynucleotide is a non-limiting example of a polymer. Examples of polynucleotides include deoxyribonucleic acid (DNA), ribonucleic acid (RNA), and their analogs, such as locked nucleic acid (LNA) and peptide nucleic acid (PNA). A polynucleotide can be a single-stranded sequence of nucleotides, such as RNA or single-stranded DNA, a double-stranded sequence of nucleotides, such as double-stranded DNA, or a mixture of single-stranded and double-stranded sequences of nucleotides. Double-stranded DNA (dsDNA) includes genomic DNA, and PCR and amplification products. Single-stranded DNA (ssDNA) can be converted to dsDNA and vice versa. A polynucleotide can include non-naturally occurring DNA, such as enantiomeric DNA, LNA, or PNA. The exact sequence of nucleotides in a polynucleotide can be known or unknown. The following are examples of polynucleotides: a gene or gene fragment (e.g., a probe, primer, expressed sequence tag (EST), or serial analysis of gene expression (SAGE) tag), genomic DNA, genomic DNA fragment, exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, ribozymes, cDNA, recombinant polynucleotides, synthetic polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes, primers, or amplified copies of any of the foregoing.

[0076] As used herein, "polymerase" is intended to mean an enzyme having an active site that assembles a polynucleotide by polymerizing nucleotides into a polynucleotide. A polymerase can bind to a primer and a single-stranded target polynucleotide and can sequentially add nucleotides to the growing primer to form a "complementary copy" polynucleotide with a sequence complementary to that of the target polynucleotide. A DNA polymerase can bind to a target polynucleotide and then move downstream of the target polynucleotide sequentially adding nucleotides to the free hydroxyl group at the 3' end of the growing polynucleotide chain. A DNA polymerase can synthesize a complementary DNA molecule from a DNA template. An RNA polymerase can synthesize an RNA molecule from a DNA template (transcription). Other RNA polymerases, such as reverse transcriptase, can synthesize cDNA molecules from an RNA template. Still other RNA polymerases can synthesize RNA molecules from an RNA template, such as RdRp. A polymerase can use a short RNA or DNA strand (primer) to initiate strand growth. Some polymerases can displace the strand upstream of the site where they add a base to the strand. Such polymerases can also be said to be strand displacing, that is, have the activity of removing the complementary strand from the template strand that is read by the polymerase.

[0077] Exemplary DNA polymerases include Bst DNA polymerase, 9°Nm DNA polymerase, Phi29 DNA polymerase, DNA polymerase I (E. coli), DNA polymerase I (Large), (Klenow) fragment, Klenow fragment (3'-5' exo-), T4 DNA polymerase, T7 DNA polymerase, Deep VentR™ (exo-) DNA polymerase, Deep VentR™ DNA polymerase, DyNAzyme™ EXT DNA, DyNAzyme™ II Hot Start DNA polymerase, Phusion™ High-Fidelity DNA polymerase, Therminator™ DNA polymerase, Therminator™ II DNA polymerase, VentR™ DNA polymerase, VentR™ (exo-) DNA polymerase, RepliPHI™ Phi29 DNA polymerase, rBst Examples of suitable polymerases include DNA polymerase, rBst DNA polymerase (Large), Fragment (IsoTherm™ DNA polymerase), MasterAmp™ AmpliTherm™ DNA polymerase, Taq DNA polymerase, Tth DNA polymerase, Tfl DNA polymerase, Tgo DNA polymerase, SP6 DNA polymerase, Tbr DNA polymerase, DNA polymerase beta, ThermoPhi DNA polymerase, and Isopol™ SD+ polymerase. In specific, non-limiting examples, the polymerase is selected from the group consisting of Bst, Bsu, and Phi29. Some polymerases have the activity of degrading their trailing strand (3' exonuclease activity). Some useful polymerases have been mutated or otherwise modified to reduce or eliminate 3' and / or 5' exonuclease activity.

[0078] Exemplary RNA polymerases include RdRps (RNA-dependent, RNA polymerases), which catalyze the synthesis of an RNA strand complementary to a given RNA template. Exemplary RdRps include poliovirus 3Dpol, vesicular stomatitis virus L, and Hepatitis C virus NS5B proteins. Exemplary RNA reverse transcriptases. A non-limiting exemplary list includes reverse transcriptases from Avian Myelomatosis Virus (AMV), Murine Moloney Leukemia Virus (MMLV), and / or Human Immunodeficiency Virus (HIV), telomerase reverse transcriptases such as (hTERT), SuperScript™ III, SuperScript™ IV reverse transcriptase, ProtoScript® II reverse transcriptase.

[0079] As used herein, the term "primer" is defined as a polynucleotide to which nucleotides can be added via a free 3'OH group. A primer may contain a 3' block that prevents polymerization until the block is removed. A primer may contain a modification at the 5' end to allow for a coupling reaction or to allow the primer to be coupled to another moiety. A primer may contain one or more moieties, such as 8-oxo-G, that can be cleaved under appropriate conditions, such as UV light, chemicals, enzymes, etc. The length of a primer may be any suitable number of bases long and may contain any suitable combination of natural and non-natural nucleotides. A target polynucleotide may contain an "amplification adaptor" or more simply an "adaptor" that hybridizes to the primer (having a sequence complementary to the primer) and can be amplified to generate a complementary copy polynucleotide by adding a nucleotide to the free 3'OH group of the primer.

[0080] As used herein, the term "plurality" is intended to mean a population of two or more distinct members. Pluralities can range in size from small, medium, large, to very large. A small size plurality can range, for example, from a few members to tens of members. A medium size plurality can range, for example, from tens of members to about 100 members or hundreds of members. A large plurality can range, for example, from about hundreds of members to about 1000 members, thousands of members, and tens of thousands of members. A very large plurality can range, for example, from tens of thousands of members to about hundreds of thousands, millions, tens of millions, or hundreds of millions or more members. Thus, pluralities can range in size from 2 to well over 100 million members, as well as between all sizes measured by number of members and larger than the exemplary ranges listed above. Exemplary polynucleotide pluralities include, for example, from about 1 x 105 or more, 5 x 10 5 A population of 10 or more, or 1 x 10 or more distinct polynucleotides is included. Thus, the definition of this term is intended to include all integer values ​​greater than 2. The upper limit of the plurality value can be set, for example, by the theoretical diversity of polynucleotide sequences in a sample.

[0081] As used herein, the term "double-stranded," when used in reference to a polynucleotide, is intended to mean that all or substantially all of the nucleotides in a polynucleotide are hydrogen bonded to each nucleotide in a complementary polynucleotide. A double-stranded polynucleotide may also be referred to as a "duplex."

[0082] As used herein, the term "single-stranded" when used in reference to a polynucleotide means that none of the nucleotides in the polynucleotide are hydrogen bonded to each nucleotide in a complementary polynucleotide.

[0083] As used herein, the term "target polynucleotide" is intended to mean a polynucleotide that is the subject of analysis or action, and may also be referred to using terms such as "library polynucleotide", "template polynucleotide", or "library template". The analysis or action includes subjecting the polynucleotide to amplification, sequencing, and / or other procedures. The target polynucleotide may include additional nucleotide sequences to the target sequence being analyzed. For example, the target polynucleotide may include one or more adapters, including amplification adapters that function as primer binding sites, that flank the target polynucleotide sequence being analyzed. In certain examples, the multiple target polynucleotides may have first and second adapters that are the same as each other, although they may have different sequences from each other. The two adapters that may flank a particular target polynucleotide sequence may have the same sequence as each other, or complementary sequences to each other, or the two adapters may have different sequences. Thus, a species in the multiple target polynucleotides may include a region of known sequence flanked by a region of unknown sequence that is evaluated, for example, by sequencing (e.g., SBS). In some instances, the target polynucleotide carries an amplification adapter at a single end, and such adapter may be located at either the 3' or 5' end of the target polynucleotide. The target polynucleotide may be used without an adapter, in which case the primer binding sequence may directly use the sequence present in the target polynucleotide.

[0084] The terms "polynucleotide" and "oligonucleotide" are used interchangeably herein. The difference in the terminology is not intended to indicate any particular difference in size, sequence, or other properties, unless otherwise specified. For clarity of explanation, when describing a particular method or composition that includes several polynucleotide species, different terms may be used to distinguish one species of polynucleotide from another species.

[0085] As used herein, the term "substrate" refers to a material used as a support for the compositions described herein. Exemplary substrate materials can include glass, silica, plastic, quartz, metal, metal oxide, organo-silicates (e.g., polyhedral organic silsesquioxanes (POSS)), polyacrylates, tantalum oxide, complementary metal oxide semiconductor (CMOS), or combinations thereof. An example of a POSS can be that described in Kehagias et al., Microelectronic Engineering 86 (2009), pp. 776-778, which is incorporated herein by reference in its entirety. In some examples, the substrate used in this application includes a silica-based substrate, such as glass, fused silica, or other silica-containing materials. In some examples, the silica-based substrate can include silicon, silicon dioxide, silicon nitride, or hydrogenated silicone. In some examples, the substrates used in this application include plastic materials or components such as polyethylene, polystyrene, poly(vinyl chloride), polypropylene, nylon, polyester, polycarbonate, and poly(methyl methacrylate). Examples of plastic materials include poly(methyl methacrylate), polystyrene, and cyclic olefin polymer substrates. In some examples, the substrate is or includes a silica-based material or a plastic material, or a combination thereof. In certain examples, the substrate has at least one surface that includes glass or a silicon-based polymer. In some examples, the substrate can include a metal. In some such examples, the metal is gold. In some examples, the substrate has at least one surface that includes a metal oxide. In one example, the surface includes tantalum oxide or tin oxide. Acrylamides, enones, or acrylates can also be utilized as substrate materials or components. Other substrate materials include, but are not limited to, gallium arsenide, indium phosphide, aluminum, ceramics, polyimides, quartz, resins, polymers, and copolymers.In some examples, the substrate and / or substrate surface can be or include quartz. In some other examples, the substrate and / or substrate surface can be or include a semiconductor, such as GaAs or ITO. The above list is intended to illustrate, but not limit, the present application. The substrate can include a single material or multiple different materials. The substrate can be a composite or laminate. In some examples, the substrate includes an organosilicate material.

[0086] The substrate can be horizontal, circular, spherical, rod-shaped, or any other suitable shape. The substrate can be rigid or flexible. In some examples, the substrate is a bead or a flow cell.

[0087] The substrate may be unpatterned, textured, or patterned on one or more surfaces of the substrate. In some examples, the substrate is patterned. Such patterns may include posts, pads, wells, ridges, channels, or other three-dimensional concave or convex structures. The pattern may be regular or irregular across the surface of the substrate. The pattern may be formed, for example, by nanoimprint lithography or by, for example, the use of metal pads to form features in a non-metallic surface.

[0088] In some examples, the substrates described herein form at least a portion of a flow cell, are located within a flow cell, or are coupled to a flow cell. A flow cell may include a flow chamber that is divided into multiple lanes or multiple sectors. Examples of flow cells that can be used in the methods and compositions described herein, as well as examples of substrates for manufacturing flow cells, include, but are not limited to, those commercially available from Illumina, Inc. (San Diego, Calif.).

[0089] As used herein, the term "electrode" is intended to mean a solid structure that conducts electricity. The electrode may include any suitable conductive material, such as gold, palladium, or platinum, or combinations thereof. In some examples, the electrode may be disposed on a substrate. In some examples, the electrode may define the substrate.

[0090] As used herein, the term "nanopore" is intended to mean a structure that includes an opening that allows a molecule to pass from a first side of the nanopore to a second side of the nanopore, where a portion of the opening of the nanopore has a width of 100 nm or less, e.g., 10 nm or less, or 2 nm or less. The opening extends through the first and second sides of the nanopore. Molecules that can pass through the opening of the nanopore can include, for example, ions, or water-soluble molecules such as amino acids or nucleotides. The nanopore can be disposed in a barrier or can be provided through a substrate. Optionally, the portion of the opening can be narrower than one or both of the first and second sides of the nanopore, in which case that portion of the opening can be referred to as a "constriction." Alternatively or additionally, the opening of the nanopore, or the constriction of the nanopore (if present), or both, can be 0.1 nm, 0.5 nm, 1 nm, 10 nm, or more. The nanopore can include multiple constrictions, e.g., at least two, or three, or four, or five, or more than four constrictions, and the nanopore includes a biological nanopore, a solid-state nanopore, or a hybrid biological and solid-state nanopore.

[0091] Biological nanopores include, for example, polypeptide nanopores and polynucleotide nanopores. "Polypeptide nanopore" is intended to mean a nanopore made from one or more polypeptides. The one or more polypeptides may include monomers, homopolymers, or heteropolymers. Polypeptide nanopore structures include, for example, α-helix bundle nanopores and β-barrel nanopores, as well as all others known in the art. Exemplary polypeptide nanopores include α-hemolysin, Mycobacterium smegmatis porin A, gramicidin A, maltoporin, OmpF, OmpC, PhoE, Tsx, F-pilus, SP1, mitochondrial porin (VDAC), Tom40, outer membrane phospholipase A, CsgG, aerolysin, and Neisseria autotransporter lipoprotein (Nalp). Mycobacterium smegmatis porin A (MspA) is a membrane porin produced by mycobacteria that allows hydrophilic molecules to enter the bacteria. MspA forms a tightly interconnected octamer and transmembrane beta barrel that resembles a goblet and contains a central constriction. For further details regarding α-hemolysin, see U.S. Pat. No. 6,015,714, the entire contents of which are incorporated herein by reference. For further details regarding SP1, see Wang et al., Chem. Commun., 49:1741-1743 (2013), the entire contents of which are incorporated herein by reference.For further details regarding MspA, see Butler et al., "Single-molecule DNA detection with an engineered MspA protein nanopore," Proc. Natl. Acad. Sci. 105:20647-20652 (2008) and Derrington et al., "Nanopore DNA sequencing with MspA," Proc. Natl. Acad. Sci. USA, 107:16060-16065 (2010), both of which are incorporated herein by reference in their entireties. Other nanopores include, for example, the MspA homologue from Norcadia farcinica, and lysenin. For further details regarding lysenin, see WO 2013 / 153359, which is incorporated herein by reference in its entirety. For further details regarding aerolysin, see Cao et al., "Single-molecule sensing of peptides and nucleic acids by engineered aerolysin nanopores," Nature Communications 10:Article number:4918 (2019), the entire contents of which are incorporated by reference herein.

[0092] "Polynucleotide nanopore" is intended to mean a nanopore made from one or more nucleic acid polymers. A polynucleotide nanopore can include, for example, a polynucleotide origami.

[0093] "Solid-state nanopore" is intended to mean a nanopore made from one or more materials that are not of biological origin. Solid-state nanopores can be formed of inorganic or organic materials. Solid-state nanopores include, for example, silicon nitride (SiN), silicon dioxide (SiCh), silicon carbide (SiC), hafnium oxide (HfCh), molybdenum disulfide (M0S2), hexagonal boron nitride (h-BN), or graphene. Solid-state nanopores may include an opening formed in a solid-state membrane, for example, a membrane that includes any such material.

[0094] "Biological and solid-state hybrid nanopore" is intended to mean a hybrid nanopore made from materials of both biological and non-biological origin. Materials of biological origin are defined above and include, for example, polypeptides and polynucleotides. Biological and solid-state hybrid nanopores include, for example, polypeptide solid-state hybrid nanopores and polynucleotide solid-state nanopores.

[0095] As used herein, a "barrier" is generally intended to mean a structure that inhibits the passage of molecules from one side of the barrier to the other side of the barrier. Molecules that are inhibited from passing may include, for example, ions or water-soluble molecules such as nucleotides and amino acids. However, when a nanopore is disposed within the barrier, the opening of the nanopore may allow the passage of molecules from one side of the barrier to the other side of the barrier. As one specific example, when a nanopore is disposed within the barrier, the opening of the nanopore may allow the passage of molecules from one side of the barrier to the other side of the barrier. Barriers include membranes of biological origin, such as lipid bilayers, and non-biological barriers, such as solid membranes or substrates.

[0096] As used herein, "biologically derived" refers to material that is derived from or isolated from a biological environment, such as an organism or cell, or a synthetically produced version of a biologically available structure.

[0097] As used herein, "solid" refers to a material that is not of biological origin.

[0098] As used herein, a "blocking moiety" is intended to mean a moiety that inhibits a polymerase from adding another nucleotide to the end of a double strand until the moiety is removed. A "blocking group" is a non-limiting example of a blocking moiety and is intended to mean a chemical group. In some instances, a nucleotide may be bound to a blocking group. Removal of a blocking group from a nucleotide may be referred to as "unblocking" the nucleotide. In instances where the 3' position of a nucleotide is bound to a blocking group, the blocking group may be referred to as a "3'-blocking group." A 3'-blocking group may inhibit a polymerase from binding another nucleotide to the nucleotide until the moiety is removed and replaced with a hydroxyl (OH) group.

[0099] As used herein, the term "methylated base" refers to a base that includes a methyl group (-CH3 or -Me) or a derivatized methyl group. For example, "methylcytosine" or "mC" refers to a cytosine in DNA that includes a methyl group or is a derivative of methylcytosine (i.e., 2'-deoxycytosine). As another example, "methyladenine" or "mA" refers to an adenine in DNA that includes a methyl group or is a derivative of methyladenine. A non-limiting example of a derivatized methyl group is an oxidized methyl group. A non-limiting example of an oxidized methyl group is hydroxymethyl (-CH2OH). An mC derivative with a hydroxymethyl group may be referred to as hydroxymethylcytosine or hmC. Another non-limiting example of an oxidized methyl group is a formyl group (-CHO). An mC derivative with a formyl group may be referred to as formylcytosine or fC. Another non-limiting example of an oxidized methyl group is carboxyl (-COOH). mC derivatives containing a carboxyl group may be referred to as carboxycytosine or caC. The methyl group of methylcytosine may be located at the 5-position of cytosine, in which case mC may be referred to as 5mC. The oxidized methyl group may be located at the 5-position of cytosine, in which case hmC may be referred to as 5hmC, fC may be referred to as 5fC, or caC may be referred to as 5caC. The methyl group of methyladenine may be located at the 6-position of adenine, in which case mA may be referred to as 6mA.

[0100] Polynucleotide sequencing using nanopores Some exemplary operations, compositions, and systems for sequencing polynucleotides using nanopores are now described with reference to Figures 1A-1H, 2A-2E, 3, 4A-4C, 5A-5B, 6, 7, 8A-8E, 9, and 10.

[0101] Referring now to FIG. 1A, a sequencing system 100 may include a nanopore 110, a polynucleotide 150 that is desired to be sequenced (which may be, for example, unknown), a polynucleotide 140 that hybridizes to a first portion 155 of the polynucleotide 150 to form a duplex 154, and circuitry 160.

[0102] Nanopore 110 may be disposed within barrier 101 and may include a first side 111, a second side 112, and an opening 113 extending through the first and second sides. Optionally, in some examples, nanopore 110 may include a constriction 114 within opening 113. Opening 113 of nanopore 110 may provide a path for fluid 120 and / or fluid 120′ to flow through barrier 101. Nanopore 110 may include a solid-state nanopore, a biological nanopore (e.g., MspA as illustrated in FIG. 1A), or a hybrid biological and solid-state nanopore. 1A, nanopore 110 may be oriented such that a first side 111 of the nanopore comprises a majority of opening 113, such that 3' ends 153 of duplex 154 may fit relatively deeply within opening 113 such that they are relatively inaccessible to fluid 120 and therefore may not be acted upon by any polymerase in the fluid. Alternatively, as illustrated in FIG. 7, nanopore 110 may be oriented such that a second side 112 of the nanopore comprises a majority of opening 113, such that 3' ends of duplex 154 may fit relatively shallowly within opening 113 and yet may be relatively inaccessible to fluid 120 and therefore may not be acted upon by any polymerase in the fluid. Although many of the examples herein may be described with reference to an asymmetric nanopore in a "forward" orientation as illustrated in FIG. 1A, it should be understood that in all such examples, the nanopore 110 may have either the orientation illustrated in FIG. 1A or the "reverse" orientation illustrated in FIG. 7. Furthermore, a nanopore may be symmetric and therefore may not necessarily be considered to have either a "forward" or a "reverse" orientation.

[0103] The barrier 101 may have any suitable structure that typically prevents the passage of molecules from one side of the barrier to the other side of the barrier, e.g., typically prevents contact between fluid 120 and fluid 120'. For example, as illustrated in FIG. 1A, the barrier 101 may include a first layer 107 and a second layer 108, one or both of which inhibit the flow of molecules across the layer. Illustratively, the barrier 101 may include a lipid bilayer including lipid layers 107 and 108. However, it will be understood that the barrier 101 may include any suitable structure, any suitable material, and any suitable number of layers. For example, the barrier 101 may include a solid barrier that may include a single layer or multiple layers. Non-limiting examples of materials that may be used for the barrier are provided elsewhere herein. Non-limiting examples and properties of barriers and nanopores are described elsewhere herein, as well as in U.S. Pat. No. 9,708,655, which is incorporated herein by reference in its entirety.

[0104] Polynucleotide 150 may include, for example, DNA or RNA. Polynucleotide 150 may be disposed through opening 113 of nanopore 110 such that first portion 155 of polynucleotide 150 is optionally located entirely on first side 111 of nanopore. A 3' end of polynucleotide 150 may be located on first side 111 of nanopore 110 and a 5' end of polynucleotide 150 may be located on second side 112 of nanopore 110. Optionally, the 3' end of polynucleotide 150 may be bound to or include a first steric lock 151. First steric lock 151 is sufficiently large that it cannot pass through nanopore 110 or a feature thereof (e.g., through constriction 114), thus retaining its end on the first side of the nanopore. Additionally or alternatively, the 5' end of polynucleotide 150 may be located on the second side 112 of nanopore 110 and may be coupled to or include second steric lock 152. Second steric lock 152 is sufficiently large to prevent passage of constriction 114 (such as an oligonucleotide hybridized to polynucleotide 150), thus retaining the second end of polynucleotide 150 on the second side of the nanopore. Thus, regardless of the polarity of the bias voltage that circuitry 160 may apply to electrodes 102 and 103 during the mode of operation illustrated in FIG. 1A, polynucleotide 150 may remain associated with nanopore 110 during this mode of operation.

[0105] Polynucleotide 140 may include, for example, DNA or RNA. Polynucleotide 140 may include, for example, a primer that hybridizes to a first portion of polynucleotide 150 to form a duplex 154. Duplex 154 between polynucleotide 140 and first portion of polynucleotide 150 may be located, optionally entirely, on the first side 111 of the nanopore and may include a 3' end 153. In this non-limiting example, 3' end 153 includes the base pair GC, desirably identifying a C as being at this position in the sequence of polynucleotide 150 in addition to the sequence of other nucleotides in polynucleotide 150. G nucleotide 121 may, but need not, be incorporated into polynucleotide 140 based on the sequence of polynucleotide 150 using a polymerase in a manner as described in more detail below with reference to FIG. 1B. A single-stranded second portion 156 of polynucleotide 150 (e.g., bases A and T) may be within opening 113, but may not be hybridized to polynucleotide 140. Bases of polynucleotides 140, 150 not specifically illustrated or labeled should be understood to be present, but are omitted for ease of illustration.

[0106] Hybridization between polynucleotide 140 and a first portion of polynucleotide 150 may be strong enough so that the duplex remains substantially intact, regardless of any bias voltage that circuit 160 may apply to electrodes 102 and 103 during a particular mode of operation. However, in a manner as further described below with reference to Figures 4A-4C, circuit 160 may be configured to have another mode of operation in which the circuit applies a sufficiently strong force to dissociate polynucleotide 140 from polynucleotide 150, making polynucleotide 150 available for resequencing using another polynucleotide 140, optionally using the same or a different set of measurement conditions. Additionally or alternatively, in a manner as further described below with reference to Figures 5A-5B, circuit 160 may be configured to have another mode of operation in which the circuit applies a sufficiently strong force to dissociate polynucleotide 150 from nanopore 110, making nanopore 110 available for sequencing a different polynucleotide.

[0107] In the mode of operation illustrated in FIG. 1A, the circuit 160 may apply a first force (F1) that positions the 3′ end 153 of the duplex 154 within the opening 113. The nanopore 110 inhibits translocation of the 3′ end 153 of the duplex 154 to the second side of the nanopore while the first force is applied. For example, at the particular moment illustrated in FIG. 1A, the circuit 160 applies a first force F1 (such as a first voltage across the nanopore 110) that moves the duplex 154 toward the second side 112 of the nanopore 110, while a constriction 114 or other feature of the nanopore 110 inhibits passage of the 3′ end 153 of the duplex to the second side of the nanopore. In the example illustrated in FIG. 1A, the duplex 154 may be wider than the constriction 114 and therefore sterically hindered from passing through the constriction 114. However, it should be understood that any suitable portion of nanopore 110 can be used to inhibit passage of 3' end 153 of duplex 154 to the second side of the nanopore.

[0108] As can be seen from the presence of duplex 154 during the mode of operation illustrated in FIG. 1A, the first force F1 is selected to be insufficient to cause dissociation of duplex 154, e.g., dehybridization of polynucleotide 140 from polynucleotide 150. Similarly, in other figures herein where a duplex is illustrated, it is understood that the force applied by circuit 160 is insufficient to cause dissociation of that duplex in the mode of operation described, however, other modes of operation may be used to intentionally dissociate the duplex. In some examples, duplex 154 may include one or more nucleotide analogs. Such nucleotide analogs may enhance the stability of duplex 154, for example, compared to natural nucleotides. For example, the nucleotide analogs may include one or more locked nucleic acids (LNAs), or may include one or more 2'-methoxy (2'-OMe) nucleotides, or may include one or more 2'-fluorinated (2'-F) nucleotides, or may include one or more peptide nucleic acids (PNAs). Such analogs may be included in, for example, polynucleotide 140. Illustratively, such analogs may be added to the primer as it is synthesized, or latent triphosphate nucleotides with these modifications may be incorporated into strand 140 by a polymerase. Such analogs may increase the Tm (melting temperature) of duplex 154. For example, the addition of LNA monomers may help to increase the Tm of duplex 154 and may be used to fine-tune the Tm of duplex 154. Or, for example, 2'-OMe may increase the Tm of an RNA:RNA duplex, causing only minor changes in RNA:DNA stability. Or, for example, 2' fluoro bases may have fluorine-modified ribose that increases binding affinity (Tm) and may also confer some relative nuclease resistance when compared to natural RNA.

[0109] The circuitry 160 may be configured to measure values ​​of electrical properties of the 3' end 153 of the duplex 154 and the single-stranded portion (e.g., second portion 156) of the polynucleotide 150 while applying the first force F1, and to use the measured values ​​to identify at least one nucleotide within the polynucleotide 150. For example, the circuitry 160 may be in operative communication with the first and second electrodes 102, 103 and may be configured to detect an ionic current passing through the nanopore 110, or a current, electrical resistance, or voltage drop across the nanopore 110 during application of the first force F1. For example, the fluid 120 may include one or more salts, such as KCl, NaCl, potassium ferrocyanide, potassium ferricyanide, or potassium glutamate. In a non-limiting example illustrated in FIG. 1A, a particular base pair (e.g., GC) at the 3' end 153 of double strand 154 and a sequence of one or more bases (e.g., A, T) in single stranded portion 156 can alter the rate at which salt in fluid 120 moves through opening 113 into fluid 120', and thus can change the current, ionic current, electrical resistance, or voltage drop across nanopore 110 in a manner that is detected by circuit 160. Additionally or alternatively, a particular base pair (e.g., GC) at the 3' end 153 of the duplex 154 and a sequence of one or more bases (e.g., A, T) in the single-stranded portion 156 may change the noise in the measurement, for example, changing the variation in the rate at which salts in the fluid 120 move through the opening 113 into the fluid 120', and thus changing the standard deviation of the current, the standard deviation of the ionic current, the standard deviation of the electrical resistance, or the standard deviation of the voltage drop across the nanopore 110 in a manner that is detected by the circuit 160. In some examples, the duplex 154 may include one or more nucleotide analogs that change the value of the electrical property of the nanopore compared to a natural nucleotide. For example, the polynucleotide 140 may include a nucleotide analog, such as a 2' modification or a base modification. Such analogs may thus be identified using the circuit 160.

[0110] 1A , the value that circuit 160 measures may be based on any suitable number of nucleotides in first portion 155 (which is hybridized to polynucleotide 140) and second portion 156 (which is single stranded and not hybridized to polynucleotide 140). For example, the measurement may be based on at least M nucleotides of the single stranded portion and D pairs of hybridized nucleotides of the double stranded portion, where M is 1 or more and D is 1 or more. M and D may have any suitable values. For example, M may be 2 or more, or M may be 3 or more, or M may be 4 or more. Illustratively, M may be about 1. Or illustratively, M may be about 2. Or illustratively, M may be about 3. Or illustratively, M may be about 4. Or illustratively, M may be about 5. Additionally or alternatively, D may be 2 or greater, or D may be 3 or greater, or D may be 4 or greater. Exemplarily, D may be about 1. Or, exemplarily, D may be about 2. Or, exemplarily, D may be about 3. Or, exemplarily, D may be about 4. Or, exemplarily, D may be about 5.

[0111] The values ​​measured by the circuit 160 may include a current, ionic current, electrical resistance, or voltage drop based on the M nucleotides and the hybridized nucleotide D-pairs. The values ​​measured by the circuit 160 may also or alternatively include noise of the current, ionic current, electrical resistance, or voltage drop across the nanopore, for example, the standard deviation of the noise, which is based on the M nucleotides and the hybridized nucleotide D-pairs. For example, referring now to FIG. 1F, numbers are used to represent bases that may affect the measurement under a given force (here, a first force F1). In the MspA nanopore as illustrated, it may be expected that bases 4, 5, 6, and 4' may primarily affect the measurement, but bases adjacent to these may also affect the measurement in a manner as described elsewhere herein. For example, bases 3 and 3' may affect the measurement. Additionally or alternatively, base 7 may affect the measurement. Additionally or alternatively, base 8 may affect the measurement. In one non-limiting example, bases 3, 3', 4, 4', 5, 6, 7, and 8 affect the measurement. In another non-limiting example, bases 3, 3', 4, 4', 5, 6, and 7 affect the measurement. In another non-limiting example, bases 4, 4', 5, 6, 7, and 8 affect the measurement. In another non-limiting example, bases 3, 3', 4, 4', 5, and 6 affect the measurement. In another non-limiting example, bases 4, 4', 5, 6, and 7 affect the measurement. In another non-limiting example, bases 4, 4', 5, and 6 affect the measurement.

[0112] It will be understood that the particular number and location of bases that affect the measurements may depend on the particular nanopore configuration used, the particular circuit configuration used, and the particular conditions under which the measurements are performed. For example, as illustrated in FIG. 1G, under a modified first force (F1'), the 3' end of the duplex 154 may be positioned differently relative to the nanopore 110, such as deeper within the opening 113 if F1' is greater than F1, or shallower within the opening 113 if F1' is less than F1, in the non-limiting example shown in FIG. 1G. The ionic current through the opening 113 may be affected differently by bases within the duplex 154 and the single-stranded portion 156 of the polynucleotide 150 for different forces. In a manner as further described below with reference to FIGS. 8A-8E, 9, and 10, measurements made under different forces may be compared to still further enhance the precision with which nucleotides are identified.

[0113] In some examples, circuit 160 is configured to identify at least one nucleotide in polynucleotide 150 as being modified, such as, but not limited to, a methylated base. For example, referring to FIG. 1H, where numbers are again used to represent bases that may affect measurements, base 5 * is methylated (or otherwise modified), it may be expected that such a group may affect the measurement differently than the unmethylated (or otherwise unmodified) base, and thus the presence of a methylated (or otherwise modified) base may be identified via its effect on the measurement.

[0114] Further details regarding the manner in which circuit 160 can be used to identify nucleotides using measurements are provided below with reference to FIGS. 8A-8E, 9, and 10.

[0115] 1A, while the first force F1 is applied using the circuit 160, the addition of nucleotides to the 3' end 153 of the duplex 154 may be inhibited. For example, the fluid 120 may be in contact with the first side 111 of the nanopore 110 and may include a plurality of nucleotides 121, 122, 123, 124, e.g., G, T, A, and C, respectively. Each of the nucleotides 121, 122, 123, 124 in the fluid 120 may be optionally modified in a manner as described in more detail below, e.g., may be bound to a respective blocking moiety or may include a nucleotide analog. The fluid 120 may further include a plurality of polymerases 105 that may be used to add nucleotides to the polynucleotide 140 using the sequence of the polynucleotide 150 in a manner as described with reference to FIG. 1B. 1A , the opening 113 of nanopore 110 may inhibit the addition of a nucleotide to the 3′ end 153 of duplex 154. For example, polymerase 105 may be sterically hindered from binding to 3′ end 153 while the 3′ end is located within opening 113.

[0116] However, the circuit 160 may be configured to switch the system 100 to a mode of operation in which the 3' end of the duplex 154 may be extended by adding a nucleotide. Such a nucleotide addition operation may be performed after forming the duplex 154 and before applying the first force F1, for example, to add nucleotide 121 before a certain time illustrated in FIG. 1A. FIG. 1B illustrates an exemplary mode for adding a nucleotide to the duplex 154. In this mode, the circuit 160 may be configured to apply a second force F2 (such as a second voltage across the nanopore 110 in a direction opposite to that of the first voltage) that moves the 3' end 153 of the duplex 154 out of the opening 113 so that the polymerase 105 can contact the 3' end of the duplex and add a nucleotide thereto, e.g., T122, based on the next nucleotide (e.g., A) in the sequence of the polynucleotide 150. It will be appreciated that the circuit 160 may instead be configured to release the first force F1, after which the released 3' end 153 of the duplex 154 may naturally diffuse away from the opening 113 so that the polymerase 105 may contact the 3' end of the duplex and add a first nucleotide thereto, without the need to use a circuit to actively apply a force to cause such movement and make the 3' end of the duplex available to the fluid 120.

[0117] It should be noted that any suitable type of polymerase 105 may be used to synthesize any suitable type of polynucleotide 140 based on any suitable type of polynucleotide 150. For example, a DNA polymerase 105 may be used to copy a DNA polynucleotide 150 to form a DNA polynucleotide 140. Or, for example, an RNA polymerase 105 may be used to copy a DNA polynucleotide 150 to form an RNA polynucleotide 140. Or, for example, an RdRP (RNA-dependent RNA polymerase) 105 may be used to copy an RNA polynucleotide 150 to form an RNA polynucleotide 140. Or, for example, a reverse transcriptase 105 may be used to copy an RNA polynucleotide 150 to form a DNA polynucleotide 140. Any DNA modification may occur on the base or sugar comprising the 3' end. Any RNA modification may occur on the base or sugar comprising the 2' end. Examples of modifications to the 2' end include, but are not limited to, 2'-O-methoxy-ethyl, 2'-OMe, 2'-F, and locked nucleic acid (LNA).

[0118] The circuit 160 may be configured to repeatedly switch the system 100 between a nucleotide addition mode (FIG. 1B) and a measurement mode (FIG. 1A). For example, after performing the measurement described with reference to FIG. 1A, another nucleotide may be added in a manner as described with reference to FIG. 1B, and another measurement may be performed. During the measurement mode, the opening or other features of the nanopore 110 may inhibit the polymerase 105 from adding another nucleotide. Thus, each cycle of nucleotide measurement may be controlled to have any desired duration, e.g., electronically controlled using the circuit 160 to provide an appropriate signal-to-noise ratio (SNR) for the type of measurement being performed. In general, the longer the measurement cycle, the better the SNR that can be obtained. In some applications, the circuit 160 may be configured to adjust the length of the measurement mode to obtain a sufficiently high SNR (e.g., an SNR above a predefined threshold) even with a lower throughput (number of bases per unit time). In other applications, the circuit 160 may be configured to adjust the length of the measurement mode to obtain a sufficiently high throughput (eg, a throughput above a predefined threshold) even with a lower SNR.

[0119] The circuit 160 may be further configured to perform any suitable number of repeated cycles of nucleotide addition and measurement. For example, as illustrated in FIG. 1C, after applying the second force F2 while the nucleotide is added, the circuit 160 may again apply the first force F1. The first force F1, in some instances, can remove the polymerase 105 from contact with the 3′ end 153 of the duplex 154 and move the now extended 3′ end 153 of the duplex 154 into the opening 113 of the nanopore 110, where a constriction 114 or other feature of the nanopore 110 inhibits translocation of the 3′ end to the second side 112 of the nanopore in a manner similar to that described with reference to FIG. 1A. Alternatively, the first force F1 may be insufficient to remove the polymerase 105 from contact with the 3' end 153 of the duplex 154, and the circuit 160 may be configured to apply another force, which may be greater than the first force, to remove the polymerase from contact with the 3' end of the duplex.

[0120] In some examples, the circuit 160 is configured to measure the binding of the polymerase 105 to the 3' end 153 of the duplex 154 as a way to confirm that the polymerase is in fact in the process of adding a nucleotide to the duplex or has already added a nucleotide to the duplex. Such confirmation may be useful, for example, when subsequent nucleotides are added that may otherwise produce the same or similar electrical property values ​​as one another and thus may otherwise be difficult to distinguish from one another based solely on their values. For example, in a manner as illustrated in FIG. 1D, the circuit 160 may apply a third force F3 to position the polymerase in contact with the 3' end of the duplex within or adjacent to the opening of the first side of the nanopore. While the third force F3 is applied, the nanopore 110 may prevent the polymerase 105 from passing into or further into the opening. For example, the polymerase 105 may be sterically hindered from entering the opening 113 at all or more than a limited amount (note that an actual polymerase may be significantly larger than that illustrated generally in this application, and may in fact in some instances be approximately the same diameter as the nanopore 110). A third force F3 may be applied so as not to peel the polynucleotide 140 from the polynucleotide 150. This may be accomplished by keeping F3 below such a peeling force, or if F3 is greater than the force required for peeling, F3 is applied in a transient manner so as not to allow time for the duplex 154 to dissociate.

[0121] During the mode of operation illustrated in FIG. 1D, the circuit 160 can measure the value of an electrical property of the polymerase 150 while applying the third force F3 and use the measured value to identify contact of the polymerase with the 3′ end of the duplex. For example, association of the polymerase 105 with the 3′ end 153 can change the rate at which salt in the fluid 120 moves through the opening 113 into the fluid 120′, and thus change the current, ionic current, electrical resistance (resistance to ion flow), or voltage drop across the nanopore 110, or the standard deviation of any such electrical characteristic, in a manner detected by the circuit 160. The third force F3 can be sufficient to move the complex including the polymerase 105 and the 3′ end 153 toward and into contact with the nanopore 110, and can be low enough so as not to remove the polymerase from the 3′ end 153. Additionally or alternatively, the circuit 160 can determine whether there is a polymerase 105 still bound to the 3' end 153 of the duplex, since if present, the 3' end of the duplex is held at the top of the pore by a polymerase that is too large to enter the pore, and primarily single-stranded DNA is in the opening 113. After the measurement is made, the circuit 160 applies a suitable force to remove the polymerase 105 from the 3' end of the duplex, and a shift in the signal may be observed to confirm removal of the polymerase as the 3' end of the duplex enters the pore. In some examples, F1 may be stronger than F3 and in the same direction as F3. The circuit 160 may apply force F1 in a manner to position the single-stranded second portion 156 of the polynucleotide 150 within the opening 113 for a measurement as illustrated in FIG. 1E, which repeats the measurement described with reference to FIG. 1A, but with the newly extended 3' end 153 and the newly shifted single-stranded second portion 156.The operations of adding a nucleotide to the 3' end 153, optionally detecting a polymerase 105 in contact with the 3' end 153, removing the polymerase, and measuring values ​​of electrical properties of the 3' end 153 of the duplex 154 and the single-stranded second portion 156 of the polynucleotide 150 (from which at least one nucleotide can be identified) can be repeated any suitable number of times, for example, repeated along substantially the entire length of the polynucleotide 150 to substantially sequence the polynucleotide 150.

[0122] Note that in examples including a constriction 114, such constriction 114 occupies only a portion of the length of the nanopore 110, but since the constriction is where the largest voltage drop occurs (as it exhibits the greatest resistance between electrodes 102 and 103), it may be the most sensitive region for base discrimination. A typical nanopore constriction is longer than a single DNA nucleotide of a single-stranded polynucleotide, and therefore the current signal that the nanopore can generate may depend on two or more nucleotides, typically three, four, five, or even six or more nucleotides. These nucleotides form what may be referred to as a "K-mer." The number of possible K-mers for four bases of DNA is 4. K Some previously known strand sequencing methods work by translocating a single strand of DNA through the constriction (i.e., "sensing zone") of a nanopore, such as MspA or CsgG, and attempting to associate the current generated by each K-mer with a sequence of K-mers. One way to perform this association is to create a table of each possible K-mer present in the nanopore constriction at any time and the associated current it generates, and then use a lookup table to find the K-mer that corresponds to the measured current. Another way to perform the association is through the use of machine learning algorithms. In practice, any method of K-mer-based strand sequencing has several limitations.

[0123] For example, two currents corresponding to a unique K-mer may appear to be the same, or may be similar enough that they are indistinguishable from one another within a given experimental setup. The larger the K-mer, the more likely such "degenerate" cases are to occur. For this reason, the α-hemolysin nanopore has a large K-mer, on the order of 10 nucleotides, and thus a K-mer of about 4 10 This presents a challenge to strand sequencing because there are many possible signals, too many to realistically resolve. Other methods have sometimes been used to try to distinguish K-mers when they are somewhat short, around six nucleotides in length, and degenerate. For example, using a translocating enzyme (e.g., a helicase or polymerase) to translocate the DNA through the pore one nucleotide at a time, allowing time to measure the electrical value for each nucleotide in succession, alleviates the K-mer problem. This so-called single-base ratcheting alleviates the K-mer problem because a given K-mer can transition to one of only four possibilities, since the base that leaves the sensing zone is replaced by A, C, G, or T, while the other base remains unchanged in the sensing zone. For example, for a 4-mer such as ACGT, if the A leaves, CGT changes register, and then when the new base comes in, the new 4-mer can only be CGTT, CGTA, CGTC, or CGTG, but in reality, there are four nucleotide types that can be formed: 4-mer, ... 4 There are (256) K-mers. Thus, the list of possible new K-mers is reduced from 256 to only 4, reducing the possibility of degenerate signals.

[0124] In addition, if a homopolymer stretch in the template DNA translocates through a nanopore that is longer than the nanopore sensing zone (e.g., the length of the K-mer for a particular nanopore type), current strand sequencing methodologies may not be able to accurately determine the length of the homopolymer because there is no indication that the template DNA has translocated. In these cases, attempts have been made to use the length of time that the homopolymer passes through the nanopore to estimate the length of the homopolymer (by multiplying the average number of nucleotides per unit of time by the time). However, because the translocating enzyme does not move at a consistent rate, time alone cannot provide sufficient accuracy to determine the length of the homopolymer. In fact, because the rate at which the translocating enzyme moves cannot be well controlled and the time per nucleotide may not be consistent, some translocation events may occur extremely rapidly and may be too short to be detected using strand sequencing. As a result, deletion errors may occur.

[0125] The present subject matter is believed to alleviate any and all of such problems associated with strand sequencing, and in fact provides significantly improved accuracy, control, and reproducibility as compared to strand sequencing.

[0126] 1A-1H, circuit 160 measures a signal based on the combination of duplex 154 and single-stranded portion 156 of polynucleotide 150, as well as the particular set of measurement conditions used. Thus, the information measured is much richer than for an exclusively single-stranded polynucleotide under a single set of measurement conditions, as in, for example, strand sequencing.

[0127] Through the use of the methods described herein to confirm the addition of a nucleotide to the 3' end of 153, the homopolymer problem is resolved. For example, each time a new base is added to the 3' end 153 of the duplex 154 (e.g., as described with reference to Figures 1B and 1D), a distinguishable signal is generated (e.g., as described with reference to Figures 1A, 1C, and 1E), through which it can be confirmed that a single nucleotide has been added. If the signals from the different consecutively added nucleotides are sufficiently similar or identical, it can be inferred that the homopolymer stretch of polynucleotide 150 has been sequenced. Furthermore, different nucleotides within a given homopolymer stretch of polynucleotide 150 may result in different measurements of electrical properties, since each such nucleotide may have a different proximity to other different nucleotides outside the homopolymer stretch. Thus, even nucleotides of the same type within a homopolymer can be individually identified using the values ​​measured by circuit 160.

[0128] Further, during this measurement as described with reference to FIG. 1A , 3′ end 153 (including the 3′ end of polynucleotide 140) may be sequestered within opening 113 or otherwise inaccessible for the addition of another nucleotide until circuitry 160 applies an opposing force that releases 3′ end 153 from opening 113 and makes the 3′ end accessible to fluids, polymerase, and nucleotides for use in adding another nucleotide to 3′ end 153.

[0129] The system and method provide excellent control of the sequencing process, so that the problem of enzyme speed is also solved. For example, the circuit 160 can electronically control the duration of the measurement operation to achieve a desired SNR (e.g., an SNR above a predefined threshold) while inhibiting the polymerase 105 from adding another nucleotide. The translocation of the polynucleotide 150 through the nanopore 110 is also electronically controlled using the circuit 160 and is therefore not subject to variable kinetics of the translocation enzyme, as in the case of strand sequencing, where very fast translocation events may go undetected, which could otherwise lead to deletions or other types of errors affecting accuracy.

[0130] For example, Figures 8A-8E illustrate exemplary values ​​of electrical characteristics that may be measured using the system of Figures 1A-1H. More specifically, Figure 8A illustrates an exemplary series of values ​​that may be measured under a first set of measurement conditions when base pairs G(153)C(paired at 150 with 153) (referred to herein as "GC"), T(153)A(paired at 150 with 153) (referred to herein as "TA"), and A(153)T(paired at 150 with 153) (referred to herein as "AT") each become located at the 3' end 153 of duplex 154, and a sequence of nucleotides AT, TG, and GC each become located at a second portion 156 of single strand of polynucleotide 150, in a manner as described with reference to Figures 1A, 1C, and 1E, respectively. As illustrated in Figure 8A, a first value is measured for the combination of GC and AT (Figure 1A), a second value is measured for the combination of TA and TG (Figure 1C), and a third value is measured for the combination of AT and GC (Figure 1E). The particular types of nucleotides in the sequence of nucleotides at the 3' end 153 of the double strand 154 and located in the single-stranded second portion 156 of the polynucleotide 150 can affect the ionic current (and / or fluctuations in the ionic current) through the nanopore 110 and thus the measurements.

[0131] Indeed, certain types of paired nucleotides within duplex 154 and away from the 3' end 153 of duplex 154, and / or certain types of unpaired nucleotides within polynucleotide 150 and away from the single-stranded second portion 156 of polynucleotide 150, may affect the ionic current (and / or ionic current fluctuations or other noise) through nanopore 110 and thus the measurements. For example, Figure 8B illustrates a sequence of nucleotides AT, TG, and GC such that base pairs GC, TA, and AT, respectively, are located at the 3' end 153 of duplex 154, in the manner described with reference to Figures 1A, 1C, and 1E, respectively. * 1 illustrates an exemplary series of values ​​that may be measured, again under the first set of measurement conditions, when each of C * is a nucleotide analog, e.g., methylated cytosine. As illustrated in Figure 8B, a first value is determined for the combination of GC and AT (Figure 1A), a second value is determined for the combination of TA and TG (Figure 1C), and a third value is determined for the combination of AT and GC * The first value in FIG. 8B may be the same as the first value in FIG. 8A, but this may be, for example, C * is three bases away from the 3' end of the duplex and therefore may not significantly affect the ionic current through nanopore 110. The second value in FIG. 8B may differ somewhat from the second value in FIG. 8A, for example, C * Compared with how an unmethylated C can affect such a current, even though C is separated by two bases from the 3′ end of the duplex. * may have some effect on the ionic current through nanopore 110. The third value in FIG. 8B may be significantly different from the third value in FIG. 8A, for example, because * is directly adjacent to the 3' end of the duplex, compared to how unmethylated C could affect such currents. *can significantly affect the ionic current through nanopore 110.

[0132] Different sets of measurement conditions may also affect the values ​​measured for different combinations of nucleotides at the 3' end 153 of duplex 154 and the second portion 156 of the single strand of polynucleotide 150. For example, applying different forces using circuit 160 in a manner as described with reference to FIG. 1G may move the 3' end 153 of duplex 154 and the second portion 156 of polynucleotide 150 to different positions relative to nanopore 110, where the nucleotides in duplex 154 and polynucleotide 150 may affect the ionic current (and / or ionic current fluctuations or other noise) differently than they would at another position (under another force). Changes in the measurement conditions may affect the measurements linearly or nonlinearly, and may indeed change the measurements in different directions for different combinations of nucleotides at the 3' end 153 of duplex 154 and within the second portion 156 of polynucleotide 150.

[0133] For example, FIG. 8C illustrates a duplex 154 having base pairs GC, TA, and AT, respectively, located at the 3′ end 153 of the duplex 154, with the sequences of nucleotides AT, TG, and GC, in the manner described with reference to FIGS. 1A, 1C, and 1E, respectively. *FIG. 8C illustrates an exemplary set of values ​​that may be measured under a second set of measurement conditions that differs from the first set of measurement conditions when each of the nucleotides 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 199, 200, 2000, 2001, In this non-limiting example, the second set of measurement conditions decreases the first value, increases the second value, and increases the third value relative to their values ​​under the first set of measurement conditions.

[0134] Even though sequencing the same polynucleotide 150 under two or more different sets of measurement conditions may provide different series of measurements that may be non-linearly related to the particular measurement conditions used, it is expected that the measurements will be so reproducible and accurate that each such series of measurements can be reliably used to identify the sequence of nucleotides in polynucleotide 150. Indeed, in a manner such as that described with reference to Figure 9, polynucleotide 150 may be sequenced multiple times under different sets of measurement conditions and / or using multiple sets of conditions during a single round of sequencing polynucleotide 150, and the sequences compared to each other to further refine the accuracy of the sequence.

[0135] As further described above, aspects of the present system and method address the issue of homopolymers. One such aspect is the contribution of the signal from the double strand as well as the signal from the surrounding nucleotides to a particular value of the signal that can be obtained within a given portion of the polymer. For example, FIG. 8D illustrates an exemplary series of values ​​that can be measured for a polynucleotide sequence CAAAG in polynucleotide 150, where the nucleotides TTT are added sequentially such that the base pair GC is first located at the 3' end of the double strand, the base pairs TA, TA, TA are sequentially located at the 3' end of the double strand, and the polynucleotide sequence AA, AA, and AG are sequentially located within the single stranded second portion 156 of polynucleotide 150. As illustrated in FIG. 8D, under a first set of measurement conditions, a first value is measured for the combination of GC and AA, a second value is measured for the combination of TA and AA, a third value is measured for the combination of TA and AA, and a fourth value is measured for the combination of TA and AG. Particular types of nucleotides in the sequence of nucleotides at the 3' end 153 of duplex 154 and located in the single-stranded second portion 156 of polynucleotide 150 may affect the ionic current (and / or fluctuations in ionic current) through nanopore 110 and therefore may affect measurements. Indeed, even particular types of paired nucleotides within duplex 154 and away from the 3' end 153 of duplex 154 and / or particular types of unpaired nucleotides within polynucleotide 150 and away from the single-stranded second portion 156 of polynucleotide 150 may affect the ionic current (and / or fluctuations in ionic current) through nanopore 110 and therefore may affect measurements.

[0136] For example, during the first and second measurements, AA is located within the single-stranded second portion 156 of polynucleotide 150, but the presence of the different base pairs GC and TA at the 3' end 153 of duplex 154 provides a signal contribution that allows the first AA to be easily distinguished from the second AA. Furthermore, during the second and third measurements, TA is located at the 3' end 153 of duplex 154 and AA is located within the single-stranded second portion 156 of polynucleotide 150, but the presence of different base pairs within duplex 154 and away from the 3' end 153 of duplex 154, and / or different unpaired nucleotides within polynucleotide 150 and away from the single-stranded second portion 156 of polynucleotide 150 provides a signal contribution that allows the first TA and AA combination to be easily distinguished from the second TA and AA combination.

[0137] It will therefore be appreciated that signal contributions from different parts of the duplex 154 and / or from additional unpaired bases in the sequence of the polynucleotide 150 can be used to distinguish different nucleotides in a homopolymeric sequence from one another. In this regard, referring back to Figures 1F and 1G, the larger the values ​​of M and / or D, the longer the sequence or "word" that can be read at a given time and the longer the homopolymeric stretch that can be read reliably, since more unpaired nucleotides and / or duplex base pairs can contribute to the signal by which the nucleotides of the homopolymer can be distinguished from one another. It will be appreciated that the series of measurements illustrated in Figure 8D can be repeated under different sets of measurement conditions to obtain different sets of values ​​that can distinguish the nucleotides in a homopolymeric sequence from one another in different ways.

[0138] Different combinations of nucleotides and / or double-stranded base pairs may affect the magnitude of the ionic current through the nanopore, but such combinations may also, or alternatively, affect the fluctuation or noise of the ionic current through the nanopore. The circuit 160 may measure such fluctuation or noise and use the measurements of those fluctuations or noise to identify the nucleotides. Thus, the fluctuations may be considered to be electrical properties of the 3' end 153 of the duplex 154 and the second portion 156 of the polynucleotide 150, but to facilitate consideration of certain distinctions between information obtained from such fluctuations and information obtained from the measurement magnitude, such fluctuations may sometimes be referred to as "standard deviations" in another measurement magnitude. For example, FIG. 8E illustrates an exemplary series of standard deviation values, which may be measured for the same series of values ​​described with reference to FIG. 8D. As illustrated in Figure 8E, under a first set of measurement conditions, a first standard deviation value is measured for the combination of GC and AA, a second standard deviation value is measured for the combination of TA and AA, a third standard deviation value is measured for the combination of TA and AA, and a fourth standard deviation value is measured for the combination of TA and AG. A particular type of nucleotide at the 3' end 153 of the duplex 154 and / or a particular type of nucleotide in the sequence of nucleotides located in the single-stranded second portion 156 of the polynucleotide 150 may affect the fluctuation of the ionic current through the nanopore 110 and thus affect the measured standard deviation value. Indeed, even a particular type of nucleotide within the duplex 154 and away from the 3' end 153 of the duplex 154 and / or a particular type of unpaired nucleotide within the polynucleotide 150 and away from the single-stranded second portion 156 of the polynucleotide 150 may affect the fluctuation of the ionic current and thus affect the measured standard deviation value. Thus, the standard deviation of the measurements may be used alone or in conjunction with the magnitude of the measurements to identify the sequence of polynucleotide 150.It will be understood that the series of measurements illustrated in FIG. 8E can be repeated under different sets of measurement conditions to obtain different sets of standard deviation values ​​that can distinguish nucleotides in a homopolymer sequence (or any other sequence) from one another in different ways.

[0139] For example, measurements such as those described with reference to Figures 8A-8E as provided herein can be used to generate a data structure, such as an N-dimensional "read map" that correlates measurements (each of which may have any desired level of precision and may be obtained under any suitable set of measurement conditions) with different sequence combinations within duplex 154 and the second portion of polynucleotide 150. For example, Figure 9 illustrates an exemplary N-dimensional read map of electrical features that may be used to identify nucleotides using the system of Figures 1A-1H. In the non-limiting example illustrated in Figure 9, the read map 900 includes a first axis corresponding to a first type of measurement, a second axis corresponding to a second type of measurement (e.g., the standard deviation of measurements of the same type as the first axis), and a third axis corresponding to a given measurement condition that may vary between different measurements. Illustratively, the measurement conditions may be or may include the application of a first force F1 as described with reference to Figures 1A, 1C, 1E, and 1F, and the measurement conditions may be varied (along the respective axis of Figure 9) by applying a modified first force F1' as described with reference to Figure 1G that positions the 3' end 153 of the duplex 154 differently relative to the nanopore 110, thus causing a change in one or more measurements and / or noise or ionic current variation, e.g., standard deviation, of the same or different measurements.

[0140] The N-th order read map 900 illustrated in Figure 9 may be generated using a calibration procedure in which a sufficient number of different polynucleotides (whose sequences are known a priori) are sequenced using measurement and nucleotide addition operations in a manner as described with reference to Figures 1A-1E. Such a calibration procedure may be performed on a system-by-system basis, or may be performed such that the points in the read map 900 apply approximately equally to each of a plurality of systems, such that each system does not need to be individually calibrated to generate its own read map. For example, in a first measurement condition, the circuit 160 may measure values ​​of one or more electrical properties of the 3' end 153 of the duplex 154 and the second portion 156 of the polynucleotide 150, such as the current, ionic current, electrical resistance, or voltage drop across the nanopore 110, and / or the standard deviation of one or more of such electrical properties, during the nucleotide addition process. The circuit 160 can store the measurements and measurement conditions in a non-volatile computer readable medium (e.g., memory), and this set of information can be considered to populate a first plane of points into the read map 900. Then, at a second (different) measurement condition, the circuit 160 can measure values ​​of one or more electrical properties of the 3' end 153 of the duplex 154 and the second portion 156 of the polynucleotide 150 during the nucleotide addition process, and / or the standard deviation of one or more of such electrical properties. The circuit 160 can store the measurements and measurement conditions in a computer readable medium (e.g., memory), and this set of information can be considered to populate a second plane of points into the read map 900. Such operations of measuring and storing the measurements and respective measurement conditions can be repeated any suitable number of times to provide a read map 900 having a desired number of dimensions.

[0141] For each point stored in the read map 900, the circuit 160 may also store the identity and respective position in the sequence of nucleotides known to have contributed to the signal since the sequence of the polynucleotide is known a priori. For example, for each point stored in the read map 900, the circuit 160 may store the identity and position in the sequence of at least nucleotides 4, 4', 5, and 6 illustrated in FIG. 1F. Optionally, the circuit 160 may also store the identity of nucleotides 3 and 3'. As a further option, the circuit 160 may also store the identity and position in the sequence of nucleotides 2 and 2'. As a further option, the circuit 160 may also store the identity and position in the sequence of nucleotides 1 and 1'. Additionally or alternatively, the circuit 160 may also store the identity and position in the sequence of nucleotide 7. As a further option, the circuit 160 may also store the identity and position in the sequence of nucleotide 8. Non-limiting examples of the numbers and positions of nucleotides that may contribute to the measured signal and therefore be stored in the read map 900 are described with reference to Figures 1F-1G.

[0142] As a purely illustrative example, for a point 911 in the read map defined by the measured value and standard deviation of the current through the nanopore 110 at a first set of measurement conditions (e.g., a first force F1) illustrated in FIG. 1A, and for a point 911' corresponding to the same measurement at a second set of measurement conditions (e.g., a modified first force F1' described with reference to FIG. 1G), the circuit 160 can store the nucleotides and their positions in a format such as GCAT, where the first enumerated nucleotide corresponds to nucleotide 4' located at the 3' end 153 of the duplex 154, the first enumerated nucleotide corresponds to nucleotide 4 which is a base pair of nucleotide 4', the third enumerated nucleotide corresponds to nucleotide 5 which is unpaired and adjacent to the base pair, and the fourth enumerated nucleotide corresponds to nucleotide 6 which is unpaired and adjacent to nucleotide 5, as defined by conventional techniques. Similarly, for a point 912 in the read map defined by the measurement and standard deviation of the current through nanopore 110 at a first set of measurement conditions (e.g., a first force F1) illustrated in Figure 1C, and for a point 912' corresponding to the same measurement at a second set of measurement conditions (e.g., a modified first force F1' described with reference to Figure 1G), circuitry 160 may store the nucleotides and their positions in a format such as TATG using the same conventional techniques. Of course, any other suitable format may be used, and any suitable number of nucleotides may be included, and their identity and positions may be indicated in any suitable manner.

[0143] It should be noted that methylated bases or other modified nucleotides may be included in the polynucleotide at positions known a priori, and a calibration procedure may be performed. In a manner as described with reference to Figures 1H and 8C, one or more of the values ​​measured from a sequence including such modified nucleotides may differ from values ​​measured from a sequence including unmodified nucleotides. The nucleotide identities stored by circuit 160 may include a suitable indication of whether and how the nucleotide is modified. For example, point 913 in read map 900 may contain an AT at the 3' end, a G at position 5, and a methylcytosine (C) at position 6. * ), the point 913 may be defined by the measured value and standard deviation of the current through the nanopore 910 at a first set of measurement conditions (e.g., a first force F1). It may be seen that point 913 may have a different position in the read map 900 from another point where the cytosine is unmodified, as well as from yet another point where the cytosine includes a different type of modification.

[0144] It will be appreciated that including multiple measurement axes (e.g., both measurements and standard deviations) in the read map 900 can increase the accuracy with which the circuit 160 identifies nucleotides, although in some instances only a single such axis corresponding to a single type of measurement is included in the read map 900. Conversely, the read map 900 can include any suitable number of dimensions, e.g., axes corresponding to any suitable number and type of measurements, standard deviations of the measurements, measurement conditions (fluid composition such as temperature, salt concentration, and / or pH, position within the flow cell, etc.), type of nanopore (e.g., variants thereof), type of polymerase (e.g., variants thereof), nucleotide modifications, etc.

[0145] The read map 900 (or other suitable data structure as described elsewhere herein) may be stored in a non-volatile computer readable memory in operative communication with the circuit 160, which may use the read map repeatedly to identify individual nucleotides in the sequence of the polynucleotide 150, for example, as nucleotides are added to the 3' end 153 of the duplex 154. For example, in a manner as described with reference to Figures 1A, 1C, and 1E, the circuit 160 may measure values ​​of one or more electrical properties of the 3' end 153 of the duplex 154 and the single-stranded second portion 156 of the polynucleotide 150 under a predetermined set of measurement conditions. In some examples, the circuit 160 may be programmed to use a set of measurement conditions for which the read map 900 contains points to facilitate comparison of values ​​measured for an unknown sequence to values ​​previously measured for sequences that were known a priori. In some examples, circuit 160 can find a set of points (e.g., a plane) in the read map that correspond to the set of measurement conditions, compare the measurements to a corresponding set of values ​​in the read map, and select a value or combination of values ​​in that set that is closest to the measurements based on such comparison. For the selected value, circuit 160 can then retrieve the identity and location of the nucleotide that generated the value during the calibration process. Illustratively, circuit 160 can determine that the measurements for the unknown polynucleotide sequence are closest in magnitude to the value for point 911 in read map 900. Based on such comparison, circuit 160 can determine that the unknown polynucleotide sequence includes the base pairs GC (positions 4' and 4, respectively) at the 3' end 153 of duplex 154, an A at position 5, and a T at position 6. Such a determination can be referred to as a "base call."

[0146] However, it should be noted that circuit 160 is not limited to using a single point in read map 900 to make a base call, although that is an option. Instead, circuit 160 may use multiple points in the read map to significantly improve the accuracy of the base call. For example, as can be seen from comparing Figures 1A, 1B, and 1C, a nucleotide added to the 3' end 153 of double strand 154 shifts single-stranded second portion 156 of polynucleotide 150 up by a single nucleotide. Thus, the sequence of single-stranded second portion 156 in Figure 1A and the sequence of single-stranded second portion 156 in Figure 1C overlap each other by at least one nucleotide, and in this example, the T shifts from position 6 in Figure 1A to position 5 in Figure 1C. The circuit 160 can compare the base calls from the measurements made at the time of FIG. 1A with the base calls from the measurements made at the time of FIG. 1C to ascertain whether the same nucleotide is present but the position is shifted by a single nucleotide, such as when (i) a single nucleotide is added in the operation illustrated in FIG. 1B, and (ii) the base calls made for FIG. 1A and FIG. 1C are both correct. Based on such comparison, if the circuit 160 determines that the base calls between the nucleotide additions match each other (e.g., include a sequence shifted by one nucleotide), the circuit 160 can proceed with the sequencing process. On the other hand, if the circuit 160 determines that the base calls do not match each other, the circuit can flag this portion of the sequence as containing an error, attempt to make the base call again, or resequence the polynucleotide in a manner as described elsewhere herein. Note that for sufficiently long homopolymer stretches, there may not necessarily be a change in the signal. In this case, other information provided by the present system and method can still be used to confirm the individual additions of nucleotides in a manner as described elsewhere herein.

[0147] Circuit 160 may also, or alternatively, use multiple measurements to find the closest point in read map 900. For example, circuit 160 may make another base call on the same sequence using a second, different set of measurement conditions where read map 900 contains points to facilitate comparison of values ​​measured for an unknown sequence to values ​​previously measured for a sequence that was known a priori. Circuit 160 may impose the second set of measurement conditions immediately after imposing the first set of measurement conditions, for example, applying a first force F1, and then applying a modified first force F1' before adding another nucleotide. Alternatively, circuit 160 may impose the second set of measurement conditions while resequencing polynucleotide 150 in a manner as described elsewhere herein. Circuit 160 may compare base calls from measurements made using the first set of measurement conditions to base calls from measurements made using the same set of measurement conditions to see if the same nucleotide is present at the same position in both base calls, which would be the case if both base calls were correct. For example, base calls made using points 911 and 911' in read map 900 match, while a base call made using point 911 and a base call made using point 912' match. Based on such a comparison, if circuit 160 determines that the base calls for the different measurement conditions match each other (e.g., contain the same sequence), circuit 160 may proceed with the sequencing process. On the other hand, if circuit 160 determines that the base calls do not match each other, the circuit may flag this portion of the sequence as containing an error, may attempt to make one or both of the base calls again, or may resequence the polynucleotide in a manner as described elsewhere herein.

[0148] Although the read map 900 is illustrated for discussion purposes, it should be understood that the points in the read map may be suitably stored in any suitable format in a non-volatile computer readable medium that correlates a given measurement condition with one or more values ​​measured under that condition and the combination of double-stranded and single-stranded sequences used to generate those values. A look-up table (LUT) is a non-limiting example of a format that may be used to store the correlation between the measurements and the known combinations of double-stranded and single-stranded sequences used to generate those values. However, it should be understood that any suitable data structure may be used to store the correlation between the measurements and the known combinations of double-stranded and single-stranded sequences used to generate those values. For example, the data structure may be generated by a machine learning algorithm and suitably stored for use by the machine learning algorithm. For example, the data structure may be generated by training a machine learning algorithm to recognize values ​​obtained under each given set of measurement conditions that are known a priori to correspond to each combination of double-stranded and single-stranded sequences, and may be stored in any suitable format in a non-volatile computer readable medium. The data structure may then be used by a trained machine learning algorithm implemented by circuit 160 to generate an output that identifies nucleotides within a sequence of polynucleotide 150 based on inputs of values ​​measured during the nucleotide addition process as described with reference to Figures 1A-1H. It will be understood that circuit 160 may be used to train the machine learning algorithm or a different circuit may be used to train the machine learning algorithm that is then implemented by circuit 160. In one particular example, the data structure may be generated by a neural network, such as a deep learning algorithm, and appropriately stored for use by the neural network.For example, the data structure can be generated by training a neural network to recognize values ​​obtained under each given set of measurement conditions and known a priori to correspond to each combination of double-stranded and single-stranded sequences, and can be stored in any suitable format in a non-volatile computer-readable medium. Thus, the data structure can include neurons of a neural network (e.g., a deep learning algorithm). The data structure can then be used by the trained neural network (e.g., a deep learning algorithm) implemented by circuit 160 to generate an output that identifies nucleotides in a sequence of polynucleotide 150 based on an input of values ​​measured during the nucleotide addition steps as described with reference to Figures 1A-1H. It will be understood that circuit 160 can be used to train the neural network (e.g., a deep learning algorithm) or a different circuit can be used to train the neural network (e.g., a deep learning algorithm) that is then implemented by circuit 160.

[0149] It should be noted that the operations as described herein, such as measuring values, adding nucleotides, and identifying nucleotides, can be performed using any suitable combination of hardware and software. For example, FIG. 10 illustrates an exemplary circuit 160 for sequencing an unknown polynucleotide in a manner as provided herein. The circuit 160 can include a processor 1040 and at least one computer-readable medium 1030. The computer-readable medium 1030 can store a plurality of measurements 1031 of electrical properties of the 3' end of a duplex comprising a first portion of an unknown polynucleotide and a second portion of a known polynucleotide within the opening of the nanopore. The computer readable medium 1030 can store a data structure 1032 (e.g., a read map 900 or a trained neuron) that correlates different measurements with different combinations of nucleotides within the 3' end 153 of the duplex 154 comprising a first portion 155 of the known polynucleotide and a second portion 156 of the known polynucleotide within the opening 113 of the nanopore 110. Data structure 1032 may further identify the respective measurement conditions under which the measurements were taken, e.g., the magnitude of the bias voltage used to apply the first force F1 or the modified first force F1' during the measurement. Data structure 1032 may further identify the nucleotide combinations that provided each of the measurements (e.g., at least positions 4 and 4', and positions 5 and 6 of the 3' end 153 of the duplex 154 illustrated in FIG. 1F).

[0150] The computer-readable medium 1030 may also store instructions for the processor 1040 to perform operations as provided herein. For example, the instructions may be for the processor 1040 to compare the multiple measurements to values ​​in a data structure (such as the read map 900 or a trained neuron) and use the comparison to determine a sequence of nucleotides in a sequence of an unknown polynucleotide, and output a representation of the determined sequence of nucleotides. Illustratively, the instructions may be provided in a sequencing module 1033. The sequencing module 1033 may include a nucleotide addition module 1034 configured to cause the processor 1040 to add nucleotides to the 3' end 153 of the duplex 154 in a manner as described with reference to Figures 1B and 1D. For example, the circuit 160 may be operably coupled to the electrodes 102 and 103. The nucleotide addition module 1034 may cause the processor 1040 to apply a second force F2 by applying a suitable voltage bias across the electrodes 102 and 103. The sequencing module 1033 may include a measurement module 1035 configured to cause the processor 1040 to measure values ​​of an electrical property of a 3′ end 153 of a duplex 154 comprising a first portion of an unknown polynucleotide and a second portion of a known polynucleotide within the opening of the nanopore during a nucleotide addition operation in a manner as described with reference to Figures 1A, 1C, and 1E, and store the measurements in the computer readable medium 1031. For example, the measurement module 1035 may cause the processor 1040 to apply a first force F1 or a modified first force F1' by applying a suitable voltage bias across the electrodes 102 and 103 and measuring values ​​of the electrical property while applying such forces.

[0151] The circuit 160 may include one or more sensors 1010, each configured to measure a value of one or more electrical properties. Each sensor 1010 may be configured to measure, for example, a voltage drop across the nanopore 110, a current through the nanopore 110, or an electrical resistance across the nanopore 110, or an intensity of light from a dye whose emission varies in response to the amount of ionic current through the nanopore 110. Optionally, each sensor 1010 may be further configured to measure a standard deviation of the metric it is measuring, for example, a standard deviation of the voltage drop across the nanopore 110, a standard deviation of the current through the nanopore 110, or a standard deviation of the electrical resistance across the nanopore 110, or a standard deviation of the intensity of light from a dye whose emission varies in response to the amount of ionic current through the nanopore 110. Alternatively, the measurement module 1035 may be configured to determine the standard deviation of such a metric based on a statistical analysis of the measurements 1031. The measurement value 1031 may further identify the measurement conditions under which the measurement value was obtained, for example, the magnitude of the bias voltage used to apply the first force F1 or the modified first force F1' during the measurement.

[0152] In examples where the measurements are electrical in nature, the sensor 1010 may include one or both of the electrodes 102 and 103, or may include one or more additional electrodes or other circuit components in contact with the fluid 120 and / or the fluid 120', respectively. In examples where the measurements are electrical in nature, the circuit components in contact with the fluid 120 and / or the fluid 120' may include field effect transistors configured to sense a voltage drop across the fluid 120 and the fluid 120'. In examples where the measurements are optical in nature, the sensor 1010 may include any suitable photodetector, such as an active-pixel sensor (APS) including an array of amplified photodetectors configured to generate an electrical signal based on light received by the photodetectors. The APS may be based on complementary metal oxide semiconductor (CMOS) technology known in the art. The CMOS-based detector may include a field effect transistor (FET), for example, a metal oxide semiconductor field effect transistor (MOSFET). In a particular example, a CMOS imager having a single-photon avalanche diode (CMOS-SPAD) may be used, for example to perform fluorescence lifetime imaging (FLIM). In other examples, the photodetector may include an avalanche photodiode, a charge-coupled device (CCD), a cryogenic photon detector, a reverse-biased light emitting diode (LED), a photoresistor, a phototransistor, a photocell, a photomultiplier tube (PMT), a quantum dot photoconductor, or a photodiode. Each sensor 1010 generates an electrical signal corresponding to a measurement and provides the signal to memory 1030 for storage at 1031.Optionally, the circuit 160 includes an amplifier 1020 each configured to amplify the electrical signal from each sensor 210 before storing the signal at 1031, and as a further option, the amplifier 1020 may be included within each sensor 1010.

[0153] The sequencing module 1033 may also include a nucleotide identification module 1036 configured to cause the processor 1040 to identify nucleotides by comparing the measurements 1031 to values ​​in the data structure 1032, e.g., in a manner as described with reference to Figures 8A-8E and 9. For example, the nucleotide identification module 1036 may cause the processor 1040 to use the measurements from 1031 (illustratively in an order corresponding to the chronological order in which they were obtained) and at least a portion of the data structure 1032 that was obtained under the same operating conditions as the measurements. For example, the nucleotide identification module 1036 may cause the processor 1040 to compare the identification of the measurement conditions stored in the measurements 1031 to the identification of the measurement conditions stored in the data structure 1032 and to ignore any values ​​in the data structure 1032 that do not match the measurement conditions of the measurements 1031. Nucleotide identification module 1036 may cause processor 1040 to perform operations to compare the measurements to values ​​in the data structure, for example, by taking the difference between the measurements and the values ​​in the data structure, by taking a ratio between the measurements and the values ​​in the data structure, by performing a statistical comparison such as a T-test, etc. Nucleotide identification module 1036 may cause processor 1040 to select the value in data structure 1032 that is most similar to the measurements based on the comparison, and to select from data structure 1032 the combination of nucleotides that previously provided the measurements (e.g., at least positions 4 and 4', and positions 5 and 6 of the 3' end of the duplex illustrated in FIG. 1F).

[0154] The nucleotide identification module 1036 can cause the processor 1040 to use the selected combination of nucleotides to construct an electronic sequence of nucleotides that corresponds to the sequence of physical nucleotides in the unknown polynucleotide 150 to which the measurements 1031 correspond. For example, using the labels illustrated in FIG. 1F, each measurement 1031 includes contributions from base pair 4, 4', and unpaired nucleotides 5 and 6. Thus, in some examples, the nucleotide identification module 1036 can cause the processor 1040 to include nucleotides at positions 4, 5, and 6 in the electronic sequence of nucleotides that correspond to the sequence of physical nucleotides in the unknown polynucleotide 150. In this regard, at least some of such nucleotides may already be included in the electronic sequence, since they also contributed to the measured value in the immediately preceding measurement step, i.e., before the addition of the single nucleotide, but at a position shifted by a single nucleotide. Thus, the electronic sequence may already include nucleotides that are now at positions 4 and 5, but that were at positions 5 and 6 in the previous measurement step. The nucleotide at position 6 can now be added to the electronic sequence; for example, if it was at position 7 in the previous measurement step, it may not have contributed enough to the value in the previous measurement step to be identifiable during that step, but is now identifiable. Nucleotide identification module 1036 may cause processor 1040 to output the electronic sequence of the nucleotide, for example, by storing the sequence in memory 1030, electronically transmitting the sequence to another device or system (not specifically illustrated), displaying the sequence or a portion thereof on a display screen (not specifically illustrated) operably coupled to circuitry 160, etc.

[0155] The nucleotide identification module 1036 may optionally be configured to have the processor 1040 use measurements of multiple types of values ​​to identify a nucleotide or confirm the identity of a nucleotide. For example, in a manner as described with reference to Figures 8A-8E and 9, for any given combination of nucleotides, the data structure 1032 may optionally include different values ​​of a particular type, each corresponding to a different measurement condition. Additionally or alternatively, for any given combination of nucleotides, the data structure 1032 may optionally include measurements of different types. The nucleotide identification module 1036 may have the processor 1040 use any combination of values ​​obtained using different types of measurements and / or different measurement conditions to identify a nucleotide or confirm the identity of a nucleotide. As one illustrative example, the nucleotide identification module 1036 may have the processor 1040 use both a given value and a standard deviation of that value or different values ​​to identify a point in the data structure 1032 that corresponds to a particular combination of a double-stranded base pair and an unpaired nucleotide. As another illustrative example, the nucleotide identification module 1036 can cause the processor 1040 to use values ​​from two other different types of measurements to identify a point in the data structure 1032 that corresponds to a particular combination of a double-stranded base pair and an unpaired nucleotide.

[0156] In some examples, the nucleotide identification module 1036 may additionally or alternatively be configured to have the processor 1040 verify the accuracy of the identification using at least the immediately preceding or subsequent measurement steps, if not earlier and / or later measurement steps. For example, in the manner described with reference to FIGS. 8A-8E, if for a given measurement step the processor 1040 identifies nucleotides A, C, and G at positions 4, 5, and 6, and if such identification is accurate, then for the next measurement step (after adding a single nucleotide using the nucleotide addition module 1034), the processor should identify nucleotides C and G at positions 4 and 5. On the other hand, if either the current or previous identification is inaccurate, then one or both of the nucleotides at positions 4 and 5 in one measurement step may not match the nucleotides at positions 5 and 6 in the immediately preceding measurement step. To increase the accuracy of the resulting electronic sequence, the nucleotide identification module 1036 can cause the processor 1040 to compare the identified nucleotide for a given measurement step with the identified nucleotide for at least one other (e.g., earlier or later) measurement step and indicate (e.g., flag) an error based on any difference between identified nucleotides that would have been the same as one another.

[0157] In response to an indication of a nucleotide identification error, the nucleotide identification module 1036 can cause the processor 1040 to take one or more corrective actions. For example, the nucleotide identification module 1036 can cause the processor 1040 to ignore one or more nucleotide identifications that are incorrect (e.g., because each given nucleotide may contribute to three or more consecutive measurements), for example, by replacing the incorrect identifications with identifications known to be correct from other measurement steps that match each other, or by indicating that the nucleotide has not been identified (e.g., using a nonce character such as "X" to indicate that the identity of the nucleotide is unknown). As another example that may be implemented in addition to or as an alternative to another corrective action, the nucleotide identification module 1036 can cause the processor 1040 to attempt to identify the nucleotide again using the stored measurements 1031 and data structure 1032. As another example that may be implemented in addition to or as an alternative to another corrective action, the nucleotide identification module 1036 may cause the processor 1040 to attempt to identify the nucleotide again by obtaining a new measurement 1031 using different measurement conditions (e.g., a modified first force F1') and a data structure 1032. As another example that may be implemented in addition to or as an alternative to another corrective action, the nucleotide identification module 1036 may cause the processor 1040 to attempt to identify the nucleotide again by obtaining a new measurement 1031 using a different measurement type (e.g., a standard deviation of the first measurement type used) and a data structure 1032; in this regard, it should be noted that a separate measurement step does not necessarily have to be performed, for example, since the standard deviation can be obtained from the stored measurement 1031. As another example that may be implemented in addition to or as an alternative to another corrective action, the nucleotide identification module 1036 may cause the processor 1040 to resequence the unknown polynucleotide 150 in a manner as described elsewhere herein.As another example that may be implemented as an alternative to such corrective actions, or if such corrective actions have been attempted but unsuccessful, the nucleotide identification module 1036 may cause the processor 1040 to indicate in the electronic sequence that a nucleotide was not identified (e.g., using a nonce character such as "X" to indicate that the identity of the nucleotide is unknown). In some examples, the nucleotide identification module 1036 may cause the processor 1040 to select among these or other actions based on the apparent nature of the error and the options available when the error was identified.

[0158] It should be noted that in some examples, the data structure 1032 and the nucleotide identification module 1036 may be suitably implemented using machine learning. For example, the data structure 1032 may be generated by training any suitable machine learning algorithm, such as a neural network (e.g., a deep learning algorithm), using the measurements, the nucleotide combinations known a priori to correspond to those measurements, and the measurement conditions under which those measurements were obtained. In this regard, the data structure 1032 may have a configuration that is readily usable by a trained machine learning algorithm, such as a trained neural network, such as a trained deep learning algorithm (the nucleotide identification module 1036 implemented by the processor 1040), to identify nucleotide combinations using the measurements, but such a configuration may not necessarily be usable by any other software, module, or algorithm to determine correlations between measurements and unknown combinations of nucleotides. For example, a machine learning algorithm (such as a neural network, e.g., a deep learning algorithm) may be trained to make base calls using the signal output of the circuit 160. Non-limiting examples of machine learning algorithms are supervised, semi-supervised, unsupervised, and reinforcement algorithms. Neural network algorithms are a subset of machine learning algorithms and may include deep learning algorithms, convolutional neural networks, recurrent neural networks, generative adversarial networks, and recurrent neural networks. Thus, the particular configuration of the data structure 1032 may include, for example, a vector space, a graph space, neurons of a neural network, etc. Alternatively, the data structure 1032 may be implemented using any suitable data structure that may be queried using the nucleotide identification module, such as a look-up table (LUT), a matrix, a flat file database structure, an SQL database structure, etc.The nucleotide identification module 1036 can suitably cause the processor 1040 to identify points in the data structure 1032 having measurements for known combinations of nucleotides that correspond to measurements for unknown combinations of nucleotides.

[0159] It should be understood that the circuitry 160 may be implemented using any suitable combination of digital electronic circuitry, integrated circuits, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), central processing units (CPUs), graphical processing units (GPUs), computer hardware, firmware, software, and / or combinations thereof. For example, one or more functionality of the circuitry 160 may be implemented in one or more computer programs, special purpose or general purpose, executable and / or interpretable on a programmable system including at least one programmable processor coupled to receive data and instructions from, and transmit data and instructions to, a storage system, at least one input device, and at least one output device. The programmable system or computing system may include clients and servers. Clients and servers are generally remote from each other and typically interact through a communications network. The relationship of clients and servers arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.

[0160] These computer programs, which may also be referred to as modules, programs, software, software applications, applications, components, or code, may include machine instructions for a programmable processor and / or may be implemented in a high-level procedural language, an object-oriented programming language, a functional programming language, a logic programming language, and / or an assembly / machine language. As used herein, the term "machine-readable medium" "computer-readable medium" refers to any computer program product, apparatus, and / or device, such as a magnetic disk, optical disk, memory, programmable logic device (PLD) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable data processor. A computer-readable medium may store such machine instructions non-transiently, such as a non-transient solid-state memory or a magnetic hard drive or any equivalent storage medium. The computer-readable medium may alternatively or additionally be capable of temporarily storing such machine instructions, such as a processor cache or other random access memory associated with one or more physical processor cores.

[0161] The computer components, software modules, functions, data stores, and data structures described herein may be directly or indirectly connected to each other to enable the flow of data necessary for their operation. It should also be noted that a module or processor includes, but is not limited to, a unit of code that performs software operations, and may be implemented, for example, as a subroutine unit of code, or as a software functional unit of code, or as an object (as in an object-oriented paradigm), or as an applet, or in a computer scripting language, or as another type of computer code. Software components and / or functionality may be located on a single computer or distributed across multiple computers and / or clouds, depending on the situation at hand.

[0162] In one non-limiting example, the circuit 160 described with reference to FIG. 10 may be implemented using a computing device architecture. In such an architecture, a bus (not specifically illustrated) may function as an information highway interconnecting the other illustrated components of the hardware. The system bus may also include at least one communication port (such as a network interface) that allows communication with external devices that are physically connected to the computing system or externally available through a wired or wireless network. The processor 1040 may be implemented using a CPU (Central Processing Unit) (e.g., one or more computer processors / data processors in a given computer or computers) that can perform the calculations and logical operations necessary to execute a program. The memory 1030 may include a non-transitory processor-readable storage medium, such as a read only memory (ROM) and / or a random access memory (RAM), in communication with the processor 1040, and may include one or more programming instructions for the operations provided herein, such as the sequencing module 1033 and its components, and may store the measurements 1031 and data structures 1032. Optionally, memory 1030 may include a magnetic disk, an optical disk, a recordable memory device, a flash memory, or other physical storage medium. To provide for interaction with a user, circuitry 160 may include or be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying obtained information to a user, and a keyboard and / or pointing device (e.g., a mouse or trackball), and / or a touch screen by which a user can provide input to the computer.

[0163] It should be noted that the sequencing module 1033 may be further configured to cause the processor 1040 to perform additional operations, such as those described with reference to Figures 2A-2E, 3, 4A-4C, 5A-5B, and 6. Such operations may be performed using additional modules not specifically illustrated in Figure 10, or may be performed using suitable modifications to the nucleotide addition module 1034, the measurement module 1035, and / or the nucleotide addition module 1036.

[0164] The systems, configurations, and operations described with reference to Figures 1A-1H may be suitably modified to further enhance the accuracy of measurements made. For example, Figures 2A-2E generally illustrate the use of an exemplary sequencing system, as well as additional exemplary compositions and operations, for sequencing polynucleotides using a nanopore. The system 200 described herein with reference to Figures 2A-2E may be configured and used similarly to the system 100 described with reference to Figures 1A-1H. For example, the system 200 may similarly include a barrier 101 including first and second electrodes 102, 103, optionally first and second layers 107, 108, a nanopore 110, polynucleotides 140, 150, a polymerase 105, nucleotides 121, 122, 123, 124, and a circuit 160, each of which may be configured similarly as described with reference to Figures 1A-1H. However, duplex 254 may be modified relative to duplex 154 in that blocking moiety 234 may reversibly associate with the duplex and reversibly inhibit polymerase 105 from adding a second nucleotide. For example, at a particular time illustrated in FIG. 2A, circuit 160 may apply a second force F2 such that polymerase 105 may add nucleotide 122 to 3' terminus 253 in a manner similar to that described with reference to FIG. 1B. Blocking moiety 234 may inhibit polymerase 105 in fluid 220 from adding another nucleotide to 3' terminus 253 until blocking moiety 234 is intentionally removed. In some examples, blocking moiety 234 may include a 3'-blocking group attached to each nucleotide 121, 122, 123, 124 in fluid 220.

[0165] Circuit 160 may be configured to detect the association of blocking moiety 234 with duplex 254. For example, while applying a first force F1 as illustrated in FIG. 2B, the circuit may also measure the value of an electrical property in a manner similar to that described with reference to FIG. 1A, 1C, and 1E. As discussed with reference to FIG. 1A, 1F, 1G, and 1H, a particular base pair (e.g., TA) at 3′ end 253 of duplex 254, a sequence of one or more bases (e.g., T, G) in single-stranded second portion 156, in combination with the presence of blocking moiety 234, may alter the rate at which salt in fluid 120 moves through opening 113 into fluid 120′ under the particular measurement conditions used, and thus alter the current, ionic current, electrical resistance (resistance to ion flow), or voltage drop across nanopore 110, or the standard deviation of any such electrical property, in a manner as detected by circuit 160. Thus, circuit 160 can use the measured values ​​to detect the presence of blocking moiety 234, and can use the measured values ​​to identify at least one nucleotide in polynucleotide 150. Additionally, or alternatively, circuit 160 can detect the absence of blocking moiety 234 from duplex 254. For example, if blocking moiety 234 is not present, its absence can change the rate at which salt in fluid 120 moves through opening 113 into fluid 120', compared to the rate in the presence of blocking moiety 234, and thus change the current, the current across ionic nanopore 110, the electrical resistance (resistance to ion flow), or the voltage drop, or the standard deviation of any such electrical property, in a manner as detected by circuit 160. In this regard, it is noted that additional dimensions of data structure 1032 can further include measurements corresponding to the presence of blocking moiety 234 attached to a nucleotide at the 3' end 253 of duplex 254.

[0166] Blocking moiety 234 may be alterable or removable in a manner that allows polymerase 105 to add a second nucleotide after the presence of the blocking moiety is confirmed, for example, using operations as described with reference to FIG. 2B. In a non-limiting example illustrated in FIG. 2C, fluid 220 is replaced with a fluid that does not contain a polymerase and does not contain a blocked nucleotide, and circuit 160 applies a second force F2 such that 3' end 253 becomes accessible to altered fluid 220' that includes reactant 235 that may alter or remove blocking moiety 234, resulting in modified blocking moiety 234' that may no longer associate with 3' end 253, for example. Illustratively, in an example where blocking moiety 234 includes a 3'-blocking group, removal of blocking moiety 234 from nucleotide 221 may include replacing the blocking moiety with another chemical moiety, such as -OH group 225 as illustrated in FIG. 2D. After blocking portion 234 has been removed, circuit 160 may again apply a first force F1 that places 3' end 253 of duplex 254 (now not associated with blocking portion 234 and thus may correspond to 3' end 153 of duplex 154) within opening 112, in a manner as illustrated in FIG. 2E, and place single-stranded second portion 156 of polynucleotide 150 within the opening. While circuit 160 again applies first force F1, constriction 114 may again inhibit passage of 3' end 253 of duplex 254 to second side 112 of nanopore 110, and circuit may measure the value of an electrical property of the nanopore to confirm that the blocking portion has been removed. Fluid 220' may then be optionally replaced with another fluid, for example a fluid that does not cause the block to be removed.Whether or not fluid 220' is replaced at this point, because blocking moiety 234 is not present at 3' end 253 of duplex 254, the absence of blocking moiety may alter the rate at which salt in fluid 120 migrates through opening 113 into fluid 120' compared to the rate in the presence of blocking moiety 234, and thus alter the current, ionic current, electrical resistance, or voltage drop across nanopore 110, and / or the standard deviation of any such electrical characterization, in a manner as detected by circuit 160. Thus, nucleotide 122 may be identified using the measured values, for example, in a manner as described with reference to Figures 1A, 1C, 1E, 8A-8E, 9, and 10. Fluid 220' may be replaced with fluid 220 containing polymerase and blocked nucleotides to repeat the process of adding blocked nucleotides to the ends of duplexes 254 (FIG. 2A), generating a signal using the blocked nucleotides at the ends of the duplexes (FIG. 2B), removing blocking group 234 (FIGS. 2C-2D), and generating a signal using the unblocked nucleotides at the ends of the duplexes (FIG. 2E). Such operations may be repeated any suitable number of times, e.g., to sequence at least a portion of polynucleotide 150, e.g., to identify the added nucleotides using the signals in a manner as described elsewhere herein.

[0167] It will therefore be appreciated that measurements such as those described with reference to FIG. 2B can be used to confirm that the blocked nucleotide has been incorporated. The blocking moiety prevents the duplex 254 from being extended by any additional nucleotides until the blocking moiety 234 is removed (FIG. 2C), thereby preventing deletion errors that may otherwise result from uncontrolled and unobserved addition of multiple nucleotides. Measurements such as those described with reference to FIG. 2E can be used to confirm that the nucleotide has been properly unblocked, in addition to identifying the nucleotide. It will be appreciated that the ability to perform such repetitive measurements in an electronically controlled stepwise manner can provide a significant improvement in accuracy, as compared to, for example, the enzyme-driven translocations used in strand sequencing.

[0168] In one illustrative example, a second force F2 is used to eject the 3' end 253 of the duplex 254 from the nanopore 110 into the fluid 220 (FIG. 2A), providing a suitable time window (e.g., about 1-50 milliseconds) for the polymerase to add a nucleotide. During this time window, a blocked nucleotide may have been added. For example, it may be useful to confirm whether this blocked nucleotide was indeed added in situations where it may be difficult to identify a nucleotide combination otherwise. For example, a relatively long homopolymer region has been sequenced, and subsequent nucleotide combinations result in measurement signals that are similar to each other. Or, for example, the nucleotide identification module 1036 may determine that a measurement for an unknown combination is too similar to a number of the measurements in the data structure 1032. Information confirming that a blocked nucleotide was added may provide a useful aid to the nucleotide identification module 1036 in appropriately processing the signal. For example, in response to identifying the presence of a blocked nucleotide at the 3' end 253 of the duplex 254, the nucleotide identification module 1036 may cause the processor 1040 to include the nucleotide in the electronic sequence even if it has been flagged as having a low quality base call or as not being identified. Or, for example, the nucleotide identification module 1036 may cause the processor 1040 to implement any other suitable remedial measures, such as repeating the measurement using a modified first force F1'. Or, for example, in response to determining that the nucleotide at the 3' end 253 of the duplex 254 did not include the blocking moiety 234, the nucleotide identification module 1036 may cause the processor 1040 to assume that the polymerase did not add the nucleotide and therefore eject the 3' end of the duplex 254 again and wait again for a time window during which the polymerase may add the blocked nucleotide. Thus, a measurement showing that blocking moiety 234 is not present at the 3' end of the duplex after a time window can be usefully interpreted as meaning that the polymerase did not act on the duplex during that time window.In response to such an instruction, circuitry 160 may repeat the nucleotide addition and sequencing operations to again attempt to identify (or confirm the identity of) the next nucleotide in the sequence of polynucleotide 150.

[0169] It will be further appreciated that measurements and operations as described with reference to Figures 2A-2E may be suitably used in combination with measurements as described with reference to Figure ID, for example, to further confirm that a polymerase is acting on the 3' end of a duplex at the appropriate time. In this regard, the measurements described with reference to Figure ID may further be used to detect a blocking moiety attached to the nucleotide on which the polymerase is acting.

[0170] In the example as described with reference to Figures 2A-2E, the nucleotides 121, 122, 123, 124 may be attached to any suitable blocking moiety 234, and any suitable reactant 235 may be used to unblock the nucleotides (remove the blocking moiety) at the 3' end 253 of the duplex 254. In one non-limiting example, the blocking moiety 234 includes an azidomethyl (AZM) group (-OCH2N3), which, when present at the 3' end of the duplex 254, may be expected to alter the current through the nanopore 110 and may be removed using tris(2-carboxyethyl)phosphine (TCEP) or tris(hydroxypropyl)phosphine (THP) to leave a 3'OH group that may be extended by another blocked nucleotide using a polymerase. A variety of other blocking moieties, and reactants for removing such blocking moieties, are known in the art and may be suitably adapted for use with the subject matter of the present invention. See, for example, U.S. Patent Publication No. 2020 / 0216891 to Francais et al., the entire contents of which are incorporated by reference herein.

[0171] Although FIGS. 2A-2E illustrate an example in which reactant 235 is positioned in the fluid on the first side 111 of nanopore 110 to remove blocking moiety 234 while 3′ end 253 of duplex 254 is ejected from nanopore 110, reactant may alternatively be positioned to remove blocking moiety while 3′ end of duplex is within nanopore. Thus, the presence of blocking moiety may be measured in a manner similar to that described with reference to FIGS. 2A-2E, and removal of blocking moiety may also be observed in real time. For example, if the presence of blocking moiety affects the measurement, a change in the measurement may be expected as the moiety is removed. Confirming removal of blocking moiety via such a change is useful information because if the moiety was not removed (as reflected by no change in the measurement), when the duplex is ejected from the nanopore (which may occur prematurely whether the moiety was removed or not), circuit 160 may determine that the next “uptake” should be ignored because it was not actually an uptake from a new cycle.

[0172] It should be noted that the need for fluid circulation (exchange of one fluid for another between certain operations) may be reduced in the example as described with reference to Figures 1A-1H compared to the example as described with reference to Figures 2A-2E. For example, to dissociate blocking moiety 234 from nucleotide 122 during the operation as described with reference to Figure 2C, fluid 220 containing nucleotides and polymerase may be removed (e.g., by rinsing with an aqueous solvent) and then replaced with fluid 220' containing reagent 235. Thereafter, to add another nucleotide to unblocked nucleotide 122, fluid 220' may be removed (e.g., by rinsing with an aqueous solvent) and then replaced with fluid 220. In the example described with reference to Figures 2A-2E, multiple fluid cycles are used because the nucleotide unblocking and nucleotide addition steps occur in the same compartment (on the same first side 111 of the nanopore), and direct contact between the nucleotide unblocking and nucleotide addition components in that compartment may cause premature unblocking (e.g., before addition) of the nucleotide, which may result in less control of the nucleotide addition process. In comparison, throughout the nucleotide addition steps described with reference to Figures 1A-1H, the polymerase 105 may move freely without the need for fluid circulation. For example, Figures 1A-1H obviate the need for addition and removal of unblocking agents, since the polymerase may freely add nucleotides from the fluid 120 to the 3' end 153 of the duplex, and such nucleotides may be distinguished from one another in a manner as described with reference to Figures 8A-8E, 9, and 10, and the examples provided further below.

[0173] It will be further understood that any suitable combination of operations as described with reference to Figures 1A-1H, 2A-2E, 7, 8A-8E, 9, and 10 may be used to sequence a polynucleotide, e.g., polynucleotide 150. For example, Figure 3 illustrates an operational flow in an exemplary method for sequencing a polynucleotide. Method 300 may use a nanopore including a first side, a second side, and an opening extending through the first and second sides, e.g., in a manner as described with reference to Figures 1A-1H and 7. Method 300 may include disposing a polynucleotide through the opening of the nanopore such that a 3' end of the polynucleotide is on the first side of the nanopore and a 5' end of the polynucleotide is on the second side of the nanopore (operation 310). Exemplary operations for disposing a polynucleotide 150 through the opening of nanopore 110 in such a manner are provided below with reference to Figure 6. The method 300 may include forming a duplex with a polynucleotide on a first side of the nanopore, the duplex including a 3' end (operation 320). Such a duplex may be formed, for example, by hybridizing a first portion 155 of nucleotide 150 to a primer (polynucleotide 140 or a portion thereof) on the first side of the nanopore, for example, in a manner as described below with reference to FIG.

[0174] Method 300 may include extending the duplex at a first side of the nanopore by adding a nucleotide to the 3' end of the duplex (operation 330). For example, circuit 160 may apply a second force F2 in response to the 3' end of the duplex being located outside the opening of the nanopore such that a polymerase may act on the 3' end of the duplex to add a nucleotide thereto, e.g., in a manner as described with reference to Figures 1B, 1D, and 2A. Method 300 may include applying a first force that positions the 3' end of the duplex within the opening (operation 340). For example, circuit 160 may apply a first force F1 in a manner as described with reference to Figures 1A, 1C, 1E, 2B, and 2E. Operation 340 of method 300 may include using the nanopore to inhibit translocation of the 3' end of the extended duplex to a second side of the nanopore while applying the first force (operation 341). For example, constriction 114 or other features of nanopore 110 may inhibit passage of 3' end 153 of duplex 154, or passage of 3' end 253 of duplex 254, from first side 111 of nanopore 110 to second side 112 of nanopore 110 in a manner as described with reference to Figures 1A, 1C, 1E, 2B, and 2E. Operation 340 of method 300 may also include measuring values ​​of electrical properties of the 3' ends of the duplexes and the single-stranded portions of the polynucleotide while applying the first force (operation 342). For example, circuit 160 may measure values ​​of such electrical properties in a manner as described with reference to Figures 1A, 1C, 1E, 2B, and 2E. Method 300 may include identifying the added nucleotides (operation 350) using the values ​​measured in operation 340. For example, circuit 160 may identify one or more nucleotides in polynucleotide 150 in a manner as described with reference to Figures 8A-8E, 9, and 10. As shown in Figure 3, operations 330-350 may be repeated any suitable number of times, for example, to substantially sequence polynucleotide 150.

[0175] It should be noted that the operations as described with reference to method 300 are compatible with and may be used in conjunction with any other operations as provided herein. For example, the added nucleotides may optionally include respective blocking moieties, and such blocking moieties may optionally be removed in a manner as described with reference to Figures 2A-2E.

[0176] It will be further understood that the same nanopore may be used in multiple different series of operations as described with reference to Figures 1A-1H, 2A-2E, 3, 7, 8A-8E, 9, and 10, which may use the same polynucleotide 150 as each other or may use different polynucleotides 150 as each other. For example, Figures 4A-4C illustrate generally the use of the sequencing system of Figures 1A-1H to resequence the same polynucleotide 150, e.g., to generate a consensus read. Referring now to Figure 4A, the system 100 is shown at a point where the sequencing of the polynucleotide 150 has been substantially completed, alternatively, the sequencing of the polynucleotide 150 may be partially completed and the nucleotide identification module 1042 has caused the processor 1040 to take corrective action to resequence the polynucleotide 150. In response to the circuit 160 applying a first force F1 as shown in FIG. 4A, the polynucleotide 150 remains hybridized to the (now extended) polynucleotide 140 in a manner as described elsewhere herein. To resequence the polynucleotide 150, the circuit can be configured to apply a sufficiently high voltage F4 to dissociate the double strand 154 or dissociate the double strand 254, i.e., to dehybridize the extended polynucleotide 140 from the polynucleotide 150 in a manner as illustrated in FIG. 4B. The polynucleotide 150 can optionally be resequenced using an operation that includes hybridizing a new (shorter) polynucleotide 140', e.g., a primer, to the polynucleotide 150, e.g., in a manner as illustrated in FIG. 4C. For example, the primer 140' can be included in the fluid 120. A series of nucleotide addition and measurement operations provided herein can then be used to partially or completely sequence the polynucleotide 150 again. Thus, the new duplex formed on the first side of the nanopore can include the same first portion 155 of polynucleotide 150 described with reference to FIG. 1A and can have a different 3' end 153 that includes the end of the new polynucleotide 140'.Such a series of operations described with reference to Figures 4A-4C may be repeated any desired number of times, for example, to sequence polynucleotide 150 multiple times until a desired level of accuracy is achieved to provide a desired confidence in the sequence. Accuracy is improved by combining multiple reads of the same template, for example, where errors are random and they can be "averaged out" through the use of consensus algorithms known in the art.

[0177] After fully or partially sequencing the polynucleotide 150 any desired number of times, alternatively after sequencing the polynucleotide 150 once, or after only partially sequencing the polynucleotide 150, the circuit 160 may optionally eject the polynucleotide 150 from contact with the nanopore 110 so that the nanopore 110 may be used again with a different polynucleotide. For example, FIGS. 5A-5B illustrate generally the use of the sequencing system of FIGS. 1A-1G to prepare a nanopore for sequencing a different polynucleotide. In the manner illustrated in FIG. 5A, the circuit 160 may be configured to apply a sufficiently high force F5 (which may be greater than force F4) to dissociate the first steric lock 151 from the polynucleotide 150. Alternatively, in the manner illustrated in FIG. 5B, the circuit 160 may be configured to apply a sufficiently high force F6 to dissociate the second steric lock 152 from the polynucleotide 150. The nanopore can then be recycled, for example, by placing a different polynucleotide 150 through it, which can be sequenced (and optionally resequenced) in the manner provided herein.

[0178] Polynucleotide 150 may be positioned within nanopore 110 and optionally locked to nanopore 210 using any suitable structure and any suitable combination of operations. For example, FIG. 6 illustrates generally the use of the sequencing system of FIGS. 1A-1H to generate and use polynucleotides for sequencing or polynucleotide synthesis. In operation A illustrated in FIG. 6, first steric lock 151 may be attached to the 3′ end of polynucleotide 150, optionally while the polynucleotide is hybridized to its naturally occurring complementary strand 150′. Illustratively, the 3′ end of polynucleotide 150 may be biotinylated or otherwise functionalized, and first steric lock 151 may include a functional group (such as neutravidin or streptavidin) that binds to the functionalized (e.g., biotinylated) 3′ end of polynucleotide 150 and remains attached thereto until a sufficiently strong force is applied in a manner as described with reference to FIG. 5A. In another example, the first steric lock 151 can include an LNA or PNA that hybridizes to the polynucleotide 150 and is long enough to remain hybridized to the polynucleotide 150 until a sufficiently strong force is applied in a manner as described with reference to FIG. 5A. In operation B illustrated in FIG. 6, the first side 111 of the nanopore 110 can be contacted with the duplex 150, 150'. In operation C illustrated in FIG. 6, the circuit 160 applies a force sufficient to cause dissociation of the complementary strand 150' from the polynucleotide 150, and the 5' end of the strand 150 is translocated through the nanopore 110. During application of such force, the first steric lock 151 can inhibit the 3' end of the polynucleotide 150 from passing through the opening to the second side of the nanopore in a manner as described with reference to FIGS. 1A-1E. The second steric lock 152 can then bind to the 5' end of the polynucleotide 150. Illustratively, the second steric lock 152 can include an LNA or PNA that hybridizes to the polynucleotide 150 and is long enough to remain hybridized to the polynucleotide 150 until a sufficiently strong force is applied in the manner described with reference to FIG. 5B.In operation D illustrated in FIG. 6, polynucleotide 140 can be hybridized to polynucleotide 150 to form a duplex that can be used to sequence polynucleotide 150 in a manner as described elsewhere herein. For example, in operations E1 and E2 illustrated in FIG. 6, polymerase 105 can add nucleotide 121 to polynucleotide 140 based on the sequence of polynucleotide 150, and optionally, in operation E2 illustrated in FIG. 6, the nucleotide can be attached to blocking moiety 234. Operations for identifying nucleotides and for adding further nucleotides are provided elsewhere herein. The first 3D lock 151 added in operation A can be removable from polynucleotide 150 in a manner as described with reference to FIG. 5A. The second 3D lock 152 added in operation C can be removable from polynucleotide 150 in a manner as described with reference to FIG. 5B.

[0179] It will be further appreciated that the systems, compositions, and operations as described with reference to Figures 1A-1H, 2A-2E, 3, 4A-4C, 5A-5B, 6, 7, 8A-8E, 9, and 10 may be suitably adapted for use in various methods of synthesizing polynucleotides, including, but not limited to, sequencing-by-synthesis (SBS).

[0180] From the above, it should be understood that the present disclosure provides what is set forth in the following clauses.

[0181] Clause 1. A method is provided for sequencing a polynucleotide using a nanopore comprising a first side, a second side, and an opening extending through the first and second sides. The method includes (a) positioning a polynucleotide through the opening of the nanopore such that a 3' end of the polynucleotide is on the first side of the nanopore and a 5' end of the polynucleotide is on the second side of the nanopore. The method also includes (b) forming a duplex with the polynucleotide on the first side of the nanopore, the duplex comprising a 3' end. The method also includes (c) extending the duplex on the first side of the nanopore by adding a first nucleotide to the 3' end of the duplex. The method also includes (d) applying a first force that positions the 3' end of the extended duplex within the opening, inhibiting translocation of the 3' end of the extended duplex to a second side of the nanopore while applying the first force, and measuring values ​​of an electrical property of the 3' end of the extended duplex and the single-stranded portion of the polynucleotide. The method also includes (e) identifying the first nucleotide using the value measured in operation (d).

[0182] Clause 2. The method of clause 1, wherein the value measured in operation (d) comprises a current, an ionic current, an electrical resistance, or a voltage drop across the nanopore.

[0183] Clause 3. The method of clause 1 or 2, wherein the value measured in operation (d) comprises noise in the current, ionic current, electrical resistance, or voltage drop across the nanopore.

[0184] Clause 4. The method of clause 3, wherein the value measured in operation (d) comprises a standard deviation of the noise.

[0185] Clause 5. The method of any one of clauses 1-4, wherein the value measured in operation (d) is based on at least M nucleotides of the single-stranded portion of the polynucleotide and D pairs of hybridized nucleotides of the extended duplex, where M is 2 or more and D is 1 or more.

[0186] Clause 6. The method of clause 5, wherein M is 3 or greater.

[0187] The method of any one of clauses 5 to 6, wherein clause 7.D is 2 or greater.

[0188] Clause 8. The method of any one of clauses 5 to 7, wherein at least one of the M nucleotides of the single-stranded portion comprises a modified base, and the method comprises using the value measured in operation (d) to identify the modified base.

[0189] Clause 9. The method of clause 8, wherein the modified base comprises a methylated base.

[0190] Clause 10. The method of any one of clauses 1-9, further comprising, in operation (d), inhibiting addition of another nucleotide to the 3' end of the extended duplex while the first force is applied.

[0191] Clause 11. The method of clause 10, wherein the nanopore inhibits addition of another nucleotide.

[0192] Clause 12. The method of any one of clauses 1-11, wherein the nanopore is oriented such that a first side of the nanopore comprises a majority of the opening.

[0193] Clause 13. The method of any one of clauses 1-11, wherein the nanopore is oriented such that the second side of the nanopore comprises a majority of the opening.

[0194] Clause 14. The method of any one of clauses 1-13, further comprising: (f) applying a modified first force that repositions the 3' end of the extended duplex within the opening and inhibiting translocation of the 3' end of the extended duplex to a second side of the nanopore while applying the modified first force; measuring values ​​of an electrical property of the 3' end of the extended duplex and the single-stranded portion of the polynucleotide; and (g) identifying the first nucleotide using the value measured in operation (f).

[0195] Clause 15. The method of any one of clauses 1-14, wherein the first nucleotide is added using a polymerase that contacts the 3' end of the duplex.

[0196] Clause 16. The method of clause 15, further comprising reversibly inhibiting the polymerase from adding a second nucleotide to the 3' end of the extending duplex.

[0197] Clause 17. The method of clause 16, wherein the blocking moiety reversibly inhibits the polymerase from adding a second nucleotide to the 3' end of the extending duplex.

[0198] Clause 18. The method of clause 17, wherein the first nucleotide is linked to a blocking moiety.

[0199] Clause 19. The method of clause 18, wherein the blocking moiety comprises a 3'-blocking group.

[0200] Clause 20. The method of clause 17, wherein the blocking moiety reversibly associates with the extended duplex.

[0201] Clause 21. The method of clause 20, further comprising detecting association of the blocking moiety with the extended duplex.

[0202] Clause 22. The method of clause 20 or 21, further comprising detecting the absence of the blocking moiety from the extended duplex.

[0203] Clause 23. The method of any one of clauses 17-22, further comprising removing the blocking moiety to allow a polymerase to add a second nucleotide to the 3' end of the extended duplex.

[0204] 24. The method of any one of clauses 15-23, wherein the first force applied in clause 24(d) removes the polymerase from contact with the 3' end of the extended duplex.

[0205] Clause 25. The method of any one of clauses 15 to 23, further comprising applying a second force to remove the polymerase from contact with the 3' end of the extended duplex, the second force being greater than the first force.

[0206] Clause 26. The method of any one of clauses 15 to 25, further comprising: (f) applying a third force that positions the polymerase in contact with the 3'-end of the duplex within or adjacent to the opening of the first side of the nanopore, inhibiting movement of the polymerase into or further into the opening using the nanopore while applying the third force, and measuring a value of an electrical property of the polymerase; and (g) using the value measured in operation (f) to identify contact of the polymerase with the 3'-end of the duplex.

[0207] Clause 27. The method of clause 26, wherein the third force is less than the first force.

[0208] Clause 28. The method of clause 26 or 27, wherein operation (f) is performed after operation (c) and before operation (d).

[0209] Clause 29. The method of any one of clauses 26-28, wherein the first nucleotide is associated with a blocking moiety.

[0210] Clause 30. The method of clause 29, wherein operation (g) further comprises confirming the presence of a blocking moiety associated with the first nucleotide using the value determined in operation (f).

[0211] Clause 31. The method of any one of clauses 15 to 30, wherein the polymerase comprises DNA.

[0212] Clause 32. The method of any one of clauses 15 to 30, wherein the polymerase comprises RNA.

[0213] Clause 33. The method of any one of clauses 15 to 30, wherein the polymerase comprises a reverse transcriptase.

[0214] Clause 34. The method of any one of clauses 1 to 33, wherein the first nucleotide is associated with a blocking moiety, and operation (e) further comprises confirming the presence of the blocking moiety associated with the nucleotide using the value measured in operation (d).

[0215] Clause 35. The method of clause 34, further comprising after act (d), removing the blocking moiety from the first nucleotide.

[0216] Clause 36. The method of clause 35, further comprising, after removing the blocking moiety, (f) again applying a first force that positions the 3' end of the extended duplex within the opening and positions the single-stranded portion of the polynucleotide within the opening, and using the nanopore to inhibit translocation of the 3' end of the extended duplex to a second side of the nanopore while again applying the first force, and measuring values ​​of the electrical properties of the 3' end of the extended duplex and the single-stranded portion of the polynucleotide, and (g) again identifying the first nucleotide using the value measured in operation (f).

[0217] Clause 37. The method of clause 36, further comprising removing the blocking moiety from the first nucleotide prior to operation (d).

[0218] Clause 38. The method of clauses 1 to 37, wherein the extended duplex comprises one or more nucleotide analogues.

[0219] Clause 39. The method of clause 38, wherein the one or more nucleotide analogues enhance the stability of the extended duplex compared to natural nucleotides.

[0220] Clause 40. The method of clause 38 or 39, wherein the one or more nucleotide analogues comprise one or more locked nucleic acids (LNA).

[0221] Clause 41. The method of any one of clauses 38 to 40, wherein the one or more nucleotide analogues comprise one or more 2'-methoxy (2'-OMe) nucleotides.

[0222] Clause 42. The method of any one of clauses 38 to 41, wherein the one or more nucleotide analogues comprise one or more 2'-fluorinated (2'-F) nucleotides.

[0223] Clause 43. The method of any one of clauses 38 to 42, wherein one or more nucleotide analogues have an altered value of an electrical property compared to a natural nucleotide.

[0224] Clause 44. The method of any one of clauses 38 to 43, wherein the first nucleotide comprises one of one or more nucleotide analogues.

[0225] Clause 45. The method of clause 45, wherein one or more nucleotide analogues comprises a 2' modification.

[0226] Clause 46. The method of clause 44 or 45, wherein one or more nucleotide analogues comprises a base modification.

[0227] Clause 47. The method of any one of clauses 1 to 46, wherein the first force is insufficient to cause dissociation of the extended duplex.

[0228] Clause 48. The method of any one of clauses 1-47, wherein the first force comprises a first voltage.

[0229] Clause 49. The method of any one of clauses 1 to 48, wherein operations (b) and (c) are performed in the absence of the first force.

[0230] Clause 50. The method of any one of clauses 1-49, wherein operation (c) is performed in the presence of a fourth force that opposes the first force.

[0231] Clause 51. The method of any one of clauses 1 to 50, wherein a first locking structure is attached to a 3' end of the polynucleotide on a first side of the nanopore, and the first locking structure inhibits translocation of the 3' end of the polynucleotide through the opening to a second side of the nanopore.

[0232] Clause 52. The method of clause 51, wherein the first locking structure is removable.

[0233] Clause 53. The method of any one of clauses 1-52, wherein a second locking structure is attached to a 5' end of the polynucleotide on a second side of the nanopore, and the second locking structure inhibits translocation of the 5' end of the polynucleotide through the opening to the first side of the nanopore.

[0234] Clause 54. The method of clause 53, wherein the second locking structure is removable.

[0235] Clause 55. The method of any one of clauses 1-54, further comprising, after operation (d), dissociating the extended duplex from the polynucleotide and forming a new duplex with the polynucleotide on the first side of the nanopore, the new duplex comprising a new 3' end.

[0236] Clause 56. The method of any one of clauses 1-55, wherein operation (a) comprises contacting a nanopore with a polynucleotide hybridized to a substantially complementary polynucleotide, and applying a sixth force that dehybridizes the substantially complementary polynucleotide from the polynucleotide.

[0237] Clause 57. The method of any one of clauses 1 to 56, wherein the nanopore comprises a solid-state nanopore.

[0238] Clause 58. The method of any one of clauses 1 to 57, wherein the nanopore comprises a biological nanopore.

[0239] Clause 59. The method of clause 58, wherein the biological nanopore comprises MspA.

[0240] Clause 60. The method of any one of clauses 1-59, wherein the first polynucleotide comprises RNA.

[0241] Clause 61. The method of any one of clauses 1 to 59, wherein the first polynucleotide comprises DNA.

[0242] Clause 62. The method of any one of clauses 1 to 61, wherein the extended duplex comprises a primer hybridized to the polynucleotide.

[0243] Clause 63. A sequencing system is provided, comprising a nanopore comprising a first side, a second side, and an opening extending through the first and second sides. The system comprises a polynucleotide disposed through the opening of the nanopore such that a 3' end of the polynucleotide is on the first side of the nanopore and a 5' end of the polynucleotide is on the second side of the nanopore. The system comprises a duplex having a polynucleotide disposed on the first side of the nanopore, the duplex comprising a 3' end at which a first nucleotide is disposed. The system is configured to apply a first force that positions the 3' end of the duplex within the opening, measure values ​​of an electrical property of the 3' end of the duplex and the single-stranded portion of the polynucleotide while applying the first force, and use the measurements to identify the first nucleotide, the nanopore inhibiting translocation of the 3' end of the duplex to the second side of the nanopore while the first force is applied.

[0244] Clause 64. The system of clause 63, wherein the value measured by the circuit comprises current, ionic current, electrical resistance, or voltage drop across the nanopore.

[0245] Clause 65. A system as described in clause 63 or 64, wherein the value measured by the circuit comprises noise in the current, ionic current, electrical resistance, or voltage drop across the nanopore.

[0246] Clause 66. The system of clause 65, wherein the value measured by the circuit includes a standard deviation of the noise.

[0247] Clause 67. A system according to any one of clauses 63 to 66, wherein the value measured by the circuit is based on at least M nucleotides of a single-stranded portion of the polynucleotide and D pairs of hybridized nucleotides of the double strand, where M is 2 or more and D is 1 or more.

[0248] Clause 68. The system of clause 67, wherein M is 3 or greater.

[0249] The system according to clause 67 or 68, wherein clause 69.D is 2 or greater.

[0250] Clause 70. The system of any one of clauses 67 to 69, wherein at least one of the M nucleotides of the single-stranded portion comprises a modified base, and the circuit is configured to identify the modified base using a value measured by the circuit.

[0251] Clause 71. The system of clause 70, wherein the modified base comprises a methylated base.

[0252] Clause 72. A system according to any one of clauses 63 to 71, wherein addition of another nucleotide to the 3' end of the duplex is inhibited while the first force is applied.

[0253] Clause 73. The system of clause 72, wherein the nanopore inhibits addition of another nucleotide.

[0254] Clause 74. A system described in any one of clauses 63 to 73, wherein the nanopore is oriented such that a first side of the nanopore comprises a majority of the opening.

[0255] Clause 75. A system described in any one of clauses 63 to 73, wherein the nanopore is oriented such that the second side of the nanopore comprises the majority of the opening.

[0256] Clause 76. The system of any one of clauses 63 to 75, wherein the circuitry is further configured to apply a modified first force that repositions the 3' end of the duplex within the opening, measure values ​​of electrical properties of the 3' end of the duplex and the single-stranded portion of the polynucleotide while applying the modified first force, and identify the first nucleotide using the measured value, and wherein the nanopore inhibits translocation of the 3' end of the duplex to a second side of the nanopore while the modified first force is applied.

[0257] Clause 77. The system of any one of clauses 63 to 76, further comprising a polymerase that contacts the 3' end of the duplex and adds a first nucleotide.

[0258] Clause 78. The system of clause 77, wherein the polymerase is reversibly inhibited from adding a second nucleotide to the 3' end of the duplex.

[0259] Clause 79. The system of clause 78, further comprising a blocking moiety that reversibly inhibits the polymerase from adding a second nucleotide to the 3' end of the duplex.

[0260] Clause 80. The system of clause 79, wherein the first nucleotide is bound to a blocking moiety.

[0261] Clause 81. The system according to clause 80, wherein the blocking moiety comprises a 3'-blocking group.

[0262] Clause 82. The system according to clause 79, wherein the blocking moiety reversibly associates with the duplex.

[0263] Clause 83. The system of clause 82, wherein the circuitry is further configured to detect association of the blocking moiety with the duplex.

[0264] Clause 84. The system of clause 82 or clause 83, wherein the circuitry is further configured to detect the absence of the blocking moiety from the duplex.

[0265] Clause 85. A system according to any one of clauses 79 to 84, wherein the blocking moiety is removable to allow the polymerase to add a second nucleotide to the 3' end of the duplex.

[0266] Clause 86. The system of any one of clauses 78 to 85, wherein the first force removes the polymerase from contact with the 3' end of the duplex.

[0267] Clause 87. The system of any one of clauses 78 to 86, wherein the circuit is further configured to apply a second force to remove the polymerase from contact with the 3' end of the double strand, the second force being greater than the first force.

[0268] Clause 88. The system of any one of clauses 78 to 87, wherein the circuitry is further configured to apply a third force that positions the polymerase in contact with the 3' end of the duplex within or adjacent to the opening on the first side of the nanopore, measure a value of an electrical property of the polymerase while applying the third force, and use the measured value to identify contact of the polymerase with the 3' end of the duplex, and wherein the nanopore inhibits movement of the polymerase into or further into the opening.

[0269] Clause 89. The system of clause 88, wherein the third force is less than the first force.

[0270] Clause 90. A system as described in clause 88 or 89, wherein the circuit is configured to apply a third force before applying the first force.

[0271] Clause 91. A system described in any one of clauses 88 to 90, wherein the first nucleotide is associated with a blocking moiety.

[0272] Clause 92. The system of clause 91, wherein the circuit is configured to use the measured value to confirm the presence of a blocking moiety associated with the first nucleotide.

[0273] Clause 93. A system described in any one of clauses 77 to 92, wherein the polymerase comprises a DNA polymerase.

[0274] Clause 94. A system described in any one of clauses 77 to 92, wherein the polymerase comprises an RNA polymerase.

[0275] Clause 95. A system described in any one of clauses 77 to 92, wherein the polymerase comprises a reverse transcriptase.

[0276] Clause 96. The system of any one of clauses 63 to 95, wherein the first nucleotide is associated with a blocking moiety, and the circuitry is further configured to use the measured value to confirm the presence of the blocking moiety associated with the nucleotide.

[0277] Clause 97. The system of clause 96, wherein the blocking moiety is removed from the first nucleotide after application of the first force.

[0278] Clause 98. The system of clause 97, wherein the circuit is configured to position the 3' end of the duplex within the opening after the blocking portion is removed, re-apply a first force that positions the single-stranded portion of the polynucleotide within the opening, measure values ​​of the electrical properties of the 3' end of the duplex and the single-stranded portion of the polynucleotide while re-applying the first force, and again identify the first nucleotide using the measured value, and wherein the nanopore inhibits translocation of the 3' end of the duplex to a second side of the nanopore.

[0279] Clause 99. The system of clause 98, wherein the blocking moiety is removed from the first nucleotide before the first force is applied.

[0280] Clause 100. A system described in clauses 63-99, wherein the duplex contains one or more nucleotide analogues.

[0281] Clause 101. The system according to clause 100, wherein one or more nucleotide analogues enhance duplex stability compared to natural nucleotides.

[0282] Clause 102. The system according to clause 100 or 101, wherein the one or more nucleotide analogues comprise one or more locked nucleic acids (LNAs).

[0283] Clause 103. The system of any one of clauses 100 to 102, wherein the one or more nucleotide analogs include one or more 2'-methoxy (2'-OMe) nucleotides.

[0284] Clause 104. The system of any one of clauses 100 to 103, wherein the one or more nucleotide analogs include one or more 2'-fluorinated (2'-F) nucleotides.

[0285] Clause 105. A system according to any one of clauses 100 to 104, wherein one or more nucleotide analogues have altered values ​​of electrical properties compared to natural nucleotides.

[0286] Clause 106. A system described in any one of clauses 100 to 105, wherein the first nucleotide comprises one of one or more nucleotide analogues.

[0287] Clause 107. The system according to clause 106, wherein one or more of the nucleotide analogues comprises a 2' modification.

[0288] Clause 108. The system according to clause 106 or clause 107, wherein one or more nucleotide analogues comprise a base modification.

[0289] Clause 109. A system described in any one of clauses 63 to 108, wherein the first force is insufficient to cause dissociation of the double strand.

[0290] Clause 110. A system described in any one of clauses 63 to 109, wherein the first force includes a first voltage.

[0291] Clause 120. A system described in any one of clauses 63 to 110, wherein in the absence of a first force, the polynucleotide is positioned through the opening and the double strand is positioned on a first side of the nanopore.

[0292] Clause 121. A system described in any one of clauses 63 to 120, wherein the circuit is configured to apply a fourth force opposing the first force.

[0293] Clause 122. A system described in any one of clauses 63 to 121, wherein a first locking structure is bound to a 3' end of the polynucleotide on a first side of the nanopore, and the first locking structure inhibits translocation of the 3' end of the polynucleotide through the opening to a second side of the nanopore.

[0294] Clause 123. The system of clause 122, wherein the first locking structure is removable.

[0295] Clause 124. A system described in any one of clauses 63 to 123, wherein a second locking structure is attached to the 5' end of the polynucleotide on a second side of the nanopore, and the second locking structure inhibits translocation of the 5' end of the polynucleotide through the opening to the first side of the nanopore.

[0296] Clause 125. The system of clause 124, wherein the second locking structure is removable.

[0297] Clause 126. A system described in any one of clauses 63 to 125, wherein the circuit is configured to dissociate the double strand from the polynucleotide after applying the first force.

[0298] Clause 127. The system of any one of clauses 63 to 126, wherein the circuit is configured to apply a sixth force that dehybridizes the substantially complementary polynucleotide from the polynucleotide so as to position the polynucleotide through the opening of the nanopore.

[0299] Clause 128. A system described in any one of clauses 63 to 127, wherein the nanopore comprises a solid-state nanopore.

[0300] Clause 129. A system described in any one of clauses 63 to 128, wherein the nanopore comprises a biological nanopore.

[0301] Clause 130. The system described in clause 129, wherein the biological nanopore comprises MspA.

[0302] Clause 131. The system of any one of clauses 63 to 130, wherein the first polynucleotide comprises RNA.

[0303] Clause 132. A system described in any one of clauses 63 to 130, wherein the first polynucleotide comprises DNA.

[0304] Clause 133. A system according to any one of clauses 63 to 132, wherein the duplex comprises a primer hybridized to the polynucleotide.

[0305] Clause 134. A method of sequencing an unknown polynucleotide is provided. The method includes providing a plurality of measurements of electrical properties of a single-stranded portion of the unknown polynucleotide and a 3' end of a duplex having the unknown polynucleotide within the opening of the nanopore as input to a nucleotide identification module. The method also includes using the nucleotide identification module to compare the plurality of measurements to values ​​in a data structure, the data structure correlating different measurements with different combinations of nucleotides within the single-stranded portion of the known polynucleotide and the 3' end of a known duplex having the known polynucleotide within the opening of the nanopore. The method also includes using the nucleotide identification module to determine a sequence of nucleotides in the sequence of the unknown polynucleotide using the comparison. The method also includes receiving a representation of the determined sequence of nucleotides as output from the nucleotide identification module.

[0306] Clause 135. The method of clause 134, wherein the nucleotide discrimination module comprises a trained machine learning algorithm.

[0307] Clause 136. The method of clause 135, wherein the nucleotide discrimination module comprises a trained deep learning algorithm.

[0308] Clause 137. The method of clause 135 or 136, wherein the data structure comprises neurons of a trained machine learning algorithm.

[0309] Clause 138. The method of any one of clauses 134 to 137, wherein the data structure includes a read map.

[0310] Clause 139. The method of clause 138, wherein the read map comprises a look-up table that stores different measurements and representations of different combinations of nucleotides within the 3' end of the known double strand and the single strand portion of the known nucleotides.

[0311] Clause 140. The method of any one of clauses 134-139, further comprising using a measurement module by a computer to generate a plurality of measurements using the opening of the nanopore.

[0312] Clause 141. The method of any one of clauses 134 to 140, further comprising using a nucleotide addition module, a measurement module, and a nucleotide identification module by a computer to generate a data structure using the opening of the nanopore.

[0313] Clause 142. A system for sequencing an unknown polynucleotide is provided. The system includes a processor and at least one computer readable medium storing a plurality of measurements of electrical properties of a single-stranded portion of an unknown polynucleotide and a 3' end of a duplex comprising the unknown polynucleotide within an opening of a nanopore. The at least one computer readable medium further stores a data structure correlating different measurements with different combinations of nucleotides within the single-stranded portion of a known polynucleotide and the 3' end of a known duplex within the opening. The at least one computer readable medium further stores instructions for causing the processor to perform operations including: comparing the plurality of measurements to values ​​in the data structure; determining a sequence of nucleotides in the sequence of the unknown polynucleotide using the comparison; and outputting a representation of the determined sequence of nucleotides.

[0314] Clause 143. The system of clause 142, wherein the nucleotide discrimination module comprises a trained machine learning algorithm.

[0315] Clause 144. The system of clause 143, wherein the nucleotide discrimination module comprises a trained deep learning algorithm.

[0316] Clause 145. The system of clause 143 or 144, wherein the data structure includes neurons of a trained machine learning algorithm.

[0317] Clause 146. A system described in any one of clauses 142 to 145, wherein the data structure includes a read map.

[0318] Clause 147. The system of clause 146, wherein the read map comprises a lookup table that stores different measurements and representations of different combinations of nucleotides within the 3' end of the known double strand and the single strand portion of the known nucleotides.

[0319] Clause 148. The system of any one of clauses 142-147, wherein the instructions further cause the processor to generate a plurality of measurements using the opening of the nanopore.

[0320] Clause 149. A system described in any one of clauses 142 to 148, wherein the instructions further cause the processor to generate a data structure using the opening of the nanopore.

[0321] Clause 150. A method is provided for locking a polynucleotide to a nanopore comprising a first side, a second side, and an opening extending through the first and second sides. The method includes (a) attaching a first locking group to a 3' end of the polynucleotide. The method also includes (b) positioning the polynucleotide through the opening of the nanopore such that the 3' end of the polynucleotide and the first locking group are on the first side of the nanopore and the 5' end of the polynucleotide is on the second side of the nanopore. The method also includes (c) attaching a second locking group to the 5' end of the polynucleotide on the second side of the nanopore.

[0322] Clause 151. The method of clause 150, wherein the first locking group comprises a locked nucleic acid (LNA) or a peptide nucleic acid (PNA).

[0323] Clause 152. The method of clause 150 or 151, wherein the second locking group comprises a locked nucleic acid (LNA) or a peptide nucleic acid (PNA).

[0324] Clause 153. The method of any one of clauses 150-152, wherein the polynucleotide is hybridized to a complementary polynucleotide prior to operation (a), the method further comprising dehybridizing the complementary polynucleotide between operations (b) and (c). EXAMPLES

[0325] The following examples are intended to be purely illustrative and not limiting.

[0326] Figure 11 illustrates a plot of values ​​measured as a function of time during uptake using an exemplary template polynucleotide. More specifically, the system 100 described with reference to Figures 1A-1H was used, with MspA used as the nanopore 110. Polynucleotides 140 and 150 each had the sequence shown in Figure 11. More specifically, polynucleotide 140 had the sequence 5'-TGGTCAGGTG TTTGCGTA (SEQ ID NO:5) and polynucleotide 150 had the sequence

[0327] [Table 1] or equivalently

[0328] [Table 2] (SEQ ID NO:6), where "X" indicates an abasic nucleotide (abasic site), and the bold indicates the nucleotide in polynucleotide 150 from which a signal was obtained as shown in Figure 11. Note that although polynucleotide 150 contained an abasic nucleotide, the abasic nucleotide was located in a portion of polynucleotide 150 that was not sequenced, and the abasic nucleotide was not used to generate a signal in any part of the cycle.

[0329] Circuit 160 was used to apply a second force F2 (here, a bias voltage of -50 mV was applied for 40 ms in plot 1101 and 60 ms in the continuation of plot 1101), during which polymerase 105 added unmodified nucleotides dTTP, dATP, dCTP, and dGTP to polynucleotide 140 based on the sequence of polynucleotide 150 in the direction indicated by arrow 141 in a manner as described with reference to Figures 1B and 1D. Circuit 160 was also used to apply a first force F1 (here, a bias voltage of 85 mV was applied for 100 ms) that positioned double-stranded 3' end 153 and single-stranded second portion 156 of polynucleotide 140 within the opening of the nanopore in a manner as described with reference to Figures 1A, 1C, and 1E. Circuit 160 alternated between applying the first force and the second force. While applying the first force, circuit 160 measured the average current through the nanopore, as shown in plot 1101 of Figure 11. The standard deviation of the average current is shown in plot 1102 of Figure 11. In plots 1101 and 1102, different value levels are shown corresponding to nucleotides that were inferred to have been added to polynucleotide 140 based on the sequence of nucleotides shown in bold in polynucleotide 150. Each point in Figure 11 corresponds to the average current (or its standard deviation) measured during one of the 100 millisecond cycles during which the first force was applied.

[0330] Referring first to plot 1101, an average current of about 10 pA at 0 seconds corresponds to polynucleotide 150 being disposed within nanopore 110 without being hybridized to polynucleotide 140. The average current then increases to about 15.5 pA from about 5-15 seconds, corresponding to annealing of polynucleotide 140 (the primer) to polynucleotide 150. The average current then increases to about 17 pA from about 15-70 seconds, presumably corresponding to the addition of a nucleotide T to polynucleotide 140 based on the next nucleotide A in polynucleotide 150. The average current then decreases to about 14 pA from about 70-80 seconds, presumably corresponding to the addition of a nucleotide A to polynucleotide 140 based on the next nucleotide T in polynucleotide 150. The average current then remains at about 14 pA from about 80-150 seconds, presumably corresponding to the addition of another nucleotide A to polynucleotide 140 based on the next nucleotide T in polynucleotide 150. The average current then increases to about 17 pA from about 150-160 seconds, which likely corresponds to the addition of another nucleotide A to polynucleotide 140 based on the next nucleotide T in polynucleotide 150. The average current then decreases to about 16.5 pA from about 160-162 seconds, which likely corresponds to the addition of a nucleotide G to polynucleotide 140 based on nucleotide C in polynucleotide 150. The average current then decreases to about 15.5 pA from about 162-170 seconds, which likely corresponds to the addition of a nucleotide C to polynucleotide 140 based on nucleotide G in polynucleotide 150. The average current then decreases to about 11.5 pA from about 170-185 seconds, which likely corresponds to the addition of a nucleotide A to polynucleotide 140 based on nucleotide T in polynucleotide 150. The average current then increases from about 185-200 s to about 18 pA, presumably corresponding to the addition of nucleotide C to polynucleotide 140 based on nucleotide G in polynucleotide 150.The average current then increases from about 200-240 seconds to about 14.5 pA, which likely corresponds to the addition of nucleotide A to polynucleotide 140 based on nucleotide T in polynucleotide 150. The average current then increases from about 240-250 seconds to about 25 pA, which likely corresponds to the addition of nucleotide G to polynucleotide 140 based on nucleotide C in polynucleotide 150. Similarly, as additional nucleotides are added, the average current continues to increase or decrease during the measurement based on the particular combination of the nucleotide at the 3' end of the duplex and the unpaired nucleotide in the second region 156 of the single strand of polynucleotide 150.

[0331] From plot 1101, it can be seen that the change in the measurement as a function of time can be used to ascertain when polynucleotide 140 (the primer) is first hybridized to polynucleotide 150. For example, the average current was observed to increase from about 10 pA to about 15.5 pA in response to hybridization of polynucleotide 140 to polynucleotide 150 (an increase of more than about 50%). From plot 1101, it can also be seen that a nucleotide can be added to polynucleotide 140 at the 3' end 153 of the duplex, and that the change in the measurement as a function of time can be used to ascertain when a nucleotide is added to polynucleotide 140. For example, the average current was observed to increase from about 15.5 pA to about 17 pA with the addition of the first T, and then decrease from about 17 pA to about 14 pA with the addition of the first A. The average current was observed to range from about 11.5 pA (4th A) to about 25 pA (2nd G), a change of over about 210%.

[0332] It can also be seen from plot 1101 that the change in measurements can be based not only on the particular nucleotide added to the 3' end 153 of the duplex, but also on other nucleotides. For example, the average current for each of the five different A's mentioned above was about 14 pA, about 14 pA, about 17 pA, about 11.5 pA, and about 14.5 pA (a variation of more than about 45%). If the average current measured was based only on the addition of A's, the same value would be expected for each such addition. For some nucleotide combinations, similar measurements were observed for the same added nucleotide (e.g., the first and second A), while for other nucleotide combinations, different measurements were observed for the same added nucleotide (e.g., the third and fourth A, or the first and second G).

[0333] Referring now to plot 1102, the standard deviation is the standard deviation of the average current as described with reference to plot 1101. The standard deviation of about 0.9 pA at 0 seconds corresponds to polynucleotide 150 being disposed within nanopore 110 without being hybridized to polynucleotide 140. The standard deviation then increases to about 1.3 pA from about 5-15 seconds, corresponding to annealing of polynucleotide 140 (the primer) to polynucleotide 150. The standard deviation then remains at about 1.3 pA from about 15-70 seconds, presumably corresponding to the addition of nucleotide T to polynucleotide 140 based on the next nucleotide A in polynucleotide 150. The standard deviation then increases to about 1.4 pA from about 70-80 seconds, presumably corresponding to the addition of nucleotide A to polynucleotide 140 based on the next nucleotide T in polynucleotide 150. The standard deviation then increases to about 1.8 pA from about 80-150 seconds, which likely corresponds to the addition of another nucleotide, A, to polynucleotide 140 based on the next nucleotide, T, in polynucleotide 150. The standard deviation then increases to about 1.9 pA from about 150-160 seconds, which likely corresponds to the addition of another nucleotide, A, to polynucleotide 140 based on the next nucleotide, T, in polynucleotide 150. The standard deviation then decreases to about 1.4 pA from about 160-162 seconds, which likely corresponds to the addition of a nucleotide, G, to polynucleotide 140 based on a nucleotide, C, in polynucleotide 150. The standard deviation then decreases to about 1.3 pA from about 162-170 seconds, which likely corresponds to the addition of a nucleotide, C, to polynucleotide 140 based on a nucleotide, G, in polynucleotide 150. The standard deviation then remains at about 1.3 pA from about 170-185 seconds, which presumably corresponds to the addition of nucleotide A to polynucleotide 140 based on nucleotide T in polynucleotide 150. The standard deviation then increases to about 1.4 pA from about 185-200 seconds, which presumably corresponds to the addition of nucleotide C to polynucleotide 140 based on nucleotide G in polynucleotide 150.The standard deviation then decreases from about 200-240 seconds to about 1.1 pA, which presumably corresponds to the addition of nucleotide A to polynucleotide 140 based on nucleotide T in polynucleotide 150. The standard deviation then increases from about 240-250 seconds to about 1.9 pA, which presumably corresponds to the addition of nucleotide G to polynucleotide 140 based on nucleotide C in polynucleotide 150. Similarly, as additional nucleotides are added, the standard deviation continued to increase or decrease during the measurement based on the particular combination of the nucleotide at the 3' end of the duplex and the unpaired nucleotide in single-stranded second region 156 of polynucleotide 150.

[0334] From plot 1102, it can be seen that the change in standard deviation as a function of time can be used to ascertain when polynucleotide 140 (the primer) is first hybridized to polynucleotide 150. For example, the standard deviation was observed to increase from about 0.9 pA to about 1.3 pA in response to hybridization of polynucleotide 140 to polynucleotide 150 (an increase of more than about 40%). From plot 1102, it can also be seen that a nucleotide can be added to polynucleotide 140 at the 3' end 153 of the duplex, and that the change in standard deviation as a function of time can be used to ascertain when a nucleotide is added to polynucleotide 140. For example, the standard deviation was observed to increase from about 1.4 pA to about 1.9 pA between the addition of the first A and the second A. The standard deviation was observed to range from about 1.1 pA (5th A) to about 1.9 pA (2nd G), a change of more than about 70%.

[0335] From the plot 1102, it can be seen that the change in standard deviation can be based not only on the particular nucleotide added to the 3' end 153 of the duplex, but also on other nucleotides. For example, the standard deviations for each of the five different As mentioned above were about 1.4 pA, about 1.8 pA, about 1.9 pA, about 1.3 pA, and about 1.1 pA (more than about 70% variation). If the standard deviation was based only on the addition of As, the same value would be expected for each such addition. For some nucleotide combinations, similar standard deviations were observed for the same added nucleotide (e.g., the second and third A, or the first and fourth A), while for other nucleotide combinations, different standard deviations were observed for the same added nucleotide (e.g., the first and fifth A).

[0336] From plots 1101 and 1102, it can also be seen that the use of multiple different types of measurements can be used to distinguish between different nucleotides. For example, some nucleotides may have similar values ​​to each other for one particular measurement type and different values ​​to each other for another measurement type. Such similar measurements may be referred to as "degenerate" herein because additional information may be required to distinguish the nucleotides from each other. In a manner as described with reference to FIG. 8E, that additional information may include the standard deviation of those measurements. Illustratively, as shown in plot 1101, the first and second A may have similar average currents to each other and thus may be considered to be "degenerate" with respect to their metrics measured under those particular conditions. However, as shown in plot 1102, the standard deviation of the average current for the first A may be easily distinguished from that for the second A. Thus, it can be seen that even if the average current alone may be insufficient to distinguish certain nucleotides, the standard deviation (or other type of measurement) may be sufficient to distinguish those nucleotides. Conversely, as shown in plot 1102, T and the first A may have similar standard deviations to each other and therefore may be considered "degenerate" with respect to their metrics measured under those particular conditions. However, as shown in plot 1101, the average current for T may be easily distinguished from the average current for the first A. Thus, it can be seen that even if the standard deviation alone measured under those particular conditions may be insufficient to distinguish certain nucleotides, the average current (or other type of measurement) may be sufficient to distinguish those nucleotides. It can also be seen that any degeneracy can be easily resolved using multiple different types of information to characterize the combination of nucleotides at the 3' end 153 of the duplex and at the single-stranded second portion 156 of the polynucleotide 150.It should be noted that such different types of information do not necessarily require making multiple types of measurements, but may instead involve obtaining different types of information from the same set of measurements, illustratively measurements and their standard deviations. More generally, it is expected that different sets of measurement conditions (e.g., different voltages, different fluid compositions such as different salt concentrations, etc.) can be used to characterize the same sequence of nucleotides in different ways, at least some of which are not "degenerate" and thus allow the nucleotides to be distinguished from one another.

[0337] It can also be seen from plots 1101 and 1102 that polynucleotide 140 can be extended far enough that interactions between the polymerase and nanopore 110 can inhibit any further extension of polynucleotide 140. For example, window 1132 illustrated in Figure 11 illustrates the approximate position of MspA nanopore 110 relative to polynucleotide 150 when circuit 160 applies second force F2 such that lock 152 is positioned relative to nanopore 110 in the manner described with reference to Figures 1B and 1D. Although the example illustrated in Figure 11 shows an abasic nucleotide within window 1132, any suitable moiety can be used within window 1132, such as natural nucleotides, non-natural nucleotides, abasic nucleotides, other suitable chemical moieties, and combinations thereof. Window 1131 illustrated in FIG. 11 illustrates the approximate position of polymerase 105 when polynucleotide 140 is sufficiently extended such that the polymerase is parked against the nanopore and can no longer access the 3′ end 153 of the duplex.

[0338] As described above with reference to Figures 1G, 4A-4C, 8A-8E, and 9, polynucleotide 150 may be sequenced and resequenced under the same set of measurement conditions or under different sets of measurement conditions. Figure 12 illustrates a plot of values ​​measured during resequencing of an exemplary polynucleotide under a set of measurement conditions. More specifically, the system 100 described with reference to Figures 1A-1H, in which MspA was used as the nanopore 110, was used in a manner similar to that described with reference to Figure 11. An excess of primer (polynucleotide 140) and polymerase were added to the fluid on the first side of the nanopore, and the primer was allowed to hybridize to polynucleotide 150. Polynucleotides 140 and 150 each had the sequence shown in Figure 12 (3'-ATTTCGT-5' for polynucleotide 150, and 5'-TAAA-3' for polynucleotide 140). Circuit 160 was used to apply a second force F2 (here, a bias voltage of -50 mV) during which polymerase 105 added modified nucleotides dTTP and dATP to polynucleotide 140 based on the sequence of polynucleotide 150 in the direction indicated by arrow 141 in a manner as described with reference to Figures 1B and 1D. dCTP was also present, but dGTP was excluded so as to inhibit extension of polynucleotide 140 past the sequence TAAA. Circuit 160 was also used to apply a first force F1 (here, a bias voltage of 80 mV) that positioned the double-stranded 3' end 153 and the single-stranded second portion 156 of polynucleotide 150 within the opening of the nanopore in a manner as described with reference to Figures 1A, 1C, and 1E. Circuit 160 alternated between the first and second forces at 100 millisecond intervals. While applying the first force, the circuit 160 measured the average current through the nanopore, as shown in plot 1201 of Figure 12. After the nucleotide TAAA was added to the polynucleotide 140, the extended polynucleotide 140 spontaneously dissociated from the polynucleotide 150. The circuit was then used to again alternate between applying a second force and the first force at 100 millisecond intervals.During application of the second force, a new primer (polynucleotide 140) was hybridized to polynucleotide 150 in a manner as described with reference to Figure 4C, and such primer was then extended to resequence polynucleotide 150. Such cycles of stripping and resequencing were repeated several times.

[0339] In plot 1200, the different value levels are coded as follows and are shown in the legend: template alone=1206, primer (hybridized to template)=1201, primer+T (hybridized to template)=1202, primer+TA (hybridized to template)=1203, primer+TAA (hybridized to template)=1204, primer+TAAA (hybridized to template)=1205. Each point in FIG. 12 corresponds to the average current measured during one of the 100 millisecond cycles during which the first force was applied. It can be seen in plot 1200 that at about 15 seconds, 120 seconds, 170 seconds, 220 seconds, and 260 seconds, the average current corresponded to the average current of the template (polynucleotide 150) without the primer (1205). After each such time point, the average current corresponded to hybridization of the primer (polynucleotide 140) to the template (1201). Then, after each such time point, the series of average currents corresponded to the successive additions to the primer of a T (1202), then an A (1203), then another A (1204), then another A (1205). It can be seen that the average currents for the addition of the first, second, and third A are different from each other (compare 1203, 1204, 1205 with each other), but in each resequencing cycle, these As have the same average currents as they did in each other cycle. Thus, from FIG. 12, it can be seen that the polynucleotide can be repeatedly resequenced by stripping off the extended polynucleotide 140 and adding and extending new polynucleotide 140, with the values ​​during each resequencing cycle correlating reliably to the sequence of nucleotides added to polynucleotide 140, and thus to the sequence of polynucleotide 150.

[0340] 13A-13C illustrate plots of values ​​measured during resequencing of an exemplary polynucleotide under different sets of measurement conditions. More specifically, the system 100 described with reference to FIGS. 1A-1H, in which MspA was used as the nanopore 110, was used in a manner similar to that described with reference to FIGS. 11 and 12. Polynucleotides 140 and 150 each had the sequence shown in FIG. 12. Circuit 160 was used to apply a second force F2 (here, a bias voltage of −50 mV) while polymerase 105 added modified nucleotides dTTP and dATP to polynucleotide 140 based on the sequence of polynucleotide 150 in the direction indicated by arrow 141 in a manner as described with reference to FIGS. 1B and 1D. dCTP was also present, but dGTP was excluded so as to inhibit extension of polynucleotide 140 past the sequence TAAA. Circuit 160 was also used to sequentially apply several different (adjusted) first forces F1, F1', F1'', F1''', generated by specific bias voltages (here, bias voltages of 75 mV, 80 mV, 85 mV, and 90 mV, each applied for 100 ms). Each of the first forces positioned the 3' end 153 of the duplex and the single-stranded second portion 156 of the polynucleotide 150 within the opening of the nanopore in a manner as described with reference to Figures 1A, 1C, and 1E. Without wishing to be bound by any theory, it is believed that the different forces caused the nucleotides in the duplex and the single-stranded second portion of the polynucleotide 150 to interact differently with the nanopore, resulting in different measurements, in a manner as described with reference to Figure 1G. After applying the four different first forces, the circuit 160 applied a second force (-50 mV for 100 ms), resulting in the addition of another nucleotide, followed by sequential application of the four different first forces.

[0341] While applying each of the first forces in a manner similar to that described elsewhere herein, the circuit 160 measured the average current through the nanopore, shown in each one of plots 1301, 1302, 1303, 1304 in FIG. 13A. After adding the nucleotide TAAA to the polynucleotide 140, the extended polynucleotide 140 dissociated from the polynucleotide 150A, a new primer (polynucleotide 140) was hybridized to the polynucleotide 150 in a manner as described with reference to FIG. 4C, and extended to resequence the polynucleotide 150. Note that each of the first forces was applied for 100 milliseconds in this example, resulting in a read cycle of, for example, 400 milliseconds (for the read at each of the four voltages). The plots of the different read voltage values ​​are illustrated vertically so that they all refer to the same read cycle start time. Optionally, after it is determined that the nucleotides have been fully characterized in their reads, a force F2 can be applied to incorporate the next nucleotide. Thus, for each cycle, a single F2 can be applied, and it will be understood that many different read voltages can be applied in any order for as long a duration as desired. Then, after the entire template has been passed, F4 can be applied to strip the template, and the sequencing process is repeated as many times as desired using fresh primers.

[0342] In plot 1301 corresponding to a first force F1 of 90 mV, different value levels are coded using different fills to correspond to hybridization of primer (polynucleotide 140) or nucleotide sequence TAAA, and each point corresponds to the average current measured as a function of time during one of the 100 millisecond cycles during which a first force F1'''' of 90 mV was applied. Plot 1302 was similarly obtained using a first force F1″ of 85 mV. Plot 1303 was similarly obtained using a first force F1′ of 80 mV. Plot 1304 was similarly obtained using a first force F1 of 75 mV. It can be seen that in each of plots 1301, 1302, 1303, and 1304, at about 0 seconds, the average current corresponds to the average current of the template (polynucleotide 150) containing no primer. This was followed by the average current corresponding to hybridization of the primer (polynucleotide 140) to the template. This was then followed by a series of average currents corresponding to the successive addition of T, A, A, and A to the primer. It can be seen from plots 1301, 1302, 1303, and 1304 that under a given first force F1, the average currents corresponding to hybridization of the primer or different nucleotides, respectively, are similar to each other in some respects and different from each other in other respects.

[0343] Plots 1311, 1312, 1313, and 1314 of FIG. 13B illustrate the average signal level of each of the average currents of FIG. 13A as a function of step (also referred to as sequence index) to facilitate comparison of the raw data in plots 1301, 1302, 1303, and 1304, respectively. The signal levels shown are averages over the entire time that a particular base was at the 3' end 153 of the duplex 154. It will be understood that the amount of time for a particular nucleotide incorporation can vary, for example, a given incorporation can occur in less than one cycle between a first force and a second force, or over the course of many such cycles. In plots 1311, 1312, 1313, and 1314, the signal levels corresponding to hybridization of the primer are higher than the signal levels corresponding to the polynucleotide 150 without such hybridization. In plots 1311, 1312, and 1313, the signal level corresponding to primer extension by T is higher than the signal level corresponding to primer hybridization, while in plot 1314, the signal level corresponding to primer extension by T is similar to (but still distinguishable from) the signal level corresponding to primer hybridization. In plots 1311, 1312, 1313, and 1314, the signal level corresponding to primer extension by T is different from the signal level corresponding to extension by each of the three As. In each of plots 1311, 1313, and 1314, the signal levels corresponding to primer extension by each of the three As are different from each other, while in plot 1312, the signal levels corresponding to primer extension by the first and second As are approximately the same as each other (degenerate).

[0344] Plot 1321 in FIG. 13C illustrates the average signal (measured current) levels of FIG. 13B as a function of bias voltage to further facilitate comparison of the raw data in plots 1301, 1302, 1303, and 1304, respectively. In plot 1321, it can be seen that each of the signal levels increases as a function of the bias voltage applied during the measurement, and at least some of the signal levels have different slopes from each other. To facilitate comparison of the different slopes in plot 1321, plot 1322 in FIG. 13C illustrates the normalized average signal levels of FIG. 13B as a function of bias voltage, where the signal levels were normalized by dividing the average signal level corresponding to the added nucleotide by the average signal level corresponding to the hybridization of the primer. It can be seen that the normalized signal level for T is significantly higher than the signal level for the first A, and varies in a different direction. It can also be seen that the normalized signal level for the first A is similar (but still distinguishable) to the normalized signal level for the second A, but varies in a different direction. It can be seen that intersection point 1331 corresponds to an approximate bias voltage (85 mV) where the normalized signal levels of the first and second A are approximately the same (degenerate) as one another. It can also be seen that the normalized signal level for the third A is significantly higher than the normalized signal levels for the first and second A, varying in the same direction as the first A, and varying in a different direction than the second A and T. It can be seen that intersection point 1332 corresponds to an approximate bias voltage (85 mV) where the normalized signal levels for the T and the third A are approximately the same (degenerate) as one another.

[0345] Thus, from FIGS. 13A-13C, it can be seen that at some first forces F1, the measurements for a particular nucleotide combination may be degenerate, while at other first forces F1', F1'', F1''', which may be applied successively during a nucleotide addition cycle using F2, the measurements for some or all nucleotide combinations may be readily distinguished. Thus, it will be understood that measurements may be performed using any suitable number of forces to obtain measurements in which any degeneracy in the measurements may be resolved. In other examples, measurements may be performed using different first forces selected such that each possible combination of nucleotides has a measurement under at least one of such first forces that is readily distinguishable from the measurements of other combinations. It can also be seen from FIGS. 13A-13C that by applying different forces using circuit 160, the 3' end of the duplex and the second portion of polynucleotide 150 may be moved to different positions relative to nanopore 110 where the nucleotides in the duplex and polynucleotide 150 may affect the measurements differently than they do at other positions (under different forces). 1G and 8C, changes in the measurement conditions may affect the measurements linearly or nonlinearly, and may indeed shift the measurements in different directions for different combinations of nucleotides at the 3' end of the duplex and within the second portion of polynucleotide 150. It should be noted that while each of plots 1301, 1302, 1303, 1304 was obtained while sequencing polynucleotide 150 using sequentially applied forces F1, F1', F1", F1"", in other configurations, circuit 160 may be configured to sequence polynucleotide 150 using a selected one of the first forces and then sequence the polynucleotide using one or more other of the first forces.

[0346] The measurements described with reference to Figures 11, 12, and 13A-13C were obtained using an MspA nanopore 110 oriented in a manner as illustrated in Figures 1A-1H such that the first side 111 of the nanopore comprises a majority of the opening 113, such that the 3' end 153 of the duplex 154 can fit relatively deeply within the opening 113. Figure 14 illustrates plots of values ​​measured during resequencing of an exemplary polynucleotide under different sets of measurement conditions, in which the sequencing system has an alternative configuration as described with reference to Figure 7. More specifically, the MspA nanopore 110 was oriented such that the second side 112 of the nanopore comprises a majority of the opening 113, such that the 3' end of the duplex 154 can fit relatively shallowly within the opening 113. It can be seen in FIG. 14 that 28 nucleotides in polynucleotide 150 were resequenced in a manner similar to that described with reference to FIGS. 11, 12, and 13A using first forces of 30 mV (trace 1401), 40 mV (trace 1402), and 50 mV (trace 1403), and that using different first forces, different current levels were observed corresponding to different combinations of nucleotides at the 3′ end 153 of the duplex and the second portion of polynucleotide 150. Thus, it can be seen from FIG. 14 that regardless of the orientation of the nanopore, the addition of nucleotides to polynucleotide 140 results in measurements that correlate to the combination of nucleotides obtained at the 3′ end 153 of the duplex and the second portion of polynucleotide 150. The numbers shown above trace 1403 correspond to the assignment of different signal levels to the different nucleotides added.

[0347] The exemplary systems described with reference to Figures 11, 12, 13A-13C, and 14 were also used to demonstrate that modified bases, such as methylated bases, can be discriminated. For example, Figures 15A-15B illustrate plots of values ​​measured during sequencing of a polynucleotide containing modified bases. Two versions of polynucleotide 150 were prepared, one with the sequence 3'-GCATTTTTTACATTTTTTACATTTTTT-5' (SEQ ID NO:1) and the other with the similar sequence

[0348] [Table 3] The cytosine, having the sequence (SEQ ID NO: 2) and indicated by bold and an asterisk, was methylated (5mC) in some measurements and unmethylated in others. Primer 5'-CGT was hybridized to each of these sequence polynucleotides 150 in the manner illustrated in FIG. 15A and extended in the manner described with reference to FIG. 11, where the circuit 160 alternates between the application of a first force of 80 mV (read voltage) and a second force of -50 mV (capture voltage). Each read was performed twice (T1 and T2).

[0349] Figure 15A illustrates the average current over the entire time that a particular base was at the 3' end 153 of the duplex 154, similar to that described for Figure 13B. For the sequence (SEQ ID NO:2) shown above the plot in Figure 15A, the approximate K-mer that is believed to be located in the sensing region of the nanopore is illustrated below each measured signal level. In one set of measurements, the sequence was methylcytosine (5mC, C * ) while in another set of measurements, methylcytosines in the sequence were replaced with cytosines (i.e., C * (In the figure, cytosine was replaced with C). For K-mers 0, 1, 2, 3, 4, 5, 6, 9, 10, 11, 12, 13, 14, 19, 20, 21, 22, 23, and 24, it can be seen that the read currents for the two sequences were similar to each other. This was attributed to the fact that cytosine or methylcytosine either affect the read current similarly to each other or are far enough away from the read head of the nanopore to not significantly affect the read current. In comparison, for K-mers 7, 8, 15, 16, 17, and 18, the read currents for the two sequences were significantly different from each other. This was due to the fact that C *This was due to the fact that the cytosine or methylcytosine at each position indicated, for example at positions 5 and 4 illustrated in FIG. 1F for the first methylcytosine in the sequence, and at positions 6, 5, 4, and 3 illustrated in FIG. 1F for the second methylcytosine in the sequence, are sufficiently close to the read head of the nanopore to affect the read current differently from one another in readily distinguishable ways.

[0350] FIG. 15B illustrates the difference in time (level duration) that a given average current was observed while a particular base was at the 3′ end 153 of the duplex 154 for the same sequence (SEQ ID NO:2), again with C * For a sequence having 5mC at the position indicated by C * The measurements with cytosine at the position indicated by is compared. The approximate K-mers believed to be located in the sensing region of the nanopore are again illustrated below each measured signal level. For K-mers 0, 1, 2, 3, 4, 9, 10, 11, 12, and 13, it can be seen that the level durations for the two sequences are similar to each other. This was due to the fact that cytosine or methylcytosine either affect the read current similarly to each other or are far enough away from the read head of the nanopore to not significantly affect the read current. In comparison, for K-mers 5, 6, 7, 8, 14, 15, 16, 17, the level durations for the two sequences were significantly different from each other. This was due to the fact that cytosines or methylcytosines, for example at positions 7, 6, 5, and 4 illustrated in FIG. 1F for the first methylcytosine in the sequence, and at positions 7, 6, 5, and 4 illustrated in FIG. 1F for the second methylcytosine in the sequence, were sufficiently close to the nanopore readhead to affect level durations differently from one another in readily distinguishable ways.

[0351] From the plots described with reference to FIGS. 11, 12, 13A-13C, 14, and 15A-15B, it can be further seen that the circuit 160 can be used to obtain measurements for each nucleotide addition with any desired level of precision. The level of precision can be increased by resolving degeneracy using multiple different first forces F1, F1', F1''. For example, while the circuit 160 applies any suitable number of first forces, the nanopore 110 inhibits the polymerase from adding another nucleotide to the 3' end 153 of the duplex in a manner as described with reference to FIGS. 1A, 1C, and 1C. As described elsewhere herein, the circuit 160 can be configured to sequentially apply any suitable number of first forces F1, F1', F1'', etc., for any suitable amount of time to obtain measurements having a sufficient SNR for the particular situation in which the sequencing system and method is being performed, without substantial risk of adding another nucleotide during such measurements. Thus, it can be seen from the results herein that the stepwise addition and characterization of single nucleotides can be highly controlled and used to sequence polynucleotides at any suitable level of speed or accuracy.

[0352] Further comments While various illustrative examples have been described above, it will be apparent to one skilled in the art that various changes and modifications can be made herein without departing from the present invention. It is intended that the appended claims cover all such changes and modifications that fall within the true spirit and scope of the present invention.

[0353] It should be understood that any respective feature / example of each of the aspects of the present disclosure described herein may be implemented together in any suitable combination, and any feature / example from any one or more of these aspects may be implemented together in any suitable combination with any of the features of the other aspects described herein, in order to achieve the benefits described herein.

Claims

1. 1. A method of sequencing a polynucleotide using a nanopore comprising a first side, a second side, and an opening extending through said first and second sides, comprising: (a) placing a polynucleotide through the opening of the nanopore such that a 3′ end of the polynucleotide is on the first side of the nanopore and a 5′ end of the polynucleotide is on the second side of the nanopore; (b) forming a duplex with the polynucleotide on the first side of the nanopore, the duplex including a 3′ end; (c) extending the duplex at the first side of the nanopore by adding a first nucleotide to the 3′ end of the duplex; (d) applying a force that positions the 3' end of the extended duplex within the opening, and while applying the force: using the nanopore to inhibit translocation of the 3′ end of the extended duplex to the second side of the nanopore; measuring values ​​of electrical properties of the 3' end of the extended duplex and the single-stranded portion of the polynucleotide, determining that the value determined in act (d) is based on at least M nucleotides of the single-stranded portion of the polynucleotide and D pairs of hybridized nucleotides of the extended duplex, where M is 2 or more and D is 1 or more; (e) identifying the first nucleotide using the value determined in act (d).

2. the value measured in operation (d) comprises a current, an ionic current, an electrical resistance, or a voltage drop across the nanopore; and / or 10. The method of claim 1, wherein the value measured in operation (d) comprises noise in the current, ionic current, electrical resistance, or voltage drop across the nanopore.

3. 2. The method of claim 1, wherein M is 3 or greater and / or D is 2 or greater.

4. 2. The method of claim 1, wherein at least one of the M nucleotides of the single-stranded portion comprises a modified base, the method comprising identifying the modified base using the value measured in operation (d), and optionally the modified base comprises a methylated base.

5. 10. The method of claim 1, wherein the method further comprises inhibiting addition of another nucleotide to the 3' end of the extended duplex while the force is applied in operation (d), and optionally, the nanopore inhibits addition of another nucleotide.

6. (f) again applying a modified force to position the 3′ end of the extended duplex within the opening, and while applying the modified force, using the nanopore to inhibit translocation of the 3′ end of the extended duplex to the second side of the nanopore, and measuring values ​​of an electrical property of the 3′ end of the extended duplex and the single-stranded portion of the polynucleotide; 10. The method of claim 1, further comprising: (g) identifying the first nucleotide using the value determined in operation (f).

7. Claim 7: i) the modified force of operation (f) positions the 3' end of the extended duplex at a different position relative to the nanopore than the force of operation (d); and / or ii) the modified force of operation (f) is applied before adding another nucleotide, or the modified force of operation (f) is applied while resequencing the polynucleotide.

8. The method of claim 6, wherein the first nucleotide is identified using a comparison of the value measured in operation (d) with the value measured in operation (f), or wherein a trained neural network identifies the first nucleotide based on input of the value measured in operation (d) and the value measured in operation (f).

9. the first nucleotide is added using a polymerase that contacts the 3' end of the duplex, and optionally The method further comprises reversibly inhibiting the polymerase from adding a second nucleotide to the 3' end of the extending duplex, and optionally a blocking moiety that reversibly inhibits the polymerase from adding the second nucleotide to the 3' end of the extending duplex, and optionally The method further comprises removing the blocking moiety to allow the polymerase to add the second nucleotide to the 3' end of the extended duplex, and optionally i) the force applied in (d) removes the polymerase from contact with the 3' end of the extended duplex, or the method further comprises applying a force different from the force applied in (d) to remove the polymerase from contact with the 3' end of the extended duplex; and / or ii) The method of claim 1, wherein the polymerase comprises a DNA polymerase, or the polymerase comprises an RNA polymerase, or the polymerase comprises a reverse transcriptase.

10. the extended duplex comprises one or more nucleotide analogues, optionally wherein the one or more nucleotide analogues enhance the stability of the extended duplex compared to natural nucleotides, or optionally wherein the one or more nucleotide analogues comprise one or more locked nucleic acids (LNAs), one or more 2'-methoxy (2'-OMe) nucleotides, or one or more 2'-fluorinated (2'-F) nucleotides, and optionally 2. The method of claim 1, wherein the one or more nucleotide analogues alter the value of the electrical property compared to a natural nucleotide, or the first nucleotide comprises one of the one or more nucleotide analogues, and further optionally, the one or more nucleotide analogues comprise a 2' modification or a base modification.

11. the force applied in act (d) is insufficient to cause dissociation of the extended duplex; or the force applied in operation (d) comprises a first voltage; or actions (b) and (c) are performed in the absence of the force applied in action (d); or The method of claim 1 , wherein operation (c) is performed in the presence of a force that opposes the force applied in operation (d).

12. a first locking structure is attached to the 3′ end of the polynucleotide on the first side of the nanopore, the first locking structure inhibiting translocation of the 3′ end of the polynucleotide through the opening to the second side of the nanopore, and optionally the first locking structure is removable; or 2. The method of claim 1, wherein a second locking structure is attached to a 5' end of the polynucleotide on the second side of the nanopore, the second locking structure inhibiting translocation of the 5' end of the polynucleotide through the opening to the first side of the nanopore, and optionally the second locking structure is removable.

13. The method of claim 1, further comprising, after operation (d), dissociating the extended duplex from the polynucleotide and forming a new duplex with the polynucleotide on the first side of the nanopore, the new duplex including a new 3' end, and optionally, the polynucleotide being resequenced in response to sequencing of the polynucleotide being substantially complete or the polynucleotide being resequenced as a corrective measure.

14. The method further comprising, after the new duplex is formed, resequencing the polynucleotide by repeating operations (c) through (e), and optionally:

14. The method of claim 13, wherein the method further comprises generating a consensus read using the identification from the sequencing of the polynucleotide in operation (e) and the identification from resequencing the polynucleotide in operation (e).

15. 2. The method of claim 1, wherein operation (a) comprises contacting the nanopore with the polynucleotide hybridized to a substantially complementary polynucleotide and applying a force that dehybridizes the substantially complementary polynucleotide from the polynucleotide.

16. 1. A sequencing system comprising: a nanopore including a first side, a second side, and an opening extending through the first and second sides; a polynucleotide positioned through the opening of the nanopore such that a 3' end of the polynucleotide is on the first side of the nanopore and a 5' end of the polynucleotide is on the second side of the nanopore; a duplex having the polynucleotide disposed on the first side of the nanopore, the duplex comprising a 3' end at which a first nucleotide is disposed; A circuit comprising: applying a force to position the 3' end of the duplex within the opening; measuring values ​​of electrical properties of the 3' end of the duplex and the single-stranded portion of the polynucleotide while applying the force; the determined value is based on at least M nucleotides of the single-stranded portion of the polynucleotide and D pairs of hybridized nucleotides of the extended duplex, where M is 2 or more and D is 1 or more; and a circuit configured to identify the first nucleotide using the determined value; The nanopore inhibits translocation of the 3' end of the duplex to the second side of the nanopore while the force is applied.

17. the value measured by the circuit comprises current, ionic current, electrical resistance, or voltage drop across the nanopore; or 17. The system of claim 16, wherein the value measured by the circuit comprises noise in the current, ionic current, electrical resistance, or voltage drop across the nanopore.

18. 17. The system of claim 16, wherein M is 3 or greater, or D is 2 or greater.

19. 17. The system of claim 16, wherein at least one of the M nucleotides of the single-stranded portion comprises a modified base, and wherein the circuit is configured to identify the modified base using the value measured by the circuit, and optionally, the modified base comprises a methylated base.

20. 17. The system of claim 16, wherein addition of another nucleotide to the 3' end of the duplex is inhibited while the force is applied, and optionally, the nanopore inhibits the addition of another nucleotide.

21. The circuit applying a modified force that repositions the 3' end of the duplex within the opening; measuring values ​​of electrical properties of the 3' end of the duplex and the single-stranded portion of the polynucleotide while applying the modified force; further configured to identify the first nucleotide using the determined value; 17. The system of claim 16, wherein the nanopore inhibits translocation of the 3' end of the duplex to the second side of the nanopore while the modified force is applied.

22. (i) the modified force positions the 3' end of the extended duplex at a different position relative to the nanopore than the force positions it; and / or ii) the modified force is applied before adding another nucleotide, or the modified force is applied while resequencing the polynucleotide.

23. The system of claim 21, wherein the first nucleotide is identified using a comparison of the value measured under the force with the value measured under the modified force, or wherein a trained neural network identifies the first nucleotide based on input of the value measured under the force and the value measured under the modified force.

24. the system further comprising a polymerase that contacts the 3' end of the duplex and adds the first nucleotide; wherein the polymerase is reversibly inhibited from adding a second nucleotide to the 3' end of the duplex, and optionally the system further comprising a blocking moiety that reversibly inhibits the polymerase from adding the second nucleotide to the 3' end of the duplex; 17. The system of claim 16, wherein the blocking moiety is removable to allow the polymerase to add the second nucleotide to the 3' end of the duplex.

25. the force removes the polymerase from contact with the 3' end of the duplex; or 25. The system of claim 24, wherein the circuitry is further configured to apply a force to remove the polymerase from contact with the 3' ends of the duplex, wherein the force used to remove the polymerase is different from the force used to position the 3' ends of the duplex within the opening.

26. 25. The system of claim 24, wherein the polymerase comprises a DNA polymerase, an RNA polymerase, or a reverse transcriptase.

27. the duplex comprises one or more nucleotide analogues, optionally wherein the one or more nucleotide analogues enhance the stability of the duplex compared to natural nucleotides, or the one or more nucleotide analogues comprise one or more locked nucleic acids (LNA), or one or more 2'-methoxy (2'-OMe) nucleotides or one or more 2'-fluorinated (2'-F) nucleotides, optionally 17. The system of claim 16, wherein the one or more nucleotide analogs alter the value of the electrical property compared to a natural nucleotide, or the first nucleotide comprises one of the one or more nucleotide analogs, and optionally the one or more nucleotide analogs comprise a 2' modification or a base modification.

28. 17. The system of claim 16, wherein the force used to position the 3' ends of the duplex within the opening is insufficient to cause dissociation of the duplex, or wherein the force used to position the 3' ends of the duplex within the opening comprises a first voltage.

29. 17. The system of claim 16, wherein the polynucleotide is positioned through the opening and the duplex is positioned on the first side of the nanopore in the absence of the force used to position the 3' ends of the duplex within the opening, or the circuit is configured to apply a force that counteracts the force used to position the 3' ends of the duplex within the opening.

30. a first locking structure is attached to the 3′ end of the polynucleotide on the first side of the nanopore, the first locking structure inhibiting translocation of the 3′ end of the polynucleotide through the opening to the second side of the nanopore, and optionally the first locking structure is removable; or 17. The system of claim 16, wherein a second locking structure is attached to the 5' end of the polynucleotide on the second side of the nanopore, the second locking structure inhibiting translocation of the 5' end of the polynucleotide through the opening to the first side of the nanopore, and optionally the second locking structure is removable.

31. 17. The system of claim 16, wherein the circuit is configured to dissociate the duplex from the polynucleotide and form a new duplex after applying the force used to position the 3' end of the duplex within the opening.

32. i) the circuitry is configured to resequence the polynucleotide after the new duplex is formed, and optionally the circuitry is configured to generate a consensus read using the identification from sequencing the polynucleotide and the identification from resequencing the polynucleotide; and / or ii) the circuitry is configured to resequence the polynucleotide in response to sequencing of the polynucleotide being substantially complete, or the circuitry is configured to resequence the polynucleotide as a corrective action.

33. 17. The system of claim 16, wherein the circuitry is configured to apply a force to dehybridize a substantially complementary polynucleotide from the polynucleotide to displace the polynucleotide through the opening of the nanopore.

34. i) the nanopore comprises a solid-state nanopore or a biological nanopore, optionally the biological nanopore comprises MspA; and / or ii) the polynucleotide comprises RNA or DNA, and / or iii) The method of any one of claims 1 to 15 or the system of any one of claims 16 to 33, wherein the double strand comprises a primer hybridized to the polynucleotide.