Controlled polynucleotide transfer in nanopore sequencing
By attaching modifications like circular loop structures with stopping constructs to nucleotides, the method controls polynucleotide migration through nanopores, improving sequencing accuracy and reducing errors, thus enabling efficient and cost-effective polynucleotide sequencing.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- ILLUMINA INC
- Filing Date
- 2024-04-26
- Publication Date
- 2026-05-19
AI Technical Summary
Existing polynucleotide sequencing techniques face challenges in achieving accurate and efficient sequencing due to rapid migration of polynucleotides through nanopores, leading to errors and reduced signal-to-noise ratios, particularly under pulsed voltage regimes.
The method involves attaching modifications, such as circular loop structures with stopping constructs, to nucleotides to control the migration speed of polynucleotides through nanopores, allowing for controlled movement and improved signal detection.
This approach enhances sequencing accuracy by ensuring single-base resolution and reduces errors, simplifies base calling algorithms, and enables high-throughput, lower-cost polynucleotide sequencing.
Smart Images

Figure 2026515799000001_ABST
Abstract
Description
[Technical Field]
[0001] Some polynucleotide sequencing techniques involve carrying out a large number of controlled reactions on a support surface or inside a predetermined reaction chamber. These controlled reactions can then be observed or detected, and subsequent analysis can help identify the characteristics of the polynucleotides involved in the reaction. Examples of such sequencing techniques include ligation sequencing, synthetic sequencing, reversible terminator chemistry, or next-generation sequencing or large-scale parallel sequencing with pyrosequencing approaches.
[0002] Some polynucleotide sequencing techniques can utilize nanopores to provide pathways for ionic currents. For example, when polynucleotides traverse a nanopore, it influences the current flowing through the nanopore. Each traversing nucleotide, or a sequence of nucleotides, through the nanopore yields a characteristic current. These characteristic currents of the traversing polynucleotides can be recorded to determine the polynucleotide sequence. [Overview of the project]
[0003] Examples provided herein offer methods for sequencing biopolymers, particularly polynucleotides, as well as systems and kits for carrying out such methods.
[0004] Each of the systems, devices, kits, and methods disclosed herein has several embodiments, and no single embodiment alone possesses any of their desirable attributes. Without limiting the scope of the claims, some notable features are briefly considered here. Numerous other embodiments are also intended to include embodiments having fewer, additional, and / or different components, processes, features, subjects, benefits, and advantages. The components, embodiments, and steps may be arranged and ordered differently. After considering this consideration, and especially after reading the section titled “Modes for Carrying Out the Invention,” you will understand how the features of the devices and methods disclosed herein are advantageous over other known devices and methods.
[0005] In one embodiment, the foregoing provides a method for determining the sequence of a target polynucleotide in a nanopore-based sequencing system, the method comprising providing a target polynucleotide comprising a nucleotide, each nucleotide being coupled to a modification configured to terminate the target polynucleotide in a nanopore. In one embodiment, the modification comprises a termination construct. In some embodiments, the modification further comprises a circular loop, the termination construct being coupled to the circular loop. In some embodiments, the circular loop further comprises a reporter element encoding a nucleotide. In some embodiments, the circular loop further comprises a spacer. The method may further comprise applying a driving voltage to move one or more portions of the target polynucleotide through a nanopore, continuously measuring the current in the nanopore during the movement, and identifying the sequence of the target polynucleotide by correlating the measured current with the identity of one or more nucleotides.
[0006] The first reporter element may contain one or more nucleotides in addition to other parts, such as modifications bound to nucleotides. The first electrical response may be an ionic current passing through the nanopore. The modifications may be covalently bonded to the nucleotide. The modifications may interact with the nanopore non-covalently. In some embodiments, the first electrical response depends on the modifications bound to the nucleotides in the first reporter element.
[0007] The method may further include applying a voltage for a duration sufficient to move a first reporter element through a nanopore. The method may further include, during a second stopping event, applying a reading voltage across the entire nanopore to identify a second reporter element in a constricted portion of the nanopore based on a second electrical response in the system, the second electrical response depending on the identity of the second reporter element. The second reporter element may include one or more nucleotides in addition to other parts, e.g., nucleotide-bound modifications. The second electrical response may be an ionic current through the nanopore. In some embodiments, nucleotide-bound modifications also move through the nanopore. In some embodiments, the voltage is configured to be just enough to move one nucleotide at a time through the nanopore. In some embodiments, the second electrical response depends on the nucleotide-bound modifications in the second reporter element. In some embodiments, the voltage is configured to be sufficient to sequentially move multiple nucleotides through the nanopore.
[0008] In some embodiments, the method may further include monitoring the electrical response in the system when a voltage is applied so as to determine that only the first reporter element moves through the nanopore, and stopping the voltage after the first reporter element has moved through the nanopore.
[0009] In some embodiments, a method for determining the sequence of a target polynucleotide in a nanopore-based sequencing system is provided herein, comprising: providing a target polynucleotide comprising a nucleotide, each nucleotide being bound to a modification, the modification comprising a stopping construct configured to stop the target polynucleotide relative to a nanopore; applying a driving voltage to move one or more portions of the target polynucleotide through the nanopore; continuously measuring the current in the nanopore during the movement; and identifying the sequence of the target polynucleotide by correlating the measured current with the identity of the nucleotide.
[0010] In another embodiment, a kit is provided for use in carrying out the disclosed method. The kit may include a nucleotide having any of the modifications described herein.
[0011] In yet another embodiment, a system or device is provided configured to determine the sequence of a target polynucleotide using one of the disclosed methods.
[0012] In some embodiments, the techniques described herein involve providing a target polynucleotide by synthesizing a daughter chain based on a template polynucleotide using a modified nucleotide, wherein the modification is covalently bonded to the nucleotide.
[0013] In some embodiments, the techniques described herein include a method for providing a target polynucleotide, comprising synthesizing a daughter chain based on a template polynucleotide using a modified nucleotide, wherein the modification is covalently bonded to the nucleotide, and cleaving the daughter chain to produce a target polynucleotide having an elongated polynucleotide chain.
[0014] In some embodiments, the technology described herein relates to a method for keeping a drive voltage constant during movement.
[0015] In some embodiments, the techniques described herein relate to a method by which the measured current depends on a reporter element or nucleotide passing through a nanopore.
[0016] In some embodiments, the techniques described herein relate to a method in which the termination construct comprises a polymer selected from the group consisting of linear synthetic hydrophilic polymers, linear synthetic hydrophobic polymers, linear polynucleotides, linear polypeptides, branched polymers, dendritic polymers, cyclic polymers, fluoroalkyls, rigid conjugated chromophores, and rigid macrorings.
[0017] In some embodiments, the techniques described herein relate to methods by which a termination construct comprises a covalent bond between a polymer and a corresponding nucleotide or cyclic loop.
[0018] In some embodiments, the techniques described herein relate to a method for selecting a covalent bond from the group consisting of amine-NHS esters, amine-imide esters, amine-pentafluorophenyl esters, amine-hydroxymethylphosphine, carboxyl-carbodiimide, thiol-maleimide, thiol-haloacetyl, thiol-pyridyl disulfide, thiol-thiosulfonate, thiol-vinylsulfone, aldehyde-hydrazide, aldehyde-alkoxyamine, hydroxyisocyanate, azido-alkyne, azido-phosphine, transcyclooctene-tetrazine, norbornene-tetrazine, azido-cyclooctin, and azido-norbornene.
[0019] In some embodiments, the techniques described herein relate to a method for selecting a linear synthetic hydrophilic polymer from the group consisting of polyethylene glycol, polyvinyl alcohol, polyacrylamide, polyvinylpyrrolidone, polystyrene sulfonate, polyethyleneimine, and combinations thereof.
[0020] In some embodiments, the technology described herein relates to methods in which the linear synthetic hydrophobic polymer is selected from the group consisting of polylactic acid, polymethyl methacrylate, polystyrene, and combinations thereof.
[0021] In some embodiments, the technology described herein relates to methods in which the linear polynucleotide is a homopolymer of natural nucleotides, a homopolymer of unnatural nucleotides, a mixed sequence polymer of natural nucleotides, or a mixed sequence polymer of unnatural nucleotides.
[0022] In some embodiments, the technology described herein relates to methods in which the linear polypeptide contains one or more types of amino acids.
[0023] In some embodiments, the technology described herein relates to methods in which the branched polymer contains two or more branches.
[0024] In some embodiments, the technology described herein relates to methods in which the cyclic polymer has three or more repeating units.
[0025] In some embodiments, the technology described herein relates to methods in which each repeating unit is a small molecule, a nucleotide, or an amino acid.
[0026] In some embodiments, the technology described herein relates to methods in which the repeating units are the same.
[0027] In some embodiments, the technology described herein relates to methods in which at least two of the repeating units are different.
[0028] In some embodiments, the technology described herein relates to a system for determining an array polynucleotide or oligonucleotide according to the methods disclosed herein.
[0029] In some embodiments, the techniques described herein relate to a cyclic loop nucleotide comprising a cyclic loop modification that bridges a nucleic acid base and a phosphate group, wherein the cyclic loop modification comprises a reporter encoding the identity of the nucleic acid base and a termination construct adjacent to the reporter.
[0030] In some embodiments, the techniques described herein relate to a cyclic loop nucleotide in which the termination construct is adjacent to the reporter.
[0031] In some embodiments, the techniques described herein relate to a cyclic loop nucleotide having one of the following structures.
[0032] [ka] In the formula, X is -O-, -CH2-, -NSO2-, -NH-
[0033] [ka] And X' is -S-, =N-SO2-, =NH-CO-, or
[0034] [ka] The bases are nucleic acid bases, L1 and L2 are linking groups, RP is a reporter that codes for nucleic acid bases, and ARC is a termination construct.
[0035] In some embodiments, the techniques described herein relating to cyclic loop nucleotides, wherein the ARC is covalently bonded to the cyclic loop via a covalent bond selected from amine-NHS esters, amine-imide esters, amine-pentafluorophenyl esters, amine-hydroxymethylphosphine, carboxyl-carbodiimide, thiol-maleimide, thiol-haloacetyl, thiol-pyridyl disulfide, thiol-thiosulfonate, thiol-vinylsulfone, aldehyde-hydrazide, aldehyde-alkoxyamine, hydroxy-isocyanate, azido-alkyne, azido-phosphine, transcyclooctene-tetrazine, norbornene-tetrazine, azido-cyclooctin, and azido-norbornene.
[0036] Further details of exemplary nanopore alignment devices that may be used in conjunction with the disclosed technology, and methods for operating such devices, can be found in U.S. Provisional Patent Application No. 63 / 200868 (International Publication No. 2022 / 005780) and No. 63 / 169041, the entirety of each disclosure incorporated herein by reference.
[0037] It will be understood that any feature of the devices and / or arrays disclosed herein can be combined together in any desired manner and / or configuration. Furthermore, it will be understood that any feature of the methods of using the devices can be combined together in any desired manner. Furthermore, it will be understood that any combination of features of the methods, and / or devices, and / or arrays can be used together and / or combined with any of the embodiments disclosed herein. Moreover, it will be understood that any feature or combination of features of any of the devices, and / or arrays, and / or methods can be combined together in any desired manner and / or combined with any of the embodiments disclosed herein.
[0038] It should be understood that all combinations of the aforementioned concepts and additional concepts, which will be discussed in more detail below, are considered to be part of the subject matter of the inventions disclosed herein and may be used to realize the benefits and advantages described herein. [Brief explanation of the drawing]
[0039] The features of the examples in this disclosure will become apparent from the following detailed description and drawings. In the drawings, similar reference numerals correspond to components that are similar but not identical. For brevity, reference numerals or features having the aforementioned functions may or may not be described in relation to other drawings in which they appear. [Figure 1] This shows an example of controlled DNA transfer using modifications. [Figure 2] This document presents an exemplary workflow for determining modified and controlled nanopore arrangement. [Figure 3] This shows an exemplary library preparation process using modified nucleotides. [Figure 4] This shows an exemplary DNA insertion process to initiate nanopore sequencing. [Figure 5] This illustrates an exemplary DNA transfer process during modification-controlled nanopore sequencing. [Figure 6] This shows an exemplary model system for screening stopped structures. [Figure 7] The results from experiments using the exemplary model system shown in Figure 6 are presented. [Figure 8] An example of a linear synthetic hydrophilic polymer termination structure is shown. [Figure 9] An example of a polynucleotide termination construct is shown. [Figure 10] An example of a branched, stopped structure is shown. [Figure 11] An example of a branched, stopped structure is shown. [Figure 12] An example of a branched, stopped structure is shown. [Figure 13] An example of a circular stop structure is shown. [Figure 14] This document illustrates embodiments of stationary structures with various characteristics. [Figure 15] Examples of common DNA base methylation sites are the extracyclic amine at position 6 of adenine and the fifth carbon on the cytosine ring. [Figure 16] This document illustrates an exemplary process for modifying dsDNA using DNA methyltransferase (MTase) and SAM analogues. [Figure 17] The following shows an exemplary structure of a modified SAM analog. [Figure 18] Figure 17 shows examples of stop structures that can be used in the exemplary structure shown. Top: Linear PEG. Bottom: Branched PEG. [Figure 19] An example of a SAM analog having a PEG4 group conjugated via a Cu-click reaction is shown. [Figure 20] This document describes an exemplary process for modifying 5-methylcytosine using CMD1. [Figure 21] This shows an exemplary circular loop modification. [Figure 22] An exemplary sequence containing multiple circular loop-modified nucleotides is shown. [Figure 23A] This shows experimental assays and average residence times of annular loop-modified constructs at different driving voltages. [Figure 23B] This shows experimental assays and average residence times of annular loop-modified constructs at different driving voltages. [Figure 23C] This shows experimental assays and average residence times of annular loop-modified constructs at different driving voltages. [Figure 24A] The experimental results and current traces of the annular loop modification structure are shown. [Figure 24B] The experimental results and current traces of the annular loop modification structure are shown. [Figure 24C] The experimental results and current traces of the annular loop modification structure are shown. [Figure 25A] The experimental results and current traces of the annular loop modification structure are shown. [Figure 25B]The experimental results and current traces of the annular loop modification structure are shown. [Figure 25C] The experimental results and current traces of the annular loop modification structure are shown. [Figure 25D] The experimental results and current traces of the annular loop modification structure are shown. [Modes for carrying out the invention]
[0040] All patents, applications, published applications, and other publications referenced herein are incorporated herein by reference in their entirety. Where any term or phrase is used herein in a manner that contradicts or otherwise contradicts the definitions contained in the patents, applications, published applications, and other publications incorporated herein by reference, the use herein shall prevail over the definitions incorporated herein by reference.
[0041] definition All technical and scientific terms used herein have the same meanings as those generally understood by those skilled in the art to which this disclosure pertains, unless otherwise defined.
[0042] As used herein, the singular forms “a,” “and,” and “the” refer to multiple objects unless the context explicitly indicates otherwise. Therefore, for example, a reference to “array” may include multiple such arrays.
[0043] The terms comprising, including, and containing, and their various forms, are synonymous and equally broad in meaning. Furthermore, unless otherwise explicitly stated, an example of having, including, or possessing one or more elements having a particular characteristic may include additional elements, regardless of whether those additional elements possess that characteristic.
[0044] As used herein, the term “nanopore” is intended to mean a hollow structure that is separate from or defined within a membrane and extends across the membrane. Nanopores allow ions, electric currents, and / or fluids to traverse from one side of a membrane to the other. For example, a membrane that inhibits the passage of ions or water-soluble molecules may include nanopore structures extending across the membrane to allow the passage of ions or water-soluble molecules from one side of the membrane to the other (through nanoscale openings extending through the nanopore structure). The diameter of the nanoscale openings extending through the nanopore structure may vary along its length (i.e., from one side of the membrane to the other), but at any point it is nanoscale (i.e., about 1 nm to about 100 nm, or less than 1000 nm). Examples of nanopores include, for example, biological nanopores, solid-state nanopores, and biological and solid-state hybrid nanopores. In some embodiments, nanopores refer to pores having an opening with a diameter of about 0.3 nm to about 2 nm at their narrowest point. For example, nanopores may be solid-state nanopores, graphene nanopores, elastomer nanopores, or native or recombinant proteins that form tunnels upon insertion into bilayers, thin films, membranes, or solid openings, also referred herein as protein pores or protein nanopores (e.g., transmembrane pores). When a protein is inserted into a membrane, the protein is a tunnel-forming protein.
[0045] As used herein, the term “diameter” is intended to mean the longest inscribed straight line in the cross-section of the nanoscale aperture, passing through the center of mass of the cross-section of the nanoscale aperture. It should be understood that the nanoscale aperture may or may not have a circular or substantially circular cross-section (a cross-section of the nanoscale aperture substantially parallel to the cis / trans electrodes). Furthermore, the cross-section may be regular or irregular in shape.
[0046] As used herein, the term “biological nanopore” is intended to mean a nanopore whose structural components are made from biologically derived materials. Biologically derived refers to materials derived from or isolated from biological environments such as organisms or cells, or from synthetically produced variants of biologically usable structures. Examples of biological nanopores include polypeptide nanopores and polynucleotide nanopores.
[0047] As used herein, the term “polypeptide nanopore” is intended to mean a protein / polypeptide that extends across a membrane and allows ions, electric currents, polymers such as DNA or peptides, or other molecules of appropriate size and charge, and / or fluids, to flow from one side of the membrane to the other. Polypeptide nanopores can be monomers, homopolymers, or heteropolymers. Structures of polypeptide nanopores include, for example, α-helix bundle nanopores and β-barrel nanopores. Examples of polypeptide nanopores include α-hemolysin, Mycobacterium smegmatis porin A (MspA), gramidiin A, maltoporin, OmpF, OmpC, PhoE, Tsx, F pili, etc. The protein α-hemolysin is naturally found in cell membranes and functions as a pore for ions or molecules transported in and out of cells. Mycobacterium smegmatis porin A (MspA) is a membrane porin produced by mycobacteria that allows hydrophilic molecules to enter the bacteria. MspA is goblet-like and forms tightly interconnected octamers and transmembrane beta barrels containing a central pore.
[0048] Polypeptide nanopores can be synthetic. Synthetic polypeptide nanopores contain protein-like amino acid sequences that do not occur naturally. Protein-like amino acid sequences may contain some amino acids that are known to exist but do not form the basis of proteins (i.e., non-proteinogenic amino acids). Protein-like amino acid sequences can be synthesized artificially and then purified / isolated, rather than being expressed in organisms.
[0049] The nanopores disclosed herein may be hybrid nanopores. “Hybrid nanopores” refers to nanopores containing both biological and non-biological materials. For example, hybrid nanopores may include biological materials such as polypeptides and / or polynucleotides adjacent to and / or conjugated to non-biological materials such as semiconductors and / or other solid-state embodiments. Therefore, examples of hybrid nanopores may include polypeptide solid-state hybrid nanopores and polynucleotide solid-state nanopores.
[0050] Applying a potential difference across nanopores can force the rearrangement of nucleic acids passing through the nanopores. One or more signals are generated in response to the movement of nucleotides through the nanopores. Thus, when a target polynucleotide, or an oligonucleotide or mononucleotide or probe derived from a target polynucleotide or mononucleotide, passes through the nanopore, the current across the membrane changes, for example, due to the blocking of the base-dependent (or probe-dependent) aspect of the constriction. The signals from this change in current can be measured using one of several methods. Each signal is specific to the species of nucleotide (or probe) in the nanopore, so that the resulting signal can be used to determine the character of the polynucleotide or oligonucleotide. For example, the identity of one or more species of nucleotides (or probes) that produce characteristic signals can be determined.
[0051] As used herein, “nucleotide” comprises a nitrogen-containing heterocyclic base, a sugar, and one or more phosphate groups. A nucleotide is a monomeric unit of a nucleic acid sequence. Examples of nucleotides include, for example, ribonucleotides or deoxyribonucleotides. In ribonucleotides (RNA), the sugar is ribose, and in deoxyribonucleotides (DNA), the sugar is deoxyribose, i.e., a sugar lacking the hydroxyl group at the 2' position of ribose. The nitrogen-containing heterocyclic base may be a purine base or a pyrimidine base. Examples of purine bases include adenine (A) and guanine (G), and their modified derivatives or analogs. Examples of pyrimidine bases include cytosine (C), thymine (T), and uracil (U), and their modified derivatives or analogs. The C-1 atom of deoxyribose is bonded to N-1 of pyrimidine or N-9 of purine. The phosphate group may be monophosphate, diphosphate, or triphosphate. While these nucleotides are natural nucleotides, it should be further understood that non-natural nucleotides, modified nucleotides, or analogues of the aforementioned nucleotides may also be used.
[0052] As used herein, the term “signal” is intended to mean an indicator representing information. Signals include, for example, electrical signals and optical signals. The term “electrical signal” refers to an indicator of electrical quality representing information. Indicators can be, for example, current, voltage, tunneling, resistance, potential, conductance, or lateral electrical effect. “Electron current” or “current” refers to the flow of electric charge. In embodiments, an electrical signal may be a current passing through a nanopore, and a current may flow when a potential difference is applied across the nanopore.
[0053] As used herein, the terms “driving force” or “driving voltage” are intended to mean an electric current that enables at least a portion of a polynucleotide or oligonucleotide to move through a nanopore. In some embodiments, an electric current may flow when a potential difference is applied across the nanopore.
[0054] As used herein, the term “retaining force” is intended to mean the resistance that slows and / or stops polynucleotides or oligonucleotides from moving through nanopores. In some embodiments, the retaining force is overcome by the application of a driving force. Thus, the driving force overcomes / neutralizes the resistance that slows and / or stops the polynucleotides, thereby allowing the polynucleotides to move through nanopores.
[0055] As used herein, the term “modification” is intended to mean a portion attached to a nucleotide. A modification may include a termination construct or a cyclic loop portion further containing a termination construct. A modification can be attached to any portion of a nucleotide, or it may be attached to a nucleotide at two positions that form a loop.
[0056] As used herein, the term “stopping construct” means a construct that can provide resistance (in the form of “retaining force”) that slows and / or stops polynucleotides, oligonucleotides, or mononucleotides from moving through nanopores, unless the resistance provided by the stopping construct is overcome by a “driving force”. The resistance provided by the stopping construct is due to the properties of the stopping construct (e.g., size, geometry, and / or non-covalent interaction with nanopores). The stopping construct can act as a ratchet or brake for polypeptide movement through nanopores.
[0057] As used herein, the term “stop” is intended to mean stopping and / or slowing down. For example, when the movement of polynucleotides through a nanopore is stopped, the relative motion of the polynucleotides with respect to the nanopore may stop or continue at a slower rate. A stopping construct that stops the movement of nucleotides through a nanopore may function to stop or slow down the movement compared to unmodified nucleotides.
[0058] As used herein, “reporter element” includes those known as “tag” or “label.” A reporter element may comprise one or more nucleotides or polymers. Four reporter elements may encode each of a nucleic acid base that produces a distinguishable signal (fingerprint / signature) when passing through a nanopore reading head. Multiple reporters or reporter signals may also encode nucleic acid bases. A reporter may consist of two or more sub-reporters.
[0059] The terms “nucleic acid” or “polynucleotide” refer to single-stranded or double-stranded deoxyribonucleotides or ribonucleotide polymers and, unless otherwise specified, include known analogues of naturally occurring nucleotides that hybridize to nucleic acids in a manner similar to naturally occurring nucleotides such as peptide nucleic acids (PNAs) and phosphorothioate DNA. Unless otherwise specified, a particular nucleic acid sequence includes its complementary sequence. Examples of nucleotides include, but are not limited to, ATP, dATP, CTP, dCTP, GTP, dGTP, UTP, TTP, dUTP, 5-methyl-CTP, 5-methyl-dCTP, ITP, dITP, 2-amino-adenosine-TP, 2-amino-deoxyadenosine-TP, 2-thiothymidine triphosphate, pyrrolo-pyrimidine triphosphate, and 2-thiocytidine, as well as alpha-thio triphosphate for all of the above, and 2'-O-methyl-ribonucleotide triphosphate for all of the above bases. Examples of modified bases include, but are not limited to, 5-Br-UTP, 5-Br-dUTP, 5-F-UTP, 5-F-dUTP, 5-propynyl-dCTP, and 5-propynyl-dUTP. Polynucleotides and oligonucleotides may be interchangeable as used herein.
[0060] As used herein, “nucleic acid base” refers to heterocyclic bases, such as adenine, guanine, cytosine, thymine, uracil, inosine, xanthine, hypoxanthine, or their heterocyclic derivatives, analogs, or tautomers. Nucleic acid bases may be naturally occurring or synthesized. Non-restrictive examples of nucleic acid bases include adenine, guanine, thymine, cytosine, uracil, xanthine, hypoxanthine, 8-azapurine, purines substituted with methyl or bromine at position 8, 9-oxo-N6-methyladenine, 2-aminoadenine, 7-deazaxanthine, 7-deazaguanine, 7-deaza-adenine, N4-ethanocytosine, 2,6-diaminopurine, N6-ethano-2,6-diaminopurine, 5-methylcytosine, 5-(C3~C6)-alkynylcytosine, 5-fluorouracil, and 5-bromouracil. These are nucleic acid bases that do not exist in nature, as described in U.S. Patents Nos. 5,432,272 and 6,150,510, and International Publications Nos. 92 / 002258, 93 / 10820, 94 / 22892, and 94 / 24144, and Fasman ("Practical Handbook of Biochemistry and Molecular Biology", pp. 385-394, 1989, CRC Press, Boca Raton, LO) (all of which are incorporated herein by reference in their entirety).
[0061] As used herein, “peptide” means two or more amino acids linked to one another by an amide bond (i.e., a “peptide bond”). A peptide may contain up to 50 amino acids. A peptide may be linear or cyclic. A peptide may be α, β, γ, δ, or higher, or a mixture of them. A peptide may contain any mixture of amino acids as defined herein, including any combination of D, L, α, β, γ, δ, or higher-order amino acids.
[0062] As used herein, “part” is one of two or more parts of something that can be divided (e.g., different parts of a tether, molecule, or probe).
[0063] As used herein, “cis” refers to the side of a nanopore opening into which the analyte or modified analyte enters. The “cis” side may depend on the applied positive or negative voltage.
[0064] As used herein, “trans” refers to the side of a nanopore opening from which the analyte or modified analyte (or fragment thereof) exits the opening.
[0065] The embodiments and examples described herein and enumerated in the claims can be understood in consideration of the above definitions.
[0066] Introduction The disclosed technology relates to a strand-based nucleic acid nanopore sequencing approach, e.g., a system and method for polynucleotide or DNA strand-based nanopore sequencing. The strand sequence can be determined by reading the base-specific changes in the ionic current through the nanopore as the polynucleotide or oligonucleotide strand passes through the constriction of the nanopore. However, the sequencing readout depends not only on the size of the constriction in the nanopore but also on the rate of movement of the polynucleotide through the nanopore (i.e., the migration rate).
[0067] The movement of polynucleotides can be driven by a voltage bias that drives the movement of bases through a nanopore reading head. In some embodiments, the voltage bias is constant and within a set range of a predetermined voltage. In some embodiments, the voltage bias cycles up to a readout or drive voltage, and the voltage is sufficient to drive the movement. The movement speed under the effective drive voltage in the nanopore is 10 6The range can be in the range of bases / second. In this range, the movement speed is too fast and can negatively affect reading accuracy. Due to the inherent limitations of current detection electronics, it is desirable to reduce the movement speed by several orders of magnitude to achieve single-base resolution. Movement can be driven by a constant voltage bias (e.g., automatic forward movement) or by transient voltage changes (e.g., pulses) that drive the movement of bases through the nanopore reading head.
[0068] Several methods for reducing migration speed rely on so-called motor enzymes (e.g., helicases and polymerases). However, these methods have several drawbacks. Firstly, migration speed is controlled by enzymes. Each migration event may occur too rapidly to obtain a suitable signal-to-noise ratio, leading to sequencing errors. Each migration event may not occur at precisely regular timings. Even when enzyme motors are engineered to have more regular motion, irregular migration can still occur. Secondly, enzymes have other drawbacks, including backstepping and energy requirements that can lead to errors and complicate the system.
[0069] In some cases, transient voltage pulses can advance the movement of polynucleotides or oligonucleotides through a nanopore reading head. In certain embodiments, when sequencing polynucleotides or oligonucleotides, stop constructs can halt their movement through the nanopores. However, under certain circumstances, such as under a pulsed voltage regime, the movement of polynucleotides through the nanopores can result in skipping or stalling. For example, skipping may result in a signal from an undetected reporter, and stalling may result in the detection of the same reporter after a pulse intended to further move the polynucleotide. In some cases, if the movement of polynucleotides results in skipping, the voltage pulse may move multiple bases and / or stop constructs, thereby generating transient or even non-existent signals as multiple bases pass through the reading head. The effect of skipping then truncates and produces errors in the measured sequence of the target polynucleotide. In some embodiments, the controlled movement of certain modified nucleic acid bases may be skipped at a rate of 1-10%. Alternatively, if stalling occurs, multiple transient voltage pulses may be required to advance the base within the nanopore reading head, thereby slowing down the overall advancement of the polynucleotide.
[0070] Under pulsed voltage regimes, an average of 3–10 voltage pulses could be applied before the bases successfully advanced across the nanopore reading head. Therefore, either skipping or stalling can increase the error rate when sequencing nucleic acids. Several methods explore the use of DNA polymerase-nanopore conjugates for sequencing template nucleic acids, where the maximum achievable read length and sequencing accuracy are influenced by the stability of the polymerase-template complex. In some examples, constant voltage bias advancement can be utilized for nucleic acid sequencing, where controlled movement of modified nucleic acid bases occurs without the need for voltage pulses, or where the background voltage is sufficient to drive the bases through the nanopores. In some embodiments where a constant voltage bias is utilized, the systems of this disclosure can detect changes in the reporter signal regardless of duration.
[0071] In some embodiments, this disclosure relates to the second consideration described above, namely, strategies for addressing migration speed. A target polynucleotide or oligonucleotide (e.g., modified DNA) may have modifications designed to slow down migration speed. The modifications may include a stopping construct. In some embodiments, the modifications may include a circular loop portion containing a stopping construct. In some embodiments, the selection of stopping constructs and an understanding of their interaction with nanopores are provided herein. In some embodiments, pore mutations or modifications to provide enhanced stopping construct-based control of nucleic acid migration are provided herein. In some embodiments, the synthesis of novel modifications having properties tuned for enhanced nucleic acid migration behavior is provided herein. In some embodiments, chemiactivations for selective conjugation of modifications are provided herein.
[0072] In some embodiments, to enable controlled single-base movement of polynucleotides or oligonucleotides (such as DNA) through nanopores and to improve the accuracy and efficiency of reading polynucleotide sequences, the concept of attaching “modifications” to daughter strand nucleotides for insertion through nanopores is provided herein. In some embodiments, the modifications enable controlled movement of polynucleotides through nanopores. The modifications may include a resting construct. In some embodiments, the modifications may include a circular loop portion, which further includes a resting construct. Without being limited by any particular theory, the size, geometry, and charge of the resting construct directly affect the movement rate of the polynucleotide. In some embodiments, due to its size and interaction with the nanopore, the resting construct slows and / or stops the movement of the polynucleotide under electrical conditions suitable for reading the polynucleotide. In some embodiments, the application of a voltage results in a driving force and / or a change in electrical conditions suitable for the movement of polynucleotides through nanopores. The driving force overcomes the retaining force that holds the polynucleotide in place as a result of the interaction between the resting construct and the nanopore.
[0073] This controlled migration method may offer several significant advantages over conventional chain-based approaches. In some embodiments, stopping and / or slowing down the migration of target polynucleotides or oligonucleotides by modification significantly improves the temporal resolution of bases, resulting in more accurate base identification. In some embodiments, the control of migration mitigates the homopolymer problem (difficulty in accurately identifying sequences of repeating bases) typically associated with chain-based nanopore sequencing methods. In some embodiments, well-regarded timing between base signals may simplify the base calling algorithm, as it may allow the algorithm to focus on relevant portions of the signal stream.
[0074] In some embodiments, controlled polynucleotide movement is achieved by attaching a modified portion to the nucleotide, which includes a stopping construct that slows or stops the movement due to its physicochemical properties when it encounters a nanopore. In some embodiments, by using cleavable sites along the polynucleotide backbone while attaching adjacent nucleic acid bases to the barcode region, the disclosed technique allows for increasing the distance between adjacent nucleic acid bases, eliminating the need to deconvolve a large number of signals. In some embodiments, the disclosed technique allows only one elongated polynucleotide to be present in the read head at any given time, successfully reducing read variability to four (A, T, C, and G), enabling lower-cost and more accurate sequencing. In some embodiments, the disclosed technique provides high-throughput, lower-cost, and more accurate polynucleotide sequencing. In some embodiments, no change in voltage bias is required to advance the nucleotide through the nanopore read head. In some embodiments, a change in voltage bias is required to advance the nucleotide through the nanopore read head. Accordingly, in some embodiments, a kit for polynucleotide sequencing is provided, containing at least a nucleotide having modifications such as those disclosed herein.
[0075] In some embodiments, modifications can be used to control polynucleotide movement by modifying any nucleotide on the strand supporting the modification, and the modification moves along with the nucleotide through the nanopore. As shown in Figure 1, protein nanopores 120 are deposited within a lipid bilayer 130. Single-stranded DNA 110 passes from the "cis" side through the nanopore 120 to the "trans" side. DNA 110 contains nucleotides to which the modification is bound. For example, nucleotide 111 ("G" base) is bound to modification 117, and the modification contains a stopping construct. Although not bound by theory, contact between the stopping construct on the nucleotide and a portion of the nanopore 120 can "stop" or "slow down" DNA movement. In the particular example shown in Figure 1, all nucleotides in DNA 110 are bound to modifications. However, in some examples, nucleotides in the head and / or tail portions of the DNA may not be modified with modifications. In some cases, not all nucleotides are modified; for example, only one nucleotide every 2, 3, 4, or 5 nucleotides is modified with a modification that includes a stopping construct.
[0076] In some embodiments, the disclosure relates to stopping constructs for stopping and / or slowing DNA movement. In some embodiments, the disclosure relates to methods for tuning the properties of stopping constructs to enhance and / or fine-tune their stopping ability. In some embodiments, the disclosure relates to a system for determining the sequence of polynucleotides using the methods disclosed herein. In some embodiments, a change in voltage bias is not required to advance the nucleotides through a nanopore reading head. In some embodiments, a change in voltage bias is performed to advance the nucleotides through a nanopore reading head.
[0077] operation Some polynucleotide sequencing techniques can utilize nanopores to provide pathways for ionic currents. For example, as polynucleotides traverse through nanopores, they influence the current flowing through the nanopores. In some embodiments, each transiting nucleotide, or a series of nucleotides, passing through a nanopore can yield a characteristic current. In some embodiments, each recording element passing through a nanopore can also yield a characteristic current. These characteristic currents of the moving polynucleotides can be recorded to determine the sequence of the polynucleotides.
[0078] Figure 2 shows an exemplary workflow 200 of the nanopore sequencing method described herein. Workflow 200 begins with step 205, namely the isolation of a sample polynucleotide (such as DNA) from a biological source using a suitable extraction method. Following the isolation of the sample polynucleotide, the workflow proceeds to step 210, where the sample polynucleotide is subjected to a library preparation process that includes using the sample polynucleotide as a template for synthesizing a new daughter chain using a modified polynucleotide having modifications. During this step, the modifications are introduced onto the newly synthesized polynucleotide chain. In some embodiments, the modifications include a cyclic loop structure with a stopping construct. The nucleotide may contain sites where one or more chemical bonds can be cleaved to form an extended polynucleotide.
[0079] After library preparation, the workflow moves to the nanopore sequencing step 215, where modified polynucleotides in the library move through nanopores, and the data during movement is collected for use in determining base identity. Following the nanopore sequencing step, the workflow moves to the data analysis step 220, where base calling is performed based on the collected data. In some embodiments, workflow 200 in Figure 2 is part of the sequencing cycle.
[0080] In some embodiments, the generation of modified polynucleotides is achieved by incorporating modified tagged nucleotides (i.e., modified nucleotides) into daughter strands synthesized by polymerase, for example, as shown in Figure 3. As shown in the exemplary library preparation process using modified nucleotides in Figure 3, polymerase 360 is synthesized from a template polynucleotide 305 of the isolated sample polynucleotide. During library generation, the sample DNA is converted into a library of “daughter strands” 310, for example, via polymerase incorporation of modified nucleotides 318 or by using a ligase. Each modified nucleotide 318 includes nucleotide 311 and modified 317. In some iterations of nanopore sequencing, daughter strands 310 can also include unique barcodes (e.g., reporter elements) on each base to improve single-base resolution and sequencing accuracy. In some iterations, a reader oligo is added to daughter strands 310 to provide direction-specific insertion into the nanopores.
[0081] Figure 4 shows that a voltage of 450 is applied to introduce target DNA 410 into the nanopore 420. In a particular example in Figure 4, a protein pore 420 having a constricted region 424 is inserted into a lipid bilayer 430. In other examples, solid-state pores are fabricated directly into a synthetic membrane, which can be modified for DNA transfer and used for nanopore sequencing. The nanopore 420 separates two chambers, indicated as cis and trans. Both chambers are filled with a suitable electrolyte solution, and a voltage of 450 can be applied across the nanopore. A target polynucleotide 410, for example, a daughter strand from a library generation process, can be added to the cis, trans, or both chambers of the nanopore. The introduction of the polynucleotide into the nanopore 420 can be achieved by applying a capture voltage of 450. In some embodiments, the nanopore may be deposited such that the constricted region of the nanopore is closer to the trans chamber. In other embodiments, the nanopore may be deposited such that the constricted region of the nanopore is closer to the cis chamber.
[0082] Figure 5 illustrates a method for detecting nucleotides and the controlled movement of polynucleotides through nanopores by an appropriate driving voltage 555. The left side of Figure 5 shows that protein nanopores 520 are deposited within the lipid bilayer 530. As daughter polynucleotides 510 move through nanopores 520, one or more stopping constructs may encounter and interact with the nanopores 520, thereby stopping the movement of polynucleotides 510.
[0083] In some embodiments, to identify nucleotide 511 (base "A") located in the constriction of the nanopore 520, a characteristic ion-sealing current dependent on nucleotide 511 or a nucleotide-bound stopping construct can be recorded by the system. In other examples, the characteristic ion-sealing current may depend on 2, 3, 4, or 5 nucleotides located near the constriction of the nanopore. The characteristic ion-sealing current may further depend on stopping constructs bound to 2, 3, 4, or 5 nucleotides located near the constriction. In some embodiments, the characteristic ion-sealing current may depend on modifications bound to nucleotide 511, such as a reporter element which is part of the modification.
[0084] After the detection of nucleotide 511, the next stopping construct encounters the nanopore, and the movement of the polynucleotide stops again. The right side of Figure 5 shows that nucleotide 521, following nucleotide 511, is located at the constriction of the pore. The recorded characteristic ion-sealing current depends on nucleotide 521. The characteristic ion-sealing current may further depend on the modifications bound to nucleotide 521.
[0085] qualification Nanopores with constriction diameters in the subnanometer range are desirable to prevent multiple chains or secondary structures from moving through the pore. Another important dimension is the length of the constriction, which determines the number of bases contributing to the blockage current. Protein pores such as MspA, with small diameters of about 1.2 nm, have been used to identify homopolymerized DNA. At its narrowest point, the MspA constriction is 0.6 nm long, yet the blockage current is determined by at least four nucleic acid bases present in the constriction. This means that a minimum of 4^4 (256) signals can arise based on the tetramer arrangement in the constriction.
[0086] In some embodiments, the modification may include a stopping construct. In some embodiments, the modification may include a component for extending or spreading adjacent nucleic acid bases in a polynucleotide, such as a cyclic loop group. Certain sites on the modification, or otherwise on the modified nucleotide, may cause one or more chemical bonds to be cleaved to produce an extended chain, which may work to isolate signal generation as the extended chain moves through a nanopore reading head. In some embodiments, the cyclic loop may further include a stopping construct. In some embodiments, the cyclic loop may further include a reporter element, or a reporter element containing two or more subreporter elements. In some embodiments, the cyclic loop may also include a spacer that can be used to adjust the spacing between adjacent nucleotides, between a stopping construct and a reporter element, or between a stopping construct of one nucleotide and another nucleotide in a target polynucleotide.
[0087] The residence time of a termination construct in a constricted region is expected to correlate with its physical and molecular size. Therefore, attaching termination constructs of various sizes and molecular weights to cyclic loop modifications of nucleotides or polynucleotide chains can influence the migration rate of polynucleotides through nanopores. Various parts that can be used as termination constructs are provided herein. Termination constructs are tunable to influence the migration rate of polynucleotides through nanopores (e.g., to slow or stop migration). This can be done by changing the size and / or geometry of the modification portion and / or by changing the nature of the interaction between the modification portion and the nanopore.
[0088] Stop structures of various sizes and geometric shapes The size of the termination construct can be adjusted by increasing i) the physical length of the modified portion and ii) the molecular weight of the modified portion. In some embodiments, the termination construct may include linear polymers. Many of these polymers are commercially available and have a wide variety of structures and chemical properties for selection. The size and weight of linear polymers can be adjusted by changing the number of repeating units in the polymer chain (and thus changing the molecular weight). In some embodiments, linear polymers may include, but are not limited to, linear hydrophilic synthetic polymers, linear hydrophobic synthetic polymers, polynucleotides, peptides, and polypeptides.
[0089] In some embodiments, hydrophilic synthetic polymers include, but are not limited to, polyethylene glycol, polyvinyl alcohol, polyacrylamide, polyvinylpyrrolidone, polystyrene sulfonate, polyethyleneimine, and combinations thereof. Therefore, a stop construct containing a hydrophilic synthetic polymer may be represented by the following structure.
[0090] [ka] In the formula, n refers to the number of repeating units in the polymer chain.
[0091] [ka] This refers to the covalent bond between the polymer and the nucleotides on the daughter chain. * is an end group on the polymer chain. The end group on the polymer chain is not particularly limited and may be any suitable polymer end group. For example, * This may include, but is not limited to, hydrogen, alkyl (e.g., methyl), halogen (e.g., bromide), thiol, amine, dithiobenzoate, dodecyl trithiocarbonate, phenylcarbamodithioate, or dimethylacetic acid. In some embodiments, the polymer may be a continuous chain within a cyclic loop, as opposed to branching from a cyclic loop. Therefore, in some embodiments, * This may be a covalent bond to another part of the annular loop. In some embodiments, n is about 10 to about 200, about 10 to about 100, or about 100 to about 200. In some embodiments,
[0092] [ka] The coupling moiety may be selected from the group consisting of amine-NHS esters, amine-imide esters, amine-pentafluorophenyl esters, amine-hydroxymethylphosphine, carboxyl-carbodiimide, thiol-maleimide, thiol-haloacetyl, thiol-pyridyl disulfide, thiol-thiosulfonate, thiol-vinylsulfone, aldehyde-hydrazide, aldehyde-alkoxyamine, hydroxy-isocyanate, azido-alkyne, azido-phosphine, transcyclooctene-tetrazine, norbornene-tetrazine, azido-cyclooctin, and azido-norbornene. In some embodiments,
[0093] [ka] The coupling portion may further include a linker between the coupling portion and the polymer. Several examples of stop structures containing hydrophilic synthetic polymers are shown in Figure 8. In some embodiments, the stop structure may include a copolymer of two or more hydrophilic synthetic polymers.
[0094] In some embodiments, the hydrophobic synthetic polymer may include polylactic acid, polymethyl methacrylate, polystyrene, and combinations thereof. Therefore, a stopping construct containing a hydrophobic synthetic polymer may be represented by the following structure.
[0095] [ka] In the formula, n refers to the number of repeating units in the polymer chain.
[0096] [ka] This refers to the covalent bond between the polymer and the nucleotides on the daughter chain. * is an end group on the polymer chain. The end group on the polymer chain is not particularly limited and may be any suitable polymer end group. For example, * n may include, but is not limited to, hydrogen, alkyl (e.g., methyl), halogen (e.g., bromide), thiol, amine, dithiobenzoate, dodecyl trithiocarbonate, phenylcarbamodithioate, and dimethylacetic acid. In some embodiments, n is about 10 to about 200, about 10 to about 100, or about 100 to about 200. In some embodiments,
[0097] [ka] The coupling moiety may be selected from the group consisting of amine-NHS esters, amine-imide esters, amine-pentafluorophenyl esters, amine-hydroxymethylphosphine, carboxyl-carbodiimide, thiol-maleimide, thiol-haloacetyl, thiol-pyridyl disulfide, thiol-thiosulfonate, thiol-vinylsulfone, aldehyde-hydrazide, aldehyde-alkoxyamine, hydroxy-isocyanate, azido-alkyne, azido-phosphine, transcyclooctene-tetrazine, norbornene-tetrazine, azido-cyclooctin, and azido-norbornene. In some embodiments,
[0098] [ka] The coupling structure may further include a linker between the coupling portion and the polymer. In some embodiments, the stopping structure may include a copolymer of two or more hydrophobic synthetic polymers, such as polylactic acid-co-polymethyl methacrylate, polystyrene-co-polymethyl acrylate, and polystyrene-co-polylactic acid.
[0099] In some embodiments, the termination construct may include DNA / RNA polynucleotides and other phosphate-containing polymers. In some embodiments, the polynucleotides may include homopolymers, mixed-sequence polymers, tetrameric repeat sequences, and polybasic-debased nucleotides. In some embodiments, the phosphate-containing polymers may include synthetic organic monomers linked by phosphate bonds. Complex mixed-base polynucleotides of various lengths can be synthesized using automated synthesizers. Therefore, termination constructs containing polynucleotides or phosphate-containing polymers may be represented by the following structures.
[0100] [ka] In the formula, n refers to the number of repeating units, B refers to repeating units of natural or unnatural nucleotides and organic molecules, B1, B2, B3, B4 refer to mixed sequences of natural or unnatural nucleotides and organic molecule repeating units, X refers to a heteroatom on a phosphate bond, where X = O or S.
[0101] [ka] This refers to the covalent bond between a peptide / polypeptide and a nucleotide on its daughter chain. * is an end group on the polymer chain. The end group on the polymer chain is not particularly limited and may be any suitable polymer end group. For example, * n may include, but is not limited to, hydroxyls, amines, phosphates, and phosphorothioates. In some embodiments, n is about 1 to about 50, about 10 to about 40, or about 10 to about 30. In some embodiments,
[0102] [ka] The coupling moiety may be selected from the group consisting of amine-NHS esters, amine-imide esters, amine-pentafluorophenyl esters, amine-hydroxymethylphosphine, carboxyl-carbodiimide, thiol-maleimide, thiol-haloacetyl, thiol-pyridyl disulfide, thiol-thiosulfonate, thiol-vinylsulfone, aldehyde-hydrazide, aldehyde-alkoxyamine, hydroxy-isocyanate, azido-alkyne, azido-phosphine, transcyclooctene-tetrazine, norbornene-tetrazine, azido-cyclooctin, and azido-norbornene. In some embodiments,
[0103] [ka] The coupling portion may further contain a linker between the coupling portion and the polymer. Several examples of termination constructs containing polynucleotides are shown in Figure 9.
[0104] In some embodiments, the peptide / polypeptide may comprise an amino acid, an amino acid sequence, or a polypeptide. The size of the peptide / polypeptide can be adjusted by: i) varying the number of amino acids in the peptide sequence, ii) judicious selection of individual amino acids having various side chains, or iii) chemical modification of amino acid residues (e.g., phosphorylation, sulfonation, conjugation of small molecules to the side chain). In some embodiments, non-limiting examples of small molecules can include polyvinyl alcohol, polyvinyl pyrrolidone, polystyrene, polystyrene sulfonate, polymethyl methacrylate, polylactic acid, D-glucose, polyethyleneimine, polyacrylamide, glycoluril, dialkoxybenzene, polyethylene glycol, polycarbonate, polymethyl methacrylate or polyamide. The polypeptide stop construct can comprise one type of amino acid or a mixture of different types / various sizes of amino acids. Thus, the stop construct containing a peptide / polypeptide can be represented by the following structure.
[0105]
Chemical formula
[0106]
Chemical formula
[0107] [ka] The coupling moiety may be selected from the group consisting of amine-NHS esters, amine-imide esters, amine-pentafluorophenyl esters, amine-hydroxymethylphosphine, carboxyl-carbodiimide, thiol-maleimide, thiol-haloacetyl, thiol-pyridyl disulfide, thiol-thiosulfonate, thiol-vinylsulfone, aldehyde-hydrazide, aldehyde-alkoxyamine, hydroxy-isocyanate, azido-alkyne, azido-phosphine, transcyclooctene-tetrazine, norbornene-tetrazine, azido-cyclooctin, and azido-norbornene. In some embodiments,
[0108] [ka] The mixture may further include a linker between the coupling moiety and the peptide / polypeptide. In some embodiments, more than four types of amino acids may be present in the mixed polypeptide subunit. In other embodiments, two or three types of amino acids may be present in the mixed polypeptide subunit.
[0109] In some embodiments, two or more different types of polymers can be linked together to form a termination construct. Thus, the termination construct may comprise copolymers of various polymer types. In some embodiments, the termination construct may comprise hydrophilic synthetic polymers, hydrophobic synthetic polymers, polynucleotides, and / or peptides / polypeptides linked in series. A linker may be present between any two linked polymers. The linker may also be covalently bonded.
[0110] [ka] It may be part of it.
[0111] The size of a linear stop construct can be easily adjusted by connecting multiple constructs in series (including different types of modifications), but the flexibility of such constructs allows for the adoption of multiple conformations in solution. Some conformations of a linear construct promote or enhance its ability to move through pores, thereby reducing their residence time in constricted areas. In some scenarios, a large linear stop construct can completely bypass a pore without slowing or stopping DNA movement. Therefore, in addition to size, it may be useful to adjust the geometry of the stop construct to improve its stopping ability. The geometry of a stop construct can be adjusted by i) increasing the cross-sectional diameter of the stop construct portion, or ii) increasing the structural rigidity of the modified portion.
[0112] In some embodiments, the cross-sectional diameter of the termination construct portion can be modified by introducing branching elements in the molecular design or by using annular portions. In some embodiments, the cross-sectional diameter of the termination construct can be in the range of approximately 1.5 nm, 2.0 nm, 2.5 nm, 3.0 nm to approximately 3.5 nm. In some embodiments, a linear termination construct can be modified to introduce one or more branching structures. In some embodiments, the cross-sectional diameter of the termination construct portion can be adjusted by varying the number of branching / polymer arms from the branching points in the structure of the termination construct. In some embodiments, the termination construct can include multiple branching points to mimic a dendritic structure. In some embodiments, the termination construct can include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more branching points to mimic a dendritic structure.
[0113] In some embodiments, the anchor structure includes a branched polymer. The branched polymer may have two, three, or four branches or polymer arms. In some embodiments, the branched polymer may have one or more branching points. For example, each branch on a first branching point (i.e., a first-generation branching point) may further include another branching point (i.e., a second-generation branching point) which may include two, three, four, or five branches or polymer arms. Such a branched polymer resembles a second-generation dendron. In some embodiments, each branch from a second-generation branching point may further include an additional branching point (i.e., a third-generation branching point) which may include two, three, four, or five branches from a third-generation dendron. In some embodiments, the dendritic anchor structure may have more than one generation of branching points. In some embodiments, the dendritic anchor structure may have two, three, four, or five generations.
[0114] Examples of branching and dendritic termination structures include the following:
[0115] [ka] In the formula, -[X] n -and-[Y]n Each of these represents a polymer selected from synthetic organic polymers, polynucleotides, peptides, or polypeptide sequences (as described above), where n refers to the number of repeating units in the polymer.
[0116] [ka] This refers to the covalent bond between the polymer and the nucleotides on the daughter chain. * is an end group on the polymer chain. The end group on the polymer chain is not particularly limited and may be any suitable polymer end group. In some embodiments, * This may include, but is not limited to, hydrogen, alkyl (e.g., methyl), halogen (e.g., bromide), thiol, amine, dithiobenzoate, dodecyl trithiocarbonate, phenylcarbamodithioate, dimethylacetic acid, hydroxyl, amine, phosphate and phosphorothioate, amide, acetyl, carboxylic acid, or amine. In some embodiments, n may be in the range of about 1 to 10, or about 1 to about 20, or about 1 to about 30, or about 1 to about 40. In some embodiments,
[0117] [ka] The coupling moiety may be selected from the group consisting of amine-NHS esters, amine-imide esters, amine-pentafluorophenyl esters, amine-hydroxymethylphosphine, carboxyl-carbodiimide, thiol-maleimide, thiol-haloacetyl, thiol-pyridyl disulfide, thiol-thiosulfonate, thiol-vinylsulfone, aldehyde-hydrazide, aldehyde-alkoxyamine, hydroxy-isocyanate, azido-alkyne, azido-phosphine, transcyclooctene-tetrazine, norbornene-tetrazine, azido-cyclooctin, and azido-norbornene. In some embodiments,
[0118] [ka] The coupling portion may further include a linker between the coupling portion and the polymer (in X). In some embodiments, [X] and [Y] may be the same. In other embodiments, [X] and [Y] may be different. Several examples of stop structures containing branched polymers are shown in Figures 10, 11, and 12.
[0119] In some embodiments, the termination construct includes a cyclic portion. In some embodiments, the cyclic portion may include repeating units of organic small molecules, nucleotides, or amino acids. In some embodiments, non-limiting examples of small molecules include polyvinyl alcohol, polyvinylpyrrolidone, polystyrene, polystyrene sulfonate, polymethyl methacrylate, polylactic acid, D-glucose, polyethylimine, polyacrylamide, glycouryl, dialkoxybenzene, polyethylene glycol, polycarbonate, polymethyl methacrylate, or polyamide. In some embodiments, the cyclic portion may consist of as few as three repeating units or as many as ten repeating units. In some embodiments, the cyclic portion may have three, four, five, six, seven, eight, nine, ten, eleven, twelve, thirteen, fourteen, or fifteen repeating units. Each repeating unit may be a small molecule, a nucleotide, or an amino acid. In some embodiments, the repeating units may be the same. In some embodiments, at least two of the repeating units are different.
[0120] Examples of such stationary structures include the following:
[0121] [ka] In the formula, X1, X2, X3, X4, X5, X6, X7, and X8 refer to small molecules (e.g., polyvinyl alcohol, polyvinylpyrrolidone, polystyrene, polystyrene sulfonate, polymethyl methacrylate, polylactic acid, D-glucose, polyethylimine, polyacrylamide, glycouryl, dialkoxybenzene, polyethylene glycol, polycarbonate, polymethyl methacrylate, or polyamide), nucleotides, or repeating amino acid units.
[0122] [ka] This refers to the covalent bond between the circular portion and the nucleotides on the daughter chain. In some embodiments, X1, X2, X3, X4, X5, X6, X7, and X8 are the same. In some iterations, X1, X2, X3, X4, X5, X6, X7, and X8 are distinct groups. In some embodiments,
[0123] [ka] The coupling moiety may be selected from the group consisting of amine-NHS esters, amine-imide esters, amine-pentafluorophenyl esters, amine-hydroxymethylphosphine, carboxyl-carbodiimide, thiol-maleimide, thiol-haloacetyl, thiol-pyridyl disulfide, thiol-thiosulfonate, thiol-vinylsulfone, aldehyde-hydrazide, aldehyde-alkoxyamine, hydroxy-isocyanate, azido-alkyne, azido-phosphine, transcyclooctene-tetrazine, norbornene-tetrazine, azido-cyclooctin, and azido-norbornene. In some embodiments,
[0124] [ka] It may further include a linker between the coupling portion and the annular portion. X1(In this case). The annular portion may have 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 repeating units, and thus X1, X2, X3, X4, X5, X6, X7, X8, X9, X 10 , X 11 , X 12 , X 13 , X 14 , X 15 It may include. Several examples of stop structures containing branched polymers are shown in Figure 13.
[0125] In some embodiments, the use of rigid termination construct molecules can also improve deceleration or termination ability by limiting the steric changes of the termination construct in nanoporous systems. While cyclic termination constructs are expected to be less flexible than their linear counterparts, rigidity can be introduced by selecting termination constructs from a wide range of conjugated chromophores. Non-limiting examples include aromatic small molecules such as perylene, pyrene, and relendimide; condensed heterocyclic fluorophores such as Alexa Fluor® 568 (ThermoFisher Scientific catalog no. A33081), Atto 565 (Atto-Tec catalog no. AD565), and Atto 647N (Atto-Tec catalog no. AD647N); porphyrins; and phthalocyanines. In some embodiments, rigid termination constructs can also be selected from rigid macrocyclic structures. Non-limiting examples include α-, β-, and γ-cyclodextrins; cucurbit[n]uryl; and pillar[n]arene.
[0126] Non-covalent interactions between the anchoring structure and nanopores While adjusting the size and geometry of stop constructs is an intuitive and straightforward approach to improving their stopping ability, indiscriminately increasing the size of stop constructs is impractical, especially when the stop construct portion needs to bind to all bases in the daughter chain. Since oligos modified with stop constructs are synthesized by polymerase during the library preparation process (at least in some embodiments), manipulating a polymerase that can incorporate bulky stop constructs while maintaining its processing capacity is expected to be a challenge. Furthermore, while stop constructs can help slow or stop migration, the passage of bulky stop constructs through pores when voltage is applied can "overlap" with adjacent bases, potentially leading to a loss of information.
[0127] Another strategy to enhance the ability of arrest constructs to stop or slow DNA movement is to enhance the interaction between the arrest construct and the surface of the nanopore. Non-limiting examples of interaction types based on the presence of charge, polarity, and aromatic residues include charge-charge, charge-polar, polar-polar, aromatic-aromatic, and nonpolar-nonpolar. Other substitutions and combinations are also being considered.
[0128] For example, non-covalent interactions may include electrostatic interactions, hydrogen bonding, and hydrophobic interactions. Electrostatic interactions may exist when the groups on the termination construct and the amino acid residues in the inner wall of the MspA nanopores are oppositely charged. In some embodiments, hydrogen bonding may occur via a dipole-dipole attraction between a hydrogen atom bonded to an electronegative atom and a charged group / long pair. In some embodiments, hydrophobic interactions may exist when an alkyl / aromatic group on the termination construct associates with a hydrophobic residue in the pore in a polar solvent.
[0129] Amino acid residues on the inner walls of protein pores and surface functional groups on solid-state pores provide good opportunities for non-covalent interactions between the termination construct and the pore surface. In some embodiments, non-covalent interactions may include, but are not limited to, electrostatic interactions, ion-dipole interactions, dipole-dipole interactions, hydrophobic interactions, or combinations thereof between the termination construct and the nanopore surface.
[0130] In some embodiments, electrostatic interactions may exist when the groups on the anchoring construct and the amino acid residues in the inner wall of the nanopore are oppositely charged. For example, in the case of protein pores (e.g., MspA), the pores can be manipulated to contain specific residues that have electrostatic interactions with the anchoring construct portion, while the anchoring construct molecule can be designed to have complementary ionic groups. In some embodiments, the ionic groups on the anchoring construct may include, but are not limited to, negatively charged carboxylate, sulfonate, phosphate functional groups, or combinations thereof. In some embodiments, the ionic groups may also include positively charged primary amines, secondary amines, tertiary amines, quaternary ammonium, guanidinium, and combinations thereof.
[0131] In some embodiments, the stopping structure can interact via electrostatic attraction between opposite polarities on the stopping structure and the nanopore surface. While not limited to any particular theory, such interaction is expected to increase the residence time of the stopping structure in the constricted area. In some embodiments, the stopping structure can also interact with the pores via electrostatic repulsion between similar polarities. While not limited to any particular theory, such interaction may induce conformational changes in the stopping structure that hinder passage through the pores and increase residence time.
[0132] In some embodiments, the termination structure can also interact with the nanopore surface via ion-dipole interactions. In some embodiments, the termination structure portion contains neutral polar groups, while the pore surface is functionalized with ionic groups. In other embodiments, the termination structure portion contains ionic groups, while the pore surface is functionalized with neutral polar groups. In some embodiments, the neutral polar functional groups on the termination structure may include, but are not limited to, alcohols, thiols, and amides. In some embodiments, the polar groups may include, but are not limited to, fluorinated moieties.
[0133] In some embodiments, the termination construct can also interact with the nanopore via dipole-dipole interactions. In some embodiments, the polar residues of the termination construct and the nanopore interact via strong H bonds. In some embodiments, the polar residues of the termination construct and the nanopore interact via weak van der Waals forces.
[0134] Examples of anchoring structures that can interact with nanopores via non-covalent bonds include the following:
[0135] [ka] In the formula, M1, M2, M3, M4, M5, M6, M, and N each refer to a small molecule, nucleotide, or amino acid, and n refers to the number of repeating units.
[0136] [ka] This refers to the covalent bond between the repeating unit and the nucleotides on the daughter chain. * is a terminal group. The terminal group on the repeating unit is not particularly limited and may be any suitable terminal group. In some embodiments, *n may include, but is not limited to, hydrogen, alkyl (e.g., methyl), halogen (e.g., bromide), thiol, amine, dithiobenzoate, dodecyl trithiocarbonate, phenylcarbamodithioate, dimethylacetic acid, hydroxyl, amine, phosphate, phosphorothioate, amide, acetyl, carboxylic acid, or amine. In some embodiments, n may be in the range of about 1 to 10, or about 1 to about 20, or about 1 to about 30, or about 1 to about 40. In some embodiments,
[0137] [ka] The coupling moiety may be selected from the group consisting of amine-NHS esters, amine-imide esters, amine-pentafluorophenyl esters, amine-hydroxymethylphosphine, carboxyl-carbodiimide, thiol-maleimide, thiol-haloacetyl, thiol-pyridyl disulfide, thiol-thiosulfonate, thiol-vinylsulfone, aldehyde-hydrazide, aldehyde-alkoxyamine, hydroxy-isocyanate, azido-alkyne, azido-phosphine, transcyclooctene-tetrazine, norbornene-tetrazine, azido-cyclooctin, and azido-norbornene. In some embodiments,
[0138] [ka] The coupling portion may further include a linker between M / M1. In some embodiments, M1, M2, M3, M4, M5, and M6 are the same. In other embodiments, M1, M2, M3, M4, M5, and M6 are different. In some embodiments, [M] and [N] may be the same. In other embodiments, [M] and [N] may be different. In some embodiments, M1, M2, M3, M4, M5, M6, M, and N may be negatively charged. In some embodiments, M1, M2, M3, M4, M5, M6, M, and N may be positively charged. In some embodiments, M1, M2, M3, M4, M5, M6, M, and N may be neutral polar groups. In some embodiments, M1, M2, M3, M4, M5, M6, M, and N may be hydrophobic groups. In some embodiments, M1, M2, M3, M4, M5, M6, M, and N may be a combination of one or more negatively charged groups, positively charged groups, polar groups, and hydrophobic groups. In some embodiments, non-limiting examples of small molecules M1, M2, M3, M4, M5, M6, M, and N may include polyvinyl alcohol, polyvinylpyrrolidone, polystyrene, polystyrene sulfonate, polymethyl methacrylate, polylactic acid, D-glucose, polyethylimine, polyacrylamide, glycouryl, dialkoxybenzene, polyethylene glycol, polycarbonate, polymethyl methacrylate, or polyamide.
[0139] Further examples of termination constructs that can interact with nanopores through non-covalent bonds include fluoroalkyl moieties.
[0140] In some embodiments, the ability of a stop construct to function as a single entity can vary depending on a number of factors, including, but not limited to, its cross-sectional diameter, charge, and combinations thereof. For example, referring to Figure 14, 9xPEG12 provides a large cross-sectional diameter of about 2.4 nm. Doubler T5 has a smaller cross-sectional diameter of about 2.0 nm compared to 9xPEG12, but provides more negative charge. The 8E stop construct provides both a large cross-sectional diameter of about 2.3 nm and more negative charge.
[0141] Examples of stopped structures In some embodiments, PEG-based stop constructs are provided. In some embodiments, the PEG-based stop constructs are linear. Non-limiting examples include PEG4, PEG8, PEG12, and PEG24, as shown in Figure 8. In some embodiments, the PEG-based stop constructs are branched. Non-limiting examples include 3xPEG4, 3xPEG12, 9xPEG12, and 9xPEG24, as shown in Figures 10 and 11. For example, a 9xPEG12 stop construct has a size of approximately 9.6 nm × 2.4 nm (2D) (determined, for example, by energy minimization modeling using ChemDraw 3D).
[0142] In some embodiments, oligo-based termination structures are provided. In some embodiments, the oligo-based termination structures are linear. Non-limiting examples include the trimer TdT and pentamer TdTdT, and the necaper, as shown in Figure 9. In some embodiments, the oligo-based termination structures are hairpin-shaped. In some embodiments, the oligo-based termination structures are branched. In some embodiments, the oligo-based termination structures are circular. Non-limiting examples include the trimer T, pentamer A, pentamer T, and heptamer T, as shown in Figure 13.
[0143] In some embodiments, branched peptide-based termination constructs are provided. In some embodiments, the branched peptide-based termination constructs include a positive residue (e.g., K), a negative residue (e.g., E), a polar residue (e.g., N), and a hydrophobic residue (e.g., W). Non-limiting examples of branched peptide-based termination constructs include 4K, 4E, 4N, 4W, and 8E, shown in Figure 14. In some embodiments, the branched peptide-based termination constructs can include 2-, 4-, 8-, and 2n-branching. In some embodiments, branched peptide-based termination constructs having a negative residue (e.g., E) are preferred. In some embodiments, the size (two-dimensional) of the branched peptide-based termination construct is in the range of about 2.1 nm to about 2.6 nm (e.g., 4K, 4E, 4N, and 4W). In some embodiments, the size (two-dimensional) of the branched peptide-based termination construct is in the range of about 3.1 nm to about 2.3 nm (e.g., 8E).
[0144] In some embodiments, other peptide-based termination constructs are provided. In some embodiments, the other peptide-based termination constructs include cyclic termination constructs. Non-limiting examples include cyclic 6N, cyclic 6E, cyclic 6Gla, cyclic 5E, cyclic 5Gla, cyclic 4E, and cyclic 4Gla. In some embodiments, the other peptide-based termination constructs include branched termination constructs. Non-limiting examples include 2Gla and 4Gla. In some embodiments, Gla may be preferred over E because Gla provides twice the amount of charge compared to E. In some embodiments, branched and cyclic termination constructs are preferred.
[0145] In some embodiments, fluoroalkyl-based termination constructs are provided. Non-limiting examples include F5, F7, F9, and F13. In some embodiments, fluorine in the fluoroalkyl-based termination construct interacts with residues in the pore.
[0146] Screening of stopped structures Figure 6 shows an exemplary model system for screening suitable stop constructs. Figure 7 shows current-time trace results from experiments using the model system shown in Figure 6.
[0147] Figure 6 shows that the model system used a branched polyethylene glycol molecule (catalog number BP-23454, BroadPharm, San Diego) as the arresting construct to test its arresting ability in MspA nanopores. The arresting construct was conjugated to the center of a synthetic oligo with a specified sequence (poly dT, followed by poly dC) and attached to the cis side of the protein pore. In this model system shown in Figure 6, the arresting construct is greatly magnified and therefore not drawn to scale, but the nanopore is deposited so that the constricted region of the nanopore is closer to the cis side. As shown in Figure 7, a capture voltage of 50 mV was applied to attract the DNA into the nanopore, and a reading voltage of approximately 50 mV was applied constantly throughout the nanopore. Figure 7 shows that after the DNA modified with the arresting construct was captured, a constant current of approximately 6.4 pA, characteristic of the DNA base (poly dT) beneath the arresting construct, was observed. This indicates that the arresting construct stably arrested the DNA translocation.
[0148] Next, a voltage of approximately 150 mV was applied for 200 msec to allow the stop construct to pass through the nanopore, during which a current signal characteristic of a single base transfer event was observed. After the stop construct passed through the nanopore, movement resumed until the large protein (neutraavidin) bound to the DNA molecule stopped moving again. Figure 7 shows that a constant current of approximately 12.6 pA, characteristic of a base (poly dC), was observed after the stop construct passed through. The DNA was expelled from the nanopore by reversing the voltage polarity to -50 mV.
[0149] Sequence determination method In some embodiments, methods are provided for determining the sequence of a target polynucleotide in a nanopore-based sequencing system.
[0150] In some embodiments, the method includes providing a target polynucleotide which is a daughter chain of a sample polynucleotide containing a modified nucleotide. Each nucleotide associates with or binds to a modification, which includes a stopping construct and optionally a reporter element or a plurality of sub-reporter elements. In some embodiments, the daughter chain of the sample polynucleotide includes a nucleotide analog, which includes a chemical structure corresponding to any nucleotide of the sample polynucleotide, and the daughter chain is a polymer. In some embodiments, the method includes applying a capture voltage to insert the daughter chain into a nanopore. Then, a drive voltage is applied to move the target polynucleotide through the nanopore. In some embodiments, the drive voltage is kept constant (e.g., constant) to move the target polynucleotide. In some embodiments, the drive voltage may be complemented by voltage pulses to move the target polynucleotide. When a first modification containing a first stopping construct interacts with the nanopore, thereby positioning the first nucleotide or the corresponding reporter element in a constricted portion of the nanopore, the movement stops or slows down. When a second modification, including a second stopping construct, interacts with the nanopore and places a second polynucleotide or corresponding reporter element into a constricted portion of the nanopore, the movement stops or slows down again.
[0151] In some embodiments of this method, the residence time of the nucleotide or reporter element in the nanopore is longer than 0.1 ms. In some embodiments of this method, the residence time of the nucleotide or reporter element in the nanopore is longer than 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 ms, or any value between the aforementioned values. In some embodiments of this method, the residence time of the nucleotide or reporter element in the nanopore is longer than 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 ms, or any value between the aforementioned values.
[0152] In some embodiments of this method, the residence time of the nucleotide or reporter element in the nanopore is less than 0.1 ms. In some embodiments of this method, the residence time of the nucleotide or reporter element in the nanopore is less than 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 ms, or any value between the aforementioned values. In some embodiments of this method, the residence time of the nucleotide or reporter element in the nanopore is less than 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 ms, or any value between the aforementioned values.
[0153] The ionic current passing through the nanopores is recorded as the target polynucleotide moves through them. The recorded changes in current correlate with different nucleotides or corresponding reporter elements located in the narrowed parts of the nanopores. Therefore, the identity of the nucleotide sequence in the target polynucleotide can be determined based on the recorded current traces.
[0154] The reporter element may contain one or more nucleotides in addition to the other parts. In some embodiments, a termination construct may be attached to the reporter element.
[0155] In some embodiments of this method, providing a daughter chain involves synthesizing a daughter chain using a modified nucleotide, the modification being covalently bonded to the nucleotide. In some embodiments of this method, the electrical response from the pore depends on the identity of the nucleotide. In some embodiments of this method, the electrical response from the pore depends on the identity of the modification attached to each nucleotide. In some embodiments of this method, the electrical response from the pore depends on the identity of the reporter element. In some embodiments of this method, the electrical response from the pore may be affected by spacers in the modification, such as the distance between adjacent nucleotides, the distance between adjacent reporter elements or termination constructs, or the distance between a termination construct and a reporter element.
[0156] In some embodiments of this method, the constricted portion has an opening with an inner diameter of about 0.6 nm to about 1.2 nm. In some embodiments of this method, the constricted portion has an opening with an inner diameter of about 0.3 nm to about 2.4 nm. In some embodiments of this method, the constricted portion has an opening with an inner diameter of about 0.3 nm to about 0.6 nm. In some embodiments of this method, the constricted portion has an opening with an inner diameter of about 0.6 nm to about 1.2 nm. In some embodiments of this method, the constricted portion has an opening with an inner diameter of about 1.2 nm to about 1.8 nm. In some embodiments of this method, the constricted portion has an opening with an inner diameter of about 1.8 nm to about 2.4 nm.
[0157] In some embodiments of this method, the termination construct comprises a polymer selected from the group consisting of linear synthetic hydrophilic polymers, linear synthetic hydrophobic polymers, linear polynucleotides, linear polypeptides, branched polymers, dendritic polymers, and cyclic polymers. In some embodiments, the termination construct can be configured based on the inner diameter of the constriction opening. For example, the termination construct may have a diameter of at least about 0.5 times, about 1 time, about 1.25 times, about 1.5 times, or about 2 times the inner diameter of the constriction opening. For example, Figure 11 shows that the 9xPEG12 termination construct has a diameter of 2.4 nm. Therefore, if the inner diameter of the constriction opening is about 2.4 nm, the termination construct is 1 times the inner diameter. In some embodiments, the inner diameter of the constriction may be about 1.2 nm, similar to the inner diameter of MspA.
[0158] In some embodiments of this method, the modified and / or termination construct includes a covalent bond between the polymer and the nucleotide. In some embodiments of this method, the covalent bond is selected from the group consisting of amine-NHS esters, amine-imide esters, amine-pentafluorophenyl esters, amine-hydroxymethylphosphine, carboxyl-carbodiimide, thiol-maleimide, thiol-haloacetyl, thiol-pyridyl disulfide, thiol-thiosulfonate, thiol-vinylsulfone, aldehyde-hydrazide, aldehyde-alkoxyamine, hydroxyisocyanate, azido-alkyne, azido-phosphine, transcyclooctene-tetrazine, norbornene-tetrazine, azido-cyclooctin, and azido-norbornene.
[0159] In some embodiments, a method for sequencing a target polynucleotide in a nanopore-based sequencing system is described herein, comprising: providing a target polynucleotide comprising a nucleotide, wherein each nucleotide is bound to a modification configured to slow the target polynucleotide relative to a nanopore; and applying a constant voltage across the nanopore to identify a first reporter element in the constricted portion of the nanopore based on a first electrical response in the system.
[0160] In some embodiments, modifications bound to nucleotides also move through nanopores. In some embodiments, providing a target polynucleotide involves synthesizing daughter strands based on a template polynucleotide using nucleotides having modifications, where the modifications are covalently bonded to the nucleotides. In some embodiments, the target polynucleotide comprises at least one additional nucleotide containing additional modifications. In some embodiments, at least two of the modifications are the same. In some embodiments, at least two of the modifications are different. In some embodiments, each nucleotide is bound to a modification specific to the type of nucleotide. In some embodiments, the first electrical response further depends on the nucleotide-specific modifications. In some embodiments, at least one modification comprises a circular loop. In some embodiments, the circular loop further comprises a stopping construct, which increases the residence time of the nucleotide in the nanopore. In some embodiments, the modification further comprises one or more spacers. In some embodiments, the modification further comprises one or more reporter elements. In some embodiments, the reporter elements correspond to and encode the identity of a nucleic acid base. In some embodiments, one or more spacers separate consecutive nucleotides and / or stopping constructs.
[0161] In some embodiments of this method, the voltage is in the range of approximately 50mV to approximately 450mV. In some embodiments of this method, the voltage is in the range of approximately 75mV to approximately 300mV. In some embodiments of this method, the voltage is in the range of approximately 100mV to approximately 200mV. In some embodiments of this method, the voltage is in the range of approximately 150mV to approximately 200mV. In some embodiments of this method, the voltage is in the range of approximately 200mV to approximately 250mV. In some embodiments of this method, the voltage is in the range of approximately 250mV to approximately 400mV.
[0162] In some embodiments of this method, the duration of the voltage bias is approximately 200 msec (milliseconds). In some embodiments of this method, the duration of the voltage is in the range of approximately 100 msec to approximately 400 msec. In some embodiments of this method, the duration of the voltage is in the range of approximately 100 msec to approximately 200 msec. In some embodiments of this method, the duration of the voltage is in the range of approximately 200 msec to approximately 300 msec. In some embodiments of this method, the duration of the voltage is in the range of approximately 300 msec to approximately 400 msec. In some embodiments of this method, the duration of the voltage is in the range of approximately 1 μs (microseconds) to approximately 10 μs, approximately 10 μs to approximately 100 μs, approximately 100 μs to approximately 500 μs, approximately 500 μs to approximately 1000 μs, approximately 1 msec to approximately 100 msec, or any value in between.
[0163] In some embodiments of this method, the residence time of nucleotides in the nanopores is longer than 0.1 ms. In some embodiments of this method, the residence time of nucleotides in the nanopores is longer than 0.5 ms.
[0164] Method for synthesizing target polynucleotides Figure 15 shows that the extracyclic amine at position 6 of adenine and the fifth carbon on the cytosine ring are two commonly present methylation sites (shown as R groups) found in DNA. Figure 16 shows that DNA methyltransferase (MTase) can catalyze the transfer of methyl groups (R groups) from S-adenosylmethionine (SAM) analogs to methylation sites found in DNA, such as the two commonly present methylations shown in Figure 15. Based on the process shown in Figure 16, in some embodiments, DNA can be directly modified to have R groups.
[0165] Figure 17 shows an exemplary structure of a SAM analog with modifications. In some embodiments, the modifications or moieties in the SAM analog can be transferred to nucleic acid bases in ssDNA using MTase, similar to the process shown in Figure 16. The modified ssDNA can then be subjected to nanopore sequencing. The modifications or moieties covalently bound to the ssDNA can act as “stop constructs” because they can stop the movement of ssDNA through the nanopores until an increase in the applied voltage drives further movement. Stop constructs used as ratchets or stop constructs may include PEGs (see Figure 18), peptides, or oligonucleotides having different lengths and branching.
[0166] Most MTases are sequence-dependent, which can limit the potential applicability of modified bases on DNA molecules. However, in preferred embodiments, the adenine DNA methyltransferase M.EcoGII can be used. M.EcoGII is sequence-independent and can methylate up to 99% of adenine in small dsDNA model substrates. In some embodiments, sequence-specific cytosine methyltransferases such as M.SssI, which recognize CG dinucleotides, can be used. In some embodiments, mutations in M.SssI, which have been engineered to be sequence-independent, can be used.
[0167] Figure 17 shows an example of a typical structure of a SAM analog, which includes a termination construct transferred onto an oligonucleotide or polynucleotide. In some embodiments, the termination construct may be in the form of a linear PEG chain (n=1-48) or a branched PEG (n=1-48, x=2-9), as shown in Figure 18. In some embodiments, the termination construct may be part of a circular loop or branched from a circular loop. In some embodiments, the PEG chain may be covalently bonded to the SAM analog by amine-NHS esters, amine-imide esters, amine-pentafluorophenyl esters, amine-hydroxymethylphosphine, carboxyl-carbodiimide, thiol-maleimide, thiol-haloacetyl, thiol-pyridyl disulfide, thiol-thiosulfonate, thiol-vinylsulfone, aldehyde-hydrazide, aldehyde-alkoxyamine, hydroxy-isocyanate, azido-alkyne, azido-phosphine, transcyclooctene-tetrazine, norbornene-tetrazine, azido-cyclooctin, and azido-norbornene. Figure 19 shows an example of a SAM analog having a PEG4 group conjugated via a Cu-click reaction. In some embodiments, other possible termination constructs include fluorophores, oligonucleotides, peptides, or any molecule that can be covalently bonded to the SAM analog using one or more of the coupling chemical reactions listed above.
[0168] Figure 20 illustrates an exemplary process for modifying naturally occurring 5-methylcytosine using CMD1, a Ten-Eleven Translocation (TET) dioxygenase homolog found in the green alga Chlamydomonas reinhardtii, and an analogue of vitamin C. Vitamin C is a natural comatrix of CMD1. In some embodiments, the vitamin C analogue may have modifications (R1 and R2) on carbons 5 and 6. In some embodiments, modifications R1 and R2 may be in the form of a linear PEG chain (n=1-48) or a branched PEG chain (n=1-48, x=2-9). Based on the process shown in Figure 20, 5-methylcytosine in a target polynucleotide can be identified using the disclosed nanopore sequencing method, which involves the modification and controlled polynucleotide transfer.
[0169] Circular loop modification Figure 21 shows a general embodiment of a modified nucleotide that can be used to form daughter strands of a sample polynucleotide. The modified nucleotide has a cyclic loop modification 2104 attached to nucleotide 2102 at two positions, thereby forming a cyclic loop. The cyclic loop modification further includes a stopping construct 2106 configured to slow or regulate the migration rate of the target polynucleotide through a nanopore reading head. The cyclic loop modification further includes reporter elements 2108 encoding each of the nucleic acid bases. Each reporter element 2108 generates a specific signal as it traverses the nanopore, and each nucleic acid base's reporter element 2108 generates substantially different signals. These signals are differentiated within a range of voltages used in the methods disclosed herein.
[0170] Figure 22 shows a simplified diagram of an elongated polynucleotide 2200 containing a stopping construct 2106 and reporter elements (represented by A, T, C, G, etc.), where each reporter element generates an identifiable signal when traversed through a nanopore. In some embodiments, a cyclic loop modification is incorporated into the daughter strand. In some embodiments, each cyclic loop nucleotide contains an original base (e.g., A, T, C, or G) and a unique barcoding / reporter region (i.e., reporter element) specific to the cleavable site. The daughter strand is then "elongated" by cleaving the cleavable site. As a result, when sequencing the daughter strand, the nanopore can "read" the barcoding / reporter region and identify the base it codes for. Reporter elements introduced into the daughter strand via polymerization are designed to completely occupy the nanopore reading head using spacers that increase the distance between adjacent reporter elements, thus reducing the number of signals to just four, i.e., one per nucleic acid base. Thus, the disclosed technique enables barcode-based decoding of individual bases.
[0171] In some embodiments, a stopping construct may be present to regulate the migration rate of the elongated circular loop polynucleotide. In some embodiments, the circular loop contains a non-barcoding spacer that allows the daughter strand to elongate after cleavage of the cleavable site. The non-barcoding spacer construct generates a signal or signal break that is distinguishable from the signal of the nucleic acid base as it passes through the nanopore, thereby allowing isolation and / or enhancement of the signal recorded from the nucleic acid base. Thus, the disclosed technique enables improved resolution of the recorded signal. In some embodiments, the spacer may contain both a barcoding / reporter region and a non-barcoding spacer construct. In some embodiments, the spacer construct may be a barcoding / reporter element.
[0172] Base calling Synthetic cyclic loop constructs were evaluated for the movement of constituent nucleotides across nanopores according to Table 1, and sequence abbreviations are detailed in Table 2. All constructs were incubated with 1:1 traptabidine in KCl buffer prior to sequencing. A polymer membrane (P5 in octane at 5 mg / mL) was formed across a 50 μm opening to separate the cis and trans compartments. Only the trans compartment contained "Lock Sh16LNA" (5 μM) while the membrane traversed symmetric buffer conditions (1 M KCl, 50 mM Hepes pH 7.5). M2NNN MspA was inserted into the membrane, and the ion current through the nanopores was recorded at a sampling rate of 10 kHz.
[0173] Figures 23A–23C show the residence time distribution of a single stationary construct using construct 1 from Table 1. As shown in Figure 23A, at high drive voltages, residence times conformed to exponential decay. In Figure 23B, residence times closely conformed to a gamma distribution when held at moderate voltages. In this example, approximately 2% of events showed a duration of less than 5 ms. In contrast, Figure 23C shows that at low positive voltages, the stationary construct remained immobilized in the pore until an arbitrary experimental cutoff of 500 ms.
[0174] Figures 24A–24C relate to the current traces and statistics of the synthetic construct derived from construct 1 in Table 1. Specifically, Figure 24A shows heterogeneous sequences of four elongated cyclic nucleotides from construct 1. Figure 24B shows the current traces from the synthetic construct containing four termination constructs and four reporter elements corresponding to A, T, C, and G. Figure 24C shows the statistical distribution for 26 migrating molecules, indicating that 99.6% of migrating events were detected for all four reporters.
[0175] Figures 25A–25D show current traces and statistics for composite structures derived from various structures listed in Table 1. Figure 25A shows the current trace of a composite structure derived from Structure 2 in Table 1, containing six stationary structures and six reporter elements. Figure 25B shows the current trace from a composite structure derived from Structure 3 in Table 1, containing eleven reporter elements. As can be seen from Figures 25A and 25B, six and eleven different signals were observed for each reporter element. Figures 25C and 25D show aggregate statistics for the six and eleven reporter elements from Figures 25A and 25B, respectively, showing that more than 95% of moving events were detected for all reporter elements.
[0176] [Table 1]
[0177] [Table 2]
[0178] Additional information It should be understood that all combinations of the aforementioned concepts and further concepts, which are discussed in more detail below, are intended to be part of the subject matter of the inventions disclosed herein (provided that such concepts do not contradict each other). Specifically, all combinations of claimed subject matter appearing at the end of this disclosure are intended to be part of the subject matter of the inventions disclosed herein. It should also be understood that terms used expressly herein and that may appear in any disclosure incorporated by reference should be given meanings that most coincide with the specific concepts disclosed herein.
[0179] Throughout this specification, references to "one example," "another example," and "an example" mean that certain elements (e.g., features, structures, and / or characteristics) described in relation to an example are included in at least one example described herein, and may or may not be present in other examples. In addition, unless explicitly indicated otherwise in the context, it should be understood that elements described in relation to any example may be combined in any preferred manner in various examples.
[0180] It should be understood that the ranges provided herein include the specified ranges and any values or subranges within those specified ranges, as if such values or subranges were explicitly enumerated. For example, the range of approximately 2 nm to approximately 20 nm should be interpreted to include not only the explicitly enumerated limits of approximately 2 nm to approximately 20 nm, but also individual values, e.g., approximately 3.5 nm, approximately 8 nm, approximately 18.2 nm, and subranges, e.g., approximately 5 nm to approximately 10 nm. Furthermore, where "approximately" and / or "substantially" are used to express a value, this means that a small variation (up to ±10%) from the specified value is included.
[0181] While several embodiments have been described in detail, it should be understood that the disclosed examples can be modified. Therefore, the above description should be considered non-limiting.
[0182] While certain embodiments are described, these embodiments are presented for illustrative purposes only and are not intended to limit the scope of this disclosure. In fact, the novel methods and systems described herein may be embodied in various other forms. Furthermore, various omissions, substitutions, and modifications of the systems and methods described herein may be made without departing from the spirit of this disclosure. The appended claims and their equivalents are intended to cover such forms or modifications as to be included in the scope and spirit of this disclosure.
[0183] Features, materials, properties, or bases described in conjunction with a particular aspect or example shall be understood to be applicable to any other aspect or example described in this section or elsewhere in this Spec. All of the steps of any method or process disclosed herein (including any appended claims, abstract and drawings) and / or may be combined in any combination, except in cases where at least some of such features and / or steps are mutually exclusive. Protection is not limited to the details of any of the aforementioned examples. Protection is extended to any novel one or any novel combination of features disclosed herein (including any appended claims, abstract and drawings), or to any novel one or any novel combination of any steps of any method or process disclosed.
[0184] Certain features described in this disclosure in the context of separate implementations may also be implemented in combination in a single implementation. Conversely, various features described in the context of a single implementation may also be implemented separately in multiple implementations or in any preferred partial combination. Furthermore, even if features are described above as functioning in a particular combination, one or more features from a claimed combination may, in some cases, be removed from the combination, and this combination may be claimed as a partial combination or a variation of a partial combination.
[0185] Furthermore, while operations may be depicted in the drawings or described herein in a specific order, it is not necessary for such operations to be performed in a specific order or sequentially, or for all operations to be performed, in order to achieve the desired result. Other operations not depicted or described may be incorporated into exemplary methods and processes. For example, one or more additional operations may be performed before, after, simultaneously with, or in between any of the described operations. Furthermore, operations may be rearranged or rearranged in other implementations. Those skilled in the art will understand that in some embodiments, the actual steps taken in the illustrated and / or disclosed processes may differ from those shown in the drawings. Depending on the embodiment, certain steps described above may be omitted, or others may be added. Furthermore, the features and attributes of the particular embodiments disclosed above may be combined in different ways to form additional embodiments, all of which fall within the scope of this disclosure. Also, the separation of various system components of the implementation described above should not be understood as requiring such separation in all implementations, and it should be understood that the described components and systems can usually be integrated together in a single product or packaged in multiple products. For example, any of the components for an energy storage system described herein may be provided separately or integrated together (e.g., packaged together or mounted together) to form an energy storage system.
[0186] For the purposes of this disclosure, certain aspects, advantages, and novel features are described herein. Not all such advantages can necessarily be achieved according to a particular embodiment. Therefore, for example, a person skilled in the art will recognize that this disclosure can be embodied or practiced in such a manner that one or a group of advantages taught herein is achieved without necessarily achieving other advantages that can be taught or proposed herein.
[0187] Unless otherwise specified, conditional language such as “can,” “could,” “might,” or “may,” unless specifically stated otherwise or understood within the context in which they are used, is generally intended to convey that a particular embodiment includes a particular feature, element, and / or step, while other embodiments do not. Therefore, such conditional language is generally not intended to imply that a feature, element, and / or step is required in some way in one or more embodiments, or that one or more embodiments necessarily include logic, with or without user input or prompting, that determines whether these features, elements, and / or steps should be included or performed in any particular embodiment.
[0188] Conjunctions such as “at least one of X, Y, and Z” are understood in the usual context, unless otherwise specified, to indicate that an item, term, etc., may be X, Y, or Z. Therefore, such conjunctions are generally not intended to imply that a particular embodiment requires the presence of at least one of X, at least one of Y, and at least one of Z.
[0189] The terms "approximately," "about," "generally," and "substantially" as used herein refer to values, quantities, or characteristics that are close to the stated values, quantities, or characteristics that still perform the desired function or achieve the desired result.
[0190] The scope of this disclosure is not intended to be limited by any specific disclosure of preferred embodiments in this section or elsewhere in this specification, but may be defined by claims presented in this section or elsewhere in this specification, or hereafter presented. The language of the claims should be interpreted broadly based on the language used in the claims, and not limited to the embodiments described herein or in the proceedings of this application, and the embodiments should be interpreted non-exclusively.
[0191] While the invention described herein relates to certain preferred embodiments, other embodiments will be apparent to those skilled in the art. Furthermore, other combinations, omissions, substitutions, and modifications will be apparent to those skilled in the art in light of the disclosure herein. Accordingly, the invention is not limited by the description of preferred embodiments but is defined by reference to the appended claims. All references cited herein are incorporated in their entirety by reference.
[0192] The terms used in the descriptions presented herein are not intended to be construed in any restrictive or limiting manner, and unless otherwise indicated, they refer to the ordinary meanings that would be understood by a person skilled in the art in light of this specification. Furthermore, embodiments may include, consist of, or essentially consist of several novel features, none of which alone bear their desired attributes, or are considered essential for carrying out the embodiments described herein. When used herein, section headings are for structural purposes only and should never be construed as limiting the subject matter described herein. All documents and similar materials cited in this application, including but not limited to patents, patent applications, articles, books, papers, and internet web pages, are expressly incorporated by reference in their entirety for any purpose. If the definitions of terms in incorporated references appear to differ from the definitions provided in this instruction, the definitions provided in this instruction shall prevail. Since “approximately” is implied before temperatures, concentrations, times, etc., considered in this instruction, it will be understood that slight and non-substantial deviations are within the scope of this instruction as described herein.
[0193] While this disclosure is in the context of specific embodiments and examples, those skilled in the art will understand that this disclosure extends beyond the specifically disclosed embodiments to the use of other alternative embodiments and / or embodiments, as well as obvious modifications and equivalents thereof. In addition, while some variations of the embodiments are shown and described in detail, other modifications that are within the scope of this disclosure will be readily apparent to those skilled in the art based on this disclosure. Various combinations or partial combinations of specific features and aspects of the embodiments can be made and are still intended to be included within the scope of this disclosure. It should be understood that various features and aspects of the disclosed embodiments can be combined with or substituted for each other to form various modes or embodiments of this disclosure. Accordingly, the scope of this disclosure disclosed herein is not intended to be limited by the specific disclosed embodiments described above.
Claims
1. A method for determining the sequence of a target polynucleotide in a nanopore-based sequencing system, wherein the method is To provide a target polynucleotide comprising a nucleotide, wherein each nucleotide comprises a modification, and the modification comprises a stopping construct configured to stop the movement of the target polynucleotide through a nanopore, Applying a driving voltage to move one or more portions of the target polynucleotide through the nanopores, To continuously measure the current in the nanopores during movement, A method comprising identifying the sequence of the target polynucleotide by correlating the measured current with the identity of one or more nucleotides.
2. The method according to claim 1, wherein providing the target polynucleotide comprises synthesizing a daughter chain based on a template polynucleotide using a modified nucleotide, wherein the modification is covalently bonded to the nucleotide.
3. The method according to claim 1, wherein the modification further includes an annular loop, and the stop structure is coupled to the annular loop.
4. The method according to claim 3, wherein the annular loop further comprises a reporter element encoding the nucleotide.
5. The method according to claim 4, wherein the annular loop further includes a spacer.
6. To provide the target polynucleotide, The process involves synthesizing daughter chains based on a template polynucleotide using a modified nucleotide, wherein the modification is covalently bonded to the nucleotide. The method according to any one of claims 3 to 5, comprising cleaving the daughter chain to produce the target polynucleotide having an extended polynucleotide chain.
7. The method according to any one of claims 1 to 6, wherein the nucleotide has a residence time of more than 5 ms in the nanopore.
8. The method according to any one of claims 1 to 6, wherein the nucleotide has a residence time in the nanopore of more than 0.1 ms.
9. The method according to claim 1, wherein the drive voltage is kept constant during the movement.
10. The method according to any one of claims 1 to 9, wherein the measured current depends on the reporter element or the nucleotide passing through the nanopore.
11. The method according to any one of claims 1 to 10, wherein the termination structure comprises a polymer selected from the group consisting of linear synthetic hydrophilic polymers, linear synthetic hydrophobic polymers, linear polynucleotides, linear polypeptides, branched polymers, dendritic polymers, cyclic polymers, fluoroalkyls, rigid conjugated chromophores, and rigid macrorings.
12. The method according to claim 11, wherein the termination construct comprises a covalent bond between the polymer and the corresponding nucleotide or the cyclic loop.
13. The method according to claim 12, wherein the covalent bond is selected from the group consisting of amine-NHS esters, amine-imide esters, amine-pentafluorophenyl esters, amine-hydroxymethylphosphine, carboxyl-carbodiimide, thiol-maleimide, thiol-haloacetyl, thiol-pyridyl disulfide, thiol-thiosulfonate, thiol-vinylsulfone, aldehyde-hydrazide, aldehyde-alkoxyamine, hydroxyisocyanate, azido-alkyne, azido-phosphine, transcyclooctene-tetrazine, norbornene-tetrazine, azido-cyclooctin, and azido-norbornene.
14. The method according to any one of claims 11 to 13, wherein the linear synthetic hydrophilic polymer is selected from the group consisting of polyethylene glycol, polyvinyl alcohol, polyacrylamide, polyvinylpyrrolidone, polystyrene sulfonate, polyethyleneimine, and combinations thereof.
15. The method according to any one of claims 11 to 13, wherein the linear synthetic hydrophobic polymer is selected from the group consisting of polylactic acid, polymethyl methacrylate, polystyrene, and combinations thereof.
16. The method according to any one of claims 11 to 13, wherein the linear polynucleotide is a homopolymer of natural nucleotides, a homopolymer of unnatural nucleotides, a polymer of a mixed sequence of natural nucleotides, or a polymer of a mixed sequence of unnatural nucleotides.
17. The method according to any one of claims 11 to 13, wherein the linear polypeptide comprises one or more types of amino acids.
18. The method according to any one of claims 11 to 13, wherein the branched polymer comprises two or more branches.
19. The method according to any one of claims 11 to 13, wherein the cyclic polymer comprises three or more repeating units.
20. The method according to claim 19, wherein each repeating unit is a small molecule, a nucleotide, or an amino acid.
21. The method according to claim 19, wherein the repeating units are the same.
22. The method according to claim 19, wherein at least two of the repeating units are different.
23. The method according to any one of claims 1 to 22, wherein the stopping structure interacts with the nanopores via non-covalent interactions.
24. The method according to claim 23, wherein the non-covalent interaction includes electrostatic interactions, ion-dipole interactions, dipole-dipole interactions, hydrophobic interactions, and combinations thereof.
25. A kit for carrying out a method for determining the sequence of a target polynucleotide in a nanopore-based sequencing system, wherein the kit comprises a nucleotide having the modification described in any one of claims 1 to 24.
26. A system for determining the sequence of a target polynucleotide using the method according to any one of claims 1 to 24.
27. A cyclic loop nucleotide comprising a cyclic loop modification that bridges a nucleic acid base and a phosphate group, wherein the cyclic loop modification comprises a reporter that encodes the identity of the nucleic acid base and a termination construct adjacent to the reporter.
28. The cyclic loop nucleotide according to claim 27, wherein the termination structure is adjacent to the reporter.
29. The cyclic loop nucleotide according to claim 27, wherein the termination structure includes a linear, branched, cyclic, or dendritic structure.
30. It has one of the following structures: 【Chemistry 1】 During the ceremony, X is -O-, -CH 2 -, -NSO 2 -, -NH-, 【Chemistry 2】 And, X' is -S-, =N-SO 2 -, =NH-CO-, or 【Transformation 3】 And, The bases are nucleic acid bases, L 1 and L 2 Each of these is a linking group, RP is a reporter that codes for nucleic acid bases, ARC is a terminating construct, the cyclic loop nucleotide according to claim 27.
31. The cyclic loop nucleotide according to claim 30, wherein the ARC is covalently bonded to the cyclic loop via a covalent bond selected from amine-NHS ester, amine-imide ester, amine-pentafluorophenyl ester, amine-hydroxymethylphosphine, carboxyl-carbodiimide, thiol-maleimide, thiol-haloacetyl, thiol-pyridyl disulfide, thiol-thiosulfonate, thiol-vinylsulfone, aldehyde-hydrazide, aldehyde-alkoxyamine, hydroxyisocyanate, azido-alkyne, azido-phosphine, transcyclooctene-tetrazine, norbornene-tetrazine, azido-cyclooctin, and azido-norbornene.