Ribose-mediated cyclic loop opening for nanopore sequencing

By employing cyclic loop nucleotides with barcoding and cleavable sites, the method addresses the challenge of concurrent base sensing in nanopore sequencers, achieving accurate and efficient single-base resolution sequencing.

WO2025183909A1PCT designated stage Publication Date: 2025-09-04ILLUMINA INC

Patent Information

Application Number
PCT/US2025/015735
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-01
Filing Date
2025-02-13
Publication Date
2025-09-04

AI Technical Summary

Technical Problem

Nanopore sequencers face challenges in accurately sequencing DNA or RNA due to the readhead sensing multiple bases concurrently, leading to complex signal permutations and reduced sequencing accuracy, and the translocation speed of single-stranded nucleotides exceeds electronic detection capabilities.

Method used

The use of cyclic loop nucleotides with unique barcoding regions and cleavable sites allows for the synthesis of a daughter strand that is sequenced, reducing the number of signals to four by ensuring each nucleobase occupies the nanopore readhead individually, and incorporating arresting constructs to control translocation speed.

Benefits of technology

This approach enhances sequencing resolution and accuracy by isolating nucleobase signals, enabling high-throughput, cost-effective single-base resolution sequencing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025015735_04092025_PF_FP_ABST
    Figure US2025015735_04092025_PF_FP_ABST
Patent Text Reader

Abstract

In one aspect, the disclosed technology relates to nanopore sequencing with a polynucleotide including a plurality of ribonucleotides, wherein each ribonucleotide comprises a cyclic loop between two positions of the ribonucleotide, wherein the cyclic loop comprises a spacer or reporter moiety corresponding to the identity of the ribonucleotide, one or more linkers, and optionally one or more arresting constructs. The cyclic loop may be selectively cleaved in a ribose-mediated pathway to cause cleavage at the 5' P-O bond of the phosphate backbone.
Need to check novelty before this filing date? Find Prior Art

Description

RIBOSE-MEDIATED CYCLIC LOOP OPENING FOR NANOPORE SEQUENCINGINCORPORATION BY REFERENCE TO ANY PRIORITY APPLICATIONS

[0001] This application claims the priority benefit of U.S. Provisional Application No. 63 / 560,215, filed March 1, 2024, the entire disclosure of which is incorporated herein by reference in its entirety.BACKGROUND

[0002] Some polynucleotide sequencing techniques involve performing a large number of controlled reactions on support surfaces or within predefined reaction chambers. The controlled reactions may then be observed or detected, and subsequent analysis may help identify properties of the polynucleotide involved in the reaction. Examples of such sequencing techniques include next-generation sequencing or massive parallel sequencing involving sequencing-by-ligation, sequencing-by-synthesis, reversible terminator chemistry, or pyrosequencing approaches.

[0003] Some polynucleotide sequencing techniques utilize a nanopore, which can provide a path for an ionic electrical current. For example, as the polynucleotide traverses through the nanopore, it influences the electrical current through the nanopore. Each passing nucleotide, or series of nucleotides, that passes through the nanopore yields a characteristic blockage current. These characteristic electrical currents of the traversing polynucleotide can be recorded to determine the sequence of the polynucleotide.SUMMARY

[0004] The readhead of nanopores (e.g., the constriction region of nanopores) usually “senses” several bases concurrently along the sample polynucleotide, such as DNA or RNA strand, increasing the challenge of accurate nanopore sequencing due to many permutations of signals arising from different sequences. For example, the MspA pore reads about 4 bases at a time, giving rise to at least 4A4= 256 different signals that need to be deconvoluted and resolved.

[0005] In one aspect, the disclosed technology provides a method that instead of directly sequencing the sample DNA or RNA, a daughter strand is synthesized using cyclic loop nucleotides. In some embodiments, each cyclic loop nucleotide contains a unique barcoding / reporter region that is specific to the original bases (e.g., A, T, U, C, or G) and a cleavable site. The daughter strand is then “elongated” by cutting the cleavable sites. Consequently, when sequencing the elongated daughter strand, the nanopore can “read” the barcoding / reporter region to identify the base that it is coding for. The linker and barcoding construct that is introduced in the daughter strand via polymerization is designed to occupy the readhead of the nanopore entirely, hence reducing the number of signals to just four, i.e., one per nucleobase. Thus, the disclosed technology allows barcode-based decoding of individual bases. In some embodiments, the cyclic loops contain non-barcoding regions that are configured to elongate the ribonucleotide after cutting the cleavable sites. The non-barcoding regions may produce a distinguishable signal or a distinguishable signal break from signals of the nucleobases when passing through the nanopore, thereby isolating and / or enhance the recorded signals from the nucleobases. Thus, the disclosed technology allows improved resolution of the recorded signal. In some embodiments, the cyclic loop contains both barcoding / reporter regions and non-barcoding linker regions. In some embodiments, the linker regions may be the barcoding / reporter element. In some embodiments, nucleotides and oligonucleotides are modified using heavy atoms (including, for example, sulfur and selenium).

[0006] In another aspect, the disclosed technology provides systems, devices, kits, and methods which allow cleavable linkages along the RNA backbone, synthesis of cyclic loop nucleotides, barcodes for individual base identification, and polymerase mutation for incorporation of modified ribonucleotides. Systems may be prepared to allow parallel reads in multiple nanopores, such as thousands or millions of nanopores. Accordingly, components of any system may be functionally duplicated to multiply sequencing throughput. Any system herein may also be adapted with microfluidics or automation.

[0007] The systems, devices, kits, and methods disclosed herein each have several aspects, no single one of which is solely responsible for their desirable attributes. Without limiting the scope of the claims, some prominent features will now be discussed briefly. Numerous other examples are also contemplated, including examples that have fewer,additional, and / or different components, steps, features, objects, benefits, and advantages. The components, aspects, and steps may also be arranged and ordered differently. After considering this discussion, and particularly after reading the section entitled “Detailed Description,” one will understand how the features of the devices and methods disclosed herein provide advantages over other known devices and methods.

[0008] Additional details of exemplary nanopore sequencing devices which can be used with the disclosed technology, and methods of operating the devices, can be found in U.S. Patent Publication PCT / US2021 / 038125 and PCT / US2022 / 020395, the entirety of each of the disclosures is incorporated herein by reference.

[0009] The embodiments disclosed herein relate to a method for determining a sequence of a polynucleotide in a nanopore-based sequencing system, the method including: providing a polynucleotide including a plurality of ribonucleotides, wherein each ribonucleotide includes a cyclic loop, the cyclic loop having a first end attached to a first position of the ribonucleotide and a second end attached to the second position of the ribonucleotide; selectively cleaving a 5' P-0 bond of the phosphate backbone on each of the plurality of ribonucleotides between the first and the second positions, thereby opening the cyclic loop to form a linear cyclic loop in the form of an elongated polymer, wherein the cleaved 5' P-0 bond oxygen is not directly connected to the cyclic loop; applying a voltage to cause the elongated polymer to insert into and translocate through a nanopore; and (i) detecting and identifying one or more reporter barcodes in the cyclic loop when the opened cyclic loop passes through the nanopore; or (ii) detecting and identifying a base on the ribonucleotide when the ribonucleotide passes through the nanopore.

[0010] In some aspects, the techniques described herein relate to a method, wherein the cleaving is mediated by one or more nuclease enzyme(s) or self-immolative groups.

[0011] In some aspects, the techniques described herein relate to a method, wherein the nuclease enzyme is a 3'-Phosphate yielding endonuclease.

[0012] In some aspects, the techniques described herein relate to a method, wherein the nuclease is selected from the group consisting of RNase A, RNase Tl, RNase T2, RNase U2, RNase I, and Phosphodiesterase II.

[0013] In some aspects, the techniques described herein relate to a method, wherein the cyclic loop includes a first linking group, a second linking group, and a spacer between the first and second linking groups.

[0014] In some aspects, the techniques described herein relate to a method, wherein the cyclic loop includes one or more linking groups, an arresting construct, and one or more reporter barcodes.

[0015] In some aspects, the techniques described herein relate to a method, wherein the spacer includes a polynucleotide having 10 to 100 repeating units, polypeptide having 10 to 100 repeating units, alkyl chains having 10 to 200 carbons, hydrophilic polymers having 10 to 100 repeating units selected from the group consisting of polyethyleneglycol, polyvinyl alcohol, polyacrylamide, polyvinylpyrrolidone, polystyrenesulfonate, and polyethyleneimine, hydrophobic polymers having 10 to 100 repeating units selected from the group consisting of polylactic acid, polymethymethacrylate, and polystyrene, and combinations thereof.

[0016] In some aspects, the techniques described herein relate to a method, wherein the spacer includes the one or more reporter barcodes, wherein the one or more reporter barcodes correspond to and identify a nucleobase of a ribonucleotide.

[0017] In some aspects, the techniques described herein relate to a method, wherein a first or second linking group independently includes a conjugating moiety selected from the group consisting of amine-NHS ester, amine-imidoester, amine-pentofluorophenyl ester, amine-hydroxymethyl phosphine, carboxyl-carbodiimide, thiol-maleimide, thiol-haloacetyl, thiol-pyridyl disulfide, thiol-thiosulfonate, thiol-vinyl sulfone, aldehyde-hydrazide, aldehydealkoxyamine, hydroxy-isocyanate, azide-alkyne, azide-phosphine, transcyclooctene-tetrazine, norbornene-tetrazine, azide-cyclooctyne, and azide-norbornene.

[0018] In some aspects, the techniques described herein relate to a method, wherein the elongated polymer is a linear elongated polymer and includes an arresting construct attached to each nucleobase or each cyclic loop, wherein the arresting construct is configured to slow, pause, or halt the translocation.

[0019] In some aspects, the techniques described herein relate to a method, wherein the arresting construct is a linear, a branched or a cyclic polymer.

[0020] In some aspects, the techniques described herein relate to a method, wherein the arresting construct includes a synthetic hydrophobic polymer, a synthetic hydrophilic polymer, an oligonucleotide / polynucleotide, a peptide / polypeptide, or combinations thereof

[0021] In some aspects, the techniques described herein relate to a method, wherein the nanopore includes a constriction having an opening with an inner diameter from about 0.6 nm to about 1.2 nm.

[0022] In some aspects, the techniques described herein relate to a compound having one of the following structures:wherein: X is -O-, -CH2-, -=N-, or -NH-; Y is -O-, -S-, or -NH-; arc is an arresting construct; m is a positive integer; Base is a nucleobase; Li is a first linking group; L2 is a second linking group; and SP is a spacer.

[0023] In some aspects, the techniques described herein relate to a compound, wherein SP includes one or more of the following moieties: (1) alkyl chains having 10 to 200 carbons, (2) oligonucleotides having 10 to 100 repeating units, (3) polypeptides having 10 to 100 repeating units, (4) hydrophilic polymers having 10 to 100 repeating units selected from the group consisting of polyethyleneglycol, polyvinyl alcohol, polyacrylamide, polyvinylpyrrolidone, polystyrenesulfonate, and polyethyleneimine, and (5) hydrophobic polymers having 10 to 100 repeating units selected from the group consisting of polylactic acid, polymethymethacrylate, and polystyrene.

[0024] In some aspects, the techniques described herein relate to a compound, wherein each of Li and L2 independently includes a conjugating moiety selected from the group consisting of amine-NHS ester, amine-imidoester, amine-pentofluorophenyl ester, aminehydroxymethyl phosphine, carboxyl-carbodiimide, thiol-maleimide, thiol-haloacetyl, thiolpyridyl disulfide, thiol-thiosulfonate, thiol-vinyl sulfone, aldehyde-hydrazide, aldehydealkoxyamine, hydroxy-isocyanate, azide-alkyne, azide-phosphine, transcyclooctene-tetrazine, norbornene-tetrazine, azide-cyclooctyne, azide-norbornene, Cyclooctyne-tetrazine, hydroxylamine-potassium acyltrifluoroborate, tetrazine-isocyanide, and cyclooctyne- tetrachlorocyclopentadienone ethylene ketal.

[0025] In some aspects, the techniques described herein relate to a compound, wherein each of Li and L2 independently further includes a first linker between the conjugating moiety and X, and a second linker between the conjugating moiety and SP.

[0026] In some aspects, the techniques described herein relate to a compound, wherein the first linker and the second linker are independently selected from the group consisting of polynucleotide having 10 to 100 repeating units, polypeptide having 10 to 100 repeating units, alkyl chains having 10 to 200 carbons, hydrophilic polymers having 10 to 100 repeating units including polyethyleneglycol, polyvinyl alcohol, polyacrylamide, polyvinylpyrrolidone, polystyrenesulfonate, or polyethyleneimine, hydrophobic polymers having 10 to 100 repeating units including poly lactic acid, polymethymethacrylate, or polystyrene, and combinations thereof.

[0027] In some aspects, the techniques described herein relate to a compound, wherein the arc is provided in the middle, before, or after the SP, and the SP includes one or more reporter barcodes, wherein the one or more reporter barcodes correspond to and identify a nucleobase of a ribonucleotide.

[0028] In some aspects, the techniques described herein relate to a compound, wherein the arresting construct is a linear, a branched or a cyclic polymer.

[0029] In some aspects, the techniques described herein relate to a compound, wherein the arresting construct includes a synthetic hydrophobic polymer, a synthetic hydrophilic polymer, an oligonucleotide / polynucleotide, a peptide / polypeptide, or combinations thereof.

[0030] In some aspects, the techniques described herein relate to an oligonucleotide including one of the following structures:wherein: X is -O-, -CH2-, -=N-, or -NH-; Y is -O-, -S-, or -NH-; Base is a nucleobase; Li is a first linking group; L2 is a second linking group; and SP is a spacer.

[0031] In some aspects, the techniques described herein relate to an oligonucleotide, wherein cleavage of a 5’ P-0 bond in the structures lb and lib is configured

[0032] In some aspects, the techniques described herein relate to an oligonucleotide, wherein SP includes one or more of the following moieties: (1) alkyl chains having 10 to 200 carbons, (2) oligonucleotides having 10 to 100 repeating units, (3)polypeptides having 10 to 100 repeating units, (4) hydrophilic polymers having 10 to 100 repeating units selected from the group consisting of polyethyleneglycol, polyvinyl alcohol, polyacrylamide, polyvinylpyrrolidone, polystyrenesulfonate, and polyethyleneimine, and (5) hydrophobic polymers having 10 to 100 repeating units selected from the group consisting of polylactic acid, polymethymethacrylate, and polystyrene.

[0033] In some aspects, the techniques described herein relate to an oligonucleotide, wherein each of Li and L2 independently includes a conjugating moiety selected from the group consisting of amine-NHS ester, amine-imidoester, amine- pentofluorophenyl ester, amine-hydroxymethyl phosphine, carboxyl-carbodiimide, thiol- maleimide, thiol-haloacetyl, thiol-pyridyl disulfide, thiol-thiosulfonate, thiol-vinyl sulfone, aldehyde-hydrazide, aldehyde-alkoxyamine, hydroxy-isocyanate, azide-alkyne, azidephosphine, transcyclooctene-tetrazine, norbornene-tetrazine, azide-cyclooctyne, azidenorbornene, cyclooctyne-tetrazine, hydroxylamine-potassium acyltrifluoroborate, tetrazineisocyanide, and cyclooctyne-tetrachlorocyclopentadienone ethylene ketal.

[0034] In some aspects, the techniques described herein relate to an oligonucleotide, wherein each of Li and L2 independently further includes a first linker between the conjugating moiety and X, and a second linker between the conjugating moiety and SP.

[0035] In some aspects, the techniques described herein relate to an oligonucleotide, wherein the first linker and the second linker are independently selected from the group consisting of a polynucleotide having 10 to 100 repeating units, polypeptide having 10 to 100 repeating units, alkyl chains having 10 to 200 carbons, hydrophilic polymers having 10 to 100 repeating units including polyethyleneglycol, polyvinyl alcohol, polyacrylamide, polyvinylpyrrolidone, polystyrenesulfonate, or polyethyleneimine, hydrophobic polymers having 10 to 100 repeating units including poly lactic acid, polymethymethacrylate, or polystyrene, and combinations thereof.

[0036] In some aspects, the techniques described herein relate to an oligonucleotide, wherein SP further includes an arresting construct modification to slow or halt the movement of the polynucleotide through a nanopore.

[0037] In some aspects, the techniques described herein relate to an oligonucleotide, wherein the Base or the SP further includes a modification.

[0038] In some aspects, the techniques described herein relate to an oligonucleotide, wherein the modification is a linear, a branched or a cyclic polymer.

[0039] In some aspects, the techniques described herein relate to an oligonucleotide, wherein the modification includes a synthetic hydrophobic polymer, a synthetic hydrophilic polymer, an oligonucleotide / polynucleotide, a peptide / polypeptide, or combinations thereof.

[0040] In some aspects, the techniques described herein relate to a method including the compounds or ribonucleotides disclosed herein.

[0041] In some aspects, the techniques described herein relate to a kit for performing the methods disclosed herein.

[0042] In some aspects, the techniques described herein relate to a system for determining a sequence of a polynucleotide according to methods disclosed herein.

[0043] In some aspects, the techniques described herein relate to a system for performing a method for determining a sequence of a polynucleotide comprising a plurality of ribonucleotides, wherein the ribonucleotides are selected from any of the compounds disclosed herein.

[0044] It should be appreciated that all combinations of the foregoing concepts and additional concepts discussed in greater detail below are contemplated as being part of the inventive subject matter disclosed herein and may be used to achieve the benefits and advantages described herein.BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Features of examples of the present disclosure will become apparent by reference to the following detailed description and drawings, in which like reference numerals correspond to similar, though perhaps not identical, components. For the sake of brevity, reference numerals or features having a previously described function may or may not be described in connection with other drawings in which they appear.

[0046] FIG. 1 schematically illustrates an example of sequencing an elongated polynucleotide with nucleotides containing cyclic loops.

[0047] FIG. 2 schematically illustrates an example of elongating a polynucleotide with nucleotides having cleavable cyclic loops.

[0048] FIG. 3 illustrates possible reaction pathways for a sample ribonucleotide in a basic environment.

[0049] FIG. 4 illustrates the desired reaction pathway of a ribonucleotide in a basic environment.

[0050] FIG. 5 illustrates the reaction pathway of an enzyme-mediate cleavage of a ribonucleotide.

[0051] FIG. 6 illustrates the conditions for the LCMS data of FIGS. 7A-7B.

[0052] FIGS. 7A-7B illustrate LCMS data from sample nucleotide reactions.

[0053] FIGS. 8-11 illustrate example nucleotides with unexpanded or uncleaved cyclic loops.DETAILED DESCRIPTION

[0054] All patents, applications, published applications and other publications referred to herein are incorporated herein by reference to the referenced material and in their entireties. If a term or phrase is used herein in a way that is contrary to or otherwise inconsistent with a definition set forth in the patents, applications, published applications and other publications that are herein incorporated by reference, the use herein prevails over the definition that is incorporated herein by reference.Definitions

[0055] All technical and scientific terms used herein have the same meaning as commonly understood to one of ordinary skill in the art to which this disclosure belongs unless clearly indicated otherwise.

[0056] As used herein, the singular forms “a”, “and”, and “the” include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to “a sequence” may include a plurality of such sequences, and so forth.

[0057] The terms comprising, including, containing and various forms of these terms are synonymous with each other and are meant to be equally broad. Moreover, unless explicitly stated to the contrary, examples comprising, including, or having an element or a plurality of elements having a particular property may include additional elements, whether or not the additional elements have that property.

[0058] As used herein, the term “modified oligonucleotide” refers to a polymeric chain of nucleobases or nucleotides assembled with moieties comprising a modified nucleobase, modified sugar rings (e.g. LN A, constraint ethyl, ethylene bridged, TN A, 2’-0me, 2’F, 2’-M0E) or nucleobases attached to a scaffold (e.g. unlock, 4’-thio, CeNA, HNA, TNA, GNA, FNA).

[0059] As used herein, the term “nanopore” is intended to mean a hollow structure discrete from, or defined in, and extending across the membrane. The nanopore permits ions, electric current, and / or fluids to cross from one side of the membrane to the other side of the membrane. For example, a membrane that inhibits the passage of ions or water-soluble molecules can include a nanopore structure that extends across the membrane to permit the passage (through a nanoscale opening extending through the nanopore structure) of the ions or water-soluble molecules from one side of the membrane to the other side of the membrane. The diameter of the nanoscale opening extending through the nanopore structure can vary along its length (i.e., from one side of the membrane to the other side of the membrane), but at any point on the nanoscale (i.e., from about 1 nm to about 100 nm, or to less than 1000 nm). Examples of the nanopore include, for example, biological nanopores, solid-state nanopores, and biological and solid-state hybrid nanopores. In some embodiments, a nanopore refers to a pore having an opening with a diameter at its most narrow point of about 0.3 nm to about 2 nm. For example, a nanopore may be a solid-state nanopore, a graphene nanopore, an elastomer nanopore, or may be a naturally-occurring or recombinant protein that forms a tunnel upon insertion into a bilayer, thin film, membrane, or solid-state aperture, also referred to as a protein pore or protein nanopore herein (e.g., a transmembrane pore). If the protein inserts into the membrane, then the protein is a tunnel-forming protein.

[0060] As used herein, the term “diameter” is intended to mean the longest straight line inscribable in a cross-section of a nanoscale opening through a centroid of the crosssection of the nanoscale opening. It is to be understood that the nanoscale opening may or may not have a circular or substantially circular cross-section (the cross-section of the nanoscale opening being substantially parallel with the cis / trans electrodes). Further, the cross-section may be regularly or irregularly shaped.

[0061] As used herein, “cis” refers to the side of a nanopore opening through which an analyte or modified analyte enters the opening or across the face of which the analyte or modified analyte moves.

[0062] As used herein, “trans” refers to the side of a nanopore opening through which an analyte or modified analyte (or fragments thereof) exits the opening or across the face of which the analyte or modified analyte does not move.

[0063] As used herein, the term “biological nanopore” is intended to mean a nanopore whose structure portion is made from materials of biological origin. Biological origin refers to a material derived from or isolated from a biological environment such as an organism or cell, or a synthetically manufactured version of a biologically available structure. Biological nanopores include, for example, polypeptide nanopores and polynucleotide nanopores.

[0064] As used herein, a “moiety” is one of two or more parts into which something may be divided, such as, for example, the various parts of a tether, a molecule or a probe.

[0065] As used herein, a “reporter” is composed of one or more reporter elements or reporter moieties. Reporters include what are known as “tags” and “labels.” The cyclic loop (when including reporter moiety) or nucleobase residue of the elongated polymer can be considered a reporter. Reporters serve to parse the identity of the target nucleic acid. Reporters may include constituent sub-reporters, and multiple reporters may be present on a single nucleotide. When present in the readhead of a nanopore, reporters can provide distinctive and sometimes unique blockage currents at given read voltages.

[0066] As used herein, a “linker” is a molecule or moiety that joins two molecules or moieties and provides spacing between the two molecules or moieties such that they are able to function in their intended manner. For example, a linker can comprise a diamine hydrocarbon chain that is covalently bound through a reactive group on one end to an oligonucleotide analog molecule and through a reactive group on another end to a solid support, such as, for example, a bead surface. Coupling of linkers to nucleotides and substrate constructs of interest can be accomplished through the use of coupling reagents that are known in the art (see, e.g., Efimov et al., Nucleic Acids Res. 27: 4416-4426, 1999). Methods of derivatizing and coupling organic molecules are well known in the arts of organic and bioorganic chemistry. A linker may also be cleavable or reversible.

[0067] As used herein, the term “heavy atom” refers to any atom used within a molecular structure that is not hydrogen. Heavy atoms used within a modified oligonucleotide may be bridging (e.g. used to connect multiple oligonucleotides), or non-bridging (e.g. not directly linked to multiple oligonucleotides).

[0068] As used herein, the term “polypeptide nanopore” is intended to mean a protein / polypeptide that extends across the membrane, and permits ions, electric current, polymers such as DNA, RNA, or peptides, or other molecules of appropriate dimension and charge, and / or fluids to flow therethrough from one side of the membrane to the other side of the membrane. A polypeptide nanopore can be a monomer, a homopolymer, or a heteropolymer. Structures of polypeptide nanopores include, for example, an a-helix bundle nanopore and a 0-barrel nanopore. Example polypeptide nanopores include a-hemolysin, Mycobacterium smegmatis porin A (MspA), gramicidin A, maltoporin, OmpF, OmpC, PhoE, Tsx, F-pilus, etc. The protein a-hemolysin is found naturally in cell membranes, where it acts as a pore for ions or molecules to be transported in and out of cells. Mycobacterium smegmatis porin A (MspA) is a membrane porin produced by Mycobacteria, which allows hydrophilic molecules to enter the bacterium. MspA forms a tightly interconnected octamer and transmembrane beta-barrel that resembles a goblet and contains a central pore.

[0069] As used herein, a “peptide” refers to two or more amino acids joined together by an amide bond (that is, a “peptide bond”). Peptides comprise up to or include 50 amino acids. Peptides may be linear or cyclic. Peptides may be a, 0, y, 5, or higher, or mixed. Peptides may comprise any mixture of amino acids as defined herein, such as comprising any combination of D, L, a, 0, y, 5, or higher amino acids.

[0070] As used herein, a “protein” refers to an amino acid sequence having 51 or more amino acids.

[0071] A polypeptide nanopore can be synthetic. A synthetic polypeptide nanopore includes a protein-like amino acid sequence that does not occur in nature. The protein-like amino acid sequence may include some of the amino acids that are known to exist but do not form the basis of proteins (i.e., non-proteinogenic amino acids). The protein-like amino acid sequence may be artificially synthesized rather than expressed in an organism and then purified / isolated.

[0072] The nanopores disclosed herein may be hybrid nanopores. A “hybrid nanopore” refers to a nanopore including materials of both biological and non-biological origins. An example of a hybrid nanopore includes a polypeptide-solid-state hybrid nanopore and a polynucleotide-solid-state nanopore.

[0073] The application of the electric potential difference across a nanopore may force the translocation of a nucleic acid through the nanopore. One or more signals are generated that correspond to the translocation of the nucleotide through the nanopore. Accordingly, as a target polynucleotide, or as a mononucleotide or a probe derived from the target polynucleotide or mononucleotide, transits through the nanopore, the current across the membrane changes due to base-dependent (or probe dependent) blockage of the constriction, for example. The signal from that change in current can be measured using any one of a variety of methods. Each signal is unique to the species of nucleotide(s) (or cyclic loops with a reporter moiety region) in the nanopore, such that the resultant signal can be used to determine a characteristic of the polynucleotide. For example, the identity of one or more species of nucleotide(s) (or probe) that produce a characteristic signal can be determined.

[0074] As used herein, a “nucleotide” includes a nitrogen containing heterocyclic base, a sugar, and one or more phosphate groups. Nucleotides are monomeric units of a nucleic acid sequence. Examples of nucleotides include, for example, ribonucleotides or deoxyribonucleotides. In ribonucleotides, the sugar is a ribose, and in deoxyribonucleotides, the sugar is a deoxyribose, i.e., a sugar lacking a hydroxyl group that is present at the 2’ position in ribose. The nitrogen-containing heterocyclic base can be a purine base or a pyrimidine base. Purine bases include adenine (A) and guanine (G), and modified derivatives or analogs thereof. Pyrimidine bases include cytosine (C), thymine (T), and uracil (U), and modified derivatives or analogs thereof. The C-l atom of deoxyribose is bonded to N-l of a pyrimidine or N-9 of a purine. The phosphate groups may be in the mono-, di-, or tri-phosphate form. These nucleotides are natural nucleotides, but it is to be further understood that non-natural nucleotides, modified nucleotides or analogs of the aforementioned nucleotides can also be used.

[0075] As used herein, “nucleobase” is a heterocyclic base such as adenine, guanine, cytosine, thymine, uracil, inosine, xanthine, hypoxanthine, or a heterocyclic derivative, analog, or tautomer thereof. A nucleobase can be naturally occurring or synthetic.Non-limiting examples of nucleobases are adenine, guanine, thymine, cytosine, uracil, xanthine, hypoxanthine, 8-azapurine, purines substituted at the 8 position with methyl or bromine, 9-oxo-N6-methyladenine, 2-aminoadenine, 7-deazaxanthine, 7-deazaguanine, 7- deaza-adenine, N4-ethanocytosine, 2,6- diaminopurine, N6-ethano-2,6-diaminopurine, 5- methylcytosine, 5-(C3-C6)- alkynylcytosine, 5-fluorouracil, 5-bromouracil, thiouracil, pseudoisocytosine, 2-hydroxy-5-methyl-4-triazolopyridine, isocytosine, isoguanine, inosine, nitroindole, LNA, 2’-0Me, 2’-F, 7,8-dimethylalloxazine, 6-dihydrothymine, 5,6- dihydrouracil, 4-methyl-indole, ethenoadenine and the non-naturally occurring nucleobases described in U.S. Pat. Nos. 5,432,272 and 6,150,510 and PCT applications WO 92 / 002258, WO 93 / 10820, WO 94 / 22892, and WO 94 / 24144, and Fasman (“Practical Handbook of Biochemistry and Molecular Biology”, pp. 385-394, 1989, CRC Press, Boca Raton, LO), all herein incorporated by reference in their entireties.

[0076] The term “nucleic acid” or “polynucleotide” refers to a deoxyribonucleotide or ribonucleotide polymer in either single- or double-stranded form, and unless otherwise limited, encompasses known analogs of natural nucleotides that hybridize to nucleic acids in manner similar to naturally occurring nucleotides, such as peptide nucleic acids (PNAs) and phosphorothiolate DNA. Unless otherwise indicated, a particular nucleic acid sequence includes the complementary sequence thereof. Nucleotides include, but are not limited to, ATP, dATP, CTP, dCTP, GTP, dGTP, UTP, TTP, dUTP, 5-methyl-CTP, 5-methyl-dCTP, ITP, diTP, 2-amino-adenosine-TP, 2-amino-deoxyadenosine-TP, 2-thiothymidine triphosphate, pyrrolo-pyrimidine triphosphate, and 2-thiocytidine, as well as the alphathiotriphosphates for all of the above, and 2'-O-methyl-ribonucleotide triphosphates for all the above bases. Modified bases include, but are not limited to, 5-Br-UTP, 5-Br-dUTP, 5-F-UTP, 5-F-dUTP, 5-propynyl dCTP, and 5-propynyl-dUTP.

[0077] As used herein, the term “signal” is intended to mean an indicator that represents information. Signals include, for example, an electrical signal and an optical signal. The term “electrical signal” refers to an indicator of an electrical quality that represents information. The indicator can be, for example, current, voltage, tunneling, resistance, potential, voltage, conductance, or a transverse electrical effect. An “electronic current” or “electric current” refers to a flow of electric charge. In an example, an electrical signal maybe an electric current passing through a nanopore, and the electric current may flow when an electric potential difference is applied across the nanopore.

[0078] As used herein, the term “driving force” is intended to mean an electrical current that allows a polynucleotide to translocate through the nanopore. In some embodiments, the electrical current may flow when an electric potential difference is applied across the nanopore.

[0079] As used herein, the term “holding force” is intended to mean a resistance that slows and / or stops a polynucleotide to translocate through the nanopore. In some embodiments, the holding force is overcome by the application of a driving force. Thus, the driving force overcomes / overrides the resistance that slows and / or stops a polynucleotide, thereby allowing the polynucleotide to translocate through the nanopore.

[0080] As used herein, the term “arresting construct” is intended to mean a moiety attached to a nucleotide. An arresting construct may provide a resistance (in the form of a “holding force”) that slows and / or stops a polynucleotide to translocate through the nanopore unless the resistance due to the modification is overcome by a “driving force.” The resistance provided by the modification is due to a property of the modification (e.g., size, geometry, and / or non-covalent interaction with the nanopore). Arresting constructs can operate as a ratchet or a brake for the polypeptide translocation through a nanopore. An arresting construct can be attached to any part of the nucleotide and can also be attached to the nucleotide at two locations forming a loop.

[0081] As used herein, a “linear elongated polymer” is intended to mean an elongated polymer that does not have significant branching such that the longest continuous chain of the elongated polymer is a linear chain between two nucleobases. The linear elongated polymer may contain arresting constructs that are appended to or branch off of the linear elongated polymer.

[0082] The aspects and examples set forth herein and recited in the claims can be understood in view of the above definitions.Overview

[0083] A common drawback with nanopore sequencers is that the nanopore is sensitive to multiple bases of a DNA or RNA strand in the nanopore, as opposed to reading single base one at a time. For example, the MspA nanopore has a constriction region whichserves as a readhead of at least 4 nucleotides (termed a “k-mer”), resulting in minimally 256 (4A4) different permutation of 4-mer sequences that needs to be deconvoluted. For a k-mer of 5 bases, the number of possible signals is 4A5= 1,024 possible signals. A longer readhead will result in an exponential increase in the number of signals to be differentiated, which complicates the sequencing readout and increases the complexity of base calling, thus reducing accuracy. Another issue with nanopore sequencers is that the speed of translocation of natural single stranded DNA or RNA is in the order of >10 million nucleotides per second, way above the rate that is compatible with electronics and detectors.

[0084] In some embodiments, by using cleavable sites along the RNA backbone while conjoining adjacent nucleobases with a barcoding region, the disclosed technology allows the distance between adjacent nucleobases to be increased and negates the need to deconvolute a large number of signals. Once the backbone is cleaved, the reporter portion of the elongated polynucleotide would occupy the entire readhead of the nanopore for highly accurate single molecule sequencing with single base resolution.

[0085] In some embodiments, the disclosed technology allows having one nucleobase of the elongated polynucleotide to reside in the readhead at any point in time, successfully reducing the diversity of reads to 4 (A, U, C, G, or A, T, C and G for DNA), enabling more accurate sequencing at a lower cost. In some embodiments, the disclosed technology provides high throughput, cheaper and more accurate RNA sequencing.System and Method

[0086] FIG. 1 schematically illustrates an example of sequencing an elongated polynucleotide. A protein nanopore 101 is deposited in a lipid bilayer 102. An elongated polynucleotide 103 translocates through the nanopore 101. The polynucleotide 103 includes cyclic loop regions 117 between successive nucleotides. By introducing a cyclic loop 117 between successive nucleotides, the k-mer length can be reduced to 1, resulting in just 4 signals (for A, T or U, C and G), reducing the complexity of base calling. Additionally, a characteristic linker / barcode may be assigned to each of the 4 individual bases to achieve base recognition. For example, the signal unit or moiety 105 includes an “A” nucleotide and a corresponding cyclic loop 117 which may contain one or more reporters that serve as the barcode for nucleotide A. Diversity of reads is reduced to 4 with a single barcode characteristic of each nucleobase residing in the nanopore readhead.

[0087] The cyclic loop may also contain one or more modifications or “arresting” constructs (arc), which may operate to stop or slow down the translocation of the daughter sequence through the nanopore. With the incorporation of arresting constructs the nucleotide sequence may be advanced through the nanopore in a “ratcheted” manner. The advancement of the nucleotide sequence through the nanopore may be performed either in a faster “autoratchet” mode (with stochastic and fairly unpredictable advancement of the nucleotide sequence) or in a slower “pulsed-ratchet” mode which promotes the advancement of nucleotides (along with their arresting constructs) in fairly pre-defined time intervals.

[0088] By “translocation,” it is meant that an analyte (e.g., a polynucleotide, such as RNA or DNA) enters one side of an opening of a nanopore and move to and out of the other side of the opening. It is contemplated that any embodiment herein comprising translocation may refer to electrophoretic translocation or non-electrophoretic translocation, unless specifically noted. An electric field may move an analyte or modified analyte. By “interacts,” it is meant that the analyte or modified analyte moves into and, optionally, through the opening, where “through the opening” (or “translocates”) means to enter one side of the opening and move to and out of the other side of the opening. In some embodiments, physical pressure causes a modified analyte to interact with, enter, or translocate (after alteration) through the opening. In some embodiments, a magnetic bead is attached to an analyte or modified analyte on the trans side, and magnetic force causes the modified analyte to interact with, enter, or translocate (after alteration) through the opening. Other methods for translocation include but not limited to gravity, osmotic forces, temperature, and other physical forces such as centripetal force.

[0089] In some embodiments, the nanopore may comprise a solid-state material, such as silicon nitride, modified silicon nitride, silicon, silicon oxide, or graphene, or a combination thereof. In some embodiments, the nanopore is protein that forms a tunnel upon insertion into a bilayer, membrane, thin film, or solid-state aperture. In some embodiments, the nanopore is comprised in a lipid bilayer. In some embodiments, the nanopore is comprised in an artificial membrane comprising a mycolic acid. The nanopore may be a Mycobacterium smegmatis porin (Msp) having a vestibule and a constriction zone that define the tunnel. The Msp porin may be a mutant MspA porin. In some embodiments, amino acids at positions 90, 91, and 93 of the mutant MspA porin are each substituted with asparagine. Some embodimentsmay comprise altering the translocation velocity or sequencing sensitivity by removing, adding, or replacing at least one amino acid of an Msp porin. A “mutant MspA porin” is a multimer complex that has at least or at most 70, 75, 80, 85, 90, 95, 98, or 99 percent or more identity, or any range derivable therein, but less than 100%, to its corresponding wild-type MspA porin and retains tunnel-forming capability. A mutant MspA porin may be recombinant protein. Optionally, a mutant MspA porin is one having a mutation in the constriction zone or the vestibule of a wild-type MspA porin. Optionally, a mutation may occur in the rim or the outside of the periplasmic loops of a wild-type MspA porin. A mutant MspA porin may be employed in any embodiment described herein.

[0090] A “vestibule” refers to the cone-shaped portion of the interior of an Msp porin whose diameter generally decreases from one end to the other along a central axis, where the narrowest portion of the vestibule is connected to the constriction zone. A vestibule may also be referred to as a “goblet.” The vestibule and the constriction zone together define the tunnel of an Msp porin. A “constriction zone” or the “readhead” refers to the narrowest portion of the tunnel of an Msp porin, in terms of diameter, that is connected to the vestibule. The length of the constriction zone may range from about 0.3 nm to about 2 nm. Optionally, the length is about, at most about, or at least about 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0, 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 2, or 3 nm, or any range derivable therein. The diameter of the constriction zone may range from about 0.3 nm to about 2 nm. Optionally, the diameter is about, at most about, or at least about 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0, 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 2, or 3 nm, or any range derivable therein. A “tunnel” refers to the central, empty portion of an Msp porin that is defined by the vestibule and the constriction zone, through which a gas, liquid, ion, or analyte may pass. A tunnel is an example of an opening of a nanopore.

[0091] Various conditions such as light and the liquid medium that contacts a nanopore, including its pH, buffer composition, detergent composition, and temperature, may affect the behavior of the nanopore, particularly with respect to its conductance through the tunnel as well as the movement of an analyte with respect to the tunnel, either temporarily or permanently.

[0092] In some embodiments, the disclosed system for nanopore sequencing comprises an Msp porin having a vestibule and a constriction zone that define a tunnel, whereinthe tunnel is positioned between a first liquid medium and a second liquid medium, wherein at least one liquid medium comprises an analyte polynucleotide, and wherein the system is operative to detect a property of the analyte. The system may be operative to detect a property of any analyte comprising subjecting an Msp porin to an electric field such that the analyte interacts with the Msp porin. The system may be operative to detect a property of the analyte comprising subjecting the Msp porin to an electric field such that the analyte electrophoretically translocates through the tunnel of the Msp porin. In some embodiments, the system comprises an Msp porin having a vestibule and a constriction zone that define a tunnel, wherein the tunnel is positioned in a lipid bilayer between a first liquid medium and a second liquid medium, and wherein the only point of liquid communication between the first and second liquid media occurs in the tunnel. Moreover, any Msp porin described herein may be comprised in any system described herein. In some embodiments, the system may further comprise an amplifier or a data acquisition device. The system may further comprise one or more temperature regulating devices in communication with the first liquid medium, the second liquid medium, or both. The system described herein may be operative to translocate an analyte through an Msp porin tunnel either electrophoretically or otherwise.

[0093] As illustrated in FIG. 2, an elongated polynucleotide may be formed from a polynucleotide having modified nucleotides, each modified nucleotide comprises a cyclic loop modification (dashed line). A daughter strand polynucleotide can be synthesized by polymerase from a template DNA, RNA, or hybrid template using modified nucleotides (e.g., modified dNTPs). In the polymerization process, the modified dNTPs with a cyclic loop is incorporated into a growing daughter strand. Once the daughter strand is made, the polynucleotide backbone is cleaved at cleavable sites, allowing the cyclic loop modification to open and result in elongation of the daughter strand polynucleotide. The cyclic loop modifications on the modified dNTPs become the cyclic loops in the elongated polynucleotide that create distance between adjacent nucleotides.

[0094] The polymerase used is an enzyme generally for joining 3 ’-OH 5’- triphosphate nucleotides, oligomers, and their analogs. Polymerases include, but are not limited to, DNA-dependent DNA polymerases, DNA-dependent RNA polymerases, RNA- dependent DNA polymerases, RNA-dependent RNA polymerases, T7 DNA polymerase, T3 DNA polymerase, T4 DNA polymerase, T7 RNA polymerase, T3 RNA polymerase, SP6 RNApolymerase, DNA polymerase I, Klenow fragment, Thermophilus aquaticus DNA polymerase, Tth DNA polymerase, VentR® DNA polymerase (New England Biolabs), Deep VentR® DNA polymerase (New England Biolabs), Bst DNA Polymerase Large Fragment, Stoeffel Fragment, 90N DNA Polymerase, 90N DNA polymerase, Pfu DNA Polymerase, Tfl DNA Polymerase, Tth DNA Polymerase, RepliPHI Phi29 Polymerase, Tii DNA polymerase, eukaryotic DNA polymerase beta, telomerase, Therminator™ polymerase (New England Biolabs), KOD HiFi™ DNA polymerase (Novagen), K0D1 DNA polymerase, Q-beta replicase, terminal transferase, AMV reverse transcriptase, M-MLV reverse transcriptase, Phi6 reverse transcriptase, HIV-1 reverse transcriptase, novel polymerases discovered by bioprospecting, and polymerases cited in US 2007 / 0048748, US 6,329,178, US 6,602,695, and US 6,395,524 (incorporated by reference). These polymerases include wild-type, mutant isoforms, and genetically engineered variants. “Encode” or “parse” are verbs referring to transferring from one format to another and refers to transferring the genetic information of target template base sequence into an arrangement of reporters or other signaling moiety.

[0095] After the polymerization process is completed, cleavage at predetermined locations opens the loops and increases the distances between adjacent nucleotides. Cleavage of the daughter strand can be designed to occur at any part of the backbone as long as it occurs within the loop structure between the two positions where the cyclic loop is attached to the nucleotide structure. Cleavage of the daughter strand along the backbone opens the loops and elongates the daughter strand, leaving the opened cyclic loop conjoining the backbone phosphate and the sugar. In embodiments where the cyclic loop contains an arresting construct configured to interact with the nanopore, the arresting constructs may slow or halt the translocation of the elongated polymer and allow the nucleotides to be read by the nanopore one at a time.

[0096] In embodiments where a reporter moiety (such as a reporter barcode) is a part of a cyclic loop. The cleaved product, which may take the form of an elongated polymer that is substantially linear, exposes a series of reporter moieties, each of which reports the identity of the base to which it corresponds. In embodiments where the cyclic loop also contains an arresting construct configured to interact with the nanopore, the elongated polymer can be sequenced in the nanopore one barcode at a time.Cleavage Pathways

[0097] FIG. 3 illustrates potential cleave pathways associated with RNA or other genetic strands having a hydroxyl group present at the 2’ position in the cyclic sugar molecule. As shown in the figure, a cyclic loop 302 is provided between the backbone phosphate and the base. The first step of the reaction pathway involves a nucleophilic attack of the oxygen at the 2’ position which leads to a bridging of the phosphate with the ribose sugar to produce a 2’,3’- cyclic phosphate. The nucleophilic attack may be an intramolecular nucleophilic attack or it may be mediated by one or more enzymes. The unsymmetrical 2’, 3 ’-cyclic phosphate molecule then can proceed to react in one of two pathways as the negatively charged oxygen (O') again forms a double bond with the phosphate atom. In either pathway, an oxygen of the phosphate acts as a nucleophile and leaves the phosphate. In the first pathway, the 5 ’-OH group of the adjacent nucleotide reacts with an acid in solution and opens up the cyclic loop 302, which is the desired pathway. In the second pathway, the oxygen attached to the cyclic loop 302 acts as a nucleophile and reacts with an acid in solution, which forms an undesirable adduct that removes the cyclic loop connection between the base and the phosphate backbone. The 2’, 3 ’-cyclic phosphate may then further undergo hydrolysis to form either 2’ or 3 ’-extended constructs, with the phosphate remaining attached to either the 2’ or 3’ oxygen.

[0098] The second pathway involves a reaction at the apical position of the phosphate backbone molecule. However, the desirable first pathway is mediated by a reaction at an axial position of the phosphate backbone molecule. Thus, the inventors have contemplated chemical configurations which promote a reaction at an equatorial position. The apical and equatorial positions of a trigonal bipyramidal structure are shown below:. Apical position*-apLe_ '"LeqeqEquatorial position l-eq l-ap

[0099] A pentacoordinated phosphorous compound, such as that shown in the second step of the reaction pathway in FIG. 3, forms a trigonal bipyramidal structure. According to Westheimer’s rules, nucleophilic entry and departure occur through an apical position. Thus, the inventors have contemplated ways to direct the cyclic loop to the equatorial position, ensuring that the cyclic loop is not disconnected from the backbone phosphate duringcleavage, such as through enzymatic cleavage. For example, bulky groups may be provided in the axial positions of the phosphate or in the cyclic loop to direct the cyclic loop to the equatorial position, as apical positions are subject to higher steric hindrance.

[0100] One potential solution is to provide stronger or weaker electron withdrawing groups or atoms on the phosphorous of the phosphate backbone. More electronegative substituents favor nucleophilic entry and departure via the apical position, which is the desired reaction pathway for opening cyclic loops in ribose-based nucleotides, such as shown in FIG. 3. Thus, providing a less electronegative element at the axial position favors nucleophilic departure via the apical position. Oxygen has a Pauling electronegativity value of 3.5, whereas Nitrogen has a value of 3.0 and Carbon has a value of 2.5. Other elements such as Sulfur has a Pauling electronegativity value of 2.5 or silicon has a value of 2.1. Thus, the element attached to the cyclic loop or the element attached to the phosphorous may be chosen based upon the electronegativity values.

[0101] Since nitrogen and carbon have lower electronegativity values than oxygen, they may optimally link or connect the cyclic loop to the phosphate backbone. Nitrogen should optimally favor connection of the cyclic loop to the equatorial position under basic conditions, resulting in the exclusive cleavage of the 5’ P-0 bond. For nitrogen to be more apicophilic than oxygen, it should be in the protonated form, but it is well known that P-N bonds are unstable under acidic conditions. Thus, an unprotonated nitrogen may be used to connect the cyclic loop to the phosphate backbone. Carbon may alternatively be used to link the cyclic loop to the phosphate backbone, as the P-C bond is stable and carbon is a poor nucleophilic leaving group, so only one cleavage pathway is possible.

[0102] Thus, under slightly basic or basic hydrolytic conditions, the backbone P-0 bond of phosphoramidate and phosphonate adducts is expected to cleave, yielding the desired open loop construct. This is shown with regard to FIG. 4, which illustrates a single pathway for cleavage of the phosphate backbone and opening of the cyclic loop with one or more identifier sequences, spacers, linkers, and / or arresting constructs. The cyclic loop, similar to the cyclic loop 302 shown in FIG. 3, connects the nucleobase with a phosphate on the phosphate backbone via a linker, X, which is selected to the desired open loop construct.

[0103] In an alternative scenario to that of FIG. 4, the 2’ oxygen may be covalently linked to a self-immolative linker which may be attached to one or more cleavage groups, Y.The self-immolative linker may be configured to depart from the 2’ oxygen when a stimulus “Z” is exposed to the nucleotide or added as a reactant. As a result, the 2’ oxygen (with a negative charge) acts as a nucleophile to trigger the cleavage sequence. Examples of self- immolative linkers may be disulfides, allyls, Val-Cit, or TCO. The same or different stimulus reactant (such as immolative linker “R” shown below) may be provided on an adjacent nucleotide. Thus, a different stimulus reactant may be configured to open different cyclic loops on certain nucleotides. The stimulus may be a reducing agent, platinum (0), Cathepsin B, or tetrazine. One example of this alternative reaction is illustrated below:Examples of linker and stimulus pairings:” disulfide, Z = reducing agentallyl, Z = Pd(0) = Cathepsin BtetrazineEnzymatic Cleavage

[0104] Similar to the reaction pathway shown above in FIG. 3, endonucleases can cleave native RNA phosphodi ester backbones. FIG. 5 shows the reaction pathway where a 3’- Phosphate yielding endonuclease deprotonates the 2’ OH group of the substrate and donates a proton to the departing 5 ’-OH group of the adjacent nucleotide, promoting in-line attack of the newly formed 2’ -hydroxide. This is the desired pathway, which does not leave undesirable adducts and ensures high fidelity in the opening of the cyclic loop. Given the strict conformational requirements of enzymes, cyclic loop modified linkages between nucleotides are expected to follow this in-line attack mechanism, cleaving exclusively at the P-05’ bond.

[0105] In addition to endonucleases, exonucleases may be used to selectively cleave the 5’ P-0 bond. Moreover, organic agents may be used that mimic RNA endonucleases or exonucleases can be used as well.

[0106] For example, RNase may be used to selectively cleave ribonucleotide strands. RNase A is an example of an RNA nuclease that can be used to remove single-stranded RNA. RNase A effectively recognizes and cuts single-stranded RNA, including RNA in RNA:DNA hybrids that are not in a perfect double-stranded complex. Moreover, RNA bulges, loops, and even single base mismatches can be recognized and cleaved by RNase A. RNase Tl, RNase T2, RNase U2 may also be advantageously used as they bind to single stranded RNA.

[0107] Any endocytic ribonuclease of appropriate substrate specificity can be used for ribonucleotide cleavage. For cleavage with a ribonuclease it is preferred to include two or more consecutive ribonucleotides, such as from 2 to 10 or from 5 to 10 consecutive ribonucleotides. The precise sequence of the ribonucleotides is generally not material, except that certain RNases have specificity for cleavage after certain residues. RNaseA cleaves after C and U residues. Hence, when cleaving with RNaseA the cleavage site must include at least one ribonucleotide which is C or U.

[0108] Table 1 below provides a non- exhaustive list of 3 ’-phosphate yielding ribonucleases that may be used, along with their cleavage specificity.Arresting Constructs

[0109] The cyclic loop embodiments herein may be provided with one or more arresting constructs. An arresting construct is generally configured to slow, pause, or halt the translocation of the polynucleotide through the nanopore. The slowing, pausing, or halting the translocation of the elongated polymer through the nanopore may allow the nucleotides to be read by the nanopore one at a time. In some embodiments, the cyclic loop halts thetranslocation of the polynucleotide through the nanopore until a forward voltage pulse is applied to advance the polynucleotide to the next reporting region or arresting construct.

[0110] The arresting construct can be constructed of one or more durable, aqueous- or solvent-soluble polymers including, but not limited to, the following segment or segments: polyethylene glycols, polyglycols, polypyridines, polyisocyanides, polyisocyanates, poly(triarylmethyl) methacrylates, polyaldehydes, polypyrrolinones, polyureas, polyglycol phosphodiesters, polyacrylates, polymethacrylates, polyacrylamides, polyvinyl esters, polystyrenes, polyamides, polyurethanes, polycarbonates, polybutyrates, polybutadienes, polybutyrolactones, polypyrrolidinones, polyvinylphosphonates, polyacetamides, polysaccharides, polyhyaluranates, polyamides, polyimides, polyesters, polyethylenes, polypropylenes, polystyrenes, polycarbonates, polyterephthalates, polysilanes, polyurethanes, polyethers, polyamino acids, polyglycines, polyprolines, N-substituted polylysine, polypeptides, side-chain N-substituted peptides, poly -N-substituted glycine, peptoids, sidechain carboxyl-substituted peptides, homopeptides, oligonucleotides, ribonucleic acid oligonucleotides, deoxynucleic acid oligonucleotides, oligonucleotides modified to prevent Watson-Crick base pairing, oligonucleotide analogs, polycytidylic acid, polyadenylic acid, polyuridylic acid, polythymidine, polyphosphate, polynucleotides, polyribonucleotides, polyethylene glycol-phosphodiesters, peptide polynucleotide analogues, threosyl- polynucleotide analogues, glycol-polynucleotide analogues, morpholino-polynucleotide analogues, locked nucleotide oligomer analogues, polypeptide analogues, branched polymers, comb polymers, star polymers, dendritic polymers, random, gradient and block copolymers, anionic polymers, cationic polymers, polymers forming stem-loops, rigid segments and flexible segments.

[0111] In some embodiments, the arresting construct is a branch off of the chain of the cyclic loop of the ribonucleotide. In elongation of the cyclic loop with the arresting construct branch the cyclic loop may be termed linear or substantially linear as the longest chain of the cyclic loop continues in a linear or substantially linear connection between two nucleotides. In some embodiments the arresting construct is sequentially provided in the direct sequence (e.g. longest chain) of the modified cyclic loop. In some embodiments the width of the arresting construct is larger than the width of the linker, spacer, reporter, or barcode on the cyclic loop, where the length of the molecule or group of molecules is a part of the sequenceof the modified cyclic loop. In some embodiments the width of the arresting construct is at least 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 75%, at least about 100% larger than the width of the linker, spacer, reporter or barcode.

[0112] In some embodiments, each reporter or spacer is adjacent to an arresting construct. In some embodiments, the arresting construct is before or after each reporter or spacer in the cyclic loop. In some embodiments, the spacer is provided with two reporter regions flanking an arresting construct in the middle of the two reporter regions.

[0113] In some embodiments, the cyclic loop may further include one or more nonreporting spacer regions which operate to increase the distances between one or more reporting elements. In the case where a voltage pulse is applied, the non-reporting spacer permits the force on the arresting construct due to the voltage pulse to fade away before the next arresting construct arrives in the pore to avoid skipping. In the alternative, the translocation of the sequence may operate via a constant voltage (“auto ratchet” mode) that does not rely on a voltage pulse for translocation, but rather automatically advances at a slow rate during the application of a constant voltage. In the auto ratchet mode, the spacer also provides a means to control translocation rates and allow distance between one or more reporters.Cleavable Cyclic Loop Nucleotides

[0114] Cleavable cyclic loop nucleotides are nucleotides / nucleotide analogs that are modified to include a cyclic loop attached to two positions of the nucleotide / nucleotide analog structure. Although “cleavable cyclic loop nucleotide” is used, it includes both modified natural nucleotide and modified nucleotide analogs.

[0115] Depending on the structure and composition of a cyclic loop, several functions or structures may be present, including one or more of the following: conjugating moieties, linkers, spacers, reporters (reporter elements, barcodes), and arresting constructs.

[0116] Conjugating moieties in a cyclic loop modification is formed by the conjugation of the spacer moiety to one or more nucleotides or the conjugation of additional modifications (such as arresting constructs) to the cyclic loop. A reactive group at each end of the spacer moiety reacts with the reactive groups on the bifunctional nucleotide to form the conjugating moieties. In some embodiments, one or more arresting constructs may be attached to the cyclic loop through the conjugating moiety. Arresting constructs are moieties configuredto slow the translocation of the polynucleotide so one or more reporter elements could have a longer dwell time within the nanopore read head in the presence of a driving voltage, permitting identification of the reporter, and thus the corresponding base. Spacers (SP) distance successive arresting constructs to allow for sufficient decay of an applied pulse voltage before the next arresting construct. Spacers may also serve to elongate any polynucleotide once particular backbone elements are cleaved. Therefore, a cyclic loop is comprised of any number of sub-elements which may serve to affect and attenuate sequencing. In some embodiments, the strength of the background electric field, the type of nanopore, and the properties of an elongated polynucleotide affect translocation speed, efficiency, and accuracy.

[0117] Further examples of the cleavable cyclic loop nucleotide compound that may be included in the nanopore sequencing system or the kit for nanopore sequencing are:wherein X is -O-, -CH2-, -=N-, or -NH-; Y is -O-, -S-, or -NH-; arc is an arresting construct; m is a positive integer; Base is a nucleobase; Li is a first linking group; L2 is a second linking group; and SP is a spacer.

[0118] In some embodiments, each of the first linking group Li and the second linking group L2 independently comprises a conjugating moiety selected from the group consisting of amine-NHS ester, amine-imidoester, amine-pentafluorophenyl ester, aminehydroxymethyl phosphine, carboxyl-carbodiimide, thiol-maleimide, thiol-haloacetyl, thiolpyridyl disulfide, thiol-thiosulfonate, thiol-vinyl sulfone, aldehyde-hydrazide, aldehydealkoxyamine, hydroxy-isocyanate, azide-alkyne, azide-phosphine, transcyclooctene-tetrazine,norbornene-tetrazine, azide-cyclooctyne, cyclooctyne-tetrazine, hydroxylamine-potassium acyltrifluoroborate, tetrazine-isocyanide, cyclooctyne-tetrachlorocyclopentadienone ethylene ketal, and azide-norbornene. Li and L2 may or may not be the same structure or sequence of molecules.

[0119] In some embodiments, each of the first linking group Li and the second linking group L2 may independently further comprises a linker. A first linker may be present between the conjugating moiety and X / X’ (alpha phosphate), and a second linker may be present between the conjugating moiety and SP. In some embodiments, the linker may be selected from the group consisting of hydrophilic polymers (e.g., polyethylene glycol, polyvinyl alcohol, polyacrylamide, polyvinylpyrrolidone, polystyrenesulfonate, polyethyleneimine), hydrophobic polymers (e.g., polylactic acid, polymethylmethacrylate, polystyrene), oligonucleotides, peptides, polypeptides, aliphatic chains (C5 to C50) and combinations thereof. In some embodiments, the first and the second linkers may independently comprise peptides, polypeptides, alkyl chains, polyethylene glycol, or combinations thereof. In some embodiments, one or more linkers may be absent.

[0120] In some embodiments, SP comprises one or more of the following moieties: (1) simple aliphatic chains, such as alkyl chains having 5 to 50 carbons, and substituted aliphatic chains (the substituent may include halo such as chloro, bromo or fluoro, alkyl such as methyl, ethyl or propyl, or aromatic groups such as phenyl or pyridyl), (2) oligonucleotides, modified oligonucleotides or polyphosphates having 1 to 100 repeating units, (3) polypeptides having 1 to 100 repeating units, (4) hydrophilic polymers having 1 to 100 repeating units, examples include polyethyleneglycol, polyvinyl alcohol, polyacrylamide, polyvinylpyrrolidone, polystyrenesulfonate, and polyethyleneimine, and (5) hydrophobic polymers having 1 to 100 repeating units, examples include polylactic acid, polymethylmethacrylate, and polystyrene. In some embodiments, the alkyl chains may be substituted or unsubstituted. In some embodiment, the number of repeating units (monomers) in SP may range from, for example 1-5, 6-10, 11-15, 16-20, 20-25, 26-50, or 50-100, or a combination of any of the foregoing ranges. In some embodiments, the total number of repeating units in SP may be 5-100, 10-100, 10-80, 10-70, 5-60 or 5-50.

[0121] The number of the repeating units and the length of the spacer SP may depend on the following factors: (1) the choice of repeating units / monomers - a monomer thatis shorter / smaller would likely require more repeats to make up a similar length as compared to a longer monomer; (2) the steric bulk of the spacer - a larger spacer monomer would likely result in steric clash with the nanopore readhead and consequently, a slower translocation speed as compared to a less bulky monomer; (3) the interactions of the spacer with the nanopore - a spacer monomer that is capable of forming stronger interactions (e.g., electrostatic interactions, H-bonding) with nanopore residues is likely to experience slower translocation speed as compared to a monomer that forms weaker interactions (e.g., non-polar interactions); (4) the charge of the selected modifications - a loop with higher net negative charge would experience a higher translocation rate (compared to a lower net negative charged loop) in the presence of an applied voltage.

[0122] In some embodiments an arresting construct may be provided before, after, or in the middle of the spacer unit. As shown above, multiple spacers may be provided, each with an arresting construct to slow or halt the translocation of the expanded nucleotide through the nanopore. In some embodiments each spacer may have an arresting construct before and after the spacer. In addition, an arresting construct may be placed between one or more regions of a spacer. As a non-limiting example, an arresting construct may be provided after a nonencoding region of the spacer and before an encoding region or subregion of the spacer. The one or more encoding regions or subregions can be used to encode the identity of the nucleobase containing the cyclic loop.Method of Making Cyclic Loop Nucleotides

[0123] Cyclic loop nucleotides can be made by conjugating a spacer moiety to a bifunctional nucleotide. In some embodiments, the conjugation of the spacer moiety and the bifunctional nucleotide involves click chemistry. The Bifunctional nucleotide designs may be derived from the below:

[0124] X may be -CH2-, -=N-, or -NH-. Y may be -O-, -S-, -Se-, or -NH-. L’ is a linker, such as those disclosed herein. Ri and R2 are reactive groups that can utilize click chemistry to conjugate with the spacer SP. In some embodiments, Ri and / or R2 comprises click chemistry reagents. In some embodiments, Ri and R2 may independently be or contain a hydroxyl, thiocyanate, aldehyde, carboxyl, azide (-N3), amine (-NH2), alkyne, bicyclononyne (BCN), dibenzocyclooctyne (DBCO), thiol (-SH), tetrazine, trans-cyclooctyne (TCO), NHS ester, imidoester, pentofluorophenyl ester, hydroxylmethyl phosphine, carbodiimide, maleimide, haloacetyl, pyridyl disulfide, thiosulfonate, vinyl sulfone, hydrazide, alkoxyamine, isocyanate, phosphine, or norbornene. In some embodiments, Ri and R2 are the same. In other embodiments, Ri and R2 may be different.

[0125] In some embodiments, the linker L’ may be independently selected from the group consisting of hydrophilic polymers, hydrophobic polymers, oligonucleotides, peptides, polypeptides, aliphatic chains (C5 to C50) and combinations thereof. In some embodiments, the hydrophilic polymers, the hydrophobic polymers, the oligonucleotides, and the polypeptides may each have 1 to 100 repeating units. The hydrophilic polymer may comprise polyethyleneglycol, polyvinyl alcohol, polyacrylamide, polyvinylpyrrolidone, polystyrenesulfonate, polyethyleneimine, or a combination thereof. The hydrophobic polymers may comprise polylactic acid, polymethylmethacrylate, or polystyrene, or a combination thereof. In some embodiments, the linker L’ may comprise a reporter encoding the associated nucleobase. In some embodiments, the linker L’ may also further comprise an arrest construct configured to interact with the nanopore to slow the translocation of the polynucleotide in which it is incorporated.

[0126] Non-limiting representative examples of sample bifunctional nucleotides include:

[0127] In order to conjugate with a spacer moiety to form the cyclic loop nucleotide, the spacer moiety comprises a spacer SP and reactive groups Ri’ and R2’ on both ends of the spacer moiety where conjugation to the bifunctional nucleotide is desired. In some embodiments, Ri’ and R2’ may be selected from the group consisting of hydroxyl, thiocyanate, aldehyde, carboxyl, azide (-N3), amine (-NH2), alkyne, bicyclononyne (BCN), dibenzocyclooctyne (DBCO), thiol (-SH), tetrazine, trans-cyclooctyne (TCO), N- Hydroxysuccinimide (NHS) ester, imidoester, pentofluorophenyl ester, hydroxylmethyl phosphine, carbodiimide, maleimide, haloacetyl, pyridyl disulfide, thiosulfonate, vinyl sulfone, hydrazide, alkoxyamine, isocyanate, phosphine, and norbornene. In some embodiments, Ri’ and R2’ are the same. In other embodiments, Ri’ and R2’ may be different.

[0128] In some embodiments, the spacer moiety may further comprise a linker L” on one or both sides of the SP, such as between SP and Ri’ and / or between SP and R2’. In some embodiments, each of the linker L” may be independently selected from the group consisting of hydrophilic polymers, hydrophobic polymers, oligonucleotides, peptides, polypeptides, aliphatic chains (C5 to C50) and combinations thereof. In some embodiments, the hydrophilic polymers, the hydrophobic polymers, the oligonucleotides, and the polypeptides may each have 1 to 100 repeating units. The hydrophilic polymer may comprise polyethyleneglycol, polyvinyl alcohol, polyacrylamide, polyvinylpyrrolidone, polystyrenesulfonate, polyethyleneimine, or a combination thereof. The hydrophobic polymers may comprise polylactic acid, polymethylmethacrylate, or polystyrene, or a combination thereof. In some embodiments, the spacer moiety may further comprise an arrest construct, which is designed to slow down the translocation of the polynucleotide through the nanopore.

[0129] To form a cyclic loop nucleotide, Ri and R2 of the bifunctional nucleotide react with Ri’ and R2’ of the SP, respectively, to form conjugating moieties. In some embodiments, the conjugating moieties (e.g., R1-R1’ and R2-R2’) may be independentlyselected from a non-exhaustive list of chemistries such as: amine-NHS ester, amine-imidoester, amine-pentofluorophenyl ester, amine-hydroxymethyl phosphine, carboxyl-carbodiimide, thiol-maleimide, thiol-haloacetyl, thiol-pyridyl disulfide, thiol-thiosulfonate, thiol-vinyl sulfone, aldehyde-hydrazide, aldehyde-alkoxyamine, hydroxy-isocyanate, azide-alkyne, azide-phosphine, transcyclooctene-tetrazine, norbornene-tetrazine, azide-cyclooctyne, and azide-norbornene. Examples of reactive groups R1 / R2, Ri ’ / FE’, and the resulting conjugating moieties formed by R1-R1’ or R2-R2’ are shown in the table below. As shown in this table, R and R’ indicates the two groups to be joined by reaction between Ri and Ri’ or R2 and R2’. One of the two groups may represent the bifunctional nucleotide, and the other group may represent the spacer moiety.

[0130] Table 2

[0131] Spacer moieties may be generated via one or more synthetic schemes, including solid-phase, solution / liquid phase, and enzymatic based synthesis. Exemplarysynthetic methods include polymer synthesized using phosphoramidite chemistry, peptide synthesis, click chemistry, and other bioconjugation methods known in the art, which can result in the formation of phosphodiester, methylphosphonate, or phosphorothioate bonds between each moiety. In some embodiments, synthesis of spacer moieties may be executed linearly. In some embodiments, synthesis of spacer moieties may be executed via branching. In some embodiments, the synthesis of spacer moieties may be executed by joining specific segments, the segments comprising one or more sub-elements. Under linear synthesis, constituent subelements (monomers) may be added to a growing chain of sub-elements which comprise a spacer moiety. In some embodiments, the initiation of synthesis begins at or near one end of a spacer moiety, wherein a first reactive group is configured to attach to a bifunctional nucleotide, and ends at or near the other end of the spacer moiety, wherein a second reactive group is configured to attach to the bifunctional nucleotide. In some embodiments, the first and second reactive groups are chemically distinct. In some embodiments, the first and second reactive groups are chemically identical.Cleavage of Cyclic Loop Nucleotide

[0132] The daughter strand can further be subject to a condition disclosed herein suitable for cleaving the cyclic loop nucleotides, elongating the daughter strand to form an elongated polynucleotide:wherein X is -O-, -CH2-, -=N-, or -NH-; Y is -O-, -S-, or -NH-; Base is a nucleobase; Li is a first linking group; L2 is a second linking group; and SP is a spacer. The bases are defined as above and the adjacent Bases on the elongated polymer strand may be different or the same. The reaction pathway for cleavage of the cyclic loop may proceed according to the reaction mechanism described herein, such as in FIG. 3 and 4. In these reaction pathways, the 2’,3’- cyclic phosphates may further undergo hydrolysis to form either 2’ or 3 ’-extended constructs with the backbone phosphate attached to either the 2’ or 3’ oxygen.Method of Making an Elongated Oligonucleotide

[0133] In some embodiments, the methods disclosed herein may relate to a method of making an elongated oligonucleotide. The method of making an elongated oligonucleotide may comprise the steps of providing an oligonucleotide comprising two or more ribonucleotides having a cyclic loop on each ribonucleotide; a first end of a cyclic loop is attached at a first position of a ribonucleotide and a second end of the cyclic loop is attached at a second position of a ribonucleotide; wherein a nucleophile is configured to deprotonate the hydrogen on the 2’ -hydroxy group of the ribonucleotide and selectively cleave the phosphate backbone of the ribonucleotide at the P-05’ bond. In some embodiments the selective cleaving is facilitated through a 2’,3’-cyclic phosphate on the first ribonucleotide. Insome embodiments, the oligonucleotide comprises two or more ribonucleotides having a cyclic loop on each ribonucleotide. In some embodiments, deprotonation of the hydrogen on the 2’- hydroxy group is facilitated through an enzyme or a mixture of enzymes. In some embodiments, the P-X bond may be a P-N bond, a P=N bond or a P-CH2- bond.

[0134] A method of making an elongated oligonucleotide may comprise the steps of providing an oligonucleotide comprising one or more ribonucleotides comprising a cyclic loop; a first end of a cyclic loop is attached to a non-bridging P-X bond of a phosphate bridging a first ribonucleotide and a second nucleotide and a second end of the cyclic loop is attached to the nucleobase of the second ribonucleotide adjacent to the first ribonucleotide; wherein a nucleophile is configured to deprotonate the hydrogen on the 2 ’-hydroxy group of the first nucleotide ribonucleotide and selectively cleave a P-05’ bond bridging the first ribonucleotide and the second nucleotide. In some embodiments, the P-X bond may be a P-N bond, a P=N bond or a P-CH2- bond. In some embodiments, the selective cleaving is facilitated through a 2’, 3’ -cyclic phosphate on the first ribonucleotide. In some embodiments, the oligonucleotide comprises two or more ribonucleotides having a cyclic loop on each ribonucleotide. In some embodiments, deprotonation of the hydrogen on the 2’ -hydroxy group is facilitated through an enzyme or a combination of enzymes. In some embodiments, deprotonation of the hydrogen on the 2’ -hydroxy group is facilitated through an intramolecular reaction.

[0135] A method of making an elongated oligonucleotide may comprise the steps of providing an oligonucleotide comprising two or more ribonucleotides having a cyclic loop on each ribonucleotide; a first end of a cyclic loop is attached to a non-bridging P-X bond of a phosphate bridging a first ribonucleotide and a second nucleotide and a second end of the cyclic loop is attached to a position adjacent to the second ribonucleotide; wherein a nucleophile is configured to deprotonate the hydrogen on the 2 ’-hydroxy group of the first nucleotide ribonucleotide and selectively cleave a P-05’ bond bridging the first ribonucleotide and the second nucleotide. In some embodiments, the second position is a position on the phosphate backbone between the first ribonucleotide and the second nucleotide. In some embodiments, the P-X bond may be a P-N bond, a P=N bond or a P-CH2- bond. In some embodiments, the selective cleaving is facilitated through a 2’,3’-cyclic phosphate on the first ribonucleotide. In some embodiments, deprotonation of the hydrogen on the 2’ -hydroxy group is facilitatedthrough an enzyme or a combination of enzymes. In some embodiments, deprotonation of the hydrogen on the 2’ -hydroxy group is facilitated through an intramolecular reaction.Cleavable Cyclic Loop Construction

[0136] The cyclic loops can be designed symmetrically or asymmetrically, depending on the conjugation chemistries on the nucleotide. In some embodiments, each of the first linking group Li and the second linking group L2 independently comprises a conjugating moiety selected from the group consisting of amine-NHS ester, amine- imidoester, amine-pentafluorophenyl ester, amine-hydroxymethyl phosphine, carboxyl-carbodiimide, thiol-maleimide, thiol-haloacetyl, thiol-pyridyl disulfide, thiol-thiosulfonate, thiol-vinyl sulfone, aldehyde-hydrazide, aldehyde-alkoxyamine, hydroxy-isocyanate, azide-alkyne, azide-phosphine, transcyclooctene-tetrazine, norbornene-tetrazine, azide-cyclooctyne, cyclooctyne-tetrazine, hydroxylamine-potassium acyltrifluoroborate, tetrazine-isocyanide, cyclooctyne-tetrachlorocyclopentadienone ethylene ketal, and azide-norbornene. Li and L2 may or may not be the same or have the same structure.

[0137] In some embodiments, each of the first linking group Li and the second linking group L2 may independently further comprises a linker. A first linker may be present between the conjugating moiety and a base in the 4-mer oligonucleotide, and a second linker may be present between the conjugating moiety and another base in the 4-mer oligonucleotide. In some embodiments, the linker may be selected from the group consisting of hydrophilic polymers (polyethylene glycol, polyvinyl alcohol, polyacrylamide, polyvinylpyrrolidone, polystyrenesulfonate, polyethyleneimine), hydrophobic polymers (polylactic acid, polymethylmethacrylate, polystyrene), oligonucleotides, peptides, polypeptides, aliphatic chains (C5 to C50) and combinations thereof. In some embodiments, the first and the second linkers may independently comprise peptides, polypeptides, alkyl chains, polyethylene glycol, or combinations thereof. In some embodiments, one or more linkers may be absent.

[0138] In some embodiments, SP (spacer) comprises a polymer. Spacers within a cyclic loop provide buffering distance between successive cyclic loops or successive subelements comprising one or more cyclic loops. In some embodiments, the polymer in the SP comprises oligonucleotides, modified oligonucleotides, hydrophilic polymers (polyethylene glycol, polyvinyl alcohol, polyacrylamide, polyvinylpyrrolidone, polystyrenesulfonate, polyethyleneimine), hydrophobic polymers (polylactic acid, polymethylmethacrylate,polystyrene), polypeptides, aliphatic chains (C5 to C50), substituted aliphatic chains (small molecules such as chloro, bromo or fluoro, alkyl such as methyl, ethyl or propyl, or aromatic groups such as phenyl or pyridyl), or combinations thereof. In some embodiments, modified oligonucleotides may include oligonucleotides that do not have a base attached to the sugar,positive integer.

[0139] In some embodiments, modified oligonucleotides may include, where “a” is a positive integer, that can be assembled into a polymer using an oligonucleotide synthesis process. In some embodiments, the spacer may comprise one or more reporter moieties that correspond to the specific nucleobase. In other embodiments, the spacer may not include one or more reporter moieties (barcode).

[0140] In some embodiments, the cyclic loop may comprise one or more barcodes, or reporters, including reporter moieties, sub-reporters, reporter elements, and sub-reporter elements. Examples include, but not limited to, nucleosidic bases, non-nucleosidic bases, peptides or other synthetic polymers such as polyethyleneglycol, polyvinylalcohol, polyacrylamide, polyvinylpyrrolidone, polyethyleneimine, etc. In some embodiments, the reporter element can comprise of macromolecules such as crown ethers, cucurbiturils, pillararenes or cyclodextrins. Conjugation of these macromolecules to the cyclic loop construct is possible through covalent conjugation chemistries such as amine-NHS ester, amine- imidoester, amine-pentofluorophenyl ester, amine-hydroxymethyl phosphine, carboxylcarbodiimide, thiol-maleimide, thiol-haloacetyl, thiol-pyridyl disulfide, thiol-thiosulfonate, thiol-vinyl sulfone, aldehyde-hydrazide, aldehyde-alkoxyamine, hydroxy-isocyanate, azidealkyne, azide-phosphine, transcyclooctene-tetrazine, norbornene-tetrazine, azide-cyclooctyne, Cyclooctyne-tetrazine, hydroxylamine-potassium acyltrifluoroborate, tetrazine-isocyanide,cyclooctyne-tetrachlorocyclopentadienone ethylene ketal, and azide-norbornene. In some embodiments, reporters can be selected from any moiety, including spacers, conjugating moiety, and arresting constructs. In some embodiments, reporters can comprise a plurality of modifications ranging anywhere from 1-10, or 11-15, or 16-20, or 21-25, 26-50, or 50-100 units in length, as long as the reporter(s) provides reproducible signals at a given voltage waveform when resident in the readhead of a nanopore.

[0141] In some embodiments, advancement of each nucleotide bearing one or more barcodes or reporter moieties, corresponds to a translocation event. In some embodiments, the reporter moiety comprises one or more sub-reporter moieties, the sub-reporter moieties configured to identify a translocation event and generate a signal when passed through the readhead of a nanopore. In some embodiments, the reporter moiety comprises two or more sub-reporter moieties, wherein each sub-reporter moiety in the two or more sub-reporter moieties are distinguishable, reproducible, and resolvable. In some embodiments, a cyclic loop may comprise a first set of reporter moieties, and a second set of reporter moieties, wherein the first set of reporter moieties are configured to generate signals to identify particular nucleotides passing through the readhead, and wherein the second set of reporter moieties are configured to generate signals to identify the passage of each nucleotide regardless of nucleotide identity. In some embodiments the spacer may encode the reporter moieties, including the first or second reporter moieties.Cyclic Loop Embodiments

[0142] Nucleotides having cyclic loops may be comprised of one or more spacers, arresting constructs, linkers, and reporters, as discussed herein. The following nucleotides depicted in FIGS. 8-11 may be implemented in some embodiments disclosed herein. In some embodiments the spacers may be provided with reporters / encoding regions or the spacers may be comprised entirely of reporters / encoding regions. In some embodiments, the spacers may be symmetrical and bisected by an arresting construct, where the symmetrical spacers may provide the same signal or a different signal from the other spacer. In some embodiments, there may be multiple spacers having encoding regions with an arresting construct provided in the middle of or flanking the ends of the spacers.

[0143] FIG. 8 shows an example nucleotide having an unexpanded or uncleaved cyclic loop including one or more linking groups 702, one or more spacers 704, and at leastone arresting construct 706. In the nucleotide shown in FIG. 8 the cyclic loop structure is symmetrical with spacer regions flanking each side of the arresting construct, or, put another way, the arresting construct bisecting two symmetrical spacer regions. The arresting construct 706 is generally branched off of the cyclic loop but the cyclic loop is generally linear, as the branch is not a part of the longest chain of subunits in the cyclic loop. Thus, when the cyclic loop is extended it may comprise a linear elongated polymer. The cyclic loop may extend from the alpha phosphate to the nucleobase (thymine shown in FIG. 8), but other configurations, such as those mentioned herein, could be applied.

[0144] In FIG. 8 the spacer regions generally comprises one or more base units in sequence. These base units may provide spacing between other parts of the cyclic loop or provide encoding or reporter signals, such as the sequence of Base 1, Base 2, Base 3, and Base 4. The sequence of the bases in the cyclic loop or in each reporter region may indicate the identity of the ribonucleotide attached to the phosphate backbone. In some embodiments the sequence of bases on each side of the arresting construct 706 may be identical, such that the reading of Base 1, Base 2, Base 3, and Base 4 is duplicative, reducing the possibility of readout error through redundancy. However, Base 1, Base 2, Base 3, and Base 4 on each side of the arresting construct may be unique, such that the combination of Base 1, Base 2, Base 3, and Base 4 on each side of the arresting construct 706 constitutes a signal for the identity of the ribonucleotide. Each region such as linking groups 702, spacers 704, and arresting construct 706 may have one or more linkers or conjugating groups in order to connect the various portions of the cyclic loop or to facilitate chemical synthesis.

[0145] The bases (Base 1, Base 2, Base 3, and Base 4) may be A, T, C, G, or U or substituted with one or more modified nucleosidic bases such as inosine, nitroindole, LNA, 2’-OMe, or 2’-F:

[0146] FIG. 9 depicts an additional example of a nucleotide having an unexpanded or uncleaved cyclic loop including one or more linking groups 802, one or more spacers 804,and one or more arresting constructs 806. The cyclic loop is connected to the alpha phosphate position and the base (thymine shown here) but other configurations discussed herein are possible. As with FIG. 8, various linkers or other conjugating chemistry may be supplied in the cyclic loop to facilitate synthesis. FIG. 9 differs from FIG. 8 in that the one or more spacers 804 includes polymeric subunits instead of bases. These polymeric subunits may encode the identity of the ribonucleotide and may be bisected by the arresting construct 806.

[0147] As depicted in FIG. 9, various non-nucloesidic moieties may be used as reporter / encoding regions or subregions or may provide spacing between encoding regions or subregions, with various non-limiting examples shown in FIG. 9. Other non-nucleosidic moieties may be used, such as spC12, spermine, phosphorothioate, or methylphosphonate:

[0148] The sequence of the non-nucleosidic moieties in the cyclic loop may indicate the identity of the ribonucleotide attached to the phosphate backbone. The sequence of non-nucleosidic moieties may be provided in conjunction with the signal of the ribonucleotide or in conjunction with signals from nucleosidic base units that may be provided in the cyclic loop. In some embodiments the symmetrical pattern of the non-nucleosidic moieties on each side of the arresting construct may provide the same signal (e.g. 1, 2, 3 and 1, 2, 3) or an inverted signal (e.g. 1, 2, 3 and 3, 2, 1) in order to provide redundancy and indicate the identity of the ribonucleotide.

[0149] FIG. 10 depicts an additional example of a nucleotide having an unexpanded or uncleaved cyclic loop including one or more linking groups 902, one or more spacers 904, and an arresting construct 906. In this embodiment, the spacers 904 comprise a sequence of modified moieties with an arresting construct 906 bisecting the spacer subregions. The modified moieties may provide a signal that may be indicative of the ribonucleotide, or they may provide spacing between the readout of other signals, such as the signal from the base of the ribonucleotide. Other modified moieties may be used as alternatives to the moieties shown in FIG. 10 such as PEG2, PEG6, PEGU, PEG12, or Lys(FITC) (for example poly- Lysine-FITC labeled).

[0150] FIG. 11 depicts an additional example of a nucleotide having an unexpanded or uncleaved cyclic loop including one or more linking groups 1002, one or more spacers 1004, and an arresting construct 1006. This embodiment illustrates an asymmetric loop where one or more nucleosidic moieties are used in conjunction with one or more non- nucleosidic moieties. In some embodiments, the nucleosidic moieties may provide reporter encoding regions and the non-nucleosidic moieties may provide non-encoding regions. In some embodiments, both nucleosidic and non-nucleosidic moieties may provide an encoding signal used to identify the ribonucleotide. Further, as shown in FIG. 11, the arresting construct may be placed to bisect any portion of the spacer region. For example, in some embodiments the arresting construct may be provided at the beginning of the spacer 1004, such as shown in FIG. 11. The molecules in the spacer 1004 may be any of those identified previously such as nucleosidic bases such as A, T, C, G, or U or modified bases inosine, nitroindole, LNA, 2’- OMe, 2’-F, non-nucleosidic moieties such as spC12, spermine, phosphorothioate, methylphosphonate, or amino acid residues or moieties such as PEG2, PEG6, PEGU, PEG12, or Lys(FITC).Sample Dinucleotide 5

[0151] A sample dinucleotide, dinucleotide 5, was synthesized as a phosphoramidate. The synthesis for dinucleotide 5 is shown below:

[0152] This sample dinucleotide 5 was used to investigate the chemoselectivity of hydrolysis under basic conditions. As a substitute for a cyclic loop, a straight hydrocarbon chain was provided on the amine attached to the phosphate backbone. This straight hydrocarbon is believed to react similarly to the various different cyclic loops, which may include optional linkers. However, as discussed above, apical positions may be subjected to higher steric hinderance. Thus, even where the cyclic loop has bulkier groups and / or higher steric hinderance the straight hydrocarbon chain would react similarly, as they both would be directed to or provided in the equatorial position. The hydrocarbon chain substitution also allowed the cleavage to be inspected via LCMS, as discussed with reference to FIG. 6-7B.

[0153] Treatment of dinucleotide 5 with trisulfonium difluorotrimethylsilicate (TASF) or tetrabutylammonium fluoride (TB AF) yielded selective cleavage, as shown in FIG. 6. Under both conditions, cleavage occurred at the 5’ P-0 bond, which was confirmed by liquid chromatography mass spectrometry (LCMS).

[0154] FIGS. 7A-7B show the LCMS results of dinucleotide 5 treated with TASF or TBAF. The combined mass of the dinucleotide is 787.32 amu (g / mol). Treatment with TASF or TBAF, both of which the LCMS spectrum is provided in FIG. 7A and 7B, show that the cleavage occurred at the 5’ P-0 bond, instead of cleavage at the P-N bond. In particular, ifcleavage occurred at the amine bond two molecules of having a mass of 574.13 amuand 101.12 amu would have been present in the mass spectrometry. However, only two molecules having a mass of 407.15 amu and 284.10 amu were observed. Thus, the cleavage did not occur at the P-N bond, but rather occurred at the 5’ P-0 bond, which is the desired product that does not result in a disconnection of the cyclic loop from either nucleotide.

[0155] FIGS. 6-7B also demonstrate that the kinetics of 5’ P-0 cleavage occur in a mildly basic environment with TASF or with a basic environment with TBAF. Thus, the basicity of the environment may be optimized or tailored based upon the desired conditions for the specific nanopore or nucleotide sequence.Additional Notes

[0156] It should be appreciated that all combinations of the foregoing concepts and additional concepts discussed in greater detail below (provided such concepts are not mutually inconsistent) are contemplated as being part of the inventive subject matter disclosed herein. In particular, all combinations of claimed subject matter appearing at the end of this disclosure are contemplated as being part of the inventive subject matter disclosed herein. It should also be appreciated that terminology explicitly employed herein that also may appear in any disclosure incorporated by reference should be accorded a meaning most consistent with the particular concepts disclosed herein.

[0157] Reference throughout the specification to “one example”, “another example”, “an example”, and so forth, means that a particular element (e.g., feature, structure, and / or characteristic) described in connection with the example is included in at least one example described herein, and may or may not be present in other examples. In addition, it is to be understood that the described elements for any example may be combined in any suitable manner in the various examples unless the context clearly dictates otherwise.

[0158] It is to be understood that the ranges provided herein include the stated range and any value or sub-range within the stated range, as if such value or sub-range were explicitly recited. For example, a range from about 2 nm to about 20 nm should be interpreted to include not only the explicitly recited limits of from about 2 nm to about 20 nm, but also to include individual values, such as about 3.5 nm, about 8 nm, about 18.2 nm, etc., and sub-ranges, such as from about 5 nm to about 10 nm, etc. Furthermore, when “about” and / or “substantially”are / is utilized to describe a value, this is meant to encompass minor variations (up to + / - 10%) from the stated value.

[0159] While several examples have been described in detail, it is to be understood that the disclosed examples may be modified. Therefore, the foregoing description is to be considered non-limiting.

[0160] While certain examples have been described, these examples have been presented by way of example only, and are not intended to limit the scope of the disclosure. Indeed, the novel methods and systems described herein may be embodied in a variety of other forms. Furthermore, various omissions, substitutions and changes in the systems and methods described herein may be made without departing from the spirit of the disclosure. The accompanying claims and their equivalents are intended to cover such forms or modifications as would fall within the scope and spirit of the disclosure.

[0161] Features, materials, characteristics, or groups described in conjunction with a particular aspect, or example are to be understood to be applicable to any other aspect or example described in this section or elsewhere in this specification unless incompatible therewith. All of the features disclosed in this specification (including any accompanying claims, abstract and drawings), and / or all of the steps of any method or process so disclosed, may be combined in any combination, except combinations where at least some of such features and / or steps are mutually exclusive. The protection is not restricted to the details of any foregoing examples. The protection extends to any novel one, or any novel combination, of the features disclosed in this specification (including any accompanying claims, abstract and drawings), or to any novel one, or any novel combination, of the steps of any method or process so disclosed.

[0162] Furthermore, certain features that are described in this disclosure in the context of separate implementations can also be implemented in combination in a single implementation. Conversely, various features that are described in the context of a single implementation can also be implemented in multiple implementations separately or in any suitable sub-combination. Moreover, although features may be described above as acting in certain combinations, one or more features from a claimed combination can, in some cases, be excised from the combination, and the combination may be claimed as a sub-combination or variation of a sub-combination.

[0163] Moreover, while operations may be depicted in the drawings or described in the specification in a particular order, such operations need not be performed in the particular order shown or in sequential order, or that all operations be performed, to achieve desirable results. Other operations that are not depicted or described can be incorporated in the example methods and processes. For example, one or more additional operations can be performed before, after, simultaneously, or between any of the described operations. Further, the operations may be rearranged or reordered in other implementations. Those skilled in the art will appreciate that in some examples, the actual steps taken in the processes illustrated and / or disclosed may differ from those shown in the figures. Depending on the example, certain of the steps described above may be removed or others may be added. Furthermore, the features and attributes of the specific examples disclosed above may be combined in different ways to form additional examples, all of which fall within the scope of the present disclosure. Also, the separation of various system components in the implementations described above should not be understood as requiring such separation in all implementations, and it should be understood that the described components and systems can generally be integrated together in a single product or packaged into multiple products. For example, any of the components for an energy storage system described herein can be provided separately, or integrated together (e.g., packaged together, or attached together) to form an energy storage system.

[0164] For purposes of this disclosure, certain aspects, advantages, and novel features are described herein. Not necessarily all such advantages may be achieved in accordance with any particular example. Thus, for example, those skilled in the art will recognize that the disclosure may be embodied or carried out in a manner that achieves one advantage or a group of advantages as taught herein without necessarily achieving other advantages as may be taught or suggested herein.

[0165] Conditional language, such as “can,” “could,” “might,” or “may,” unless specifically stated otherwise, or otherwise understood within the context as used, is generally intended to convey that certain examples include, while other examples do not include, certain features, elements, and / or steps. Thus, such conditional language is not generally intended to imply that features, elements, and / or steps are in any way required for one or more examples or that one or more examples necessarily include logic for deciding, with or without user inputor prompting, whether these features, elements, and / or steps are included or are to be performed in any particular example.

[0166] Conjunctive language such as the phrase “at least one of X, Y, and Z,” unless specifically stated otherwise, is otherwise understood with the context as used in general to convey that an item, term, etc. may be either X, Y, or Z. Thus, such conjunctive language is not generally intended to imply that certain examples require the presence of at least one of X, at least one of Y, and at least one of Z.

[0167] Language of degree used herein, such as the terms “approximately,” “about,” “generally,” and “substantially” represent a value, amount, or characteristic close to the stated value, amount, or characteristic that still performs a desired function or achieves a desired result.

[0168] The scope of the present disclosure is not intended to be limited by the specific disclosures of preferred examples in this section or elsewhere in this specification, and may be defined by claims as presented in this section or elsewhere in this specification or as presented in the future. The language of the claims is to be interpreted broadly based on the language employed in the claims and not limited to the examples described in the present specification or during the prosecution of the application, which examples are to be construed as non-exclusive.

[0169] Although the foregoing invention has been described in terms of certain preferred embodiments, other embodiments will be apparent to those of ordinary skill in the art. Additionally, other combinations, omissions, substitutions and modification will be apparent to the skilled artisan, in view of the disclosure herein. Accordingly, the present invention is not intended to be limited by the recitation of the preferred embodiments, but is instead to be defined by reference to the appended claims.

[0170] The terminology used in the description presented herein is not intended to be interpreted in any limited or restrictive manner and unless otherwise indicated refers to the ordinary meaning as would be understood by one of ordinary skill in the art in view of the specification. Furthermore, embodiments may comprise, consist of, consist essentially of, several novel features, no single one of which is solely responsible for its desirable attributes or is believed to be essential to practicing the embodiments herein described. As used herein, the section headings are for organizational purposes only and are not to be construed as limitingthe described subject matter in any way. All literature and similar materials cited in this application, including but not limited to, patents, patent applications, articles, books, treatises, and internet web pages are expressly incorporated by reference in their entirety for any purpose. When definitions of terms in incorporated references appear to differ from the definitions provided in the present teachings, the definition provided in the present teachings shall control. It will be appreciated that there is an implied “about” prior to the temperatures, concentrations, times, etc. discussed in the present teachings, such that slight and insubstantial deviations are within the scope of the present teachings herein.

[0171] Although this disclosure is in the context of certain embodiments and examples, those of ordinary skill in the art will understand that the present disclosure extends beyond the specifically disclosed embodiments to other alternative embodiments and / or uses of the embodiments and obvious modifications and equivalents thereof. In addition, while several variations of the embodiments have been shown and described in detail, other modifications, which are within the scope of this disclosure, will be readily apparent to those of ordinary skill in the art based upon this disclosure. It is also contemplated that various combinations or sub-combinations of the specific features and aspects of the embodiments may be made and still fall within the scope of the disclosure. It should be understood that various features and aspects of the disclosed embodiments can be combined with, or substituted for, one another in order to form varying modes or embodiments of the disclosure. Thus, it is intended that the scope of the present disclosure herein disclosed should not be limited by the particular disclosed embodiments described above.

Claims

WHAT IS CLAIMED IS:

1. A method for determining a sequence of a polynucleotide in a nanopore-based sequencing system, the method comprising: providing a polynucleotide comprising a plurality of ribonucleotides, wherein each ribonucleotide comprises a cyclic loop, the cyclic loop having a first end attached to a first position of the ribonucleotide and a second end attached to the second position of the ribonucleotide; selectively cleaving a 5’ P-0 bond of the phosphate backbone on each of the plurality of ribonucleotides between the first and the second positions, thereby opening the cyclic loop to form a linear cyclic loop in the form of an elongated polymer, wherein the cleaved 5’ P-0 bond oxygen is not directly connected to the cyclic loop; applying a voltage to cause the elongated polymer to insert into and translocate through a nanopore; and(i) detecting and identifying one or more reporter barcodes in the cyclic loop when the opened cyclic loop passes through the nanopore; or(ii) detecting and identifying a base on the ribonucleotide when the ribonucleotide passes through the nanopore.

2. The method of Claim 1, wherein the cleaving is mediated by one or more nuclease enzyme(s) or self-immolative groups.

3. The method of Claim 2, wherein the nuclease enzyme is a 3 ’ -Phosphate yielding endonuclease.

4. The method of Claim 2, wherein the nuclease is selected from the group consisting of RNase A, RNase Ti, RNase T2, RNase U2, RNase I, and Phosphodiesterase II.

5. The method of Claims 1-3, wherein the cyclic loop comprises a first linking group, a second linking group, and a spacer between the first and second linking groups.

6. The method of any one of Claims 1-4, wherein the cyclic loop comprises one or more linking groups, an arresting construct, and one or more reporter barcodes.

7. The method of claim 5, wherein the spacer comprises a polynucleotide having 10 to 100 repeating units, polypeptide having 10 to 100 repeating units, alkyl chains having 10 to 200 carbons, hydrophilic polymers having 10 to 100 repeating units selected form the group consisting of polyethyleneglycol, polyvinyl alcohol, polyacrylamide, polyvinylpyrrolidone, polystyrenesulfonate, and polyethyleneimine, hydrophobic polymers having 10 to 100 repeating units selected from the group consisting of polylactic acid, polymethymethacrylate, and polystyrene, and combinations thereof.

8. The method of Claim 5, wherein the spacer comprises one or more reporter barcodes, wherein the one or more reporter barcodes correspond to and identify a nucleobase of a ribonucleotide.

9. The method of Claim 5, wherein a first or second linking group independently comprises a conjugating moiety selected from the group consisting of amine-NHS ester, amine-imidoester, amine-pentofluorophenyl ester, amine-hydroxymethyl phosphine, carboxyl-carbodiimide, thiol-maleimide, thiol-haloacetyl, thiol-pyridyl disulfide, thiolthiosulfonate, thiol-vinyl sulfone, aldehyde-hydrazide, aldehyde-alkoxyamine, hydroxyisocyanate, azide-alkyne, azide-phosphine, transcyclooctene-tetrazine, norbornene-tetrazine, azide-cyclooctyne, azide-norbornene, cyclooctyne-tetrazine, hydroxylamine-potassium acyltrifluoroborate, tetrazine-isocyanide, and cyclooctyne-tetrachlorocyclopentadienone ethylene ketal.

10. The method of any one of Claims 1 to 9, wherein the elongated polymer is a linear elongated polymer and comprises an arresting construct attached to each nucleobase or each cyclic loop, wherein the arresting construct is configured to slow, pause, or halt the translocation.

11. The method of Claim 10, wherein the arresting construct is a linear, a branched or a cyclic polymer.

12. The method of Claim 11, wherein the arresting construct comprises a synthetic hydrophobic polymer, a synthetic hydrophilic polymer, an oligonucleotide / polynucleotide, a peptide / polypeptide, or combinations thereof.

13. The method of any one of Claims 1 to 12, wherein the nanopore comprises a constriction having an opening with an inner diameter from about 0.6 nm to about 1.2 nm.

14. A compound having one of the following structures:wherein:X is -O-, -CH2-, -=N-, or -NH-;Y is -O-, -S-, or -NH-;ARC is an arresting construct; m is a positive integer;Base is a nucleobase;Li is a first linking group;L2 is a second linking group; and SP is a spacer.

15. The compound of Claim 14, wherein SP comprises one or more of the following moieties:(1) alkyl chains having 10 to 200 carbons,(2) oligonucleotides having 10 to 100 repeating units,(3) polypeptides having 10 to 100 repeating units,(4) hydrophilic polymers having 10 to 100 repeating units selected from the group consisting of polyethyleneglycol, polyvinyl alcohol, polyacrylamide, polyvinylpyrrolidone, polystyrenesulfonate, and polyethyleneimine, and(5) hydrophobic polymers having 10 to 100 repeating units selected from the group consisting of polylactic acid, polymethymethacrylate, and polystyrene.

16. The compound of Claim 14 or 15, wherein each of Li and L2 independently comprises a conjugating moiety selected from the group consisting of amine-NHS ester, amine-imidoester, amine-pentofluorophenyl ester, amine-hydroxymethyl phosphine, carboxyl-carbodiimide, thiol-maleimide, thiol-haloacetyl, thiol-pyridyl disulfide, thiolthiosulfonate, thiol-vinyl sulfone, aldehyde-hydrazide, aldehyde-alkoxyamine, hydroxyisocyanate, azide-alkyne, azide-phosphine, transcyclooctene-tetrazine, norbornene-tetrazine, azide-cyclooctyne, cyclooctyne-tetrazine, hydroxylamine-potassium acyltrifluoroborate, tetrazine-isocyanide, cyclooctyne-tetrachlorocyclopentadienone ethylene ketal, and azidenorbornene.

17. The compound of Claim 16, wherein each of Li and L2 independently further comprises a first linker between the conjugating moiety and X, and a second linker between the conjugating moiety and SP.

18. The compound of Claim 17, wherein the first linker and the second linker are independently selected from the group consisting of polynucleotide having 10 to 100 repeating units, polypeptide having 10 to 100 repeating units, alkyl chains having 10 to 200 carbons, hydrophilic polymers having 10 to 100 repeating units comprising polyethyleneglycol, polyvinyl alcohol, polyacrylamide, polyvinylpyrrolidone, polystyrenesulfonate, or polyethyleneimine, hydrophobic polymers having 10 to 100 repeating units comprising polylactic acid, polymethymethacrylate, or polystyrene, and combinations thereof.

19. The compound of any one of Claims 14 to 18, wherein the arc is provided in the middle of, before, or after the SP, and the SP comprises one or more reporter barcodes,wherein the one or more reporter barcodes correspond to and identify a nucleobase of a ribonucleotide.

20. The compound of any one of Claims 14 to 18, wherein the arresting construct is a linear, a branched or a cyclic polymer.

21. The compound of Claim 20, wherein the arresting construct comprises a synthetic hydrophobic polymer, a synthetic hydrophilic polymer, an oligonucleotide / polynucleotide, a peptide / polypeptide, or combinations thereof.

22. An oligonucleotide comprising one of the following structures:Hb) •wherein:X is -O-, -CH2-, -=N-, or -NH-;Y is -O-, -S-, or -NH-;Base is a nucleobase;Li is a first linking group;L2 is a second linking group; and SP is a spacer.

23. The oligonucleotide of Claim 22, wherein cleavage of a 5’ P-0 bond in the structures lb and lib is configured to generate:

24. The oligonucleotide of any one of Claims 22 to 23, wherein SP comprises one or more of the following moieties:(1) alkyl chains having 10 to 200 carbons,(2) oligonucleotides having 10 to 100 repeating units,(3) polypeptides having 10 to 100 repeating units,(4) hydrophilic polymers having 10 to 100 repeating units selected from the group consisting of polyethyleneglycol, polyvinyl alcohol, polyacrylamide, polyvinylpyrrolidone, polystyrenesulfonate, and polyethyleneimine, and(5) hydrophobic polymers having 10 to 100 repeating units selected from the group consisting of polylactic acid, polymethymethacrylate, and polystyrene.

25. The oligonucleotide of any one of Claims 22 to 24, wherein each of Li and L2 independently comprises a conjugating moiety selected from the group consisting of amine- NHS ester, amine- imidoester, amine-pentofluorophenyl ester, amine-hydroxymethyl phosphine, carboxyl-carbodiimide, thiol-maleimide, thiol-haloacetyl, thiol-pyridyl disulfide, thiol-thiosulfonate, thiol-vinyl sulfone, aldehyde-hydrazide, aldehyde-alkoxyamine, hydroxyisocyanate, azide-alkyne, azide-phosphine, transcyclooctene-tetrazine, norbornene-tetrazine, azide-cyclooctyne, cyclooctyne-tetrazine, hydroxylamine-potassium acyltrifluoroborate,tetrazine-isocyanide, cyclooctyne-tetrachlorocyclopentadienone ethylene ketal, and azidenorbornene.

26. The oligonucleotide of Claim 25, wherein each of Li and L2 independently further comprises a first linker between the conjugating moiety and X, and a second linker between the conjugating moiety and SP.

27. The oligonucleotide of Claim 26, wherein the first linker and the second linker are independently selected from the group consisting of a polynucleotide having 10 to 100 repeating units, polypeptide having 10 to 100 repeating units, alkyl chains having 10 to 200 carbons, hydrophilic polymers having 10 to 100 repeating units comprising polyethyleneglycol, polyvinyl alcohol, polyacrylamide, polyvinylpyrrolidone, polystyrenesulfonate, or polyethyleneimine, hydrophobic polymers having 10 to 100 repeating units comprising polylactic acid, polymethymethacrylate, or polystyrene, and combinations thereof.

28. The oligonucleotide of any one of Claims 22 to 27, wherein SP further comprises an arresting construct modification to slow or halt the movement of the polynucleotide through a nanopore.

29. The oligonucleotide of any one of Claims 23 to 27, wherein the Base or the SP further comprises a modification.

30. The oligonucleotide of Claim 29, wherein the modification is a linear, a branched or a cyclic polymer.

31. The oligonucleotide of Claims 29 or 30, wherein the modification comprises a synthetic hydrophobic polymer, a synthetic hydrophilic polymer, an oligonucleotide / polynucleotide, a peptide / polypeptide, or combinations thereof.

32. The method of Claim 1, wherein the polynucleotide comprises one of the compounds of any one of Claims 14 to 21.

33. The method of Claim 1, wherein each of the nucleotides of the plurality of ribonucleotides are selected from the compounds according to any one of Claims 14 to 21.

34. A kit for performing a method for determining a sequence of a polynucleotide in a nanopore-based sequencing system, the kit comprising the compound according to any one of Claims 14 to 21.

35. A system for determining a sequence of a polynucleotide, the system configured to perform a method according to any one of Claims 1 to 13.

36. A system for performing a method for determining a sequence of a polynucleotide comprising a plurality of ribonucleotides, each ribonucleotide selected from any of the compounds according to any one of Claims 14-21.

Citation Information

Patent Citations

  • Mutant polymerases for sequencing and genotyping

    US20070048748A1

  • Method for incorporating into a DNA or RNA oligonucleotide using nucleotides bearing heterocyclic bases

    US5432272A

  • Modified oligonucleotides, their preparation and their use

    US6150510A

  • DNA polymerase mutant having one or more mutations in the active site

    US6329178B1

  • Thermostable polymerases having altered fidelity and method of identifying and using same

    US6395524B2

Cited By

  • Modified nucleotides and related methods for nanopore sequencing

    WO2026043922A1