Cleavable Circular Loop Nucleotides for Nanopore Sequencing
By introducing circular cyclic nucleotides and cleavable connectors into nanopore sequence technology, the problem of simultaneous induction of multiple bases is solved, achieving higher sequence decoding accuracy and accuracy.
Patent Information
- Application Number
- JP2024564938
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-05-04
- Filing Date
- 2023-05-04
- Publication Date
- 2025-05-13
AI Technical Summary
The existing nanopore sequence technology faces the problem of simultaneous induction of multiple bases, resulting in signal confusion and increasing the complexity and error of sequence accuracy.
By using circular loop nucleotides and cleavable connectors, linking constructs are introduced between multiple nucleotides to form cleavable loop chains, and only the encoded signals in the link construct are read in the nanopore read head.
It effectively reduces the types of signals sensed in nanopore read heads, improves the accuracy and accuracy of sequence decoding, and reduces sequence complexity and cost.
Smart Images

Figure 2025515100000001_ABST
Abstract
Description
[Technical field]
[0001] Some polynucleotide sequencing techniques involve performing a large number of controlled reactions on a support surface or inside a given reaction chamber. The controlled reactions can then be observed or detected, and subsequent analysis can help identify the properties of the polynucleotides involved in the reactions. Examples of such sequencing techniques include sequencing by ligation, sequencing by synthesis, next generation sequencing or massively parallel sequencing involving reversible terminator chemistry, or pyrosequencing approaches.
[0002] Some polynucleotide sequencing techniques utilize nanopores, which can provide a pathway for ionic current. For example, as a polynucleotide traverses through a nanopore, it affects the current flow through the nanopore. Each passing nucleotide, or series of nucleotides, that passes through the nanopore results in a characteristic blocking current. These characteristic currents of the traversing polynucleotide can be recorded to determine the sequence of the polynucleotide. Summary of the Invention
[0003] The nanopore read head (e.g., the constricted region of the nanopore) typically "senses" several bases simultaneously along the sample DNA strand, increasing the challenge of accurate nanopore sequencing due to the many permutations of signals resulting from different sequences. For example, an MspA pore reads approximately 4 bases at a time, resulting in at least 4^4=256 different signals that need to be deconvoluted and resolved.
[0004] In one aspect, the disclosed technology provides a method to synthesize a daughter strand using circular loop nucleotides instead of directly sequencing the sample DNA. In some embodiments, each circular loop nucleotide contains a unique barcoding / reporter region specific to the original base (e.g., A, T, C, or G) and a cleavable site. The daughter strand is then "extended" by cleaving the cleavable site. As a result, when sequencing the daughter strand, the nanopore can "read" the barcoding / reporter region and identify the base it encodes. The linker and barcoding constructs introduced into the daughter strand via polymerization are designed to fully occupy the nanopore's read head, thus reducing the number of signals to only four, i.e., one per nucleic acid base. Thus, the disclosed technology allows for barcode-based decoding of individual bases. In some embodiments, the circular loop contains a non-barcoding linker construct that allows the daughter strand to be extended after cleavage of the cleavable site. The non-barcoding linker constructs can generate a distinguishable signal or a distinguishable signal break from the signal of the nucleobase when passing through the nanopore, thereby isolating and / or enhancing the recorded signal from the nucleobase. Thus, the disclosed technology allows for improved resolution of the recorded signal. In some embodiments, the linker constructs can contain both barcoding / reporter regions and non-barcoding linker constructs. In some embodiments, the linker constructs can be barcoding / reporter elements. In some embodiments, the nucleotides and oligonucleotides are modified using heavy atoms (including, for example, sulfur and selenium).
[0005] In another aspect, the disclosed technology provides systems, devices, kits, and methods that allow for cleavable bonds along the DNA backbone, synthesis of cleavable circular loop nucleotides, barcodes for individual base identification, and polymerase mutations for incorporation of modified nucleotides. Systems can be prepared to allow parallel reading at multiple nanopores, such as thousands or millions of nanopores. Thus, components of any system can be functionally replicated to increase sequencing throughput. Any system can also be adapted for microfluidics or automation.
[0006] The systems, devices, kits, and methods disclosed herein each have several aspects, no one of which is solely responsible for their desirable attributes. Without limiting the scope of the claims, we will now briefly discuss some prominent features. Numerous other embodiments are also contemplated, including embodiments having fewer, additional, and / or different components, steps, features, objects, benefits, and advantages. The components, aspects, and steps may be arranged and ordered differently. After considering this discussion, and especially after reading the section entitled "Detailed Description of the Invention," one will understand how the features of the devices and methods disclosed herein are advantageous over other known devices and methods.
[0007] Further details of exemplary nanopore sequencing devices that may be used with the disclosed technology, and methods of operating the devices, can be found in U.S. Provisional Patent Applications Nos. 63 / 200868 and 63 / 169041, the disclosures of each of which are incorporated by reference in their entireties.
[0008] Disclosed herein is a compound having the following structure: [ka] (Wherein, X is -O-, -CH 2 -, -NH- [ka] and X' is =N-SO 2 -, =NH-CO-, or [ka] Y is -O-, -S-, -NH-, or -Se-; L 1 is a first linking group, and L 2 is a second linking group, and SP is a spacer. The present invention includes compounds having one of the following structures:
[0009] Also disclosed herein is a compound having the structure: [ka] [ka] (Wherein, X is -O-, -CH 2 -, -NH- [ka] and X' is =N-SO 2 -, =NH-CO-, or [ka] Y is -O-, -S-, -NH-, or -Se-; R 1 , R 2 , and R 3 One of them is allyl and the other is H, and L 1 is a first linking group, L 2 is a second linking group, and SP is a spacer. The oligonucleotide includes an oligonucleotide having one of the following structures:
[0010] In some embodiments, structure (VII) has the following structure: [ka] This can be further expressed by:
[0011] In some embodiments, structure (VIII) has the following structure: [ka] This can be further expressed by:
[0012] In some embodiments, structure (X) has the following structure: [ka] This can be further expressed by:
[0013] Also disclosed herein is a compound having the structure: [ka] (In the formula, Y is -O-, -S-, -NH-, or -Se-; Y 1 , Y 2 and Y 3 One of the groups is -S- or -Se-, and the other is -O- or -NH-. The oligonucleotide comprises one of:
[0014] In some embodiments, the SP is selected from the group consisting of the following moieties: (1) an alkyl chain having 5-50 carbons; (2) an oligonucleotide or modified oligonucleotide having 1-100 repeat units; (3) a polypeptide having 1-100 repeat units; (4) a hydrophilic polymer having 1-100 repeat units selected from the group consisting of polyethylene glycol, polyvinyl alcohol, polyacrylamide, polyvinylpyrrolidone, polystyrene sulfonate, and polyethyleneimine; and (5) a hydrophobic polymer having 1-100 repeat units selected from the group consisting of polylactic acid, polymethyl methacrylate, and polystyrene. Contains one or more of the following:
[0015] In some embodiments, the hydrophilic polymer is selected from the group consisting of polyethylene glycol, polyvinyl alcohol, polyacrylamide, polyvinylpyrrolidone, polystyrene sulfonate, and polyethyleneimine. In some embodiments, the hydrophobic polymer is selected from the group consisting of polylactic acid, polymethyl methacrylate, and polystyrene.
[0016] In some embodiments, L 1 and L 2 each independently comprises a conjugated moiety selected from the group consisting of amine-NHS ester, amine-imido ester, amine-pentafluorophenyl ester, amine-hydroxymethylphosphine, carboxyl-carbodiimide, thiol-maleimide, thiol-haloacetyl, thiol-pyridyl disulfide, thiol-thiosulfonate, thiol-vinyl sulfone, aldehyde-hydrazide, aldehyde-alkoxyamine, hydroxy-isocyanate, azido-alkyne, azido-phosphine, transcyclooctene-tetrazine, norbornene-tetrazine, azido-cyclooctyne, and azido-norbornene.
[0017] In some embodiments, L 1 and L 2 further comprises, independently, a first linker between the conjugate moiety and X / X', and a second linker between the conjugate moiety and SP.
[0018] In some embodiments, the first linker and the second linker are independently selected from the group consisting of a hydrophilic polymer, a hydrophobic polymer, an oligonucleotide, a peptide, a polypeptide, an aliphatic chain (C5-C50), and combinations thereof. In some embodiments, the hydrophilic polymer, the hydrophobic polymer, the oligonucleotide, and the polypeptide may each have 1-100 repeating units, 2-100 repeating units, 5-100 repeating units, 10-100 repeating units, 2-50 repeating units, 2-30 repeating units, or any range between 1-100. The hydrophilic polymer may include polyethylene glycol, polyvinyl alcohol, polyacrylamide, polyvinylpyrrolidone, polystyrene sulfonate, polyethyleneimine, or combinations thereof. The hydrophobic polymer may include polylactic acid, polymethyl methacrylate, or polystyrene, or combinations thereof. In some embodiments, the aliphatic chain may include an alkyl, an alkenyl, an alkynyl, or combinations thereof.
[0019] In some embodiments, the SP further comprises a stop construct. In some embodiments, the base further comprises a stop construct. In some embodiments, the stop construct is a linear, branched, or cyclic polymer. In some embodiments, the stop construct comprises a synthetic hydrophobic polymer, a synthetic hydrophilic polymer, an oligonucleotide / polynucleotide, a peptide / polypeptide, or a combination thereof.
[0020] In some embodiments, L 1 , SP, and L 2 is a subelement of the cyclic loop. In some embodiments, the cyclic loop is symmetric. In some embodiments, the cyclic loop is asymmetric. In some embodiments, the cyclic loop s was synthesized using one or more of the following: solid phase synthesis, solution phase synthesis, and enzymatic synthesis. In some embodiments, the cyclic loop was synthesized using one or more of the following: linear synthesis, branched synthesis, or segmented synthesis.
[0021] Also disclosed herein is a method for determining a sequence of a polynucleotide in a nanopore-based sequencing system, the method comprising: providing a polynucleotide comprising a plurality of nucleotides, where each nucleotide comprises a linker construct, the linker construct having a first end attached to a first position of the nucleotide and a second end attached to a second position of the nucleotide; cleaving a severable bond on each of the plurality of nucleotides between the first position and the second position, thereby extending the polynucleotide to form an extended polymer; applying a voltage to cause the extended polymer to insert into and translocate through a nanopore; and (i) detecting and identifying a reporter moiety as the linker construct passes through the nanopore, or (ii) detecting and identifying a base on the nucleotide as the nucleotide passes through the nanopore.
[0022] In some embodiments, the linker construct comprises a first linking group, a second linking group, and a spacer between the first linking group and the second linking group.
[0023] In some embodiments, the spacer comprises an oligonucleotide, a modified oligonucleotide, or a polyphosphate having 1-100 repeating units, a polypeptide having 1-100 repeating units, an alkyl chain having 5-50 carbons, a hydrophilic polymer having 1-100 repeating units selected from the group consisting of polyethylene glycol, polyvinyl alcohol, polyacrylamide, polyvinylpyrrolidone, polystyrene sulfonate, and polyethyleneimine, a hydrophobic polymer having 1-100 repeating units selected from the group consisting of polylactic acid, polymethyl methacrylate, and polystyrene, and combinations thereof.
[0024] In some embodiments, the spacer comprises a reporter moiety, where the reporter moiety corresponds to and identifies a nucleotide.
[0025] In some embodiments, the first and second linking groups L 1and L 2 each independently comprises a conjugated moiety selected from the group consisting of amine-NHS ester, amine-imido ester, amine-pentafluorophenyl ester, amine-hydroxymethylphosphine, carboxyl-carbodiimide, thiol-maleimide, thiol-haloacetyl, thiol-pyridyl disulfide, thiol-thiosulfonate, thiol-vinyl sulfone, aldehyde-hydrazide, aldehyde-alkoxyamine, hydroxy-isocyanate, azido-alkyne, azido-phosphine, transcyclooctene-tetrazine, norbornene-tetrazine, azido-cyclooctyne, and azido-norbornene.
[0026] In some embodiments, the extending polymer further comprises a stop construct attached to each nucleobase or each linker construct, where the stop construct is configured to slow down, pause, or stop migration.
[0027] In some embodiments, the stopping construct (i.e., modification) is a linear, branched, or cyclic polymer. In some embodiments, the stopping construct comprises a synthetic hydrophobic polymer, a synthetic hydrophilic polymer, an oligonucleotide / polynucleotide, a peptide / polypeptide, or a combination thereof. In some embodiments, the nanopore comprises a constriction having an opening with an inner diameter of about 0.6 nm to about 1.2 nm.
[0028] In some embodiments, the reporter moiety comprises one or more sub-reporter moieties, where the one or more sub-reporter moieties correspond to and identify a nucleotide or a transfer event. In some embodiments, the reporter moiety comprises two or more sub-reporter moieties, where each sub-reporter moiety in the two or more sub-reporter moieties is distinguishable, reproducible, and resolvable. In some embodiments, the two or more sub-reporter moieties comprise a crown ether, a cucurbituril, a pillararene, or a cyclodextrin.
[0029] Disclosed herein includes a kit for carrying out a method of sequencing a polynucleotide in a nanopore-based sequencing system, the kit comprising a compound disclosed herein.
[0030] Disclosed herein are systems for determining the sequence of a polynucleotide, including systems configured to carry out the methods disclosed herein.
[0031] Disclosed herein includes a system for performing a method for determining the sequence of a polynucleotide comprising a plurality of nucleotides, wherein the nucleotides are selected from any of the compounds disclosed herein.
[0032] It is to be understood that all combinations of the foregoing concepts and additional concepts discussed in more detail below are considered to be part of the inventive subject matter disclosed herein and may be used to realize the benefits and advantages described herein. [Brief description of the drawings]
[0033] Features of examples of the present disclosure will become apparent upon reference to the following detailed description and the drawings in which like reference numbers correspond to similar, but possibly not identical, components. For purposes of brevity, reference numbers or features having previously described functions may or may not be described in conjunction with the other drawings in which they appear.
[0034] [Figure 1] 1 shows a schematic diagram of an example of sequencing an extended polynucleotide.
[0035] [Diagram 2] 1 shows a schematic of an example of polynucleotide extension with a cleavable circular loop nucleotide.
[0036] [Diagram 3] Examples of cleavable bonds are shown diagrammatically.
[0037] [Figure 4] 1 shows a schematic of an example of the use of allyl-cleavable chemistry in the disclosed sequencing methods.
[0038] [Diagram 5] Schematic showing an example of an allyl-based cleavable circular loop nucleotide with 10 Ts as the barcode region, a phosphate-base linked (PBL) circular loop, and "modifications" on the bases.
[0039] [Figure 6] 1 shows a schematic example of sequencing an extended polynucleotide having modifications.
[0040] [Figure 7] 1 shows experimental data for polymerase incorporation of allyl-dTTP.
[0041] [Figure 8] 1 shows experimental data for cleavage of allyl groups.
[0042] [Figure 9] Schematic diagram of the first design example of a fully functional circular loop nucleotide (CLN-1).
[0043] [Figure 10] 1 shows an exemplary synthesis process for a fully functional cyclic loop nucleotide (CLN-1A).
[0044] [Figure 11] 1 shows an exemplary synthesis process for a fully functional cyclic loop nucleotide (CLN-1B).
[0045] [Figure 12] 1 shows a schematic of an exemplary synthesis process for a fully functional cyclic loop nucleotide (CLN-1C).
[0046] [Figure 13] 1 shows an exemplary synthesis process for a fully functional cyclic loop nucleotide (CLN-1D).
[0047] [Figure 14] Schematic diagram of a second example of the design of a fully functional circular loop nucleotide (CLN-2).
[0048] [Figure 15] 1 shows an exemplary synthesis process for a fully functional cyclic loop nucleotide (CLN-2A).
[0049] [Figure 16] 1 shows an exemplary synthesis process for a fully functional circular loop nucleotide (CLN-2B).
[0050] [Figure 17] 1 shows an exemplary synthesis process for a fully functional cyclic loop nucleotide (CLN-2C).
[0051] [Figure 18] Schematic diagram of a third example design of a fully functional cyclic loop nucleotide (CLN-3), where X can be O, NH, NSO2 or CH2, and Y can be O, S or NH, with an alpha phosphate-allyl linked asymmetric loop and a stop construct on the nucleobase.
[0052] [Figure 19] 1 shows an exemplary synthesis process for a fully functional cyclic loop nucleotide (CLN-3A).
[0053] [Figure 20] 1 shows an exemplary synthesis process for a fully functional cyclic loop nucleotide (CLN-3B).
[0054] [Figure 21]1 shows an exemplary synthesis process for a fully functional cyclic loop nucleotide (CLN-3C).
[0055] [Figure 22] 1 shows a schematic of a fourth example design of a fully functional cyclic loop nucleotide (CLN-4), where X can be O, NH, NSO2 or CH2, Y can be O, S or NH, with an alpha phosphate-allyl linked asymmetric loop and a stop construct on the loop.
[0056] [Figure 23] 1 shows an exemplary synthesis process for a fully functional cyclic loop nucleotide (CLN-4A).
[0057] [Figure 24] 1 shows an exemplary synthesis process for a fully functional cyclic loop nucleotide (CLN-4B).
[0058] [Diagram 25] 1 shows an exemplary synthesis process for a fully functional cyclic loop nucleotide (CLN-4C).
[0059] [Figure 26] 1 shows a schematic of a fifth example design of a fully functional cyclic loop nucleotide (CLN-5), where X can be O, NH, NSO2 or CH2, Y can be O, S or NH, with an alpha phosphate linked asymmetric loop and a stop construct on the loop.
[0060] [Figure 27] 1 shows an exemplary synthesis process for a fully functional cyclic loop nucleotide (CLN-5A).
[0061] [Figure 28] 1 shows an exemplary synthesis process for a fully functional cyclic loop nucleotide (CLN-5B).
[0062] [Figure 29] 1 shows an exemplary synthesis process for a fully functional cyclic loop nucleotide (CLN-5C).
[0063] [Diagram 30] 1 shows an exemplary synthesis process for a fully functional cyclic loop nucleotide (CLN-5D).
[0064] [Diagram 31] 1 shows a schematic diagram of a sixth design example of a cyclic loop nucleotide (CLN-6), where X can be O, NH, NSO2 or CH2, and Y can be O, S or NH, with a peptide-based cyclic loop and a stop construct on the nucleobase.
[0065] [Diagram 32] 13A-13C are schematic diagrams illustrating an example of a seventh design of a cyclic loop nucleotide (SNL-1), where X can be O, NH, NSO2, or CH2, with an alpha phosphate linked symmetric nucleoside loop and a stop construct on the loop.
[0066] [Diagram 33] 13A-13C are schematic diagrams illustrating an example of an eighth design of a cyclic loop nucleotide (SNNL-1), where X can be O, NH, or CH2, with an alpha phosphate-linked symmetric non-nucleoside loop and a stop construct on the loop.
[0067] [Diagram 34] 13A-13C are schematic diagrams illustrating a ninth example design of a cyclic loop nucleotide (SPL-1), where X can be O, NH, NSO2, or CH2, with an alpha phosphate-linked symmetric peptide loop and a stop construct on the loop.
[0068] [Diagram 35] 13A-13C are schematic diagrams illustrating a tenth example of a cyclic loop nucleotide (SNL-2) design, where X can be O, NH, NSO2, or CH2, with an alpha phosphate-allyl linked symmetric nucleoside loop and a stop construct on the loop.
[0069] [Diagram 36] 13 is a schematic diagram of an example of an eleventh design of a cyclic loop nucleotide (SNNL-2), where X can be O, NSO2, or CH2, with an alpha phosphate-allyl linked symmetric non-nucleoside loop and a stop construct on the loop.
[0070] [Figure 37] 13A-13C are schematic diagrams illustrating an example of a twelfth design of a cyclic loop nucleotide (SPL-2), where X can be O, NH, NSO2, or CH2, with an alpha phosphate-allyl linked symmetric peptide loop and a stop construct on the loop.
[0071] [Figure 38] 13 shows a schematic diagram of a thirteenth design example of a cyclic loop nucleotide (ASL-1), where X can be O, NH, NSO2, or CH2, with an alpha phosphate linked asymmetric loop and a stop construct on the loop.
[0072] [Figure 39] 13 is a schematic diagram of a fourteenth example design of a cyclic loop nucleotide (ASL-2), where X can be O, NH, NSO2, or CH2, with an alpha phosphate-C5' allyl linked asymmetric loop and a stop construct on the loop.
[0073] [Diagram 40] 13A-13C are schematic diagrams illustrating an example of a fifteenth design of a cyclic loop nucleotide (ASL-3), where X can be O, NH, NSO2, or CH2, and has an alpha phosphate-C5' allyl linked asymmetric loop and a stop construct on the nucleobase.
[0074] [Diagram 41] FIG. 16 shows a schematic diagram of a sixteenth design example of a circular loop tetramer oligonucleotide (tetramer-SNL-1) having a tetramer oligo with a symmetric nucleoside loop and a stop construct on the loop.
[0075] [Diagram 42]FIG. 17 shows a schematic diagram of a seventeenth design example of a circular loop tetramer oligonucleotide (tetramer-SNNL-1) having a tetramer oligo with a symmetric non-nucleoside loop and a stop construct on the loop.
[0076] [Diagram 43] FIG. 18 shows a schematic diagram of an eighteenth design example of a circular loop tetramer oligonucleotide (tetramer-SPL-1) having a tetramer oligo with a symmetric peptide loop and a stop construct on the loop.
[0077] [Diagram 44] FIG. 19 shows a schematic diagram of a 19th design example of a circular loop tetramer oligonucleotide (tetramer-ASL-1) having a tetramer oligo with an asymmetric loop and a stop construct on the loop.
[0078] [Diagram 45] 1 shows a schematic of an exemplary synthetic process for generating a tetrameric oligonucleotide that can be used to create a cleavable circular loop tetrameric oligonucleotide.
[0079] [Figure 46] Schematic diagram of an example chemical scheme (Scheme I) for the synthesis of a bridged PS nucleotide, the product of which is Nucleotide 1.
[0080] [Figure 47A] 1 shows a non-limiting characterization of phosphorothiolate (bridged-PS) nucleotides made according to Scheme I.
[0081] [Figure 47B] 1 shows a non-limiting characterization of phosphorothiolate (bridged-PS) nucleotides made according to Scheme I.
[0082] [Figure 48] An alternative method for generating bridged PS nucleotides is outlined in Scheme II.
[0083] [Figure 49] FIG. 1 shows experimental data from the incorporation of a polyA tail formed by cross-linking PS (phosphorothiolate) deoxyribonucleotides.
[0084] [Figure 50A] 1 is a reaction scheme for cleavage of phosphorus-sulfur bonds by silver nitrate via bridged PS nucleotides.
[0085] [Figure 50B] 1 shows experimental results obtained from cleavage product analysis after cleavage of phosphorus-sulfur bonds by silver nitrate via cross-linked PS nucleotides.
[0086] [Figure 51] 1 shows the gel results of a cleavage reaction with cross-linked PS nucleotide 1 and extended primer using DNA polymerase Pol1901.
[0087] [Figure 52] 1 shows a schematic diagram of the synthesis of 5'SDMT phosphoramidite "9" by multiple methods.
[0088] [Figure 53A] 1 shows one embodiment of attaching a spacer moiety to a tetramer oligonucleotide.
[0089] [Figure 53B] 1 shows one embodiment of a circular loop tetramer oligonucleotide.
[0090] [Figure 54A] FIG. 53B shows HPLC, IR and MS characterization of the circular loop tetramer oligonucleotide shown in FIG. 53B. [Figure 54B] FIG. 53B shows HPLC, IR and MS characterization of the circular loop tetramer oligonucleotide shown in FIG. 53B. [Fig. 54C] FIG. 53B shows HPLC, IR and MS characterization of the circular loop tetramer oligonucleotide shown in FIG. 53B.
[0091] [Figure 55] Non-limiting experimental ligation results for ligating up to 10 looped tetrameric oligonucleotides are shown.
[0092] [Figure 56] 1 shows non-limiting experimental results from sequential ligation and cleavage of specific polynucleotides.
[0093] [Figure 57A] 1 is a reaction scheme for the synthesis of one embodiment of a non-bridged PS nucleotide.
[0094] [Figure 57B] 1 is a reaction scheme for the synthesis of another embodiment of a non-bridged PS nucleotide.
[0095] [Figure 58] 1 shows the incorporation of non-bridged PS nucleotides using various polymerases.
[0096] [Figure 59] 1 shows a schematic diagram of the attachment of a spacer moiety to a functionalized tetrameric oligonucleotide to form a circular loop tetrameric oligonucleotide.
[0097] [Figure 60A] Schematic diagram of the characterization of the 4-mer oligonucleotide shown in FIG. [Figure 60B] Schematic diagram of the characterization of the 4-mer oligonucleotide shown in FIG. [Figure 60C] Schematic diagram of the characterization of the 4-mer oligonucleotide shown in FIG. [Figure 60D] Schematic diagram of the characterization of the 4-mer oligonucleotide shown in FIG.
[0098] [Figure 61]A ligase is shown that performs up to 10 sequential ligation events using the circular loop tetramer oligonucleotide shown in FIG.
[0099] [Figure 62A] 1 illustrates the use of iodine to cleave the bridging PO bond at the phosphonothioate site.
[0100] [Figure 62B] Characterization before and after PO bond cleavage is shown.
[0101] [Figure 63] 1 is a reaction scheme for the synthesis of imino-P substituted nucleotides.
[0102] [Figure 64] 1 is a reaction scheme for the synthesis of cyclic loop nucleotides having imino-p substitutions according to an embodiment.
[0103] [Figure 65] FIG. 1 shows HPLC and LCMS characterization of bifunctional imino-P nucleotides.
[0104] [Figure 66] 1 shows a schematic of one embodiment of a method for synthesizing an imino-P allyl bifunctional nucleotide.
[0105] [Figure 67] 1 shows a schematic diagram of another embodiment of a method for synthesizing an imino-P allyl bifunctional nucleotide.
[0106] [Figure 68] 1 shows a schematic diagram of another embodiment of a method for synthesizing an imino-P allyl bifunctional nucleotide.
[0107] [Figure 69] 1 shows a schematic diagram of another embodiment of a method for synthesizing an imino-P allyl bifunctional nucleotide.
[0108] [Figure 70A] 1 shows a schematic of one embodiment of a method for synthesizing an imino-P allyl bifunctional nucleotide.
[0109] [Figure 70B] Exemplary methods for activating αP monophosphate are provided.
[0110] [Figure 71] 1 shows an exemplary synthesis of bifunctional imino-P allyl nucleotides with different reactive groups.
[0111] [Figure 72] Schematic representation of deprotection of the 3'-OTBDPS group to form the 3'-OH, generating nucleotide 6.
[0112] [Figure 73] HF-TEA and TBAF deprotection methods are presented to assess conversion efficiency and yield to the desired triphosphate product.
[0113] [Fig. 74A] Crude HPLC, analytical HPLC, and LCMS spectra of purified nucleotides are shown. [Fig. 74B] Crude HPLC, analytical HPLC, and LCMS spectra of purified nucleotides are shown. [Fig. 74C] Crude HPLC, analytical HPLC, and LCMS spectra of purified nucleotides are shown.
[0114] [Figure 75] The stereoselective reduction of a carbonyl group is shown diagrammatically.
[0115] [Figure 76] 1 shows a schematic diagram of an enzymatic kinetic diastereoselection process from a racemic precursor.
[0116] [Figure 77]The chiral derivatization of the isomers that allows for final column separation is shown diagrammatically.
[0117] [Figure 78] 1 shows a schematic of chiral ligand promoted stereoselective alkyl addition.
[0118] [Figure 79] 1 shows a schematic of the stereoselective enzymatic synthesis of triphosphates.
[0119] [Figure 80] 1 shows a schematic diagram of a method for controlling chirality at an alpha phosphorus atom.
[0120] [Figure 81] 1 shows the synthesis of an exemplary Staudinger mutant.
[0121] [Figure 82] The synthesis of an exemplary azide mutant is shown.
[0122] [Figure 83A] 1 shows a schematic representation of an embodiment of a pathway leading to the formation of an exemplary circular loop nucleotide structure. [Figure 83B] 1 shows a schematic representation of an embodiment of a pathway leading to the formation of an exemplary circular loop nucleotide structure.
[0123] [Fig. 84A] HPLC, LCMS, and FTIR characterization of the cyclic loop structure of Figure 83A is shown. [Fig. 84B] HPLC, LCMS, and FTIR characterization of the cyclic loop structure of Figure 83A is shown. [Fig. 84C] HPLC, LCMS, and FTIR characterization of the cyclic loop structure of Figure 83A is shown.
[0124] [Figure 85]1 shows the results of bifunctional nucleotide 6 tested in an incorporation assay with Dpo4 enzyme compared to native dTTP.
[0125] [Figure 86] 1 shows the results of bifunctional nucleotide 10 tested in an incorporation assay with Dpo4 enzyme compared to native dTTP.
[0126] [Figure 87] 1 shows experimental results on the kinetics of incorporation of looped nucleotides.
[0127] [Figure 88] 1 shows a stability assay performed on bifunctional nucleotide 6 and looped nucleotide 10.
[0128] [Figure 89] 1 shows a stability assay performed on bifunctional nucleotide 6 and looped nucleotide 10.
[0129] [Figure 90] 1 shows an exemplary synthesis of a spacer portion with a stop construct.
[0130] [Figure 91] 13 shows another exemplary synthesis of a spacer portion with a stop construct.
[0131] [Figure 92] 13 shows another exemplary synthesis of a spacer portion with a stop construct.
[0132] [Figure 93] 1 shows an asymmetric circular loop.
[0133] [Figure 94] 1 shows a symmetric circular loop. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0134] All patents, applications, published applications, and other publications mentioned herein are incorporated herein by reference for the material referenced and in their entirety. Where a term or phrase is used herein in a manner that is contrary to or otherwise inconsistent with a definition set forth in a patent, application, published application, or other publication incorporated herein by reference, the usage herein takes precedence over the definition incorporated herein by reference.
[0135] definition All technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs, unless defined otherwise.
[0136] As used herein, the singular forms "a," "and," and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to "an array" can include a plurality of such arrays, and so forth.
[0137] The terms comprising, including, containing, and various forms of these terms are synonymous and intended to be equally broad. Furthermore, unless expressly stated otherwise, an example comprising, including, or having an element or elements having a particular characteristic can include the additional elements, regardless of whether the additional elements have that characteristic.
[0138] As used herein, the term "modified oligonucleotide" refers to a polymeric chain of nucleobases or nucleotides assembled from moieties that contain modified nucleobases, modified sugar rings (e.g., LNA, constrained ethyl, ethylene bridge, TNA, 2'-Ome, 2'F, 2'-MOE), or nucleobases attached to a scaffold (e.g., unlocked, 4'-thio, CeNA, HNA, TNA, GNA, FNA).
[0139] As used herein, the term "phosphoramidite analog" refers to any polymer synthesized using phosphoramidite or related chemistry that results in the formation of phosphodiester, methylphosphonate, or phosphorothioate linkages between moieties.
[0140] As used herein, the term "modified polyamide" refers to a polymer assembled from individual moieties each having at least one amino group and one carboxylic acid group, resulting in the formation of an amide bond.
[0141] As used herein, the term "nanopore" is intended to mean a hollow structure that is distinct from or defined within a membrane and extends across the membrane. A nanopore allows ions, current, and / or fluids to cross from one side of the membrane to the other side of the membrane. For example, a membrane that inhibits the passage of ions or water-soluble molecules can include a nanopore structure that extends across the membrane to allow the passage of ions or water-soluble molecules from one side of the membrane to the other side of the membrane (through a nanoscale opening that extends through the nanopore structure). The diameter of the nanoscale opening that extends through the nanopore structure can vary along its length (i.e., from one side of the membrane to the other side of the membrane), but is at any point in the nanoscale (i.e., from about 1 nm to about 100 nm, or less than 1000 nm). Examples of nanopores include, for example, biological nanopores, solid-state nanopores, and biological and solid-state hybrid nanopores. In some embodiments, a refers to a pore having an opening with a diameter of about 0.3 nm to about 2 nm at its narrowest point. For example, the nanopore may be a solid nanopore, a graphene nanopore, an elastomeric nanopore, or a natural or recombinant protein that forms a tunnel upon insertion into a bilayer, thin film, membrane, or solid opening, also referred to herein as a protein pore or protein nanopore (e.g., a transmembrane pore). If the protein is inserted into a membrane, the protein is a tunnel-forming protein.
[0142] As used herein, the term "diameter" is intended to mean the longest straight line that can be inscribed in the cross-section of a nanoscale aperture that passes through the center of mass of the cross-section of the nanoscale aperture. It is understood that a nanoscale aperture may or may not have a circular or substantially circular cross-section (a cross-section of a nanoscale aperture that is substantially parallel to the cis / trans electrodes). Furthermore, the cross-section may be of a regular or irregular shape.
[0143] As used herein, "cis" refers to the side of the nanopore opening on which the analyte or modified analyte enters the opening or across the face on which the analyte or modified analyte travels.
[0144] As used herein, "trans" refers to the side of the nanopore opening through which the analyte or modified analyte (or a fragment thereof) exits the opening or across which the analyte or modified analyte does not move.
[0145] As used herein, the term "biological nanopore" is intended to mean a nanopore whose structural portion is made from materials of biological origin. Biological origin refers to materials derived or isolated from a biological environment, such as an organism or cell, or synthetically produced versions of biologically available structures. Biological nanopores include, for example, polypeptide nanopores and polynucleotide nanopores.
[0146] As used herein, a "moiety" is one of two or more parts into which something can be divided (eg, various portions of a tether, molecule, or probe).
[0147] As used herein, a "reporter" is composed of one or more reporter elements or reporter moieties. Reporters include those known as "tags" and "labels." A linker construct (if it contains a reporter moiety) or a nucleobase residue of an extended polymer can be considered a reporter. A reporter serves to resolve the identity of a target nucleic acid. A reporter may include constituent sub-reporters, and multiple reporters may be present on a single nucleotide. When present in the readhead of a nanopore, a reporter provides a characteristic, and sometimes unique, blocking current at a given read voltage.
[0148] As used herein, a "linker" is a molecule or moiety that connects two molecules or moieties and provides spacing between the two molecules or moieties so that they can function in the manner intended. For example, a linker can include a diamine hydrocarbon chain that is covalently attached to an oligonucleotide analog molecule through a reactive group at one end and to a solid support (such as a bead surface) through a reactive group at the other end. Attachment of the linker to the nucleotide and substrate construct of interest can be accomplished through the use of coupling reagents known in the art (see, for example, Efimov et al., Nucleic Acids Res. 27:4416-4426, 1999). Methods for derivatizing and attaching organic molecules are well known in the art of organic chemistry and bioorganic chemistry. The linker can also be cleavable or reversible.
[0149] As used herein, the term "heavy atom" refers to any atom used in a molecular structure that is not hydrogen. Heavy atoms used in modified oligonucleotides may be bridging (e.g., used to connect multiple oligonucleotides) or non-bridging (e.g., not directly linked to multiple oligonucleotides).
[0150] As used herein, the term "polypeptide nanopore" is intended to mean a protein / polypeptide that extends across a membrane and allows ions, electrical current, polymers such as DNA or peptides, or other molecules of appropriate size and charge, and / or fluids to flow from one side of the membrane to the other side of the membrane. Polypeptide nanopores can be monomeric, homopolymeric, or heteropolymeric. Structures of polypeptide nanopores include, for example, α-helical bundle nanopores and β-barrel nanopores. Examples of polypeptide nanopores include α-hemolysin, Mycobacterium smegmatis porin A (MspA), gramidiin A, maltoporin, OmpF, OmpC, PhoE, Tsx, F fimbria, and the like. The protein α-hemolysin is naturally found in cell membranes and functions as a pore for ions or molecules to be transported in and out of the cell. Mycobacterium smegmatis porin A (MspA) is a membrane porin produced by mycobacteria that allows hydrophilic molecules to enter the bacteria. MspA forms a tightly interconnected octamer and transmembrane beta-barrel that resembles a goblet and contains a central pore.
[0151] As used herein, "peptide" refers to two or more amino acids linked together by amide bonds (i.e., "peptide bonds"). A peptide contains up to 50 amino acids. A peptide may be linear or cyclic. A peptide may be alpha, beta, gamma, delta, or higher, or may be mixed. A peptide may contain any mixture of amino acids as defined herein, such as including any combination of D, L, alpha, beta, gamma, delta, or higher order amino acids.
[0152] As used herein, "protein" refers to an amino acid sequence having 51 or more amino acids.
[0153] The polypeptide nanopore can be synthetic. The synthetic polypeptide nanopore comprises a protein-like amino acid sequence that does not occur in nature. The protein-like amino acid sequence may comprise some of the amino acids that are known to exist but do not form the basis of a protein (i.e., non-proteinogenic amino acids). The protein-like amino acid sequence can be artificially synthesized rather than expressed in an organism and then purified / isolated.
[0154] The nanopores disclosed herein may be hybrid nanopores. "Hybrid nanopore" refers to a nanopore that contains materials of both biological and non-biological origin. Examples of hybrid nanopores include polypeptide solid-state hybrid nanopores and polynucleotide solid-state nanopores.
[0155] Application of a potential difference across the nanopore can force the movement of the nucleic acid through the nanopore. One or more signals are generated corresponding to the movement of the nucleotide through the nanopore. Thus, when a target polynucleotide, or a mononucleotide or a probe derived from a target polynucleotide or mononucleotide, passes through the nanopore, the current across the membrane changes, for example, due to a base-dependent (or probe-dependent) blockage of the constriction. The signal from that change in current can be measured using any of a variety of methods. Each signal is specific to the species of nucleotide (or linker construct with reporter moiety) in the nanopore, such that the resulting signal can be used to determine a characteristic of the polynucleotide. For example, the identity of one or more species of nucleotide (or probe) that produces a characteristic signal can be determined.
[0156] As used herein, a "nucleotide" comprises a nitrogen-containing heterocyclic base, a sugar, and one or more phosphate groups. A nucleotide is a monomeric unit of a nucleic acid sequence. Examples of nucleotides include, for example, ribonucleotides or deoxyribonucleotides. In ribonucleotides (RNA), the sugar is ribose, and in deoxyribonucleotides (DNA), the sugar is deoxyribose, i.e., a sugar lacking the hydroxyl group present at the 2' position of the ribose. The nitrogen-containing heterocyclic base can be a purine base or a pyrimidine base. Purine bases include adenine (A) and guanine (G), as well as modified derivatives or analogs thereof. Pyrimidine bases include cytosine (C), thymine (T), and uracil (U), as well as modified derivatives or analogs thereof. The C-1 atom of the deoxyribose is attached to the N-1 of the pyrimidine or the N-9 of the purine. The phosphate group can be in mono-, di-, or triphosphate form. These nucleotides are naturally occurring nucleotides, however it should be further understood that non-naturally occurring nucleotides, modified nucleotides or analogs of the aforementioned nucleotides can also be used.
[0157] As used herein, a "nucleobase" is a heterocyclic base, such as adenine, guanine, cytosine, thymine, uracil, inosine, xanthine, hypoxanthine, or a heterocyclic derivative, analog, or tautomer thereof. Nucleobases can be naturally occurring or synthetic. Non-limiting examples of nucleobases include adenine, guanine, thymine, cytosine, uracil, xanthine, hypoxanthine, 8-azapurine, purine substituted with methyl or bromine at the 8-position, 9-oxo-N6-methyladenine, 2-aminoadenine, 7-deazaxanthine, 7-deazaguanine, 7-deaza-adenine, N4-ethanocytosine, 2,6-diaminopurine, N6-ethano-2,6-diaminopurine, 5-methylcytosine, 5-(C3-C6)-alkynylcytosine, 5-fluorouracil, 5-bromouracil. , thiouracil, pseudoisocytosine, 2-hydroxy-5-methyl-4-triazolopyridine, isocytosine, isoguanine, inosine, 7,8-dimethylalloxazine, 6-dihydrothymine, 5,6-dihydrouracil, 4-methyl-indole, ethenoadenine, and the non-naturally occurring nucleobases described in U.S. Pat. Nos. 5,432,272 and 6,150,510, and WO 92 / 002258, WO 93 / 10820, WO 94 / 22892, and WO 94 / 24144, and in Fasman (Practical Handbook of Biochemistry and Molecular Biology, pp. 385-394, 1989, CRC Press, Boca Raton, LO), all of which are incorporated herein by reference in their entireties.
[0158] The term "nucleic acid" or "polynucleotide" refers to deoxyribonucleotide or ribonucleotide polymers in single- or double-stranded form, and includes known analogs of natural nucleotides that hybridize to nucleic acids in a manner similar to naturally occurring nucleotides, such as peptide nucleic acids (PNAs) and phosphorothioate DNA, unless otherwise specified. A particular nucleic acid sequence includes its complementary sequence, unless otherwise specified. Nucleotides include, but are not limited to, ATP, dATP, CTP, dCTP, GTP, dGTP, UTP, TTP, dUTP, 5-methyl-CTP, 5-methyl-dCTP, ITP, dITP, 2-amino-adenosine-TP, 2-amino-deoxyadenosine-TP, 2-thiothymidine triphosphate, pyrrolo-pyrimidine triphosphate, and 2-thiocytidine, as well as alpha thiotriphosphate for all of the above, and 2'-O-methyl-ribonucleotide triphosphate for all of the above bases. Modified bases include, but are not limited to, 5-Br-UTP, 5-Br-dUTP, 5-F-UTP, 5-F-dUTP, 5-propynyl dCTP, and 5-propynyl-dUTP.
[0159] As used herein, the term "signal" is intended to mean an indication that represents information. Signals include, for example, electrical signals and optical signals. The term "electrical signal" refers to an indication of an electrical quality that represents information. The indication can be, for example, current, voltage, tunneling, resistance, potential, voltage, conductance, or lateral electrical effect. "Electronic current" or "electric current" refers to the flow of charge. In an example, the electrical signal can be a current passing through a nanopore, and the current can flow when a potential difference is applied across the nanopore.
[0160] As used herein, the term "driving force" is intended to mean an electric current that allows a polynucleotide to translocate through a nanopore. In some embodiments, when a potential difference is applied across the nanopore, a current can flow.
[0161] As used herein, the term "retention force" is intended to mean a resistance that slows and / or stops a polynucleotide from moving through a nanopore. In some embodiments, the retention force is overcome by application of a driving force. Thus, the driving force overcomes / negates the resistance that slows and / or stops the polynucleotide, thereby allowing the polynucleotide to move through the nanopore.
[0162] The term "modification" as used herein is intended to mean a moiety attached to a nucleotide. The modification may provide resistance (in the form of a "holding force") that slows down and / or stops the polynucleotide from moving through the nanopore unless the resistance from the modification is overcome by a "driving force". The resistance provided by the modification is due to the properties of the modification (e.g., size, geometry, and / or non-covalent interactions with the nanopore). The modification may act as a ratchet or brake for the polypeptide movement through the nanopore. The modification may be attached to any part of the nucleotide, or may be attached to the nucleotide at two positions that form a loop. The modification may also be referred to as a stop construct.
[0163] The aspects and examples described and claimed herein can be understood in light of the above definitions.
[0164] overview A common drawback of nanopore sequencers is that the nanopore is sensitive to multiple bases of the DNA strand within the nanopore, as opposed to reading one base at a time. For example, the MspA nanopore has a constriction region that serves as a read head of at least four nucleotides (called a "k-mer"), resulting in a minimum of 256 (4^4) different permutations of the tetrameric sequence that need to be deconvoluted. For a 5-base k-mer, the number of possible signals is 4^5=1024. A longer read head results in an exponential increase in the number of signals to be differentiated, which complicates the sequencing readout and increases the complexity of the base calling, thus reducing the accuracy. Another problem with nanopore sequencers is that the migration speed of native single-stranded DNA is on the order of >10 million nucleotides per second, far beyond the speed that can be accommodated by the electronics and detectors.
[0165] In some embodiments, by using cleavable sites along the DNA backbone while connecting adjacent nucleobases with the barcode region, the disclosed technology allows for increasing the distance between adjacent nucleobases, eliminating the need to deconvolute multiple signals. Once the backbone is cleaved, the reporter moiety of the extended polynucleotide occupies the entire nanopore read head for high-precision single molecule sequencing at single base resolution. It has been observed that DNA backbone cleavage can be affected by instability resulting from the introduction of certain alpha phosphate substitutions. Thus, stable alpha phosphate substitutions are preferred when synthesizing daughter oligonucleotide strands, especially when the polynucleotide or oligonucleotide is cleaved to extend the leading strand.
[0166] In some embodiments, the disclosed technology allows for having one nucleobase of an extended polynucleotide present at any one time in the read head, successfully reducing the read diversity to four (A, T, C, and G), allowing for lower cost, more accurate sequencing. In some embodiments, the disclosed technology provides high throughput, cheaper, and more accurate DNA sequencing.
[0167] Systems and methods FIG. 1 shows a schematic example of sequencing an extended polynucleotide. A protein nanopore 101 is deposited in a lipid bilayer 102. An extended polynucleotide 103 translocates through the nanopore 101. The polynucleotide 103 includes linker construct regions between consecutive nucleotides. By introducing linker constructs between consecutive nucleotides, the k-mer length can be reduced to 1, resulting in only four signals (for A, T, C and G), reducing the complexity of base calling. A distinctive linker / barcode can be assigned to each of the four individual bases to achieve base recognition. For example, a signal unit or portion 105 includes an "A" nucleotide and a corresponding "linker construct 1", which may include a reporter that serves as a barcode for nucleotide A. The read multiplicity is reduced to 4 using a single barcode characteristic for each nucleic acid base present in the nanopore readhead.
[0168] By "translocation" it is meant that an analyte (e.g., a polynucleotide such as DNA) enters one side of the opening of the nanopore and moves out to the other side of the opening. It is contemplated that any embodiment herein that includes translocation may refer to electrophoretic or non-electrophoretic translocation, unless otherwise specified. An electric field may translocate the analyte or modified analyte. By "interacting" it is meant that the analyte or modified analyte moves into the opening and optionally through the opening, and by "passing the opening" (or "translocating") it is meant that the analyte enters one side of the opening and moves out to the other side of the opening. Optionally, methods that do not use electrophoretic translocation are contemplated. In some embodiments, physical pressure causes the modified analyte to interact with, enter, or (after modification) through the opening. In some embodiments, a magnetic bead is coupled to the analyte or modified analyte on the trans side, and a magnetic force causes the modified analyte to interact with, enter, or (after modification) through the opening. Other methods for movement include, but are not limited to, other physical forces such as gravity, osmotic forces, temperature, and centripetal forces.
[0169] In some embodiments, the nanopore may comprise a solid material such as silicon nitride, modified silicon nitride, silicon, silicon oxide, or graphene, or a combination thereof. In some embodiments, the nanopore is a protein that forms a tunnel upon insertion into a bilayer, membrane, thin film, or solid aperture. In some embodiments, the nanopore is comprised in a lipid bilayer. In some embodiments, the nanopore is comprised in an artificial membrane that comprises mycolic acid. The nanopore may be a Mycobacterium smegmatis porin (Msp) having a vestibule and a constriction zone that define a tunnel. The Msp porin may be a mutant MspA porin. In some embodiments, the amino acids at positions 90, 91, and 93 of the mutant MspA porin are each substituted with asparagine. Some embodiments may include altering the translocation rate or sequencing sensitivity by removing, adding, or substituting at least one amino acid of the Msp porin. A "mutant MspA porin" is a multimeric complex having at least or at most 70, 75, 80, 85, 90, 95, 98, or 99 percent or more identity, or any range derivable therein, but less than 100% identity, to its corresponding wild-type MspA porin and retaining tunnel-forming ability. The mutant MspA porin may be a recombinant protein. Optionally, the mutant MspA porin has a mutation in the constriction zone or vestibule of wild-type MspA porin. Optionally, the mutation may occur at the edge or outside of the periplasmic loop of wild-type MspA porin. The mutant MspA porin may be used in any of the embodiments described herein.
[0170] "Vestibule" refers to the inner conical portion of the Msp porin, the diameter of which generally decreases from one end to the other along the central axis, with the narrowest portion of the vestibule being connected to the constriction zone. The vestibule may be referred to as a "goblet." The vestibule and the constriction zone together define the tunnel of the Msp porin. "Constriction zone" or "read head" refers to the narrowest portion of the tunnel of the Msp porin, in terms of diameter, that is connected to the vestibule. The length of the constriction zone may range from about 0.3 nm to about 2 nm. Optionally, the length is about, at most about, or at least about 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0, 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 2, or 3 nm, or any range derivable therein. The diameter of the constriction zone may range from about 0.3 nm to about 2 nm. Optionally, the diameter is about, at most about, or at least about 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0, 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 2, or 3 nm, or any range derivable therein. "Tunnel" refers to the central, empty portion of the Msp porin defined by the vestibule and the constriction zone through which gases, liquids, ions, or analytes may pass. A tunnel is an example of an opening in a nanopore.
[0171] Various conditions, such as light and the liquid medium in contact with the nanopore (including its pH, buffer composition, detergent composition, and temperature), can temporarily or permanently affect the behavior of the nanopore, particularly with respect to its conductance through the tunnel and the movement of analytes relative to the tunnel.
[0172] In some embodiments, the disclosed system for nanopore sequencing includes an Msp porin having a vestibule and a constriction zone defining a tunnel, where the tunnel is located between a first liquid medium and a second liquid medium, at least one liquid medium includes an analyte polynucleotide, and the system is operative to detect a characteristic of the analyte. The system may be operative to detect any characteristic of the analyte including subjecting the Msp porin to an electric field such that the analyte interacts with the Msp porin. The system may be operative to detect any characteristic of the analyte including subjecting the Msp porin to an electric field such that the analyte electrophoretically migrates through the tunnel of the Msp porin. In some embodiments, the system includes an Msp porin having a vestibule and a constriction zone defining a tunnel, where the tunnel is located within a lipid bilayer between the first liquid medium and the second liquid medium, and the only point of liquid communication between the first liquid medium and the second liquid medium occurs within the tunnel. Additionally, any Msp porin described herein may be included in any system described herein. In some embodiments, the system may further comprise an amplifier or a data collection device. The system may further comprise one or more temperature regulation devices in communication with the first liquid medium, the second liquid medium, or both. The systems described herein may be operable to move analytes through the Msp porin tunnels, either electrophoretically or otherwise.
[0173] As shown in FIG. 2, an extended polynucleotide 203 can be formed from a polynucleotide having modified nucleotides 210, each of which includes a circular loop modification 211. A daughter strand polynucleotide 220 can be synthesized by a polymerase from a template DNA using the modified nucleotides (e.g., modified dNTPs) 210. In the polymerization process, the modified dNTPs 210 with the circular loop 211 are incorporated into the growing daughter strand 220. Once the daughter strand 220 is created, the polynucleotide backbone is cleaved at the cleavable site, allowing the circular loop modification 211 to open and result in the extension of the daughter strand polynucleotide 220. The circular loop modifications on the modified dNTPs become linker constructs 206 in the extended polynucleotide 203 that create distance between adjacent nucleotides.
[0174] The polymerase used is generally an enzyme for joining 3'-OH 5'-triphosphate nucleotides, oligomers and their analogs. Polymerases include DNA-dependent DNA polymerase, DNA-dependent RNA polymerase, RNA-dependent DNA polymerase, RNA-dependent RNA polymerase, T7 DNA polymerase, T3 DNA polymerase, T4 DNA polymerase, T7 RNA polymerase, T3 RNA polymerase, SP6 RNA polymerase, DNA polymerase I, Klenow fragment, Thermophilus aquaticus DNA polymerase, Tth DNA polymerase, VentR® DNA polymerase (New England Biolabs), Deep VentR® DNA polymerase (New England Biolabs), Bst DNA polymerase large fragment, Stoeffel fragment, 90N DNA polymerase, 90N DNA polymerase, Pfu DNA polymerase, TfI DNA polymerase, Tth DNA polymerase, RepliPHI Phi29 polymerase, Tii DNA polymerase, eukaryotic DNA polymerase beta, telomerase, Therminator™ polymerase (New England Biolabs). Examples of polymerases that may be used include, but are not limited to, Fabry-Perot (Biolabs), KOD HiFi™ DNA polymerase (Novagen), KOD1 DNA polymerase, Q-beta replicase, terminal transferase, AMV reverse transcriptase, M-MLV reverse transcriptase, Phi6 reverse transcriptase, HIV-1 reverse transcriptase, novel polymerases discovered by bioprospecting, and polymerases cited in US Patent Application Publication No. 2007 / 0048748, US Patent Nos. 6,329,178, 6,602,695, and 6,395,524 (incorporated by reference). These polymerases include wild-type, mutant isoforms, and engineered variants. "Encode" or "parse" is a verb that refers to transferring from one format to another, and refers to transferring the genetic information of the target template sequence to the reporter configuration.
[0175] After the polymerization process is complete, cleavage at a predetermined position opens the loop and increases the distance between adjacent nucleotides. Cleavage of the daughter strand can be designed to occur at any part of the backbone, as long as it occurs within the loop structure between the two positions where the linker construct is attached to the nucleotide structure. Cleavage of the daughter strand along the backbone opens the loop and extends the daughter strand, leaving the linker construct linking the backbone phosphate and sugar. In embodiments where the linker construct contains a modification configured to interact with a nanopore, the modification can slow or stop the movement of the extending polymer, allowing the nucleotides to be read one by one by the nanopore.
[0176] In embodiments where a reporter moiety (such as a reporter barcode) is part of the linker construct, the cleaved product, i.e., the extended polymer, exposes a series of reporter moieties, each of which reports the identity of the base to which it corresponds. In embodiments where the linker construct also contains modifications configured to interact with the nanopore, the extended polymer can sequence one barcode at a time within the nanopore.
[0177] Cleavable cyclic loop nucleotides A cleavable cyclic loop nucleotide is a nucleotide / nucleotide analogue that has been modified to include a linker construct attached to two positions of the nucleotide / nucleotide analogue structure. The term "cleavable cyclic loop nucleotide" is used, which includes both modified natural nucleotides and modified nucleotide analogues. Useful nucleotides and nucleotide analogues as described herein include the following compounds: [ka] In the formula, X is NHR, OR, or CH 2R, where R is H, alkyl, aryl, heteroaryl, cycloalkyl, heterocycloalkyl, and the base is selected from the group consisting of adenine, cytosine, guanine, thymine, and uracil. The linker construct may be attached to a nucleotide / nucleotide analog to form a cleavable circular loop nucleotide.
[0178] The linker construct may be attached at one end to a nucleotide / nucleotide analogue by either a PNN, PO, or PC bond or other PX heteroatom bond, and the other end may be attached to (i) any position on the nucleobase in all of the above structures, (ii) the 5'-carbon of the ribose sugar in structures NT-1, NT-3, and NT-4, (iii) the N atom adjacent to the 5'-carbon in structure NT-4, (iv) the 5'-allyl modification in structure NT-2, or (v) any position on the ribose ring. Examples of covalent linking chemistries include amine-NHS ester, amine-imido ester, amine-pentafluorophenyl ester, amine-hydroxymethylphosphine, carboxyl-carbodiimide, thiol-maleimide, thiol-haloacetyl, thiol-pyridyl disulfide, thiol-thiosulfonate, thiol-vinyl sulfone, aldehyde-hydrazide, aldehyde-alkoxyamine, hydroxy-isocyanate, azido-alkyne, azido-phosphine, transcyclooctene-tetrazine, norbornene-tetrazine, azido-cyclooctyne, and azido-norbornene. When the linker construct is attached to the nucleotide / nucleotide analog, it forms the cyclic loop portion of the cleavable cyclic loop nucleotide. Thus, the cyclic loop is represented by -L. 1 -SP-L 2 - Including parts.
[0179] Depending on the structure and composition of the circular loop, several functions or structures may be present, including one or more of the following: conjugation moieties, linkers, spacers, reporters (reporter elements, barcodes), and stop constructs.
[0180] The conjugate moiety in the circular loop modification is formed by the attachment of a spacer moiety to one or more nucleotides or by the attachment of an additional modification (such as a stop construct) to the circular loop. The reactive groups at each end of the spacer moiety react with the reactive groups on the bifunctional nucleotide to form the conjugate moiety. In some embodiments, one or more stop constructs can also be attached to the circular loop via a conjugate moiety. The stop construct is a moiety configured to slow down the movement of the polynucleotide, so that one or more reporter elements can have a longer residence time in the nanopore readhead in the presence of a driving voltage, allowing identification of the reporter, and therefore the corresponding base. The spacer (SP) separates successive stop constructs to allow sufficient decay of the pulse voltage applied before the next stop construct. The spacer can also serve to extend any polynucleotide once a particular backbone element is cleaved. Thus, the circular loop is composed of any number of subelements that can act to affect and attenuate sequencing. In some embodiments, the strength of the background electric field, the type of nanopore, and the properties of the extended polynucleotide affect the rate of movement, efficiency, and accuracy.
[0181] Examples of cleavable cyclic loop nucleotides include: [ka] In the formula, X is -O-, -CH 2 -, -NH- [ka] and X' is =N-SO 2 -, =NH-CO-, or [ka] Y is -O-, -S-, -Se-, or -NH-; L 1 is a first linking group, and L 2is a second linking group, SP is a spacer, and the base is selected from the group consisting of adenine, cytosine, guanine, thymine, and uracil.
[0182] Further examples of cleavable circular loop nucleotide compounds that can be included in a nanopore sequencing system or kit for nanopore sequencing are as follows: [ka] In the formula, X is -O-, -CH 2 -, -NH- [ka] and X' is =N-SO 2 -, =NH-CO-, or [ka] Y is -O-, -S-, -Se-, or -NH-; L 1 is a first linking group, and L 2 is a second linking group, SP is a spacer, and the base is selected from the group consisting of adenine, cytosine, guanine, thymine, and uracil.
[0183] In some embodiments, the first linking group L 1 and a second linking group L 2 Each of L independently comprises a conjugated moiety selected from the group consisting of amine-NHS ester, amine-imido ester, amine-pentafluorophenyl ester, amine-hydroxymethylphosphine, carboxyl-carbodiimide, thiol-maleimide, thiol-haloacetyl, thiol-pyridyl disulfide, thiol-thiosulfonate, thiol-vinyl sulfone, aldehyde-hydrazide, aldehyde-alkoxyamine, hydroxy-isocyanate, azido-alkyne, azido-phosphine, transcyclooctene-tetrazine, norbornene-tetrazine, azido-cyclooctyne, and azido-norbornene. 1 and L 2may or may not be the same.
[0184] In some embodiments, the first linking group L 1 and a second linking group L 2 may each independently further comprise a linker. A first linker may be present between the conjugated moiety and X / X' (alpha phosphate) and a second linker may be present between the conjugated moiety and SP. In some embodiments, the linker may be selected from the group consisting of hydrophilic polymers (e.g., polyethylene glycol, polyvinyl alcohol, polyacrylamide, polyvinylpyrrolidone, polystyrene sulfonate, polyethyleneimine), hydrophobic polymers (e.g., polylactic acid, polymethyl methacrylate, polystyrene), oligonucleotides, peptides, polypeptides, aliphatic chains (C5-C50), and combinations thereof. In some embodiments, the first and second linkers may independently comprise a peptide, a polypeptide, an alkyl chain, a polyethylene glycol, or combinations thereof. In some embodiments, one or more linkers may be absent.
[0185] In some embodiments, SP comprises one or more of the following moieties: (1) simple aliphatic chains, such as alkyl chains having 5-50 carbons, and substituted aliphatic chains (wherein the substituents may include halo, such as chloro, bromo, or fluoro, alkyl, such as methyl, ethyl, or propyl, or aromatic groups, such as phenyl or pyridyl), (2) oligonucleotides, modified oligonucleotides, or polyphosphates having 1-100 repeat units, (3) polypeptides having 1-100 repeat units, (4) hydrophilic polymers having 1-100 repeat units, such as polyethylene glycol, polyvinyl alcohol, polyacrylamide, polyvinylpyrrolidone, polystyrene sulfonate, and polyethyleneimine, (5) hydrophobic polymers having 1-100 repeat units, such as polylactic acid, polymethyl methacrylate, and polystyrene. In some embodiments, the alkyl chains may be substituted or unsubstituted. In some embodiments, the number of repeating units (monomers) in the SP can be, for example, in the range of 1-5, 6-10, 11-15, 16-20, 20-25, 26-50, or 50-100, or any combination of the foregoing ranges. In some embodiments, the total number of repeating units in the SP can be 5-100, 10-100, 10-80, 10-70, 5-60, or 5-50.
[0186] The number of repeat units and the length of the spacer SP may depend on the following factors: (1) Selection of repeat units / monomers - shorter / smaller monomers may require more repeats to construct a similar length compared to longer monomers. (2) Steric bulk of the spacer - larger spacer monomers are more likely to result in steric clashes with the nanopore read head, resulting in slower translocation rates compared to less bulky monomers. (3) Interaction of the spacer with the nanopore - spacer monomers that can form stronger interactions (e.g., electrostatic interactions, H-bonds) with the nanopore residues are more likely to experience slower translocation rates compared to monomers that form weaker interactions (e.g., non-polar interactions). (4) Charge of the selected modification - loops with a higher net negative charge will experience a higher translocation rate (compared to lower net negative charge loops) in the presence of an applied voltage.
[0187] In some embodiments, phosphoramidite analogs can be assembled into polymers (polyphosphates) using oligonucleotide synthesis processes (e.g., phosphoramidite methods). Examples of polyphosphates can include the following: [ka] [ka] In the formula, X 1 =O - , OMe, or S - and a is 1 to 100. In some embodiments, a is 1 to 5, 6 to 10, 11 to 15, 16 to 20, 20 to 25, 26 to 50, or 50 to 100, or any combination of the foregoing ranges.
[0188] In some embodiments, the polypeptide can be a homopolypeptide or a heteropolypeptide. [ka] wherein a is 1 to 100. In some embodiments, a is 1 to 5, 6 to 10, 11 to 15, 16 to 20, 20 to 25, 26 to 50, or 50 to 100, or a combination of any of the foregoing ranges. The polypeptides can include both natural and unnatural amino acid residues, including non-exhaustive examples of residues selected from the following: L / D-Natural Amino Acids [ka] L / D-unnatural amino acids [ka] [ka] Non-amino acids [ka]
[0189] In some embodiments, the modified oligonucleotides in the SP may include modified nucleotides and / or modified nucleobases. In some embodiments, examples of modified nucleotides and modified nucleobases include, but are not limited to, the following: Modified Nucleotides [ka] Modified Nucleobases [ka]
[0190] In some embodiments, the polyamide compound may be a homopolyamide or a heteropolyamide. [ka] The polyamide compound may include one or more of the following residues: polyamide [ka] [ka]
[0191] In some embodiments, the spacer may include a reporter moiety that corresponds to a particular nucleobase, while in other embodiments, the spacer may not include a reporter moiety.
[0192] In some embodiments, the cleavable cyclic loop nucleotide may further comprise a stop construct configured to interact with a nanopore. The stop construct may comprise a linear, branched or cyclic polymer, where the polymer is selected from a synthetic hydrophobic polymer, a synthetic hydrophilic polymer, an oligonucleotide / polynucleotide, a peptide / polypeptide, and combinations thereof. In some embodiments, the stop construct may be a side branch attached to the cleavable cyclic loop nucleotide. In some embodiments, the stop construct may be a side branch attached to the nucleobase of the nucleotide / nucleotide analog. In some embodiments, the stop construct may be a spacer SP or a linking group L. 1 Or L 2 In some embodiments, the stop construct may be incorporated into or be part of a circular loop structure. In some embodiments, the stop construct may be adjacent to the reporter in the circular loop structure.
[0193] When the stop construct is attached to the spacer SP portion of the circular loop, the stop construct can be attached to any part of the spacer, for example, the middle of the spacer strand, either end of the spacer strand, or anywhere in between. In embodiments where the stop construct is attached to the middle of the circular loop structure, the circular loop can be a symmetric loop. In embodiments where the stop construct is attached to another part of the circular loop structure, the circular loop can be an asymmetric loop.
[0194] The stop construct comprises a third linking group, L 3In some embodiments, L 3 may be a moiety selected from the group consisting of hydrophilic polymers (polyethylene glycol, polyvinyl alcohol, polyacrylamide, polyvinylpyrrolidone, polystyrene sulfonate, polyethyleneimine), hydrophobic polymers (polylactic acid, polymethyl methacrylate, polystyrene), oligonucleotides, peptides, polypeptides, aliphatic chains (C5-C50), aromatic groups (phenyl or pyridyl), and combinations thereof.
[0195] The cyclic loop comprises one or more sub-elements including linkers, conjugated moieties, spacers, stop constructs, and reporters. The arrangement of the sub-elements of the cyclic loop may be asymmetric (i.e., asymmetric cyclic loops-ACLs) or symmetric with respect to the order of the sub-elements (i.e., symmetric cyclic loops-SCLs). Figure 93 shows non-limiting examples of ACLs, where each cyclic loop comprises any number of conjugated moieties, spacers, reporters, and stop constructs (ARCs). Asymmetric cyclic loops are polymer loops in which the sequence or overall composition of the individual sub-elements is not symmetrically arranged. Thus, when oriented inside the readhead of a nanopore, ACLs can affect the physicochemical properties, residence time, retention force, and voltage required for translocation. In contrast, symmetric cyclic loops comprise cyclic loops in which the subunits are symmetrically arranged. Figure 94 shows non-limiting examples of SCLs, where each SCL comprises any number of conjugated moieties (handles), spacers, reporters, and stop constructs (ARCs). The ACL and SCL structures may feature any number of the sub-elements described herein, and may omit certain sub-elements, without limitation.
[0196] Model alpha-phosphate substituted nucleotides: Imino-P and other substitutions In some embodiments, the constituent group in the nucleotide may have one or more substitutions compared to naturally occurring nucleotides. In some embodiments, one or more phosphate groups comprising the nucleotide are substituted. In some embodiments, one or more atoms in the nucleotide structure may be substituted with a heteroatom, a heteroatom isotope, or a homoatom isotope thereof. In some embodiments, a stable sulfonylimino-P substitution ("imino-P") for an alpha phosphate substituted nucleotide is provided herein. A bifunctional nucleotide comprising an imino-P substitution at the alpha phosphate is shown above. [ka]
[0197] Other useful nucleotides with substitutions at the alpha phosphate include: [ka]
[0198] In some embodiments, the nucleotides described herein are storage stable at 25° C. for 2 days. In some embodiments, the nucleotides described herein are storage stable at −25° C. for 2 days. In some embodiments, the nucleotides described herein are storage stable at −25 to 25° C. for 2 days. In some embodiments, the nucleotides described herein are storage stable at −25 to 25° C. for 2 days. In some embodiments, the nucleotides described herein are storage stable between −50 to 50° C. for 2 days. In some embodiments, the nucleotides described herein are storage stable at 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, and 50° C. for 2 days. In some embodiments, the nucleotides described herein are storage stable at 25° C. for 0 to 2 days. In some embodiments, the nucleotides described herein are storage stable at −25° C. for 0 to 2 days. In some embodiments, the nucleotides described herein are storage stable at −25° C. for 0 to 2 days. In some embodiments, the nucleotides described herein are storage stable at -25 to 25°C for 0 to 2 days. In some embodiments, the nucleotides described herein are storage stable at -50 to 50°C for 0 to 2 days. In some embodiments, the nucleotides described herein are storage stable at 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, and 50°C for 0 to 2 days. In some embodiments, the nucleotides described herein are storage stable at 25°C. In some embodiments, the nucleotides described herein are storage stable at -25° C. In some embodiments, the nucleotides described herein are storage stable at -25 to 25° C.In some embodiments, the nucleotides described herein are storage stable at -25 to 25° C. In some embodiments, the nucleotides described herein are storage stable at -50 to 50° C. In some embodiments, the nucleotides described herein are storage stable at 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, and 50° C.
[0199] Cuttable area The cleavable cyclic loop nucleotide comprises a specific cleavable site in a bond that can be broken under controlled conditions, such as, for example, a phosphorothiolate bond, a photocleavable bond, a phosphoramidite bond, a phosphoramide bond, a 3'-O-BD-ribofuranosyl-2' bond, a thioether bond, a selenoether bond, a sulfoxide bond, a disulfide bond, a deoxyribosyl-5'-3' phosphodiester bond, or a ribosyl-5'-3' phosphodiester bond, as well as conditions for selective cleavage of other cleavable bonds known in the art. The selectively cleavable bond can be an intra-tether bond or between or within the probe or nucleobase residues, or a bond formed by hybridization between the probe and the template strand. The selectively cleavable bond is not limited to a covalent bond, but can be a non-covalent bond or association, such as those based on hydrogen bonds, hydrophobic bonds, ionic bonds, pi-bond ring stacking interactions, van der Waals interactions, and the like.
[0200] For example, in some embodiments, the cleavable site includes the PY bond / linkage in cleavable circular loop nucleotide structures (I) and (II), the PN bond / linkage in structure (III), and the OC (5'-C of the ribose sugar) bond / linkage in structures (IV) and (V) and (XII). These bonds are shown below as bolded bonds and can be cleaved under conditions known in the art. [ka] In the formula, X, X', Y, L 1 , L 2 , SP and the base are as defined above.
[0201] Figure 3 shows a schematic example of a cleavable bond. The dashed arc indicates a possible linker construct attached to two sections / positions of the same nucleotide. As disclosed herein, there are several attachment points on the nucleotide to which a linker construct (e.g., a circular loop) can be attached, and the cleavage site must be located within the two attachment points.
[0202] There are numerous cleavable bonds that can be easily incorporated into the existing structure of an oligonucleotide without causing a significant increase in the length of the DNA backbone. Referring to FIG. 3, the cleavable linkages / bonds in various nucleotides / nucleotide analogs are circled ( [ka] (shown as ). Below are the conditions for cleaving these bonds for the nucleotides / nucleotide analogs shown in Figure 3 (from left to right): 1) The phosphoramidate-circled bond can be cleaved under acidic conditions (e.g., 10 mM sodium citrate-HCl or 80% acetic acid at pH 4.0). 2) Phosphorothiolate-circled bonds can be cleaved with 50 mM silver nitrate or iodine in aqueous acetone / pyridine (1:1). 3) Allyl-circled bonds can be cleaved with Pd(0) salts. In one example, Figure 4 shows a schematic of an allyl cleavable chemistry. 4) The phosphodiester-cyclic bond can be cleaved via enzymatic cleavage (eg, endo-, exo-nucleases or RNases, basic conditions in the case of RNA ribose).
[0203] Methods for making circular loop nucleotides Circular loop nucleotides can be made by attaching a spacer moiety to a bifunctional nucleotide. In some embodiments, the attachment of the spacer moiety to the bifunctional nucleotide involves click chemistry. The bifunctional nucleotide design can be derived from the following: [ka]
[0204] R 1 and R 2 is a reactive group that can be attached to the spacer SP using click chemistry. In some embodiments, R 1 and R 2 comprises a click chemistry reagent. In some embodiments, R 1 and R 2 are independently hydroxyl, thiocyanate, aldehyde, carboxyl, azide (-N 3 ), amines (-NH 2 ), alkyne, bicyclononyne (BCN), dibenzocyclooctyne (DBCO), thiol (-SH), tetrazine, trans-cyclooctyne (TCO), NHS ester, imidoester, pentofluorophenyl ester, hydroxylmethylphosphine, carbodiimide, maleimide, haloacetyl, pyridyl disulfide, thiosulfonate, vinyl sulfone, hydrazide, alkoxyamine, isocyanate, phosphine, and norbornene. 1 and R 2 is the same. In another embodiment, R 1 and R 2 may be different.
[0205] X is -O-, -CH 2 -, -NH-, [ka] X' may be =N-SO 2 -, =NH-CO-, or [ka] Y may be -O-, -S-, -NH-, or -Se-. L' is a linker. In some embodiments, the linker L' may be independently selected from the group consisting of a hydrophilic polymer, a hydrophobic polymer, an oligonucleotide, a peptide, a polypeptide, an aliphatic chain (C5-C50), and combinations thereof. In some embodiments, the hydrophilic polymer, the hydrophobic polymer, the oligonucleotide, and the polypeptide may each have 1-100 repeat units. The hydrophilic polymer may comprise polyethylene glycol, polyvinyl alcohol, polyacrylamide, polyvinylpyrrolidone, polystyrene sulfonate, polyethyleneimine, or combinations thereof. The hydrophobic polymer may comprise polylactic acid, polymethylmethacrylate, or polystyrene, or combinations thereof. In some embodiments, the linker L' may comprise a reporter that encodes the associated nucleobase. In some embodiments, the linker L' may also further comprise a stop construct configured to interact with the nanopore to slow the translocation of the polynucleotide into which it is incorporated.
[0206] Representative examples of bifunctional nucleotides include: [ka]
[0207] To bind to the spacer portion to form a circular loop nucleotide, the spacer portion is provided with a spacer SP and a reactive group R at both ends of the spacer portion that is desired to be bound to the bifunctional nucleotide. 1 ' and R 2 In some embodiments, R 1 ' and R 2 ' is a hydroxyl, thiocyanate, aldehyde, carboxyl, azide (-N 3 ), amines (-NH 2), alkyne, bicyclononyne (BCN), dibenzocyclooctyne (DBCO), thiol (-SH), tetrazine, trans-cyclooctyne (TCO), N-hydroxysuccinimide (NHS) ester, imidoester, pentofluorophenyl ester, hydroxylmethylphosphine, carbodiimide, maleimide, haloacetyl, pyridyl disulfide, thiosulfonate, vinyl sulfone, hydrazide, alkoxyamine, isocyanate, phosphine, and norbornene. 1 ' and R 2 In another embodiment, R 1 ' and R 2 ' may be different.
[0208] In some embodiments, the spacer moiety is SP and R 1 ' and / or SP and R 2 The spacer portion may further comprise a linker L″ on one or both sides of the SP, such as between SP′ and SP′. In some embodiments, each of the linkers L″ may be independently selected from the group consisting of hydrophilic polymers, hydrophobic polymers, oligonucleotides, peptides, polypeptides, aliphatic chains (C5-C50), and combinations thereof. In some embodiments, the hydrophilic polymers, hydrophobic polymers, oligonucleotides, and polypeptides may each have 1-100 repeat units. The hydrophilic polymer may comprise polyethylene glycol, polyvinyl alcohol, polyacrylamide, polyvinylpyrrolidone, polystyrene sulfonate, polyethyleneimine, or combinations thereof. The hydrophobic polymer may comprise polylactic acid, polymethylmethacrylate, or polystyrene, or combinations thereof. In some embodiments, the spacer portion may further comprise a stop construct designed to slow down the movement of the polynucleotide through the nanopore.
[0209] To form a circular loop nucleotide, the bifunctional nucleotide R 1 and R 2 are the R of SP, respectively. 1 ' and R 2' to form a conjugate moiety. In some embodiments, the conjugate moiety (e.g., R 1 -R 1 ' and R 2 -R 2 The reactive groups R′ may be independently selected from a non-exhaustive list of chemistries such as amine-NHS ester, amine-imido ester, amine-pentofluorophenyl ester, amine-hydroxymethylphosphine, carboxyl-carbodiimide, thiol-maleimide, thiol-haloacetyl, thiol-pyridyl disulfide, thiol-thiosulfonate, thiol-vinyl sulfone, aldehyde-hydrazide, aldehyde-alkoxyamine, hydroxy-isocyanate, azido-alkyne, azido-phosphine, transcyclooctene-tetrazine, norbornene-tetrazine, azido-cyclooctyne, and azido-norbornene. 1 / R 2 , R 1 ' / R 2 ', and R 1 -R 1 ' or R 2 -R 2 Examples of the resulting conjugated moieties formed by R′ are shown in the table below. As shown in the table, R and R′ are 1 and R 1 Between ' and R 2 and R 2 ' denotes two groups that are joined by a reaction between . One of the two groups may represent a bifunctional nucleotide, and the other group may represent a spacer moiety. [Table 1]
[0210] The spacer moiety may be generated via one or more synthesis schemes, including solid phase, solution / liquid phase, and enzyme-based synthesis. Exemplary synthesis methods include polymers synthesized using phosphoramidite chemistry, peptide synthesis, click chemistry, and other bioconjugation methods known in the art, which may result in the formation of phosphodiester, methylphosphonate, or phosphorothioate bonds between each moiety. In some embodiments, the synthesis of the spacer moiety may be performed linearly. In some embodiments, the synthesis of the spacer moiety may be performed via branching. In some embodiments, the synthesis of the spacer moiety may be performed via the joining of specific segments that include one or more subelements. Under linear synthesis, constituent subelements (monomers) may be added to a growing chain of subelements that include the spacer moiety. In some embodiments, the initiation of synthesis begins at or near one end of the spacer moiety where a first reactive group is configured to bind to a bifunctional nucleotide and a second reactive group is configured to bind to a bifunctional nucleotide, and ends at or near the other end of the spacer moiety. In some embodiments, the first and second reactive groups are chemically distinct. In some embodiments, the first and second reactive groups are chemically identical.
[0211] In contrast, branched synthesis (i.e., branched) includes synthetic schemes in which two or more branches (i.e., arms) of a spacer moiety are generated from any starting subunit, thereby generating a spacer moiety available on a solid or liquid support.
[0212] In some embodiments, the spacer moiety is synthesized using solid phase synthesis. In solid phase synthesis, the polymer spacer moiety is synthesized on a solid support in a stepwise manner, with subsequent subelements (monomers) being added to the growing polymer chain. In some embodiments, the spacer moiety is synthesized using solution phase synthesis. In solution phase synthesis, the spacer moiety is synthesized stepwise in solution, with the subelements (monomers) assembled onto the growing polymer chain. In some embodiments, the spacer moiety is synthesized using enzymatic synthesis. In enzymatic synthesis, enzymes are used to synthesize the polymer spacer moiety in a stepwise process. Enzymes can be used to link specific subelements (monomers) with specific chemical bonds. In some embodiments, the spacer moiety is synthesized using a combination of one or more of the following: solid phase synthesis, solution phase synthesis, and enzymatic synthesis.
[0213] In some embodiments, exemplary symmetric spacer portions having a stop construct may be synthesized using a solid phase synthesis route as shown in Figure 90. In some embodiments, exemplary asymmetric spacer portions having a stop construct may be synthesized using a solid phase synthesis route as shown in Figure 91.
[0214] In some embodiments, an exemplary symmetric spacer moiety with a stop construct can be synthesized by linking segments in a two-step process, where a polymer chain containing a cyclic loop, compound 1, is attached to the stop construct mPEG2 in step 1, and an azide conjugate moiety is attached in step 2 to yield compound 3 (see Figure 92).
[0215] daughter chains Compounds (Ia), (IIa), (IIIa), (IVa), (Va) and (XIIa) can be used as cyclic loop nucleotides to synthesize a daughter strand. The daughter strand comprises at least one of compounds (I), (II), (III), (IV), (V) and (XII). In some embodiments, the daughter strand is formed by linking together multiple compounds (I), multiple compounds (II), multiple compounds (III), multiple compounds (IV), multiple compounds (V) or multiple compounds (XII) into a polynucleotide or oligonucleotide. The daughter strand comprises one of the following structures: [ka] [ka] In the formula, X, X', Y, L 1 , L 2 , SP, and the bases are as defined above, and adjacent bases on the daughter strand may be different or the same.
[0216] In some embodiments, structure (VII) can be further represented by the following structure: [ka]
[0217] In some embodiments, structure (VIII) can be further represented by the following structure: [ka]
[0218] In some embodiments, structure (X) can be further represented by the following structure: [ka]
[0219] Cleavage of the circular loop nucleotide The daughter strand may be further subjected to conditions disclosed herein suitable to cleave the circular loop nucleotides and extend the daughter strand to form an extended polynucleotide: [ka] In the formula, X, X', Y, L 1 , L 2 , SP, and the base are as defined above, and adjacent bases on the extending polymer chain can be different or the same.
[0220] Cleavable circular loop tetramer oligonucleotides In some embodiments, the two ends of the linker construct can be attached to two nucleobases of the tetrameric oligonucleotide to form a cleavable circular loop tetrameric oligonucleotide. The allylic cleavable bond can be located anywhere along the oligonucleotide backbone between the two attachment points, and upon cleavage, open the circular loop. In some embodiments, the attachment points can be the first and second bases, the first and third bases, the first and fourth bases, the second and third bases, the second and fourth bases, or the third and fourth bases. For example, when the attachment points are the first and second bases, the allylic cleavage site can be located on the backbone between the first and second bases. When the attachment points are the first and third bases, the allylic cleavage site can be located on the backbone between the first and third bases. When the attachment points are the first and fourth bases, the allylic cleavage site can be located on the backbone between the first and fourth bases. When the attachment points are the second and third bases, the allylic cleavage site can be located on the backbone between the second and third bases. If the points of attachment are the second and fourth bases, the allylic cleavage site can be located on the backbone between the second and fourth bases. If the points of attachment are the third and fourth bases, the allylic cleavage site can be located on the backbone between the third and fourth bases.
[0221] In some embodiments, when the points of attachment are the first and second bases, the allyl group is at the C of the second nucleotide. 5If the attachment points are the first and third bases, the allyl group may be located on the C' of the second or third nucleotide. 5 If the attachment points are at the first and fourth bases, the allyl group can be located at the C of the second, third, or fourth nucleotide. 5 If the attachment points are the second and third bases, the allyl group can be located on the C of the third nucleotide. 5 If the attachment points are the second and fourth bases, the allyl group can be located on the C of the third or fourth nucleotide. 5 If the attachment points are the third and fourth bases, the allyl group can be located on the C of the fourth nucleotide. 5 In some embodiments, heavy atom substituted nucleotides can be used to generate the resulting 4-mer oligonucleotide.
[0222] Examples of cleavable circular loop tetramer oligonucleotides include the following: [ka] In the formula, L 1 is a first linking group, and L 2 is a second linking group, and R 1 , R 2 , and R 3 one of is allyl and the other is H, SP is a spacer, and the base is selected from the group consisting of adenine, cytosine, guanine, thymine, and uracil.
[0223] The cyclic loop can be designed to be symmetrical or asymmetrical depending on the linking chemistry on the nucleotide. In some embodiments, the first linking group L 1 and a second linking group L 2Each of L independently comprises a conjugated moiety selected from the group consisting of amine-NHS ester, amine-imido ester, amine-pentafluorophenyl ester, amine-hydroxymethylphosphine, carboxyl-carbodiimide, thiol-maleimide, thiol-haloacetyl, thiol-pyridyl disulfide, thiol-thiosulfonate, thiol-vinyl sulfone, aldehyde-hydrazide, aldehyde-alkoxyamine, hydroxy-isocyanate, azido-alkyne, azido-phosphine, transcyclooctene-tetrazine, norbornene-tetrazine, azido-cyclooctyne, and azido-norbornene. 1 and L 2 may or may not be the same.
[0224] In some embodiments, the first linking group L 1 and a second linking group L 2 Each of the linkers may independently further comprise a linker. A first linker may be present between the conjugate moiety and a base in the tetrameric oligonucleotide, and a second linker may be present between the conjugate moiety and another base in the tetrameric oligonucleotide. In some embodiments, the linker may be selected from the group consisting of hydrophilic polymers (polyethylene glycol, polyvinyl alcohol, polyacrylamide, polyvinylpyrrolidone, polystyrene sulfonate, polyethyleneimine), hydrophobic polymers (polylactic acid, polymethylmethacrylate, polystyrene), oligonucleotides, peptides, polypeptides, aliphatic chains (C5-C50), and combinations thereof. In some embodiments, the first and second linkers may independently comprise a peptide, a polypeptide, an alkyl chain, polyethylene glycol, or a combination thereof. In some embodiments, one or more polynucleotides may be used.
[0225] In some embodiments, the SP (spacer) comprises a polymer. The spacer in the cyclic loop provides a buffer distance between successive cyclic loops or between successive subelements comprising one or more cyclic loops. In some embodiments, the polymer in the SP comprises an oligonucleotide, a modified oligonucleotide, a hydrophilic polymer (polyethylene glycol, polyvinyl alcohol, polyacrylamide, polyvinylpyrrolidone, polystyrene sulfonate, polyethyleneimine), a hydrophobic polymer (polylactic acid, polymethylmethacrylate, polystyrene), a polypeptide, an aliphatic chain (C5-C50), a substituted aliphatic chain (small molecule such as chloro, bromo or fluoro, alkyl such as methyl, ethyl or propyl, or aromatic group such as phenyl or pyridyl), or a combination thereof. In some embodiments, the modified oligonucleotide is an oligonucleotide that does not have a base attached to the sugar, e.g., [ka] may include. In some embodiments, the modified oligonucleotide comprises: [ka] and can be assembled into polymers using oligonucleotide synthesis processes. In some embodiments, the spacer can include a reporter moiety that corresponds to a particular nucleobase. In other embodiments, the spacer may not include a reporter moiety (barcode).
[0226] In some embodiments, the circular loop may include one or more barcodes or reporter moieties, sub-reporters, reporter elements, and reporters including sub-reporter elements. Examples include, but are not limited to, nucleoside bases, non-nucleoside bases, peptides, or other synthetic polymers (e.g., polyethylene glycol, polyvinyl alcohol, polyacrylamide, polyvinylpyrrolidone, polyethyleneimine, etc.). In some embodiments, the reporter element may include a polymer (e.g., crown ether, cucurbituril, pillararenes, or cyclodextrins). Coupling of these macromolecules to cyclic loop structures is possible through covalent chemistries such as amine-NHS ester, amine-imido ester, amine-pentafluorophenyl ester, amine-hydroxymethylphosphine, carboxyl-carbodiimide, thiol-maleimide, thiol-haloacetyl, thiol-pyridyl disulfide, thiol-thiosulfonate, thiol-vinyl sulfone, aldehyde-hydrazide, aldehyde-alkoxyamine, hydroxy-isocyanate, azido-alkyne, azido-phosphine, transcyclooctene-tetrazine, norbornene-tetrazine, azido-cyclooctyne, and azido-norbornene. In some embodiments, the reporter can be selected from any moiety including spacers, conjugated moieties, and termination constructs. In some embodiments, the reporter can contain multiple modifications ranging from any of 1-10, or 11-15, or 16-20, or 21-25, 26-50, or 50-100 units in length, so long as the reporter provides a reproducible signal with a given voltage waveform when present to the readhead of the nanopore.
[0227] In some embodiments, the advancement of each nucleotide bearing one or more barcodes or reporter moieties corresponds to a translocation event. In some embodiments, the reporter moiety comprises one or more sub-reporter moieties, where the sub-reporter moieties are configured to identify the translocation event and generate a signal when passed through a readhead of the nanopore. In some embodiments, the reporter moiety comprises two or more sub-reporter moieties, where each sub-reporter moiety in the two or more sub-reporter moieties is distinguishable, reproducible, and resolvable. In some embodiments, the circular loop may comprise a first set of reporter moieties and a second set of reporter moieties, where the first set of reporter moieties is configured to generate a signal to identify a particular nucleotide passing through the readhead, where the second set of reporter moieties is configured to generate a signal to identify the passage of each nucleotide regardless of nucleotide identity.
[0228] Heavy Atom Substituted Oligonucleotides In some embodiments, heavy atoms (e.g., non-hydrogen atoms) can be substituted into modified nucleotides. Naturally occurring nucleotides based on triphosphate and phosphodiester bonds can be limiting for various enzymatic processes (e.g., ligation, incorporation). Heavy atoms (including, for example, sulfur and selenium) as substitutions in nucleotides can enable novel biochemistry and applications for cleaving, ligating, incorporation, or otherwise manipulating oligonucleotides. Heavy atom modified nucleotides can be incorporated or otherwise linked into the resulting modified polynucleotide.
[0229] In some embodiments, phosphorothiolate (bridged PS) and phosphorothioate (non-bridged PS) substitutions are described, as well as methods for generating such substitutions. Table 2 shows various cleavage reagents (including iodine and silver nitrate) known to cleave phosphorus-sulfur bonds. In some embodiments, selenium can also be used in place of sulfur when generating heavy atom modified nucleotides or oligonucleotides. Selenium modified nucleic acid analogs can also potentially be synthesized and cleaved according to the embodiments presented herein. Thus, disclosed herein include the following oligonucleotides: [ka] In the formula, Y is -O-, -S-, -NH-, or -Se-; Y 1 , Y 2 and Y 3 one of them is -S- or -Se-, and the rest are -O- or -NH-; [Table 2]
[0230] Some examples of heavy atom substituted nucleotides are shown below. [ka]
[0231] The number of bridging PS substitutions for any K-mer oligonucleotide can be any number between 1 and K-1. For example, for a 4-mer oligonucleotide, there can be 1-3 bridging PS modifications. In some embodiments, the number of bridging PS substitutions that can be introduced is 1-10, 1-20, 1-30, 1-40, 1-50, 1-60, 1-70, 1-80, 1-90, and 1-100 substitutions, or any value between the aforementioned ranges of values. In some embodiments, for an oligonucleotide of length K, the number of bridging PS substitutions can be 0 to K-1.
[0232] For phosphorothioates (non-crosslinked PS), iodine can selectively cleave the crosslinked PO bond at the site of the phosphorothioate in the presence of nucleophiles, such as amines. However, the subsequent crosslinked PO bond cleavage is not specific and either the 5' or 3' end can potentially be cleaved. Furthermore, the efficiency of cleavage can be affected by the possible conversion of phosphorothioates to phosphates. EXAMPLES
[0233] 5 shows a schematic of an example of an allyl-based cleavable cyclic loop nucleotide 501 having 10 Ts as a spacer region 502, a phosphate-linked (PBL) cyclic loop, and a stop construct 509 attached to the spacer region 502. One end of the spacer region 502 is attached to the alpha phosphate group of the nucleotide via a first linkage group 503, and the other end of the spacer region 502 is attached to the base of the nucleoside via a second linkage group 504. The stop construct 509 is attached to the spacer region 502 via a third linkage group 505.
[0234] The stop construct 509 can be constructed from one or more durable water-soluble or solvent-soluble polymers, including, but not limited to, the following segment(s): polyethylene glycol, polyglycol, polypyridine, polyisocyanide, polyisocyanate, poly(triarylmethyl)methacrylate, polyaldehyde, polypyrrolinone, polyurea, polyglycol phosphodiester, polyacrylate, polymethacrylate, polyacrylamide, polyvinyl ester, polystyrene, polyamide, polyurethane, polycarbonate, polybutyrate, polybutadiene, polybutyrolactone, polypyrrolidinone, polyvinylphosphonate, polyacetamide, polysaccharide, polyhyaluronic acid, polyamide, polyimide, polyester, polyethylene, polypropylene, polystyrene, polycarbonate, polyterephthalate, polysilane, polyurethane, polyether, polyamino acid, polyglycine, polyproline, N-substituted polylysine, polypeptide, side chain N-substituted peptide, poly-N-substituted glycine, peptoid, side chain carboxylate, poly ... carboxyl substituted peptides, homopeptides, oligonucleotides, ribonucleic acid oligonucleotides, deoxynucleic acid oligonucleotides, oligonucleotides modified to prevent Watson-Crick base pairing, oligonucleotide analogues, polycytidylic acid, polyadenylic acid, polyuridylic acid, polythymidine, polyphosphate, polynucleotides, polyribonucleotides, polyethylene glycol-phosphodiesters, peptide polynucleotide analogues, threosyl-polynucleotide analogues, glycol-polynucleotide analogues, morpholino-polynucleotide analogues, locked nucleotide oligomer analogues, polypeptide analogues, branched polymers, comb polymers, star polymers, dendritic polymers, random, gradient and block copolymers, anionic polymers, cationic polymers, polymers forming stem loops, rigid segments and flexible segments.
[0235] FIG. 6 shows a schematic example of sequencing an extended polymer 603 with a stop construct. The extended polymer 603 is formed after cleaving the O-C bond 506 of the allyl-based cleavable circular loop nucleotide 501 as shown in FIG. 5. In FIG. 6, a protein nanopore 601 is deposited in a lipid bilayer 602. The extended polymer 603 translocates through the nanopore 601. The polynucleotide 603 includes a spacer region 502 between successive nucleotides. In this embodiment, the spacer region 502 includes a reporter moiety, such as a reporter barcode (10 Ts) corresponding to the nucleotide "T". As the extended polymer 603 translocates through the nanopore 601, the stop construct 509 slows or temporarily stops translocation when an interaction occurs between the stop construct 509 and the nanopore 601, allowing the read head to read the reporter barcode, thereby identifying the nucleotide.
[0236] In other embodiments, the stop construct can be attached at different positions on the linker construct (e.g., attached to the first or second linking group, or to the nucleobase). By adjusting the position of stop construct attachment, the type / size of the stop construct, and the length of the spacer region / linker construct, it is possible to affect which region of the extending polymer is present at the nanopore constriction when translocation is slowed or stopped. In some embodiments, the nucleobase can be located at the constriction when translocation is slowed or stopped. Thus, the read head can also read the nucleobase itself, regardless of whether a reporter moiety is included in the spacer.
[0237] Figure 9 shows the structure of the cations where X is O, NH, or NSO. 2 or CH 2FIG. 10 shows an example of a first design of a fully functional cyclic loop nucleotide (CLN-1) with an alpha phosphate linked symmetric loop and a stop construct on the loop, where X=O and Y=O. FIG. 11 shows an example of a synthesis process of a fully functional cyclic loop nucleotide (CLN-1B) shown in FIG. 9 with X=NH and Y=O. FIG. 12 shows an example of a synthesis process of a fully functional cyclic loop nucleotide (CLN-1A) shown in FIG. 9 with X=CH. 2 and Y=O. Figure 13 shows an exemplary synthesis process for a fully functional cyclic loop nucleotide (CLN-1C) shown in Figure 9, where X=O and Y=S. Figure 13 shows an exemplary synthesis process for a fully functional cyclic loop nucleotide (CLN-1D) shown in Figure 9, where X=O and Y=S, which is compatible with phosphorothiolate cleavage.
[0238] Figure 14 shows the structure of the cations where X is O, NH, or NSO. 2 or CH 2 FIG. 15 shows an example of a synthesis process for a fully functional cyclic loop nucleotide (CLN-2A) shown in FIG. 14 where X=O and Y=O. FIG. 16 shows an example of a synthesis process for a fully functional cyclic loop nucleotide (CLN-2B) shown in FIG. 14 where X=NH and Y=O. FIG. 17 shows an example of a synthesis process for a fully functional cyclic loop nucleotide (CLN-2C) shown in FIG. 14 where X=CH. 2 15A-15C are schematic diagrams illustrating an exemplary synthesis process for a fully functional circular loop nucleotide (CLN-2C) shown in FIG. 14, in which Y=O and Y=O.
[0239] Figure 18 shows the structure of the cations where X is O, NH, or NSO. 2 or CH 2FIG. 19 shows an example of a third design of fully functional cyclic loop nucleotide (CLN-3) with an alpha phosphate-allyl linked asymmetric loop and a stop construct on the nucleobase, where X=O and Y=O. FIG. 20 shows an example of a synthetic process for fully functional cyclic loop nucleotide (CLN-3B) shown in FIG. 18 with X=NH and Y=O. FIG. 21 shows an example of a synthetic process for fully functional cyclic loop nucleotide (CLN-3A) shown in FIG. 18 with X=CH. 2 and Y=O, as shown in FIG. 18 (CLN-3C).
[0240] Figure 22 shows the structure of the cations where X is O, NH, or NSO. 2 or CH 2 FIG. 23 shows an example of a fourth design of fully functional cyclic loop nucleotide (CLN-4) with an alpha phosphate-allyl linked asymmetric loop and a stop construct on the loop. FIG. 23 shows an example of a synthesis process of fully functional cyclic loop nucleotide (CLN-4A) shown in FIG. 22 with X=O and Y=O. FIG. 24 shows an example of a synthesis process of fully functional cyclic loop nucleotide (CLN-4B) shown in FIG. 22 with X=NH and Y=O. FIG. 25 shows an example of a synthesis process of fully functional cyclic loop nucleotide (CLN-4C) shown in FIG. 22 with X=CH. 2 23A-23C show schematic diagrams of an exemplary synthesis process for a fully functional circular loop nucleotide (CLN-4C) shown in FIG. 22, in which A is cyclic nucleotide and Y=O.
[0241] Figure 26 shows the structure of the cations where X is O, NH, or NSO. 2 or CH 2FIG. 27 shows an example of a fifth design of fully functional cyclic loop nucleotide (CLN-5) with an alpha phosphate linked asymmetric loop and a stop construct on the loop, where X=O and Y=O. FIG. 28 shows an example of a synthetic process for fully functional cyclic loop nucleotide (CLN-5B) shown in FIG. 26 with X=NH and Y=O. FIG. 29 shows an example of a synthetic process for fully functional cyclic loop nucleotide (CLN-5C) shown in FIG. 27 with X=CH. 2 and Y=O. Figure 30 is a schematic diagram of an exemplary synthesis process for a fully functional cyclic loop nucleotide (CLN-5D) shown in Figure 26, where X=O and Y=S, which is compatible with phosphorothioate cleavage.
[0242] FIG. 31 shows a schematic diagram of a sixth example design of a cyclic loop nucleotide (CLN-6), where X is O, NH, NSO 2 or CH 2 and Y can be O, S or NH, with a peptide-based cyclic loop and a stop construct on the nucleobase. The peptide-based cyclic loop can include various amino acid residues, such as arginine, histidine, lysine, glutamic acid, aspartic acid, cysteine, tyrosine, asparagine, tryptophan, leucine, and alanine. The amino acid side chain Z on the cyclic loop is shown in FIG. 31.
[0243] Some of the further examples below relate to symmetric cyclic loop nucleotides and cyclic loop k-mer oligonucleotides, which may include nucleoside bases, non-nucleoside residues, sp9, sp18, abasic nucleotides, C3 moiety residues, or peptide residues. Some of the further examples below relate to asymmetric cyclic loop nucleotides and cyclic loop k-mer oligonucleotides, which may include nucleoside bases and non-nucleoside residues.
[0244] Figure 32 shows the structure of the cations where X is O, NH, or NSO.2 or CH 2 32 shows a schematic diagram of a seventh example of a design of a cyclic loop nucleotide (SNL-1) that may be, with an alpha phosphate-linked symmetric nucleoside loop and a stop construct on the loop. In some embodiments, the cyclic loop may include one or more different nucleobases. For example, the nucleoside may include inosine, nitroindole, and the like. In some embodiments, other nucleoside bases with modified sugars, such as locked nucleic acid (LNA), 2'-Ome, or 2'-F, may be used in the cyclic loop. Non-limiting nucleoside base structures are also shown in FIG. 32.
[0245] FIG. 33 shows a schematic diagram of an example of an eighth design of a cyclic loop nucleotide (SNNL-1), where X is O, NH, or CH 2 and has an alpha phosphate linked symmetric non-nucleoside loop and a stop construct on the loop. The cyclic loop may include one or more of the following moieties: [ka] Alkyl (C3-C12) phosphates, polyethylene glycol phosphates, other modified non-nucleoside moieties such as spermine, phosphorothioates, or methylphosphonates may be used in the cyclic loop in some embodiments.
[0246] Figure 34 shows the structure of the cations where X is O, NH, or NSO. 2 or CH 2 13A-C are schematic diagrams showing an example of a ninth design of a cyclic loop nucleotide (SPL-1) that may be: α-phosphate-linked symmetric peptide loop and a stop construct on the loop. Other modified amino acid residues or moieties that are compatible for use in peptide synthesis, such as PEG2, 6, 11, 12, or Lys(FITC), may be used in some embodiments.
[0247] Figure 35 shows the structure of the nucleus where X is O, NH, or NSO. 2 or CH 210 is a schematic diagram of an example of a tenth design of a cyclic loop nucleotide (SNL-2) that may be: α-phosphate-allyl linked symmetric nucleoside loop and stop construct on the loop. Other modified nucleoside bases, such as inosine, nitroindole, LNA, 2'-Ome, or 2'-F, may be used in some embodiments.
[0248] FIG. 36 shows a schematic diagram of an example of an eleventh design of a cyclic loop nucleotide (SNNL-2), where X is an O, NSO with an α-phosphate-allyl linked symmetric non-nucleoside loop and a stop construct on the loop. 2 or CH 2 Other modified non-nucleoside moieties, such as alkyl (C3-C12), spermine, phosphorothioate, or methylphosphonate, may be used in some embodiments.
[0249] Figure 37 shows the structure of the nucleus where X is O, NH, or NSO. 2 or CH 2 13 shows a schematic of an example of a twelfth design of a cyclic loop nucleotide (SPL-2) that may be: α-phosphate-allyl linked symmetric peptide loop and a stop construct on the loop. Other modified amino acid residues or moieties that are compatible for use in peptide synthesis, such as PEG2, 6, 11, 12, or Lys(FITC), may be used in some embodiments.
[0250] Figure 38 shows the structure of the nucleus where X is O, NH, or NSO. 2 or CH 213 is a schematic diagram of an example of a thirteenth design of a cyclic looped nucleotide (ASL-1) that may be a cyclic looped nucleotide having an alpha phosphate-linked asymmetric loop and a stop construct on the loop. The asymmetric loop may be formed from a combination of nucleoside bases, non-nucleoside moieties, and amino acid residues as shown in the figure. Other modified nucleoside bases, such as inosine, nitroindole, LNA, 2'-Ome, or 2'-F, may be used in some embodiments. Other modified non-nucleoside moieties, such as alkyl (C3-C12), spermine, phosphorothioate, or methylphosphonate, may be used in some embodiments. Other modified amino acid residues or moieties that are compatible with use in peptide synthesis, such as PEG2, 6, 11, 12, or Lys(FITC), may be used in some embodiments.
[0251] FIG. 39 shows a schematic diagram of a fourteenth design example of a cyclic loop nucleotide (ASL-2) having an alpha phosphate-C5′ allyl linked asymmetric loop and a stop construct on the loop, where X is O, NH, NSO 2 or CH 2 The asymmetric loop may be formed from a combination of nucleoside bases, non-nucleoside moieties, and amino acid residues, as shown in Figure 39. Other modified nucleoside bases, such as inosine, nitroindole, LNA, 2'-Ome, or 2'-F, may be used in some embodiments. Other modified non-nucleoside moieties, such as spC12, spermine, phosphorothioate, or methylphosphonate, may be used in some embodiments. Other modified amino acid residues or moieties that are compatible for use in peptide synthesis, such as PEG2, 6, 11, 12, or Lys(FITC), may be used in some embodiments.
[0252] Figure 40 shows the formula for X being O, NH, and NSO. 2 or CH 240, which is a schematic diagram of an example of a fifteenth design of a cyclic loop nucleotide (ASL-3) having an alpha phosphate-C5' allyl linked asymmetric loop and a stop construct on the nucleobase. The asymmetric loop can be formed from a combination of nucleoside bases, non-nucleoside moieties, and amino acid residues as shown in FIG. 40. Other modified nucleoside bases, such as inosine, nitroindole, LNA, 2'-Ome, or 2'-F, can be used in some embodiments. Other modified non-nucleoside moieties, such as alkyl (C3-C12), spermine, phosphorothioate, or methylphosphonate, can be used in some embodiments. Other modified amino acid residues or moieties that are compatible with use in peptide synthesis, such as PEG2, 6, 11, 12, or Lys(FITC), can be used in some embodiments.
[0253] Figure 41 shows a schematic of a sixteenth design example of a circular loop tetramer oligonucleotide (tetramer-SNL-1) with a tetramer oligo with a symmetric nucleoside loop and a stop construct on the loop. The loop binding position may be 1-2, 1-3, 1-4, 2-3, 2-4, or 3-4 bases. The cleavage chemistry may be RNA cleavage or allylic cleavage; the cleavage position may be the 2nd, 3rd, or 4th base.
[0254] Figure 42 shows a schematic of a seventeenth design example of a circular loop tetramer oligonucleotide (tetramer-SNNL-1) with a tetramer oligo with a symmetric non-nucleoside loop and a stop construct on the loop. The loop binding position may be 1-2, 1-3, 1-4, 2-3, 2-4, or 3-4 bases. The cleavage chemistry may be RNA cleavage or allylic cleavage; the cleavage position may be the 2nd, 3rd, or 4th base.
[0255] Figure 43 shows a schematic of an example of an eighteenth design of a circular loop tetramer oligonucleotide (tetramer-SPL-1) with a tetramer oligo with a symmetric peptide loop and a stop construct on the loop. The loop binding position may be 1-2, 1-3, 1-4, 2-3, 2-4, or 3-4 bases. The cleavage chemistry may be RNA cleavage or allyl cleavage; the cleavage position may be the 2nd, 3rd, or 4th base.
[0256] Figure 44 shows a schematic of a nineteenth design example of a circular loop tetramer oligonucleotide (tetramer-ASL-1) with a tetramer oligo with an asymmetric loop and a stop construct on the loop. The asymmetric loop can be formed from a combination of nucleoside bases, non-nucleoside moieties, and amino acid residues. The loop binding position can be 1-2, 1-3, 1-4, 2-3, 2-4, or 3-4 bases. The cleavage chemistry can be RNA cleavage or allylic cleavage; the cleavage position can be the 2nd, 3rd, or 4th base.
[0257] FIG. 45 shows generally an exemplary synthesis process for one embodiment of a tetrameric oligonucleotide that can be used to generate a cleavable circular loop tetrameric oligonucleotide.
[0258] Example 1: Synthesis of nucleotides with 5'-allyl termination constructs and incorporation into modified polynucleotides The synthesis of nucleotides bearing 5'-allyl stop constructs and the subsequent demonstration that the modified nucleotides can be incorporated by a polymerase followed by cleavage of the allyl group is described.
[0259] Figure 7 shows the unexpectedly successful experimental results of polymerase incorporation of allyl dTTP. The polymerase incorporated "allyl dTTP" and "ext allyl dTTP" into the template and primer combination for up to 10 consecutive incorporations, as indicated by the arrows in Figure 7. The reaction conditions and the definitions of "allyl dTTP" and "ext allyl dTTP" are shown in the right panel of Figure 7. Additionally, "A" and "B" refer to the optically pure diastereomers arising from the C5' allyl carbon center.
[0260] FIG. 7 shows 12 lanes in a gel image. Lanes 1 (starting from the left), 3, 5, 7, 9 and 11 were loaded in a 1 hour reaction run; lanes 2, 4, 6, 8, 10 and 12 were loaded in a 22 hour reaction run. Lanes 1 and 2 show controls containing native dTTP without a stop construct (note, bands above +10 are seen with the use of homopolymer A templates, a common observation). Lanes 3 and 4 show: one isomer of allyl dTTP was used ("Allyl A dTTP"), the isomer arising from the C-stereogenic center (marked with a "+" in the structure shown in the right panel of FIG. 7). Incorporation up to +7 was observed after 1 hour. The extended duration of 22 hours indicates exo activity. Lanes 5 and 6 show the use of one isomer of ext allyl dTTP ("Ext Allyl A dTTP"). Incorporation up to +7 was observed after 22 hours. The slower incorporation is likely due to steric gain. Lanes 7 and 8 show the use of another isomer of allyl dTTP ("Allyl B dTTP"). Incorporation up to +10 was observed after 22 hours. Lanes 9 and 10 show the use of another isomer of ext allyl dTTP ("Ext Allyl B dTTP"). Incorporation up to +7 was observed after 22 hours. Lanes 11 and 12 show the negative control where no nucleotide was added.
[0261] FIG. 8 shows the unexpectedly successful experimental results of cleavage of the allyl group. After incorporation of the nucleotide, the sample was treated with 100 mM Pd / THP solution at 37° C. for 1 hour to cleave the allyl bond. Lanes 1, 3, 5, 7, and 9 were loaded with samples after incorporation but not treated with Pd / THP. Lanes 2, 4, 6, 8, and 10 were loaded with samples after treatment with Pd / THP. Lanes 1 and 2 show the absence of allyl groups and therefore no change with or without treatment with Pd / THP. Lanes 3 vs. 4, 5 vs. 6, 7 vs. 8, and 9 vs. 10 show that treatment with Pd / THP reduces the higher band back to baseline, demonstrating successful cleavage of the allyl group.
[0262] Example 2: Use of Cross-Linked PS Nucleotides to Generate Modified Polynucleotides and Subsequent Cleavage The synthesis of modified nucleotides with bridging PS substitutions and the subsequent demonstration that the modified nucleotides can be incorporated by polymerases with subsequent cleavage of the phosphorus-sulfur bond are described.
[0263] a. Synthesis of cross-linked PS nucleotides [ka]
[0264] Heavy atom substituted nucleotides having bridged PS substitutions are made according to the synthetic scheme shown in Figure 46. Figures 47A and 47B show mass and NMR spectra of examples of such nucleotides. [ka]
[0265] Bifunctional nucleotides with bridged PS substitutions can also be made according to the reaction scheme in FIG.
[0266] b. Incorporation of heavy atom substituted nucleotides to generate modified polynucleotides
[0267] Figure 49 shows the results of incorporation of a polyA tail formed by cross-linking PS (phosphorothiolate) deoxyribonucleotides (nucleotide 1) to generate a DNA oligonucleotide with a continuous phosphorothiolate backbone. In particular, incorporation was tested using multiple polymerases including Pol812, 1901, hPrimpol, and BSU. A positive signal for each heavy atom substituted nucleotide incorporated into the oligonucleotide strand was observed for each tested polymerase.
[0268] C. Cleavage of phosphorothiolate heavy atom modified nucleotides
[0269] Figure 50A is a reaction scheme for cleavage of the phosphorus-sulfur bond by silver nitrate through a cross-linked PS nucleotide 1. Silver nitrate was added to nucleotide 1, and 15 minutes later DTT was added and incubated for 1 hour. Cleavage by use of silver nitrate is specific to the PS bond, thereby generating the 5'-S nucleoside (shown in Figure 50A). Figure 50B shows HPLC chromatographs of the nucleotide (control) and cleavage products after addition of DTT, as well as ESI-MS data of the cleavage products.
[0270] The above cleavage reaction was applied to the primer extended using cross-linked PS nucleotide 1 and DNA polymerase Pol1901 (see Figure 51, lane 4). Gel analysis showed that silver nitrate and silver nitrate + DTT could cleanly cleave the oligonucleotide as shown in Figure 51, lanes 5 and 6, respectively.
[0271] Example 3: Generation of tetrameric oligonucleotides with bridging PS substitutions The synthesis and characterization of oligonucleotides with PS cross-linking substitutions are described.
[0272] a. Synthesis of Cross-Linked PS Oligonucleotides
[0273] The bridged PS substitutions found in tetrameric oligonucleotides are structurally located similarly to bridged PS nucleotides. The number of bridged PS substitutions that a k-mer oligonucleotide can introduce into a growing strand each time can be any number between 1 and k-1. For example, for a tetrameric oligonucleotide, there can be 1 to 3 bridged PS substitutions. PS substitutions were introduced into tetrameric oligonucleotides via 5'SDMT phosphoramidite 9. Figure 52 shows the schematic synthesis of 5'SDMT phosphoramidite 9 by three methods. 5'SDMT phosphoramidite 9 is then introduced into a tetrameric oligonucleotide and ligated to generate a substitution into the tetrameric oligonucleotide via 5'SDMT phosphoramidite 9. As shown in Figure 53A, post-synthetic modifications can be added, such as attaching a circular loop onto the modified oligonucleotide product. A tetrameric oligonucleotide with a bridged PS substitution and a circular loop attached to the first and last nucleobases of the tetrameric oligonucleotide (oligonucleotide 10) is shown in Figure 53B. Figures 54A-C show the HPLC, IR, and MS characterization of the resulting circular loop tetramer oligonucleotides.
[0274] b. Ligation / cleavage of cross-linked PS oligonucleotides
[0275] Wild-type ligase Ame was used to ligate up to 10 circular loop modified tetrameric oligonucleotides 10. Figure 55 shows that wild-type ligase Ame was able to ligate up to 10 circular loop tetrameric oligonucleotides 10 after 2.5 hours, with improved ligation efficiency at 22 hours. The signal from the ligated oligonucleotides can be observed most strongly in lane 6, which corresponds to 22 hours.
[0276] FIG. 56 shows that upon successive ligations, cleavage of the newly synthesized strands generates longer strands that migrate slower (higher) in the gel due to the effect of the incorporated PEG loops (and subsequent elongation following selective cleavage of the polynucleotide backbone).
[0277] Thus, Example 3 shows that the synthesis of circular loop tetramer oligonucleotide 10 was successful, and that the circular loop remained closed during the ligation process. Furthermore, ligation was successful in generating a newly synthesized strand, and finally, the newly synthesized strand can be selectively and efficiently cleaved using silver nitrate to generate an extended strand.
[0278] Example 4: Use of non-crosslinked PS nucleotides to generate modified polynucleotides and subsequent cleavage a. Synthesis of non-crosslinked PS nucleotides [ka] Modified nucleotides must be synthesized prior to enzymatic incorporation. The phosphorothioate (non-bridged PS) triphosphate nucleotides shown above were prepared according to the synthesis scheme in Figure 57A. The phosphorothioate (non-bridged PS) triphosphate nucleotides functionalized with the alpha phosphate were prepared according to the synthesis scheme in Figure 57B.
[0279] b. Incorporation of non-crosslinked PS nucleotides
[0280] Non-bridged PS nucleotides were incorporated by several wild-type polymerases shown in Figure 58, including 812, Dpo4, DbH, DinB, Asfy, Mu, λi, 8kbeta, human PrimPoli, SulfoPrimPol, 300 Cyan PrimPol, 190 Cyan PrimPol, BSU, RB49, and PolC.
[0281] Example 5: Generation of tetrameric polynucleotides with non-bridging PS substitutions The synthesis and characterization of polynucleotides having PS non-bridging substitutions is described herein.
[0282] a. Synthesis of non-crosslinked PS oligonucleotides 59 and 60A-D show schematic diagrams of tetrameric oligonucleotide 13 synthesized by generating phosphorothioate modifications at desired sites using a sulfurizing agent such as DDT and separating the resulting diastereomers by HPLC.
[0283] b. Ligation of non-crosslinked PS oligonucleotides
[0284] FIG. 61 shows that wild-type Ame and PsyA ligases were able to perform up to 10 sequential ligation events with the circular loop tetramer oligonucleotide 13 within 20 hours.
[0285] c. Cleavage of non-crosslinked PS oligonucleotides
[0286] Figure 62A shows the use of iodine to selectively cleave the bridge PO bond at the site of the phosphorothioate in the presence of a nucleophile such as an amine. However, the subsequent bridge PO bond cleavage is not specific and either the 5' or 3' end can potentially be cleaved, giving four potential products (Figure 62B).
[0287] Example 6: Nucleotides with imino-P substitutions at the alpha phosphate Figure 63 shows a schematic of a model synthesis of imino-P substitution on the alpha phosphate of a nucleotide. Imino-P nucleotides were synthesized using 50 mM KOH at pH 5.5 and 7.5. 2 PO 4 It was stable in buffer for over 2 days at 25° C. No decomposition of Imino-P was observed after 2 days at both pH 5.5 and 7.5.
[0288] Figure 64 shows the synthesis of cyclic loop nucleotides with imino-p substitutions according to embodiments herein. Three steps are shown: precursor synthesis (introduction of an allyl cleavable group), imino-P synthesis, and deprotection of the triphosphate nucleotide, followed by an exemplary cyclic loop attachment using a bis-azido PEG20 linker. Imino-P nucleotides were synthesized using 50 mM KHPO at pH 5.5 and 7.5. 2 PO 4It was stable in buffer for 2 days at 25° C. Figure 65 shows the HPLC and LCMS characterization of imino-P nucleotides.
[0289] Example 7: Access to Imino-P Allyl Bifunctional Triphosphates Various methods for accessing imino-p-bifunctional triphosphates were evaluated.
[0290] Figure 66 shows a schematic of one embodiment of a method for synthesizing an imino-P allyl bifunctional nucleotide. A 5'-OH allyl nucleoside 1 is activated by a phosphitylating agent to form a first intermediate 2, which is subsequently activated and coupled to a pyrophosphate to form a second phosphite triester intermediate 3. A Staudinger reaction with an azide moiety occurs to convert the unstable phosphite triester 3 to a third imino-P phosphotriester intermediate 4. Deprotection then removes all orthogonal protecting groups to form the imino-P allyl bifunctional triphosphate 6.
[0291] Figure 67 shows a schematic of one embodiment of a method for synthesizing an imino-P allyl bifunctional nucleotide. A 5'-OH allyl nucleoside 1 is coupled to an activated phosphitylating agent to form a first cyclic phosphite triester intermediate 2, followed by a Staudinger reaction with an azide functionality to form a second imino-P phosphotriester intermediate 4. Subsequent deprotection removes all orthogonal protecting groups to form the imino-P allyl bifunctional triphosphate 6.
[0292] Figure 68 shows a schematic of one embodiment of a method for synthesizing an imino-P allyl bifunctional nucleotide. A 5'-OH allyl nucleoside 1 is coupled to a phosphitylating agent and hydrolyzed to form a first H-phosphonate intermediate 2, which is subjected to successive BSA treatments and a Staudinger reaction with an azide function to form a second imino-P phosphotriester intermediate 3. Subsequent coupling with pyrophosphate forms a third intermediate 4. Deprotection removes all orthogonal protecting groups to form the imino-P allyl bifunctional triphosphate 6.
[0293] Figure 69 shows a schematic of one embodiment of a method for synthesizing an imino-P allyl bifunctional nucleotide. The 5'-OH allyl nucleoside 1 is activated and coupled with a phosphitylating agent to form a first phosphite triester intermediate 2. It then undergoes a Staudinger reaction with an azide functionality to form a second imino-P phosphotriester intermediate 3. Deprotection is performed to form a third intermediate 4, which is reactivated (see Table 2) and coupled to a pyrophosphate moiety to form a triphosphate intermediate 5. Deprotection removes all orthogonal protecting groups to form the imino-P allyl bifunctional triphosphate 6.
[0294] Figure 70 shows a schematic of one embodiment of a method for synthesizing an imino-P allyl bifunctional nucleotide. The method involves activation of the pyrophosphate moiety (as opposed to the nucleoside monophosphate), thereby exchanging the roles of the nucleophilic and electrophilic groups. Activation of the pyrophosphate 7 forms a first reactive pyrophosphate intermediate 8, which is then coupled to the nucleoside monophosphate 3 to form a second nucleoside triphosphate intermediate 5. Deprotection removes all orthogonal protecting groups to form the imino-P allyl bifunctional triphosphate 6.
[0295] FIG. 71 illustrates an exemplary synthesis of bifunctional imino-P allyl nucleotides with different reactive groups.
[0296] Figure 72 shows a schematic of the deprotection of the 3'-OTBDPS group where a 3'-OH is formed, generating nucleotide 6. Additionally, deprotection methods can be utilized to deprotect groups such as 3'-OTBDPS, including those selected from the following: HF, TEA, HF-pyridine, TBAF, TAST, DBU, acetyl chloride / dry MeOH, selectfluor, or lithium acetate.
[0297] Figure 73 is a table showing the HF-TEA and TBAF deprotection methods to assess the conversion efficiency and yield of the desired triphosphate product. Using the optimized HF-TEA deprotection conditions in entry 3, approximately 77% conversion of 3'-OTBDPS bifunctional nucleotide 5 to 3'-OH bifunctional nucleotide 6 was achieved.
[0298] The crude HPLC, analytical HPLC and LCMS spectra of the purified nucleotide produced using entry 3 of the table in Figure 73 are shown in Figures 74A-C, respectively. Deprotection of 3'-OTBDPS by treatment with 1MTBAF for 2 hours at 4°C was also successful. The recovered nucleotide 6 was similarly characterized by LCMS and HPLC to confirm identity and purity.
[0299] Example 8: Selection of diastereomeric mixtures. [ka] Desymmetrization of the substituents around the alpha phosphorus and 5'-carbon atoms resulted in the creation of two new chiral centers ( * (denoted by the symbol) is introduced, resulting in the formation of four possible diastereomers. Methods for controlling chirality and selecting specific isomers are described herein.
[0300] FIG. 75 illustrates diagrammatically the stereoselective reduction of the carbonyl group and thus the control of chirality at the 5′-carbon atom.
[0301] FIG. 76 shows a schematic diagram of the enzymatic kinetic diastereoselection process from racemic precursors.
[0302] FIG. 77 shows a schematic of the chiral derivatization of the isomers to allow for final column separation.
[0303] FIG. 78 illustrates generally chiral ligand-promoted stereoselective alkyl addition (eg, nucleophilic addition to an aldehyde group), as well as representative examples of chiral ligands.
[0304] FIG. 79 shows a schematic of the stereoselective enzymatic synthesis of triphosphates.
[0305] FIG. 80 shows a schematic of a method for controlling chirality at the alpha phosphorus atom, in particular the stereoselective Staudinger reaction induced by a chiral auxiliary. FIG. 81 shows an exemplary list of Staudinger variants. Several variations of the Staudinger reaction can be considered to modify the substituents at the alpha phosphorus center. One embodiment described herein generates an "imino-P" substitution involving the Staudinger reaction between a phosphite and a sulfonyl azide moiety. One skilled in the art would be able to extend the Staudinger reaction to other variations such as phosphites with aryl azides, phosphites with guanidine azides, and H-phosphonates with acyl azides. Further exemplary synthetic methods for preparing azide variants are provided in FIG. 82.
[0306] Example 9: Cyclic Loop Nucleotide Synthesis Route Figures 83A and 83B show schematic diagrams of various pathways leading to the formation of an exemplary looped nucleotide structure 10 using three alternative pathways: Path A: 3→11→9→10, Path B: 3→5→9→10, and Path C: 3→5→6→10. Figures 84A-C show characterized HPLC, LCMS, and FTIR characterization data for the circular looped nucleotide structure 10. LCMS confirms the mass of the looped nucleotide 10, the HPLC spectrum shows the purity of the isolated product, and FTIR confirms the absence of unreacted azide (appears at approximately 2100 cm-1 corresponding to the -N=N=N stretch).
[0307] Example 10: Preliminary incorporation assay of cyclic loop alpha phosphate-substituted nucleotides Figure 85 shows the results of bifunctional nucleotide 6 tested in an incorporation assay using Dpo4 enzyme compared to native dTTP. The results show that bifunctional nucleotide 6 was incorporated at least twice under the conditions evaluated.
[0308] Similarly, looped nucleotide 10 was evaluated as shown in Figure 86. Compared to the incorporation of nucleotide 6, the incorporation of the sterically bulky looped nucleotide 10 was comparatively slower. As shown in Figure 87, increasing the concentration of looped nucleotide 10 by 5-fold (from 20 μM to 100 μM) increased the catalytic Mn 2+ A five-fold increase in the concentration of (from 2 mM to 10 mM), a two-fold increase in the concentration of Dpo4 (from 1 μM to 2 μM), and an increase in temperature (from room temperature to 37°C) clearly improved incorporation of looped nucleotide 10.
[0309] Example 11: Preliminary stability data for looped alpha phosphate substituted nucleotides. Figures 88-89 show stability assays performed on bifunctional nucleotide 6 and looped nucleotide 10, which were subjected to stability testing at 25°C and 60°C in 50 mM Tris pH 7.5. Both nucleotides 6 and 10 were stable for up to 2 days at 25°C, but about 55% and about 74% of nucleotides 6 and 10, respectively, were degraded at 60°C. About 72% of the α PC modified nucleotides were degraded over 2 days in 50 mM Tris pH 7.5 at 25°C. This indicates that imino-P modification can significantly stabilize α phosphate substituted nucleotide triphosphates. This enhanced stability is important for the manufacture, scalability and productivity of similar nucleotides for sequencing purposes.
[0310] Additional Notes It should be understood that all combinations of the foregoing concepts and further concepts discussed in more detail below (unless such concepts are mutually inconsistent) are contemplated as part of the inventive subject matter disclosed herein. In particular, all combinations of claimed subject matter appearing at the end of this disclosure are contemplated as part of the inventive subject matter disclosed herein. It should also be understood that the terms expressly used in this specification and which may also appear in any disclosures incorporated by reference should be given the meaning most consistent with the particular concepts disclosed herein.
[0311] References throughout this specification to "one example," "another example," "an example," etc. mean that particular elements (e.g., features, structures, and / or characteristics) described in connection with an example are included in at least one example described herein and may or may not be present in other examples. In addition, unless the context clearly dictates otherwise, it should be understood that the described elements with respect to any example may be combined in any suitable manner in the various examples.
[0312] Ranges provided herein should be understood to include the stated range and any value or subrange within that stated range, as if such value or subrange were expressly recited. For example, a range of about 2 nm to about 20 nm should be interpreted to include not only the explicitly recited limits of about 2 nm to about 20 nm, but also individual values, such as about 3.5 nm, about 8 nm, 18.2 nm, etc., and subranges, such as about 5 nm to about 10 nm, etc. Additionally, when "about" and / or "substantially" are used to express a value, this is meant to encompass slight variations (up to ±10%) from the stated value.
[0313] Although several embodiments have been described in detail, it should be understood that the disclosed examples may be modified, and therefore the foregoing description should be considered as non-limiting.
[0314] Although certain specific examples are described, these examples are presented as examples only and are not intended to limit the scope of the disclosure. Indeed, the novel methods and systems described herein may be embodied in a variety of other forms. Furthermore, various omissions, substitutions, and modifications of the systems and methods described herein may be made without departing from the spirit of the disclosure. The appended claims and their equivalents are intended to cover such forms or modifications as fall within the scope and spirit of the disclosure.
[0315] It shall be understood that features, materials, properties, or groups described in connection with a particular embodiment or example are applicable to any other embodiment or example described in this section or elsewhere in this specification, unless incompatible therewith. All of the steps of any method or process as disclosed herein (including any accompanying claims, abstract and drawings), and / or as disclosed, may be combined in any combination, except combinations where at least some of such features and / or steps are mutually exclusive. Protection is not limited to the details of any preceding example. Protection extends to any novel one or any novel combination of features disclosed herein (including any accompanying claims, abstract and drawings), or to any novel one or any novel combination of steps of any method or process as disclosed.
[0316] Certain features that are described in the present disclosure in the context of separate implementations may also be implemented in combination in a single implementation. Conversely, various features that are described in the context of a single implementation may also be implemented in multiple implementations separately or in any suitable subcombination. Furthermore, even if features may be described above as functioning in a particular combination, one or more features from a claimed combination may in some cases be deleted from the combination, and the combination may be claimed as a subcombination or a variation of the subcombination.
[0317] Furthermore, although operations may be depicted in the figures or described herein in a particular order, it is not necessary that such operations be performed in the particular order shown or sequentially, or that all operations be performed, to achieve desirable results. Other operations not depicted or described may be incorporated into the example methods and processes. For example, one or more additional operations may be performed before, after, simultaneously with, or between any of the described operations. Furthermore, operations may be rearranged or reordered in other implementations. Those skilled in the art will appreciate that in some embodiments, the actual steps taken in the illustrated and / or disclosed processes may differ from those shown in the figures. Depending on the embodiment, certain steps described above may be removed, or others may be added. Furthermore, features and attributes of certain embodiments disclosed above may be combined in different ways to form additional embodiments, all of which fall within the scope of the present disclosure. Also, it should be understood that the separation of various system components of the implementations described above should not be understood as requiring such separation in all implementations, and that the described components and systems may typically be integrated together in a single product or packaged within multiple products. For example, any of the components for the energy storage systems described herein can be provided separately or integrated together (e.g., packaged together or attached together) to form an energy storage system.
[0318] For purposes of this disclosure, certain aspects, advantages, and novel features have been described herein. Not necessarily all such advantages may be achieved in accordance with a particular embodiment. Thus, for example, a person skilled in the art will recognize that the present disclosure may be embodied or carried out in a manner that achieves one advantage or group of advantages taught herein, without necessarily achieving other advantages that may be taught or suggested herein.
[0319] Unless otherwise stated, conditional language such as "can," "could," "may," or "might," unless otherwise stated or understood within the context in which it is used, is intended to convey that, in general, a particular embodiment includes certain features, elements, and / or steps and other embodiments do not. Thus, such conditional language is not intended to imply that features, elements, and / or steps are in any way required by one or more embodiments, or that one or more embodiments necessarily include logic that determines whether those features, elements, and / or steps should be included or performed in any particular embodiment, with or without user input or prompting.
[0320] Conjunctive language, such as the phrase "at least one of X, Y, and Z," is understood in the context as it is commonly used, unless otherwise noted, to convey that an item, term, etc. can be either X, Y, or Z. Thus, such conjunctive language is not generally intended to imply that a particular embodiment requires the presence of at least one of X, at least one of Y, and at least one of Z.
[0321] As used herein, language of degree, such as the terms "approximately," "about," "generally," and "substantially," refers to a value, amount, or characteristic that approaches a stated value, amount, or characteristic that still performs a desired function or achieves a desired result.
[0322] The scope of the present disclosure is not intended to be limited by the specific disclosure of preferred embodiments in this section or elsewhere herein, but may be defined by the claims presented in this section or elsewhere herein, or presented in the future. The language of the claims is to be interpreted broadly based on the language used in the claims, and is not to be limited to the embodiments described in this specification or during the prosecution of this application, which embodiments are to be construed as non-exclusive.
[0323] Although the foregoing invention has been described with respect to certain preferred embodiments, other embodiments will be apparent to those skilled in the art. Moreover, other combinations, omissions, substitutions and modifications will be apparent to those skilled in the art in light of the disclosure herein. Thus, the present invention is not limited by the description of the preferred embodiments, but is instead defined by reference to the appended claims. All references cited herein are incorporated by reference in their entirety.
[0324] The terms used in the description presented herein are not intended to be interpreted in any limiting or restrictive manner, and refer to their ordinary meaning as would be understood by a person skilled in the art in light of the present specification, unless otherwise indicated. Furthermore, the embodiments may include, consist of, or consist essentially of several novel features, none of which are solely responsible for its desirable attributes or considered essential to carrying out the embodiments described herein. As used herein, the section headings are for organizational purposes only and should not be interpreted in any way as limiting the subject matter described. All literature and similar materials cited in this application, including but not limited to patents, patent applications, articles, books, papers, and Internet web pages, are expressly incorporated by reference in their entirety for any purpose. In the event that the definition of a term in the incorporated references differs from the definition provided in the present teachings, the definition provided in the present teachings shall prevail. Since "about" is implied before the temperature, concentration, time, etc. discussed in the present teachings, it will be understood that slight and insubstantial deviations are within the scope of the present teachings herein.
[0325] Although the present disclosure is in the context of specific embodiments and examples, those skilled in the art will understand that the present disclosure extends beyond the specifically disclosed embodiments to the use of other alternative embodiments and / or embodiments, as well as obvious modifications and equivalents thereof. In addition, while several variations of the embodiments have been shown and described in detail, other modifications that are within the scope of the present disclosure will be readily apparent to those skilled in the art based on the present disclosure. It is also contemplated that various combinations or subcombinations of specific features and aspects of the embodiments can be made and still fall within the scope of the present disclosure. It should be understood that various features and aspects of the disclosed embodiments can be combined with or substituted for one another to form various modes or embodiments of the present disclosure. It is therefore intended that the scope of the present disclosure disclosed herein should not be limited by the specific disclosed embodiments described above.
Claims
1. The structure: 【Chemistry 1】 (In the formula, X is -O-, -CH 2 -, -NH-, 【Chemistry 2】 and X' is =N-SO 2 -, =NH-CO-, or 【Chemistry 3】 and Y is -O-, -S-, -NH-, or -Se-; L 1 is a first linking group, L 2 is a second linking group, SP is a spacer. A compound having one of the following:
2. The SP: (1) an alkyl chain having 5 to 50 carbon atoms; (2) Oligonucleotides, modified oligonucleotides or polyphosphates having 1 to 100 repeating units; (3) A polypeptide having 1 to 100 repeat units; (4) a hydrophilic polymer having 1 to 100 repeat units, and (5) A hydrophobic polymer having 1 to 100 repeating units.
3. 3. The compound of claim 2, wherein the hydrophilic polymer is selected from the group consisting of polyethylene glycol, polyvinyl alcohol, polyacrylamide, polyvinylpyrrolidone, polystyrene sulfonate, and polyethyleneimine.
4. The compound of claim 2 , wherein the hydrophobic polymer is selected from the group consisting of polylactic acid, polymethyl methacrylate, and polystyrene.
5. L 1 and L 2 each independently comprises a conjugated moiety selected from the group consisting of amine-NHS ester, amine-imido ester, amine-pentafluorophenyl ester, amine-hydroxymethylphosphine, carboxyl-carbodiimide, thiol-maleimide, thiol-haloacetyl, thiol-pyridyl disulfide, thiol-thiosulfonate, thiol-vinyl sulfone, aldehyde-hydrazide, aldehyde-alkoxyamine, hydroxy-isocyanate, azido-alkyne, azido-phosphine, transcyclooctene-tetrazine, norbornene-tetrazine, azido-cyclooctyne, and azido-norbornene.
6. L 1 and L 2 6. The compound of claim 5, wherein each of further independently comprises a first linker between the conjugate moiety and X / X', and a second linker between the conjugate moiety and SP.
7. 7. The compound of claim 6, wherein the first linker and the second linker are independently selected from the group consisting of a hydrophilic polymer, a hydrophobic polymer, an oligonucleotide, a peptide, a polypeptide, an aliphatic chain (C5-C50), and combinations thereof.
8. The compound of any one of claims 1 to 7, wherein the SP further comprises a stop construct.
9. The compound according to any one of claims 1 to 7, wherein the base further comprises a stop construct.
10. 10. The compound of claim 8 or 9, wherein the termination construct is a linear, branched or cyclic polymer.
11. The compound of claim 10 , wherein the termination construct comprises a synthetic hydrophobic polymer, a synthetic hydrophilic polymer, an oligonucleotide / polynucleotide, a peptide / polypeptide, or a combination thereof.
12. The structure: 【Chemistry 4】 【Chemistry 5】 (In the formula, X is -O-, -CH 2 -, -NH-, 【Chemistry 6】 and X' is =N-SO 2 -, =NH-CO-, or 【Chemistry 7】 and Y is -O-, -S-, -NH-, or -Se-; R 1 , R 2 , and R 3 one of which is allyl and the other is H; L 1 is a first linking group, L 2 is a second linking group, SP is a spacer. An oligonucleotide comprising one of:
13. Structure (VII) has the following structure: 【Chemistry 8】 The oligonucleotide of claim 12, which can be further represented by:
14. Structure (VIII) has the following structure: 【Chemistry 9】 The oligonucleotide of claim 12, which can be further represented by:
15. Structure (X) is the following structure: 【Chemistry 10】 The oligonucleotide of claim 12, which can be further represented by:
16. The structure: 【Chemistry 11】 (In the formula, Y is -O-, -S-, -NH-, or -Se-; and Y 1 , Y 2 and Y 3 One of the groups is -S- or -Se-, and the other is -O- or -NH-. An oligonucleotide comprising one of:
17. The SP: (1) an alkyl chain having 5 to 50 carbon atoms; (2) an oligonucleotide having 1 to 100 repeat units, a modified oligonucleotide, or a polyphosphate; (3) A polypeptide having 1 to 100 repeat units; (4) A hydrophilic polymer having 1 to 100 repeating units selected from the group consisting of polyethylene glycol, polyvinyl alcohol, polyacrylamide, polyvinylpyrrolidone, polystyrene sulfonate, and polyethyleneimine; and (5) A hydrophobic polymer having 1 to 100 repeating units selected from the group consisting of polylactic acid, polymethyl methacrylate, and polystyrene; The oligonucleotide according to any one of claims 12 to 16, comprising one or more of:
18. L 1 and L 2 each independently comprises a conjugate moiety selected from the group consisting of amine-NHS ester, amine-imido ester, amine-pentafluorophenyl ester, amine-hydroxymethylphosphine, carboxyl-carbodiimide, thiol-maleimide, thiol-haloacetyl, thiol-pyridyl disulfide, thiol-thiosulfonate, thiol-vinyl sulfone, aldehyde-hydrazide, aldehyde-alkoxyamine, hydroxy-isocyanate, azido-alkyne, azido-phosphine, transcyclooctene-tetrazine, norbornene-tetrazine, azido-cyclooctyne, and azido-norbornene.
19. L 1 and L 2 20. The oligonucleotide of claim 18, wherein each of said further comprises, independently, a first linker between said conjugate moiety and X, and a second linker between said conjugate moiety and SP.
20. 20. The oligonucleotide of claim 19, wherein the first linker and the second linker are independently selected from the group consisting of hydrophilic polymers (e.g., polyethylene glycol, polyvinyl alcohol, polyacrylamide, polyvinylpyrrolidone, polystyrene sulfonate, polyethyleneimine), hydrophobic polymers (e.g., polylactic acid, polymethyl methacrylate, polystyrene), oligonucleotides, peptides, polypeptides, aliphatic chains (C5-C50), and combinations thereof.
21. The oligonucleotide of any one of claims 12 to 20, wherein the SP further comprises a stop construct.
22. The oligonucleotide of any one of claims 12 to 20, wherein the base further comprises a stop construct.
23. 23. The oligonucleotide of claim 21 or 22, wherein the stop construct is a linear, branched or circular polymer.
24. 24. The oligonucleotide of claim 23, wherein the termination construct comprises a synthetic hydrophobic polymer, a synthetic hydrophilic polymer, an oligonucleotide / polynucleotide, a peptide / polypeptide, or a combination thereof.
25. L 1 , SP and L 2 The oligonucleotide of any one of claims 12 to 24, wherein:
26. 26. The oligonucleotide of claim 25, wherein the circular loop is symmetric or asymmetric.
27. 27. The oligonucleotide of claim 26, wherein the circular loop is synthesized using one or more of solid phase synthesis, solution phase synthesis, and enzymatic synthesis.
28. 27. The oligonucleotide of claim 26, wherein the circular loop is synthesized using one or more of linear synthesis, branched synthesis, or segmented synthesis.
29. 1. A method for determining a sequence of a polynucleotide in a nanopore-based sequencing system, the method comprising: providing a polynucleotide comprising a plurality of nucleotides, wherein each nucleotide comprises a linker construct, the linker construct having a first end attached to the first position of the nucleotide and a second end attached to the second position of the nucleotide; cleaving a scissionable bond on each of the plurality of nucleotides between the first position and the second position, thereby extending the polynucleotide to form an extended polynucleotide; applying a voltage to cause the elongated polymer to insert into and translocate through the nanopore; (i) detecting and identifying a reporter moiety when the linker construct passes through the nanopore; or (ii) detecting and identifying a base on said nucleotide as said nucleotide passes through said nanopore.
30. 30. The method of claim 29, wherein the linker construct comprises a first linking group, a second linking group, and a spacer between the first linking group and the second linking group.
31. 30. The method of claim 29, wherein the spacer comprises an oligonucleotide having 1-100 repeating units, a modified oligonucleotide having 1-100 repeating units, a polyphosphate having 1-100 repeating units, a polypeptide having 1-100 repeating units, an alkyl chain having 5-50 carbons, a hydrophilic polymer having 1-100 repeating units selected from the group consisting of polyethylene glycol, polyvinyl alcohol, polyacrylamide, polyvinylpyrrolidone, polystyrene sulfonate, and polyethyleneimine, a hydrophobic polymer having 1-100 repeating units selected from the group consisting of polylactic acid, polymethylmethacrylate, and polystyrene, and combinations thereof.
32. 30. The method of claim 29, wherein the spacer comprises the reporter moiety, and the reporter moiety corresponds to and identifies a nucleotide.
33. 33. The method of any one of claims 29 to 32, wherein each of the first and second linking groups independently comprises a conjugated moiety selected from the group consisting of amine-NHS ester, amine-imido ester, amine-pentafluorophenyl ester, amine-hydroxymethylphosphine, carboxyl-carbodiimide, thiol-maleimide, thiol-haloacetyl, thiol-pyridyl disulfide, thiol-thiosulfonate, thiol-vinyl sulfone, aldehyde-hydrazide, aldehyde-alkoxyamine, hydroxy-isocyanate, azido-alkyne, azido-phosphine, transcyclooctene-tetrazine, norbornene-tetrazine, azido-cyclooctyne, and azido-norbornene.
34. 34. The method of any one of claims 29 to 33, wherein the extending polymer further comprises a stopping construct attached to each nucleobase or each linker construct, the stopping construct being configured to slow down, pause or stop the movement.
35. 35. The method of claim 34, wherein the termination construct is a linear, branched or circular polymer.
36. 35. The method of claim 34, wherein the termination construct comprises a synthetic hydrophobic polymer, a synthetic hydrophilic polymer, an oligonucleotide / polynucleotide, a peptide / polypeptide, or a combination thereof.
37. 37. The method of any one of claims 35-36, wherein the nanopore comprises a constriction having an opening with an inner diameter of about 0.6 nm to about 1.2 nm.
38. The method of claim 29, wherein the polynucleotide comprises one of the oligonucleotides according to any one of claims 12 to 24.
39. 30. The method of claim 29, wherein each of said plurality of nucleotides is selected from the compounds according to any one of claims 1 to 11.
40. 30. The method of claim 29, wherein the reporter moiety comprises two or more sub-reporter moieties, and each sub-reporter moiety in the two or more sub-reporter moieties is distinguishable, reproducible and resolvable.
41. 41. The method of claim 40, wherein the two or more sub-reporter moieties comprise a crown ether, a cucurbituril, a pillararene, or a cyclodextrin.
42. A kit for carrying out a method of sequencing a polynucleotide in a nanopore-based sequencing system, said kit comprising a compound according to any one of claims 1 to 11.
43. A system for determining a sequence of a polynucleotide, said system being configured to carry out the method of any one of claims 29 to 34.
44. A system for carrying out a method for determining the sequence of a polynucleotide comprising a plurality of nucleotides, wherein the nucleotides are selected from any of the compounds according to claims 1 to 11.