Compositions of modified nucleoside triphosphates

CN122535618APending Publication Date: 2026-08-07HE SEQUENCING SOLUTIONS
View PDF 14 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HE SEQUENCING SOLUTIONS
Filing Date
2025-01-07
Publication Date
2026-08-07

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The present invention relates to a diastereomer of a nucleoside triphosphate suitable for extension sequencing. The diastereomer provides better acceptance and integration by DNA polymerase and better performance in an extension sequencing workflow. The invention also relates to a sequencing method using the diastereomer of the nucleoside triphosphate.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Sequence lists are merged by reference

[0002] This application hereby submits by reference the sequence list which is incorporated in the accompanying computer-readable format. Technical Field

[0003] This invention relates to compositions comprising an excess of a specific diastereomeric nucleoside triphosphate having a 5'-phosphatase ester. The invention also relates to methods for generating complementary strands and methods for sequencing using these compositions. Background Technology

[0004] Over the past two decades, biomembranes have become important tools in a variety of biomedical applications. This includes the use of lipid bilayers in nanopore-based sequencing applications, where nanopores provide constant and reproducible physical pores through which target molecules can be guided and sequenced.

[0005] One method for nanopore-based sequencing of nucleic acids, for example, involves an extended sequencing approach that transcribes the sequence of nucleic acids into easily measurable polymer molecules called Xpandomers. Much like polymerase chain reaction (PCR), Xpandomer synthesis is based on the natural function of DNA replication, where extensible nucleoside triphosphates (XNTPs) act as substrates for replication.

[0006] Xpandomer synthesis is based on four easily distinguishable XNTPs, including high signal-to-noise ratio reporter genes, with one reporter gene per DNA base. Engineered polymerases integrate these modified nucleotides into Xpandomer, thereby generating copies of the target nucleic acid template from a library. As the Xpandomer molecule passes through a nanopore, the unique electrical signal of each reporter gene is easily recognized, enabling high-precision and high-throughput nanopore-based nucleic acid sequencing. See, for example, U.S. Patent No. 7,939,259, entitled "High Throughput Nucleic Acid Sequencing by Expansion"; and PCT Publication WO 2020 / 236526 A1, entitled "Translocation control elements, reporter codes, and further means for translocation control for use in nanopore sequencing," both of which are hereby incorporated in their entirety.

[0007] Modified nucleoside triphosphates, such as dNTP-2c, with two clickable (e.g., terminal alkyne) groups are used as building blocks for reagents in nanopore sequencing, particularly in extended sequencing techniques. In this technique, such building blocks are typically clicked onto a ligand containing a reporter gene and translocation control elements to generate XNTPs. Structures and methods for extended sequencing, as well as reagents used therein, are disclosed in WO 2016 / 081871, WO 2020 / 236526, and WO 2020 / 172479.

[0008] The synthesis of dNTP-2c has previously been performed using a commercially available DNA / RNA synthesizer via solid-phase synthesis via a dedicated synthetic method. Since the α-phosphatase in dNTP-2c (and the XNTP derived therefrom) is chiral, in this method, dNTP-2c (and the XNTP derived therefrom) is obtained as a 1:1 diastereomeric mixture of the two isomers. Summary of the Invention

[0009] This disclosure relates to a nucleoside triphosphate with a 5' phosphoramidate that can be used in extended sequencing. The α-phosphoramidate in this type of nucleoside triphosphate is chiral and can have two different stereoconfigurations:

[0010] or

[0011] The inventors have surprisingly discovered that a stereoconfiguration of XNTPs (the “active” isomer) provides the desired functional performance in extended sequencing due to better polymerase integration and Xpandomer production compared to another “inactive” isomer or a mixture of both. As illustrated in the examples, only the active isomer allows for the generation of the full-length complementary (Xpandomer) product. Surprisingly, the inactive isomer not only fails to produce any full-length product, but its presence in a mixture with the active isomer negatively impacts the yield of the full-length product.

[0012] The specific use of compositions containing an excess of an active isomer relative to an inactive isomer (e.g., wherein at least 80%, preferably at least 90%, or even 100%, of the nucleoside triphosphate represents the active isomer) thus enables efficient and cost-effective Xpandomer production for extended sequencing. Based on these findings, it is desirable to remove inactive isomers from the final solution used in Xpandomer synthesis, or even avoid their generation from the outset. Therefore, this disclosure provides a composition containing a nucleoside triphosphate having a chiral 5'-phosphatase, wherein one isomer is present in excess of the other.

[0013] Exemplary implementations of this disclosure are as follows:

[0014] 1. A composition comprising nucleoside triphosphates having the following structure:

[0015] or

[0016] Where NB is a nucleobase; R 1 Contains or is composed of hydrocarbons; R 2 Independently, it is H, OH, or any 2'-ribose modification; R 3 It is H or any protecting group; and R 4 Contains or is composed of hydrocarbons; G 1 and G 2 Independently represents the terminal clickable group; L 1 and L 2 Independently representing the linking group; and T is a chain molecule;

[0017] Wherein at least 80%, preferably at least 90%, such as at least 95%, at least 99%, or 100%, of the said nucleoside triphosphate has the following stereoconfiguration at the α-phosphatase:

[0018]

[0019] 2. The composition according to item 1, wherein at least 90% of the nucleoside triphosphate has the following stereoconfiguration at the α-phosphoramidate ester:

[0020]

[0021] 3. The composition according to item 1 or 2, wherein 100% of the nucleoside triphosphate has the following stereoconfiguration at the α-phosphoramidate ester:

[0022]

[0023] 4. The composition according to any one of the preceding items, wherein NB is selected from cytosine, thymine, 7-deadenine, and 7-deadenine.

[0024] 5. The composition according to any one of the foregoing items, wherein R 1 When the nucleobase is a pyrimidine nucleobase, it is attached to position 5 of the nucleobase, and when the nucleobase is a purine nucleobase, it is attached to position 7 of the nucleobase.

[0025] 6. The composition according to any one of the preceding items, wherein the NB has one of the following structures:

[0026]

[0027] 7. The composition according to any one of the foregoing items, wherein the composition comprises a mixture of four different types of nucleoside triphosphates.

[0028] 8. The composition according to item 7, wherein the four different types of nucleoside triphosphates comprise four different types of nucleobases.

[0029] 9. The composition according to item 7 or 8, wherein the four different types of nucleoside triphosphates are base-paired with guanine, adenine, thymine and cytosine, respectively.

[0030] 10. The composition according to any one of items 7 to 9, wherein the four different types of nucleoside triphosphates each comprise the four different types of nucleobases of item 6.

[0031] 11. The composition according to any one of the preceding items, wherein R 1 It contains or is composed of unsaturated hydrocarbons.

[0032] 12. The composition according to any one of the preceding items, wherein R 1 It is composed of hydrocarbons such as alkynyl groups.

[0033] 13. The composition according to any one of the preceding items, wherein R 1 It is acyclic.

[0034] 14. The composition according to any one of the preceding items, wherein R 1 It's a linear chain.

[0035] 15. The composition according to any one of the preceding items, wherein R 1It contains 1 to 20 carbon atoms, such as 1 to 10 carbon atoms or 5 to 10 carbon atoms, such as 6 or 8 carbon atoms.

[0036] 16. The composition according to any one of the preceding items, wherein R 1 It is a hex-1-ynyl or oct-1-ynyl group.

[0037] 17. The composition according to any one of the preceding items, wherein R 1 With G 1 Together they are oct-1,7-diynyl or dec-1,9-diynyl groups.

[0038] 18. The composition according to any one of the preceding items, wherein two R 2 Both are H.

[0039] 19. The composition according to any one of the preceding items, wherein R 3 It is H.

[0040] 20. The composition according to any one of the preceding items, wherein R 4 It contains saturated hydrocarbons or is composed of saturated hydrocarbons.

[0041] 21. The composition according to any one of the preceding items, wherein R 4 It contains 1 to 20 carbon atoms, such as 3 to 15 carbon atoms, 3 to 10 carbon atoms, such as 4 carbon atoms.

[0042] 22. The composition according to any one of the foregoing items, wherein R 4 It is acyclic.

[0043] 23. The composition according to any one of the foregoing items, wherein R 4 It's a linear chain.

[0044] 24. The composition according to any one of the preceding items, wherein R 4 It is composed of hydrocarbons.

[0045] 25. The composition according to any one of the preceding items, wherein R 4 It is a n-butyl group.

[0046] 26. The composition according to any one of the preceding items, wherein R 4 With G 2 Together they form a hex-5-ynyl group.

[0047] 27. The composition according to any one of items 1 to 23, wherein R 4 It contains or is composed of two or more hydrocarbons, which are linked by atoms or atomic groups other than carbon, such as phosphorus atoms and / or oxygen atoms.

[0048] 28. The composition according to any one of the preceding items, wherein the terminal clickable group is a terminal alkyne group or a terminal azide group, preferably a terminal alkyne group.

[0049] 29. The composition according to any one of the foregoing items, wherein G 1 and G 2 Indicates the same type of terminal clickable groups.

[0050] 30. The composition according to any one of the preceding items, wherein L 1 and L 2 Independently represents the linking group formed via a click reaction.

[0051] 31. The composition according to any one of the preceding items, wherein L 1 and L 2 Each of them is a 1,2,3-triazole.

[0052] 32. The composition according to any one of the foregoing items, wherein at least 80%, preferably at least 90%, such as at least 95%, at least 99%, or 100% of the nucleoside triphosphate has a structure selected from the following:

[0053]

[0054]

[0055] 33. The composition according to any one of the foregoing items, wherein at least 80%, preferably at least 90%, such as at least 95%, at least 99%, or 100% of the nucleoside triphosphate has a structure selected from the following:

[0056]

[0057]

[0058]

[0059] 34. The composition according to any one of items 1-29, wherein at least 80%, preferably at least 90%, such as at least 95%, at least 99%, or 100%, of the nucleoside triphosphates has a structure selected from the following:

[0060]

[0061]

[0062] 35. The composition according to any one of items 1-29, wherein at least 80%, preferably at least 90%, such as at least 95%, at least 99%, or 100%, of the nucleoside triphosphates has a structure selected from the following:

[0063]

[0064]

[0065]

[0066] 36. A method for generating a complementary strand of a nucleic acid, the method comprising contacting the nucleic acid with a composition according to any one of items 1 to 33.

[0067] 37. The method according to item 36, wherein the nucleic acid is contained in a nucleic acid library.

[0068] 38. The method according to item 36 or 37, wherein the nucleic acid is DNA.

[0069] 39. The method according to any one of items 36 to 38, wherein the composition further comprises a nucleic acid polymerase.

[0070] 40. The method according to any one of items 36 to 39, wherein the composition further comprises a buffer such as TrisCl and / or a polymerase cofactor such as MnCl2.

[0071] 41. The method according to any one of items 36 to 40, the method comprising hybridizing primers with the nucleic acid while or subsequently contacting the nucleic acid with the composition.

[0072] 42. A method for determining the sequence of nucleic acids, the method comprising the following steps performed in sequence:

[0073] 1) Generate the complementary strand of the nucleic acid by the method according to any one of items 36 to 41;

[0074] 2) Selectively cleave the PN bonds within the nucleoside triphosphate to generate an extended complementary chain;

[0075] 3) Sequencing the extended complementary strand.

[0076] 4) Determine the sequence of the nucleic acid based on the sequence of the extended complementary strand.

[0077] 43. The method according to item 42, wherein in step 3), the PN bond is selectively cleaved under acidic conditions.

[0078] 44. The method according to item 42 or 43, wherein the extended complementary strand is sequenced by nanopore-based sequencing in step 4).

[0079] 45. The method according to item 44, wherein the nanopore-based sequencing comprises:

[0080] (a) Providing a chip for sequencing based on the nanopore, the chip comprising:

[0081] (i) An electrochemical resistance barrier disposed above an aperture on the surface of a chip, wherein the barrier separates the cis side and the trans side;

[0082] (ii) A nanopore inserted into the barrier, wherein the nanopore has an inlet side on the cis side of the barrier and an outlet side on the trans side of the barrier;

[0083] (b) Contact the cis side of the barrier with the extended complementary chain;

[0084] (c) Apply a voltage across the barrier of the chip to cause the extended complementary chain to translocate to the inverted side;

[0085] (d) Determine one or more changes in the electrical properties of the nanopore associated with the occupation of the extended complementary chain during translocation; and

[0086] (e) Determine the sequence of the extended complementary chain based on one or more variations in the electrical properties of the nanopore. Attached Figure Description

[0087] Figure 1: The HPLC chromatogram shows two independent peaks for the active and inactive diastereomers of four different dNTP-2c molecules.

[0088] Figure 2: Graphical representation of data selected from Table 2. Ratio: The ratio of active to inactive isomers present during Xpandomer synthesis; % Full length: The percentage of full-length complementary chains among all detected complementary chain products.

[0089] Figure 3: Gel electrophoresis of Xpandomer after synthesis using the active or inactive isomer. The active isomer (lane 37) yielded the full-length Xpandomer product, while the inactive isomer (lane 38) did not yield the full-length Xpandomer product. Detailed Implementation

[0090] The invention will now be described in detail by reference only using the following definitions and examples. All patents and publications cited herein, including all sequences disclosed therein, are expressly incorporated by reference.

[0091] Unless otherwise defined herein, all scientific and technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. Singleton (Singleton et al., Dictionary of microbiology and molecular biology, 2nd edition, 1994, John Wiley and Sons, New York), Hale (Hale and Marham, The Harper Collins dictionary of biology, 1991, Harper Perennial, NY), and Walker (Walker and Cox, The Language of Biotechnology: A Dictionary of Terms. 1988, American Chemical Society, Washington, DC ISBN-0-8412-1499-1) provide general dictionaries for many of the terms used in this invention that are employed by those skilled in the art. Practitioners should pay particular attention to the definitions and terminology used in the art by Sambrook (Sambrook et al., Molecular Cloning: A Laboratory Manual, 1989, Cold Spring Harbor Laboratory Press) and Ausubel (Ausubel et al., Current Protocols in Molecular Biology, 1993, John Wiley & Sons, Inc.). It should be understood that the invention is not limited to the specific methods, procedures, and reagents described, as these are variable.

[0092] As used herein, unless the context clearly indicates otherwise, the singular forms “a,” “an,” and “the” include the plural objects referred to.

[0093] In the structures shown in this article, when not all the natural valences of the atoms are filled by named groups, it should be understood that the unfilled valences are filled by hydrogen. When the wavy line in the structure intersects a bond, the intersecting bond is the location where the structure connects to the rest of the molecule.

[0094] When a structure describes a molecule having one or more negatively charged oxygen atoms, the structure also encompasses molecules having oxygen atoms bonded to H+ and / or any organic or inorganic cations. When a structure describes a molecule having one or more hydroxyl groups, the structure also encompasses molecules having oxygen atoms from the hydroxyl groups bonded to H+ and / or any organic or inorganic cations.

[0095] Throughout this specification, references to "an embodiment," "a particular embodiment," and variations thereof indicate that a specific feature, structure, or characteristic described in connection with that embodiment is included in at least one embodiment. Therefore, the phrases "in an embodiment" or "in a particular embodiment" appearing in different places throughout this specification do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.

[0096] Unless otherwise stated, nucleic acids are written from left to right with a 5' to 3' orientation; amino acid sequences are written from left to right with an orientation from amino to carboxyl.

[0097] The headings provided herein are not intended to limit the various aspects or embodiments of the invention, which can be obtained by referring to the entire specification. Therefore, the terms defined below are more fully defined by referring to the entire specification.

[0098] I. Terminology

[0099] Percentage of Identity: The term "% identity" in the context of nucleic acid or amino acid sequences refers to the level of sequence identity between a nucleic acid sequence and a reference nucleic acid sequence or between an amino acid sequence and a reference amino acid sequence when aligned using a sequence alignment program. For example, as used herein, 80% identity means that the sequence has greater than 80% sequence identity in length to the reference sequence. Exemplary levels of sequence identity include, but are not limited to, 80% or higher, 85% or higher, 90% or higher, 95% or higher, and 98% or higher sequence identity with a reference sequence (e.g., the wild-type sequence of any of the polypeptides described herein). Exemplary computer programs that can be used to determine the identity between two sequences include, but are not limited to, BLAST program suites (e.g., BLASTN, BLASTX, and TBLASTX, BLASTP, and TBLASTN), which are publicly available on the Internet. See also Altschul I and Altschul II. The BLASTN program is typically used for sequence searching when evaluating a given nucleic acid sequence relative to nucleic acid sequences in GenBank DNA sequences and other public databases. The BLASTX program can be used to search for nucleic acid sequences translated against amino acid sequences in GenBank protein sequences and other public databases across all reading frames. The BLASTP program can be used to search for amino acid sequences translated against amino acid sequences in GenBank protein sequences and other public databases. BLASTN, BLASTX, and BLASTP all run with the default parameters of an open vacancy penalty of 11.0 and an extended vacancy penalty of 1.0, and utilize the BLOSUM-62 matrix. (See, for example, Altschul II). In some exemplary embodiments, to determine the “% identity” between two or more sequences, the selected sequences are aligned using, for example, the CLUSTAL-W program in MacVector version 13.0.7, which operates with default parameters including an open vacancy penalty of 10.0, an extended vacancy penalty of 0.1, and a BLOSUM 30 similarity matrix.

[0100] Phosphate esters: "Phosphate esters" include "organophosphate esters" and their variants such as "aminophosphate esters" (which are synonyms for "phosphatidyl esters"). Phosphate esters may include side chains such as -R in the nucleoside triphosphates disclosed herein. 4 -G 2The first, second, and third phosphate esters, counted from the 5' end of the nucleoside, are also referred to as α-phosphate, β-phosphate, and γ-phosphate, respectively (or, in the case that the α-phosphate is a phosphoramidate, it may also be referred to as "α-phosphoramidate"). The type of a given phosphate ester may also be derived from the structure provided herein.

[0101] Scalable NTPs: "Scalable NTPs" or "XNTPs" refer to 5'-phosphate-modified non-natural nucleoside triphosphate (NTP) molecules (typically non-natural 2'-deoxynucleoside triphosphate molecules) that are compatible with template-dependent enzyme polymerization. Each XNTP has two distinct functional regions: a selectively cleavable bond (e.g., a phosphoramidite bond) linking the 5'-α-phosphate to the sugar contained in the nucleoside, and a tandem within the XNTP that allows expansion to be controlled by cleaving the cleavable bond (e.g., a tandem linking the 5'-α-phosphate and the nucleobase). Thus, XNTPs can exist in a restricted configuration (when the cleavable bond remains intact) or an extended configuration (when the cleavable bond has been cleaved, for example, via acid treatment).

[0102] dNTP-2c: "dNTP-2c" refers to a non-natural dNTP molecule modified with a 5' phosphate ester that can serve as an intermediate in XNTP synthesis. dNTP-2c contains two clickable groups, such as a terminal alkyne, one as part of the modification at the 5' α-phosphate ester and the other as part of the modification at the nucleobase. These two clickable groups allow for the addition of a chain linker between the α-phosphate ester and the nucleobase to form an XNTP.

[0103] Xpandomer: "Xpandomer" or "Xp" refers to a molecule composed of at least two XNTPs. Xpandomer can be obtained, for example, by using XNTPs as polymerase substrates to mediate the synthesis of complementary strands with template nucleic acid polymerases. Extended configurations of Xpandomer can be obtained by cleaving the phosphoramidite bonds in the XNTPs, for example, via acid treatment.

[0104] II. Nucleoside triphosphates with clickable groups

[0105] In some embodiments, this disclosure relates to nucleoside triphosphates comprising two clickable groups, one attached to an α-phosphatase and the other to a nucleobase. This allows the α-phosphatase to be linked to the nucleobase via a chain reaction to produce a scalable NTP. Such nucleoside triphosphates typically have the following structure:

[0106]

[0107] Where NB is a nucleobase; R 1 Contains or is composed of hydrocarbons; R 2 Independently, it is H, OH, or any 2'-ribose modification; R 3 It is H or any protecting group; and R 4 Contains or is composed of hydrocarbons; and G 1 and G 2 Independently representing the terminal clickable group. Preferably, the nucleoside triphosphate is a modified 2'-deoxynucleoside triphosphate (dNTP).

[0108] Nucleoside triphosphates contain a stereocenter at the phosphorus atom and can therefore exist as two distinct isomers with different stereoconfigurations at the α-phosphoramidate ester, as shown below:

[0109] or

[0110] Therefore, this disclosure provides a composition comprising the nucleoside triphosphate (having two clickable groups) disclosed herein. In some embodiments, this disclosure provides a composition comprising a nucleoside triphosphate having the following structure:

[0111]

[0112] Where NB is a nucleobase; R 1 Contains or is composed of hydrocarbons; R 2 Independently, it is H, OH, or any 2'-ribose modification; R 3 It is H or any protecting group; and R 4 Contains or is composed of hydrocarbons; G 1 and G 2 Independently represents a terminal clickable group; and T is a chain molecule;

[0113] Wherein at least 80%, preferably at least 90%, such as at least 95%, at least 99%, or 100%, of the said nucleoside triphosphate has the following stereoconfiguration at the α-phosphatase:

[0114] or

[0115] Preferably,

[0116]

[0117] Preferably, at least 90%, and more preferably 100%, of the nucleoside triphosphates in the composition have the following stereoconfiguration at the α-phosphoramidate:

[0118] or

[0119] Preferably,

[0120]

[0121] Therefore, for example, at least 80%, preferably at least 90%, such as at least 95%, at least 99%, or 100% of the nucleoside triphosphates in the composition can have the following structure:

[0122] or

[0123] Preferably,

[0124]

[0125] The clickable group can be any group that allows selective reactions with complementary clickable groups via click chemistry. Click chemistry and suitable clickable groups are well known in the art, see, for example, Fantoni et al., 2021, Chemical Reviews, 121 (12): 7122-7154; and Klöcker et al., 2020, Chem. Soc. Rev., 49:8749-8773. Examples of click reactions include alkyne + azide reaction (CuAAC), copper-free click tension-promoted azide-alkyne click (SPAAC) reaction (e.g., DBCO + azide), and reverse electron-demanding Diels-Alder cycloaddition reaction (IEDDA). Thus, for example, the terminal clickable group can be a terminal alkyne group or an azide group, preferably an alkyne group. In a preferred embodiment, G 1 and G 2 They are the same type of terminal clickable groups. Preferably, G 1 and G 2 Both refer to terminal alkyne groups.

[0126] NB is a nucleobase and will typically be a pyrimidine or purine nucleobase. This includes naturally occurring nucleobases such as adenine, guanine, cytosine, or thymine, as well as nucleobases with modifications that do not interfere with base pairing with complementary nucleobases. For example, a pyrimidine nucleobase may be modified at position 5, and a purine nucleobase may be modified at position 7. Non-limiting examples of nucleobases are adenine, guanine, thymine, cytosine, uracil, xanthine, hypoxanthine, 8-azapurine, purines substituted with a methyl or bromine at position 8, 9-oxo-N6-methyladenine, 2-aminoadenine, 7-deazoxanthine, 7-deazoguanine, 7-deazoadenine, N4-ethanolylcytosine, 2,6-diaminopurine, N6-ethanolyl-2,6-diaminopurine. Purines, 5-methylcytosine, 5-(C3–C10)-alkynylcytosine, 5-fluorouracil, 5-bromouracil, thiouracil, pseudoisocytosine, 2-hydroxy-5-methyl-4-triazolopyridine, isocytosine, isoguanine, hypoxanthine, 7,8-dimethylpyrazine, 6-dihydrothymidine, 5,6-dihydrouracil, 4-methyl-indole, alcohol adenine, and U.S. Patent No. 5,432,272 Nucleotides not found in any of the following documents are incorporated herein by reference in their entirety: Nos. 6,150,510 and 6,150,510, and published PCT applications WO 92 / 002258, WO 93 / 10820, WO 94 / 22892 and WO 94 / 24144, and non-naturally occurring nucleotides described in Fasman's "Practical Handbook of Biochemistry and Molecular Biology", pp. 385-394, 1989, CRC Press, Boca Raton, La. In one embodiment, the nucleotides are selected from adenine, guanine, uracil, and cytosine, and modified versions of these nucleotides such as those disclosed herein (e.g., 7-deadenine or 7-deadenine). NB is preferably selected from cytosine, thymine, 7-deadenine, and 7-deadenine. In a preferred embodiment, NB has one of the following structures:

[0127]

[0128] R 1 This is usually done so that it does not interfere with base pairing with complementary nucleobases. For example, R 1The nucleobase is attached to position 5 of the nucleobase when the nucleobase is a pyrimidine nucleobase, and to position 7 of the nucleobase when the nucleobase is a purine nucleobase (where the naturally occurring nitrogen at position 7 can be replaced by carbon, for example, as in 7-deadenine or 7-deadenine). Nucleosides with modifications at position 5 (pyrimidine base) or position 7 (purine base) and their synthesis are well known, see, for example, Kozak et al., 2020 (Russ.Chem. Rev., 2020, 89 (3) 281-310) and Matyugina et al., 2021 (Russ.Chem. Rev., 2021, 90 (11) 1454-1491). For specific methods of synthesizing nucleosides with the bases shown above, see also WO 2016 / 081871.

[0129] R 1 It contains (substituted or unsubstituted, preferably unsubstituted) hydrocarbons or is composed of hydrocarbons. For extended sequencing applications, R 1 It typically consists of hydrocarbons (substituted or unsubstituted, preferably unsubstituted). The hydrocarbons can be saturated or unsaturated, preferably unsaturated. For example, R 1 It may contain (substituted or unsubstituted, preferably unsubstituted) alkyl, alkenyl or ynyl groups, preferably ynyl groups, or be composed of alkyl, alkenyl or ynyl groups, preferably ynyl groups.

[0130] Preferably, R 1 Contains 1 to 100 carbon atoms, preferably 1 to 30 carbon atoms or 1 to 20 carbon atoms, such as 3 to 20 carbon atoms, 3 to 10 carbon atoms, or 5 to 10 carbon atoms, such as 6 or 8 carbon atoms. Typically, R 1 It will be acyclic. Preferably, R 1 It is a linear chain. In some implementations, R 1 The molecular weight is 1500 g / mol or lower, 1000 g / mol or lower, 500 g / mol or lower, 200 g / mol or lower, or 100 g / mol or lower.

[0131] In some implementation schemes, R 1 It is a substituted or unsubstituted, branched or unbranched, saturated or unsaturated alkyl group containing 1 to 100 carbon atoms, which optionally includes one or more oxygen heteroatoms, nitrogen heteroatoms, phosphorus heteroatoms or sulfur heteroatoms (e.g., including ethers, thioethers, phosphate diesters or phosphate triesters or PEG, heterocycles such as triazoles or imidazoles).

[0132] In some implementation schemes, R 1 Yes –R w –Z, where R w It is a substituted or unsubstituted, branched or unbranched, saturated or unsaturated alkyl group having between 1 and 100 carbon atoms, optionally including one or more oxygen heteroatoms, nitrogen heteroatoms, phosphorus heteroatoms or sulfur heteroatoms, and wherein Z is alkyl, alkenyl, alkynyl, acyl, –Het or –CH2–Het, wherein “Het” is a substituted or unsubstituted 5- or 6-membered heterocyclic moiety.

[0133] As a preferred example, R 1 Composed of straight-chain hydrocarbons such as alkynyl groups, and G 1 It is a terminal alkyne group. Preferably, R 1 It is a hex-1-ynyl group or an oct-1-ynyl group. In the most preferred example, G 1 With R 1 It is an octyl-1,7-diynyl or decyl-1,9-diynyl group.

[0134] R 2 Independently modified by H, OH, or any 2'-ribose. In some embodiments, both R... 2 Both are H or one R 2 It is H and another R 2 It is OH. Preferably, two R 2 All are H. 2'-ribose modification is known in the art and includes, for example, tert-butyldimethylsilyl groups and triisopropylsiloxymethyl ether groups, as well as 2'-O-methyl groups or 2'-fluorine groups.

[0135] R 3 It is H or any protecting group, preferably H. Protecting groups are known in the art. Examples of protecting groups include acetyl, benzoyl, benzyl, methoxyethoxymethyl ether, dimethoxytriphenylmethyl, ethoxymethyl ether, methoxytriphenylmethyl, p-methoxybenzyl ether, p-methoxyphenyl ether, methylthiomethyl ether, neopentanoyl, tert-butyl ether, tetrahydropyranyl, tetrahydrofuran, triphenylmethyl, silyl ether (e.g., trimethylsilyl, tert-butyldimethylsilyl, triisopropylsiloxymethyl or triisopropylsilyl ether), methyl ether and ethoxyethyl ether.

[0136] R 4 It contains or is composed of hydrocarbons. For extended sequencing applications, R... 4It is typically composed of hydrocarbons. The hydrocarbons can be substituted or unsubstituted, preferably unsubstituted. The hydrocarbons can be saturated or unsaturated, preferably saturated. For example, R... 4 It may contain alkyl, alkenyl or ynyl groups, preferably alkyl or composed of alkyl, alkenyl or ynyl groups, preferably alkyl.

[0137] Typically, R 4 Contains 1 to 100 carbon atoms, preferably 1 to 30 carbon atoms or 1 to 20 carbon atoms, such as 1 to 15 carbon atoms, 3 to 15 carbon atoms, 3 to 10 carbon atoms or 3 to 6 carbon atoms, such as 4 carbon atoms. Typically, R 4 It will be acyclic. Preferably, R 4 It is a linear chain. In the preferred embodiment, R 4 It contains or is composed of straight-chain (saturated) alkyl groups. In some embodiments, R 4 The molecular weight is 1500 g / mol or lower, 1000 g / mol or lower, 500 g / mol or lower, 200 g / mol or lower, or 100 g / mol or lower.

[0138] In some implementation schemes, R 4 It comprises or consists of branched, straight-chain, cyclic or heterocyclic, substituted or unsubstituted, saturated or unsaturated hydrocarbons, optionally including one or more heteroatoms, optionally selected from nitrogen, oxygen, phosphorus and sulfur. For example, the cyclic or heterocyclic hydrocarbons may be 5- or 6-membered. For example, the cyclic or heterocyclic hydrocarbons may be aromatic.

[0139] In some implementation schemes, R 4 It is a substituted or unsubstituted, branched or unbranched, saturated or unsaturated alkyl group containing 1 to 100 carbon atoms, which optionally includes one or more oxygen heteroatoms, nitrogen heteroatoms, phosphorus heteroatoms or sulfur heteroatoms (e.g., including ethers, thioethers, phosphate diesters or phosphate triesters or PEG, heterocycles such as triazoles or imidazoles).

[0140] In some implementation schemes, R 4 Yes –R w –Z, where R wIt is a substituted or unsubstituted, branched or unbranched, saturated or unsaturated alkyl group having between 1 and 100 carbon atoms, optionally including one or more oxygen heteroatoms, nitrogen heteroatoms, phosphorus heteroatoms or sulfur heteroatoms, and wherein Z is alkyl, alkenyl, alkynyl, acyl, –Het or –CH2–Het, wherein “Het” is a substituted or unsubstituted 5- or 6-membered heterocyclic moiety;

[0141] In a preferred embodiment, R 4 Composed of straight-chain hydrocarbons such as alkyl groups, and G 2 It is a terminal alkyne group. As a preferred example, R... 4 With G 2 It is a hex-5-ynyl group. Therefore, as a preferred example, R 4 With G 2 It has the following structure:

[0142]

[0143] R 4 It may also contain or consist of two or more hydrocarbons, said two or more hydrocarbons being linked by atoms or atomic groups other than carbon, such as phosphorus atoms and / or oxygen atoms. For example, two hydrocarbons (each independently containing 1 to 10, 2 to 10, or 2 to 6 carbon atoms) such as alkyl groups may be linked by oxygen atoms. In another example, two or three hydrocarbons (each independently containing 1 to 10, 2 to 10, or 2 to 6 carbon atoms) such as alkyl groups may be linked by phosphate diesters or phosphate triesters, respectively. Thus, in embodiments, R 4 With G 2 It has the following structure:

[0144]

[0145] During production, R containing phosphate esters 4 Phosphate esters typically have protecting groups, such as β-cyano-ethyl groups, at the hydroxyl functional groups. These protecting groups can be removed upon addition of pyrophosphate to produce triphosphate esters.

[0146] For example, R 4 It can be by

[0147] a) Straight-chain hydrocarbons, or

[0148] b) It consists of two straight-chain hydrocarbons linked by a phosphate diester.

[0149] Where R 4Contains 3 to 15 carbon atoms, and G 2 It is a terminal alkyne group.

[0150] In a preferred embodiment, at least 80%, preferably at least 90%, such as at least 95%, at least 99%, or 100%, of the nucleoside triphosphates in the composition have a structure selected from the following:

[0151]

[0152]

[0153] In a more preferred embodiment, at least 80%, preferably at least 90%, such as at least 95%, at least 99%, or 100%, of the nucleoside triphosphates in the composition have a structure selected from the following:

[0154]

[0155]

[0156]

[0157] For example, racemic mixtures of nucleoside triphosphates having two clickable groups, as disclosed herein, can be produced by the method disclosed in WO 2016 / 081871. Certain types of R 4 For example, those containing phosphate esters may include a protecting group, such as an ethyl cyanide group, at the phosphate ester until a 5' triphosphate ester has been formed. The given diastereomeric pure isomer can then be separated from the other isomers by, for example, high-performance liquid chromatography (HPLC). Specific conditions for separating two isomers by preparative HPLC are given in Example 1. As described in Example 1, the active configuration can first be identified as the elution in HPLC. This configuration can be further functionally identified, for example, by separating the two enantiomers and testing whether the scalable nucleoside triphosphate synthesized using the given enantiomer is suitable for Xpandomer synthesis as described in Example 2.

[0158] III. Scalable NTP

[0159] The nucleoside triphosphates disclosed in this paper can be used, for example, for extended sequencing. In this case, α-phosphatases are typically linked to nucleobases, for example, via a chain linker molecule (to form an extensible NTP). For example, R 1 It can be linked to R, for example, via a chain molecule. 4Therefore, this disclosure also relates to expandable nucleoside triphosphates, and is thus suitable for, for example, extended sequencing. Such nucleoside triphosphates can be obtained, for example, by linking the α-phosphatase and nucleobase in the nucleoside triphosphate to two clickable groups via a click reaction. Such nucleoside triphosphates typically have the following structure:

[0160]

[0161] Where T is a chain molecule; NB is a nucleobase; R 1 Contains or is composed of hydrocarbons; R 2 Independently, it is H, OH, or any 2'-ribose modification; R 3 It is H or any protecting group; R 4 Contains or is composed of hydrocarbons, and L 1 and L 2 Independently representing the linking group. NB, R from the context of a nucleoside triphosphate with two clickable groups. 1 R 2 R 3 and R 4 Further descriptions also apply.

[0162] Nucleoside triphosphates contain a stereocenter at the phosphorus atom and can therefore exist as two distinct diastereomers with different stereoconfigurations at the α-phosphoramidate ester, as shown below:

[0163]

[0164] Therefore, this disclosure provides a composition comprising (scalable) nucleoside triphosphates suitable for extended sequencing as disclosed herein. In some embodiments, this disclosure provides a composition comprising (scalable) nucleoside triphosphates having the following structure:

[0165]

[0166] Where NB is a nucleobase; R 1 Contains or is composed of hydrocarbons; R 2 Independently, it is H, OH, or any 2'-ribose modification; R 3 It is H or any protecting group; and R 4 Contains or is composed of hydrocarbons; L 1 and L 2 Independently representing the linking group; and T is a chain molecule;

[0167] Wherein at least 80%, preferably at least 90%, such as at least 95%, at least 99%, or 100%, of the said nucleoside triphosphate has the following stereoconfiguration at the α-phosphatase:

[0168] or

[0169] Preferably,

[0170]

[0171] Therefore, for example, at least 80%, preferably at least 90%, such as at least 95%, at least 99%, or 100%, of the nucleoside triphosphate in the composition can have the following structure:

[0172] or

[0173] Preferably,

[0174]

[0175] The composition may also comprise a mixture of different types of nucleoside triphosphates. In some embodiments, the composition comprises a mixture of four different types of nucleoside triphosphates. Typically, the four different types of nucleoside triphosphates comprise four different types of nucleobases. Preferably, the four different types of nucleoside triphosphates are base-paired with guanine, adenine, thymine, and cytosine, respectively. Thus, in some embodiments, the composition comprises four different types of nucleoside triphosphates, each comprising a unique nucleobase, for example selected from 7-deadenine, 7-deadenine, thymine, and cytosine. Preferably, the four different types of nucleoside triphosphates each comprise a unique nucleobase selected from:

[0176]

[0177] There are no particular restrictions on ligand molecules, but they typically include a reporter gene (for example, to achieve specific recognition of the attached nucleobase via nanopore-based sequencing). Ligand molecules can be, for example, symmetric synthetic reporter ligands (SSRTs) disclosed in WO 2020 / 236526 A1. Such ligand molecules typically have the following structure: linker A – reporter gene – linker B. Linker A can be attached to an α-phosphatase and linker B can be attached to a nucleobase, and vice versa.For example, linker A and linker B can be polymers containing two or more repeating units selected from: spermidine (Q), hexaethylene glycol (D), 2-((4-((3-(benzoyloxy)-2-(((1-(3-(benzoyloxy)-2-((benzoyloxy)methyl)-2-((phosphodiesteroxy)methyl)propyl)-1H-1,2,3-triazol-4-yl)methoxy)methyl)-2-((benzoyloxy)methyl)propoxy)methyl)-1H-1,2,3-triazol-1-yl)methyl)-2-O-phosphodiester-propane-1,3-dibenzoate, 1,3-O- bis(phosphodiester-2,2-bis(1-Me-4-(Me-O-PEG2-O-Bz)-1,2,3-triazole))propane, 1,3-O-bis(phosphodiester-2-(4-(Me-O-PEG5)-1-(Et-O-Ac)-1,2,3-triazole)-propane, 1,3-O-bis(phosphodiester-2s-O-(4-(Me-O-PEG7)-1-(Et-OBz)-1,2,3-triazole)-propane, 1,3-O-bis(phosphodiester-2s-O-(4-(Me-O-PEG7)-1-(Et-OBz)-1,2,3-triazole)-propane, 1,3-O-bis(phosphodiester-2s-O-(4-(Me-O-PEG3)-1-(Et-2,2,2-Tris-(M e-O-Bz))-1,2,3-triazole-propane, 1,3-O-bis(phosphate diester-2-(4-(Me-O-PEG5)-1-(Et-2,2,2-tri-(Me-O-Ac))-1,2,3-triazole)-propane, 1,2-O-bis(phosphate diester)-3-(4-(Me-O-PEG3-O-Bz)-1-(1,2,3-triazole))-propane, 1,3-O-bis(phosphate diester-2,2-bis(4-(Me-O-PEG2-O-Me)-1-(Et-O-Bz)-1,2,3-triazole)-propane, 1,3-O-bis( Phosphodiester-2,2-bis(4-(Me-O-PEG3-O-Me)-1-(Et-2,2,2-tri-(Me-O-Bz))-1,2,3-triazole)-propane, 1,2-O-bis(phosphodiester)-3-(4-methylpiperazin-1-yl)-propane, 1,3-O-bis(phosphodiester-2,2-bis(4-(Me-O-PEG3-O-Me)-1-(Et-O-Bz)-1,2,3-triazole)-propane, and 1,1'-O-bis(phosphodiester)-N(p-tolyl)-diethanolamine, preferably spermidine. In some embodiments, linker A and linker B are reverse copies of each other.

[0178] For example, a reporter gene can be a polymer containing two or more repeating units selected from: hexaethylene glycol (D), ethane (L), triethylene glycol (X), 1,3-O-bis(phosphodiester)-2S-O-mPEG4-propane, 1,3-O-bis(phosphodiester)-2-(4-Me-O-PEG3)-1-(Et-O-Ac)-1,2,3-triazole)-propane, 1,3-O-bis(phosphodiester)-2,2-bis(Me-O-mPEG2)-propane, 1,3-O-bis(phosphodiester)-2S-O-(PEG4-O-Bz)-propane, 1,3-O-bis(phosphodiester)-2S-O-mPEG6-propane, 1,3-O-bis(phosphate diester-2s-O-(4-(Me-O-PEG3)-1-(Et-2,2,2-tris(Me-O-Bz))-1,2,3-triazole)-propane, 1,3-O-bis(phosphate diester-2s-O-(4-(Me-O-PEG3)-1-(Me-acetate)-1,2,3-triazole)-propane, 1,3-O-bis(phosphate diester)-2s-O-(4-(Me-O-PEG2)-1-(Et-OBz)-1,2,3-triazole)-propane, 1,3-O-bis(phosphate diester)-2-(4-Et-1-(Et-O-mPEG1)-1,2,3-triazole)-propane, 2,3-O-bis( 1-(1-dimethoxyquinazolinidone)-propane, 2,3-O-bis(phosphate diester)-1-(N9-(3,6-dimethoxycarbazole)-propane, 1,1'-O-bis(phosphate diester)-2,2'-(sulfonylbis(phenyl-4-yl))-diethanol, 1,1'-O-bis(phosphate diester)-2,2'-bipyridin-4,4'-yl)-diethanol, 2,3-O-bis(phosphate diester)-1-(N1-(4,6-dimethoxy-3-methylindole)-propane, 3-(1,2-O-bis(phosphate diester)-propyl)-8,8-dimethylhexahydro-3H-3a,6-methylenebenzo[c]isothiazolium 2,2-dioxide, 2,3-O-bis( Phosphate diester)-1-(N1-(6-azathymidine))-propane, 1,5-O-bis(phosphate diester)-hexahydrofuran[2,6]furan, 1,1'-O-bis(phosphate diester)-octahydro-2,6-dimethyl-3,8:4,7-dimethylbridge-2,6-naphthidine-4,8-diyl)-diethanol, 2,3-O-bis(phosphate diester)-1-(N1-(2-methyl-5-nitroindole)-propane, 2,3-O-bis(phosphate diester)-1-(N1-(2-methyl-5-nitroindole))-propane, 2,3-O-bis(phosphate diester)-1-(5-benzofuran)-propane, 1,2-O-bis(phosphate diester)-3-O-mPEG2-propane, 1,3-O-bis(phosphodiester)-2-(4-Et-1-(Et-O-mPEG3)-1,2,3-triazole)-propane and 1,3-O-bis(phosphodiester)-3-O-mPEG4-propane (see WO 2020 / 236526 A1). In some embodiments, the reporter gene comprises or consists of two inverted copies of the same polymer, which contains two or more repeating units selected above. The two inverted copies of the same polymer in the reporter gene may be linked via branching elements that are further linked to translocation control elements (TCEs), as described in WO 2020 / 236526 A1.

[0179] L 1 and L 2 The linking group is indicated independently. There are no particular restrictions on the linking group, and it includes substituted or unsubstituted hydrocarbons, including, for example, 1,2,3-triazoles.

[0180] Preferably, the chain molecules attach via a click chemistry reaction. In other words, G 1 and G 2 The terminal clickable groups in [the chain] can be used to link chain molecules via click reactions. These terminal clickable groups are typically of the same type, for example, each being an alkyne group. When the chain is attached via a click reaction, L […]. 1 and L 2 Each is a product of the click reaction, such as 1,2,3-triazole. For example, when G... 1 and G 2 When both are terminal alkyne groups, the chain molecule with terminal azide groups attached to both ends can react with G. 1 and G 2 The reaction produces two 1,2,3-triazoles (or vice versa, i.e., G). 1 and G 2 It is a terminal azide group and is chain-attached to two terminal alkyne groups. In a preferred embodiment, L 1 and L 2 It is 1,2,3-triazole.

[0181] In a preferred embodiment, at least 80%, preferably at least 90%, such as at least 95%, at least 99%, or 100%, of the nucleoside triphosphates in the composition have a structure selected from the following:

[0182]

[0183]

[0184] The first two structures can be obtained by: taking a nucleoside triphosphate with two clickable groups as disclosed herein (wherein, R 1 Connected to G 1 = Oct-1,7-diynyl or Dec-1,9-diynyl group, and R 4 Connected to G 2 =Hexyl-5-ynyl group) reacts with the chain attached to two terminal azide groups.

[0185] In a more preferred embodiment, at least 80%, preferably at least 90%, such as at least 95%, at least 99%, or 100%, of the nucleoside triphosphates in the composition have a structure selected from the following:

[0186]

[0187]

[0188]

[0189] The composition may also be provided as a master mixture or a reaction solution. The composition may therefore further contain additional reagents for complementary strand synthesis. For example, the composition may further contain a nucleic acid polymerase (see Section IV for details). Furthermore, the composition may further contain a buffer such as TrisCl. In some embodiments, the composition further contains at least one of the following (including each of the following): TrisOAc, NH4OAc, PEG, a water-miscible organic solvent such as DMF or NMP, polyphosphate 60, N-methylsuccinimide (NMS) and (MnCl2), a single-strand binding protein (SSB), and urea. For example, the SSB may be Kod SSB (from *Thermococcus kodakarensis*). The composition may further contain a polymerase enhancer molecule (PEM) such as that described in EP 3 735 409 B1. The reaction solution typically also contains at least one nucleic acid. In these embodiments, the composition typically contains four different types of nucleoside triphosphates, wherein the four different types of nucleoside triphosphates are base-paired with guanine, adenine, thymine, and cytosine, respectively.

[0190] (Scalable) nucleoside triphosphates can be, for example, by utilizing two clickable G groups. 1 and G 2 α-phosphatase is obtained by linking it to the nucleobase of a nucleoside triphosphate. α-phosphatase can be linked to the nucleobase via a chain molecule, which can be attached via a click chemistry reaction. For example, when G... 1 and G 2When the terminal alkyne group is attached to two terminal azide groups, the double-click reaction at each end attaches the chain molecule to R via two 1,2,3-triazoles. 1 and R 4 The reaction can be shown as follows:

[0191]

[0192] Among them G 1 and G 2 Independently representing terminal clickable groups, G 1a and G 2a Independently representing terminal clickable groups, L 1 and L 2 This indicates that G is used to make G 1 With G 1a reaction and G 2 With G 2a The reaction forms linking groups, and T, NB, R 1 R 2 R 3 and R 4 As defined above.

[0193] Nucleoside triphosphates can be further purified, for example, by HPLC. Purification can be carried out, for example, after reaction with pyrophosphate and / or after linking α-phosphoramidate to a nucleobase.

[0194] IV. Methods

[0195] This disclosure also provides a method for generating complementary strands of nucleic acids, the method comprising contacting the nucleic acid with a composition comprising a nucleoside triphosphate as disclosed herein. The complementary strand generated by this method is typically an Xpandomer (in a restricted configuration).

[0196] There are no particular limitations on the type of nucleic acid, and it includes either DNA or RNA. In a preferred embodiment, the nucleic acid is DNA, such as genomic DNA or cDNA. Typically, the DNA is cell-free DNA.

[0197] Nucleic acids can be part of a nucleic acid library. For example, the library could be a library of genomic DNA or cDNA.

[0198] The generation of complementary strands is typically initiated by primers. Primer design and generation are known in the art. There are no particular restrictions on the primers used, and they can be designed, for example, to hybridize with nucleic acids at a certain location to allow the generation of complementary strands of the nucleic acids, including the full-length target portion. When sequencing a nucleic acid library, for example, standard primers that bind to all target nucleic acids in the library can be used, or a random primer mixture can be used.

[0199] In some embodiments, the method may include hybridizing primers with nucleic acids, followed by contacting the nucleic acids with a composition comprising nucleoside triphosphates as disclosed herein.

[0200] If desired, nucleic acids may also be denatured, for example, to facilitate primer hybridization. Methods for denaturation are not particularly limited and include, for example, applying heat (e.g., 90°C to 100°C). Therefore, in some embodiments, the method may include denaturing the nucleic acid and then hybridizing the primer with the nucleic acid, followed by contacting the nucleic acid with a composition comprising a nucleoside triphosphate as disclosed herein.

[0201] Typically, the complementary strand is generated using a polymerase such as a (DNA-dependent) DNA polymerase. Therefore, the composition used preferably further comprises a nucleic acid polymerase. Typically, a polymerase containing a mutation that spatially allows the use of XNTPs as substrates will be used. Suitable types of polymerases for incorporating XNTPs include the transdamage DNA polymerase (i.e., class Y polymerase) family, which includes, for example, DPO4 polymerase. Due to their relatively large substrate-binding sites, transdamage DNA polymerases exhibit more flexible substrate recognition than conventional (e.g., replication) polymerases, which have evolved to adapt to naturally occurring, large-volume DNA damage. Suitable polymerases include, for example, modified DPO4 polymerases described in WO 2017 / 087281, WO 2018 / 204707, or WO 2019 / 118372, which are incorporated herein by reference in their entirety. Suitable examples, such as SEQ ID NO: 1 to 5, are provided herein. Therefore, in some embodiments, the polymerase has at least 95% sequence identity with the amino acid sequence of SEQ ID NO: 1, such as at least 98% or 99%. This polymerase has DNA-dependent DNA polymerase activity and, more specifically, is capable of using XNTPs as a polymerization substrate.

[0202] In addition, the composition typically further comprises a buffer such as TrisCl. In some embodiments, the composition comprises at least one of the following (including each of the following): TrisOAc, NH4OAc, PEG, a water-miscible organic solvent such as dimethylformamide (DMF) or N-methylpyrrolidone (NMP), polyphosphate 60, N-methylsuccinimide (NMS) and MnCl2, single-chain binding protein (SSB), and urea. For example, the SSB may be Kod SSB.

[0203] The composition may further comprise a polymerase enhancer molecule (PEM) such as that described in EP 3 735 409 B1. For example, the PEM may be a compound of the following formula, as described in the claims of EP 3 735 409 B1, that increases the processability, rate, or fidelity of a nucleic acid polymerase reaction:

[0204]

[0205] Independently in each occurrence: m is 1, 2, or 3; n is 0, 1, or 2; p is 0, 1, or 2; Ar1 ​​is optionally a substituted aryl group; Ar2 is selected from 5-membered monocyclic aromatic rings and 6-membered monocyclic aromatic rings, as well as 9-membered fused bicyclic and 10-membered fused bicyclic rings, wherein the 9-membered fused bicyclic and 10-membered fused bicyclic rings comprise two fused 5-membered monocyclic aromatic rings and / or 6-membered monocyclic aromatic rings, wherein at least one of the two monocyclic rings is an aromatic ring, wherein Ar2 is optionally selected from halides, C1-C6 alkyl groups, C1-C6 haloalkyl groups, and ECO2R. 0 ,E-CONH2,E-CHO,EC(O)NH(OH),EN(R 0 )2 and E-OR 0 One or more substituents are substituted, wherein E is selected from direct bonds and C1-C6 alkylene groups; and R 0 The group is selected from H, C1-C6 alkyl and C1-C6 haloalkyl, M is selected from hydrogen, halogen and C1-C4 alkyl; and L is a linking group; or a solvate, hydrate, tautomer, chelate or salt thereof.

[0206] The composition typically contains four different types of nucleoside triphosphates, which are base-paired with guanine, adenine, thymine and cytosine, respectively.

[0207] This disclosure also provides a method for sequencing nucleic acids using scalable nucleoside triphosphates as disclosed herein or compositions containing scalable nucleoside triphosphates.

[0208] Therefore, this disclosure provides a method for determining the sequence of nucleic acids, the method comprising the following sequential steps:

[0209] 1) Generate complementary strands of nucleic acids using the methods disclosed herein for generating complementary strands;

[0210] 2) Selectively cleave the PN bonds within the nucleoside triphosphate to generate an extended complementary chain;

[0211] 3) Sequencing the extended complementary strand.

[0212] 4) Determine the sequence of the nucleic acid based on the sequence of the extended complementary strand.

[0213] The nucleoside triphosphates used in such methods allow for the formation of extended complementary chains. This is typically achieved by using nucleoside triphosphates, where α-phosphatase is linked to the nucleobase via a tie-chain molecule. When the phosphatase (PN) bond is cleaved, the α-phosphatase and the nucleobase remain linked via the tie-chain molecule, thereby generating an extended complementary chain.

[0214] In some implementations, after step 1), the complementary strand is separated from the nucleic acid, for example, by denaturation. The complementary strand may optionally be purified before proceeding to step 2).

[0215] For example, under acidic conditions, the PN bond can be selectively cleaved in step 2). This can be achieved by adding an acid such as DCl. The cleavage in step 2) typically yields an extended configuration of Xpandomer. The product of step 2) can optionally be purified before proceeding to step 3).

[0216] In a preferred embodiment, the extended complementary strand is sequenced in step 3) by nanopore-based sequencing. Methods for nanopore-based sequencing are known in the art, see, for example, WO 2020 / 236526. For example, nanopore-based sequencing may include:

[0217] (a) Providing a chip for sequencing based on the nanopore, the chip comprising:

[0218] (i) An electrochemical resistance barrier disposed above an aperture on the surface of a chip, wherein the barrier separates the cis side and the trans side;

[0219] (ii) A nanopore inserted into the barrier, wherein the nanopore has an inlet side on the cis side of the barrier and an outlet side on the trans side of the barrier;

[0220] (b) Contact the cis side of the barrier with the extended complementary chain;

[0221] (c) Apply a voltage across the barrier of the chip to cause the extended complementary chain to translocate to the inverted side;

[0222] (d) Determine one or more changes in the electrical properties of the nanopore associated with the occupation of the extended complementary chain during translocation; and

[0223] (e) Determine the sequence of the extended complementary chain based on one or more variations in the electrical properties of the nanopore.

[0224] Barriers are typically lipid bilayer membranes such as DPhPE / hexadecane bilayers. Nanopores, such as a. hemolysine nanopores, can be inserted into the membrane via electroporation in a buffer (e.g., a buffer of 2 M NH4Cl and 100 mM HEPES, pH 7.4). Before introducing Xpandomer to the cis side for sequencing, the cis pores can be perfused with a buffer containing 0.4 M NH4Cl, 600 mM GuanCl, 100 mM HEPES, pH 7.4, and 5% glycerol, and the trans pores can be perfused with a buffer containing 0.4 M NH4Cl, 600 mM GuanCl, 5% ethyl acetate, 10 mM HEPES, pH 7.4.

[0225] V. Example

[0226] Example 1: Production of diastereomers of dNTP-2c molecules

[0227] Racemic mixtures of nucleoside triphosphates containing two clickable groups as disclosed herein were prepared by the method disclosed in WO2016 / 081871. Four separate racemic mixtures were produced for four different types of nucleoside triphosphates having four different nucleobases corresponding to C, T, A, and G, respectively, each having a -R nucleobase. 1 -G 1 Furthermore, the α-phosphatidyl ester disclosed herein has -R 4 -G 2 For each racemic mixture, the two isomers were then separated from each other by preparative HPLC.

[0228] HPLC System: Agilent 1290 Infinity II Preparative LC System

[0229] Column: 50 x 250 mmWaters Xbridge C18

[0230] Protective bollard: Waters Xbridge C18

[0231] Mobile phase: MeOH / H2O premixed to remove heat of mixing. 1M TEAB was mixed on the HPLC instrument by pumping in at 10%.

[0232] Table 1 shows the preparation workflow used.

[0233] Table 1

[0234]

[0235] Figure 1 shows exemplary HPLC chromatograms for four different types of nucleoside triphosphates. The chromatograms show two distinct peaks for the active and inactive isomers of the four different types of nucleoside triphosphates. The two isomers are separated from each other by collecting appropriate fractions from the two peaks. This allows for the production of compositions in which one of the two isomers is present in excess.

[0236] Example 2: Synthesis and sequencing of Xpandomer using active and inactive isomers

[0237] To generate Xpandomer copies of the DNA template, solid-state primer expansion reactions were performed using equimolar amounts of each XNTP, 4 pmol of template, and 20 pmol of E-oligo primers (solid-state Xpandomer synthesis of extended oligonucleotides covalently bound to chip substrates is described in WO 2020 / 172479 A1, which is incorporated herein by reference in its entirety). The 50 µL expansion reaction consisted of the following reagents: 50 mM TrisCl, pH 8.84, 200 mM NH4OAc, 50 mM GuC120% PEG8K, 10% N-methylpyrrolidone (NMP), 15 nmol polyphosphate PP-60.23, 2.5 µg Kod single-strand binding protein (SSB), 0.1 M urea, 15 mM PEM additive, and 13 µg purified recombinant DNA polymerase C4760 (SEQ ID NO: 2, a variant of DPO4 polymerase; other suitable variants include SEQ ID NO: 1 and 3 to 5). The expansion reaction was carried out at 37 °C for 60 min.

[0238] Next, the Xpandomer product was sequenced using the SBX protocol. In brief, the constrained Xpandomer product was washed in buffer B.064 (1% Tween-20 / 3% SDS / 5mM HEPES, pH 8.0 / 100mM NaPO4 / 15% DMF) and cleaved by adding 200 µl of buffer C.001 (7.5M DCl) and incubating at 23°C for 30 min to produce linear Xpandomer. The sample was then neutralized by adding 2000 µl of buffer B.064 and incubating at room temperature for 2 min. The Xpandomer sample was then amine-modified by adding 500 µmol of succinic anhydride to buffer B.064 and incubating at 23°C for 5 min. The sample was then washed in buffer D.102 (50% ACN), and Xpandomer was released from the substrate by photocutting and eluted in 60 µl of elution buffer.

[0239] Protein nanopores were prepared by inserting α-hemolysin into a DPhPE / hexadecane bilayer membrane buffered in 2 M NH4Cl and 100 mM HEPES at pH 7.4. Cis-wells were perfused with buffer AG242 containing 0.4 M NH4Cl, 600 mM GuanCl, 100 mM HEPES, pH 7.4, and 5% glycerol, while trans-wells were perfused with buffer AB080 containing 0.4 M NH4Cl, 600 mM GuanCl, 5% ethyl acetate, 10 mM HEPES, pH 7.4. Xpandomer samples were heated to 70°C for 2 min, completely cooled, and vortexed before adding 2 µL aliquots to the cis-wells. Voltage parameters were run as follows: 70 mV / 625 mV / 6 µs / 1.0 ms (readout voltage / pulse voltage / pulse voltage duration / pulse frequency). Data were acquired using LabVIEW software.

[0240] A mixture of four XNTPs (complementary to the four naturally occurring nucleobases A, G, C, and T, respectively) was used for Xp synthesis and subsequently extended sequencing in the following forms: 1) diastereomeric pure active isomers, 2) diastereomeric pure inactive isomers, or 3) a mixture of active and inactive isomers in a 90:10 or 75:25 ratio. The results are summarized in Table 2 below:

[0241] Table 2

[0242]

[0243] Using the active isomer at concentrations of 100 µM or higher yielded approximately 40% of the full-length Xpandomer product, while the 75 µM active isomer yielded 33% of the full-length product, thus providing a correlation between the full-length product and active isomer concentrations up to 100 µM. This is also graphically illustrated in Figure 2 (see the data series with a 100:0 ratio). Table 2 further shows that the % full-length value did not change significantly at a concentration of 150 µM active isomer compared to 100 µM, suggesting a possible saturation effect between the 75 µM and 100 µM active isomers. In contrast, using the inactive isomer at 100 µM yielded no full-length Xpandomer product at all. This difference between using diastereomeric pure active isomers and inactive isomers is confirmed by Figure 3, which shows a representational image of the product from the Xpandomer synthesis after gel electrophoresis using active or inactive isomers. Figure 3 The presence of the active isomer demonstrates that it allows for the generation of a full-length Xpandomer, as evidenced by the presence of a band in lane 37. In contrast, the use of the inactive isomer does not generate any full-length Xpandomer, as evidenced by the absence of any bands in lane 38.

[0244] Table 2 also shows that the sequence error rate (deletion, substitution, or insertion-deletion) of the truncated Xpandomer product obtained using the inactive isomer is higher than that of the Xpandomer product synthesized in the presence of the active isomer.

[0245] Furthermore, as further demonstrated in Table 2, the use of the active isomer achieved a higher percentage of full-length Xpandomer compared to a mixture of the two isomers. This was even when the concentration of the active isomer in the mixture was the same as that in the diastereomeric pure formulation: while 100 µM of the pure active isomer produced 40% of the full-length product, the 111 µM 90:10 mixture (containing 100 µM of the active isomer and 11 µM of the inactive isomer) or the 133 µM 75:25 mixture (containing 100 µM of the active isomer and 33 µM of the inactive isomer) produced only 37% and 32% of the full-length product, respectively. Similarly, while 75 µM of the pure active isomer produced 33% of the full-length product, the 100 µM 75:25 mixture (containing 75 µM of the active isomer and 25 µM of the inactive isomer) produced only 29% of the full-length product. This is illustrated graphically in Figure 2. The figure plots the concentration of the active isomer versus the obtained % full-length Xpandomer product. It is evident that the % full-length value decreases at a given concentration of the active isomer in the presence of the inactive isomer, depending on the amount of inactive isomer present: although only a slight negative impact is observed in a 90:10 mixture, the negative impact is more pronounced in a 75:25 mixture.

[0246] The negative impact of inactive isomers on Xpandomer length is also supported by the average length of the obtained products. See Table 2, which shows an average length of 281 nt with a pure active isomer of 100 µM, and a slightly shorter average length of 279 nt with an active isomer containing 100 µM plus 11 or 33 µM of inactive isomer.

[0247] Overall, the data indicate that the yield of full-length Xpandomer product is directly dependent on the presence of the active isomer. The inactive isomer fails to produce any high-quality full-length Xpandomer product, and surprisingly, the presence of the active isomer even has a concentration-dependent negative impact on the yield of full-length Xpandomer product. Without being bound by any theory, one possible explanation for this negative impact is that—despite a strong preference for the active isomer—the polymerase may occasionally incorporate the inactive isomer, which could lead to premature termination of Xpandomer elongation.

[0248] In summary, it is desirable to use compositions containing an excess of the active isomer, for example at least 80%, preferably at least 90%, or even 100% of the active isomer, in the Xpandomer synthesis. Conversely, compositions containing an excess of the inactive isomer may have some utility, such as serving as a negative control for the Xpandomer synthesis.

[0249] sequence

[0250] SEQ ID NO: 1 (DPO4 C4552)

[0251] MIVLFVDFDYFYAQVEEVLNPSLKGKPVVVCVFSGRFEDSGVVATANYEARKFGVYAGIPIVEAKKILPNAVYLPWRDLVYWGVSERIMNLLREYSEKIEIASIDEAYLDISDKVRDYREAYNLGLEIKNKILEKEKITVTVGISKNKVFAAVAGRMAKPNGIKVIDDEE VKRLIRELDIADVQGIPYFTAEKLKKLGINKLVDTLSIEFDKLKGMIGEAKAKYLISLARDEYNEPIRTRVRKSIGRTVTMKRNSRNLEEIKPYLFRAIEESYYKLDKRIPKAIHVVAWKSYWNSQYRWSWFPHGISKETAYSESVQLLQQILKKDKRKIRRIGVRFSKF

[0252] SEQ ID NO: 2 (DPO4 C4760)

[0253] MIVLFVDFDYFYAQVEEVLNPSLKGKPVVVCVFSGRFEDSGVVATANYEARKFGVYAGIPIVRAKKILPNAVYLPWRDLVYWGVSERIMNLLREYSEKIEIASIDEAYLDISDKVRDYREAYNLGLEIKNKILEKEKITVTVGISKNKVFAAVAGRMAKPNGIKVIDDEE VKRLIRELDIADVQGIPYFTAEKLKKLGINKLVDTLSIEFDKLKGMIGEAKAKYLISLARDEYNEPIRTRVRKSIGRTVTMKRNSRNLEEIKPYLFRAIEESYYKLDKRIPKAIHVVAWKSYWNSQYRWSWFPHGISKETAYSESVQLLQQILKKDKRKIRRIGVRFSKF

[0254] SEQ ID NO: 3 (DPO4 C4842)

[0255] MIVLFVDFDYFYAQVEEVLNPSLKGKPVVVCVFSGRFEDSGVVATANYEARKFGVYAGIPIVRAKKILPNAVYLPWRDLVYWGVSERIMNLLREYSEKIEIASIDEAYLDISDKVRDYREAYNLGLEIKNKILEKEKITVTVGISKNKVFAAVAGRMAKPNGIKIVIDDEEVKRLIRELDIADVQGIPYFTAEKLKKLGINKLVDTLSIEFDKLKGMIGEAKAKYLISLARDEYNEPIRTRVRRSIGRTVTKMKRNSRNLEEIKPYLFRAIEESYYKLDKRIPKAIHVVAWKSYWNSQYRWSWFPHGISKETAYSESVQLLQQILKKDKRKIRRIGVRFSKF

[0256] SEQ ID NO: 4 (DPO4 C4852)

[0257] MIVLFVDFDYFYAQVEEVLNPSLKGKPVVVCVFSGRFEDSGVVATANYEARKFGVYAGIPIVRAKKILPNAVYLPWRDLVYWGVSERIMNLLREYSEKIEIASIDEAYLDISDKVRDYREAYNLGLEIKNKILEKEKITVTVGISKNKVFAAVAGRMAKPNGIKIVIDDEEVKRLIRELDIADVQGIPYFTAEKLKKLGINKLVDTLSIEFDKLKGMIGEAKAKYLISLARDEYNEPIRTRVRKSIGRTVTKMKRDSRNLEEIKPYLFRAIEESYYKLDKRIPKAIHVVAWKSYWNSQYRWSWFPHGISKETAYSESVQLLQQILKKDKRKIRRIGVRFSKF

[0258] SEQ ID NO: 5 (DPO4 C4862)

[0259] MIVLFVDFDYFYAQVEEVLNPSLKGKPVVVCVFSGRFEDSGVVATANYEARKFGVYAGIPIKRAKKILPNAVYLPWRDLVYWGVSERIMNLLREYSEKIEIASIDEAYLDISDKVRDYREAYNLGLEIKNKILEKEKITVTVGISKNKVFAAVAGRMAKPNGIKIVIDDEEVKRLIRELDIADVQGIPYFTAEKLKKLGINKLVDTLSIEFDKLKGMIGEAKAKYLISLARDEYNEPIRTRVRKSIGRTVTKMKRNSRNLEEIKPYLFRAIEESYYKLDKRIPKAIHVVAWKSYWNSQYRWSWFPHGISKETAYSESVQLLQQILKKDKRKIRRIGVRFSKF

Claims

1. A composition comprising nucleoside triphosphates having the following structure: or Where NB is a nucleobase; R 1 Contains or is composed of hydrocarbons; R 2 Independently, it is H, OH, or any 2'-ribose modification; R 3 It is H or any protecting group; and R 4 Contains or is composed of hydrocarbons; G 1 and G 2 Independently represents the terminal clickable group; L 1 and L 2 Independently represents the linking group; and T is a chain molecule; Wherein at least 80%, preferably at least 90%, such as at least 95%, at least 99%, or 100%, of the said nucleoside triphosphate has the following stereoconfiguration at the α-phosphatase: 。 2. The composition according to claim 1, wherein at least 90% of the nucleoside triphosphate has the following stereoconfiguration at the α-phosphoramidate ester: 。 3. The composition according to claim 1 or 2, wherein 100% of the nucleoside triphosphate has the following stereoconfiguration at the α-phosphoramidate ester: 。 4. The composition according to any one of the preceding claims, wherein NB is selected from cytosine, thymine, 7-deadenine, and 7-deadenine.

5. The composition according to any one of the preceding claims, wherein R 1 When the nucleobase is a pyrimidine nucleobase, it is attached to position 5 of the nucleobase, and when the nucleobase is a purine nucleobase, it is attached to position 7 of the nucleobase.

6. The composition according to any one of the preceding claims, wherein the NB has one of the following structures: 。 7. The composition according to any one of the preceding claims, wherein the composition comprises a mixture of four different types of nucleoside triphosphates.

8. The composition according to claim 7, wherein the four different types of nucleoside triphosphates comprise four different types of nucleobases.

9. The composition according to claim 7 or 8, wherein the four different types of nucleoside triphosphates are base-paired with guanine, adenine, thymine and cytosine, respectively.

10. The composition according to any one of claims 7 to 9, wherein the four different types of nucleoside triphosphates each comprise the four different types of nucleobases of claim 6.

11. The composition according to any one of the preceding claims, wherein R 1 It contains or is composed of unsaturated hydrocarbons.

12. The composition according to any one of the preceding claims, wherein R 1 It is composed of hydrocarbons such as alkynyl groups.

13. The composition according to any one of the preceding claims, wherein R 1 It is acyclic.

14. The composition according to any one of the preceding claims, wherein R 1 It's a linear chain.

15. The composition according to any one of the preceding claims, wherein R 1 It contains 1 to 20 carbon atoms, such as 1 to 10 carbon atoms or 5 to 10 carbon atoms, such as 6 or 8 carbon atoms.

16. The composition according to any one of the preceding claims, wherein R 1 It is a hex-1-ynyl or oct-1-ynyl group.

17. The composition according to any one of the preceding claims, wherein R 1 With G 1 Together they are oct-1,7-diynyl or dec-1,9-diynyl groups.

18. The composition according to any one of the preceding claims, wherein the two R 2 Both are H.

19. The composition according to any one of the preceding claims, wherein R 3 It's H.

20. The composition according to any one of the preceding claims, wherein R 4 It contains saturated hydrocarbons or is composed of saturated hydrocarbons.

21. The composition according to any one of the preceding claims, wherein R 4 It contains 1 to 20 carbon atoms, such as 3 to 15 carbon atoms, 3 to 10 carbon atoms, such as 4 carbon atoms.

22. The composition according to any one of the preceding claims, wherein R 4 It is acyclic.

23. The composition according to any one of the preceding claims, wherein R 4 It's a linear chain.

24. The composition according to any one of the preceding claims, wherein R 4 It is composed of hydrocarbons.

25. The composition according to any one of the preceding claims, wherein R 4 It is a n-butyl group.

26. The composition according to any one of the preceding claims, wherein R 4 With G 2 Together they form a hex-5-ynyl group.

27. The composition according to any one of claims 1 to 23, wherein R 4 It contains or is composed of two or more hydrocarbons, which are connected by atoms or atomic groups other than carbon, such as phosphorus atoms and / or oxygen atoms.

28. The composition according to any one of the preceding claims, wherein the terminal clickable group is a terminal alkyne group or a terminal azide group, preferably a terminal alkyne group.

29. The composition according to any one of the preceding claims, wherein G 1 and G 2 Indicates the same type of terminal clickable groups.

30. The composition according to any one of the preceding claims, wherein L 1 and L 2 Independently represents the linking group formed via a click reaction.

31. The composition according to any one of the preceding claims, wherein L 1 and L 2 Each of them is a 1,2,3-triazole.

32. The composition according to any one of the preceding claims, wherein at least 80%, preferably at least 90%, such as at least 95%, at least 99%, or 100%, of the nucleoside triphosphate has a structure selected from the following: 。 33. The composition according to any one of the preceding claims, wherein at least 80%, preferably at least 90%, such as at least 95%, at least 99%, or 100%, of the nucleoside triphosphate has a structure selected from the following: 。 34. The composition according to any one of claims 1 to 29, wherein at least 80%, preferably at least 90%, such as at least 95%, at least 99%, or 100%, of the nucleoside triphosphate has a structure selected from the following: 。 35. The composition according to any one of claims 1 to 29, wherein at least 80%, preferably at least 90%, such as at least 95%, at least 99%, or 100% of the nucleoside triphosphate has a structure selected from the following: 。 36. A method for generating a complementary strand of a nucleic acid, the method comprising contacting the nucleic acid with a composition according to any one of claims 1 to 33.

37. The method of claim 36, wherein the nucleic acid is contained in a nucleic acid library.

38. The method according to claim 36 or 37, wherein the nucleic acid is DNA.

39. The method according to any one of claims 36 to 38, wherein the composition further comprises a nucleic acid polymerase.

40. The method according to any one of claims 36 to 39, wherein the composition further comprises a buffer such as TrisCl and / or a polymerase cofactor such as MnCl2.

41. The method according to any one of claims 36 to 40, the method comprising hybridizing primers with the nucleic acid while or subsequently contacting the nucleic acid with the composition.

42. A method for determining the sequence of nucleic acids, the method comprising the following steps performed in sequence: 1) The complementary strand of the nucleic acid is generated by the method according to any one of claims 36 to 41; 2) Selectively cleave the PN bonds within the nucleoside triphosphate to generate an extended complementary chain; 3) Sequencing the extended complementary strand. 4) Determine the sequence of the nucleic acid based on the sequence of the extended complementary strand.

43. The method of claim 42, wherein in step 3), the PN bonds are selectively cleaved under acidic conditions.

44. The method of claim 42 or 43, wherein the extended complementary strand is sequenced in step 4) by nanopore-based sequencing.

45. The method of claim 44, wherein the nanopore-based sequencing comprises: (a) A chip for nanopore-based sequencing is provided, the chip comprising: (i) An electrochemical resistance barrier disposed above an aperture on the surface of the chip, wherein the barrier separates the cis side and the trans side; (ii) A nanopore inserted into the barrier, wherein the nanopore has an inlet side on the cis side of the barrier and an outlet side on the trans side of the barrier; (b) Contact the cis side of the barrier with the extended complementary chain; (c) Apply a voltage across the barrier of the chip to cause the extended complementary chain to translocate to the inverted side; (d) Determine one or more changes in the electrical properties of the nanopore associated with the extended complementary chain occupying the nanopore during the translocation; and (e) Determine the sequence of the extended complementary chain based on one or more variations in the electrical properties of the nanopore.

Citation Information

Patent Citations

  • Enhancement of nucleic acid polymerization by aromatic compounds

    EP3735409B1

  • Method for incorporating into a DNA or RNA oligonucleotide using nucleotides bearing heterocyclic bases

    US5432272A

  • Modified oligonucleotides, their preparation and their use

    US6150510A

  • High throughput nucleic acid sequencing by expansion

    US7939259B2

  • Nuclease resistant, pyrimidine modified oligonucleotides that detect and modulate gene expression

    WO1992002258A1