Nucleosides and nucleotides having 3 '-hydroxy blocking groups and their use in polynucleotide sequencing methods
By using nucleosides and nucleotides with 3'-OH acetal or thiocarbamate blocking groups, the problem of insufficient nucleotide stability in existing technologies is solved, thereby improving sequencing stability and data quality and extending read time.
Patent Information
- Application Number
- CN202511228365.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2018-12-26
- Filing Date
- 2019-12-23
- Publication Date
- 2025-12-16
AI Technical Summary
Existing 3'-hydroxy protecting groups are not stable enough during nucleotide sequencing, affecting sequencing accuracy and efficiency. They are also difficult to maintain long-term stability in solution and signal stability when operating on sequencing instruments.
Nucleosides and nucleotides with 3'-OH acetal or thiocarbamate blocking groups were used to improve stability in solution and achieve low pre-phase and signal attenuation during sequencing, thereby improving data quality.
It improves the stability of nucleotides and signals during sequencing, extends read time, and improves sequencing data quality.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
[0001] This application is a divisional application of a Chinese patent application (Filing date: 23 December 2019; Application No. 201980042583.5 (International Application No. PCT / EP2019 / 086926); Invention title: Nucleotides, nucleosides with 3’-hydroxyl blocking groups and their use in polynucleotide sequencing methods).
[0002] BACKGROUND TECHNICAL FIELD
[0003] The present disclosure relates generally to nucleotides, nucleosides or oligonucleotides comprising a 3’-hydroxyl protecting group and their use in polynucleotide sequencing methods. Methods of making 3’-hydroxyl protected nucleotides, nucleosides or oligonucleotides are also disclosed.
[0004] Related Art
[0005] Advances in molecular research have been due, in part, to improvements in techniques for characterizing molecules or their biological responses. In particular, research on nucleic acids, DNA and RNA, has benefited from the development of techniques for sequence analysis and studies of hybridization events.
[0006] One example of a technique that has improved nucleic acid research is the development of fabricated arrays of immobilized nucleic acids. These arrays typically consist of a high density matrix of polynucleotides immobilized on a solid support material. See, for example, Fodor et al., Trends Biotech. 12: 19-26, 1994, which describes a method for assembling nucleic acids using a chemically sensitized glass surface to which nucleotide phosphoramidites are attached, protected by a mask but exposed at defined regions to allow for appropriate modification. Fabricated arrays can also be made by techniques in which known polynucleotides are “spotted” onto predetermined locations on a solid support (e.g., Stimpson et al., Proc. Natl. Acad. Sci. 92: 6379-6383, 1995).
[0007] Methods for determining the nucleotide sequence of nucleic acids bound to an array are referred to as “sequencing by synthesis” or “SBS”. This technique for determining DNA sequence ideally requires the controlled incorporation (i.e., one at a time) of the correct complementary nucleotide opposite the nucleic acid being sequenced. Since each nucleotide residue is sequenced one at a time, accurate sequencing can be performed by adding nucleotides in multiple cycles, preventing a series of uncontrolled incorporations from occurring. The incorporated nucleotide is read using an appropriate label attached to it before the label moiety is removed and the next round of sequencing proceeds.
[0008] To ensure that only a single incorporation occurs, each tagged nucleotide, which is added to the growing chain to ensure that only one nucleotide is incorporated, includes a structural modification (protecting group or blocking group). After the nucleotide with the protecting group is added, the protecting group is removed under reaction conditions that do not interfere with the integrity of the DNA being sequenced. The sequencing cycle can then continue with the incorporation of the next protected, tagged nucleotide.
[0009] To be useful for DNA sequencing, nucleotides, typically nucleoside triphosphates, generally require a 3'-hydroxyl protecting group to prevent the polymerase enzyme used to incorporate the nucleotide into a polynucleotide chain from continuing to replicate once the base on the nucleotide has been added. The types of groups that can be added to a nucleotide and still be suitable are limited. The protecting group should prevent additional nucleotide molecules from being added to the polynucleotide chain while being easily removed from the sugar moiety without causing damage to the polynucleotide chain. In addition, the modified nucleotide needs to be compatible with the polymerase enzyme or another suitable enzyme used to incorporate it into a polynucleotide chain. Thus, the ideal protecting group must exhibit long-term stability, be efficiently incorporated by the polymerase, prevent secondary or further incorporation of the nucleotide, and have the ability to be removed under mild conditions, preferably in aqueous conditions, without destroying the structure of the polynucleotide.
[0010] Reversible protecting groups have been described previously. For example, Metzker et al. (Nucleic Acids Research, 22 (20): 4259-4267, 1994) disclose the synthesis and use of eight 3'-modified 2-deoxyribonucleoside 5'-triphosphates (3'-modified dNTPs) and test incorporation activity in two DNA template assays. WO 2002 / 029003 describes a sequencing method that can include the use of an allyl protecting group to block the 3'-OH group on the growing strand of DNA in a polymerase reaction.
[0011] In addition, the development of a number of reversible protecting groups and methods for their deprotection under DNA compatible conditions have been previously reported in International Application Publications WO 2004 / 018497 and WO 2014 / 139596, each of which is incorporated herein by reference in its entirety.
[0012] SUMMARY
[0013] Some embodiments of the present disclosure relate to a ribose or deoxyribose containing nucleoside or nucleotide having a removable 3'-OH blocking group that forms a structure covalently linked to the 3'-carbon atom wherein:
[0014] each R1a and R 1b independently H, C1-C6 alkyl, C1-C6 haloalkyl, C1-C6 alkoxy, C1-C6 haloalkoxy, cyano, halogen, optionally substituted phenyl, or optionally substituted aralkyl;
[0015] each R 2a and R 2b independently H, C1-C6 alkyl, C1-C6 haloalkyl, cyano, or halogen;
[0016] or, R 1a and R 2a together with the atoms to which they are attached form an optionally substituted five- to eight-membered heterocyclyl;
[0017] R 3 is H, optionally substituted C 2- C6 alkenyl, optionally substituted C 3- C7 cycloalkenyl, optionally substituted C 2- C6 alkynyl, or optionally substituted (C1-C6 alkylene)Si(R 4 )3; and
[0018] each R 4 independently H, C1-C6 alkyl, or optionally substituted C 6- C 10 aryl. In some embodiments, when each R 1a and R 1b is H or C1-C6 alkyl, R 2a and R 2b are both H, then R 3 is substituted C 2- C6 alkenyl, optionally substituted C 3- C7 cycloalkenyl, optionally substituted C 2- C6 alkynyl, or optionally substituted (C1-C6 alkylene)Si(R 4 )3. In some embodiments, when each R 1a , R 1b , R 2a , and R 2b is H, then R 3 is not H.
[0019] Some embodiments of the disclosure relate to a nucleoside or nucleotide comprising a ribose or deoxyribose sugar having a removable 3'-OH blocking group forming a structure covalently linked to the 3'-carbon atom wherein:
[0020] each R 5 and R 6independently H, C1-C6alkyl, C 2- C6alkenyl, C 2- C6alkynyl, C1-C6haloalkyl, C 2- C8alkoxyalkyl, optionally substituted -(CH2) m -phenyl, optionally substituted -(CH2) n -(5- or 6-membered heteroaryl), optionally substituted -(CH2) k -C 3- C7carbocyclyl, or optionally substituted -(CH2) p -(3- to 7-membered heterocyclyl);
[0021] each -(CH2) m -, -(CH2) n -, -(CH2) k - and -(CH2) p - is optionally substituted; and
[0022] each m, n, k, and p is independently 0, 1, 2, 3, or 4.
[0023] Some embodiments of the present disclosure relate to oligonucleotides or polynucleotides comprising a 3’-OH blocked nucleotide molecule described herein.
[0024] Some embodiments of the present disclosure relate to methods of preparing a growing polynucleotide complementary to a target single-stranded polynucleotide in a sequencing reaction comprising incorporating a nucleotide molecule described herein into the growing complementary polynucleotide, wherein incorporation of the nucleotide prevents any subsequent nucleotide from being introduced into the growing complementary polynucleotide. In some embodiments, the incorporation of the nucleotide is accomplished by a polymerase, terminal deoxynucleotidyl transferase (TdT), or reverse transcriptase. In one embodiment, the incorporation is accomplished by a polymerase (e.g., a DNA polymerase).
[0025] Some further embodiments of the present disclosure relate to methods of determining a sequence of a target single-stranded polynucleotide comprising:
[0026] (a) incorporating a nucleotide described herein comprising a 3’-OH blocking group and a detectable label into a replication polynucleotide strand complementary to at least a portion of the target polynucleotide strand;
[0027] (b) detecting the identity of the nucleotide incorporated into the replication polynucleotide strand; and
[0028] (c) chemically removing the label and the 3’-OH blocking group from the nucleotide incorporated into the replication polynucleotide strand.
[0029] In some embodiments, the sequencing method further comprises (d) washing away the chemically removed label and 3' blocking group from the replicated polynucleotide strand. In some embodiments, this washing step also removes unincorporated nucleotides. In some such embodiments, the 3' blocking group and detectable label of an incorporated nucleotide are removed prior to incorporation of the next complementary nucleotide. In some further embodiments, the 3' blocking group and detectable label are removed in a single chemical reaction step. In some embodiments, the sequencing incorporation described herein is performed at least 50 times, at least 100 times, at least 150 times, at least 200 times, or at least 250 times.
[0030] Some further embodiments of the disclosure relate to kits comprising a plurality of the nucleotide or nucleoside molecules described herein and packaging materials therefor. The nucleotides, nucleosides, oligonucleotides, or kits set forth herein can be used to detect, measure, or identify a biological system (including, e.g., a process or component thereof). Exemplary techniques that can use the nucleotides, oligonucleotides, or kits include sequencing, expression analysis, hybridization analysis, genetic analysis, RNA analysis, cellular assays (e.g., cell binding or cell function analysis), or protein assays (e.g., protein binding assays or protein activity assays). The use can be implemented on an automated instrument for performing the particular technique, such as an automated sequencing instrument. The sequencing instrument can contain two or more lasers operating at different wavelengths to distinguish different detectable labels. BRIEF DESCRIPTION OF DRAWINGS
[0031] Figure 1 A line graph showing the stability of various 3' blocked nucleotides over time in a buffered solution at 65 °C.
[0032] Figure 2A A line graph showing the percentage (%) of remaining nucleotides (starting material) over time comparing the deblocking rate of nucleotides with 3'-AOM blocking groups to nucleotides with 3'-0-azidomethyl (-CH2N3) blocking groups in solution.
[0033] Figure 2B A line graph showing the percentage (%) of 3' deblocked nucleotides over time comparing the deblocking rate of 3' blocked nucleotides with various acetal blocking groups in solution.
[0034] Figure 3A and 3B Sequencing results on an Illumina MiniSeq® instrument using fully functionalized nucleotides (ffN) with 3'-AOM blocking groups in the incorporation mix are shown.
[0035] Figure 3CGraph showing sequencing error rates using fully functionalized nucleotides (ffNs) with 3'-AOM blocking group in incorporation mix compared to standard ffNs with 3'-O-azido blocking group.
[0036] Figure 4A and 4B Each shows a comparison of key sequencing metrics including phasing, pre-phasing, and error rates for fully functionalized nucleotides with 3'-AOM and 3'-O-azidomethyl blocking groups using two different DNA polymerases (Pol 812 and Pol 1901), respectively.
[0037] Figure 5 Graph showing the change in sequencing stability over time for fully functionalized nucleotides with 3'-AOM or 3'-O-azidomethyl blocking groups in a buffered solution at 45 °C.
[0038] Figure 6 Graph showing the change in stability over time for nucleosides with various 3' blocking groups in a buffered solution at 65 °C.
[0039] Figure 7 Graph showing the change in percentage (%) of remaining 3' blocked nucleotides over time comparing the rate of cleavage (deblocking) of the thiocarbamate 3' blocking group dimethylthiocarbamate (DMTC) to the 3'-O-azidomethyl (3'-O-CH2N3) blocking group under two different conditions (Oxone® or NaIO4). DETAILED DESCRIPTION
[0040] Embodiments of the present disclosure relate to nucleosides and nucleotides with 3'-OH acetal or thiocarbamate blocking groups for sequencing applications, e.g., sequencing by synthesis (SBS). These blocking groups provide better stability in solution compared to those known in the art. In particular, the 3'-OH blocking groups have improved stability during the synthesis of fully functionalized nucleotides (ffNs) and also have good stability in solution during formulation, storage, and operation on sequencing instruments. Additionally, the 3'-OH blocking groups described herein also enable low pre-phasing, lower signal decay to improve data quality, which allows for longer reads from sequencing applications.
[0041] Definitions
[0042] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. The use of the term “comprising” and other forms of the term (e.g., “include,” “includes,” and “included”) is not restrictive. The use of the term “having” and other forms of the term (e.g., “have,” “has,” and “had”) is not restrictive. As used herein, whether in transitional phrases or in the body of the claims, the terms “comprise(s)” and “comprising” should be interpreted in an open-ended sense. That is, the above terms are to be interpreted synonymously with the phrases “at least having” or “at least comprising.” For example, when used in the context of a process, the term “comprising” means that the process includes at least the listed steps, but may include other steps. When used in the context of a compound, composition, or device, the term “comprising” means that the compound, composition, or device includes at least the listed features or components, but may also include other features or components.
[0043] As used in this article, common organic abbreviations are defined as follows:
[0044] Temperature (°C)
[0045] dATP deoxyadenosine triphosphate
[0046] dCTP deoxycytidine triphosphate
[0047] dGTP deoxyguanosine triphosphate
[0048] dTTP deoxythymidine triphosphate
[0049] ddNTP dideoxynucleotide triphosphate
[0050] ffN fully functionalized nucleotides
[0051] RT room temperature
[0052] SBS sequencing-by-synthesis
[0053] SM raw materials
[0054] As used herein, the term "array" refers to a population of different probe molecules attached to one or more substrates such that the different probe molecules can be distinguished from one another according to relative position. An array can include different probe molecules, each at a different addressable position on a substrate. Alternatively or additionally, an array can include mutually separate substrates each bearing different probe molecules, where the different probe molecules can be identified according to the position of the substrate on the surface to which it is attached or according to the position of the substrate in a liquid. Exemplary arrays in which mutually separate substrates are positioned on a surface include, but are not limited to, those including beads in wells, such as described in U.S. Patent 6,355,431 Bl, US 2002 / 0102578, and PCT Publication WO 00 / 63437. Exemplary formats that can be used in the present application to distinguish beads in a liquid array, for example, using a microfluidic device such as a fluorescence activated cell sorter (FACS), are described in, for example, U.S. Patent 6,524,793. Other examples of arrays that can be used in the present application include, but are not limited to, those described in U.S. Patents 5,429,807; 5,436,327; 5,561,071; 5,583,211; 5,658,734; 5,837,858; 5,874,219; 5,919,523; 6,136,269; 6,287,768; 6,287,776; 6,288,220; 6,297,006; 6,291,193; 6,346,413; 6,416,949; 6,482,591; 6,514,751; and 6,610,482; as well as WO 93 / 17126; WO 95 / 11995; WO 95 / 35505; EP 742 287; and EP 799 897.
[0055] As used herein, the term "covalently attached" or "covalently bonded" refers to the formation of a chemical bond characterized by the sharing of electron pairs between atoms. For example, a covalently attached polymer coating refers to a polymer coating that forms a chemical bond with the functionalized surface of a substrate, as opposed to via other means such as adhesion or electrostatic interactions. It will be appreciated that a polymer covalently attached to a surface can also be bonded by means other than covalent attachment.
[0056] As used herein, any "R" group represents a substituent that can be attached to the atom shown. R groups can be substituted or unsubstituted. If two "R" groups are described as "together with the atom to which they are attached" forming a ring or ring system, it is meant that the atom, the intervening bond, and the collective of the two R groups are the enumerated ring. For example, when the following substructure is present:
[0057]
[0058] and R 1 and R 2 is defined as selected from hydrogen and alkyl or R 1 and R 2 together with the atom to which they are attached form an aryl or carbocyclyl, it means that R 1 and R 2 may be selected from hydrogen or alkyl or a substructure having the following structure:
[0059]
[0060] wherein A is an aryl ring or carbocyclyl containing the double bond shown.
[0061] It should be understood that the nomenclature of some groups can include a mono-radical or di-radical depending on the context. For example, when a substituent requires two points of attachment to the rest of the molecule, it should be understood that the substituent is a di-radical. For example, a substituent identified as an alkyl group requiring two points of attachment includes di-radicals such as -CH2-, -CH2CH2-, -CH2CH(CH3)CH2-, and the like. Other group nomenclature clearly indicates that the group is a di-radical, such as “alkylene” or “alkenylene”.
[0062] As used herein, the term “halogen” or “halo” refers to any radiation-stable atom of column 7 of the Periodic Table of Elements, for example, fluorine, chlorine, bromine, or iodine, with fluorine and chlorine being preferred.
[0063] As used herein, “C a to C b“C1-C6alkyl” includes C1, C2, C3, C4, C5, and C6alkyl, C2-C6alkyl, C1-C3alkyl, and the like. Similarly, C1-C6alkylene includes C1, C2, C3, C4, C5, and C6alkylene, C2-C6alkylene, C1-C3alkylene, and the like. C1-C6alkenyl includes C2, C3, C4, C5, and C6alkenyl, C2-C5alkenyl, C3-C4alkenyl, and the like. C1-C6alkynyl includes C2, C3, C4, C5, and C6alkynyl, C2-C5alkynyl, C3-C4alkynyl, and the like. C3-C8cycloalkyl each includes a hydrocarbon ring containing 3, 4, 5, 6, 7, and 8 carbon atoms or a range defined by either of the two numbers, such as C3-C7cycloalkyl or C5-C6cycloalkyl. 2- C6alkenyl, C2-C5alkenyl, C3-C4alkenyl, and the like. C1-C6alkynyl includes C2, C3, C4, C5, and C6alkynyl, C2-C5alkynyl, C3-C4alkynyl, and the like. C3-C8cycloalkyl each includes a hydrocarbon ring containing 3, 4, 5, 6, 7, and 8 carbon atoms or a range defined by either of the two numbers, such as C3-C7cycloalkyl or C5-C6cycloalkyl. 2- C6alkynyl, C2-C5alkynyl, C3-C4alkynyl, and the like. C3-C8cycloalkyl each includes a hydrocarbon ring containing 3, 4, 5, 6, 7, and 8 carbon atoms or a range defined by either of the two numbers, such as C3-C7cycloalkyl or C5-C6cycloalkyl.
[0064] As used herein, “alkyl” refers to a straight-chain or branched-chain hydrocarbon chain that is fully saturated (i.e., contains no double or triple bonds), i.e., an alkyl group. The alkyl group can have from 1 to 20 carbon atoms (whenever it appears, a numerical range, such as “1 to 20,” refers to each integer within the given range— e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, etc. —and includes the endpoints as naturally included in the range; “1 to 20 carbon atoms” means that the alkyl group can consist of 1 carbon atom, 2 carbon atoms, 3 carbon atoms, etc., up to and including 20 carbon atoms, although the definition also contemplates that the alkyl group contains numbers outside of this range, e.g., 23 carbon atoms, 50 carbon atoms, 1,000 carbon atoms, etc.). The alkyl group can also be a medium size alkyl group having from 1 to 9 carbon atoms. The alkyl group can also be a lower alkyl group having from 1 to 6 carbon atoms. The alkyl group can be designated as “Ci-C4alkyl” or similar designations nomenclature. By way of example only, “C1-C6alkyl” denotes an alkyl chain of between one and six carbon atoms, i.e., the alkyl chain is selected from the group consisting of methyl, ethyl, propyl, isopropyl, n-butyl, isobutyl, sec-butyl, and t-butyl. Typical alkyl groups include, but are not limited to, methyl, ethyl, propyl, isopropyl, butyl, isobutyl, t-butyl, pentyl, hexyl, and the like.
[0065] As used herein, "alkoxy" refers to the group -OR, where R is alkyl as defined above, e.g., "C1-C9 alkoxy", including but not limited to methoxy, ethoxy, n-propyloxy, 1-methylethoxy (isopropoxy), n-butyloxy, isobutyloxy, sec-butyloxy, and t-butyloxy, and the like.
[0066] As used herein, "alkenyl" refers to a straight or branched hydrocarbon chain that contains one or more double bonds. The alkenyl group can have 2 to 20 carbon atoms, although the present definition encompasses alkenyl groups wherein no number range is specified. The alkenyl group can also be a medium size alkenyl having 2 to 9 carbon atoms. The alkenyl group can also be a lower alkenyl having 2 to 6 carbon atoms. The alkenyl group can be designated as "C 2- C6alkenyl" or like designations. By way of example only, "C 2- C6alkenyl" means that there are two to six carbon atoms in the alkenyl chain, i.e., the alkenyl chain is selected from the group consisting of ethenyl, propen-1-yl, propen-2-yl, propen-3-yl, buten-1-yl, buten-2-yl, buten-3-yl, buten-4-yl, 1-methylpropen-1-yl, 2-methylpropen-1-yl, 1-ethyl-ethen-1-yl, 2-methyl-propen-3-yl, buta-1,3-dienyl, buta-1,2-dienyl, and buta-1,2-dien-4-yl. Typical alkenyl groups include, but are in no way limited to, ethenyl, propenyl, butenyl, pentenyl, and hexenyl, and the like.
[0067] As used herein, "alkynyl" refers to a straight or branched hydrocarbon chain that contains one or more triple bonds. The alkynyl group can have 2 to 20 carbon atoms, although the present definition encompasses alkynyl groups wherein no number range is specified. The alkynyl group can also be a medium size alkynyl having 2 to 9 carbon atoms. The alkynyl group can also be a lower alkynyl having 2 to 6 carbon atoms. The alkynyl group can be designated as "C 2- C6alkynyl" or like designations. By way of example only, "C 2- C6alkynyl" means that there are two to six carbon atoms in the alkynyl chain, i.e., the alkynyl chain is selected from the group consisting of ethynyl, propyn-1-yl, propyn-2-yl, butyn-1-yl, butyn-3-yl, butyn-4-yl, and 2-butynyl. Typical alkynyl groups include, but are in no way limited to, ethynyl, propynyl, butynyl, pentynyl, and hexynyl, and the like.
[0068] As used herein, "heteroalkyl" refers to a straight-chain or branched- chain hydrocarbon chain that contains one or more heteroatoms (i.e., an element other than carbon, including but not limited to nitrogen, oxygen, and sulfur) in the chain backbone. The heteroalkyl can have 1 to 20 carbon atoms, although the present definition also covers the occurrence of the term "heteroalkyl" where no numerical range is specified. The heteroalkyl can also be a medium size heteroalkyl having 1 to 9 carbon atoms. The heteroalkyl can also be a lower heteroalkyl having 1 to 6 carbon atoms. The heteroalkyl can be designated as "Ci-C6heteroalkyl" or similar designations. The heteroalkyl can contain one or more heteroatoms. By way of example only, "Ci-C6heteroalkyl" means that there are between one and six carbon atoms in the heteroalkyl chain and there is one additional heteroatom in the chain backbone of the chain. 4- "C6heteroalkyl" means that there are between four and six carbon atoms in the heteroalkyl chain and there is one additional heteroatom in the chain backbone of the chain.
[0069] The term "aromatic" refers to a ring or ring system having a conjugated pi-electron system and includes carbocyclic aromatic groups (e.g., phenyl) and heterocyclic aromatic groups (e.g., pyridine). The term includes single rings or fused ring polycyclic (i.e., rings that share pairs of adjacent atoms) groups, provided that the overall ring system is aromatic.
[0070] As used herein, "aryl" refers to an aromatic ring or ring system (i.e., two or more fused rings that share two adjacent carbon atoms) that contains only carbon in the ring backbone. When the aryl is a ring system, every ring in the system is aromatic. The aryl can have 6 to 18 carbon atoms, although the present definition also covers occurrences where no numerical range is specified for the term "aryl." In some embodiments, the aryl has 6 to 10 carbon atoms. The aryl can be designated as "C 6- C 10 aryl," "C6or C 10 aryl," or similar designations. Examples of aryl groups include, but are not limited to, phenyl, naphthyl, azulenyl, and anthracenyl.
[0071] "Arylalkyl" or "arylalkyl" is an aryl group as a substituent connected via an alkylene group, e.g., "C 7-14 arylalkyl" and the like, including but not limited to benzyl, 2-phenylethyl, 3-phenylpropyl, and naphthylalkyl. In some cases, the alkylene group is a lower alkylene group (i.e., Ci-C6alkylene).
[0072] As used herein, "heteroaryl" refers to an aromatic ring or ring system (i.e., two or more fused rings sharing two adjacent atoms) containing one or more heteroatoms (i.e., elements other than carbon, including but not limited to nitrogen, oxygen, and sulfur) in the ring backbone. When the heteroaryl is a ring system, each ring in the system is aromatic. The heteroaryl can have 5-18 ring members (i.e., the number of atoms comprising the ring backbone, including carbon atoms and heteroatoms), although the present definition also encompasses occurrences of the term "heteroaryl" where no numerical range is specified. In some embodiments, the heteroaryl has 5 to 10 ring members or 5 to 7 ring members. The heteroaryl can be designated as a "5-7 membered heteroaryl," "5-10 membered heteroaryl," or similar designation. Examples of heteroaryl rings include, but are not limited to, furanyl, thienyl, phthalazinyl, pyrrolyl, oxazolyl, thiazolyl, imidazolyl, pyrazolyl, isoxazolyl, isothiazolyl, triazolyl, thiadiazolyl, pyridyl, pyridazinyl, pyrimidinyl, pyrazinyl, triazinyl, quinolinyl, isoquinolinyl, benzimidazolyl, benzoxazolyl, benzothiazolyl, indolyl, isoindolyl, and benzothienyl.
[0073] "Heteroaralkyl" or "heteroarylalkyl" is a heteroaryl group as a substituent linked via an alkylene group. Examples include, but are not limited to, 2-thienylmethyl, 3-thienylmethyl, furanylmethyl, thienylethyl, pyrrolylalkyl, pyridylalkyl, isoxazolylalkyl, and imidazolylalkyl. In some cases, the alkylene group is lower alkylene (i.e., C1-C6 alkylene).
[0074] As used herein, "carbocyclyl" refers to a non-aromatic ring or ring system comprising only carbon atoms in the ring system backbone. When the carbocyclyl is a ring system, two or more rings can be fused, bridged, or spiro-linked together. The carbocyclyl can have any degree of saturation, provided that at least one ring in the ring system is not aromatic. Thus, carbocyclyl includes cycloalkyl, cycloalkenyl, and cycloalkynyl. The carbocyclyl group can have 3 to 20 carbon atoms, although the present definition also encompasses occurrences of the term "carbocyclyl" where no numerical range is specified. The carbocyclyl group can also be a medium size carbocyclyl having 3 to 10 carbon atoms. The carbocyclyl group can also be a carbocyclyl having 3 to 6 carbon atoms. The carbocyclyl group can be designated as a "C3-C6 carbocyclyl" or similar designation. Examples of carbocyclyl rings include, but are not limited to, cyclopropyl, cyclobutyl, cyclopentyl, cyclohexyl, cyclohexenyl, indane, bicyclo[2.2.2]octyl, adamantyl, and spiro[4.4]nonanyl. 3- C6 carbocyclyl" or similar designation. Examples of carbocyclyl rings include, but are not limited to, cyclopropyl, cyclobutyl, cyclopentyl, cyclohexyl, cyclohexenyl, indane, bicyclo[2.2.2]octyl, adamantyl, and spiro[4.4]nonanyl.
[0075] As used herein, "cycloalkyl" refers to a fully saturated carbocyclyl ring or ring system. Examples include cyclopropyl, cyclobutyl, cyclopentyl, and cyclohexyl.
[0076] As used herein, "heterocyclyl" refers to a non-aromatic ring or ring system that includes at least one heteroatom in the ring backbone. Heterocyclyl groups can be linked together in a fused, bridged, or spiro-connected manner. Heterocyclyl groups can have any degree of saturation, provided that at least one ring in the ring system is not aromatic. The heteroatoms can be present in non-aromatic or aromatic rings of the ring system. The heterocyclyl group can have from 3 to 20 ring members (i.e., the number of atoms comprising the ring backbone, including carbon atoms and heteroatoms), although the present definition also encompasses occurrences of the term "heterocyclyl" where no number range is specified. The heterocyclyl group can also be a medium size heterocyclyl group having from 3 to 10 ring members. The heterocyclyl group can also be a heterocyclyl group having from 3 to 6 ring members. The heterocyclyl group can be designated as a "3-6 membered heterocyclyl" or similar designation. In preferred six-membered monocyclic heterocyclyl groups, the heteroatoms are selected from up to three of O, N, or S, and in preferred five-membered monocyclic heterocyclyl groups, the heteroatoms are selected from one or two heteroatoms selected from O, N, or S. Examples of heterocyclyl rings include, but are not limited to, azepinyl, acridinyl, carbazolyl, cinnolinyl, dioxolanyl, imidazolinyl, imidazolidinyl, morpholinyl, oxiranyl, oxepanyl, thiepanyl, piperidinyl, piperazinyl, dioxopiperazinyl, pyrrolidinyl, pyrrolidinonyl, pyrrolidinedionyl, 4-piperidonyl, pyrazolinyl, pyrazolidinyl, 1,3-dioxinyl, 1,3-dioxanyl, 1,4-dioxinyl, 1,4-dioxanyl, 1,3-oxathianyl, 1,4-oxathianyl, 1,4-oxathianyl, 2H-1,2-oxazinyl, trioxanyl, hexahydro-1,3,5-triazinyl, 1,3-dioxolyl, 1,3-dioxolane, 1,3-dithiolyl, 1,3-dithiolane, isoxazolinyl, isoxazolidinyl, oxazolinyl, oxazolidinyl, oxazolonyl, thiazolinyl, thiazolidinyl, 1,3-oxathiolanyl, indolyl, isoindolyl, tetrahydrofurfuryl, tetrahydropyranyl, tetrahydrothiophenyl, tetrahydrothiopyranyl, tetrahydro-1,4-thiazinyl, thiomorpholinyl, dihydrobenzofuranyl, benzimidazolidinyl, and tetrahydroquinoline.
[0077] "O-carboxy" group means "-OC(=0)R", where R is selected from hydrogen, C1-C6 alkyl, C 2- C6 alkenyl, C 2- C6 alkynyl, C 3- C7 carbocyclyl, C 6- C 10 aryl, 5-10 membered heteroaryl, and 3-10 membered heterocyclyl.
[0078] "C-carboxyl group" refers to the "-C(=O)OR" group, where R is selected from hydrogen, C1-C6 alkyl, C6, and C6 groups as defined in this article. 2- C6 alkenyl, C 2- C6 ynyl group, C 3- C7 carbon cyclogroup, C 6- C 10 Aryl, 5-10 membered heteroaryl and 3-10 membered heterocyclic groups. Non-limiting examples include methylcarboxyl (i.e., -C(=O)OH).
[0079] "Sulfonyl" refers to "-SO2R", where R is selected from hydrogen, C1-C6 alkyl, C2, and C2 as defined herein. 2- C6 alkenyl, C 2- C6 ynyl group, C 3- C7 carbon cyclogroup, C 6- C 10 Aryl, 5-10 heteroaryl and 3-10 heterocyclic groups.
[0080] The "sulfinyl" group refers to the "-S(=O)OH" group.
[0081] "S-sulfonamide group" refers to "-SO2NR". A R B ", where R A and R B Each is independently selected from hydrogen, C1-C6 alkyl, C... as defined herein. 2- C6 alkenyl, C 2- C6 ynyl group, C 3- C7 carbon cyclogroup, C 6- C 10 Aryl, 5-10 heteroaryl and 3-10 heterocyclic groups.
[0082] The "N-sulfonamide group" refers to "-N(R)". A SO2R B "group, wherein R A and R B Each is independently selected from hydrogen, C1-C6 alkyl, C... as defined herein. 2- C6 alkenyl, C 2- C6 ynyl group, C 3- C7 carbon cyclogroup, C 6- C 10 Aryl, 5-10 heteroaryl and 3-10 heterocyclic groups.
[0083] The "C-amide group" refers to "-C(=O)NR". A R B "group, wherein R A and R BEach is independently selected from hydrogen, C1-C6 alkyl, C... as defined herein. 2- C6 alkenyl, C 2- C6 ynyl group, C 3- C7 carbon cyclogroup, C 6- C 10 Aryl, 5-10 heteroaryl and 3-10 heterocyclic groups.
[0084] The "N-amide group" refers to "-N(R)". A )C(=O)R B "group, wherein R A and R B Each is independently selected from hydrogen, C1-C6 alkyl, C... as defined herein. 2- C6 alkenyl, C 2- C6 ynyl group, C 3- C7 carbon cyclogroup, C 6- C 10 Aryl, 5-10 heteroaryl and 3-10 heterocyclic groups.
[0085] The "amino" group refers to "-NR". A R B "group, wherein R A and R B Each is independently selected from hydrogen, C1-C6 alkyl, C... as defined herein. 2- C6 alkenyl, C 2- C6 ynyl group, C 3- C7 carbon cyclogroup, C 6- C 10 Aryl, 5-10 membered heteroaryl, and 3-10 membered heterocyclic groups. Non-limiting examples include free amino groups (i.e., -NH2).
[0086] An "aminoalkyl" group refers to an amino group linked via an alkylene group.
[0087] The "alkoxyalkyl" group refers to an alkoxy group linked via an alkylene group, such as "C". 2- "C8 alkoxyalkyl", etc.
[0088] As used herein, a substituent group is derived from an unsubstituted parent group in which one or more hydrogen atoms has been exchanged for another atom or group. Unless otherwise indicated, when a group is deemed to be “substituted” it is meant that the group is substituted with one or more substituents independently selected from C1-C6alkyl, C1-C6alkenyl, C1-C6alkynyl, C1-C6heteroalkyl, C3-C7carbocyclyl (optionally substituted with halogen, C1-C6alkyl, C1-C6alkoxy, C1-C6haloalkyl, and C1-C6haloalkoxy), C3-C7-carbocyclyl-C1-C6-alkyl (optionally substituted with halogen, C1-C6alkyl, C1-C6alkoxy, C1-C6haloalkyl, and C1-C6haloalkoxy), 3-10 membered heterocyclyl (optionally substituted with halogen, C1-C6alkyl, C1-C6alkoxy, C1-C6haloalkyl, and C1-C6haloalkoxy), 3-10 membered heterocyclyl-C1-C6-alkyl (optionally substituted with halogen, C1-C6alkyl, C1-C6alkoxy, C1-C6haloalkyl, and C1-C6haloalkoxy), aryl (optionally substituted with halogen, C1-C6alkyl, C1-C6alkoxy, C1-C6haloalkyl, and C1-C6haloalkoxy), aryl(C1-C6)alkyl (optionally substituted with halogen, C1-C6alkyl, C1-C6alkoxy, C1-C6haloalkyl, and C1-C6haloalkoxy), 5-10 membered heteroaryl (optionally substituted with halogen, C1-C6alkyl, C1-C6alkoxy, C1-C6haloalkyl, and C1-C6haloalkoxy), 5-10 membered heteroaryl(C1-C6)alkyl (optionally substituted with halogen, C1-C6alkyl, C1-C6alkoxy, C1-C6haloalkyl, and C1-C6haloalkoxy), halogen, -CN, hydroxyl, C1-C6alkoxy, C1-C6alkoxy(C1-C6)alkyl (i.e., an ether), aryloxy, sulfhydryl and mercapto, halo(C1-C6)alkyl (e.g., -CF3), halo(C1-C6)alkoxy (e.g., -OCF3), (C1-C6)alkylsulfanyl, arylsulfanyl, amino, amino(C1-C6)alkyl, nitro, O-carbamoyl, N-carbamoyl, O-thiocarbamoyl, N-thiocarbamoyl, C-amido, N-amido, S-sulfonamido, N-sulfonamido, C-carboxy, O-carboxy, acyl, cyanato, isocyanato, thiocyanato, isothiocyanato, sulfinyl, sulfonyl, -SO3H, sulfoxyl, -OSO2C1-4alkyl, and oxo (=O). Where a group is described as “optionally substituted,” the group can be substituted with the above recited substituents.
[0089] As used herein, the term “hydroxyl” refers to an -OH group.
[0090] As used herein, the term "cyano" refers to a "-CN" group.
[0091] As used herein, the term "azido" refers to a -N3 group.
[0092] As used herein, "nucleotides" include a nitrogen-containing heterocyclic base, a sugar, and one or more phosphate groups. They are the monomeric units of nucleic acid sequences. In RNA, the sugar is ribose, while in DNA it is deoxyribose, i.e., a sugar lacking the hydroxyl group present in ribose. The nitrogen-containing heterocyclic base can be a purine or pyrimidine base. Purine bases include adenine (A) and guanine (G), and modified derivatives or analogs thereof. Pyrimidine bases include cytosine (C), thymine (T) and uracil (U), and modified derivatives or analogs thereof. The C-1 atom of the deoxyribose sugar is bonded to the N-1 of a pyrimidine or the N-9 of a purine.
[0093] As used herein, "nucleosides" are structurally similar to nucleotides, but lack a phosphate moiety. One example of a nucleoside analog is one in which a label is attached to the base and there is no phosphate group attached to the sugar molecule. The term "nucleoside" is used herein in its general sense as understood by those skilled in the art. Examples include, but are not limited to, ribonucleosides comprising a ribose moiety and deoxyribonucleosides comprising a deoxyribose moiety. A modified pentose moiety is one in which an oxygen atom has been replaced by a carbon and / or a carbon has been replaced by a sulfur or oxygen atom. "Nucleosides" are monomers that can have a substituted base and / or sugar moiety. Additionally, nucleosides can be incorporated into larger DNA and / or RNA polymers and oligomers.
[0094] As understood by those skilled in the art, the term "purine base" is used herein in its ordinary sense and includes its tautomers. Similarly, the term "pyrimidine base" is used herein in its ordinary sense as understood by those skilled in the art and includes its tautomers. A non-limiting list of optionally substituted purine bases includes purine, adenine, guanine, hypoxanthine, xanthine, allopurinol, 7-alkylguanine (e.g., 7-methylguanine), theobromine, caffeine, uric acid, and isoguanine. Examples of pyrimidine bases include, but are not limited to, cytosine, thymine, uracil, 5,6-dihydrouracil, and 5-alkylcytosine (e.g., 5-methylcytosine).
[0095] As used herein, when an oligonucleotide or polynucleotide is described as "comprising" a nucleoside or nucleotide described herein, it is meant that the nucleoside or nucleotide described herein forms a covalent bond with the oligonucleotide or polynucleotide. Similarly, when a nucleoside or nucleotide is described as part of an oligonucleotide or polynucleotide, for example, "incorporated" into an oligonucleotide or polynucleotide, it is meant that the nucleoside or nucleotide described herein forms a covalent bond with the oligonucleotide or polynucleotide. In some such embodiments, a covalent bond is formed between the 3' hydroxyl group of the oligonucleotide or polynucleotide described herein and the 5' phosphate group of the nucleotide, as a phosphodiester bond between the 3' carbon atom of the oligonucleotide or polynucleotide and the 5' carbon atom of the nucleotide.
[0096] As used herein, "derivative" or "analog" refers to a synthetic nucleotide or nucleoside derivative having a modified base moiety and / or a modified sugar moiety. Such derivatives and analogs are discussed in, for example, Scheit in Nucleotide Analogs (John Wiley & Son, 1980) and Uhlman et al. in Chemical Reviews 90:543-584, 1990. Nucleotide analogs can also include modified phosphodiester linkages, including phosphorothioate, phosphorodithioate, alkyl-phosphonate, phosphoramidate and aminoalkylphosphoramidate linkages. "Derivative," "analog," and "modified" as used herein are used interchangeably and are encompassed by the terms "nucleotide" and "nucleoside" as defined herein.
[0097] As used herein, the term "phosphate" is used in its ordinary sense as understood by those skilled in the art and includes its protonated forms (e.g. and As used herein, the terms "monophosphate," "diphosphate," and "triphosphate" are used in their ordinary sense as understood by those skilled in the art and include protonated forms.
[0098] As used herein, the term "protecting group" refers to any atom or group of atoms added to a molecule to prevent an existing group in the molecule from undergoing an undesirable chemical reaction. Sometimes, "protecting group" and "blocking group" are used interchangeably.
[0099] As used herein, the prefix "photo" or "photo-" means relating to light or electromagnetic radiation. The term can encompass all or part of the electromagnetic spectrum, including but not limited to one or more ranges of the radio, microwave, infrared, visible, ultraviolet, X-ray, or gamma-ray portions commonly referred to as the optical spectrum. The portion of the spectrum can be a portion that is blocked by a metallic region of a surface, such as those described herein. Alternatively or additionally, the portion of the spectrum can be a portion that passes through a gap region of a region surface, such as made of glass, plastic, silicon dioxide, or other materials set forth herein. In particular embodiments, radiation capable of passing through metal can be used. Alternatively or additionally, radiation masked by glass, plastic, silicon dioxide, or other materials set forth herein can be used.
[0100] As used herein, the term "phasing" refers to a phenomenon in SBS that is caused by incomplete removal of 3' terminators and fluorophores and is not completed by the polymerase incorporation of the partial DNA strand within the cluster in a given sequencing cycle. Pre-phasing is caused by incorporation of a nucleotide without an effective 3' terminator, where the incorporation event is advanced by 1 cycle due to a failed termination. Phasing and pre-phasing result in the measured signal intensity within a particular cycle to include the signal of the current cycle as well as noise from previous and subsequent cycles. As the cycle number increases, the portion of the sequence of each cluster affected by phasing and pre-phasing increases, thereby hindering the identification of the correct base. The presence of trace amounts of unprotected or unblocked 3'-OH nucleotides during sequencing by synthesis (SBS) can cause pre-phasing. The unblocked 3'-OH nucleotides can be generated during production or can be generated during storage and reagent handling. Thus, it is surprising to find nucleotide analogs that reduce the incidence of pre-phasing and provide a significant advantage over existing nucleotide analogs in SBS applications. For example, the provided nucleotide analogs can result in faster SBS cycle times, lower phasing and pre-phasing values, and longer sequencing read lengths.
[0101] 3'-hydroxy acetal blocking group
[0102] Some embodiments of the present disclosure relate to a ribose or deoxyribose containing nucleotide or nucleoside having a removable 3'-OH protecting or blocking group that forms a structure covalently linked to the 3'-carbon atom wherein:
[0103] each R 1a and R 1bindependently H, C1-C6alkyl, C1-C6haloalkyl, C1-C6alkoxy, C1-C6haloalkoxy, cyano, halogen, optionally substituted phenyl, or optionally substituted aralkyl;
[0104] each R 2a and R 2b are independently H, C1-C6alkyl, C1-C6haloalkyl, cyano, or halogen;
[0105] or, R 1a and R 2a together with the atom to which they are attached form an optionally substituted five- to eight-membered heterocyclyl;
[0106] R 3 is H, optionally substituted C 2- C6alkyl, optionally substituted C 3- C7cycloalkenyl, optionally substituted C 2- C6alkynyl, or optionally substituted (C1-C6alkylene)Si(R 4 )3; and
[0107] each R 4 is independently H, C1-C6alkyl, or optionally substituted C 6- C 10 aryl; provided that when each R 1a , R 1b , R 2a , and R 2b is H, then R 3 is not H.
[0108] Some further embodiments of the disclosure relate to compounds having the structure of Formula (I): (I), wherein R’ is H, monophosphate, diphosphate, triphosphate, thiophosphate, phosphate analog, -O- linked to an active phosphorus group, or -O- protected with a protecting group; R’’ is H or OH; B is a nucleobase; each R 1a , R 1b , R 2a , R 2b , and R 3 is as defined above. In some further embodiments, B is , , , , , , or . In some further embodiments, the nucleobase is covalently bound to a detectable label (e.g., a fluorescent dye), optionally through a linking group, e.g., B is , , or In some such embodiments, R' is triphosphate. In some such embodiments, R" is H.
[0109] In some embodiments of the acetal blocking group described herein, at least one of R 1a and R 1b is H. In some such embodiments, each of R 1a and R 1b is H. In some other embodiments, at least one of R 1a and R 1b is C1-C6 alkyl, e.g., methyl, ethyl, isopropyl, or tert-butyl. In some embodiments, each of R 2a and R 2b is independently H, halogen, or C1-C6 alkyl. In some such embodiments, at least one of R 2a and R 2b is H or C1-C6 alkyl. In some such embodiments, each of R 2a and R 2b is H. In some such embodiments, each of R 2a and R 2b is C1-C6 alkyl, e.g., methyl, ethyl, isopropyl, or tert-butyl. In one embodiment, each of R 2a and R 2b is methyl. In some such embodiments, each of R 2a and R 2b is independently C1-C6 alkyl or halogen. In some such embodiments, R 2a is H, and R 2b is halogen or C1-C6 alkyl.
[0110] In some embodiments of the acetal blocking group described herein, R 3 is optionally substituted C 2- C6 alkenyl. In some such embodiments, R 3 is C 2- C6 alkenyl (e.g., ethenyl, propenyl) optionally substituted with one or more substituents independently selected from the group consisting of halogen, C1-C6 alkyl, C1-C6 haloalkyl, and combinations thereof. In some further embodiments, R 3 is , , , or In some other embodiments, R 3 is optionally substituted C2- C6alkynyl. In some such embodiments, R 3 is C 2- C6alkynyl (e.g., ethynyl, propynyl) optionally substituted with one or more substituents independently selected from the group consisting of halogen, C1-C6alkyl, C1-C6haloalkyl, and combinations thereof. In one embodiment, R 3 is optionally substituted ethynyl ). In some other embodiments, R 3 is optionally substituted (C1-C6alkylene)Si(R 4 )3. In some such embodiments, at least one R 4 is C 1-4 alkyl. In some further embodiments, each R 4 is C1-C4alkyl, e.g., methyl, ethyl, isopropyl, or tert-butyl. In one embodiment, R 3 is -(CH2)-SiMe3. In some alternative embodiments, R 3 is C1-C6alkyl.
[0111] In some alternative embodiments, R 1a and R 2a together with the atoms to which they are attached form a five- to seven-membered heterocyclyl. In some such embodiments, R 1a and R 2a together with the atoms to which they are attached form a six-membered heterocyclyl. In some such embodiments, the six-membered heterocyclyl has the structure . In some further embodiments, at least one of R 1b , R 2b , and R 3 is H. In some other embodiments, at least one of R 1b , R 2b , and R 3 is C1-C6alkyl. In one embodiment, each of R 1b , R 2b , and R 3 is H.
[0112] In some further embodiments, the compound of formula (I) is also represented by formula (la):
[0113] (Ia), wherein each R 2c and R2dIndependently, it is H, a halogen (e.g., fluorine, chlorine), a C1-C6 alkyl group (e.g., methyl, ethyl, or isopropyl), or a C1-C6 haloalkyl group (e.g., -CHF2, -CH2F, or –CF3). In some such embodiments, R 1a and R 1b One of them is H. In some such implementations, each R 1a and R 1b For H. In some other implementations, at least R 1a and R 1b One of them is a C1-C6 alkyl group, such as methyl, ethyl, isopropyl, or tert-butyl. In some embodiments, each R 2a and R 2b Independently, it is H, halogen, or C1-C6 alkyl. In some such embodiments, each R 2a and R 2b For H. In some such implementations, each R 2c and R 2d Independently, it is H, halogen, or C1-C6 alkyl. In some such embodiments, each R 2c and R 2d It is a C1-C6 alkyl group, such as methyl, ethyl, isopropyl, or tert-butyl. In one embodiment, each R 2c and R 2d For methyl. In some such embodiments, each R 2c and R 2d Independently halogenated. In some such implementations, R 2c For H, and R 2d It is H, a halogen (fluorine, chlorine), or a C1-C6 alkyl group (e.g., methyl, ethyl, isopropyl, or tert-butyl). In a further embodiment, each R 1a and R 1b For H; R 2a For H; R 2b H, halogen, or methyl; R 2c For H; and R 2d It can be H, halogen, methyl, ethyl, isopropyl or tert-butyl.
[0114] Non-limiting embodiments of the blocking group described herein include those having a structure selected from:
[0115] , , , , , and covalently attached to the 3'-carbon of a ribose or deoxyribose sugar.
[0116] 3'-hydroxy thioformate blocking group
[0117] Some additional embodiments of the present disclosure relate to a ribose or deoxyribose containing nucleoside or nucleotide having a removable 3'-OH blocking group that forms a structure covalently attached to the 3'-carbon atom wherein:
[0118] each R 5 and R 6 is independently H, C1-C6 alkyl, C 2- C6 alkenyl, C 2- C6 alkynyl, C1-C6 haloalkyl, C 2- C8 alkoxyalkyl, optionally substituted -(CH2) m -phenyl, optionally substituted -(CH2) n -(5 or 6 membered heteroaryl), optionally substituted -(CH2) k -C 3- C7 carbocyclyl, or optionally substituted -(CH2) p -(3 to 7 membered heterocyclyl);
[0119] or, R 5 and R 6 together with the atoms to which they are attached form an optionally substituted five to seven membered heterocyclyl;
[0120] each -(CH2) m -, -(CH2) n -, -(CH2) k -, and -(CH2) p - is optionally substituted; and
[0121] each m, n, k, and p is independently 0, 1, 2, 3, or 4.
[0122] Some additional embodiments relate to a compound of Formula (II):
[0123] (II), wherein R' is H, a monophosphate, diphosphate, triphosphate, thiophosphate, phosphate analog, -O- linked to an active phosphorus group, or -O- protected with a protecting group; R" is H or OH; B is a nucleobase; each R 5 and R 6 is as defined above. In some further embodiments, B is , , In some further embodiments, the nucleoside base is covalently bound to a detectable label (e.g., a fluorescent dye), optionally via a linking group, such as B. , or In some such embodiments, R' is a triphosphate. In some such embodiments, R'' is H.
[0124] In some embodiments of the thiocarbamate blocking group described herein, at least R 5 and R 6 One of them is H. In some such implementations, each R 5 and R 6 For H. In some such implementations, R 5 For H and R 6 It is a C1-C6 alkyl group, such as methyl, ethyl, isopropyl, or tert-butyl. In some such embodiments, R 5 For H and R 6 C 2- C6 alkenyl (e.g., vinyl or allyl) or C 2- C6 ynyl group (e.g., ethynyl or propynyl). In some such embodiments, R 5 For H and R 2 –(CH2) is an optional substitution. m –Phenyl, optionally substituted –(CH2) n –(5 or 6-membered heteroaryl), optionally substituted –(CH2) k –C 3- C7 carbocyclic group or optionally substituted –(CH2) p –(3 to 7-membered heterocyclic groups). In some further embodiments, the C 3- The C7 carbon cyclogroup can be C 3- C7 cycloalkyl or C 3- C7 cycloalkenyl. The 3- to 7-membered heterocyclic group may contain zero or one double bond in the ring structure. In a further embodiment, R... 5 For H and R 6 –(CH2) is an optional substitution. m –Phenyl, optionally substituted –(CH2) n –6-membered heteroaryl, optionally substituted –(CH2) k –C5 or C6 carbocyclic group or optionally substituted –(CH2) p –(5 or 6-membered heterocyclic group). In some embodiments, m, n, k, or p is 0. In other embodiments, m, n, k, or p is 1 or 2. In some other embodiments, at least R 5 and R 6one of R and R is C1-C6alkyl, for example, methyl, ethyl, isopropyl, or tert-butyl. In some further embodiments, R 5 and R 6 are each C1-C6alkyl. In one embodiment, R 5 and R 6 are each methyl.
[0125] In some alternative embodiments, R 5 and R 6 together with the atom to which they are attached form an optionally substituted five- to seven-membered heterocyclyl group. In some such embodiments, R 5 and R 6 together with the atom to which they are attached form an optionally substituted piperidinyl group.
[0126] Non-limiting embodiments of 3’-O-thioformacetal blocking groups described herein include those having a structure selected from the group consisting of: , and (DMTC) covalently attached to the 3’-carbon of a ribose or deoxyribose sugar.
[0127] Further embodiments of the disclosure relate to oligonucleotides or polynucleotides comprising a nucleoside or nucleotide described herein.
[0128] In any embodiment of a blocking group described herein, when a group is described as “optionally substituted,” it can be unsubstituted or substituted.
[0129] In any embodiment of a nucleotide or nucleoside described herein having a 3’-hydroxyl blocking group, the nucleoside or nucleotide can be covalently attached to a detectable label (e.g., a fluorophore), optionally via a linking group. The linking group can be cleavable or non-cleavable. In some such embodiments, the detectable label (e.g., a fluorophore) is covalently attached to the nucleoside base of the nucleoside or nucleotide via a cleavable linking group. In some other embodiments, the detectable label (e.g., a fluorophore) is covalently attached to the 3’ oxygen of the nucleoside or nucleotide via a cleavable linking group. In some further embodiments, such a cleavable linking group can comprise an azido moiety or a disulfide moiety, an acetal moiety or a thioformacetal moiety. In some embodiments, the 3’ hydroxyl blocking group and the cleavable linking group (and attached label) can be removed under the same or substantially the same chemical reaction conditions, e.g., the blocking group and the detectable label can be removed in a single chemical reaction. In other embodiments, the blocking group and the detectable label are removed in two separate steps.
[0130] In some embodiments, the nucleotides or nucleosides described herein comprise a 2’ deoxyribose. In other aspects, the 2’ deoxyribose contains one, two, or three phosphate groups at the 5’ position of the sugar ring. In another aspect, the nucleotides described herein are nucleotide triphosphates.
[0131] In some embodiments, the 3’ blocked nucleotides or nucleosides described herein provide superior stability in solution during storage or in reagent handling during sequencing applications compared to the same nucleotides or nucleosides protected with standard 3’-OH blocking groups disclosed in the prior art (e.g., 3’-O-azidomethyl protecting groups). For example, the acetal or thiocarbamate blocking groups disclosed herein can confer at least 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 300%, 400%, 500%, 600%, 700%, 800%, 900%, 1000%, 1500%, 2000%, 2500%, or 3000% improved stability compared to azidomethyl protected 3’-OH over the same period of time under the same conditions, resulting in reduced predetermined phase values and longer sequencing read lengths. In some embodiments, the stability is measured at ambient temperature or at a temperature below ambient temperature (e.g., 4-10 °C). In other embodiments, the stability is measured at an elevated temperature, e.g., 40 °C, 45 °C, 50 °C, 55 °C, 60 °C, or 65 °C. In some such embodiments, the stability is measured in solution in an alkaline pH environment, e.g., at pH 9.0, 9.2, 9.4, 9.6, 9.8, or 10.0. In some such embodiments, the stability is measured in the presence or absence of enzymes such as polymerases (e.g., DNA polymerases), terminal deoxynucleotidyl transferases, or reverse transcriptases.
[0132] In some embodiments, the 3’ blocked nucleotides or nucleosides described herein provide superior deblocking rates in solution during the chemical cleavage step of sequencing applications compared to the same nucleotides or nucleosides protected with standard 3’-OH blocking groups (e.g., 3’-O-azidomethyl protecting groups) disclosed in the prior art. For example, the acetal or thiocarbamate blocking groups disclosed herein can confer at least 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 150%, 200%, 300%, 400%, 500%, 600%, 700%, 800%, 900%, 1000%, 1500%, or 2000% improvement in deblocking rates compared to azidomethyl protected 3’-OH using standard deblocking reagents (e.g., tris(hydroxypropyl)phosphine), thereby reducing the total time of sequencing cycles. In some embodiments, the deblocking rates are measured at ambient temperature or at temperatures below ambient temperature (e.g., 4-10 °C). In other embodiments, the deblocking rates are measured at elevated temperatures, e.g., 40 °C, 45 °C, 50 °C, 55 °C, 60 °C, or 65 °C. In some such embodiments, the deblocking rates are measured in solution at basic pH environments, e.g., at pH 9.0, 9.2, 9.4, 9.6, 9.8, or 10.0. In some such embodiments, the molar ratio of deblocking reagent to substrate (i.e., 3’ blocked nucleoside or nucleotide) is about 10:1, about 5:1, about 2:1, or about 1:1.
[0133] In some embodiments, palladium deblocking reagents (e.g., Pd(0) are used to remove 3’ acetal blocking groups (e.g., AOM blocking groups). Pd can form a chelating complex with both oxygen atoms of the AOM group as well as the double bond of the allyl group, such that the deblocking reagent is directly adjacent to the functionality to be removed and can result in accelerated deblocking rates.
[0134] Deprotection of 3'-OH blocking group
[0135] The 3’-acetal blocking groups described herein can be removed or cleaved under a variety of chemical conditions. For acetal blocking groups containing a vinyl or alkenyl moiety Non-limiting cleavage conditions include Pd(II) complexes, such as Pd(OAc)2or allyl palladium(II) chloride dimer, in the presence of a phosphine ligand, such as tris(hydroxymethyl)phosphine (THMP) or tris(hydroxypropyl)phosphine (THP or THPP). For those blocking groups containing an alkynyl group (e.g., ethynyl), they can also be removed by Pd(II) complexes (e.g., Pd(OAc)2or allyl palladium(II) chloride dimer) in the presence of a phosphine ligand (e.g., THP or THMP).
[0136] Palladium cleavage reagent
[0137] In some embodiments, the acetal blocking group described herein can be cleaved by a palladium catalyst. In some such embodiments, the palladium catalyst is water soluble. In some such embodiments, it is a Pd(0) complex (e.g., (3,3',3"-phosphinyltris(phenylsulfonato)palladium(0) nonasodium salt nonahydrate). In some cases, the Pd(0) complex can be generated in situ by reduction of a Pd(II) complex with a reagent such as an olefin, alcohol, amine, phosphine, or metal hydride. Suitable palladium sources include Na2PdCl4, Pd(CH3CN)2Cl2, (PdCl(C3H5))2, [Pd(C3H5)(THP)]Cl, [Pd(C3H5)(THP)2]Cl, Pd(OAc)2, Pd(Ph3)4, Pd(dba)2, Pd(Acac)2, PdCl2(COD), and Pd(TFA)2. In one such embodiment, the Pd(0) complex is generated in situ from Na2PdCl4. In another embodiment, the palladium source is allyl palladium(II) chloride dimer [(PdCl(C3H5))2]. In some embodiments, the Pd(0) complex is generated in situ by mixing a Pd(II) complex with a phosphine in aqueous solution. Suitable phosphines include water-soluble phosphines such as tris(hydroxypropyl)phosphine (THP), tris(hydroxymethyl)phosphine (THMP), 1,3,5-triaza-7-phosphaadamantane (PTA), bis(p-sulfonatophenyl)phenylphosphine dihydrate potassium salt, tris(carboxyethyl)phosphine (TCEP), and triphenylphosphine-3,3',3"-trisulfonic acid trisodium salt.
[0138] In some embodiments, the Pd(0) is prepared in situ by mixing a Pd(II) complex [(PdCl(C3H5))2] with THP. The molar ratio of Pd(II) complex to THP is about 1 :2, 1 :3, 1 :4, 1 :5, 1 :6, 1 :7, 1 :8, 1 :9, or 1 : 10. In some further embodiments, one or more reducing agents, such as ascorbic acid or a salt thereof (e.g., sodium ascorbate), can be added. In some embodiments, the cleavage mixture can comprise an additional buffering agent, such as a primary amine, a secondary amine, a tertiary amine, a carbonate, a phosphate, or a borate, or a combination thereof. In some further embodiments, the buffering agent comprises ethanolamine (EA), tris(hydroxymethyl)aminomethane (Tris), glycine, sodium carbonate, sodium phosphate, sodium borate, 2-dimethylaminoethanol (DMEA), 2-diethylaminoethanol (DEEA), N,N,N',N'-tetramethylethylenediamine (TEMED), or N,N,N',N'-tetraethylethylenediamine (TEEDA), or a combination thereof. In one embodiment, the buffering agent is DEEA. In another embodiment, the buffering agent comprises one or more inorganic salts, such as a carbonate, a phosphate, or a borate, or a combination thereof. In one embodiment, the inorganic salt is a sodium salt.
[0139] Alternatively, alkynyl moieties containing a blocking group can also be cleaved in the presence of (NH4)2MoS4. Non-limiting cleavage conditions for alkynyl moieties include Cu(II) complex with THPTA ligand (tris(3-hydroxypropyltriazolylmethyl)amine) and ascorbic acid. Non-limiting cleavage conditions for blocking groups containing a six-membered heterocycle (e.g., tetrahydropyran) include cyclodextrin or Ln(OTf)3(lanthanum triflate). Non-limiting cleavage conditions for blocking groups containing an alkyl silyl group (e.g., -CH2SiMe3) include LiBF4(lithium tetrafluoroborate). Other acetal blocking groups can be removed by LiBF4or Bi(OTf)3(bismuth triflate), such as -O(CH2)0-C1-C6alkyl. Non-limiting exemplary conditions for cleaving the various blocking groups described are illustrated in Scheme 1 below.
[0140] Scheme 1. Illustration of 3'-deblocking conditions
[0141]
[0142]
[0143] The 3'-0-thioformyl blocking group described herein can be removed or cleaved under various chemical conditions. Non-limiting exemplary conditions for cleaving the thioformyl blocking group described herein include NaIO4and Oxone® (potassium peroxomonosulfate).
[0144] Additionally, the azido group in -CH2N3 can be converted to an amino group by contacting the molecule with a phosphine. Alternatively, the azido group in -CH2N3 can be converted to an amino group by contacting the molecule with a thiol, particularly a water-soluble thiol such as dithiothreitol (DTT). In one embodiment, the phosphine is THP.
[0145] Compatibility with linearization
[0146] To maximize the throughput of nucleic acid sequencing reactions, it is advantageous to be able to sequence multiple template molecules in parallel. Parallel processing of multiple templates can be achieved using nucleic acid array technology. These arrays are typically composed of a high density matrix of polynucleotides immobilized on a solid support.
[0147] WO 98 / 44151 and WO 00 / 18957 both describe nucleic acid amplification methods that allow the amplification products to be immobilized on a solid support to form an array composed of clusters or “colonies” formed from multiple identical immobilized polynucleotide strands and multiple identical immobilized complementary strands. This type of array is referred to herein as a “clustered array”. The nucleic acid molecules present in the DNA colonies on a clustered array prepared according to these methods can provide templates for sequencing reactions, for example as described in WO 98 / 44152. The products of the solid phase amplification reactions such as described in WO 98 / 44151 and WO 00 / 18957 are so-called “bridged” structures formed by annealing of an immobilized polynucleotide strand and an immobilized complementary strand, both strands being attached at the 5’ end to a solid support. To provide a more suitable template for nucleic acid sequencing, it is preferred to remove substantially all or at least a portion of one of the immobilized strands in the “bridged” structure to produce a template that is at least partially single stranded. Thus, the portion of the template that is single stranded can be used to hybridize to a sequencing primer. The process of removing all or a portion of one of the immobilized strands in a “bridged” double stranded nucleic acid structure is referred to as “linearization”. Linearization can be by a variety of means including, but not limited to, enzymatic cleavage, photochemical cleavage, or chemical cleavage. Non-limiting examples of linearization methods are disclosed in PCT Publication WO 2007 / 010251, U.S. Patent Publication 2009 / 0088327, U.S. Patent Publication 2009 / 0118128, and U.S. Application 62 / 671,816, the entire contents of which are incorporated by reference.
[0148] In some embodiments, the conditions for deprotection or removal of the 3'-OH blocking group are also compatible with the linearization process. In some further embodiments, the deprotection conditions are compatible with a chemical linearization process, which includes the use of a Pd complex and a phosphine, such as Pd(OAc)2and THP. In some embodiments, the Pd complex is a Pd(II) complex, which generates Pd(0) in situ in the presence of a phosphine.
[0149] Unless otherwise indicated, reference to a nucleotide is also intended to apply to a nucleoside.
[0150] Labeled nucleotides
[0151] According to one aspect of the disclosure, the described 3'-OH blocked nucleotides further comprise a detectable label, and such nucleotides are referred to as labeled nucleotides. The label (e.g., a fluorescent dye) can be bound through an optional linking group by a variety of methods, including hydrophobic attraction, ionic attraction, and covalent linkage. In some aspects, the dye is bound to the substrate through covalent linkage. More particularly, the covalent linkage is by means of a linking group. In some cases, such labeled nucleotides are also referred to as "modified nucleotides."
[0152] Labeled nucleosides and nucleotides can be used to label polynucleotides formed by enzymatic synthesis, such as, by way of non-limiting example, PCR amplification, isothermal amplification, solid phase amplification, polynucleotide sequencing (e.g., solid phase sequencing), nick translation reactions.
[0153] In some embodiments, the dye can be covalently linked to an oligonucleotide or nucleotide through the nucleobase. For example, the labeled nucleotide or oligonucleotide can have a label attached to the C5 position of a pyrimidine base or the C7 position of a 7-deazapurine base through a linking group moiety.
[0154] Unless otherwise indicated, reference to a nucleotide is also intended to apply to a nucleoside. The present application will be further described with reference to the mention of DNA, unless otherwise indicated, although the description also applies to RNA, PNA, and other nucleic acids.
[0155] Linking group
[0156] In some embodiments described herein, the purine or pyrimidine base of the nucleotides or nucleoside molecules described herein can be linked to a detectable label as described above. In some such embodiments, the linking group used is cleavable. The use of a cleavable linking group can ensure that, after detection of the label, it is removed if desired, thereby avoiding any interference with any subsequently incorporated labelled nucleotides or nucleosides. In some embodiments, the cleavable linking group comprises an azido moiety, -O-C 2- a C6alkenyl moiety (e.g., -O-allyl), a disulfide moiety, an acetal moiety (the same as or similar to the 3' acetal blocking group described herein), or a thiocarbamate moiety (the same as or similar to the 3' acetal blocking group described herein).
[0157] In some other embodiments, the linking group used is not cleavable. Since there is no need for any subsequent incorporation of nucleotides in each case where a labelled nucleotide of the application is incorporated, there is no need to remove the label from the nucleotide.
[0158] Cleavable linking groups are known in the art and can be attached to nucleoside bases and labels using routine chemistry. The linking group can be cleaved by any suitable method, including exposure to acid, base, nucleophiles, electrophiles, radicals, metals, reducing or oxidizing agents, light, temperature, enzymes, and the like. The linking groups described herein can also be cleaved with the same catalysts that cleave the bond of the 3'-0-blocking group. Suitable linking groups can be adapted to standard chemical protecting groups, as disclosed in Greene & Wuts, Protective Groups in Organic Synthesis, John Wiley & Sons. Other suitable cleavable linking groups for use in solid phase synthesis are disclosed in Guillier et al. (Chem. Rev. 100:2092-2157, 2000).
[0159] Where the detectable label is linked to the base, the linking group can be attached to any position of the nucleoside base, provided that Watson-Crick base pairing can still occur. In the case of purine bases, the linking group is preferably attached via the 7-position of the purine or preferably a deazapurine analogue, via an 8-modified purine, via an N-6 modified adenosine or an N-2 modified guanine. For pyrimidines, attachment is preferably via the 5-position of cytosine, thymine or uracil and the N-4 position of cytosine.
[0160] In some embodiments, the linking group can comprise a spacer unit. The length of the linking group is not critical, so long as the label is held at a sufficient distance from the nucleotide to not interfere with any interactions between the nucleotide and an enzyme, such as a polymerase.
[0161] In some embodiments, the linking group can consist of a functional group similar to the 3'-OH protecting group. This will make the deprotection and deblocking process more efficient, as only one treatment is required to remove both the label and the protecting group.
[0162] The use of the term "cleavable linking group" does not imply that removal of the entire linking group is required. The cleavage site can be located at a position on the linking group that ensures that a portion of the linking group remains attached to the dye and / or substrate moiety after cleavage. As non-limiting examples, the cleavable linking group can be an electrophilic cleavable linking group, a nucleophilic cleavable linking group, a photocleavable linking group, cleavable under reducing conditions, oxidizing conditions (e.g., disulfide-containing or azido-containing linking groups), cleavable via the use of a safety-catch linking group, and cleavable by an elimination mechanism. The use of a cleavable linking group to attach a dye compound to a substrate moiety can ensure that the label is removed after detection, if desired, thereby avoiding any interfering signals in downstream steps.
[0163] Useful linking groups can be found in PCT Publication WO 2004 / 018493 (incorporated herein by reference), examples of which include linking groups that can be cleaved using a water-soluble phosphine or a water-soluble transition metal catalyst formed from a transition metal and at least partially water-soluble ligand (e.g., Pd(II) complexes and THP). The latter forms at least partially water-soluble transition metal complexes in aqueous solution. Such cleavable linking groups can be used to attach the base of a nucleotide to a label, such as a dye described herein.
[0164] Particular linking groups include those disclosed in PCT Publication WO 2004 / 018493 (incorporated herein by reference), such as those comprising a moiety of the formula:
[0165]
[0166] (wherein X is selected from O, S, NH, and NQ, wherein Q is C1- 10 substituted or unsubstituted alkyl group, Y is selected from O, S, NH, and N(allyl), T is hydrogen or C1-C 10 substituted or unsubstituted alkyl group, denotes the location of attachment of the moiety to the remainder of the nucleotide or nucleoside). In some aspects, the linking group attaches the base of the nucleotide to a label, such as a dye compound described herein.
[0167] Other examples of linking groups include those disclosed in U.S. Publication 2016 / 0040225 (incorporated herein by reference), such as those comprising a moiety of the formula:
[0168]
[0169]
[0170] The linking group moieties shown herein can comprise all or part of the linking group structure between the nucleotide / nucleoside and the label.
[0171] Further examples of linking groups ("L") include moieties of the formula:
[0172] or where B is a nucleoside base; Z is -N3 (azido), -0-C1-C6alkyl, -0-C 2- C6alkenyl, or -0-C 2- C6alkynyl; and Fl comprises a fluorescent label, which can comprise additional linking group structure. Those of ordinary skill in the art understand that the label is covalently bonded to the linking group through reaction of a functional group of the label (e.g., a carboxyl group) with a functional group of the linking group (e.g., an amino group).
[0173] In particular embodiments, the length of the linking group between the fluorescent dye (fluorophore) and the guanine base can be varied, e.g., by the insertion of a polyethylene glycol spacer, such that the fluorescence intensity is higher compared to the same fluorophore linked to the guanine base through other linkage known in the art. Exemplary linking groups and their properties are set forth in PCT Publication WO2007020457 (incorporated herein by reference). The design of the linking group, particularly its increased length, can improve the brightness of the fluorophore attached to the guanine base of a guanosine nucleotide when incorporated into a polynucleotide (e.g., DNA). Thus, when the dye is used in any analytical method that requires detection of the fluorescent dye label attached to a guanine-containing nucleotide, it is advantageous for the linking group to comprise a spacer group of the formula -((CH2)2O) n - (where n is an integer between 2 and 50) as described in WO 2007 / 020457.
[0174] Nucleosides and nucleotides can be labeled at the site of the sugar or the nucleobase. As known in the art, a "nucleotide" consists of a nitrogenous base, a sugar, and one or more phosphate groups. In RNA the sugar is ribose, and in DNA it is deoxyribose (i.e., a sugar lacking the hydroxyl group present in ribose). The nitrogenous base is a derivative of a purine or a pyrimidine. The purines are adenine (A) and guanine (G), and the pyrimidines are cytosine (C) and thymine (T) or, in the case of RNA, uracil (U). The C-1 atom of deoxyribose is bonded to the N-1 of a pyrimidine or the N-9 of a purine. Nucleotides are also phosphates of nucleosides, which are esterified at the C-3 or C-5 attached hydroxyl group of the sugar. Nucleotides are usually mono-, di-, or tri-phosphates.
[0175] A "nucleoside" is structurally similar to a nucleotide, but lacks the phosphate moiety. An example of a nucleoside analog is one in which the label is attached to the base without a phosphate group on the sugar molecule.
[0176] While the bases are generally referred to as purines or pyrimidines, one skilled in the art will appreciate that derivatives and analogs are possible that do not change the ability of the nucleotide or nucleoside to Watson-Crick base pair. A "derivative" or "analog" refers to a compound or molecule whose core structure is the same as or very similar to the parent compound, but which has a chemical or physical modification (e.g., a different or additional side group that links the derivative nucleotide or nucleoside to another molecule). For example, the base can be a deazapurine. In particular embodiments, the derivative should be able to Watson-Crick pair. "Derivatives" and "analogues" also include, for example, synthetic nucleotide or nucleoside derivatives having modified base moieties and / or modified sugar moieties. Such derivatives and analogs are discussed in, for example, Scheit, Nucleotide analogs (John Wiley & Son, 1980) and Uhlman et al., Chemical Reviews 90:543-584, 1990. Nucleotide analogs can also include modified phosphodiester linkages, including phosphorothioate, phosphorodithioate, alkyl-phosphonate, phosphoramidate linkages, and the like.
[0177] The dye can be attached (e.g., via a linking group) to any position of the nucleotide base. In particular embodiments, Watson-Crick base pairing can still be performed on the resulting analog. Particular nucleoside base labeling sites include the C5 position of a pyrimidine base or the C7 position of a 7-deazapurine base. As noted above, the dye can be covalently attached to the nucleoside or nucleotide using a linking group.
[0178] In particular embodiments, the labeled nucleoside or nucleotide can be enzyme incorporable and enzyme extendable. Thus, the linking group moiety can have sufficient length to link the nucleotide to a compound such that the compound does not significantly interfere with the overall binding and recognition of the nucleotide by nucleic acid replicase enzymes. Thus, the linking group can also comprise a spacer unit. The spacer unit separates, for example, the nucleotide base from the cleavage site or label.
[0179] The nucleoside or nucleotide labeled with a dye described herein can have the following formula:
[0180]
[0181] wherein the dye is a dye compound; B is a nucleoside base, such as uracil, thymine, cytosine, adenine, guanine, etc.; L is an optional linking group, which can or can not be present; R’ can be H, monophosphate, diphosphate, triphosphate, phosphorothioate, phosphate analog, linked to an active phosphorus group -O- or -O- protected by a blocking group; R’’’ can be H, OH, phosphoramidite, or a 3’-OH blocking group described herein, and R’’ is H or OH. When R’’’ is a phosphoramidite, R’ is an acid cleavable hydroxyl protecting group that allows for subsequent monomer coupling under automated synthesis conditions.
[0182] In particular embodiments, both the linking group (between the dye and the nucleotide) and the blocking group are present and are separate moieties. In particular embodiments, both the linking group and the blocking group can be cleavable under substantially similar conditions. Thus, the de-blocking process can be more efficient, as only one treatment is required to remove both the dye compound and the blocking group. However, in some embodiments, the linking group and the blocking group need not be cleavable under similar conditions, but are individually cleavable under different conditions.
[0183] The present disclosure also includes polynucleotides incorporating the dye compound. Such polynucleotides can be DNA or RNA consisting of deoxyribonucleotides or ribonucleotides, respectively, linked by phosphodiester linkages. The polynucleotide can comprise naturally occurring nucleotides in combination with at least one modified nucleotide described herein (e.g., labeled with a dye compound), non-naturally occurring (or modified) nucleotides other than the labeled nucleotides described herein, or any combination thereof. Polynucleotides according to the present disclosure can also include non-natural backbone linkages and / or non-nucleotide chemical modifications. Chimeric structures consisting of a mixture of ribonucleotides and deoxyribonucleotides comprising at least one labeled nucleotide are also contemplated.
[0184] Non-limiting exemplary labeled nucleotides described herein include:
[0185]
[0186]
[0187]
[0188]
[0189]
[0190] wherein L represents a linking group, and R represents a sugar residue as described above, or a sugar residue substituted at the 5' position with one, two, or three phosphates.
[0191] In some embodiments, non-limiting exemplary fluorescent dye conjugates are shown below:
[0192]
[0193]
[0194] wherein PG represents a 3'-hydroxyl blocking group described herein. In any embodiment of a labeled nucleotide described herein, the nucleotide is a nucleoside triphosphate.
[0195] Kit
[0196] The present disclosure also provides kits comprising one or more 3' blocked nucleosides and / or nucleotides described herein, e.g., a 3' blocked nucleotide of Formula (I), (la), or (II). Such kits typically include at least one 3' blocked nucleotide or nucleoside labeled with a dye and at least one other component. The other component can be one or more components determined in the methods described herein or in the Example section below. Some non-limiting examples of components that can be combined into a kit of the present disclosure are set forth below.
[0197] In a particular embodiment, a kit can include at least one labeled 3' blocked nucleotide or nucleoside and labeled or unlabeled nucleotides or nucleosides. For example, a nucleotide labeled with a dye can be provided in combination with unlabeled or natural nucleotides and / or with fluorescently labeled nucleotides or any combination thereof. The combination of nucleotides can be provided as separate individual components (e.g., one nucleotide type per vessel or tube) or as a mixture of nucleotides (e.g., two or more nucleotides mixed in the same vessel or tube).
[0198] When the kit comprises a plurality, in particular two or three, or more particularly four, 3' blocked nucleotides labeled with a dye compound, the different nucleotides can be labeled with different dye compounds, or one nucleotide can be dark, which does not have a dye compound. When the different nucleotides are labeled with different dye compounds, the kit is characterized in that the dye compounds are spectrally distinguishable fluorescent dyes. As used herein, the term "spectrally distinguishable fluorescent dyes" refers to fluorescent dyes that, when two or more such dyes are present in one sample, emit fluorescent energy at wavelengths that can be distinguished by a fluorescence detection device (e.g., a commercial capillary-based DNA sequencing platform). When two nucleotides labeled with fluorescent dye compounds are provided in kit form, some embodiments are characterized in that the spectrally distinguishable fluorescent dyes can be excited at the same wavelength, e.g., by the same laser. When four 3' blocked nucleotides (A, C, T, and G) labeled with fluorescent dye compounds are provided in kit form, some embodiments are characterized in that two spectrally distinguishable fluorescent dyes can both be excited at one wavelength, and the other two spectrally distinguishable dyes can both be excited at another wavelength. The particular excitation wavelengths are 488 nm and 532 nm.
[0199] In one embodiment, the kit comprises a first 3' blocked nucleotide labeled with a first dye and a second nucleotide labeled with a second dye, wherein the maximum absorbance difference of the dyes is at least 10 nm, in particular 20 nm to 50 nm. More particularly, the Stokes shift of the two dye compounds is between 15-40 nm, wherein "Stokes shift" is the distance between the peak absorbance wavelength and the peak emission wavelength.
[0200] In an alternative embodiment, the kits of the present disclosure can comprise 3’ blocked nucleotides, wherein the same base is labeled with two or more different dyes. A first nucleotide (e.g., a 3’ blocked T nucleoside triphosphate or a 3’ blocked G nucleoside triphosphate) can be labeled with a first dye. A second nucleotide (e.g., a 3’ blocked C nucleoside triphosphate) can be labeled with a second dye that is spectrally distinct from the first dye, e.g., an “green” dye having an absorbance less than 600 nm, and a “blue” dye having an absorbance less than 500 nm, e.g., 400 nm to 500 nm, in particular 450 nm to 460 nm. A third nucleotide (e.g., a 3’ blocked A nucleoside triphosphate) can be labeled as a mixture of the first and second dyes, or a mixture of the first, second, and third dyes, and a fourth nucleotide (e.g., a 3’ blocked G nucleoside triphosphate or a 3’ blocked T nucleoside triphosphate) can be “dark” and unlabeled. In one example, the nucleotides 1-4 can be labeled as “blue,” “green,” “blue / green,” and dark. To further simplify the instrument, the four nucleotides can be labeled with two dyes that are excited under a single laser, so the labels for nucleotides 1-4 can be “blue 1,” “blue 2,” “blue 1 / blue 2,” and dark.
[0201] In particular embodiments, the kits can comprise four labeled 3’ blocked nucleotides (e.g., A, C, T, G), wherein each type of nucleotide comprises the same 3’ blocking group and fluorescent label, and wherein each fluorescent label has a different fluorescence maximum and each fluorescent label is distinguishable from the other three labels. The kits can have two or more fluorescent labels with similar maximum absorbance but different Stokes shifts. In some other embodiments, one of the nucleotides is unlabeled.
[0202] While the kits are exemplified herein with configurations having different nucleotides labeled with different dye compounds, it is understood that the kits can include 2, 3, 4, or more different nucleotides with the same dye compound. In some embodiments, the kits further include an enzyme and a buffer suitable for the action of the enzyme. In some such embodiments, the enzyme is a polymerase, a terminal deoxynucleotidyl transferase, or a reverse transcriptase. In particular embodiments, the enzyme is a DNA polymerase, such as DNA polymerase 812 (Pol 812) or DNA polymerase 1901 (Pol 1901). The amino acid sequences of the Pol 812 and Pol 1901 polymerases are described in, e.g., U.S. Patent Application 16 / 670,876, filed October 31, 2019, and U.S. Patent Application 16 / 703,569, filed December 4, 2019, which are incorporated by reference herein.
[0203] Other components included in such kits can include buffers and the like. Any of the nucleotide components of the present disclosure, including mixtures of different nucleotides, can be provided in a kit in concentrated form for dilution prior to use. In such embodiments, a suitable dilution buffer can also be included. Likewise, one or more components determined in the methods set forth herein can be included in kits of the present disclosure.
[0204] Sequencing method
[0205] Labeled nucleotides or nucleosides according to the present disclosure can be used in any analytical method, for example, methods that include detection of a fluorescent label attached to a nucleotide or nucleoside, whether the labeled nucleotide or nucleoside is present alone or incorporated or attached in a larger molecular structure or conjugate. In this context, the term "incorporated polynucleotide" can mean that the 5' phosphate participates in the formation of a phosphodiester bond to the 3'-OH group of a second (modified or unmodified) nucleotide, which second nucleotide can itself form part of a longer polynucleotide chain. The 3' end of a nucleotide described herein can or can not participate in the formation of a phosphodiester bond to the 5' phosphate of another (modified or unmodified) nucleotide. Thus, in one non-limiting embodiment, the present disclosure provides a method of detecting a nucleotide incorporated into a polynucleotide, the method comprising: (a) incorporating at least one nucleotide of the present disclosure into a polynucleotide, and (b) detecting the nucleotide incorporated into the polynucleotide by detecting a fluorescent signal of a dye compound attached to the nucleotide.
[0206] The method can comprise a synthesis step (a) of incorporating one or more nucleotides according to the present disclosure into a polynucleotide; and a detection step (b) of detecting the one or more nucleotides incorporated into the polynucleotide by detecting or quantitatively measuring fluorescence of the one or more nucleotides incorporated into the polynucleotide.
[0207] Some embodiments of the present application relate to a sequencing method comprising: (a) incorporating at least one labeled nucleotide described herein into a polynucleotide; and (b) detecting the labeled nucleotide incorporated into the polynucleotide by detecting a fluorescent signal of a new fluorescent dye attached to the nucleotide.
[0208] Some embodiments of the present disclosure relate to a method of determining a sequence of a target single-stranded polynucleotide, comprising:
[0209] (a) incorporating a nucleotide described herein comprising a 3'-OH blocking group and a detectable label into a replication polynucleotide strand complementary to at least a portion of the target polynucleotide strand;
[0210] (b) detecting the identity of the nucleotide incorporated into the replication polynucleotide strand; and
[0211] (c) chemically removing the label and 3' blocking group from the nucleotide incorporated into the replicated polynucleotide strand.
[0212] In some embodiments, the sequencing method further comprises (d) washing the chemically removed label and 3' blocking group from the replicated polynucleotide strand. In some such embodiments, the 3' blocking group and the detectable label are removed prior to incorporation of the next complementary nucleotide. In some further embodiments, the 3' blocking group and detectable label are removed in a single chemical reaction step. In some embodiments, the washing step (d) also removes unincorporated nucleotides. In some further embodiments, a palladium scavenger is also used in the washing step following chemical cleavage of the label and 3' blocking group.
[0213] In some embodiments, steps (a) through (d) are repeated until the sequence of the portion of the template polynucleotide strand is determined. In some such embodiments, steps (a) through (d) are repeated at least 50 times, at least 75 times, at least 100 times, at least 150 times, at least 200 times, at least 250 times, or at least 300 times.
[0214] In some embodiments, the label and 3' blocking group are removed in two separate chemical reactions. In some such embodiments, removing the label from the incorporated nucleotide in the replicated polynucleotide strand comprises contacting the replicated strand comprising the incorporated nucleotide with a first cleavage solution. In some such embodiments, the first cleavage solution comprises a phosphine, such as a trialkylphosphine. Non-limiting examples of trialkylphosphines include tris(hydroxypropyl)phosphine (THP), tris(2-carboxyethyl)phosphine (TCEP), tris(hydroxymethyl)phosphine (THMP), or tris(hydroxyethyl)phosphine (THEP). In one embodiment, the first cleavage solution contains THP. In some such embodiments, removing the 3' blocking group from the incorporated nucleotide in the replicated polynucleotide strand comprises contacting the replicated strand comprising the incorporated nucleotide with a second cleavage solution. In some such embodiments, the second cleavage solution comprises a palladium (Pd) catalyst. In some further embodiments, the Pd catalyst is a Pd(0) catalyst. In some such embodiments, the Pd(0) is prepared in situ by mixing a Pd(II) complex [(PdCl(C3H5))2] with THP. The molar ratio of Pd(II) complex to THP is about 1 :2, 1 :3, 1 :4, 1 :5, 1 :6, 1 :7, 1 :8, 1 :9, or 1 : 10. In one embodiment, the molar ratio of Pd:THP is 1 :5. In some further embodiments, one or more reducing agents, such as ascorbic acid or a salt thereof (e.g., sodium ascorbate), can be added. In some embodiments, the second cleavage solution can comprise one or more buffers, such as a primary amine, a secondary amine, a tertiary amine, a carbonate, a phosphate, or a borate, or a combination thereof. In some further embodiments, the buffer comprises ethanolamine (EA), tris(hydroxymethyl)aminomethane (Tris), glycine, sodium carbonate, sodium phosphate, sodium borate, 2-dimethylaminoethanol (DMEA), 2-diethylaminoethanol (DEEA), N,N,N',N'-tetramethylethylenediamine (TEMED), or N,N,N',N'-tetraethylethylenediamine (TEEDA), or a combination thereof. In one embodiment, the buffer reagent is DEEA. In another embodiment, the buffer comprises one or more inorganic salts, such as a carbonate, a phosphate, or a borate, or a combination thereof. In one embodiment, the inorganic salt is a sodium salt. In some other embodiments, the second cleavage solution contains NaIO4or Oxone®. In some further embodiments, the 3' blocked nucleotide contains an AOM group, the second cleavage solution contains a palladium (Pd) catalyst and one or more buffers described herein (e.g., a tertiary amine, such as DEEA) and has a pH of about 9.0 to 10.0 (e.g., 9.6 or 9.8).
[0215] In some alternative embodiments, the label and 3'-OH blocking group are removed in a single chemical reaction. In some such embodiments, the label is attached to the nucleotide through a cleavable linker that comprises the same moiety as the 3' blocking group, e.g., the linker and 3' blocking group can both comprise an acetal moiety as described herein or a thiocarbamate moiety In some such embodiments, the single chemical reaction is performed in a cleavage solution containing the Pd catalyst described above.
[0216] In some further embodiments, the nucleotides used in the incorporation step (a) are fully functionalized A, C, T, and G nucleotide triphosphates, each comprising a 3’ blocking group described herein. In some such embodiments, the nucleotides herein provide superior stability in solution during sequencing as compared to the same nucleotides protected with standard 3’-O-azidomethyl blocking groups. For example, the acetal or thiocarbamate blocking groups disclosed herein can confer at least 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 300%, 400%, 500%, 600%, 700%, 800%, 900%, 1000%, 1500%, 2000%, 2500%, or 3000% improved stability as compared to azidomethyl-protected 3’-OH over the same period of time under the same conditions, resulting in reduced phasing and longer sequencing read lengths. In some embodiments, the stability is measured at ambient temperature or at a temperature below ambient temperature, e.g., 4-10 °C. In other embodiments, the stability is measured at an elevated temperature, e.g., 40 °C, 45 °C, 50 °C, 55 °C, 60 °C, or 65 °C. In some such embodiments, the stability is measured in solution at an alkaline pH environment, e.g., at pH 9.0, 9.2, 9.4, 9.6, 9.8, or 10.0. In some further embodiments, the phasing of the 3’-blocked nucleotides having the 3’ blocking groups described herein is less than about 0.25, 0.24, 0.23, 0.22, 0.21, 0.20, 0.19, 0.18, 0.17, 0.16, 0.15, 0.14, 0.13, 0.12, 0.11, 0.10, 0.09, 0.08, 0.07, 0.06, or 0.05 after 50, 100, or 150 SBS cycles. In some further embodiments, the phasing of the 3’-blocked nucleotides having the 3’ blocking groups described herein is less than about 0.25, 0.24, 0.23, 0.22, 0.21, 0.20, 0.19, 0.18, 0.17, 0.16, 0.15, 0.14, 0.13, 0.12, 0.11, 0.10, 0.09, 0.08, 0.07, 0.06, or 0.05 after 50, 100, or 150 SBS cycles. In one embodiment, each ffN comprises the 3’-AOM group.
[0217] In some embodiments, the 3'-blocked nucleotides described herein provide superior unblocking rates in solution during the chemical break-off step of a sequencing run compared to the same nucleotide protected by a standard 3'-O-azidomethyl blocking group. For example, the acetal (e.g., AOM) or thiocarbamate blocking groups disclosed herein can confer improved unblocking rates of at least 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 150%, 200%, 300%, 400%, 500%, 600%, 700%, 800%, 900%, 1000%, 1500%, or 2000% compared to 3'-OH protected by an azidomethyl blocking agent such as tris(hydroxypropyl)phosphine, thereby reducing the total sequencing cycle time. In some embodiments, the unblocking time per nucleotide is reduced by approximately 5%, 10%, 20%, 30%, 40%, 50%, or 60%. For example, under certain chemical reaction conditions, the deblocking times of 3′-AOM and 3′-O-azidomethyl groups are approximately 4-5 seconds and 9-10 seconds, respectively. In some embodiments, the half-life (t) of the AOM blocking group is... 1 / 2 It is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 times faster than the azidomethyl blocking group. In some such embodiments, the AOM's t 1 / 2 The t of azidomethyl is approximately 1 minute. 1 / 2 Approximately 11 minutes. In some embodiments, the unblocking rate is measured at or below ambient temperature (e.g., 4-10°C). In other embodiments, the unblocking rate is measured at elevated temperatures, such as 40°C, 45°C, 50°C, 55°C, 60°C, or 65°C. In some such embodiments, the unblocking rate is measured in a solution at an alkaline pH environment, for example, at pH 9.0, 9.2, 9.4, 9.6, 9.8, or 10.0. In some such embodiments, the molar ratio of the unblocking agent to the substrate (i.e., the 3' protected nucleoside or nucleotide) is about 10:1, about 5:1, about 2:1, about 1:1, about 1:2, about 1:5, or about 1:10. In one embodiment, each ffN contains the 3'-AOM group.
[0218] In any embodiment of the method described herein, the labeled nucleotide is a nucleoside triphosphate. In any embodiment of the method described herein, the target polynucleotide chain is attached to a solid support, such as a flow cell.
[0219] In one embodiment, in the synthesis step, at least one nucleotide is incorporated into the polynucleotide by the action of a polymerase enzyme. In some such embodiments, the polymerase enzyme can be DNA polymerase Pol 812 or Pol 1901. However, other methods of joining nucleotides to polynucleotides can be used, such as chemical oligonucleotide synthesis or ligation of labeled oligonucleotides to unlabeled oligonucleotides. Thus, the term "incorporation" when referring to nucleotides and polynucleotides can encompass polynucleotide synthesis by chemical methods as well as enzymatic methods.
[0220] In one particular embodiment, the synthesis step is performed and optionally comprises incubating a template polynucleotide strand with a reaction mixture comprising the labeled 3' blocked nucleotide of the disclosure. A polymerase enzyme can also be provided under conditions that allow a phosphodiester linkage to form between a free 3'-OH group on a polynucleotide strand that is annealed to the template polynucleotide strand and the 5' phosphate group on the nucleotide. Thus, the synthesis step can comprise the formation of a polynucleotide strand following the direction of base pairing of the nucleotide to the complementary bases of the template strand.
[0221] In all embodiments of the method, the detection step can be performed while the polynucleotide strand incorporating the labeled nucleotide is annealed to the template strand, or after a denaturation step in which the two strands are separated. Other steps, such as chemical or enzymatic reaction steps or purification steps, can be included between the synthesis step and the detection step. In particular, the target strand incorporating the labeled nucleotide can be isolated or purified and then further processed or used in subsequent analyses. For example, a target polynucleotide labeled with the nucleotides described herein in the synthesis step can be subsequently used as a labeled probe or primer. In other embodiments, the product of the synthesis step described herein can be subjected to further reaction steps, and if desired, the products of these subsequent steps can be purified or isolated.
[0222] Those familiar with standard molecular biology techniques will be familiar with suitable conditions for the synthesis step. In one embodiment, the synthesis step can resemble a standard primer extension reaction using nucleotide precursors, including the nucleotides described herein, to form an extended target strand complementary to a template strand in the presence of a suitable polymerase. In other embodiments, the synthesis step can itself constitute part of an amplification reaction, resulting in a labeled double-stranded amplification product composed of annealed complementary strands derived from replication of the target and template polynucleotide strands. Other exemplary synthesis steps include nick translation, strand displacement polymerization, random primed DNA labeling, and the like. A particularly useful polymerase for the synthesis step is an enzyme capable of catalyzing incorporation of the nucleotides described herein. A variety of naturally occurring or modified polymerases can be used. For example, a thermostable polymerase can be used for synthesis reactions performed using thermal cycling conditions, while a thermostable polymerase can not be required for isothermal primer extension reactions. Suitable thermostable polymerases capable of incorporating the nucleotides according to the present disclosure include those described in WO 2005 / 024010 or WO 06 / 120433, each of which is incorporated herein by reference. In synthesis reactions performed at lower temperatures (e.g., 37°C), the polymerase need not necessarily be a thermostable polymerase, and thus the choice of polymerase depends on many factors, such as reaction temperature, pH, strand displacement activity, and the like.
[0223] In particular non-limiting embodiments, the present disclosure encompasses methods of nucleic acid sequencing, resequencing, whole genome sequencing, single nucleotide polymorphism scoring, any other application involving detection of the labeled nucleotides or nucleosides described herein when incorporated into a polynucleotide. Any of a variety of other applications that would benefit from the use of nucleotide-labeled polynucleotides comprising fluorescent dyes can use the dye-labeled nucleotides or nucleosides described herein.
[0224] In one particular embodiment, the present disclosure provides the use of labeled nucleotides according to the present disclosure in a polynucleotide sequencing-by-synthesis (SBS) reaction. Sequencing-by-synthesis generally involves the sequential addition of one or more nucleotides or oligonucleotides in the 5' to 3' direction to a growing polynucleotide chain using a polymerase or ligase enzyme to form an extended polynucleotide chain complementary to the template nucleic acid to be sequenced. The identity of the base present in the one or more added nucleotides can be determined in a detection or "imaging" step. The identity of the added base can be determined after each nucleotide incorporation step. The sequence of the template can then be inferred using conventional Watson-Crick base pairing rules. For example, in scoring single nucleotide polymorphisms, it can be useful to determine the identity of a single base using the labeled nucleotides described herein, and such single base extension reactions are within the scope of the present disclosure.
[0225] In one embodiment of the present disclosure, the sequence of a template polynucleotide is determined by detecting the incorporation of one or more 3' blocked nucleotides described herein into a nascent strand complementary to the template polynucleotide to be sequenced by detecting a fluorescent label attached to the incorporated nucleotide. Sequencing of the template polynucleotide can be primed with a suitable primer (or prepared as a hairpin construct which will contain the primer as part of the hairpin) and the nascent strand is extended in a stepwise fashion by the addition of nucleotides to the 3' end of the primer under polymerase catalyzed reactions.
[0226] In particular embodiments, each of the different nucleotide triphosphates (A, T, G, and C) can be labeled with a different fluorophore and also contain a blocking group at the 3' position to prevent uncontrolled polymerization. Optionally, one of the four nucleotides can be unlabeled (dark). The polymerase incorporates the nucleotide into a nascent strand complementary to the template polynucleotide, and the blocking group prevents further incorporation of nucleotides. Any unbound nucleotides can be washed away, and the fluorescent signal from each incorporated nucleotide can be optically "read" by a suitable means (e.g., a charge-coupled device using laser excitation and appropriate emission filters). The 3'-blocking group and fluorescent dye compound can then be removed (deprotected) simultaneously or sequentially to expose the nascent strand for further incorporation of nucleotides. Typically, the identity of the incorporated nucleotide will be determined after each incorporation step, but this is not strictly necessary. Similarly, U.S. Patent 5,302,509 (which is incorporated herein by reference) discloses a method of sequencing a polynucleotide immobilized on a solid support.
[0227] As described above, the method utilizes the incorporation of fluorescently labeled 3'-blocked nucleotides A, G, C, and T into a growing strand complementary to an immobilized polynucleotide in the presence of a DNA polymerase. The polymerase incorporates a base complementary to the target polynucleotide, but is prevented from further addition by the 3'-blocking group. The label of the incorporated nucleotide can then be determined, and the blocking group removed by chemical cleavage to allow further polymerization to occur. The nucleic acid template to be sequenced in a sequencing-by-synthesis reaction can be any polynucleotide for which sequencing is desired. The nucleic acid template used in the sequencing reaction will typically comprise a double-stranded region with a free 3'-OH group that serves as a primer or as a starting point for further addition of other nucleotides in the sequencing reaction. The region of the template to be sequenced will overhang this free 3'-OH group on the complementary strand. The overhanging end region of the template to be sequenced can be single-stranded or double-stranded, provided that there is a "gap" in the strand complementary to the strand of the template to be sequenced to provide a free 3'-OH group for initiation of the sequencing reaction. In such embodiments, sequencing can proceed by strand displacement. In some embodiments, a primer with a free 3'-OH group can be added as a separate component (e.g., a short oligonucleotide) that hybridizes to a single-stranded region of the template to be sequenced. Alternatively, the primer and the template strand to be sequenced can each form part of a partially self-complementary nucleic acid strand that is capable of forming an intramolecular duplex (e.g., a hairpin loop structure). Hairpin polynucleotides and methods of attaching them to solid supports are disclosed in PCT publications WO 01 / 57248 and WO 2005 / 047301, each of which is incorporated herein by reference. Nucleotides can be added to the growing primer sequentially, resulting in the synthesis of a polynucleotide strand in the 5' to 3' direction. The identity of the bases that have been added can be determined (particularly, but not necessarily, after each addition of a nucleotide) to provide sequence information for the nucleic acid template. Thus, the nucleotide is incorporated into the nucleic acid strand (or polynucleotide) via formation of a phosphodiester linkage with the 5' phosphate group of the nucleotide to the free 3'-OH group of the nucleic acid strand.
[0228] The nucleic acid template to be sequenced can be DNA or RNA, or even a hybrid molecule composed of deoxynucleotides and ribonucleotides. The nucleic acid template can comprise naturally-occurring and / or non-naturally-occurring nucleotides and natural or non-natural backbone linkages, provided that they do not prevent replication of the template in the sequencing reaction.
[0229] In some embodiments, the nucleic acid template to be sequenced can be attached to a solid support via any suitable attachment method known in the art, for example, by covalent attachment. In some embodiments, the template polynucleotide can be directly attached to a solid support (e.g., a silica-based support). However, in other embodiments of the disclosure, the surface of the solid support can be modified in some way to allow for direct covalent attachment of the template polynucleotide, or the template polynucleotide can be immobilized by a hydrogel or a polyelectrolyte multilayer, which can itself be non-covalently attached to the solid support.
[0230] Embodiments and alternatives of sequencing by synthesis
[0231] Some embodiments include pyrosequencing techniques. Pyrosequencing detects the release of inorganic pyrophosphate (PPi) as specific nucleotides are incorporated into a nascent strand, (Ronaghi, M., Karamohamed, S., Pettersson, B., Uhlen, M., and Nyren, P. (1996) "Real-time DNA sequencing using detection of pyrophosphate release." Analytical Biochemistry 242(1), 84-9; Ronaghi, M. (2001) "Pyrosequencing sheds light on DNA sequencing." Genome Res. 11(1), 3-11; Ronaghi, M., Uhlen, M., and Nyren, P. (1998) "A sequencing method based on real-time pyrophosphate." Science 281(5375), 363; U.S. Patents 6,210,891; 6,258,568; and 6,274,320, the disclosures of which are incorporated by reference herein in their entireties). In pyrosequencing, the released PPi can be detected by being immediately converted to adenosine triphosphate (ATP) by ATP sulfurylase, and the level of ATP generated is detected via the photons produced by luciferase. The nucleic acid to be sequenced can be attached to features in an array, and the array can be imaged to capture the chemiluminescent signal generated as nucleotides are incorporated on the features of the array. After the array is treated with a particular nucleotide type (e.g., A, T, C, or G), an image can be obtained. The images obtained after the addition of each nucleotide type can differ depending on the features detected in the array. These differences in the images reflect the different sequence content of the features on the array. However, the relative position of each feature will remain the same in the images. The images can be stored, processed, and analyzed using the methods set forth herein. For example, the images obtained after treating the array with each of the different nucleotide types can be processed in the same manner as the images obtained from the different detection channels of the reversible terminator-based sequencing methods exemplified herein.
[0232] In another exemplary SBS type, cyclic sequencing is accomplished by stepwise addition of reversible terminator nucleotides that comprise, for example, cleavable or photobleachable dye labels as described in WO 04 / 018497 and U.S. Patent 7,057,026, the disclosures of which are incorporated herein by reference. This method is commercialized by Solexa (now Illumina, Inc.) and is also described in WO 91 / 06678 and WO 07 / 123,744, each of which is incorporated herein by reference. The availability of fluorescently labeled terminators (both terminators are reversible) and cleavage of the fluorescent label facilitates efficient cyclic reversible termination (CRT) sequencing. The polymerase can also be co-engineered to efficiently incorporate and extend from these modified nucleotides.
[0233] Preferably, in a reversible terminator-based sequencing embodiment, the label does not substantially inhibit extension under SBS reaction conditions. However, the detection label can be removed by cleavage or degradation. An image can be captured after the label is incorporated into the nucleic acid features of the array. In a particular embodiment, each cycle involves simultaneous delivery of four different nucleotide types to the array, and each nucleotide type has a spectrally distinct label. Four images can then be acquired, each using a detection channel selective for one of the four different labels. Alternatively, different nucleotide types can be added sequentially, and an image of the array can be acquired between each addition step. In such an embodiment, each image will show nucleic acid features that have incorporated a particular type of nucleotide. Because the sequence content of each feature is different, different features will be present or absent in different images. However, the relative positions of the features will remain unchanged in the images. Images acquired from this reversible terminator-SBS method can be stored, processed, and analyzed as described herein. After the image capture step, the label can be removed, and the reversible terminator moiety can be removed for subsequent nucleotide addition and detection cycles. Detection of label removal in a particular cycle and prior to subsequent cycles can provide an advantage of reducing background signal and cross-talk between cycles. Examples of useful labels and removal methods are as follows.
[0234] Some embodiments can utilize detection of four different nucleotides using fewer than four different labels. For example, SBS can be performed utilizing the methods and systems described in the incorporated material of U.S. Patent Publication 2013 / 0079232. As a first example, a pair of nucleotide types can be detected at the same wavelength, but distinguished based on a difference in intensity of one member of the pair relative to the other, or based on a change in one member of the pair (e.g., via chemical modification, photochemical modification, or physical modification) that results in a clear signal appearance or disappearance compared to the signal detected for the other member of the pair. As a second example, three of the four different nucleotide types can be detected under certain conditions, while the fourth nucleotide type lacks a label that is detectable under those conditions, or is detected to a minimal extent under those conditions (e.g., minimal detection due to background fluorescence, etc.). Incorporation of the first three nucleotide types into the nucleic acid can be determined based on the presence of their respective signals, and incorporation of the fourth nucleotide type into the nucleic acid can be determined from the absence or minimal detection of any signal. As a third example, one nucleotide type can include labels that are detected in two different channels, while the other nucleotide types are detected in no more than one channel. The foregoing three exemplary configurations are not considered to be mutually exclusive, and can be used in various combinations. An exemplary embodiment that incorporates all three examples is a fluorescence-based SBS method that uses a first nucleotide type that is detected in a first channel (e.g., dATP with a label that is detected in the first channel when excited by a first excitation wavelength), a second nucleotide type that is detected in a second channel (e.g., dCTP with a label that is detected in the second channel when excited by a second excitation wavelength), a third nucleotide type that is detected in both the first and second channels (e.g., dTTP with at least one label that is detected in both channels when excited by the first and / or second excitation wavelengths), and a fourth nucleotide type that lacks a label (e.g., dGTP without a label that is detected in either channel).
[0235] In addition, as described in the incorporated material of U.S. Publication Patent 2013 / 0079232, sequencing data can be obtained using a single channel. In this so-called single-dye sequencing method, a first nucleotide type is labeled, but the label is removed after the first image is generated, and a second nucleotide type is labeled only after the first image is generated. A third nucleotide type retains its label in both the first and second images, while a fourth nucleotide type is unlabeled in both images.
[0236] Some embodiments can utilize sequencing by ligation techniques. Such techniques utilize DNA ligase to incorporate oligonucleotides and identify incorporation of these oligonucleotides. The oligonucleotides typically have different labels associated with the identity of a particular nucleotide in the sequence to which the oligonucleotide hybridizes. As with other SBS methods, after treating an array of nucleic acid features with labeled sequencing reagents, an image can be obtained. Each image will show nucleic acid features that have had a particular type of label bound. Because the sequence content of each feature is different, different features will be present or absent in different images, but the relative position of the features will remain the same in the images. Images obtained from sequencing by ligation methods can be stored, processed, and analyzed as described herein. Exemplary SBS systems and methods that can be used with the methods and systems described herein are discussed in U.S. Patents 6,969,488, 6,172,218, and 6,306,597, the disclosures of which are incorporated by reference herein in their entireties.
[0237] Some embodiments can utilize nanopore sequencing (Deamer, D. W. & Akeson, M. "Nanopores and nucleic acids: prospects for ultrarapid sequencing." Trends Biotechnol. 18, 147-151 (2000); Deamer, D. and D. Branton, "Characterization of nucleic acids by nanopore analysis", Acc. Chem. Res. 35:817-825 (2002); Li, J., M. Gershow, D. Stein, E. Brandin, and J. A. Golovchenko, "DNA molecules and configurations in a solid-state nanopore microscope" Nat. Mater. 2:611-615 (2003), the disclosures of which are incorporated by reference in their entireties). In such embodiments, a target nucleic acid is passed through a nanopore. The nanopore can be a synthetic pore or a biological membrane protein, such as a-hemolysin. As the target nucleic acid passes through the nanopore, each base pair can be identified by measuring fluctuations in the electrical conductivity of the pore. (U.S. Patent 7,001,792; Soni, G. V. & Meller, "A. Progress toward ultrafast DNA sequencing using solid-state nanopores." Clin. Chem. 53, 1996-2001 (2007); Healy, K. "Nanopore-based single-molecule DNA analysis." Nanomed. 2, 459-481 (2007); Cockroft, S. L., Chu, J., Amorin, M. & Ghadiri, M. R. "A single-molecule nanopore device detects DNA polymerase activity with single-nucleotide resolution." J. Am. Chem. Soc. 130, 818-820 (2008), the disclosures of which are incorporated by reference in their entireties). As described herein, data obtained from nanopore sequencing can be stored, processed, and analyzed.In particular, data can be processed as images in accordance with the exemplary processing of optical images and other images set forth herein.
[0238] Some other embodiments of sequencing methods involve the use of the 3’- blocked nucleotides described herein in nanoball sequencing technology, such as those described in U.S. Patent 9,222,132, the disclosure of which is incorporated herein by reference. By a rolling circle amplification (RCA) process, a large number of discrete DNA nanoballs can be generated. The nanoball mixture is then distributed onto a patterned slide surface containing features that allow individual nanoballs to be associated with each location. During the generation of the DNA nanoballs, the DNA is fragmented and linked to a first of four adapter sequences. The template is amplified, circularized, and cleaved with a type II endonuclease. A second set of adapters is added, followed by amplification, circularization, and cleavage. The remaining two adapters repeat this process. The final product is a circular template with four adapters, each separated by a template sequence. The library molecules go through a rolling circle amplification step, generating a large number of concatemers called DNA nanoballs, which are then deposited in a flow cell. Goodwin et al., “Coming of age: ten years of next-generation sequencing technologies,” Nat Rev Genet. 2016; 17(6):333-51.
[0239] Some embodiments can utilize methods involving real-time monitoring of DNA polymerase activity. Incorporation of nucleotides can be detected by, for example, fluorescence resonance energy transfer (FRET) interactions between fluorophore-labeled polymerases and gamma-phosphate labeled nucleotides as described in U.S. Patent 5,426,038, for example, U.S. Patents 7,329,492 and 7,211,414, the entire contents of which are incorporated herein by reference, or by zero-mode waveguides as described in U.S. Patent 7,315,019, incorporated herein by reference, and the use of fluorescent nucleotide analogs and engineered polymerases as described in U.S. Patent 7,405,281 and U.S. Patent 2008 / 0108082, incorporated herein by reference. Illumination can be confined to volumes on the order of zeptoliters around surface-tethered polymerases, such that incorporation of fluorescently labeled nucleotides can be observed in low background (Levene, MJ, et al., "Zero-mode waveguides for single-molecule analysis at high concentrations." Science 299, 682-686 (2003); Lundquist, PM, et al., "Parallel confocal detection of single molecules in real time." Opt. Lett. 33, 1026-1028 (2008); Korlach, J., et al., "Selective aluminum passivation for targeted immobilization of single DNA polymerase molecules in zero-mode waveguide nano structures." Proc. Natl. Acad. Sci. USA 105, 1176-1181 (2008), the disclosures of which are incorporated herein by reference in their entireties). Images obtained from such methods can be stored, processed, and analyzed as described herein.
[0240] Some SBS embodiments include detecting protons released upon incorporation of nucleotides into extension products. For example, sequencing based on detection of released protons can use electrical detectors and related technology, which is commercially available from Ion Torrent (a subsidiary of Life Technologies, Guilford, CT), or can use sequencing methods and systems described in U.S. Published Patents 2009 / 0026082; 2009 / 0127589; 2010 / 0137143; and 2010 / 0282617, the entire contents of which are incorporated herein by reference. The methods set forth herein using kinetic exclusion to amplify target nucleic acids can be readily applied to substrates for detecting protons. More specifically, the methods set forth herein can be used to generate clonal populations of amplicons for detecting protons.
[0241] The SBS methods described above can advantageously be performed in a multiplexed fashion, such that multiple different target nucleic acids are operated on simultaneously. In particular embodiments, different target nucleic acids can be processed in a common reaction vessel or on the surface of a particular substrate. This allows for convenient delivery of sequencing reagents, removal of unreacted reagents, and detection of incorporation events in a multiplexed fashion. In embodiments using surface-bound target nucleic acids, the target nucleic acids can be in an array format. In an array format, the target nucleic acids can typically be bound to a surface in a spatially distinguishable fashion. The target nucleic acids can be bound by direct covalent linkage, linkage to a bead or other particle, or binding to a polymerase or other molecule attached to the surface. The array can include a single copy of a target nucleic acid at each site (also referred to as a feature), or multiple copies of the same sequence can be present at each site or feature. Multiple copies can be generated by amplification methods such as bridge amplification or emulsion PCR, as described in further detail below.
[0242] The methods set forth herein can use arrays having features of various densities, e.g., at least about 10 features / cm 2 , 100 features / cm 2 , 500 features / cm 2 , 1,000 features / cm 2 , 5,000 features / cm 2 , 10,000 features / cm 2 , 50,000 features / cm 2 , 100,000 features / cm 2 , 1,000,000 features / cm 2 , 5,000,000 features / cm 2 or higher.
[0243] The methods set forth herein are advantageous in that they provide rapid and efficient detection of multiple target nucleic acids in parallel. Accordingly, the present disclosure provides integrated systems capable of preparing and detecting nucleic acids using techniques known in the art, such as those exemplified above. Thus, the integrated systems of the present disclosure can include fluidic components capable of delivering amplification reagents and / or sequencing reagents to one or more immobilized DNA fragments, including components such as pumps, valves, vessels, fluidic lines, and the like. Flow cells can be configured in the integrated system and / or used to detect target nucleic acids. Exemplary flow cells are described in, e.g., U.S. Patent Publication 2010 / 0111768 and U.S. Patent 13 / 273,666, each of which is incorporated herein by reference. As exemplified for flow cells, one or more fluidic components of the integrated system can be used in amplification methods and detection methods. By way of nucleic acid sequencing embodiments, one or more fluidic components of the integrated system can be used in amplification methods set forth herein and for delivery of sequencing reagents in sequencing methods, such as those exemplified above. Alternatively, the integrated system can include separate fluidic systems to perform amplification methods and detection methods. Examples of integrated sequencing systems capable of producing amplified nucleic acids and also capable of determining nucleic acid sequences include, but are not limited to, the MiSeq® platform (Illumina, Inc., San Diego, CA) and the devices described in U.S. Patent 13 / 273,666, incorporated herein by reference. ™ platform (Illumina, Inc., San Diego, CA) and the devices described in U.S. Patent 13 / 273,666, incorporated herein by reference.
[0244] Arrays in which polynucleotides have been attached directly to silica-based supports are, for example, those disclosed in WO 00 / 06770 (incorporated herein by reference), in which polynucleotides are immobilized on glass supports by reaction of pendant epoxy groups on the glass with internal amines on the polynucleotides. In addition, polynucleotides can be attached to solid supports by reaction of sulfur-based nucleophiles with the solid support, for example, as described in WO 2005 / 047301 (incorporated herein by reference). Further examples of solid-supported template polynucleotides are those in which the template polynucleotide is attached to a hydrogel loaded on a silica-based solid support or other solid support, for example, as described in WO 00 / 31148, WO 01 / 01143, WO 02 / 12566, WO 03 / 014392, U.S. Patent 6,465,178, and WO 00 / 53812, each of which is incorporated herein by reference.
[0245] A particular surface on which a template polynucleotide can be immobilized is a polyacrylamide hydrogel. Polyacrylamide hydrogels are described in the above cited references and in WO 2005 / 065814, which is incorporated herein by reference. Particular hydrogels that can be used include those in WO 2005 / 065814 and U.S. Published Patent 2014 / 0079923. In one embodiment, the hydrogel is PAZAM (poly N-(5-azidomethyl)acrylamide-co-acrylamide).
[0246] DNA template molecules can be attached to beads or microparticles, for example, as described in U.S. Patent 6,172,218 (incorporated herein by reference). Attachment to beads or microparticles can be used for sequencing applications. Libraries of beads can be prepared in which each bead contains a different DNA sequence. Exemplary libraries and methods of their creation are described in Nature, 437, 376-380 (2005); Science, 309, 5741, 1728-1732 (2005), each of which is incorporated herein by reference. Sequencing of arrays of such beads using the nucleotides set forth herein is within the scope of the present disclosure.
[0247] Templates to be sequenced can form part of an "array" on a solid support, in which case the array can take any convenient form. Thus, the methods of the present disclosure are applicable to all types of high-density arrays, including single-molecule arrays, cluster arrays, and bead arrays. The labeled nucleotides of the present disclosure can be used for template sequencing on substantially any type of array, including but not limited to those formed by immobilizing nucleic acid molecules on a solid support.
[0248] However, the labeled nucleotides of the present invention are particularly advantageous in the context of sequencing of cluster arrays. In a cluster array, different regions (often referred to as sites or features) on the array contain multiple polynucleotide template molecules. Typically, the multiple polynucleotide molecules cannot be individually resolved by optical means, but are detected as a whole. Depending on the manner in which the array is formed, each site on the array can contain multiple copies of a single polynucleotide molecule (e.g., the site is homogenous for a particular single- or double-stranded nucleic acid species) or even multiple copies of a small number of different polynucleotide molecules (e.g., multiple copies of two different nucleic acid species). Cluster arrays of nucleic acid molecules can be produced using techniques generally known in the art. By way of example, WO 98 / 44151 and WO 00 / 18957, each of which is incorporated herein, describe methods of amplification of nucleic acids in which both the template and the amplification product are held immobilized on a solid support to form an array composed of clusters or "colonies" of immobilized nucleic acid molecules. The nucleic acid molecules present on a cluster array prepared according to these methods are suitable templates for sequencing using nucleotides labeled with the dye compounds of the present invention.
[0249] The labeled nucleotides of the present disclosure can also be used for sequencing of templates on a single molecule array. As used herein, the term "single molecule array" or "SMA" refers to a population of polynucleotide molecules distributed (or arrayed) on a solid support, wherein the spacing of any individual polynucleotide from all other molecules of the population is such that the individual polynucleotide molecules can be individually resolved. Thus, in some embodiments, the target nucleic acid molecules immobilized on the surface of the solid support can be resolved optically. This means that one or more distinct signals each representing one polynucleotide will occur within the resolvable area of the particular imaging device used.
[0250] Single molecule detection can be achieved, wherein the spacing between adjacent polynucleotide molecules on the array is at least 100 nm, more particularly at least 250 nm, still more particularly at least 300 nm, even more particularly at least 350 nm. Thus, each molecule can be individually resolved and detected as a single molecule fluorescent spot, and the fluorescence of said single molecule fluorescent spot also exhibits single step photobleaching.
[0251] The terms "individually resolved" and "individual resolution" are used herein to designate that, when visualized, one molecule on the array can be distinguished from its neighboring molecules. The separation between individual molecules on the array will be determined in part by the particular technology used to resolve the individual molecules. The general features of a single molecule array will be understood by reference to published applications WO 00 / 06770 and WO 01 / 57248, each incorporated herein by reference. While one use of the nucleotides of the present disclosure is in a sequencing by synthesis reaction, the use of the nucleotides is not limited to such methods. Indeed, the nucleotides can be advantageously used in any sequencing method that requires detection of a fluorescent label attached to a nucleotide incorporated into a polynucleotide.
[0252] In particular, the labeled nucleotides of the present disclosure can be used in automated fluorescent sequencing protocols, in particular fluorescent dye-terminator cycle sequencing based on the chain termination sequencing method of Sanger and colleagues. Such methods typically use enzymes and cycle sequencing to incorporate fluorescently labeled dideoxynucleotides into a primer extension sequencing reaction. The so-called Sanger sequencing method and related protocols (Sanger-type) utilize random chain termination with labeled dideoxynucleotides.
[0253] Thus, the present disclosure also encompasses labeled nucleotides which are dideoxynucleotides lacking a hydroxyl group at both the 3' and 2' positions, such dideoxynucleotides being suitable for use in Sanger-type sequencing methods and the like.
[0254] It should be recognized that the labeled nucleotides of the present disclosure that incorporate a 3’-blocking group can also be used in Sanger methods and related protocols, as the same effect obtained by using dideoxynucleotides can be achieved by using nucleotides with 3’-OH blocking groups: both prevent the incorporation of subsequent nucleotides. When using nucleotides with 3’ blocking groups according to the present application in Sanger-type sequencing methods, it should be understood that the dye compound or detectable label attached to the nucleotide need not be attached through a cleavable linking group, as each instance of incorporation of a labeled nucleotide of the present disclosure does not require subsequent incorporation of a nucleotide, and thus removal of the label from the nucleotide is not required.
[0255] In any embodiment of the methods described herein, the nucleotides used in sequencing applications are the 3’ blocked nucleotides described herein, e.g., nucleotides of formula (I), (la), or (II). In any embodiment, the 3’ blocked nucleotide is a nucleoside triphosphate. Examples
[0256] Further embodiments are disclosed in further detail in the following examples, which are in no way intended to limit the scope of the claims.
[0257] Example 1. Preparation of 3'-acetal blocked nucleosides
[0258] In this example, various 3’-acetal protected T nucleosides were prepared according to Scheme 2.
[0259]
[0260] Scheme 2
[0261]
[0262] Preparation of T1 : To a nitrogen-purged 100 mL oven-dried flask was added 5-iodo-2'-deoxyuridine (5.0 g, 14.12 mmol). This was co-evaporated with 30 mL of pyridine three times, then placed under nitrogen. Anhydrous pyridine (25 mL) was added and the reaction was stirred at room temperature until a homogenous solution was obtained (~ 15 minutes). The mixture was cooled to 0 °C in an ice-water bath and tert-butyldiphenylsilyl chloride (4.04 mL, 15.5 mmol) was added dropwise with vigorous stirring (~ 1 hour). The reaction was kept at 0 °C for 8 hours until all starting material was consumed (TLC). Saturated aqueous ammonium chloride solution (~ 15 mL) was added and the reaction was allowed to warm to room temperature. The mixture was diluted with ethyl acetate (100 mL) and washed with saturated aqueous ammonium chloride solution (200 mL). The organic layer was separated and the aqueous layer was extracted with ethyl acetate (4 x 50 mL). The organic layers were combined, dried (MgS04) and concentrated in vacuo to give ~ 8 g of a clear yellow oil. This crude product T1 was purified by flash column chromatography on silica gel as a white crystalline solid. Yield 6.94 g (83%). LC-MS (electrospray negative) 591.08 [M-H]-
[0263] Preparation of T2: T1 (6.23 g, 10.5 mmol), copper(I)iodide (200 mg, 1.05 mmol) and palladium(II) bis(triphenylphosphine) dichloride (369 mg, 0.526 mmol) were added to an oven-dried nitrogen-purged brown 500 mL three-necked flask under nitrogen. The flask was protected from light and anhydrous degassed DMF (200 mL) was added. To this solution was added 2,2,2-trifluoro-N-prop-2-ynylacetamide (4.74 g, 31.6 mmol) followed by degassed triethylamine (2.92 mL, 21.0 mmol). When no more starting material was observed by TLC analysis, the reaction was stirred at room temperature under nitrogen for 6 hours. The volatiles were removed in vacuo (~ 15 minutes) and the DMF was removed under high vacuum (~ 1 hour) to give a brown residue. This was dissolved in ethyl acetate (200 mL) and extracted with 0.1 M aqueous EDTA (2 x 200 mL). The aqueous layers were combined and further extracted with ethyl acetate (200 mL). The organic phases were combined, dried (MgS04) and the volatiles removed in vacuo (~ 30 minutes) and further dried under high vacuum (~ 1 hour) to give ~ 8 g of a crude brown / yellow oil. This mixture was purified by flash column chromatography on silica gel as an off-white solid. Yield: 6.0 g (85%). LC-MS (electrospray negative) 614.19 [M-H]-
[0264] Preparation of T3: To an oven-dried nitrogen-purged 100 mL flask containing the starting nucleoside T2 (2.0 g, 3.25 mmol) under nitrogen, dry DMSO (6.9 ml, 97.5 mmol) was added in one portion at room temperature and stirred until a homogenous solution was formed. Acetic acid (11.1 mL, 195 mmol) and acetic anhydride (15.1 mL, 162.09 mmol) were added dropwise in that order (ca. 5 min each). The mixture was warmed to 50 °C and stirred until complete consumption of the starting nucleoside was detected by TLC (EtOAc / petroleum ether 3:2) (~5 h). The reaction was then concentrated to half volume and cooled to ca. 0.5 °C with an ice bath. The work-up was initiated by slow addition of cold (~0.5 °C) NaHC03(saturated aqueous solution) (45 mL), followed by further stirring until no more foaming was observed (~15 min). The solution was allowed to warm to room temperature, and the aqueous phase was then extracted into EtOAc (3 x 100 mL). The combined organic layers were dried over MgS04, filtered, and the volatiles were evaporated under reduced pressure and further under high vacuum. The crude product T3 was purified by flash chromatography on silica gel as an off-white solid. Yield: 1.79 g (82%). LC-MS (electrospray negative) 674.20 [M-H] - .
[0265] Preparation of T4: To a solution of the starting nucleoside T3 (1.79 g, 2.649 mmol) in dry CH2Cl2(50 mL) under N2was added cyclohexene (1.34 mL, 13.2 mmol). The mixture was cooled to 0 °C with an ice bath and distilled sulfuryl chloride (322 µL, 3.97 mmol) was added slowly (ca. 20 min) under a nitrogen atmosphere. After 20 min stirring at this temperature, TLC (EtOAc: petroleum ether = 3:2 v / v) indicated that the starting nucleoside had been completely consumed. The chloride intermediate was then quenched by direct dropwise addition of the freshly distilled corresponding unsaturated alcohol (5 eq) as shown in Scheme 3. The resulting solution was stirred at room temperature for 2 h, then the volatiles were evaporated under reduced pressure. The oily residue was partitioned between EtOAc: brine (3:2) (125 mL). The organic layer was separated and the aqueous layer was further extracted into EtOAc (2 x 50 mL). The combined organic extracts were dried over MgSO4, filtered, and the volatiles were evaporated under reduced pressure. The oily residue was partitioned between EtOAc: brine (3:2) (125 mL). The organic layer was separated and the aqueous layer was further extracted into EtOAc (2 x 50 mL). The combined organic extracts were dried over MgSO4, filtered, and the volatiles were evaporated under reduced pressure. The crude product T4 was purified by flash chromatography on silica gel to give the final product as a yellow oil. Yield: AOM: 1.20 g (69%); PrOM: 1.29 g (71%); DPrOM: 1.34 g (71%).
[0266] 3’-AOM: Yellow oil. LC-MS (electrospray negative ion) [M-H] 684.24.
[0267] 3’-PrOM: Yellow oil. LC-MS (electrospray negative ion) [M-H] 682.22.
[0268] 3’-DPrOM: Yellow oil. LC-MS (electrospray negative ion) [M-H] 710.25.
[0269]
[0270] Scheme 3.
[0271] Preparation of T5: To the starting material T4 (1.04 g, 1.516 mmol) in a 50 mL round bottom flask was added anhydrous THF (9 mL) at room temperature under nitrogen. TBAF (1.0 M in THF, 1.7 mL, 1.70 mmol) was then added dropwise and the solution was stirred until all starting material was consumed (TLC) (~2 hours). The solution turned orange during the course of the reaction. The volatiles were removed in vacuo to give an orange residue which was dissolved in EtOAc (100 mL) and separated with NaHC03(saturated aqueous solution) (60 mL). The layers were separated and the aqueous layer was extracted with EtOAc (60 mL). The organic layers were combined, dried (MgS04), filtered and evaporated to give the crude product as a yellow oil. The crude product was purified by flash chromatography on silica gel to give a clear yellow oil. Yield: AOM: 637 mg (94%); PrOM: 526 mg (78%); DPrOM: 617 mg (86%).
[0272] 3'-AOM: clear yellow oil. LC-MS (electrospray negative ion) [M-H] 446.12.
[0273] 3'-PrOM: clear yellow oil. (526 mg 78%). LC-MS (electrospray negative ion): [M-H] 444.10.
[0274] 3'-DPrOM: clear yellow oil. (617 mg 86%). LC-MS (electrospray negative ion): [M-H] 472.13.
[0275] In addition, two other 3' blocked T nucleosides (3'-eAOM T and 3'-iAOM T) were prepared in a similar manner to that described above. 3'-iAOM T: LC-MS (ES): (negative ion) m / z 325.5 (M-H + ), (positive ion) 327.3 (M+H + ). 3'-eAOM T: LC-MS (ES): (positive ion) m / z 341.3 (M+1H + ).
[0276]
[0277] Example 2. 3'-OH blocking group stability test
[0278] In this example, stability tests of 5'-mP 3'-AOM T nucleotides were performed in parallel with standard 5'-mP 3'-O-azidomethyl T nucleotides in incorporation buffer solution.
[0279]
[0280] Preparation of buffer solutions
[0281] One mL of 0.1 mM of each 5'-monophosphate 3' protected T nucleotide in a solution of 100 mL of ethanolamine buffer (pH 9.8), 100 mM NaCl, and 2.5 mM EDTA was incubated in a 65 °C heat block for 2 weeks. At set time points, 40 μL aliquots were taken and analyzed by HPLC to determine the percentage of blocked nucleotides remaining and unblocked nucleotides formed.
[0282] The results of stability tests related to 5'-monophosphate 3'-blocked nucleotides with AOM, PrOM, DPrOM acetal protecting groups, and standard azidomethyl blocking groups are shown in Figure 1 The 3' blocked nucleoside monophosphates with AOM, PrOM, and DPrOM blocking groups were observed to provide 30-50 fold or more improvement in solution in reducing deblocking rates. This experiment mimics the performance when the corresponding fully functionalized nucleotides (ffN) are stored in the incorporation mixtures in sequencing instrument cartridges. The stability improvements provided by these acetal protecting groups will also result in lower pre-phasing rates in sequencing runs. Finally, it improves the shelf life of the incorporation mix reagents.
[0283] Example 3. 3'-AOM deblocking test
[0284] In this example, deblocking tests of 5'-mP 3'-AOM T and standard 5'-mP 3'-O-azidomethyl T nucleotides were run separately in unique solutions of each blocking group. Conditions were developed to mimic Illumina's standard deblocking reagents as closely as possible and following the same methodology. In all tests, the concentrations of active deblocking reagents, buffers, and nucleotides were kept the same, but the identity of each component was unique. In this way, rate differences observed between individual deblocking chemistries cannot be attributed to differences in formulation concentrations.
[0285]
[0286] Standard azidomethyl deblocking conditions
[0287] Nucleotide: 5'-monophosphate 3'-O-azidomethyl T. Active Deblocking Blocking Reagent: Tris(hydroxylpropyl)phosphine (THP) (1 M in 18 mΩ water). (Optional) Additive: Sodium ascorbate (0.1 mM in 18 mΩ water). Final concentration = 1 mM. Buffer: Ethanolamine pH 9.8 (2 M in 18 mΩ water). Quench: H2O2.
[0288] AOM deblocking conditions
[0289] Nucleotide: 5'-monophosphate 3'-0-azidomethyl T. In a glass vial under nitrogen, a stock solution of 3'-AOM T was diluted to 0.1 mM in 100 mM ethanolamine buffer (pH 9.8). A stock solution of sodium ascorbate additive was added to a final concentration of 0.1 mM and the solution was stirred for 5 minutes. To start the assay, 40 μL aliquots were taken at the indicated time points and quenched with 6 μL of a 1 :3 mixture of EDTA / H2O2(0.025:0.075 M). HPLC analysis was performed by measuring the area of the starting nucleoside peak, the 3'-OH peak, and any other nucleotide peaks that appeared in the HPLC chromatogram. No other nucleotide-based side products were observed. Deblocking The blocking reagent (Pd / THP = 1 / 5; sodium ascorbate; ethanolamine) was added to the stirred solution at room temperature to a final concentration of 1 mM THP. At the indicated time points, 40 μL aliquots were taken and quenched with 6 μL of a 1 :3 mixture of EDTA / H2O2(0.025:0.075 M). HPLC analysis was performed by measuring the area of the starting nucleoside peak, the 3'-OH peak, and any other nucleotide peaks that appeared in the HPLC chromatogram. No other nucleotide-based side products were observed.
[0290] The results of the comparison are shown in Figure 2A . It was observed that AOM provides a 10-fold rate improvement in deblocking rate in solution compared to the standard azidomethyl blocking group. The purpose of this experiment was to simulate the performance of the corresponding ffN in sequencing in the deblocking step. The significant increase in deblocking rate would allow for the adoption of a flush-through deblocking step in place of the 10- to 20-second incubation time typically used by some Illumina sequencing platforms. Thus, the deblocking rate would have a significant impact on the sequencing-by-synthesis (SBS) cycle time.
[0291] Similar experimental conditions were used for deblocking assays of 3'-eAOM T and 3'-iAOM T. As a separate change, the ratio of Pd catalyst to substrate was reduced to 5:1 in order to observe a smaller difference in deblocking rate. 3'-AOM T was used as a reference and the results are shown in Figure 2B These results indicate that eAOM and iAOM have a 2- to 3-fold slower deblocking rate than AOM under the specific concentration of Pd catalyst deblocking reagent. It can be expected that the difference in deblocking rate between the substituted and unsubstituted versions of the AOM blocking group would be smaller when the ratio of Pd catalyst to substrate is higher.
[0292] Example 4. Optimization of palladium cleavage mix in sequencing
[0293] The Pd / THP catalyst used in the deblocking reaction described in Example 2 is very sensitive to air. When exposed to air, it shows a large loss of activity. In this example, an oxidation stress assay was developed to evaluate the air sensitivity of different preparations of the palladium cleavage cocktail.
[0294] The Pd cleavage cocktail was aliquoted into 5 mL glass vials at 0.5 mL and left open to air at room temperature for 3 hours. The residual activity of the oxidized cleavage cocktail was evaluated by measuring the cleavage of 3'-AOM T as follows. A stock solution of 3'-AOM T was diluted to 0.1 mM in 100 mM cleavage cocktail buffer. A stock solution of sodium ascorbate was added to a final concentration of 1 mM, and the oxidized cleavage cocktail was then diluted to a final concentration of 1 / 20. After 1 hour, 40 μΐ^of the solution was immediately quenched with 10 μΐ^of a 1 : 1 mixture of EDTA / H2O2(0.25:0.25 M) and analyzed by HPLC. In this experiment, various buffer reagents were screened, including: primary amines (e.g., ethanolamine, Tris, and glycine); tertiary amines (e.g., 2-dimethylaminoethanol ((DMEA), 2-diethylaminoethanol (DEEA), N,N,N',N'-tetramethylethylenediamine (TEMED), or N,N,N',N'-tetraethylethylenediamine (TEEDA)); and various inorganic salts (e.g., borates, carbonates, phosphates). It was observed that inorganic buffers (e.g., sodium borate, sodium carbonate, sodium phosphate) provided the best air stability, and the palladium complex retained a high % activity. Additionally, tertiary amines also significantly improved the stability of the Pd cleavage cocktail compared to primary amines.
[0295] Based on these findings, two palladium cleavage mixtures were prepared. In the first example, a stock solution of 250 mM borate buffered aqueous solution was diluted with water (14 mL) (pH 9.6, 20 mL), followed by the addition of THP (1 M in 100 mM Tris, pH 9, 5 mL, 5.0 mmol) and allyl palladium (II) chloride dimer (183 mg, 0.5 mmol). The mixture was stirred vigorously at room temperature for a few minutes, followed by the addition of 1 M aqueous sodium ascorbate (0.5 mL, 0.5 mmol), 5 M aqueous NaCl (10 mL), and 10% v / v Tween 20 (0.5 mL). In the second example, a stock solution of 2 M DEEA buffered aqueous solution (pH 9.6, 0.6 mL) was diluted with water (7.6 mL), followed by the addition of THP (1 M in 100 mM Tris, pH 9, 1.2 mL, 1.2 mmol) and solid allyl palladium (II) chloride dimer (43.9 mg, 0.12 mmol). The mixture was stirred vigorously at room temperature for a few minutes, followed by the addition of 1 M aqueous sodium ascorbate (0.12 mL, 0.12 mmol), 5 M aqueous NaCl (2.4 mL), and 10% v / v Tween 20 (0.12 mL).
[0296] Example 5. Preparation of fully functionalized nucleotides and their use in sequencing applications
[0297] In this example, the preparation of various fully functionalized nucleotides (ffNs) with 3'-AOM blocking groups is described in detail. These ffNs were also used in the sequencing-by-synthesis on the Illumina MiniSeq® platform.
[0298] Scheme 4. Synthesis of 3'-AOM-ffC-LN3-SO7181
[0299]
[0300] Synthesis of Intermediate AOM C2: Nucleoside CI (0.5 g, 0.64 mmol) was dissolved in dry DCM (12 mL) under N2and the mixture was cooled to 0 °C. Cyclohexene (0.32 mL, 3.21 mmol) was added, followed by dropwise addition of SO2CI2(1.0 M in DCM, 1.27 mL, 1.27 mmol). Additional cyclohexene (0.32 mL, 3.21 mmol) was added, then the reaction was quickly transferred to a rotary evaporator to remove all volatiles under reduced pressure. The solid residue was additionally dried under high vacuum for 10 min, then dissolved in dry DCM (5 mL) under N2. The mixture was cooled to 0 °C, and ice-cold allyl alcohol (5 mL) was added dropwise. The reaction was stirred at 0 °C for 2 h, then quenched by addition of saturated aqueous NaHC03(50 mL) and DCM (30 mL). The two phases were separated, and the aqueous layer was extracted with EtOAc (2 x 50 mL). The organic layers were combined, dried over MgS04, filtered, and the volatiles evaporated under reduced pressure. The crude product was purified by flash chromatography on silica gel using EtOAc / pet. ether to give AOM C2 as a white solid (264 mg, 52% yield). LC-MS (electrospray negative ion): [M-H] 787, [M+CI] 823.
[0301] Synthesis of Intermediate AOM C3: AOM C2 (246 mg, 0.31 mmol) was dissolved in dry THF (9.5 mL) under N2and the mixture was cooled to 0 °C. Acetic acid (0.054 mL, 0.94 mmol) was added, followed by dropwise addition of TBAF (1.0 M in THF, 5 wt.% water, 0.99 mL, 0.94 mmol). The reaction was stirred at 0 °C for 5 h, then diluted with EtOAc (20 mL) and poured into 0.05 M aqueous HC1 (20 mL). The two layers were separated, and the aqueous layer was extracted with EtOAc (2 x 20 mL). The organic layers were combined, dried over MgS04, filtered, and the volatiles evaporated under reduced pressure. The crude product was purified by flash chromatography on silica gel using DCM / EtOAc to give AOM C3 as a pale yellow solid (114 mg, 66% yield). LC-MS (electrospray negative ion): [M-H] 549, [M+H20-H] 567, [M+CI] 585, (positive ion electrospray): [M+H] 551, [M+H20+H] 569.
[0302] Synthesis of Intermediate AOM C4: AOM C3 (0.114 g, 0.21 mmol), freshly activated 4A molecular sieves, proton sponge (0.066 g, 0.31 mmol), and a magnetic stirrer were placed under N2and anhydrous trimethylphosphite (1.0 mL) was added. The reaction mixture was cooled to -10 °C and freshly distilled POCl3(23 µL, 0.25 mmol) was added dropwise. The reaction was stirred at -10 °C for 1 h. A solution of di-tri-n-butylammonium pyrophosphate (0.5 M in DMF, 1.7 mL, 0.85 mmol) and anhydrous tri-n-butylamine (0.41 mL, 1.74 mmol) were pre-mixed and added to the ice-cold activated nucleoside solution in one portion. The mixture was stirred vigorously at room temperature for 5 min. The reaction mixture was poured into a separate flask containing vigorously stirred 2 M TEAB aqueous solution (~10 mL). The reaction flask was rinsed with a small amount of H2O and the wash was added to the 2 M TEAB solution. The combined mixture was then stirred at room temperature for 4 h and the solvent was evaporated under reduced pressure. The residue was dissolved in aqueous NH3solution (35%, ~10 mL) and stirred at room temperature overnight. The reaction was concentrated in vacuo and purified by flash chromatography on DEAE-Sephadex. The product was further purified by preparative HPLC to give pure AOM C4 (62 µmol, 30% yield, determined by UV-Vis spectroscopy, λ max = 294 nm, ε = 8600 M -1 cm -1 -1). LC-MS (electrospray negative): [M-H] 589.
[0303] Synthesis of 3’-AOM-ffC-LN3-SO7181: LN3-SO7181 (0.0205 mmol) was dissolved in anhydrous DMA (4 mL) under N2. N,N-diisopropylethylamine (28.6 µL, 0.164 mmol) was added followed by TSTU (0.1 M in DMA, 234 µL, 0.0234 mmol). The reaction was stirred at room temperature under N2for 1 h. Meanwhile, an aqueous solution of AOM C4 (0.0101 mmol) was evaporated to dryness under reduced pressure, re-suspended in 0.1 M TEAB aqueous solution (400 µL) and added to the LN3-SO7181 solution. The reaction was stirred at room temperature for 17.5 h and then quenched with 0.1 M TEAB aqueous solution (4 mL). The crude product was purified by flash chromatography on DEAE-Sephadex. The product was further purified by preparative HPLC to give pure 3’-AOM-ffC-LN3-SO7181 (6.81 µmol, 67% yield, determined by UV-Vis spectroscopy, λ max= 644 nm, ε = 200000 M -1 cm- 1 ). LC-MS (electrospray negative): [M-H] 1561, [M-2H] 781, [M-3H] 520.
[0304] Scheme 5. Synthesis of 3'-AOM-ffA
[0305]
[0306]
[0307]
[0308]
[0309] Synthesis of intermediate AOM A2: Nucleoside A1 (716 mg, 0.95 mmol) was dissolved in 10 mL of dry dichloromethane under N2atmosphere, cyclohexene (481 µL, 4.75 mmol) was added and the solution was cooled to ~ -15°C. Sulfuryl chloride (distilled, 92 µL, 1.14 mmol) was added dropwise and the reaction was stirred for 20 minutes. After all the starting material was consumed, additional cyclohexene (481 µL, 4.75 mmol) was added and the reaction mixture was evaporated to dryness under reduced pressure. The residue was quickly purged with nitrogen and then allyl alcohol (5 mL, ~ 100 mmol) was added with stirring at 0°C. The reaction was stirred at 0°C for 1 hour and then quenched with 50 mL of saturated aqueous NaHCO3solution. The mixture was extracted with 2x100 mL of ethyl acetate. The combined organic phases were washed with 100 mL of water and 100 mL of brine, then dried over MgSO4, filtered and evaporated to dryness. The residue was purified by flash chromatography on silica gel using petroleum ether / EtOAc. 60% yield (435 mg, 0.57 mmol). LC-MS (ES and CI): (positive ion) m / z 763 (M+H + ); (negative ion) m / z 761 (M-H + ).
[0310] Synthesis of intermediate AOM A3: Nucleoside AOM A2 (476 mg, 0.62 mmol) was dissolved in dry THF (5 mL) under N2 atmosphere, then 1.0 M TBAF (750 μL, 0.75 mmol) in THF was added. The solution was stirred at room temperature for 1.5 hours. The solution was diluted with 50 mL EtOAc, then washed with 100 mL NaH2P04saturated solution (pH = 3) and with 100 mL brine. The organic phase was dried over MgS04, filtered and evaporated to dryness. The residue was purified by flash chromatography on silica gel using EtOAc / MeOH. 90% yield (292 mg, 0.55 mmol). LC-MS (ES and CI): (positive ion) m / z 525 (M+H + ); (negative ion) m / z 523 (M-H + ).
[0311] Synthesis of intermediate AOM A4: Nucleoside AOM A3 (285 mg, 0.544 mmol) was dried under reduced pressure with P205for 18 hours. Under nitrogen, dry triethylphosphite (2 mL) and some freshly activated 4A molecular sieves were added to it, then the reaction flask was cooled to 0°C in an ice bath. Freshly distilled POCl3(61 μL, 0.65 mmol) was added dropwise, then Proton Sponge®(175 mg, 0.816 mmol) was added. After the addition was complete, the reaction was further stirred at 0°C for 15 minutes. Then, a 0.5 M solution of di-tri-n-butylammonium pyrophosphate (5.4 mL, 2.72 mmol) in dry DMF was quickly added, followed immediately by tri-n-butylamine (540 μL, 2.3 mmol). The reaction was kept in an ice-water bath for another 10 minutes, then it was quenched by pouring into 1 M aqueous triethylammonium bicarbonate (TEAB, 20 mL) and stirred at room temperature for 4 hours. All solvents were evaporated under reduced pressure. A 35% aqueous ammonia solution (20 mL) was added to the above residue and the mixture was stirred at room temperature for at least 5 hours. Then the solvents were evaporated under reduced pressure. The crude product was first purified by ion exchange chromatography on DEAE-Sephadex A25 (100 g). The column was eluted with a gradient of aqueous triethylammonium bicarbonate. The fractions containing the triphosphate were combined and the solvents were evaporated to dryness under reduced pressure. The crude material was further purified by preparative HPLC using a YMC-Pack-Pro C18 column, eluted with 0.1 M TEAB and acetonitrile. The triethylammonium salt of compound AOM A4 was obtained. 56% yield (306 μmol). LC-MS (ES and CI): (negative ion) m / z 612 (M-H + ); (positive ion) m / z 614 (M+H+ ), 715 (M+Et3NH + ).
[0312] General procedure for ffA synthesis: The dye linker (0.020 mmol) was dissolved in 2 mL of anhydrous N,N’-dimethylacetamide (DMA). N,N’-Diisopropylethylamine (28.4 µL, 0.163 mmol) was added, followed by a 0.1 M solution of anhydrous DMA of N,N,N’,N’-tetramethyl-O-(N-succinimidyl)urea tetrafluoroborate (TSTU, 232 µL, 0.023 mmol). The reaction was stirred at room temperature under nitrogen for 1 hour. Meanwhile, an aqueous solution of AOM A4 triphosphate (0.01 mmol) was evaporated to dryness under reduced pressure and resuspended in 200 µL of a 0.1 M aqueous solution of triethylammonium bicarbonate (TEAB). The activated dye-linker solution was added to the triphosphate and the reaction was stirred at room temperature for 18 hours. The crude product was first purified by ion-exchange chromatography on DEAE-Sephadex A25 (25 g). Fractions containing the triphosphate were pooled and the solvent was evaporated to dryness under reduced pressure. The crude material was further purified by preparative scale RP-HPLC using a YMC-Pack-Pro C18 column. 3’-AOM-ffA-LN3-NR7180A: 38% yield (3.8 µmol). LC-MS (ES): (negative ion) m / z 1459 (M-H + ), 729 (M-2H + ), 486 (M-3H + ). 3’-AOM-ffA-LN3-BL-NR550S0: 37% yield (3.7 µmol). LC-MS (ES): (negative ion) m / z 1771 (M-H + ), 885 (M-2H + ), 589 (M-3H + ). 3’-AOM ffA-LN3-BL-NR 6 50C5: 51% yield (51 µmol). LC-MS (ES): (negative ion) m / z 1917 (M-H + ), 958 (M-2H + ), 645 (M-3H + ).
[0313] Scheme 6. Synthesis of 3'-AOM-pppG
[0314]
[0315] Synthesis of intermediate AOM G4: The known nucleoside dG3 (100 mg, 0.143 mmol) was dissolved in 10 mL of dry dichloromethane under N2atmosphere, cyclohexene (72 µL, 0.714 mmol) was added and the solution was cooled to -12 °C. Sulfuryl chloride (distilled, (1 M in DCM), 171 μL, 0.171 mmol) was added dropwise and the reaction was stirred for 10 minutes. Additional cyclohexene (72 μL, 0.714 mmol) was added and the reaction was stirred at -12 °C for 30 minutes. The reaction was evaporated to dryness under reduced pressure, the residue was purged with nitrogen and ice-cold neat allyl alcohol (distilled, 0.8 mL, 12 mmol) was added with stirring at -12 °C. The reaction was stirred at -12 °C for 60 minutes and then quenched with 2 mL of saturated aqueous NaHCO3solution. The mixture was separated with ethyl acetate (2 mL) and the aqueous layer was extracted with ethyl acetate. The combined organic phases were washed with 4 mL of water and 4 mL of brine, dried over MgSO4, filtered and evaporated to a crude oil. The residue was purified by flash chromatography on silica gel to give AOM G4 as a clear oil. 36% yield (50.9 mg, 0.072 mmol). LC-MS (ES and CI): (positive ion) m / z 710 [M+H] + ; (negative ion) m / z 708 [M-H] - .
[0316] Synthesis of intermediate AOM G5: The nucleoside AOM-G4 (111 mg, 0.156 mmol) was dissolved in dry THF (5 mL) under N2atmosphere. Acetic acid (27 µL, 0.468 mmol) was added followed by 1.0 M TBAF (296 µL, 0.296 mmol) in THF. The solution was stirred at room temperature for 5 hours. The solution was diluted with 10 mL of EtOAc and washed with 10 mL of 0.05 M aqueous HC1, the organic phase was separated. The aqueous phase was extracted with ethyl acetate. The combined organic phases were dried over MgSO4, filtered and evaporated to dryness. The residue was purified by flash chromatography on silica gel to give AOM G5 as a white solid. 44% yield (32.4 mg, 0.068 mmol). LC-MS (ES and CI): (positive ion) m / z 472 [M+H] + ; (negative ion) m / z 470 [M-H] - .
[0317] Synthesis of 3'-AOM-pppG: Nucleoside AOM-G5 (79 mg, 0.168 mmol) and fresh activated 4 A molecular sieves were dried under reduced pressure with P2O5for 18 h. Proton Sponge®(175 mg, 0.816 mmol) and anhydrous triethyl phosphite (0.8 mL) were added under nitrogen and stirred at room temperature for 1 h. The reaction flask was cooled to 0 °C in an ice bath and freshly distilled POCl3(19 µL, 0.202 mmol) was added dropwise and the reaction stirred at 0 °C for 15 min. Then, a solution of di-tri-n-butylammonium pyrophosphate (1.68 mL, 0.84 mmol) in anhydrous DMF (0.5 M) was added rapidly, followed immediately by tri-n-butylamine (168 µL, 0.705 mmol). The reaction was removed from the ice / water bath and stirred vigorously for 5 min, then quenched by pouring into 1 M aqueous triethylammonium bicarbonate (TEAB, 6 mL) and stirred at room temperature for 18 h. All solvents were evaporated under reduced pressure. The residue was dissolved in 35% aqueous ammonia solution (10 mL) and stirred at room temperature for at least 5 h. The solvent was then evaporated under reduced pressure and further co-evaporated with water. The crude product was first purified by ion exchange chromatography on DEAE-Sephadex A25 (50 g). The column was eluted with a linear gradient of aqueous triethylammonium. Fractions containing the triphosphate were collected and the solvent was evaporated to dryness under reduced pressure. The crude material was further purified by preparative scale HPLC using a YMC-Pack-Pro C18 column. The triethylammonium salt of 3'-AOM-pppG was obtained. 24% yield (39.7 µmol). LC-MS (ES and CI): (negative ion) m / z 576 [M-H] - ; (positive ion) m / z 578 [M+H] + .
[0318] 3'-AOM-ffT-LN3-NR550S0 was synthesized in a similar manner to that described in the preparation of 3'-AOM ffA and ffC.
[0319] Scheme 7. Synthesis of 3'-AOM-ffT-LN3'-NR550S0
[0320]
[0321] Synthesis of Intermediate T1 : 5-Iodo-2'-deoxyuridine (3 g, 8.4 mmol) and palladium (II) acetate (1.6 g, 7.14 mmol) were dissolved in dry degassed DMF, followed by the addition of N-allyl trifluoroacetamide (6.4 mL, 42 mmol). The solution was placed under vacuum, then purged with nitrogen 3 times, followed by the addition of degassed triethylamine (2.3 mL, 16.8 mmol). The solution was heated to 80 °C for 2 hours. The black mixture was cooled to room temperature, then diluted with 50 mL of methanol. About 0.5 g of activated charcoal was added, and the solution was filtered over celite, then evaporated under reduced pressure to give a brown thick oil. The crude product was purified by chromatography on silica gel using EtOAc / MeOH elution. Yield: (2.27 g, 5.99 mmol). LC-MS (ES and CI): (negative ion) m / z 378 (M-H + ) at 2.1 min.
[0322] Synthesis of Intermediate T2: 5-[3-(2,2,2-trifluoroacetamido)-allyl]-2'- deoxyuridine (T1 ) (2.55 g, 6.72 mmol) was dissolved in dry DMF. Imidazole (1.37 g, 20.1 mmol) was added, followed by the addition of 4-(dimethylamino)pyridine (410 mg, 3.36 mmol). The reaction was cooled to 0 °C, then tert-butyl(chloro)diphenylsilane (1.92 mL, 7.39 mmol) was added slowly in three portions at 30 min intervals. The reaction was stirred at 0 °C for 6 hours. The solvent was then evaporated, and the residue was re-suspended in 200 mL of EtOAc, and washed with 2 x 200 mL of saturated aqueous NaHC03solution and 200 mL of water, then 100 mL of brine. The organic phase was dried over MgS04, filtered and evaporated to dryness. The crude product was purified by flash chromatography on silica gel using DCM / EtOAc. 68% yield (2.806 g, 4.54 mmol). LC-MS (ES and CI): (positive ion) m / z 618 (M+H + ); (negative ion) m / z 616 (M-H + ) at 2.1 min.
[0323] Synthesis of Intermediate T3: 5'-O-(tert-butyldiphenylsilyl)-5-[3-(2,2,2- trifluoroacetamido)-allyl]-2'-deoxyuridine (T2) (2.8 g, 4.53 mmol) was dissolved in 10 mL of anhydrous DMSO (136 mmol) followed by the addition of glacial acetic acid (16 mL, 272 mmol) and acetic anhydride (16 mL, 158 mmol). The reaction was heated to 50 °C for 6 hours, then quenched with 200 mL of saturated aqueous NaHCO3solution. After the solution stopped bubbling, it was extracted with 2x 150 mL EtOAc. The organic phases were combined and washed with 2x 200 mL of saturated aqueous NaHCO3solution, 200 mL water, and 100 mL brine. The organic phase was dried over MgSO4, filtered, and evaporated to dryness. The crude product was purified by flash chromatography on silica gel using DCM / EtOAc. 77% yield (2.375 g, 3.51 mmol). LC-MS (ES and CI): (positive ion) m / z 678 (M+H + ); (negative ion) m / z 676 (M-H + ).
[0324] Synthesis of Intermediate T4: 5'-O-(tert-butyldiphenylsilyl)-3'-O-methylmethylthio-5- [3-(2,2,2-trifluoroacetamido)-allyl]-2'-deoxyuridine (T3) (310 mg, 0.45 mmol) was dissolved in 5 mL of anhydrous dichloromethane under a N2atmosphere, cyclohexene (228 μL, 2.25 mmol) was added, and the solution was cooled to approximately -15 °C. Sulfuryl chloride (distilled, 55 μL, 0.675 mmol) was added dropwise, and the reaction was stirred for 20 minutes. After all starting material was consumed, additional cyclohexene (228 μL, 2.25 mmol) was added, and the reaction was evaporated to dryness under reduced pressure. The residue was flash-purged with nitrogen, then ice-cold allyl alcohol (2.5 mL) was added with stirring at 0 °C. The reaction was stirred at 0 °C for 35 minutes, then quenched with 25 mL of saturated aqueous NaHCO3solution, then further diluted with 100 mL of saturated aqueous NaHCO3solution. The mixture was extracted with 2x 50 mL ethyl acetate. The combined organic phases were dried over MgSO4, filtered, and evaporated to dryness. The residue was purified by flash chromatography on silica gel using DCM / EtOAc. 69% yield (214 mg, 0.311 mmol). LC-MS (ES and CI): (positive ion) m / z 688 (M+H + ); (negative ion) m / z 686 (M-H + ).
[0325] Synthesis of Intermediate T5: 5'-O-(tert-butyldiphenylsilyl)-3'-O-allyloxymethyl-5-[3-(2,2,2-trifluoroacetamido)-allyl]-2'-deoxyuridine (T4) (210 mg, 0.305 mmol) was dissolved in dry THF (3 mL) under N2atmosphere. 1.0 M TBAF in THF (367 μL, 0.367 mmol) was added. The solution was stirred at room temperature for 3 hours. The solution was diluted with 50 mL EtOAc and washed with 50 mL saturated NaH2PO4(pH=3), 50 mL water. The organic phase was dried over MgSO4, filtered and evaporated to dryness. The residue was purified by flash chromatography on silica gel using DCM / EtOAc. 95% yield (130 mg, 0.289 mmol). LC-MS (ES and CI): (negative ion) m / z 448 (M-H + ), 484 (M+Cl - ).
[0326] Synthesis of Intermediate T6: 3'-O-allyloxymethyl-5-[3-(2,2,2- trifluoroacetamido)-allyl]-2'-deoxyuridine (T5) (120 mg, 0.267 mmol) was dried under reduced pressure with P2O5for 18 hours. To this was added under nitrogen anhydrous triethyl phosphite (1 mL) and some freshly activated 4A molecular sieves, then the reaction flask was cooled to 0 °C. Freshly distilled POCl3(30 μL, 0.32 mmol) was added dropwise, followed by Proton Sponge®(85 mg, 0.40 mmol). After the addition was complete, the reaction was further stirred at 0 °C for 15 minutes. Then, a solution of di-tri-n-butylammonium pyrophosphate (2.7 mL, 1.33 mmol) in anhydrous DMF was added rapidly, followed immediately by tri-n-butylamine (270 μL, 1.2 mmol). The reaction was kept in an ice-water bath for another 10 minutes, then it was quenched by pouring into 1 M aqueous triethylammonium bicarbonate (TEAB, 10 mL) and stirred at room temperature for 4 hours. All solvents were evaporated under reduced pressure. A 35% aqueous ammonia solution (10 mL) was added to the above residue and the mixture was stirred at room temperature for 18 hours. Then the solvent was evaporated under reduced pressure, the residue was re-suspended in 10 mL of 0.1 M TEAB and filtered. The filtrate was first purified by ion exchange chromatography on DEAE-Sephadex A25 (100 g). The column was eluted with triethylammonium bicarbonate aqueous solution (TEAB). Fractions containing triphosphates were combined and the solvent was evaporated to dryness under reduced pressure. The crude material was further purified by preparative scale HPLC using a YMC-Pack-Pro C18 column. Compound T6 was obtained as a triethylammonium salt. 33% yield (89 μmol). LC-MS (ES and CI): (negative ion) m / z 592 (M-H + ), 295 (M-2H + ).
[0327] Synthesis of 3'-AOM-ffT-LN3'-NR550S0: Dry known compound LN3-NR550S0 (0.015 mmol) was dissolved in dry DMA (2 mL) under N2. N,N-diisopropylethylamine (17 μL, 0.1 mmol) was added, followed by TSTU (0.1 M in DMA, 180 μL, 0.018 mmol). The reaction was stirred at room temperature for 1 hour under N2. Meanwhile, T6 aqueous solution (0.01 mmol) was evaporated to dryness under reduced pressure, resuspended in 0.1 M TEAB aqueous solution (200 μL) and added to the LN3-NR550S0 solution. The reaction was stirred at room temperature for 18 hours, then quenched with 0.1 M TEAB aqueous solution (4 mL). The crude product was purified by flash chromatography on DEAE-Sephadex. The product was further purified by preparative HPLC to give pure 3'-AOM-ffT-LN3'-NR550S0. 67% yield (41 μmol, determined by UV-Vis spectroscopy, λ = 550 nm, ε = 125000 M max cm -1 -1 LC-MS (ES): (negative ion) m / z 1521 (M-H + ), 761 (M-2H + ), 507 (M-3H + ).
[0328] Sequencing by synthesis experiments
[0329] Subsequent sequencing experiments were performed with ffN using the Illumina MiniSeq® instrument. All standard commercial reagents were used except for the new incorporation mix including these ffN. The standard 2x150 recipe was used. In addition to the standard sequencing by synthesis (SBS) protocol, a 5 seconds incubation in the palladium cleavage mix (Pd:THP = 1 / 5 in DEEA as described in example 4) was added to unblock the 3'-AOM.
[0330] In a first experiment, the following ffN were used in the incorporation mix: 3'-AOM-ffT-LN3-NR550S0, 3'-AOM-ffA-LN3-BL-NR550S0, 3'-AOM-ffA-LN3-BL-NR650C5, 3'-AOM-ffA-LN3-NR7180A, 3'-AOM-ffC-LN3-SO7181 and 3'-AOM-pppG (dark G). The sequencing results for read 1 are summarized as follows.
[0331]
[0332] PF: Percent of clusters passing through filter after 26 cycles
[0333] In a second experiment, unlabeled 3'-AOM-pppT was synthesized in a similar manner to the preparation of 3'-AOM-pppG described above (LC-MS (ES): (negative ion) m / z 551 (M-H + )). It was used for sequencing in the presence of the commercial green ffG-LN3-PEG12-ATTO532 (used on Illumina 4-channel system) and the same ffA and ffC as described in the first experiment above. The results are summarized below. The phasing and pre-phasing values were significantly improved and no signal decay was observed Figure 3A ). In addition, the error rates for read 1 and read 2 were both reduced as well.
[0334]
[0335]
[0336] In another experiment, the incorporation mix contained 3'-AOM-ffT-LN3'-NR550S0, 3'-AOM-ffA-LN3-BL-NR550S0, 3'-AOM-ffA-LN3-BL-NR650C5, 3'-AOM-ffA-LN3-NR7180A, 3'-AOM-ffC-LN3-SO7181, and 3'-AOM-pppG (dark G). Similar to previous runs, a 5 minute incubation of the cleavage mix containing palladium catalyst (Pd / THP = 1 : 10; 100 mM DEEA described in Example 4) was added to the standard SBS cycles. The standard MiniSeq® DNA polymerase was used, but the incorporation time was doubled. No signal decay phenotype was observed Figure 3B ). In addition, these sequencing results were compared to a commercial MiniSeq® run with ffNs bearing the standard azidomethyl blocking group (average of 3; N = 3). The error rates were observed to be nearly identical Figure 3C ). The sequencing results are summarized below.
[0337]
[0338] In addition, the main sequencing metrics for ffNs with 3'-AOM blocking groups were compared to the main sequencing metrics produced by the standard MiniSeq® commercial kit including DNA polymerase Pol 812, and the results are shown in Figure 4A . Due to the improved stability of the 3'-AOM-ffNs, very low pre-phasing was observed. However, even with a 2-fold incorporation time, the phasing was still improved.
[0339] In another experiment, a different DNA polymerase (Pol 1901) was used instead of the DNA polymerase in the commercial kit (Pol 812). Pol 1901 allowed for a standard 1x incorporation time during sequencing, instead of the aforementioned 2x incorporation time. Furthermore, the incubation time in the Pd fragmentation mixture was reduced by half compared to the standard run. This resulted in a 10% time saving throughout the SBS chemical cycle. Sequencing metrics were significantly improved and exceeded values obtained from the standard commercial kit containing a 3'-O-azidomethyl blocking group. Figure 4B ).
[0340] 3' blocking group stability test in sequencing
[0341] To demonstrate the improved stability of ffNs with 3'-AOM, they were compared in parallel with standard MiniSeq® ffNs containing 3'-O-azidomethyl groups. Both groups of ffNs were incubated for several days at 45°C in a standard incorporation mixture formulation (excluding DNA polymerase only). For each time point, fresh polymerase was added directly to complete the incorporation mixture before loading into MiniSeq®. Sequencing conditions as previously described were used. The predetermined phase % is a direct indicator of the percentage of 3'OH-ffN in the mixture and is therefore directly related to the stability of the 3' blocking group. The predetermined phase values for both groups of ffNs were recorded and plotted. Figure 5 At 45 °C, ffN containing 3'-AOM appeared to be 6 times more stable than standard ffN with a 3'-O-azidomethyl group. Sequencing parameters also confirmed the trend observed during stability assays in solution – the 3'-AOM blocking group was significantly more stable than the 3'-O-azidomethyl group.
[0342] Example 6. Preparation of 3'-O-thioformate blocked nucleosides
[0343] In this embodiment, various 3'-O-thiocarbamate-protected T-nucleotides were prepared according to scheme 8.
[0344] Scheme 8. Synthesis of 3'-O-dimethylthioformate T nucleoside
[0345]
[0346] Preparation of T-7: To an oven-dried 100 mL flask purged and maintained with nitrogen, 5'-O-(4,4'-dimethoxytrityl) thymidine (1.0 g, 1.836 mmol) was added. This was co-evaporated with anhydrous DMF (3 x 20 mL) and placed under nitrogen. Anhydrous DCM (9.2 mL) and 4-dimethylaminopyridine (224 mg, 0.184 mmol) were added and stirred at room temperature until a homogenous solution was formed. Then 1,1'-thiocarbonyldiimidazole (360 mg, 2.02 mmol) was quickly added under a stream of nitrogen, the reaction was resealed and stirred at room temperature for 2 hours until all starting material was consumed. The reaction mixture was filtered through a pad of silica gel and the filter cake was washed with EtOAc (10 mL). The volatiles were removed in vacuo and the crude residue was used without further purification.
[0347] Preparation of T-8: The compound T-7 from the previous step was used immediately after being dried in vacuo. The residue was placed in a 25 mL round bottom flask under nitrogen and dimethylamine (2M in THF, 7.3 mL, 14.6 mmol) was added and the reaction was stirred for 2 hours until all starting material was consumed according to TLC. All volatiles were removed in vacuo to form a clear crude residue which was purified by flash column chromatography on silica gel to give T-8 as a white solid. Yield: 1.15 g (99%). LC-MS (electrospray negative) 630.23 [M-H]-
[0348] Preparation of T-9: The starting nucleoside T-8 (320 mg, 0.504 mmol) was dissolved in a small amount of acetonitrile in an air open 50 mL round bottom flask. A 5:1 solution of AcOH / H2O (12.5 mL:2.5 mL) was added in one portion and the reaction was stirred at room temperature until all starting material was consumed (2-4 hours). All volatiles were evaporated under vacuum and the residue was co-evaporated in toluene (2 x 60 mL) to give the crude product as an off-white solid. The crude product was purified by flash column chromatography to give T-9 as a white solid. Yield: 123 mg (74%). LC-MS (electrospray negative) [M-H] 328.10.
[0349] Nucleosides with two other thio-carbamate protecting groups were also prepared following similar synthetic procedures using the corresponding MeNH2or NH3. The general reaction scheme is shown below:
[0350]
[0351] 3'-O-thioformate blocking group stability test
[0352] Stability experiments of 5'-mP 3'-DMTCT nucleotides were performed in parallel with standard 5'-mP 3'-O-azidomethyl T nucleotides incorporated in a buffer.
[0353]
[0354] For both 5'-mP 3'-DMTC T and 5'-mP 3'-O-azidomethyl T, the final solution volume was 1 mL and the final concentration of the respective nucleotide was 0.1 mM. The other components of the aqueous buffer included ethanolamine (EA), ethanolamine HC1, NaCl (100 mM), and EDTA (2.5 mM). The concentrations of the buffer solution were as follows: 0.5 M EA buffer, 0.5 M NaCl, 0.01 M EDTA.
[0355] Stability test methodology
[0356] Two hundred microliters of 10x buffer solution was added to a 1.7 mL polypropylene snap lock microtube and diluted with an accurate volume of 18 mW water. The respective nucleotide was then added, the vial was sealed and mixed by inversion, gentle agitation, or pumping with a micropipette. A 40 microliter aliquot was taken and analyzed by HPLC as a starting (or t=0) value. The vial was then placed in a preheated heating block set to 65 °C, covered with a thick layer of aluminum foil, and heated for one month. Periodically (week 1: once per day. Weeks 2-4: 1 time every 2 days) a 40 microliter aliquot was taken and analyzed by HPLC to determine the percentage of starting material and the percentage of de-blocked (3'-OH) nucleotide in the sample. The HPLC analysis was performed by measuring the area of the starting nucleotide peak and the 3'OH peak. These values were used to calculate the percentage of de-blocked nucleotide, which was graphed and used to compare the stability between samples in the incorporation buffer. Figure 6 The results of the comparison of the stability of three different thio-carbamate 3' blocked nucleotides with the nucleotide blocked by the 3'-O-azidomethyl blocking group at 65 °C are shown. It was observed that while the nucleotides with 3'-0-C(=S)NH2 or 3'-0-C(=S)NHCH3 were less stable than the standard 3'-O-azidomethyl group protected nucleotide, the nucleotide with 3'-DMTC conferred improved stability over the 9 day experiment. Thus, DMTC demonstrated improved stability over the standard azidomethyl blocking group.
[0357] Example 7. 3'-O-thioformate blocking group deprotection test
[0358] In this example, deblocking or deblocking trials of 5'-mP 3'-DMTC T and standard 5'-mP 3'-O-azidomethyl T nucleotides were performed separately in unique solutions of each blocking group. Conditions were developed to mimic Illumina's standard deblocking reagents as closely as possible and following the same methodology. In all trials, the concentrations of active deblocking reagents, buffers, and nucleosides were kept the same, but the identity of each component was unique. In this way, rate differences between the individual deblocking chemistries observed could not be attributed to differences in formulation concentrations.
[0359]
[0360] Deblocking test general methodology
[0361] Each reaction component was formulated as a concentrated stock in 18 mΩ water, stored appropriately, and aliquots were combined in the following specific order. Reactions were initiated by the addition of pre-formulated deblocking reagents. Final concentrations: nucleoside (0.1 mM), active deblocking reagent (1 mM), additive (unique to deblocking reagent), buffer (100 mM). Final volume: 2000 µL.
[0362] In a 3 mL glass vial, the pre-formulated buffer solution was added, followed by the pre-formulated additive solution. The appropriate volume of 18 mΩ water was added and stirred for 10 minutes. An aliquot of nucleotide solution was then added and stirred for 5 minutes. A 40 µL aliquot was then removed, quenched with the appropriate quenching agent, and analyzed by HPLC as the reference peak (or t = 0 minutes). The deblocking reagent was then added to the stirred solution in a single addition and timing was initiated. At the specified time points, 40 µL aliquots were removed and immediately quenched with the appropriate quenching agent and then analyzed by HPLC to determine the amount of deblocked nucleotide that had occurred at these specified time points. The results were graphically plotted and used to compare deblocking efficiency and efficacy.
[0363] DMTC deblocking
[0364] Nucleotide: 5'-mP 3'-DMTCT. Active deblocking reagent: NaIO4 (0.1 M in 18 mΩ water) or Oxone® (0.1 M in 18 mΩ water). Additive: none. NaIO4 buffer: pH 6.75 phosphate buffer (1 M in 18 mΩ water). Oxone® buffer: pH 8.65 phosphate buffer (1 M in 18 mΩ water). Quenching agent: sodium thiosulfate. 3'-O-azidomethyl deblocking conditions were the same as described in Example 3.
[0365] HPLC analysis was performed by measuring the area of the starting nucleoside peak, the 3'-OH peak, and any other nucleotide peaks present in the HPLC chromatogram. These values were used to calculate the percentage of starting nucleotide and deblocked nucleotide, displayed graphically, and used to compare deblocking rates, efficiencies, and potency between samples. Comparison results are shown in Figure 7 It was observed that deblocking of DMTC with NaIO4was ineffective. However, when DMTC was cleaved with Oxone®, the percentage of starting material remaining for nucleotides with the DMTC blocking group was significantly less. Overall, it was demonstrated that the deblocking rate of DMTC (using Oxone®) was superior to that of the standard azidomethyl blocking group.
[0366] The present application is particularly directed to the following embodiments:
[0367] 1. A ribose or deoxyribose containing nucleoside or nucleotide having a removable 3'-OH blocking group, the removable 3'-OH blocking group forming a structure covalently linked to the 3'-carbon atom wherein:
[0368] each R 1a and R 1b is independently H, C1-C6 alkyl, C1-C6 haloalkyl, C1-C6 alkoxy, C1-C6 haloalkoxy, cyano, halogen, optionally substituted phenyl, or optionally substituted aralkyl;
[0369] each R 2a and R 2b is independently H, C1-C6 alkyl, C1-C6 haloalkyl, cyano, or halogen;
[0370] alternatively, R 1a and R 2a together with the atom to which they are attached form an optionally substituted five- to eight-membered heterocyclyl;
[0371] R 3 is H, optionally substituted C 2- C6 alkenyl, optionally substituted C 3- C7 cycloalkenyl, optionally substituted C 2- C6 alkynyl, or optionally substituted (C1-C6 alkylene)Si(R 4 )3; and
[0372] each R 4 is independently H, C1-C6 alkyl, or optionally substituted C 6- C 10 aryl; provided that when each R 1a , R 1b , R 2a , and R 2b is H, then R3 is not H.
[0373] 2. The nucleoside or nucleotide of claim 1, wherein each R 1a and R 1b is H.
[0374] 3. The nucleoside or nucleotide of claim 2, wherein each R 1a and R 1b is H.
[0375] 4. The nucleoside or nucleotide of any one of claims 1-3, wherein each R 2a and R 2b is independently H, halo, or Ci-C6alkyl.
[0376] 5. The nucleoside or nucleotide of claim 4, wherein each R 2a and R 2b is H.
[0377] 6. The nucleoside or nucleotide of claim 4, wherein each R 2a and R 2b is independently Ci-C6alkyl or halo.
[0378] 7. The nucleoside or nucleotide of claim 6, wherein each R 2a and R 2b is methyl.
[0379] 8. The nucleoside or nucleotide of claim 4, wherein R 2a is H, and R 2b is halo or Ci-C6alkyl.
[0380] 9. The nucleoside or nucleotide of any one of claims 1 to 7, wherein R 3 is C 2- C6alkynyl optionally substituted with one or more substituents independently selected from the group consisting of halo, Ci-C6alkyl, Ci-C6haloalkyl, and combinations thereof.
[0381] 10. The nucleoside or nucleotide of claim 9, wherein R 3 is .
[0382] 11. The nucleoside or nucleotide of any one of claims 1 to 7, wherein R 3 is C 2- C6alkenyl optionally substituted with one or more substituents independently selected from the group consisting of halo, Ci-C6alkyl, Ci-C6haloalkyl, and combinations thereof.
[0383] 12. The nucleoside or nucleotide of claim 11, wherein R 3 is 、 、.
[0384] 13. The nucleoside or nucleotide of any one of claims 1 to 7, wherein R 3 is optionally substituted (Ci-C6alkylene)Si(R 4 )3 and wherein each R 4 is Ci-C6alkyl.
[0385] 14. The nucleoside or nucleotide of claim 13, wherein R 3 is -(CH2)-SiMe3.
[0386] 15. The nucleoside or nucleotide of claim 1, wherein R 1a and R 2a , together with the atoms to which they are attached, form a six-membered heterocyclyl group.
[0387] 16. The nucleoside or nucleotide of claim 15, wherein the six-membered heterocyclyl group has the structure .
[0388] 17. The nucleoside or nucleotide of claim 15 or 16, wherein each R 1b , R 2b and R 3 is H.
[0389] 18. The nucleoside or nucleotide of claim 1, wherein the 3'-OH blocking group comprises a structure selected from the group consisting of:
[0390] covalently attached to the 3'-carbon of the ribose or deoxyribose sugar 、 、 .
[0391] 19. A nucleoside or nucleotide comprising a ribose or deoxyribose sugar having a removable 3'-OH blocking group, the removable 3'-OH blocking group forming a structure covalently attached to the 3'-carbon atom , wherein:
[0392] each R 5 and R 6 is independently H, Ci-C6alkyl, C 2- C6alkenyl, C 2- C6alkynyl, Ci-C6haloalkyl, C 2- C8alkoxyalkyl, optionally substituted -(CH2) m -phenyl, optionally substituted -(CH2) n -(5 or 6-membered heteroaryl), optionally substituted -(CH2) k -C 3-C7carbocyclyl or optionally substituted -(CH2) p -(3- to 7-membered heterocyclyl);
[0393] each -(CH2) m -, -(CH2) n -, -(CH2) k -, and -(CH2) p is optionally substituted; and
[0394] each m, n, k, and p is independently 0, 1, 2, 3, or 4.
[0395] 20. The nucleoside or nucleotide of claim 19, wherein R 5 and R 6 are each H or C1-C6 alkyl.
[0396] 21. The nucleoside or nucleotide of claim 20, wherein each R 5 and R 6 is H.
[0397] 22. The nucleoside or nucleotide of claim 20, wherein R 5 is H and R 6 is C1-C6 alkyl, C 2- C6 alkenyl, C 2- C6 alkynyl, optionally substituted -(CH2) m -phenyl, or optionally substituted -(CH2) n -6-membered heteroaryl, and wherein m and n are each 0 or 1.
[0398] 23. The nucleoside or nucleotide of claim 20, wherein each R 5 and R 6 is C1-C6 alkyl.
[0399] 24. The nucleoside or nucleotide of claim 23, wherein each R 5 and R 6 is methyl.
[0400] 25. The nucleoside or nucleotide of claim 19, wherein the 3'-OH blocking group comprises a structure selected from the group consisting of: .
[0401] 26. The nucleoside or nucleotide of any one of claims 1 to 25, wherein the nucleoside or nucleotide is covalently linked to a detectable label, optionally via a cleavable linking group.
[0402] 27. The nucleoside or nucleotide of claim 26, wherein the detectable label is covalently linked to the nucleoside base of the nucleoside or nucleotide via a cleavable linking group.
[0403] 28. The nucleoside or nucleotide of claim 26, wherein the detectable label is covalently linked to the 3'-oxygen of the nucleoside or nucleotide via a cleavable linking group.
[0404] 29. The nucleoside or nucleotide of claim 27 or 28, wherein the linking group is a cleavable linking group comprising an azido moiety, a -O-allyl moiety, a disulfide moiety, an acetal moiety, or a thiocarbamate moiety.
[0405] 30. The nucleoside or nucleotide of any one of claims 27 to 29, wherein the 3'-OH blocking group and the cleavable linking group can be removed under the same chemical reaction conditions.
[0406] 31. The nucleoside or nucleotide of any one of claims 1 to 30, comprising a 2'-deoxyribose.
[0407] 32. The nucleoside or nucleotide of claim 31, wherein the nucleotide is a nucleoside triphosphate.
[0408] 33. An oligonucleotide comprising the nucleotide of any one of claims 1 to 32.
[0409] 34. A method of making a growing polynucleotide complementary to a target single-stranded polynucleotide in a sequencing reaction, comprising incorporating the nucleotide of any one of claims 1 to 32 into the growing complementary polynucleotide, wherein the incorporation of the nucleotide prevents any subsequent nucleotide from being introduced into the growing complementary polynucleotide.
[0410] 35. The method of claim 34, wherein the incorporation of the nucleotide is accomplished by a polymerase, terminal deoxynucleotidyl transferase, or reverse transcriptase.
[0411] 36. A method of determining the sequence of a target single-stranded polynucleotide, comprising:
[0412] (a) incorporating the nucleotide of any one of claims 26 to 32 into a replicated polynucleotide strand complementary to at least a portion of the target polynucleotide strand;
[0413] (b) detecting the identity of the nucleotide incorporated into the replicated polynucleotide strand; and
[0414] (c) chemically removing the label and the 3' blocking group from the nucleotide incorporated into the replicated polynucleotide strand.
[0415] 37. The method of claim 36, further comprising (d) washing the chemically removed label and 3' blocking group from the replicated polynucleotide strand.
[0416] 38. The method of claim 37, further comprising repeating steps (a) through (d) until the sequence of the portion of the template polynucleotide strand is determined.
[0417] 39. The method of claim 37, wherein steps (a) through (d) are repeated at least 50 times.
[0418] 40. The method of any one of claims 36 to 39, wherein the label and 3' blocking group are removed from the nucleotide incorporated into the replicated polynucleotide strand in a single chemical reaction.
[0419] 41. The method of claim 40, wherein step (c) comprises contacting the incorporated nucleotide with a cleavage solution comprising a palladium catalyst.
[0420] 42. The method of any one of claims 36 to 39, wherein the label and 3' blocking group are removed from the nucleotide incorporated into the replicated polynucleotide strand in two separate chemical reactions.
[0421] 43. The method of claim 41, wherein step (c) comprises contacting the incorporated nucleotide with a cleavage solution comprising a phosphine and a cleavage solution comprising a palladium catalyst.
[0422] 44. The method of claim 41 or 43, wherein the phosphine is tris(hydroxymethyl)phosphine, tris(hydroxyethyl)phosphine, or tris(hydroxypropyl)phosphine.
[0423] 45. The method of any one of claims 41, 43, or 44, wherein the cleavage solution comprising a palladium catalyst further comprises one or more buffers selected from the group consisting of primary amines, secondary amines, tertiary amines, carbonates, phosphates, and borates, and combinations thereof.
[0424] 46. The method of claim 45, wherein the buffer is selected from the group consisting of ethanolamine (EA), tris(hydroxymethyl)aminomethane (Tris), glycine, carbonates, phosphates, borates, 2-dimethylaminomethanol (DMEA), 2-diethylaminomethanol (DEEA), N,N,N',N'-tetramethylethylenediamine (TEMED), and N,N,N',N'-tetraethylethylenediamine (TEEDA), and combinations thereof.
[0425] 47. A kit comprising one or more nucleosides or nucleotides of any one of claims 1 to 32.
[0426] 48. The kit of claim 47, further comprising an enzyme and a buffer suitable for action of the enzyme.
[0427] 49. The kit of claim 48, wherein the enzyme is a polymerase, a terminal deoxynucleotidyl transferase, or a reverse transcriptase.
[0428] 50. The kit of claim 49, wherein the polymerase is a DNA polymerase.
Claims
1. A nucleoside or nucleotide containing ribose or deoxyribose, having a removable 3′-OH blocking group that forms a structure covalently linked to a 3′-carbon atom. ,in: Each R 1a and R 1b Independently, it is H, C1-C6 alkyl, C1-C6 haloalkyl, C1-C6 alkoxy, C1-C6 haloalkoxy, cyano, halogen, optionally substituted phenyl or optionally substituted aralkyl; Each R 2a and R 2b It is independently H, C1-C6 alkyl, C1-C6 haloalkyl, cyano or halogen; Or, R 1a and R 2a Together with the atoms to which they are attached, they form optional substituted five- to eight-membered heterocyclic groups; R 3 H, C with optional substitution 2- C6 alkenyl, optionally substituted C 3- C7 cycloalkenyl, optionally substituted C 2- C6 ynyl or optionally substituted (C1-C6 alkylene)Si(R) 4 )3; and Each R 4 Independently H, C1-C6 alkyl or optionally substituted C 6- C 10 Aryl; the condition is that when each R 1a R 1b R 2a and R 2b If H is the value of R, then R 3 Not H.
2. The nucleoside or nucleotide of claim 1, wherein R 1a and R 1b At least one of them is H.
3. The nucleoside or nucleotide of claim 2, wherein each R 1a and R 1b For H.
4. The nucleoside or nucleotide of any one of claims 1-3, wherein each R 2a and R 2b It can be independently H, halogen, or C1-C6 alkyl.
5. The nucleoside or nucleotide of claim 4, wherein each R 2a and R 2b For H.
6. The nucleoside or nucleotide of claim 4, wherein each R 2a and R 2b It is independently a C1-C6 alkyl or halogen.
7. The nucleoside or nucleotide of claim 6, wherein each R 2a and R 2b It is a methyl group.
8. The nucleoside or nucleotide of claim 4, wherein R 2a For H, and R 2b It is a halogen or a C1-C6 alkyl group.
9. The nucleoside or nucleotide of any one of claims 1 to 7, wherein R 3 C is optionally substituted by one or more substituents independently selected from the following 2- C6 alkynyl: halogen, C1-C6 alkyl, C1-C6 haloalkyl and combinations thereof.
10. The nucleoside or nucleotide of claim 9, wherein R 3 for .
Citation Information
Patent Citations
Modified nucleic acid probes
EP0742287A2
Methods and compositions for selecting tag nucleic acids and probe arrays
EP0799897A1
Polymerases, compositions, and methods of use
US11104888B2
Alternative substrates and formats for bead-based array of arrays TM
US20020102578A1
Polymerase enzymes and reagents for enhanced nucleic acid sequencing
US20080108082A1