Nucleosides and nucleotides with 3' acetal blocking groups

Nucleotides with a 3'-acetal blocking group address the challenge of controlled nucleotide incorporation by ensuring stability and efficient removal, improving sequencing accuracy and read length.

JP7797426B2Active Publication Date: 2026-01-13ILLUMINA CAMBRIDGE LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2022578720
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-06-22
Filing Date
2021-06-21
Publication Date
2026-01-13
Estimated Expiration
2041-06-21

AI Technical Summary

Technical Problem

Existing nucleotide sequencing methods face challenges in achieving controlled incorporation of nucleotides to prevent uncontrolled serial incorporation, requiring a protecting group that is stable, removable under mild conditions, and compatible with polymerase enzymes, while maintaining the integrity of the polynucleotide chain.

Method used

Nucleotides with a 3'-acetal blocking group covalently attached via a cleavable linker, allowing for efficient incorporation and removal in a single reaction step, enhancing stability and reducing prephasing and signal decay during sequencing.

Benefits of technology

The 3'-acetal blocking group provides improved stability and reduces prephasing and signal decay, leading to higher data quality and longer reads in sequencing applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007797426000095
    Figure 0007797426000095
  • Figure 0007797426000096
    Figure 0007797426000096
  • Figure 0007797426000097
    Figure 0007797426000097
Patent Text Reader

Abstract

The present invention relates to a nucleoside or nucleotide comprising a nucleobase attached to a detectable label via a cleavable linker, wherein the nucleoside or nucleotide comprises a ribose or 2' deoxyribose moiety and a 3'-OH blocking group, and the cleavable linker comprises a moiety of the structure: [Formula 1] TIFF2023531009000088.tif21128In formula, Each of X and Y is independently O or S; R 1a , R 2b , R 2 , R 3a and R 3b is independently H, halogen, unsubstituted or substituted C1-C6 alkyl, or C1-C6 haloalkyl. Also provided herein are methods for preparing such nucleotide and nucleoside molecules, and the use of fully functionalized nucleotides containing a 3' acetal blocking group for sequencing applications.
Need to check novelty before this filing date? Find Prior Art

Description

Incorporation by reference of priority applications

[0001] This application claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 042,240, filed June 22, 2020, the entire contents of which are incorporated by reference in their entirety. Background technology [Technical Field]

[0002] The present disclosure relates generally to nucleotides, nucleosides, or oligonucleotides that contain a 3' acetal blocking group and their use in polynucleotide sequencing methods. Methods for preparing the 3' blocked nucleotides, nucleosides, or oligonucleotides are also disclosed. 2. Description of Related Art

[0003] Advances in molecular research are driven, in part, by improvements in the techniques used to characterize molecules or their biological responses. In particular, the study of nucleic acids, DNA and RNA, has benefited from advanced techniques used for sequence analysis and the study of hybridization events.

[0004] An example of a technology that has improved the study of nucleic acids is the development of fabricated arrays of immobilized nucleic acids. These arrays typically consist of a high-density matrix of polynucleotides immobilized on a solid support material. See, for example, Fodor et al., Trends Biotech. 12:19-26, 1994, which describes a method for assembling nucleic acids using a chemically sensitized glass surface that is protected by a mask but exposed in predetermined areas to allow the attachment of appropriately modified nucleotide phosphoramidites. Fabricated arrays can also be produced by a technique of "spotting" known polynucleotides onto a solid support at predetermined locations (e.g., Stimpson et al., Proc. Natl. Acad. Sci. 92:6379-6383, 1995).

[0005] One method for determining the nucleotide sequence of array-bound nucleic acids is called "sequencing by synthesis" or "SBS." This technique for sequencing DNA ideally requires the controlled (i.e., one at a time) incorporation of the correct complementary nucleotide opposite the nucleic acid being sequenced. This allows for accurate sequencing by adding nucleotides in multiple cycles, as each nucleotide residue is sequenced one at a time, thus preventing uncontrolled serial incorporation. The incorporated nucleotide is read using the appropriate label attached to it before removing the label moiety and subsequent rounds of sequencing.

[0006] To ensure that only a single incorporation occurs, a structural modification (a "protecting group" or "blocking group") is included in each labeled nucleotide added to the growing strand to ensure that only one nucleotide is incorporated. After the nucleotide bearing the protecting group is added, the protecting group is removed under reaction conditions that do not interfere with the integrity of the DNA being sequenced. The sequencing cycle can then continue with the incorporation of the next protected, labeled nucleotide.

[0007] To be useful in DNA sequencing, nucleotides, usually nucleotide triphosphates, generally require a 3'-hydroxy protecting group to prevent the polymerase used to incorporate them into a polynucleotide chain from continuing to replicate as the base on the nucleotide is added. There are many limitations on the types of groups that can be added to a nucleotide and still be suitable. The protecting group should prevent additional nucleotide molecules from being added to the polynucleotide chain while simultaneously being easily removable from the sugar moiety without causing damage to the polynucleotide chain. Furthermore, the modified nucleotide must be compatible with the polymerase or another appropriate enzyme used to incorporate it into the polynucleotide chain. Therefore, an ideal protecting group exhibits long-term stability, is efficiently incorporated by the polymerase enzyme, blocks secondary or further nucleotide incorporation, and is capable of being removed under mild conditions, preferably aqueous conditions, that do not damage the polynucleotide structure.

[0008] Reversible protecting groups have been previously described. For example, Metzker et al. (Nucleic Acids Research, 22(20):4259-4267, 1994) disclose the synthesis and use of eight 3'-modified 2-deoxyribonucleoside 5'-triphosphates (3'-modified dNTPs) and their testing in two DNA template assays for incorporation activity. International Publication No. 2002 / 029003 describes a sequencing method that may involve the use of an allyl protecting group to cap the 3'-OH group on a growing strand of DNA in a polymerase reaction.

[0009] Additionally, the development of many reversible protecting groups and methods for deprotecting them under DNA-compatible conditions has been previously reported in International Publication Nos. WO2004 / 018497 and WO2014 / 139596, each of which is incorporated herein by reference in its entirety. Summary of the Invention

[0010] Some embodiments of the present disclosure relate to a nucleotide or nucleoside comprising a nucleobase attached to a detectable label via a cleavable linker, wherein the nucleoside or nucleotide comprises a ribose or 2' deoxyribose moiety and a 3'-OH blocking group, and the cleavable linker comprises a moiety of the following structure: [ka] wherein each of X and Y is independently O or S; R 1a , R 1b , R 2 , R 3a and R 3b Each of is independently H, halogen, unsubstituted or substituted C1-C6 alkyl, or C1-C6 haloalkyl.

[0011] Some embodiments of the present disclosure relate to oligonucleotides or polynucleotides that include the 3'-OH blocked labeled nucleotides described herein.

[0012] Some embodiments of the present disclosure relate to a method for preparing a growing polynucleotide complementary to a target single-stranded polynucleotide in a sequencing reaction, comprising incorporating a nucleotide molecule described herein into the growing complementary polynucleotide, wherein the incorporation of the nucleotide prevents the introduction of any subsequent nucleotide into the growing complementary polynucleotide. In some embodiments, the incorporation of the nucleotide is achieved by a polymerase, terminal deoxynucleotidyl transferase (TdT), or reverse transcriptase. In one embodiment, the incorporation is achieved by a polymerase (e.g., a DNA polymerase).

[0013] Some further embodiments of the present disclosure relate to a method for determining the sequence of a target single-stranded polynucleotide, the method comprising: (a) incorporating a nucleotide as described herein into a copy polynucleotide strand that is complementary to at least a portion of a target polynucleotide strand; (b) detecting the identity of the nucleotide incorporated into the copy polynucleotide strand; (c) chemically removing the label and the 3'-OH blocking group from the nucleotide incorporated into the copy polynucleotide strand; Includes.

[0014] In some embodiments, the detecting step comprises determining the identity of the nucleotide incorporated into the copy polynucleotide strand by making one or more measurements of a fluorescent signal from the detectable label. In some embodiments, the sequencing method further comprises (d) using a post-cleavage wash solution to wash the chemically removed label and 3'-OH blocking group away from the copy polynucleotide strand. In some embodiments, such a wash step also removes unincorporated nucleotides. In other embodiments, the method may comprise a separate wash step prior to step (b) to wash unincorporated nucleotides away from the copy polynucleotide strand. In some such embodiments, the 3'-OH blocking group and the detectable label of the incorporated nucleotide are removed before introducing the next complementary nucleotide. In some further embodiments, the 3'-OH blocking group and the detectable label are removed in a single chemical reaction step. In some embodiments, the sequential incorporation described herein is performed at least 50 times, at least 100 times, at least 150 times, at least 200 times, or at least 250 times.

[0015] Some further embodiments of the present disclosure relate to kits comprising a plurality of nucleotide or nucleoside molecules described herein and packaging materials therefor. The nucleotides, nucleosides, oligonucleotides, or kits described herein can be used to detect, measure, or identify biological systems (e.g., including processes or components thereof). Exemplary techniques in which the compounds, nucleotides, oligonucleotides, or kits can be used include sequencing, expression analysis, hybridization analysis, genetic analysis, RNA analysis, cellular assays (e.g., cell binding or cell function assays), or protein assays (e.g., protein binding assays or protein activity assays). The use may also be in automated instruments for performing specific techniques, such as automated sequencing instruments. Sequencing instruments may include two or more lasers operating at different wavelengths to distinguish between different detectable labels. [Brief explanation of the drawings]

[0016] [Figure 1] FIG. 1 is a diagram comparing the 3′-OH deblocking efficiency of [(allyl)PdCl] 2 with Na 2 PdCl 4 using various ratios of tris(hydroxylpropyl)phosphine (THP).

[0017] [Figure 2] FIG. 1 is a diagram showing the percent prephasing values ​​as a function of time for a standard fully functionalized nucleotide (ffN) having an LN3 linker moiety and a 3′-O-azidomethyl blocking group compared to an ffN having an AOL linker moiety and a 3′-AOM blocking group.

[0018] [Figure 3] FIG. 1 shows a comparison of phasing values ​​on Illumina's MiniSeq® instrument using fully functionalized nucleotides (ffNs) with a 3′-AOM blocking group with and without palladium scavenger in the post-cleavage wash step.

[0019] [Figure 4] FIG. 1 shows primary sequencing metrics, including phasing, prephasing, and errors, on an Illumina MiniSeq® instrument using fully functionalized nucleotides (ffNs) with 3′-AOM blocking groups and AOL linker moieties, compared to the same sequencing metrics using standard ffNs with 3′-O-azidomethyl blocking groups, when a palladium scavenger was used.

[0020] [Figure 5] FIG. 1 shows a comparison of primary sequencing metrics, including phasing and prephasing, on an Illumina MiniSeq® instrument using fully functionalized nucleotides (ffNs) with a 3′-AOM blocking group and an AOL linker moiety when glycine or ethanolamine, respectively, is used in the incorporation mixture.

[0021] [Figure 6] FIG. 1 shows primary sequencing metrics including phasing, prephasing, and error rates on an Illumina MiniSeq® instrument using fully functionalized nucleotides (ffNs) with an AOL linker moiety when using a 3′-AOM blocking group and glycine in the incorporation mixture, compared to the same sequencing metrics using standard ffNs with a 3′-O-azidomethyl blocking group.

[0022] [Figure 7A] FIG. 1 shows the error rate and Q30 sequencing metrics, respectively, of 2×300 sequencing runs on an Illumina iSeq™ instrument using fully functionalized nucleotides (ffNs) with a 3′-AOM blocking group and an AOL linker moiety compared to the same sequencing metrics using standard ffNs with a 3′-O-azidomethyl blocking group. [Figure 7B]FIG. 1 shows the error rate and Q30 sequencing metrics, respectively, of 2×300 sequencing runs on an Illumina iSeq™ instrument using fully functionalized nucleotides (ffNs) with a 3′-AOM blocking group and an AOL linker moiety compared to the same sequencing metrics using standard ffNs with a 3′-O-azidomethyl blocking group.

[0023] [Figure 8A] FIG. 1 shows the error rate and Q30 sequencing metrics, respectively, of 2×150 sequencing runs on an Illumina iSeq™ instrument using fully functionalized nucleotides (ffNs) with a 3′-AOM blocking group and an AOL linker moiety compared to the same sequencing metrics using standard ffNs with a 3′-O-azidomethyl blocking group. [Figure 8B] FIG. 1 shows the error rate and Q30 sequencing metrics, respectively, of 2×150 sequencing runs on an Illumina iSeq™ instrument using fully functionalized nucleotides (ffNs) with a 3′-AOM blocking group and an AOL linker moiety compared to the same sequencing metrics using standard ffNs with a 3′-O-azidomethyl blocking group.

[0024] [Figure 9A] Figure 1 shows primary sequencing metrics (phasing, signal decay rate, error rate) as a function of blue laser power at constant green laser power. Sequencing experiments were performed on an Illumina NovaSeq™ instrument using fully functionalized nucleotides (ffNs) with a 3'-AOM blocking group and an AOL linker moiety when using the palladium scavenger L-cysteine, compared to the same sequencing metrics using standard protocols and ffNs with a 3'-O-azidomethyl blocking group. [Figure 9B]Figure 1 shows primary sequencing metrics (phasing, signal decay rate, error rate) as a function of blue laser power at constant green laser power. Sequencing experiments were performed on an Illumina NovaSeq™ instrument using fully functionalized nucleotides (ffNs) with a 3'-AOM blocking group and an AOL linker moiety when using the palladium scavenger L-cysteine, compared to the same sequencing metrics using standard protocols and ffNs with a 3'-O-azidomethyl blocking group. [Figure 9C] Figure 1 shows primary sequencing metrics (phasing, signal decay rate, error rate) as a function of blue laser power at constant green laser power. Sequencing experiments were performed on an Illumina NovaSeq™ instrument using fully functionalized nucleotides (ffNs) with a 3'-AOM blocking group and an AOL linker moiety when using the palladium scavenger L-cysteine, compared to the same sequencing metrics using standard protocols and ffNs with a 3'-O-azidomethyl blocking group. [Figure 9D] Figure 1 shows primary sequencing metrics (phasing, signal decay rate, error rate) as a function of blue laser power at constant green laser power. Sequencing experiments were performed on an Illumina NovaSeq™ instrument using fully functionalized nucleotides (ffNs) with a 3'-AOM blocking group and an AOL linker moiety when using the palladium scavenger L-cysteine, compared to the same sequencing metrics using standard protocols and ffNs with a 3'-O-azidomethyl blocking group. [Figure 9E]Figure 1 shows primary sequencing metrics (phasing, signal decay rate, error rate) as a function of blue laser power at constant green laser power. Sequencing experiments were performed on an Illumina NovaSeq™ instrument using fully functionalized nucleotides (ffNs) with a 3'-AOM blocking group and an AOL linker moiety when using the palladium scavenger L-cysteine, compared to the same sequencing metrics using standard protocols and ffNs with a 3'-O-azidomethyl blocking group.

[0025] [Figure 10] Figure 1 shows primary sequencing metrics (%PF, error rate, Q30, and signal decay) on an Illumina iSeq™ instrument (1 x 150 cycles) using fully functionalized (ffN) with a 3'-AOM blocking group and an AOL linker moiety, where the same palladium cleavage mixture was also used in the first step of chemical linearization in SBS, compared to SBS of the same ffN using standard enzymatic linearization. DETAILED DESCRIPTION OF THE INVENTION

[0026] Embodiments of the present disclosure relate to nucleosides and nucleotides having a 3' acetal blocking group for sequencing applications, such as sequencing by synthesis (SBS). In some embodiments, the nucleoside or nucleotide contains a label covalently attached thereto via a cleavable linker comprising an acetal moiety, which allows for cleavage of the 3' acetal blocking group and the label in a single reaction step. The 3' acetal blocking group also provides improved stability during synthesis of the fully functionalized nucleotide (ffN), as well as greater stability in solution during formulation, storage, and operation on a sequencing instrument. Furthermore, the 3' acetal blocking groups described herein can also achieve reduced prephasing, lower signal decay for improved data quality, and enable longer reads from sequencing applications. definition

[0027] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. The use of the term "including" and other forms, such as "include," "includes," and "included," is not limiting. Furthermore, the use of the term "having" and other forms, such as "have," "has," and "had," is not limiting. As used herein, in transitional phrases or claim text, the terms "comprise" and "comprising" should be interpreted as having an open-ended meaning. That is, the above terms should be interpreted as synonymous with the phrases "having at least" or "including at least." For example, when used in the context of a process, the term "comprising" means that the process includes at least the recited steps, but may include additional steps. When used in the context of a compound, composition, or device, the term "comprising" means that the compound, composition, or device includes at least the recited features or components, but may include additional features or components.

[0028] When a range of values ​​is provided, it is understood that the upper and lower limits, and each intervening value between the upper and lower limits of the range, are encompassed within an embodiment.

[0029] As used herein, common organic abbreviations are defined as follows: [Table 1]

[0030] As used herein, the term "array" refers to a collection of different probe molecules attached to one or more substrates such that the different probe molecules can be differentiated from one another according to their relative positions. An array can include different probe molecules, each located at a different addressable position on the substrate. Alternatively or additionally, an array can include separate substrates, each bearing a different probe molecule, and the different probe molecules can be identified according to the position of the substrate on the surface to which the substrate is attached or according to the position of the substrate in a liquid. Exemplary arrays in which separate substrates are located on a surface include, but are not limited to, those containing beads in wells, such as those described in U.S. Pat. No. 6,355,431, U.S. Patent Application Publication No. 2002 / 0102578, and International Publication No. WO 00 / 63437. An exemplary format that can be used in the present invention to differentiate beads in a liquid array using a microfluidic device, such as a fluorescence-activated cell sorter (FACS), is described, for example, in U.S. Pat. No. 6,524,793. Further examples of arrays that can be used in the present invention include, but are not limited to, those described in U.S. Patent Nos. 5,429,807, 5,436,327, 5,561,071, 5,583,211, 5,658,734, 5,837,858, 5,874,219, 5,919,523, 6,136,269, 6,287,768, and 6,287,776. Nos. 6,288,220, 6,297,006, 6,291,193, 6,346,413, 6,416,949, 6,482,591, 6,514,751 and 6,610,482, and WO 93 / 17126, WO 95 / 11995, WO 95 / 35505, EP 742287 and EP 799897.

[0031] As used herein, the terms "covalently attached" or "covalently bonded" refer to the formation of a chemical bond characterized by the sharing of electron pairs between atoms. For example, a covalently bonded polymer coating refers to a polymer coating that forms a chemical bond with the functionalized surface of a substrate, as compared to attaching to the surface by other means, such as adhesion or electrostatic interactions. It will be understood that a polymer covalently attached to a surface can be attached by means in addition to covalent bonds.

[0032] As used herein, any "R" group(s) represents a substituent that may be attached to the indicated atom. The R group may be substituted or unsubstituted. When two "R" groups are described as forming a ring or ring system "together with the atoms to which they are attached," it means that the collective unit of the atoms, intervening bonds, and two R groups is the recited ring. For example, if the following substructure is present: [ka] R 1 and R 2 is defined as selected from the group consisting of hydrogen and alkyl, or R 1 and R 2 form an aryl or carbocyclyl together with the atom to which they are attached, but R 1 and R 2 can be selected from hydrogen or alkyl, or the substructure has the following structure: [ka] wherein A is an aryl ring or a carbocyclyl containing the indicated double bond.

[0033] It is understood that certain radical nomenclature can include either a monoradical or a diradical, depending on the context. For example, if the substituent requires two points of attachment to the rest of the molecule, the substituent is understood to be a diradical. For example, a substituent specified as alkyl, which requires two points of attachment, includes di-radicals such as -CH2-, -CH2CH2-, -CH2CH(CH3)CH2-, etc. Other radical nomenclature clearly indicates that the radical is a diradical, such as "alkylene" or "alkenylene."

[0034] As used herein, the term "halogen" or "halo" means any one of Group 7 of the Periodic Table of the Elements, for example, fluorine, chlorine, bromine, or iodine, with fluorine and chlorine being preferred.

[0035] As used herein, "C" is a generalized expression in which "a" and "b" are integers. a ~C b" refers to the number of carbon atoms in an alkyl, alkenyl, or alkynyl group, or the number of ring atoms in a cycloalkyl or aryl group. That is, alkyl, alkenyl, alkynyl, cycloalkyl rings, and aryl rings can contain "a" to "b" carbon atoms at both ends. For example, a "C1-C4 alkyl" group refers to all alkyl groups having 1 to 4 carbons, i.e., CH3-, CH3CH2-, CH3CH2CH2-, (CH3)2CH-, CH3CH2CH2CH2-, CH3CH2CH(CH3)-, and (CH3)3C-, and a C3-C4 cycloalkyl group refers to all cycloalkyl groups having 3 to 4 carbon atoms, i.e., cyclopropyl and cyclobutyl. Similarly, a "4- to 6-membered heterocyclyl" group refers to all heterocyclyl groups having 4 to 6 total ring atoms, such as azetidine, oxetane, oxazoline, pyrrolidine, piperidine, piperazine, morpholine, etc. When "a" and "b" are not specified with respect to an alkyl, alkenyl, alkynyl, cycloalkyl, or aryl group, the broadest range described by those definitions is to be assumed. As used herein, the term "C1-C6" includes C1, C2, C3, C4, C5, and C6, as well as ranges defined by either of the two numbers. For example, C1-C6 alkyl includes C1, C2, C3, C4, C5, and C6 alkyl, C2-C6 alkyl, C1-C3 alkyl, etc. Similarly, C2-C6 alkenyl includes C2, C3, C4, C5, and C6 alkenyl, C2-C5 alkenyl, C3-C4 alkenyl, etc., and C2-C6 alkynyl includes C2, C3, C4, C5, and C6 alkynyl, C2-C5 alkynyl, C3-C4 alkynyl, etc. C3-C8 cycloalkyl each includes hydrocarbon rings containing 3, 4, 5, 6, 7, and 8 carbon atoms, or ranges defined by either of two numbers, such as C3-C7 cycloalkyl or C5-C6 cycloalkyl.

[0036] As used herein, "alkyl" refers to a straight or branched hydrocarbon chain that is fully saturated (ie, contains no double or triple bonds). An alkyl group may have 1 to 20 carbon atoms (whenever presented herein, a numerical range such as "1 to 20" refers to each integer within the given range. For example, "1 to 20 carbon atoms" means that the alkyl group can consist of 1 carbon atom, 2 carbon atoms, 3 carbon atoms, etc., up to 20 carbon atoms, but this definition also covers occurrences of the term "alkyl" (where no numerical range is specified). An alkyl group may also be a medium-sized alkyl having 1 to 9 carbon atoms. An alkyl group may also be a lower alkyl having 1 to 6 carbon atoms. An alkyl group may be designated as "C1-C4 alkyl" or similar designation. By way of example only, "C1-C6 alkyl" indicates that there are 1 to 6 carbon atoms in the alkyl chain, i.e., the alkyl chain is selected from the group consisting of methyl, ethyl, propyl, isopropyl, n-butyl, iso-butyl, sec-butyl, and t-butyl. Typical alkyl groups include, but are not limited to, methyl, ethyl, propyl, isopropyl, butyl, isobutyl, tertiary butyl, pentyl, hexyl, and the like.

[0037] As used herein, "alkoxy" refers to the formula -OR where R is alkyl as defined above, such as "C1-C9 alkoxy", including, but not limited to, methoxy, ethoxy, n-propoxy, 1-methylethoxy (isopropoxy), n-butoxy, isobutoxy, sec-butoxy, and tert-butoxy.

[0038] As used herein, "alkenyl" refers to a straight or branched hydrocarbon chain containing one or more double bonds. Alkenyl groups can have from 2 to 20 carbon atoms, although this definition also covers occurrences of the term "alkenyl" where no numerical range is specified. Alkenyl groups can also be medium-sized alkenyls having from 2 to 9 carbon atoms. Alkenyl groups can also be lower alkenyls having from 2 to 6 carbon atoms. Alkenyl groups may also be designated as "C2-C6 alkenyl" or similar designations. By way of example only, "C2-C6 alkenyl" indicates that there are 2 to 6 carbon atoms in the alkenyl chain, i.e., the alkenyl chain is selected from the group consisting of ethenyl, propen-1-yl, propen-2-yl, propen-3-yl, buten-1-yl, buten-2-yl, buten-3-yl, buten-4-yl, 1-methyl-propen-1-yl, 2-methyl-propen-1-yl, 1-ethyl-ethen-1-yl, 2-methyl-propen-3-yl, buta-1,3-dienyl, buta-1,2-dienyl, and buta-1,2-dien-4-yl. Typical alkenyl groups include, but are not limited to, ethenyl, propenyl, butenyl, pentenyl, and hexenyl.

[0039] As used herein, "alkynyl" refers to a straight or branched hydrocarbon chain containing one or more triple bonds. An alkynyl group can have 2 to 20 carbon atoms, although this definition also covers occurrences of the term "alkynyl" where no numerical range is specified. An alkynyl group may also be a medium-sized alkynyl having 2 to 9 carbon atoms. An alkynyl group may also be a lower alkynyl having 2 to 6 carbon atoms. An alkynyl group may be designated as "C2-C6 alkynyl" or similar designations. By way of example only, "C2-C6 alkynyl" indicates that there are 2 to 6 carbon atoms in the alkynyl chain, i.e., the alkynyl chain is selected from the group consisting of ethynyl, propyn-1-yl, propyn-2-yl, butyn-1-yl, butyn-3-yl, butyn-4-yl, and 2-butynyl. Typical alkynyl groups include, but are not limited to, ethynyl, propynyl, butynyl, pentynyl, hexynyl, and the like.

[0040] As used herein, "heteroalkyl" refers to a straight or branched hydrocarbon chain containing one or more heteroatoms, which include atoms other than carbon in the chain backbone, including, but not limited to, nitrogen, oxygen, and sulfur. Heteroalkyl groups may have 1 to 20 carbon atoms, although this definition also covers occurrences of the term "heteroalkyl" where no numerical range is specified. Heteroalkyl groups may also be medium-sized heteroalkyls having 1 to 9 carbon atoms. Heteroalkyl groups may also be lower heteroalkyls having 1 to 6 carbon atoms. Heteroalkyl groups may be designated as "C1-C6 heteroalkyl" or similar notations. Heteroalkyl groups may contain one or more heteroatoms. By way of example only, "C4-C6 heteroalkyl" indicates that there are 4 to 6 carbon atoms in the heteroalkyl chain and, in addition, there are one or more heteroatoms in the chain backbone.

[0041] The term "aromatic" refers to a ring or ring system having a conjugated π-electron system and including both carbocyclic aromatic (e.g., phenyl) and heterocyclic aromatic (e.g., pyridine) groups. The term includes monocyclic or fused polycyclic (i.e., rings that share adjacent pairs of atoms) groups, so long as the entire ring system is aromatic.

[0042] As used herein, "aryl" refers to an aromatic ring or ring system (i.e., two or more fused rings that share two adjacent carbon atoms) that contains only carbon in the ring backbone. When aryl is a ring system, all rings in the system are aromatic. An aryl group can have 6 to 18 carbon atoms, although this definition also covers occurrences of the term "aryl" where no numerical range is specified. In some embodiments, an aryl group has 6 to 10 carbon atoms. An aryl group is defined as "C6-C 10 aryl," "C6 or C 10 Examples of aryl groups include, but are not limited to, phenyl, naphthyl, azulenyl, and anthracenyl.

[0043] "Aralkyl" or "arylalkyl" refers to "C 7~14 An aryl group bonded as a substituent via an alkylene group, such as "aralkyl," includes, but is not limited to, benzyl, 2-phenylethyl, 3-phenylpropyl, and naphthylalkyl. In some cases, the alkylene group is a lower alkylene group (i.e., a C1-C6 alkylene group).

[0044] As used herein, "heteroaryl" refers to an aromatic ring or ring system (i.e., two or more fused rings that share two adjacent atoms) containing one or more heteroatoms, i.e., elements other than carbon, in the ring backbone, including, but not limited to, nitrogen, oxygen, and sulfur. When heteroaryl is a ring system, all rings in the system are aromatic. Heteroaryl groups can have 5 to 18 ring members (i.e., the number of atoms comprising the ring backbone, including carbon atoms and heteroatoms), although this definition also covers occurrences of the term "heteroaryl" where no numerical range is specified. In some embodiments, heteroaryl groups have 5 to 10 ring members or 5 to 7 ring members. Heteroaryl groups can be designated as "5- to 7-membered heteroaryl," "5- to 10-membered heteroaryl," or similar designations. Examples of heteroaryl rings include, but are not limited to, furyl, thienyl, phthalazinyl, pyrrolyl, oxazolyl, thiazolyl, imidazolyl, pyrazolyl, isoxazolyl, isothiazolyl, triazolyl, thiadiazolyl, pyridinyl, pyridazinyl, pyrimidinyl, pyrazinyl, triazinyl, quinolinyl, isoquinolinyl, benzimidazolyl, benzoxazolyl, benzothiazolyl, indolyl, isoindolyl, and benzothienyl.

[0045] A "heteroaralkyl" or "heteroarylalkyl" is a heteroaryl group bonded as a substituent via an alkylene group. Examples include, but are not limited to, 2-thienylmethyl, 3-thienylmethyl, furylmethyl, thienylethyl, pyrrolylalkyl, pyridylalkyl, isoxazolylalkyl, and imidazolylalkyl. In some cases, the alkylene group is a lower alkylene group (i.e., a C1-C6 alkylene group).

[0046] As used herein, "carbocyclyl" refers to a non-aromatic ring or ring system containing only carbon atoms in the ring system backbone. When a carbocyclyl is a ring system, two or more rings can be joined together in a fused, bridged, or spiro-connected manner. A carbocyclyl can have any degree of saturation, provided that at least one ring in the ring system is not aromatic. Thus, carbocyclyl includes cycloalkyl, cycloalkenyl, and cycloalkynyl. A carbocyclyl group can have 3 to 20 carbon atoms, although this definition also covers occurrences of the term "carbocyclyl" where no numerical range is specified. A carbocyclyl group can also be a medium-sized carbocyclyl having 3 to 10 carbon atoms. A carbocyclyl group can also be a carbocyclyl having 3 to 6 carbon atoms. A carbocyclyl group can also be designated as "C3-C6 carbocyclyl" or similar designations. Examples of carbocyclyl rings include, but are not limited to, cyclopropyl, cyclobutyl, cyclopentyl, cyclohexyl, cyclohexenyl, 2,3-dihydro-indene, bicycle[2.2.2]octanyl, adamantyl, and spiro[4.4]nonanyl.

[0047] As used herein, "cycloalkyl" means a fully saturated carbocyclyl ring or ring system. Examples include cyclopropyl, cyclobutyl, cyclopentyl, and cyclohexyl.

[0048] As used herein, "heterocyclyl" refers to a non-aromatic ring or ring system containing at least one heteroatom in the ring backbone. Heterocyclyls may be joined together in fused, bridged, or spiro-linked configurations. Heterocyclyls may have any degree of saturation, provided that at least one ring in the ring system is not aromatic. The heteroatom may be present in either the non-aromatic or aromatic ring in the ring system. Heterocyclyl groups may have 3 to 20 ring members (i.e., the number of atoms comprising the ring backbone, including carbon atoms and heteroatoms), although this definition also covers occurrences of the term "heterocyclyl" where no numerical range is specified. Heterocyclyl groups may also be medium-sized heterocyclyls having 3 to 10 ring members. Heterocyclyl groups may also be heterocyclyls having 3 to 6 ring members. Heterocyclyl groups may also be designated as "3- to 6-membered heterocyclyl" or similar notations. In preferred 6-membered monocyclic heterocyclyls, the heteroatoms are selected from one to three of O, N, or S. In preferred 5-membered monocyclic heterocyclyls, the heteroatoms are selected from one or two heteroatoms selected from O, N, or S.Examples of heterocyclyl rings include azepinyl, acridinyl, carbazolyl, cinnolinyl, dioxolanyl, imidazolinyl, imidazolidinyl, morpholinyl, oxiranyl, oxepanyl, thiapanyl, piperidinyl, piperazinyl, dioxapiperazinyl, pyrrolidinyl, pyrrolidionyl, 4-piperidonyl, pyrazolinyl, pyrazolidinyl, 1,3-dioxinyl, 1,3-dioxanyl, 1,4-dioxinyl, 1,4-dioxanyl, 1,3-oxathinyl, 1,4-oxathinyl, 1,4-oxathianyl, 2H-1,2-oxazinyl, trioxanyl, hexahydro-1,3, These include, but are not limited to, 5-triazinyl, 1,3-dioxolyl, 1,3-dioxolanyl, 1,3-dithiolyl, 1,3-dithiolanyl, isoxazolinyl, isoxazolidinyl, oxazolinyl, oxazolidinyl, oxazolidinonyl, thiazolinyl, thiazolidinyl, 1,3-oxathiolanyl, indolinyl, isoindolinyl, tetrahydrofuranyl, tetrahydropyranyl, tetrahydrothiophenyl, tetrahydrothiopyranyl, tetrahydro-1,4-thiazinyl, thiamorpholinyl, dihydrobenzofuranyl, benzimidazolidinyl, and tetrahydroquinoline.

[0049] As used herein, "alkoxyalkyl" or "(alkoxy)alkyl" refers to an alkoxy group bonded via an alkylene group, such as a C2-C8 alkoxyalkyl, or a (C1-C6 alkoxy)C1-C6 alkyl, such as -(CH2) 1-3 -OCH3.

[0050] As used herein, "-O-alkoxyalkyl" or "-O-(alkoxy)alkyl" refers to an alkoxy group attached via an -O-(alkylene) group, such as -O-(C1-C6 alkoxy)C1-C6 alkyl, e.g., -O-(CH2) 1-3 -OCH3.

[0051] As used herein, "(heterocyclyl)alkyl" refers to a heterocyclic or heterocyclyl group as defined above, bonded as a substituent via an alkylene group as defined above. The alkylene and heterocyclyl groups of a (heterocyclyl)alkyl can be substituted or unsubstituted. Examples include, but are not limited to, tetrahydro-2H-pyran-4-yl)methyl, (piperidin-4-yl)ethyl, (piperidin-4-yl)propyl, (tetrahydro-2H-thiopyran-4-yl)methyl, and (1,3-thiazinane-4-yl)methyl.

[0052] As used herein, "(cycloalkyl)alkyl" or "(carbocyclyl)alkyl" refers to a cycloalkyl or carbocyclyl group (as defined herein) attached as a substituent via an alkylene group. Examples include, but are not limited to, cyclopropylmethyl, cyclobutylmethyl, cyclopentylethyl, and cyclohexylpropyl.

[0053] An "O-carboxy" group refers to an "-OC(=O)R" group, where R is hydrogen, C1-C6 alkyl, C2-C6 alkenyl, C2-C6 alkynyl, C3-C7 carbocyclyl, C6-C7 alkyl, C8-C9 alkyl, C9-C10 alkyl, C11-C12 alkyl, C12-C14 alkyl, C13-C16 alkyl, C14-C16 alkyl, C15-C16 alkyl, C16-C18 alkyl, C17-C18 alkyl, C18-C19 alkyl, C19-C20 ... 10 It is selected from aryl, 5- to 10-membered heteroaryl, and 3- to 10-membered heterocyclyl.

[0054] A "C-carboxy" group refers to a "-C(=O)OR" group, where R is hydrogen, C1-C6 alkyl, C2-C6 alkenyl, C2-C6 alkynyl, C3-C7 carbocyclyl, C6-C7 alkyl, C6-C8 alkyl, C6-C9 alkyl, C6-C10 alkyl, C6-C11 alkyl, C6-C12 alkyl, C6-C13 alkyl, C6-C14 alkyl, C6-C15 alkyl, C6-C16 alkyl, C6-C17 alkyl, C6-C18 alkyl, C6-C19 ... 10 It is selected from the group consisting of aryl, 5- to 10-membered heteroaryl, and 3- to 10-membered heterocyclyl. Non-limiting examples include carboxyl (i.e., -C(=O)OH).

[0055] A "sulfonyl" group refers to a "-SO2R" group, where R is hydrogen, C1-C6 alkyl, C2-C6 alkenyl, C2-C6 alkynyl, C3-C7 carbocyclyl, C6-C7 alkyl, C6-C8 alkyl, C6-C9 alkyl, C6-C10 alkyl, C6-C11 alkyl, C6-C12 alkyl, C6-C13 alkyl, C6-C14 alkyl, C6-C15 alkyl, C6-C16 alkyl, C6-C17 alkyl, C6-C18 alkyl, C6-C19 ... 10 It is selected from aryl, 5- to 10-membered heteroaryl, and 3- to 10-membered heterocyclyl.

[0056] A "sulfino" group refers to a "-S(=O)OH" group.

[0057] The "S-sulfonamide" group is "-SONR A R B " group, where R A and R B are each independently selected from hydrogen, C1-C6 alkyl, C2-C6 alkenyl, C2-C6 alkynyl, C3-C7 carbocyclyl, C6-C 10 It is selected from aryl, 5- to 10-membered heteroaryl, and 3- to 10-membered heterocyclyl.

[0058] An "N-sulfonamido" group is defined as "-N(R A )SO2R B " refers to the group R A and R b are each independently selected from hydrogen, C1-C6 alkyl, C2-C6 alkenyl, C2-C6 alkynyl, C3-C7 carbocyclyl, C6-C 10 It is selected from aryl, 5- to 10-membered heteroaryl, and 3- to 10-membered heterocyclyl.

[0059] A "C-amido" group is defined as "-C(=O)NR A R B " refers to the group R A and R B are each independently selected from hydrogen, C1-C6 alkyl, C2-C6 alkenyl, C2-C6 alkynyl, C3-C7 carbocyclyl, C6-C 10 It is selected from aryl, 5- to 10-membered heteroaryl, and 3- to 10-membered heterocyclyl.

[0060] An "N-amido" group is defined as "-N(RA )C(=O)R B " refers to the group R A and R B are each independently selected from hydrogen, C1-C6 alkyl, C2-C6 alkenyl, C2-C6 alkynyl, C3-C7 carbocyclyl, C6-C 10 It is selected from aryl, 5- to 10-membered heteroaryl, and 3- to 10-membered heterocyclyl.

[0061] The "amino" group is "-NR A R B " refers to the group R A and R B are each independently selected from hydrogen, C1-C6 alkyl, C2-C6 alkenyl, C2-C6 alkynyl, C3-C7 carbocyclyl, C6-C 10 It is selected from aryl, 5- to 10-membered heteroaryl, and 3- to 10-membered heterocyclyl. Non-limiting examples include free amino (i.e., -NH2).

[0062] An "aminoalkyl" group refers to an amino group linked via an alkylene group.

[0063] An "(alkoxy)alkyl" group refers to an alkoxy group bonded via an alkylene group, such as "(C1-C6 alkoxy)C1-C6 alkyl."

[0064] As used herein, the term "hydroxy" refers to an --OH group.

[0065] As used herein, the term "cyano" group refers to a "-CN" group.

[0066] As used herein, the term "azido" refers to the group --N3.

[0067] As used herein, the term "propargylamine" refers to an amino group substituted with a propargyl group. [ka] When propargylamine is used in context as a divalent moiety, [ka] wherein R is as defined herein. A is hydrogen, and is C1-C6 alkyl, C2-C6 alkenyl, C2-C6 alkynyl, C3-C7 carbocyclyl, C6-C 10 Includes aryl, 5- to 10-membered heteroaryl, and 3- to 10-membered heterocyclyl.

[0068] As used herein, the term "propargylamide" refers to a propargyl group [ka] When propargylamide is used in the context of a divalent moiety, it refers to a C-amido or N-amido group substituted with [ka] and R, as defined herein. A is hydrogen, and is C1-C6 alkyl, C2-C6 alkenyl, C2-C6 alkynyl, C3-C7 carbocyclyl, C6-C 10 Includes aryl, 5- to 10-membered heteroaryl, and 3- to 10-membered heterocyclyl.

[0069] As used herein, the term "allylamine" refers to an amino group substituted with an allyl group (CH=CH-CH-). When allylamine is used in the context of a divalent moiety, it is -CH=CH-CH-NR A wherein R A is, as defined herein, hydrogen, C1-C6 alkyl, C2-C6 alkenyl, C2-C6 alkynyl, C3-C7 carbocyclyl, C6-C 10 aryl, 5- to 10-membered heteroaryl, and 3- to 10-membered heterocyclyl.

[0070] The term "allylamido" as used herein refers to a C-amido or N-amido group substituted with an allyl group (CH=CH-CH-). When allylamide is used in this context as a divalent moiety, it includes -CH=CH-CH-NR, as defined herein. A -C(=O)- or -CH=CH-CH2-C(=O)-NR A wherein R A is hydrogen, C1-C6 alkyl, C2-C6 alkenyl, C2-C6 alkynyl, C3-C7 carbocyclyl, C6-C 10 aryl, 5- to 10-membered heteroaryl, and 3- to 10-membered heterocyclyl.

[0071] When a group is described as "optionally substituted," it can be either unsubstituted or substituted. Similarly, when a group is described as "substituted," the substituents may be selected from one or more of the listed substituents. As used herein, a substituent is derived from an unsubstituted parent group in which one or more hydrogen atoms have been exchanged for another atom or group. Unless otherwise stated, when a group is deemed "substituted," it means that the group is substituted with one or more substituents independently selected from the following: C1-C6 alkyl, C1-C6 alkenyl, C1-C6 alkynyl, C1-C6 heteroalkyl, C3-C7 carbocyclyl (optionally substituted with halo, C1-C6 alkyl, C1-C6 alkoxy, C1-C6 haloalkyl, and C1-C6 haloalkoxy), C3-C7 carbocyclyl-C1-C6-alkyl (optionally substituted with halo, C1-C6 alkyl, C1-C6 alkoxy, C1-C6 haloalkyl, and and C1-C6 haloalkoxy), 3-10 membered heterocyclyl (optionally substituted with halo, C1-C6 alkyl, C1-C6 alkoxy, C1-C6 haloalkyl, and C1-C6 haloalkoxy), 3-10 membered heterocyclyl-C1-C6-alkyl (optionally substituted with halo, C1-C6 alkyl, C1-C6 alkoxy, C1-C6 haloalkyl, and C1-C6 haloalkoxy), aryl (optionally substituted with halo, C1-C6 alkyl, C1-C6 alkoxy, C1-C6 haloalkyl, and C1-C6 haloalkoxy), (aryl)C1-C6 alkyl (optionally substituted with halo, C1-C6 alkyl, C1-C6 alkoxy, C1-C6 haloalkyl, and C1-C6 haloalkoxy), 5-10 membered heteroaryl (optionally substituted with halo, C1-C6 alkyl, C1-C6 alkoxy, C1- substituted with C6 haloalkyl, and C1-C6 haloalkoxy), (5-10 membered heteroaryl)C1-C6 alkyl (optionally substituted with halo, C1-C6 alkyl, C1-C6 alkoxy, C1-C6 haloalkyl, and C1-C6 haloalkoxy), halo, -CN, hydroxy, C1-C6 alkoxy, (C1-C6 alkoxy)C1-C6 alkyl, -O(C1-C6 alkoxy)C1-C6 alkyl;(C1-C6 haloalkoxy)C1-C6 alkyl; -O(C1-C6 haloalkoxy)C1-C6 alkyl; aryloxy, sulfhydryl (mercapto), halo(C1-C6)alkyl (e.g., -CF3), halo(C1-C6)alkoxy (e.g., -OCF3), C1-C6 alkylthio, arylthio, amino, amino(C1-C6)alkyl, nitro, O-carbamyl, N-carbamyl, O-thiocarbamyl, N-thiocarbamyl, C-amido, N-amido, S-sulfonamido, N-sulfonamido, C-carboxy, O-carboxy, acyl, cyanato, isocyanato, thiocyanato, sulfinyl, sulfonyl, -SO3H, sulfino, -OSO2C; 1~4 Alkyl, monophosphate, diphosphate, triphosphate, and oxo (=O). When a group is described as being "optionally substituted," the group, if substituted, can be substituted with the above substituents.

[0072] Where a substituent is depicted as a diradical (i.e., having two points of attachment to the rest of the molecule), it is understood that the substituent can be attached in any directional configuration unless otherwise indicated. Thus, for example, -AE- or [ka] Substituents depicted as: represent substituents oriented such that A is attached at the left-most point of attachment of the molecule, as well as when A is attached at the right-most point of attachment of the molecule. [ka] where L is defined as an optionally present linker moiety; when L is not present (or absent), such a group or substituent is [ka] is equivalent to

[0073] As used herein, a "nucleotide" comprises a nitrogenous heterocyclic base, a sugar, and one or more phosphate groups. They are the monomeric units of nucleic acid sequences. In RNA, the sugar is ribose, and in DNA, it is deoxyribose, i.e., a sugar lacking the hydroxyl group present in ribose. The nitrogenous heterocyclic base can be a purine or a pyrimidine base. Purine bases include adenine (A) and guanine (G), as well as modified derivatives or analogs thereof. Pyrimidine bases include cytosine (C), thymine (T), and uracil (U), as well as modified derivatives or analogs thereof. The C-1 atom of deoxyribose is linked to the N-1 atom of a pyrimidine or the N-9 atom of a purine.

[0074] As used herein, a "nucleoside" is structurally similar to a nucleotide but lacks a phosphate moiety. An example of a nucleoside analog is one in which a label is linked to the base and there is no phosphate group attached to the sugar molecule. The term "nucleoside" is used herein in its ordinary sense as understood by those skilled in the art. Examples include, but are not limited to, ribonucleosides containing a ribose moiety and deoxyribonucleosides containing a deoxyribose moiety. Modified pentose moieties are those in which an oxygen atom is replaced with a carbon and / or a carbon is replaced with a sulfur or oxygen atom. A "nucleoside" is a monomer that may have a substituted base and / or sugar moiety. Furthermore, nucleosides can be incorporated into larger DNA and / or RNA polymers and oligomers.

[0075] The term "purine base" is used herein in its ordinary sense as understood by those skilled in the art and includes its tautomers. Similarly, the term "pyrimidine base" is used herein in its ordinary sense as understood by those skilled in the art and includes its tautomers. A non-limiting list of optionally substituted purine bases includes purine, deazapurine, adenine, 7-deazaadenine, guanine, 7-deazaguanine, hypoxanthine, xanthine, alloxanthine, 7-alkylguanine (e.g., 7-methylguanine), theobromine, caffeine, uric acid, and isoguanine. Examples of pyrimidine bases include, but are not limited to, cytosine, thymine, uracil, 5,6-dihydrouracil, and 5-alkylcytosine (e.g., 5-methylcytosine).

[0076] As used herein, when an oligonucleotide or polynucleotide is described as "comprising" or "labeled with" a nucleoside or nucleotide described herein, it means that the nucleoside or nucleotide described herein forms a covalent bond with the oligonucleotide or polynucleotide. Similarly, when a nucleoside or nucleotide is described as being part of an oligonucleotide or polynucleotide, such as being "incorporated into" an oligonucleotide or polynucleotide, it means that the nucleoside or nucleotide described herein forms a covalent bond with the oligonucleotide or polynucleotide. In some such embodiments, the covalent bond is formed between the 3' hydroxy group of the oligonucleotide or polynucleotide and the 5' phosphate group of a nucleotide described herein, or as a phosphodiester bond between the 3' carbon atom of the oligonucleotide or polynucleotide and the 5' carbon atom of the nucleotide.

[0077] As used herein, the term "cleavable linker" does not imply that the entire linker must be removed. The cleavage site can be located on the linker at a position that ensures that a portion of the linker remains attached to the detectable label and / or nucleoside or nucleotide moiety after cleavage.

[0078] As used herein, "derivative" or "analog" refers to a synthetic nucleotide or nucleoside derivative having a modified base moiety and / or a modified sugar moiety. Such derivatives and analogs are discussed, for example, in Scheit, Nucleotide Analogs (John Wiley & Son, 1980) and Uhlman et al., Chemical Reviews 90:543-584, 1990. Nucleotide analogs can also include modified phosphodiester linkages, including phosphorothioate, phosphorodithioate, alkyl-phosphonate, phosphoranilidate, and phosphoramidate linkages. As used herein, "derivative," "analog," and "modified" can be used interchangeably and are encompassed by the terms "nucleotide" and "nucleoside" as defined herein.

[0079] As used herein, the term "phosphate" is used in its ordinary sense as understood by one of ordinary skill in the art and includes its protonated form (e.g., [ka] As used herein, the terms "monophosphate," "diphosphate," and "triphosphate" are used in their ordinary sense as understood by those of skill in the art and include protonated forms.

[0080] As used herein, the terms "protecting group" and "protecting group" refer to any atom or group of atoms added to a molecule to prevent existing groups in the molecule from undergoing undesired chemical reactions. Sometimes, "protecting group" and "blocking group" can be used interchangeably.

[0081] As used herein, the term "phasing" refers to a phenomenon in sequencing by sequence sequencing (SBS) caused by incomplete removal of the 3' terminator and fluorophore, and the incorporation of a portion of the DNA strand within a cluster by the polymerase in a given sequencing cycle due to a failure to complete the sequence. Prephasing is caused by the incorporation of a nucleotide without an effective 3' terminator, and the incorporation event advances one cycle due to the failure to complete the sequence. Phase and pre-phase ensure that the measured signal intensity for a particular cycle is composed of signal from the current cycle and noise from the previous and next cycles. As the number of cycles increases, the proportion of sequences per cluster affected by phase and pre-phase increases, preventing the identification of the correct base. Prephasing can be caused by the presence of trace amounts of unprotected or unblocked 3'-OH nucleotides during sequencing by synthesis (SBS). Unprotected 3'-OH nucleotides can be generated during the manufacturing process or, in some cases, during storage and reagent handling processes. Therefore, the discovery of nucleotide analogs that reduce the incidence of prephasing is surprising and offers significant advantages in SBS applications over existing nucleotide analogs. For example, provided nucleotide analogs may result in faster SBS cycle times, lower phasing and prephasing values, and longer sequencing read lengths.

[0082] Nucleosides or nucleotides with 3' acetal blocking groups Some embodiments of the present disclosure relate to nucleotide or nucleoside molecules comprising a nucleobase attached to a detectable label via a cleavable linker and a ribose or deoxyribose moiety, wherein the cleavable linker has the structure: [ka] Each of X and Y is independently O or S, and R 1a , R 1b , R 2 , R 3a and R 3bis independently H, halogen, unsubstituted or substituted C1-C6 alkyl, or C1-C6 haloalkyl. In some embodiments, the ribose or deoxyribose moiety comprises a 3'-OH protecting group as described herein. In some embodiments, the cleavable linker is 1 Or L 2 , or both, and L 1 or L 2 are described in detail below.

[0083] In some embodiments, a nucleoside or nucleotide described herein has the structure of Formula (I): [ka] wherein B is a nucleobase; R 4 is H or OH, R 5 is H, a 3'-OH blocking group, or a phosphoramidite; R 6 is H, monophosphate, diphosphate, triphosphate, thiophosphate, a phosphate ester analog, a reactive phosphorus-containing group, or a hydroxy protecting group; L is [ka] and L 1 and L 2 Each of is independently an optionally present linker moiety.

[0084] In some embodiments of the cleavable linker moieties described herein, each of X and Y is O. In some other embodiments, X is S and Y is O, or X is O and Y is S. In some embodiments, R 1a , R 1b , R 2 , R 3a and R 3b is H. In other embodiments, R 1a , R1b , R 2 , R 3a and R 3b At least one of R is halogen (e.g., fluoro, chloro) or unsubstituted C1-C6 alkyl (e.g., methyl, ethyl, isopropyl, isobutyl, or t-butyl). In some such examples, R 1a and R 1b Each of is H and R 2 , R 3a and R 3b At least one of is unsubstituted C1-C6 alkyl or halogen (e.g., R 2 is unsubstituted C1-C6 alkyl, and R 3a and R 3b each of is H, or R 2 is H and R 3a and R 3b wherein one or both of are halogen or unsubstituted C1-C6 alkyl. In one embodiment, the cleavable linker or L is [ka] (the "AOL" linker portion).

[0085] In some embodiments of the nucleosides or nucleotides described herein, the nucleobase ("B" in Formula (I)) is a purine (adenine or guanine), a deazapurine, or a pyrimidine (e.g., cytosine, thymine, or uracil). In some further embodiments, the deazapurine is a 7-deazapurine (e.g., 7-deazaadenine or 7-deazaguanine). Non-limiting examples of B are: [ka] or optionally substituted derivatives and analogs thereof. In some further embodiments, the labeled nucleobase has the structure [ka] Includes.

[0086] In some embodiments of the nucleosides or nucleotides described herein, the ribose or deoxyribose moiety comprises a 3′-OH blocking group (i.e., R 5 is a 3'-OH blocking group). In some embodiments, the 3'-OH blocking group or R 5 and [ka] In the formula, R a , R b , R c , R d and R e Each of is independently H, halogen, unsubstituted or substituted C1-C6 alkyl, or C1-C6 haloalkyl. a and R b is H and R c , R d and R e At least one of R is independently halogen (e.g., fluoro, chloro) or unsubstituted C1-C6 alkyl (e.g., methyl, ethyl, isopropyl, isobutyl, or t-butyl). c is unsubstituted C1-C6 alkyl, and R d and R e Each of R is H. c is H and R d and R e One or both of R is halogen or unsubstituted C1-C6 alkyl. 5 Other non-limiting embodiments of [ka] In one embodiment, R 5 teeth, [ka] which, together with the 3' oxygen, is attached to the 3' carbon atom of a ribose or deoxyribose moiety [ka] ("AOM") group. In other embodiments, the 3'-OH blocking group or R 5 may include an azide moiety (e.g., -CH2N3 or azidomethyl). Further embodiments of 3'-OH blocking groups are described in U.S. Patent Application Publication No. 2020 / 0216891, which is incorporated by reference in its entirety, and which include a 3'-OH blocking group attached to the 3' carbon atom of a ribose or deoxyribose moiety: [ka] Further examples of 3' acetal blocking groups include:

[0087] In some other embodiments of the nucleosides or nucleotides described herein, R 5 is a phosphoramidite in formula (I). In such embodiments, R 6 is an acid-cleavable hydroxy protecting group that allows subsequent monomer coupling under automated synthesis conditions.

[0088] In some embodiments of the nucleosides or nucleotides described herein, L 1 exists, and L 1 In some further embodiments, L comprises a moiety selected from the group consisting of propargylamine, propargylamide, allylamine, allylamide, and optionally substituted variants thereof. 1 teeth, [ka] In some further embodiments, the asterisk * is the L relative to the nucleobase (e.g., C5 position of a pyrimidine base or C7 position of a 7-deazapurine base). 1 The attachment points are shown.

[0089] In some embodiments, a nucleotide described herein is a fully functionalized nucleotide (ffN) and comprises a dye compound covalently attached to a nucleobase via a 3'-OH blocking group described herein and a cleavable linker described herein, wherein the cleavable linker has the structure [ka] L 1 Including, * L 1indicates the point of attachment to a nucleobase (e.g., the C5 position of a cytosine, thymine, or uracil base, or the C7 position of a 7-deazaadenine or 7-deazaguanine). In some examples, ffNs having an allylamine or allylamide linker moiety described herein are also referred to as ffN-DB or ffN-(DB), where "DB" refers to the double bond of the linker moiety. In some examples, sequencing runs using ffN sets (including ffA, ffT, ffC, ffG) in which one or more ffNs are ffN-DB provide superior incorporation rates of ffN compared to ffN sets using propargylamine or propargylamide linker moieties (also known as ffN-PA or ffN-(PA)) described herein. For example, an ffNs-DB set using an allylamine or allylamide linker moiety and a 3'-AOM blocking group described herein can result in at least a 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 300%, 400%, or 500% improvement in incorporation rate, thereby improving phasing values, compared to an ffNs-PA set using a 3'-O-azidomethyl blocking group under the same conditions for the same period of time. In other embodiments, incorporation rate / speed is measured by the surface reaction rate Vmax on the surface of a substrate (e.g., a flow cell or cBot system). For example, an ffNs-DB set having a 3'-AOM blocking group exhibits a Vmax value (ms ) that is at least 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 300%, 400%, or 500% higher than an ffNs-PA set having a 3'-O-azidomethyl blocking group under the same conditions for the same period of time. -1In some embodiments, the incorporation rate / speed is measured at ambient or subambient temperatures (e.g., 4-10°C). In other embodiments, the incorporation rate / speed is measured at elevated temperatures, such as 40°C, 45°C, 50°C, 55°C, 60°C, or 65°C. In some such embodiments, the incorporation rate / speed is measured in a basic pH environment, for example, in a solution at pH 9.0, 9.2, 9.4, 9.6, 9.8, or 10.0. In some such embodiments, the incorporation rate / speed is measured in the presence of an enzyme, such as a polymerase (e.g., DNA polymerase), terminal deoxynucleotidyl transferase, or reverse transcriptase. In some embodiments, the ffN-DB is the ffT-DB, ffC-DB, or ffA-DB. In one embodiment, the ffN-DB set having the improved phasing values ​​described herein includes the ffT-DB, ffC, ffA, and ffG. In another embodiment, the ffN-DB set having the improved phasing values ​​described herein includes ffT-DB, ffC-DB, ffA, and ffG. In yet another embodiment, the ffN-DB set having the improved phasing values ​​described herein includes ffT-DB, ffC-DB, ffA-DB, and ffG.

[0090] In some further embodiments, when the nucleobase of a nucleotide described herein is thymine or an optionally substituted derivative or analog thereof (i.e., the nucleotide is T), L 1 comprises an allylamine or allylamide moiety, or an optionally substituted variant thereof. In a particular example, L 1 teeth, [ka] and * is the L for the C5 position of the thymine base 1 In some embodiments, the T nucleotide described herein is attached directly to the C5 position of the thymine base (i.e., ffT-DB). [ka] The ffT-DB is a fully functionalized T nucleotide (ffT) labeled with a dye molecule via a cleavable linker, comprising: In some examples, when the ffT-DB is used for sequencing applications in the presence of a palladium catalyst, it can substantially improve sequencing metrics such as phasing, prephasing, and error rate. For example, when an ffT-DB having a 3'-AOM blocking group as described herein is used, it can result in at least 50%, 100%, 200%, 300%, 400%, 500%, 600%, 700%, 800%, 900%, 1000%, 1500%, 2000%, 2500%, or 3000% improvement in one or more sequencing metrics described herein compared to when a standard ffT-PA having a 3'-O-azidomethyl blocking group is used.

[0091] Some further embodiments of the nucleosides or nucleotides described herein include those of formula (Ia), (Ia'), (Ib), (Ic), (Ic'), or (Id). [ka] [ka]

[0092] In some further embodiments of the nucleosides or nucleotides described herein, L 2 exists, and L 2 teeth, [ka] wherein each of n and m is independently an integer of 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10, and the phenyl moiety is optionally substituted. In some such embodiments, n is 5 and L 2 is unsubstituted. In some further embodiments, m is 4.

[0093] In any embodiment of the nucleosides or nucleotides described herein, a cleavable linker or L 1 / L 2 is a disulfide moiety or an azide moiety ( [ka] etc.), or combinations thereof. Further non-limiting examples of linker moieties include L 1 or L 2 can be incorporated into [ka] Additional linker moieties are disclosed in WO 2004 / 018493 and U.S. Patent Application Publication No. 2016 / 0040225, which are incorporated herein by reference.

[0094] In any embodiment of the nucleoside or nucleotide described herein, the nucleoside or nucleotide includes a 2' deoxyribose moiety (i.e., R 4 is formula (I) and (Ia)-(Id) are H). In some further embodiments, the 2' deoxyribose contains one, two, or three phosphate groups at the 5' position of the sugar ring. In some further embodiments, the nucleotides described herein are nucleotide triphosphates (i.e., R in formulas (I) and (Ia)-(Id) are H). 6 forms a triphosphate).

[0095] In any of the embodiments of the nucleosides or nucleotides described herein, the detectable label can comprise a fluorescent dye.

[0096] Further embodiments of the present disclosure relate to oligonucleotides or polynucleotides comprising the nucleosides or nucleotides described herein. For example, oligonucleotides or polynucleotides incorporating a nucleotide of formula (Ia') may have the following structure: [ka] In some such embodiments, the oligonucleotide or polynucleotide hybridizes to a template or target polynucleotide. In some such embodiments, the template polynucleotide is immobilized on a solid support.

[0097] Further embodiments of the present disclosure relate to solid supports comprising an array of a plurality of immobilized templates or target polynucleotides, at least a portion of which are hybridized to oligonucleotides or polynucleotides comprising nucleosides or nucleotides described herein.

[0098] In any embodiment of the nucleotides or nucleosides described herein, the 3'-OH blocking group and the cleavable linker (and attached label) may be removable under the same or substantially the same chemical reaction conditions, e.g., the 3'-OH blocking group and the detectable label may be removed in a single chemical reaction. In other embodiments, the 3'-OH blocking group and the detectable label are removed in two separate steps.

[0099] In some embodiments, the 3'-blocked nucleotides or nucleosides described herein provide superior stability during storage in solution or lyophilized form, or during reagent handling during sequencing applications, compared to the same nucleotides or nucleosides protected with standard 3'-OH blocking groups disclosed in the prior art, e.g., 3'-O-azidomethyl protecting groups. For example, the acetal blocking groups disclosed herein can confer improved stability of at least 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 300%, 400%, 500%, 600%, 700%, 800%, 900%, 1000%, 1500%, 2000%, 2500%, or 3000% compared to azidomethyl-protected 3'-OH over the same conditions and for the same period of time, thereby reducing prephasing and increasing sequencing read lengths. In some embodiments, stability is measured at ambient or subambient temperatures (e.g., 4-10°C). In other embodiments, stability is measured at elevated temperatures, such as 40°C, 45°C, 50°C, 55°C, 60°C, or 65°C. In some such embodiments, stability is measured in a basic pH environment, e.g., a solution at pH 9.0, 9.2, 9.4, 9.6, 9.8, or 10.0. In some such embodiments, stability is measured with or without the presence of an enzyme, such as a polymerase (e.g., a DNA polymerase), terminal deoxynucleotidyl transferase, or reverse transcriptase.

[0100] In some embodiments, the 3'-blocked nucleotides or nucleosides described herein provide superior deblocking rates in solution during the chemical cleavage step of sequencing applications compared to the same nucleotide or nucleoside protected with a standard 3'-OH blocking group disclosed in the prior art, e.g., a 3'-O-azidomethyl protecting group. For example, the acetal blocking groups disclosed herein may confer improved deblocking rates of at least 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 150%, 200%, 300%, 400%, 500%, 600%, 700%, 800%, 900%, 1000%, 1500%, or 2000% compared to an azidomethyl-protected 3'-OH using a standard deblocking reagent (such as tris(hydroxypropyl)phosphine), thereby reducing the overall time for a sequencing cycle. In some embodiments, the deblocking rate is measured at ambient or subambient temperatures (e.g., 4-10°C). In other embodiments, the deblocking rate is measured at elevated temperatures, such as 40°C, 45°C, 50°C, 55°C, 60°C, or 65°C. In some such embodiments, the deblocking rate is measured in a basic pH environment, e.g., a solution at pH 9.0, 9.2, 9.4, 9.6, 9.8, or 10.0. In some such embodiments, the molar ratio of deblocking reagent to substrate (i.e., 3'-blocked nucleoside or nucleotide) is about 10:1, about 5:1, about 2:1, or about 1:1.

[0101] In some embodiments, a palladium deblocking reagent (e.g., Pd(0)) is used to remove the 3' acetal blocking group (e.g., AOM blocking group). Pd forms a chelation complex with the two oxygen atoms of the AOM group and the double bond of the allyl group, allowing for removal of the deblocking reagent in close proximity to the functional group, which can accelerate the deblocking rate. For example, after cleavage of the Pd of a linker described herein and the 3' blocking group of an incorporated nucleotide having formula (Ia), (Ia'), (Ib), (Ic), (Ic'), or (Id), the remaining linker construct on the copy polynucleotide can comprise the following structure: [ka] The squiggle indicates the attachment of oxygen to the phosphodiester bond of the remaining copy polynucleotide strand. For example, [ka] It included: The arylamide or propargylamide moiety can be further cleaved by Pd catalysis. Furthermore, the remaining linker construct attached to the detectable label has the following structure: [ka] for example, [ka]

[0102] Cleavage conditions for cleavable linkers The cleavable linkers described herein can be removed or cleaved under a variety of chemical conditions. Non-limiting cleavage conditions include palladium catalysts such as Pd(II) complexes (e.g., Pd(OAc), allylPd(II) chloride dimer [(allyl)PdCl], or NaPdCl) in the presence of a water-soluble phosphine ligand, such as tris(hydroxypropyl)phosphine (THP or THPP) or tris(hydroxymethyl)phosphine (THMP). In some embodiments, the 3' acetal blocking group can be cleaved under the same or substantially the same cleavage conditions as those of the cleavable linker.

[0103] Palladium catalyst In some embodiments, the 3' acetal blocking groups and cleavable linkers described herein can be cleaved by a palladium catalyst. In some such embodiments, the Pd catalyst is water-soluble. In some such embodiments, the Pd catalyst is a Pd(0) complex (e.g., tris(3,3',3"-phosphinidinetris(benzenesulfonato)palladium(0)) salt nonasodium nonahydrate). In some cases, Pd(0) can be generated in situ from the reduction of a Pd(II) complex with a reagent such as an alkene, alcohol, amine, phosphine, or metal hydride. Suitable palladium sources include Pd(CHCN)Cl, [PdCl(allyl)], [Pd(allyl)(THP)]Cl, [Pd(allyl)(THP)]Cl, Pd(OAc), Pd(PPh), Pd(dba), Pd(Acac), PdCl(COD), and Pd(TFA). In one such embodiment, the Pd(0) complex is N In another embodiment, the palladium source is allylpalladium(II) chloride dimer [(allyl)PdCl] or [PdCl(C3H5)]2. In some embodiments, the Pd(0) catalyst is generated in aqueous solution by mixing a Pd(II) complex with a phosphine. Suitable phosphines include water-soluble phosphines such as tris(hydroxypropyl)phosphine (THP), tris(hydroxymethyl)phosphine (THMP), 1,3,5-triaza-7-phosphaadamantane (PTA), bis(p-sulfonatophenyl)phenylphosphine dihydrate potassium salt, tris(carboxyethyl)phosphine (TCEP), and triphenylphosphine-3,3',3"-trisulfonic acid trisodium salt.

[0104] In some embodiments, the palladium catalyst is prepared in situ by mixing [(allyl)PdCl] with THP. The molar ratio of [(allyl)PdCl] to THP can be about 1:2, 1:3, 1:4, 1:5, 1:6, 1:7, 1:8, 1:9, or 1:10. In one embodiment, the molar ratio of [(allyl)PdCl] to THP is 1:10. In some other embodiments, the palladium catalyst is prepared in situ by mixing the water-soluble Pd reagent NaPdCl with THP. The molar ratio of NaPdCl to THP can be about 1:2, 1:3, 1:4, 1:5, 1:6, 1:7, 1:8, 1:9, or 1:10. In one embodiment, the molar ratio of NaPdCl to THP is about 1:3. In another embodiment, the molar ratio of NaPdCl to THP is about 1:3.5. In some further embodiments, one or more reducing agents, such as ascorbic acid or a salt thereof (e.g., sodium ascorbate), may be added. In some embodiments, the cleavage mixture may contain an additional buffering reagent, such as a primary amine, a secondary amine, a tertiary amine, a natural amino acid, an unnatural amino acid, a carbonate, a phosphate, or a borate, or a combination thereof. In some further embodiments, the buffering reagent comprises ethanolamine (EA), tris(hydroxymethyl)aminomethane (Tris), glycine, sodium carbonate, sodium phosphate, sodium borate, dimethylethanolamine (DMEA), diethylethanolamine (DEEA), N,N,N',N'-tetramethylethylenediamine (TMEDA), N,N,N',N'-tetraethylethylenediamine (TEEDA), or 2-piperidineethanol, or a combination thereof. In one embodiment, the one or more buffering reagents comprises DEEA. In another embodiment, the one or more buffer reagents contain one or more inorganic salts, such as carbonate, phosphate, or borate, or a combination thereof. In one embodiment, the inorganic salt is a sodium salt.

[0105] In other embodiments, the cleavage conditions of the cleavable linker are different from the cleavage conditions of the 3'-OH blocking group. For example, if the 3'-blocking group is 3'-O-azidomethyl, the -CH2N3 moiety can be converted to an amino group with a phosphine. Alternatively, the azido group in -CH2N3 can be converted to an amino group by contacting such a molecule with a thiol, particularly a water-soluble thiol such as dithiothreitol (DTT). In one embodiment, the phosphine is THP.

[0106] Compatibility with linearization In order to maximize the throughput of nucleic acid sequencing reactions, it is advantageous to be able to sequence multiple template molecules in parallel. Parallel processing of multiple templates can be achieved by using nucleic acid sequencing technology. These arrays typically consist of a high-density matrix of polynucleotides immobilized on a solid support material.

[0107] Both WO 98 / 44151 and WO 00 / 18957 describe methods of nucleic acid amplification that allow amplification products to be immobilized on a solid support to form arrays consisting of clusters or "colonies" formed from a plurality of identical immobilized polynucleotide strands and a plurality of identical immobilized complementary strands. This type of array is referred to herein as a "clustered array." Nucleic acid molecules present in DNA colonies on clustered arrays prepared according to these methods can provide templates for sequencing reactions, such as those described in WO 98 / 44152. The product of solid-phase amplification reactions such as those described in WO 98 / 44151 and WO 00 / 18957 is a so-called "bridged" structure formed by annealing pairs of immobilized polynucleotide strands and immobilized complementary strands, with both strands attached to the solid support at their 5' ends. To provide a template more suitable for nucleic acid sequencing, it is preferable to remove substantially all or at least a portion of one of the immobilized strands of the "bridged" structure to generate a template that is at least partially single-stranded. Thus, the portion of the template that is single-stranded is available for hybridization to a sequencing primer. The process of removing all or part of one immobilized strand in a "bridged" double-stranded nucleic acid structure is called "linearization." Linearization can be achieved by a variety of methods, including, but not limited to, enzymatic cleavage, photochemical cleavage, or chemical cleavage. Non-limiting examples of linearization methods are disclosed in International Publication No. 2007 / 010251, U.S. Patent Application Publication No. 2009 / 0088327, U.S. Patent Application Publication No. 2009 / 0118128, and U.S. Patent Application Publication No. 2019 / 0352327, which are incorporated by reference in their entireties.

[0108] Specifically, amplification (e.g., bridge amplification or exclusion amplification) forms an array consisting of clusters or "colonies" formed from multiple identical immobilized target polynucleotide strands and multiple identical immobilized complementary strands. The target strand and complementary strand form an at least partially double-stranded polynucleotide complex, with both strands immobilized at their 5' ends to a solid support. To generate at least partially single-stranded templates, the double-stranded polynucleotide is contacted with an aqueous solution of a palladium catalyst, which cleaves one strand at a cleavage site containing an allyl-modified nucleoside (e.g., an allyl-modified T nucleoside), removing at least a portion of one of the immobilized strands. Thus, the portion of the single-stranded template is available for hybridization to a sequencing primer to initiate the first round of SBS (Read 1). In some embodiments, the allyl-modified nucleoside is located in the P5 primer sequence. This method is referred to as first chemical linearization, as opposed to standard enzymatic linearization, in which such removal or cleavage is facilitated by an enzymatic cleavage reaction using the enzyme USER to cleave the U position on the P5 primer.

[0109] In some embodiments, the conditions for cleaving the cleavable linker and / or deprotecting or removing the 3'-OH blocking group are also compatible with the linearization process. In some further embodiments, such cleavage conditions are compatible with chemical linearization processes involving the use of Pd complexes and phosphines. In some embodiments, the Pd complex is a Pd(II) complex (e.g., Pd(OAc)2, [(Allyl)PdCl]2, or Na2PdCl4) that generates Pd(0) in situ in the presence of a phosphine (e.g., THP). A chemical linearization process using a Pd catalyst to cleave allyl-modified T nucleosides in a P5 primer sequence is described in detail in U.S. Patent Application Publication No. 2019 / 0352327, the entire contents of which are incorporated by reference. In further embodiments, the Pd cleavage mixture disclosed herein (e.g., [Pd(Allyl)Cl)2 and THP in a buffer solution containing DEEA) can be used directly in the first actinic linearization step. The reduction in the number of reagents allows for further simplification of equipment (fluids and cartridges).

[0110] Unless otherwise indicated, reference to a nucleotide is intended to be applicable to a nucleoside as well.

[0111] Labeled nucleotides According to one embodiment of the present disclosure, the described 3'-OH blocked nucleotide also includes a detectable label, and such a nucleotide is referred to as a labeled nucleotide or a fully functionalized nucleotide (ffN). The label (e.g., a fluorescent dye) is conjugated via a cleavable linker by various means, including hydrophobic attraction, ionic attraction, and covalent bonding. In some embodiments, the dye is conjugated to the nucleotide by a covalent bond via the cleavable linker. In some cases, such a labeled nucleotide is also referred to as a "modified nucleotide." Those skilled in the art will understand that a label can be covalently attached to a linker by reacting a functional group (e.g., carboxyl) of the label with a functional group (e.g., amino) of the linker.

[0112] Labeled nucleosides and nucleotides are useful for labeling polynucleotides formed by enzymatic synthesis, non-limiting examples of which include PCR amplification, isothermal amplification, solid-phase amplification, polynucleotide sequencing (e.g., solid-phase sequencing), nick-translation reactions, and the like.

[0113] In some embodiments, the dye may be covalently attached to the oligonucleotide or nucleotide via the nucleotide base. For example, a labeled nucleotide or oligonucleotide may have a label attached to the C5 position of a pyrimidine base or a label attached via a cleavable linker moiety to the C7 position of a 7-deazapurine base.

[0114] Unless otherwise indicated, reference to nucleotides is intended to be applicable to nucleosides as well. This application is also further described with reference to DNA, but the description is also applicable to RNA, PNA, and other nucleic acids unless otherwise indicated.

[0115] Nucleosides and nucleotides may be labeled at sites on the sugar or nucleobase. As known in the art, a "nucleotide" consists of a nitrogenous base, a sugar, and one or more phosphate groups. In RNA, the sugar is ribose, while in DNA, the sugar is deoxyribose, i.e., a sugar lacking the hydroxyl group present in ribose. The nitrogenous base is a derivative of a purine or pyrimidine. Purines are adenine (A) and guanine (G), and pyrimidines are cytosine (C) and thymine (T), or, in the context of RNA, uracil (U). The C-1 atom of deoxyribose is linked to the N-1 atom of a pyrimidine or the N-9 atom of a purine. Nucleotides are also phosphate esters of nucleosides, esterified to the hydroxyl group attached to C-3 or C-5 of the sugar. Nucleosides are usually monophosphates, diphosphates, or triphosphates.

[0116] A "nucleoside" is structurally similar to a nucleotide, but lacks the phosphate moiety. An example of a nucleotide analog is one in which the label is linked to the base and there is no phosphate group attached to the sugar molecule.

[0117] While bases are typically referred to as purines or pyrimidines, those skilled in the art will appreciate that derivatives and analogs are available that do not alter the ability of a nucleotide or nucleoside to undergo Watson-Crick base pairing. A "derivative" or "analog" refers to a compound or molecule whose core structure is the same as or closely similar to the parent compound, but that has chemical or physical modifications, such as different or additional side groups that allow the derivative nucleotide or nucleoside to be linked to other molecules. For example, the base may be a deazapurine. In certain embodiments, the derivative should be able to undergo Watson-Crick pairing. "Derivative" and "analog" also include synthetic nucleotide or nucleoside derivatives, for example, with modified base moieties and / or modified sugar moieties. Such derivatives and analogs are discussed, for example, in Scheit, Nucleotide Analogs (John Wiley & Son, 1980) and Uhlman et al., Chemical Reviews 90:543-584, 1990. Nucleotide analogs can also contain modified phosphodiester linkages, including phosphorothioate, phosphorodithioate, alkyl-phosphonate, phosphoranilidate, phosphoramidite linkages, and the like.

[0118] In certain embodiments, the labeled nucleoside or nucleotide may be enzymatically incorporable and enzymatically extendable. Thus, the linker portion may be long enough to connect the nucleotide to the compound so that the compound does not significantly interfere with the overall binding and recognition of the nucleotide by the nucleic acid replication enzyme. Thus, the linker may also include a spacer unit. The spacer may, for example, distance the nucleotide base from the cleavage site or label.

[0119] The present disclosure also encompasses polynucleotides incorporating dye compounds. Such polynucleotides may be DNA or RNA composed of deoxyribonucleotides or ribonucleotides, respectively, joined by phosphodiester linkages. Polynucleotides can include naturally occurring nucleotides, non-naturally occurring (or modified) nucleotides other than the labeled nucleotides described herein, or any combination thereof, in combination with at least one modified nucleotide as described herein (e.g., labeled with a dye compound). Polynucleotides according to the present disclosure can also include non-naturally occurring backbone linkages and / or non-nucleotide chemical modifications. Chimeric structures composed of a mixture of ribonucleotides and deoxyribonucleotides containing at least one labeled nucleotide are also contemplated.

[0120] Non-limiting exemplary labeled nucleotides described herein include: [ka] where L is a cleavable linker (L as described herein). 2 wherein R represents a ribose or deoxyribose moiety as defined above, or a ribose or deoxyribose moiety having the 5' position substituted with one, two, or three phosphates.

[0121] In some embodiments, non-limiting fluorescent dye conjugates are shown below: [ka] [ka] wherein PG represents a 3'-OH blocking group as described herein, n is an integer of 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10, and k is 0, 1, 2, 3, 4, or 5. In one embodiment, -O-PG is AOM. In another embodiment, -O-PG is -O-azidomethyl. In one embodiment, n is 5. [ka] refers to the point of attachment of the dye with a cleavable linker as a result of the reaction between an amino group on the linker moiety and a carboxyl group on the dye.

[0122] Sequencing methods The labeled nucleotides or nucleosides according to the present disclosure can be used in any analytical method involving the detection of a fluorescent label attached to the nucleotide or nucleoside, regardless of whether the nucleotide or nucleoside is analyzed by itself or after it is incorporated into or associated with a larger molecular structure or conjugate. In this context, the term "incorporated into a polynucleotide" can mean that the 5' phosphate is phosphodiester-linked to the 3'-OH group of a second (modified or unmodified) nucleotide, which itself may form part of a longer polynucleotide chain. The 3' end of the nucleotide described herein may or may not be phosphodiester-linked to the 5' phosphate of a further (modified or unmodified) nucleotide. Thus, in one non-limiting embodiment, the present disclosure provides a method for detecting a nucleotide incorporated into a polynucleotide, comprising: (a) incorporating at least one nucleotide of the present disclosure into a polynucleotide; and (b) detecting the nucleotide(s) incorporated into the polynucleotide by detecting a fluorescent signal from a detectable label (e.g., a fluorescent compound) attached to the nucleotide(s). The method may comprise: (a) a synthesis step in which one or more nucleotides according to the present disclosure are incorporated into the polynucleotide; and (b) a detection step in which one or more nucleotides incorporated into the polynucleotide are detected by detecting or quantitatively measuring their fluorescence.

[0123] Additional aspects of the present disclosure include methods of preparing a growing polynucleotide complementary to a target single-stranded polynucleotide in a sequencing reaction, comprising incorporating a nucleotide as described herein into the growing complementary polynucleotide, wherein the incorporation of the nucleotide prevents the introduction of any subsequent nucleotide into the growing complementary polynucleotide.

[0124] Some embodiments of the present disclosure relate to a method for determining the sequence of a target single-stranded polynucleotide, comprising: (a) a 3'-OH blocking group as described herein [ka] incorporating a nucleotide (e.g., dATP, dCTP, dGTP, dTTP, or dUTP) containing a nucleotide (attached to the 3' oxygen) and a detectable label as described herein into a copy polynucleotide strand that is complementary to at least a portion of a target polynucleotide strand; (b) detecting the identity of the nucleotide incorporated into the copy polynucleotide strand; (c) chemically removing the label and the 3'-OH blocking group from the nucleotide incorporated into the copy polynucleotide strand; Includes.

[0125] In some aspects, the sequencing method further comprises (d) washing the chemically removed label and 3'-OH blocking group away from the copy polynucleotide strand by using a post-cleavage wash solution. In some such embodiments, the 3'-OH blocking group and detectable label are removed before introducing the next complementary nucleotide. In some embodiments, washing step (d) also removes unincorporated nucleotides. In other embodiments, the method may comprise a separate wash step prior to step (b) to wash unincorporated nucleotides away from the copy polynucleotide strand.

[0126] In some embodiments, steps (a)-(d) are repeated until the sequence of a portion of the target polynucleotide strand is determined, hi some such embodiments, steps (a)-(d) are repeated at least 50 times, at least 75 times, at least 100 times, at least 150 times, at least 200 times, at least 250 times, or at least 300 times.

[0127] Incorporation mixture In some embodiments of the methods described herein, step (a), also referred to as the incorporation step, comprises contacting a mixture comprising one or more nucleotides (e.g., dATP, dCTP, dGTP, and dTTP or dUTP) with the copy polynucleotide / target polynucleotide complex in an incorporation solution comprising a polymerase and one or more buffers. In some such embodiments, the polymerase is a DNA polymerase, e.g., Pol812, Pol1901, Pol1558, or Pol963. The amino acid sequences of Pol812, Pol1901, Pol1558, or Pol963 DNA polymerases are described, for example, in U.S. Patent Application Publication Nos. 2020 / 0131484 and 2020 / 0181587, both of which are incorporated herein by reference. In some embodiments, the one or more buffers comprise a primary amine, a secondary amine, a tertiary amine, a natural amino acid, or an unnatural amino acid, or a combination thereof. In further embodiments, the buffer comprises ethanolamine or glycine, or a combination thereof. In one embodiment, the buffer comprises or is glycine. In some embodiments, the use of glycine in the incorporation mixture can improve phasing values ​​compared to a standard buffer, such as ethanolamine (EA), under the same conditions. For example, the use of glycine provides a reduction or decrease in phasing values ​​of at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 300%, 400%, or 500% compared to ethanolamine used under the same conditions. In some cases, the use of glycine provides a % phasing value of less than about 0.15%, 0.1%, 0.09%, 0.08%, 0.07%, 0.06%, 0.05%, 0.04%, 0.03%, 0.02%, or 0.01% in an SBS sequencing run of at least 50 cycles. In a further embodiment, the use of glycine provides a % phasing value of less than about 0.08% in Read 1 of an SBS sequencing run of at least 150 cycles.

[0128] cutting mixture In some embodiments of the methods described herein, step (c), also referred to as the cleavage step, comprises contacting the incorporated nucleotides and copy polynucleotide strand with a cleavage solution comprising a palladium catalyst described herein. In some such embodiments, the 3'-OH blocking group and the detectable label are removed in a single step of the reaction. In one such embodiment, the 3'-blocking group is AOM and the cleavable linker comprises an AOL moiety, both of which are removed or cleaved in a single step of the chemical reaction. In some further embodiments, the cleavage solution (also referred to as a cleavage mix) comprises a Pd catalyst described herein.

[0129] In some further embodiments, the Pd catalyst is a Pd(0) catalyst. In some such embodiments, Pd(0) is prepared in situ by mixing a Pd(II) reagent with one or more phosphine ligands. In some such embodiments, the palladium catalyst can be prepared in situ by mixing [(allyl)PdCl] with THP. The molar ratio of [(allyl)PdCl] to THP can be about 1:2, 1:3, 1:4, 1:5, 1:6, 1:7, 1:8, 1:9, or 1:10. In one embodiment, the molar ratio of [(allyl)PdCl] to THP is 1:10 (i.e., the molar ratio of Pd:THP is 1:5). In some other embodiments, the palladium catalyst can be prepared in situ by mixing the aqueous Pd(II) reagent NaPdCl with THP. The molar ratio of Na2PdCl4 to THP can be about 1:2, 1:3, 1:4, 1:5, 1:6, 1:7, 1:8, 1:9, or 1:10. In one embodiment, the molar ratio of Na2PdCl4 to THP is about 1:3. In another embodiment, the molar ratio of Na2PdCl4 to THP is about 1:3.5. Other non-limiting examples of Pd catalysts include Pd(CH3CN)2Cl2.

[0130] In some further embodiments, one or more reducing agents, such as ascorbic acid or a salt thereof (e.g., sodium ascorbate), may be added. In some embodiments, the cleavage solution may include one or more buffer reagents, such as a primary amine, a secondary amine, a tertiary amine, a carbonate, a phosphate, or a borate, or a combination thereof. In some further embodiments, the buffer reagent includes ethanolamine (EA), tris(hydroxymethyl)aminomethane (Tris), glycine, sodium carbonate, sodium phosphate, sodium borate, dimethylethanolamine (DMEA), diethylethanolamine (DEEA), N,N,N',N'-tetramethylethylenediamine (TEMED), N,N,N',N'-tetraethylethylenediamine (TEEDA), or 2-piperidineethanol, or a combination thereof. In one embodiment, the buffer reagent includes or is DEEA. In another embodiment, the buffer reagent contains one or more inorganic salts, such as a carbonate, a phosphate, or a borate, or a combination thereof. In one embodiment, the inorganic salt is a sodium salt. In a further embodiment, the cleavage solution comprises a palladium (Pd) catalyst (e.g., [(allyl)PdCl] / THP or NaPdCl / THP) and one or more buffering reagents described herein (e.g., a tertiary amine such as DEEA), and has a pH of about 9.0 to about 10.0 (e.g., 9.6 or 9.8).

[0131] In other embodiments, the label and the 3'-blocking group are removed in two separate chemical reactions. In some cases, removing the label from the nucleotide incorporated into the copy polynucleotide strand comprises contacting the copy strand containing the incorporated nucleotide with a first cleavage solution containing a Pd catalyst described herein. In some cases, removing the 3'-OH blocking group from the nucleotide incorporated into the copy polynucleotide strand comprises contacting the copy strand containing the incorporated nucleotide with a second cleavage solution. In some such embodiments, the second cleavage solution contains one or more phosphines, such as trialkylphosphines. Non-limiting examples of trialkylphosphines include tris(hydroxypropyl)phosphine (THP), tris-(2-carboxyethyl)phosphine (TCEP), tris(hydroxymethyl)phosphine (THMP), or tris(hydroxyethyl)phosphine (THEP). In one embodiment, the 3'-OH blocking group is 3'-O-azidomethyl, and the second cleavage solution contains THP.

[0132] In some embodiments, the cleavage solutions described herein can also be used in conventional chemical linearization processes described herein. In particular, chemical linearization of clustered polynucleotides in preparation for sequencing is accomplished by palladium-catalyzed cleavage of one or more first strands of double-stranded polynucleotides immobilized on a solid support, thus generating single-stranded (or at least partially single-stranded) templates available for hybridization to a sequencing primer and subsequent sequencing applications (e.g., first-round sequencing by synthesis (Read 1)). In some embodiments, each double-stranded polynucleotide comprises a first strand and a second strand. The first strand is generated by extending a first extension primer immobilized on a solid support. In some embodiments, the first strand comprises a cleavage site that can be cleaved by a palladium complex (e.g., a Pd(0) complex). In certain embodiments, the cleavage site is located in the first extension primer portion of the first strand. In further embodiments, the cleavage site comprises a thymine nucleoside or nucleotide analog having an allyl functionality. In some embodiments of the methods described herein, the target single-stranded polynucleotide is formed by chemically cleaving a complementary strand from a double-stranded polynucleotide. In further embodiments, both the complementary strand in the duplex and the target polynucleotide are immobilized on a solid support at their 5' ends. In some further embodiments, the chemical cleavage of the complementary strand is carried out under the same reaction conditions as those used to chemically remove the detectable label and 3'-OH blocking group from the nucleotide incorporated into the copy polynucleotide strand (i.e., step (c) of the methods described herein). In one embodiment, the first chemical linearization utilizes the same cleavage mix described herein. Palladium (Pd) scavenger

[0133] Pd, most often in its inactive Pd(II) form, has the ability to adhere to DNA, potentially interfering with the binding between DNA and polymerase, leading to increased phasing. A post-cleavage wash composition containing a Pd scavenger compound can be used following the deblocking step. For example, International Publication No. WO 2020 / 126593 discloses Pd scavengers such as 3,3'-dithiodipropionic acid (DDPA) and lipoic acid (LA), which may be included in the scanning composition and / or post-cleavage wash composition. The use of these scavengers in the post-cleavage wash solution aims to scavenge Pd(0) and convert it to the inactive Pd(II) form, thereby improving prephasing values ​​and sequencing metrics, reducing signal degradation, and extending sequencing read lengths.

[0134] In some embodiments of the methods described herein, step (a) of the method comprises contacting nucleotides with the copy polynucleotide strand in an incorporation solution comprising a polymerase, at least one palladium scavenger, and one or more buffers. In some embodiments, the Pd scavenger in the incorporation solution is a Pd(0) scavenger. In some such embodiments, the Pd scavenger is selected from the group consisting of -O-allyl, -S-allyl, -NR-allyl, and -N + and R'-aryl, wherein R is H, unsubstituted or substituted C1-C6 alkyl, unsubstituted or substituted C2-C6 alkenyl, unsubstituted or substituted C2-C6 alkynyl, unsubstituted or substituted C6-C 10 Aryl, unsubstituted or substituted 5-10 membered heteroaryl, unsubstituted or substituted C3-C 10 carbocyclyl, or unsubstituted or substituted 5-10 membered heterocyclyl, and R' is H, unsubstituted C1-C6 alkyl, or substituted C1-C6 alkyl.

[0135] In some such embodiments, the Pd(0) scavenger in the incorporation solution comprises one or more -O-allyl moieties. In some further embodiments, the Pd(0) scavenger is [ka] or a combination thereof. Alternative Pd(0) scavengers are disclosed in US Patent No. 63 / 190983, which is incorporated by reference in its entirety.

[0136] In some embodiments, the concentration of the Pd(0) scavenger comprising one or more allyl moieties in the incorporation solution is about 0.1 mM to about 100 mM, 0.2 mM to about 75 mM, about 0.5 mM to about 50 mM, about 1 mM to about 20 mM, or about 2 mM to about 10 mM. In further embodiments, the concentration of the Pd(0) scavenger is about 0.5 mM, 1 mM, 1.5 mM, 2 mM, 2.5 mM, 3 mM, 3.5 mM, 4 mM, 4.5 mM, 5 mM, 5.5 mM, 6 mM, 6.5 mM, 7 mM, 7.5 mM, 8 mM, 8.5 mM, 9 mM, 9.5 mM, 10 mM, 12.5 mM, 15 mM, 17.5 mM, or 20 mM. In a further embodiment, the pH of the incorporation solution is about 9-10.

[0137] In some embodiments, the molar ratio of the palladium catalyst (in the starting solution) to the palladium scavenger comprising one or more allyl moieties is about 1:100, 1:50, 1:20, 1:10, or 1:5.

[0138] In some other embodiments of the methods described herein, the Pd(0) scavenger comprises one or more allyl moieties described herein in the scan solution used in step (b) when performing one or more fluorescence measurements to detect the identity of the incorporated nucleotide in the copy polynucleotide. In still other embodiments, the Pd(0) scavenger comprises one or more allyl moieties that can be present in both the incorporation solution and the scan solution.

[0139] In some further embodiments of the methods described herein, a post-cleavage wash step is used after labeling to remove the 3' blocking group. In some such embodiments, one or more palladium scavengers are also used in the wash step after labeling and cleavage of the 3' blocking group. In some further embodiments, the one or more Pd scavengers in the post-cleavage wash solution include a Pd(II) scavenger. In some such embodiments, the palladium scavenger includes an isocyanoacetic acid (ICNA) salt, cysteine ​​or a salt thereof, or a combination thereof. In one embodiment, the palladium scavenger includes or is potassium isocyanoacetate or sodium isocyanoacetate. In another embodiment, the palladium scavenger includes or is cysteine ​​or a salt thereof (e.g., L-cysteine ​​or L-cysteine ​​HCl salt). Other non-limiting examples of palladium scavengers in the post-cleavage wash solution may include ethyl isocyanoacetate, methyl isocyanoacetate, N-acetyl-L-cysteine, potassium ethylxanthate (PEX or KS-C(=S)-OEt), potassium isopropylxanthate, glutathione, lipoic acid, ethylenediaminetetraacetic acid (EDTA), iminodiacetic acid, nitrilodiacetic acid, trimercapto-S-triazine, dimethyldithiocarbamate, dithiothreitol, mercaptoethanol, allyl alcohol, propargyl alcohol, thiols, tertiary amines and / or tertiary phosphines, or combinations thereof.

[0140] In a further embodiment, the concentration of the Pd(II) scavenger, such as L-cysteine, in the post-cleavage washing solution is about 0.1 mM to about 100 mM, 0.2 mM to about 75 mM, about 0.5 mM to about 50 mM, about 1 mM to about 20 mM, or about 2 mM to about 10 mM. In a further embodiment, the concentration of the Pd(II) scavenger, such as L-cysteine, is about 1 mM, 2 mM, 3 mM, 4 mM, 5 mM, 6 mM, 6.5 mM, 7 mM, 8 mM, 9 mM, 10 mM, 12.5 mM, 15 mM, 17.5 mM, or 20 mM. In one embodiment, the concentration of the Pd scavenger, such as L-cysteine ​​or a salt thereof, in the post-cleavage washing solution is about 10 mM.

[0141] In some other embodiments of the methods described herein, all Pd scavengers (e.g., both Pd(0) and Pd(II) scavengers) are present in the incorporation solution and / or the scan solution, and the method does not include a specific post-cleavage washing step to remove any traces of remaining Pd species.

[0142] In some embodiments of the methods described herein, the use of a Pd scavenger (e.g., a Pd(0) scavenger having one or more allyl moieties) may reduce the prephasing value by at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 300%, 400%, 500%, 600%, 700%, 800%, 900%, or 1000% compared to the same sequencing performed under the same conditions without the use of a palladium scavenger. In some such embodiments, the Pd(0) scavenger may reduce the sequencing prephasing value to less than about 0.1%, 0.09%, 0.08%, 0.07%, 0.06%, 0.05%, 0.04%, 0.03%, 0.02%, or 0.01% in an SBS sequencing run of at least 50 cycles. In some embodiments, the prephasing value refers to the value measured after 50, 75, 100, 125, 150, 200, 250, or 300 cycles.

[0143] In some further embodiments, a palladium scavenger (e.g., a Pd(II) scavenger such as L-cysteine ​​or a salt thereof) may reduce prephasing or phasing values ​​by at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 300%, 400%, 500%, 600%, 700%, 800%, 900%, or 1000% compared to the same sequencing performed under the same conditions without the use of a palladium scavenger. In some such embodiments, the use of Pd scavengers provides a % phasing value of less than about 0.2%, 0.15%, 0.1%, 0.09%, 0.08%, 0.07%, 0.06%, 0.05%, 0.04%, 0.03%, 0.02%, or 0.01% in an SBS sequencing run of at least 50 cycles. In some embodiments, the phasing value refers to a value measured after 50, 75, 100, 125, 150, 200, 250, or 300 cycles. In further embodiments, the use of one or more Pd scavengers provides a % phasing value of less than about 0.05% in Read 1 of an SBS sequencing run of at least 150 cycles.

[0144] In some embodiments, the post-wash solutions described herein can also be used in a separate wash step prior to the detection step (i.e., step (b) in the methods described herein) to wash away any unincorporated nucleotides from step (a).

[0145] In some further embodiments, the nucleotides used in incorporation step (a) are fully functionalized A, C, T, and G nucleotide triphosphates, each containing a 3'-blocking group (e.g., 3'-AOM) and a cleavable linker (e.g., a cleavable linker containing an AOL linker moiety) described herein. In some such embodiments, the nucleotides herein provide superior stability in solution during a sequencing run compared to the same nucleotides protected with a standard 3'-O-azidomethyl blocking group. For example, the 3' acetal blocking groups disclosed herein can confer improved stability of at least 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 300%, 400%, 500%, 600%, 700%, 800%, 900%, 1000%, 1500%, 2000%, 2500%, or 3000% compared to azidomethyl-protected 3'-OH fragments under the same conditions for the same period of time, thereby reducing prephasing and increasing sequencing read lengths. In some embodiments, stability is measured at ambient or subambient temperatures (e.g., 4-10°C). In other embodiments, stability is measured at elevated temperatures, such as 40°C, 45°C, 50°C, 55°C, 60°C, or 65°C. In some such embodiments, stability is measured in a basic pH environment, e.g., a solution at pH 9.0, 9.2, 9.4, 9.6, 9.8, or 10.0. In some further embodiments, the prephasing value with a 3' blocking nucleotide described herein is less than about 0.25, 0.24, 0.23, 0.22, 0.21, 0.20, 0.19, 0.18, 0.17, 0.16, 0.15, 0.14, 0.13, 0.12, 0.11, 0.10, 0.09, 0.08, 0.07, 0.06, or 0.05 after more than 50, 100, or 150 cycles of SBS.In some further embodiments, the phasing value due to the 3'-blocking nucleotide is less than about 0.25, 0.24, 0.23, 0.22, 0.21, 0.20, 0.19, 0.18, 0.17, 0.16, 0.15, 0.14, 0.13, 0.12, 0.11, 0.10, 0.09, 0.08, 0.07, 0.06, or 0.05 after more than 50, 100, or 150 cycles of SBS. In one embodiment, each ffN comprises a 3'-AOM group.

[0146] In some embodiments, the 3'-blocked nucleotides described herein provide superior deblocking rates in solution during the chemical cleavage step of a sequencing run compared to the same nucleotide protected with a standard 3'-O-azidomethyl blocking group. For example, the 3' acetal (e.g., AOM) blocking groups disclosed herein may confer improved deblocking rates of at least 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 150%, 200%, 300%, 400%, 500%, 600%, 700%, 800%, 900%, 1000%, 1500%, or 2000% compared to an azidomethyl-protected 3'-OH using a standard deblocking reagent (such as tris(hydroxypropyl)phosphine), thereby reducing the overall time for a sequencing cycle. In some embodiments, the deblocking time for each nucleotide is reduced by about 5%, 10%, 20%, 30%, 40%, 50%, or 60%. For example, the deblocking times for 3'-AOM and 3'-O-azidomethyl are about 4-5 seconds and about 9-10 seconds, respectively, under certain chemical reaction conditions. In some embodiments, the half-life (t 1 / 2 ) is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 times faster than the azidomethyl blocking group. In some such embodiments, the t 1 / 2 is about 1 minute, while the t 1 / 2is about 11 minutes. In some embodiments, the deblocking rate is measured at ambient or subambient temperatures (e.g., 4-10°C). In other embodiments, the deblocking rate is measured at elevated temperatures, such as 40°C, 45°C, 50°C, 55°C, 60°C, or 65°C. In some such embodiments, the deblocking rate is measured in a basic pH environment, e.g., a solution at pH 9.0, 9.2, 9.4, 9.6, 9.8, or 10.0. In some such embodiments, the molar ratio of deblocking reagent to substrate (i.e., 3'-blocked nucleoside or nucleotide) is about 10:1, about 5:1, about 2:1, about 1:1, about 1:2, about 1:5, or about 1:10. In one embodiment, each ffN contains a 3'-AOM blocking group and an AOL linker moiety.

[0147] In any embodiment of the methods described herein, the labeled nucleotide is a nucleotide triphosphate having a 2' deoxyribose. In any embodiment of the methods described herein, the target polynucleotide strand is bound to a solid support, such as a flow cell.

[0148] In one embodiment, during the synthesis process, at least one nucleotide is incorporated into the polynucleotide by the action of a polymerase enzyme. In some such embodiments, the polymerase can be DNA polymerase Pol812 or Pol1901. However, other methods of joining nucleotides to polynucleotides can be used, such as chemical oligonucleotide synthesis or ligation of labeled oligonucleotides to unlabeled oligonucleotides. Thus, the term "incorporate," when used with respect to nucleotides and polynucleotides, can encompass chemical and enzymatic polynucleotide synthesis.

[0149] In certain embodiments, a synthesis step can be performed, which can optionally include incubating a template polynucleotide strand with a reaction mixture containing the labeled 3'-blocked nucleotides of the present disclosure. A polymerase can also be provided under conditions that allow the formation of a phosphodiester bond between the free 3'-OH group on the polynucleotide strand annealed to the template polynucleotide strand and the 5' phosphate group on the nucleotide. Thus, the synthesis step can include the formation of a polynucleotide strand achieved by complementary base pairing of the nucleotide to the template strand.

[0150] In all embodiments of the present method, the detection step may be performed while the polynucleotide strand incorporating the labeled nucleotide is annealing to the template or target strand, or after the denaturation step in which the two strands are separated. Additional steps, such as chemical or enzymatic reaction steps or purification steps, may be included between the synthesis step and the detection step. Specifically, the target strand incorporating the labeled nucleotide may be isolated or purified and then further processed or used for subsequent analysis. For example, a target polynucleotide labeled with a nucleotide(s) described herein in the synthesis step may then be used as a labeled probe or primer. In other embodiments, the product of the synthesis step described herein may be subjected to additional reaction steps, and the products of these subsequent steps may be purified or isolated, if desired.

[0151] Suitable conditions for the synthesis step are well known to those familiar with standard molecular biology techniques. In one embodiment, the synthesis step may resemble a standard primer extension reaction using nucleotide precursors containing the nucleotides described herein to form an extended target strand complementary to a template strand in the presence of a suitable polymerase enzyme. In other embodiments, the synthesis step may itself be part of an amplification reaction, generating a labeled, double-stranded amplification product consisting of annealed complementary strands derived from copying the target and template polynucleotide strands. Other exemplary synthesis steps include nick translation, strand displacement polymerization, random-primed DNA labeling, and the like. Polymerase enzymes particularly useful for the synthesis step are those capable of catalyzing the incorporation of the nucleotides described herein. A variety of naturally occurring or modified polymerases can be used. For example, thermostable polymerases can be used in synthesis reactions performed using thermal cycling conditions, but thermostable polymerases may not be desirable for isothermal primer extension reactions. Suitable thermostable polymerases capable of incorporating nucleotides according to the present disclosure include those described in WO 2005 / 024010 or WO 06 / 120433, each of which is incorporated herein by reference. For synthesis reactions performed at low temperatures, such as 37°C, the polymerase enzyme does not necessarily need to be a thermostable polymerase; therefore, the choice of polymerase will depend on several factors, such as reaction temperature, pH, strand displacement activity, etc.

[0152] In certain non-limiting embodiments, the present disclosure includes methods for nucleic acid sequencing, resequencing, whole genome sequencing, single nucleotide polymorphism scoring, as well as any other application involving detection of the labeled nucleotides or nucleosides described herein when incorporated into a polynucleotide. The labeled nucleotides or nucleosides bearing the dyes described herein can also be used in any of a variety of other applications that benefit from the use of polynucleotides labeled with nucleotides that contain fluorescent dyes.

[0153] In certain embodiments, the present disclosure provides for the use of labeled nucleotides according to the present disclosure in polynucleotide sequencing by synthesis (SBS) reactions. Sequencing by synthesis generally involves sequentially adding one or more nucleotides or oligonucleotides in a 5' to 3' direction to a growing polynucleotide chain using a polymerase or ligase to form a growing polynucleotide chain complementary to a template nucleic acid to be sequenced. The identity of the base present in one or more of the added nucleotides can be determined in a detection or "imaging" step. The identity of the added base can be determined after each nucleotide incorporation step. The sequence of the template can then be inferred using conventional Watson-Crick base-pairing rules. Using labeled nucleotides described herein to determine the identity of a single base can be useful, for example, in scoring single nucleotide polymorphisms, and such single-base extension reactions are within the scope of the present disclosure.

[0154] In one embodiment of the present disclosure, the sequence of a template polynucleotide is determined by detecting the incorporation of one or more 3'-blocking nucleotides into a nascent strand complementary to the template polynucleotide to be sequenced through detection of a fluorescent label(s) attached to the incorporated nucleotides. The sequencing of the template polynucleotide can be primed with a suitable primer (or a primer prepared as a hairpin structure containing the primer as part of the hairpin), and the nascent strand is extended stepwise by adding nucleotides to the 3' end of the primer in a polymer-catalyzed reaction.

[0155] In certain embodiments, each of the different nucleotide triphosphates (A, T, G, and C) may be labeled with a unique fluorophore and contain a blocking group at the 3' position to prevent uncontrolled polymerization. Alternatively, one of the four nucleotides may be unlabeled (dark). The polymerase enzyme incorporates the nucleotide into the nascent strand complementary to the template polynucleotide, while the blocking group prevents further incorporation of the nucleotide. Any unincorporated nucleotides can be washed away, and the fluorescent signal from each incorporated nucleotide can be optically "read" by suitable means, such as a charge-coupled device using laser excitation and suitable emission filters. The 3'-blocking group and fluorescent dye compound can then be simultaneously or sequentially removed (deprotected) to expose the nascent strand to incorporation of additional nucleotides. Typically, the identity of the incorporated nucleotide is determined after each incorporation step, although this is not strictly required. Similarly, US Pat. No. 5,302,509, the disclosure of which is incorporated herein by reference in its entirety, discloses a method for sequencing polynucleotides immobilized on a solid support.

[0156] The method, as exemplified above, incorporates fluorescently labeled 3'-blocked nucleotides A, G, C, and T into a growing strand complementary to an immobilized polynucleotide in the presence of a DNA polymerase. The polymerase incorporates the base complementary to the target polynucleotide, but further addition is prevented by the 3'-blocking group. The label of the incorporated nucleotide can then be determined, after which the blocking group can be removed by chemical cleavage to allow further polymerization. The nucleic acid template to be sequenced in a sequencing-by-synthesis reaction can be any polynucleotide for which sequencing is desired. Nucleic acid templates for sequencing reactions typically contain a double-stranded region with a free 3'-OH group that serves as a primer or initiation point for the addition of additional nucleotides in the sequencing reaction. The region of the template to be sequenced has this free 3'-OH group overhanging on the complementary strand. The overhang region of the template to be sequenced may be single-stranded, but can also be double-stranded, so long as a "nick" is present on the strand complementary to the template strand to be sequenced, providing a free 3'-OH group for initiating the sequencing reaction. In such embodiments, sequencing may proceed by strand displacement. In certain embodiments, a primer having a free 3'-OH group may be added as a separate component (e.g., a short oligonucleotide) that hybridizes to a single-stranded region of the template to be sequenced. Alternatively, the primer and template strand to be sequenced may each form part of a partially self-complementary nucleic acid strand that can form an intramolecular duplex, such as a hairpin loop structure. Hairpin polynucleotides and methods by which they can be attached to solid supports are disclosed in WO 01 / 57248 and WO 2005 / 047301, each of which is incorporated herein by reference. Nucleotides can be added sequentially to the growing primer to synthesize a polynucleotide chain in the 5' to 3' direction. The nature of the added base may, but need not, be determined particularly after each nucleotide addition, and such determination provides sequence information about the nucleic acid template.Thus, a nucleotide is incorporated into a nucleic acid strand (or polynucleotide) by linking the nucleotide to a free 3'-OH group of the nucleic acid strand through the formation of a phosphodiester bond with the 5' phosphate group of the nucleotide.

[0157] The nucleic acid template to be sequenced may be DNA or RNA, or even a hybrid molecule composed of deoxynucleotides and ribonucleotides. The nucleic acid template may contain naturally occurring and / or non-naturally occurring nucleotides and naturally occurring or non-naturally occurring backbone linkages, as long as they do not prevent copying of the template in the sequencing reaction.

[0158] In certain embodiments, the nucleic acid template to be sequenced can be attached to a solid support via any suitable linking method known in the art, for example, via covalent bonding. In certain embodiments, the template polynucleotide can be directly attached to a solid support (e.g., a silica-based support). However, in other embodiments of the present disclosure, the surface of the solid support can be modified in some way to allow either direct covalent attachment of the template polynucleotide or immobilization of the template polynucleotide through a hydrogel or polyelectrolyte layer (which itself is attached to the solid support by means other than covalent bonding).

[0159] Sequencing-by-Synthesis Embodiments and Alternatives Some embodiments include pyrosequencing techniques. Pyrosequencing detects the release of inorganic pyrophosphate (PPi) when a specific nucleotide is incorporated into a nascent strand (Ronaghi, M., Karamohamed, S., Pettersson, B., Uhlen, M., and Nyren, P. (1996) "Real-time DNA sequencing using detection of pyrophosphate release." Analytical Biochemistry 242(1), 84-9; Ronaghi, M. (2001) "Pyrosequencing sheds light on DNA sequencing." Genome Res. 11(1), 3-11; Ronaghi, M., Uhlen, M., and Nyren, P. (1998) "A sequencing method based on real-time pyrophosphate." Science 281 (5375), 363; U.S. Patent Nos. 6,210,891, 6,258,568, and 6,274,320, which are incorporated herein by reference in their entireties. In pyrosequencing, released PPi can be detected by its immediate conversion to adenosine triphosphate (ATP) by ATP sulfurase, and the level of generated ATP is detected via luciferase-generated photons. The nucleic acid to be sequenced can be attached to features in an array, and the array can be imaged to capture the chemiluminescent signal produced by the incorporation of nucleotides into the features of the array. Images can be obtained after treating the array with a specific nucleotide type (e.g., T, C, or G). The images obtained after the addition of each nucleotide type differ in terms of which features in the array are detected. These differences in the images reflect the different sequence content of the features on the array. However, the relative positions of each feature remain unchanged in the image. Images can be stored, processed, and analyzed using the methods described herein.For example, images obtained after treating the array with each different nucleotide type can be processed in the same manner as exemplified herein for images obtained from different detection channels for reversible terminator-based sequencing methods.

[0160] In another exemplary type of SBS, cycle sequencing is achieved by stepwise addition of reversible terminator nucleotides containing cleavable or photobleachable dye labels, as described, for example, in International Publication No. 04 / 018497 and U.S. Patent No. 7,057,026, the disclosures of which are incorporated by reference. This approach has been commercialized by Solexa (now Illumina, Inc.) and is also described in International Publication Nos. 91 / 06678 and 07 / 123,744, each of which is incorporated by reference herein. The availability of fluorescently labeled terminators, both of which can be reversed and from which the fluorescent label is cleaved, facilitates efficient cyclic reversible termination (CRT) sequencing. Polymerases can also be co-engineered to efficiently incorporate and extend from these modified nucleotides.

[0161] Preferably, in reversible terminator-based sequencing embodiments, the label does not substantially inhibit extension under SBS reaction conditions. However, the detection label may be removable, for example, by cleavage or degradation. Images can be captured after incorporation of the label into arrayed nucleic acid features. In certain embodiments, each cycle involves simultaneous delivery of four different nucleotide types to the array, with each nucleotide type bearing a spectrally distinct label. Four images can then be obtained, each using a detection channel selective for one of the four different labels. Alternatively, different nucleotide types can be added sequentially, with images of the array being obtained between each addition step. In such embodiments, each image shows nucleic acid features incorporating a particular type of nucleotide. Because the sequence content of each feature differs, different features are present or absent in different images. However, the relative positions of the features remain unchanged within the image. Images obtained from such reversible terminator-SBS methods can be stored, processed, and analyzed as described herein. Following the image capture step, the label can be removed, and the reversible terminator moiety can be removed for subsequent cycles of nucleotide addition and detection. Removal of the label after detection in a particular cycle and before the subsequent cycle has the advantage of reducing background signal and crosstalk between cycles. Examples of useful labeling and removal methods are described below.

[0162] Some embodiments may utilize detection of four different nucleotides using fewer than four different labels. For example, SBS may be performed using the methods and systems described in the incorporated document, U.S. Patent Application Publication No. 2013 / 0079232. As a first example, pairs of nucleotide types may be detected at the same wavelength but may be distinguished based on differences in intensity for one member of the pair or based on a change to one member of the pair (e.g., via chemical, photochemical, or physical modification) that results in the appearance or disappearance of a distinct signal compared to the signal detected for the other member of the pair. As a second example, three of the four different nucleotide types may be detected under certain conditions, while the fourth nucleotide type may have no detectable label under those conditions or may be minimally detected under those conditions (e.g., minimal detection due to background fluorescence, etc.). Incorporation of the first three nucleotide types into a nucleic acid may be determined based on the presence of their corresponding signals, and incorporation of the fourth nucleotide type into a nucleic acid may be determined based on the absence or minimal detection of any signal. As a third example, one nucleotide type can include a label that is detected in two different channels, while the other nucleotide type is detected in no more than one of the channels. The three exemplary configurations above are not considered mutually exclusive and can be used in various combinations.An exemplary embodiment combining all three examples is a fluorescence-based SBS method that uses a first nucleotide type that is detected in a first channel (e.g., dATP having a label that is detected in the first channel when excited by a first excitation wavelength), a second nucleotide type that is detected in a second channel (e.g., dCTP having a label that is detected in the second channel when excited by a second excitation wavelength), a third nucleotide type that is detected in both the first and second channels (e.g., dTTP having at least one label that is detected in both channels when excited by the first and / or second excitation wavelength), and an unlabeled fourth nucleotide type that is not detected or is minimally detected in either channel (e.g., unlabeled dGTP).

[0163] Furthermore, as described in incorporated U.S. Patent Application Publication No. 2013 / 0079232, sequencing data can be obtained using a single channel. In such so-called single-dye sequencing methods, a first nucleotide type is labeled but the label is removed after the first image is generated, and a second nucleotide type is labeled only after the first image is generated. A third nucleotide type retains its label in both the first and second images, and a fourth nucleotide type remains unlabeled in both images.

[0164] Some embodiments may utilize sequencing by ligation techniques. Such techniques utilize DNA ligase to incorporate oligonucleotides and identify their incorporation. The oligonucleotides typically have different labels that correlate with the identity of specific nucleotides in the sequence to which the oligonucleotides hybridize. As with other SBS methods, images can be obtained after treating an array of nucleic acid sequences with labeled sequencing reagents. Each image shows nucleic acid features that incorporate a specific type of label. Because the sequence content of each feature varies, different features may or may not be present in different images, but the relative positions of the features remain constant within the image. Images obtained from ligation-based sequencing methods can be stored, processed, and analyzed as described herein. Exemplary SBS systems and methods that can be utilized with the methods and systems described herein are described in U.S. Patent Nos. 6,969,488, 6,172,218, and 6,306,597, the disclosures of which are incorporated herein by reference in their entireties.

[0165] Some embodiments can utilize nanopore sequencing (Deamer, DW & Akeson, M. "Nanopores and nucleic acids: prospects for ultrarapid sequencing." Trends Biotechnol. 18, 147-151 (2000); Deamer, D. and D. Branton, "Characterization of nucleic acids by nanopore analysis." Acc. Chem. Res. 35:817-825 (2002); Li, J., M. Gershow, D. Stein, E. Brandin, and J. A. Golovchenko, "DNA molecules and configurations in a solid-state nanopore microscope." Nat. Mater. 2:611-615 (2003), the disclosures of which are incorporated herein by reference in their entireties). In such embodiments, the target nucleic acid passes through a nanopore. The nanopore can be a synthetic pore or a biological membrane protein, such as α-hemolysin. As the target nucleic acid passes through the nanopore, each base pair can be identified by measuring fluctuations in the pore's electrical conductance. (U.S. Pat. No. 7,001,792; Soni, GV & Meller, "A. Progress toward ultrafast DNA sequencing using solid-state nanopores." Clin. Chem. 53, 1996-2001 (2007); Healy, K. "Nanopore-based single-molecule DNA analysis." Nanomed. 2, 459-481 (2007); Cockroft, SL, Chu, J., Amorin, M. & Ghadiri, MR "A single-molecule nanopore device detects DNA polymerase activity with single-nucleotide resolution." J. Am Chem. Soc. 130, 818-820 (2008), the disclosures of which are incorporated herein by reference in their entireties.)The data obtained from nanopore sequencing can be stored, processed, and analyzed as described herein. Specifically, the data can be processed as images according to the exemplary processing of optical and other images described herein.

[0166] Some other embodiments of the sequencing method involve the use of the 3'-blocked nucleotides described herein in nanoball sequencing technology, such as that described in U.S. Patent No. 9,222,132, the disclosure of which is incorporated by reference. Through the process of rolling circle amplification (RCA), a large number of discrete DNA nanoballs can be generated. The nanoball mixture is then dispensed onto a patterned slide surface containing features that allow a single nanoball to associate with each location. In DNA nanoball generation, DNA is fragmented and ligated to the first of four adapter sequences. The template is amplified, circularized, and cleaved with a type II endonuclease. A second set of adapters is added, followed by amplification, circularization, and cleavage. This process is repeated for the remaining two adapters. The final product is a circular template with four adapters, each separated by the template sequence. Library molecules undergo a rolling circle amplification process to generate a large number of concatemers, called DNA nanoballs, which are then deposited onto a flow cell. Goodwin et al., “Coming of age: ten years of next-generation sequencing technologies,” Nat Rev Genet.2016;17(6):333-51.

[0167] Some embodiments can utilize methods involving real-time monitoring of DNA polymerase activity. Nucleotide incorporation can be detected via fluorescence resonance energy transfer (FRET) interactions between a fluorophore-containing polymerase and a γ-phosphate-labeled nucleotide, for example, as described in U.S. Patent Nos. 7,329,492 and 7,211,414, both of which are incorporated herein by reference, or nucleotide incorporation can be detected using zero-mode waveguides, for example, as described in U.S. Patent No. 7,315,019, both of which are incorporated herein by reference, and fluorescent nucleotide analogs and engineered polymerases, for example, as described in U.S. Patent No. 7,405,281 and U.S. Patent Application Publication No. 2008 / 0108082, both of which are incorporated herein by reference. Illumination can be restricted to a zeptoliter-scale volume around the surface-tethered polymerase so that incorporation of fluorescently labeled nucleotides can be observed with low background (Levene, MJ et al., "Zero-mode waveguides for single-molecule analysis at high concentrations." Science 299, 682-686 (2003); Lundquist, PM et al., "Parallel confocal detection of single molecules in real time." Opt. Lett. 33, 1026-1028 (2008); Korlach, J. et al., "Selective aluminum passivation for targeted immobilization of single DNA polymerase molecules in zero-mode waveguide nanostructures." Proc. Natl. Acad. Sci. USA 105, 1176-1181 (2008), the disclosures of which are incorporated herein by reference in their entireties). Images obtained from such methods can be stored, processed, and analyzed as described herein.

[0168] Some SBS embodiments involve the detection of protons released upon incorporation of a nucleotide into an extension product. For example, sequencing based on the detection of released protons is described in commercial electronic detectors and related technology available from Ion Torrent (Guilford, Connecticut, a Life Technologies subsidiary), or in U.S. Patent Application Publication Nos. 2009 / 0026082, 2009 / 0127589, 2010 / 0137143, and 2010 / 0282617, all of which are incorporated herein by reference. The methods described herein for amplifying target nucleic acids using equilibrium exclusion can be readily adapted to substrates used to detect protons. More specifically, the methods described herein can be used to generate clonal populations of amplicons used to detect protons.

[0169] The SBS method described above can be advantageously performed in a multiplex format, allowing multiple different target nucleic acids to be manipulated simultaneously. In certain embodiments, the different target nucleic acids can be processed in a common reaction vessel or on the surface of a specific substrate. This allows for convenient delivery of sequencing reagents, removal of unreacted reagents, and multiplexed detection of incorporation events. In embodiments using surface-bound target nucleic acids, the target nucleic acids can be in an array format. In an array format, the target nucleic acids can typically be bound to a surface in a spatially distinguishable manner. The target nucleic acids can be bound by direct covalent binding, attachment to beads or other particles, or binding to a polymerase or other molecule attached to the surface. Arrays can contain a single copy of the target nucleic acid at each site (also called a feature), or multiple copies with the same sequence can be present at each site or feature. Multiple copies can be generated by amplification methods such as bridge amplification or emulsion PCR, which are described in more detail below.

[0170] The methods described herein can be used to fabricate, for example, at least about 10 features / cm 2 , 100 features / cm 2 , 500 features / cm2 , 1,000 features / cm 2 , 5,000 features / cm 2 , 10,000 features / cm 2 , 50,000 features / cm 2 , 100,000 features / cm 2 , 1,000,000 features / cm 2 , 5,000,000 features / cm 2 Arrays having features of any of a variety of densities, including 1000 sq. ft., ... or more, can be used.

[0171] An advantage of the methods described herein is that they provide rapid and efficient detection of multiple target nucleic acids in parallel. Accordingly, the present disclosure provides an integrated system capable of preparing and detecting nucleic acids using techniques known in the art, such as those exemplified above. Accordingly, the integrated system of the present disclosure can include fluidic components capable of delivering amplification and / or sequencing reagents to one or more immobilized DNA fragments, including components such as pumps, valves, reservoirs, and fluid lines. A flow cell can be configured and / or used in the integrated system for detecting target nucleic acids. Exemplary flow cells are described, for example, in U.S. Patent Application Publication No. 2010 / 0111768 and U.S. Patent Application No. 13 / 273,666, each of which is incorporated herein by reference. As exemplified for the flow cell, one or more of the fluidic components of the integrated system can be used in the amplification and detection methods. Taking the nucleic acid sequencing embodiment as an example, one or more of the fluidic components of the integrated system can be used for the delivery of sequencing reagents in the amplification methods described herein and in the sequencing methods exemplified above. Alternatively, an integrated system may include separate fluidic systems for performing the amplification method and for performing the detection method. Examples of integrated sequencing systems capable of producing amplified nucleic acids and sequencing the nucleic acids include, but are not limited to, the MiSeq™ platform (Illumina Inc., San Diego, CA) and the devices described in U.S. Patent Application No. 13 / 273,666, which is incorporated herein by reference.

[0172] Arrays in which polynucleotides are directly attached to silica-based supports are disclosed, for example, in WO 00 / 06770 (incorporated herein by reference), in which polynucleotides are immobilized on glass supports by reaction between pendant epoxide groups on the glass and internal amino groups on the polynucleotides. Furthermore, polynucleotides can be attached to solid supports by reaction of sulfur-based nucleophiles with the solid support, as described, for example, in WO 2005 / 047301 (incorporated herein by reference). Yet another example of a solid-supported template polynucleotide is one in which the template polynucleotide is attached to a hydrogel supported on a silica-based or other solid support. Examples are described, for example, in WO 00 / 31148, WO 01 / 01143, WO 02 / 12566, WO 03 / 014392, U.S. Pat. No. 6,465,178, and WO 00 / 53812, each of which is incorporated herein by reference.

[0173] Specific surfaces onto which template polynucleotides can be immobilized include polyacrylamide hydrogels. Polyacrylamide hydrogels are described in the above references and in WO 2005 / 065814, which are incorporated herein by reference. Specific hydrogels that can be used include those described in WO 2005 / 065814 and U.S. Patent Application Publication No. 2014 / 0079923. In one embodiment, the hydrogel is PAZAM (poly(N-(5-azidoacetamidylpentyl)acrylamide-co-acrylamide)).

[0174] DNA template molecules can be attached to beads or microparticles, for example, as described in U.S. Patent No. 6,172,218, which is incorporated herein by reference. Attachment to beads or microparticles can be useful for sequencing applications. Bead libraries can be prepared, with each bead containing a different DNA sequence. Exemplary libraries and methods for their creation are described in Nature, 437, 376-380 (2005); Science, 309, 5741, 1728-1732 (2005), each of which is incorporated herein by reference. Sequencing of such bead arrays using the nucleotides described herein is within the scope of this disclosure.

[0175] The templates to be sequenced may form part of an "array" on a solid support, where the array may take any convenient form. Thus, the methods of the present disclosure are applicable to all types of high-density arrays, including single molecule arrays, clustered arrays, and bead arrays. The labeled nucleotides of the present disclosure may be used to sequence templates on essentially any type of array, including, but not limited to, those formed by immobilizing nucleic acid molecules on a solid support.

[0176] However, the labeled nucleotides of the present disclosure are particularly advantageous in the context of clustered array sequencing. In clustered arrays, distinct regions (often called sites or features) on the array contain multiple polynucleotide template molecules. Generally, the multiple polynucleotide molecules are not individually resolvable by optical means, but instead are detected as an ensemble. Depending on how the array is formed, each site on the array may contain multiple copies of one individual polynucleotide molecule (e.g., the site is homogeneous for a particular single-stranded or double-stranded nucleic acid species), or even multiple copies of a small number of different polynucleotide molecules (e.g., multiple copies of two different nucleic acid species). Clustered arrays of nucleic acid molecules can be generated using techniques generally known in the art. By way of example, WO 98 / 44151 and WO 00 / 18957, each incorporated herein, describe methods for amplifying nucleic acids, in which both the template and the amplification product remain immobilized on a solid support to form an array consisting of clusters or "colonies" of immobilized nucleic acid molecules. The nucleic acid molecules present on the clustered arrays prepared according to these methods are suitable templates for sequencing using nucleotides labeled with the dye compounds of the present disclosure.

[0177] The labeled nucleotides of the present disclosure are also useful for sequencing templates on single molecule arrays. As used herein, the term "single molecule array" or "SMA" refers to a collection of polynucleotide molecules distributed (or arrayed) on a solid support, where the separation distance of any individual polynucleotide from all other polynucleotides in the collection is such that each individual polynucleotide molecule can be individually distinguished. Thus, in some embodiments, target nucleic acid molecules immobilized on the surface of a solid support can be distinguished by optical means. This means that one or more distinct signals, each representing a polynucleotide, are generated within an area that can be distinguished by the particular imaging device used.

[0178] Single molecule detection can be achieved where the spacing between adjacent polynucleotide molecules on the array is at least 100 nm, more specifically at least 250 nm, even more specifically at least 300 nm, and even more specifically at least 350 nm, such that each molecule is individually resolvable and detectable as a single molecule fluorescent dot, and the fluorescence from this single molecule fluorescent dot also exhibits single-step photobleaching.

[0179] The terms "individually resolved" and "individual resolution" are used herein to designate the ability to distinguish one molecule on an array from its neighboring molecules when visualized. The separation between individual molecules on an array is determined, in part, by the particular technique used to resolve the individual molecules. The general characteristics of single molecule arrays can be understood by reference to International Publication Nos. 00 / 06770 and 01 / 57248, each of which is incorporated herein by reference. One application of the nucleotides of the present disclosure is in sequencing-by-synthesis reactions, although their utility is not limited to such methods. Indeed, the nucleotides can be advantageously used in any sequencing method requiring detection of a fluorescent label attached to a nucleotide incorporated into a polynucleotide.

[0180] Specifically, the labeled nucleotides of the present disclosure can be used in automated fluorescent sequencing protocols, particularly fluorescent dye terminator cycle sequencing, which is based on the chain termination sequencing method of Sanger and coworkers. Such methods generally incorporate fluorescently labeled dideoxynucleotides in primer extension sequencing reactions using enzymes and cycle sequencing. So-called Sanger sequencing and related protocols (Sanger-type) utilize randomized chain terminations with labeled dideoxynucleotides.

[0181] Thus, the present disclosure also encompasses labeled nucleotides that are dideoxynucleotides lacking hydroxyl groups at both the 3' and 2' positions, such deoxynucleotides being suitable for use in Sanger-type sequencing and the like.

[0182] It will be appreciated that labeled nucleotides of the present disclosure incorporating a 3'-blocking group are useful in Sanger sequencing and related protocols because the same effect achieved by using dideoxynucleotides can be achieved by using nucleotides with a 3'-OH blocking group (both can prevent the incorporation of subsequent nucleotides). It will be understood that when nucleotides with a 3'-blocking group according to the present disclosure are used in Sanger-type sequencing, in each instance in which a labeled nucleotide of the present disclosure is incorporated, the dye compound or detectable label attached to the nucleotide does not need to be linked via a cleavable linker, and therefore the label does not need to be removed from the nucleotide, since there is no need for subsequent incorporation.

[0183] In any embodiment of the SBS methods described herein, the nucleotide used in the sequencing application is a 3'-blocked nucleotide described herein, e.g., a nucleotide of formula (I) and (Ia)-(Id). In any embodiment, the 3'-blocked nucleotide is a nucleotide triphosphate.

[0184] In certain sequencing methods, the incorporated nucleotides are unlabeled. One or more fluorescent labels can be introduced after incorporation by using a labeled affinity reagent containing one or more fluorescent dyes. For example, one, two, three, or four different types of nucleotides (e.g., dATP, dCTP, dGTP, and dTTP or dUTP) in the incorporation buffer of step (a) may be unlabeled. Each of the four types of nucleotides (e.g., dNTPs) has a 3'-OH blocking group (e.g., 3'-AOM) as described herein to ensure that only one base can be added by the polymerase to the 3' end of the copy polynucleotide. After incorporation of the unlabeled nucleotide, an affinity reagent that specifically binds to the incorporated dNTP is then introduced to provide a labeled extension product containing the incorporated dNTP. The use of unlabeled nucleotides and affinity reagents in sequencing by synthesis is disclosed in U.S. Publication No. 2013 / 0079232. The modified sequencing method of the present disclosure using unlabeled nucleotides comprises the following steps: (a'-1) a 3'-OH blocking group as described herein [ka] incorporating an unlabeled nucleotide (e.g., dATP, dCTP, dGTP, dTTP, or dUTP) comprising (attached to the 3' oxygen) into a copy polynucleotide strand complementary to at least a portion of the target polynucleotide strand to create an extended copy polynucleotide; (a'-2) contacting the extended copy polynucleotide with a pair of affinity reagents under conditions such that one affinity reagent specifically binds to the incorporated unlabeled nucleotide to provide a labeled extended copy polynucleotide; (b') detecting the identity of the nucleotide incorporated into the copy polynucleotide strand by performing one or more fluorescence measurements of the labeled extended copy polynucleotide; (c') chemically removing the detectable label from the extended copy polynucleotide and the 3'-OH blocking group from the nucleotide incorporated into the copy polynucleotide strand; may include:

[0185] The affinity reagent may comprise a small molecule or protein tag that can bind to the hapten portion of the nucleotide (e.g., streptavidin-biotin, anti-DIG and DIG, anti-DNP and DNP), an antibody (including, but not limited to, antibody binding fragments, single-chain antibodies, bispecific antibodies, etc.), an aptamer, a knottin, an affimer, or any other known agent that binds with appropriate specificity and affinity to the incorporated nucleotide. In further embodiments, a single affinity reagent can be labeled with multiple copies of the same fluorescent dye. In some embodiments, the Pd catalyst also removes the labeled affinity reagent. For example, the hapten portion of the unlabeled nucleotide can be cleaved by the Pd catalyst using a cleavable linker described herein. [ka] (e.g., an AOL linker). In some embodiments, the method further comprises a post-cleavage wash step (d) described herein. In some embodiments, the method further comprises repeating steps (a'-1) through (c') or (a'-1) through (d) until the sequence of at least a portion of the target polynucleotide strand is determined. In some embodiments, the cycle is repeated at least 50 times, at least 100 times, at least 150 times, at least 200 times, at least 250 times, or at least 300 times.

[0186] kit The present disclosure also provides kits comprising one or more 3'-blocked nucleosides and / or nucleotides described herein, e.g., 3'-blocked nucleotides of Formulae (I) and (Ia)-(Id). Such kits generally comprise at least one 3'-blocked nucleotide or nucleoside comprising a detectable label (e.g., a fluorescent dye) with at least one additional component. The additional component may be one or more of the components identified in the methods described herein or in the Examples section below. Some non-limiting examples of components that can be combined into kits of the present disclosure are described below. In some further embodiments, the kits may comprise four types of labeled nucleotides (A, C, T, and G) of fully functionalized nucleotides described herein, each type of nucleotide comprising a 3'-AOM blocking group and an AOL linker moiety as described herein. In further embodiments, G is unlabeled and does not comprise an AOL linker. In still further embodiments, one or more of the remaining three nucleotides (i.e., A, C, and T) comprise an L-type nucleotide, which is an allylamine or allylamide linker moiety. 1 In one embodiment, the kit comprises unlabeled ffG, labeled ffA(s), labeled ffC, and labeled ffT-DB as described herein. In another embodiment, the kit comprises unlabeled ffG, labeled ffA(s), labeled ffC-DB, and labeled ffT-DB as described herein.

[0187] In certain embodiments, the kit can include at least one labeled 3'-blocking nucleotide or nucleoside along with labeled or unlabeled nucleotides or nucleosides. For example, dye-labeled nucleotides can be supplied in combination with unlabeled or natural nucleotides and / or with fluorescently labeled nucleotides, or any combination thereof. Nucleotide combinations can be provided as separate individual components (e.g., one nucleotide type per container or tube) or as nucleotide mixtures (e.g., two or more nucleotides mixed in the same container or tube).

[0188] When the kit includes multiple, particularly two, three, or more specifically four, 3'-blocked nucleotides labeled with dye compounds, the different nucleotides may be labeled with different dye compounds or may be dark and contain no dye compounds. When different nucleotides are labeled with different dye compounds, it is a feature of the kit that the dye compounds are spectrally distinguishable fluorescent dyes. As used herein, the term "spectrally distinguishable fluorescent dyes" refers to fluorescent dyes that emit fluorescent energy at wavelengths that can be distinguished by a fluorescent detection device (e.g., a commercially available capillary-based DNA sequencing platform) when two or more dyes are present in a sample. When two nucleotides labeled with fluorescent dye compounds are provided in kit form, it is a feature of some embodiments that the spectrally distinguishable fluorescent dyes can be excited at the same wavelength, e.g., by the same laser. When the four 3' blocked nucleotides (A, C, T, and G) labeled with fluorescent dye compounds are supplied in kit form, it is a feature of some embodiments that two of the spectrally distinguishable fluorescent dyes can be excited at one wavelength and the other two spectrally distinguishable dyes can both be excited at another wavelength. Specific excitation wavelengths are 488 nm and 532 nm.

[0189] In one embodiment, the kit includes a first 3'-blocked nucleotide labeled with a first dye and a second nucleotide labeled with a second dye, where the dyes have a difference in absorbance maximum of at least 10 nm, particularly 20-50 nm. More specifically, the two dye compounds have a Stokes shift of 15-40 nm. The "Stokes shift" is the distance between the peak absorption wavelength and the peak emission wavelength.

[0190] In alternative embodiments, the disclosed kits may include 3'-blocked nucleotides in which the same base is labeled with two or more different dyes. The first nucleotide (e.g., a 3'-blocked T nucleotide triphosphate or a 3'-blocked G nucleotide triphosphate) may be labeled with a first dye. The second nucleotide (e.g., a 3'-blocked C nucleotide triphosphate) may be labeled with a second, spectrally distinct dye from the first dye, such as a "green" dye that absorbs below 600 nm, and a "blue" dye that absorbs below 500 nm, e.g., between 400 and 500 nm, particularly between 450 and 460 nm. The third nucleotide (e.g., a 3'-blocked A nucleotide triphosphate) may be labeled with a mixture of the first and second dyes, or a mixture of the first, second, and third dyes, and the fourth nucleotide (e.g., a 3'-blocked G nucleotide triphosphate or a 3'-blocked T nucleotide triphosphate) may be "dark" and may not contain a label. In one example, nucleotides 1-4 can be labeled "blue," "green," "blue / green," and dark. To further simplify the instrumentation, the four nucleotides can be labeled with two dyes excited by a single laser, and the labels for nucleotides 1-4 can be "blue 1," "blue 2," "blue 1 / blue 2," and dark.

[0191] In certain embodiments, the kit may include four labeled 3'-blocked nucleotides (e.g., A, C, T, G), each type of nucleotide containing the same 3'-blocking group and fluorescent label, each fluorescent label having a distinct fluorescence maximum, and each fluorescent label being distinguishable from the other three labels. The kit may also be such that two or more of the fluorescent labels have absorption maxima but different Stokes shifts. In some other embodiments, one type of nucleotide is unlabeled.

[0192] While the kits are exemplified herein with respect to configurations having different nucleotides labeled with different dye compounds, it will be understood that the kits may include two, three, four, or more different nucleotides having the same dye compound. In some embodiments, the kits also include an enzyme and a buffer suitable for enzyme action. In some such embodiments, the enzyme is a polymerase, terminal deoxynucleotidyl transferase, or reverse transcriptase. In particular embodiments, the enzyme is a DNA polymerase, such as DNA polymerase 812 (Pol812) or DNA polymerase 1901 (Pol1901). In some further embodiments, the kits may include an incorporation mixture described herein. In further embodiments, a kit containing an incorporation mixture described herein also includes at least one Pd scavenger (e.g., a Pd(0) scavenger described herein comprising one or more allyl moieties). Pd(0) scavengers include -O-allyl, -S-allyl, -NR-allyl, and -N-allyl. + and R'-aryl, wherein R is H, unsubstituted or substituted C1-C6 alkyl, unsubstituted or substituted C2-C6 alkenyl, unsubstituted or substituted C2-C6 alkynyl, unsubstituted or substituted C6-C 10 Aryl, unsubstituted or substituted 5-10 membered heteroaryl, unsubstituted or substituted C3-C 10 carbocyclyl, or unsubstituted or substituted 5-10 membered heterocyclyl, and R' is H, unsubstituted C1-C6 alkyl, or substituted C1-C6 alkyl. In some such embodiments, the Pd(0) scavenger in the integration solution comprises one or more -O-allyl moieties. In some further embodiments, the Pd(0) scavenger is [ka] or a combination thereof. Alternative Pd(0) scavengers are disclosed in U.S. Patent No. 63 / 190983, which is incorporated by reference in its entirety. In one embodiment, the Pd(0) scavenger in the incorporation mixture is: [ka] In another embodiment, the Pd(0) scavenger in the incorporation mixture comprises or is: [ka] Contains or is.

[0193] Other components included in such kits include buffers, etc. The nucleotides of the present disclosure and any other nucleotide components, including mixtures of different nucleotides, may be provided in the kit in a concentrated form that is diluted before use. In such embodiments, a suitable dilution buffer may also be included. For example, the integration mixture kit may include one or more buffers selected from primary amines, secondary amines, tertiary amines, natural amino acids, or unnatural amino acids, or combinations thereof. In further embodiments, the buffer in the integration mixture includes ethanolamine or glycine, or a combination thereof.

[0194] Again, one or more of the components identified in the methods described herein can be included in the kits of the present disclosure. In some further embodiments, the kits can include a palladium catalyst described herein. In some embodiments, the Pd catalyst is generated by mixing a Pd(II) complex (i.e., a Pd precatalyst) with one or more water-soluble phosphines described herein. In some such embodiments, the kit containing the Pd catalyst is a cleavage mixture kit. In further embodiments, the cleavage mixture kit can include Pd(allyl)Cl]2 or Na2PdCl4 and the water-soluble phosphine THP to generate the active Pd(0) species. The molar ratio of the Pd(II) complex (e.g., Pd(allyl)Cl]2 or Na2PdCl4) to the water-soluble phosphine (e.g., THP) can be about 1:2, 1:3, 1:4, 1:5, 1:6, 1:7, 1:8, 1:9, or 1:10. In further embodiments, the cleavage mixture kit may also include one or more buffer reagents selected from the group consisting of primary amines, secondary amines, tertiary amines, carbonates, phosphates, and borates, and combinations thereof. Non-limiting examples of buffer reagents in the cleavage mixture kit are selected from the group consisting of ethanolamine (EA), tris(hydroxymethyl)aminomethane (Tris), glycine, carbonates, phosphates, borates, dimethylethanolamine (DMEA), diethylethanolamine (DEEA), N,N,N',N'-tetramethylethylenediamine (TEMED), and N,N,N',N'-tetraethylethylenediamine (TEEDA), 2-piperidineethanol, and combinations thereof. In one embodiment, the cleavage mixture kit includes DEEA. In another embodiment, the cleavage mixture kit contains 2-piperidineethanol.

[0195] In some further embodiments, the kit may include one or more palladium scavengers (e.g., Pd(II) scavengers described herein). In some such embodiments, the kit is a post-cleavage wash buffer kit. Non-limiting examples of Pd scavengers in the post-cleavage wash buffer kit include isocyanoacetic acid (ICNA) salts, ethyl isocyanoacetate, methyl isocyanoacetate, cysteine ​​or a salt thereof, L-cysteine ​​or a salt thereof, N-acetyl-L-cysteine, potassium ethyl xanthate, potassium isopropyl xanthate, glutathione, lipoic acid, ethylenediaminetetraacetic acid (EDTA), iminodiacetic acid, nitrilodiacetic acid, trimercapto-S-triazine, dimethyldithiocarbamate, dithiothreitol, mercaptoethanol, allyl alcohol, propargyl alcohol, thiol, tertiary amine, and / or tertiary phosphine, or a combination thereof. In one embodiment, the post-cleavage wash buffer kit includes L-cysteine ​​or a salt thereof.

[0196] In any embodiment of the kits described herein, the Pd scavenger (e.g., a Pd(0) or Pd(II) scavenger described herein) is in a separate container / compartment from the Pd catalyst. [Example]

[0197] Additional embodiments are disclosed in more detail in the following examples, which are not intended to limit the scope of the claims. Example 1. Synthesis of fully functionalized nucleotides with 3' AOM and AOL linker moieties [ka]

[0198] Synthesis of intermediate AOL LN2: The acetal compound LN1 (2.43 g, 9.6 mmol) was dissolved in anhydrous CHCl (100 mL) under N, and the solution was cooled to 0 °C in an ice bath. 2,4,6-Trimethylpyridine (7.6 mL, 57.5 mmol) was added, followed by the dropwise addition of trimethylsilyl trifluoromethanesulfonate (7.0 mL, 38.7 mmol). The mixture was stirred at 0 °C for 2 h, then allyl alcohol (13 mL, 191.1 mmol) was added, and the reaction was refluxed overnight. The reaction was quenched with a 98:2 mixture of MeOH / H2O, and the resulting solution was further stirred at room temperature for 3 h. The mixture was diluted with CHCl (100 mL) and water (200 mL), and the aqueous layer was acidified to pH 2–3 with 2 N HCl. The aqueous layer was separated, and the organic layer was further extracted with acidified water. The organic layer was dried over MgSO4, filtered, and the volatiles were evaporated under reduced pressure. The crude product was purified by flash chromatography on silica gel to give AOL LN2 as a colorless oil (2.56 g, 86%).

[0199] Synthesis of intermediate AOL LN3: To a solution of AOL LN2 (2.17 g, 7.0 mmol) in ethanol (17.5 mL) was added 4 M aqueous NaOH (17.5 mL, 70 mmol), and the mixture was stirred at room temperature for 3 h. After this time, all volatiles were removed under reduced pressure, and the residue was dissolved in 75 mL of water. The solution was acidified to pH 2-3 with 2 N HCl and then extracted with dichloromethane (DCM). The combined organic fractions were dried over MgSO4, filtered, and the volatiles were evaporated under reduced pressure. AOL LN3 was obtained without further purification as a colorless oil (1.65 g, 84%) that solidified upon storage at 20 °C. LC-MS (ES): (negative ion) m / z 281 (MH + (positive ion) m / z 305 (M+Na + ).

[0200] Synthesis of intermediate AOL LN4: A solution of AOL LN3 (1.62 g, 5.74 mmol) in anhydrous DMF (20 mL) was stirred under vacuum for 5 minutes and then cooled to 0 °C using an ice bath. N,N-Diisopropylethylamine (1.2 mL, 6.89 mmol) was added dropwise under N2, followed by PyBOP (3.30 g, 6.34 mmol). The reaction was stirred at 0 °C for 30 minutes, then a solution of N-(5-aminopentyl)-2,2,2-trifluoroacetamide hydrochloride (1.62 g, 6.90 mmol) in anhydrous DMF (3.0 mL) was added, immediately followed by additional N,N-diisopropylethylamine (1.4 mL, 8.04 mmol). The reaction was removed from the ice bath and stirred at room temperature for 4 hours. The volatiles were removed under reduced pressure, and the residue was dissolved in EtOAc (150 mL). The solution was extracted with 20 mM aqueous KHSO, water, and saturated aqueous NaHCO. The organic layer was dried over MgSO, filtered, and the volatiles were evaporated under reduced pressure. The crude product was purified by flash chromatography on silica gel to give AOL LN as a colorless oil (2.16 g, 82%). LC-MS (ES): (negative ion) m / z 461 (MH + ), 497(M-Cl - ).

[0201] Synthesis of AOL linker moiety To a solution of AOL LN4 (350 mg, 0.76 mmol) in CH3CN (13 mL) was added TEMPO (48 mg, 0.31 mmol), followed by NaH2PO4 in water (6.5 mL). .A solution of 2H2O (762 mg, 4.88 mmol) and NaClO2 (275 mg, 3.04 mmol) was added. Aqueous NaClO (14% available chlorine, 0.83 mL, 1.94 mmol) was added, and the solution immediately turned dark brown. The reaction was stirred at room temperature for 6 h and then quenched with 100 mM aqueous Na2SO3 until the mixture was colorless. The acetonitrile was removed under reduced pressure, and the residue was diluted with water and basified with triethylamine. The aqueous phase was extracted with EtOAc (10 mL) and then concentrated under reduced pressure. The crude product was purified by reverse-phase flash chromatography on C18 to give AOL (triethylammonium salt, 310 mg, 71%) as a colorless oil. LC-MS (ES): (negative ion) m / z 475 (MH + (positive ion) m / z 499 (M+Na + ), 578(M+Et3NH + ).

[0202] AOL-NH 2 Synthesis of the linker part To a solution of AOL (446 mg, 0.94 mmol) in methanol (10 mL) was added aqueous NH3 (35%, 40 mL), and the mixture was stirred at room temperature for 5.5 h. After this time, all volatiles were removed under reduced pressure, and the crude product was purified by reverse-phase flash chromatography on C18 to give AOL NH2 as a white solid (quantitative). 1 H NMR(400MHz,DMSO-d6):δ(ppm)8.86(t,J=5.5Hz,1H,CONH),8.28(s,3H,NH3 + ),7.85(s,1H,Ar-H),7.41(d,J=7.6Hz,1H,Ar-H),7.31(t,J=7.9Hz,1H,Ar-H),7.03(ddd,J=8.1,2.5,1.1 Hz,1H,Ar-H),5.87(ddt,J=17.2,10.5,5.3Hz,1H,OCH2CHCH2),5.24(dq,J=17.2,1.7Hz,1H,OCH2CHCH2,H a ),5.09(dq,J=10.5,1.5Hz,1H,OCH2CHCH2,H b),5.02(dd,J=6.7,2.4Hz,1H,OCHO),4.41(dd,J=12.2,2.5Hz,1H,OCH2,H a ), 4.18-3.99 (m, 3H, OCH2CHCH2 and OCH2, H b ),3.94-3.81(m,2H,OCH2COOH),3.49-3.39(m,1H,CH2,H a ), 3.21-3.10(m,1H,CH2,H b ),2.86-2.70(m,2H,CH2),1.81-1.39(m,6H,CH2). 13 C NMR (101 MHz, DMSO-d): δ (ppm) 172.9, 166.1, 158.0, 136.2, 135.2, 129.3, 120.4, 119.3, 116.0, 111.3, 99.0, 68.8, 67.7, 66.8, 38.7, 38.4, 27.8, 26.3, 23.0. LC-MS (ESI): (negative ion) 379 (M−H); (positive ion) m / z 381 (M+H) + ).

[0203] General procedure for dye-AOL linker coupling: The dye carboxylate (0.15 mmol) was dissolved in 6 mL of anhydrous N,N'-dimethylformamide (DMF). N,N-diisopropylethylamine (136 μL, 0.78 mmol) was added, followed by N,N,N',N'-tetramethyl-O-(N-succinimidyl)uronium tetrafluoroborate as a 0.5 M solution in anhydrous DMF (TSTU, 300 μL, 0.15 mmol). The reaction was stirred under nitrogen at room temperature for 1 h. A solution of AOLNH2 (0.10 mmol) in water (400 μL) was added to the activated dye solution, and the reaction was stirred at room temperature for 3 h. The crude product was purified by preparative RP-HPLC. [ka]

[0204] Characterization of AOL-SO7181: 54% yield (54 μmol). LC-MS (ES): (negative ion) m / z 1002 (MH + (positive ion) m / z 1004 (M+H + ). [ka]

[0205] Characterization of AOL-AF550POPOS0: 88% yield (88 μmol). LC-MS (ES): (negative ion) m / z 1034 (MH + ), 516(M-2H + (positive ion) m / z 1036 (M+H + ), 1137(M+Et3NH + ). [ka]

[0206] Identification of AOL-NR7180A: Yield 24% (23.9 μmol). LC-MS (ES): (Positive ion) m / z=880 (M+H) + . [ka]

[0207] Identification of AOL-NR550S0: Yield 22% (21.9 μmol). LC-MS (ES): (Positive ion) m / z=963 (M+H) + (negative ion) m / z=961 (MH + ). [ka]

[0208] Synthesis of intermediate A2: Compound A1 (319 mg, 0.419 mmol) was dissolved in 0.8 mL of anhydrous DCM under a N atmosphere and then added to pentamethylcyclopentadienyltris(acetonitrile)ruthenium(II) hexafluorophosphate ([RuCp * [(MeCN)3]PF6, 42 mg, 0.08 mmol), followed by triethoxysilane (231 μL, 1.25 mmol) were added. The reaction was stirred under N2 at room temperature for 1 h. The solution was then diluted with DCM and filtered through a plug of silica gel, which was washed with ethyl acetate. The solution was evaporated under reduced pressure and dried under vacuum for 10 min, then the residue was dissolved in 2 mL of anhydrous THF. Copper iodide (15 mg, 0.08 mmol) and a 1 M solution of TBAF in THF (920 μL, 0.919 mmol) were added. The reaction was stirred at room temperature for 2.5 h, then diluted with EtOAc and extracted with saturated NH4Cl. The aqueous phase was extracted with EtOAc. The pooled organic phases were dried over MgSO4, filtered, and evaporated to dryness. The product was purified by flash column chromatography on silica gel. Yield: 125 mg (0.237 mmol). LC-MS (ES and CI): (positive ion) m / z 527 (M+H + ).

[0209] Synthesis of intermediate A3: Nucleoside A2 (155 mg, 0.294 mmol) was dried over PO for 18 hours under reduced pressure. Triethyl phosphate anhydride (1 mL) and some freshly activated 4 Å molecular sieves were added to it under nitrogen, and the reaction flask was then cooled to 0 °C in an ice bath. Freshly distilled POCl (33 μL, 0.353 mmol) was added dropwise, followed by Proton Sponge® (113 mg, 0.53 mmol). After the addition, the reaction was stirred for an additional 15 minutes at 0 °C. A 0.5 M solution of the pyrophosphate as the bis-tri-n-butylammonium salt (2.94 mL, 1.47 mmol) in anhydrous DMF was then added quickly, followed immediately by tri-n-butylamine (294 μL, 1.32 mmol). The reaction was kept in an ice-water bath for an additional 10 minutes, then quenched by pouring it into 1M aqueous triethylammonium bicarbonate (TEAB, 10 mL) and stirring at room temperature for 4 hours. All solvents were evaporated under reduced pressure. 35% aqueous ammonia (10 mL) was added to the above residue, and the mixture was stirred at room temperature for at least 5 hours. The solvent was then evaporated under reduced pressure. The crude product was first purified by ion-exchange chromatography on DEAE-Sephadex A25 (50 g). The column was eluted with aqueous triethylammonium bicarbonate (TEAB). Fractions containing the triphosphate were pooled, and the solvent was evaporated to dryness under reduced pressure. The crude material was further purified by preparative-scale HPLC. Compound A3 was obtained as the triethylammonium salt. Yield: 134 μmol (46%). LC-MS (ESI): (negative ion) m / z 614 (MH + ).

[0210] Furthermore, the structure [ka] The 5'-triphosphate-3'-AOM-A nucleotide and the corresponding ffA were also prepared. Detailed synthesis is described in US Patent Application Publication No. 16 / 724,088.

[0211] General synthesis of nucleotide triphosphate-AOL linkers: Compound AOL (0.120 mmol) was coevaporated with 2 × 2 mL of anhydrous N,N'-dimethylformamide (DMF) and then dissolved in 3 mL of anhydrous N,N'-dimethylacetamide (DMA). N,N-Diisopropylethylamine (70 μL, 0.4 mmol) was added, followed by N,N,N',N'-tetramethyl-O-(N-succinimidyl)uronium tetrafluoroborate (TSTU, 36 mg, 0.120 mmol). The reaction was stirred under nitrogen at room temperature for 1 hour. Meanwhile, an aqueous solution of nucleotide triphosphate (0.08 mmol) was evaporated to dryness under reduced pressure and resuspended in 300 μL of 0.1 M aqueous triethylammonium bicarbonate (TEAB). The activated linker solution was added to the triphosphate salt, and the reaction was stirred at room temperature for 18 hours and monitored by RP-HPLC. The solution was concentrated, and then 10 mL of concentrated aqueous NH4OH was added. The reaction was stirred at room temperature for 24 hours, and then evaporated under reduced pressure. The crude product was first purified by ion exchange chromatography on DEAE-Sephadex A25 (50 g) eluting with aqueous triethylammonium bicarbonate (TEAB). Fractions containing the triphosphate were pooled, and the solvent was evaporated to dryness under reduced pressure. The crude material was further purified by preparative-scale HPLC. [ka]

[0212] Characterization of pppA(DB)-(3'AOM)-AOL: Yield: 60 μmol (75%). LC-MS (ES): (negative ion) m / z 976 (MH + ), 488(M-2H + ). [ka]

[0213] Identification of pppA(DB)-(3'AOM)-AOL: Yield: 68 μmol (72%). 1H NMR(400MHz,D2O): δ(ppm)7.94(d,J=1.7Hz,1H,H-2),7.32(d,J=1.6Hz,1H,H-8),7.10-6.96(m,2 H,Ar),6.95-6.87(m,2H,Ar),6.41-6.31(m,1H,1'-CH),6.01-5.81(m,2H,CHアリル),5.29(ddq,J=1 7.2,6.1,1.4Hz,2H,CHHアリル),5.20(ddt,J=10.5,3.8,1.1Hz,2H,CHHアリル),5.02(td,J=4.4,2.4Hz ,1H,O-CH2-Oリンカー),4.92-4.81(m,2H,3'-O-CH2-O),4.53(dd,J=4.9,2.4Hz,1H,3'-CH),4.39(dd, J=16.1,4.0Hz,1H,O-CHH リンカー),4.32-4.19(m,2H,O-CHH リンカー,4'-CH),4.17-4.01(m,8H,5'-CH2, CH2Oリンカー,CH2-Oアリル),3.25-3.11(m,2H,CH2-NHCO),3.04(q,J=7.3Hz,18H,Et3N),2.92-2.82(m,2 H,CH2-Nリンカー),2.55-2.41(m,2H,2'-CH2),1.59(p,J=7.6Hz,2H,CH2-CH2-Nリンカー),1.47(p,J=7.1 Hz,2H,CH2-CH2-NHCO),1.31(tt,J=8.3,4.4Hz,2H,CH2-CH2-CH2-Nリンカー),1.15(t,J=7.3Hz,26H). 31 P NMR(162MHz,D2O):δ(ppm)-6.18(d,J=20.6Hz, γ P), -11.32(d, J=19.3Hz, α P), -22.20(t, J=19.9Hz, β P).

change

[0214] Synthesis of intermediate C1 5-Iodo-5'-O-(tert-butyldiphenylsilyl)-2'-deoxycytidine (3 g, 5.07 mmol) was dissolved in 30 mL of anhydrous pyridine, followed by the dropwise addition of chlorotrimethylsilane (1.29 mL, 10.1 mmol). The reaction was stirred at room temperature for 1 hour, then placed in an ice bath, and benzoyl chloride (648 μL, 5.6 mmol) was added slowly, dropwise. The reaction was removed from the ice bath and stirred at room temperature for 1 hour. Upon completion, the solution was placed in an ice bath and quenched with 50 mL of cold water. 50 mL of methanol and 20 mL of pyridine were then added, and the suspension was stirred at room temperature overnight. The solvent was evaporated under reduced pressure, and the residue was dissolved in 200 mL of EtOAc and extracted with 2×200 mL of saturated NaHCO₃ and 100 mL of brine. The organic phase was dried over MgSO₄, filtered, and evaporated to dryness. The crude material was purified by flash chromatography on silica gel to give C1. Yield: 2.535 g (3.64 mmol, 73%). LC-MS (ESI): (positive ion) m / z 696 (M+H + ), 797(M+Et3NH + ).

[0215] Synthesis of intermediate C2 N-Benzoyl-5-iodo-5'-O-(tert-butyldiphenylsilyl)-2'-deoxycytidine (C1) (695 mg, 1 mmol) and palladium(II) acetate (190 mg, 0.85 mmol) were dissolved in dry, degassed DMF (10 mL), followed by the addition of N-allyltrifluoroacetamide (7.65 mL, 5 mmol). The solution was placed under vacuum and purged with nitrogen three times, followed by the addition of degassed triethylamine (278 μL, 2 mmol). The solution was heated to approximately 80°C and protected from light for 1 hour. The resulting black mixture was cooled to room temperature, then diluted with 50 mL of EtOAc and extracted with 100 mL of water. The aqueous phase was then extracted with EtOAc. The organic phases were pooled, dried over MgSO4, filtered, and evaporated to dryness. The crude material was purified by flash chromatography on silica gel to give C2. Yield: 305 mg (0.42 mmol, 42%). LC-MS (ES and CI): (positive ion) m / z 721 (M+H +), 797(M+Et3NH + ).

[0216] Synthesis of intermediate C3 N-Benzoyl-5-[3-(2,2,2-trifluoroacetamido)-allyl]-5'-O-(tert-butyldiphenylsilyl)2'-deoxycytidine (C2) (350 mg, 0.486 mmol) was dissolved in 1.1 mL of anhydrous DMSO (14.5 mmol), followed by the addition of glacial acetic acid (1.7 mL, 29.1 mmol) and acetic anhydride (1.7 mL, 17 mmol). The reaction was heated to 60 °C for 6 h and then quenched with 50 mL of saturated NaHCO. After the solution stopped bubbling, it was extracted with EtOAc. The organic phases were pooled and washed with saturated aqueous NaHCO, water, and brine. The organic phase was dried over MgSO, filtered, and evaporated to dryness. The crude material was purified by flash chromatography on silica gel to give C3. Yield: 226 mg (0.289 mmol, 60%). LC-MS (ESI): (positive ion) m / z 781 (M+H + ), 882(M+Et3NH + ).

[0217] Synthesis of intermediate C4 N-Benzoyl-5-[3-(2,2,2-trifluoroacetamido)-allyl]-5'-O-(tert-butyldiphenylsilyl)-3'-O-methylthiomethyl-2'-deoxycytidine (C3) (210 mg, 0.27 mmol) was dissolved in 5 mL of anhydrous DCM under a N atmosphere, cyclohexene (136 μL, 1.35 mmol) was added, and the solution was cooled to approximately −10 °C. Freshly distilled 1 M solution of sulfuryl chloride in anhydrous DCM (320 μL, 0.32 mmol) was added dropwise, and the reaction was stirred for 20 min. After all starting material was consumed, an extra portion of cyclohexene was added (136 μL, 1.35 mmol), and the reaction was evaporated to dryness under reduced pressure. The residue was quickly purged with nitrogen, then dissolved in 2.5 mL of ice-cold anhydrous DCM, and ice-cold allyl alcohol (2.5 mL) was added with stirring at 0 °C. The reaction was stirred at 0 °C for 3 h, then quenched with saturated aqueous NaHCO3, and then further diluted with saturated aqueous NaHCO3. The mixture was extracted with EtOAc. The pooled organic phases were dried over MgSO4, filtered, and evaporated to dryness. The residue was purified by flash chromatography on silica gel to give C4. Yield: 58% (124 mg, 0.157 mmol). LC-MS (ESI): (positive ion) m / z 791 (M+H + ).

[0218] Synthesis of intermediate C5 N-Benzoyl-5-[3-(2,2,2-trifluoroacetamido)-allyl]-5'-O-(tert-butyldiphenylsilyl)-3'-O-allyloxymethyl-2'-deoxycytidine (C4) (120 mg, 0.162 mmol) was dissolved in dry THF (5 mL) under a N atmosphere and then placed at 0 °C. Glacial acetic acid (29 μL, 0.486 mmol) was added, followed immediately by a solution of 1.0 M TBAF in THF (486 μL, 0.486 mmol). The solution was stirred at 0 °C for 3 h. The solution was diluted with EtOAc and then extracted with 0.025 N HCl and brine. The organic phase was dried over MgSO, filtered, and evaporated to dryness. The residue was purified by flash chromatography on silica gel to give C5. Yield: 50 mg (0.090 mmol, 55%). LC-MS (ESI): (positive ion) m / z 553 (M+H + (negative ion) m / z 551 (MH + ), 587(M+Cl - ).

[0219] Synthesis of intermediate C6 N-Benzoyl-3'-O-allyloxymethyl-5-[3-(2,2,2-trifluoroacetamido)-allyl]-2'-deoxycytidine (C5) (50 mg, 0.0.09 mmol) was dried over PO under reduced pressure for 18 hours. Triethyl phosphate anhydride (1 mL) and some freshly activated 4 Å molecular sieves were added to it under nitrogen, and the reaction flask was then cooled to 0 °C in an ice bath. Freshly distilled POCl (10 μL, 0.108 mmol) was added dropwise, followed by Proton Sponge® (29 mg, 0.135 mmol). After the addition, the reaction was stirred for an additional 15 minutes at 0 °C. A 0.5 M solution of pyrophosphate as the bis-tri-n-butylammonium salt (1 mL, 0.45 mmol) in anhydrous DMF was then added quickly, followed immediately by tri-n-butylamine (100 μL, 0.4 mmol). The reaction was kept in an ice-water bath for an additional 10 minutes, and then quenched by pouring it into 1 M aqueous triethylammonium bicarbonate (TEAB, 5 mL) and stirring at room temperature for 4 hours. All solvents were evaporated under reduced pressure. 35% aqueous ammonia (5 mL) was added to the above residue, and the mixture was stirred at room temperature for 18 hours. The solvent was then evaporated under reduced pressure, and the residue was resuspended in 10 mL of 0.1 M TEAB and filtered. The filtrate was first purified by ion exchange chromatography on DEAE-Sephadex A25 (50 g). The column was eluted with aqueous triethylammonium bicarbonate. Fractions containing the triphosphate were pooled, and the solvent was evaporated to dryness under reduced pressure. The crude material was further purified by preparative-scale HPLC to give compound C6 as the triethylammonium salt. Yield: ε 290 =5041M -1 cm -1 40.6 μmol (45%) based on 1H NMR(400MHz,D2O):δ(ppm)8.23(s,1H,H-6),6.53(dd,J=15.5,0.9Hz,1H,Ar-CH =),6.42-6.24(m,2H,Ar-CH=CH-,1'-CH),5.98(ddt,J=17.3,10.4,5.9Hz,1H,O -CH2-CH=),5.37(dq,J=17.3,1.6Hz,1H,CHH=),5.29(ddt,J=10.4,1.6,1.1Hz, 1H,CHH=),4.89(s,2H,O-CH2-O),4.60(dt,J=6.2,3.1Hz,1H,3'-CH),4.39(t,J= 2.7Hz,1H,4'-CH),4.35(dq,J=12.0,3.8Hz,1H,5'-CHH),4.28-4.21(m,1H,5'- CHH),4.20(ddt,J=6.0,2.7,1.4Hz,1H,=CH-CH2-O),3.73(dt,J=7.2,1.4Hz,2H, CH2-NH2),3.18(q,J=7.3Hz,20H,Et3N),2.59(ddd,J=14.1,6.1,3.3Hz,1H,2'- CHH),2.37(ddd,J=14.2,7.2,6.1Hz,1H,2'-CHH),1.27(t,J=7.3Hz,31H,Et3N). 31 P NMR(162MHz,D2O):δ(ppm)-6.06(d,J=20.7Hz, γ P),-11.24(d,J=19.1Hz, α P),-21.95(t,J=19.7Hz, β P). LC-MS (ESI): (negative ion) m / z 591 (MH + ).

[0220] Furthermore, 5'-triphosphate-3'-AOM-C nucleotides, 5'-triphosphate-3'-AOM-T(DB) nucleotides of the following structures: [ka] The corresponding ffC and ffT (DB) were also prepared. Finally, 5'-triphosphate-3'-AOM-G (also called ffG-(3'-AOM)) [ka] was also prepared. Detailed synthesis is described in U.S. Publication No. 2020 / 0216891. General synthesis of fully functionalized nucleotides bearing an AOL linker moiety

[0221] Dye-COOH (0.02 mmol) or Dye-AOL (0.02 mmol) was coevaporated with 2 x 2 mL of anhydrous N,N'-dimethylformamide (DMF) and then dissolved in 2 mL of anhydrous N,N'-dimethylacetamide (DMA). N,N-Diisopropylethylamine (17 μL, 0.1 mmol) was added, followed by N,N,N',N'-tetramethyl-O-(N-succinimidyl)uronium tetrafluoroborate (TSTU, 6 mg, 0.02 mmol). The reaction was stirred under nitrogen at room temperature for 1 h. Meanwhile, an aqueous solution of nucleotide triphosphate (0.01 mmol) was evaporated to dryness under reduced pressure and resuspended in 200 μL of 0.1 M aqueous triethylammonium bicarbonate (TEAB). The activated dye solution was added to the nucleotide triphosphate, and the reaction was stirred at room temperature for 18 h and monitored by RP-HPLC. The crude product was first purified by ion-exchange chromatography on DEAE-Sephadex A25 (25 g) eluting with a linear gradient of aqueous triethylammonium bicarbonate (TEAB, 0.1 M to 1 M). Fractions containing the triphosphate were pooled and the solvent was evaporated to dryness under reduced pressure. The crude material was further purified by preparative-scale HPLC. [ka]

[0222] Identification of ffA-(3'AOM)-AOL-BL-NR650C5: Yield 66% (6.6 μmol). LC-MS (ES): (negative ion) m / z 1931 (MH + ), 965(M-2H + ). [ka]

[0223] Identification of ffA-(3'AOM)-AOL-BL-NR550s0: Yield 65% (6.5 μmol). LC-MS (ES): (negative ion) m / z 1784 (MH + ), 891(M-2H + ). [ka]

[0224] Identification of ffA(DB)-(3'AOM)-AOL-BL-NR650C5: Yield 31% (3.1 μmol). LC-MS (ES): (negative ion) m / z 1933 (MH + ), 965(M-2H + ), 644(M-3H + ). [ka]

[0225] Identification of ffA(DB)-(3'AOM)-AOL-NR7181A: Yield 21% (2.12 μmol). LC-MS (ES): (negative ion) m / z = 1475 (MH + ). [ka]

[0226] Identification of ffA-(3'-AOM)-AOL-NR7180A: Yield 41% (4.1 μmol). LC-MS (ES): (negative ion) m / z 1472 (MH + ), 736(M-2H + ). [ka]

[0227] Identification of ffA(DB)-(3'-AOM)-AOL-BL-NR550S0: Yield 21% (2.1 μmol). LC-MS (ES): (negative ion) m / z 1786 (MH+ ), 892(M-2H + ), 594(M-3H + ). [ka]

[0228] Identification of ffC(DB)-(3'AOM)-AOL-SO7181: Yield 48%, (4.87 μmol). LC-MS (ES): (negative ion) m / z 1577 (MH + ), 788(M-2H + ), 525(M-3H + ). [ka]

[0229] Identification of ffC-(3'-AOM)-AOL-SO7181: Yield 56% (5.6 μmol). LC-MS (ES): (negative ion) m / z 1575 (MH + ), 787(M-2H + ). [ka]

[0230] Identification of ffT(DB)-3'AOM-AOL-AF550POPOS0: Yield 46% (4.6 μmol). LC-MS (ES): (negative ion) m / z 1609 (MH + ), 804(M-2H + ), 536(M-3H + ). [ka]

[0231] Identification of ffT(DB)-(3'AOM)-AOL-NR550s0: Yield 38% (3.8 μmol). LC-MS (ES): (negative ion) m / z = 1535 (MH + ).

[0232] Example 2. Solution cleavage efficiency of different palladium reagent formulations Figure 1 shows a comparison of the cleavage efficiency of three different formulations of palladium reagent: 1) 10 mM [(allyl)PdCl], 100 mM THP, 100 mM ethanolamine buffer, 10 mM sodium ascorbate; 2) 20 mM NaPdCl, 60 mM THP, 100 mM N,N'-diethylethanolamine, 10 mM sodium ascorbate, 1 M NaCl, 0.1% Tween 20; and 3) 20 mM NaPdCl, 70 mM THP, 100 mM N,N'-diethylethanolamine, 10 mM sodium ascorbate, 1 M NaCl, 0.1% Tween 20. Cleavage efficiency was determined by measuring the relative rate of cleavage of 3'-AOM-nucleotide substrates. Briefly, a stock solution of palladium reagent was added to a 0.1 mM solution of 3'-AOM nucleotide substrate in 100 mM buffer to a final concentration of 1 mM in Pd species. The solution was incubated at room temperature, and reaction kinetics was monitored by withdrawing aliquots from the reaction at set time points, quenching with a 1:1 solution of EDTA / HO (0.1:0.1 M), and analyzing by HPLC for the formation of 3'-OH nucleotides and the disappearance of 3'-AOM nucleotide substrate. As shown in Figure 1, the cleavage efficiency of NaPdCl is comparable to that of [(allyl)PdCl] using only 3 or 3.5 equivalents of THP, compared with 10 equivalents of THP with [(allyl)PdCl].

[0233] Example 3: AOM-AOL ffN stability testing in solution and performance in sequencing Figure 2 shows the prephasing performance of fully functionalized nucleotides (ffNs) (including labeled ffT-DB, labeled ffA, and labeled ffC, and unlabeled ffG) with 3'-AOM blocking groups and AOL linker moieties compared to stressed standard MiniSeq® ffNs. Two sets of ffNs were incubated at 45°C for several days in a standard integration mix formulation, excluding DNA polymerase. At each time point, fresh polymerase was added to complete the integration mix before filling in the MiniSeq®. The previously described sequencing conditions were used. The % prephasing is a direct indicator of the proportion of 3'OH-ffNs present in the mixture and therefore directly correlates with the stability of the 3'-blocking groups. Prephasing values ​​for both ffN sets were recorded and plotted (Figure 2). Compared to the standard, AOM-AOL-ffNs showed no increased prephasing and appeared to be substantially more stable than standard ffNs with 3'-O-azidomethyl blocking groups and LN3 linker moieties.

[0234] Example 4. Use of palladium scavengers in sequencing reactions Figure 3 shows a comparison of phasing values ​​on an Illumina MiniSeq® instrument using fully functionalized nucleotides (ffNs) with a 3'-AOM blocking group and a standard LN3 linker moiety (including labeled ffT-DB, labeled ffA, labeled ffC, and unlabeled ffG) with and without potassium isocyanoacetate in the post-cleavage wash step. Sequencing experiments were performed on an Illumina MiniSeq® cartridge using a cartridge in which the standard incorporation mix was replaced with a freshly prepared incorporation mix containing ffNs with a 3'-AOM blocking group and a standard LN3 linker moiety, and a freshly prepared solution of palladium cleavage reagent (10 mM [(allyl)PdCl], 100 mM THP, 100 mM ethanolamine buffer, 10 mM sodium ascorbate, 1 M NaCl, 0.1% Tween 20) was added to the empty wells. Potassium isocyanoacetate was added to the standard Miniseq® post-cleavage wash solution to a final concentration of 10 mM. Sequencing experiments were performed using a standard sequencing-by-synthesis (SBS) protocol plus a 2 x 151 cycle recipe that included a 5-second incubation with a solution of palladium cleavage reagent. As shown in Figure 3, the % fading was shown to decrease from 0.183 to 0.075 when 10 mM potassium isocyanoacetate was used in the post-cleavage wash solution.

[0235] Figure 4 shows primary sequencing metrics, including phasing, prephasing, and error rates, on an Illumina MiniSeq® instrument using fully functionalized nucleotides (ffNs) with a 3'-AOM blocking group and an AOL linker moiety (including labeled ffT-DB, labeled ffA, and labeled ffC, and unlabeled ffG), compared to the same sequencing metrics using standard ffNs with a 3'-O-azidomethyl blocking group and an LN3 linker moiety, using a palladium scavenger. Sequencing experiments were performed on an Illumina MiniSeq® system by running a 2x151-cycle recipe using a standard cartridge, in which the incorporation mix and standard cleavage reagent were replaced with a freshly prepared incorporation mix containing fully functionalized nucleotides (ffNs) with a 3'-AOM blocking group and an AOL linker moiety and a freshly prepared solution of palladium cleavage reagent (10 mM [(allyl)PdCl], 100 mM THP, 100 mM N,N'-diethylethanolamine buffer, 10 mM sodium ascorbate, 1 M NaCl, 0.1% Tween 20), respectively. Potassium isocyanoacetate (ICNA) was added to the standard MiniSeq® post-cleavage wash solution to a final concentration of 10 mM. Control sequencing experiments with standard ffNs using a 3'-O-azidomethyl blocking group were performed using the standard MiniSeq® kit and recipe. The results showed improved prephasing and, more importantly, error rates, demonstrating the full efficiency of this AOM-AOL SBS chemistry with a single cleavage step.

[0236] Example 5. Use of Glycine in Sequencing Reactions Figure 5 shows a comparison of primary sequencing metrics, including phasing and prephasing, on an Illumina MiniSeq® instrument using fully functionalized nucleotides (ffNs) with a 3'-AOM blocking group and an AOL linker moiety when glycine or ethanolamine, respectively, are used in the incorporation mixture. In this example, the full set of ffNs includes 3'-AOM-ffAs-AOL (labeled), 3'-AOM-ffG-AOL (unlabeled), 3'-AOM-ffT(DB)-AOL (labeled), and 3'-AOM-ffC(DB)-AOL (labeled). As shown in Figure 5, when glycine was used in the incorporation buffer, a significant decrease in phasing values ​​was observed compared to the standard incorporation buffer containing ethanolamine. When glycine was used, there was a slight increase in prephasing values, but this increase was not considered significant. Sequencing experiments were performed on a standard MiniSeq® instrument using cartridges, replacing the standard incorporation and cleavage reagents with a freshly prepared incorporation mixture containing fully functionalized nucleotides (ffNs) with a 3'-AOM blocking group and an AOL linker moiety in either 50 mM ethanolamine or 50 mM glycine buffer, respectively, and a freshly prepared solution of palladium cleavage reagent (10 mM [(allyl)PdCl], 100 mM THP, 100 mM N,N'-diethylethanolamine buffer, 10 mM sodium ascorbate, 1 M NaCl, 0.1% Tween 20). Potassium isocyanoacetate (ICNA) was added to the standard MiniSeq post-cleavage wash solution to a final concentration of 10 mM. The standard 2x151 cycle MiniSeq recipe was used.

[0237] Figure 6 shows primary sequencing metrics, including phasing, prephasing, and error rates, on an Illumina MiniSeq instrument using fully functionalized nucleotides (ffNs) with 3'-AOM and AOL linker moieties compared to the same sequencing metrics using standard ffNs and 3'-O-azidomethyl blocking groups and LN3 linker moieties. In this example, the full set of ffNs includes 3'-AOM-ffA-AOL (labeled), 3'-AOM-ffG-AOL (unlabeled), 3'-AOM-ffT(DB)-AOL (labeled), and 3'-AOM-ffC(DB)-AOL (labeled). Sequencing experiments were performed on a standard MiniSeq® instrument loaded with cartridges using a 2 × 151-cycle recipe, replacing the standard incorporation mix and standard cleavage reagent with a freshly prepared incorporation mix containing fully functionalized nucleotides (ffNs) with 3'-AOM blocking groups in 50 mM glycine buffer and a freshly prepared solution of palladium cleavage reagent (10 mM [(allyl)PdCl], 100 mM THP, 100 mM N,N'-diethylethanolamine buffer, 10 mM sodium ascorbate, 1 M NaCl, 0.1% Tween 20), respectively. Potassium isocyanoacetate was added to the standard MiniSeq® post-cleavage wash solution to a final concentration of 10 mM. Results showed further improvement in error rate compared to that achieved with standard ffNs. The improved phasing of the AOM-AOL series compared to the metrics in Figure 5 is attributed to the use of glycine buffer and the use of 3'-AOM-ffC(DB)-AOL.

[0238] Example 6. Sequencing by synthesis for iSeq™ Figures 7A and 7B show a comparison of primary sequencing metrics, including error rate and Q30 score, from 2 x 300 cycles of sequencing-by-synthesis performed on an Illumina iSeq™ instrument using fully functionalized nucleotides (ffNs) with 3'-AOM blocking groups and AOL linker moieties. In this example, the complete set of AOM ffNs includes 3'-AOM-ffA(DB)-AO-dye1,3'-AOM-ffG (unlabeled), 3'-AOM-ffT(DB)-AOL-NR550S0, and 3'-AOM-ffC(DB)-AOL-dye2. The complete set of azidomethyl (AZM) ffNs includes the same ffNs with a 3'-azidomethyl blocking group and includes a propargylamide and LN3 linker. Dye 1 is disclosed in U.S. Patent No. 63 / 127061 and, when conjugated to ffA, provides the structural moiety [ka] Dye 2 is a coumarin dye disclosed in US 2018 / 0094140, which when conjugated with ffC has the structural moiety: [ka] NR550S0 is a known green pigment.

[0239] Sequencing experiments were performed on a standard iSeq™ instrument using cartridges in which the standard incorporation mix and standard cleavage reagent were replaced with a freshly prepared incorporation mix containing fully functionalized nucleotides (ffNs) with 3'-AOM blocking groups and AOL linker moieties in either 50 mM ethanolamine or 50 mM glycine buffer, using 300% concentration of Polymerase 1901 (Pol1901) (360 μg / mL), and a freshly prepared solution of palladium cleavage reagent (10 mM [(allyl)PdCl], 100 mM THP, 100 mM N,N'-diethylethanolamine buffer, 10 mM sodium ascorbate, 1 M NaCl, 0.1% Tween 20), respectively. L-cysteine ​​was added to the standard iSeq™ post-cleavage wash solution to a final concentration of 10 mM. A 2x301 cycle iSeq™ recipe was used with a two excitation / one emission protocol. Specifically, the iSeq™ instrument captured the first image with green excitation light (approximately 520 nm) and the second image with blue excitation light (approximately 450 nm). A standard sequencing recipe was used to perform 2x301 SBS cycles (integration, followed by imaging, followed by cleavage). Sequencing metrics are summarized in the table below. [Table 2]

[0240] The AOM ffN set was observed to yield superior performance, providing superior error rates and Q30s for both Read1 and Read2. Phasing values ​​using the AOM ffN set were comparable to those produced by the AZM ffN set. However, the AOM ffN set produced significantly lower prephasing values.

[0241] Figures 8A and 8B show a comparison of primary sequencing metrics, including error rate and Q30 score, from 2 x 150 cycles of sequencing-by-synthesis performed on an Illumina iSeq™ instrument using fully functionalized nucleotides (ffNs) with 3'-AOM blocking groups and AOL linker moieties. In this example, the complete set of AOM ffNs includes 3'-AOM-ffA(DB)-AO-dye1,3'-AOM-ffG (unlabeled), 3'-AOM-ffT(DB)-AOL-NR550S0, and 3'-AOM-ffC(DB)-AOL-dye2. The complete set of AZM ffNs includes the same ffNs with a 3'-azidomethyl blocking group and includes a propargylamide and LN3 linker. Sequencing experiments were performed on a standard iSeq™ instrument using cartridges, with the standard incorporation and cleavage reagents replaced with a freshly prepared incorporation mixture containing fully functionalized nucleotides (ffNs) bearing a 3'-AOM blocking group and an AOL linker moiety in either 50 mM ethanolamine or 50 mM glycine buffer using 300% concentration of Pol1901 (360 μg / mL), and a freshly prepared solution of palladium cleavage reagent (10 mM [(allyl)PdCl], 100 mM THP, 100 mM N,N'-diethylethanolamine buffer, 10 mM sodium ascorbate, 1 M NaCl, 0.1% Tween 20). L-cysteine ​​was added to the standard iSeq™ post-cleavage wash solution to a final concentration of 10 mM. A 2 × 301-cycle iSeq™ recipe was used with a two-excitation / one-emission protocol. Specifically, the iSeq™ instrument captured the first image with green excitation light (approximately 520 nm) and the second image with blue excitation light (approximately 450 nm). Using a standard sequencing recipe, 2 x 301 SBS cycles (integration, followed by imaging, followed by cleavage) were performed.

[0242] The integration mixture contact time for AZM ffN was approximately 24.1 seconds, while the integration mixture contact time for AOM ffN was approximately 29.1 seconds. However, the longer integration time for AOM ffN was compensated for by a faster deblocking time. The cleavage mixture contact time was approximately 5.8 seconds, as opposed to approximately 10.2 seconds for the AZM ffN set. Therefore, the total incubation times for the AZM and AOM ffN sets were approximately 34.4 and 34.9 seconds, respectively. Sequencing metrics are summarized in the table below. [Table 3]

[0243] Furthermore, it was observed that the AOM ffN set yielded superior performance, providing superior error rates and Q30 for both Read 1 and Read 2. Furthermore, the AOM ffN set produced lower prephasing values.

[0244] Example 7. Sequencing by synthesis on NovaSeq™ with increasing blue laser power It has been observed that prolonged exposure to blue light in SBS sequencing results in high levels of signal delay and fading as a result of increased light dose and power density. This example compares the performance of AOM ffN and AZM ffN in blue laser titration sequencing experiments.

[0245] In this experiment, SBS of 1x151 runs was compared using an AOM ffN set configured for modified blue / green excitation NovaSeq™ with a standard ffN set containing a 3' azidomethyl blocking group and LN3 linker. The flow cell used was a 490 nm pitch BEER2 flow cell. The power configurations for the blue laser output were 600 mW, 800 mW, 1000 mW, 1400 mW, 1800 mW, and 2400 mW. The green laser power was constant at 1000 mW. The following standard AZM ffNs were used: green ffT (LN3-AF550POPOS0), dark G, red ffC (LN3-SO7181), blue ffC (sPA-Blue Dye A), blue ffA (LN3-BL-Blue Dye A), and green ffA (LN3-BL-NR550S0). The following ffNs were used in the AOM ffN set: green ffT (ffT(DB)-AOL-AF550POPOS0), dark G, red ffC (ffC(DB)-AOL-SO7181), blue ffC (ffC(DB)-AOL-blue dye A), blue ffA (ffA(DB)-AOL-BL-blue dye A), and green ffA (ffA(DB)-AOL-BL-NR550S0). The structures of the blue dyes labeled AOM ffC and ffA are shown below. [ka]

[0246] The following modifications were made to the AOM ffN set SBS run. First, 10 mM L-cysteine ​​was added to the post-cleavage wash solution. The cleavage mixture contained the following components: Na2PdCl4 in DEEA buffer. Two 10-second wait steps were added to the post-cleavage wash step. In addition, a static integration wait time of 60 seconds was used (as opposed to 38 seconds for AZM ffN). Figure 9A shows that both ffN sets had similar phasing values ​​at lower blue laser powers, but the AOM ffN set was less sensitive to increased blue laser power ramps (as indicated by a gentler phasing slope compared to that of the AZM phasing slope). Furthermore, it was observed that the AOM ffN had much less prephasing. As shown in Figure 9B, the AOM ffN set exhibited significantly less signal decay at higher blue laser powers. Figures 9C and 9D show the average error rate as a function of cycle number. Figure 9D is an expanded view of Figure 9C. The results show that at lower blue laser powers, the AZM ffN set had a lower error rate in the early cycles, but at higher blue laser powers, the AOM ffN set performed much better in later cycles. Figure 9E summarizes the average error rate over 151 cycles. Again, the AOM ffN set outperformed the AZM ffN set set at higher laser powers (e.g., at 1400 mW, 1800 mW, and 2400 mW).

[0247] Example 8. First chemical linearization using Pd cleavage mixture In this example, the Pd cleavage mixture used for SBS was tested in the first chemical linearization step after the clustering step. The experiment compared chemical linearization with standard enzymatic linearization, in which cleavage of one of the double-stranded polynucleotides was facilitated by USER, cleaving the U position on the P5 primer. One 150-cycle SBS was performed on an Illumina iSeq™ instrument using fully functionalized nucleotides (ffNs) bearing a 3'-AOM blocking group and an AOL linker moiety. In this example, the complete set of AOM ffNs included 3'-AOM-ffA(DB)-AO-dye1, 3'-AOM-ffG (unlabeled), 3'-AOM-ffT(DB)-AOL-NR550S0, and 3'-AOM-ffC(DB)-AOL-dye2, as described in Example 6. The iSeq™ instrument was configured to capture the first image using green excitation light (approximately 520 nm) and the second image using blue excitation light (approximately 450 nm) (using a two-excitation / one-emission protocol). The flow cell used in the iSeq™ instrument was grafted with modified P5 / P7 primers to enable initial chemical linearization of the P5 primer. The actinic linearization step was performed in a Pd cleavage mixture (10 mM [Pd(allyl)Cl]2 and 100 mM THP in a buffer solution containing DEEA) incubated at 63°C for 30 seconds. SBS sequencing metrics using the two different linearization methods are shown in Figure 10. When the Pd cleavage mixture was used in the first chemical linearization step, all primary sequencing metrics were observed to fall within the standard observation range. This experiment confirmed that a single reagent mixture can be used for two separate steps of sequencing, i.e., the linearization step and the SBS cleavage step, allowing for further instrument (fluidics and cartridge) simplification. The present application also includes the following aspects. [Aspect 1] A nucleotide or nucleoside comprising a nucleobase attached to a detectable label via a cleavable linker, said nucleoside or nucleotide comprising a ribose or 2' deoxyribose moiety and a 3'-OH blocking group, said cleavable linker comprising a moiety of the structure: [ka] During the ceremony, Each of X and Y is independently O or S; R 1a 、R 1b 、R 2 、R 3a and R 3b each independently represents H, halogen, unsubstituted or substituted C 1 ~C 6 Alkyl, or C 1 ~C 6 A nucleotide or nucleoside that is haloalkyl. [Aspect 2] A nucleoside or nucleotide according to embodiment 1, comprising the structure of formula (I):

change

change

change

change

change

change

change

change

change

change

Claims

1. 1. A polynucleotide sequencing material comprising a nucleoside or nucleotide comprising a nucleic acid base linked to a detectable label via a cleavable linker, The nucleoside or nucleotide has the formula (I): 【Chemistry 1】 wherein B is the nucleobase; the label is a detectable label; R 4 is H or OH, R 5 teeth, 【Chemistry 2】 and R a , R b , R c , R d and R e each independently represents H, halogen, unsubstituted C 1 ~C 6 Alkyl, or C 1 ~C 6 is haloalkyl, R 6 is H or a hydroxy protecting group, or -OR 6 is a monophosphate, diphosphate, triphosphate, or thiophosphate; L is, 【Transformation 3】 and Each of X and Y is O or S; R 1a , R 1b , R 2 , R 3a and R 3b each independently represents H, halogen, unsubstituted C 1 ~C 6 Alkyl, or C 1 ~C 6 is haloalkyl, L 1 and L 2 each of which is independently optionally present; L 1 is 【Chemistry 4】 and L2 is 【Transformation 5】 wherein each of m and n is independently an integer of 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10. having the structure The polynucleotide sequencing material.

2. 2. The polynucleotide sequencing material of claim 1, wherein the detectable label is a fluorescent dye.

3. 3. The polynucleotide sequencing material of claim 1, wherein each of X and Y is O.

4. R 1a , R 1b , R 2 , R 3a and R 3b The polynucleotide sequencing material according to any one of claims 1 to 3, wherein each of

5. R 1a , R 1b , R 2 , R 3a and R 3b At least one of the groups is halogen or unsubstituted C 1 ~C 6 The polynucleotide sequencing material according to any one of claims 1 to 3, wherein the polynucleotide is alkyl.

6. The polynucleotide sequencing material according to any one of claims 1 to 5, wherein B is a purine, a deazapurine, or a pyrimidine.

7. R 5 but, 【Transformation 6】 2. The polynucleotide sequencing material of claim 1, wherein:

8. L 1 The polynucleotide sequencing material according to any one of claims 1 to 7, wherein

9. Formula (Ia), (Ia'), (Ib), (Ic), (Ic') or (Id): 【Transformation 7】 【Transformation 8】 2. The polynucleotide sequencing material of claim 1, comprising the structure:

10. L 2 2. The polynucleotide sequencing material of claim 1, wherein:

11. 11. The polynucleotide sequencing material of claim 10, wherein n is 5.

12. 11. The polynucleotide sequencing material of claim 10, wherein m is 4.

13. R 4 is H, and -OR 6 2. The polynucleotide sequencing material of claim 1, wherein is a triphosphate.

Citation Information

Patent Citations

  • Deoxyuridine compounds, methods for preparing them and their use in medicine

    EP0104857A1

  • 2'-deoxy-5-substituted uridine derivative, its preparation and antitumor agent containing it

    JP1984036696A

  • modified nucleotide

    JP2006509040A

  • Modified nucleoside or modified nucleotide

    JP2016512206A

  • Nucleosides and nucleotides having a 3'-hydroxy blocking group and their use in polynucleotide sequencing methods

    JP2022515944A