Nucleosides and nucleotides with 3 'acetal capping groups
The nucleotide incorporation control problem in DNA sequencing is solved by using nucleotides and nucleosides containing 3' acetal end groups, combined with technology that combines cleavable linkers and palladium catalysts, and a more efficient and accurate sequencing process is achieved.
Patent Information
- Application Number
- CN202510319781.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2020-06-22
- Filing Date
- 2021-06-21
- Publication Date
- 2025-06-20
AI Technical Summary
The prior art is difficult to effectively control the incorporation of nucleotides in DNA sequencing, resulting in the introduction of uncontrolled nucleotide molecules, affecting the accuracy and efficiency of sequencing.
Using nucleotides and nucleosides containing 3' acetal capping groups, covalently attached to the detectable labels through a cleavable linker, ensuring that only one nucleotide is incorporated in the sequencing reaction, and the capping groups and labels are removed by cleavage of the palladium catalyst in a subsequent step.
It improves the accuracy and efficiency of sequencing, reduces phase and predetermined phase phenomena, extends the length of sequencing reads, and improves the stability and unblocking rate of nucleotides.
Smart Images

Figure CN120174072A_ABST
Abstract
Description
[0001] This application is a divisional application of the patent application with the application number 202180044416.1, the filing date of June 21, 2021, and the invention title of "Nucleosides and Nucleotides with 3'-Acetal Capping Groups".
[0002] Priority applications incorporated by reference
[0003] This application claims the priority benefit of U.S. Provisional Application No. 63 / 042,240, filed on June 22, 2020, which is incorporated herein by reference in its entirety. BACKGROUND OF THE INVENTION TECHNICAL FIELD
[0005] The present disclosure generally relates to nucleotides, nucleosides or oligonucleotides comprising a 3'-acetal capping group, and their use in polynucleotide sequencing methods. Methods for preparing 3'-capped nucleotides, nucleosides or oligonucleotides are also disclosed. BACKGROUND ART
[0006] Progress in molecular research has, to some extent, been driven by improvements in the techniques used to characterize molecules or their biological reactions. In particular, nucleic acid DNA and RNA research has benefited from the development of techniques for sequence analysis and the study of hybridization events.
[0007] One example of a technique that has improved nucleic acid research is the development of assembled arrays of immobilized nucleic acids. These arrays typically consist of a high-density matrix of polynucleotides immobilized on a solid support material. See, for example, Fodor et al., Trends Biotech. 12:19-26, 1994, which describes a method for assembling nucleic acids using a chemically sensitized glass surface that is protected by a mask but exposed in defined areas to allow attachment of appropriately modified nucleotide phosphoramidites. Assembled arrays can also be fabricated by techniques that "spot" known polynucleotides onto predetermined positions on a solid support (e.g., Stimpson et al., Proc. Natl. Acad. Sci. 92:6379-6383, 1995).
[0008] One method for determining the nucleotide sequence of a nucleic acid that binds to an array is called "sequencing by synthesis" or "SBS". In an ideal scenario, this technique for determining the nucleotide sequence of DNA requires the controlled (i.e., one at a time) incorporation of the correct complementary nucleotide opposite the nucleic acid being sequenced. This enables accurate sequencing by adding nucleotides in multiple cycles, as only one nucleotide residue is sequenced at a time, thereby preventing uncontrolled incorporation into the nucleotide sequence. The incorporated nucleotide is read using an appropriate label attached to the incorporated nucleotide, after which the labeled portion is removed and the subsequent next cycle of sequencing is performed.
[0009] To ensure incorporation of only a single nucleotide, a structural modification (“protecting group” or “blocking group”) is included in each labeled nucleotide, which is added to the growing chain to ensure incorporation of only one nucleotide. After addition of the nucleotide with the protecting group, the protecting group is removed under reaction conditions that do not interfere with the integrity of the DNA being sequenced. The sequencing cycle can then continue with incorporation of the next protected labeled nucleotide.
[0010] For use in DNA sequencing, nucleotides (typically nucleotide triphosphates) generally require a 3′-hydroxyl protecting group to prevent the polymerase used to incorporate them into the polynucleotide chain from continuing replication once a base has been added to the nucleotide. There are many restrictions on the types of groups that can be added to the nucleotide and still be suitable. The protecting group should prevent additional nucleotide molecules from being added to the polynucleotide chain while being easily removable from the sugar moiety without damaging the polynucleotide chain. In addition, the modified nucleotide needs to be compatible with the polymerase or another suitable enzyme used to incorporate it into the polynucleotide chain. Thus, an ideal protecting group must exhibit long-term stability, be efficiently incorporated by the polymerase, block second or further incorporation of the nucleotide, and be removable under mild conditions (preferably under aqueous conditions) without damaging the polynucleotide structure.
[0011] Reversible protecting groups have been described previously. For example, Metzker et al. (Nucleic Acids Research, 22(20):4259-4267, 1994) disclosed the synthesis and use of eight 3′-modified 2-deoxyribonucleoside 5′-triphosphates (3′-modified dNTPs) and tested their incorporation activity in two DNA template assays. WO 2002 / 029003 describes a sequencing method that can include using an allyl protecting group to block the 3′-OH group on a growing DNA strand in a polymerase reaction.
[0012] In addition, the development of many reversible protecting groups and methods for removing these protecting groups under DNA-compatible conditions have been reported previously in International Application Publication Nos. WO 2004 / 018497 and WO 2014 / 139596, each of which is hereby incorporated by reference in its entirety. SUMMARY OF THE INVENTION
[0013] Some embodiments of the present disclosure relate to nucleotides or nucleosides comprising a nucleobase attached via a cleavable linker to a detectable label, wherein the nucleoside or nucleotide comprises a ribose or 2′-deoxyribose moiety and a 3′-OH blocking group, and wherein the cleavable linker comprises a moiety having the following structure:
[0014]
[0015] wherein X and Y are each independently O or S; and R 1a 、R 1b 、R 2 、R 3a and R 3b are each independently H, a halogen, an unsubstituted or substituted C1-C6 alkyl, or a C1-C6 haloalkyl.
[0016] Some embodiments of the present disclosure relate to oligonucleotides or polynucleotides comprising the 3′-OH-terminated labeled nucleotides described herein.
[0017] Some embodiments of the present disclosure relate to a method for preparing a growing polynucleotide complementary to a target single-stranded polynucleotide in a sequencing reaction, the method comprising incorporating a nucleotide molecule described herein into the growing complementary polynucleotide, wherein incorporation of the nucleotide prevents introduction of any subsequent nucleotides into the growing complementary polynucleotide. In some embodiments, incorporation of the nucleotide is accomplished by a polymerase, terminal deoxynucleotidyl transferase (TdT), or reverse transcriptase. In one embodiment, the incorporation is accomplished by a polymerase (e.g., a DNA polymerase).
[0018] Still other embodiments of the present disclosure relate to a method for determining the sequence of a target single-stranded polynucleotide, the method comprising:
[0019] (a) incorporating a nucleotide described herein into a copy polynucleotide chain complementary to at least a portion of the target polynucleotide chain;
[0020] (b) detecting the identity of the nucleotide incorporated into the copy polynucleotide chain; and
[0021] (c) chemically removing the label and the 3′-OH capping group from the nucleotide incorporated into the copy polynucleotide chain.
[0022] In some embodiments, the detection step includes determining the identity of the nucleotides incorporated into the copy polynucleotide chain by making one or more measurements of the fluorescence signal from the detectable label. In some embodiments, the sequencing method further includes (d) washing away from the copy polynucleotide chain the chemically removed label and 3′-OH capping group using a post-cleavage wash solution. In some embodiments, this washing step also removes unincorporated nucleotides. In other embodiments, the method can include a separate washing step for washing away unincorporated nucleotides from the copy polynucleotide chain prior to step (b). In some such embodiments, the 3′-OH capping group and detectable label of the incorporated nucleotide are removed prior to introducing the next complementary nucleotide. In still other embodiments, the 3′-OH capping group and detectable label are removed in a single step of the chemical reaction. In some embodiments, the sequential incorporation described herein is performed at least 50 times, at least 100 times, at least 150 times, at least 200 times, or at least 250 times.
[0023] Still other embodiments of the present disclosure relate to kits containing the polynucleotide or nucleoside molecules described herein and their packaging materials. The nucleotides, nucleosides, oligonucleotides, or kits described herein can be used to detect, measure, or identify biological systems (including, for example, their processes or components). Exemplary techniques that can employ these nucleotides, oligonucleotides, or kits include sequencing, expression analysis, hybridization analysis, genetic analysis, RNA analysis, cell assays (e.g., cell binding or cell function analysis), or protein assays (e.g., protein binding assays or protein activity assays). The use can be on an automated instrument (such as an automated sequencing instrument) for performing a particular technique. The sequencing instrument can include two or more lasers operating at different wavelengths to distinguish different detectable labels. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 is a line graph comparing the 3′-OH deblocking efficiency of [(allyl)PdCl]2 and Na2PdCl4 when using various ratios of tris(hydroxypropyl)phosphine (THP).
[0025] Figure 2 is a line graph showing the change over time of the percentage value of a predetermined phase for a standard fully functionalized nucleotide (ffN) having an LN3 linker moiety and a 3′-O-azidomethyl capping group compared to an ffN having an AOL linker moiety and a 3′-AOM capping group.
[0026] Figure 3 Shows the Illumina in the post-cleavage wash step with and without a palladium scavenger when using a fully functionalized nucleotide (ffN) with a 3′-AOM capping group Comparison of phasing values on the instrument.
[0027] Figure 4 Shows the main sequencing metrics (including phasing, pre-phasing, and error rate) on an Illumina instrument when using fully functionalized nucleotides (ffN) with a 3'-AOM capping group and an AOL linker moiety, in the presence of a palladium scavenger, and the same sequencing metrics for comparison when using standard ffN with a 3'-O-azidomethyl capping group. Shows the main sequencing metrics (including phasing and pre-phasing) on an Illumina instrument when using fully functionalized nucleotides (ffN) with a 3'-AOM capping group and an AOL linker moiety, in the presence of glycine or ethanolamine in the incorporation mixture, and the comparison of these metrics.
[0028] Figure 5 Shows the main sequencing metrics (including phasing, pre-phasing, and error rate) on an Illumina instrument when using fully functionalized nucleotides (ffN) with a 3'-AOM capping group and an AOL linker moiety, in the presence of glycine in the incorporation mixture, and the same sequencing metrics for comparison when using standard ffN with a 3'-O-azidomethyl capping group. Shows the comparison of the main sequencing metrics (including phasing and pre-phasing) on an Illumina instrument when using fully functionalized nucleotides (ffN) with a 3'-AOM capping group and an AOL linker moiety, in the presence of glycine in the incorporation mixture.
[0029] Figure 6 Shows the main sequencing metrics (including phasing, pre-phasing, and error rate) on an Illumina instrument when using fully functionalized nucleotides (ffN) with a 3'-AOM capping group and an AOL linker moiety, in the presence of glycine in the incorporation mixture, and the same sequencing metrics for comparison when using standard ffN with a 3'-O-azidomethyl capping group. Shows the main sequencing metrics (including phasing, pre-phasing, and error rate) on an Illumina instrument when using fully functionalized nucleotides (ffN) with a 3'-AOM capping group and an AOL linker moiety, in the presence of glycine in the incorporation mixture, and the same sequencing metrics for comparison when using standard ffN with a 3'-O-azidomethyl capping group.
[0030] Figure 7A and Figure 7B Respectively show the error rate and Q30 sequencing metrics for 2×300 sequencing runs on an Illumina iSeq instrument when using fully functionalized nucleotides (ffN) with a 3'-AOM capping group and an AOL linker moiety, and the same sequencing metrics for comparison when using standard ffN with a 3'-O-azidomethyl capping group. TM Respectively show the error rate and Q30 sequencing metrics for 2×300 sequencing runs on an Illumina iSeq instrument when using fully functionalized nucleotides (ffN) with a 3'-AOM capping group and an AOL linker moiety, and the same sequencing metrics for comparison when using standard ffN with a 3'-O-azidomethyl capping group.
[0031] Figure 8A and Figure 8B Respectively show the error rate and Q30 sequencing metrics for 2×150 sequencing runs on an Illumina iSeq instrument when using fully functionalized nucleotides (ffN) with a 3'-AOM capping group and an AOL linker moiety, and the same sequencing metrics for comparison when using standard ffN with a 3'-O-azidomethyl capping group. TM Respectively show the error rate and Q30 sequencing metrics for 2×150 sequencing runs on an Illumina iSeq instrument when using fully functionalized nucleotides (ffN) with a 3'-AOM capping group and an AOL linker moiety, and the same sequencing metrics for comparison when using standard ffN with a 3'-O-azidomethyl capping group.
[0032] Figures 9A to 9EShows the variation of the main sequencing metrics (phasing, percentage of signal decay, error rate) with the blue laser power when the green laser power is constant. These sequencing experiments were performed using fully functionalized nucleotides (ffN) with 3'-AOM capping groups and AOL linker moieties, in the presence of the palladium scavenger L-cysteine, on Illumina's NovaSeq TM instrument, and compared with the same sequencing metrics when using the standard protocol and ffN with 3'-O-azidomethyl capping groups.
[0033] Figure 10 Shows the main sequencing metrics (%PF, error rate, Q30, and signal decay) on Illumina's iSeq TM instrument (1×150 cycles) when using fully functionalized nucleotides (ffN) with 3'-AOM capping groups and AOL linker moieties, and also using the same palladium cleavage mixture in the first step of the chemical linearization of SBS, and compared with SBS of the same ffN when using standard enzymatic linearization. DETAILED DESCRIPTION
[0034] Embodiments of the present disclosure relate to nucleosides and nucleotides having 3′ acetal capping groups for sequencing applications (e.g., sequencing by synthesis (SBS)). In some embodiments, the nucleoside or nucleotide comprises a label covalently attached thereto via a cleavable linker, wherein the cleavable linker comprises an acetal moiety that allows cleavage of the 3’ acetal capping group and the label in a single reaction step. The 3′ acetal capping group provides improved stability during the synthesis of fully functionalized nucleotides (ffN), and also provides great stability during formulation, storage in solution, and operation on a sequencing instrument. In addition, the 3′ acetal capping groups described herein can also achieve low pre-phasing and lower signal decay, thereby improving data quality, which enables longer reads from sequencing applications.
[0035] Definitions
[0036] Unless otherwise defined, all technical and scientific terms used herein shall have the same meaning as commonly understood by one of ordinary skill in the art. The use of the term "comprising" and other forms (such as "include", "includes", "included") is not restrictive. The use of the term "having" and other forms (such as "have", "has", "had") is not restrictive. As used in this specification, whether in a transitional phrase or in the body of a claim, the terms "comprise(s)" and "comprising" shall be interpreted to have an open-ended meaning. That is, the above terms shall be interpreted synonymously with the phrase "at least having" or "at least including". For example, when used in the context of a process, the term "comprising" means that the process includes at least the recited steps, but may also include additional steps. When used in the context of a compound, composition or device, the term "comprising" means that the compound, composition or device includes at least the recited features or components, but may also include additional features or components.
[0037] Where ranges of values are provided, it is to be understood that the upper and lower limits of the range, and each intermediate value therebetween, are encompassed within these embodiments.
[0038] As used herein, common organic abbreviations are defined as follows:
[0039] ℃ Temperature value in degrees Celsius
[0040] dATP Deoxyadenosine triphosphate
[0041] dCTP Deoxycytidine triphosphate
[0042] dGTP Deoxyguanosine triphosphate
[0043] dTTP Deoxythymidine triphosphate
[0044] ddNTP Dideoxynucleotide triphosphate
[0045] ffN Fully functionalized nucleotide
[0046] RT Room temperature
[0047] SBS Sequencing by synthesis
[0048] SM Starting material
[0049] As used herein, the term "array" refers to a population of different probe molecules that are attached to one or more substrates such that the different probe molecules can be distinguished from one another based on their relative positions. An array can include different probe molecules each located at a different addressable position on a single substrate. Alternatively, or in addition, an array can include separate substrates each bearing different probe molecules, where the different probe molecules can be identified based on the position of the substrate on the surface to which the substrate is attached or based on the position of the substrate in a liquid. Exemplary arrays in which the separate substrates are located on a surface include, but are not limited to, those arrays that include beads in wells, as described, for example, in U.S. Patent No. 6,355,431B1, US2002 / 0102578, and PCT Publication No. WO 00 / 63437. For example, exemplary formats that can be used in the present invention to distinguish beads in a liquid array using a microfluidic device such as a fluorescence-activated cell sorter (FACS) are described in U.S. Patent No. 6,524,793. Additional examples of arrays that can be used in the present invention include, but are not limited to, those described in the following patents: U.S. Patent Nos. 5,429,807, 5,436,327, 5,561,071, 5,583,211, 5,658,734, 5,837,858, 5,874,219, 5,919,523, 6,136,269, 6,287,768, 6,287,776, 6,288,220, 6,297,006, 6,291,193, 6,346,413, 6,416,949, 6,482,591, 6,514,751, and 6,610,482, WO 93 / 17126, WO 95 / 11995, WO 95 / 35505, EP 742 287, and EP 799 897.
[0050] As used herein, the term "covalently linked" or "covalently bonded" refers to the formation of a chemical bond characterized by the sharing of electron pairs between atoms. For example, a covalently linked polymer coating refers to a polymer coating that forms a chemical bond with a functionalized surface of a substrate, as compared to adhering to the surface via other means (e.g., adhesion or electrostatic interactions). It should be understood that a polymer covalently linked to a surface can also be bonded via means other than covalent linkage.
[0051] As used herein, any "R" group represents a substituent that can be attached to a specified atom. The R group can be substituted or unsubstituted. If two "R" groups are described as "together with the atoms to which they are attached" forming a ring or ring system, it means that the collection of these atoms, the intervening bonds, and the two R groups is the ring that is mentioned. For example, when the following substructure is present:
[0052]
[0053] and R 1 and R 2 are limited to be selected from the group consisting of hydrogen and alkyl, or R 1 and R 2 together with the atoms to which they are attached form an aryl or carbocyclic group, which means that R 1 and R 2 can be selected from hydrogen or alkyl, or alternatively, the substructure has the structure:
[0054]
[0055] wherein A is an aromatic or carbocyclic group containing the depicted double bond.
[0056] It should be understood that, depending on the context, certain group naming conventions may include mono-groups or di-groups. For example, when a substituent requires two connection points to the rest of the molecule, it should be understood that the substituent is a di-group. For example, substituents of alkyl groups identified as requiring two connection points include di-groups such as –CH2–, –CH2CH2–, –CH2CH(CH3)CH2–, etc. Other group naming conventions clearly indicate that the group is a di-group, such as “alkylene” or “alkenylene”.
[0057] As used herein, the term “halogen” or “halo-group” means any of the radioactively stable atoms in column 7 of the periodic table of the elements, for example, fluorine, chlorine, bromine or iodine, where fluorine and chlorine are preferred.
[0058] As used herein, “C a to C b"refers to the number of carbon atoms in an alkyl, alkenyl or alkynyl group, or the number of ring atoms in a cycloalkyl or aryl group. That is, the rings of alkyl, alkenyl, alkynyl, cycloalkyl and aryl groups can contain from "a" to "b" (including the end values) carbon atoms. For example, a "C1-C4 alkyl" group refers to all alkyl groups having 1 to 4 carbons, namely CH3-, CH3CH2-, CH3CH2CH2-, (CH3)2CH-, CH3CH2CH2CH2-, CH3CH2CH(CH3)- and (CH3)3C-; a C3-C4 cycloalkyl group refers to all cycloalkyl groups having 3 to 4 carbon atoms, namely cyclopropyl and cyclobutyl. Similarly, a "4- to 6-membered heterocyclic group" refers to all heterocyclic groups having 4 to 6 total ring atoms, such as azetidine, oxetane, oxazoline, pyrrolidine, piperidine, piperazine, morpholine, etc. If "a" and "b" are not specified for an alkyl, alkenyl, alkynyl, cycloalkyl or aryl group, the broadest ranges described in these definitions will be assumed. As used herein, the term "C1-C6" includes C1, C2, C3, C4, C5 and C6, and the ranges defined by any two of these numbers. For example, C1-C6 alkyl includes C1 alkyl, C2 alkyl, C3 alkyl, C4 alkyl, C5 alkyl and C6 alkyl, C2-C6 alkyl, C1-C3 alkyl, etc. Similarly, C2-C6 alkenyl includes C2 alkenyl, C3 alkenyl, C4 alkenyl, C5 alkenyl and C6 alkenyl, C2-C5 alkenyl, C3-C4 alkenyl, etc.; and C2-C6 alkynyl includes C2 alkynyl, C3 alkynyl, C4 alkynyl, C5 alkynyl and C6 alkynyl, C2-C5 alkynyl, C3-C4 alkynyl, etc. C3-C8 cycloalkyl each includes hydrocarbon rings containing 3, 4, 5, 6, 7 and 8 carbon atoms or ranges defined by any two numbers, such as C3-C7 cycloalkyl or C5-C6 cycloalkyl.
[0059] As used herein, "alkyl" refers to a straight-chain or branched-chain hydrocarbon chain that is completely saturated (i.e., does not contain double or triple bonds). An alkyl group can have 1 to 20 carbon atoms (whenever a numerical range such as "1 to 20" appears herein, it refers to each integer within the given range; for example, "1 to 20 carbon atoms" means that an alkyl group can be composed of 1 carbon atom, 2 carbon atoms, 3 carbon atoms, etc., up to and including 20 carbon atoms, but this definition also encompasses the term "alkyl" where no numerical range is specified). An alkyl group can also be a medium-sized alkyl group having 1 to 9 carbon atoms. An alkyl group can also be a lower alkyl group having 1 to 6 carbon atoms. An alkyl group can be named as "C 1- C4 alkyl" or a similar name. By way of example only, "C 1-"C6 alkyl" means that there are one to six carbon atoms in the alkyl chain, i.e., the alkyl chain is selected from the group consisting of methyl, ethyl, propyl, isopropyl, n-butyl, isobutyl, sec-butyl, and tert-butyl. Typical alkyl groups include but are not limited to methyl, ethyl, propyl, isopropyl, butyl, isobutyl, tert-butyl, pentyl, hexyl, etc.
[0060] As used herein, "alkoxy" refers to the formula –OR (where R is an alkyl as defined above), such as "C 1- C9 alkoxy", including but not limited to methoxy, ethoxy, n-propoxy, 1-methylethoxy (isopropoxy), n-butoxy, isobutoxy, sec-butoxy, and tert-butoxy, etc.
[0061] As used herein, "alkenyl" refers to a straight-chain or branched-chain hydrocarbon chain containing one or more double bonds. The alkenyl group can have 2 to 20 carbon atoms, but this definition also encompasses the term "alkenyl" when no numerical range is specified therein. The alkenyl group can also be a medium-sized alkenyl having 2 to 9 carbon atoms. The alkenyl group can also be a lower alkenyl having 2 to 6 carbon atoms. The alkenyl group can be named as "C 2- C6 alkenyl" or a similar name. By way of example only, "C 2- C6 alkenyl" means that there are two to six carbon atoms in the alkenyl chain, i.e., the alkenyl chain is selected from the group consisting of vinyl, prop-1-en-1-yl, prop-2-en-1-yl, prop-3-en-1-yl, but-1-en-1-yl, but-2-en-1-yl, but-3-en-1-yl, but-4-en-1-yl, 1-methylprop-1-en-1-yl, 2-methylprop-1-en-1-yl, 1-ethylvinyl-1-yl, 2-methylprop-3-en-1-yl, buta-1,3-dienyl, buta-1,2-dienyl, and buta-1,2-dien-4-yl. Typical alkenyl groups include but are not limited to vinyl, propenyl, butenyl, pentenyl, and hexenyl, etc.
[0062] As used herein, "alkynyl" refers to a straight-chain or branched-chain hydrocarbon chain containing one or more triple bonds. The alkynyl group can have 2 to 20 carbon atoms, but this definition also encompasses the term "alkynyl" when no numerical range is specified therein. The alkynyl group can also be a medium-sized alkynyl having 2 to 9 carbon atoms. The alkynyl group can also be a lower alkynyl having 2 to 6 carbon atoms. The alkynyl group can be named as "C 2- C6 alkynyl" or a similar name. By way of example only, "C 2- C6 alkynyl" means that there are two to six carbon atoms in the alkynyl chain, i.e., the alkynyl chain is selected from the group consisting of ethynyl, prop-1-yn-1-yl, prop-2-yn-1-yl, but-1-yn-1-yl, but-3-yn-1-yl, but-4-yn-1-yl, and 2-butynyl. Typical alkynyl groups include but are not limited to ethynyl, propynyl, butynyl, pentynyl, and hexynyl, etc.
[0063] As used herein, "heteroalkyl" refers to a straight-chain or branched hydrocarbon chain containing one or more heteroatoms (i.e., elements other than carbon, including but not limited to nitrogen, oxygen, and sulfur) in the chain backbone. A heteroalkyl group may have 1 to 20 carbon atoms, but this definition also encompasses the term "heteroalkyl" when no numerical range is specified therein. A heteroalkyl group may also be a medium-sized heteroalkyl having 1 to 9 carbon atoms. A heteroalkyl group may also be a lower heteroalkyl having 1 to 6 carbon atoms. A heteroalkyl group may be named "C 1- C6 heteroalkyl" or a similar name. A heteroalkyl group may contain one or more heteroatoms. By way of example only, "C 4- C6 heteroalkyl" means that there are four to six carbon atoms in the heteroalkyl chain, and in addition there is one or more heteroatoms in the backbone of the chain.
[0064] The term "aromatic" refers to a ring or ring system having a conjugated π-electron system, and includes both carbocyclic aromatic groups (e.g., phenyl) and heterocyclic aromatic groups (e.g., pyridine). The term includes monocyclic groups or fused polycyclic (i.e., rings sharing adjacent atom pairs) groups, provided that the entire ring system is aromatic.
[0065] As used herein, "aryl" refers to an aromatic ring or ring system containing only carbon in the ring backbone (i.e., two or more fused rings sharing two adjacent carbon atoms). When the aryl is a ring system, each ring in the ring system is aromatic. An aryl group may have 6 to 18 carbon atoms, but this definition also encompasses the term "aryl" when no numerical range is specified therein. In some embodiments, the aryl group has 6 to 10 carbon atoms. An aryl group may be named "C 6- C 10 aryl", "C6 or C 10 aryl" or a similar name. Examples of aryl groups include but are not limited to phenyl, naphthyl, azulyl, and anthracenyl.
[0066] "Aralkyl" or "arylalkyl" is an aryl group attached as a substituent via an alkylene group, such as "C 7-14 aralkyl", etc., including but not limited to benzyl, 2-phenylethyl, 3-phenylpropyl, and naphthylalkyl. In some cases, the alkylene group is a lower alkylene group (i.e., a C 1- C6 alkylene group).
[0067] As used herein, "heteroaryl" refers to an aromatic ring or ring system (i.e., two or more fused rings sharing two adjacent atoms) containing one or more heteroatoms (i.e., elements other than carbon, including but not limited to nitrogen, oxygen, and sulfur) in the ring skeleton. When heteroaryl is a ring system, each ring in the ring system is aromatic. A heteroaryl group can have 5 to 18 ring members (i.e., the number of atoms (including carbon atoms and heteroatoms) constituting the ring skeleton), but this definition also encompasses the term "heteroaryl" when no numerical range is specified therein. In some embodiments, the heteroaryl group has 5 to 10 ring members or 5 to 7 ring members. A heteroaryl group can be named "5- to 7-membered heteroaryl", "5- to 10-membered heteroaryl", or a similar name. Examples of heteroaryl rings include but are not limited to furyl, thienyl, phthalazinyl, pyrrolyl, oxazolyl, thiazolyl, imidazolyl, pyrazolyl, isoxazolyl, isothiazolyl, triazolyl, thiadiazolyl, pyridyl, pyridazinyl, pyrimidinyl, pyrazinyl, triazinyl, quinolinyl, isoquinolinyl, benzimidazolyl, benzoxazolyl, benzothiazolyl, indolyl, isoindolyl, and benzothienyl.
[0068] "Heteroaralkyl" or "heteroarylalkyl" is a heteroaryl group attached as a substituent via an alkylene group. Examples include but are not limited to 2-thienylmethyl, 3-thienylmethyl, furylmethyl, thienylethyl, pyrrolylalkyl, pyridylalkyl, isoxazolylalkyl, and imidazolylalkyl. In some cases, the alkylene group is a lower alkylene group (i.e., C 1- C6 alkylene group).
[0069] As used herein, "carbocyclic" means a non-aromatic cyclic ring or ring system containing only carbon atoms in the ring system skeleton. When carbocyclic is a ring system, two or more rings can be joined together in a fused, bridged, or spiro manner. A carbocyclic group can have any degree of saturation, provided that at least one ring in the ring system is not aromatic. Thus, carbocyclic includes cycloalkyl, cycloalkenyl, and cycloalkynyl. A carbocyclic group can have 3 to 20 carbon atoms, but this definition also encompasses the term "carbocyclic" when no numerical range is specified therein. A carbocyclic group can also be a medium-sized carbocyclic group having 3 to 10 carbon atoms. A carbocyclic group can also be a carbocyclic group having 3 to 6 carbon atoms. A carbocyclic group can be named "C 3- C6 carbocyclic" or a similar name. Examples of carbocyclic rings include but are not limited to cyclopropyl, cyclobutyl, cyclopentyl, cyclohexyl, cyclohexenyl, 2,3-dihydro-indene, bicyclo[2.2.2]octyl, adamantyl, and spiro[4.4]nonyl.
[0070] As used herein, "cycloalkyl" means a fully saturated carbocyclic ring or ring system. Examples include cyclopropyl, cyclobutyl, cyclopentyl, and cyclohexyl.
[0071] As used herein, "heterocyclic group" refers to a non-aromatic cyclic ring or ring system containing at least one heteroatom in the ring skeleton. Heterocyclic groups can be joined together in a fused, bridged, or spiro fashion. Heterocyclic groups can have any degree of saturation provided that at least one ring in the ring system is not aromatic. Heteroatoms can be present in non-aromatic or aromatic rings in the ring system. Heterocyclic groups can have from 3 to 20 ring members (i.e., the number of atoms (including carbon atoms and heteroatoms) that make up the ring skeleton), but this definition also encompasses the term "heterocyclic group" when no numerical range is specified therein. Heterocyclic groups can also be medium-sized heterocyclic groups having from 3 to 10 ring members. Heterocyclic groups can also be heterocyclic groups having from 3 to 6 ring members. Heterocyclic groups can be named "3- to 6-membered heterocyclic group" or similar names. In a preferred six-membered monocyclic heterocyclic group, the heteroatoms are selected from one to up to three of O, N, or S, and in a preferred five-membered monocyclic heterocyclic group, the heteroatoms are selected from one or two heteroatoms selected from O, N, or S. Examples of heterocyclic rings include, but are not limited to, azepinyl, acridinyl, carbazolyl, cinnolinyl, dioxolanyl, imidazolinyl, imidazolidinyl, morpholinyl, oxiranyl, oxepanyl, thiepanyl, piperidinyl, piperazinyl, dioxopiperazinyl, pyrrolidinyl, pyrrolidinyl, pyrrolidinedionyl, 4-piperidinonyl, pyrazolinyl, pyrazolidinyl, 1,3-dioxinyl, 1,3-dioxanyl, 1,4-dioxinyl, 1,4-dioxanyl, 1,3-oxathiolanyl, 1,4-oxathiolene, 1,4-oxathiolanyl, 2H-1,2-oxazinyl, trioxanyl, hexahydro-1,3,5-triazinyl, 1,3-meta-dioxolene, 1,3-dioxolanyl, 1,3-dithienyl, 1,3-dithianyl, isoxazolinyl, isoxazolidinyl, oxazolinyl, oxazolidinyl, oxazolidinone, thiazolinyl, thiazolidinyl, 1,3-oxathiolanyl, dihydroindolyl, isoindolinyl, tetrahydrofuranyl, tetrahydropyranyl, tetrahydrothienyl, tetrahydrothiopyranyl, tetrahydro-1,4-thiazinyl, thiomorpholinyl, dihydrobenzofuranyl, benzimidazolidinyl, and tetrahydroquinoline.
[0072] As used herein, "alkoxyalkyl" or "(alkoxy)alkyl" refers to an alkoxy group linked via an alkylene group, such as C 2- C8 alkoxyalkyl, or (C1-C6 alkoxy)C1-C6 alkyl, e.g., –(CH2) 1-3 -OCH3.
[0073] As used herein, "-O-alkoxyalkyl" or "-O-(alkoxy)alkyl" refers to an alkoxy group linked via a –O-(alkylene) group, such as –O-(C1-C6 alkoxy)C1-C6 alkyl, e.g., –O-(CH2) 1-3 -OCH3.
[0074] As used herein, “(heterocyclic group)alkyl” refers to a heterocyclic or heterocyclic group as defined above attached as a substituent via an alkylene group as defined above. The alkylene group and heterocyclic group of (heterocyclic group)alkyl may be substituted or unsubstituted. Examples include, but are not limited to, (tetrahydro-2H-pyran-4-yl)methyl, (piperidin-4-yl)ethyl, (piperidin-4-yl)propyl, (tetrahydro-2H-thiopyran-4-yl)methyl, and (1,3-thiazinan-4-yl)methyl.
[0075] As used herein, “(cycloalkyl)alkyl” or “(carbocyclic group)alkyl” refers to a cycloalkyl or carbocyclic group (as defined herein) attached as a substituent via an alkylene group. Examples include, but are not limited to, cyclopropylmethyl, cyclobutylmethyl, cyclopentylethyl, and cyclohexylpropyl.
[0076] The “O-carboxy” group refers to an “-OC(=O)R” group in which R is selected from hydrogen, C 1- C6 alkyl, C 2- C6 alkenyl, C 2- C6 alkynyl, C 3- C7 carbocyclic group, C 6- C 10 aryl, 5- to 10-membered heteroaryl, and 3- to 10-membered heterocyclic group, as defined herein.
[0077] The “C-carboxy” group refers to a “-C(=O)OR” group in which R is selected from the group consisting of hydrogen, C 1- C6 alkyl, C 2- C6 alkenyl, C 2- C6 alkynyl, C 3- C7 carbocyclic group, C 6- C 10 aryl, 5- to 10-membered heteroaryl, and 3- to 10-membered heterocyclic group, as defined herein. Non-limiting examples include carboxy (i.e., -C(=O)OH).
[0078] The “sulfonyl” group refers to a “-SO2R” group in which R is selected from hydrogen, C 1- C6 alkyl, C 2- C6 alkenyl, C 2- C6 alkynyl, C 3- C7 carbocyclic group, C 6- C 10 aryl, 5- to 10-membered heteroaryl, and 3- to 10-membered heterocyclic group, as defined herein.
[0079] The “sulfinic acid” group refers to the “-S(=O)OH” group.
[0080] The “S-sulfinylamino” group refers to one in which R A and RB each independently selected from hydrogen, C 1- C6 alkyl, C 2- C6 alkenyl, C 2- C6 alkynyl, C 3- C7 carbocyclic group, C 6- C 10 aryl, 5- to 10-membered heteroaryl, and 3- to 10-membered heterocyclic group of “-SO2NR A R B ”, as defined herein.
[0081] The “N-sulfinylamino” group means a group of “-N(R A and R B each independently selected from hydrogen, C 1- C6 alkyl, C 2- C6 alkenyl, C 2- C6 alkynyl, C 3- C7 carbocyclic group, C 6- C 10 aryl, 5- to 10-membered heteroaryl, and 3- to 10-membered heterocyclic group of “-N(R A )SO2R B ”, as defined herein.
[0082] The “C-acylamino” group means a group of “-C(═O)NR A and R B each independently selected from hydrogen, C 1- C6 alkyl, C 2- C6 alkenyl, C 2- C6 alkynyl, C 3- C7 carbocyclic group, C 6- C 10 aryl, 5- to 10-membered heteroaryl, and 3- to 10-membered heterocyclic group of “-C(═O)NR A R B ”, as defined herein.
[0083] The “N-acylamino” group means a group of “-N(R A and R B each independently selected from hydrogen, C 1- C6 alkyl, C 2- C6 alkenyl, C 2- C6 alkynyl, C 3- C7 carbocyclic group, C 6- C 10 aryl, 5- to 10-membered heteroaryl, and 3- to 10-membered heterocyclic group of “-N(R A )C(═O)R B ”, as defined herein.
[0084] The “amino” group means a group in which R A and RB Each independently selected from hydrogen, C 1- C6 alkyl, C 2- C6 alkenyl, C 2- C6 alkynyl, C 3- C7 carbocyclic group, C 6- C 10 aryl, 5- to 10-membered heteroaryl, and 3- to 10-membered heterocyclic group of the “-NR A R B ” group, as defined herein. Non-limiting examples include the free amino group (i.e., -NH2).
[0085] The “aminoalkyl” group refers to an amino group linked via an alkylene group.
[0086] The “(alkoxy)alkyl” group refers to an alkoxy group linked via an alkylene group, such as “(C 1- C6 alkoxy)C 1- C6 alkyl” and the like.
[0087] As used herein, the term “hydroxy” refers to the –OH group.
[0088] As used herein, the term “cyano” refers to the “-CN” group.
[0089] As used herein, the term “azido” refers to the –N3 group.
[0090] As used herein, the term “propargylamine” refers to an amino group substituted with a propargyl group (HC≡C-CH2-). When propargylamine is used as a divalent moiety in the context, it includes -C≡C-CH2-NR A -, where R A is hydrogen, C1-C6 alkyl, C2-C6 alkenyl, C2-C6 alkynyl, C 3- C7 carbocyclic group, C 6- C 10 aryl, 5- to 10-membered heteroaryl, and 3- to 10-membered heterocyclic group, as defined herein.
[0091] As used herein, the term “propargylamide” refers to a C-acylamino or N-acylamino group substituted with a propargyl group (HC≡C-CH2-). When propargylamide is used as a divalent moiety in the context, it includes -C≡C-CH2-NR A -C(=O)- or -C≡C-CH2-C(=O)-NR A -, where R A is hydrogen, C 1- C6 alkyl, C 2- C6 alkenyl, C 2- C6 alkynyl, C 3- C7 carbocyclic group, C6- C 10 Aryl, 5- to 10-membered heteroaryl, and 3- to 10-membered heterocyclic group, as defined herein.
[0092] As used herein, the term "allylamine" refers to an amino group substituted with an allyl group (CH2=CH-CH2-). When allylamine is used as a divalent moiety in context, it includes -CH=CH-CH2-NR A -, where R A is hydrogen, C 1- C6 alkyl, C 2- C6 alkenyl, C 2- C6 alkynyl, C 3- C7 carbocyclic group, C 6- C 10 aryl, 5- to 10-membered heteroaryl, and 3- to 10-membered heterocyclic group, as defined herein.
[0093] As used herein, the term "allylamide" refers to a C-acylamino or N-acylamino group substituted with an allyl group (CH2=CH-CH2-). When allylamide is used as a divalent moiety in context, it includes -CH=CH-CH2-NR A -C(=O)- or -CH=CH-CH2-C(=O)-NR A -, where R A is hydrogen, C 1- C6 alkyl, C 2- C6 alkenyl, C 2- C6 alkynyl, C 3- C7 carbocyclic group, C 6- C 10 aryl, 5- to 10-membered heteroaryl, and 3- to 10-membered heterocyclic group, as defined herein.
[0094] When a group is described as "optionally substituted", it may be unsubstituted or substituted. Similarly, when a group is described as "substituted", the substituent(s) may be selected from one or more of the indicated substituents. As used herein, a substituted group is derived from an unsubstituted parent group in which one or more hydrogen atoms have been exchanged with another atom or group. Unless otherwise specified, when a group is considered to be "substituted", this means that the group is substituted with one or more substituents independently selected from: C1-C6 alkyl, C1-C6 alkenyl, C1-C6 alkynyl, C1-C6 heteroalkyl, C3-C7 carbocyclic group (optionally substituted with halo, C1-C6 alkyl, C1-C6 alkoxy, C1-C6 haloalkyl, and C1-C6 haloalkoxy), C3-C7 carbocyclic-C1-C6-alkyl (optionally substituted with halo, C1-C6 alkyl, C1-C6 alkoxy, C1-C6 haloalkyl, and C1-C6 haloalkoxy), 3- to 10-membered heterocyclic group (optionally substituted with halo, C1-C6 alkyl, C1-C6 alkoxy, C1-C6 haloalkyl, and C1-C6 haloalkoxy), 3- to 10-membered heterocyclic-C1-C6-alkyl (optionally substituted with halo, C1-C6 alkyl, C1-C6 alkoxy, C1-C6 haloalkyl, and C1-C6 haloalkoxy), aryl (optionally substituted with halo, C1-C6 alkyl, C1-C6 alkoxy, C1-C6 haloalkyl, and C1-C6 haloalkoxy), (aryl)C1-C6 alkyl (optionally substituted with halo, C1-C6 alkyl, C1-C6 alkoxy, C1-C6 haloalkyl, and C1-C6 haloalkoxy), (5- to 10-membered heteroaryl)C1-C6 alkyl (optionally substituted with halo, C1-C6 alkyl, C1-C6 alkoxy, C1-C6 haloalkyl, and C1-C6 haloalkoxy), 5- to 10-membered heteroaryl(C1-C6)alkyl (optionally substituted with halo, C1-C6 alkyl, C1-C6 alkoxy, C1-C6 haloalkyl, and C1-C6 haloalkoxy), halo, -CN, hydroxy, C1-C6 alkoxy, (C1-C6 alkoxy)C1-C6 alkyl, -O(C1-C6 alkoxy)C1-C6 alkyl; (C1-C6 haloalkoxy)C1-C6 alkyl; -O(C1-C6 haloalkoxy)C1-C6 alkyl; aryloxy, mercapto (hydrogen sulfide group), halo(C1-C6)alkyl (e.g.,
[0095] –CF3), halo(C1-C6)alkoxy (e.g., –OCF3), C1-C6 alkylthio, arylthio, amino, amino(C1-C6)alkyl, nitro, O-carbamoyl, N-carbamoyl, O-thiocarbamoyl, N-thiocarbamoyl, C-acylamino, N-acylamino, S-sulfinylamino, N-sulfinylamino, C-carboxy, O-carboxy, acyl, cyanate, isocyanate, thiocyanate, isothiocyanate, sulfinyl, sulfonyl, -SO3H, sulfino, -OSO2C 1-4 alkyl, monophosphate, diphosphate, triphosphate, and oxo(=O). Wherever a group is described as "optionally substituted", the group may be substituted by the above substituents.
[0096] When a substituent is described as a diyl (i.e., having two attachment points to the rest of the molecule), it is understood that, unless otherwise specified, the substituent may be attached in any orientational configuration. Thus, for example, a substituent depicted as -AE- or includes substituents oriented such that A is attached at the leftmost attachment point of the molecule, and those in which A is attached at the rightmost attachment point of the molecule. Further, if a group or substituent is depicted as and L is defined as an optionally present linker moiety; then when L is absent (or not present), such a group or substituent is equivalent to
[0097] As used herein, "nucleotide" includes a nitrogenous heterocyclic base, a sugar, and one or more phosphate groups. They are the monomeric units of nucleic acid sequences. In RNA, the sugar is ribose, and in DNA it is deoxyribose, i.e., a sugar lacking the hydroxyl group present in ribose. The nitrogenous heterocyclic base may be a purine or pyrimidine base. Purine bases include adenine (A) and guanine (G) and their modified derivatives or analogs. Pyrimidine bases include cytosine (C), thymine (T), and uracil (U) and their modified derivatives or analogs. The C-1 atom of deoxyribose is bonded to the N-1 of pyrimidine or the N-9 of purine.
[0098] As used herein, "nucleoside" is structurally similar to a nucleotide but lacks the phosphate moiety. An example of a nucleoside analog is a nucleoside analog in which a label is attached to the base and no phosphate group is attached to the sugar molecule. The term "nucleoside" is used herein in its conventional meaning as understood by those skilled in the art. Examples include, but are not limited to, ribonucleosides containing a ribose moiety and deoxyribonucleosides containing a deoxyribose moiety. A modified pentose moiety is a pentose moiety in which an oxygen atom has been replaced by a carbon and / or a carbon has been replaced by a sulfur or oxygen atom. A "nucleoside" is a monomer that may have a substituted base and / or sugar moiety. Additionally, nucleosides can be incorporated into larger DNA and / or RNA polymers and oligomers.
[0099] The term "purine base" is used herein in its ordinary meaning as understood by those skilled in the art and includes its tautomers. Similarly, the term "pyrimidine base" is used herein in its ordinary meaning as understood by those skilled in the art and includes its tautomers. Non-limiting lists of optionally substituted purine bases include purine, deazapurine, adenine, 7-deazaadenine, guanine, 7-deazaguanine, hypoxanthine, xanthine, alloxanthine, 7-alkylguanine (e.g., 7-methylguanine), theobromine, caffeine, uric acid, and isoguanine. Examples of pyrimidine bases include, but are not limited to, cytosine, thymine, uracil, 5,6-dihydrouracil, and 5-alkylcytosine (e.g., 5-methylcytosine).
[0100] As used herein, when an oligonucleotide or polynucleotide is described as "comprising" a nucleoside or nucleotide described herein or labeled with a nucleoside or nucleotide described herein, this means that the nucleoside or nucleotide described herein forms a covalent bond with the oligonucleotide or polynucleotide. Similarly, when a nucleoside or nucleotide is described as part of an oligonucleotide or polynucleotide, such as "incorporated into" an oligonucleotide or polynucleotide, this means that the nucleoside or nucleotide described herein forms a covalent bond with the oligonucleotide or polynucleotide. In some such embodiments, the covalent bond is formed between the 3'-hydroxyl of the oligonucleotide or polynucleotide and the 5'-phosphate group of the nucleotide of the phosphodiester bond between the 3'-carbon of the oligonucleotide or polynucleotide and the 5'-carbon of the nucleotide.
[0101] As used herein, the term "cleavable linker" is not intended to imply that the entire linker needs to be removed. The cleavage site can be located on the linker at a position that ensures that a portion of the linker remains attached to the detectable label and / or nucleoside or nucleotide moiety after cleavage.
[0102] As used herein, "derivative" or "analog" means a synthetic nucleotide or nucleoside derivative having a modified base moiety and / or a modified sugar moiety. Such derivatives and analogs are discussed, for example, in Scheit, Nucleotide Analogs (John Wiley & Son, 1980) and Uhlman et al., Chemical Reviews 90:543-584, 1990. Nucleotide analogs can also include modified phosphodiester bonds, including phosphorothioate bonds, dithiophosphonate bonds, alkylphosphonate bonds, anilinophosphonate bonds, and aminophosphonate bonds. As used herein, "derivative", "analog", and "modified" can be used interchangeably and are encompassed by the terms "nucleotide" and "nucleoside" as defined herein.
[0103] As used herein, the term "phosphate ester" is used in its ordinary meaning as understood by those skilled in the art, and includes its protonated forms (e.g., ). As used herein, the terms "monophosphate", "diphosphate", and "triphosphate" are used in their ordinary meanings as understood by those skilled in the art, and include protonated forms.
[0104] As used herein, the term "protecting group / protecting groups" refers to any atom or group of atoms added to a molecule to prevent existing groups in the molecule from undergoing undesired chemical reactions. Sometimes, "protecting group" and "blocking group" may be used interchangeably.
[0105] As used herein, the term "phasing" refers to a phenomenon in SBS that is caused by incomplete removal of 3'-terminators and fluorophores and failure of the polymerase to complete incorporation of a portion of the DNA strands within a cluster in a given sequencing cycle. Premature phasing is caused by incorporation of a nucleotide that does not have a functional 3'-terminator, where the incorporation event occurs 1 cycle earlier due to termination failure. Phasing and premature phasing result in the measured signal intensity for a particular cycle being composed of the signal from the current cycle and noise from the previous and next cycles. As the number of cycles increases, the sequence fraction of each cluster affected by phasing and premature phasing increases, hindering identification of the correct base. Premature phasing can be caused by the presence of trace amounts of unprotected or unblocked 3'-OH nucleotides during sequencing by synthesis (SBS). Unprotected 3'-OH nucleotides may be generated during the manufacturing process or during storage and reagent handling. Accordingly, the discovery of nucleotide analogs that reduce the incidence of premature phasing is surprising and provides a significant advantage over existing nucleotide analogs in SBS applications. For example, the provided nucleotide analogs can result in shorter SBS cycle times, reduced phasing and premature phasing values, and longer sequencing read lengths.
[0106] Nucleosides or nucleotides with 3′ acetal capping groups
[0107] Some embodiments of the present disclosure relate to a nucleotide or nucleoside molecule comprising a nucleobase attached via a cleavable linker to a detectable label, and a ribose or deoxyribose moiety, wherein the cleavable linker comprises a moiety having the following structure:
[0108]
[0109] wherein X and Y are each independently O or S; and R 1a 、R 1b 、R 2 、R 3a and R 3bEach independently is H, a halogen, an unsubstituted or substituted C1-C6 alkyl group, or a C1-C6 haloalkyl group. In some embodiments, the ribose or deoxyribose moiety comprises a 3'-OH protecting group as described herein. In some embodiments, the cleavable linker may also include L 1 or L 2 , or both, where L 1 and L 2 are described in detail below.
[0110] In some embodiments, the nucleoside or nucleotide described herein comprises or has the structure of formula (I):
[0111]
[0112] where B is a nucleobase;
[0113] R 4 is H or OH;
[0114] R 5 is H, a 3'-OH capping group, or a phosphoramidite;
[0115] R 6 is H, a monophosphate, diphosphate, triphosphate, thiophosphate, phosphate ester analogue, a reactive phosphorus-containing group, or a hydroxy protecting group;
[0116] L is and
[0117] L 1 and L 2 are each independently an optionally present linker moiety.
[0118] In some embodiments of the cleavable linker moieties described herein, X and Y are each O. In some other embodiments, X is S and Y is O, or X is O and Y is S. In some embodiments, R 1a 、R 1b 、R 2 、R 3a and R 3b are each H. In other embodiments, at least one of R 1a 、R 1b 、R 2 、R 3a and R 3b is a halogen (e.g., fluorine, chlorine) or an unsubstituted C1-C6 alkyl group (e.g., methyl, ethyl, isopropyl, isobutyl, or tert-butyl). In some such cases, R 1a and R 1b are each H, and R 2 、R 3a and R 3bat least one of which is an unsubstituted C1-C6 alkyl or a halogen (e.g., R 2 is an unsubstituted C1-C6 alkyl, and R 3a and R 3b are each H; or R 2 is H, and one or both of R 3a and R 3b are a halogen or an unsubstituted C1-C6 alkyl). In one embodiment, the cleavable linker or L comprises ("AOL" linker moiety).
[0119] In some embodiments of the nucleosides or nucleotides described herein, the nucleobase ("B" in formula (I)) is a purine (adenine or guanine), a deazapurine, or a pyrimidine (e.g., cytosine, thymine, or uracil). In some further embodiments, the deazapurine is a 7-deazapurine (e.g., 7-deazaadenine or 7-deazaguanine). Non-limiting examples of B include or their optionally substituted derivatives and analogs. In some additional embodiments, the labeled nucleobase comprises the structure
[0120] In some embodiments of the nucleosides or nucleotides described herein, the ribose or deoxyribose moiety comprises a 3'-OH capping group (i.e., R 5 in formula (I) is a 3'-OH capping group). In some embodiments, the 3'-OH capping group or R 5 is and wherein R a , R b , R c , R d and R e are each independently H, a halogen, an unsubstituted or substituted C1-C6 alkyl, or a C1-C6 haloalkyl. In some further embodiments, R a and R b are each H, and at least one of R c , R d and R e is independently a halogen (e.g., fluorine, chlorine) or an unsubstituted C1-C6 alkyl (e.g., methyl, ethyl, isopropyl, isobutyl, or tert-butyl). For example, R c is an unsubstituted C1-C6 alkyl, and R d and R e are each H. In another example, R c is H, and one or both of R d and R e are a halogen or an unsubstituted C1-C6 alkyl. R5 Other non-limiting embodiments include
[0121] In one embodiment, R 5 is which together with the 3'-oxy forms ("AOM") group, and the AOM group is attached to the 3'-carbon atom of the ribose or deoxyribose moiety. In other embodiments, the 3'-OH capping group or R 5 may include an azido moiety (e.g., -CH2N3 or azidomethyl). Additional embodiments of the 3'-OH capping group are described in U.S. Patent Publication No. 2020 / 0216891A1, which is incorporated herein by reference in its entirety and includes additional examples of 3'-acetal capping groups such as These capping groups are attached to the 3'-carbon atom of the ribose or deoxyribose moiety.
[0122] In some other embodiments of the nucleosides or nucleotides described herein, R in formula (I) 5 is phosphoramidite. In such embodiments, R 6 is an acid-labile hydroxyl protecting group that allows subsequent monomer coupling under automated synthesis conditions.
[0123] In some embodiments of the nucleosides or nucleotides described herein, L 1 is present, and L 1 includes moieties selected from the group consisting of propargylamine, propargylamide, allylamine, allylamide, and optionally substituted variants thereof. In still other embodiments, L 1 includes or is
[0124] In still other embodiments, the asterisk * indicates the point of attachment of L 1 to the nucleobase (e.g., the C5 position of a pyrimidine base or the C7 position of a 7-deazapurine base).
[0125] In some embodiments, the nucleotides described herein are fully functionalized nucleotides (ffN) that contain the 3'-OH capping group described herein and a dye compound covalently attached to the nucleobase via a cleavable linker, wherein the cleavable linker contains an L having the structure and * indicates L 1 , and * indicates L 1Attachment points to nucleobases (e.g., the C5 position of cytosine, thymine, or uracil bases, or the C7 position of 7-deazaadenine or 7-deazaguanine). In some cases, the ffN having an allylamine or allylamide linker moiety described herein is also referred to as ffN-DB or ffN-(DB), where "DB" refers to the double bond in the linker moiety. In some cases, a sequencing run using a set of ffN in which one or more of the ffN are ffN-DB (including ffA, ffT, ffC, and ffG) provides a higher ffN incorporation rate compared to a set of ffN having a propargylamine or propargylamide linker moiety (also referred to as ffN-PA or ffN-(PA)) described herein. For example, a set of ffN-DB having an allylamine or allylamide linker moiety and a 3'-AOM capping group can increase the incorporation rate by at least 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 300%, 400%, or 500% compared to a set of ffN-PA having a 3'-O-azidomethyl capping group when processed for the same period of time under the same conditions, thereby improving the phasing value. In other embodiments, the incorporation rate / velocity is measured by the surface kinetics Vmax on the surface of a substrate (e.g., a flow cell or a cBot system). For example, a set of ffN-DB having a 3'-AOM capping group can increase the Vmax value (ms -1)At least 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 300%, 400%, or 500% increase. In some embodiments, the incorporation rate / velocity is measured at ambient temperature or a temperature below ambient temperature (such as 4 to 10 °C). In other embodiments, the incorporation rate / velocity is measured at an elevated temperature (such as 40 °C, 45 °C, 50 °C, 55 °C, 60 °C, or 65 °C). In some such embodiments, the incorporation rate / velocity is measured in a solution in an alkaline pH environment (e.g., pH 9.0, 9.2, 9.4, 9.6, 9.8, or 10.0). In some such embodiments, the incorporation rate / velocity is measured in the presence of an enzyme such as a polymerase (e.g., a DNA polymerase), terminal deoxynucleotidyl transferase, or reverse transcriptase. In some embodiments, ffN-DB is ffT-DB, ffC-DB, or ffA-DB. In one embodiment, the set of ffN-DBs having improved phasing values described herein includes ffT-DB, ffC, ffA, and ffG. In another embodiment, the set of ffN-DBs having improved phasing values described herein includes ffT-DB, ffC-DB, ffA, and ffG. In yet another embodiment, the set of ffN-DBs having improved phasing values described herein includes ffT-DB, ffC-DB, ffA-DB, and ffG.
[0126] In still other embodiments, when the nucleobase of the nucleotide described herein is thymine or its optionally substituted derivatives and analogs (i.e., the nucleotide is T), L 1 comprises an allylamine moiety or an allylamide moiety, or their optionally substituted variants. In a particular example, L 1 comprises or is and * indicates the attachment point of L 1 to the C5 position of the thymine base. In some embodiments, the T nucleotide described herein is a fully functionalized T nucleotide (ffT) labeled with a dye molecule via a cleavable linker, the cleavable linker comprising Attached directly to the C5 position of the thymine base (i.e., ffT-DB). In some cases, when ffT-DB is used in the presence of a palladium catalyst in sequencing applications, it can significantly improve sequencing metrics such as phasing, pre-phasing, and error rate. For example, when using ffT-DB with a 3′-AOM capping group as described herein, compared to the case when using a standard ffT-PA with a 3′-O-azidomethyl capping group, the ffT-DB can improve one or more of the sequencing metrics described herein by at least 50%, 100%, 200%, 300%, 400%, 500%, 600%, 700%, 800%, 900%, 1000%, 1500%, 2000%, 2500%, or 3000%.
[0127] Some additional embodiments of the nucleosides or nucleotides described herein include those having formula (Ia), (Ia′), (Ib), (Ic), (Ic′), or (Id):
[0128]
[0129]
[0130] In some additional embodiments of the nucleosides or nucleotides described herein, L 2 is present, and L 2 comprises wherein n and m are each independently an integer from 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10, and the phenyl moiety is optionally substituted. In some such embodiments, n is 5, and the phenyl moiety of L 2 is unsubstituted. In some additional embodiments, m is 4.
[0131] In any embodiment of the nucleosides or nucleotides described herein, the cleavable linker or L 1 / L 2 may also contain a disulfide moiety or an azide moiety (such as ), or a combination thereof. Additional non-limiting examples of linker moieties that can be incorporated into L 1 or L 2 include:
[0132]
[0133] . Additional linker moieties are disclosed in WO 2004 / 018493 and U.S. Patent Publication No. 2016 / 0040225, which are incorporated herein by reference.
[0134] In any embodiment of the nucleosides or nucleotides described herein, the nucleoside or nucleotide comprises a 2'-deoxyribose moiety (i.e., R in formulas (I) and (Ia) to (Id)) 4 is H). In other aspects, the 2'-deoxyribose bears one, two, or three phosphate groups at the 5'-position of the sugar ring. In other aspects, the nucleotides described herein are nucleotide triphosphates (i.e., R in formulas (I) and (Ia) to (Id)) 6 forms a triphosphate ester).
[0135] In any embodiment of the nucleosides or nucleotides described herein, the detectable label can include a fluorescent dye.
[0136] Additional embodiments of the present disclosure relate to oligonucleotides or polynucleotides comprising the nucleosides or nucleotides described herein. For example, an oligonucleotide or polynucleotide in which a nucleotide of formula (Ia') is incorporated comprises the following structure:
[0137]
[0138] In some such embodiments, the oligonucleotide or polynucleotide hybridizes to a template or target polynucleotide. In some such embodiments, the template polynucleotide is immobilized on a solid support.
[0139] Additional embodiments of the present disclosure relate to solid supports that include an array of a plurality of immobilized template or target polynucleotides, at least a portion of such immobilized template or target polynucleotides hybridizing to an oligonucleotide or polynucleotide comprising the nucleosides or nucleotides described herein.
[0140] In any embodiment of the nucleotides or nucleosides described herein, the 3'-OH capping group and the cleavable linker (and attached label) can be capable of being removed under the same or substantially the same chemical reaction conditions. For example, the 3'-OH capping group and the detectable label can be removed in a single chemical reaction. In other embodiments, the 3'-OH capping group and the detectable label are removed in two separate steps.
[0141] In some embodiments, compared to the same nucleotide or nucleoside protected with a standard 3′-OH capping group (such as a 3′-O-azidomethyl protecting group) disclosed in the prior art, the 3′-capped nucleotides or nucleosides described herein provide excellent stability in solution or in lyophilized form during storage or during reagent handling in sequencing applications. For example, the acetal capping group disclosed herein can increase stability by at least 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 300%, 400%, 500%, 600%, 700%, 800%, 900%, 1000%, 1500%, 2000%, 2500%, or 3000% compared to a 3′-OH protected with an azidomethyl group when treated under the same conditions for the same period of time, thereby reducing the predetermined phase value and allowing for longer sequencing read lengths. In some embodiments, stability is measured at ambient temperature or at a temperature below ambient temperature (such as 4 to 10 °C). In other embodiments, stability is measured at elevated temperatures (such as 40 °C, 45 °C, 50 °C, 55 °C, 60 °C, or 65 °C). In some such embodiments, stability is measured in a solution at an alkaline pH environment (e.g., pH 9.0, 9.2, 9.4, 9.6, 9.8, or 10.0). In some such embodiments, stability is measured in the presence or absence of an enzyme, such as a polymerase (e.g., a DNA polymerase), terminal deoxynucleotidyl transferase, or reverse transcriptase.
[0142] In some embodiments, compared to the same nucleotides or nucleosides protected with standard 3′-OH capping groups (such as 3′-O-azidomethyl protecting groups) disclosed in the prior art, the 3′-capped nucleotides or nucleosides described herein provide excellent uncapping rates in solution during the chemical cleavage step of sequencing applications. For example, the acetal capping groups disclosed herein can increase the uncapping rate by at least 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 150%, 200%, 300%, 400%, 500%, 600%, 700%, 800%, 900%, 1000%, 1500% or 2000% compared to 3′-OH protected with azidomethyl when using standard uncapping reagents (such as tris(hydroxypropyl)phosphine), thereby shortening the total time of the sequencing cycle. In some embodiments, the uncapping rate is measured at ambient temperature or at a temperature below ambient temperature (such as 4 to 10 °C). In other embodiments, the uncapping rate is measured at an elevated temperature (such as 40 °C, 45 °C, 50 °C, 55 °C, 60 °C or 65 °C). In some such embodiments, the uncapping rate is measured in a solution at an alkaline pH environment (e.g., pH 9.0, 9.2, 9.4, 9.6, 9.8 or 10.0). In some such embodiments, the molar ratio of the uncapping reagent to the substrate (i.e., the 3′-capped nucleoside or nucleotide) is about 10:1, about 5:1, about 2:1 or about 1:1.
[0143] In some embodiments, a palladium uncapping reagent (e.g., Pd(0)) is used to remove the 3′-acetal capping group (e.g., AOM capping group). Pd can form a chelate complex with the two oxygen atoms of the AOM group and the double bond of the allyl group, such that the uncapping reagent directly adjacent to the functional group is removed and the uncapping rate can be accelerated. For example, after Pd cleaves the linker and the 3′-capping group of the incorporated nucleotide having formula (Ia), (Ia′), (Ib), (Ic), (Ic′) or (Id) described herein, the remaining linker construct on the copy polynucleotide can have the following structure:
[0144]
[0145] The wavy line refers to the connection of oxygen to the remaining phosphodiester bond of the copy polynucleotide chain. For example,
[0146]
[0147] The allyl amido or propargyl amido moiety can be further cleaved by the Pd catalyst. In addition, the remaining linker construct attached to the detectable label has the following structure:
[0148] For example
[0149] Cleavage conditions of cleavable linkers
[0150] The cleavable linkers described herein can be removed or cleaved under various chemical conditions. Non-limiting cleavage conditions include palladium catalysts in the presence of water-soluble phosphine ligands such as tris(hydroxypropyl)phosphine (THP or THPP) or tris(hydroxymethyl)phosphine (THMP), such as Pd(II) complexes (e.g., Pd(OAc)2, allylpalladium(II) chloride dimer [(allyl)PdCl]2, or Na2PdCl4). In some embodiments, the 3′-acetal capping group can be cleaved under cleavage conditions that are the same as or substantially the same as those for the cleavage of the cleavable linker.
[0151] Palladium catalysts
[0152] In some embodiments, the 3′-acetal capping group and the cleavable linker described herein can be cleaved by a palladium catalyst. In some such embodiments, the Pd catalyst is water-soluble. In some such embodiments, the Pd catalyst is a Pd(0) complex (e.g., tris(3,3′,3″-phosphinidynetris(benzenesulfonic acid)palladium(0) nonahydrate sodium salt nonahydrate). In some cases, the Pd(0) complex can be generated in situ by reducing a Pd(II) complex with a reagent such as an alkene, alcohol, amine, phosphine, or metal hydride. Suitable palladium sources include Pd(CH3CN)2Cl2, [PdCl(allyl)]2, [Pd(allyl)(THP)]Cl, [Pd(allyl)(THP)2]Cl, Pd(OAc)2, Pd(PPh3)4, Pd(dba)2, Pd(Acac)2, PdCl2(COD), and Pd(TFA)2. In one such embodiment, the Pd(0) complex is generated in situ from Na2PdCl4. In another embodiment, the palladium source is allylpalladium(II) chloride dimer [(allyl)PdCl]2 or [PdCl(C3H5)]2. In some embodiments, the Pd(0) catalyst is generated in an aqueous solution by mixing a Pd(II) complex with a phosphine. Suitable phosphines include water-soluble phosphines such as tris(hydroxypropyl)phosphine (THP), tris(hydroxymethyl)phosphine (THMP), 1,3,5-triaza-7-phosphaadamantane (PTA), dipotassium bis(p-sulfonatophenyl)phenylphosphine dihydrate, tris(carboxyethyl)phosphine (TCEP), and trisodium triphenylphosphine-3,3′,3′-trisulfonate.
[0153] In some embodiments, the palladium catalyst is prepared by in-situ mixing of [(allyl)PdCl]2 with THP. The molar ratio of [(allyl)PdCl]2 to THP can be about 1:2, 1:3, 1:4, 1:5, 1:6, 1:7, 1:8, 1:9, or 1:10. In one embodiment, the molar ratio of [(allyl)PdCl]2 to THP is 1:10. In some other embodiments, the palladium catalyst is prepared by in-situ mixing of the water-soluble Pd reagent Na2PdCl4 with THP. The molar ratio of Na2PdCl4 to THP can be about 1:2, 1:3, 1:4, 1:5, 1:6, 1:7, 1:8, 1:9, or 1:10. In one embodiment, the molar ratio of Na2PdCl4 to THP is about 1:3. In another embodiment, the molar ratio of Na2PdCl4 to THP is about 1:3.5. In some additional embodiments, one or more reducing agents, such as ascorbic acid or its salts (e.g., sodium ascorbate), can be added. In some embodiments, the cleavage mixture can contain additional buffer reagents, such as primary amines, secondary amines, tertiary amines, natural amino acids, unnatural amino acids, carbonates, phosphates, or borates, or combinations thereof. In some other embodiments, the buffer reagents include ethanolamine (EA), tris(hydroxymethyl)aminomethane (Tris), glycine, sodium carbonate, sodium phosphate, sodium borate, dimethylethanolamine (DMEA), diethylethanolamine (DEEA), N,N,N′,N′-tetramethylethylenediamine (TMEDA), N,N,N′,N′-tetraethylethylenediamine (TEEDA), or 2-piperidinoethanol, or combinations thereof. In one embodiment, one or more buffer reagents include DEEA. In another embodiment, one or more buffer reagents contain one or more inorganic salts, such as carbonates, phosphates, or borates, or combinations thereof. In one embodiment, the inorganic salt is a sodium salt.
[0154] In other embodiments, the cleavage conditions of the cleavable linker are different from the cleavage conditions of the 3′-OH capping group. For example, in addition, when the 3′ capping group is 3′-O-azidomethyl, the –CH2N3 moiety can be converted to an amino group by phosphine. Alternatively, the azide group in –CH2N3 can be converted to an amino group by contacting these molecules with a thiol, particularly a water-soluble thiol such as dithiothreitol (DTT). In one embodiment, the phosphine is THP.
[0155] Compatible with linearization
[0156] To maximize the throughput of nucleic acid sequencing reactions, it is advantageous to be able to sequence multiple template molecules in parallel. Parallel processing of multiple templates can be achieved by using nucleic acid array technology. These arrays typically consist of a high-density matrix of polynucleotides immobilized on a solid support material.
[0157] Both WO 98 / 44151 and WO 00 / 18957 describe nucleic acid amplification methods that allow amplification products to be immobilized on a solid support to form arrays consisting of clusters or "colonies", where the clusters or "colonies" are formed from multiple identical immobilized polynucleotide strands and multiple identical immobilized complementary strands. Arrays of this type are referred to herein as "clustered arrays". The nucleic acid molecules present in the DNA colonies on the clustered arrays prepared according to these methods can provide templates for sequencing reactions, for example as described in WO 98 / 44152. The products of solid-phase amplification reactions (such as those described in WO 98 / 44151 and WO 00 / 18957) are so-called "bridged" structures that are formed by annealing pairs of immobilized polynucleotide strands and immobilized complementary strands (both strands attached to the solid support at the 5' end). To provide a more suitable template for nucleic acid sequencing, it is preferred to remove substantially all or at least a portion of one of the immobilized strands of the "bridged" structure to produce a template that is at least partially single-stranded. Thus, the single-stranded portion of the template will be available for hybridization with a sequencing primer. The process of removing all or a portion of one of the immobilized strands from the "bridged" double-stranded nucleic acid structure is referred to as "linearization". There are various ways to perform linearization, including but not limited to enzymatic cleavage, photochemical cleavage, or chemical cleavage. Non-limiting examples of linearization methods are disclosed in PCT Publication No. WO 2007 / 010251, U.S. Patent Publication No. 2009 / 0088327, U.S. Patent Publication No. 2009 / 0118128, and U.S. Patent Publication No. 2019 / 0352327, which are incorporated herein by reference in their entireties.
[0158] In particular, amplification (e.g., bridge amplification or exclusive amplification) forms an array consisting of clusters or "colonies", where the clusters or "colonies" are formed by multiple identical immobilized target polynucleotide strands and multiple identical immobilized complementary strands. The target strand and the complementary strand form a polynucleotide complex that is at least partially double-stranded, where the two strands are immobilized at their 5' ends to a solid support. The double-stranded polynucleotide is contacted with an aqueous solution of a palladium catalyst, which cleaves one strand at a cleavage site containing an allyl-modified nucleoside (e.g., an allyl-modified T nucleoside) to remove at least a portion of one of the immobilized strands, thereby producing a template that is at least partially single-stranded. Thus, the single-stranded portion of the template will be available for hybridization with a sequencing primer to initiate the first round of SBS (read 1). In some embodiments, the allyl-modified nucleoside is located in the P5 primer sequence. This method is referred to as first chemical linearization, as compared to standard enzymatic linearization, where this removal or cleavage is facilitated by an enzymatic cleavage reaction that cleaves the U site on the P5 primer using the enzyme USER.
[0159] In some embodiments, the conditions for cleaving the cleavable linker and / or deprotecting the 3'-OH capping group or removing the 3'-OH capping group are also compatible with these linearization methods. In some additional embodiments, such cleavage conditions are compatible with a chemical linearization method that includes the use of a Pd complex and a phosphine. In some embodiments, the Pd complex is a Pd(II) complex (e.g., Pd(OAc)2, [(allyl)PdCl]2, or Na2PdCl4) , which in situ generates Pd(0) in the presence of a phosphine (e.g., THP). The chemical linearization method of cleaving an allyl-modified T nucleoside in the P5 primer sequence using a Pd catalyst is described in detail in U.S. Patent Publication No. 2019 / 0352327, which is incorporated herein by reference in its entirety. In additional embodiments, the Pd cleavage mixture disclosed herein (e.g., a buffer solution containing DEEA of [Pd(allyl)Cl)2] and THP) can be directly used in the first chemical linearization step. The reduction in the number of reagents enables further simplification of the instrument (flow channels and cartridges).
[0160] Unless otherwise specified, reference to a nucleotide is also intended to apply to a nucleoside.
[0161] Labeled nucleotides
[0162] According to one aspect of the present disclosure, the described 3′-OH capped nucleotides further include a detectable label, and such nucleotides are referred to as labeled nucleotides or fully functionalized nucleotides (ffN). The label (e.g., a fluorescent dye) is conjugated via a cleavable linker in a variety of ways, including hydrophobic attraction, ionic attraction, and covalent linkage. In some aspects, the dye is conjugated to the nucleotide via covalent linkage achieved through a cleavable linker. In some cases, such labeled nucleotides are also referred to as "modified nucleotides". Those of ordinary skill in the art understand that the label can be covalently attached to the linker by reacting the functional group of the label (e.g., carboxyl group) with the functional group of the linker (e.g., amino group).
[0163] Labeled nucleosides and nucleotides can be used to label polynucleotides formed by enzymatic synthesis in, for example (by way of non-limiting example), PCR amplification, isothermal amplification, solid-phase amplification, polynucleotide sequencing (e.g., solid-phase sequencing), nick translation reactions, and the like.
[0164] In some embodiments, the dye can be covalently attached to an oligonucleotide or nucleotide via a nucleobase. For example, a labeled nucleotide or oligonucleotide can have a label attached to the C5 position of a pyrimidine base or the C7 position of a 7-deazapurine base via a cleavable linker moiety.
[0165] Unless otherwise specified, reference to nucleotides is also intended to apply to nucleosides. This application will further describe with reference to DNA, however, this description will also apply to RNA, PNA, and other nucleic acids, unless otherwise specified as not applicable.
[0166] Nucleosides and nucleotides can be labeled at sites on the sugar or nucleobase. As is known in the art, a "nucleotide" consists of a nitrogenous base, a sugar, and one or more phosphate groups. In RNA, the sugar is ribose, while in DNA, the sugar is deoxyribose, i.e., a sugar lacking the hydroxyl group present in ribose. The nitrogenous base is a derivative of purine or pyrimidine. Purines are adenine (A) and guanine (G), and pyrimidines are cytosine (C) and thymine (T), or in the case of RNA, uracil (U). The C-1 atom of deoxyribose is bonded to the N-1 of pyrimidine or the N-9 of purine. A nucleotide is also a phosphate ester of a nucleoside, where esterification occurs at the hydroxyl group attached to the C-3 or C-5 of the sugar. Nucleotides are typically monophosphate, diphosphate, or triphosphate.
[0167] A "nucleoside" is structurally similar to a nucleotide but lacks the phosphate moiety. An example of a nucleoside analogue is a nucleoside analogue in which the label is linked to the base and no phosphate group is linked to the sugar molecule.
[0168] Although bases are commonly referred to as purines or pyrimidines, one of ordinary skill in the art will understand that derivatives and analogs are available that do not alter the ability of the nucleotide or nucleoside to undergo Watson-Crick base pairing. A "derivative" or "analog" means a compound or molecule whose core structure is the same as or very similar to that of the parent compound, but which has chemical or physical modifications that allow the derived nucleotide or nucleoside to be linked to another molecule, such as different or additional side groups. For example, a base can be a deazapurine. In certain embodiments, the derivative should be capable of undergoing Watson-Crick pairing. "Derivatives" and "analogs" also include, for example, synthetic nucleotide or nucleoside derivatives having a modified base moiety and / or a modified sugar moiety. Such derivatives and analogs are discussed, for example, in Scheit, Nucleotide analogs (John Wiley & Son, 1980) and Uhlman et al., Chemical Reviews 90:543-584, 1990. Nucleotide analogs can also include modified phosphodiester bonds, including phosphorothioate bonds, dithiophosphonate bonds, alkylphosphonate bonds, anilinophosphonate bonds, phosphoramidate bonds, and the like.
[0169] In certain embodiments, the labeled nucleoside or nucleotide can be enzymatically incorporated and enzymatically extended. Thus, the linker moiety can have a length sufficient to link the nucleotide to the compound such that the compound does not significantly interfere with the overall binding and recognition of the nucleotide by nucleic acid replicating enzymes. Thus, the linker can also include spacer units. The spacer distance is, for example, the distance of a nucleobase from the cleavage site or the label.
[0170] The present disclosure also encompasses polynucleotides incorporating dye compounds. Such polynucleotides can be DNA or RNA composed of deoxyribonucleotides or ribonucleotides joined by phosphodiester bonds, respectively. The polynucleotide can include a combination of naturally occurring nucleotides, non-naturally occurring (or modified) nucleotides other than the labeled nucleotides described herein, or any combination thereof, with at least one modified nucleotide as shown herein (e.g., labeled with a dye compound). Polynucleotides according to the present disclosure can also include non-natural backbone linkages and / or non-nucleotide chemical modifications. Chimeric structures composed of mixtures of ribonucleotides and deoxyribonucleotides containing at least one labeled nucleotide are also contemplated.
[0171] Non-limiting exemplary labeled nucleotides as described herein include:
[0172]
[0173]
[0174] wherein L represents a cleavable linker (optionally including L as described herein) 2 ), R represents a ribose or deoxyribose moiety as described above, or a ribose or deoxyribose moiety substituted at the 5'-position with one, two, or three phosphates.
[0175] In some embodiments, non-limiting exemplary fluorescent dye conjugates are shown below:
[0176]
[0177]
[0178] wherein PG represents a 3'-OH capping group as described herein; n is an integer from 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10; k is 0, 1, 2, 3, 4, or 5. In one embodiment, –O–PG is AOM. In another embodiment, –O–PG is –O–azidomethyl. In one embodiment, n is 5. refers to the point of attachment of the dye to the cleavable linker as a result of the reaction between the amino group as the linker moiety and the carboxyl group of the dye.
[0179] Sequencing methods
[0180] The labeled nucleotides or nucleosides according to the present disclosure can be used in any analytical method (such as methods including detecting a fluorescent label attached to a nucleotide or nucleoside), either used by itself or incorporated into or associated with a larger molecular structure or conjugate. In this context, the term "incorporated into a polynucleotide" can mean that the 5'-phosphate is joined to the 3'-OH group of a second (modified or unmodified) nucleotide by a phosphodiester bond, and this second nucleotide itself can form part of a longer polynucleotide chain. The 3'-end of the nucleotide set forth herein can or cannot be joined to the 5'-phosphate of another (modified or unmodified) nucleotide by a phosphodiester bond. Thus, in a non-limiting embodiment, the present disclosure provides a method for detecting a nucleotide incorporated into a polynucleotide, the method comprising: (a) incorporating at least one nucleotide of the present disclosure into a polynucleotide, and (b) detecting the nucleotide incorporated into the polynucleotide by detecting a fluorescent signal from a detectable label (e.g., a fluorescent compound) attached to the nucleotide. The method can include a synthesis step (a) in which one or more nucleotides according to the present disclosure are incorporated into a polynucleotide; and a detection step (b) in which the nucleotides are detected by detecting or quantitatively measuring the fluorescence of one or more nucleotides incorporated into the polynucleotide.
[0181] Additional aspects of the disclosure include a method of preparing a growing polynucleotide complementary to a target single-stranded polynucleotide in a sequencing reaction, the method comprising incorporating a nucleotide as described herein into the growing complementary polynucleotide, wherein incorporation of the nucleotide prevents introduction of any subsequent nucleotide into the growing complementary polynucleotide.
[0182] Some embodiments of the disclosure relate to a method for determining the sequence of a target single-stranded polynucleotide, the method comprising:
[0183] (a) incorporating a nucleotide (e.g., dATP, dCTP, dGTP, dTTP, or dUTP) comprising a 3′-OH capping group as described herein (attached to a 3′-oxy group) and a detectable label as described herein into a copy polynucleotide strand that is complementary to at least a portion of the target polynucleotide strand;
[0184] (b) detecting the identity of the nucleotide incorporated into the copy polynucleotide strand; and
[0185] (c) chemically removing the label and the 3′-OH capping group from the nucleotide incorporated into the copy polynucleotide strand.
[0186] In some embodiments, the sequencing method further comprises (d) washing the chemically removed label and 3′-OH capping group from the copy polynucleotide strand using a post-cleavage wash solution. In some such embodiments, the 3′-OH capping group and the detectable label are removed prior to introduction of the next complementary nucleotide. In some embodiments, the washing step (d) also removes unincorporated nucleotides. In other embodiments, the method may include a separate washing step for washing unincorporated nucleotides from the copy polynucleotide strand prior to step (b).
[0187] In some embodiments, steps (a) through (d) are repeated until the sequence of the portion of the target polynucleotide strand is determined. In some such embodiments, steps (a) through (d) are repeated at least 50 times, at least 75 times, at least 100 times, at least 150 times, at least 200 times, at least 250 times, or at least 300 times.
[0188] Incorporation mixtures
[0189] In some embodiments of the methods described herein, step (a) (also referred to as the incorporation step) includes contacting a mixture containing one or more nucleotides (e.g., dATP, dCTP, dGTP, dTTP, or dUTP) with a copy polynucleotide / target polynucleotide complex in an incorporation solution that includes a polymerase and one or more buffers. In some such embodiments, the polymerase is a DNA polymerase, such as Pol 812, Pol 1901, Pol 1558, or Pol 963. The amino acid sequences of the Pol 812, Pol 1901, Pol 1558, or Pol 963 DNA polymerases are described, for example, in U.S. Patent Publication Nos. 2020 / 0131484A1 and 2020 / 0181587A1, both of which are incorporated herein by reference. In some embodiments, the one or more buffers include a primary amine, a secondary amine, a tertiary amine, a natural amino acid, or an unnatural amino acid, or a combination thereof. In additional embodiments, the buffer includes ethanolamine or glycine, or a combination thereof. In one embodiment, the buffer includes glycine, or is glycine. In some embodiments, using glycine in the incorporation mixture can improve the phasing value as compared to a standard buffer (such as ethanolamine (EA)) under the same conditions. For example, using glycine as compared to using ethanolamine under the same conditions reduces or decreases the phasing value by at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 300%, 400%, or 500%. In some cases, using glycine provides a % phasing value of less than about 0.15%, 0.1%, 0.09%, 0.08%, 0.07%, 0.06%, 0.05%, 0.04%, 0.03%, 0.02%, or 0.01% in an SBS sequencing run having at least 50 cycles. In additional embodiments, using glycine provides a % phasing value of less than about 0.08% in read 1 of an SBS sequencing run having at least 150 cycles.
[0190] Cleavage mixtures
[0191] In some embodiments of the methods described herein, step (c) (also referred to as the cleavage step) includes contacting the incorporated nucleotides and the copy polynucleotide chain with a cleavage solution that includes the palladium catalyst described herein. In some such embodiments, the 3′-OH capping group and the detectable label are removed in a single step of the reaction. In one such embodiment, the 3′ capping group is AOM, and the cleavable linker includes an AOL moiety, both of which are removed or cleaved in a single step of the chemical reaction. In still other embodiments, the cleavage solution (also referred to as the cleavage mixture) includes the Pd catalyst described herein.
[0192] In some other embodiments, the Pd catalyst is a Pd(0) catalyst. In some such embodiments, Pd(0) is prepared by in-situ mixing of a Pd(II) reagent with one or more phosphine ligands. In some such embodiments, the palladium catalyst can be prepared by in-situ mixing of [(allyl)PdCl]2 with THP. The molar ratio of [(allyl)PdCl]2 to THP can be about 1:2, 1:3, 1:4, 1:5, 1:6, 1:7, 1:8, 1:9 or 1:10. In one embodiment, the molar ratio of [(allyl)PdCl]2 to THP is 1:10 (i.e., the molar ratio of Pd:THP is 1:5). In some other embodiments, the palladium catalyst can be prepared by in-situ mixing of the water-soluble Pd(II) reagent Na2PdCl4 with THP. The molar ratio of Na2PdCl4 to THP can be about 1:2, 1:3, 1:4, 1:5, 1:6, 1:7, 1:8, 1:9 or 1:10. In one embodiment, the molar ratio of Na2PdCl4 to THP is about 1:3. In another embodiment, the molar ratio of Na2PdCl4 to THP is about 1:3.5. Other non-limiting examples of the Pd catalyst include Pd(CH3CN)2Cl2.
[0193] In some additional embodiments, one or more reducing agents can be added, such as ascorbic acid or its salts (e.g., sodium ascorbate). In some embodiments, the lysis solution can contain one or more buffer reagents, such as primary amines, secondary amines, tertiary amines, carbonates, phosphates or borates, or combinations thereof. In some other embodiments, the buffer reagents include ethanolamine (EA), tris(hydroxymethyl)aminomethane (Tris), glycine, sodium carbonate, sodium phosphate, sodium borate, dimethylethanolamine (DMEA), diethylethanolamine (DEEA), N,N,N′,N′-tetramethylethylenediamine (TEMED), N,N,N′,N′-tetraethylethylenediamine (TEEDA) or 2-piperidinoethanol, or combinations thereof. In one embodiment, the buffer reagent includes DEEA, or is DEEA. In another embodiment, the buffer contains one or more inorganic salts, such as carbonates, phosphates or borates, or combinations thereof. In one embodiment, the inorganic salt is a sodium salt. In some other embodiments, the lysis solution contains a palladium (Pd) catalyst (e.g., [(allyl)PdCl]2 / THP or Na2PdCl4 / THP) and one or more of the buffer reagents described herein (e.g., a tertiary amine such as DEEA), and has a pH of about 9.0 to about 10.0 (e.g., 9.6 or 9.8).
[0194] In other embodiments, the label and the 3′-blocking group are removed in two separate chemical reactions. In some cases, removing the label from the nucleotide incorporated into the copy polynucleotide chain comprises contacting the copy strand comprising the incorporated nucleotide with a first cleavage solution containing a Pd catalyst as described herein. In some cases, removing the 3′-OH blocking group from the nucleotide incorporated into the copy polynucleotide chain comprises contacting the copy strand comprising the incorporated nucleotide with a second cleavage solution. In some such embodiments, the second cleavage solution contains one or more phosphines, such as trialkylphosphines. Non-limiting examples of trialkylphosphines include tris(hydroxypropyl)phosphine (THP), tris-(2-carboxyethyl)phosphine (TCEP), tris(hydroxymethyl)phosphine (THMP), or tris(hydroxyethyl)phosphine (THEP). In one embodiment, the 3′-OH blocking group is 3′-O-azidomethyl and the second cleavage solution contains THP.
[0195] In some embodiments, the cleavage solutions described herein can also be used in the prior chemical linearization step described herein. In particular, one or more first strands of a double-stranded polynucleotide immobilized on a solid support are cleaved by palladium-catalyzed cleavage to produce a single-stranded (or at least partially single-stranded) template that will be available for hybridization with a sequencing primer and subsequent sequencing applications (e.g., the first round (read 1) of sequencing-by-synthesis), thereby effecting chemical linearization of the clustered polynucleotides prepared for sequencing. In some embodiments, each double-stranded polynucleotide comprises a first strand and a second strand. The first strand is produced by extending a first extension primer immobilized to the solid support. In some embodiments, the first strand comprises a cleavage site that is cleavable by a palladium complex (e.g., a Pd(0) complex). In a specific embodiment, the cleavage site is located in the first extension primer portion of the first strand. In an additional embodiment, the cleavage site comprises a thymidine or nucleotide analogue having allyl functionality. In some embodiments of the methods described herein, the target single-stranded polynucleotide is formed by chemically cleaving the complementary strand from the double-stranded polynucleotide. In additional embodiments, the complementary strand and the target polynucleotide in the double strand are both immobilized at their 5′ ends on the solid support. In some further embodiments, chemically cleaving the complementary strand is carried out under the same reaction conditions as chemically removing the detectable label and the 3′-OH blocking group from the nucleotide incorporated into the copy polynucleotide chain (i.e., step (c) of the methods described herein). In one embodiment, the first chemical linearization utilizes the same cleavage mixture described herein.
[0196] Palladium (Pd) scavengers
[0197] Pd has the ability to adhere to DNA primarily in its inactive Pd(II) form, which may interfere with the binding between DNA and polymerase, resulting in increased phasing. After the unblocking step, a post-cleavage washing composition comprising a Pd scavenger compound can be used. For example, PCT Publication No. WO 2020 / 126593 discloses that Pd scavengers (such as 3,3'-dithiodipropionic acid (DDPA) and lipoic acid (LA)) can be included in the scanning composition and / or the post-cleavage washing composition. The purpose of using these scavengers in the post-cleavage washing solution is to remove Pd(0) and convert Pd(0) into an inactive Pd(II) form, thereby improving pre-phasing values and sequencing indicators, reducing signal degradation, and extending sequencing read lengths.
[0198] In some embodiments of the methods described herein, step (a) of the method comprises contacting the nucleotides with the copy polynucleotide chain in a spiking solution comprising a polymerase, at least one palladium scavenger, and one or more buffers. In some embodiments, the Pd scavenger spiked into the solution is a Pd(0) scavenger. In some such embodiments, the Pd scavenger comprises one or more allyl moieties independently selected from the group consisting of: -O-allyl, -S-allyl, -NR-allyl, and -N-allyl. + RR′-allyl, wherein R is H, unsubstituted or substituted C1-C6 alkyl, unsubstituted or substituted C2-C6 alkenyl, unsubstituted or substituted C2-C6 alkynyl, unsubstituted or substituted C6-C 10 aryl, unsubstituted or substituted 5- to 10-membered heteroaryl, unsubstituted or substituted C3-C 10 carbocyclyl, or unsubstituted or substituted 5- to 10-membered heterocyclyl; and R′ is H, unsubstituted C1-C6 alkyl or substituted C1-C6 alkyl.
[0199] In some such embodiments, the Pd(0) scavenger incorporated into the solution comprises one or more -O-allyl moieties. In other embodiments, the Pd(0) scavenger comprises or is
[0200] or a combination thereof. Alternative Pd(0) scavengers are disclosed in US Serial No. 63 / 190983, which is incorporated by reference in its entirety.
[0201] In some embodiments, the concentration of the Pd(0) scavenger comprising one or more allyl moieties in the incorporation solution is from about 0.1 mM to about 100 mM, 0.2 mM to about 75 mM, about 0.5 mM to about 50 mM, about 1 mM to about 20 mM, or about 2 mM to about 10 mM. In additional embodiments, the concentration of the Pd(0) scavenger is about 0.5 mM, 1 mM, 1.5 mM, 2 mM, 2.5 mM, 3 mM, 3.5 mM, 4 mM, 4.5 mM, 5 mM, 5.5 mM, 6 mM, 6.5 mM, 7 mM, 7.5 mM, 8 mM, 8.5 mM, 9 mM, 9.5 mM, 10 mM, 12.5 mM, 15 mM, 17.5 mM, or 20 mM. In additional embodiments, the pH of the incorporation solution is from about 9 to 10.
[0202] In some embodiments, the molar ratio of the palladium catalyst (in the starting solution) to the palladium scavenger comprising one or more allyl moieties is about 1:100, 1:50, 1:20, 1:10, or 1:5.
[0203] In some other embodiments of the methods described herein, when one or more fluorescence measurements are made to detect the identity of the nucleotide incorporated into the copy polynucleotide, the Pd(0) scavenger comprising one or more allyl moieties described herein is used in the scanning solution of step (b). In still other embodiments, the Pd(0) scavenger comprising one or more allyl moieties can be present in both the incorporation solution and the scanning solution.
[0204] In some additional embodiments of the methods described herein, the post-cleavage wash step is used after removal of the label and 3′ capping group. In some such embodiments, one or more palladium scavengers are also used in the wash step after cleavage of the label and 3′ capping group. In some additional embodiments, one or more Pd scavengers in the post-cleavage wash solution comprise Pd(II) scavengers. In some such embodiments, the palladium scavenger comprises an isocyanoacetate (ICNA) salt, cysteine or a salt thereof, or a combination thereof. In one embodiment, the palladium scavenger comprises or is potassium isocyanoacetate or sodium isocyanoacetate. In another embodiment, the palladium scavenger comprises or is cysteine or a salt thereof (e.g., L-cysteine or L-cysteine HCl salt). Other non-limiting examples of palladium scavengers in the post-cleavage wash solution can include: ethyl isocyanoacetate, methyl isocyanoacetate, N-acetyl-L-cysteine, potassium ethyl xanthate (PEX or KS-C(═S)-OEt), potassium isopropyl xanthate, glutathione, lipoic acid, ethylenediaminetetraacetic acid (EDTA), iminodiacetic acid, nitrilodiacetic acid, trimercapto-S-triazine, dimethyldithiocarbamate, dithiothreitol, mercaptoethanol, allyl alcohol, propargyl alcohol, thiols, tertiary amines, and / or tertiary phosphines, or a combination thereof.
[0205] In some additional embodiments, the concentration of the Pd(II) scavenger (such as L-cysteine) in the post-cleavage wash solution is from about 0.1 mM to about 100 mM, 0.2 mM to about 75 mM, about 0.5 mM to about 50 mM, about 1 mM to about 20 mM, or about 2 mM to about 10 mM. In some additional embodiments, the concentration of the Pd(II) scavenger (such as L-cysteine) is about 1 mM, 2 mM, 3 mM, 4 mM, 5 mM, 6 mM, 6.5 mM, 7 mM, 8 mM, 9 mM, 10 mM, 12.5 mM, 15 mM, 17.5 mM, or 20 mM. In one embodiment, the concentration of the Pd scavenger (such as L-cysteine or a salt thereof) in the post-cleavage wash solution is about 10 mM.
[0206] In some other embodiments of the methods described herein, all Pd scavengers (e.g., both Pd(0) scavengers and Pd(II) scavengers) are in the incorporation solution and / or the scan solution, and the method does not employ a specific post-cleavage wash step to remove any trace amounts of remaining Pd species.
[0207] In some embodiments of the methods described herein, the use of a Pd scavenger (e.g., a Pd(0) scavenger having one or more allyl moieties) can reduce a predetermined phasing value by at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 300%, 400%, 500%, 600%, 700%, 800%, 900%, or 1000% compared to the same sequencing run under the same conditions without the use of a palladium scavenger. In some such embodiments, the Pd(0) scavenger can reduce the predetermined phasing value of the sequencing run to less than about 0.1%, 0.09%, 0.08%, 0.07%, 0.06%, 0.05%, 0.04%, 0.03%, 0.02%, or 0.01% in an SBS sequencing run having at least 50 cycles. In some embodiments, the predetermined phasing value refers to the value measured after 50 cycles, 75 cycles, 100 cycles, 125 cycles, 150 cycles, 200 cycles, 250 cycles, or 300 cycles.
[0208] In still other embodiments, a palladium scavenger (e.g., a Pd(II) scavenger such as L-cysteine or a salt thereof) can reduce a predetermined phasing value or a phasing value by at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 300%, 400%, 500%, 600%, 700%, 800%, 900%, or 1000% compared to the same sequencing run under the same conditions without the use of a palladium scavenger. In some such embodiments, the use of a Pd scavenger provides a % phasing value of less than about 0.2%, 0.15%, 0.1%, 0.09%, 0.08%, 0.07%, 0.06%, 0.05%, 0.04%, 0.03%, 0.02%, or 0.01% in an SBS sequencing run having at least 50 cycles. In some embodiments, the phasing value refers to the value measured after 50 cycles, 75 cycles, 100 cycles, 125 cycles, 150 cycles, 200 cycles, 250 cycles, or 300 cycles. In additional embodiments, the use of one or more Pd scavengers provides a % phasing value of less than about 0.05% in read 1 of an SBS sequencing run having at least 150 cycles.
[0209] In some embodiments, the post-cleavage wash solution described herein can also be used in a separate wash step prior to the detection step (i.e., step (b) in the methods described herein) to wash away any unincorporated nucleotides from step (a).
[0210] In some other embodiments, the nucleotides used in incorporation step (a) are fully functionalized A, C, T, and G nucleotide triphosphates, each containing a 3′ capping group (e.g., 3′-AOM) and a cleavable linker (e.g., a cleavable linker containing an AOL linker moiety) as described herein. In some such embodiments, the nucleotides herein provide excellent stability in solution during a sequencing run compared to the same nucleotides protected with a standard 3′-O-azidomethyl capping group. For example, the 3′ acetal capping group disclosed herein can increase stability by at least 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 300%, 400%, 500%, 600%, 700%, 800%, 900%, 1000%, 1500%, 2000%, 2500%, or 3000% compared to the 3′-OH protected with azidomethyl under the same conditions for the same period of time, thereby reducing the predetermined phase value and allowing for longer sequencing read lengths. In some embodiments, stability is measured at ambient temperature or at a temperature below ambient temperature, such as 4 to 10 °C. In other embodiments, stability is measured at an elevated temperature, such as 40 °C, 45 °C, 50 °C, 55 °C, 60 °C, or 65 °C. In some such embodiments, stability is measured in a solution in a basic pH environment (e.g., pH 9.0, 9.2, 9.4, 9.6, 9.8, or 10.0). In some other embodiments, after more than 50, 100, or 150 SBS cycles, the predetermined phase value of the 3′-capped nucleotides herein is less than about 0.25, 0.24, 0.23, 0.22, 0.21, 0.20, 0.19, 0.18, 0.17, 0.16, 0.15, 0.14, 0.13, 0.12, 0.11, 0.10, 0.09, 0.08, 0.07, 0.06, or 0.05. In some other embodiments, after more than 50, 100, or 150 SBS cycles, the phasing value of the 3′-capped nucleotides is less than about 0.25, 0.24, 0.23, 0.22, 0.21, 0.20, 0.19, 0.18, 0.17, 0.16, 0.15, 0.14, 0.13, 0.12, 0.11, 0.10, 0.09, 0.08, 0.07, 0.06, or 0.05. In one embodiment, each ffN contains a 3′-AOM group.
[0211] In some embodiments, compared to the same nucleotides protected with a standard 3′-O-azidomethyl capping group, the 3′-capped nucleotides described herein provide an excellent uncapping rate in solution during the chemical cleavage step of a sequencing run. For example, the 3′-acetal (e.g., AOM) capping group disclosed herein can increase the uncapping rate by at least 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 150%, 200%, 300%, 400%, 500%, 600%, 700%, 800%, 900%, 1000%, 1500%, or 2000% compared to the 3′-OH protected with azidomethyl when using a standard uncapping reagent such as tris(hydroxypropyl)phosphine, thereby shortening the total time of the sequencing cycle. In some embodiments, the uncapping time for each nucleotide is shortened by about 5%, 10%, 20%, 30%, 40%, 50%, or 60%. For example, under specific chemical reaction conditions, the uncapping times for 3′-AOM and 3′-O-azidomethyl are about 4 to 5 seconds and about 9 to 10 seconds, respectively. In some embodiments, the half-life (t 1 / 2 ) of the AOM capping group is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 times faster than that of the azidomethyl capping group. In some such embodiments, the t 1 / 2 of AOM is about 1 minute, while the t 1 / 2 of azidomethyl is about 11 minutes. In some embodiments, the uncapping rate is measured at ambient temperature or a temperature below ambient temperature, such as 4 to 10 °C. In other embodiments, the uncapping rate is measured at an elevated temperature, such as 40 °C, 45 °C, 50 °C, 55 °C, 60 °C, or 65 °C. In some such embodiments, the uncapping rate is measured in a solution at an alkaline pH environment (e.g., pH 9.0, 9.2, 9.4, 9.6, 9.8, or 10.0). In some such embodiments, the molar ratio of the uncapping reagent to the substrate (i.e., the 3′-capped nucleoside or nucleotide) is about 10:1, about 5:1, about 2:1, about 1:1, about 1:2, about 1:5, or about 1:10. In one embodiment, each ffN contains a 3′-AOM capping group and an AOL linker moiety.
[0212] In any embodiment of the methods described herein, the labeled nucleotide is a nucleotide triphosphate having 2′-deoxyribose. In any embodiment of the methods described herein, the target polynucleotide chain is attached to a solid support, such as a flow cell.
[0213] In one embodiment, at least one nucleotide is incorporated into a polynucleotide by the action of a polymerase during a synthesis step. In some such embodiments, the polymerase can be DNA polymerase Pol 812 or Pol 1901. However, other methods of joining nucleotides to polynucleotides can be used, such as chemically synthesizing oligonucleotides or ligating labeled oligonucleotides to unlabeled oligonucleotides. Thus, when the term "incorporated" is used with respect to nucleotides and polynucleotides, it can encompass polynucleotide synthesis by chemical as well as enzymatic methods.
[0214] In a specific embodiment, a synthesis step is performed, which can optionally include incubating a template polynucleotide strand with a reaction mixture comprising a labeled 3'-blocked nucleotide of the present disclosure. A polymerase can also be provided under conditions that permit the formation of a phosphodiester bond between a free 3'-OH group on a polynucleotide strand annealed to the template polynucleotide strand and a 5'-phosphate group on a nucleotide. Thus, the synthesis step can include directing the formation of a polynucleotide strand by complementary base pairing of nucleotides with the template strand.
[0215] In all embodiments of these methods, the detection step can be performed when a polynucleotide strand into which a labeled nucleotide has been incorporated is annealed to a template strand or a target strand, or after a denaturation step that separates the two strands. Additional steps can be included between the synthesis step and the detection step, such as a chemical reaction step or an enzymatic reaction step, or a purification step. Specifically, a target strand incorporating a labeled nucleotide can be isolated or purified and then further processed or used for subsequent analysis. By way of example, a target polynucleotide labeled with a nucleotide as described herein during the synthesis step can subsequently be used as a labeled probe or primer. In other embodiments, the product of the synthesis step described herein can be subjected to further reaction steps, and if desired, the products of these subsequent steps can be purified or isolated.
[0216] Suitable conditions for the synthesis step will be well known to those skilled in standard molecular biology techniques. In one embodiment, the synthesis step can be analogous to a standard primer extension reaction that uses nucleotide precursors (including nucleotides as described herein) in the presence of a suitable polymerase to form an extended target strand complementary to the template strand. In other embodiments, the synthesis step itself can form part of an amplification reaction that produces a labeled double-stranded amplification product composed of annealed complementary strands derived from the replication of the target polynucleotide strand and the template polynucleotide strand. Other exemplary synthesis steps include nick translation, strand displacement polymerization, random primed DNA labeling, etc. Particularly useful polymerases for the synthesis step are those that can catalyze the incorporation of nucleotides as set forth herein. A variety of naturally occurring or modified polymerases can be used. By way of example, thermostable polymerases can be used for synthesis reactions carried out using thermal cycling conditions, while thermostable polymerases may not be desirable for isothermal primer extension reactions. Suitable thermostable polymerases capable of incorporating nucleotides according to the present disclosure include those described in WO 2005 / 024010 or WO 06 / 120433, each of which is incorporated herein by reference. In synthesis reactions carried out at lower temperatures such as 37°C, the polymerase need not be a thermostable polymerase, and thus the choice of polymerase will depend on many factors such as reaction temperature, pH, strand displacement activity, etc.
[0217] In specific non-limiting embodiments, the present disclosure encompasses the following methods: nucleic acid sequencing, resequencing, whole genome sequencing, single nucleotide polymorphism scoring, any other application involving the detection of labeled nucleotides or nucleosides as set forth herein when incorporated into a polynucleotide. Any of a variety of other applications that benefit from the use of polynucleotides labeled with nucleotides containing a fluorescent dye can use nucleotides or nucleosides labeled with the dyes set forth herein.
[0218] In one specific embodiment, the present disclosure provides the use of labeled nucleotides according to the present disclosure in a polynucleotide sequencing by synthesis (SBS) reaction. Sequencing by synthesis generally involves the sequential addition of one or more nucleotides or oligonucleotides in the 5' to 3' direction to a growing polynucleotide chain using a polymerase or ligase to form an extended polynucleotide chain complementary to the template nucleic acid to be sequenced. The identity of the base present in one or more of the added nucleotides can be determined in a detection or "imaging" step. The identity of the added base can be determined after each nucleotide incorporation step. The sequence of the template can then be inferred using conventional Watson-Crick base pairing rules. It may be useful to use the labeled nucleotides set forth herein to determine the identity of individual bases, for example, in single nucleotide polymorphism scoring, and such single base extension reactions are within the scope of the present disclosure.
[0219] In one embodiment of the present disclosure, the sequence of the template polynucleotide is determined by detecting the incorporation of one or more 3′-blocked nucleotides into the nascent strand complementary to the template polynucleotide to be sequenced, which detection process is accomplished by detecting a fluorescent label attached to the incorporated nucleotide. Sequencing of the template polynucleotide can be initiated with a suitable primer (or prepared as a hairpin construct, which will include the primer as part of the hairpin), and the nascent strand is extended in a nucleotide-by-nucleotide manner by adding nucleotides to the 3′-end of the primer in a polymerase-catalyzed reaction.
[0220] In a specific embodiment, each of the different nucleotide triphosphates (A, T, G, and C) can be labeled with a unique fluorophore and also contain a blocking group at the 3′-position to prevent uncontrolled polymerization. Alternatively, one of the four nucleotides can be unlabeled (dark). The polymerase incorporates the nucleotide into the nascent strand complementary to the template polynucleotide, and the blocking group prevents further incorporation of the nucleotide. Any unincorporated nucleotides can be washed away, and the fluorescent signal from each incorporated nucleotide can be "read" optically by a suitable device such as a charge-coupled device using laser excitation and a suitable emission filter. The 3′-blocking group and the fluorescent dye compound can then be removed (deprotected) either simultaneously or sequentially to expose the nascent strand for further nucleotide incorporation. Generally, the identity of the incorporated nucleotide will be determined after each incorporation step, but this is not strictly necessary. Similarly, U.S. Patent No. 5,302,509, which is incorporated herein by reference, discloses methods for sequencing polynucleotides immobilized on a solid support.
[0221] As exemplified above, the method utilizes the incorporation of fluorescently labeled 3'-blocked nucleotides A, G, C, and T into a growing strand complementary to a immobilized polynucleotide in the presence of a DNA polymerase. The polymerase incorporates a base complementary to the target polynucleotide, but is prevented from further addition by the 3'-blocking group. The label of the incorporated nucleotide can then be determined, and the blocking group removed by chemical cleavage to allow further polymerization to occur. The nucleic acid template to be sequenced in the sequencing-by-synthesis reaction can be any polynucleotide desired to be sequenced. The nucleic acid template for the sequencing reaction will typically contain a double-stranded region with a free 3'-OH group that serves as a primer or starting point for the addition of additional nucleotides in the sequencing reaction. This region of the template to be sequenced will have the free 3'-OH group dangling on the complementary strand. This dangling region of the template to be sequenced can be single-stranded, but can also be double-stranded provided that there is a "nick" in the strand complementary to the template strand to be sequenced to provide a free 3'-OH group for initiating the sequencing reaction. In such embodiments, sequencing can proceed by strand displacement. In certain embodiments, a primer with a free 3'-OH group can be added as a separate component (e.g., a short oligonucleotide) that hybridizes to the single-stranded region of the template to be sequenced. Alternatively, the primer and template strands to be sequenced can each form part of a partially self-complementary nucleic acid strand capable of forming an intramolecular duplex (such as a hairpin loop structure). Hairpin polynucleotides and methods by which they can be attached to a solid support are disclosed in PCT Publication Nos. WO 01 / 57248 and WO 2005 / 047301, each of which is incorporated herein by reference. Nucleotides can be added sequentially to the growing primer, resulting in the synthesis of a polynucleotide strand in the 5' to 3' direction. The identity of the base that has been added can be determined, particularly but not necessarily after each addition of a nucleotide, to provide sequence information of the nucleic acid template. Thus, the nucleotides are incorporated into the nucleic acid strand (or polynucleotide) in such a way that the nucleotide is joined to the free 3'-OH group of the nucleic acid strand via formation of a phosphodiester bond with the 5’ phosphate group of the nucleotide.
[0222] The nucleic acid template to be sequenced can be DNA or RNA, or even a hybrid molecule composed of deoxynucleotides and ribonucleotides. The nucleic acid template can contain naturally occurring and / or non-naturally occurring nucleotides as well as natural or non-natural backbone linkages, provided that these do not prevent replication of the template in the sequencing reaction.
[0223] In certain embodiments, the nucleic acid template to be sequenced can be attached to a solid support via any suitable ligation method known in the art (e.g., via covalent attachment). In certain embodiments, the template polynucleotide can be attached directly to a solid support (e.g., a silica-based support). However, in other embodiments of the present disclosure, the surface of the solid support can be modified in such a way as to allow direct covalent attachment of the template polynucleotide, or immobilization of the template polynucleotide via a hydrogel or polyelectrolyte multilayer, which itself can be non-covalently attached to the solid support.
[0224] Implementations and alternatives of sequencing by synthesis
[0225] Some embodiments include pyrosequencing techniques. Pyrosequencing detects the release of inorganic pyrophosphate (PPi) when a specific nucleotide is incorporated into a nascent strand (Ronaghi, M., Karamohamed, S., Pettersson, B., Uhlen, M., and Nyren, P. (1996), “Real-time DNA sequencing using detection of pyrophosphate release.”, Analytical Biochemistry 242(1), 84-9; Ronaghi, M. (2001) “Pyrosequencing sheds light on DNA sequencing.” Genome Res., 11(1), 3-11; Ronaghi, M., Uhlen, M., and Nyren, P. (1998) “A sequencing method based on real-time pyrophosphate.” Science 281(5375), 363; U.S. Patent Nos. 6,210,891, 6,258,568, and 6,274,320, the disclosures of which are incorporated herein by reference in their entireties). In pyrosequencing, the released PPi can be detected by its immediate conversion to ATP by ATP sulfurylase, and the level of ATP produced is detected by photons generated by luciferase. The nucleic acid to be sequenced can be attached to features in an array, and the array can be imaged to capture chemiluminescence signals resulting from nucleotide incorporation at the features of the array. Images can be obtained after treating the array with a specific nucleotide type (e.g., A, T, C, or G). The images obtained after adding each nucleotide type will differ in terms of which features in the array are detected. These differences in the images reflect the different sequence content of the features on the array. However, the relative positions of each feature will remain unchanged in the images. The images can be stored, processed, and analyzed using the methods described herein. For example, the images obtained after treating the array with each different nucleotide type can be processed in the same manner as illustrated herein for images obtained from different detection channels for reversible terminator-based sequencing methods.
[0226] In another exemplary type of SBS, cycle sequencing is accomplished by stepwise addition of reversible terminator nucleotides that include, for example, cleavable or photobleachable dye labels as described in, for example, WO 04 / 018497 and U.S. Patent No. 7,057,026, the disclosures of which are incorporated herein by reference. This method was commercialized by Solexa (now Illumina, Inc.) and is also described in WO 91 / 06678 and WO 07 / 123,744, each of which is incorporated herein by reference. The availability of fluorescently labeled terminators (wherein not only can termination be reversed, but also the fluorescent label can be cleaved) facilitates efficient cycle reversible termination (CRT) sequencing. Polymerases can also be co-engineered to efficiently incorporate and extend from these modified nucleotides.
[0227] Preferably, in a reversible terminator-based sequencing embodiment, the label is substantially non-inhibitory to extension under the SBS reaction conditions. However, the detection label can be removable, for example, by cleavage or degradation. An image can be captured after incorporation of the label into the arrayed nucleic acid features. In a particular embodiment, each cycle involves delivering four different nucleotide types simultaneously to the array, and each nucleotide type has a spectrally distinct label. Four images can then be obtained, each using a detection channel selective for one of the four different labels. Alternatively, different nucleotide types can be added sequentially, and an image of the array can be obtained between each addition step. In such embodiments, each image will show the nucleic acid features in which a particular type of nucleotide has been incorporated. Due to the different sequence content of each feature, different features will be present or absent in different images. However, the relative positions of the features will remain invariant in the images. Images obtained by such reversible terminator-SBS methods can be stored, processed, and analyzed as described herein. After the image capture step, the label can be removed, and the reversible terminator moiety can be removed to allow subsequent cycles of nucleotide addition and detection. Removal of these labels after they have been detected in a particular cycle and prior to subsequent cycles can provide the advantage of reducing background signal and crosstalk between cycles. Examples of available labels and removal methods are set forth below.
[0228] Some embodiments can utilize fewer than four different labels to detect four different nucleotides. For example, the methods and systems described in the materials of incorporated U.S. Patent Publication No. 2013 / 0079232 can be used to perform SBS. As a first example, a pair of nucleotide types can be detected at the same wavelength, but distinguished based on the intensity difference of one member of the pair relative to the other member, or based on a change in one member of the pair that results in a distinct signal appearance or disappearance compared to the signal of the other member of the pair detected (e.g., by chemical modification, photochemical modification, or physical modification). As a second example, three of the four different nucleotide types can be detected under specific conditions, while the fourth nucleotide type lacks a label that can be detected or is minimally detected under those conditions (e.g., minimal detection due to background fluorescence, etc.). The incorporation of the first three nucleotide types into the nucleic acid can be determined based on the presence of their respective signals, and the incorporation of the fourth nucleotide type into the nucleic acid can be determined based on the absence of any signal or the minimal detection of any signal. As a third example, one nucleotide type can include a label detected in two different channels, while the other nucleotide types are detected in no more than one channel. The three above-described exemplary configurations are not considered mutually exclusive and can be used in various combinations. An exemplary embodiment combining all three examples is a fluorescence-based SBS method that uses a first nucleotide type detected in a first channel (e.g., dATP having a label detected in the first channel when excited by a first excitation wavelength), a second nucleotide type detected in a second channel (e.g., dCTP having a label detected in the second channel when excited by a second excitation wavelength), a third nucleotide type detected in both the first channel and the second channel (e.g., dTTP having at least one label detected in both channels when excited by the first excitation wavelength and / or the second excitation wavelength), and a fourth nucleotide type lacking a label detected or minimally detected in either channel (e.g., dGTP having no label).
[0229] In addition, as described in the materials of incorporated U.S. Publication 2013 / 0079232, sequencing data can be obtained using a single channel. In such so-called single-dye sequencing methods, the first nucleotide type is labeled, but the label is removed after generating the first image, and only the second nucleotide type is labeled after generating the first image. The third nucleotide type retains its label in both the first image and the second image, and the fourth nucleotide type remains unlabeled in both images.
[0230] Some embodiments may utilize edge - to - edge sequencing techniques. Such techniques utilize DNA ligases to incorporate oligonucleotides and determine the incorporation of such oligonucleotides. Oligonucleotides typically have different labels related to the identity of specific nucleotides in the sequences to which the oligonucleotides hybridize. As with other SBS methods, an image may be obtained after treating an array of nucleic acid features with labeled sequencing reagents. Each image will show the nucleic acid features into which a particular type of label has been incorporated. Due to the different sequence content of each feature, different features will be present or absent in different images, but the relative positions of the features will remain constant in the images. Images obtained by a ligation - based sequencing method may be stored, processed, and analyzed as described herein. Exemplary SBS systems and methods that may be used with the methods and systems described herein are described in U.S. Patent Nos. 6,969,488, 6,172,218, and 6,306,597, the disclosures of which are incorporated herein by reference in their entireties.
[0231] Some embodiments can utilize nanopore sequencing (Deamer, D. W. and Akeson, M. “Nanopores and nucleic acids: prospects for ultrarapid sequencing.” Trends Biotechnol. 18, 147-151 (2000); Deamer, D. and D. Branton, “Characterization of nucleic acids by nanopore analysis”, Acc. Chem. Res. 35:817-825 (2002); Li, J., M. Gershow, D. Stein, E. Brardin and J. A. Golovchenko, “DNA molecules and configurations in a solid-state nanopore microscope”, Nat. Mater., 2:611-615 (2003), the disclosures of which are incorporated herein by reference in their entireties). In such embodiments, the target nucleic acid passes through the nanopore. The nanopore can be a synthetic pore or a biomembrane protein such as α-hemolysin. When the target nucleic acid passes through the nanopore, each base pair can be identified by measuring fluctuations in the electrical conductivity of the pore. (U.S. Patent No. 7,001,792; Soni, G. V. and Meller, “A. Progress toward ultrafast DNA sequencing using solid-state nanopores.” Clin. Chem. 53, 1996-2001 (2007); Healy, K., “Nanopore-based single-molecule DNA analysis.”, Nanomed. 2, 459-481 (2007); Cockroft, S. L., Chu, J., Amorin, M. and Ghadiri, M. R., “A single-molecule nanopore device detects DNA polymerase activity with single-nucleotide resolution.”, J. Am. Chem. Soc. 130, 818-820 (2008), the disclosures of which are incorporated herein by reference in their entireties). Data obtained from nanopore sequencing can be stored, processed, and analyzed as described herein. Specifically, based on the exemplary processing of the optical images and other images described herein, the data can be processed as if it were an image.
[0232] Some other embodiments of the sequencing method involve the use of the 3′-blocked nucleotides described herein in nanoparticle sequencing techniques, such as those described in U.S. Patent No. 9,222,132, the disclosure of which is incorporated herein by reference. Through a rolling circle amplification (RCA) process, a large number of discrete DNA nanoparticles can be generated. The nanoparticle mixture is then distributed onto the surface of a patterned slide containing features that allow individual nanoparticles to be associated with each position. During the DNA nanoparticle generation process, the DNA is fragmented and then ligated to the first of four adaptor sequences. The template is amplified, circularized, and then cleaved with a type II endonuclease. A second set of adaptors is added, followed by amplification, circularization, and cleavage. The process is repeated for the remaining two adaptors. The final product is a circular template with four adaptors, each separated by a template sequence. Library molecules undergo a rolling circle amplification step, generating a large number of concatemers called DNA nanoparticles, which are then deposited on a flow cell. Goodwin et al., “Coming of age: ten years of next-generation sequencing technologies,” Nat Rev Genet. 2016;17(6):333-51.
[0233] Some embodiments may utilize methods involving real-time monitoring of DNA polymerase activity. Nucleotide incorporation can be detected by fluorescence resonance energy transfer (FRET) interactions between a polymerase carrying a fluorophore and a γ-phosphate-labeled nucleotide (as described, for example, in U.S. Patent Nos. 7,329,492 and 7,211,414, both of which are incorporated herein by reference), or nucleotide incorporation can be detected using zero-mode waveguides (as described, for example, in U.S. Patent No. 7,315,019, which is incorporated herein by reference), and nucleotide incorporation can be detected using fluorescent nucleotide analogs and engineered polymerases (as described, for example, in U.S. Patent Nos. 7,405,281 and U.S. Patent Publication No. 2008 / 0108082, both of which are incorporated herein by reference). Illumination can be limited to a zeptoliter-scale volume surrounding surface-tethered polymerase, such that incorporation of fluorescently labeled nucleotides can be observed at low background (Levene, M. J. et al., “Zero-mode waveguides for single-molecule analysis at high concentrations.”, Science 299, 682-686 (2003); Lundquist, P. M. et al., “Parallel confocal detection of single molecules in real time.”, Opt. Lett. 33, 1026-1028 (2008); Korlach, J. et al., “Selective aluminum passivation for targeted immobilization of single DNA polymerase molecules in zero-mode waveguide nano structures.”, Proc. Natl. Acad. Sci. USA 105, 1176-1181 (2008), the disclosures of which are incorporated herein in their entireties by reference). Images obtained by such methods can be stored, processed, and analyzed as described herein.
[0234] Some SBS embodiments include detecting protons released upon nucleotide incorporation into an extension product. For example, sequencing based on the detection of released protons can use an electrical detector and related technologies commercially available from Ion Torrent (a subsidiary of Life Technologies, Guilford, CT), or the sequencing methods and systems described in U.S. Patent Publication Nos. 2009 / 0026082, 2009 / 0127589, 2010 / 0137143, and 2010 / 0282617, all of which are incorporated herein by reference. The methods described herein for amplifying a target nucleic acid using kinetic exclusion can be readily applied to a substrate for detecting protons. More specifically, the methods described herein can be used to generate an amplified clone population for detecting protons.
[0235] The SBS methods described above can advantageously be performed in a variety of formats such that multiple different target nucleic acids are manipulated simultaneously. In a specific embodiment, the different target nucleic acids can be processed in a common reaction vessel or on the surface of a particular substrate. This allows for convenient delivery of sequencing reagents, removal of unreacted reagents, and detection of incorporation events in a variety of ways. In embodiments using surface-bound target nucleic acids, the target nucleic acids can be in an array format. In an array format, the target nucleic acids are typically capable of binding to the surface in a spatially distinguishable manner. The target nucleic acids can bind by direct covalent attachment, attachment to beads or other particles, or binding to a polymerase or other molecule attached to the surface. The array can include a single copy of the target nucleic acid at each site (also referred to as a feature), or multiple copies having the same sequence can be present at each site or feature. The multiple copies can be generated by amplification methods such as bridge amplification or emulsion PCR, as described in further detail below.
[0236] The methods described herein can use an array having features at any of a variety of densities, including, for example: at least about 10 features / cm 2 、100 features / cm 2 、500 features / cm 2 、1,000 features / cm 2 、5,000 features / cm 2 、10,000 features / cm 2 、50,000 features / cm 2 、100,000 features / cm 2 、1,000,000 features / cm 2 、5,000,000 features / cm 2 , or higher densities.
[0237] Advantages of the methods described herein are that they provide rapid and efficient detection of multiple target nucleic acids in parallel. Accordingly, the present disclosure provides an integrated system capable of preparing and detecting nucleic acids using techniques known in the art (such as those exemplified above). Thus, the integrated system of the present disclosure can include fluidic components capable of delivering amplification reagents and / or sequencing reagents to one or more immobilized DNA fragments, the system including components such as pumps, valves, reservoirs, fluid lines, and the like. A flow cell in the integrated system can be configured for and / or for detecting target nucleic acids. Exemplary flow cells are described, for example, in U.S. Patent Publication No. 2010 / 0111768 and U.S. Patent Application No. 13 / 273,666, each of which is incorporated herein by reference. As exemplified for the flow cell, one or more fluidic components of the integrated system can be used for amplification methods and detection methods. By way of example of a nucleic acid sequencing embodiment, one or more fluidic components of the integrated system can be used for the amplification methods described herein and for delivering sequencing reagents in sequencing methods (such as those exemplified above). Alternatively, the integrated system can include separate fluidic systems to perform the amplification method and to perform the detection method. Examples of integrated sequencing systems capable of generating amplified nucleic acids and also capable of determining nucleic acid sequences include, but are not limited to, the MiSeq TM platform (Illumina, Inc., San Diego, CA), and the device described in U.S. Patent Application No. 13 / 273,666, which is incorporated herein by reference.
[0238] Arrays in which polynucleotides have been directly attached to silica-based supports are, for example, those disclosed in WO 00 / 06770 (incorporated herein by reference), in which the polynucleotides are immobilized on the glass support by reaction between epoxy side groups on the glass and internal amino groups on the polynucleotides. In addition, polynucleotides can be attached to solid supports by reaction of thiol nucleophiles with the solid support, for example, as described in WO 2005 / 047301 (incorporated herein by reference). Another example of a solid-supported template polynucleotide is one in which the template polynucleotide is attached to a hydrogel carried on a silica-based solid support or other solid support, for example, as described in WO 00 / 31148, WO 01 / 01143, WO02 / 12566, WO 03 / 014392, U.S. Patent No. 6,465,178, and WO 00 / 53812, each of which is incorporated herein by reference.
[0239] A specific surface to which the template polynucleotide can be immobilized is a polyacrylamide hydrogel. Polyacrylamide hydrogels are described in the references cited above and in WO 2005 / 065814, which patent is incorporated herein by reference. Specific hydrogels that can be used include those described in WO 2005 / 065814 and U.S. Patent Publication No. 2014 / 0079923. In one embodiment, the hydrogel is PAZAM (poly(N-(5-azidoacetamidopentyl)acrylamide-co-acrylamide)).
[0240] The DNA template molecule can be attached to beads or microparticles, for example, as described in U.S. Patent No. 6,172,218 (incorporated herein by reference). Attachment to beads or microparticles can be used in sequencing applications. A bead library can be prepared, where each bead contains a different DNA sequence. Exemplary libraries and methods for their generation are described in: Nature, 437, 376-380 (2005); Science, 309, 5741, 1728-1732 (2005), each of which is incorporated herein by reference. Sequencing of such arrays of beads using the nucleotides shown herein is within the scope of the present disclosure.
[0241] The template to be sequenced can form part of an "array" on a solid support, in which case the array can take any convenient form. Thus, the methods of the present disclosure are applicable to all types of high-density arrays, including single-molecule arrays, cluster arrays, and bead arrays. The labeled nucleotides of the present disclosure can be used to sequence templates on substantially any type of array (including but not limited to those formed by immobilizing nucleic acid molecules on a solid support).
[0242] However, the labeled nucleotides of the present disclosure are particularly advantageous in the context of sequencing clustered arrays. In a clustered array, different regions on the array (commonly referred to as sites or features) contain multiple polynucleotide template molecules. Generally speaking, the multiple polynucleotide molecules cannot be individually resolved by optical means but are detected as a whole. Depending on how the array is formed, each site on the array can contain multiple copies of a single polynucleotide molecule (e.g., the site is homogeneous for a particular single-stranded nucleic acid species or double-stranded nucleic acid species) or even multiple copies of a small number of different polynucleotide molecules (e.g., multiple copies of two different nucleic acid species). Clustered arrays of nucleic acid molecules can be produced using techniques well known in the art. By way of example, WO 98 / 44151 and WO 00 / 18957 (each incorporated herein) describe methods for amplifying nucleic acids, where both the template and the amplification product remain immobilized on a solid support to form an array consisting of clusters or "colonies" of immobilized nucleic acid molecules. The nucleic acid molecules present on the clustered arrays prepared according to these methods are suitable templates for sequencing using nucleotides labeled with the dye compounds of the present disclosure.
[0243] The labeled nucleotides of the present disclosure can also be used to sequence templates on single molecule arrays. As used herein, the term "single molecule array" or "SMA" refers to a population of polynucleotide molecules distributed (or arranged) on a solid support, where the spacing of any individual polynucleotide from all other polynucleotides in the population is such that it is possible to individually resolve the individual polynucleotide molecules. Thus, in some embodiments, the target nucleic acid molecules immobilized on the surface of the solid support can be resolved by optical means. This means that one or more different signals (each signal representing a polynucleotide) will appear within a resolvable region of the particular imaging device used.
[0244] Single molecule detection can be achieved where the spacing between adjacent polynucleotide molecules on the array is at least 100 nm, more specifically at least 250 nm, still more specifically at least 300 nm, and even more specifically at least 350 nm. Thus, each molecule can be individually resolved and detected as a single molecule fluorescence spot, and the fluorescence from the single molecule fluorescence spot also exhibits single-step photobleaching.
[0245] The terms "individually resolvable" and "individually resolved" are used herein to stipulate that when visualized, it is possible to distinguish one molecule on the array from its neighboring molecules. The spacing between the individual molecules on the array will be determined in part by the particular technique used to resolve the individual molecules. The general features of single-molecule arrays will be understood by reference to the published patent applications WO 00 / 06770 and WO 01 / 57248, each of which is incorporated herein by reference. While one use of the nucleotides of the present disclosure is for sequencing-by-synthesis reactions, the utility of these nucleotides is not limited to such methods. In fact, these nucleotides can advantageously be used in any sequencing method that requires detection of a fluorescent label attached to a nucleotide incorporated into a polynucleotide.
[0246] In particular, the labeled nucleotides of the present disclosure can be used in automated fluorescent sequencing protocols, especially fluorescence dye-terminator cycle sequencing based on the chain-termination sequencing method of Sanger and colleagues. Such methods typically use enzymes and cycle sequencing to incorporate fluorescently labeled dideoxynucleotides into primer extension sequencing reactions. The so-called Sanger sequencing method and related protocols (Sanger-type) utilize randomized chain termination with labeled dideoxynucleotides.
[0247] Accordingly, the present disclosure also encompasses labeled nucleotides that are dideoxynucleotides lacking a hydroxyl group at both the 3'-position and the 2'-position, such dideoxynucleotides being suitable for use in Sanger-type sequencing methods and the like.
[0248] It should be understood that the labeled nucleotides of the present disclosure incorporating a 3'-blocking group can also be used in Sanger methods and related protocols, since the same effect can be achieved by using nucleotides having a 3'-OH blocking group as by using dideoxynucleotides: both prevent the incorporation of subsequent nucleotides. In the case where nucleotides according to the present disclosure and having a 3'-blocking group are to be used in Sanger-type sequencing methods, it should be understood that the dye compound or detectable label attached to the nucleotide need not be linked via a cleavable linker, since in each case where a labeled nucleotide of the present disclosure is incorporated; subsequent nucleotide incorporation is not required, and thus removal of the label from the nucleotide is not required.
[0249] In any embodiment of the SBS methods described herein, the nucleotides used in sequencing applications are the 3'-blocked nucleotides described herein, e.g., the nucleotides of formula (I) and (Ia) to (Id). In any embodiment, the 3'-blocked nucleotides are nucleotide triphosphates.
[0250] In some sequencing methods, the nucleotides incorporated are unlabeled. After incorporation, one or more fluorescent labels can be introduced by using an affinity reagent containing a label of one or more fluorescent dyes. For example, one, two, three or each of the four different types of nucleotides (e.g., dATP, dCTP, dGTP and dTTP or dUTP) incorporated into the buffer of step (a) can be unlabeled. Each of these four types of nucleotides (e.g., dNTP) has a 3'-OH blocking group (e.g., 3'-AOM) as described herein to ensure that only a single base can be added to the 3' end of the copy polynucleotide by a polymerase. After incorporation of unlabeled nucleotides, an affinity reagent specifically bound to the dNTP incorporated is introduced to provide an extension product containing the label of the dNTP incorporated. U.S. Patent Publication No. 2013 / 0079232 discloses the use of unlabeled nucleotides and affinity reagents in sequencing by synthesis. The improved sequencing method using unlabeled nucleotides disclosed in the present invention may include the following steps:
[0251] (a'-1) will contain the 3'-OH end-capping group described herein incorporating an unlabeled nucleotide (e.g., dATP, dCTP, dGTP, dTTP, or dUTP) (attached to the 3′ oxygen group) into a copy polynucleotide chain complementary to at least a portion of the target polynucleotide chain to produce an extended copy polynucleotide;
[0252] (a'-2) contacting the extended copy polynucleotide with a set of affinity reagents under conditions where one of the affinity reagents specifically binds to the incorporated unlabeled nucleotide to provide a labeled extended copy polynucleotide;
[0253] (b') detecting the identity of the nucleotide incorporated into the copy polynucleotide chain by performing one or more fluorescence measurements of the labeled extended copy polynucleotide; and
[0254] (c') chemically removing the detectable label from the extended copy polynucleotide and removing the 3'-OH blocking group from the nucleotides incorporated into the copy polynucleotide chain.
[0255] Affinity reagents can include: small molecule or protein tags that can bind to the hapten moiety of a nucleotide (such as streptavidin-biotin, anti-DIG and DIG, anti-DNP and DNP), antibodies (including but not limited to binding fragments of antibodies, single-chain antibodies, bispecific antibodies, etc.), aptamers, knottins, Affimer proteins, or any other known reagent that binds to the incorporated nucleotide with suitable specificity and affinity. In additional embodiments, an affinity reagent can be labeled with multiple copies of the same fluorescent dye. In some embodiments, the Pd catalyst also removes the labeled affinity reagent. For example, the hapten moiety of an unlabeled nucleotide can be attached to the nucleobase via a cleavable linker (e.g., AOL linker) as described herein, and the cleavable linker can be cleaved by the Pd catalyst. In some embodiments, the method further includes the post-cleavage wash step (d) described herein. In some embodiments, the method further includes repeating steps (a’-1) to (c’) or (a’-1) to (d) until the sequence of at least a portion of the target polynucleotide chain is determined. In some embodiments, the cycle is repeated at least 50 times, at least 100 times, at least 150 times, at least 200 times, at least 250 times, or at least 300 times.
[0256] Kits
[0257] The present disclosure also provides a kit that includes one or more of the 3′-blocked nucleosides and / or nucleotides described herein, e.g., the 3′-blocked nucleotides of formulas (I) and (Ia) to (Id). Such kits will generally include at least one 3′-blocked nucleotide or nucleoside having a detectable label (e.g., a fluorescent dye), and at least one additional component. The additional component can be one or more of the components identified in the methods shown herein or in the Examples section below. Some non-limiting examples of components that can be incorporated into the kits of the present disclosure are shown below. In some additional embodiments, the kit can include four types of labeled nucleotides, i.e., the fully functionalized nucleotides (A, C, T, and G) described herein, where each type of nucleotide contains the 3′-AOM blocking group and the AOL linker moiety described herein. In additional embodiments, G is unlabeled and does not contain an AOL linker. In further embodiments, one or more of the remaining three nucleotides (i.e., A, C, and T) contain L 1 which is an allylamine or allylamide linker moiety. In one embodiment, the kit includes the unlabeled ffG, labeled ffA, labeled ffC, and labeled ffT-DB described herein. In another embodiment, the kit includes the unlabeled ffG, labeled ffA, labeled ffC-DB, and labeled ffT-DB described herein.
[0258] In one specific embodiment, the kit can include at least one labeled 3′-blocked nucleotide or nucleoside, and labeled or unlabeled nucleotides or nucleosides. For example, nucleotides labeled with a dye can be provided in combination with unlabeled or native nucleotides and / or with fluorescently labeled nucleotides or any combination thereof. The combination of nucleotides can be provided as separate individual components (e.g., one nucleotide type per container or tube) or as a nucleotide mixture (e.g., two or more nucleotides mixed in the same container or tube).
[0259] In cases where the kit includes multiple, particularly two or three, or more particularly four 3′-blocked nucleotides labeled with dye compounds, the different nucleotides can be labeled with different dye compounds, or one can be dark and unlabeled with a dye compound. In cases where the different nucleotides are labeled with different dye compounds, a feature of the kit is that the dye compounds are spectrally distinguishable fluorescent dyes. As used herein, the term "spectrally distinguishable fluorescent dyes" refers to fluorescent dyes that emit fluorescent energy at wavelengths that can be distinguished by a fluorescence detection device (e.g., a commercial capillary-based DNA sequencing platform) when two or more such dyes are present in a sample. When two nucleotides labeled with fluorescent dye compounds are provided in kit form, some embodiments are characterized in that the spectrally distinguishable fluorescent dyes are capable of being excited at the same wavelength, such as by the same laser. When four 3′-blocked nucleotides (A, C, T, and G) labeled with fluorescent dye compounds are provided in kit form, some embodiments are characterized in that two of the spectrally distinguishable fluorescent dyes are each capable of being excited at one wavelength, while the other two spectrally distinguishable dyes are each capable of being excited at another wavelength. Specific excitation wavelengths are 488 nm and 532 nm.
[0260] In one embodiment, the kit includes a first 3′-blocked nucleotide labeled with a first dye, and a second nucleotide labeled with a second dye, wherein the dyes have a maximum absorbance difference of at least 10 nm, particularly 20 nm to 50 nm. More specifically, the two dye compounds have a Stokes shift between 15 nm and 40 nm, where "Stokes shift" is the distance between the peak absorption wavelength and the peak emission wavelength.
[0261] In an alternative embodiment, the kits of the present disclosure can include 3′-blocked nucleotides in which the same base is labeled with two or more different dyes. A first nucleotide (e.g., a 3′-blocked T nucleotide triphosphate or a 3′-blocked G nucleotide triphosphate) can be labeled with a first dye. A second nucleotide (e.g., a 3′-blocked C nucleotide triphosphate) can be labeled with a second dye that is spectrally different from the first dye, such as a "green" dye that absorbs light at a wavelength less than 600 nm and a "blue" dye that absorbs light at a wavelength less than 500 nm, such as wavelengths from 400 nm to 500 nm, particularly 450 nm to 460 nm. A third nucleotide (e.g., a 3′-blocked A nucleotide triphosphate) can be labeled as a mixture of the first dye and the second dye, or a mixture of the first dye, the second dye, and a third dye, and a fourth nucleotide (e.g., a 3′-blocked G nucleotide triphosphate or a 3′-blocked T nucleotide triphosphate) can be "dark" and not contain a label. In one example, nucleotides 1 to 4 can be labeled "blue", "green", "blue / green", and dark. To further simplify the instrument, the four nucleotides can be labeled with two dyes excited by a single laser, so the labels for nucleotides 1 to 4 can be "blue 1", "blue 2", "blue 1 / blue 2", and dark.
[0262] In a specific embodiment, the kit can include four labeled 3′-blocked nucleotides (e.g., A, C, T, G), where each type of nucleotide contains the same 3′-blocking group and a fluorescent label, and where each fluorescent label has a different maximum fluorescence value and each fluorescent label is capable of being distinguished from the other three labels. The kit can be such that two or more of these fluorescent labels have similar maximum absorbances but different Stokes shifts. In some other embodiments, one type of nucleotide is unlabeled.
[0263] Although the present disclosure illustrates the kit by way of example with configurations of different nucleotides labeled with different dye compounds, it should be understood that the kit can include two, three, four, or more different nucleotides with the same dye compound. In some embodiments, the kit further includes an enzyme and a buffer suitable for the enzyme to function. In some such embodiments, the enzyme is a polymerase, terminal deoxynucleotidyl transferase, or reverse transcriptase. In a specific embodiment, the enzyme is a DNA polymerase, such as DNA polymerase 812 (Pol 812) or DNA polymerase 1901 (Pol 1901). In still other embodiments, the kit can include the incorporation mixture described herein. In additional embodiments, the kit containing the incorporation mixture described herein further includes at least one Pd scavenger (e.g., the Pd(0) scavenger comprising one or more allyl moieties described herein). The Pd(0) scavenger includes one or more allyl moieties each independently selected from the group consisting of: –O-allyl, –S-allyl, –NR-allyl, and –N + RR′-allyl, wherein R is H, unsubstituted or substituted C1-C6 alkyl, unsubstituted or substituted C2-C6 alkenyl, unsubstituted or substituted C2-C6 alkynyl, unsubstituted or substituted C6-C 10 aryl, unsubstituted or substituted 5- to 10-membered heteroaryl, unsubstituted or substituted C3-C 10 carbocyclic group, or unsubstituted or substituted 5- to 10-membered heterocyclic group; and R′ is H, unsubstituted C1-C6 alkyl, or substituted C1-C6 alkyl. In some such embodiments, the Pd(0) scavenger in the incorporation solution comprises one or more –O-allyl moieties. In still other embodiments, the Pd(0) scavenger includes or is
[0264] or a combination thereof. Alternative Pd(0) scavengers are disclosed in U.S. Serial No. 63 / 190983, which is incorporated herein by reference in its entirety. In one embodiment, the Pd(0) scavenger in the incorporation mixture includes or is In another embodiment, the Pd(0) scavenger in the incorporation mixture includes or is
[0265] Other components to be included in such kits can include buffers and the like. The nucleotides of the present disclosure, as well as any other nucleotide components, including mixtures of different nucleotides, can be provided in the kit in a concentrated form and diluted prior to use. In such embodiments, a suitable dilution buffer can also be included. For example, the incorporation mixture kit can include one or more buffering agents selected from the group consisting of primary amines, secondary amines, tertiary amines, natural or unnatural amino acids, or combinations thereof. In additional embodiments, the buffering agent in the incorporation mixture includes ethanolamine or glycine, or combinations thereof.
[0266] Similarly, one or more of the components identified in the methods shown herein can be included in the kits of the present disclosure. In still other embodiments, the kit can include the palladium catalysts described herein. In some embodiments, the Pd catalyst is generated by mixing a Pd(II) complex (i.e., a Pd precatalyst) with one or more of the water-soluble phosphines described herein. In some such embodiments, the kit containing the Pd catalyst is a lysis mixture kit. In additional embodiments, the lysis mixture kit can contain [Pd(allyl)Cl]2 or Na2PdCl4 and the water-soluble phosphine THP for generating the active Pd(0) species. The molar ratio of the Pd(II) complex (e.g., [Pd(allyl)Cl]2 or Na2PdCl4) to the water-soluble phosphine (e.g., THP) can be about 1:2, 1:3, 1:4, 1:5, 1:6, 1:7, 1:8, 1:9, or 1:10. In additional embodiments, the lysis mixture kit can further contain one or more buffer reagents selected from the group consisting of primary amines, secondary amines, tertiary amines, carbonates, phosphates, borates, and combinations thereof. Non-limiting examples of buffer reagents in the lysis mixture kit are selected from the group consisting of ethanolamine (EA), tris(hydroxymethyl)aminomethane (Tris), glycine, carbonates, phosphates, borates, dimethylethanolamine (DMEA), diethylethanolamine (DEEA), N,N,N′,N′-tetramethylethylenediamine (TEMED), N,N,N′,N′-tetraethylethylenediamine (TEEDA), 2-piperidinoethanol, and combinations thereof. In one embodiment, the lysis mixture kit contains DEEA. In other embodiments, the lysis mixture kit contains 2-piperidinoethanol.
[0267] In some other embodiments, the kit may include one or more palladium scavengers (e.g., the Pd(II) scavengers described herein). In some such embodiments, the kit is a post-lysis wash buffer kit. Non-limiting examples of Pd scavengers in a post-lysis wash buffer kit include isocyanoacetate (ICNA) salts, ethyl isocyanoacetate, methyl isocyanoacetate, cysteine or its salts, L-cysteine or its salts, N-acetyl-L-cysteine, potassium ethyl xanthate, potassium isopropyl xanthate, glutathione, lipoic acid, ethylenediaminetetraacetic acid (EDTA), iminodiacetic acid, nitrilodiacetic acid, trithiolo-S-triazine, dimethyldithiocarbamate, dithiothreitol, mercaptoethanol, allyl alcohol, propargyl alcohol, thiols, tertiary amines, and / or tertiary phosphines, or combinations thereof. In one embodiment, the post-lysis wash buffer kit includes L-cysteine or its salts.
[0268] In any embodiment of the kits described herein, the Pd scavenger (e.g., the Pd(0) or Pd(II) scavenger described herein) is in a container / compartment separate from the Pd catalyst.
[0269] This application contains the following items:
[0270] 1. A nucleoside or nucleotide comprising a nucleobase attached via a cleavable linker to a detectable label, wherein the nucleoside or nucleotide comprises a ribose or 2'-deoxyribose moiety and a 3'-OH capping group, and wherein the cleavable linker comprises a moiety having the following structure:
[0271]
[0272] wherein
[0273] X and Y are each independently O or S; and
[0274] R 1a 、R 1b 、R 2 、R 3a and R 3b are each independently H, halogen, unsubstituted or substituted C1-C6 alkyl, or C1-C6 haloalkyl.
[0275] 2. The nucleoside or nucleotide according to item 1, comprising a structure of formula (I):
[0276]
[0277] wherein B is a nucleobase;
[0278] R 4 is H or OH;
[0279] R5 is a 3′-OH capping group;
[0280] R 6 is H, monophosphate, diphosphate, triphosphate, thiophosphate, a phosphate ester analogue, a reactive phosphorus-containing group, or a hydroxyl protecting group;
[0281] The detectable label is a fluorescent dye;
[0282] L is and
[0283] L 1 and L 2 are each independently an optionally present linker moiety.
[0284] 3. The nucleoside or nucleotide according to item 1 or 2, wherein X and Y are each O.
[0285] 4. The nucleoside or nucleotide according to any one of items 1 to 3, wherein R 1a 、R 1b 、R 2 、R 3a and R 3b are each H.
[0286] 5. The nucleoside or nucleotide according to any one of items 1 to 3, wherein at least one of R 1a 、R 1b 、R 2 、R 3a and R 3b is halogen or unsubstituted C1-C6 alkyl.
[0287] 6. The nucleoside or nucleotide according to item 5, wherein R 1a and R 1b are each H, and
[0288] R 2 、R 3a and R 3b is unsubstituted C1-C6 alkyl.
[0289] 7. The nucleoside or nucleotide according to any one of items 2 to 6, wherein B is purine, deazapurine or pyrimidine.
[0290] 8. The nucleoside or nucleotide according to any one of items 2 to 7, wherein R 5 is
[0291] and wherein R a 、R b 、R c 、R d and Re Each independently is H, a halogen, an unsubstituted or substituted C1-C6 alkyl group, or a C1-C6 haloalkyl group.
[0292] 9. The nucleoside or nucleotide according to item 8, wherein R 5 is
[0293] 10. The nucleoside or nucleotide according to any one of items 2 to 9, wherein L 1 is present, and
[0294] L 1 comprises a moiety selected from the group consisting of propargylamine, propargylamide, allylamine, allylamide, and optionally substituted variants of these substances.
[0295] 11. The nucleoside or nucleotide according to item 10, wherein L 1 comprises
[0296] 12. The nucleoside or nucleotide according to item 11, which comprises a structure of formula (Ia), (Ia′), (Ib), (Ic),
[0297] (Ic′) or (Id):
[0298]
[0299]
[0300] 13. The nucleoside or nucleotide according to any one of items 2 to 12, wherein L 2 is present, and
[0301] L 2 comprises wherein m and n are each independently an integer 1,
[0302] 2, 3, 4, 5, 6, 7, 8, 9 or 10, and the phenyl moiety is optionally substituted.
[0303] 14. The nucleoside or nucleotide according to item 13, wherein n is 5.
[0304] 15. The nucleoside or nucleotide according to item 13 or 14, wherein m is 4.
[0305] 16. The nucleoside or nucleotide according to any one of items 1 to 14, wherein the nucleotide is a nucleotide triphosphate comprising a 2′-deoxyribose moiety.
[0306] 17. An oligonucleotide or polynucleotide comprising the nucleotide according to any one of items 1 to 16.
[0307] 18. The oligonucleotide or polynucleotide according to item 17, wherein the oligonucleotide or polynucleotide hybridizes with a template polynucleotide.
[0308] 19. The oligonucleotide or polynucleotide according to item 18, wherein the template polynucleotide is immobilized on a solid support.
[0309] 20. The oligonucleotide or polynucleotide according to item 19, wherein the solid support comprises an array of immobilized template polynucleotides.
[0310] 21. A method for preparing a growing polynucleotide complementary to a target single-stranded polynucleotide in a sequencing reaction, comprising incorporating the nucleotide according to any one of items 1 to 16 into the growing complementary polynucleotide, wherein incorporating the nucleotide prevents any subsequent nucleotide from being introduced into the growing complementary polynucleotide.
[0311] 22. The method according to item 21, wherein incorporating the nucleotide is accomplished by a polymerase, terminal deoxynucleotidyl transferase, or reverse transcriptase.
[0312] 23. A method for determining the sequence of a target single-stranded polynucleotide, comprising:
[0313] (a) Incorporating the nucleotide according to any one of items 1 to 16 into a copy polynucleotide chain complementary to at least a portion of the target polynucleotide chain;
[0314] (b) Detecting the identity of the nucleotide incorporated into the copy polynucleotide chain; and
[0315] (c) Chemically removing a detectable label and a 3′-OH capping group from the nucleotide incorporated into the copy polynucleotide chain.
[0316] 24. The method according to item 23, further comprising (d) washing the chemically removed label and the 3′-OH capping group from the copy polynucleotide chain using a post-cleavage wash solution.
[0317] 25. The method according to item 24, further comprising repeating steps (a) to (d) until the sequence of at least a portion of the target polynucleotide chain is determined.
[0318] 26. The method according to item 25, wherein steps (a) to (d) are repeated at least 50 times, at least 100 times, or at least 150 times.
[0319] 27. The method according to any one of items 23 to 26, wherein step (c) comprises contacting the incorporated nucleotide with a cleavage solution comprising a palladium catalyst.
[0320] 28. The method according to item 27, wherein the palladium catalyst is a palladium(0) catalyst in-situ generated from a palladium complex and a water-soluble phosphine.
[0321] 29. The method according to item 28, wherein the palladium complex comprises [Pd(allyl)Cl]2, Na2PdCl4, [Pd(allyl)(THP)]Cl, [Pd(allyl)(THP)2]Cl, Pd(CH3CN)2Cl2, Pd(OAc)2, Pd(PPh3)4, Pd(dba)2, Pd(Acac)2, PdCl2(COD) or Pd(TFA)2, or a combination thereof.
[0322] 30. The method according to item 29, wherein the palladium complex comprises [Pd(allyl)Cl]2 or Na2PdCl4.
[0323] 31. The method according to any one of items 27 to 30, wherein the detectable label and the 3′-OH capping group from the nucleotide incorporated into the copy polynucleotide chain are removed in a single chemical reaction.
[0324] 32. The method according to any one of items 27 to 31, wherein the cleavage solution further comprises one or more buffer reagents selected from the group consisting of primary amines, secondary amines, tertiary amines, carbonates, phosphates, borates, and combinations thereof.
[0325] 33. The method according to item 32, wherein the buffer reagent in the cleavage solution is selected from the group consisting of ethanolamine (EA), tris(hydroxymethyl)aminomethane (Tris), glycine, carbonates, phosphates, borates, dimethylethanolamine (DMEA), diethylethanolamine (DEEA), N,N,N′,N′-tetramethylethylenediamine (TEMED), N,N,N′,N′-tetraethylethylenediamine (TEEDA), 2-piperidinoethanol, and combinations thereof.
[0326] 34. The method according to any one of items 27 to 33, wherein step (a) comprises contacting the nucleotide with the copy polynucleotide chain in an incorporation solution comprising a polymerase, at least one palladium scavenger, and one or more buffers.
[0327] 35. The method according to item 34, wherein the buffer in the incorporation solution comprises primary amines, secondary amines, tertiary amines, natural amino acids or unnatural amino acids, or combinations thereof.
[0328] 36. The method according to item 35, wherein the buffer in the incorporation solution comprises ethanolamine or glycine, or a combination thereof.
[0329] 37. The method according to any one of items 34 to 36, wherein the palladium scavenger in the incorporation solution comprises one or more allyl moieties independently selected from the group consisting of: –O-allyl, –S-allyl, –NR-allyl, and –N + RR′-allyl,
[0330] wherein R is H, unsubstituted or substituted C1-C6 alkyl, unsubstituted or substituted C2-C6 alkenyl, unsubstituted or substituted C2-C6 alkynyl, unsubstituted or substituted C6-C 10 aryl, unsubstituted or substituted 5- to 10-membered heteroaryl, unsubstituted or substituted C3-C 10 carbocyclic group, or unsubstituted or substituted 5- to 10-membered heterocyclic group; and
[0331] R′ is H, unsubstituted C1-C6 alkyl or substituted C1-C6 alkyl.
[0332] 38. The method according to item 37, wherein the palladium scavenger in the incorporation solution is:
[0333]
[0334] 39. The method according to any one of items 27 to 38, wherein the post-cleavage wash solution comprises one or more palladium scavengers.
[0335] 40. The method according to item 39, wherein the one or more palladium scavengers in the post-cleavage solution include isocyanoacetate (ICNA) salts, ethyl isocyanoacetate, methyl isocyanoacetate, cysteine or its salts, L-cysteine or its salts, N-acetyl-L-cysteine, potassium ethylxanthate, potassium isopropylxanthate, glutathione, lipoic acid, ethylenediaminetetraacetic acid (EDTA), iminodiacetic acid, nitrilodiacetic acid, trithiols-S-triazine, dimethyldithiocarbamate, dithiothreitol, mercaptoethanol, allyl alcohol, propargyl alcohol, thiols, tertiary amines, and / or tertiary phosphines, or combinations thereof.
[0336] 41. The method according to any one of items 23 to 40, wherein the target single-stranded polynucleotide is formed by chemically cleaving the complementary strand from a double-stranded polynucleotide.
[0337] 42. The method according to item 41, wherein the chemical cleavage of the complementary strand is carried out under the same reaction conditions as those for chemically removing the detectable label and the 3′-OH capping group from the nucleotide incorporated into the copy polynucleotide strand.
[0338] 43. A kit comprising one or more nucleosides or nucleotides according to any one of items 1 to 16.
[0339] 44. The kit according to item 43, further comprising an enzyme, at least one Pd(0) scavenger, and one or more buffers.
[0340] 45. The kit according to item 44, wherein the enzyme is a DNA polymerase, a terminal deoxynucleotidyl transferase, or a reverse transcriptase.
[0341] 46. The kit according to item 44, wherein the Pd(0) scavenger comprises one or more allyl moieties each independently selected from the group consisting of: –O-allyl, –S-allyl, –NR-allyl, and –N + RR′-allyl,
[0342] wherein R is H, unsubstituted or substituted C1-C6 alkyl, unsubstituted or substituted C2-C6 alkenyl, unsubstituted or substituted C2-C6 alkynyl, unsubstituted or substituted C6-C 10 aryl, unsubstituted or substituted 5- to 10-membered heteroaryl, unsubstituted or substituted C3-C 10 carbocyclic group, or unsubstituted or substituted 5- to 10-membered heterocyclic group; and
[0343] R′ is H, unsubstituted C1-C6 alkyl, or substituted C1-C6 alkyl.
[0344] 47. The kit according to item 46, wherein the palladium scavenger is:
[0345]
[0346] 48. The kit according to any one of items 43 to 47, further comprising a palladium catalyst.
[0347] 49. The kit according to item 48, wherein the palladium catalyst is a Pd(0) catalyst in situ generated from a Pd(II) complex and one or more water-soluble phosphines.
[0348] 50. The kit according to item 49, wherein the Pd(II) complex is [Pd(allyl)Cl]2 or Na2PdCl4.
[0349] 51. The kit according to item 49 or 50, further comprising one or more Pd(II) scavengers, wherein the Pd(II) scavenger comprises an isocyanoacetate (ICNA) salt, ethyl isocyanoacetate, methyl isocyanoacetate, cysteine or its salt, L-cysteine or its salt, N-acetyl-L-cysteine, potassium ethyl xanthate, potassium isopropyl xanthate, glutathione, lipoic acid, ethylenediaminetetraacetic acid (EDTA), iminodiacetic acid, nitrilodiacetic acid, trithiolo-S-triazine, dimethyldithiocarbamate, dithiothreitol, mercaptoethanol, allyl alcohol, propargyl alcohol, thiol, tertiary amine and / or tertiary phosphine, or a combination thereof.
[0350] 52. The kit according to item 51, wherein the Pd(II) scavenger is L-cysteine or its salt.
[0351] Examples
[0352] Additional embodiments are disclosed in more detail in the following examples, which are not intended to limit the scope of the claims in any way.
[0353] Example 1. Synthesis of fully functionalized nucleotides with 3′ AOM and AOL linker moieties
[0354] Scheme 1. Synthesis of AOL linker moiety
[0355]
[0356] Synthesis of intermediate AOLLN2 :
[0357] Under N2, the acetal compound LN1 (2.43 g, 9.6 mmol) was dissolved in anhydrous CH2Cl2 (100 mL), and the solution was cooled to 0 °C with an ice bath. 2,4,6-Trimethylpyridine (7.6 mL, 57.5 mmol) was added, and then trimethylsilyl trifluoromethanesulfonate (7.0 mL, 38.7 mmol) was added dropwise. The mixture was stirred at 0 °C for 2 hours, then allyl alcohol (13 L, 191.1 mmol) was added, and the reaction mixture was refluxed overnight. The reaction was quenched with a 98:2 mixture of MeOH / H2O, and the resulting solution was stirred for an additional 3 hours at RT. The mixture was diluted with CH2Cl2 (100 mL) and water (200 mL), and the aqueous layer was acidified to pH 2 to 3 with 2N HCl. The aqueous layer was separated, and the organic layer was additionally extracted with acidic water. The organic layer was dried over MgSO4, filtered, and the volatiles were evaporated under reduced pressure. The crude product was purified by flash chromatography on silica gel to give AOL LN2 as a colorless oil (2.56 g, 86%).
[0358] Synthesis of intermediate AOLLN3 :
[0359] To a solution of AOL LN2 (2.17 g, 7.0 mmol) in ethanol (17.5 mL) was added an aqueous solution of 4 M NaOH (17.5 mL, 70 mmol). The mixture was stirred at RT for 3 h. Thereafter, all volatiles were removed under reduced pressure and the residue was dissolved in 75 mL of water. The solution was acidified to pH 2 - 3 with 2 N HCl and then extracted with dichloromethane (DCM). The combined organic fractions were dried over MgSO4, filtered and the volatiles were evaporated under reduced pressure. AOL LN3 was obtained as a colorless oil (1.65 g, 84%) without further purification. It solidified upon storage at -20 °C. LC-MS (ES): (negative ion) m / z 281 (M - H + ); (positive ion) m / z 305 (M + Na + ).
[0360] Synthesis of intermediate AOLLN4 :
[0361] A solution of AOL LN3 (1.62 g, 5.74 mmol) in anhydrous DMF (20 mL) was stirred under vacuum for 5 min and then cooled to 0 °C in an ice bath. N,N-Diisopropylethylamine (1.2 mL, 6.89 mmol) was added dropwise under N2, followed by PyBOP (3.30 g, 6.34 mmol). The reaction mixture was stirred at 0 °C for 30 min and then a solution of N-(5-aminopentyl)-2,2,2-trifluoroacetamide hydrochloride (1.62 g, 6.90 mmol) in anhydrous DMF (3.0 mL) was added, followed immediately by an additional amount of N,N-diisopropylethylamine (1.4 mL, 8.04 mmol). The reaction mixture was removed from the ice bath and stirred at RT for 4 h. The volatiles were removed under reduced pressure and the residue was dissolved in EtOAc (150 mL). The solution was extracted with aqueous 20 mM KHSO4, water and saturated aqueous NaHCO3. The organic layer was dried over MgSO4, filtered and the volatiles were evaporated under reduced pressure. The crude product was purified by flash chromatography on silica gel to give AOL LN4 as a colorless oil (2.16 g, 82%). LC-MS (ES): (negative ion) m / z 461 (M - H + ), 497 (M - Cl - ).
[0362] Synthesis of AOL linker moiety
[0363] To a solution of AOL LN4 (350 mg, 0.76 mmol) in CH3CN (13 mL) was added TEMPO (48 mg, 0.31 mmol), followed by the addition of an aqueous solution (6.5 mL) of NaH2PO4·2H2O (762 mg, 4.88 mmol) and NaClO2 (275 mg, 3.04 mmol). An aqueous solution of NaClO (14% available chlorine, 0.83 mL, 1.94 mmol) was added and the solution immediately turned dark brown. The reaction mixture was stirred at RT for 6 h and then quenched with 100 mM aqueous Na2S2O3 until the mixture became colorless. Acetonitrile was removed under reduced pressure, the residue was diluted with water and then basified with triethylamine. The aqueous phase was extracted with EtOAc (10 mL) and then concentrated under reduced pressure. The crude product was purified by reverse-phase flash chromatography on a C18 column to give AOL as a colorless oil (triethylammonium salt, 310 mg, 71%). LC-MS (ES): (negative ion) m / z 475 (M-H + ); (positive ion) m / z 499 (M+Na + ), 578 (M+Et3NH + ).
[0364] Synthesis of AOL-NH2 linker moiety
[0365] To a solution of AOL (446 mg, 0.94 mmol) in methanol (10 mL) was added aqueous NH3 (35%, 40 mL) and the mixture was stirred at RT for 5.5 h. Thereafter, all volatiles were removed under reduced pressure and the crude product was purified by reverse-phase flash chromatography on a C18 column to give AOL NH2 as a white solid (quantitative). 1 1H NMR (400 MHz, DMSO-d6): δ (ppm) 8.86 (t, J = 5.5 Hz, 1H, CONH), 8.28 (s, 3H, NH3 + ), 7.85 (s, 1H, Ar-H), 7.41 (d, J = 7.6 Hz, 1H, Ar-H), 7.31 (t, J = 7.9 Hz, 1H, Ar-H), 7.03 (ddd, J = 8.1, 2.5, 1.1 Hz, 1H, Ar-H), 5.87 (ddt, J = 17.2, 10.5, 5.3 Hz, 1H, OCH2CHCH2), 5.24 (dq, J = 17.2, 1.7 Hz, 1H, OCH2CHCH2, H a ), 5.09 (dq, J = 10.5, 1.5 Hz, 1H, OCH2CHCH2, H b ), 5.02 (dd, J = 6.7, 2.4 Hz, 1H, OCHO), 4.41 (dd, J = 12.2, 2.5 Hz, 1H, OCH2, Ha ), 4.18–3.99 (m, 3H, OCH2CHCH2 and OCH2, H b ), 3.94–3.81 (m, 2H, OCH2COOH), 3.49–3.39 (m, 1H, CH 2, H a ), 3.21–3.10 (m, 1H, CH 2, H b ), 2.86–2.70 (m, 2H, CH2), 1.81–1.39 (m, 6H, CH2). 13 C NMR (101 MHz, DMSO-d6): δ (ppm) 172.9, 166.1, 158.0, 136.2, 135.2, 129.3, 120.4, 119.3, 116.0, 111.3, 99.0, 68.8, 67.7, 66.8, 38.7, 38.4, 27.8, 26.3, 23.0. LC-MS (ESI): (negative ion) 379 (M-H); (positive ion) m / z 381 (M+H + ).
[0366] General procedure for dye-AOL linker coupling :
[0367] Dissolve the dye carboxylate (0.15 mmol) in 6 mL of anhydrous N,N’-dimethylformamide (DMF). Add N,N-diisopropylethylamine (136 μL, 0.78 mmol), and then add a 0.5 M anhydrous DMF solution of N,N,N',N'-tetramethyl-O-(N-succinimidyl)uronium tetrafluoroborate (TSTU, 300 μL, 0.15 mmol). At RT, stir the reactants under nitrogen for 1 hour. Add an aqueous solution (400 μL) of AOL NH2 (0.10 mmol) to the activated dye solution, and stir the reactants at RT for 3 hours. Take the crude product and purify it by preparative RP-HPLC.
[0368]
[0369] Characterization of AOL-SO7181: Yield 54% (54 μmol). LC-MS (ES): (negative ion) m / z 1002 (M-H + ); (positive ion) m / z 1004 (M+H + ).
[0370]
[0371] Characterization of AOL - AF550POPOS0: Yield 88% (88 μmol). LC - MS(ES): (negative ion) m / z 1034 (M - H + ), 516 (M - 2H + ); (positive ion) m / z 1036 (M + H + ), 1137 (M + Et3NH + ).
[0372]
[0373] Characterization of AOL - NR7180A: Yield 24% (23.9 μmol). LC - MS(ES): (positive ion) m / z = 880 (M + H) + .
[0374]
[0375] Characterization of AOL - NR550S0: Yield 22% (21.9 μmol). LC - MS(ES): (positive ion) m / z = 963 (M + H) + . (negative ion) m / z = 961 (M - H + ).
[0376] Scheme 2. Synthesis of 5′-triphosphate-3′-AOM-A(DB) nucleotide
[0377]
[0378] Synthesis of intermediate A2 :
[0379] Compound A1 (319 mg, 0.419 mmol) was dissolved in 0.8 mL of anhydrous DCM under N2 atmosphere, then pentamethylcyclopentadienyltris(acetonitrile)ruthenium(II) hexafluorophosphate ([RuCp*(MeCN)3]PF6, 42 mg, 0.08 mmol) was added, and then triethoxysilane (231 μL, 1.25 mmol) was added. At RT, the reactants were stirred under N2 for 1 h. Then the solution was diluted with DCM, filtered through a silica plug, and the silica plug was washed with ethyl acetate. The solution was evaporated under reduced pressure and dried under vacuum for 10 min, then the residue was dissolved in 2 mL of anhydrous THF. Copper(I) iodide (15 mg, 0.08 mmol) and a THF solution of 1 M TBAF (920 μL, 0.919 mmol) were added. The reactants were stirred at RT for 2.5 h, diluted with EtOAc, and then extracted with saturated NH4Cl. The aqueous phase was extracted with EtOAc. The combined organic phases were dried over MgSO4, filtered, and evaporated to dryness. The residue was purified by flash chromatography on silica gel. Yield: 125 mg (0.237 mmol). LC-MS (ES and CI): (positive ion) m / z 527 (M+H + ).
[0380] Synthesis of intermediate A3 :
[0381] The nucleoside A2 (155 mg, 0.294 mmol) was dried with P2O5 under reduced pressure for 18 h. Under nitrogen, anhydrous triethyl phosphate (1 mL) and some freshly activated molecular sieves were added thereto, and then the reaction flask was cooled to 0 °C in an ice bath. Freshly distilled POCl3 (33 μL, 0.353 mmol) was added dropwise, and then Proton (113 mg, 0.53 mmol). After addition, the reaction mixture was further stirred at 0 °C for 15 min. Then, a solution of 0.5 M di - tri - n - butylammonium pyrophosphate (2.94 mL, 1.47 mmol) in anhydrous DMF was added rapidly, followed immediately by tri - n - butylamine (294 μL, 1.32 mmol). The reaction mixture was kept in an ice - water bath for an additional 10 min and then quenched by pouring into 1 M triethylammonium bicarbonate aqueous solution (TEAB, 10 mL) and stirred at RT for 4 h. All solvents were evaporated under reduced pressure. 35% aqueous ammonia solution (10 mL) was added to the above residue and the mixture was stirred at RT for at least 5 h. Then the solvent was evaporated under reduced pressure. The crude product was first purified by ion - exchange chromatography on DEAE - Sephadex A25 (50 g). The column was eluted with aqueous triethylammonium bicarbonate (TEAB). The fractions containing triphosphate were combined and the solvent was evaporated to dryness under reduced pressure. The crude material was further purified by preparative HPLC. Compound A3 was obtained as the triethylammonium salt. Yield: 134 μmol (46%). LC - MS (ESI): (negative ion) m / z 614 (M - H + ).
[0382] In addition, 5′ - triphosphate - 3′ - AOM - A nucleotide and the corresponding ffA with the structure were also prepared. The detailed synthetic procedures are described in US Patent Application No. 16 / 724,088.
[0383] General synthesis process of nucleotide triphosphate-AOL linker :[[]]END]]
[0384] The compound AOL (0.120 mmol) was co-evaporated with 2 × 2 mL of anhydrous N,N'-dimethylformamide (DMF), and then dissolved in 3 mL of anhydrous N,N'-dimethylacetamide (DMA). N,N-Diisopropylethylamine (70 μL, 0.4 mmol) was added, followed by N,N,N',N'-tetramethyl-O-(N-succinimidyl)uronium tetrafluoroborate (TSTU, 36 mg, 0.120 mmol). At RT, the reactants were stirred under nitrogen for 1 hour. Meanwhile, an aqueous solution of nucleoside triphosphate (0.08 mmol) was evaporated to dryness under reduced pressure and then resuspended in 300 μL of 0.1 M triethylammonium bicarbonate (TEAB) aqueous solution. The activated linker solution was added to the nucleoside triphosphate, and the reactants were stirred at RT for 18 hours and then monitored by RP-HPLC. The solution was concentrated, and then 10 mL of concentrated aqueous NH4OH was added. The reactants were stirred at RT for 24 hours and then evaporated under reduced pressure. The crude product was first purified by ion-exchange chromatography on DEAE-Sephadex A25 (50 g) eluted with an aqueous solution of triethylammonium bicarbonate (TEAB). The fractions containing triphosphate were combined and the solvent was evaporated to dryness under reduced pressure. The crude material was further purified by preparative HPLC.
[0385]
[0386] Characterization of pppA(DB)-(3’AOM)-AOL: Yield 60 μmol (75%). LC-MS (ES): (negative ion) m / z 976 (M-H + ), 488 (M-2H + ).
[0387]
[0388] Characterization of pppA(DB)-(3’AOM)-AOL: Yield 68 μmol (72%). 11H NMR (400 MHz, D2O): δ (ppm) 7.94 (d, J = 1.7 Hz, 1H, H-2), 7.32 (d, J = 1.6 Hz, 1H, H-8), 7.10–6.96 (m, 2H, Ar), 6.95–6.87 (m, 2H, Ar), 6.41–6.31 (m, 1H, 1’-CH), 6.01–5.81 (m, 2H, CH allyl), 5.29 (ddq, J = 17.2, 6.1, 1.4 Hz, 2H, CHH allyl), 5.20 (ddt, J = 10.5, 3.8, 1.1 Hz, 2H, CHH allyl), 5.02 (td, J = 4.4, 2.4 Hz, 1H, O-CH2-O linker), 4.92–4.81 (m, 2H, 3’-O-CH2-O), 4.53 (dd, J = 4.9, 2.4 Hz, 1H, 3’-CH), 4.39 (dd, J = 16.1, 4.0 Hz, 1H, O-CHH linker), 4.32–4.19 (m, 2H, O-CHH linker, 4’-CH), 4.17–4.01 (m, 8H, 5’-CH2, CH2O linker, CH2-O allyl), 3.25–3.11 (m, 2H, CH2-NHCO), 3.04 (q, J = 7.3 Hz, 18H, Et3N), 2.92–2.82 (m, 2H, CH2-N linker), 2.55–2.41 (m, 2H, 2’-CH2), 1.59 (p, J = 7.6 Hz, 2H, CH2-CH2-N linker), 1.47 (p, J = 7.1 Hz, 2H, CH2-CH2-NHCO), 1.31 (tt, J = 8.3, 4.4 Hz, 2H, CH2-CH2-CH2-N linker), 1.15 (t, J = 7.3 Hz, 26H). 31 31P NMR (162 MHz, D2O): δ (ppm) -6.18 (d, J = 20.6 Hz, γ P), -11.32 (d, J = 19.3 Hz, α P), -22.20 (t, J = 19.9 Hz, β P).
[0389] Scheme 3. Synthesis of 5’-triphosphate-3’-AOM-C(DB) nucleotide
[0390]
[0391] Synthesis of intermediate C1
[0392] 5-Iodo-5'-O-(tert-butyldiphenylsilyl)-2'-deoxycytidine (3 g, 5.07 mmol) was dissolved in 30 mL of anhydrous pyridine, and then trimethylchlorosilane (1.29 mL, 10.1 mmol) was added dropwise. The reaction mixture was stirred at RT for 1 h, then placed in an ice bath, and benzoyl chloride (648 μL, 5.6 mmol) was added dropwise slowly. The reaction mixture was taken out of the ice bath and then stirred at RT for 1 h. After completion, the solution was placed in an ice bath and quenched with 50 mL of cold water, then 50 mL of methanol and 20 mL of pyridine were added, and the suspension was stirred overnight at RT. The solvent was evaporated under reduced pressure, and the residue was dissolved in 200 mL of EtOAc and extracted with 2 × 200 mL of saturated NaHCO3 and 100 mL of brine. The organic phase was dried over MgSO4, filtered and evaporated to dryness. The crude product was purified by flash chromatography on silica gel to give C1. Yield: 2.535 g (3.64 mmol, 73%). LC-MS (ESI): (positive ion) m / z 696 (M+H + ), 797 (M+Et3NH + ).
[0393] Synthesis of intermediate C2
[0394] N-Benzoyl-5-iodo-5'-O-(tert-butyldiphenylsilyl)-2'-deoxycytidine (C1) (695 mg, 1 mmol) and palladium(II) acetate (190 mg, 0.85 mmol) were dissolved in dry degassed DMF (10 mL), and then N-allyltrifluoroacetamide (7.65 mL, 5 mmol) was added. The solution was placed under vacuum and purged with nitrogen three times, and then degassed triethylamine (278 μL, 2 mmol) was added. The solution was heated to about 80 °C and protected from light for 1 h. The resulting black mixture was cooled to room temperature, then diluted with 50 mL of EtOAc and extracted with 100 mL of water. Then the aqueous phase was extracted with EtOAc. The combined organic phases were dried over MgSO4, filtered and evaporated to dryness. The crude product was purified by flash chromatography on silica gel to give C2. Yield: 305 mg (0.42 mmol, 42%). LC-MS (ES and CI): (positive ion) m / z 721 (M+H + ), 797 (M+Et3NH + ).
[0395] Synthesis of intermediate C3
[0396] Dissolve N-benzoyl-5-[3-(2,2,2-trifluoroacetamido)-allyl]-5'-O-(tert-butyldiphenylsilyl)-2'-deoxycytidine (C2) (350 mg, 0.486 mmol) in 1.1 mL of anhydrous DMSO (14.5 mmol), then add glacial acetic acid (1.7 mL, 29.1 mmol) and acetic anhydride (1.7 mL, 17 mmol). Heat the reaction mixture to 60 °C for 6 hours, then quench with 50 mL of saturated aqueous NaHCO3. After the solution stops foaming, extract with EtOAc. Combine the organic phases and wash with saturated aqueous NaHCO3, water, and brine. Dry the organic phase over MgSO4, filter, and evaporate to dryness. Purify the crude product by flash chromatography on silica gel to obtain C3. Yield: 226 mg (0.289 mmol, 60%). LC-MS (ESI): (positive ion) m / z 781 (M+H + ), 882 (M+Et3NH + ).
[0397] Synthesis of intermediate C4
[0398] Under N2 atmosphere, dissolve N-benzoyl-5-[3-(2,2,2-trifluoroacetamido)-allyl]-5'-O-(tert-butyldiphenylsilyl)-3'-O-methylthiomethyl-2'-deoxycytidine (C3) (210 mg, 0.27 mmol) in 5 mL of anhydrous DCM, add cyclohexene (136 μL, 1.35 mmol), and then cool the solution to about -10 °C. Add dropwise a freshly distilled 1 M solution of sulfuryl chloride in anhydrous DCM (320 μL, 0.32 mmol), and stir the reaction mixture for 20 minutes. After all the starting materials have been consumed, add an additional portion of cyclohexene (136 μL, 1.35 mmol), and evaporate the reaction mixture to dryness under reduced pressure. Quickly purge the residue with nitrogen, then dissolve the residue in 2.5 mL of ice-cold anhydrous DCM, and add ice-cold allyl alcohol (2.5 mL) with stirring at 0 °C. Stir the reaction mixture at 0 °C for 3 hours, then quench with saturated aqueous NaHCO3 and further dilute with saturated aqueous NaHCO3. Extract the mixture with EtOAc. Combine the organic phases and dry over MgSO4, filter, and evaporate to dryness. Purify the residue by flash chromatography on silica gel to obtain C4. Yield: 58% (124 mg, 0.157 mmol). LC-MS (ESI): (positive ion) m / z 791 (M+H + ).
[0399] Synthesis of intermediate C5
[0400] Under N2 atmosphere, N-benzoyl-5-[3-(2,2,2-trifluoroacetamido)-allyl]-5'-O-(tert-butyldiphenylsilyl)-3'-O-allyloxymethyl-2'-deoxycytidine (C4) (120 mg, 0.162 mmol) was dissolved in anhydrous THF (5 mL), and then placed at 0 °C. Glacial acetic acid (29 μL, 0.486 mmol) was added, and then a 1.0 M solution of TBAF in THF (486 μL, 0.486 mmol) was added immediately. The solution was stirred at 0 °C for 3 hours. The solution was diluted with EtOAc and then extracted with 0.025 N HCl and brine. The organic phase was dried over MgSO4, filtered and evaporated to dryness. The residue was purified by flash column chromatography on silica gel to give C5. Yield: 50 mg (0.090 mmol, 55%). LC-MS (ESI): (positive ion) m / z 553 (M+H + ). (negative ion) m / z 551 (M-H + ), 587 (M+Cl - ).
[0401] Synthesis of intermediate C6
[0402] N-benzoyl-3'-O-allyloxymethyl-5-[3-(2,2,2-trifluoroacetamido)-allyl]-2'-deoxycytidine (C5) (50 mg, 0.0.09 mmol) was dried with P2O5 under reduced pressure for 18 hours. Under nitrogen, anhydrous triethyl phosphate (1 mL) and some freshly activated molecular sieves were added thereto, and then the reaction flask was cooled to 0 °C in an ice bath. Freshly distilled POCl3 (10 μL, 0.108 mmol) was added dropwise, and then Proton (29 mg, 0.135 mmol). After addition, the reaction mixture was further stirred at 0 °C for 15 minutes. Then, a solution of 0.5 M di - tri - n - butylammonium pyrophosphate (1 mL, 0.45 mmol) in anhydrous DMF was added rapidly, followed immediately by tri - n - butylamine (100 μL, 0.4 mmol). The reaction mixture was kept in an ice - water bath for an additional 10 minutes and then quenched by pouring it into 1 M aqueous triethylammonium bicarbonate (TEAB, 5 mL) and stirred at RT for 4 hours. All solvents were evaporated under reduced pressure. 35% aqueous ammonia solution (5 mL) was added to the above residue and the mixture was stirred at RT for 18 hours. The solvent was evaporated under reduced pressure, and the residue was resuspended in 10 mL of 0.1 M TEAB and then filtered. The filtrate was first purified by ion - exchange chromatography on DEAE - Sephadex A25 (50 g). The column was eluted with aqueous triethylammonium bicarbonate solution. The fractions containing triphosphate were combined and the solvent was evaporated to dryness under reduced pressure. The crude product was further purified by preparative HPLC. Compound C6 was obtained as the triethylammonium salt. Yield: 40.6 μmol (45%), based on ε 290 = 5041 M -1 cm -1 . 1 1H NMR (400 MHz, D2O): δ (ppm) 8.23 (s, 1H, H - 6), 6.53 (dd, J = 15.5, 0.9 Hz, 1H, Ar - CH=), 6.42–6.24 (m, 2H, Ar - CH=CH -, 1’ - CH), 5.98 (ddt, J = 17.3, 10.4, 5.9 Hz, 1H, O - CH2 - CH=), 5.37 (dq, J = 17.3, 1.6 Hz, 1H, CHH=), 5.29 (ddt, J = 10.4, 1.6, 1.1 Hz, 1H, CHH=), 4.89 (s, 2H, O - CH2 - O), 4.60 (dt, J = 6.2, 3.1 Hz, 1H, 3’ - CH), 4.39 (t, J = 2.7 Hz, 1H, 4’ - CH), 4.35 (dq, J = 12.0, 3.8 Hz, 1H, 5’ - CHH), 4.28–4.21 (m, 1H, 5’ - CHH), 4.20 (ddt, J = 6.0, 2.7, 1.4 Hz, 1H, =CH - CH2 - O), 3.73 (dt, J = 7.2, 1.4 Hz, 2H, CH2 - NH2), 3.18 (q, J = 7.3 Hz, 20H, Et3N), 2.59 (ddd, J = 14.1, 6.1, 3.3 Hz, 1H, 2’ - CHH), 2.37 (ddd, J = 14.2, 7.2, 6.1 Hz, 1H, 2’ - CHH), 1.27 (t, J = 7.3 Hz, 31H, Et3N). 31³¹P NMR (162 MHz, D₂O): δ (ppm) -6.06 (d, J = 20.7 Hz, γ ³¹P), -11.24 (d, J = 19.1 Hz, α ³¹P), -21.95 (t, J = 19.7 Hz, β ³¹P). LC-MS (ESI): (negative ion) m / z 591 (M-H + ).
[0403] In addition, 5′-triphosphate-3′-AOM-C nucleotide and 5′-triphosphate-3′-AOM-T(DB) nucleotide with the following structures were also prepared:
[0404] and the corresponding ffC and ffT(DB). Finally, 5′-triphosphate-3′-AOM-G (also known as ffG-(3′-AOM)) was also prepared The detailed synthesis steps are described in U.S. Patent Publication No. 2020 / 0216891.
[0405] General synthesis process of fully functionalized nucleotides with AOL linker moiety
[0406] The dye-COOH (0.02 mmol) or dye-AOL (0.02 mmol) was co-evaporated with 2 × 2 mL of anhydrous N,N′-dimethylformamide (DMF), and then dissolved in 2 mL of anhydrous N,N′-dimethylacetamide (DMA). N,N-Diisopropylethylamine (17 μL, 0.1 mmol) was added, and then N,N,N′,N′-tetramethyl-O-(N-succinimidyl)uronium tetrafluoroborate (TSTU, 6 mg, 0.02 mmol) was added. At RT, the reactants were stirred under nitrogen for 1 hour. Meanwhile, an aqueous solution of nucleotide triphosphate (0.01 mmol) was evaporated to dryness under reduced pressure and then resuspended in 200 μL of 0.1 M triethylammonium bicarbonate (TEAB) aqueous solution. The activated dye solution was added to the nucleotide triphosphate, and the reactants were stirred at RT for 18 hours and then monitored by RP-HPLC. First, the crude product was purified by ion exchange chromatography on DEAE-Sephadex A25 (25 g) with a linear gradient elution using an aqueous solution of triethylammonium bicarbonate (TEAB, 0.1 M to 1 M). The fractions containing triphosphate were combined and the solvent was evaporated to dryness under reduced pressure. The crude material was further purified by preparative HPLC.
[0407]
[0408] Characterization of ffA-(3’AOM)-AOL-BL-NR650C5: Yield 66% (6.6 μmol). LC-MS(ES): (negative ion) m / z 1931 (M-H + ), 965 (M-2H + ).
[0409]
[0410] Characterization of ffA-(3’AOM)-AOL-BL-NR550s0: Yield 65% (6.5 μmol). LC-MS(ES): (negative ion) m / z 1784 (M-H + ), 891 (M-2H + ).
[0411]
[0412] Characterization of ffA(DB)-(3’AOM)-AOL-BL-NR650C5: Yield 31% (3.1 μmol). LC-MS(ES): (negative ion) m / z 1933 (M-H + ), 965 (M-2H + ), 644 (M-3H + ).
[0413]
[0414] Characterization of ffA(DB)-(3’AOM)-AOL-NR7181A: Yield 21% (2.12 μmol). LC-MS(ES): (negative ion) m / z = 1475 (M-H + ).
[0415]
[0416] Characterization of ffA-(3'-AOM)-AOL-NR7180A: Yield 41% (4.1 μmol). LC-MS(ES): (negative ion) m / z 1472 (M-H + ), 736 (M-2H + ).
[0417]
[0418] Characterization of ffA(DB)-(3'-AOM)-AOL-BL-NR550S0: Yield 21% (2.1 μmol). LC-MS(ES): (negative ion) m / z 1786 (M-H + ), 892 (M-2H +), 594 (M - 3H + ).
[0419]
[0420] Characterization of ffC(DB)-(3’AOM)-AOL-SO7181: Yield 48% (4.87 μmol). LC-MS(ES): (negative ion) m / z 1577 (M - H + ), 788 (M - 2H + ), 525 (M - 3H + ).
[0421]
[0422] Characterization of ffC-(3'-AOM)-AOL-SO7181: Yield 56% (5.6 μmol). LC-MS(ES): (negative ion) m / z1575 (M - H + ), 787 (M - 2H + ).
[0423]
[0424] Characterization of ffT(DB)-3’AOM-AOL-AF550POPOS0: Yield 46% (4.6 μmol). LC-MS(ES): (negative ion) m / z 1609 (M - H + ), 804 (M - 2H + ), 536 (M - 3H + ).
[0425]
[0426] Characterization of ffT(DB)-(3’AOM)-AOL-NR550s0: Yield 38% (3.8 μmol). LC-MS(ES): (negative ion) m / z = 1535 (M - H + ).
[0427] Example 2. Solution cleavage efficiency of different palladium reagent formulations
[0428] Figure 1Shows a comparison of the solution cleavage efficiencies of the following three different palladium reagent formulations: 1) 10 mM [(allyl)PdCl]2, 100 mM THP, 100 mM ethanolamine buffer, 10 mM sodium ascorbate; 2) 20 mM Na2PdCl4, 60 mM THP, 100 mM N,N'-diethylethanolamine, 10 mM sodium ascorbate, 1 M NaCl, 0.1% Tween 20; 3) 20 mM Na2PdCl4, 70 mM THP, 100 mM N,N'-diethylethanolamine, 10 mM sodium ascorbate, 1 M NaCl, 0.1% Tween 20. The cleavage efficiency was determined by measuring the relative cleavage rate of the 3'-AOM-nucleotide substrate. Briefly, a stock solution of the palladium reagent was added to a 100 mM buffer solution of 0.1 mM 3'-AOM-nucleotide substrate such that the final concentration of the Pd species was 1 mM. The solution was incubated at RT and then the reaction kinetics were monitored by taking aliquots from the reactants at set time points, quenching them with a 1:1 solution of EDTA / H2O2 (0.1:0.1 M), and then analyzing these aliquots by HPLC for the formation of 3'-OH nucleotides and the disappearance of the 3'-AOM nucleotide substrate. As Figure 1 shown, the cleavage efficiency of Na2PdCl4 with only 3 or 3.5 equivalents of THP was comparable to that of [(allyl)PdCl]2, which used 10 equivalents of THP.
[0429] Example 3: Stability study of AOM-AOL ffN in solution and performance in sequencing
[0430] Figure 2 Shows the predetermined phase properties of fully functionalized nucleotides (ffN) with 3'-AOM capping groups and AOL linker moieties (including labeled ffT-DB, labeled ffA, and labeled ffC, as well as unlabeled ffG) compared to a stressed standard ffN. Two sets of ffN were incubated at 45 °C for several days in a standard incorporation mixture formulation excluding only DNA polymerase. For each time point, fresh polymerase was added to obtain a complete incorporation mixture, which was then directly loaded onto sequencing conditions described previously. The predetermined phase % is a direct indication of the percentage of 3’OH-ffN present in the mixture and is thus directly related to the stability of the 3’ capping group. The predetermined phase values for the two sets of ffN were recorded and plotted ( Figure 2 ). Compared to the standard, the AOM-AOL-ffN did not show any increase in the predetermined phase and appeared to be much more stable than the standard ffN with 3’-O-azidomethyl capping groups and LN3 linker moieties.
[0431] Example 4. Use of palladium scavengers in sequencing reactions
[0432] Figure 3 Shows the phasing values on an Illumina instrument with and without potassium isocyanoacetate in the wash step after cleavage when using fully functionalized nucleotides (ffN) with 3'-AOM capping groups and standard LN3 linker moieties (including labeled ffT-DB, labeled ffA, and labeled ffC, as well as unlabeled ffG). The sequencing experiment was performed on an Illumina instrument using a cartridge, where the standard incorporation mixture reagent was replaced with a freshly prepared incorporation mixture containing ffN with 3'-AOM capping groups and standard LN3 linker moieties, and where a freshly prepared palladium cleavage reagent solution (10 mM [(allyl)PdCl]2, 100 mM THP, 100 mM ethanolamine buffer, 10 mM sodium ascorbate, 1 M NaCl, 0.1% Tween 20) was added to empty wells. Potassium isocyanoacetate was added to the standard post-cleavage wash solution to a final concentration of 10 mM. The sequencing experiment was performed using a 2×151-cycle protocol that included, in addition to the standard sequencing-by-synthesis (SBS) protocol, a 5-second incubation with the palladium cleavage reagent solution. As shown, it was observed that when 10 mM potassium isocyanoacetate was used in the post-cleavage wash solution, the % phasing decreased from 0.183 to 0.075. Figure 3 Shows the main sequencing metrics (including phasing, pre-phasing, and error rate) on an Illumina instrument when using fully functionalized nucleotides (ffN) with 3'-AOM capping groups and AOL linker moieties (including labeled ffT-DB, labeled ffA, and labeled ffC, as well as unlabeled ffG) in the presence of a palladium scavenger, and the same sequencing metrics for comparison when using standard ffN with 3'-O-azidomethyl capping groups and LN3 linker moieties. The sequencing experiment was performed on an Illumina
[0433] Figure 4 instrument by running a 2×151-cycle protocol using a standard cartridge, where the incorporation mixture reagent and standard cleavage reagent were replaced with freshly prepared incorporation mixtures containing fully functionalized nucleotides (ffN) with 3'-AOM capping groups and AOL linker moieties, and a freshly prepared palladium cleavage reagent solution (10 mM [(allyl)PdCl]2, 100 mM THP, 100 mM N,N’-diethylethanolamine buffer, 10 mM sodium ascorbate, 1 M NaCl, 0.1% Tween 20), respectively. Potassium isocyanoacetate (ICNA) was added to the standard post-cleavage wash solution. post-cleavage wash solution. In the post-cleavage wash solution, a final concentration of 10 mM was reached. Using standard kits and procedures, control sequencing experiments were performed with standard ffN having a 3'-O-azidomethyl capping group. The results showed improvements in the phasing and more importantly the error rate, thus demonstrating the full efficiency of the AOM-AOL SBS chemical reaction in combination with a single cleavage step.
[0434] Example 5. Use of glycine in sequencing reactions
[0435] Figure 5 Shown are comparisons of the major sequencing metrics (including phasing and pre-phasing) on an Illumina instrument when using fully functionalized nucleotides (ffN) with a 3'-AOM capping group and an AOL linker moiety, in the presence of either glycine or ethanolamine in the incorporation mixture. In this example, the entire set of ffN included 3'-AOM-ffA-AOL (labeled), 3'-AOM-ffG-AOL (unlabeled), 3'-AOM-ffT(DB)-AOL (labeled), and 3'-AOM-ffC(DB)-AOL (labeled). As shown, when glycine was used in the incorporation buffer, a significant decrease in the phasing value was observed compared to a standard incorporation buffer containing ethanolamine. Although there was a slight increase in the pre-phasing value when glycine was used, this was not considered a significant increase. The sequencing experiments were performed on a standard Figure 5 instrument using a cartridge, where the standard incorporation mixture reagent and the standard cleavage reagent were replaced respectively with freshly prepared solutions of an incorporation mixture containing fully functionalized nucleotides (ffN) with a 3'-AOM capping group and an AOL linker moiety in 50 mM ethanolamine or 50 mM glycine buffer, and a freshly prepared palladium cleavage reagent solution (10 mM [(allyl)PdCl]2, 100 mM THP, 100 mM N,N'-diethylethanolamine buffer, 10 mM sodium ascorbate, 1 M NaCl, 0.1% Tween 20). Potassium isocyanatoacetate (ICNA) was added to the standard MiniSeq post-cleavage wash solution to a final concentration of 10 mM. A standard MiniSeq procedure with 2×151 cycles was used. Shown are comparisons of the major sequencing metrics (including phasing and pre-phasing) on an Illumina instrument when using fully functionalized nucleotides (ffN) with a 3'-AOM and an AOL linker moiety.
[0436] Figure 6 Shown are comparisons of the major sequencing metrics (including phasing and pre-phasing) on an Illumina instrument when using fully functionalized nucleotides (ffN) with a 3'-AOM and an AOL linker moiety. The main sequencing metrics on the instrument (including phasing, pre-phasing, and error rate), and the same sequencing metrics when compared to those obtained using standard ffNs with 3'-O-azidomethyl capping groups and LN3 linker moieties. In this example, the entire set of ffNs includes 3'-AOM-ffA-AOL (labeled), 3'-AOM-ffG-AOL (unlabeled), 3'-AOM-ffT(DB)-AOL (labeled), and 3'-AOM-ffC(DB)-AOL (labeled). The sequencing experiment was performed on a standard instrument using a 2×151-cycle protocol, where the standard incorporation mixture reagent and the standard lysis reagent were replaced with freshly prepared solutions of an incorporation mixture containing fully functionalized nucleotides (ffNs) with 3'-AOM capping groups in 50 mM glycine buffer, and a freshly prepared palladium lysis reagent solution (10 mM [(allyl)PdCl]2, 100 mM THP, 100 mM N,N'-diethylethanolamine buffer, 10 mM sodium ascorbate, 1 M NaCl, 0.1% Tween 20). Potassium isocyanoacetate was added to the standard post-lysis wash solution to a final concentration of 10 mM. The results showed a further improvement in the error rate compared to that obtained using standard ffNs. It is believed that the improvement in phasing of the AOM-AOL series compared to the Figure 5 metrics is due to the use of glycine buffer and 3'-AOM-ffC(DB)-AOL.
[0437] Example 6. In iSeq T M Synthetic and sequencing on the fly
[0438] Figure 7A and Figure 7B shows a comparison of the main sequencing metrics (including error rate and Q30 score) for 2×300 cycles of sequencing-by-synthesis performed on an Illumina iSeq TM instrument using fully functionalized nucleotides (ffNs) with 3'-AOM capping groups and AOL linker moieties. In this example, the entire set of AOM ffNs includes 3'-AOM-ffA(DB)-AO-dye 1, 3'-AOM-ffG (unlabeled), 3'-AOM-ffT(DB)-AOL-NR550S0, and 3'-AOM-ffC(DB)-AOL-dye 2. The entire set of azidomethyl (AZM) ffNs includes the same ffNs with 3'-azidomethyl capping groups, propargylamido, and LN3 linkers. Dye 1 is the chromene quinoline dye disclosed in U.S. Serial No. 63 / 127061 and has a structural moiety when conjugated to ffA Dye 2 is a coumarin dye disclosed in U.S. Patent Publication No. 2018 / 0094140 and has a structural moiety when conjugated with ffC NR550S0 is a known green dye.
[0439] Sequencing experiments were performed using cartridges on a standard iSeq TM instrument, where the standard incorporation mixture reagent and the standard lysis reagent were replaced with freshly prepared incorporation mixtures containing fully functionalized nucleotides (ffN) with 3'-AOM capping groups and AOL linker moieties in a solution of 50 mM ethanolamine or 50 mM glycine buffer using 300% concentration of polymerase 1901 (Pol 1901) (360 μg / mL), and a freshly prepared palladium lysis reagent solution (10 mM [(allyl)PdCl]2, 100 mM THP, 100 mM N,N'-diethylethanolamine buffer, 10 mM sodium ascorbate, 1 M NaCl, 0.1% Tween 20). L-cysteine was added to the standard iSeq TM post-lysis wash solution to a final concentration of 10 mM. An iSeq with 2 × 301 cycles was employed in a 2-excitation / 1-emission protocol TM procedure. In particular, the iSeq TM instrument was set to take a first image with green excitation light (approx. 520 nm) and a second image with blue excitation light (approx. 450 nm). 2 × 301 SBS cycles (incorporation, followed by imaging, then lysis) were performed using the standard sequencing procedure. The sequencing metrics are summarized in the table below.
[0440]
[0441]
[0442] It was observed that the AOM ffN group provided good performance, thus providing superior error rates and Q30 for both Read 1 and Read 2. The phasing values when using the AOM ffN group were comparable to those generated by the AZM ffN group. However, the AOM ffN group produced significantly lower pre-phasing values.
[0443] Figure 8A and Figure 8B demonstrates that when using fully functionalized nucleotides (ffN) with 3'-AOM capping groups and AOL linker moieties, on Illumina's iSeq TMComparison of the main sequencing metrics (including error rate and Q30 score) for 2×150 cycles of on-instrument synthesis and sequencing. In this example, the entire set of AOM ffN includes 3'-AOM-ffA(DB)-AO-Dye 1, 3'-AOM-ffG (unlabeled), 3'-AOM-ffT(DB)-AOL-NR550S0, and 3'-AOM-ffC(DB)-AOL-Dye 2. The entire set of AZM ffN includes the same ffN with a 3'-azidomethyl capping group, propargyl amido group, and LN3 linker. The sequencing experiment was performed on a standard iSeq TM instrument using a cartridge, where the standard incorporation mixture reagent and standard lysis reagent were replaced with a freshly prepared incorporation mixture containing fully functionalized nucleotides (ffN) with 3'-AOM capping groups and AOL linker moieties in a solution of 50 mM ethanolamine or 50 mM glycine buffer using 300% concentration of Pol 1901 (360 μg / mL), and a freshly prepared palladium lysis reagent solution (10 mM [(allyl)PdCl]2, 100 mM THP, 100 mM N,N'-diethylethanolamine buffer, 10 mM sodium ascorbate, 1 M NaCl, 0.1% Tween 20). L-Cysteine was added to the standard iSeq TM post-lysis wash solution to a final concentration of 10 mM. An iSeq TM protocol with 2×301 cycles was employed in a 2-excitation / 1-emission scheme. In particular, the iSeq TM instrument was set to take a first image with green excitation light (approx. 520 nm) and a second image with blue excitation light (approx. 450 nm). 2×301 SBS cycles (incorporation, followed by imaging, then lysis) were performed using the standard sequencing protocol.
[0444] The incorporation mixture contact time for AZM ffN was approximately 24.1 seconds, while the incorporation mixture contact time for AOM ffN was approximately 29.1 seconds. However, the longer incorporation time for AOM ffN was offset by a shorter uncapping time. The lysis mixture contact time was approximately 5.8 seconds, compared to approximately 10.2 seconds for the AZM ffN group. Thus, the total incubation times for the AZM ffN group and the AOM ffN group were approximately 34.4 seconds and 34.9 seconds, respectively. The sequencing metrics are summarized in the table below.
[0445]
[0446]
[0447] Similarly, it was observed that the AOM ffN group provided good performance, resulting in superior error rates and Q30 for both Read 1 and Read 2. Additionally, the AOM ffN group produced lower predetermined phase values.
[0448] Example 7. In NovaSeq with blue laser power titration T M Perform sequencing while synthesizing on
[0449] It has been observed that prolonged exposure to blue light in SBS sequencing leads to high levels of signal delay and phasing due to increased light dose and power density. This example compares the performance of AOM ffN and AZM ffN in a blue laser titration sequencing experiment.
[0450] In this experiment, 1×151 SBS runs using the AOM ffN group on the improved blue / green excitation NovaSeq TM were compared to those when using the standard ffN group with a 3'-azidomethyl capping group and LN3 linker. The flow cell used was a 490nm pitch BEER2 flow cell. The power configurations of the blue laser power were: 600mW, 800mW, 1000mW, 1400mW, 1800mW, and 2400mW. The green laser power was kept constant at 1000mM. The following standard AZM ffN were used: green ffT (LN3-AF550POPOS0), dark G, red ffC (LN3-SO7181), blue ffC (sPA-blue dye A), blue ffA (LN3-BL-blue dye A), green ffA (LN3-BL-NR550S0). For the AOM ffN group, the following ffN were used: green ffT (ffT(DB)-AOL-AF550POPOS0), dark G, red ffC (ffC(DB)-AOL-SO7181), blue ffC (ffC(DB)-AOL-blue dye A), blue ffA (ffA(DB)-AOL-BL-blue dye A), green ffA (ffA(DB)-AOL-BL-NR550S0). The structures of the blue dyes labeled AOM ffC and ffA are shown below:
[0451]
[0452] For running SBS with the AOM ffN group, the following modifications were made. First, 10 mM L-cysteine was added to the post-lysis wash solution. The lysis mixture included the following components: DEEA buffer and Na2PdCl4 therein. Two 10-second wait steps were added to the post-lysis wash step. Additionally, a 60-second static incorporation wait time was used (in contrast to 38 seconds for AZM ffN). Figure 9AIt is shown that these two ffN groups have similar phasing values at lower blue laser powers, but the AOM ffN group is less sensitive to increasing blue laser power titration (represented by a gentler phasing slope compared to the AZM phasing slope). Additionally, a much lower pre-phasing value for AOM ffN was observed. As Figure 9B shown, the AOM ffN group has significantly less signal attenuation at higher blue laser powers. Figure 9C And Figure 9D shows the variation of the average error rate with the number of cycles. Figure 9D is Figure 9C an enlarged view. The results show that while the AZM ffN has a lower error rate in the early cycles at lower blue laser powers, the AOM ffN group performs much better in the later cycles at higher blue laser powers. Figure 9E summarizes the average error rate for 151 cycles. Again, the AOM ffN group performs better than the AZM ffN group at high laser powers (e.g., at 1400 mW, 1800 mW, and 2400 mW).
[0453] Example 8. First chemical linearization using Pd cleavage mixture
[0454] In this example, the Pd cleavage mixture used in SBS was tested in the first chemical linearization step after the clustering step. This experiment compared chemical linearization with standard enzymatic linearization, in which USER promotes cleavage of one of the double-stranded polynucleotides at the U site on the P5 primer. Fully functionalized nucleotides (ffN) with 3'-AOM capping groups and AOL linker moieties were used in 1×150 SBS cycles on Illumina's iSeq TM instrument. In this example, the entire set of AOM ffN includes 3'-AOM-ffA(DB)-AO-dye 1, 3'-AOM-ffG (unlabeled), 3'-AOM-ffT(DB)-AOL-NR550S0, and 3'-AOM-ffC(DB)-AOL-dye 2, as described in Example 6. The iSeq TM instrument was set to take a first image with green excitation light (≈520 nm) and a second image with blue excitation light (≈450 nm) (using a 2-excitation / 1-emission scheme). The modified P5 / P7 primers were grafted onto the flow cell used on the iSeq TM instrument to allow for the first chemical linearization of the P5 primer. This chemical linearization step was carried out in a Pd cleavage mixture (a solution of 10 mM [Pd(allyl)Cl]2 and 100 mM THP in a buffer containing DEEA), which was incubated at 63 °C for 30 s. Figure 10Shows SBS sequencing metrics using two different linearization methods. When the Pd cleavage mixture was used in the first chemical linearization step, all major sequencing metrics were observed to fall within the standard observation range. This experiment confirmed that a single reagent mixture can be used for two separate sequencing steps - the linearization step and the SBS cleavage step, which enables further simplification of the instrument (flow cell and cartridge).
Claims
1. A method for determining the sequences of multiple different target single-stranded polynucleotides, comprising: (a) Contact a solid support with a solution containing a sequencing primer under hybridization conditions, wherein the solid support comprises a plurality of different target polynucleotides immobilized thereon; and the sequencing primer is complementary to at least a portion of the target polynucleotide; (b) Under conditions suitable for DNA polymerase-mediated primer extension, contact the solid support with an aqueous incorporation solution containing a DNA polymerase and one or more of four different types of nucleotides, and incorporate one type of nucleotide into the sequencing primer to produce an extended copy polynucleotide, wherein each of the nucleotides comprises a 3′ capping group, and at least one type of incorporated nucleotide comprises a detectable label linked by a cleavable linker; (c) Image the extended copy polynucleotide and perform one or more fluorescence measurements; (d) Remove the 3′ capping group from the nucleotide incorporated into the extended copy polynucleotide; (e) After removing the 3′ capping group from the incorporated nucleotide, wash the extended copy polynucleotide with a post-cleavage wash solution; and (f) Repeat steps (b) to (e) until the sequence of at least a portion of the target polynucleotide chain is determined; wherein one or more of the four different types of nucleotides in the aqueous incorporation solution have a structure selected from formula (Ia), (Ia′), (Ib), (Ib′), (Ib″), (Ic), (Ic′) or (Id): wherein R 4 is H, and -OR 6 is triphosphate.
2. The method according to claim 1, wherein steps (b) to (e) are repeated at least 50 times, at least 100 times, or at least 150 times.
3. The method according to claim 1, wherein step (d) comprises contacting the incorporated nucleotide with a cleavage solution comprising a palladium catalyst.
4. The method according to claim 3, wherein the palladium catalyst is a palladium(0) catalyst in-situ generated from a palladium complex and a water-soluble phosphine.
5. The method according to claim 4, wherein the palladium complex comprises [Pd(allyl)Cl]2, Na2PdCl4, [Pd(allyl)(THP)]Cl, [Pd(allyl)(THP)2]Cl, Pd(CH3CN)2Cl2, Pd(OAc)2, Pd(PPh3)4, Pd(dba)2, Pd(Acac)2, PdCl2(COD), or Pd(TFA)2, or a combination thereof.
6. The method according to claim 4, wherein the palladium complex comprises [Pd(allyl)Cl]2 or Na2PdCl4.
7. The method according to claim 1, wherein the cleavage solution further comprises one or more buffer reagents selected from the group consisting of primary amines, secondary amines, tertiary amines, carbonates, phosphates, borates, and combinations thereof.
8. The method according to claim 7, wherein the buffer reagent in the lysis solution is selected from the group consisting of ethanolamine (EA), tris(hydroxymethyl)aminomethane (Tris), glycine, carbonate, phosphate, borate, dimethylethanolamine (DMEA), diethylethanolamine (DEEA), N,N,N′,N′-tetramethylethylenediamine (TEMED), N,N,N′,N′-tetraethylethylenediamine (TEEDA), 2-piperidinoethanol, and combinations thereof.
9. The method according to claim 1, wherein the aqueous incorporation solution in step (b) further comprises at least one palladium(0) scavenger and one or more buffering agents.
10. The method according to claim 9, wherein the buffering agent in the aqueous incorporation solution comprises ethanolamine or glycine, or a combination thereof.
11. The method according to claim 9, wherein the palladium(0) scavenger in the aqueous incorporation solution comprises one or more allyl moieties each independently selected from the group consisting of: –O-allyl, –S-allyl, –NR-allyl, and –N + RR′-allyl, wherein R is H, unsubstituted or substituted C1-C6 alkyl, unsubstituted or substituted C2-C6 alkenyl, unsubstituted or substituted C2-C6 alkynyl, unsubstituted or substituted C6-C 10 aryl, unsubstituted or substituted 5- to 10-membered heteroaryl, unsubstituted or substituted C3-C 10 carbocyclic group, or unsubstituted or substituted 5- to 10-membered heterocyclic group; and R′ is H, unsubstituted C1-C6 alkyl, or substituted C1-C6 alkyl.
12. The method according to claim 11, wherein the palladium scavenger in the aqueous incorporation solution is:
13. The method according to claim 1, wherein the post-lysis wash solution in step (e) comprises one or more palladium(II) scavengers.
14. The method according to claim 13, wherein the one or more palladium(II) scavengers in the post-cleavage wash solution include isocyanoacetate (ICNA) salts, ethyl isocyanoacetate, methyl isocyanoacetate, cysteine or its salts, L-cysteine or its salts, N-acetyl-L-cysteine, potassium ethyl xanthate, potassium isopropyl xanthate, glutathione, lipoic acid, ethylenediaminetetraacetic acid (EDTA), iminodiacetic acid, nitrilodiacetic acid, trithiolo-S-triazine, dimethyldithiocarbamate, dithiothreitol, mercaptoethanol, allyl alcohol, propargyl alcohol, thiols, tertiary amines and / or tertiary phosphines, or combinations thereof.
15. The method according to claim 1, wherein the detectable label and the 3′-OH capping group from the nucleotides incorporated into the copy polynucleotide chain are removed in a single chemical reaction.
16. The method according to claim 1, wherein L 2 comprises wherein m and n are each independently an integer of 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10, and the phenyl moiety is optionally substituted.
17. The method according to claim 16, wherein n is 5 and m is 4.
18. The method according to claim 2, wherein one or more of the four different types of nucleotides in the aqueous incorporation solution are dATP, dGTP, dCTP, and dTTP, having the following structures: and where R 4 is H, and -OR 6 is triphosphate.
19. The method according to claim 18, wherein L 2 comprises wherein m and n are each independently an integer of 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10, and the phenyl moiety is optionally substituted.
20. The method according to claim 19, wherein n is 5 and m is 4.
21. The method according to claim 18, wherein the aqueous incorporation solution in step (b) comprises at least one palladium(0) scavenger and one or more buffers.
22. The method according to claim 18, wherein the post-cleavage wash solution in step (e) comprises one or more palladium(II) scavengers.
Citation Information
Patent Citations
Modified nucleic acid probes
EP0742287A2
Methods and compositions for selecting tag nucleic acids and probe arrays
EP0799897A1
Alternative substrates and formats for bead-based array of arrays TM
US20020102578A1
Polymerase enzymes and reagents for enhanced nucleic acid sequencing
US20080108082A1
Methods and apparatus for measuring analytes using large scale FET arrays
US20090026082A1