Nucleosides and nucleotides having a 3' acetal blocking group
Nucleotides with a 3' acetal blocking group and cleavable linker address sequencing challenges by ensuring efficient and stable nucleotide incorporation and removal, enhancing sequencing accuracy and data quality.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- ILLUMINA CAMBRIDGE LTD
- Filing Date
- 2025-12-24
- Publication Date
- 2026-06-02
AI Technical Summary
Existing nucleotide sequencing methods face challenges in achieving accurate and efficient sequencing due to limitations in the types of protecting groups that can be added to nucleotides, which need to prevent further nucleotide incorporation while being easily removable without damaging the polynucleotide chain and compatible with polymerase enzymes.
Nucleotides and nucleosides with a 3' acetal blocking group covalently bonded via a cleavable linker, allowing for simultaneous removal of the blocking group and label under mild conditions, enhancing stability and efficiency in sequencing processes.
Improves sequencing accuracy and data quality by reducing prephasing, signal attenuation, and enabling longer reads with improved stability during synthesis, formulation, and operation in sequencing instruments.
Smart Images

Figure 2026090242000087 
Figure 2026090242000088 
Figure 2026090242000089
Abstract
Description
[Technical Field]
[0001] (By reference to the priority application) This application claims priority to U.S. Provisional Patent Application No. 63 / 042,240, filed on 22 June 2020, which is incorporated in its entirety by reference. [Background technology]
[0002] (Background technology) This disclosure generally relates to nucleotides, nucleosides, or oligonucleotides, including 3' acetal blocking groups and their use in polynucleotide sequencing methods. Methods for preparing 3'-blocked nucleotides, nucleosides, or oligonucleotides are also disclosed.
[0003] (Explanation of related technologies) Advances in molecular research are partly driven by improvements in the techniques used to characterize molecules or their biological reactions. In particular, the study of nucleic acids, DNA and RNA, benefits from development techniques used in sequence analysis and the study of hybridization events.
[0004] An example of a technique that has improved nucleic acid research is the development of fabricated arrays of immobilized nucleic acids. These arrays typically consist of a high-density matrix of polynucleotides immobilized on a solid support material. See, for example, Fodor et al., Trends Biotech. 12:19-26, 1994, which describes a method for assembling nucleic acids using a chemisensitized glass surface that is protected by a mask but exposed in certain areas to allow for the attachment of appropriately modified nucleotide phosphoramidites. Fabricated arrays can also be fabricated by techniques for "spotting" known polynucleotides on a solid support at predetermined locations (e.g., Stimpson et al., Proc. Natl. Acad. Sci. 92:6379-6383, 1995).
[0005] One method for determining the nucleotide sequence of nucleic acids bound to an array is called "synthetic sequencing" or "SBS." This technique for sequencing DNA ideally requires the controlled (i.e., one at a time) incorporation of the correct complementary nucleotides on the opposite side of the nucleic acid being sequenced. This ensures that each nucleotide residue is sequenced one at a time, allowing for accurate sequencing by adding nucleotides in multiple cycles and thus preventing uncontrolled incorporations. The incorporated nucleotides are read using the appropriate label attached to them before the label portion and subsequent sequencing rounds are removed.
[0006] To ensure that only a single integration occurs, a structural modification ("protecting group" or "blocking group") is included in each labeled nucleotide added to the growth chain to ensure that only one nucleotide is integrated. After adding the nucleotide with the protecting group, the protecting group is removed under reaction conditions that do not impair the integrity of the sequenced DNA. The sequencing cycle can then continue with the integration of the next protected labeled nucleotide.
[0007] Nucleotides, which are typically nucleotide triphots, generally require a 3'-hydroxy protecting group to prevent the polymerase used to incorporate them into a polynucleotide chain from continuing to replicate as a base is added to the nucleotide. There are many limitations on the types of groups that can be added to nucleotides and are still suitable. The protecting group should prevent the addition of additional nucleotide molecules to the polynucleotide chain while simultaneously being easily removable from the sugar moiety without causing damage to the polynucleotide chain. Furthermore, the modified nucleotide must be compatible with the polymerase or another suitable enzyme used to incorporate it into the polynucleotide chain. Therefore, an ideal protecting group exhibits long-term stability, is efficiently incorporated by polymerase enzymes, causes blockage of secondary or further nucleotide incorporation, and is capable of being removed under mild conditions without damaging the polynucleotide structure, preferably under aqueous conditions.
[0008] Reversible protecting groups have been previously described. For example, Metzker et al., (Nucleic Acids Research, 22(20):4259-4267, 1994) disclose the synthesis and use of eight 3'-modified 2-deoxyribonucleoside 5'-triphosphates (3'-modified dNTPs) and their testing in two DNA template assays for incorporating their activity. International Publication No. 2002 / 029003 describes a sequencing method that may include the use of allyl protecting groups to cap 3'-OH groups on the DNA growth strand in polymerase reactions.
[0009] Furthermore, the development of numerous reversible protecting groups and methods for deprotecting them under DNA compatibility conditions has been previously reported in International Publication WO2004 / 018497 and International Publication 2014 / 139596, each of which is incorporated herein by reference in whole. [Overview of the Initiative]
[0010] Some embodiments of the present disclosure relate to nucleotides or nucleosides comprising nucleic acid bases bound to a detectable label via a cleavable linker, wherein the nucleoside or nucleotide comprises a ribose or 2'-deoxyribose moiety and a 3'-OH blocking group, and the cleavable linker comprises the following structural portion: [ka] In the formula, each of X and Y is independently either O or S, and R 1a , R 1b , R 2 , R 3a and R 3b Each of these is independently H, a halogen, an unsubstituted or substituted C1-C6 alkyl, or a C1-C6 haloalkyl.
[0011] Some embodiments of this disclosure relate to oligonucleotides or polynucleotides comprising 3'-OH block-labeled nucleotides as described herein.
[0012] Some embodiments of this disclosure relate to a method for preparing a growing polynucleotide complementary to a target single-stranded polynucleotide in a sequencing reaction, comprising incorporating a nucleotide molecule described herein into the growing complementary polynucleotide, wherein the nucleotide incorporation prevents the introduction of any subsequent nucleotides into the growing complementary polynucleotide. In some embodiments, the nucleotide incorporation is achieved by polymerase, terminal deoxynucleotidyltransferase (TdT), or reverse transcriptase. In one embodiment, the incorporation is achieved by polymerase (e.g., DNA polymerase).
[0013] Some further embodiments of the present disclosure relate to a method for determining the sequence of a target single-stranded polynucleotide, the method being: (a) Incorporating the nucleotides described herein into a copy polynucleotide chain complementary to at least a portion of the target polynucleotide chain, (b) detecting the identity of the nucleotides incorporated into the copy polynucleotide chain, (c) Chemically removing the label and the 3'-OH blocking group from the nucleotide incorporated into the copy polynucleotide chain, Includes.
[0014] In some embodiments, the detection step includes determining the identity of the nucleotides incorporated into the copy polynucleotide chain by measuring one or more fluorescent signals from a detectable label. In some embodiments, the sequencing method further includes (d) washing away the chemically removed label and 3'-OH blocking group from the copy polynucleotide chain using a post-cleavage washing solution. In some embodiments, such a washing step also removes unincorporated nucleotides. In other embodiments, the method may include a separate washing step prior to step (b) for washing away unincorporated nucleotides from the copy polynucleotide chain. In some such embodiments, the 3'-OH blocking group and the detectable label of the incorporated nucleotide are removed before introducing the next complementary nucleotide. In some further embodiments, the 3'-OH blocking group and the detectable label are removed in a single step of the chemical reaction. In some embodiments, the sequential incorporation described herein is performed at least 50 times, at least 100 times, at least 150 times, at least 200 times, or at least 250 times.
[0015] Some further embodiments of this disclosure relate to a kit comprising a plurality of nucleotide or nucleoside molecules described herein and packaging materials therefor. The nucleotides, nucleosides, oligonucleotides, or kits described herein may be used to detect, measure, or identify biological systems (e.g., processes or their components). Exemplary techniques in which the compounds, nucleotides, oligonucleotides, or kits may be used include sequencing, expression analysis, hybridization analysis, gene analysis, RNA analysis, cell assays (e.g., cell binding or cell function analysis), or protein assays (e.g., protein binding assays or protein activity assays). Their use may also be automated instruments for performing specific techniques, such as automated sequencing instruments. Sequencing instruments may include two or more lasers operating at different wavelengths to distinguish different detectable labels. [Brief explanation of the drawing]
[0016] [Figure 1] This graph compares the 3'-OH deblocking efficiency of [(allyl)PdCl]2 with that of Na2PdCl4 when using various ratios of tris(hydroxypropyl)phosphine (THP).
[0017] [Figure 2] This graph shows the prephasizing rate as a function of time for a standard fully functionalized nucleotide (ffN) having an LN3 linker moiety and a 3'-O-azidomethyl blocking group, compared to ffN having an AOL linker moiety and a 3'-AOM blocking group.
[0018] [Figure 3] This figure shows a comparison of phasing values in an Illumina MiniSeq® instrument using fully functionalized nucleotides (ffN) with 3'-AOM blocking groups, with and without the use of a palladium scavenger in the post-cleaning process.
[0019] [Figure 4] This figure compares the primary sequencing metrics, including phasing, prephasing, and errors, using fully functionalized nucleotides (ffNs) with 3'-AOM blocking groups and AOL linker moieties with palladium scavenging on Illumina's MiniSeq® instrument, compared to the same sequencing metrics using standard ffNs with 3'-O-azidomethyl blocking groups.
[0020] [Figure 5] This figure shows a comparison of primary sequencing metrics, including phasing and prephasing, on Illumina's MiniSeq® instrument using fully functionalized nucleotides (ffNs) having a 3'-AOM blocking group and an AOL linker moiety, when glycine or ethanolamine is used in the integration mixture, respectively.
[0021] [Figure 6] This figure shows the primary sequencing metrics, including phasing, prephasing, and error rates, on an Illumina MiniSeq® instrument using fully functionalized nucleotides (ffNs) with a 3'-AOM blocking group and glycine in the incorporation mixture, compared to the same sequencing metrics using standard ffNs with a 3'-O-azidomethyl blocking group.
[0022] [Figure 7A] This figure shows the error rate and Q30 sequencing metric for 2 × 300 sequencing runs on an Illumina iSeq™ instrument using fully functionalized nucleotides (ffN) with a 3'-AOM blocking group and an AOL linker moiety, compared to the same sequencing metric using standard ffN with a 3'-O-azidomethyl blocking group. [Figure 7B]This figure shows the error rate and Q30 sequencing metric for 2 × 300 sequencing runs on an Illumina iSeq™ instrument using fully functionalized nucleotides (ffN) with a 3'-AOM blocking group and an AOL linker moiety, compared to the same sequencing metric using standard ffN with a 3'-O-azidomethyl blocking group.
[0023] [Figure 8A] This figure shows the error rate and Q30 sequencing metric for 2 × 150 sequencing runs on an Illumina iSeq™ instrument using fully functionalized nucleotides (ffN) having a 3'-AOM blocking group and an AOL linker moiety, compared to the same sequencing metric using standard ffN with a 3'-O-azidomethyl blocking group. [Figure 8B] This figure shows the error rate and Q30 sequencing metric for 2 × 150 sequencing runs on an Illumina iSeq™ instrument using fully functionalized nucleotides (ffN) having a 3'-AOM blocking group and an AOL linker moiety, compared to the same sequencing metric using standard ffN with a 3'-O-azidomethyl blocking group.
[0024] [Figure 9A] This figure shows the first-order sequencing metric (fading, signal attenuation rate, error rate) as a function of blue laser output when the green laser output is constant. Sequencing experiments were performed on an Illumina NovaSeq™ instrument using fully functionalized nucleotides (ffN) with 3'-AOM blocking groups and AOL linker moieties, using palladium scavenger L-cysteine, compared to the same sequencing metric using a standard protocol and ffN with a 3'-O-azidomethyl blocking group. [Figure 9B]This figure shows the first-order sequencing metric (fading, signal attenuation rate, error rate) as a function of blue laser output when the green laser output is constant. Sequencing experiments were performed on an Illumina NovaSeq™ instrument using fully functionalized nucleotides (ffN) with 3'-AOM blocking groups and AOL linker moieties, using palladium scavenger L-cysteine, compared to the same sequencing metric using a standard protocol and ffN with a 3'-O-azidomethyl blocking group. [Figure 9C] This figure shows the first-order sequencing metric (fading, signal attenuation rate, error rate) as a function of blue laser output when the green laser output is constant. Sequencing experiments were performed on an Illumina NovaSeq™ instrument using fully functionalized nucleotides (ffN) with 3'-AOM blocking groups and AOL linker moieties, using palladium scavenger L-cysteine, compared to the same sequencing metric using a standard protocol and ffN with a 3'-O-azidomethyl blocking group. [Figure 9D] This figure shows the first-order sequencing metric (fading, signal attenuation rate, error rate) as a function of blue laser output when the green laser output is constant. Sequencing experiments were performed on an Illumina NovaSeq™ instrument using fully functionalized nucleotides (ffN) with 3'-AOM blocking groups and AOL linker moieties, using palladium scavenger L-cysteine, compared to the same sequencing metric using a standard protocol and ffN with a 3'-O-azidomethyl blocking group. [Figure 9E]This figure shows the first-order sequencing metric (fading, signal attenuation rate, error rate) as a function of blue laser output when the green laser output is constant. Sequencing experiments were performed on an Illumina NovaSeq™ instrument using fully functionalized nucleotides (ffN) with 3'-AOM blocking groups and AOL linker moieties, using palladium scavenger L-cysteine, compared to the same sequencing metric using a standard protocol and ffN with a 3'-O-azidomethyl blocking group.
[0025] [Figure 10] This figure shows the primary sequencing metrics (%PF, error rate, Q30, and signal decay) on an Illumina iSeq™ instrument (1 × 150 cycles) using fully functionalized (ffN) with a 3'-AOM blocking group and an AOL linker moiety, compared to SBS of the same ffN using standard enzymatic linearization. The same palladium cleavage mixture was also used in the first step of chemical linearization of SBS. [Modes for carrying out the invention]
[0026] Embodiments of this disclosure relate to nucleosides and nucleotides having a 3' acetal blocking group for sequencing applications, such as synthetic sequencing (SBS). In some embodiments, the nucleoside or nucleotide includes a label covalently bonded thereto via a cleavable linker containing an acetal moiety that allows for cleavage of the 3' acetal blocking group and the label in a single step of the reaction. The 3' acetal blocking group also provides improved stability during the synthesis of fully functionalized nucleotides (ffNs), as well as greater stability in solution during formulation, storage, and operation in sequencing instruments. Furthermore, the 3' acetal blocking groups described herein can also achieve lower prephasing, lower signal attenuation for improved data quality, and longer reads from sequencing applications. definition
[0027] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art. The use of the term "including," and other forms such as "include," "includes," and "included," is not limited. Furthermore, the use of the term "having," and other forms such as "have," "has," and "had," is not limited. When used herein, in transitional clauses or in the text of claims, the terms "comprise" and "comprising" should be interpreted as having an open-ended meaning; that is, the above terms should be interpreted as synonymous with the phrases "at least have" or "at least include." For example, when used in the context of a process, the term "comprising" means that the process includes at least the listed steps, but may include additional steps. When used in the context of compounds, compositions, or devices, the term "comprising" means that the compound, composition, or device includes at least the listed features or components, but may also include additional features or components.
[0028] If a range of values is provided, it is understood that the upper and lower limits, as well as each intermediate value between the upper and lower limits of the range, are included within the embodiment.
[0029] As used herein, common organic abbreviations are defined as follows: [Table 1]
[0030] As used herein, the term “array” refers to a collection of different probe molecules bound to one or more substrates such that the different probe molecules can be differentiated from one another according to their relative positions. An array may include different probe molecules, each located at different identifiable positions on the substrate. Alternatively or additionally, an array may include separate substrates, each having a different probe molecule, which can be identified according to the position of the substrate on the surface to which it is attached, or according to the position of the substrate in a liquid. Exemplary arrays in which separate substrates are located on a surface include, but are not limited to, those containing beads in wells, as described in U.S. Patent No. 6,355,431, U.S. Patent Application Publication No. 2002 / 0102578, and International Publication No. WO00 / 63437. Exemplary formats that can be used in the present invention to distinguish beads in a liquid array using a microfluidic device such as a fluorescence-activated cell sorter (FACS) are described, for example, in U.S. Patent No. 6,524,793. Further examples of arrays that can be used in the present invention include, but are not limited to, U.S. Patent Nos. 5,429,807, 5,436,327, 5,561,071, 5,583,211, 5,658,734, 5,837,858, 5,874,219, 5,919,523, 6,136,269, 6,287,768, and 6,287,776. This includes U.S. Patent Nos. 6,288,220, 6,297,006, 6,291,193, 6,346,413, 6,416,949, 6,482,591, 6,514,751 and 6,610,482, as well as International Publication Nos. 93 / 17126, 95 / 11995, 95 / 35505, European Patent No. 742287 and European Patent No. 799897.
[0031] As used herein, the term "covalently attached" (or "covalently bonded") refers to the formation of a chemical bond characterized by the sharing of electron pairs between atoms. For example, a covalently attached polymer coating refers to a polymer coating that forms a chemical bond with the functionalized surface of a substrate, as compared to attachment to the surface by other means, such as adhesion or electrostatic interaction. It will be understood that a polymer covalently attached to a surface may be attached by means in addition to covalent bonding.
[0032] As used herein, any "R" group(s) represents a substituent capable of bonding to the indicated atom. The R group(s) can be substituted or unsubstituted. When two "R" groups are described as forming a ring or ring system "together with the atom to which they are attached", it means a ring in which the atoms, intervening bonds, and the collective unit of the two R groups are enumerated. For example, when the following substructure is present,
Chemical formula
Chemical formula
[0033] It should be understood that certain radical nomenclature can include either monoradicals or diradicals depending on the context. For example, a substituent is understood to be a diradical if it requires two attachment sites to the rest of the molecule. Examples of substituents identified as alkyls that require two attachment sites include diradicals such as -CH2-, -CH2CH2-, and -CH2CH(CH3)CH2-. Other radical nomenclature explicitly indicates that the radical is a diradical, such as an "alkylene" or "alkenylene".
[0034] As used herein, the term "halogen" or "halo" means one of the elements of Group 7 of the periodic table, for example, fluorine, chlorine, bromine, or iodine, with fluorine and chlorine being preferred.
[0035] As used herein, "a" and "b" are integers. a ~C b"C1-C4 alkyl" refers to the number of carbon atoms in an alkyl, alkenyl, or alkynyl group, or the number of ring atoms in a cycloalkyl or aryl group. That is, alkyl, alkenyl, alkynyl, cycloalkyl rings, and aryl rings can contain carbon atoms at both ends, "a" to "b". For example, "C1-C4 alkyl" groups refer to all alkyl groups having 1 to 4 carbon atoms, i.e., CH3-, CH3CH2-, CH3CH2CH2-, (CH3)2CH-, CH3CH2CH2CH2-, CH3CH2CH(CH3)-, and (CH3)3C-, and "C3-C4 cycloalkyl" refers to all cycloalkyl groups having 3 to 4 carbon atoms, i.e., cyclopropyl and cyclobutyl. Similarly, "4-6 membered heterocyclyl" groups refer to all heterocyclyl groups having 4 to 6 whole ring atoms, such as azetidine, oxetane, oxazoline, pyrrolidine, piperidine, piperazine, morpholine, etc. If “a” and “b” are not specified with respect to alkyl, alkenyl, alkynyl, cycloalkyl, or aryl groups, the broadest range described in those definitions should be assumed. As used herein, the term “C1-C6” includes the range defined by C1, C2, C3, C4, C5, and C6, and either of two numbers. For example, C1-C6 alkyl includes C1, C2, C3, C4, C5, and C6 alkyl, C2-C6 alkyl, C1-C3 alkyl, etc. Similarly, C2-C6 alkenyl includes C2, C3, C4, C5, and C6 alkenyl, C2-C5 alkenyl, C3-C4 alkenyl, etc., and C2-C6 alkynyl includes C2, C3, C4, C5, and C6 alkynyl, C2-C5 alkynyl, C3-C4 alkynyl, etc. C3-C8 cycloalkyls each include a hydrocarbon ring containing 3, 4, 5, 6, 7, and 8 carbon atoms, or a range defined by either two numbers, such as C3-C7 cycloalkyl or C5-C6 cycloalkyl.
[0036] As used herein, “alkyl” refers to a fully saturated (i.e., non-double or non-triple) straight or branched hydrocarbon chain. Alkyl groups may have 1 to 20 carbon atoms (wherein always, when shown herein, numerical ranges such as “1 to 20” refer to each integer within a given range. For example, “1 to 20 carbon atoms” means that an alkyl group may consist of up to 20 carbon atoms, such as 1 carbon atom, 2 carbon atoms, 3 carbon atoms, etc., but this definition also means that it covers occurrences of the term “alkyl” (without specifying a numerical range). Alkyl groups may also be medium alkyl groups having 1 to 9 carbon atoms. Alkyl groups may also be lower alkyl groups having 1 to 6 carbon atoms. Alkyl groups may be designated as “C1 to C4 alkyl” or similar notation. For example, “C1 to C6 alkyl” indicates that there are 1 to 6 carbon atoms in the alkyl chain, i.e., the alkyl chain is selected from the group consisting of methyl, ethyl, propyl, isopropyl, n-butyl, iso-butyl, sec-butyl, and t-butyl. Typical alkyl groups include, but are not limited to, methyl, ethyl, propyl, isopropyl, butyl, isobutyl, tertiary butyl, pentyl, and hexyl.
[0037] As used herein, "alkoxy" refers to a C1-C9 alkoxy or other alkyl group of the formula -OR, where R is an alkyl group as defined above, and includes, but is not limited to, methoxy, ethoxy, n-propoxy, 1-methylethoxy (isopropoxy), n-butoxy, isobutoxy, sec-butoxy, and tert-butoxy.
[0038] As used herein, “alkenyl” refers to a linear or branched hydrocarbon chain containing one or more double bonds. An alkenyl group may have 2 to 20 carbon atoms, but this definition also covers the occurrence of the term “alkenyl” without a specified numerical range. An alkenyl group may also be a medium alkenyl having 2 to 9 carbon atoms. An alkenyl group may also be a lower alkenyl having 2 to 6 carbon atoms. An alkenyl group may be designated as “C2-C6 alkenyl” or similar notation. As just one example, "C2-C6 alkenyl" indicates that there are 2 to 6 carbon atoms in the alkenyl chain. That is, the alkenyl chain is selected from the group consisting of ethenyl, propen-1-yl, propen-2-yl, propen-3-yl, buten-1-yl, buten-2-yl, buten-3-yl, buten-4-yl, 1-methyl-propen-1-yl, 2-methyl-propen-1-yl, 1-ethyl-ethen-1-yl, 2-methyl-propen-3-yl, buta-1,3-dienyl, buta-1,2-dienyl, and buta-1,2-dien-4-yl. Typical alkenyl groups include ethenyl, propenyl, butenyl, pentenyl, and hexenyl, but are not limited to these.
[0039] As used herein, “alkynyl” refers to a linear or branched hydrocarbon chain containing one or more triple bonds. An alkynyl group may have 2 to 20 carbon atoms, but this definition also covers the occurrence of the term “alkynyl” without a specified numerical range. An alkynyl group may also be a medium alkynyl having 2 to 9 carbon atoms. An alkynyl group may also be a lower alkynyl having 2 to 6 carbon atoms. An alkynyl group may be designated as “C2-C6 alkynyl” or a similar notation. For example, “C2-C6 alkynyl” indicates that there are 2 to 6 carbon atoms in the alkynyl chain, i.e., the alkynyl chain is selected from the group consisting of ethynyl, propyne-1-yl, propyne-2-yl, butyne-1-yl, butyne-3-yl, butyne-4-yl, and 2-butynyl. Typical alkynyl groups include, but are not limited to, ethynyl, propynyl, butynyl, pentynyl, and hexynyl.
[0040] As used herein, “heteroalkyl” refers to a linear or branched hydrocarbon chain containing one or more heteroatoms, the heteroatoms being atoms other than carbon in the main chain, such as nitrogen, oxygen, and sulfur. A heteroalkyl group may have 1 to 20 carbon atoms, but this definition also covers the occurrence of the term “heteroalkyl” without a specified numerical range. A heteroalkyl group may also be a medium-sized heteroalkyl group having 1 to 9 carbon atoms. A heteroalkyl group may also be a lower heteroalkyl group having 1 to 6 carbon atoms. A heteroalkyl group may be designated as “C1-C6 heteroalkyl” or a similar notation. A heteroalkyl group may contain one or more heteroatoms. For example, “C4-C6 heteroalkyl” indicates that there are 4 to 6 carbon atoms in the heteroalkyl chain, and one or more heteroatoms in the main chain.
[0041] The term "aromatic" refers to a ring or ring system that has a conjugated π-electron system and contains both carbocyclic aromatic groups (e.g., phenyl) and heterocyclic aromatic groups (e.g., pyridine). This term includes monocyclic or fused polycyclic (i.e., rings sharing adjacent pairs of atoms) groups, as long as the entire ring system is aromatic.
[0042] As used herein, “aryl” refers to an aromatic ring or ring system (i.e., two or more fused rings sharing two adjacent carbon atoms) containing only carbon atoms in its ring skeleton. If an aryl is a ring system, all rings in the system are aromatic rings. While an aryl group may have 6 to 18 carbon atoms, this definition also covers occurrences of the term “aryl” where no numerical range is specified. In some embodiments, an aryl group has 6 to 10 carbon atoms. 10 "Aryl," "C6 or C" 10 It may be specified as "aryl" or a similar notation. Examples of aryl groups include, but are not limited to, phenyl, naphthyl, azlenyl, and anthracenyl.
[0043] "Aralkyl" or "arylalkyl" is "C 7~14 These are aryl groups bonded as substituents via an alkylene group, such as "aralkyl," and include, but are not limited to, benzyl, 2-phenylethyl, 3-phenylpropyl, and naphthylalkyl groups. In some cases, the alkylene group is a lower alkylene group (i.e., a C1-C6 alkylene group).
[0044] As used herein, “heteroaryl” refers to an aromatic ring or ring system (i.e., two or more fused rings sharing two adjacent atoms) containing one or more heteroatoms, i.e., elements other than carbon, including but not limited to nitrogen, oxygen, and sulfur in its ring skeleton. If the heteroaryl is a ring system, all rings in the system are aromatic rings. A heteroaryl group may have 5 to 18 ring members (i.e., the number of atoms constituting the ring skeleton, including carbon atoms and heteroatoms), but this definition also covers occurrences of the term “heteroaryl” where no numerical range is specified. In some embodiments, a heteroaryl group has 5 to 10 ring members or 5 to 7 ring members. A heteroaryl group may be designated as a “5-7 membered heteroaryl,” a “5-10 membered heteroaryl,” or similar notation. Examples of heteroaryl rings include, but are not limited to, furyl, thienyl, phthalazinyl, pyrrolyl, oxazolyl, thiazolyl, imidazolyl, pyrazolyl, isoxazolyl, isothiazolyl, triazolyl, thiadiazolyl, pyridinyl, pyridadinyl, pyrimidinyl, pyrazinyl, triazinyl, quinolinyl, isoquinolinyl, benzimidazolyl, benzoxazolyl, benzothiazolyl, indolyl, isoindolyl, and benzothienyl.
[0045] A "heteroaralkyl" or "heteroarylalkyl" is a heteroaryl group bonded as a substituent via an alkylene group. Examples include, but are not limited to, 2-thienylmethyl, 3-thienylmethyl, furylmethyl, thienylethyl, pyrrolylalkyl, pyridylalkyl, isoxazolylylalkyl, and imidazolylalkyl. In some cases, the alkylene group is a lower alkylene group (i.e., a C1-C6 alkylene group).
[0046] As used herein, "carbocyclyl" means a non-aromatic ring or ring system containing only carbon atoms in its ring framework. When a carbocyclyl is a ring system, two or more rings may be bonded together by condensation, bridging, or spirobonding. A carbocyclyl may have any degree of saturation, provided that at least one ring in the ring system is not aromatic. Thus, examples of carbocyclyls include cycloalkyls, cycloalkenyls, and cycloalkynyls. A carbocyclyl group may have 3 to 20 carbon atoms, but this definition also covers the occurrence of the term "carbocyclyl" without a specified numerical range. A carbocyclyl group may also be a medium-sized carbocyclyl having 3 to 10 carbon atoms. A carbocyclyl group may also be a carbocyclyl having 3 to 6 carbon atoms. A carbocyclyl group may be designated as "C3-C6 carbocyclyl" or similar notation. Examples of carbocyclyl rings include, but are not limited to, cyclopropyl, cyclobutyl, cyclopentyl, cyclohexyl, cyclohexenyl, 2,3-dihydroindene, bisicle[2.2.2]octanyl, adamantyl, and spiro[4.4]nonanyl.
[0047] As used herein, "cycloalkyl" means a fully saturated carbocyclyl ring or ring system. Examples include cyclopropyl, cyclobutyl, cyclopentyl, and cyclohexyl.
[0048] As used herein, “heterocyclyl” means a non-aromatic ring or ring system containing at least one heteroatom in its ring skeleton. Heterocyclyls may be joined together integrally by condensation, bridging, or spirobonding. Heterocyclyls may have any degree of saturation, provided that at least one ring in the ring system is not aromatic. The heteroatom may be present in either the non-aromatic or aromatic ring in the ring system. A heterocyclyl group may have 3 to 20 ring members (i.e., the number of atoms constituting the ring skeleton, including carbon atoms and heteroatoms), but this definition also covers occurrences of the term “heterocyclyl” without a specified numerical range. A heterocyclyl group may also be a medium-sized heterocyclyl having 3 to 10 ring members. A heterocyclyl group may also be a heterocyclyl having 3 to 6 ring members. A heterocyclyl group may be designated as a “3-6 membered heterocyclyl” or similar notation. In a preferred six-membered monocyclic heterocycline, the heteroatoms are selected from one to three of O, N, or S. In a preferred five-membered monocyclic heterocycline, the heteroatoms are selected from one or two of the heteroatoms selected from O, N, or S.Examples of heterocyclyl rings include azepinyl, acridinyl, carbazolyl, cinolinyl, dioxolanil, imidazolinyl, imidazolidinyl, morpholinyl, oxylanil, oxepanil, thiapanil, piperidinyl, piperazinyl, dioxapiperazinyl, pyrrolidinyl, pyrrolidinyl, pyrrolidionyl, 4-piperidonyl, pyrazolinyl, pyrazolidinyl, 1,3-dioxynyl, 1,3-dioxanyl, 1,4-dioxynyl, 1,4-dioxanyl, 1,3-oxatinyl, 1,4-oxatinyl, 1,4-oxathianyl, 2H-1,2-oxazinyl, trioxanil, hexahydro-1,3 Examples include, but are not limited to, 5-triazinyl, 1,3-dioxolyl, 1,3-dioxolanil, 1,3-dithiolyl, 1,3-dithiolanil, isoxazolinyl, isoxazolidinyl, oxazolinyl, oxazolidinyl, oxazolidinyl, oxazolidinyl, thiazolinyl, thiazolidinyl, 1,3-oxathiolanil, indolinyl, isoindolinyl, tetrahydrofuranil, tetrahydropyranil, tetrahydrothiophenyl, tetrahydrothiopyranil, tetrahydro-1,4-thiadinyl, thiamorpholinyl, dihydrobenzofuranil, benzimidazolidinyl, and tetrahydroquinoline.
[0049] As used herein, "alkoxyalkyl" or "(alkoxy)alkyl" refers to an alkoxy group bonded via an alkylene group, such as a C2-C8 alkoxyalkyl or a (C1-C6 alkoxy)C1-C6 alkyl, such as -(CH2) 1-3 - Refers to OCH3
[0050] As used herein, "-O-alkoxyalkyl" or "-O-(alkoxy)alkyl" refers to an alkoxy group bonded via an -O-(alkylene) group, such as -O-(C1~C6alkoxy)C1~C6alkyl, such as -O-(CH2) 1-3 - Refers to OCH3
[0051] As used herein, "(heterocyclyl)alkyl" refers to a heterocyclic or heterocyclyl group as defined above, bonded as a substituent via an alkylene group as defined above. The alkylene and heterocyclyl groups of (heterocyclyl)alkyl may be substituted or unsubstituted. Examples include, but are not limited to, tetrahydro-2H-pyran-4-yl)methyl, (piperidine-4-yl)ethyl, (piperidine-4-yl)propyl, (tetrahydro-2H-thiopyran-4-yl)methyl, and (1,3-thiadinan-4-yl)methyl.
[0052] As used herein, "(cycloalkyl)alkyl" or "(carbocykyl)alkyl" refers to a cycloalkyl or carbocykyl group (as defined herein) bonded as a substituent via an alkylene group. Examples include, but are not limited to, cyclopropylmethyl, cyclobutylmethyl, cyclopentylethyl, and cyclohexylpropyl.
[0053] The "O-carboxyl" group refers to the "-OC(=O)R" group, where R is defined as hydrogen, C1-C6 alkyl, C2-C6 alkenyl, C2-C6 alkynyl, C3-C7 carbocykyl, or C6-C as defined herein. 10 The selection is made from aryls, 5-10 membered heteroaryls, and 3-10 membered heterocyclines.
[0054] The "C-carboxyl" group refers to the "-C(=O)OR" group, where R is hydrogen, C1-C6 alkyl, C2-C6 alkenyl, C2-C6 alkynyl, C3-C7 carbocykyl, or C6-C as defined herein. 10 The group is selected from aryls, 5-10 membered heteroaryls, and 3-10 membered heterocyclines. A non-restrictive example is carboxyl (i.e., -C(=O)OH).
[0055] The "sulfonyl" group refers to the "-SO2R" group, where R is hydrogen, C1-C6 alkyl, C2-C6 alkenyl, C2-C6 alkynyl, C3-C7 carbocyclyl, or C6-C as defined herein. 10 The selection is made from aryls, 5-10 membered heteroaryls, and 3-10 membered heterocyclines.
[0056] The "sulfino" group refers to the "-S(=O)OH" group.
[0057] The "S-sulfonamide" group is "-SO2NR A R B " refers to the base, and in the formula, R A and R B Each of these independently represents hydrogen, C1-C6 alkyl, C2-C6 alkenyl, C2-C6 alkynyl, C3-C7 carbocykyl, and C6-C as defined herein. 10 The selection is made from aryls, 5-10 membered heteroaryls, and 3-10 membered heterocyclines.
[0058] The "N-sulfonamide" group is "-N(R A )SO2R B " refers to the base, R A and R b Each of these independently comprises hydrogen, C1-C6 alkyl, C2-C6 alkenyl, C2-C6 alkynyl, C3-C7 carbocykyl, and C6-C as defined herein. 10 The selection is made from aryls, 5-10 membered heteroaryls, and 3-10 membered heterocyclines.
[0059] The "C-amide" group is "-C(=O)NR A R B " refers to the base, R A and R B Each of these independently represents hydrogen, C1-C6 alkyl, C2-C6 alkenyl, C2-C6 alkynyl, C3-C7 carbocykyl, and C6-C as defined herein. 10 The selection is made from aryls, 5-10 membered heteroaryls, and 3-10 membered heterocyclines.
[0060] The "N-amide" group is "-N(RA )C(=O)R B " refers to the base, R A and R B Each of these independently represents hydrogen, C1-C6 alkyl, C2-C6 alkenyl, C2-C6 alkynyl, C3-C7 carbocykyl, and C6-C as defined herein. 10 The selection is made from aryls, 5-10 membered heteroaryls, and 3-10 membered heterocyclines.
[0061] The "amino" group is "-NR A R B " refers to the base, R A and R B Each of these independently represents hydrogen, C1-C6 alkyl, C2-C6 alkenyl, C2-C6 alkynyl, C3-C7 carbocykyl, and C6-C as defined herein. 10 The compounds are selected from aryls, 5- to 10-membered heteroaryls, and 3- to 10-membered heterocyclines. A non-limiting example is free aminos (i.e., -NH2).
[0062] The "aminoalkyl" group refers to an amino group linked via an alkylene group.
[0063] The "(alkoxy)alkyl" group refers to an alkoxy group bonded via an alkylene group, such as "(C1-C6 alkoxy)C1-C6 alkyl".
[0064] As used herein, the term "hydroxy" refers to the -OH group.
[0065] As used herein, the term "cyano" group refers to the "-CN" group.
[0066] As used herein, the term "azide" refers to the -N3 group.
[0067] As used herein, the term "propargylamine" refers to an amino group substituted with a propargyl group. [ka] When propargylamine is used in context as the divalent moiety, [ka] Includes, and in the formula, where defined herein, R A C1-C6 alkyl, C2-C6 alkenyl, C2-C6 alkynyl, C3-C7 carbocykyl, C6-C 10 This includes aryls, 5- to 10-membered heteroaryls, and 3- to 10-membered heterocyclines.
[0068] As used herein, the term "propargylamide" refers to the propargyl group [ka] This refers to a C-amide or N-amide group substituted with [a specific substituted group]. When propargylamide is used in context as a divalent moiety, it means that [the substituted group] [ka] Includes, and as defined herein, R A C1-C6 alkyl, C2-C6 alkenyl, C2-C6 alkynyl, C3-C7 carbocykyl, C6-C 10 This includes aryls, 5-10 membered heteroaryls, and 3-10 membered heterocyclines.
[0069] As used herein, the term "allylamine" refers to an amino group substituted with an allyl group (CH2=CH-CH2-). When allylamine is used in context as a divalent moiety, it is -CH=CH-CH2-NR A - includes, in the formula, R A These are defined as hydrogen, C1-C6 alkyl, C2-C6 alkenyl, C2-C6 alkynyl, C3-C7 carbocykyl, C6-C 10 These include aryls, 5-10 membered heteroaryls, and 3-10 membered heterocyclines.
[0070] As used herein, the term "allylamide" refers to a C-amide or N-amide group substituted with an allyl group (CH2=CH-CH2-). When allylamide is used in context as a divalent moiety, it has the -CH=CH-CH2-NR as defined herein. A -C(=O)- or -CH=CH-CH2-C(=O)-NR A - is included, in the formula, R A C1-C6 alkyl, C2-C6 alkenyl, C2-C6 alkynyl, C3-C7 carbocyryl, C6-C 10 These include aryls, 5-10 membered heteroaryls, and 3-10 membered heterocyclines.
[0071] Where a group is described as “optionally substituted,” it may be either unsubstituted or substituted. Similarly, where a group is described as “substituted,” the substituent may be selected from one or more of the substituents indicated. As used herein, substituents are derived from an unsubstituted parent group in which one or more hydrogen atoms are exchanged for another atom or group. Unless otherwise stated, where a group is considered “substituted,” it means that the group is substituted with one or more substituents independently selected from the following: C1-C6 alkyl, C1-C6 alkenyl, C1-C6 alkynyl, C1-C6 heteroalkyl, C3-C7 carbocyclyl (optionally substituted with halo, C1-C6 alkyl, C1-C6 alkoxy, C1-C6 haloalkyl, and C1-C6 haloalkoxy), C3-C7 carbocyclyl-C1-C6-alkyl (optionally substituted with halo, C1-C6 alkyl, C1-C6 alkoxy, C1-C6 haloalkyl, and (substituted with C1-C6 haloalkoxy), 3-10 member heterocyclyl (optionally substituted with halo, C1-C6 alkyl, C1-C6 alkoxy, C1-C6 haloalkyl, and C1-C6 haloalkoxy), 3-10 member heterocyclyl-C1-C6-alkyl (optionally substituted with halo, C1-C6 alkyl, C1-C6 alkoxy, C1-C6 haloalkyl, and C1-C6 haloalkoxy), aryl (Optionally substituted with halo, C1-C6 alkyl, C1-C6 alkoxy, C1-C6 haloalkyl, and C1-C6 haloalkoxy), (aryl)C1-C6 alkyl (Optionally substituted with halo, C1-C6 alkyl, C1-C6 alkoxy, C1-C6 haloalkyl, and C1-C6 haloalkoxy), 5-10 member heteroaryl (Optionally substituted with halo, C1-C6 alkyl, C1-C6 alkoxy, C1- (substituted with C6 haloalkyl and C1-C6 haloalkoxy), (5-10 member heteroaryl)C1-C6 alkyl (optionally substituted with halo, C1-C6 alkyl, C1-C6 alkoxy, C1-C6 haloalkyl and C1-C6 haloalkoxy), halo, -CN, hydroxy, C1-C6 alkoxy, (C1-C6 alkoxy)C1-C6 alkyl, -O(C1-C6 alkoxy)C1-C6 alkyl;(C1~C6 haloalkoxy)C1~C6 alkyl;-O(C1~C6 haloalkoxy)C1~C6 alkyl;aryloxy, sulfhydryl (mercapto), halo(C1~C6)alkyl (e.g., -CF3), halo(C1~C6)alkoxy (e.g., -OCF3), C1~C6 alkylthio, arylthio, amino, amino(C1~C6)alkyl, nitro, O-carbamyl, N-carbamyl, O-thiocarbamyl, N-thiocarbamyl, C-amide, N-amide, S-sulfonamide, N-sulfonamide, C-carboxy, O-carboxy, acyl, cyanato, isocyanato, thiocyanato, sulfinyl, sulfonyl, -SO3H, sulfino, -OSO2C; 1~4 Alkyl, monophosphate, diphosphate, triphosphosphate, and oxo (=O). If a group is described as "optionally substituted," then if that group is substituted, it may be substituted with one of the above substituents.
[0072] When a substituent is shown as a diradical (i.e., having two bonding points to the rest of the molecule), it should be understood that unless otherwise specified, the substituent can be bonded in any directional configuration. For example, -AE- or [ka] The substituents shown as indicate substituents that are oriented so that A is attached at the leftmost bond point of the molecule, as well as the case where A is attached at the rightmost bond point of the molecule. Furthermore, the group or substituent is [ka] When shown as such, L is defined as a linker portion that may or may be present; if L is absent (or not present), such a group or substituent is [ka] It is equivalent to this.
[0073] As used herein, “nucleotide” comprises a nitrogen-containing heterocyclic base, a sugar, and one or more phosphate groups. They are monomeric units of nucleic acid sequences. In RNA, the sugar is ribose, and in DNA, it is deoxyribose, i.e., a sugar lacking the hydroxyl group present in ribose. The nitrogen-containing heterocyclic base may be a purine base or a pyrimidine base. Examples of purine bases include adenine (A) and guanine (G), and their modified derivatives or analogs. Examples of pyrimidine bases include cytosine (C), thymine (T), and uracil (U), and their modified derivatives or analogs. The C-1 atom of deoxyribose is bonded to N-1 of the pyrimidine or N-9 of the purine.
[0074] As used herein, “nucleoside” refers to a molecule structurally similar to a nucleotide but lacking a phosphate moiety. Examples of nucleoside analogs include those in which the label is linked to a base and there is no phosphate group attached to a sugar molecule. The term “nucleoside” is used herein in the ordinary sense as understood by those skilled in the art. Examples include, but are not limited to, ribonucleosides containing a ribose moiety and deoxyribonucleosides containing a deoxyribose moiety. A modified pentose moiety is a pentose moiety in which an oxygen atom is replaced by a carbon atom and / or a carbon atom is replaced by a sulfur or oxygen atom. A “nucleoside” is a monomer that may have a substituted base and / or sugar moiety. Furthermore, nucleosides can be incorporated into larger DNA and / or RNA polymers and oligomers.
[0075] The term “purine base” is used herein in the ordinary sense as understood by those skilled in the art, and includes its tautomers. Similarly, the term “pyrimidine base” is used herein in the ordinary sense as understood by those skilled in the art, and includes its tautomers. An unrestricted list of optionally substituted purine bases includes purines, deazapurines, adenines, 7-deazaadenines, guanines, 7-deazaguanines, hypoxanthines, xanthines, alloxanthines, 7-alkylguanines (e.g., 7-methylguanine), theobromine, caffeine, uric acid, and isoguanines. Examples of pyrimidine bases include, but are not limited to, cytosine, thymine, uracil, 5,6-dihydrouracil, and 5-alkylcytosine (e.g., 5-methylcytosine).
[0076] When used herein, if an oligonucleotide or polynucleotide is described as “containing” or “labeled with” a nucleoside or nucleotide described herein, it means that the nucleoside or nucleotide described herein forms a covalent bond with the oligonucleotide or polynucleotide. Similarly, if a nucleoside or nucleotide is described as part of an oligonucleotide or polynucleotide, such as “integrated” into the oligonucleotide or polynucleotide, it means that the nucleoside or nucleotide described herein forms a covalent bond with the oligonucleotide or polynucleotide. In some such embodiments, the covalent bond is formed between the 3' hydroxyl group of the oligonucleotide or polynucleotide and the 5' phosphate group of the nucleotide described herein, as a phosphodiester bond between the 3' carbon atom of the oligonucleotide or polynucleotide and the 5' carbon atom of the nucleotide.
[0077] As used herein, the term “cleavable linker” does not mean that the entire linker must be removed. The cleavage site can be located on the linker in such a way that a portion of the linker remains bound to a detectable label and / or nucleoside or nucleotide moiety after cleavage.
[0078] As used herein, “derivative” or “analog” means a synthetic nucleotide or nucleoside derivative having a modified base moiety and / or a modified sugar moiety. Such derivatives and analogs are discussed, for example, in Scheit, Nucleotide Analogs (John Wiley & Son, 1980) and Uhlman et al., Chemical Reviews 90:543-584, 1990. Nucleoside analogs may also include modified phosphodiester linkages, including phosphorothioates, phosphorodithioates, alkyl phosphonates, phosphoranilidates, and phosphoramidate linkages. As used herein, “derivative,” “analog,” and “modified” may be used interchangeably and are encompassed within the terms “nucleotide” and “nucleoside” as defined herein.
[0079] As used herein, the term "phosphate" is used in the ordinary sense as understood by those skilled in the art, and refers to its protonated form (e.g., [ka] ) includes. As used herein, the terms “monophosphate,” “diphosphate,” and “triphosphate” are used in the ordinary sense as understood by those skilled in the art, and include protonated forms.
[0080] As used herein, the terms “protecting group” and “blocking group” refer to any atom or atomic group added to a molecule to prevent existing groups in the molecule from undergoing undesirable chemical reactions. Sometimes, “protecting group” and “blocking group” may be used interchangeably.
[0081] As used herein, the term “phasing” refers to a phenomenon in synthetic sequencing (SBS) caused by incomplete removal of the 3' terminator and fluorophore, and incomplete integration of a portion of the DNA strand within a cluster by polymerase in a given sequencing cycle. Prephasing is caused by the integration of a nucleotide without an effective 3' terminator, and the integration event advances one cycle due to the incomplete completion. Phase and prephasing ensure that the measured signal intensity for a particular cycle consists of the signal from the current cycle, as well as noise from previous and next cycles. As the number of cycles increases, the proportion of sequences per cluster affected by phase and prephasing increases, hindering the identification of the correct base. Prephasing can be caused by the presence of trace amounts of unprotected or unblocked 3'-OH nucleotides in synthetic sequencing (SBS). Unprotected 3'-OH nucleotides can be generated during the manufacturing process, or possibly during the storage and reagent handling processes. Therefore, the discovery of nucleotide analogs that reduce the incidence of prephasing is remarkable and offers significant advantages in SBS applications compared to existing nucleotide analogs. For example, the provided nucleotide analogs may result in faster SBS cycle times, lower phasing and prephasing values, and longer sequencing read lengths.
[0082] Nucleosides or nucleotides having a 3'-acetal blocking group Some embodiments of the present disclosure relate to nucleotide or nucleoside molecules comprising a cleavable linker and a nucleic acid base bound to a detectable label via a ribose or deoxyribose moiety, wherein the cleavable linker has the following structure: [ka] Each of X and Y is independently O or S, and R 1a , R 1b , R 2 , R 3a and R 3bEach of these is independently H, a halogen, an unsubstituted or substituted C1-C6 alkyl, or a C1-C6 haloalkyl. In some embodiments, the ribose or deoxyribose moiety comprises the 3'-OH protecting group described herein. In some embodiments, the cleavable linker is L 1 Or L 2 , or may further include both, L 1 or L 2 Details are provided below.
[0083] In some embodiments, the nucleoside or nucleotide described herein has the structure of formula (I), [ka] In the formula, B is a nucleic acid base. R 4 is H or OH, R 5 This is an H, 3'-OH blocking group, or a phosphoramidite. R 6 This is H, a monophosphate, a diphosphate, a triphosphosphate, a thiophosphate, a phosphate ester analog, a reactive phosphorus-containing group, or a hydroxyprotecting group. L is [ka] And, L 1 and L 2 Each of these is an independent, optional linker component.
[0084] In some embodiments of the severable linker portion described herein, each of X and Y is O. In some other embodiments, X is S and Y is O, or X is O and Y is S. In some embodiments, R 1a , R 1b , R 2 , R 3a and R 3b Each of these is H. In other embodiments, R 1a , R1b 、 R 2 、 R 3a and R 3b at least one of which is halogen (e.g., fluoro, chloro) or unsubstituted C1-C6 alkyl (e.g., methyl, ethyl, isopropyl, isobutyl or t-butyl). In some such examples, R 1a and R 1b each is H, and R 2 、 R 3a and R 3b at least one of which is unsubstituted C1-C6 alkyl or halogen (e.g., R 2 is unsubstituted C1-C6 alkyl, R 3a and R 3b each is H, or R 2 is H, and R 3a and R 3b one or both of which are halogen or unsubstituted C1-C6 alkyl). In one embodiment, the cleavable linker or L is
Chemical formula
[0085] In some embodiments of the nucleosides or nucleotides described herein, the nucleobase ("B" in formula (I)) is a purine (adenine or guanine), a deazapurine or a pyrimidine (e.g., cytosine, thymine or uracil). In some further embodiments, the deazapurine is 7-deazapurine (e.g., 7-deazaadenine or 7-deazaguanine). Non-limiting examples of B are
Chemical formula
Chemical formula
[0086] In some embodiments of the nucleosides or nucleotides described herein, the ribose or deoxyribose moiety comprises a 3'-OH blocking group (i.e., R in formula (I) 5 is a 3'-OH blocking group). In some embodiments, the 3'-OH blocking group or R 5 is
Chemical formula
Chemical formula
Chemical formula
Chemical formula
[0087] In some other embodiments of the nucleosides or nucleotides described herein, R 5 In formula (I), is a phosphoramidite. In such embodiments, R 6 This is an acid-cleavable hydroxy protecting group that enables subsequent monomer coupling under automated synthesis conditions.
[0088] In some embodiments of the nucleoside or nucleotide described herein, L 1 L 1 This includes a portion selected from the group consisting of propargylamine, propargylamide, allylamine, allylamide, and optionally substituted variants thereof. In some further embodiments, L 1 teeth, [ka] Includes or is. In some further embodiments, asterisk * This is L for nucleic acid bases (e.g., the C5 position of pyrimidine bases or the C7 position of 7-deazapurine bases). 1 This shows the connection point.
[0089] In some embodiments, the nucleotides described herein are fully functionalized nucleotides (ffNs) comprising a dye compound covalently bonded to a nucleic acid base via a 3'-OH blocking group and a cleavable linker as described herein, wherein the cleavable linker is structural [ka] L 1 Includes, * L 1This indicates a binding site to a nucleic acid base (e.g., the C5 position of a cytosine, thymine, or uracil base, or the C7 position of 7-deazaadenine or 7-deazaguanine). In some cases, ffNs having the allylamine or allylamide linker moiety described herein are also called ffN-DB or ffN-(DB), where "DB" refers to the double bond of the linker moiety. In some cases, performing sequencing using ffN sets (including ffA, ffT, ffC, ffG) in which one or more ffNs are ffN-DB provides a superior incorporation rate of ffNs compared to ffN sets using propargylamine or propargylamide linker moieties described herein (also known as ffN-PA or ffN-(PA)). For example, the ffNs-DB set using the allylamine or allylamide linker moiety and 3'-AOM blocking group described herein results in an improvement in the integration rate of at least 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 300%, 400%, or 500% compared to the ffNs-PA set using the 3'-O-azidomethyl blocking group under the same conditions and for the same period, thereby improving the phasing value. In other embodiments, the integration rate / speed is measured by the surface reaction rate Vmax on the surface of the substrate (e.g., flow cell or cBot system). For example, ffNs-DB sets having a 3'-AOM blocking group showed at least 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 300%, 400%, or 500% higher Vmax values (ms) compared to ffNs-PA sets having a 3'-O-azidomethyl blocking group under the same conditions and for the same period. -1This may result in improvements to the integration rate / speed. In some embodiments, the integration rate / speed is measured at ambient temperature or a temperature lower than ambient temperature (e.g., 4-10°C). In other embodiments, the integration rate / speed is measured at high temperatures such as 40°C, 45°C, 50°C, 55°C, 60°C, or 65°C. In some such embodiments, the integration rate / speed is measured in a basic pH environment, for example, in a solution with pH 9.0, 9.2, 9.4, 9.6, 9.8, or 10.0. In some such embodiments, the integration rate / speed is measured in the presence of an enzyme such as polymerase (e.g., DNA polymerase), terminal deoxynucleotidyltransferase, or reverse transcriptase. In some embodiments, ffN-DB is ffT-DB, ffC-DB, or ffA-DB. In one embodiment, the ffN-DB set having the improved phasing value described herein includes ffT-DB, ffC, ffA, and ffG. In another embodiment, the ffN-DB set having the improved fading value described herein includes ffT-DB, ffC-DB, ffA, and ffG. In yet another embodiment, the ffN-DB set having the improved fading value described herein includes ffT-DB, ffC-DB, ffA-DB, and ffG.
[0090] In some further embodiments, when the nucleic acid base of the nucleotide described herein is thymine or its derivatives and analogs in which the nucleic acid base is optionally substituted (i.e., the nucleotide is T), L 1 This includes an allylamine moiety or an allylamide moiety, or a variant thereof that is optionally substituted. In certain examples, L 1 teeth, [ka] It includes or is one of the following: * This is L relative to the C5 position of the thymine base. 1 The binding site is shown. In some embodiments, the T nucleotide described herein is directly bound to the C5 position of the thymine base (i.e., ffT-DB), [ka] This is a fully functionalized T nucleotide (ffT) labeled with a dye molecule via a cleavable linker, including the following: In some cases, when ffT-DB is used for sequencing applications in the presence of a palladium catalyst, sequencing metrics such as phasing, prephasing, and error rates can be substantially improved. For example, when ffT-DB having a 3'-AOM blocking group as described herein is used, it can result in improvements of at least 50%, 100%, 200%, 300%, 400%, 500%, 600%, 700%, 800%, 900%, 1000%, 1500%, 2000%, 2500%, or 3000% to one or more sequencing metrics as described herein, compared to when standard ffT-PA having a 3'-O-azidomethyl blocking group is used.
[0091] Some further embodiments of the nucleosides or nucleotides described herein include formulas (Ia), (Ia'), (Ib), (Ic), (Ic'), or (Id). [ka] [ka]
[0092] In some further embodiments of the nucleosides or nucleotides described herein, L 2 L 2 teeth, [ka] The phenyl portion is optionally substituted, where n and m are independently integers 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10. In some such embodiments, n is 5, and L 2 The phenyl portion is unsubstituted. In some further embodiments, m is 4.
[0093] In any embodiment of the nucleoside or nucleotide described herein, a cleavable linker or L 1 / L 2 is the disulfide portion or the azide portion ( [ka] etc.), Or it may further include combinations thereof. Further non-limiting examples of linker parts are L 1 or L 2 It can be incorporated into [ka] It included. Further linker portions are disclosed in International Publication No. 2004 / 018493 and U.S. Patent Application Publication No. 2016 / 0040225, which are incorporated herein by reference.
[0094] In any embodiment of the nucleoside or nucleotide described herein, the nucleoside or nucleotide includes a 2' deoxyribose moiety (i.e., R 4 (I) is of formula (I), where (Ia)-(Id) is H. In some further embodiments, the 2'-deoxyribose contains one, two, or three phosphate groups at the 5' position of the sugar ring. In some further embodiments, the nucleotides described herein are nucleotide triphots (i.e., R in formulas (I) and (Ia)-(Id)). 6 (It forms triphothates).
[0095] In any embodiment of the nucleoside or nucleotide described herein, the detectable label may include a fluorescent dye.
[0096] Further embodiments of this disclosure relate to oligonucleotides or polynucleotides comprising nucleosides or nucleotides as described herein. For example, an oligonucleotide or polynucleotide incorporating the nucleotide of formula (Ia') has the following structure: [ka] It contained. In some such embodiments, an oligonucleotide or polynucleotide hybridizes to a template or target polynucleotide. In some such embodiments, the template polynucleotide is immobilized on a solid support.
[0097] Further embodiments of the present disclosure relate to a solid support comprising a plurality of immobilized templates or arrays of target polynucleotides, wherein at least a portion of such immobilized templates or target polynucleotides are hybridized to oligonucleotides or polynucleotides comprising nucleosides or nucleotides as described herein.
[0098] In any embodiment of the nucleotides or nucleosides described herein, the 3'-OH blocking group and the cleavable linker (and attached label) may be removed under the same or substantially the same chemical reaction conditions, for example, the 3'-OH blocking group and the detectable label may be removed in a single chemical reaction. In other embodiments, the 3'-OH blocking group and the detectable label are removed in two separate steps.
[0099] In some embodiments, the 3'-blocked nucleotides or nucleosides described herein offer superior stability during storage in solution or lyophilized form, or during reagent handling between sequencing applications, compared to the same nucleotides or nucleosides protected with a standard 3'-OH blocking group disclosed in the prior art, such as a 3'-O-azidomethyl protecting group. For example, the acetal blocking groups disclosed herein can impart at least 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 300%, 400%, 500%, 600%, 700%, 800%, 90%, 1000%, 1500%, 2000%, 2500%, or 3000% improved stability compared to azidomethyl-protected 3'-OH under the same conditions and for the same duration, thereby reducing prephasing values and increasing sequencing read lengths. In some embodiments, stability is measured at ambient temperature or below ambient temperature (e.g., 4-10°C). In other embodiments, stability is measured at high temperatures such as 40°C, 45°C, 50°C, 55°C, 60°C, or 65°C. In some such embodiments, stability is measured in a basic pH environment, for example, in a solution with pH 9.0, 9.2, 9.4, 9.6, 9.8, or 10.0. In some such embodiments, stability is measured with or without the presence of enzymes such as polymerase (e.g., DNA polymerase), terminal deoxynucleotidyltransferase, or reverse transcriptase.
[0100] In some embodiments, the 3'-blocked nucleotides or nucleosides described herein provide superior deblocking rates in solution during the chemical cleavage step for sequencing applications compared to the same nucleotides or nucleosides protected with standard 3'-OH blocking groups disclosed in the prior art, such as 3'-O-azidomethyl protecting groups. For example, the acetal blocking groups disclosed herein may impart improved deblocking rates of at least 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 150%, 200%, 300%, 400%, 500%, 600%, 700%, 800%, 90%, 1000%, 1500%, or 2000% compared to azidomethyl-protected 3'-OH using standard deblocking reagents (such as tris(hydroxypropyl)phosphine), thereby reducing the overall time for sequencing cycles. In some embodiments, the deblocking rate is measured at ambient temperature or a temperature lower than ambient temperature (e.g., 4-10°C). In other embodiments, the deblocking rate is measured at high temperatures such as 40°C, 45°C, 50°C, 55°C, 60°C, or 65°C. In some such embodiments, the deblocking rate is measured in a basic pH environment, for example, in a solution with pH 9.0, 9.2, 9.4, 9.6, 9.8, or 10.0. In some such embodiments, the molar ratio of the deblocking reagent to the substrate (i.e., 3'-blocked nucleoside or nucleotide) is about 10:1, about 5:1, about 2:1, or about 1:1.
[0101] In some embodiments, a palladium deblocking reagent (e.g., Pd(0)) is used to remove a 3' acetal blocking group (e.g., an AOM blocking group). Pd forms a chelate complex with the two oxygen atoms of the AOM group and the double bond of the allyl group, allowing for the removal of the deblocking reagent very close to the functional group and accelerating the deblocking rate. For example, after the cleavage of the Pd of the linker described herein and the 3' blocking group of the incorporated nucleotide having formula (Ia), (Ia'), (Ib), (Ic), (Ic'), or (Id), the remaining linker construct on the copy polynucleotide may include the following structures. [ka] The winding curves indicate the bonding of oxygen to the remaining phosphodiester bonds in the copy polynucleotide chain. For example, [ka] It included. The allylamide or propargylamide moiety can be further cleaved by a Pd catalyst. Furthermore, the remaining linker construct, bound to a detectable label, has the following structure: [ka] for example, [ka]
[0102] Cutting conditions for a severable linker The cleavable linkers described herein can be removed or cleaved under a variety of chemical conditions. Non-limiting cleavage conditions include palladium catalysts such as Pd(II) complexes (e.g., Pd(OAc)2, allyl Pd(II) chloride dimer [(allyl)PdCl]2, or Na2PdCl4) in the presence of water-soluble phosphine ligands, such as tris(hydroxypropyl)phosphine (THP or THPP) or tris(hydroxymethyl)phosphine (THMP). In some embodiments, the 3' acetal blocking group can be cleaved under the same or substantially the same cleavage conditions as those of the cleavable linker.
[0103] Palladium catalyst In some embodiments, the 3'-acetal blocking group and cleavable linker described herein can be cleaved by a palladium catalyst. In some such embodiments, the Pd catalyst is water-soluble. In some such embodiments, the Pd catalyst is a Pd(0) complex (e.g., tris(3,3',3”-phosphininidine tris(benzenesulfonate)palladium(0)) salt nonasodium notahydrate). In some cases, Pd(0) can be produced in situ from the reduction of a Pd(II) complex with a reagent such as an alkene, alcohol, amine, phosphine, or metal hydride. Suitable palladium sources include Pd(CH3CN)2Cl2, [PdCl(allyl)]2, [Pd(allyl)(THP)]Cl, [Pd(allyl)(THP)2]Cl, Pd(OAc)2, Pd(PPh3)4, Pd(dba)2, Pd(Acac)2, PdCl2(COD), and Pd(TFA)2. In one such embodiment, the Pd(0) complex is N It is generated in situ from a2PdCl4. In another embodiment, the palladium source is the allylpalladium(II) chloride dimer [(allyl)PdCl]2 or [PdCl(C3H5)]2. In some embodiments, the Pd(0) catalyst is generated in aqueous solution by mixing the Pd(II) complex with a phosphine. Suitable phosphines include water-soluble phosphines such as tris(hydroxypropyl)phosphine (THP), tris(hydroxymethyl)phosphine (THMP), 1,3,5-triaza-7-phosphine-adamantane (PTA), bis(p-sulfonatophenyl)phenylphosphine dihydrate potassium salt, tris(carboxyethyl)phosphine (TCEP), and triphenylphosphine-3,3',3”-trisulfonic acid trisodium salt.
[0104] In some embodiments, the palladium catalyst is prepared by mixing [(allyl)PdCl]2 with THP in situ. The molar ratio of [(allyl)PdCl]2 to THP may be about 1:2, 1:3, 1:4, 1:5, 1:6, 1:7, 1:8, 1:9, or 1:10. In one embodiment, the molar ratio of [(allyl)PdCl]2 to THP is 1:10. In some other embodiments, the palladium catalyst is prepared by mixing the water-soluble Pd reagent Na2PdCl4 with THP in situ. The molar ratio of Na2PdCl4 to THP may be about 1:2, 1:3, 1:4, 1:5, 1:6, 1:7, 1:8, 1:9, or 1:10. In one embodiment, the molar ratio of Na2PdCl4 to THP is about 1:3. In another embodiment, the molar ratio of Na2PdCl4 to THP is about 1:3. In some further embodiments, one or more reducing agents, such as ascorbic acid or a salt thereof (e.g., sodium ascorbate), may be added. In some embodiments, the cleavage mixture may contain further buffering reagents, such as primary amines, secondary amines, tertiary amines, natural amino acids, unnatural amino acids, carbonates, phosphates or borates, or combinations thereof. In some further embodiments, the buffering reagents include ethanolamine (EA), tris(hydroxymethyl)aminomethane (Tris), glycine, sodium carbonate, sodium phosphate, sodium borate, dimethylethanolamine (DMEA), diethylethanolamine (DEEA), N,N,N',N'-tetramethylethylenediamine (TMEDA), N,N,N',N'-tetraethylethylenediamine (TEEDA), or 2-piperidineethanol, or combinations thereof. In one embodiment, one or more buffering reagents include DEEA. In another embodiment, one or more buffering reagents contain one or more inorganic salts, such as carbonates, phosphates, or borates, or combinations thereof. In one embodiment, the inorganic salt is a sodium salt.
[0105] In other embodiments, the cleavage conditions for the cleavable linker differ from those for the cleavage conditions for the 3'-OH blocking group. For example, if the 3'-blocking group is further 3'-O-azidomethyl, the -CH2N3 moiety can be converted to an amino group by a phosphine. Alternatively, the azido group in -CH2N3 can be converted to an amino group by contacting such a molecule with a thiol, particularly a water-soluble thiol such as dithiothreitol (DTT). In one embodiment, the phosphine is THP.
[0106] Compatibility with linearization To maximize the throughput of nucleic acid sequencing reactions, it is advantageous to be able to sequence multiple template molecules in parallel. Parallel processing of multiple templates can be achieved through the use of nucleic acid sequencing techniques. These arrays typically consist of a high-density matrix of polynucleotides immobilized on a solid support material.
[0107] International Publication Nos. 98 / 44151 and 00 / 18957 both describe methods for nucleic acid amplification that enable the immobilization of amplification products onto a solid support in order to form an array consisting of clusters or “colonies” formed from multiple identical immobilized polynucleotide chains and multiple identical immobilized complementary chains. This type of array is referred to herein as a “clustered array.” Nucleic acid molecules present in DNA colonies on clustered arrays prepared according to these methods can provide templates for sequencing reactions, such as those described in International Publication No. 98 / 44152. The products of solid-phase amplification reactions, such as those described in International Publication Nos. 98 / 44151 and 00 / 18957, are so-called “crosslinked” structures formed by the annealing of pairs of immobilized polynucleotide chains and immobilized complementary chains, with both chains bound to the solid support at their 5' ends. To provide a template more suitable for nucleic acid sequencing, it is preferable to remove substantially all or at least part of one of the immobilized chains of the “crosslinked” structure to produce a template that is at least partially single-stranded. Therefore, the single-stranded template portion is available for hybridization to sequencing primers. The process of removing all or part of one immobilized strand in a "crosslinked" double-stranded nucleic acid structure is called "linearization." Linearization can be performed in various ways, including but not limited to enzymatic cleavage, photochemical cleavage, or chemical cleavage. Non-limiting examples of linearization methods are disclosed in International Publication 2007 / 010251, U.S. Patent Application Publication 2009 / 0088327, U.S. Patent Application Publication 2009 / 0118128, and U.S. Patent Application Publication 2019 / 0352327, which are incorporated in their entirety by reference.
[0108] In particular, amplification (e.g., bridge amplification or exclusion amplification) forms an array consisting of clusters or "colonies" formed from multiple identical immobilized target polynucleotide chains and multiple identical immobilized complementary chains. The target chain and complementary stand form at least partially double-stranded polynucleotide complexes, and both chains are immobilized to a solid support at their 5' ends. To generate a template that is at least partially single-stranded, the double-stranded polynucleotide is brought into contact with an aqueous solution of a palladium catalyst, which removes at least a portion of one of the immobilized chains by cleaving one of the chains at a cleavage site containing an allyl-modified nucleoside (e.g., an allyl-modified T nucleoside). Thus, the single-stranded portion of the template is available for hybridization to a sequencing primer to initiate the first round of SBS (Read 1). In some embodiments, the allyl-modified nucleoside is located in the P5 primer sequence. This method is called first chemical linearization in comparison to standard enzymatic linearization, in which such removal or cleavage is facilitated by an enzymatic cleavage reaction using the enzyme USER to cleave the U position on the P5 primer.
[0109] In some embodiments, conditions for cleaving a cleavable linker and / or deprotecting or removing a 3'-OH blocking group are also compatible with the linearization process. In some further embodiments, such cleavage conditions are compatible with a chemical linearization process involving the use of a Pd complex and phosphine. In some embodiments, the Pd complex is a Pd(II) complex (e.g., Pd(OAc)2, [(Allyl)PdCl]2, or Na2PdCl4) that produces Pd(0) in situ in the presence of phosphine (e.g., THP). A chemical linearization process for cleaving an allyl-modified T nucleoside in a P5 primer sequence using a Pd catalyst is described in detail in U.S. Patent Application Publication 2019 / 0352327, which is incorporated in whole by reference. In further embodiments, the Pd cleavage mixture disclosed herein (e.g., [Pd(allyl)Cl)2 and THP in a buffer solution containing DEEA) can be used directly in the first chemical linearization step. The reduction in the number of reagents allows for further simplification of equipment (fluids and cartridges).
[0110] Unless otherwise indicated, references to nucleotides are also intended to be applicable to nucleosides.
[0111] Labeled nucleotides According to one aspect of this disclosure, the described 3'-OH-blocked nucleotides also include detectable labels, and such nucleotides are referred to as labeled nucleotides or fully functionalized nucleotides (ffNs). The label (e.g., a fluorescent dye) is conjugated via a linker that can be cleaved by various means, including hydrophobic attraction, ionic attraction, and covalent bonding. In some aspects, the dye is conjugated to the nucleotide by covalent bonding via a cleavable linker. In some cases, such labeled nucleotides are also referred to as “modified nucleotides.” Those skilled in the art will understand that the label can be covalently bonded to a linker by reacting the functional group of the label (e.g., a carboxyl) with the functional group of the linker (e.g., an amino).
[0112] Labeled nucleosides and nucleotides are useful for labeling polynucleotides formed by enzymatic synthesis. Non-limiting examples of such synthesis include PCR amplification, isothermal amplification, solid-phase amplification, polynucleotide sequencing (e.g., solid-phase sequencing), and nick translation reactions.
[0113] In some embodiments, the dye may be covalently bonded to an oligonucleotide or nucleotide via a nucleotide base. For example, the labeled nucleotide or oligonucleotide may have a label attached to the C5 position of the pyrimidine base, or a label attached to the C7 position of the 7-deazapurine base via a cleavable linker moiety.
[0114] Unless otherwise indicated, references to nucleotides are also intended to be applicable to nucleosides. This application is also described further with reference to DNA, but that description is also applicable to RNA, PNA and other nucleic acids unless otherwise indicated.
[0115] Nucleosides and nucleotides may be labeled at a site on a sugar or nucleic acid base. As is known in the art, a “nucleotide” consists of a nitrogen base, a sugar, and one or more phosphate groups. In RNA, the sugar is ribose, while in DNA, the sugar is deoxyribose, i.e., a sugar lacking the hydroxyl group present in ribose. The nitrogen base is a derivative of a purine or pyrimidine. Purines are adenine (A) and guanine (G), and pyrimidines are cytosine (C) and thymine (T), or, in the context of RNA, uracil (U). The C-1 atom of deoxyribose is bonded to N-1 of the pyrimidine or N-9 of the purine. Nucleosides are also phosphate esters of nucleosides, with esterification occurring at the hydroxyl group bonded to C-3 or C-5 of the sugar. Nucleosides are usually monophosphates, diphosphates, or triphosphosphates.
[0116] A "nucleoside" is structurally similar to a nucleotide but lacks a phosphate group. Examples of nucleotide analogs are those in which a label is linked to a base and there is no phosphate group attached to a sugar molecule.
[0117] Bases are typically called purines or pyrimidines, but those skilled in the art will understand that derivatives and analogues are available that do not alter the ability of the nucleotide or nucleoside to undergo Watson-Crick base pairing. “Derivative” or “analog” means a compound or molecule whose core structure is the same as or closely similar to the parent compound, but which has undergone chemical or physical modifications, such as different or additional side groups (which allow the derivative nucleotide or nucleoside to be linked to other molecules). For example, the base may be a deazapurine. In certain embodiments, the derivative should be capable of Watson-Crick pairing. “Derivatives” and “analogues” also include, for example, synthetic nucleotide or nucleoside derivatives having modified base moieties and / or modified sugar moieties. Such derivatives and analogues are discussed, for example, in Scheit, Nucleotide analogs (John Wiley & Son, 1980) and Uhlman et al., Chemical Reviews 90:543-584, 1990. Nucleotide analogs may also include modified phosphodiester linkages, such as phosphorothioates, phosphorodithioates, alkylphosphonates, phosphoranilidates, and phosphoramidite linkages.
[0118] In certain embodiments, the labeled nucleoside or nucleotide may be enzymatically embeddable and enzymatically elongable. Therefore, the linker portion may be long enough to connect the nucleotide to the compound, and as a result, the compound does not significantly interfere with the overall binding and recognition of the nucleotide by nucleic acid replication enzymes. Thus, the linker may also include spacer units. Spacers, for example, move nucleotide bases away from cleavage sites or labels.
[0119] This disclosure also includes polynucleotides incorporating dye compounds. Such polynucleotides may be DNA or RNA consisting of deoxyribonucleotides or ribonucleotides joined by phosphodiester links, respectively. Polynucleotides may include naturally occurring nucleotides, non-naturally occurring (or modified) nucleotides other than the labeled nucleotides described herein, or any combination thereof, in combination with at least one modified nucleotide (e.g., labeled with a dye compound) as described herein. Polynucleotides according to this disclosure may also include non-naturally occurring skeletal links and / or non-nucleotide chemical modifications. Chimeric structures consisting of mixtures of ribonucleotides and deoxyribonucleotides containing at least one labeled nucleotide can also be conceived.
[0120] Examples of non-limiting exemplary labeled nucleotides described herein include: [ka] In the formula, L is a severable linker (L as specified herein). 2 R represents (optionally containing), where R represents the ribose or deoxyribose moiety described above, or a ribose or deoxyribose moiety having a 5' position substituted with one, two, or three phosphate groups.
[0121] In some embodiments, non-limiting fluorescent dye conjugates are shown below. [ka] [ka] In the formula, PG represents a 3'-OH blocking group as described herein, n is an integer of 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10, and k is 0, 1, 2, 3, 4, or 5. In one embodiment, -O-PG is AOM. In another embodiment, -O-PG is -O-azidomethyl. In one embodiment, n is 5. [ka] This refers to the linkage point of the dye having a cleavable linker, resulting from the reaction between the amino group of the linker moiety and the carboxyl group of the dye.
[0122] Sequence determination method The labeled nucleotides or nucleosides described herein can be used in any analytical method, including the detection of a fluorescent label attached to the nucleotide or nucleoside, regardless of whether the nucleotide or nucleoside is analyzed on its own or in conjunction with or incorporated into a larger molecular structure or conjugate. In this context, the term “incorporated into a polynucleotide” may mean that the 5' phosphate is phosphodiester-bonded to the 3'-OH group of a second (modified or unmodified) nucleotide, which itself may form part of a longer polynucleotide chain. The 3' ends of the nucleotides described herein may or may not be phosphodiester-bonded to the 5' phosphate of further (modified or unmodified) nucleotides. Accordingly, in one non-limiting embodiment, the Disclosure provides a method for detecting nucleotides incorporated into a polynucleotide, comprising (a) incorporating at least one nucleotide of the Disclosure into a polynucleotide, and (b) detecting the nucleotide(s) incorporated into the polynucleotide by detecting a fluorescent signal from a detectable label (e.g., a fluorescent compound) bound to the nucleotide(s). The Method may comprise (a) a synthesis step of incorporating one or more nucleotides of the Disclosure into a polynucleotide, and (b) a detection step of detecting one or more nucleotides incorporated into the polynucleotide by detecting or quantitatively measuring their fluorescence.
[0123] Additional aspects of the present disclosure include a method for preparing a growing polynucleotide complementary to a target single-stranded polynucleotide in a sequencing reaction, comprising incorporating the nucleotides described herein into the growing complementary polynucleotide, wherein the incorporation of the nucleotides prevents the introduction of any subsequent nucleotides into the growing complementary polynucleotide.
[0124] Some embodiments of this disclosure relate to methods for determining the sequence of a target single-stranded polynucleotide. (a) 3'-OH blocking group as described herein [ka] Incorporating a nucleotide (e.g., dATP, dCTP, dGTP, dTTP, or dUTP) containing (bonded to the 3' oxygen) and a detectable label as described herein into a copy polynucleotide chain complementary to at least a portion of the target polynucleotide chain, (b) detecting the identity of the nucleotides incorporated into the copy polynucleotide chain, (c) Chemically removing the label and the 3'-OH blocking group from the nucleotide incorporated into the copy polynucleotide chain, Includes.
[0125] In some embodiments, the sequencing method further includes (d) washing away chemically removed labels and 3'-OH blocking groups from the copy polynucleotide chain by using a post-cleavage washing solution. In some such embodiments, the 3'-OH blocking groups and detectable labels are removed before introducing the next complementary nucleotide. In some embodiments, washing step (d) also removes unincorporated nucleotides. In other embodiments, the method may include a separate washing step prior to step (b) for washing away unincorporated nucleotides from the copy polynucleotide chain.
[0126] In some embodiments, steps (a)-(d) are repeated until the sequence of a portion of the target polynucleotide strand is determined. In some such embodiments, steps (a)-(d) are repeated at least 50 times, at least 75 times, at least 100 times, at least 150 times, at least 200 times, at least 250 times, or at least 300 times.
[0127] Incorporated mixture In some embodiments of the methods described herein, step (a), also referred to as the integration step, comprises contacting a mixture comprising one or more nucleotides (e.g., dATP, dCTP, dGTP, and dTTP or dUTP) with a copy polynucleotide / target polynucleotide complex in an integration solution comprising a polymerase and one or more buffers. In some such embodiments, the polymerase is a DNA polymerase, e.g., Pol812, Pol1901, Pol1558, or Pol963. The amino acid sequences of the Pol812, Pol1901, Pol1558, or Pol963 DNA polymerases are described, for example, in U.S. Patent Application Publications 2020 / 0131484 and 2020 / 0181587, both of which are incorporated herein by reference. In some embodiments, the one or more buffers include primary amines, secondary amines, tertiary amines, natural amino acids, or non-natural amino acids, or combinations thereof. In further embodiments, the buffer comprises ethanolamine or glycine, or a combination thereof. In one embodiment, the buffer comprises or is glycine. In some embodiments, the use of glycine in the incorporated mixture can improve the fading value compared to a standard buffer such as ethanolamine (EA) under the same conditions. For example, the use of glycine provides a reduction or decrease in the fading value of at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 300%, 400%, or 500% compared to ethanolamine used under the same conditions. In some cases, the use of glycine provides a fading value of less than approximately 0.15%, 0.1%, 0.09%, 0.08%, 0.07%, 0.06%, 0.05%, 0.04%, 0.03%, 0.02%, or 0.01% in at least 50 cycles of SBS sequencing runs. In a further embodiment, the use of glycine provides a fading value of less than approximately 0.08% in Read1 of at least 150 cycles of SBS sequencing execution.
[0128] cutting mixture In some embodiments of the methods described herein, step (c), also referred to as the cleavage step, involves contacting the incorporated nucleotides and copy polynucleotide strands with a cleavage solution comprising a palladium catalyst described herein. In some such embodiments, the 3'-OH blocking group and the detectable label are removed in a single step of the reaction. In one such embodiment, the 3' blocking group is AOM, the cleavable linker comprises an AOL moiety, and both are removed or cleaved in a single step of the chemical reaction. In some further embodiments, the cleavage solution (also called the cleavage mix) comprises a Pd catalyst described herein.
[0129] In some further embodiments, the Pd catalyst is a Pd(0) catalyst. In some such embodiments, Pd(0) is prepared by mixing a Pd(II) reagent with one or more phosphine ligands in situ. In some such embodiments, the palladium catalyst can be prepared by mixing [(allyl)PdCl]2 with THP in situ. The molar ratio of [(allyl)PdCl]2 to THP can be about 1:2, 1:3, 1:4, 1:5, 1:6, 1:7, 1:8, 1:9, or 1:10. In one embodiment, the molar ratio of [(allyl)PdCl]2 to THP is 1:10 (i.e., the molar ratio of Pd:THP is 1:5). In some other embodiments, the palladium catalyst can be prepared by mixing the water-soluble Pd(II) reagent Na2PdCl4 with THP in situ. The molar ratio of Na2PdCl4 to THP can be about 1:2, 1:3, 1:4, 1:5, 1:6, 1:7, 1:8, 1:9, or 1:10. In one embodiment, the molar ratio of Na2PdCl4 to THP is about 1:3. In another embodiment, the molar ratio of Na2PdCl4 to THP is about 1:3.5. Other non-limiting examples of Pd catalysts include Pd(CH3CN)2Cl2.
[0130] In some further embodiments, one or more reducing agents, such as ascorbic acid or a salt thereof (e.g., sodium ascorbate), may be added. In some embodiments, the cleavage solution may contain one or more buffering reagents, such as primary amines, secondary amines, tertiary amines, carbonates, phosphates, or borates, or combinations thereof. In some further embodiments, the buffering reagents include ethanolamine (EA), tris(hydroxymethyl)aminomethane (Tris), glycine, sodium carbonate, sodium phosphate, sodium borate, dimethylethanolamine (DMEA), diethylethanolamine (DEEA), N,N,N',N'-tetramethylethylenediamine (TEMED), N,N,N',N'-tetraethylethylenediamine (TEEDA), or 2-piperidineethanol, or combinations thereof. In one embodiment, the buffering reagent contains or is DEEA. In another embodiment, the buffering reagent contains one or more inorganic salts, such as carbonates, phosphates, or borates, or combinations thereof. In one embodiment, the inorganic salt is a sodium salt. In a further embodiment, the cleavage solution comprises a palladium (Pd) catalyst (e.g., [(allyl)PdCl]2 / THP or Na2PdCl4 / THP) and one or more buffering reagents described herein (e.g., tertiary amines such as DEEA), and has a pH of about 9.0 to about 10.0 (e.g., 9.6 or 9.8).
[0131] In other embodiments, the label and the 3'-blocking group are removed by two separate chemical reactions. In some cases, removing the label from a nucleotide incorporated into a copy polynucleotide chain involves contacting the copy chain containing the incorporated nucleotide with a first cleavage solution containing the Pd catalyst described herein. In some cases, removing the 3'-OH blocking group from a nucleotide incorporated into a copy polynucleotide chain involves contacting the copy chain containing the incorporated nucleotide with a second cleavage solution. In some such embodiments, the second cleavage solution contains one or more phosphines, such as trialkylphosphines. Non-limiting examples of trialkylphosphines include tris(hydroxypropyl)phosphine (THP), tris-(2-carboxyethyl)phosphine (TCEP), tris(hydroxymethyl)phosphine (THMP), or tris(hydroxyethyl)phosphine (THEP). In one embodiment, the 3'-OH blocking group is 3'-O-azidomethyl, and the second cleavage solution contains THP.
[0132] In some embodiments, the cleavage solutions described herein may also be used in conventional chemical linearization steps described herein. In particular, chemical linearization of clustered polynucleotides in preparation for sequencing is achieved by palladium-catalyzed cleavage of one or more first strands of a double-stranded polynucleotide immobilized on a solid support, thus producing a single-stranded (or at least partially single-stranded) template that is available for hybridization to sequencing primers and subsequent sequencing applications (e.g., first round sequencing by synthesis (Read1)). In some embodiments, each double-stranded polynucleotide comprises a first strand and a second strand. The first strand is produced by extending a first extension primer immobilized on a solid support. In some embodiments, the first strand includes a cleavage site that can be cleaved by a palladium complex (e.g., a Pd(0) complex). In certain embodiments, the cleavage site is located on the first extension primer portion of the first strand. In further embodiments, the cleavage site includes a thymine nucleoside or nucleotide analog having an allyl functional group. In some embodiments of the method described herein, the target single-stranded polynucleotide is formed by chemically cleaving a complementary strand from a double-stranded polynucleotide. In further embodiments, both the complementary strand in the double-stranded polynucleotide and the target polynucleotide are immobilized on a solid support at their 5' ends. In some further embodiments, the chemical cleavage of the complementary strand is carried out under the same reaction conditions as chemically removing detectable labels and 3'-OH blocking groups from nucleotides incorporated into the copy polynucleotide chain (i.e., step (c) of the method described herein). In one embodiment, the first chemical linearization utilizes the same cleavage mix described herein. Palladium (Pd) Scavenger
[0133] Pd, in most cases, possesses the ability to adhere to DNA in its inactive Pd(II) form, which can interfere with the binding between DNA and polymerase, causing increased phasing. Post-cleavage washing compositions containing Pd scavenger compounds can be used following a deblocking step. For example, International Publication WO2020 / 126593 discloses Pd scavengers such as 3,3'-dithiodipropionic acid (DDPA) and lipoic acid (LA), which may be included in scanning compositions and / or post-cleavage washing compositions. The use of these scavengers in post-cleavage washing solutions is intended to scavenge Pd(0), convert Pd(0) to the inactive Pd(II) form, thereby improving pre-phasing values and sequencing metrics, reducing signal degradation, and extending sequencing read lengths.
[0134] In some embodiments of the method described herein, step (a) of the method comprises contacting a nucleotide with a copy polynucleotide chain in an integration solution comprising polymerase, at least one palladium scavenger, and one or more buffers. In some embodiments, the Pd scavenger in the integration solution is a Pd(0) scavenger. In some such embodiments, the Pd scavenger is -O-allyl, -S-allyl, -NR-allyl, and -N + The formula comprises one or more allyl moieties independently selected from the group consisting of RR'-allyls, where R is H, an unsubstituted or substituted C1-C6 alkyl, an unsubstituted or substituted C2-C6 alkenyl, an unsubstituted or substituted C2-C6 alkynyl, or an unsubstituted or substituted C6-C 10 Aryl, unsubstituted or substituted 5-10 member heteroaryl, unsubstituted or substituted C3-C 10 The molecule is a carbocyclyl, or an unsubstituted or substituted 5-10 member heterocyclyl, where R' is H, an unsubstituted C1-C6 alkyl, or a substituted C1-C6 alkyl.
[0135] In some such embodiments, the Pd(0) scavenger in the incorporated solution comprises one or more -O-allyl moieties. In some further embodiments, the Pd(0) scavenger is [ka] Or include or be a combination thereof. An alternative Pd(0) scavenger is disclosed in U.S. Patent No. 63 / 190983, which is incorporated in whole by reference.
[0136] In some embodiments, the concentration of the Pd(0) scavenger containing one or more allyl moieties in the incorporated solution is about 0.1 mM to about 100 mM, 0.2 mM to about 75 mM, about 0.5 mM to about 50 mM, about 1 mM to about 20 mM, or about 2 mM to about 10 mM. In further embodiments, the concentration of the Pd(0) scavenger is about 0.5 mM, 1 mM, 1.5 mM, 2 mM, 2.5 mM, 3 mM, 3.5 mM, 4 mM, 4.5 mM, 5 mM, 5.5 mM, 6 mM, 6.5 mM, 7 mM, 7.5 mM, 8 mM, 8.5 mM, 9 mM, 9.5 mM, 10 mM, 12.5 mM, 15 mM, 17.5 mM, or 20 mM. In further embodiments, the pH of the incorporated solution is approximately 9 to 10.
[0137] In some embodiments, the molar ratio of the palladium catalyst (in the starting solution) to the palladium scavenger containing one or more allyl moieties is about 1:100, 1:50, 1:20, 1:10, or 1:5.
[0138] In some other embodiments of the method described herein, the Pd(0) scavenger includes one or more allyl moieties described herein in the scan solution used in step (b) when performing one or more fluorescence measurements to detect the identity of the incorporated nucleotide in the copy polynucleotide. In yet another embodiment, the Pd(0) scavenger includes one or more allyl moieties that may be present in both the incorporation solution and the scan solution.
[0139] In some further embodiments of the methods described herein, a post-cleavage washing step is used after labeling to remove the 3' blocking groups. In some such embodiments, one or more palladium scavengers are also used in the post-cleavage washing step for labeling and removal of the 3' blocking groups. In some further embodiments, one or more Pd scavengers in the post-cleavage washing solution include Pd(II) scavengers. In some such embodiments, the palladium scavenger includes isocyanoacetate (ICNA) salt, cysteine or a salt thereof, or a combination thereof. In one embodiment, the palladium scavenger includes or is potassium isocyanoacetate or sodium isocyanoacetate. In another embodiment, the palladium scavenger includes or is cysteine or a salt thereof (e.g., L-cysteine or L-cysteine HCl salt). Other non-limiting examples of palladium scavengers in the post-cleavage washing solution include ethyl isocyanoacetate, methyl isocyanoacetate, N-acetyl-L-cysteine, potassium ethylxanthonate (PEX or KS-C(=S)-OEt), potassium isopropylxanthonate, glutathione, lipoic acid, ethylenediaminetetraacetic acid (EDTA), iminodiacetic acid, nitrilodiacetic acid, trimercapto-S-triazine, dimethyldithiocarbamate, dithiothreitol, mercaptoethanol, allyl alcohol, propargyl alcohol, thiols, tertiary amines and / or tertiary phosphines, or combinations thereof.
[0140] In further embodiments, the concentration of Pd(II) scavengers such as L-cysteine in the post-cutting washing solution is approximately 0.1 mM to approximately 100 mM, 0.2 mM to approximately 75 mM, approximately 0.5 mM to approximately 50 mM, approximately 1 mM to approximately 20 mM, or approximately 2 mM to approximately 10 mM. In further embodiments, the concentration of Pd(II) scavengers such as L-cysteine is approximately 1 mM, 2 mM, 3 mM, 4 mM, 5 mM, 6 mM, 6.5 mM, 7 mM, 8 mM, 9 mM, 10 mM, 12.5 mM, 15 mM, 17.5 mM, or 20 mM. In one embodiment, the concentration of Pd scavengers such as L-cysteine or its salt in the post-cutting washing solution is approximately 10 mM.
[0141] In some other embodiments of the method described herein, all Pd scavengers (e.g., both Pd(0) and Pd(II) scavengers) are in the incorporation solution and / or scan solution, and the method does not involve any specific post-cutting wash steps to remove any trace amounts of residual Pd species.
[0142] In some embodiments of the methods described herein, the use of a Pd scavenger (e.g., a Pd(0) scavenger having one or more allele moieties) can reduce the prephasing value by at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 300%, 400%, 500%, 600%, 700%, 800%, 900%, or 1000% compared to the same sequencing performed under the same conditions without the use of a palladium scavenger. In some such embodiments, the Pd(0) scavenger can reduce the prephasizing value of sequencing to less than approximately 0.1%, 0.09%, 0.08%, 0.07%, 0.06%, 0.05%, 0.04%, 0.03%, 0.02%, or 0.01% over at least 50 cycles of SBS sequencing execution. In some embodiments, the prephasizing value refers to the value measured after 50, 75, 100, 125, 150, 200, 250, or 300 cycles.
[0143] In some further embodiments, a palladium scavenger (e.g., a Pd(II) scavenger such as L-cysteine or a salt thereof) can reduce the pre-phasing or phasing value by at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 300%, 400%, 500%, 600%, 700%, 800%, 900%, or 1000% compared to the same sequencing determination performed under the same conditions without the use of a palladium scavenger. In some such embodiments, the use of a Pd scavenger provides a fading value of less than approximately 0.2%, 0.15%, 0.1%, 0.09%, 0.08%, 0.07%, 0.06%, 0.05%, 0.04%, 0.03%, 0.02%, or 0.01% in at least 50 cycles of SBS sequencing runs. In some embodiments, the fading value refers to the value measured after 50, 75, 100, 125, 150, 200, 250, or 300 cycles. In further embodiments, the use of one or more Pd scavengers provides a fading value of less than approximately 0.05% in Read1 of at least 150 cycles of SBS sequencing runs.
[0144] In some embodiments, the post-washing solution described herein may also be used in a separate washing step prior to the detection step (i.e., step (b) in the method described herein) to wash away any unintegrated nucleotides from step (a).
[0145] In some further embodiments, the nucleotides used in the integration step (a) are fully functionalized A, C, T, and G nucleotide triphodes, each containing a 3' blocking group (e.g., 3'-AOM) and a cleavable linker (e.g., a cleavable linker containing an AOL linker moiety) as described herein. In some such embodiments, the nucleotides herein offer superior stability in solution during sequencing runs compared to the same nucleotides protected with a standard 3'-O-azidomethyl blocking group. For example, the 3'-acetal blocking groups disclosed herein can impart at least 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 300%, 400%, 500%, 600%, 700%, 800%, 90%, 1000%, 1500%, 2000%, 2500%, or 3000% improved stability compared to azidomethyl-protected 3'-OH under the same conditions and for the same duration, thereby reducing prephasing values and increasing sequencing read lengths. In some embodiments, stability is measured at ambient temperature or below ambient temperature (e.g., 4-10°C). In other embodiments, stability is measured at high temperatures such as 40°C, 45°C, 50°C, 55°C, 60°C, or 65°C. In some such embodiments, stability is measured in a basic pH environment, for example, in a solution with pH 9.0, 9.2, 9.4, 9.6, 9.8, or 10.0. In some further embodiments, the prephasing values by the 3'-blocking nucleotides described herein are less than approximately 0.25, 0.24, 0.23, 0.22, 0.21, 0.20, 0.19, 0.18, 0.17, 0.16, 0.15, 0.14, 0.13, 0.12, 0.11, 0.10, 0.09, 0.08, 0.07, 0.06, or 0.05 after SBS exceeding 50, 100, or 150 cycles.In some further embodiments, the phasing values due to the 3'-blocking nucleotide are approximately 0.25, 0.24, 0.23, 0.22, 0.21, 0.20, 0.19, 0.18, 0.17, 0.16, 0.15, 0.14, 0.13, 0.12, 0.11, 0.10, 0.09, 0.08, 0.07, 0.06, or less than 0.05 after SBS exceeding 50, 100, or 150 cycles. In one embodiment, each ffN contains a 3'-AOM group.
[0146] In some embodiments, the 3'-blocking nucleotides described herein provide superior deblocking rates in solution during the chemical cleavage step of a sequencing run compared to the same nucleotide protected with a standard 3'-O-azidomethyl blocking group. For example, the 3'-acetal (e.g., AOM) blocking groups disclosed herein may impart an improved deblocking rate of at least 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 150%, 200%, 300%, 400%, 500%, 600%, 700%, 800%, 90%, 1000%, 1500%, or 2000% compared to azidomethyl-protected 3'-OH using a standard deblocking reagent (e.g., tris(hydroxypropyl)phosphine), thereby reducing the overall time for the sequencing cycle. In some embodiments, the deblocking time of each nucleotide is reduced by about 5%, 10%, 20%, 30%, 40%, 50%, or 60%. For example, the deblocking times of 3'-AOM and 3'-O-azidomethyl are about 4-5 seconds and about 9-10 seconds, respectively, under specific chemical reaction conditions. In some embodiments, the half-life (t) of the AOM blocking group is reduced. 1 / 2 ) is at least 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 times faster than the azidomethyl blocking group. In some such embodiments, the t of AOM 1 / 2 The time is approximately 1 minute, while the t of azidomethyl 1 / 2is about 11 minutes. In some embodiments, the deblocking rate is measured at ambient temperature or a temperature lower than ambient temperature (such as 4 - 10 °C). In other embodiments, the deblocking rate is measured at a high temperature such as 40 °C, 45 °C, 50 °C, 55 °C, 60 °C, or 65 °C. In some such embodiments, the deblocking rate is measured in a basic pH environment, such as a solution with pH 9.0, 9.2, 9.4, 9.6, 9.8, or 10.0. In some such embodiments, the molar ratio of the deblocking reagent to the substrate (i.e., the 3'-blocked nucleoside or nucleotide) is about 10:1, about 5:1, about 2:1, about 1:1, about 1:2, about 1:5, or about 1:10. In one embodiment, each ffN contains a 3'-AOM blocking group and an AOL linker moiety.
[0147] In any embodiment of the methods described herein, the labeled nucleotide is a nucleotide triphosphate having 2'-deoxyribose. In any embodiment of the methods described herein, the target polynucleotide strand is bound to a solid support such as a flow cell.
[0148] In one embodiment, in the synthesis step, at least one nucleotide is incorporated into the polynucleotide by the action of a polymerase enzyme. In some such embodiments, the polymerase can be DNA polymerase Pol812 or Pol1901. However, other methods of joining nucleotides to polynucleotides, such as chemical oligonucleotide synthesis or ligation of a labeled oligonucleotide to an unlabeled oligonucleotide, can be used. Thus, when the term "incorporate" is used with respect to nucleotides and polynucleotides, it can encompass polynucleotide synthesis by chemical and enzymatic methods.
[0149] In certain embodiments, a synthesis step may be performed, optionally comprising incubating a template polynucleotide chain with a reaction mixture containing the labeled 3' block nucleotides of the present disclosure. The polymerase may also be provided under conditions that allow for the formation of phosphodiester bonds between the free 3'-OH group on the polynucleotide chain annealed to the template polynucleotide chain and the 5' phosphate group on the nucleotide. Thus, the synthesis step may include the formation of a polynucleotide chain achieved by complementary base pairing of the nucleotides to the template chain.
[0150] In all embodiments of this method, the detection step may be performed while the polynucleotide chain into which the labeled nucleotide is incorporated is annealed to a template or target chain, or after a denaturation step in which the two chains are separated. Further steps, such as chemical or enzymatic reaction steps or purification steps, may be included between the synthesis step and the detection step. Specifically, the target chain into which the labeled nucleotide is incorporated may be isolated or purified and then further processed or used for subsequent analysis. For example, a target polynucleotide labeled with the nucleotide(s) described herein in the synthesis step may then be used as a labeled probe or primer. In other embodiments, the product of the synthesis step described herein may be subjected to further reaction steps, and if desired, the products of these subsequent steps may be purified or isolated.
[0151] The conditions suitable for the synthesis process are well known to those familiar with standard molecular biological techniques. In one embodiment, the synthesis process may be similar to a standard primer extension reaction using a nucleotide precursor containing the nucleotides described herein, forming an extension target chain complementary to the template chain in the presence of a suitable polymerase enzyme. In other embodiments, the synthesis process may itself constitute part of an amplification reaction, producing a labeled double-stranded amplification product consisting of annealed complementary chains derived from copying the target and template polynucleotide chains. Other exemplary synthesis processes include nick translation, chain substitution polymerization, and random priming DNA labeling. Polymerase enzymes particularly useful for the synthesis process are those capable of catalyzing the incorporation of the nucleotides described herein. A variety of naturally derived or modified polymerases can be used. For example, thermostable polymerases can be used in synthesis reactions carried out using thermal cycling conditions, although thermostable polymerases may be undesirable for isothermal primer extension reactions. Suitable thermostable polymerases capable of incorporating the nucleotides described herein include those described in International Publication No. 2005 / 024010 or International Publication No. 06 / 120433, each incorporated herein by reference. In synthesis reactions carried out at low temperatures such as 37°C, the polymerase enzyme does not necessarily need to be a thermostable polymerase, and therefore the selection of the polymerase depends on several factors such as reaction temperature, pH, and chain displacement activity.
[0152] In certain non-limiting embodiments, the Disclosure includes methods for nucleic acid sequencing, rearrangement, whole genome sequencing, and single nucleotide polymorphism scoring, as well as any other applications involving the detection of the labeled nucleotides or nucleosides described herein when incorporated into polynucleotides. Labeled nucleotides or nucleosides having the dyes described herein can also be used in any of the various other applications that would be beneficial to the use of polynucleotides labeled with nucleotides containing fluorescent dyes.
[0153] In certain embodiments, the Disclosure provides the use of labeled nucleotides according to the Disclosure in synthetic polynucleotide sequencing (SBS) reactions. Synthetic sequencing generally involves sequentially adding one or more nucleotides or oligonucleotides in the 5'-3' direction to a growing polynucleotide chain using a polymerase or ligase to form a growing polynucleotide chain complementary to the template nucleic acid to be sequenced. The identity of the bases present in one or more of the added nucleotides may be determined in a detection or "imaging" step. The identity of the added bases may be determined after each nucleotide incorporation step. The sequence of the template can then be inferred using conventional Watson-Crick base pairing rules. While the use of labeled nucleotides described herein to determine the identity of a single base may be useful, for example, in scoring single nucleotide polymorphisms, such single base extension reactions are within the scope of the Disclosure.
[0154] In one embodiment of the present disclosure, the sequence of a template polynucleotide is determined by detecting the incorporation of one or more 3' block nucleotides into a nascent chain complementary to the template polynucleotide, which is sequenced by detecting a fluorescent label(s) attached to the incorporated nucleotides. Sequence determination of the template polynucleotide can be performed by priming with a suitable primer (or a primer prepared as a hairpin structure containing the primer as part of a hairpin), and the nascent chain is extended stepwise by adding nucleotides to the 3' end of the primer in a polymer-catalyzed reaction.
[0155] In certain embodiments, each of the different nucleotide triphodes (A, T, G, and C) may be labeled with a unique fluorophore and may also contain a blocking group at the 3' position to prevent uncontrolled polymerization. Alternatively, one of the four nucleotides may be unlabeled (dark). The polymerase enzyme incorporates the nucleotides into a nascent chain complementary to the template polynucleotide, while the blocking group prevents further incorporation of nucleotides. Any unincorporated nucleotides can be washed away, and the fluorescent signal from each incorporated nucleotide can be optically "read" by suitable means, such as a charge-coupled element using laser excitation and a suitable light-emitting filter. The 3'-blocking group and the fluorescent dye compound can then be removed simultaneously or sequentially (deprotected) to allow the nascent chain to be subjected to further nucleotide incorporation. Typically, the identity of the incorporated nucleotides is determined after each incorporation step, but this is not strictly necessary. Similarly, U.S. Patent No. 5,302,509, whose disclosure is incorporated in its entirety by reference, discloses a method for sequencing polynucleotides immobilized on a solid support.
[0156] The method exemplified above involves incorporating fluorescently labeled 3'-blocked nucleotides A, G, C, and T into a growth chain complementary to an immobilized polynucleotide in the presence of DNA polymerase. The polymerase incorporates the base complementary to the target polynucleotide, but further addition is prevented by the 3'-blocking group. The label of the incorporated nucleotide is then determined, and the blocking group may subsequently be removed by chemical cleavage to allow further polymerization. The nucleic acid template to be sequenced in the synthetic sequencing reaction may be any polynucleotide to be sequenced. The nucleic acid template for the sequencing reaction typically includes a double-stranded region having a free 3'-OH group that functions as a primer or starting point for further nucleotide addition in the sequencing reaction. The region of the sequencing template overhangs this free 3'-OH group onto the complementary strand. The overhang region of the sequencing template may be single-stranded, but may also be double-stranded, insofar as a "break exists" on a chain complementary to the sequencing template chain, providing a free 3'-OH group to initiate the sequencing reaction. In such embodiments, sequencing may proceed by chain substitution. In certain embodiments, a primer having a free 3'-OH group may be added as a separate component (e.g., a short oligonucleotide) that hybridizes to the single-stranded region of the sequencing template. Alternatively, the sequencing primer and template chain may each form part of a partially self-complementary nucleic acid chain capable of forming an intramolecular double helix, such as a hairpin loop structure. Hairpin polynucleotides and methods by which they can be bound to a solid support are disclosed in International Publication Nos. 01 / 57248 and 2005 / 047301, respectively, which are incorporated herein by reference. Nucleotides can be sequentially added to a growing primer to synthesize a polynucleotide chain in the 5'-3' direction. The properties of the added bases may be determined after each nucleotide is added, but this is not necessarily required. This determination provides sequence information about the nucleic acid template.Therefore, nucleotides are incorporated into nucleic acid chains (or polynucleotides) by attaching them to the free 3'-OH groups of the nucleic acid chain via the formation of a phosphodiester bond with the 5' phosphate group of the nucleotide.
[0157] The nucleic acid template for sequencing may be DNA or RNA, or even a hybrid molecule consisting of deoxynucleotides and ribonucleotides. The nucleic acid template may include naturally occurring nucleotides and / or non-naturally occurring nucleotides and naturally occurring or non-naturally occurring skeletal links, as long as this does not prevent the template from being copied in the sequencing reaction.
[0158] In certain embodiments, the sequenced nucleic acid template may be bound to a solid support by any preferred linking method known in the art, such as covalent bonding. In certain embodiments, the template polynucleotide may be directly bound to the solid support (e.g., a silica-based support). However, in other embodiments of the present disclosure, the surface of the solid support may be modified in any way to allow either direct covalent bonding of the template polynucleotide or immobilization of the template polynucleotide through a hydrogel or polyelectrolyte layer (which itself is attached to the solid support by means other than covalent bonding).
[0159] Embodiments and alternative forms of sequence determination by synthesis One embodiment of this technique is pyrosequencing. Pyrosequencing detects the release of inorganic pyrophosphate (PPi) when specific nucleotides are incorporated into the nascent DNA chain. (Ronaghi, M., Karamohamed, S., Pettersson, B., Uhlen, M. and Nyren, P. (1996) "Real-time DNA sequencing using detection of pyrophosphate release." Analytical Biochemistry 242(1), 84-9; Ronaghi, M. (2001) "Pyrosequencing sheds light on DNA sequencing." Genome Res., 11(1), 3-11; Ronaghi, M., Uhlen, M. and Nyren, P. (1998) "A sequencing method based on real-time pyrophosphate." Science) 281(5375),363; U.S. Patents 6,210,891, 6,258,568 and 6,274,320 (the entirety of which is incorporated herein by reference). In pyrosequencing, released PPi can be detected by immediate conversion to adenosine triphosphate (ATP) by ATP sulfurase, and the level of the generated ATP is detected via photons produced by luciferase. The nucleic acids to be sequenced can be attached to feature regions in an array, and the array can be imaged to capture the chemiluminescent signals produced by incorporating nucleotides into the feature regions of the array. Images can be obtained after processing the array with a specific nucleotide type (e.g., T, C, or G). The images obtained after the addition of each nucleotide type differ in terms of which feature regions in the array are detected. These differences in the images reflect the different sequence content of the feature regions on the array. However, the relative position of each feature region remains unchanged in the image. The images can be stored, processed, and analyzed using the methods described herein.For example, images obtained after processing an array with each different nucleotide type can be processed in the same manner as those obtained from different detection channels for a reversible terminator-based sequencing method, as illustrated herein.
[0160] In another exemplary type of SBS, cyclic sequencing is achieved by stepwise adding reversible terminator nucleotides containing cleavable or photobleachable dye labels, such as those described in International Publication No. 04 / 018497 and U.S. Patent No. 7,057,026, whose disclosures are incorporated herein by reference. This technique has been commercialized by Solexa (now Illumina, Inc.) and is also described in International Publication No. 91 / 06678 and International Publication No. 07 / 123,744, which are each incorporated herein by reference. The availability of fluorescently labeled terminators with both ends reversible and cleaved fluorescent labels facilitates efficient cyclic reversible terminal (CRT) sequencing. Polymerases can also be co-operated to efficiently incorporate and extend these modified nucleotides.
[0161] Preferably, in a reversible terminator-based sequencing embodiment, the label does not substantially inhibit extension under SBS reaction conditions. However, the detection label may be removable, for example, by cleavage or decomposition. Images can be captured after the incorporation of the label into the arrayed nucleic acid feature region. In a particular embodiment, each cycle involves the simultaneous delivery of four different nucleotide types to the array, each nucleotide type having a spectrally different label. Four images can then be obtained by using a selective detection channel for each of the four different labels. Alternatively, different nucleotide types can be added sequentially, and an image of the array can be obtained between each addition step. In such an embodiment, each image shows a nucleic acid feature region incorporating a particular type of nucleotide. Because the sequence content of each feature region is different, different feature regions may or may not be present in different images. However, the relative positions of the feature regions remain unchanged within the image. Images obtained from such a reversible terminator-SBS method can be stored, processed, and analyzed as described herein. Following the image capture step, the label can be removed, and the reversible terminator portion can be removed for subsequent nucleotide addition and detection cycles. Removing the marker after detection in a specific cycle and before the subsequent cycle has the advantage of reducing background signal and crosstalk between cycles. Examples of useful marker and removal methods are described below.
[0162] Several embodiments can utilize the detection of four different nucleotides using fewer than four different labels. For example, SBS can be carried out using the method and system described in the incorporated material, U.S. Patent Application Publication No. 2013 / 0079232. As a first example, pairs of nucleotide types can be detected at the same wavelength but can be distinguished based on a difference in intensity for one member of the pair, or based on a change in one member of the pair (e.g., through chemical modification, photochemical modification, or physical modification) that causes a noticeable signal to appear or disappear compared to the signal detected for the other members of the pair. As a second example, three of the four different nucleotide types can be detected under specific conditions, while the fourth nucleotide type has no detectable label under those conditions or is minimally detectable under those conditions (e.g., minimal detection by background fluorescence). Incorporation of the first three nucleotide types into the nucleic acid can be determined based on the presence of their corresponding signals, and incorporation of the fourth nucleotide type into the nucleic acid can be determined based on the absence of any signal or minimal detection. As a third example, one nucleotide type may include a label that is detected by two different channels, while other nucleotide types are detected by one or fewer channels. The three exemplary configurations described above are not considered mutually exclusive and can be used in various combinations.An exemplary embodiment combining all three examples is a fluorescence-based SBS method using a first nucleotide type detected in a first channel (e.g., dATP with a label detected in the first channel when excited by a first excitation wavelength), a second nucleotide type detected in a second channel (e.g., dCTP with a label detected in the second channel when excited by a second excitation wavelength), a third nucleotide type detected in both the first and second channels (e.g., dTTP with at least one label detected in both channels when excited by the first and / or second excitation wavelengths), and an unlabeled fourth nucleotide type (e.g., unlabeled dGTP) that is not detected or minimally detected in any channel.
[0163] Furthermore, as described in the incorporated document, U.S. Patent Application Publication No. 2013 / 0079232, sequencing data can be obtained using a single channel. In such a so-called one-dye sequencing method, a first nucleotide type is labeled but the label is removed after the first image is generated, and a second nucleotide type is labeled only after the first image is generated. A third nucleotide type retains its label in both the first and second images, and a fourth nucleotide type remains unlabeled in both images.
[0164] Several embodiments can utilize sequencing by ligation techniques. Such techniques utilize DNA ligases to incorporate oligonucleotides and identify the incorporation of such oligonucleotides. Oligonucleotides typically have different labels that correlate with the identity of specific nucleotides in the sequence into which the oligonucleotide hybridizes. As with other SBS methods, an image can be obtained after processing an array of nucleic acid sequences with a labeled sequencing reagent. Each image shows nucleic acid feature regions incorporating a specific type of label. Because the sequence content of each feature region is different, different images may or may not contain different feature regions, but the relative positions of the feature regions remain constant within the image. Images obtained from ligation-based sequencing methods can be stored, processed, and analyzed as described herein. Exemplary SBS systems and methods that can be used in conjunction with the methods and systems described herein are described in U.S. Patents 6,969,488, 6,172,218, and 6,306,597, the disclosures of which are incorporated herein by reference in their entirety.
[0165] Several embodiments can utilize nanopore sequencing (Deamer, DW & Akeson, M. "Nanopores and nucleic acids: prospects for ultrarapid sequencing." Trends Biotechnol. 18, 147-151 (2000); Deamer, D. and D. Branton, "Characterization of nucleic acids by nanopore analysis," Acc. Chem. Res. 35: 817-825 (2002); Li, J., M. Gershow, D. Stein, E. Brandin, and J. Golovchenko, "DNA molecules and configurations in a solid-state nanopore microscope," Nat. Mater. 2: 611-615 (2003). These disclosures are incorporated herein by reference in their entirety). In such embodiments, the target nucleic acid passes through the nanopore. The nanopore may be a synthetic pore such as α-hemolysin or a biological membrane protein. As target nucleic acids pass through nanopores, each base pair can be identified by measuring the variation in the electrical conductance of the pores. (U.S. Patent No. 7,001,792; Soni, GV & Meller, "A. Progress toward ultrafast DNA sequencing using solid-state nanopores." Clin. Chem. 53, 1996-2001 (2007); Healy, K., "Nanopore-based single-molecule DNA analysis." Nanomed. 2, 459-481 (2007); Cockroft, SL, Chu, J., Amorin, M. & Ghadiri, MR, "A single-molecule nanopore device detects DNA polymerase activity with single-nucleotide resolution." J. Am Chem. Soc. 130, 818-820 (2008). These disclosures are incorporated herein by reference in their entirety.)Data obtained from nanopore arrangement determination can be stored, processed, and analyzed as described herein. Specifically, the data can be processed as images according to exemplary processing of optical and other images as described herein.
[0166] Some other embodiments of the sequencing method involve the use of 3' block nucleotides in nanoball sequencing techniques as described herein, such as those described in U.S. Patent No. 9,222,132, whose disclosure is incorporated by reference. A large number of discrete DNA nanoballs can be generated through a rolling circle amplification (RCA) process. The nanoball mixture is then distributed onto a patterned slide surface that includes features allowing a single nanoball to associate with each position. In DNA nanoball generation, DNA is fragmented and ligated to the beginning of four adapter sequences. The template is amplified, circularized, and cleaved with a type II endonuclease. A second set of adapters is added and subsequently amplified, circularized, and cleaved. This process is repeated for the remaining two adapters. The final product is a circular template with four adapters, each separated by a template sequence. The library molecules undergo the rolling circle amplification step to generate a large number of concatemers called DNA nanoballs, which are then deposited onto a flow cell. Goodwin et al., “Coming of age: ten years of next-generation sequencing technologies,” Nat Rev Genet.2016;17(6):333-51.
[0167] Some embodiments may utilize methods involving real-time monitoring of DNA polymerase activity. Nucleotide incorporation can be detected, for example, via fluorescence resonance energy transfer (FRET) interactions between fluorophore-containing polymerases and γ-phosphate-labeled nucleotides, as described in U.S. Patents 7,329,492 and 7,211,414, both of which are incorporated herein by reference; or nucleotide incorporation can be detected using zero-mode waveguides, as described in U.S. Patent 7,315,019, both of which are incorporated herein by reference, and fluorescent nucleotide analogs and manipulated polymerases, as described in U.S. Patent 7,405,281 and U.S. Patent Application Publication 2008 / 0108082, both of which are incorporated herein by reference. Illumination can be limited to a zeptolite-scale volume around the surface-tethered polymerase so that the incorporation of fluorescently labeled nucleotides can be observed with low background (Levene, M.J. et al. "Zero-mode waveguides for single-molecule analysis at high concentrations." Science 299, 682-686 (2003); Lundquist, P. et al. "Parallel confocal detection of single molecules in real time." Opt. Lett. 33, 1026-1028 (2008); Korlach, J. et al. "Selective aluminum passivation for targeted immobilization of single DNA polymerase molecules in zero-mode waveguide nanostructures." Proc. Natl. Acad. Sci. USA 105, 1176-1181 (2008), these disclosures are incorporated herein by reference in their entirety). Images obtained by such methods can be stored, processed, and analyzed as described herein.
[0168] Some SBS embodiments involve the detection of protons released during the incorporation of nucleotides into the extension product. For example, sequencing based on the detection of released protons is described in the electrodetectors and related technologies commercially available from Ion Torrent Corporation (Guilford, Connecticut, a subsidiary of Life Technologies), or in U.S. Patent Applications Publications 2009 / 0026082, 2009 / 0127589, 2010 / 0137143, and 2010 / 0282617, all of which are incorporated herein by reference. The methods herein for amplifying target nucleic acids using binding equilibrium exclusion can be readily applied to substrates used for proton detection. More specifically, the methods herein can be used to produce a clonal population of amplicons used for proton detection.
[0169] The above SBS method can be advantageously implemented in a multiplex format so that multiple different target nucleic acids are manipulated simultaneously. In certain embodiments, different target nucleic acids can be processed on a common reaction vessel or on the surface of a specific substrate. This allows for convenient delivery of sequencing reagents, removal of unreacted reagents, and detection of embedded events in multiplex formats. In embodiments using surface-bound target nucleic acids, the target nucleic acids may be in array format. In array format, the target nucleic acids can typically be bound to a surface in a spatially distinguishable manner. The target nucleic acids can be bound by direct covalent bonding, attachment to beads or other particles, or binding to polymerase or other molecules attached to the surface. The array may contain a single copy of the target nucleic acid at each site (also called a feature site), or multiple copies having the same sequence may be present at each site or feature site. Multiple copies can be produced by amplification methods such as bridge amplification or emulsion PCR, which are described in more detail below.
[0170] The methods described herein include, for example, at least about 10 feature parts / cm². 2 , 100 feature parts / cm 2 , 500 feature parts / cm2 1,000 feature parts / cm 2 5,000 feature parts / cm 2 10,000 feature parts / cm 2 50,000 feature parts / cm 2 100,000 feature parts / cm 2 1,000,000 feature parts / cm 2 5,000,000 feature parts / cm 2 Arrays can be used that have feature portions of various densities, including , or more.
[0171] An advantage of the methods described herein is that they provide the rapid and efficient parallel detection of multiple target nucleic acids. Therefore, this disclosure provides an integrated system that allows the preparation and detection of nucleic acids using techniques known in the art, such as those exemplified above. Accordingly, the integrated system of this disclosure may include a fluid component capable of delivering amplification reagents and / or sequencing reagents to one or more immobilized DNA fragments, and the system may include components such as pumps, valves, reservoirs, and fluid lines. Flow cells may constitute and / or be used in the integrated system for detecting target nucleic acids. Exemplary flow cells are described, for example, in U.S. Patent Application Publication 2010 / 0111768 and U.S. Patent Application 13 / 273,666, each incorporated herein by reference. As exemplified with respect to flow cells, one or more fluid components of the integrated system may be used in amplification and detection methods. Taking an embodiment of nucleic acid sequencing as an example, one or more fluid components of the integrated system may be used for the delivery of sequencing reagents in the amplification method described herein and the sequencing method as exemplified above. Alternatively, the integrated system may include separate fluid systems for performing amplification methods and detection methods. Examples of integrated sequencing systems capable of producing amplified nucleic acids and determining the sequence of nucleic acids include, but are not limited to, the MiSeq® platform (Illumina Inc., San Diego, CA) and the device described in U.S. Patent Application No. 13 / 273,666, incorporated herein by reference.
[0172] Arrays in which polynucleotides are directly bound to a silica-based support are disclosed, for example, in International Publication No. 00 / 06770 (incorporated herein by reference), where the polynucleotides are immobilized on the glass support by a reaction between pendant epoxide groups on the glass and internal amino groups on the polynucleotides. Furthermore, polynucleotides can be bound to a solid support by a reaction between a sulfur-based nucleophile and the solid support, as described, for example, in International Publication No. 2005 / 047301 (incorporated herein by reference). Another example of a solid-supported template polynucleotide is one in which the template polynucleotide is attached to a hydrogel supported on a silica-based or other solid support. Examples of this are described, for example, in International Publication Nos. 00 / 31148, 01 / 01143, 02 / 12566, 03 / 014392, U.S. Patent No. 6,465,178, and International Publication No. 00 / 53812, each incorporated herein by reference.
[0173] Specific surfaces on which template polynucleotides can be immobilized include polyacrylamide hydrogels. Polyacrylamide hydrogels are described in the above references and International Publication No. 2005 / 065814, and are incorporated herein by reference. Specific hydrogels that may be used include those described in International Publication No. 2005 / 065814 and U.S. Patent Application Publication No. 2014 / 0079923. In one embodiment, the hydrogel is PAZAM (poly(N-(5-azidoacetamidylpentyl)acrylamide-co-acrylamide)).
[0174] DNA template molecules can be attached to beads or microparticles, for example, as described in U.S. Patent No. 6,172,218 (which is incorporated herein by reference). Attachment to beads or microparticles may be useful for sequencing applications. Bead libraries can be prepared, each containing a different DNA sequence. Exemplary libraries and methods for their preparation are described in Nature, 437, 376-380 (2005); Science, 309, 5741, 1728-1732 (2005), which are incorporated herein by reference, respectively. Sequencing of arrays of such beads using the nucleotides described herein is within the scope of this disclosure.
[0175] The sequencing template may form part of an "array" on a solid support, in which case the array can take any convenient form. Therefore, the method of this disclosure is applicable to all types of high-density arrays, including single-molecule arrays, clustered arrays, and bead arrays. The labeled nucleotides of this disclosure may be used to sequence a template on essentially any type of array, such arrays being formed by the immobilization of nucleic acid molecules on a solid support, but are not limited to these.
[0176] However, the labeled nucleotides of this disclosure are particularly advantageous in the context of sequencing clustered arrays. In a clustered array, distinct regions on the array (often referred to as sites or feature regions) contain multiple polynucleotide template molecules. Generally, multiple polynucleotide molecules are not individually degradable by optical means, but rather detected as an ensemble. Depending on how the array is formed, each site on the array may contain multiple copies of one individual polynucleotide molecule (e.g., the site is homogeneous with respect to a particular single-stranded or double-stranded nucleic acid species), or even multiple copies of a small number of different polynucleotide molecules (e.g., multiple copies of two different nucleic acid species). Clustered arrays of nucleic acid molecules can be generated using techniques generally known in the art. For example, International Publication No. 98 / 44151 and International Publication No. 00 / 18957, each incorporated herein, describe methods for amplifying nucleic acids, in which both the template and amplification product remain immobilized on a solid support to form an array consisting of clusters or "colonies" of immobilized nucleic acid molecules. Nucleic acid molecules present on clustered arrays prepared according to these methods are suitable templates for sequencing using nucleotides labeled with the dye compounds of this disclosure.
[0177] The labeled nucleotides of this disclosure are also useful for sequencing templates on single-molecule arrays. As used herein, the term “single-molecule array” or “SMA” refers to a collection of polynucleotide molecules distributed (or sequenced) on a solid support, such that the separation distance of any individual polynucleotide from all other polynucleotides in the collection allows for the individual identification of each polynucleotide molecule. Thus, in some embodiments, target nucleic acid molecules immobilized on the surface of a solid support can be identified by optical means. This means that one or more distinct signals, each representing a single polynucleotide, occur within an area that can be identified by the particular imaging device used.
[0178] Single-molecule detection can be achieved when the spacing between adjacent polynucleotide molecules on the array is at least 100 nm, more specifically at least 250 nm, even more specifically at least 300 nm, and even more specifically at least 350 nm. Thus, each molecule is individually resolvable (distinguishable) and detectable as a single-molecule fluorescence site, and the fluorescence from this single-molecule fluorescence site also exhibits a single-step photobleach.
[0179] The terms “individually decomposed” and “individual resolution” are used herein to specify that, when visualized, it is possible to distinguish one molecule on an array from its adjacent molecules. The separation between individual molecules on an array is determined in part by specific techniques used to decompose the individual molecules. The general characteristics of single-molecule arrays are understood by referring to International Publications 00 / 06770 and 01 / 57248, respectively, which are incorporated herein by reference. One application of the nucleotides of this disclosure is in synthetic sequencing reactions, but the usefulness of the nucleotides is not limited to such methods. In fact, these nucleotides can be advantageously used in any sequencing method that requires the detection of a fluorescent label bound to a nucleotide incorporated into a polynucleotide.
[0180] Specifically, the labeled nucleotides of this disclosure can be used in automated fluorescence sequencing protocols, particularly in fluorescent dye-terminator cycle sequencing based on the chain termination sequencing method of Sanger and collaborators. Such methods generally use enzymes and cycle sequencing to incorporate fluorescently labeled dideoxynucleotides in a primer extension sequencing reaction. The so-called Sanger sequencing method and related protocols (Sanger-type) utilize chain ends randomized with labeled dideoxynucleotides.
[0181] Therefore, this disclosure also includes labeled nucleotides, which are dideoxynucleotides lacking hydroxyl groups at both the 3' and 2' positions, and such deoxynucleotides are suitable for use in Sanger sequencing and the like.
[0182] It will be recognized that the labeled nucleotides of this disclosure incorporating a 3'-blocking group are useful in Sanger sequencing and related protocols. This is because the same effect achieved by using dideoxynucleotides is obtained by using nucleotides having a 3'-OH blocking group (both can prevent the incorporation of subsequent nucleotides). When the nucleotides having a 3'-blocking group according to this disclosure are used in Sanger sequencing, it will be understood that in each instance in which the labeled nucleotide of this disclosure is incorporated, the dye compound or detectable label attached to the nucleotide does not need to be linked via a cleavable linker, and therefore the label does not need to be removed from the nucleotide, as it does not need to be incorporated afterward.
[0183] In any embodiment of the SBS method described herein, the nucleotide used for sequencing purposes is a 3' block nucleotide as described herein, for example, nucleotides of formulas (I) and (Ia)-(Id). In any embodiment, the 3' block nucleotide is a nucleotide triphot.
[0184] In certain sequencing methods, the incorporated nucleotides are unlabeled. One or more fluorescent labels can be introduced after incorporation by using a label affinity reagent containing one or more fluorescent dyes. For example, each of the one, two, three, or four different types of nucleotides (e.g., dATP, dCTP, dGTP, and dTTP or dUTP) in the incorporation buffer in step (a) may be unlabeled. Each of the four types of nucleotides (e.g., dNTPs) has a 3'-OH blocking group (e.g., 3'-AOM) as described herein to ensure that only one base can be added to the 3' end of the copy polynucleotide by polymerase. After incorporation of the unlabeled nucleotides, an affinity reagent that specifically binds to the incorporated dNTPs is then introduced to provide a labeled extension product containing the incorporated dNTPs. The use of unlabeled nucleotides and affinity reagents in synthetic sequencing is disclosed in U.S. Publication No. 2013 / 0079232. The modified sequencing method of this disclosure using unlabeled nucleotides comprises the following steps: (a'-1) 3'-OH blocking group as described herein [ka] A step of producing an extended copy polynucleotide by incorporating an unlabeled nucleotide (e.g., dATP, dCTP, dGTP, dTTP, or dUTP) (bound to the 3' oxygen) containing into a copy polynucleotide chain complementary to at least a portion of the target polynucleotide chain, (a'-2) A step of contacting an extended copy polynucleotide with a pair of affinity reagents under conditions in which one affinity reagent specifically binds to the incorporated unlabeled nucleotide to provide a labeled extended copy polynucleotide, (b') A step of detecting the identity of nucleotides incorporated into the copy polynucleotide chain by performing fluorescence measurements on one or more labeled extended copy polynucleotides, (c') A step of labeling the extended copy polynucleotide with a detectable label, and chemically removing the 3'-OH blocking group from the nucleotide incorporated into the copy polynucleotide chain, It may include.
[0185] Affinity reagents may include hapten moieties of nucleotides (e.g., streptavidin-biotin, anti-DIG and DIG, anti-DNP and DNP), antibodies (but not limited to antibody binding fragments, single-chain antibodies, bispecific antibodies, etc.), aptamers, Nottins, Affimers, or small molecules or protein tags capable of binding to any other known factor that binds with appropriate specificity and affinity to the incorporated nucleotide. In further embodiments, one affinity reagent may be labeled with multiple copies of the same fluorescent dye. In some embodiments, the Pd catalyst also removes the labeled affinity reagent. For example, the hapten moiety of an unlabeled nucleotide can be cleaved by the Pd catalyst, as described herein for cleavable linkers. [ka] The nucleic acid bases may be bound via (e.g., an AOL linker). In some embodiments, the method further comprises the post-cleavage washing step (d) described herein. In some embodiments, the method further comprises repeating steps (a'-1) to (c') or (a'-1) to (d) until the sequence of at least a portion of the target polynucleotide chain is determined. In some embodiments, the cycle is repeated at least 50, at least 100, at least 150, at least 200, at least 250, or at least 300 times.
[0186] kit This disclosure also provides kits comprising one or more 3'-block nucleosides and / or nucleotides described herein, e.g., 3'-block nucleotides of formulas (I) and (Ia) to (Id). Such kits generally comprise at least one 3'-block nucleotide or nucleoside comprising a detectable label (e.g., a fluorescent dye) having at least one further component. The further component may be one or more of the components identified in the Methods Described herein or in the Examples section below. Some non-limiting examples of components that can be combined in the kits of this disclosure are described below. In some further embodiments, the kit may comprise four types of labeled nucleotides (A, C, T, and G) of the fully functionalized nucleotides described herein, each type of nucleotide comprising the 3'-AOM blocking group and AOL linker moiety described herein. In further embodiments, G is unlabeled and does not contain the AOL linker. In further embodiments, one or more of the remaining three nucleotides (i.e., A, C, and T) are L, which is an allylamine or allylamide linker moiety. 1 This includes. In one embodiment, the kit includes unlabeled ffG, labeled ffA(plural), labeled ffC, and labeled ffT-DB as described herein. In another embodiment, the kit includes unlabeled ffG, labeled ffA(plural), labeled ffC-DB, and labeled ffT-DB as described herein.
[0187] In certain embodiments, the kit may include at least one labeled 3' block nucleotide or nucleoside, along with labeled or unlabeled nucleotides or nucleosides. For example, dye-labeled nucleotides may be supplied unlabeled or in combination with natural nucleotides and / or fluorescently labeled nucleotides, or any combination thereof. Combinations of nucleotides may be supplied as separate individual components (e.g., one nucleotide type per container or tube) or as a nucleotide mixture (e.g., two or more nucleotides mixed in the same container or tube).
[0188] If the kit contains multiple, particularly two, three, or more specifically four, 3' block nucleotides labeled with a dye compound, the different nucleotides may be labeled with different dye compounds or may be dark in color and not contain any dye compound. A key feature of this kit is that, when different nucleotides are labeled with different dye compounds, the dye compounds are spectroscopically distinguishable fluorescent dyes. As used herein, the term “spectroscopically distinguishable fluorescent dye” refers to a fluorescent dye that emits fluorescence energy at wavelengths that can be distinguished by a fluorescence detector (e.g., a commercially available capillary-based DNA sequencing platform) when two or more dyes are present in a single sample. In some embodiments, when two nucleotides labeled with a fluorescent dye compound are supplied in kit form, the spectroscopically distinguishable fluorescent dyes can be excited at the same wavelength, such as by the same laser. In some embodiments, when four 3' block nucleotides (A, C, T, and G) labeled with fluorescent dye compounds are supplied in kit form, two of the spectroscopically distinguishable fluorescent dyes may be excited at one wavelength, while the other two spectroscopically distinguishable dyes may both be excited at a different wavelength. The specific excitation wavelengths are 488 nm and 532 nm.
[0189] In one embodiment, the kit comprises a first 3' block nucleotide labeled with a first dye and a second nucleotide labeled with a second dye, wherein the dyes have a difference in maximum absorbance of at least 10 nm, particularly between 20 nm and 50 nm. More specifically, the two dye compounds have a Stokes shift of 15 to 40 nm. Note that "Stokes shift" is the distance between the peak absorption wavelength and the peak emission wavelength.
[0190] In alternative embodiments, the kits of the present disclosure may include 3' block nucleotides in which the same base is labeled with two or more different dyes. A first nucleotide (e.g., a 3' block T nucleotide triphot or a 3' block G nucleotide triphot) may be labeled with a first dye. A second nucleotide (e.g., a 3' block C nucleotide triphot) may be labeled with a second spectrally different dye from the first dye, e.g., a “green” dye that absorbs below 600 nm, or a “blue” dye that absorbs below 500 nm, e.g., 400 nm to 500 nm, particularly 450 nm to 460 nm. A third nucleotide (e.g., a 3' block A nucleotide triphot) may be labeled as a mixture of the first and second dyes, or a mixture of the first, second, and third dyes, and a fourth nucleotide (e.g., a 3' block G nucleotide triphot or a 3' block T nucleotide triphot) may be “dark” and may not be labeled. For example, nucleotides 1-4 can be labeled with "blue," "green," "blue / green," and dark. To further simplify the apparatus, the four nucleotides can be labeled with two dyes excited by a single laser, and the labeling of nucleotides 1-4 can be "blue 1," "blue 2," "blue 1 / blue 2," and dark.
[0191] In certain embodiments, the kit may comprise four labeled 3'-blocking nucleotides (e.g., A, C, T, G), where each type of nucleotide comprises the same 3'-blocking group and fluorescent label, each fluorescent label having a distinct fluorescence maximum value, and each fluorescent label being distinguishable from the other three labels. The kit may also have two or more of the fluorescent labels having different Stokes shifts, although they are absorption maximum values. In some other embodiments, one type of nucleotide is unlabeled.
[0192] With respect to configurations having different nucleotides labeled with different dye compounds, kits are exemplified herein, but it will be understood that kits may contain two, three, four or more different nucleotides having the same dye compound. In some embodiments, the kit also includes an enzyme and a buffer suitable for the action of the enzyme. In some such embodiments, the enzyme is a polymerase, a terminal deoxynucleotidyltransferase, or a reverse transcriptase. In certain embodiments, the enzyme is a DNA polymerase such as DNA polymerase 812 (Pol812) or DNA polymerase 1901 (Pol1901). In some further embodiments, the kit may include the integration mixture described herein. In further embodiments, the kit containing the integration mixture described herein also includes at least one Pd scavenger (e.g., a Pd(0) scavenger described herein comprising one or more allyl moieties). The Pd(0) scavengers include -O-allyl, -S-allyl, -NR-allyl, and -N + The formula comprises one or more allyl moieties independently selected from the group consisting of RR'-allyls, where R is H, an unsubstituted or substituted C1-C6 alkyl, an unsubstituted or substituted C2-C6 alkenyl, an unsubstituted or substituted C2-C6 alkynyl, or an unsubstituted or substituted C6-C 10 Aryl, unsubstituted or substituted 5-10 member heteroaryl, unsubstituted or substituted C3-C 10 The Pd(0) scavenger is a carbocyclyl, or an unsubstituted or substituted 5- to 10-membered heterocyclyl, where R' is H, an unsubstituted C1-C6 alkyl, or a substituted C1-C6 alkyl. In some such embodiments, the Pd(0) scavenger in the incorporated solution comprises one or more -O-allyl moieties. In some further embodiments, the Pd(0) scavenger is [ka] Or it includes or is a combination thereof. An alternative Pd(0) scavenger is disclosed in U.S. Patent No. 63 / 190983, which is incorporated by reference in whole. In one embodiment, the Pd(0) scavenger in the incorporated mixture is [ka] In another embodiment, the Pd(0) scavenger in the incorporated mixture is [ka] It includes or is.
[0193] Other components included in such kits include buffers. The nucleotides of this disclosure, and any other nucleotide components including mixtures of different nucleotides, may be provided in the kit in a concentrated form that is diluted before use. In such embodiments, a suitable dilution buffer may also be included. For example, the integrated mixture kit may include one or more buffers selected from primary amines, secondary amines, tertiary amines, natural or non-natural amino acids, or combinations thereof. In further embodiments, the buffer in the integrated mixture may include ethanolamine or glycine, or combinations thereof.
[0194] Herein too, one or more of the components identified by the methods described herein can be included in the kit of this disclosure. In some further embodiments, the kit may include the palladium catalyst described herein. In some embodiments, the Pd catalyst is produced by mixing a Pd(II) complex (i.e., a Pd pre-catalyst) with one or more water-soluble phosphines described herein. In some such embodiments, the kit containing the Pd catalyst is a cleavage mix kit. In further embodiments, the cleavage mix kit may include Pd(allyl)Cl]2 or Na2PdCl4 and the water-soluble phosphine THP to produce an active Pd(0) species. The molar ratio of the Pd(II) complex (e.g., Pd(allyl)Cl]2 or Na2PdCl4) to the water-soluble phosphine (e.g., THP) may be about 1:2, 1:3, 1:4, 1:5, 1:6, 1:7, 1:8, 1:9, or 1:10. In further embodiments, the cleavage mixing kit may also include one or more buffering reagents selected from the group consisting of primary amines, secondary amines, tertiary amines, carbonates, phosphates, and borates, and combinations thereof. Non-limiting examples of buffering reagents in the cleavage mixing kit are selected from the group consisting of ethanolamine (EA), tris(hydroxymethyl)aminomethane (Tris), glycine, carbonates, phosphates, borates, dimethylethanolamine (DMEA), diethylethanolamine (DEEA), N,N,N',N'-tetramethylethylenediamine (TEMED) and N,N,N',N'-tetraethylethylenediamine (TEEDA), 2-piperidineethanol, and combinations thereof. In one embodiment, the cleavage mixing kit contains DEEA. In other embodiments, the cleavage mixing kit contains 2-piperidineethanol.
[0195] In some further embodiments, the kit may include one or more palladium scavengers (e.g., Pd(II) scavengers as described herein). In some such embodiments, the kit is a post-cutting wash buffer kit. Non-limiting examples of Pd scavengers in a post-cutting wash buffer kit include isocyanoacetate (ICNA) salts, ethyl isocyanoacetate, methyl isocyanoacetate, cysteine or salts thereof, L-cysteine or salts thereof, N-acetyl-L-cysteine, potassium ethylxanthogenic acid, potassium isopropylxanthogenic acid, glutathione, lipoic acid, ethylenediaminetetraacetic acid (EDTA), iminodiacetic acid, nitrilodiacetic acid, trimercapto-S-triazine, dimethyldithiocarbamate, dithiothreitol, mercaptoethanol, allyl alcohol, propargyl alcohol, thiols, tertiary amines and / or tertiary phosphines, or combinations thereof. In one embodiment, the post-cutting wash buffer kit includes L-cysteine or a salt thereof.
[0196] In any embodiment of the kit described herein, the Pd scavenger (e.g., the Pd(0) or Pd(II) scavenger described herein) is located in a separate container / compartment from the Pd catalyst. [Examples]
[0197] Additional embodiments are disclosed in more detail in the following embodiments, but these embodiments are not intended to limit the scope of the claims. Example 1. Synthesis of fully functionalized nucleotides having 3'AOM and AOL linker moieties [ka]
[0198] Synthesis of intermediate AOL LN2: Acetal compound LN1 (2.43 g, 9.6 mmol) was dissolved in anhydrous CH2Cl2 (100 mL) under N2, and the solution was cooled to 0°C in an ice bath. 2,4,6-Trimethylpyridine (7.6 mL, 57.5 mmol) was added, followed by the dropwise addition of trimethylsilyl trifluoromethanesulfonate (7.0 mL, 38.7 mmol). The mixture was stirred at 0°C for 2 hours, then allyl alcohol (13 mL, 191.1 mmol) was added, and the reaction was refluxed overnight. The reaction was quenched with a 98:2 mixture of MeOH / H2O, and the resulting solution was further stirred at room temperature for 3 hours. The mixture was diluted with CH2Cl2 (100 mL) and water (200 mL), and the aqueous layer was acidified to pH 2-3 with 2N HCl. The aqueous layer was separated, and the organic layer was further extracted with acidic water. The organic layer was dried over MgSO4, filtered, and volatile substances were evaporated under reduced pressure. The crude product was purified by flash chromatography using silica gel to obtain AOL LN2 as a colorless oil (2.56 g, 86%).
[0199] Synthesis of intermediate AOL LN3: A solution of AOL LN2 (2.17 g, 7.0 mmol) in ethanol (17.5 mL) was mixed with 4 M NaOH aqueous solution (17.5 mL, 70 mmol), and the mixture was stirred at room temperature for 3 hours. Afterward, all volatile substances were removed under reduced pressure, and the residue was dissolved in 75 mL of water. The solution was acidified to pH 2-3 with 2N HCl, and then extracted with dichloromethane (DCM). The combined organic fraction was dried over MgSO4, filtered, and volatile substances were evaporated under reduced pressure. AOL LN3 was obtained as a colorless oil that solidified upon storage at 20°C without further purification (1.65 g, 84%). LC-MS(ES):(Negative ion)m / z281(MH) + ); (Positive ions) m / z 305 (M+Na + ).
[0200] Synthesis of intermediate AOL LN4: A solution of AOL LN3 (1.62 g, 5.74 mmol) in anhydrous DMF (20 mL) was stirred under vacuum for 5 minutes, and then cooled to 0°C using an ice bath. N,N-diisopropylethylamine (1.2 mL, 6.89 mmol) was added dropwise under N2, followed by the dropwise addition of PyBOP (3.30 g, 6.34 mmol). The reaction mixture was stirred at 0°C for 30 minutes, then a solution of N-(5-aminopentyl)-2,2,2-trifluoroacetamide hydrochloride (1.62 g, 6.90 mmol) in anhydrous DMF (3.0 mL) was added, followed immediately by the addition of N,N-diisopropylethylamine (1.4 mL, 8.04 mmol). The reaction mixture was removed from the ice bath and stirred at room temperature for 4 hours. Volatile substances were removed under reduced pressure, and the residue was dissolved in RINKAN (150 mL). The solution was extracted with 20 mM KHSO4 aqueous solution, water, and saturated NaHCO3 aqueous solution. The organic layer was dried over MgSO4, filtered, and volatile substances were evaporated under reduced pressure. The crude product was purified by flash chromatography on silica gel to obtain AOL LN4 as a colorless oil (2.16 g, 82%). LC-MS(ES): (negative ion) m / z 461(MH) + ),497(M-Cl - ).
[0201] Synthesis of the AOL linker portion To a solution of AOL LN4 (350 mg, 0.76 mmol) in CH3CN (13 mL), add TEMPO (48 mg, 0.31 mmol), then add NaH2PO4 in water (6.5 mL). .Solutions of 2H2O (762 mg, 4.88 mmol) and NaClO2 (275 mg, 3.04 mmol) were added. An aqueous solution of NaClO (14% available chlorine, 0.83 mL, 1.94 mmol) was added, and the solution immediately turned dark brown. The reaction mixture was stirred at room temperature for 6 hours, and then quenched with a 100 mM aqueous solution of Na2S2O3 until the mixture became colorless. Acetonitrile was removed under reduced pressure, the residue was diluted with water, and basicized with triethylamine. The aqueous phase was extracted with RINKAN (10 mL) and then concentrated under reduced pressure. The crude product was purified by reverse-phase flash chromatography on C18 to obtain AOL as a colorless oil (triethylammonium salt, 310 mg, 71%). LC-MS(ES): (negative ion) m / z 475 (MH) + ); (Positive ions) m / z 499 (M+Na) + ),578(M+Et3NH + ).
[0202] AOL-NH 2 Linker part synthesis A solution of AOL (446 mg, 0.94 mmol) in methanol (10 mL) was mixed with NH3 aqueous solution (35%, 40 mL), and the mixture was stirred at room temperature for 5.5 hours. After this time, all volatile substances were removed under reduced pressure, and the crude product was purified by reverse-phase flash chromatography on C18 to obtain AOL NH2 as a white solid (quantitative). 1 H NMR(400MHz,DMSO-d6):δ(ppm)8.86(t,J=5.5Hz,1H,CONH),8.28(s,3H,NH3 + ),7.85(s,1H,Ar-H),7.41(d,J=7.6Hz,1H,Ar-H),7.31(t,J=7.9Hz,1H,Ar-H),7.03(ddd,J=8.1,2.5,1.1 Hz,1H,Ar-H),5.87(ddt,J=17.2,10.5,5.3Hz,1H,OCH2CHCH2),5.24(dq,J=17.2,1.7Hz,1H,OCH2CHCH2,H a ),5.09(dq,J=10.5,1.5Hz,1H,OCH2CHCH2,H b),5.02(dd,J=6.7,2.4Hz,1H,OCHO),4.41(dd,J=12.2,2.5Hz,1H,OCH2,H a ),4.18-3.99(m,3H,OCH2CHCH2 and OCH2,H b ),3.94-3.81(m,2H,OCH2COOH),3.49-3.39(m,1H,CH2,H a ),3.21-3.10(m,1H,CH2,H b ),2.86-2.70(m,2H,CH2),1.81-1.39(m,6H,CH2). 13 ¹³C NMR (101MHz, DMSO-d6): δ(ppm) 172.9, 166.1, 158.0, 136.2, 135.2, 129.3, 120.4, 119.3, 116.0, 111.3, 99.0, 68.8, 67.7, 66.8, 38.7, 38.4, 27.8, 26.3, 23.0. LC-MS (ESI): (Negative ions) 379 (MH); (Positive ions) m / z 381 (M+H + ).
[0203] General procedure for dye-AOL linker coupling: 0.15 mmol of the dye carboxylate was dissolved in 6 mL of anhydrous N,N'-dimethylformamide (DMF). 136 μL, 0.78 mmol of N,N-diisopropylethylamine was added, followed by the addition of N,N,N',N'-tetramethyl-O-(N-succinimidyl)uronium tetrafluoroborate as a 0.5 M solution in anhydrous DMF (TSTU, 300 μL, 0.15 mmol). The reaction mixture was stirred under nitrogen at room temperature for 1 hour. A solution of AOL NH2 (0.10 mmol) in water (400 μL) was added to the activated dye solution, and the reaction mixture was stirred at room temperature for 3 hours. The crude product was purified by preparative scale RP-HPLC. [ka]
[0204] Identification of AOL-SO7181: 54% yield (54 μmol). LC-MS(ES): (Negative ion) m / z 1002 (MH) + ); (Positive ions) m / z 1004(M+H + ). [ka]
[0205] Identification of AOL-AF550POPOS0: 88% yield (88 μmol). LC-MS(ES): (Negative ion) m / z 1034 (MH) + ), 516 (M-2H + ); (Positive ions) m / z 1036 (M+H + ),1137(M+Et3NH + ). [ka]
[0206] Identification of AOL-NR7180A: Yield 24% (23.9 μmol). LC-MS(ES): (Positive ion) m / z = 880 (M + H) + . [ka]
[0207] Identification of AOL-NR550S0: Yield 22% (21.9 μmol). LC-MS(ES): (Positive ion) m / z = 963 (M + H) + (Negative ions) m / z = 961 (MH) + ). [ka]
[0208] Synthesis of intermediate A2: Compound A1 (319 mg, 0.419 mmol) was dissolved in 0.8 mL of anhydrous DCM under an N2 atmosphere, and then pentamethylcyclopentadienyltris(acetonitrile)ruthenium(II) hexafluorophosphate ([RuCp * (MeCN)3]PF6, 42 mg, 0.08 mmol), followed by triethoxysilane (231 μL, 1.25 mmol). The reaction mixture was stirred under N2 at room temperature for 1 hour. The solution was then diluted with DCM, filtered through a silica gel plug, and washed with ethyl acetate. The solution was evaporated under reduced pressure and dried under vacuum for 10 minutes, and the residue was dissolved in 2 mL of anhydrous THF. Copper iodide (15 mg, 0.08 mmol) and 1 M TBAF solution in THF (920 μL, 0.919 mmol) were added. The reaction mixture was stirred at room temperature for 2.5 hours, then diluted with siRNA and extracted with saturated NH4Cl. The aqueous phase was extracted with siRNA. The pooled organic phase was dried over MgSO4, filtered, and evaporated to dryness. The product was purified by flash column on silica gel. Yield: 125 mg (0.237 mmol). LC-MS (ES and CI): (Positive ion) m / z 527 (M+H + ).
[0209] Synthesis of intermediate A3: Nucleoside A2 (155 mg, 0.294 mmol) was dried under reduced pressure over P2O5 for 18 hours. Anhydrous triethyl phosphate (1 mL) and some freshly activated 4 Å molecular sieves were added under nitrogen, and the reaction flask was then cooled to 0°C in an ice bath. Freshly distilled POCl3 (33 μL, 0.353 mmol) was added dropwise, followed by the addition of Proton Sponge® (113 mg, 0.53 mmol). After the addition, the reaction mixture was stirred further at 0°C for 15 minutes. Then, a 0.5 M solution of pyrophosphate as bis-tri-n-butylammonium salt (2.94 mL, 1.47 mmol) in anhydrous DMF was rapidly added, followed immediately by the addition of tri-n-butylamine (294 μL, 1.32 mmol). The reaction mixture was held in an ice bath for a further 10 minutes, then quenched by pouring it into a 1 M aqueous solution of triethylammonium bicarbonate (TEAB, 10 mL), and stirred at room temperature for 4 hours. All solvent was evaporated under reduced pressure. A 35% aqueous solution of ammonia (10 mL) was added to the above residue, and the mixture was stirred at room temperature for at least 5 hours. The solvent was then evaporated under reduced pressure. The crude product was first purified by ion-exchange chromatography using DEAE-Sephadex A25 (50 g). The column was eluted with aqueous triethylammonium bicarbonate (TEAB). The fraction containing triphot was pooled, and the solvent was evaporated to dryness under reduced pressure. The crude product was further purified by preparative scale HPLC. Compound A3 was obtained as the triethylammonium salt. Yield: 134 μmol (46%). LC-MS (ESI): (negative ion) m / z 614 (MH) + ).
[0210] Furthermore, structure [ka] The 5'-triphosphonate-3'-AOM-A nucleotide and the corresponding ffA were also prepared. Detailed synthesis is described in U.S. Patent Application Publication No. 16 / 724,088.
[0211] General synthesis of nucleotide triphosphate-AOL linkers: Compound AOL (0.120 mmol) was co-evaporated with 2 × 2 mL of anhydrous N,N'-dimethylformamide (DMF), and then dissolved in 3 mL of anhydrous N,N'-dimethylacetamide (DMA). N,N-diisopropylethylamine (70 μL, 0.4 mmol) was added, followed by N,N,N',N'-tetramethyl-O-(N-succinimidyl)uronium tetrafluoroborate (TSTU, 36 mg, 0.120 mmol). The reaction mixture was stirred under nitrogen at room temperature for 1 hour. Meanwhile, an aqueous solution of nucleotide triphotphosphate (0.08 mmol) was evaporated to dryness under reduced pressure and resuspended in 300 μL of 0.1 M triethylammonium bicarbonate (TEAB) aqueous solution. The activated linker solution was added to the triphotphosphate salt, and the reaction mixture was stirred at room temperature for 18 hours and monitored by RP-HPLC. The solution was concentrated, and then 10 mL of concentrated NH4OH aqueous solution was added. The reaction mixture was stirred at room temperature for 24 hours, and then evaporated under reduced pressure. The crude product was first purified by ion-exchange chromatography using DEAE-Sephadex A25 (50 g) eluted with aqueous triethylammonium bicarbonate (TEAB). The fraction containing triphot was pooled, and the solvent was evaporated to dryness under reduced pressure. The crude product was further purified by preparative scale HPLC. [ka]
[0212] Identification of pppA(DB)-(3'AOM)-AOL: Yield: 60 μmol, (75%). LC-MS(ES): (Anion) m / z 976 (MH) + ), 488 (M-2H + ). [ka]
[0213] Identification of pppA(DB)-(3'AOM)-AOL: Yield: 68 μmol, (72%). 1H NMR(400MHz,D2O): δ(ppm)7.94(d,J=1.7Hz,1H,H-2),7.32(d,J=1.6Hz,1H,H-8),7.10-6.96(m,2 H,Ar),6.95-6.87(m,2H,Ar),6.41-6.31(m,1H,1'-CH),6.01-5.81(m,2H,CHアリル),5.29(ddq,J=1 7.2,6.1,1.4Hz,2H,CHHアリル),5.20(ddt,J=10.5,3.8,1.1Hz,2H,CHHアリル),5.02(td,J=4.4,2.4Hz ,1H,O-CH2-Oリンカー),4.92-4.81(m,2H,3'-O-CH2-O),4.53(dd,J=4.9,2.4Hz,1H,3'-CH),4.39(dd, J=16.1,4.0Hz,1H,O-CHH リンカー),4.32-4.19(m,2H,O-CHH リンカー,4'-CH),4.17-4.01(m,8H,5'-CH2, CH2Oリンカー,CH2-Oアリル),3.25-3.11(m,2H,CH2-NHCO),3.04(q,J=7.3Hz,18H,Et3N),2.92-2.82(m,2 H,CH2-Nリンカー),2.55-2.41(m,2H,2'-CH2),1.59(p,J=7.6Hz,2H,CH2-CH2-Nリンカー),1.47(p,J=7.1 Hz,2H,CH2-CH2-NHCO),1.31(tt,J=8.3,4.4Hz,2H,CH2-CH2-CH2-Nリンカー),1.15(t,J=7.3Hz,26H). 31 P NMR(162MHz,D2O):δ(ppm)-6.18(d,J=20.6Hz, γ P), -11.32(d, J=19.3Hz, α P), -22.20(t, J=19.9Hz, β P).
change
[0214] Synthesis of intermediate C1 5-iodo-5'-O-(tert-butyldiphenylsilyl)-2'-deoxycytidine (3 g, 5.07 mmol) was dissolved in 30 mL of anhydrous pyridine, and then chlorotrimethylsilane (1.29 mL, 10.1 mmol) was added dropwise. The reaction mixture was stirred at room temperature for 1 hour, then placed in an ice bath, and benzoyl chloride (648 μL, 5.6 mmol) was slowly added dropwise. The reaction mixture was removed from the ice bath and stirred at room temperature for 1 hour. Once complete, the solution was placed in an ice bath, quenched with 50 mL of cold water, then 50 mL of methanol and 20 mL of pyridine were added, and the suspension was stirred at room temperature overnight. The solvent was evaporated under reduced pressure, and the residue was dissolved in 200 mL of siRNA and extracted with 2 x 200 mL of saturated NaHCO3 and 100 mL of brine. The organic phase was dried over MgSO4, filtered, and evaporated to dryness. The crude product was purified by flash chromatography on silica gel to obtain C1. Yield: 2.535g (3.64 mmol, 73%). LC-MS (ESI): (Positive ion) m / z 696 (M+H + ),797(M+Et3NH + ).
[0215] Synthesis of intermediate C2 N-benzoyl-5-iodo-5'-O-(tert-butyldiphenylsilyl)-2'-deoxycytidine (C1) (695 mg, 1 mmol) and palladium(II) acetate (190 mg, 0.85 mmol) were dissolved in 10 mL of dried, degassed DMF, and then N-allyltrifluoroacetamide (7.65 mL, 5 mmol) was added. The solution was placed under vacuum and purged three times with nitrogen, and then degassed triethylamine (278 μL, 2 mmol) was added. The solution was heated to approximately 80°C and protected from light for 1 hour. The resulting black mixture was cooled to room temperature, then diluted with 50 mL of ethylacetate, and then extracted with 100 mL of water. The aqueous phase was then extracted with ethylacetate. The organic phase was pooled, dried over MgSO4, filtered, and evaporated to dryness. The crude product was purified by flash chromatography on silica gel to obtain C2. Yield: 305 mg (0.42 mmol, 42%). LC-MS (ES and CI): (Positive ion) m / z 721 (M+H +),797(M+Et3NH + ).
[0216] Synthesis of intermediate C3 N-benzoyl-5-[3-(2,2,2-trifluoroacetamide)-allyl]-5'-O-(tert-butyldiphenylsilyl)2'-deoxycytidine (C2) (350 mg, 0.486 mmol) was dissolved in 1.1 mL of anhydrous DMSO (14.5 mmol), and then glacial acetic acid (1.7 mL, 29.1 mmol) and anhydrous acetic acid (1.7 mL, 17 mmol) were added. The reaction mixture was heated at 60°C for 6 hours and then quenched with 50 mL of saturated NaHCO3. After the solution stopped foaming, it was extracted with ELISA. The organic phase was pooled and washed with saturated NaHCO3 aqueous solution, water, and brine. The organic phase was dried over MgSO4, filtered, and evaporated to dryness. The crude product was purified by flash chromatography on silica gel to obtain C3. Yield: 226 mg (0.289 mmol, 60%). LC-MS(ESI): (Positive ions) m / z 781(M+H + ),882(M+Et3NH + ).
[0217] Synthesis of intermediate C4 N-benzoyl-5-[3-(2,2,2-trifluoroacetamide)-allyl]-5'-O-(tert-butyldiphenylsilyl)-3'-O-methylthiomethyl-2'-deoxycytidine (C3) (210 mg, 0.27 mmol) was dissolved in 5 mL of anhydrous DCM under an N2 atmosphere, and cyclohexene (136 μL, 1.35 mmol) was added. The solution was cooled to approximately -10°C. A 1 M solution of freshly distilled sulfuryl chloride in anhydrous DCM (320 μL, 0.32 mmol) was added dropwise, and the reaction mixture was stirred for 20 minutes. After all the starting materials had been consumed, the excess cyclohexene (136 μL, 1.35 mmol) was added, and the reaction mixture was evaporated to dryness under reduced pressure. The residue was rapidly purged with nitrogen, then dissolved in 2.5 mL of ice-cold anhydrous DCM, and ice-cold allyl alcohol (2.5 mL) was added while stirring at 0°C. The reaction mixture was stirred at 0°C for 3 hours, then quenched with saturated NaHCO3 aqueous solution, and then further diluted with saturated NaHCO3 aqueous solution. The mixture was extracted with RINKAN. The pooled organic phase was dried over MgSO4, filtered, and evaporated to dryness. The residue was purified by flash chromatography on silica gel to obtain C4. Yield: 58% (124 mg, 0.157 mmol). LC-MS (ESI): (positive ion) m / z 791 (M+H + ).
[0218] Synthesis of intermediate C5 N-benzoyl-5-[3-(2,2,2-trifluoroacetamide)-allyl]-5'-O-(tert-butyldiphenylsilyl)-3'-O-allyloxymethyl-2'-deoxycytidine (C4) (120 mg, 0.162 mmol) was dissolved in dry THF (5 mL) under an N2 atmosphere and then left at 0°C. Glacial acetic acid (29 μL, 0.486 mmol) was added, followed immediately by the addition of a solution of 1.0 M TBAF (486 μL, 0.486 mmol) in THF. The solution was stirred at 0°C for 3 hours. The solution was diluted with ELISA and then extracted with 0.025 N HCl and brine. The organic phase was dried over MgSO4, filtered, and evaporated to dryness. The residue was purified by flash chromatography on silica gel to obtain C5. Yield: 50 mg (0.090 mmol, 55%). LC-MS(ESI): (Positive ions) m / z 553 (M+H + );(Negative ions) m / z 551 (MH) + ),587(M+Cl - ).
[0219] Synthesis of intermediate C6 N-benzoyl-3'-O-allyloxymethyl-5-[3-(2,2,2-trifluoroacetamide)-allyl]-2'-deoxycytidine (C5) (50 mg, 0.0.09 mmol) was dried on P2O5 under reduced pressure for 18 hours. Anhydrous triethyl phosphate (1 mL) and some newly activated 4 Å molecular sieves were added under nitrogen, and the reaction flask was then cooled to 0°C in an ice bath. Freshly distilled POCl3 (10 μL, 0.108 mmol) was added dropwise, followed by the addition of Proton Sponge® (29 mg, 0.135 mmol). After the addition, the reaction mixture was stirred further at 0°C for 15 minutes. Next, a 0.5 M solution of pyrophosphate as bis-tri-n-butylammonium salt (1 mL, 0.45 mmol) in anhydrous DMF was rapidly added, followed immediately by the addition of tri-n-butylamine (100 μL, 0.4 mmol). The reaction mixture was held in an ice bath for a further 10 minutes, then quenched by pouring it into a 1 M aqueous solution of triethylammonium bicarbonate (TEAB, 5 mL), and stirred at room temperature for 4 hours. All solvent was evaporated under reduced pressure. A 35% aqueous solution of ammonia (5 mL) was added to the above residue, and the mixture was stirred at room temperature for 18 hours. The solvent was then evaporated under reduced pressure, and the residue was resuspended in 10 mL of 0.1 M TEAB and filtered. The filtrate was first purified by ion-exchange chromatography using DEAE-Sephadex A25 (50 g). The column was eluted with aqueous triethylammonium bicarbonate. The fraction containing the triphosphate was pooled, and the solvent was evaporated under reduced pressure to dryness. The crude material was further purified by preparative scale HPLC. Compound C6 was obtained as the triethylammonium salt. Yield: ε 290 =5041M -1 cm -1 Based on this, 40.6 μmol (45%). 1H NMR(400MHz,D2O):δ(ppm)8.23(s,1H,H-6),6.53(dd,J=15.5,0.9Hz,1H,Ar-CH =),6.42-6.24(m,2H,Ar-CH=CH-,1'-CH),5.98(ddt,J=17.3,10.4,5.9Hz,1H,O -CH2-CH=),5.37(dq,J=17.3,1.6Hz,1H,CHH=),5.29(ddt,J=10.4,1.6,1.1Hz, 1H,CHH=),4.89(s,2H,O-CH2-O),4.60(dt,J=6.2,3.1Hz,1H,3'-CH),4.39(t,J= 2.7Hz,1H,4'-CH),4.35(dq,J=12.0,3.8Hz,1H,5'-CHH),4.28-4.21(m,1H,5'- CHH),4.20(ddt,J=6.0,2.7,1.4Hz,1H,=CH-CH2-O),3.73(dt,J=7.2,1.4Hz,2H, CH2-NH2),3.18(q,J=7.3Hz,20H,Et3N),2.59(ddd,J=14.1,6.1,3.3Hz,1H,2'- CHH),2.37(ddd,J=14.2,7.2,6.1Hz,1H,2'-CHH),1.27(t,J=7.3Hz,31H,Et3N). 31 P NMR(162MHz,D2O):δ(ppm)-6.06(d,J=20.7Hz, γ P), -11.24(d, J=19.1Hz, α P), -21.95(t, J=19.7Hz, β P). LC-MS(ESI): (Negative ions) m / z 591 (MH) + ).
[0220] Furthermore, the following structures are 5'-triphosphonate-3'-AOM-C nucleotide, 5'-triphosphonate-3'-AOM-T(DB) nucleotide, [ka] In addition, corresponding ffC and ffT(DB) were prepared. Finally, 5'-triphosphate-3'-AOM-G (also called ffG-(3'-AOM)) [ka] It was also prepared. The detailed synthesis is described in U.S. Publication No. 2020 / 0216891. General synthesis of fully functionalized nucleotides with AOL linker moieties
[0221] Dye-COOH (0.02 mmol) or Dye-AOL (0.02 mmol) was co-evaporated with 2 x 2 mL of anhydrous N,N'-dimethylformamide (DMF), and then dissolved in 2 mL of anhydrous N,N'-dimethylacetamide (DMA). N,N-diisopropylethylamine (17 μL, 0.1 mmol) was added, followed by N,N,N',N'-tetramethyl-O-(N-succinimidyl)uronium tetrafluoroborate (TSTU, 6 mg, 0.02 mmol). The reaction mixture was stirred under nitrogen at room temperature for 1 hour. Meanwhile, an aqueous solution of nucleotide triphotphosphate (0.01 mmol) was evaporated to dryness under reduced pressure and resuspended in 200 μL of 0.1 M triethylammonium bicarbonate (TEAB) aqueous solution. The activated dye solution was added to the nucleotide triphotphosphate, and the reaction mixture was stirred at room temperature for 18 hours and monitored by RP-HPLC. The crude product was first purified by ion-exchange chromatography using DEAE-Sephadex A25 (25 g) with elution under a linear gradient of aqueous triethylammonium bicarbonate (TEAB, 0.1 M-1 M). The fraction containing triphot phosphate was pooled, and the solvent was evaporated to dryness under reduced pressure. The crude product was further purified by preparative scale HPLC. [ka]
[0222] Identification of ffA-(3'AOM)-AOL-BL-NR650C5: Yield 66% (6.6 μmol). LC-MS(ES): (Negative ion) m / z 1931 (MH) + ),965(M-2H + ). [ka]
[0223] Identification of ffA-(3'AOM)-AOL-BL-NR550s0: Yield 65% (6.5 μmol). LC-MS(ES): (Negative ion) m / z 1784 (MH) + ),891(M-2H + ). [ka]
[0224] Identification of ffA(DB)-(3'AOM)-AOL-BL-NR650C5: Yield 31% (3.1 μmol). LC-MS(ES): (Negative ion) m / z 1933 (MH) + ),965(M-2H + ), 644 (M-3H + ). [ka]
[0225] Identification of ffA(DB)-(3'AOM)-AOL-NR7181A: Yield 21% (2.12 μmol). LC-MS(ES): (Negative ion) m / z = 1475 (MH) + ). [ka]
[0226] Identification of ffA-(3'-AOM)-AOL-NR7180A: Yield 41% (4.1 μmol). LC-MS(ES): (Negative ion) m / z 1472 (MH) + ),736(M-2H + ). [ka]
[0227] Identification of ffA(DB)-(3'-AOM)-AOL-BL-NR550S0: Yield 21% (2.1 μmol). LC-MS(ES): (Negative ion) m / z 1786 (MH)+ ),892(M-2H + ),594(M-3H + ). [ka]
[0228] Identification of ffC(DB)-(3'AOM)-AOL-SO7181: Yield 48%, (4.87 μmol). LC-MS(ES): (Negative ion) m / z 1577 (MH) + ),788(M-2H + ),525(M-3H + ). [ka]
[0229] Identification of ffC-(3'-AOM)-AOL-SO7181: Yield 56% (5.6 μmol). LC-MS(ES): (Negative ion) m / z 1575 (MH) + ),787(M-2H + ). [ka]
[0230] Identification of ffT(DB)-3'AOM-AOL-AF550POPOS0: Yield 46% (4.6 μmol). LC-MS(ES): (Negative ion) m / z 1609 (MH) + ),804(M-2H + ),536(M-3H + ). [ka]
[0231] Identification of ffT(DB)-(3'AOM)-AOL-NR550s0: Yield 38% (3.8 μmol). LC-MS(ES): (Negative ion) m / z = 1535 (MH) + ).
[0232] Example 2. Solution cleavage efficiency of different palladium reagent formulations Figure 1 shows a comparison of the solution cleavage efficiencies of three different formulations of palladium reagent: 1) 10 mM [(allyl)PdCl]2, 100 mM THP, 100 mM ethanolamine buffer, 10 mM sodium ascorbate; 2) 20 mM Na2PdCl4, 60 mM THP, 100 mM N,N'-diethylethanolamine, 10 mM sodium ascorbate, 1 M NaCl, 0.1% Tween20; and 3) 20 mM Na2PdCl4, 70 mM THP, 100 mM N,N'-diethylethanolamine, 10 mM sodium ascorbate, 1 M NaCl, 0.1% Tween20. Cleavage efficiency was determined by measuring the relative rate of cleavage of the 3'-AOM-nucleotide substrate. In short, a 0.1 mM solution of 3'-AOM nucleotide substrate in 100 mM buffer was to which stock palladium reagent was added to a final concentration of 1 mM Pd species. The solution was incubated at room temperature, a fixed amount was taken from the reaction at set points, quenched with a 1:1 EDTA / H2O2 solution (0.1:0.1 M), and the reaction rate was monitored by HPLC analysis for the formation of 3'-OH nucleotides and the disappearance of the 3'-AOM nucleotide substrate. As shown in Figure 1, the cleavage efficiency of Na2PdCl4 is comparable to the cleavage efficiency of [(allyl)PdCl]2 when using only 3 or 3.5 equivalents of THP, compared to 10 equivalents of THP for [(allyl)PdCl]2.
[0233] Example 3: Performance in AOM-AOL ffN stability testing and sequencing in solution Figure 2 shows the prephasizing performance of fully functionalized nucleotides (ffNs) (including labeled ffT-DB, labeled ffA, and labeled ffC, as well as unlabeled ffG) with a 3'-AOM blocking group and an AOL linker moiety, compared to standard MiniSeq® ffNs subjected to stress. Two sets of ffNs were incubated at 45°C for several days in a standard integration mixture formulation excluding DNA polymerase. At each time point, fresh polymerase was added to complete the integration mixture before loading into MiniSeq®. The sequencing conditions described above were used. Prephasizing % is a direct indicator of the proportion of 3'OH-ffNs present in the mixture and therefore directly correlates with the stability of the 3' block group. Prephasizing values for both sets of ffNs were recorded and plotted (Figure 2). Compared to the standard, AOM-AOL-ffNs did not show an increase in prephasizing and appeared to be substantially more stable than standard ffNs with a 3'-O-azidomethyl blocking group and an LN3 linker moiety.
[0234] Example 4. Use of palladium scavenger in sequencing reaction Figure 3 shows a comparison of phasing values on an Illumina MiniSeq® instrument using fully functionalized nucleotides (ffNs) (including labeled ffT-DB, labeled ffA, labeled ffC, and unlabeled ffG) with a 3'-AOM blocking group and a standard LN3 linker moiety, with and without the use of potassium isocyanoacetate in the post-cleavage washing step. Sequencing experiments were performed on an Illumina MiniSeq® using a cartridge in which a newly prepared solution of palladium cleavage reagent (10 mM [(allyl)PdCl]2, 100 mM THP, 100 mM ethanolamine buffer, 10 mM sodium ascorbate, 1 M NaCl, 0.1% Tween 20) was added to empty wells, after replacing the standard incorporation mixture with a newly prepared incorporation mixture containing ffNs with a 3'-AOM blocking group and a standard LN3 linker moiety. Potassium isocyanoacetate was added to a standard Miniseq® post-clear washing solution to a final concentration of 10 mM. Sequencing experiments were performed using a 2 × 151 cycle recipe, including a 5-second incubation with the palladium cleavage reagent solution, in addition to the standard synthetic sequencing (SBS) protocol. As shown in Figure 3, the use of 10 mM potassium isocyanoacetate in the post-clear washing solution reduced the phasing percentage from 0.183 to 0.075.
[0235] Figure 4 shows the primary sequencing metrics, including phasing, prephasing, and error rates, using fully functionalized nucleotides (ffNs) with 3'-AOM blocking groups and AOL linker moieties (including labeled ffT-DB, labeled ffA, and labeled ffC, as well as unlabeled ffG) on Illumina's MiniSeq® instrument, when using palladium scavenger, compared to the same sequencing metrics using standard ffNs with 3'-O-azidomethyl blocking groups and LN3 linker moieties. Sequencing experiments were performed on Illumina MiniSeq® by running a 2x151 cycle recipe using a standard cartridge, where the embedded mixture and standard cleavage reagents were replaced with newly prepared solutions of a palladium cleavage reagent (10 mM [(allyl)PdCl]2, 100 mM THP, 100 mM N,N'-diethylethanolamine buffer, 10 mM sodium ascorbate, 1 M NaCl, 0.1% Tween 20), respectively, containing fully functionalized nucleotides (ffN) with a 3'-AOM blocking group and an AOL linker moiety. Potassium isocyanoacetate (ICNA) was added to the standard MiniSeq® post-cleavage washing solution to a final concentration of 10 mM. Controlled sequencing experiments with standard ffNs using a 3'-O-azidomethyl blocking group were performed using a standard MiniSeq® kit and recipe. The results showed improvements in prephasing and, more importantly, error rates, demonstrating the complete efficiency of this AOM-AOL SBS chemistry through a single cutting process.
[0236] Example 5. Use of glycine in sequencing reactions Figure 5 shows a comparison of primary sequencing metrics, including phasing and prephasing, on Illumina's MiniSeq® instrument using fully functionalized nucleotides (ffNs) with a 3'-AOM blocking group and an AOL linker moiety, when glycine or ethanolamine is used in the incorporation mixture, respectively. In this example, the complete set of ffNs includes 3'-AOM-ffAs-AOL (labeled), 3'-AOM-ffG-AOL (unlabeled), 3'-AOM-ffT(DB)-AOL (labeled), and 3'-AOM-ffC(DB)-AOL (labeled). As shown in Figure 5, a significant decrease in phasing values was observed when glycine was used as the incorporation buffer compared to the standard incorporation buffer containing ethanolamine. When glycine was used, there was a slight increase in prephasing values, but it was not considered a significant increase. Sequencing experiments were performed using a standard MiniSeq® instrument with cartridges. Standard integration mixtures and standard cleavage reagents were replaced with newly prepared integration mixtures containing fully functionalized nucleotides (ffN) with 3'-AOM blocking groups and AOL linker moieties in either 50 mM ethanolamine or 50 mM glycine buffer, respectively, and with newly prepared solutions of palladium cleavage reagent (10 mM [(allyl)PdCl]2, 100 mM THP, 100 mM N,N'-diethylethanolamine buffer, 10 mM sodium ascorbate, 1 M NaCl, 0.1% Tween20). Potassium isocyanoacetate (ICNA) was added to the standard MiniSeq post-cleavage washing solution to a final concentration of 10 mM. A standard 2x151 cycle MiniSeq recipe was used.
[0237] Figure 6 shows primary sequencing metrics, including phasing, prephasing, and error rates, on Illumina's MiniSeq® instrument using fully functionalized nucleotides (ffNs) with 3'-AOM and AOL linker moieties, compared to the same sequencing metrics using standard ffNs and 3'-O-azidomethyl blocking groups and LN3 linker moieties. In this example, the complete set of ffNs includes 3'-AOM-ffA-AOL (labeled), 3'-AOM-ffG-AOL (unlabeled), 3'-AOM-ffT(DB)-AOL (labeled), and 3'-AOM-ffC(DB)-AOL (labeled). Sequencing experiments were performed using a standard MiniSeq® instrument with a cartridge-loaded setup and a 2×151 cycle recipe. Standard integration mixtures and standard cleavage reagents were replaced with newly prepared solutions of a palladium cleavage reagent (10 mM [(allyl)PdCl]2, 100 mM THP, 100 mM N,N'-diethylethanolamine buffer, 10 mM sodium ascorbate, 1 M NaCl, 0.1% Tween20) containing fully functionalized nucleotides (ffN) with 3'-AOM blocking groups in 50 mM glycine buffer, respectively. Potassium isocyanoacetate was added to the standard MiniSeq® post-cleavage washing solution to a final concentration of 10 mM. The results showed further improvement in error rate compared to those achieved with standard ffN. The improvement in fading of the AOM-AOL series compared to the metrics in Figure 5 is thought to be due to the use of glycine buffer and 3'-AOM-ffC(DB)-AOL.
[0238] Example 6. Synthetic sequencing of iSeq(trademark) Figures 7A and 7B show a comparison of primary sequencing metrics, including error rate and Q30 score, of sequencing performed on an Illumina iSeq™ instrument for 2 × 300 cycles of synthesis using fully functionalized nucleotides (ffNs) and 3'-AOM blocking groups and AOL linker moieties. In this example, the complete set of AOM ffNs includes 3'-AOM-ffA(DB)-AO-Dye 1, 3'-AOM-ffG (unlabeled), 3'-AOM-ffT(DB)-AOL-NR550S0, and 3'-AOM-ffC(DB)-AOL-Dye 2. The complete set of azidomethyl (AZM) ffNs includes the same ffNs having 3'-azidomethyl blocking groups and includes propargylamide and LN3 linkers. Dye 1 is disclosed in U.S. 63 / 127061 and, when conjugated with ffA, has a structural moiety. [ka] Dye 2 is a coumarin dye disclosed in U.S. Patent No. 2018 / 0094140, and when conjugated with ffC, the structural part [ka] It has [this characteristic]. NR550S0 is a known green dye.
[0239] Sequencing experiments were performed using a standard iSeq® instrument with cartridges. Standard integration mixtures and standard cleavage reagents were replaced, respectively, with a newly prepared integration mixture containing fully functionalized nucleotides (ffN) having a 3'-AOM blocking group and an AOL linker moiety in either 50 mM ethanolamine or 50 mM glycine buffer, using 300% polymerase 1901 (Pol1901) (360 ug / mL), and a newly prepared solution of palladium cleavage reagent (10 mM [(allyl)PdCl]2, 100 mM THP, 100 mM N,N'-diethylethanolamine buffer, 10 mM sodium ascorbate, 1 M NaCl, 0.1% Tween 20). L-cysteine was added to the standard iSeq® cleavage post-wash solution to a final concentration of 10 mM. A 2×30¹ cycle iSeq® recipe was used with a two-excitation / one-emission protocol. Specifically, the iSeq® instrument acquired the first image with green excitation light (approximately 520 nm) and the second image with blue excitation light (approximately 450 nm). A 2×30¹ cycle SBS cycle (incorporation, followed by imaging, followed by cleavage) was performed using a standard sequencing recipe. Sequencing metrics are summarized in the table below. [Table 2]
[0240] The AOM ffN set was observed to deliver superior performance, providing excellent error rates and Q30 for both Read1 and Read2. The fading values using the AOM ffN set were comparable to those generated by the AZM ffN set. However, the AOM ffN set produced significantly lower pre-fading values.
[0241] Figures 8A and 8B show a comparison of primary sequencing metrics, including error rate and Q30 score, of sequencing performed on an Illumina iSeq™ instrument for 2 × 150 cycles of synthesis using fully functionalized nucleotides (ffNs) with a 3'-AOM blocking group and an AOL linker moiety. In this example, the complete set of AOM ffNs includes 3'-AOM-ffA(DB)-AO-dye 1, 3'-AOM-ffG (unlabeled), 3'-AOM-ffT(DB)-AOL-NR550S0, and 3'-AOM-ffC(DB)-AOL-dye 2. The complete set of AZM ffNs includes the same ffNs with a 3'-azidomethyl blocking group, and also includes propargylamide and an LN3 linker. Sequencing experiments were performed using a standard iSeq® instrument with cartridges, replacing the standard incorporation mixture and standard cleavage reagent with a newly prepared incorporation mixture containing fully functionalized nucleotides (ffN) having a 3'-AOM blocking group and an AOL linker moiety in either 50 mM ethanolamine or 50 mM glycine buffer, using 300% concentration Pol1901 (360 ug / mL), and a newly prepared solution of palladium cleavage reagent (10 mM [(allyl)PdCl]2, 100 mM THP, 100 mM N,N'-diethylethanolamine buffer, 10 mM sodium ascorbate, 1 M NaCl, 0.1% Tween 20). L-cysteine was added to the standard iSeq® cleavage post-wash solution to a final concentration of 10 mM. A 2 × 301 cycle iSeq® recipe was used with a two excitation / one release protocol. Specifically, the iSeq® instrument acquired the first image with green excitation light (approximately 520 nm) and the second image with blue excitation light (approximately 450 nm). Two × 30¹ SBS cycles (embedding, followed by imaging, followed by cleavage) were performed using a standard sequencing recipe.
[0242] The integration mixture contact time for AZM ffN was approximately 24.1 seconds, while the integration mixture contact time for AOM ffN was approximately 29.1 seconds. However, the longer integration time for AOM ffN was compensated for by a faster deblocking time. The cleavage mixture contact time was approximately 5.8 seconds, in contrast to approximately 10.2 seconds for the AZM ffN set. Therefore, the total incubation times for the AZM and AOM ffN sets were approximately 34.4 and 34.9 seconds, respectively. The sequencing metrics are summarized in the table below. [Table 3]
[0243] Furthermore, the AOM ffN set was observed to deliver superior performance, providing excellent error rates and Q30 for both Read1 and Read2. Additionally, the AOM ffN set generated lower prephasing values.
[0244] Example 7. Synthesis-based sequencing of NovaSeq™ with gradually increasing blue laser power. Prolonged exposure to blue light in SBS sequencing has been observed to cause high levels of signal delay and fading as a result of increased light dose and power density. This example compares the performance of AOM ffN and AZM ffN in a blue laser gradual sequencing experiment.
[0245] In this experiment, 1 × 15¹ runs of SBS were compared to a standard ffN set with a 3'-azidomethyl blocking group and LN3 linker, using an AOM ffN set configured with a modified blue / green excitation NovaSeq™. The flow cell used was a 490 nm pitch BEER2 flow cell. The power configurations for the blue laser output were 600 mW, 800 mW, 1000 mW, 1400 mW, 1800 mW, and 2400 mW. The green laser power was kept constant at 1000 mM. The following standard AZM ffNs were used: green ffT (LN3-AF550POPOS0), dark G, red ffC (LN3-SO7181), blue ffC (sPA-blue dye A), blue ffA (LN3-BL-blue dye A), and green ffA (LN3-BL-NR550S0). The following ffNs were used in the AOM ffN set: green ffT (ffT(DB)-AOL-AF550POPOS0), dark G, red ffC (ffC(DB)-AOL-SO7181), blue ffC (ffC(DB)-AOL-blue pigment A), blue ffA (ffA(DB)-AOL-BL-blue pigment A), and green ffA (ffA(DB)-AOL-BL-NR550S0). The structures of the blue pigments labeled AOMffC and ffA are shown below. [ka]
[0246] The following modifications were made to the AOM ffN set SBS run. First, 10 mM L-cysteine was added to the post-cleaning wash solution. The cleaning mixture contained the following components: Na2PdCl4 in DEEA buffer. Two 10-second waiting steps were added to the post-cleaning wash step. In addition, a static incorporation waiting time of 60 seconds was used (in contrast to 38 seconds for AZM ffN). Figure 9A shows that both ffN sets had similar fading values at lower blue laser power, but the AOM ffN set was less sensitive to increasing blue laser power (indicated by a gentler fading gradient compared to that of AZM). Furthermore, it was also observed that the AOM ffN had much lower pre-fading. As shown in Figure 9B, the AOM ffN set showed significantly less signal attenuation at higher blue laser power. Figures 9C and 9D show the mean error rate as a function of the number of cycles. Figure 9D is an enlarged view of Figure 9C. The results show that at lower blue laser powers, the AZM ffN had a lower error rate in the initial cycles, but at higher blue laser powers, the AOM ffN set performed much better in later cycles. Figure 9E summarizes the average error rate over 151 cycles. Here again, the AOM ffN set outperformed the AZM ffN set set at higher laser powers (e.g., 1400mW, 1800mW, and 2400mW).
[0247] Example 8. First chemical linearization using a Pd cleavage mixture In this example, the Pd cleavage mixture used for SBS was tested in a first chemical linearization step after the clustering step. The experiment compared chemical linearization with standard enzymatic linearization in which the cleavage of one of the double-stranded polynucleotides is facilitated by USER to cleave the U position on the P5 primer. 1 × 150 cycles of SBS were performed using an Illumina iSeq™ instrument with fully functionalized nucleotides (ffN) having a 3'-AOM blocking group and an AOL linker moiety. In this example, as described in Example 6, the complete set of AOM ffN includes 3'-AOM-ffA(DB)-AO-Dye 1, 3'-AOM-ffG (unlabeled), 3'-AOM-ffT(DB)-AOL-NR550S0, and 3'-AOM-ffC(DB)-AOL-Dye 2. The iSeq® instrument was configured to capture a first image using green excitation light (approximately 520 nm) and a second image using blue excitation light (approximately 450 nm) (using a two-excitation / one-emission protocol). The flow cell used with the iSeq® instrument was grafted with modified P5 / P7 primers to enable the initial chemical linearization of the P5 primer. The chemical linearization step was performed in a Pd cleavage mixture (10 mM [Pd(allyl)Cl]2 and 100 mM THP in a buffer solution containing DEEA) incubated at 63°C for 30 seconds. Figure 10 shows the SBS sequencing metrics using the two different linearization methods. When the Pd cleavage mixture was used in the first chemical linearization step, it was observed that all primary sequencing metrics fell within the standard observation range. This experiment confirms that a single reagent mixture can be used in two separate sequencing steps, namely the linearization step and the SBS cleavage step, allowing for further simplification of the instrument (fluids and cartridges).
Claims
1. A nucleotide or nucleoside comprising a nucleic acid base bound to a detectable label via a cleavable linker, wherein the nucleoside or nucleotide comprises a ribose or 2'-deoxyribose moiety and a 3'-OH blocking group, and the cleavable linker comprises a portion of the following structure: 【Chemistry 1】 During the ceremony, Each of X and Y is independently either O or S. R 1a , R 1b , R 2 , R 3a and R 3b each of which is independently H, halogen, unsubstituted or substituted C 1 ~C 6 alkyl, or C 1 ~C 6 haloalkyl, a nucleotide or a nucleoside.
2. A nucleoside or nucleotide according to claim 1, comprising the structure of formula (I), 【Chemistry 2】 In the formula, B is the nucleic acid base, R 4 However, it is either H or OH, R 5 However, the 3'-OH blocking group is R 6 However, it is H, a monophosphate, a diphosphate, a tripphosphate, a thiophosphate, a phosphate ester analog, a reactive phosphorus-containing group, or a hydroxyprotecting group. The detectable label is a fluorescent dye. L, 【Transformation 3】 And, L 1 and L 2 The nucleoside or nucleotide according to claim 1, wherein each of them is an independently and optionally present linker portion.
3. A nucleoside or nucleotide according to claim 1 or 2, wherein each of X and Y is O.
4. R 1a , R 1b , R 2 , R 3a and R 3b A nucleoside or nucleotide according to any one of claims 1 to 3, wherein each of them is H.
5. R 1a , R 1b , R 2 , R 3a and R 3b At least one of them is a halogen or unsubstituted C 1 ~C 6 A nucleoside or nucleotide according to any one of claims 1 to 3, wherein the nucleoside is alkyl.
6. R 1a and R 1b Each of them is H, and R 2 , R 3a and R 3b One of them is unsubstituted C 1 ~C 6 The nucleoside or nucleotide according to claim 5, wherein it is alkyl.
7. A nucleoside or nucleotide according to any one of claims 2 to 6, wherein B is a purine, deazapurine, or pyrimidine.
8. R 5 but, 【Chemistry 4】 And R a , R b , R c , R d and R e Each of these independently consists of H, halogen, unsubstituted or substituted C. 1 ~C 6 Alkyl, or C 1 ~C 6 A nucleoside or nucleotide according to any one of claims 2 to 7, wherein the nucleoside is a haloalkyl.
9. R 5 but, 【Transformation 5】 The nucleoside or nucleotide according to claim 8.
10. L 1 There exists, L 1 The nucleoside or nucleotide according to any one of claims 2 to 9, wherein the portion comprises a portion selected from the group consisting of propargylamine, propargylamide, allylamine, allylamide, and optionally substituted variants thereof.
11. L 1 but, 【Transformation 6】 A nucleoside or nucleotide according to claim 10, comprising:
12. A nucleoside or nucleotide according to claim 11, comprising the structure of formula (Ia), (Ia'), (Ib), (Ic), (Ic'), or (Id). 【Transformation 7】 【Transformation 8】
13. L 2 There exists, L 2 but, 【Chemistry 9】 A nucleoside or nucleotide according to any one of claims 2 to 12, comprising, where each of m and n is independently an integer of 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10, and the phenyl portion is optionally substituted.
14. The nucleoside or nucleotide according to claim 13, wherein n is 5.
15. The nucleoside or nucleotide according to claim 13 or 14, wherein m is 4.
16. The nucleoside or nucleotide according to any one of claims 1 to 14, wherein the nucleotide is a nucleotide triphot containing a 2'-deoxyribose moiety.
17. An oligonucleotide or polynucleotide comprising the nucleotide according to any one of claims 1 to 16.
18. The oligonucleotide or polynucleotide according to claim 17, wherein the oligonucleotide or polynucleotide hybridizes to a template polynucleotide.
19. The oligonucleotide or polynucleotide according to claim 18, wherein the template polynucleotide is immobilized on a solid support.
20. The oligonucleotide or polynucleotide according to claim 19, wherein the solid support comprises an array of a plurality of immobilized template polynucleotides.
21. A method for preparing a growing polynucleotide complementary to a target single-stranded polynucleotide in a sequencing reaction, comprising incorporating a nucleotide according to any one of claims 1 to 16 into the growing complementary polynucleotide, wherein the incorporation of the nucleotide prevents the introduction of any subsequent nucleotides into the growing complementary polynucleotide.
22. The method according to claim 21, wherein the incorporation of the nucleotide is achieved by polymerase, terminal deoxynucleotidyltransferase, or reverse transcriptase.
23. A method for determining the sequence of a target single-stranded polynucleotide, (a) Incorporating the nucleotide according to any one of claims 1 to 16 into a copy polynucleotide chain complementary to at least a portion of the target polynucleotide chain, (b) To detect the identity of the nucleotides incorporated into the copy polynucleotide chain, (c) Chemically removing the detectable label and the 3'-OH blocking group from the nucleotide incorporated into the copy polynucleotide chain, Methods that include...
24. (d) The method of claim 23, further comprising washing away the chemically removed label and the 3'-OH blocking group from the copy polynucleotide chain using a post-cleavage washing solution.
25. The method according to claim 24, further comprising repeating steps (a) to (d) until the sequence of at least a portion of the target polynucleotide chain is determined.
26. The method according to claim 25, wherein steps (a) to (d) are repeated at least 50 times, at least 100 times, or at least 150 times.
27. The method according to any one of claims 23 to 26, wherein step (c) comprises contacting the incorporated nucleotide with a cleavage solution containing a palladium catalyst.
28. The method according to claim 27, wherein the palladium catalyst is a palladium (0) catalyst in situ generated from a palladium complex and a water-soluble phosphine.
29. The palladium complex is [Pd(allyl)Cl] 2 Na 2 PdCl 4 [Pd(allyl)(THP)]Cl, [Pd(allyl)(THP)] 2 ]Cl, Pd(CH 3 CN) 2 Cl 2 , Pd(OAc) 2 , Pd(PPh 3 ) 4 , Pd(dba) 2 , Pd(Acac) 2 , PdCl 2 (COD), or Pd(TFA) 2 The method according to claim 28, including, or a combination thereof.
30. The palladium complex is [Pd(allyl)Cl] 2 or Na 2 PdCl 4 The method according to claim 29, including the method described in claim 29.
31. The method according to any one of claims 27 to 30, wherein the detectable label and the 3'-OH blocking group from the nucleotide incorporated into the copy polynucleotide chain are removed by a single chemical reaction.
32. The method according to any one of claims 27 to 31, wherein the cleavage solution further comprises one or more buffering reagents selected from the group consisting of primary amines, secondary amines, tertiary amines, carbonates, phosphates, and borates, and combinations thereof.
33. The method according to claim 32, wherein the buffering reagent in the cutting solution is selected from the group consisting of ethanolamine (EA), tris(hydroxymethyl)aminomethane (Tris), glycine, carbonate, phosphate, borate, dimethylethanolamine (DMAA), diethylethanolamine (DEEA), N,N,N',N'-tetramethylethylenediamine (TEMED) and N,N,N',N'-tetraethylethylenediamine (TEEDA), 2-piperidineethanol, and combinations thereof.
34. The method according to any one of claims 27 to 33, wherein step (a) comprises contacting the nucleotide with the copy polynucleotide chain in an integration solution comprising a polymerase, at least one palladium scavenger, and one or more buffers.
35. The method according to claim 34, wherein the buffer in the incorporated solution comprises a primary amine, a secondary amine, a tertiary amine, a natural amino acid, or a non-natural amino acid, or a combination thereof.
36. The method according to claim 35, wherein the buffer in the incorporated solution comprises ethanolamine, glycine, or a combination thereof.
37. The palladium scavenger in the incorporated solution is -O-allyl, -S-allyl, -NR-allyl, and -N + It comprises one or more allelic moieties independently selected from the group consisting of RR'-allyles, In the formula, R is H, unsubstituted or substituted C. 1 ~C 6 Alkyl, unsubstituted, or substituted C 2 ~C 6 Alkenyl, unsubstituted, or substituted C 2 ~C 6 Alkynyl, unsubstituted, or substituted C 6 ~C 10 Aryl, unsubstituted or substituted 5- to 10-membered heteroaryl, unsubstituted or substituted C 3 ~C 10 Carbocyclyl, or unsubstituted or substituted 5-10 member heterocyclyl, R' is H, unsubstituted C 1 ~C 6 Alkyl or substituted C 1 ~C 6 The method according to any one of claims 34 to 36, wherein the alkyl group is alkyl.
38. The palladium scavenger in the aforementioned incorporated solution 【Chemistry 10】 The method according to claim 37.
39. The method according to any one of claims 27 to 38, wherein the post-cutting cleaning solution comprises one or more palladium scavengers.
40. The method according to claim 39, wherein the one or more palladium scavengers in the post-cleavage solution include isocyanoacetate (ICNA) salt, isocyanoethyl acetate, isocyanoacetate methyl, cysteine or a salt thereof, L-cysteine or a salt thereof, N-acetyl-L-cysteine, potassium ethylxanthogenic acid, potassium isopropylxanthogenic acid, glutathione, lipoic acid, ethylenediaminetetraacetic acid (EDTA), iminodiacetic acid, nitrilodiacetic acid, trimercapto-S-triazine, dimethyldithiocarbamate, dithiothreitol, mercaptoethanol, allyl alcohol, propargyl alcohol, thiol, tertiary amine and / or tertiary phosphine, or a combination thereof.
41. The method according to any one of claims 23 to 40, wherein the target single-stranded polynucleotide is formed by chemically cleaving a complementary strand from a double-stranded polynucleotide.
42. The method according to claim 41, wherein the chemical cleavage of the complementary chain is carried out under the same reaction conditions as chemically removing the detectable label and the 3'-OH blocking group from the nucleotide incorporated into the copy polynucleotide chain.
43. A kit comprising one or more nucleosides or nucleotides as described in any one of claims 1 to 16.
44. The kit according to claim 43, further comprising an enzyme, at least one Pd(0) scavenger, and one or more buffers.
45. The kit according to claim 44, wherein the enzyme is DNA polymerase, terminal deoxynucleotidyltransferase, or reverse transcriptase.
46. The Pd(0) scavengers are, independently of each other, -O-allyl, -S-allyl, -NR-allyl, and -N + It comprises one or more allelic moieties selected from the group consisting of RR'-allyles, In the formula, R is H, unsubstituted or substituted C 1 ~C 6 alkyl, unsubstituted or substituted C 2 ~C 6 alkenyl, unsubstituted or substituted C 2 ~C 6 alkynyl, unsubstituted or substituted C 6 ~C 10 aryl, unsubstituted or substituted 5- to 10-membered heteroaryl, unsubstituted or substituted C 3 ~C 10 carbocyclic, or unsubstituted or substituted 5- to 10-membered heterocyclic, and R' is H, unsubstituted C 1 ~C 6 Alkyl or substituted C 1 ~C 6 The kit according to claim 44, wherein the alkyl group is present.
47. The palladium scavenger mentioned above 【Chemistry 11】 The kit according to claim 46.
48. The kit according to any one of claims 43 to 47, further comprising a palladium catalyst.
49. The kit according to claim 48, wherein the palladium catalyst is a Pd(0) catalyst in situ generated from a Pd(II) complex and one or more water-soluble phosphines.
50. The Pd(II) complex is [Pd(allyl)Cl] 2 or Na 2 PdCl 4 The kit according to claim 49, wherein the kit is as described above.
51. The kit according to claim 49 or 50, further comprising one or more Pd(II) scavengers, wherein the Pd(II) scavengers include isocyanoacetate (ICNA) salt, isocyanoethyl acetate, isocyanomethyl acetate, cysteine or a salt thereof, L-cysteine or a salt thereof, N-acetyl-L-cysteine, potassium ethylxanthogenic acid, potassium isopropylxanthogenic acid, glutathione, lipoic acid, ethylenediaminetetraacetic acid (EDTA), iminodiacetic acid, nitrilodiacetic acid, trimercapto-S-triazine, dimethyldithiocarbamate, dithiothreitol, mercaptoethanol, allyl alcohol, propargyl alcohol, thiol, tertiary amine and / or tertiary phosphine, or a combination thereof.
52. The kit according to claim 51, wherein the Pd(II) scavenger is L-cysteine or a salt thereof.