Modified Nucleotide Linkers

Fluorophores attached via linkers improve nucleotide incorporation rates in sequencing by synthesis, addressing inefficiencies in existing sequencing methods and enhancing sequencing accuracy.

JP7726707B2Active Publication Date: 2025-08-20ILLUMINA CAMBRIDGE LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2021148730
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2014-08-08
Filing Date
2021-09-13
Publication Date
2025-08-20
Estimated Expiration
2035-08-06

AI Technical Summary

Technical Problem

Existing nucleotide sequencing methods face challenges in increasing the rate of nucleotide incorporation during sequencing by synthesis, leading to inefficiencies in determining nucleotide sequences.

Method used

The use of fluorophores covalently attached via linkers, such as those represented by Formulas I, II, or III, to enhance nucleotide incorporation efficiency in sequencing processes.

Benefits of technology

Enhances the rate of nucleotide incorporation, improving the accuracy and efficiency of nucleotide sequencing by synthesis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007726707000029
    Figure 0007726707000029
  • Figure 0007726707000030
    Figure 0007726707000030
  • Figure 0007726707000031
    Figure 0007726707000031
Patent Text Reader

Abstract

To provide novel modified nucleotide linkers for increasing the efficiency of nucleotide incorporation in Sequencing by Synthesis applications and methods of preparing these modified nucleotide linkers.SOLUTION: The present invention discloses a nucleoside or nucleotide covalently attached to a fluorophore through a linker. The linker comprises a structure of formula (I) or (II), or combination of both.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Some embodiments of the present application relate to increasing nucleotide incorporation in DNA sequencing. and other diagnostic applications, e.g., sequencing by synthesis. It relates to nucleoside or nucleotide linkers. [Background technology]

[0002] Advances in molecular research are due, in part, to the availability of information used to characterize molecules or their biological responses. In particular, the study of nucleic acids, DNA and RNA, has been driven by advances in sequencing and The present invention has benefited from the development of techniques used to study the interaction between nucleotides and hybridization events.

[0003] Examples of technologies that have improved the study of nucleic acids include fabricated arrays of immobilized nucleic acids. The array typically consists of polymers immobilized on a solid support material. It has a high density matrix of nucleotides. See, e.g., Fodor et al., Trends Biotech. 12: 19-26, 1994 (which includes the use of masks to protect against exposure in designated areas and the use of appropriate Chemically sensitized glass capable of binding appropriately modified nucleotide phosphoramidites. (See, e.g., [Crossref], [PubMed], [Web of Science ® ... Alternatively, known polynucleotides may be spotted at predetermined locations on a solid support ("spotting"). ") techniques (see, for example, Stimpson et al., Proc. Nat. l. Acad. Sci. 92: 6379-6383, 1995).

[0004] One method for determining the nucleotide sequence of nucleic acids bound to an array is "sequencing by This technique for determining the nucleotide sequence of DNA is called "Sequence Sequence Synthesis" or "Sequence Sequence Sequence Synthesis." Ideally, there is a controlled sequence of exactly complementary nucleotides on opposite sides of the nucleic acid to be sequenced. This requires incorporation of each nucleotide residue in the sequence one by one. Nucleotides are added in multiple cycles as determined, resulting in an uncontrolled series of nucleotides. Preventing incorporation of nucleotides allows accurate sequencing. Each of the peptides has an appropriate label attached thereto, allowing for removal of the label and subsequent selection of the next peptide. -Read out before Kenshin Round.

[0005] Therefore, in the context of nucleic acid sequencing reactions, there are many applications of nucleotide sequences that can increase the efficiency of sequencing methods. It is desirable to be able to increase the rate of nucleotide incorporation during sequencing by synthesis. It would be nice. Summary of the Invention

[0006] Some embodiments disclosed herein include a fluorophore covalently attached via a linker. wherein the linker is a nucleoside or nucleotide of Formula I or Formula II, or refers to a nucleoside or nucleotide, including structures that combine both.

[0007] [ka] (In the formula, R 1 and R 2 each independently represents hydrogen or an optionally substituted C 1-6 Alkyl Selected from; R 3is hydrogen, optionally substituted C 1-6 Alkyl, -NR 5 -C(=O)R 6 , or -NR 7 -C(=O)-O R 8 Selected from; R 4 is hydrogen or optionally substituted C 1-6 alkyl; R 5 and R 7 each independently represents hydrogen, optionally substituted C 1-6 Alkyl, substituted optionally substituted phenyl or optionally substituted C 7-12 aralkyl; R 6 and R 8 each independently represents an optionally substituted C 1-6 Alkyl, substituted optionally substituted phenyl, optionally substituted C 7-12 Aralkyl, optionally substituted C 3-7 S cycloalkyl, or optionally substituted 5-10 membered heteroaryl; [ka] wherein each methylene repeat unit is optionally substituted; X is selected from methylene (CH), oxygen (O), or sulfur (S); m is an integer from 0 to 20; n is an integer from 1 to 20; p is an integer from 1 to 20.

[0008] In some embodiments, a fluorophore-labeled nucleoside comprising the structure of Formula I: Or the nucleotide does not include the following structure:

[0009] [ka]

[0010] Some embodiments disclosed herein include a fluorophore covalently attached via a linker. a nucleoside or nucleotide comprising the linker, wherein the linker comprises the structure of formula III; Relating to nucleotides or nucleotides.

[0011] [ka] (In the formula, L 1 is empty or is any one of the linkers, protecting groups, parts, or combinations thereof; L 2 is optionally substituted C 1-20 Alkylene, optionally substituted C 1-20 Heteroarchitecture C, optionally substituted, interrupted by a substituted aromatic group 1-20 Alkylene, or substituted Optionally, a substituted aromatic group interrupted by C 1-20 heteroalkylene; L 3 is optionally substituted C 1-20 Alkylene, or optionally substituted C 1-20 Hetero alkylene; R A is hydrogen, cyano, hydroxy, halogen, C 1-6 Alkyl, C 1-6 Alkoxy, C 1-6 Halo Alkyl, C 1-6 haloalkoxy, or azido; [ka] wherein at least one of the repeat units contains an azide group; Z is oxygen (O) or NR B Selected from; R B and R Ceach independently represents hydrogen or an optionally substituted C 1-6 Select from alkyl Selected; k is an integer from 1 to 50.

[0012] Some embodiments disclosed herein include labeled nucleosides or labeled nucleotides. The kit contains a linker between the fluorophore and the nucleoside or nucleotide. wherein the linker is any one of Formula I, Formula II, or Formula III, or a combination thereof. The present invention relates to a kit including a combination structure.

[0013] Some embodiments disclosed herein include nucleosides comprising a fluorophore and a linker. A reagent for modifying a side or nucleotide, wherein the linker is represented by Formula I, Formula II, or relates to a reagent comprising the structure of any one of Formula III, or a combination thereof.

[0014] Some embodiments disclosed herein include nucleosides incorporated into polynucleotides. 1. A method for detecting (a) Linking a labeled nucleoside or labeled nucleotide containing a linker to a polynucleotide and a step of capturing; (b) the labeled nucleoside or labeled nucleotide incorporated in step (a) and detecting a fluorescent signal from the The linker may be any one of Formula I, Formula II, or Formula III, or a combination thereof. In some embodiments, the method further comprises the steps of: providing a partially hybridized nucleic acid strand and a partially hybridized nucleic acid strand, is at least one nucleoside or nucleotide complementary to the corresponding position of the template strand. Incorporation of a single nucleoside or nucleotide into the hybridized strand, step (b) by identifying the base of the incorporated nucleoside or nucleotide; The identity of the complementary nucleoside or nucleotide of the template strand is indicated.

[0015] Some embodiments disclosed herein include a method for sequencing a template nucleic acid molecule, comprising: A step of incorporating one or more labeled nucleotides into a nucleic acid strand complementary to a template nucleic acid. Top and; To determine the sequence of the template nucleic acid molecule, one or more incorporated labeled nucleosides are determining the identities of the bases present in the sequence; The identity of the bases present in the one or more labeled nucleotides is determined by the and detecting a fluorescent signal produced by the nucleotide sequence, at least one incorporated labeled nucleotide comprises a linker as described above; The linker may be any one of Formula I, Formula II, or Formula III, or a combination thereof. In some embodiments, the method includes a structure of one or more nucleotides. The identity of the bases present is determined after each nucleotide incorporation step. [Brief explanation of the drawings]

[0016] [Figure 1] Figure 1A shows the partial linker structure of a standard labeled nucleotide. Figure 1B shows the labeled nucleotide of Figure 1A, with two possible linkers 125 and 130 inserted into the standard linker of Figure 1A. [Figure 2] 1B is a graph of nucleic acid incorporation rates using the labeled nucleotide of FIG. 1A and the modified labeled nucleotide of FIG. 1B. [Figure 3]Figure 3A shows a structural formula of an additional linker inserted into the standard linker of Figure 1A. Figure 3B shows a structural formula of an additional linker inserted into the standard linker of Figure 1A. Figure 3C shows a structural formula of an additional linker inserted into the standard linker of Figure 1A. Figure 3D shows a structural formula of an additional linker inserted into the standard linker of Figure 1A. Figure 3E shows a structural formula of an additional linker inserted into the standard linker of Figure 1A. [Figure 4] FIG. 3B is a data table of two dye sequencing runs used to assess the effect of the 125 insertion in FIG. 1B and the 315 insertion in FIG. 3B on sequencing quality. [Figure 5] 5A is a graph of the error rate for read 1 of the sequencing run of FIG. 4 using linker insertion 125 and linker insertion 315. FIG. 5B is a graph of the error rate for read 2 of the sequencing run of FIG. 4 using linker insertion 125 and linker insertion 315. [Figure 6] 1B and 3A are data tables of sequencing runs used to assess the effect of the 125 insertion in FIG. 1B and the 310 insertion in FIG. 3A on sequencing quality. [Figure 7] 7A is a graph of the error rate for read 1 of the sequencing run of FIG. 6 using linker insertion 125. FIG. 7B is a graph of the error rate for read 2 of the sequencing run of FIG. 6 using linker insertion 125. [Figure 8] Figure 8A shows an example of a standard LN3 linker structure. Figure 8B shows an example of a modified version of the LN3 linker structure of Figure 8A. Figure 8C shows an example of a modified version of the LN3 linker structure of Figure 8A. Figure 8D shows an example of a modified version of the LN3 linker structure of Figure 8A. Figure 8E shows the insertion of a protecting moiety into the linker of Figure 8D. [Figure 9]Figure 9A is a chromatogram showing the appearance of impurities in ffA with an SS linker. Figure 9B is a table comparing the stability of ffA with an SS linker and an AEDI linker. Figure 9C is a chromatogram showing a comparison of ffA with an SS linker and an AEDI linker at 22 hours IMX60°, where the SS linker again shows impurities. [Figure 10A] FIG. 1 shows an unexpected increase in the rate of nucleotide incorporation in solutions with varying linkers, and is a graph showing the incorporation rate at 1 μM. [Figure 10B] FIG. 1 shows the unexpected increase in the rate of nucleotide incorporation in solutions with varying linkers, with tabulated results. [Figure 10C] FIG. 1 shows an unexpected increase in the rate of nucleotide incorporation in solutions with varying linkers, and is a schematic representation of the AEDI and SS linkers with NR550S0. [Figure 11A] FIG. 1 shows scatter plots for V10 combinations with different A-550S0 (same concentrations). [Figure 11B] FIG. 1 shows the Kcat of FFA linkers in solution. [Figure 12A] FIG. 1 shows sequencing metrics for M111, human 550, 2×151 cycles. [Figure 12B] FIG. 1 shows sequencing metrics for M111, human 550, 2×151 cycles. DETAILED DESCRIPTION OF THE INVENTION

[0017] Some embodiments disclosed herein include a fluorophore covalently attached via a linker. The linker is a nucleoside or nucleotide represented by the following formula I or formula II: or a combination of both, the definitions of the variables of which are defined above. Concerning nucleotides or nucleotides.

[0018] [ka]

[0019] In some embodiments of the structure of Formula I, R 1 is hydrogen. In some other embodiments, R 1 is optionally substituted C 1-6 In some such embodiments, R 1 Hamechi It is.

[0020] R of Formula I described herein 1 In any embodiment of the present invention, R 2 is hydrogen. In the embodiment, R 2 is optionally substituted C 1-6 In some such embodiments, the alkyl is alkyl. So, R 2 is methyl. In one embodiment, R 1 and R 2 Both are methyl. In terms of form, R 1 and R 2 are both hydrogen.

[0021] In some embodiments of the structure of Formula I, m is 0. In some other embodiments, m is 1. be.

[0022] In some embodiments of the structure of formula I, n is 1.

[0023] In some embodiments of the structure of Formula I, the structure of Formula I can also be represented by Formula Ia or Formula Ib It is possible.

[0024] [ka]

[0025] In some embodiments described herein, Formula Ia is referred to as "AEDI" and Formula Ib is referred to as "SS "

[0026] In some embodiments of the structure of Formula II, R 3 is hydrogen. In some other embodiments, Te, R 3 is optionally substituted C 1-6 In some such embodiments, , R 3 is methyl. In some embodiments, R 3 Ha-NR 5 -C(=O)R 6 Some of these In such an embodiment, R 5 is hydrogen. In some such embodiments, R 6 is a substitution C 1-6 alkyl, e.g., methyl. In some embodiments, R 3 teeth -NR 7 -C(=O)OR 8 In some such embodiments, R 7 is hydrogen. In such embodiments, R 8 is optionally substituted C 1-6 Alkyl, e.g., t-butyl is.

[0027] R of Formula II described herein 3 In any embodiment of the present invention, R 4 is hydrogen. In embodiments, R 4 is optionally substituted C 1-6 Some of these compounds are alkyl. In the embodiment, R 4 is methyl. In one embodiment, R 3 and R 4Both are methyl In another embodiment, R 3 and R 4 and R are hydrogen. 3 is -NH(C=O)CH3, and R 4 is hydrogen. In another embodiment, R 3 -NH(C=O)O t Bu (Boc ) and R 4 is hydrogen.

[0028] In some embodiments of the structure of Formula II, X is methylene, which is substituted In another embodiment, X is oxygen (O). In yet another embodiment, X is sulfur (S )

[0029] In some embodiments of the structure of Formula II, p is 1. In some other embodiments, p is 2 is.

[0030] In some embodiments of the structure of Formula II, the structure of Formula II can also be represented by Formula IIa, Formula IIb, It can be represented by formula IIc, formula IId, formula IIe, or formula IIf.

[0031] [ka]

[0032] In some embodiments described herein, Formula IIa is referred to as "ACA" and Formula IIb is Formula IIc is referred to as "BocLys", Formula IId is referred to as "dMeO", and Formula IIe is referred to as "dMeS" and formula IIf is referred to as "DMP".

[0033] a fluorophore via a linker comprising a structure of Formula I or Formula II described herein; In any embodiment of the labeled nucleoside or nucleotide, the nucleoside or The nucleotide can be attached to the left of the linker either directly or via an additional linker. It is Noh.

[0034] Some embodiments disclosed herein include a fluorophore covalently attached via a linker. a nucleoside or nucleotide comprising a linker having the structure of Formula III, The number definitions refer to nucleosides or nucleotides as defined above.

[0035] [ka]

[0036] In some embodiments of the structure of Formula III, L 1 is empty. In some other embodiments L 1 is of formula I or formula II, in particular formula Ia, formula Ib, formula II, formula IIa, formula IIb, formula II c, Formula IId, Formula IIe, or Formula IIf. In the embodiment of the present invention, L 1 can be a protective moiety that includes a molecule that prevents DNA damage In some such embodiments, the protecting moiety is trolox, gallic acid, p-nitrobenzyl benzoate, or benzoyl benzoate. (pNB), or ascorbate, or a combination thereof.

[0037] In some embodiments of the structure of Formula III, L 2 is optionally substituted C 1-20 Alkire In some further embodiments, L 2 is optionally substituted C 4-10 Alkire In some such embodiments, L 2is heptylene. In terms of form, L 2 is optionally substituted C 1-20 Some of these are heteroalkylenes. In such embodiments, optionally substituted C 1-20 Heteroalkylene may be one or more In some such embodiments, C 1-20 A small number of heteroalkylenes At least one carbon atom is substituted with oxo (=O). Leave it, L 2 is optionally substituted C 3-6 In some embodiments, L 2 is a substitution C 6-10 Substituted aromatic groups such as aryl groups or 5 groups containing 1 to 3 heteroatoms In some such embodiments, L is interrupted by a 10- to 10-membered substituted heteroaryl group. 2 Ha In some such embodiments, the phenyl group is interrupted by a substituted phenyl group. Toro, Cyano, Halo, Hydroxy, C 1-6 Alkyl, C 1-6 Alkoxy, C 1-6 haloalkyl, C 1-6 one or more (at most) selected from haloalkoxy, or sulfonyl hydroxide; In some further such embodiments, phenyl is substituted with at least four substituents. The yl group can be nitro, cyano, halo, or sulfonyl hydroxide (i.e., -S(=O)2OH). ) is substituted with 1 to 4 substituents selected from

[0038] In some embodiments of the structure of formula III, [ka] R Ais selected from hydrogen or azide. In some such embodiments, one R A When one is an azide and the other is hydrogen, k is 2.

[0039] In some embodiments of the structure of Formula III, L 3 is optionally substituted C 1-20 Alkire In some further embodiments, L 3 is optionally substituted C 1-6 Alkylene In some such embodiments, L 3 is ethylene. In L 3 is optionally substituted C 1-20 Heteroalkylene. In an embodiment, optionally substituted C 1-20 Heteroalkylene is one or more In some such embodiments, L 1 is optionally substituted C 1-6 a alkylene oxides, such as C 1-3 It is an alkylene oxide.

[0040] In some embodiments of the structure of formula III, R B is hydrogen. , R C is hydrogen. In some further embodiments, R B and R C are both hydrogen.

[0041] In some embodiments of the structure of Formula III, the structure of Formula III can also be represented by Formula IIIa, Formula I IIb, or can be represented by formula IIIc.

[0042] [ka] (In the formula, R D are nitro, cyano, halo, hydroxy, C 1-6 Alkyl, C 1-6 Alkoxy, C 1-6 Haloalkyl, C 1-6 haloalkoxy, or sulfonyl hydroxide. In a further embodiment of the group, R D is nitro, cyano, halo, or sulfonylhydroxy Selected from:

[0043] Labeled with a fluorophore via a linker comprising the structure of Formula III, as described herein. In any embodiment of the nucleoside or nucleotide, the fluorophore may be a linker. It can be connected to the left side of the hub directly or via an additional connection.

[0044] Any of the linkers described herein that include a structure of Formula I, Formula II, or Formula III In the embodiments of the present invention, the term "optionally substituted" is used to define a variable. In this case, such variables may not be substituted.

[0045] definition Unless otherwise defined, all technical and scientific terms used herein are commonly understood by those skilled in the art. The term "including" and other forms, e.g. Use of "include", "includes", and "included" The term "having" as well as other forms such as "having" The use of "have," "has," and "had" is not limiting. As used herein, the terms "comprise(s)" and "comprises" are used in both the claim transitional phrase and the body of the text. The term "comprising" shall be construed to have an open-ended meaning. That is, the above term should be interpreted as "having at least " or "including at least" should be interpreted synonymously. For example, when used in the context of a process, the term "comprising" The process includes at least the steps listed, but may also include additional steps. When used in the context of a compound, composition, or device, the term "comprises" "comprising" means that the compound, composition, or device exhibits at least the recited characteristics. or components, but may also include additional features or components.

[0046] The section headings used herein are for organizational purposes only and do not limit the subject matter described. It should not be construed as limiting.

[0047] As used herein, common organic abbreviations are defined as follows: Ac Acetyl Ac2O acetic anhydride aq. water-based Bn Benzyl Bz Benzoyl BOC or Boc tert-butoxycarbonyl Bu n-butyl cat. catalytic °C Celsius temperature CHAPS 3-[(3-cholamidopropyl)dimethylammonio]-1-propanesulfonate dATP deoxyadenosine triphosphate dCTP deoxycytidine triphosphate dGTP deoxyguanosine triphosphate dTTP deoxythymidine triphosphate ddNTP(s) Dideoxynucleotides DCM methylene chloride DMA Dimethylacetamide DMAP 4-dimethylaminopyridine DMF N,N'-dimethylformamide DMSO dimethyl sulfoxide DSC N,N'-Succinimidyl Carbonate EDTA Ethylenediaminetetraacetic acid Et Ethyl EtOAc ethyl acetate ffN fully functional nucleotides ffA fully functional adenosine nucleotide g grams GPC Gel Permeation Chromatography h or hr Time Hunig's base N,N-diisopropylethylamine iPr Isopropyl Kpi 10 mM potassium phosphate buffer, pH 7.0 IPA Isopropyl Alcohol LCMS Liquid Chromatography-Mass Spectrometry LDA Lithium diisopropylamide m or min MeCN acetonitrile mL milliliter PEG polyethylene glycol PG protecting group Ph phenyl pNB p-nitro-benzyl ppt precipitate rt room temperature SBS Sequencing by Synthesis -S(O)2OH sulfonyl hydroxide TEA Triethylamine TEAB Tetraethylammonium Bromide TFA trifluoroacetic acid Tert, t THF tetrahydrofuran TLC thin layer chromatography TSTU O-(N-Succinimidyl)-N,N,N',N'-tetramethyluronium tetrafluoroborate Rath μL microliter

[0048] As used herein, the term "array" refers to a collection of different molecules bound to one or more substrates. a population of probe molecules, the different probe molecules being arranged in a mutually distinct manner according to their relative positions; An array refers to a collection of probe molecules that can be individually distinguished from one another. may include different probe molecules each located at a different addressable location of Alternatively, or in addition, the array may comprise individual substrates each carrying a different probe molecule. the different probe molecules can be different from the location of the substrate on the surface to which the substrate is bound; Alternatively, the substrates can be identified by their position in the liquid. Exemplary arrays that may be used include, but are not limited to, those disclosed in U.S. Pat. No. 6,355,431, U.S. Patent Application Publication No. 2002 / 0102578, and and beads in the wells, as described in PCT Publication WO 00 / 63437. Liquid arrays, such as fluorescence activated cell sorters (FACS) The present invention can be used to distinguish beads in microfluidic devices. Exemplary formats that can be used are described, for example, in U.S. Pat. No. 6,524,793. Further examples of arrays that can be used in the present invention include limited Although not necessarily the case, U.S. Patent Nos. 5,429,807, 5,436,327, and Specification No. 5561071, Specification No. 5583211, Specification No. 5658734 , Specification No. 5837858, Specification No. 5874219, Specification No. 5919523 Specifications, Specification No. 6136269, Specification No. 6287768, No. 6287776 Specification No. 6288220, Specification No. 6297006, No. 62911 Specification No. 93, Specification No. 6346413, Specification No. 6416949, Specification No. 648 2591, 6514751, 6610482, International Publication Patent Nos. 93 / 17126, 95 / 11995, 95 / 35505, European Patent Examples include those described in US Pat. Nos. 742287 and 799897. .

[0049] As used herein, the terms "covalently attached" or "covalently bonded" "Covalently bonded" refers to a chemical bond characterized by the sharing of electron pairs between atoms. For example, a covalently bonded polymer coating can be formed by other means, e.g., attachment. or chemical binding to the functionalized surface of the substrate compared to binding to the surface via electrostatic interactions. Covalently bonded polymers to a surface are not covalently bonded to the surface. It will be understood that coupling is also possible by means of

[0050] As used herein, "C" is a set of integers where "a" and "b" are integers. a " ~ "C b " or "C a-b " refers to the number of carbon atoms in a particular group. That is, a group can have at least "a" and no more than "b" carbon atoms. Thus, for example, a "C1-C4 alkyl" group or a "C1-4 The "alkyl" group is All alkyl groups with 4 carbons, i.e., CH3-, CH3CH2-, CH3CH2CH2-, (CH3)2C H-, CH3CH2CH2CH2-, CH3CH2CH(CH3)-, and (CH3)3C-.

[0051] The term "halogen" or "halo" as used herein refers to fluorine, chlorine, bromine, or means any one of the radiation-stable atoms in the seventh column of the periodic table, such as iodine or iodine. Of these, fluorine and chlorine are preferred.

[0052] As used herein, "alkyl" refers to a group that is fully saturated (i.e., has no double or triple bonds). An alkyl group is a straight or branched hydrocarbon chain (including bonds) with 1 to 20 carbon atoms. may have elementary atoms (wherever this appears in this specification, a numerical range such as "1 to 20" refers to each integer in the given range, for example, "1 to 20 carbon atoms" ... The alkyl group has 1 carbon atom, 2 carbon atoms, 3 carbon atoms, or less than 20 carbon atoms. This definition also applies when no numerical range is specified. (This also applies to occurrences of the word "alkyl"). Alkyl groups also have 1 to 9 carbon atoms. The alkyl group can also be a medium size alkyl group having 1 to 4 carbon atoms. The alkyl group can be a lower alkyl having "C" atoms. 1-4 Alkyl or similar designation. 1-4 "Alkyl" is an alkyl The alkyl chain has 1 to 4 carbon atoms, that is, the alkyl chain is methyl, ethyl, or propyl. The group consisting of propyl, isopropyl, n-butyl, isobutyl, sec-butyl, and t-butyl Exemplary alkyl groups include, but are not limited to, Methyl, ethyl, propyl, isopropyl, butyl, isobutyl, tert-butyl, pentyl Alkyl groups can be substituted or unsubstituted. do.

[0053] As used herein, "alkoxy" refers to the formula -OR, where R is an alkyl group such as "C 1-9 Alkoxy and alkyl as defined above, including, but not limited to, methoxy, ethoxy, , n-propoxy, 1-methylethoxy (isopropoxy), n-butoxy, isobutoxy, s Examples include ec-butoxy and tert-butoxy.

[0054] As used herein, "heteroalkyl" refers to a group containing one or more heteroatoms, i.e., containing elements other than carbon in the chain backbone, including, but not limited to, nitrogen, oxygen, and sulfur. Heteroalkyl groups are straight or branched hydrocarbon chains containing 1 to 20 carbon atoms. However, this definition also applies to occurrences of the term "heteroalkyl" where no numerical range is specified. Heteroalkyl groups are also sometimes used in the art, and are generally medium-sized groups having 1 to 9 carbon atoms. The heteroalkyl group can also be a heteroalkyl group having 1 to 4 carbon atoms. The heteroalkyl group can be a lower heteroalkyl having the formula "C 1-4 Heteroa Heteroalkyl groups may be designated as "heteroalkyl" or similar designations. Heteroalkyl groups may have one or more heteroatoms. It can contain any number of heteroatoms. 1-4 "Heteroalkyl" means hetero The alkyl chain has 1 to 4 carbon atoms, and in addition, one or more heterocyclic groups in the chain backbone. Indicates the presence of a B atom.

[0055] As used herein, "alkylene" refers to a branched or straight chain alkyl group containing only carbon and hydrogen. refers to a fully saturated diradical chemical group in a chain, which has two points of attachment (i.e., alkanediyl The alkylene group has 1 to 20 carbon atoms. However, this definition also applies to the occurrence of the term alkylene where no numerical range is specified. Alkylene groups are also medium-sized alkylene groups having 1 to 9 carbon atoms. The alkylene group can also be a lower alkylene group having 1 to 4 carbon atoms. The alkyl group can be "C 1-4 "alkylene" or similar designation. For example, 1-4 "Alkylene" refers to an alkylene group with 1 to 4 carbon atoms in the alkylene chain. The presence of atoms, that is, alkylene chains are methylene, ethylene, ethane-1,1-diyl, Propylene, propane-1,1-diyl, propane-2,2-diyl, 1-methyl-ethylene, butyl butane-1,1-diyl, butane-2,2-diyl, 2-methyl-propane-1,1-diyl, 1-methyl -propylene, 2-methyl-propylene, 1,1-dimethyl-ethylene, 1,2-dimethyl-ethylene, and 1-ethyl-ethylene. If the alkylene chain is interrupted by an aromatic group, it is connected to the carbon atoms of the alkylene chain through two points of attachment. - Insertion of an aromatic group between the carbon bonds or one of the alkylene chains through one point of attachment Refers to the attachment of an aromatic group to the end. For example, n-butylene is interrupted by a phenyl group. , exemplary structures include: [ka] Examples include:

[0056] As used herein, the term "heteroalkylene" refers to one or more alkylene groups. The backbone atoms are atoms other than carbon, such as oxygen, nitrogen, sulfur, phosphorus, or a combination thereof. The length of the heteroalkylene chain is 2 to 20,000. Exemplary heteroalkylenes include, but are not limited to: -OCH2-, -OCH(CH3)-, -OC(CH3)2-, -OCH2CH2-, -CH(CH3)O-, -CH2OCH2-, -CH2OCH2CH2-, -SCH2-, -SCH(CH3)-, -SC(CH3)2-, -SCH2CH2-, -CH2SCH2CH2-, -NHCH2-, -NHCH(CH3)-, - Examples include -NHC(CH3)2-, -NHCH2CH2-, -CH2NHCH2-, and -CH2NHCH2CH2-. As used herein, when a heteroalkylene is interrupted by an aromatic group, it is One carbon-carbon bond or carbon-heteroatom bond of the heteroalkylene chain via a point Aromatic groups can be inserted between or at one end of a heteroalkylene chain via a single point of attachment. For example, n-propylene oxide is interrupted by a phenyl group. In this case, exemplary structures include: [ka] Examples include:

[0057] As used herein, "alkenyl" refers to a straight or branched hydrocarbon chain having one or more alkyl groups. An alkenyl group may be unsubstituted or substituted. It is possible.

[0058] As used herein, "alkynyl" refers to a straight or branched hydrocarbon chain having one or more alkynyl groups. An alkynyl group refers to an alkyl group containing one or more triple bonds. It can be said that:

[0059] As used herein, "cycloalkyl" refers to a fully saturated (both double and triple bonds) alkyl group. It refers to a monocyclic or polycyclic hydrocarbon ring system (or rings containing two or more rings). The rings may be joined together in a fused fashion. The cycloalkyl group may have 3 to 10 ring atoms. In some embodiments, the cycloalkyl group can contain 1 to 2 atoms in the ring. It can contain 3 to 8 atoms. The cycloalkyl group can be unsubstituted or substituted. Typical cycloalkyl groups include, but are not limited to: , cyclopropyl, cyclobutyl, cyclopentyl, cyclohexyl, cycloheptyl, and cyclooctyl.

[0060] The term "aromatic" refers to a ring or ring system having a π-electron conjugated system, including carbocyclic aromatic rings. These include both aromatic (e.g., phenyl) and heteroaromatic (e.g., pyridine) groups. The term includes monocyclic or fused polycyclic (i.e., adjacent rings) rings, provided that all ring systems are aromatic. This includes groups in which the rings share one atom pair.

[0061] As used herein, "aryl" refers to an aromatic ring or rings containing only carbon in the ring backbone. refers to an aromatic ring system (i.e., two or more fused rings sharing two adjacent carbon atoms). When an aryl is a ring system, all rings in the system are aromatic. An aryl group has 6 to 18 carbon atoms. Although the term "aryl" does not specify a numerical range, this definition does not apply to the term "aryl" which does not specify a numerical range. In some embodiments, the aryl group is a 6-10 carbon atom group. The aryl group has "C 6-10 aryl," "C6 or C 10 aryl" or similar Examples of aryl groups include, but are not limited to, , phenyl, naphthyl, azulenyl, and anthracenyl.

[0062] "Aralkyl" or "arylalkyl" refers to the "C 7-14 Aralkyl Aryl groups bonded as substituents via an alkyl group, including but not limited to benzoyl groups. Examples include diphenyl, 2-phenylethyl, 3-phenylpropyl, and naphthylalkyl. In some cases, the alkylene group may be a lower alkylene group (i.e., C 1-4 alkylene group).

[0063] As used herein, "heteroaryl" refers to an aromatic ring or ring system (i.e., refers to two or more fused rings that share two adjacent atoms, which can be one or more Heteroatoms, i.e., atoms other than carbon, including but not limited to nitrogen, oxygen, and sulfur In the case where the heteroaryl is a ring system, all rings in the system contain the elements Heteroaryl groups are aromatic and have 5 to 18 ring members (i.e., carbon atoms and heteroatoms). The number of atoms constituting the ring skeleton (including the number of atoms constituting the ring skeleton) can be any number, but this definition does not apply to cases where a numerical range is specified. This also applies to occurrences of the term "heteroaryl" where the term is not present. The heteroaryl group has 5 to 10 ring members or 5 to 7 ring members. "5- to 7-membered heteroaryl," "5- to 10-membered heteroaryl," or similar designations. Examples of heteroaryl rings include, but are not limited to, furyl, Thienyl, phthalazinyl, pyrrolyl, oxazolyl, thiazolyl, imidazolyl, pyrazolyl aryl, isoxazolyl, isothiazolyl, triazolyl, thiadiazolyl, pyridinyl , pyridazinyl, pyrimidinyl, pyrazinyl, triazinyl, quinolinyl, isoquinolinyl benzimidazolyl, benzoxazolyl, benzothiazolyl, indolyl, isopropyl and benzothienyl.

[0064] A "heteroaralkyl" or "heteroarylalkyl" is a heteroaryl group substituted via an alkylene group. Examples include, but are not limited to, heteroaryl groups bonded as substituents. 2-thienylmethyl, 3-thienylmethyl, furylmethyl, thienylethyl, pyrrolylalkyl alkyl, pyridylalkyl, isoxazolylalkyl, and imidazolylalkyl. In some cases, the alkylene group may be a lower alkylene group (i.e., C 1-4 alkylene group) be.

[0065] As used herein, "cycloalkyl" refers to a fully saturated carbocyclyl ring or means a carbocyclyl ring system. Examples include cyclopropyl, cyclobutyl, cyclopentyl, Examples include butyl, butyl, and cyclohexyl.

[0066] An "O-carboxy" group refers to an "-OC(=O)R" group, where R is as defined herein. Sea urchin, hydrogen, C 1-6 Alkyl, C 2-6 Alkenyl, C 2-6 Alkynyl, C 3-7Carbocyclyl, C6 -10 It is selected from aryl, 5- to 10-membered heteroaryl, and 5- to 10-membered heterocyclyl.

[0067] A "C-carboxy" group refers to a "-C(=O)OR" group, where R is as defined herein. to, hydrogen, C 1-6 Alkyl, C 2-6 Alkenyl, C 2-6 Alkynyl, C 3-7 Carbocyclyl, C 6-1 aryl, 5- to 10-membered heteroaryl, and 5- to 10-membered heterocyclyl. A non-limiting example is carboxyl (ie, -C(=O)OH).

[0068] A "cyano" group refers to a "-CN" group.

[0069] An "azido" group refers to an "-N3" group.

[0070] The "O-carbamyl" group is "-OC(=O)NR A R B " group, where R A and R B are independent of each other. and, as defined herein, hydrogen, C 1-6 Alkyl, C 2-6 Alkenyl, C 2-6 Alkini Lu, C 3-7 Carbocyclyl, C 6-10 Aryl, 5- to 10-membered heteroaryl, and 5- to 10-membered heteraryl is selected from: cyclocyclyl;

[0071] The "N-carbamyl" group is "-N(R A )OC(=O)R B " group, where R A and R B are independent and, as defined herein, hydrogen, C 1-6Alkyl, C 2-6 Alkenyl, C 2-6 Archi Nill, C 3-7 Carbocyclyl, C 6-10 Aryl, 5- to 10-membered heteroaryl, and 5- to 10-membered hexaaryl is selected from tetracyclyl.

[0072] The "C-amide" group is "-C(=O)NR A R B " group, where R A and R B are each independently As defined herein, hydrogen, C 1-6 Alkyl, C 2-6 Alkenyl, C 2-6 Alkynyl, C 3-7 Carbocyclyl, C 6-10 Aryl, 5- to 10-membered heteroaryl, and 5- to 10-membered heteroaryl Selected from Krill.

[0073] The "N-amide" group is "-N(R A )C(=O)R B " group, where R A and R B are each independently , as defined herein, hydrogen, C 1-6 Alkyl, C 2-6 Alkenyl, C 2-6 Alkynyl , C 3-7 Carbocyclyl, C 6-10 Aryl, 5- to 10-membered heteroaryl, and 5- to 10-membered heteroaryl cyclyl.

[0074] The "amino" group is "-NR A R B " group, where R A and R B are each independently defined in this specification. hydrogen, C, as defined by 1-6 Alkyl, C 2-6 Alkenyl, C2-6 Alkynyl, C 3-7 Cal Bocyclyl, C 6-10 Aryl, 5- to 10-membered heteroaryl, and 5- to 10-membered heterocyclyl Non-limiting examples include free amino (i.e., -NH2).

[0075] As used herein, the term "Trolox" refers to 6-hydroxy-2,5,7,8-tetrahydrofuran. It refers to tetramethylchroman-2-carboxylic acid.

[0076] As used herein, the term "ascorbate" refers to ascorbate salt.

[0077] As used herein, the term "gallic acid" refers to 3,4,5-trihydroxybenzoic acid.

[0078] As used herein, a substituent is a group in which one or more hydrogen atoms are exchanged with another atom or group. Derived from an unsubstituted parent group unless otherwise indicated. As long as a group is considered to be "substituted," it means that the group is substituted with C1-C6 alkyl, C1-C6 Alkenyl, C1-C6 alkynyl, C1-C6 heteroalkyl, C3-C7 carbocyclyl (halo, C1- Substituted with C6 alkyl, C1-C6 alkoxy, C1-C6 haloalkyl, and C1-C6 haloalkoxy C3-C7-carbocyclyl-C1-C6-alkyl (halo, C1-C6 alkyl, C1 -Optionally substituted with C6 alkoxy, C1-C6 haloalkyl, and C1-C6 haloalkoxy ), 5-10 membered heterocyclyl (halo, C1-C6 alkyl, C1-C6 alkoxy, C1-C6 haloal alkyl, and C1-C6 haloalkoxy), 5-10 membered heterocyclyl-C 1-C6-Alkyl (halo, C1-C6 alkyl, C1-C6 alkoxy, C1-C6 haloalkyl, and C1 -C6 haloalkoxy), aryl (halo, C1-C6 alkyl, C1-C6 optionally substituted with alkoxy, C1-C6 haloalkyl, and C1-C6 haloalkoxy , aryl(C1-C6) alkyl (halo, C1-C6 alkyl, C1-C6 alkoxy, C1-C6 haloalkoxy and C1-C6 haloalkoxy), 5-10 membered heteroaryl (halo (b) C1-C6 alkyl, C1-C6 alkoxy, C1-C6 haloalkyl, and C1-C6 haloalkoxy substituted with), 5-10 membered heteroaryl(C1-C6)alkyl (halo, C1-C6 alkoxy) substituted with alkyl, C1-C6 alkoxy, C1-C6 haloalkyl, and C1-C6 haloalkoxy; (optionally), halo, cyano, hydroxy, C1-C6 alkoxy, C1-C6 alkoxy(C1-C6) Alkyl (i.e., ether), aryloxy, sulfhydryl (mercapto), halo (C1-C6) alkyl (e.g., -CF3), halo(C1-C6) alkoxy (e.g., -OCF3), C1-C6 alkyl Alkylthio, arylthio, amino, amino(C1-C6)alkyl, nitro, O-carbamyl, N -Carbamyl, O-thiocarbamyl, N-thiocarbamyl, C-amido, N-amido, S-sulfone Amide, N-sulfonamide, C-carboxy, O-carboxy, acyl, cyanate, isocyanate Nato, thiocyanato, isothiocyanato, sulfinyl, sulfonyl, and oxo(= O) Whenever a group is described as being "optionally substituted," the group may be substituted with any of the above substituents. It can be substituted.

[0079] Some radical nomenclature may include either monoradicals or diradicals, depending on the context. For example, if a substituent requires two points of attachment to the rest of the molecule, When a group is present, it is understood that the substituent is a diradical. For example, Substituents specified as alkyl include -CH2-, -CH2CH2-, and -CH2CH(CH3)CH2-. Similarly, diradicals such as amine, which requires two points of attachment, are identified as Suitable radicals include -NH- and diradicals such as -N(CH3)-. Other radical nomenclature is clearly To be sure, the radical is a diradical such as "alkylene" or "alkenylene." show.

[0080] The substituent is drawn as a diradical (i.e., has two points of attachment to the rest of the molecule). ) wherever possible, the substituents may be attached in any directed configuration unless otherwise indicated. Thus, for example, -AE- or [ka] The substituent drawn has A oriented so that it is attached to the leftmost attachment point of the molecule. groups, and when A is attached to the rightmost attachment point of the molecule.

[0081] When the compounds disclosed herein have at least one stereocenter, they may be attached to the respective mirror image. As enantiomers and diastereomers, or mixtures of such isomers, including racemates The separation of individual isomers or the selective synthesis of individual isomers is well known to those skilled in the art. Unless otherwise indicated, all such isomers are and mixtures thereof are included within the scope of the compounds disclosed herein.

[0082] As used herein, a "nucleotide" includes a heterocyclic base, a sugar, and one or more They contain nitrogen atoms containing several phosphate groups. They are the monomer units of nucleic acid sequences. In A, the sugar is ribose, and in DNA, it is deoxyribose, i.e., present in ribose It is a sugar lacking a hydroxyl group. Nitrogen containing heterocyclic bases are purine bases or pyridines. The purine bases include adenine (A) and guanine (G), and Pyrimidine bases include cytosine (C), thymine (T), thiamin (H), thiamin (T ... These include thiamin (T), thiamin (T), and uracil (U), as well as modified derivatives or analogs thereof. The C-1 atom of oxyribose is linked to the N-1 of a pyrimidine or the N-9 of a purine.

[0083] As used herein, a "nucleoside" is structurally similar to a nucleotide, but The phosphate moiety is missing. An example of a nucleoside analog is one in which the label is attached to the base and the sugar moiety is attached to the base. The term "nucleoside" is used herein to mean a nucleotide that is a nucleotide sequence that will be understood by those skilled in the art. The term "common name" is used in its ordinary and understood sense. Examples include, but are not limited to: Ribonucleosides containing a ribose moiety and deoxyribonucleosides containing a deoxyribose moiety A modified pentose moiety is one in which the oxygen atom is replaced by a carbon atom and / or is a pentose moiety in which a carbon atom is replaced by a sulfur atom or an oxygen atom. A "nucleoside" is a monomer that may have a substituted base and / or sugar moiety. can be incorporated into larger DNA and / or RNA polymers and oligomers.

[0084] As used herein, the term "polynucleotide" generally refers to nucleic acids, including DNA (e.g., , genomic DNA, cDNA), RNA (e.g., mRNA), synthetic oligonucleotides, and synthetic nucleic acids Polynucleotides may contain natural or unnatural bases or combinations thereof. and natural or non-natural backbone linkages, e.g., phosphoro Thioates, PNAs, or 2'-O-methyl-RNAs, or combinations thereof, may be included.

[0085] As used herein, the term "phasing" refers to a phenomenon in SBS, This results in incomplete removal of the 3' terminator and fluorophore, and In the single cycle, the polymerase completes the incorporation of a portion of the DNA strand within the cluster. Pre-phasing is caused by the inability to catalyze effective 3' It is caused by the incorporation of a non-terminator nucleotide, and the incorporation event is The phasing and prephasing are extracted for a specific cycle. The intensity measured is the signal of the current cycle and the noise of the preceding and following cycles. As the number of cycles increases, each clock affected by fading Prephasing increases the number of sequence fragments in the cluster, preventing the identification of the correct base. This can be caused by the presence of traces of 3'-OH nucleotides, either protected or unblocked. Protective 3'-OH nucleotides can be generated during the manufacturing process or can be used for storage and handling of reagents. This can be generated during the handling process, thus allowing for faster SBS cycle times and lower Low phasing and prephasing values and longer sequencing reads Nucleotide analogs or linker group modifications that result in longer sequences can be used in SBS applications. brings great advantages.

[0086] As used herein, the term "protective moiety" includes, but is not limited to, a protective moiety that protects against DNA damage. These include molecules that can protect against (e.g., photodamage or other chemical damage). Examples include vitamin C, vitamin E derivatives, phenolic acids, polyphenols, and The term "protective moiety" is defined as follows: In this context, it refers to one or more functional groups of the protecting moiety and a linker as described herein. For example, if a protected moiety is "gallic acid" ", it is not gallic acid itself with a free carboxyl group, but the acid of gallic acid. It may refer to midos and esters.

[0087] Detectable Label Some embodiments described herein involve the use of conventional detectable labels. This can be done by any suitable method, including fluorescence spectroscopy or other optical methods. A preferred label is a fluorophore, which emits at a defined wavelength after absorption of energy. Many suitable fluorescent labels are known, see, for example, Welch et al. (Chem. Eur. J. 5(3):951-960, 1999) is a dansyl-functionalized amine that can be used in the present invention. (Cytometry 28:206-211, 1997) discloses fluorescent moieties. The use of the labels Cy3 and Cy5 has been described and may also be used in the present invention. The labeling was also performed as described by Prober et al. (Science 238:336-341, 1987); Connell et al. (BioTechn iques 5(4):342-384, 1987), Ansorge et al. (Nucl. Acids Res. 15(11):4593-4602, 19 87), and Smith et al. (Nature 321:674, 1986). Possible fluorescent labels include, but are not limited to, fluorescein, rhodamine (TM), R, Texas Red® and Rox), Alexa, bodipy, ac These include lysine, coumarin, pyrene, benzanthracene, and cyanine.

[0088] Multiple labels, such as bi-fluorophore FRET cassettes (Tet. Lett. 46:8867-8871, 2000 ) can also be used in this application. Multi-fluorescent dendrimeric system (J. Am. Chem. Soc. 123:8101-8108, 2001) can also be used. Fluorescent labels are preferred, but other It will be apparent to one skilled in the art that detectable labels in the form of, for example, quantum dots, (Empodocles et al., Nature 399:126-130, 1999), gold nanoparticles (Reichert et al., An al. Chem. 72:6025-6029, 2000), and microbeads (Lacoste et al., Proc. Natl Acad. Sci USA 97(17):9461-9466, 2000) and other fine particles can also be used. .

[0089] Multi-component labels may also be used in this application. Multi-component labels may contain additional compounds for detection. The most common multi-component labels used in biology are biomolecules. The biotin-streptavidin system is a target that binds to nucleotide bases. Streptavidin is then added separately to allow detection. Other multicomponent systems are available. For example, dinitrophenol can be used for detection. There are commercially available fluorescent antibodies that can

[0090] Unless otherwise indicated, reference to a nucleotide is also applicable to a nucleoside This application will also be further described in relation to DNA, but that description may also be used interchangeably. Unless otherwise indicated, the present invention may be applicable to RNA, PNA, and other nucleic acids.

[0091] Sequencing Methods The nucleosides or nucleotides described herein are compatible with various sequencing techniques. In some embodiments, the nucleotide sequence of the target nucleic acid can be determined by The determining process may be an automated process.

[0092] The nucleotide analogs presented herein can be used in methods such as sequencing-by-synthesis (SBS) methods. Briefly, SBS involves the step of cleaving a target nucleic acid into a single molecule. Initiated by contacting the nucleic acid with one or more labeled nucleotides, DNA polymerase, etc. It is possible to extend the primer using the target nucleic acid as a template. These features may incorporate detectably labeled nucleotides. The labeled nucleotides further allow for further processing once the nucleotides have been added to the primer. It is possible to provide a reversible termination feature that terminates the extension of the primer. A nucleotide analog having a target terminator moiety is introduced into the target region by delivering a deblocking agent to the target region. It is possible to add a primer so that subsequent extension cannot occur until the component is removed. Therefore, in embodiments using reversible termination, the deblocking reagent is added to the flow cell (detector). Washing can be performed between the various delivery steps. The cycle is then repeated n times to extend the primer by n nucleotides. , thereby making it possible to detect sequences of length n. An exemplary SBS procedure, fluidics system, which can be easily adapted for use with arrays, is shown. The system and detection platform are described, for example, in Bentley et al., Nature 456:53-59 (20 08), International Publication Nos. 04 / 018497; 91 / 06678; 07 / 1237 44; U.S. Patent Nos. 7,057,026; 7,329,492; 721 1414; 7315019, or 7405281; and U.S. Patent Application Publication No. 2008 / 0108082, each of which is incorporated herein by reference.

[0093] Use other sequencing procedures that use cyclic reactions, such as pyrosequencing Pyrosequencing is the process by which specific nucleotides are incorporated into nascent nucleic acid strands. detects the release of inorganic pyrophosphate (PPi) when the protein is absorbed (Ronaghi, et al., Analytical Bio chemistry 242(1), 84-9 (1996);Ronaghi, Genome Res. 11(1), 3-11 (2001);Ronaghi et al. Science 281(5375), 363 (1998); U.S. Pat. No. 6,210,891; U.S. Pat. No. 6,210,891 Nos. 58568 and 6274320, each of which is incorporated by reference. In pyrosequencing, the released PPi is converted to ATP sulfur. It can be detected by converting it into adenosine triphosphate (ATP) by acetyltransferase. and the resulting ATP can be detected via luciferase-generated photons. Therefore, the sequencing reaction is monitored via a luminescence detection system. The excitation radiation source used in fluorescence-based detection systems is pyrochlore. It is not necessary for the pyrosequencing procedure. Useful fluidic systems, detectors, and procedures that can be used are described, for example, in WIPO patents Application No. PCT / US11 / 57111, U.S. Patent Application Publication No. 2005 / 0191698 No. 7,595,883, and No. 7,244,559. Nos. 10 / 139,999, ...

[0094] See, e.g., Shendure et al. Science 309:1728-1732 (2005); U.S. Pat. No. 5,599,675 No. 5,750,341, each of which is incorporated herein by reference. Sequencing-by-ligation reactions, including those described in (incorporated herein) are also useful. Some embodiments are described, for example, in Bains et al., Journal of Theoretical Biology 135(3), 30 3-7 (1988);Drmanac et al., Nature Biotechnology 16, 54-58 (1998);Fodor et al., Science 251(4995), 767-773 (1995); and International Publication No. WO 1989 / 10977 (which (each of which is incorporated herein by reference) Sequencing-by-ligation and sequencing-by-hybridization procedures In both cases, nucleic acids present in gel-containing wells (or other recessed features) are transferred to oligonucleotides. The method is described herein or is subject to repeated cycles of nucleotide delivery and detection. The fluidic system for the SBS method in the reference cited in the book is a sequencing-by-ligation system. or can be easily adapted for delivery of reagents for sequencing-by-hybridization procedures. Typically, the oligonucleotides are fluorescently labeled and can be used in the methods described herein or in the present invention. Using a fluorescence detector, similar to those described for SBS procedures in the references cited herein. It is possible to detect

[0095] Some embodiments utilize methods involving real-time monitoring of DNA polymerase activity. For example, nucleotide incorporation can be achieved by using fluorophore-bearing polymers. Fluorescence resonance energy transfer (FRET) interaction between γ-phosphate-labeled nucleotides and γ-phosphate-labeled nucleotides Detection can be achieved via the FRET-based system or by using the zero mode wavelength. Techniques and reagents for sequencing are described, for example, in Levene et al. Science 299, 682-686. (2003);Lundquist et al. Opt. Lett.33, 1026-1028 (2008);Korlach et al. Proc. N atl. Acad. Sci. USA 105, 1176-1181 (2008), the disclosures of which are incorporated herein by reference. No. 60 / 699,493, filed on Oct. 1, 2003, which is incorporated

[0096] Some SBS embodiments include photons emitted upon incorporation of a nucleotide into an extension product. For example, sequencing based on the detection of emitted photons is being developed by Ion Torrent ( Commercially available from AquaLife Technologies, a subsidiary of AquaLife Technologies, Inc., Guilford, CT Electrical detectors and related techniques, or U.S. Patent Application Publication No. 2009 / 0026082 Specifications; Specification No. 2009 / 0127589; Specification No. 2010 / 0137143 or US Patent Application Publication No. 2010 / 0282617 (each of which is incorporated herein by reference). The sequencing methods and systems described in is.

[0097] Exemplary Modified Linkers Additional embodiments are disclosed in more detail in the following examples, which are also set forth in the claims. It is not intended to be limiting in any way.

[0098] FIG. 1A shows a partial structural formula of labeled nucleotide 100. Labeled nucleotide 100 is a complete functional group. a functional adenosine nucleotide (ffA) 110, a standard linker portion 115, and a fluorescent dye 120. The standard linker portion 115 is used for the synthesis of labeled nucleotides for sequencing-by-synthesis (SBS). In one example, the fluorescent dye 120 is NR550S0 In this example, labeled nucleotide 100 can be written as "ffA-NR550S0." In another example, the fluorescent dye 120 is SO7181 and the labeled nucleotide 100 is described as "ffA-SO7181." It is possible.

[0099] FIG. 1B shows the target of FIG. 1A with two possible structural modifications to the standard linker portion 115. In one example, the standard linker portion 115 is a carbonyl ( It contains an AEDI insertion 125 between the -C(=O)-) and amino (i.e., -NH-) moieties. The decorated labeled nucleotide can be written as "ffA-AEDI-NR550S0". The linker portion 115 contains an SS insertion 130, and the modified labeled nucleotide is designated "ffA-SS-NR550S0." It is possible.

[0100] FIG. 2 shows a modified labeled nucleotide 100 of FIG. 1A using the labeled nucleotide 100 of FIG. 1B. A graph of the nucleotide incorporation rate is shown. The assay was performed at 55°C in 40 mM ethanolamine. (pH 9.8), 9 mM MgCl , 40 mM NaCl, 1 mM EDTA, 20 nM primer:template DNA and 30 μL in 0.2% CHAPS, 1 mM nucleotides with 10 μg / mL of polymerase 812 (MiSeq Kit V2). The enzyme was bound to the DNA and then briefly (up to 10 seconds) placed in a quench flow machine. The mixture is rapidly mixed with nucleotides at 20°C and then quenched with 500 mM EDTA. The resulting samples were then analyzed on a denaturing gel and converted to DNA+1. The percentage of DNA converted was determined and plotted against time to determine the percentage of DNA converted for each nucleotide. The first-order rate constant is calculated using the data tabulated in Table 1 below. The incorporation rate of the labeled nucleotide (ffa-SS-NR550S0) was compared with that of the labeled nucleotide containing the standard linker 115. The uptake rate was approximately twice as fast as that of eosinophil (ffa-NR550S0). The incorporation rate of the labeled nucleotide containing 25 (ffa-AEDI-NR550S0) was The data also show that the speed of the standard linker 115 and the fluorescent dye SO718 was approximately four times faster than that of the standard linker 115 and the fluorescent dye SO718. The incorporation rate of the labeled nucleotide containing 1 (ffA-SO7181) was This shows that it was about four times faster than the previous method.

[0101] [Table 1]

[0102] 3A-3F show additional insertions 310, 315, 320, 325, 330, 331, 332, 333, 334, 335, 336, 337, 338, 339, 340, 341, 342, 343, 344, 345, 346, 347, 348, 3 The structural formulas of 30 and 335 are shown. In ACA insertion 310, compared to insertion 125, a dimethyl substituent is added. The sulfur-sulfur (SS) bond is then replaced with a carbon-carbon bond. For example, it is not necessary for a two-dye SBS or a four-dye SBS.

[0103] AcLys insertion 315 replaces insertion 125 with an acetyl-protected lysine.

[0104] For the BocLys insertion 320, a tert-butoxycarbonyl-protected lysine was used to insert 125. Replace.

[0105] In the dMeO insertion 325, the sulfur-sulfur (SS) bond was replaced by an oxygen-carbon (O-CH2) bond compared to the insertion 125. Replace it with the correct one.

[0106] In the dMeS insertion 330, the sulfur-sulfur (SS) bond was replaced by a sulfur-carbon (S-CH2) bond compared to the insertion 125. Replace it with the correct one.

[0107] In DMP insertion 335, sulfur-sulfur (SS) bonds were replaced with sulfur-carbon (CH2) bonds compared to insertion 125. and exchange it for.

[0108] In various examples described herein, insertions containing dimethyl substitution patterns (e.g., AE DI insertion 125, dMeO insertion 325, and dMeS insertion 330) showed the nucleotide incorporation rate during SBS. was found to be fast.

[0109] In various examples described herein, the length of the carbon chain in the insertion may also be varied. can be done.

[0110] Figure 4 shows the sequencing quality of the AEDI insertion 125 in Figure 1B and the AcLys insertion 315 in Figure 3B. Figure 1 shows a data table for a two-dye sequencing run used to evaluate the effect of Sequencing was performed on a Miseq Hybrid Platform using a human 550 bp template for two 150 cycle runs. The new dye set, V10 / cyan-peg4, A-AEDI550S0, V10 / cyan-pe g4 A-AcLys550S0, and V10 / cyan A-AcLys550S0 were used in the standard commercial dye sets V4 and and Nova platform V5.75 improved dye set. V10 / cyan-peg4 A-AEDI550S0 sample The V10 / cyan-peg4 A-AcLys550S0 sample and the V10 / cyan A-AcLys550S0 sample were For each, the phasing value (Ph R1) is the same as for the samples without AEDI insertion and without AcLys insertion. Therefore, the fading values of the additional insertions 125 and 315 were included. Labeled nucleotides showed improved sequencing quality.

[0111] Figures 5A and 5B show the error rate graphs for read 1 of the sequencing run in Figure 4, respectively. The graph shows the error rate for rough and lead 2. For lead 1, V10 / cyan-peg4 A-AEDI550S The error rates of V10 / cyan-peg4 A-AcLys550S0 and V10 / cyan-peg4 A-AcLys550S0 were similar to those of the dye V10 / Cyan-peg4 without the insertion. In lead 2, the difference was even more pronounced with the AcLys insertion 315. The final error rate was reduced by 30% compared to the no-insertion dye set. 15 was found to significantly improve sequencing quality.

[0112] Figure 6 evaluates the effect of AEDI insertion 125 and ACA insertion 310 on sequencing quality. The data table shows the sequencing run used to sequence the human 550 bp template. The mold was run twice for 150 cycles on the Miseq hybrid platform. The dye sets, V10 / cyan-peg4 A-AEDI550S0 and V10 / cyan-peg4 A-ACALys550S0, were used in standard commercial dye sets. The comparison was made with the commercial dye set V4 and the improved dye set of the Nova platform V5.75. 10 / cyan-peg4 A-AEDI550S0 sample and V10 / cyan-peg4 A-ACA550S0 sample, respectively The phasing values (Ph R1) were lower compared to the samples without AEDI and ACA insertions. Thus, it demonstrated an improvement in sequencing quality.

[0113] Figures 7A and 7B show graphs of the error rate for read 1 of the sequencing run in Figure 6, respectively. The graphs show the error rates of lead 1 and lead 2. The new insertion AEDI (V10 / cyan-peg4 A-AEDI) The error rate for read 1 in the set containing the insert-free dye, V10 / Cyan-peg4, was similar to that of the insert-free dye, V10 / Cyan-peg4. The insert ACA (V10 / Cyan-peg4 ACA550S0) was lower than the standard V10 / Cyan-peg4. AEDI again demonstrated its commitment to improving sequencing quality. These data also suggest that the structure of the insert itself can have an effect on improving sequencing quality. It was shown that it gives

[0114] FIG. 8A shows the structural formula of a standard LN3 linker 800. LN3 linker 800 is a linker having a first substituted amide group. a functional moiety 810, a second azide-substituted PEG functional moiety 815, and a third ester functional moiety 820, These are preferably linker structures for connecting the dye molecule 825 to the nucleotide 830. The first functional moiety 810 is used, for example, to attach a dye molecule 825 to the LN3 linker 800. The second functional group 815 can, for example, cleave the dye molecule 825 from the LN3 linker 800. The third functional moiety 820 is a cleavable functional group that can be used to cleave the nucleic acid. 8B can be used to attach a leukocyte 830 to an LN3 linker 800. Some modifications to anchor 800 are shown, where the phenoxy moiety 850 is replaced with -NO2, -CN, halo or -SO3H. In addition, the ester moiety Moiety 820 is exchanged with amide moiety 855.

[0115] Figure 9A is a graph showing the appearance of impurities in ffA with an SS linker. In the purified ffA-SS-NR550S0, impurities appeared overnight under slightly basic conditions (pH 8-9). In 1 M TEAB / CH3CN at room temperature overnight. Figure 9B shows the ffA with the SS linker and the AEDI linker. The stability of ffA with the AEDI linker was significantly improved compared to the SS linker. The stability of ffA-LN3-NR550S0 was improved. Disulfide by-products were detected. :(NR550S0-S-)2. Figure 9C shows the results of SS linker ffA and AEDI linker ffA at 60°C for 22 hours. The comparison of the two is shown, again showing impurities at the SS linker. The internal control is ffA-LN3-NR550S0. Di-P: diphosphate.

[0116] Figures 10A, 10B, and 10C show the nucleotide uptake in solution with varying linkers. Figure A shows the uptake rate at 1 μM. The results are The dye and linker significantly influence the uptake kinetics, as shown in Table 10. Figure 10B clearly shows the advantage of the AEDI linker in uptake kinetics. Schematic representation of AEDI and SS linkers along with 0S0.

[0117] FIG. 11A shows a scatter plot for V10 combinations with different A-550S0 (same concentrations). The scatter plot is tile 1, cycle 2. Figure 11B shows the Kcat FFA linker in solution. On the surface, the AEDI linker was found to be faster than the no-linker in terms of uptake rate. The "A" cloud (in the scatter plot) is slightly centered. The AcLys linker is Slower than without a linker: "A" cloud is closer to the x-axis. In solution, the ACA linker is the slowest, while the A It can be seen that the Kcat of EDI and ACLy are similar, followed by BocLys.

[0118] Figures 12A and 12B show the sequencing criteria for M111, human 550, and 2x151 cycles. The use of ffN in combination with different A-550S0 (same concentration). It can be seen that both the AEDI and the AEDI provide similar and good sequencing results. The Kcat values for the AcLys and AcLys solutions are similar, but the AEDI sequencing results are slightly better. .

[0119] The LN3 linker 800 includes an optional phenoxy moiety 835, which is shown in Figure 8C. The N3 linker 800 can also be removed from the N3 linker 800. The N3 linker 800 can also be removed from the N3 linker 800. portion 840 and optional ether portion 845, both of which are shown in FIG. 8D The LN3 linker 800 can be removed from the amide moiety 810, the phenoxy moiety 835, and other The purpose of removing certain functional groups is to prevent them from reducing the efficiency of incorporation during nucleotide incorporation. This is to test whether the compound has any negative interactions with enzymes that may be involved.

[0120] Figure 8E shows the insertion or addition of a protecting moiety 860 to the linker. Protecting moiety 860 is a functional group. 8A-8C). (e.g., can bind to 850). Protective moiety 860 can be, for example, a molecule that prevents DNA damage. DNA damage, including photodamage or other chemical damage, can be a cumulative effect of SBS (i.e., Substantially reducing or eliminating DNA damage is one of the It can provide efficient SBS and longer sequencing reads. In one embodiment, the protective moiety 860 is a compound selected from the group consisting of Trolox, gallic acid, 2-mercaptoethanol (B It is possible to select from triplet state quenchers such as cyclohexylmethyl ether (ME). In embodiments, protecting moiety 860 is 4-nitrobenzyl alcohol, or ascorbic acid. from quenching or protecting reagents, such as salts of ascorbic acid, such as sodium ascorbate; In some other embodiments, the protecting agent is a labeled nucleoside. Rather than forming a covalent bond with the nucleotide or nucleotide, the nucleotide is physically mixed in the buffer. However, this approach requires higher concentrations of the protectant and is less effective. Alternatively, a protecting moiety covalently attached to the nucleoside or nucleotide may be used. However, some further details in Figures 8C, 8D, and 8E are not available. In some embodiments, ester moiety 821 can also be replaced with amide moiety 855. The phenoxy moiety 835 can be further substituted.

[0121] In any of the examples shown in FIGS. 8A-8E, the AEDI insert 125 and SS insert 130 of FIG. 1B, and 3A-3F, the insertion 300 may be, for example, a first functional moiety 810 of a linker 800 and a dye moiety 825 can be inserted between [Example]

[0122] example Additional embodiments are disclosed in more detail in the following examples, which are incorporated by reference in their entirety. It is not intended to be limiting in any way.

[0123] ffA-LN 3 -General reaction procedure for preparing AEDI-NR550S0: [ka]

[0124] In a 50 mL round-bottom flask, Boc-AEDI-OH (1 g, 3.4 mmol) was dissolved in DCM (15 mL) and TFA (1. 3 mL, 17 mL) was added to the solution at room temperature. The reaction mixture was stirred for 2 hours. TLC (DCM:MeOH=9:1) showed , indicating complete consumption of Boc-AEDI-OH. The reaction mixture was evaporated to dryness. TEAB (2M, ∼15 mL) was then added to the residue and the pH was monitored until neutral. The mixture was then diluted with HO / CHCN (1:1 The solution was dissolved in 100 mL of TEAB (~15 mL) and evaporated to dryness. The procedure was repeated three times to remove excess TEAB salt. The white solid residue was treated with CH3CN (20 mL) and stirred for 0.5 h. The solution was filtered and the solid was removed with CH3CN The pure AEDI-OH-TFA salt was obtained (530 mg, 80%). 1 H NMR (400 MHz, DO, δ (ppm)) ;3.32(t, J=6.5 Hz, 2H, NH2-CH2);2.97(t, J=6.5 Hz, 2H, S-CH2);1.53(s, 6H , 2x CH3). 13 C NMR (400 MHz, DO, δ(ppm)): 178.21(s, CO); 127.91, 117.71(2 s, TFA);51.87(s,C-(CH3)2);37.76(t,CH2-NH2);34.23(t,S-CH2);23.84(q , 2x CH3). 19 F NMR (400 MHz, D2O, δ(ppm)): -75.64.

[0125] In a 50 mL round-bottom flask, dissolve the dye NR550S0 (114 mg, 176 μmol) in DMF (anhydrous, 20 mL). The procedure was repeated three times. Anhydrous DMA (10 mL) and Hunig's base (92 μL) were added. L, 528 μmol, 3 equiv.) was then pipetted into a round-bottom flask. (1.3 equiv.) was added in one portion. The reaction mixture was kept at room temperature. After 30 min, TLC (CHCN:H2O = 85 :15) Analysis showed that the reaction was complete. AEDI-OH (68 mg, 352 μmol, 2 equivalents) was added to the reaction mixture and stirred at room temperature for 3 hours. TLC (CH3CN:H2O = 8:2) showed activity of The analytical HPLC also showed complete consumption of the ester, with red spots appearing below the active ester. , indicating complete consumption of the active ester and formation of the product. The reaction was run in TEAB buffer (0.1 M, 10 The volatile solvent was removed by evaporation under reduced pressure (HV) and purified on an Axia column. NR550S0-AEDI-OH was obtained as a result of the synthesis. Yield: 60%.

[0126] In a 25 mL round-bottom flask, NR550S0-AEDI-OH (10 μmol) was dissolved in DMF (anhydrous, 5 mL) and evaporated. The mixture was evaporated to dryness. The procedure was repeated three times. Anhydrous DMA (5 mL) and DMAP (1.8 mL, 15 μmol, 1.5 equiv.) were added. Amount) was then added to the round-bottom flask. DSC (5.2 mg, 20 μmol, 2 equiv.) was added in one portion. The mixture was kept at room temperature, and after 30 minutes, TLC (CH3CN:H2O=8:2) analysis indicated that the reaction was complete. Hunig's base (3.5 μL, 20 μmol) was pipetted into the reaction mixture. Next, a solution of pppA-LN3 (20 μmol, 2 equiv. in 0.5 mL of HO) and Et3N (5 μL) was added to the reaction mixture. The mixture was stirred at room temperature overnight. TLC (CH3CN:H2O = 8:2) showed complete consumption of the active ester. On the other hand, analytical HPLC also showed complete consumption of the active ester. The reaction was quenched with TEAB buffer (0.1 M, 10 mL) and DEAE Sep was added. The column was loaded onto a Hadex column (25g Biotage column). It eluted at ent.

[0127] A: 0.1 M TEAB buffer (10% CH3CN) B: 1 M TEAB buffer (10% CH3CN)

[0128] gradient: [Table 2]

[0129] The desired product was eluted from 45% to 100% of 1M TEAB buffer. The fragments containing the product were reassembled. The extracts were combined, evaporated and purified by HPLC (YCL column, 8 mL / min). Yield: 53%.

[0130] In summary, the present invention provides a nucleoside covalently attached to a fluorophore via a linker. or a nucleotide, wherein the linker is of Formula I or Formula II, or a combination of both. The term may refer to a nucleotide or a nucleotide, including a combination structure.

[0131] [ka] (In the formula, R 1 and R 2 each independently represents hydrogen or an optionally substituted C 1-6 Alkyl Selected from; R 3 is hydrogen, optionally substituted C 1-6 Alkyl, -NR 5 -C(=O)R 6 , or -NR 7 -C(=O)-O R 8 Selected from; R 4 is hydrogen or optionally substituted C 1-6 alkyl; R 5 and R 7 each independently represents hydrogen, optionally substituted C 1-6 Alkyl, substituted optionally substituted phenyl or optionally substituted C 7-12 aralkyl; R 6 and R 8 each independently represents an optionally substituted C 1-6 Alkyl, substituted optionally substituted phenyl, optionally substituted C 7-12 Aralkyl, optionally substituted C 3-7 S cycloalkyl, or optionally substituted 5-10 membered heteroaryl; [ka] wherein each methylene repeat unit is optionally substituted; X is selected from methylene (CH), oxygen (O), or sulfur (S); m is an integer from 0 to 20; n is an integer from 1 to 20, p indicates that the fluorophore-labeled nucleoside or nucleotide does not have the following structure: If not, it is an integer between 1 and 20: [ka]

[0132] In the case of some of the nucleosides or nucleotides described above, the structure of formula I may also be represented by formula Ia or represented by formula Ib:

[0133] [ka]

[0134] Furthermore, the structure of Formula II can be represented by Formula IIa, Formula IIb, Formula IIc, Formula IId, Formula IIe, or can also be expressed as the formula IIf.

[0135] [ka]

[0136] Specifically, the present invention provides a nucleoside covalently attached to a fluorophore via a linker. or a nucleotide, wherein the linker is of Formula I or Formula II, or a combination of both. The term may refer to a nucleoside or nucleotide, including any combination thereof.

[0137] [ka] (In the formula, R 1 is optionally substituted C 1-6 alkyl; R 2 is hydrogen or optionally substituted C 1-6 alkyl; R 3 may be substituted C 1-6 Alkyl, -NR 5 -C(=O)R 6 , or -NR 7 -C(=O)-OR 8 from Selected; R 4 is hydrogen or optionally substituted C 1-6 alkyl; R 5 and R 7 each independently represents hydrogen, optionally substituted C 1-6 Alkyl, substituted optionally substituted phenyl or optionally substituted C 7-12 aralkyl; R 6 and R 8 each independently represents an optionally substituted C 1-6 Alkyl, substituted optionally substituted phenyl, optionally substituted C 7-12 Aralkyl, optionally substituted C 3-7 S cycloalkyl, or optionally substituted 5-10 membered heteroaryl; [ka] wherein each methylene repeat unit is optionally substituted; X is selected from methylene (CH), oxygen (O), or sulfur (S); m is an integer from 0 to 20; n is an integer from 1 to 20, p is an integer from 1 to 20.

[0138] The structure of Formula I is also suitably represented by Formula Ia:

[0139] [ka]

[0140] Furthermore, the structure of Formula II can be represented by Formula IIb, Formula IIc, Formula IId, Formula IIe, or Formula IIf. It is also represented as

[0141] [ka]

Claims

1. A labeled nucleoside or nucleotide covalently attached to a fluorophore via a linker, wherein the linker comprises the structure of formula (IIc), and the nucleoside or nucleotide and the fluorophore are attached to the linker with the structure of formula (1): 【Chemical 1】 【Chemistry 2】 (In formula (1), Nu represents a nucleoside or a nucleotide, and the dye represents a fluorophore.)

2. A kit comprising the labeled nucleoside or labeled nucleotide of claim 1.

3. A reagent for modifying a nucleoside or nucleotide described in claim 1, comprising a fluorophore described in claim 1 and a linker described in claim 1.

4. 1. A method for detecting a labeled nucleoside or labeled nucleotide incorporated into a polynucleotide, comprising: (a) incorporating a labeled nucleoside or labeled nucleotide of claim 1 into a polynucleotide; (b) detecting a fluorescent signal from the labeled nucleoside or labeled nucleotide incorporated in step (a).

5. 5. The method of claim 4, further comprising the steps of providing a template nucleic acid strand and a partially hybridized nucleic acid strand, wherein step (a) incorporates into the hybridized strand at least one labeled nucleoside or labeled nucleotide that is complementary to a nucleoside or nucleotide at a corresponding position in the template strand, and step (b) indicates the identity of the complementary nucleoside or nucleotide in the template strand by identifying the base of the incorporated labeled nucleoside or labeled nucleotide.

6. 1. A method for sequencing a template nucleic acid molecule, comprising: incorporating one or more labeled nucleotides into a nucleic acid strand complementary to the template nucleic acid; determining the identity of a base present in one or more incorporated labeled nucleotides to determine the sequence of the template nucleic acid molecule; the identity of the base present in the one or more labeled nucleotides is determined by detecting a fluorescent signal produced by the labeled nucleotide; A method wherein at least one incorporated labeled nucleotide is a labeled nucleotide according to claim 1.

Citation Information

Patent Citations

  • Modified nucleotide linkers

    JP2020015722A