Fluorescent dyes containing bis-boron fused heterocycles and their use in sequencing

JP2024516191A5Pending Publication Date: 2025-05-27ILLUMINA CAMBRIDGE LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023565315
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-05-05
Filing Date
2022-05-02
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

Existing fluorescent dyes for nucleic acid sequencing face challenges such as difficulty in finding dyes with resolved absorption and emission spectra, poor photostability under high-power excitation, and incompatibility with sequencing reagents, limiting the efficiency and accuracy of multiplexed fluorescence detection.

Method used

Development of bis-boron fused heterocyclic dyes with improved chemical stability and tunable fluorescence properties, suitable for blue light excitation, which can be conjugated to nucleotides for enhanced fluorescence intensity and spectral distinction.

Benefits of technology

The bis-boron fused heterocyclic dyes provide strong fluorescence and chemical stability, enabling accurate and efficient nucleic acid sequencing with improved optical resolution and reduced sequencing errors, suitable for high-throughput methods like solid-phase sequencing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

This application relates to substituted dyes containing bis-boron fused heterocycles and their use as fluorescent labels. These compounds can be used as fluorescent labels for nucleotides in nucleic acid sequencing applications.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present disclosure relates to fluorescent dyes containing bis-boron fused heterocycles and their use as fluorescent labels of nucleotides in nucleic acid sequencing applications. [Background technology]

[0002] Non-radioactive detection of fluorescently labeled nucleic acids is an important technique in molecular biology. Many of the procedures used in recombinant DNA technology have previously been developed, e.g. 32 These methods have relied on the use of nucleotides or polynucleotides radioactively labeled with P. Radioactive compounds allow for sensitive detection of nucleic acids and other molecules of interest. However, the use of radioisotopes has significant limitations, such as their cost, limited shelf life, poor sensitivity, and more importantly, safety considerations. Eliminating the need for radioactive labeling reduces both safety risks and environmental impacts, as well as costs associated with, for example, reagent disposal. Methods suitable for non-radioactive fluorescence detection include, by way of non-limiting example, automated DNA sequencing, hybridization methods, real-time detection of polymerase chain reaction products, and immunoassays.

[0003] For many applications, it is desirable to use multiple spectrally distinguishable fluorescent labels to achieve independent detection of multiple spatially overlapping analytes. Such multiplexing methods can reduce the number of reaction vessels, simplifying experimental protocols and facilitating the generation of application-specific reagent kits. For example, in multicolor automated DNA sequencing systems, multiplexing fluorescent detection allows the analysis of multiple nucleotide bases in a single electrophoretic lane, thereby increasing throughput compared to single-color methods and reducing uncertainties associated with lane-to-lane electrophoretic mobility variations.

[0004] However, multiplexed fluorescence detection can be problematic, and there are several important factors that constrain the selection of suitable fluorescent labels. First, it can be difficult to find dye compounds with substantially resolved absorption and emission spectra for a given application. In addition, when several fluorescent dyes are used together, it can be complicated to generate fluorescent signals in distinguishable spectral regions by simultaneous excitation, since the absorption bands of the dyes are usually widely separated, so it is difficult to achieve comparable fluorescence excitation efficiency even for two dyes. Many excitation methods use high-power light sources such as lasers, and therefore the dyes must be photostable enough to withstand such excitation. A final consideration, particularly important for molecular biology methods, is the extent to which the fluorescent dye must be compatible with reagent chemicals, such as DNA synthesis solvents and reagents, buffers, polymerase enzymes, and ligase enzymes.

[0005] As sequencing technology advances, the need for additional fluorescent dye compounds, their nucleic acid conjugates, and multiple dye sets that meet all the above constraints and are particularly suitable for high-throughput molecular methods such as solid-phase sequencing is developing.

[0006] Fluorescent dye molecules with improved fluorescence properties, such as suitable fluorescence intensity, shape, and wavelength maximum of the fluorescence band, can improve the speed and accuracy of nucleic acid sequencing. Strong fluorescence signals are especially important when measurements are performed in aqueous biological buffers and at high temperatures, since the fluorescence intensity of most organic dyes is significantly lower under such conditions. Furthermore, the nature of the base to which the dye is attached also affects the fluorescence maximum, fluorescence intensity, and other spectral properties of the dye. The sequence-specific interaction between the nucleic acid base and the fluorescent dye can be tailored by the specific design of the fluorescent dye. Optimizing the structure of the fluorescent dye can improve the efficiency of nucleotide incorporation, reduce the level of sequencing errors, and reduce the amount of reagents used in nucleic acid sequencing, thereby reducing the cost of nucleic acid sequencing.

[0007] Several optical and technological developments have already significantly improved image quality, but ultimately were limited by insufficient optical resolution. In general, the optical resolution of optical microscopes is limited to distinguishing objects separated by a distance of approximately half the wavelength of the light used. In practical terms, only objects that are very far from each other (at least 200-350 nm) can be distinguished by optical microscopy. One way to improve image resolution and increase the number of objects that can be distinguished per unit surface area is to use excitation light with a shorter wavelength. For example, with the same optical system, if the wavelength is shortened by Δλ about 100 nm, the resolution will be better (about Δ50 nm / (about 15%)), images with less distortion will be recorded, and the density of objects on the recognizable area will increase by about 35%.

[0008] Certain nucleic acid sequencing methods use laser light to excite and detect dye-labeled nucleotides. These instruments use longer wavelength light, such as a red laser, with an appropriate dye excitable at 660 nm. To detect more densely packed nucleic acid sequencing clusters while maintaining useful resolution, a shorter wavelength blue light source (450-460 nm) can be used. In this case, the optical resolution would not be limited by the emission wavelength of the long wavelength red fluorescent dye, but rather by the emission of the dye excitable by the next longer wavelength light source, e.g., a green laser (532 nm). Thus, there is a need for blue dye labels for use in fluorescence detection in sequencing applications.

[0009] Although blue dye chemistry and related laser technology have been improved, suitable commercially available blue dyes with strong fluorescence for nucleotide labeling are still very rare. However, certain blue dyes, especially coumarin dyes, are not stable for long periods in aqueous environments. For example, in basic conditions, coumarin dyes can be easily attacked by nucleophiles, thus resulting in disturbance or degradation of the dye. Boron-containing fluorescent dyes such as BODIPY, BOPHY, BOPPY, BOPYPY, BOAHY and BOPAHY have been reported in several scientific and patent publications. For example, J Am Chem Soc 2014, 136(15), 5623-5626, Organic Letters 2014, 16(11), 3048-3051, Organic Letters 2018, 20(15), 4462-4466, Chinese Chemical Letters 2019, 30:2271-2273, Organic Letters 2020, 22(12), 4588-4592, Chem Communications 2020, 56(43), 5791-5794, Chemistry A Eur J 2020, 26(4), 863-872, International Publication No. 2015 / 77427 and Chinese Patent No. 108516985(A). However, it remains challenging to design new boron-containing dyes with suitable adsorption, good chemical stability and Stokes shift as nucleic acid labels for sequencing applications. Summary of the Invention

[0010] Described herein is a new class of dyes containing bis-boron fused heterocycles that have improved chemical stability and strong fluorescence under blue light excitation (e.g., blue LED or laser, e.g., about 450 nm to about 460 nm). These dyes also have highly tunable absorption and emission properties that are suitable for nucleic acid labeling.

[0011] One aspect of the disclosure relates to a compound of formula (I), or a salt or mesomeric form thereof:

[0012] [ka] R 1 , R 2 , R 3 and R 4 each of which is independently H, unsubstituted or substituted C1-C6 alkyl, C1-C6 alkoxy, C2-C6 alkenyl, C2-C6 alkynyl, C1-C6 haloalkyl, C1-C6 haloalkoxy, C1-C6 hydroxyalkyl, (C1-C6 alkoxy)(C1-C6 alkyl), unsubstituted or substituted amino, halo, cyano, hydroxy, nitro, sulfonyl, sulfino, sulfo, sulfonate, S-sulfonamido, N-sulfonamido, unsubstituted or substituted C3-C 10 Carbocyclyl, unsubstituted or substituted C6-C 10 aryl, unsubstituted or substituted 5- to 10-membered heteroaryl, or unsubstituted or substituted 3- to 10-membered heterocyclyl; Each R a , R b , R c and R d are independently halo, cyano, C1-C6 alkyl, C1-C6 haloalkyl, C1-C6 alkoxy, C1-C6 haloalkoxy, C6-C 10 Aryl, C6-C 10 Aryloxy or -OC(=O)R 5 and Alternatively, R a and R b Both are -OC(=O)R 5 If, then, two R 5 together with the atoms to which they are attached form an unsubstituted or substituted 6- to 10-membered heterocyclyl, R c and R d Both are -OC(=O)R 5 If so, then two R 5 together with the atom to which they are attached form an unsubstituted or substituted 6- to 10-membered heterocyclyl; R 5 is unsubstituted or substituted C1-C6 alkyl; Ring A is one or more R 6 is a 6-10 membered heteroaryl optionally substituted with Each R 6 are independently unsubstituted or substituted C1-C6 alkyl, C1-C6 alkoxy, C2-C6 alkenyl, C2-C6 alkynyl, C1-C6 haloalkyl, C1-C6 haloalkoxy, C1-C6 hydroxyalkyl, (C1-C6 alkoxy)(C1-C6 alkyl), -NR 7 R 8 , halo, cyano, carboxy, hydroxy, nitro, sulfonyl, sulfino, sulfo, sulfonate, S-sulfonamido, N-sulfonamido, unsubstituted or substituted C3-C 10 Carbocyclyl, unsubstituted or substituted C6-C 10 aryl, unsubstituted or substituted 5- to 10-membered heteroaryl, or unsubstituted or substituted 3- to 10-membered heterocyclyl; Each R 7 and R 8 is independently H, unsubstituted or substituted C1-C6 alkyl, or R 7 and R 8 together with the nitrogen atom to which they are attached form an unsubstituted or substituted 3- to 10-membered heterocyclyl; However, R 1 , R 2 , R 3 , R 4 and at least one of rings A contains a carboxyl group.

[0013] In some embodiments, the compound of formula (I) may also have the structure of formula (Ia) or (Ib):

[0014] [ka] or a salt or mesomeric form thereof, where m is 0, 1, 2, or 3. In further embodiments, the compound may have the structure of formula (Ic), (Id) or (Ie):

[0015] [ka] or a salt or mesomeric form thereof.

[0016] In some aspects, the compounds of the present disclosure are labeled or conjugated with a substrate moiety, such as, for example, a nucleoside, a nucleotide, a polynucleotide, a polypeptide, a carbohydrate, a ligand, a particle, a cell, a semi-solid surface (e.g., a gel), or a solid surface. Labeling or conjugation can be performed via a carboxyl group, which can react with an amino or hydroxy group on a moiety (such as a nucleotide) or linker attached thereto to form an amide or ester, using methods known in the art.

[0017] Another aspect of the present disclosure relates to dye compounds that include a linker group to allow for covalent attachment to, for example, a substrate moiety. The attachment can be at any position on the dye. In some embodiments, the attachment can be at any position on the dye, for example, at any position on the substrate moiety ... 1 , R 2 , R 3 , R 4 and ring A.

[0018] A further aspect of the disclosure provides a labeled nucleoside or nucleotide compound defined by the formula: NL-dye wherein N is a nucleoside or nucleotide; L is an optional linker moiety, The dye is a moiety of a fluorescent compound of formula (I) according to the present disclosure, where a functional group (e.g., a carboxyl group) of the compound of formula (I) (e.g., (Ia), (Ib), (Ic), (Id), or (Ie)) reacts with an amino or hydroxyl group of a linker moiety or a nucleoside / nucleotide to form a covalent bond.

[0019] Some additional aspects of the present disclosure pertain to oligonucleotides or polynucleotides labeled with a compound of Formula (I) (e.g., (Ia), (Ib), (Ic), (Id), or (Ie)).

[0020] Some additional aspects of the present disclosure relate to kits that contain dye compounds (free or labeled form) that can be used in various immunological assays, oligonucleotide or nucleic acid labeling, or DNA sequencing by synthesis. In yet another aspect, the present disclosure provides kits that contain dye "sets" that are particularly suitable for cycles of sequencing by synthesis on an automated instrument platform. In some aspects, the kits contain one or more nucleotides, at least one of which is a labeled nucleotide as described herein.

[0021] A further aspect of the present disclosure is a method for determining the sequence of a plurality of target polynucleotides, comprising: (a) contacting a solid support with a solution containing sequencing primers under hybridization conditions, where the solid support comprises a plurality of different target polynucleotides immobilized thereon, and the sequencing primers are complementary to at least a portion of the target polynucleotides; (b) contacting the solid support with an aqueous solution comprising a DNA polymerase and one or more of four different types of nucleotides under conditions suitable for DNA polymerase-mediated primer extension, where at least one of the types of nucleotides is a labeled nucleotide as described herein; (c) incorporating one type of nucleotide into a sequencing primer to generate an extended copy polynucleotide; (d) performing one or more fluorescence measurements of the extended copy polynucleotide to determine the identity of the incorporated nucleotide. [Brief description of the drawings]

[0022] [Figure 1] 1 shows the emission spectra of ffA-spA-I-4 and ffC labeled with reference dye A when excited by blue light (450 nm). [Diagram 2]The percentage of fluorescent signal remaining as a function of time for dyes I-1 and I-3 compared to fully C-labeled with reference dye A under the same conditions is shown. [Diagram 3] The percent phasing of the integration mix containing ffA-spA-I-4 is shown compared to two reference integration mixes on the MiSeq™. [Figure 4A] Scatter plot obtained for integration mix containing ffA-spA-I-3 at cycle 26 using blue light at 1× and 5× doses. [Figure 4B] Scatter plot obtained for integration mix containing ffA-spA-I-3 at cycle 26 using blue light at 1× and 5× doses. [Figure 4C] Scatter plot obtained for integration mix containing ffA-spA-I-4 at cycle 26 using blue light at 1× and 5× doses. [Figure 4D] Scatter plot obtained for integration mix containing ffA-spA-I-4 at cycle 26 using blue light at 1× and 5× doses. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0023] Embodiments of the present disclosure relate to dyes containing bis-boron fused heterocycles with enhanced fluorescence intensity, tunable Stokes shift and improved chemical stability. In some embodiments, the Stokes shift of the dyes described herein ranges from about 15 nm to 50 nm (e.g., about 20 nm). The bis-boron-containing dyes described herein can be used in Illumina sequencing platforms, such as MiSeq™, with two-channel detection (green light excitation and blue light excitation).

[0024] definition The section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described.

[0025] It should be noted that, as used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless expressly and unambiguously limited to one reference. It will be apparent to those skilled in the art that various modifications and variations can be made to the various embodiments described herein without departing from the spirit or scope of the present teachings. Accordingly, it is intended that the various embodiments described herein encompass other modifications and variations that are within the scope of the appended claims and their equivalents.

[0026] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. The use of the term "including" as well as other forms such as "include", "includes" and "included" is not limiting. Furthermore, the use of the term "having" as well as other forms such as "have", "has" and "had" is not limiting. As used herein, whether in a transitional phrase or in the body of a claim, the terms "comprise" and "comprising" should be interpreted as having an open-ended meaning. That is, the above terms should be interpreted as synonymous with the phrase "having at least" or "comprising at least". For example, when used in the context of a process, the term "comprising" means that the process includes at least the recited steps, but may include additional steps. When used in the context of a compound, composition, or device, the term "comprising" means that the compound, composition, or device includes at least the recited features or components, but may include additional features or components.

[0027] As used herein, common organic abbreviations are defined as follows: ℃ Celsius temperature dATP deoxyadenosine triphosphate dCTP deoxycytidine triphosphate dGTP Deoxyguanosine triphosphate dTTP deoxythymidine triphosphate ddNTP Dideoxynucleotide triphosphate ffA fully functionalized A nucleotide ffC fully functionalized C nucleotide ffG fully functionalized G nucleotide ffN fully functionalized nucleotides ffT Fully functionalized T nucleotide h time RT room temperature SBS Sequencing by Synthesis USM Universal scan mix

[0028] As used herein, the term "array" refers to a collection of different probe molecules that are attached to one or more substrates such that the different probe molecules can be distinguished from one another according to their relative positions. An array can include probe molecules that are each located at a different addressable position on a substrate. Alternatively, or in addition, an array can include separate substrates, each carrying a different probe molecule, and the different probe molecules can be identified according to the position of the substrate on the surface to which the substrate is attached, or according to the position of the substrate in a liquid. Exemplary arrays in which separate substrates are located on a surface include, but are not limited to, those that include beads in wells, as described, for example, in U.S. Pat. No. 6,355,431 (B1), U.S. Patent Application Publication No. 2002 / 0102578, and PCT Publication No. 00 / 63437. An exemplary format that can be used in the present invention to distinguish beads in a liquid array uses a microfluidic device, such as, for example, a fluorescence activated cell sorter (FACS), and is described, for example, in U.S. Pat. No. 6,524,793. Further examples of arrays that can be used in the present invention include those described in U.S. Patent Nos. 5,429,807, 5,436,327, 5,561,071, 5,583,211, 5,658,734, 5,837,858, 5,874,219, 5,919,523, 6,136,269, 6,287,768, 6,287,776, 6,288,220, 6,288,230, 6,288,240, 6,288,250, 6,288,260, 6,288,270, 6,288,282, 6,288,290, 6,288,300, 6,288,310, 6,288,320, 6,288,330, 6,288,340, 6,288,350, 6,288,360, 6,288,370, 6,288,382, 6,288,390, 6,288,400, 6,288,410, 6,288,520, 6,288,530, 6,288,540, 6,288,550, 6,288,600, 6,288,710, 6,288,720, 6,288,730, 6,288,740, 6,288,750, 6,288,760, 6,288,770, 6,288,780, 6 Nos. 6,297,006, 6,291,193, 6,346,413, 6,416,949, 6,482,591, 6,514,751, and 6,610,482, as well as WO 93 / 17126, WO 95 / 11995, WO 95 / 35505, European Patent Nos. 742287, and 799897.

[0029] As used herein, the terms "covalently attached" or "covalently bonded" refer to the formation of a chemical bond characterized by the sharing of electron pairs between atoms. For example, a covalently attached polymer coating refers to a polymer coating that forms a chemical bond with the functionalized surface of a substrate, as compared to attachment to the surface by other means, such as adhesion or electrostatic interactions. It will be understood that a polymer covalently attached to a surface can be attached by means in addition to covalent bonds.

[0030] As used herein, the term "halogen" or "halo" refers to any one of Group 7 of the Periodic Table of the Elements, e.g., fluorine, chlorine, bromine, or iodine, with fluorine and chlorine being preferred.

[0031] As used herein, "C" is a formula in which "a" and "b" are integers. a -C b " refers to the number of carbon atoms in an alkyl, alkenyl, or alkynyl group, or the number of ring atoms in a cycloalkyl or aryl group. That is, the alkyl, alkenyl, alkynyl, cycloalkyl ring, and aryl ring can contain "a" to "b" carbon atoms at both ends. For example, a "C1-C4 alkyl" group refers to all alkyl groups having 1 to 4 carbons, i.e., CH3-, CH3CH2-, CH3CH2CH2-, (CH3)2CH-, CH3CH2CH2CH2-, CH3CH2CH(CH3)-, and (CH3)3C-, and a C3-C4 cycloalkyl group refers to all cycloalkyl groups having 3 to 4 carbon atoms, i.e., cyclopropyl and cyclobutyl. Similarly, a "4- to 6-membered heterocyclyl" group refers to all heterocyclyl groups having 4 to 6 total ring atoms, such as azetidine, oxetane, oxazoline, pyrrolidine, piperidine, piperazine, morpholine, and the like. When "a" and "b" are not specified with respect to an alkyl, alkenyl, alkynyl, cycloalkyl, or aryl group, the broadest range described in these definitions should be assumed. 1-"C6" includes C1, C2, C3, C4, C5, and C6, as well as ranges defined by either of the two numbers. For example, C1-C6 alkyl includes C1, C2, C3, C4, C5, and C6 alkyl, C2-C6 alkyl, C1-C3 alkyl, etc. Similarly, C2-C6 alkenyl includes C2, C3, C4, C5, and C6 alkenyl, C2-C5 alkenyl, C3-C4 alkenyl, etc., and C2-C6 alkynyl includes C2, C3, C4, C5, and C6 alkynyl, C2-C5 alkynyl, C3-C4 alkynyl, etc. C3-C8 cycloalkyl includes hydrocarbon rings containing 3, 4, 5, 6, 7, and 8 carbon atoms, respectively, or ranges defined by either of the two numbers, such as C3-C7 cycloalkyl or C5-C6 cycloalkyl.

[0032] As used herein, "alkyl" refers to a straight or branched hydrocarbon chain that is fully saturated (ie, contains no double or triple bonds). An alkyl group may have 1 to 20 carbon atoms (wherever indicated herein, a numerical range such as "1 to 20" refers to each integer within the given range. For example, "1 to 20 carbon atoms" means that the alkyl group may consist of 1 carbon atom, 2 carbon atoms, 3 carbon atoms, etc., up to 20 carbon atoms, but this definition also covers occurrences of the term "alkyl" (where no numerical range is specified). The alkyl group may also be a medium sized alkyl having 1 to 9 carbon atoms. The alkyl group may also be a lower alkyl having 1 to 6 carbon atoms. By way of example only, "C1-6 alkyl" or "C1-C6 alkyl" indicates that there are 1 to 6 carbon atoms in the alkyl chain, i.e., the alkyl chain is selected from the group consisting of methyl, ethyl, propyl, isopropyl, n-butyl, iso-butyl, sec-butyl, and t-butyl. Typical alkyl groups include, but are not limited to, methyl, ethyl, propyl, isopropyl, butyl, isobutyl, tertiary butyl, pentyl, hexyl, and the like.

[0033] As used herein, "alkoxy" refers to "C1-9 alkoxy" or "C 1- "C9 alkoxy" refers to the formula -OR where R is alkyl as defined above, including, but not limited to, methoxy, ethoxy, n-propoxy, 1-methylethoxy (isopropoxy), n-butoxy, isobutoxy, sec-butoxy, and tert-butoxy.

[0034] As used herein, "-OAc" or "-O-acyl" refers to acetyloxy, which has the structure -OC(=O)CH3.

[0035] As used herein, "alkenyl" refers to a straight or branched hydrocarbon chain containing one or more double bonds. Alkenyl groups can have from 2 to 20 carbon atoms, although this definition also covers occurrences of the term "alkenyl" where no numerical range is specified. Alkenyl groups can also be medium sized alkenyls having from 2 to 9 carbon atoms. Alkenyl groups can also be lower alkenyls having from 2 to 6 carbon atoms. By way of example only, "C 2- C6 alkenyl" or "C 2-6 "Alkenyl" indicates that there are 2 to 6 carbon atoms in the alkenyl chain, i.e., the alkenyl chain is selected from the group consisting of ethenyl, propen-1-yl, propen-2-yl, propen-3-yl, buten-1-yl, buten-2-yl, buten-3-yl, buten-4-yl, 1-methyl-propen-1-yl, 2-methyl-propen-1-yl, 1-ethyl-ethen-1-yl, 2-methyl-propen-3-yl, buta-1,3-dienyl, buta-1,2-dienyl, and buta-1,2-dien-4-yl. Exemplary alkenyl groups include, but are not limited to, ethenyl, propenyl, butenyl, pentenyl, and hexenyl.

[0036] As used herein, "alkynyl" refers to a straight or branched hydrocarbon chain containing one or more triple bonds. Alkynyl groups can have from 2 to 20 carbon atoms, although this definition also covers occurrences of the term "alkynyl" where no numerical range is specified. Alkynyl groups can also be medium sized alkynyls having from 2 to 9 carbon atoms. Alkynyl groups can also be lower alkynyls having from 2 to 6 carbon atoms. By way of example only, "C 2-6 alkynyl" or "C 2- "C6 alkynyl" indicates that there are 2 to 6 carbon atoms in the alkynyl chain, i.e., the alkynyl chain is selected from the group consisting of ethynyl, propyn-1-yl, propyn-2-yl, butyn-1-yl, butyn-3-yl, butyn-4-yl, and 2-butynyl. Exemplary alkynyl groups include, but are not limited to, ethynyl, propynyl, butynyl, pentynyl, and hexynyl.

[0037] The term "aromatic" refers to a ring or ring system having a conjugated pi-electron system and including both carbocyclic aromatic (e.g., phenyl) and heterocyclic aromatic (e.g., pyridine) groups. The term includes monocyclic or fused polycyclic (i.e., rings which share adjacent pairs of atoms) groups, provided that the entire ring system is aromatic.

[0038] As used herein, "aryl" refers to an aromatic ring or ring system (i.e., two or more fused rings sharing two adjacent carbon atoms) that contains only carbon in the ring backbone. When aryl is a ring system, all rings in the system are aromatic rings. Aryl groups can have from 6 to 18 carbon atoms, although this definition also covers occurrences of the term "aryl" where no numerical range is specified. In some embodiments, aryl groups have from 6 to 10 carbon atoms. An aryl group is defined as "C 6- C 10 Aryl, C6 or C 10 Examples of aryl groups include, but are not limited to, phenyl, naphthyl, azulenyl, and anthracenyl.

[0039] "Aralkyl" or "arylalkyl" means "C 7-14 Aryl groups bonded as a substituent via an alkylene group, such as "aralkyl," include, but are not limited to, benzyl, 2-phenylethyl, 3-phenylpropyl, and naphthylalkyl. In some cases, the alkylene group can be joined to a lower alkylene group (i.e., 1-6 alkylene group).

[0040] As used herein, "heteroaryl" refers to an aromatic ring or ring system (i.e., two or more fused rings sharing two adjacent atoms) containing one or more heteroatoms, i.e., elements other than carbon, including but not limited to nitrogen, oxygen, and sulfur, in the ring backbone. When heteroaryl is a ring system, all rings in the system are aromatic rings. Heteroaryl groups can have 5 to 18 ring members (i.e., the number of atoms that make up the ring backbone, including carbon atoms and heteroatoms), although this definition also covers occurrences of the term "heteroaryl" where no numerical range is specified. In some embodiments, heteroaryl groups have 5 to 10 ring members or 5 to 7 ring members. Heteroaryl groups can be designated as "5- to 7-membered heteroaryl," "5- to 10-membered heteroaryl," or similar designations. Examples of heteroaryl rings include, but are not limited to, furyl, thienyl, phthalazinyl, pyrrolyl, oxazolyl, thiazolyl, imidazolyl, pyrazolyl, isoxazolyl, isothiazolyl, triazolyl, thiadiazolyl, pyridinyl, pyridazinyl, pyrimidinyl, pyrazinyl, triazinyl, quinolinyl, isoquinolinyl, benzimidazolyl, benzoxazolyl, benzothiazolyl, indolyl, isoindolyl, and benzothienyl.

[0041] A "heteroaralkyl" or "heteroarylalkyl" is a heteroaryl group bonded as a substituent via an alkylene group. Examples include, but are not limited to, 2-thienylmethyl, 3-thienylmethyl, furylmethyl, thienylethyl, pyrrolylalkyl, pyridylalkyl, isoxazolylalkyl, and imidazolylalkyl. In some cases, an alkylene group may be joined to a lower alkylene group (i.e., C 1-6 alkylene group).

[0042] As used herein, "carbocyclyl" refers to a non-aromatic ring or ring system containing only carbon atoms in the ring system backbone. When a carbocyclyl is a ring system, two or more rings may be joined together in a fused, bridged, or spiro-connected manner. A carbocyclyl may have any degree of saturation, provided that at least one ring in the ring system is not aromatic. Thus, carbocyclyl includes cycloalkyl, cycloalkenyl, and cycloalkynyl. A carbocyclyl group may have 3 to 20 carbon atoms, although this definition also covers occurrences of the term "carbocyclyl" where no numerical range is specified. A carbocyclyl group may also be a medium-sized carbocyclyl having 3 to 10 carbon atoms. A carbocyclyl group may also be a carbocyclyl having 3 to 6 carbon atoms. A carbocyclyl group is defined as a "C 3-6 Carbocyclyl, C 3- C6 carbocyclyl" or similar designations. Examples of carbocyclyl rings include, but are not limited to, cyclopropyl, cyclobutyl, cyclopentyl, cyclohexyl, cyclohexenyl, 2,3-dihydro-indene, bicycle[2.2.2]octanyl, adamantyl, and spiro[4.4]nonanyl.

[0043] As used herein, "cycloalkyl" means a fully saturated carbocyclyl ring or ring system. Examples include cyclopropyl, cyclobutyl, cyclopentyl, and cyclohexyl.

[0044] As used herein, "heterocyclyl" refers to a non-aromatic ring or ring system containing at least one heteroatom in the ring backbone. Heterocyclyls may be bonded together in fused, bridged, or spiro-linked fashions. Heterocyclyls may have any degree of saturation, provided that at least one ring in the ring system is not aromatic. The heteroatoms may be present in either the non-aromatic or aromatic rings in the ring system. Heterocyclyl groups may have 3-20 ring members (i.e., the number of atoms that make up the ring backbone, including carbon atoms and heteroatoms), although this definition also covers occurrences of the term "heterocyclyl" where no numerical range is specified. Heterocyclyl groups may also be medium-sized heterocyclyls having 3-10 ring members. Heterocyclyl groups may also be heterocyclyls having 3-6 ring members. Heterocyclyl groups may also be designated as "3-6 membered heterocyclyl" or similar designations. In preferred 6-membered monocyclic heterocyclyls, the heteroatoms are selected from one to three of O, N or S. In preferred 5-membered monocyclic heterocyclyls, the heteroatoms are selected from one or two heteroatoms selected from O, N or S.Examples of heterocyclyl rings include azepinyl, acridinyl, carbazolyl, cinnolinyl, dioxolanyl, imidazolinyl, imidazolidinyl, morpholinyl, oxiranyl, oxepanyl, thiapanyl, piperidinyl, piperazinyl, dioxapiperazinyl, pyrrolidinyl, pyrrolidionyl, pyrrolidionyl, 4-piperidonyl, pyrazolinyl, pyrazolidinyl, 1,3-dioxinyl, 1,3-dioxanyl, 1,4-dioxinyl, 1,4-dioxanyl, 1,3-oxathinyl, 1,4-oxathinyl, 1,4-oxathiyl, 2H-1,2-oxazinyl, trioxanyl, hexahydro-1,3, Examples include, but are not limited to, 5-triazinyl, 1,3-dioxolyl, 1,3-dioxolanyl, 1,3-dithiolyl, 1,3-dithiolanyl, isoxazolinyl, isoxazolidinyl, oxazolinyl, oxazolidinyl, oxazolidinonyl, thiazolinyl, thiazolidinyl, 1,3-oxathiolanyl, indolinyl, isoindolinyl, tetrahydrofuranyl, tetrahydropyranyl, tetrahydrothiophenyl, tetrahydrothiopyranyl, tetrahydro-1,4-thiazinyl, thiamorpholinyl, dihydrobenzofuranyl, benzimidazolidinyl, and tetrahydroquinoline.

[0045] As used herein, "alkoxyalkyl" or "(alkoxy)alkyl" refers to any alkyl group selected from the group consisting of C 2- C alkoxyalkyl, or (C-C alkoxy)C-C alkyl, such as -(CH) 1-3 Refers to an alkoxy group bonded via an alkylene group, such as -OCH3.

[0046] An "O-carboxy" group refers to a "-OC(=O)R" group, where R is hydrogen, as well as C, as defined herein. 1-6 Alkyl, C 2-6 Alkenyl, C 2-6 Alkynyl, C 3-7 Carbocyclyl, C 6-10 It is selected from aryl, 5- to 10-membered heteroaryl, and 3- to 10-membered heterocyclyl.

[0047] A "C-carboxy" group refers to a "-C(=O)OR" group, where R is hydrogen and C(=O) is an alkyl group, as defined herein. 1-6 Alkyl and C 2-6 Alkenyl and C 2-6 Alkynyl and C 3-7 Carbocyclyl and C 6-10 It is selected from the group consisting of aryl, 5-10 membered heteroaryl, and 3-10 membered heterocyclyl. Non-limiting examples include carboxyl (i.e., -C(=O)OH).

[0048] A "sulfonyl" group refers to a "-SO2R" group, where R is hydrogen, as defined herein, as well as C 1-6 Alkyl, C 2-6 Alkenyl, C 2-6 Alkynyl, C 3-7 Carbocyclyl, C 6-10 It is selected from aryl, 5- to 10-membered heteroaryl, and 3- to 10-membered heterocyclyl.

[0049] A "sulfino" group refers to a "-S(=O)OH" group.

[0050] A "sulfo" group refers to a "-S(=O)2OH" or a "-SO3H" group.

[0051] The "sulfonate" group is "-SO3 - " group.

[0052] The "sulfate" group is "-SO4 - " group.

[0053] The "S-sulfonamide" group is "-SONR A R B " group, R A and R B is hydrogen, C as defined herein 1-6 Alkyl, C 2-6 Alkenyl, C 2-6 Alkynyl, C 3-7 Carbocyclyl, C 6-10Each is independently selected from aryl, 5- to 10-membered heteroaryl, and 3- to 10-membered heterocyclyl.

[0054] An “N-sulfonamide” group is defined as “-N(R A )SO2R B " group, where R A and R.R. b is hydrogen, as defined herein, C 1-6 Alkyl, C 2-6 Alkenyl, C 2-6 Alkynyl, C 3~7 Carbocyclyl, C 6-10 Each is independently selected from aryl, 5- to 10-membered heteroaryl, and 3- to 10-membered heterocyclyl.

[0055] A "C-amide" group is defined as "-C(=O)NR A R B " group, where R A and R B is hydrogen, as defined herein, C 1-6 Alkyl, C 2-6 Alkenyl, C 2-6 Alkynyl, C 3-7 Carbocyclyl, C 6-10 Each is independently selected from aryl, 5- to 10-membered heteroaryl, and 3- to 10-membered heterocyclyl.

[0056] An "N-amide" group is defined as "-N(R A )C(=O)R B In the formula, R A and R B is hydrogen, as defined herein, C 1-6 Alkyl, C 2-6 Alkenyl, C 2-6 Alkynyl, C 3-7 Carbocyclyl, C 6-10 Each is independently selected from aryl, 5- to 10-membered heteroaryl, and 3- to 10-membered heterocyclyl.

[0057] The "amino" group is "-NR A R B " In the formula, RA and R B is hydrogen, as defined herein, C 1-6 Alkyl, C 2-6 Alkenyl, C 2-6 Alkynyl, C 3-7 Carbocyclyl, C 6-10 Each independently is selected from aryl, 5- to 10-membered heteroaryl, and 3- to 10-membered heterocyclyl. Non-limiting examples include free amino (i.e., -NH2).

[0058] An "aminoalkyl" group refers to an amino group linked via an alkylene group.

[0059] An "alkoxyalkyl" group refers to an alkoxy group linked via an alkylene group, for example, "C 2- C8 alkoxyalkyl.

[0060] When a group is described as "optionally substituted", it may be unsubstituted or substituted. Similarly, when a group is described as "substituted", the substituents may be selected from one or more of the indicated substituents. As used herein, a substituent is derived from an unsubstituted parent group in which one or more hydrogen atoms have been exchanged for another atom or group. Unless otherwise stated, when a group is considered to be "substituted", it is meant that the group is selected from C1-C6 alkyl, C1-C6 alkenyl, C1-C6 alkynyl, C3-C7 carbocyclyl (optionally substituted with halo, C1-C6 alkyl, C1-C6 alkoxy, C1-C6 haloalkyl, and C1-C6 haloalkoxy), C3-C7-carbocyclyl-C1-C6-alkyl (optionally substituted with halo, C1-C6 alkyl, C1-C6 alkoxy, C1-C6 haloalkyl, and C1-C6 haloalkoxy), C3-C7-carbocycyl-C1-C6-alkyl (optionally substituted with halo, C1-C6 alkyl, C1-C6 alkoxy, C1-C6 haloalkyl, and C1-C6 haloalkoxy). haloalkyl, and C1-C6 haloalkoxy), 3-10 membered heterocyclyl (optionally substituted with halo, C1-C6 alkyl, C1-C6 alkoxy, C1-C6 haloalkyl, and C1-C6 haloalkoxy), 3-10 membered heterocyclyl-C1-C6-alkyl (optionally substituted with halo, C1-C6 alkyl, C1-C6 alkoxy, C1-C6 haloalkyl, and C1-C6 haloalkoxy), aryl (optionally substituted with aryl(C1-C6)alkyl (optionally substituted with halo, C1-C6 alkyl, C1-C6 alkoxy, C1-C6 haloalkyl, and C1-C6 haloalkoxy); aryl(C1-C6)alkyl (optionally substituted with halo, C1-C6 alkyl, C1-C6 alkoxy, C1-C6 haloalkyl, and C1-C6 haloalkoxy); 5-10 membered heteroaryl (optionally substituted with halo, C1-C6 alkyl, C1-C6 alkoxy, C1-C6 haloalkyl, and C1-C6 haloalkoxy); aryloxy, -C1-C6 alkyl, -C1-C6 alkoxy, -C1-C6 haloalkyl, and -C1-C6 haloalkoxy; 5-10 membered heteroaryl(C1-C6)alkyl (optionally substituted with halo, C1-C6 alkyl, C1-C6 alkoxy, C1-C6 haloalkyl, and C1-C6 haloalkoxy); halo, -CN, hydroxy, C1-C6 alkoxy, C1-C6 alkoxy(C1-C6)alkyl (i.e., ether), aryloxy, sulfhydryl (mercapto), halo(C1-C6)alkyl (e.g., -CF3);It means substituted with one or more substituents independently selected from halo(C1-C6)alkoxy (e.g., -OCF3), C1-C6 alkylthio, arylthio, amino, amino(C1-C6)alkyl, nitro, O-carbamyl, N-carbamyl, O-thiocarbamyl, N-thiocarbamyl, C-amido, N-amido, S-sulfonamido, N-sulfonamido, C-carboxy, O-carboxy, acyl, cyanate, isocyanate, thiocyanate, isothiocyanate, sulfinyl, sulfonyl, -SO3H, sulfonate, sulfate, sulfino, -OSO2C1-C4 alkyl, and oxo (=O). When a group is described as being "optionally substituted," it means that, if the group is substituted, it can be substituted with the above-listed substituents. In some embodiments, when the alkyl, alkenyl, alkynyl, aryl, heteroaryl, carbocyclyl, or heterocyclyl group is substituted, each is selected from halo, -CN, -SO3, - , -OSO3 - , -SO3H, -SR A , -OR A , -NR B R C , oxo, -CONR B R C , -SO2NR B R C , -COOH, and -COOR B wherein R is independently substituted with one or more substituents selected from the group consisting of A , R B , and R C are each independently substituted with H, alkyl, substituted alkyl, alkenyl, substituted alkenyl, alkynyl, substituted alkynyl, aryl, or substituted aryl.

[0061] As will be appreciated by those of skill in the art, the compounds described herein may be in ionized form, e.g., -CO2 - -SO3 - or -O-SO3 - The compound may have a positively or negatively charged substituent, e.g., -SO3 -In some embodiments, the compound may contain a negatively or positively charged counterion such that the compound is overall neutral. In other embodiments, the compound may exist in a salt form, with the counterion being provided by a conjugate acid or base.

[0062] It is understood that certain radical nomenclature can include either monoradicals or diradicals depending on the context. For example, if a substituent requires two points of attachment to the remainder of the molecule, the substituent is understood to be a diradical. For example, a substituent specified as alkyl, which requires two points of attachment, includes di-radicals such as -CH2-, -CH2CH2-, -CH2CH(CH3)CH2-. Other radical nomenclature clearly indicates that the radical is a diradical, such as "alkylene" or "alkenylene."

[0063] When two "adjacent" R groups are said to form a ring together with the atoms to which they are attached, it is meant that the collective unit of atoms, the intervening bond, and the two R groups is the recited ring. For example, if the following substructure is present:

[0064] [ka] R 1 and R 2 is defined as selected from the group consisting of hydrogen and alkyl, or R 1 and R 2 form an aryl or carbocyclyl together with the atom to which they are attached, but R 1 and R 2 can be selected from hydrogen or alkyl, or alternatively, the substructure means having the structure:

[0065] [ka] wherein A is an aryl ring or a carbocyclyl containing the indicated double bond.

[0066] If a substituent is depicted as a diradical (i.e., having two points of attachment to the remainder of the molecule), it is understood that the substituent can be attached in any directional configuration unless otherwise indicated. Thus, for example, -AE- or

[0067] [ka] Substituents depicted as: represent substituents oriented such that A is attached at the left-most attachment point of the molecule, as well as cases where A is attached at the right-most attachment point of the molecule.

[0068] [ka] where L is defined as an optionally present linker moiety, and when L is absent (or not present), such group or substituent is

[0069] [ka] is equivalent to

[0070] The compounds described herein can be represented in several mesomeric forms. When a single structure is drawn, any of the related mesomeric forms are intended. The bis-boron-containing dyes described herein are represented by a single structure, but can be represented in any of the related mesomeric forms as well. An exemplary mesomeric structure is shown in the following formula (Ia):

[0071] [ka]

[0072] In each case where a single mesomeric form of the compounds described herein is shown, alternative mesomeric forms are also contemplated.In addition, the positive charge on the nitrogen atom of the compound (when there are four bonds connected to the nitrogen atom) and the negative charge on the boron atom (when there are four bonds connected to the boron atom) may not be shown in certain compound structures for simplicity.

[0073] As used herein, a "nucleotide" comprises a nitrogen-containing heterocyclic base, a sugar, and one or more phosphate groups. They are the monomeric units of nucleic acid sequences. In RNA, the sugar is ribose, and in DNA, it is deoxyribose, i.e., a sugar lacking the hydroxyl group present in ribose. The nitrogen-containing heterocyclic base can be a purine, deazapurine, or pyrimidine base. Purine bases include adenine (A) and guanine (G), as well as modified derivatives or analogs thereof, such as 7-deazaadenine or 7-deazaguanine. Pyrimidine bases include cytosine (C), thymine (T), and uracil (U), as well as modified derivatives or analogs thereof. The C-1 atom of the deoxyribose is attached to the N-1 of the pyrimidine or the N-9 of the purine.

[0074] As used herein, a "nucleoside" is structurally similar to a nucleotide, but lacks a phosphate moiety. An example of a nucleoside analog is one in which a label is linked to the base and there is no phosphate group attached to the sugar molecule. The term "nucleoside" is used herein in its ordinary sense as understood by those of skill in the art. Examples include, but are not limited to, ribonucleosides, which contain a ribose moiety, and deoxyribonucleosides, which contain a deoxyribose moiety. Modified pentose moieties are those in which an oxygen atom is replaced with a carbon and / or a carbon is replaced with a sulfur or oxygen atom. A "nucleoside" is a monomer that may have a substituted base and / or sugar moiety. Additionally, nucleosides can be incorporated into larger DNA and / or RNA polymers and oligomers.

[0075] The term "purine base" is used herein in its ordinary sense as understood by those of skill in the art, and includes its tautomers. Similarly, the term "pyrimidine base" is used herein in its ordinary sense as understood by those of skill in the art, and includes its tautomers. A non-limiting list of optionally substituted purine bases includes purine, adenine, guanine, deazapurine, 7-deazaadenine, 7-deazaguanine, hypoxanthine, xanthine, alloxanthine, 7-alkylguanine (e.g., 7-methylguanine), theobromine, caffeine, uric acid, and isoguanine. Examples of pyrimidine bases include, but are not limited to, cytosine, thymine, uracil, 5,6-dihydrouracil, and 5-alkylcytosine (e.g., 5-methylcytosine).

[0076] As used herein, when an oligonucleotide or polynucleotide is described as "comprising" a nucleoside or nucleotide described herein, it means that the nucleoside or nucleotide described herein forms a covalent bond with the oligonucleotide or polynucleotide. Similarly, when a nucleoside or nucleotide is described as being part of an oligonucleotide or polynucleotide, such as being "incorporated" into an oligonucleotide or polynucleotide, it means that the nucleoside or nucleotide described herein forms a covalent bond with the oligonucleotide or polynucleotide. In some such embodiments, the covalent bond is formed between the 3' hydroxy group of the oligonucleotide or polynucleotide and the 5' phosphate group of the nucleotide described herein, as a phosphodiester bond between the 3' carbon atom of the oligonucleotide or polynucleotide and the 5' carbon atom of the nucleotide.

[0077] As used herein, the term "cleavable linker" does not imply that the entire linker must be removed. The cleavage site can be located at a position on the linker that ensures that a portion of the linker remains attached to the detectable label and / or the nucleoside or nucleotide moiety after cleavage.

[0078] As used herein, "derivative" or "analog" refers to a synthetic nucleotide or nucleoside derivative having a modified base moiety and / or a modified sugar moiety. Such derivatives and analogs are discussed, for example, in Scheit, Nucleotide Analogs (John Wiley & Son, 1980) and Uhlman et al., Chemical Reviews 90:543-584, 1990. Nucleotide analogs can also include modified phosphodiester linkages, including phosphorothioate, phosphorodithioate, alkyl-phosphonate, phosphoranilidate, and phosphoramidate linkages. As used herein, "derivative," "analog," and "modified" can be used interchangeably and are encompassed by the terms "nucleotide" and "nucleoside" as defined herein.

[0079] As used herein, the term "phosphate" is used in its ordinary sense as understood by one of ordinary skill in the art and includes its protonated form (e.g.,

[0080] [ka] As used herein, the terms "monophosphate," "diphosphate," and "triphosphate" are used in their ordinary sense as understood by those of skill in the art and include protonated forms.

[0081] As used herein, the term "phasing" refers to a phenomenon in SBS caused by incomplete removal of 3' terminators and fluorophores and / or failure to complete incorporation of some of the DNA strands in a cluster by the polymerase in a given sequencing cycle. Prephasing is caused by incorporation of a nucleotide without an effective 3' terminator, and the incorporation event advances one cycle due to the failure to terminate. Phasing and prephasing cause the signal intensity measured at a particular cycle to consist of signal from the current cycle as well as noise from the previous and next cycles. As the number of cycles increases, the proportion of sequences per cluster affected by phasing and prephasing increases, preventing the identification of the correct base. Prephasing can be caused by the presence of traces of unprotected or unblocked 3'-OH nucleotides during sequencing by synthesis (SBS). Unprotected 3'-OH nucleotides can be generated during the manufacturing process or, in some cases, during storage and reagent handling processes. Thus, the discovery of nucleotide analogs that reduce the incidence of prephasing is surprising and offers a significant advantage over existing nucleotide analogs in SBS applications. For example, the provided nucleotide analogs may result in faster SBS cycle times, lower phasing and prephasing values, and longer sequencing read lengths.

[0082] Dyes containing a bis-boron fused heterocycle of formula (I) Some aspects of the present disclosure relate to bis-boron-containing dyes of formula (I), and salts and mesomeric forms thereof:

[0083] [ka] or a salt or mesomeric form thereof In the formula, R 1 , R 2 , R 3 and R 4each of which is independently H, unsubstituted or substituted C1-C6 alkyl, C1-C6 alkoxy, C2-C6 alkenyl, C2-C6 alkynyl, C1-C6 haloalkyl, C1-C6 haloalkoxy, C1-C6 hydroxyalkyl, (C1-C6 alkoxy)(C1-C6 alkyl), unsubstituted or substituted amino, halo, cyano, hydroxy, nitro, sulfonyl, sulfino, sulfo, sulfonate, S-sulfonamido, N-sulfonamido, unsubstituted or substituted C3-C 10 Carbocyclyl, unsubstituted or substituted C6-C 10 aryl, unsubstituted or substituted 5- to 10-membered heteroaryl, or unsubstituted or substituted 3- to 10-membered heterocyclyl; Each R a , R b , R c and R d are independently halo, cyano, C1-C6 alkyl, C1-C6 haloalkyl, C1-C6 alkoxy, C1-C6 haloalkoxy, C6-C 10 Aryl, C6-C 10 Aryloxy or -OC(=O)R 5 and Alternatively, R a and R b Both are -OC(=O)R 5 If, then, two R 5 together with the atom to which they are attached form an unsubstituted or substituted 6- to 10-membered heterocyclyl, R c and R d Both are -OC(=O)R 5 If, then, two R 5 together with the atom to which they are attached form an unsubstituted or substituted 6- to 10-membered heterocyclyl; R 5 is unsubstituted or substituted C1-C6 alkyl; Ring A is one or more R 6 is a 6-, 7-, 8-, 9- or 10-membered heteroaryl optionally substituted with Each R 6are independently unsubstituted or substituted C1-C6 alkyl, C1-C6 alkoxy, C2-C6 alkenyl, C2-C6 alkynyl, CC1-C6 haloalkyl, C1-C6 haloalkoxy, C1-C6 hydroxyalkyl, (C1-C6 alkoxy)(C1-C6 alkyl), -NR 7 R 8 , halo, cyano, carboxy, hydroxy, nitro, sulfonyl, sulfino, sulfo, sulfonate, S-sulfonamido, N-sulfonamido, unsubstituted or substituted C3-C 10 Carbocyclyl, unsubstituted or substituted C6-C 10 aryl, unsubstituted or substituted 5- to 10-membered heteroaryl, or unsubstituted or substituted 3- to 10-membered heterocyclyl; Each R 7 and R 8 is independently H, unsubstituted or substituted C1-C6 alkyl, or R 7 and R 8 together with the nitrogen atom to which they are attached form an unsubstituted or substituted 3- to 10-membered heterocyclyl; However, R a , R b , R c and R d When each of R is fluoro, 1 , R 2 , R 3 , R 4 and at least one of rings A contains a carboxyl group. 1 , R 2 , R 3 , R 4 At least one of R and ring A contains a carboxyl group. 1 , R 2 , R 3 , R 4 and one of rings A contains a carboxyl group.

[0084] In some embodiments of the compound of Formula (I), ring A is selected from the group consisting of one or more R 6In some embodiments, the 6-membered heteroaryl comprises one or two nitrogen atoms. In further embodiments, the 6-membered heteroaryl is pyridyl, pyrimidyl, or pyrazinyl. In further embodiments, the compound of formula (I) is a compound of formula (Ia) or (Ib):

[0085] [ka] or a salt or mesomeric form thereof, where m is 0, 1, 2, or 3.

[0086] In a further embodiment, the compound of formula (Ia) is also represented by formula (Ic) and the compound of formula (Ib) is also represented by formula (Id) or (Ie)

[0087] [ka] or a salt or mesomeric form thereof.

[0088] In some embodiments of the compounds of Formula (I) and (Ia)-(Ie), each R 6 is independently halo, cyano, carboxyl, unsubstituted or substituted C1-C6 alkyl, unsubstituted phenyl, phenyl substituted with carboxyl, unsubstituted 5-membered heteroaryl, 5-membered heteroaryl substituted with carboxyl, or -NR 7 R 8 In some embodiments, R 6 is halo (e.g., fluoro, chloro, or bromo). In some embodiments, R 6 is unsubstituted furan or carboxyl substituted furan. In some further embodiments, R 6 is a substituted C1-C6 alkyl (e.g., methyl, ethyl, n-propyl, isopropyl, n-butyl, isobutyl, n-pentyl, isopentyl, n-hexyl, or isohexyl), independently selected from halo, -CN, -SO3 -, -SO3H, -NH2, -NH(C1-C6 alkyl), -N(C1-C6 alkyl)2, -C(=O)OH, and -C(=O)O(C1-C6 alkyl). In some further embodiments, R 6 Ha-NR 7 R 8 and R 7 is H and R 8 is C1-C6 alkyl substituted with one or more substituents selected from the group consisting of carboxyl, sulfo, and sulfonate, or R 7 and R 8 together with the nitrogen atom to which they are attached form a 3- to 10-membered heterocyclyl (e.g., a 4-, 5-, 6-, or 7-membered heterocyclyl containing a nitrogen atom, or a nitrogen and a second heteroatom such as oxygen or sulfur) substituted with carboxy. In some further embodiments, R 6 teeth,

[0089] [ka] wherein each of the ring structures is optionally substituted with carboxyl.

[0090] In some embodiments of the compounds of Formula (I) and (Ia)-(Ie), R 1 , R 2 and R 3 Each of is H. In some other embodiments, R 1 , R 2 and R 3 Each of R is independently an unsubstituted C1-C6 alkyl (e.g., methyl, ethyl, n-propyl, isopropyl, n-butyl, isobutyl, n-pentyl, isopentyl, n-hexyl, or isohexyl). 1 and R 3 is methyl, and R 2 is ethyl. In some other embodiments, R 1 , R 2 and R 3Two of R are H or unsubstituted C1-C6 alkyl; 1 , R 2 and R 3 In a further embodiment, one of R is halo, phenyl, 5- or 6-membered heteroaryl, carboxyl, or C1-C6 alkyl substituted with carboxyl. 1 and R 3 is methyl, and R 2 is bromo, chloro, fluoro, phenyl, carboxyl, or -CH2-COOH.

[0091] In some embodiments of the compounds of Formula (I) and (Ia)-(Ie), R 4 is H or unsubstituted C1-C6 alkyl. In some other embodiments, R 4 are each C1-C6 alkyl or phenyl substituted with carboxyl.

[0092] In some embodiments of the compounds of Formula (I) and (Ia)-(Ie), R a and R b Each of is independently fluoro, cyano, methyl, trifluoromethyl, methoxy, phenyl, phenoxy, or -O-acyl. (-OC(=O)CH3 or -OAc). In a further embodiment, R a and R b Both of R are fluoro, methyl, trifluoromethyl, methoxy, or -O-acyl. a and R b Both are -OC(=O)R 5 and two R 5 are the structures along with the atoms to which they are attached.

[0093] [ka] and the methylene portion of the structure may be optionally substituted with one or two substituents selected from fluoro, methyl, trifluoromethyl, methoxy, phenyl, or phenoxy. c and R d Each of R is independently fluoro, cyano, methyl, trifluoromethyl, methoxy, phenyl, phenoxy, or -O-acyl. c and R d Both of R are fluoro, methyl, trifluoromethyl, methoxy, or -O-acyl. c and R d Both are -OC(=O)R 5 and two R 5 are the structures along with the atoms to which they are attached.

[0094] [ka] wherein the methylene portion of the structure may be optionally substituted with one or two substituents selected from fluoro, methyl, trifluoromethyl, methoxy, phenyl, or phenoxy.

[0095] In any embodiment of the compounds of formula (I) and (Ia)-(Ie), C3-C 10 Carbocyclyl (e.g., C3-C 10 Cycloalkyl), C6-C 10 When an aryl, a 5- to 10-membered heteroaryl, or a 3- to 10-membered heterocyclyl is substituted, it is substituted with one or more R 6 When a group is defined as a substituted C1-C6 alkyl, it is C1, C2, C3, C4, C5 or C6 alkyl (e.g., methyl, ethyl, n-propyl, isopropyl, n-butyl, isobutyl, n-pentyl, isopentyl, n-hexyl or isohexyl) and may be substituted with carboxyl, carboxylate, sulfo, sulfonate, -C(O)O(C1-C6 alkyl) or -C(O)NR eR f and each R e and R f is independently H or C1-C6 alkyl substituted with carboxyl, carboxylate, sulfo, or sulfonate.

[0096] Further embodiments of the compounds of formula (I) include, but are not limited to, the following:

[0097] [ka]

[0098] [ka]

[0099] [ka] and their salts and mesomeric forms. Non-limiting examples of the corresponding C1-C6 alkyl carboxylic acid esters (such as methyl esters, ethyl esters, isopropyl esters, and t-butyl esters formed from the carboxylic acid group of the compound).

[0100] In any of the embodiments of the bis-boron fused heterocyclic compounds described herein, the compound may be further modified to introduce a photoprotective moiety, such as a cyclooctatetraene moiety, covalently attached thereto.

[0101] Labeled nucleotides According to one aspect of the present disclosure, the dye compounds described herein are suitable for attachment to substrate moieties, particularly substrate moieties that include a linker group that allows attachment to the substrate moiety. The substrate moiety can be virtually any molecule or substance to which the dyes of the present disclosure can be conjugated, including, but not limited to, nucleosides, nucleotides, polynucleotides, carbohydrates, ligands, particles, solid surfaces, organic and inorganic polymers, chromosomes, nuclei, living cells, and combinations or aggregates thereof. The dyes can be conjugated by various means, including hydrophobic attraction, ionic attraction, and covalent binding, with an optional linker. In some aspects, the dyes are conjugated to the substrate by covalent bonds. More specifically, the covalent bonds are through linker groups. In some cases, such labeled nucleotides are also referred to as "modified nucleotides".

[0102] Some aspects of the present disclosure relate to nucleotides labeled with dyes of formula (I) (including (Ia)-(Ie)) described herein, or salts of their mesomeric forms or derivatives thereof containing the photoprotective moiety COT described herein. The labeled nucleotide or oligonucleotide may be attached to the dye compounds disclosed herein via a carboxy group (-CO2H) or an alkyl-carboxy group to form an amide or alkyl-amide bond. In some further embodiments, the carboxyl group may be in the form of an activated form of the carboxyl group, e.g., an amide or ester, which may be used to attach to an amino or hydroxyl group of a nucleotide or oligonucleotide. As used herein, the term "activated ester" refers to a carboxyl group derivative that can react under mild conditions, e.g., with a compound containing an amino group. Non-limiting examples of activated esters include, but are not limited to, p-nitrophenyl, pentafluorophenyl, and succinimide esters.

[0103] For example, the dye compounds of formula (I) (including (Ia) to (Ie)) are 1 , R 2 , R3 , R 4 and may be attached to the nucleotide via one of rings A of formula (I). In some such embodiments, R 1 , R 2 , R 3 , R 4 and one of rings A contains a carboxyl group, and the bond forms an amide moiety between the carboxyl functionality of the compound of formula (I) and the amino functionality of the nucleotide or nucleotide linker.

[0104] In some embodiments, the dye compound may be covalently attached to the nucleotide via the nucleotide base. In some such embodiments, the labeled nucleotide may have a label attached to the C5 position of the pyrimidine base, or a label attached to the C7 position of the 7-deazapurine base, optionally through a linker moiety. For example, the nucleobase may be 7-deazaadenine, and the dye is attached to the 7-deazaadenine at the C7 position, optionally through a linker. The nucleobase may be 7-deazaguanine, and the dye is attached to the 7-deazaguanine at the C7 position, optionally through a linker. The nucleobase may be cytosine, and the dye is attached to the cytosine at the C5 position, optionally through a linker. As another example, the nucleobase may be thymine or uracil, and the dye is attached to the thymine or uracil at the C5 position, optionally through a linker.

[0105] 3' blocking group The labeled nucleotide or oligonucleotide may also have a blocking group covalently attached to the ribose or deoxyribose sugar of the nucleotide. The blocking group may be attached at any position on the ribose or deoxyribose sugar. In certain embodiments, the blocking group is at the 3' position of the ribose or deoxyribose sugar of the nucleotide. Various 3' blocking groups are disclosed in WO 2004 / 018497 and WO 2014 / 139596, which are incorporated herein by reference. For example, the blocking group may be azidomethyl (-CH2N3) or substituted azidomethyl (e.g., -CH(CHF2)N3 or CH(CH2F)N3), or an allyl connected to the 3' oxygen atom of the ribose or deoxyribose sugar. In some embodiments, the 3' blocking group is azidomethyl, forming 3'-OCH2N3 with the 3' carbon of the ribose or deoxyribose.

[0106] In some other embodiments, the 3' blocking group and the 3' oxygen atom are covalently attached to the 3' carbon of the ribose or deoxyribose.

[0107] [ka] to form an acetal group of the formula In the formula, each R 1a and R 1b are independently H, C1-C6 alkyl, C 1- C6 haloalkyl, C 1- C6 alkoxy, C 1- C6 haloalkoxy, cyano, halogen, optionally substituted phenyl, or optionally substituted aralkyl; Each R 2a and R 2b are independently H, C1-C6 alkyl, C1-C6 haloalkyl, cyano, or halogen; Alternatively, R 1a and R 2a together with the atom to which they are attached form an optionally substituted 5- to 8-membered heterocyclyl group; R F is H, an optionally substituted C2-C6 alkenyl, an optionally substituted C3-C7 cycloalkenyl, an optionally substituted C2-C6 alkynyl, or an optionally substituted (C1-C6 alkylene)Si(R 3a )3, Each R 3a are independently H, C1-C6 alkyl, or optionally substituted C6-C 10 It is aryl.

[0108] Additional 3' hydroxy blocking groups are disclosed in U.S. Patent Application Publication No. 2020 / 0216891(A1), which is incorporated by reference in its entirety. Acetal Blocking Groups

[0109] [ka] are non-limiting examples, each covalently attached to the 3' carbon of ribose or deoxyribose.

[0110] Deprotection of 3' blocking group In some embodiments, the azidomethyl 3' hydroxy protecting group can be removed or deprotected by using a water-soluble phosphine reagent. Non-limiting examples include tris(hydroxymethyl)phosphine (THMP), tris(hydroxyethyl)phosphine (THEP), or tris(hydroxylpropyl)phosphine (THP or THPP). The 3' blocking groups described herein can be removed or cleaved under a variety of chemical conditions. Acetal blocking groups containing vinyl or alkenyl moieties

[0111] [ka] In, non-limiting cleavage conditions include Pd(II) complexes such as Pd(OAc)2 or allylPd(II) chloride dimer in the presence of a phosphine ligand, such as tris(hydroxymethyl)phosphine (THMP) or tris(hydroxylpropyl)phosphine (THP or THPP). For those blocking groups that contain an alkynyl group (e.g., ethynyl), they can also be removed by Pd(II) complexes (e.g., Na2PdCl4, K2PdCl4, Pd(OAc)2 or allylPd(II) chloride dimer) in the presence of a phosphine ligand (e.g., THP or THMP).

[0112] Palladium cleavage reagent In some embodiments, the 3' blocking groups described herein can be cleaved by a palladium catalyst. In some such embodiments, the Pd catalyst is water soluble. In some such embodiments, is a Pd(0) complex (e.g., tris(3,3',3''-phosphinidinetris(benzenesulfonato)palladium(0)) nonasodium salt nonahydrate). In some cases, Pd(0) can be generated in situ from the reduction of a Pd(II) complex with a reagent such as an alkene, alcohol, amine, phosphine, or metal hydride. Suitable palladium sources include Na2PdCl4, K2PdCl4, Pd(CH3CN)2Cl2, (PdCl(C3H5))2, [Pd(C3H5)(THP)]Cl, [Pd(C3H5)(THP)2]Cl, Pd(OAc)2, Pd(Ph3)4, Pd(dba)2, Pd(Acac)2, PdCl2(COD), and Pd(TFA)2. In one such embodiment, the Pd(0) complex is generated in situ from Na2PdCl4. In another embodiment, the palladium source is allylpalladium(II) chloride dimer [(PdCl(C3H5))2]. In some embodiments, the Pd(0) complex is generated in aqueous solution by mixing a Pd(II) complex with a phosphine. Suitable phosphines include water-soluble phosphines such as tris(hydroxypropyl)phosphine (THP), tris(hydroxymethyl)phosphine (THMP), 1,3,5-triaza-7-phosphaadamantane (PTA), bis(p-sulfonatophenyl)phenylphosphine dihydrate potassium salt, tris(carboxyethyl)phosphine (TCEP), and triphenylphosphine-3,3',3''-trisulfonic acid trisodium salt.

[0113] In some embodiments, Pd(0) is prepared in situ by mixing a Pd(II) complex [(PdCl(C3H5))2] with THP. The molar ratio of Pd(II) complex to THP can be about 1:2, 1:3, 1:4, 1:5, 1:6, 1:7, 1:8, 1:9, or 1:10. In some further embodiments, one or more reducing agents, such as ascorbic acid or a salt thereof (e.g., sodium ascorbate), may be added. In some embodiments, the cleavage mixture may contain additional buffering reagents, such as primary amines, secondary amines, tertiary amines, carbonates, phosphates, or borates, or combinations thereof. In some further embodiments, the buffer reagent comprises ethanolamine (EA), tris(hydroxymethyl)aminomethane (Tris), glycine, sodium carbonate, sodium phosphate, sodium borate, 2-dimethylethanolamine (DMEA), 2-diethylethanolamine (DEEA), N,N,N',N'-tetramethylethylenediamine (TEMED), or N,N,N',N'-tetraethylethylenediamine (TEEDA), or combinations thereof. In one embodiment, the buffer reagent is DEEA. In another embodiment, the buffer reagent contains one or more inorganic salts, such as carbonates, phosphates, or borates, or combinations thereof. In one embodiment, the inorganic salt is a sodium salt.

[0114] Linker The dye compounds disclosed herein may include a reactive linker group at one of the substitution positions for covalently attaching the compound to a substrate or another molecule. A reactive linking group is a moiety that can form a bond (e.g., covalent or non-covalent), especially a covalent bond. In certain embodiments, the linker may be a cleavable linker. The use of the term "cleavable linker" does not imply that the entire linker must be removed. The cleavage site may be located at a position on the linker that ensures that a portion of the linker remains attached to the dye and / or substrate moiety after cleavage. The cleavable linker may be, by way of non-limiting examples, an electrophilically cleavable linker, a nucleophilically cleavable linker, a photocleavable linker, a linker that is cleavable under reducing conditions (e.g., a disulfide or azide-containing linker), a linker that is cleavable under oxidative conditions, a linker that is cleavable by the use of a safety lock linker, or a linker that is cleavable by a removal mechanism. By using a cleavable linker to attach the dye compound to the substrate moiety, the label can be removed after detection, if desired, to avoid any interfering signals in downstream steps.

[0115] Useful linker groups can be found in PCT Publication WO 2004 / 018493 (hereby incorporated by reference), examples of which include linkers that can be cleaved using water-soluble phosphines, or water-soluble transition metal catalysts formed from transition metals and at least partially water-soluble ligands. In aqueous solution, the latter forms at least partially water-soluble transition metal complexes. Such cleavable linkers can be used to connect the base of a nucleotide to a label, such as the dyes described herein.

[0116] Particular linkers include those disclosed in PCT Publication No. WO 2004 / 018493 (herein incorporated by reference), such as those that include a moiety of the formula:

[0117] [ka] In the formula, X is selected from the group including O, S, NH, and NQ, Q is a C1-10 substituted or unsubstituted alkyl group, Y is selected from the group including O, S, NH, and N(allyl), and T is hydrogen or C1-C 10 is a substituted or unsubstituted alkyl group, * indicates where the moiety is attached to the remainder of the nucleotide or nucleoside. In some aspects, a linker connects the base of the nucleotide to a label, such as, for example, a dye compound described herein.

[0118] Additional examples of linkers include those disclosed in U.S. Patent Application Publication No. 2016 / 0040225 (herein incorporated by reference), such as those that include a moiety of the formula:

[0119] [ka] (In the formula, * indicates where the moiety is attached to the remainder of the nucleotide or nucleoside. The linker moieties shown herein may include the entire or partial linker structure between the nucleotide / nucleotide and the label. The linker moieties shown herein may include the entire or partial linker structure between the nucleotide / nucleotide and the label.

[0120] Additional examples of linkers include those of the formula:

[0121] [ka] where B is a nucleobase, Z is -N3 (azide), -O-C1-C6 alkyl, -O-C2-C6 alkenyl, or -O-C2-C6 alkynyl, and Fl comprises a dye moiety that may contain an additional linker structure. One of skill in the art will appreciate that the dye compounds described herein are covalently attached to a linker by reacting a functional group (e.g., carboxyl) of the dye compound with a functional group (e.g., amino) of the linker. In one embodiment, the cleavable linker is

[0122] [ka] (the "AOL" linker moiety), where Z is --O-allyl.

[0123] In certain embodiments, the length of the linker between the fluorescent dye (fluorophore) and the guanine base can be modified, for example by introducing a polyethylene glycol spacer group, thereby increasing the fluorescence intensity compared to the same fluorophore attached to the guanine base via other linkages known in the art. Exemplary linkers and their properties are described in PCT Publication WO 2007 / 020457, which is incorporated herein by reference. Such design of linkers, particularly their increased length, can allow for improved brightness of fluorophores attached to the guanine base of guanosine nucleotides when incorporated into polynucleotides such as DNA. Thus, when the dye is for use in any analytical method requiring the detection of a fluorescent dye label attached to a guanine-containing nucleotide, the linker can be of the formula -((CH2)2O) n It is advantageous if the compound contains a spacer group of the formula - (wherein n is an integer from 2 to 50).

[0124] Nucleosides and nucleotides may be labeled at sites on the sugar or nucleobase. As known in the art, a "nucleotide" consists of a nitrogenous base, a sugar, and one or more phosphate groups. In RNA, the sugar is ribose, while in DNA, the sugar is deoxyribose, i.e., a sugar lacking the hydroxyl group present in ribose. The nitrogenous bases are derivatives of purines or pyrimidines. Purines are adenine (A) and guanine (G), and pyrimidines are cytosine (C) and thymine (T), or, in the context of RNA, uracil (U). The C-1 atom of the deoxyribose is attached to the N-1 of the pyrimidine or the N-9 of the purine. Nucleotides are also phosphate esters of nucleosides, with esterification occurring at the hydroxyl group attached to the C-3 or C-5 of the sugar. Nucleosides are usually monophosphates, diphosphates, or triphosphates.

[0125] A "nucleoside" is structurally similar to a nucleotide, but lacks the phosphate moiety. An example of a nucleotide analog is one in which a label is linked to the base and there is no phosphate group attached to the sugar molecule.

[0126] Although the bases are commonly referred to as purines or pyrimidines, one of skill in the art will appreciate that derivatives and analogs are available that do not alter the ability of the nucleotide or nucleoside to undergo Watson-Crick base pairing. By "derivative" or "analog" is meant a compound or molecule whose core structure is the same as or closely similar to the parent compound, but that has been modified chemically or physically, e.g., with different or additional side groups that allow the derivative nucleotide or nucleoside to be linked to other molecules. For example, the base may be a deazapurine. In certain embodiments, the derivative should be able to undergo Watson-Crick pairing. "Derivative" and "analog" also include synthetic nucleotide or nucleoside derivatives, e.g., with modified base moieties and / or modified sugar moieties. Such derivatives and analogs are discussed, for example, in Scheit, Nucleotide analogs (John Wiley & Son, 1980) and Uhlman, Chemical Reviews 90:543-584, 1990. Nucleotide analogs can also contain modified phosphodiester linkages, including phosphorothioate, phosphorodithioate, alkyl-phosphonate, phosphoranilidate, phosphoramidate linkages, and the like.

[0127] The dye may be attached at any position on the nucleotide base, for example, via a linker. In certain embodiments, Watson-Crick base pairing can still be performed on the resulting analog. Particular nucleobase labeling sites include the C5 position of pyrimidine bases, or the C7 position of 7-deazapurine bases. As noted above, a linker group can be used to covalently attach the dye to the nucleoside or nucleotide.

[0128] In certain embodiments, the labeled nucleotide or oligonucleotide can be enzymatically incorporated and enzymatically extendable.Thus, the linker portion can be long enough to connect the nucleotide to the compound, so that the compound does not significantly interfere with the overall binding and recognition of the nucleotide by nucleic acid replicating enzyme.Thus, the linker can also include a spacer unit.The spacer, for example, distances the nucleotide base from the cleavage site or the label.

[0129] A nucleoside or nucleotide labeled with the dyes described herein can have the formula:

[0130] [ka]

[0131] wherein the dye is a dye containing a fused bis-boron heterocycle (label) as described herein, (after covalent attachment between the functional group of the dye and the functional group of the linker "L"), B is a nucleobase, e.g., uracil, thymine, cytosine, adenine, 7-deazaadenine, guanine, 7-deazaguanine, etc., and L is an optional linker group that may or may not be present, R' can be H, -OR' is a monophosphate, diphosphate, triphosphate, thiophosphate, phosphate ester analog, -O- attached to a reactive phosphorus-containing group, or -O- protected by a blocking group, R" is H or OH, R'" is H, a 3' hydroxy blocking group as described herein, or -OR'" forms a phosphoramidite. wherein -OR'" is a phosphoramidite and R' is an acid cleavable hydroxy protecting group that allows for subsequent monomer coupling under automated synthesis conditions. In some further embodiments, B is

[0132] [ka] or optionally substituted derivatives and analogs thereof. In some further embodiments, the labeled nucleobase has the structure

[0133] [ka] Includes.

[0134] In certain embodiments, the blocking group is separate and independent from the dye compound, i.e., not attached to the dye compound. Alternatively, the dye may include all or a portion of the 3'-OH blocking group. Thus, R''' may be a 3' hydroxy blocking group that may or may not include a dye compound.

[0135] In yet another alternative embodiment, there is no blocking group on the 3' carbon of the pentose sugar, e.g., the dye (or dye and linker structure) attached to the base may be of sufficient size or structure to act as a block to the incorporation of additional nucleotides. Thus, regardless of whether the dye is attached to the 3' position of the sugar, blocking may be due to steric hindrance or may be due to a combination of size, charge, and structure.

[0136] In yet another alternative embodiment, a blocking group can be present on the 2' or 4' carbon of the pentose sugar and can be of sufficient size or structure to act as a block against the incorporation of additional nucleotides.

[0137] The use of blocking groups allows for the control of polymerization, e.g., by terminating elongation when a labeled nucleotide is incorporated. If the blocking effect is reversible, e.g., but not limited to, by changing the chemical conditions or by removing the chemical block, elongation can be halted at a particular point and then allowed to continue.

[0138] In certain embodiments, the linker (the linker between the dye and the nucleotide) and the blocking group are both present and are separate moieties. In certain embodiments, the linker and the blocking group are both cleavable under the same or substantially similar conditions. Thus, the deprotection and deblocking process may be more efficient since only a single treatment is required to remove both the dye compound and the blocking group. However, in some embodiments, the linker and the blocking group do not need to be cleavable under similar conditions, instead being individually cleavable under separate conditions.

[0139] The present disclosure also encompasses polynucleotides incorporating dye compounds. Such polynucleotides may be DNA or RNA composed of deoxyribonucleotides or ribonucleotides, respectively, linked by phosphodiester linkages. Polynucleotides may include naturally occurring nucleotides, non-naturally occurring (or modified) nucleotides other than the labeled nucleotides described herein, or any combination thereof, in combination with at least one modified nucleotide as described herein (e.g., labeled with a dye compound). Polynucleotides according to the present disclosure may also include non-naturally occurring backbone linkages and / or non-nucleotide chemical modifications. Chimeric structures composed of a mixture of ribonucleotides and deoxyribonucleotides containing at least one labeled nucleotide are also contemplated.

[0140] Non-limiting exemplary labeled nucleotides described herein include:

[0141] [ka] where L represents a linker and R represents a ribose or deoxyribose moiety as described above, or a ribose or deoxyribose moiety substituted at the 5' position with a monophosphate, diphosphate, or triphosphate.

[0142] In some embodiments, non-limiting fluorescent dye conjugates are shown below:

[0143] [ka]

[0144] [ka] wherein PG represents a 3'OH blocking group as described herein, p is an integer of 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10, and k is 0, 1, 2, 3, 4, or 5. In one embodiment, -O-PG is AOM. In another embodiment, -O-PG is -O-azidomethyl. In one embodiment, k is 5. In some further embodiments, p is 1, 2, or 3, and k is 5.

[0145] [ka] refers to the point of attachment of the dye with a cleavable linker as a result of a reaction between an amino group of the linker moiety and a carboxyl group of the dye. In any embodiment of the labeled nucleotides described herein, the nucleotide is a nucleotide triphosphate.

[0146] Additional aspects of the present disclosure relate to oligonucleotides or polynucleotides comprising the labeled nucleotides described herein. In some embodiments, the oligonucleotides or polynucleotides are hybridized to and / or complementary to at least a portion of a target polynucleotide. In some embodiments, the target polynucleotide is immobilized on a solid support. In some further embodiments, the solid support comprises an array of a plurality of immobilized target polynucleotides. In further embodiments, the solid support comprises a patterned flow cell.

[0147] Additional aspects of the present disclosure relate to protein tags or antibodies that contain one or more bis-boron dyes described herein. In particular, the protein tags or antibodies may contain multiple copies of the same dye for increased fluorescence intensity. The protein tags or antibodies can be used as affinity reagents that superficially bind to certain types of unlabeled 3' blocked nucleotides.

[0148] kit Provided herein is a kit that includes a first type of nucleotide labeled with a bis-boron dye of the present disclosure (i.e., a first label). In some embodiments, the kit also includes a second type of labeled nucleotide labeled with a second compound (i.e., a second label) that is different from the bis-boron dye in the first labeled nucleotide. In some further embodiments, the kit can include a third type of nucleotide, which is labeled with a third compound (i.e., a third label) that is different from the first and second labels. In some further embodiments, the kit can further include a fourth type of nucleotide. In some such embodiments, the fourth type of nucleotide is unlabeled (dark). In other embodiments, the fourth nucleotide is labeled with a compound different from the first, second, and third nucleotides, and each label has a distinct absorbance maximum that is distinguishable from the other labels. In some embodiments, the nucleotides can be used in sequencing applications that involve the use of two light sources with different wavelengths. In some embodiments, the first light source has a wavelength of about 500 nm to about 550 nm, about 510 to about 540 nm, or about 520 to about 530 nm (e.g., 520 nm). The second light source has a wavelength of about 400 nm to about 480 nm, about 420 nm to about 470 nm, or 450 nm to about 460 nm (e.g., 450 nm). In further embodiments, the first label, second label, and third label each have an emission spectrum that can be collected into two separate collection filters or channels.

[0149] In some embodiments, the kit may include four types of labeled nucleotides (A, C, G, and T or U), where the first type of nucleotide is labeled with a compound disclosed herein. In such a kit, each of the four types of nucleotides may be labeled with a compound that is the same as or different from the label on the other three nucleotides. Alternatively, the first type of the four types of nucleotides is a labeled nucleotide described herein (i.e., labeled with a bis-boron dye described herein), the second type of nucleotide carries a second label, the third type of nucleotide carries a third label, and the fourth type of nucleotide is unlabeled (dark). As another example, the first type of the four types of nucleotides is a labeled nucleotide as described herein, the second type of nucleotide carries a second label, the third type of nucleotide comprises a mixture of the third type of nucleotide carrying two labels (e.g., the third type of nucleotide carrying a first label and the third type of nucleotide carrying a second label), and the fourth type of nucleotide is unlabeled (dark).In this particular example, one or both of the two labels of the third type of nucleotide may be a label that is structurally different from the first or second label, but can be excited under the same wavelength of the light source, but has a stronger emission signal intensity (e.g., the third type of nucleotide carrying a third label and the third type of nucleotide carrying a fourth label, where the third label can be excited under the same wavelength as the first label and the fourth label can be excited under the same wavelength as the second label). In these examples, one or more of the labeled compounds may have a distinct absorbance and / or emission maximum such that the compound is distinguishable from the other compounds. For example, each compound may have a different absorbance and / or emission maximum such that each of the compounds is spectrally distinguishable from the other three compounds (or two compounds if the fourth nucleotide is unlabeled). It will be understood that portions of the absorbance and / or emission spectrum other than the maximum may differ, and these differences may be utilized to distinguish the compounds.The kit may be such that two or more compounds have distinct absorbance maxima. The bis-boron dyes described herein typically absorb light in the region below 500 nm. For example, these bis-boron dyes may have an absorption wavelength of about 410 nm to about 480 nm, about 420 nm to about 470 nm, or about 440 nm to about 460 nm.

[0150] The bis-boron compounds, nucleotides, or kits described herein may be used to detect, measure, or identify biological systems, including, for example, processes or components thereof. Exemplary techniques that the compounds, nucleotides, or kits may be used for include sequencing, expression analysis, hybridization analysis, genetic analysis, RNA analysis, cell assays (e.g., cell binding or cell function analysis), or protein assays (e.g., protein binding assays or protein activity assays). The use may be an automated instrument for performing a particular technique, such as an automated sequencing instrument. The sequencing instrument may include two light sources operating at different wavelengths.

[0151] In certain embodiments, the labeled nucleotides described herein may be provided in combination with unlabeled or natural nucleotides, or any combination thereof. The combination of nucleotides may be provided as separate individual components (e.g., one nucleotide type per container or tube) or as a nucleotide mixture (e.g., two or more nucleotides mixed in the same container or tube).

[0152] When the kit includes multiple, particularly two, or three, or more specifically four nucleotides, the different nucleotides may be labeled with different dye compounds or may be dark, without dye compounds. When the different nucleotides are labeled with different dye compounds, it is a feature of the kit that the dye compounds are spectrally distinguishable fluorescent dyes. As used herein, the term "spectrally distinguishable fluorescent dyes" refers to fluorescent dyes that emit fluorescent energy at wavelengths that can be distinguished by a fluorescent detection device (e.g., a commercially available capillary-based DNA sequencing platform) when two or more dyes are present in a sample. When two nucleotides labeled with fluorescent dye compounds are provided in kit form, it is a feature of some embodiments that the spectrally distinguishable fluorescent dyes can be excited at the same wavelength, e.g., by the same light source. When four nucleotides labeled with fluorescent dye compounds are provided in kit form, it is a feature of some embodiments that two of the spectrally distinguishable fluorescent dyes can be excited at one wavelength, and the other two spectrally distinguishable dyes can both be excited at another wavelength. The specific excitation wavelength for the dye is 450 to 460 nm, 490 to 500 nm, or 520 nm or longer (for example, 532 nm).

[0153] In an alternative embodiment, the kit of the present disclosure may contain nucleotides in which the same base is labeled with two different compounds. The first nucleotide may be labeled with a compound of the present disclosure, for example, with a "blue" dye that absorbs below 500 nm. The second nucleotide may be labeled with a spectrally distinguishable compound, for example, with a "green" dye that absorbs below 600 nm but above 500 nm. The third nucleotide may be labeled as a mixture of a compound of the present disclosure and a spectrally distinguishable compound, and the fourth nucleotide may be "dark" and contain no label. Thus, for convenience, nucleotides 1-4 may be labeled "blue", "green", "blue / green", and dark. To further simplify the instrumentation, the four nucleotides may be labeled with two dyes excited by a single light source, and thus the labels of nucleotides 1-4 may be "blue 1", "blue 2", "blue 1 / blue 2", and dark.

[0154] Although the kits are exemplified herein in a configuration having different nucleotides labeled with different dye compounds, it will be understood that the kits may include two, three, four or more different nucleotides having the same dye compound.

[0155] In addition to the labeled nucleotides, the kit may also include at least one additional component. The additional component may be one or more of the components identified in the methods described herein or in the Examples section below. Some non-limiting examples of components that can be combined in the kits of the present disclosure are described below. In some embodiments, the kit further includes a DNA polymerase (a mutant DNA polymerase of 9°N polymerase, such as those disclosed in WO 2005 / 024010) and one or more buffer compositions. One buffer composition may include an antioxidant, such as ascorbic acid or sodium ascorbate, and may be used to protect the dye compound from photodamage during detection. Additional buffer compositions may include reagents that may be used to cleave the 3' blocking group and / or the cleavable linker. For example, a water-soluble phosphine or a water-soluble transition metal catalyst formed from at least partially water-soluble ligands, such as transition metal and palladium complexes. Various components of the kit may be provided in concentrated form that is diluted before use. In such embodiments, a suitable dilution buffer may also be included. Again, one or more of the components identified in the methods described herein can be included in the kits of the present disclosure. In any of the embodiments of the nucleotides or labeled nucleotides described herein, the nucleotide contains a 3' blocking group.

[0156] Alternatively, the kit may include one or more different types of unlabeled 3' blocked nucleotides and one or more affinity reagents (e.g., protein tags and antibodies), where at least one affinity reagent is labeled with multiple copies of a bis-boron dye described herein.

[0157] Sequencing methods Nucleotides containing dye compounds according to the present disclosure can be used in any analytical method involving detection of a fluorescent label attached to such nucleotides, regardless of whether the nucleotide or nucleoside is analyzed by itself or when the nucleotide or nucleoside is incorporated or associated with a larger molecular structure or conjugate. In this context, the term "incorporated into a polynucleotide" can mean that the 5' phosphate is linked to the 3' hydroxyl group of a second nucleotide, which itself may form part of a longer polynucleotide chain, by a phosphodiester linkage. The 3' end of the nucleotide described herein may or may not be linked to the 5' phosphate of a further nucleotide by a phosphodiester linkage. Thus, in a non-limiting embodiment, the present disclosure provides a method for detecting a labeled nucleotide incorporated into a polynucleotide, comprising: (a) incorporating at least one labeled nucleotide of the present disclosure into a polynucleotide; and (b) determining the identity of the nucleotide incorporated into the polynucleotide by detecting a fluorescent signal from a dye compound attached to the nucleotide.

[0158] The method may include (a) a synthesis step, in which one or more labeled nucleotides according to the present disclosure are incorporated into a polynucleotide, and (b) a detection step, in which one or more labeled nucleotides incorporated into the polynucleotide are detected by detecting or quantitatively measuring their fluorescence.

[0159] Some embodiments of the present application are methods of determining the sequence of a target polynucleotide (e.g., a single-stranded target polynucleotide), comprising: (a) contacting a primer polynucleotide with one or more labeled nucleotides (e.g., nucleoside triphosphates A, G, C, and T), where at least one of the labeled nucleotides is a labeled nucleotide described herein, and the primer polynucleotide is complementary to at least a portion of the target polynucleotide; (b) incorporating the labeled nucleotide into the primer polynucleotide; and (c) performing one or more fluorescence measurements to determine the identity of the incorporated nucleotide. In some such embodiments, a primer polynucleotide / target polynucleotide complex is formed by contacting a target polynucleotide with a primer polynucleotide complementary to at least a portion of the target polynucleotide. In some embodiments, the method further comprises (d) removing a label moiety and a 3' hydroxy blocking group from the nucleotide incorporated into the primer polynucleotide. In some further embodiments, the method may also comprise (e) washing the removed label moiety and 3' blocking group from the primer polynucleotide strand. In some embodiments, steps (a)-(d) or steps (a)-(e) are repeated until at least a portion of the target polynucleotide strand is sequenced. In some cases, steps (a)-(d) or steps (a)-(e) are repeated for at least 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 250, or 300 cycles. In some embodiments, the label moiety and the 3' blocking group from the nucleotide incorporated into the primer polynucleotide strand are removed in a single chemical reaction. In some further embodiments, the method is performed on an automated sequencing instrument, the automated sequencing instrument comprising two light sources operating at different wavelengths. In some embodiments, sequencing is performed after completion of the repeated cycles of the sequencing steps described herein.

[0160] Some embodiments of the present disclosure relate to a method for determining a sequence of a target polynucleotide (e.g., a single-stranded target polynucleotide), comprising: (a) contacting a primer polynucleotide with an incorporation mixture comprising one or more of four different types of nucleotide conjugates, wherein a first type of nucleotide conjugate comprises a first label, a second type of nucleotide conjugate comprises a second label, and a third type of nucleotide conjugate comprises a third label, each of the first label, the second label, and the third label being spectrally distinct from one another, and wherein the primer polynucleotide is spectrally distinct from at least one of the target polynucleotides. (b) contacting a nucleotide conjugate from the mixture with a primer polynucleotide complementary to a portion of the nucleotide conjugate; (c) performing a first imaging event using a first excitation light source to detect a first emission signal from the extended polynucleotide; and (d) performing a second imaging event using a second excitation light source to detect a second emission signal from the extended polynucleotide, wherein the first excitation light source and the second excitation light source have different wavelengths, and the first emission signal and the second emission signal are detected or collected in a single emission detection channel. In some embodiments, the bis-boron dyes described herein can be used as any one of the first, second, or third labels described in the method. In some embodiments, the method does not include chemical modification of any nucleotide conjugate in the mixture after the first imaging event and before the second imaging event. In some further embodiments, the incorporation mixture further comprises a fourth type of nucleotide, and the fourth type of nucleotide is unlabeled or labeled with a fluorescent moiety that does not emit a signal from either the first or second imaging event.In this sequencing method, the identity of each incorporated nucleotide conjugate is determined based on the detection pattern of the first imaging event and the second imaging event.For example, the incorporation of the first type of nucleotide conjugate is determined by the signal state in the first imaging event and the dark state in the second imaging event.The incorporation of the second type of nucleotide conjugate is determined by the dark state in the first imaging event and the signal state in the second imaging event. The incorporation of the third type of nucleotide conjugate is determined by the signal state in both the first imaging event and the second imaging event. The incorporation of the fourth type of nucleotide conjugate is determined by the dark state in both the first imaging event and the second imaging event. In further embodiments, steps (a)-(d) are performed in repeated cycles (e.g., at least 30, 50, 100, 150, 200, 250, 300, 400, or 500 times), and the method further comprises successively determining the sequence of at least a portion of the single-stranded target polynucleotide based on the identity of each successively incorporated nucleotide conjugate. In some embodiments, the first excitation light source has a shorter wavelength than the second excitation light source. In some such embodiments, the first excitation light source has a wavelength of about 400 nm to about 480 nm, about 420 nm to about 470 nm, or about 450 nm to about 460 nm (i.e., "blue light"). In one embodiment, the first excitation light source has a wavelength of about 450 nm. The second excitation light source has a wavelength of about 500 nm to about 550 nm, about 510 nm to about 540 nm, or about 520 nm to about 535 nm (i.e., "green light"). In one embodiment, the second excitation light source has a wavelength of about 520 nm. In other embodiments, the first excitation light source has a longer wavelength than the second excitation light source. In some such embodiments, the first excitation light source has a wavelength of about 500 nm to about 550 nm, about 510 nm to about 540 nm, or about 520 nm to about 535 nm (i.e., "green light"). In one embodiment, the second excitation light source has a wavelength of about 520 nm. The second excitation light source has a wavelength of about 400 nm to about 480 nm, about 420 nm to about 470 nm, or about 450 nm to about 460 nm (i.e., "blue light"). In one embodiment, the second excitation light source has a wavelength of about 450 nm.

[0161] Some embodiments of the present disclosure relate to a method for determining the sequence of a plurality of target polynucleotides (e.g., a plurality of different target polynucleotides), comprising: (a) contacting a solid support with a solution comprising a sequencing primer under hybridization conditions, the solid support comprising a plurality of different target polynucleotides immobilized thereon, the sequencing primer being complementary to at least a portion of the target polynucleotides; (b) contacting the solid support with an aqueous solution comprising a DNA polymerase and one or more of four different types of nucleotides under conditions suitable for DNA polymerase-mediated primer extension, where at least one type of nucleotide is a labeled nucleotide as described herein; (c) incorporating one type of nucleotide into the sequencing primer to generate an extended copy polynucleotide; and (d) performing one or more fluorescence measurements of the extended copy polynucleotide to determine the identity of the incorporated nucleotide. In some embodiments, the method further comprises (e) removing a 3' blocking group from the incorporated nucleotide into the extended copy polynucleotide. In some such embodiments, step (e) also removes the label of the incorporated nucleotide. In some embodiments, the method further comprises (f) washing the solid support after said removal of the label and 3' blocking group from the incorporated nucleotide. In further embodiments, the method comprises repeating steps (b)-(f) until the sequence of at least a portion of the target polynucleotide is determined. In some such embodiments, steps (b)-(f) are repeated at least 50, 100, 150, 200, 250, or 300 cycles. In further embodiments, the label and 3' blocking group from the incorporated nucleotide in the extended copy polynucleotide are removed in a single chemical reaction. In some embodiments, step (d) comprises two imaging and fluorescence measurements. In further embodiments, the method is performed on an automated sequencing instrument, and the automated sequencing instrument comprises two light sources operating at different wavelengths.In some such embodiments, one light source operates at a wavelength of about 400 nm to about 480 nm, about 420 nm to about 470 nm, or about 450 nm to about 460 nm (i.e., "blue light"). In further embodiments, the other light source operates at a wavelength of about 500 nm to about 550 nm, about 510 nm to about 540 nm, or about 520 nm to about 535 nm (i.e., "green light"). In some embodiments, the four different types of nucleotides are dATP, dCTP, dGTP, and dTTP or dUTP, or non-natural nucleotide analogs thereof. In certain embodiments, the aqueous solution comprising DNA polymerase and one or more of four different types of nucleotides comprises or is an incorporation mixture having a first type of nucleotide carrying a first label (labeled with a bis-boron dye as described herein), a second type of nucleotide carrying a second label, a third type of nucleotide carrying a mixture of two labels, and a fourth type of nucleotide that is unlabeled (dark). For example, the third type of nucleotide may be a mixture of a third type of nucleotide carrying a first label and a third type of nucleotide carrying a second label. In such an embodiment, the incorporation of the first type of nucleotide may be determined by the signal state in the first imaging event and the dark state in the second imaging event. The incorporation of the second type of nucleotide may be determined by the signal state in the first imaging event / fluorescence measurement and the dark state in the second imaging event / fluorescence measurement. The incorporation of the third type of nucleotide is determined by the signal state in both the first and second imaging events / fluorescence measurement. Incorporation of the fourth type of nucleotide conjugate is determined by dark conditions in both the first and second imaging events / fluorescence measurements. In another embodiment, the incorporation mixture comprises a first type of nucleotide carrying a first label (labeled with a bis-boron dye described herein), a second type of nucleotide carrying a second label, a third type of nucleotide having a third label, and a fourth type of unlabeled nucleotide.In this case, each of the first label, the second label, and the third label is spectrally distinct from one another, the first label being excitable by a first light source, the second label being excitable by a second light source, and the third label being excitable by both the first and second light sources. As a result, the incorporation of the four types of nucleotides can also be distinguished based on the same signal patterns described herein.

[0162] In some embodiments of the sequencing method described herein, at least one nucleotide is incorporated into a polynucleotide (such as a single-stranded primer polynucleotide described herein) in a synthesis step by the action of a polymerase enzyme. However, other methods of attaching nucleotides to polynucleotides can be used, such as chemical oligonucleotide synthesis or ligation of a labeled oligonucleotide to an unlabeled oligonucleotide. Thus, when the term "incorporate" is used in reference to nucleotides and polynucleotides, it can include chemical and enzymatic polynucleotide synthesis.

[0163] In certain embodiments, a synthesis step is performed and may optionally include incubating a template or target polynucleotide strand with a reaction mixture that includes the fluorescently labeled nucleotides of the present disclosure. A polymerase may also be provided under conditions that allow the formation of a phosphodiester linkage between a free 3' hydroxyl group on a polynucleotide strand annealed to the template or target polynucleotide strand and a 5' phosphate group on a labeled nucleotide. Thus, the synthesis step may include the formation of a polynucleotide strand directed by complementary base pairing of the nucleotide to the template / target strand.

[0164] In all embodiments of the method, the detection step may be performed while the polynucleotide strand into which the labeled nucleotide is incorporated is annealed to the template / target strand, or after the denaturation step in which the two strands are separated. Additional steps may be included between the synthesis step and the detection step, such as chemical or enzymatic reaction steps or purification steps. In particular, the polynucleotide strand incorporating the labeled nucleotide may be isolated or purified and then further processed or used for subsequent analysis. As an example, the target polynucleotide incorporating the labeled nucleotide described herein in the synthesis step may then be used as a labeled probe or primer. In other embodiments, the product of the synthesis step described herein may be subjected to additional reaction steps, and the products of these subsequent steps may be purified or isolated, if desired.

[0165] Suitable conditions for the synthesis step are well known to those familiar with standard molecular biology techniques. In one embodiment, the synthesis step may be similar to a standard primer extension reaction using nucleotide precursors containing the labeled nucleotides described herein to form an extended polynucleotide strand (primer polynucleotide strand) complementary to the template / target strand in the presence of a suitable polymerase enzyme. In other embodiments, the synthesis step may itself be part of an amplification reaction that generates a labeled double-stranded amplification product consisting of a primer and an annealed complementary strand derived from copying the template polynucleotide strand. Other exemplary synthesis steps include nick translation, strand displacement polymerization, random priming DNA labeling, and the like. Particularly useful polymerase enzymes for the synthesis step are those that are capable of catalyzing the incorporation of the labeled nucleotides described herein. A variety of naturally occurring or mutant / modified polymerases can be used. As an example, thermostable polymerases can be used for synthesis reactions carried out using thermal cycling conditions, but thermostable polymerases may not be desirable for isothermal primer extension reactions. Suitable thermostable polymerases that can incorporate labeled nucleotides according to the present disclosure include those described in WO 2005 / 024010 or WO 06120433, each of which is incorporated herein by reference. For synthesis reactions carried out at low temperatures, such as 37°C, the polymerase enzyme does not necessarily need to be a thermostable polymerase, and thus the choice of polymerase depends on several factors, such as reaction temperature, pH, strand displacement activity, etc. Exemplary polymerases include, but are not limited to, Pol812, Pol1901, Pol1558, or Pol963. The amino acid sequences of Pol812, Pol1901, Pol1558, or Pol963 DNA polymerases are described, for example, in U.S. Patent Publication Nos. 2020 / 0131484A1 and 2020 / 0181587A1, both of which are incorporated herein by reference.

[0166] In certain non-limiting embodiments, the present disclosure includes methods of nucleic acid sequencing, resequencing, whole genome sequencing, single nucleotide polymorphism scoring, as well as any other application involving detection of modified nucleotides or nucleosides labeled with the dyes described herein when incorporated into a polynucleotide.

[0167] Certain embodiments of the present disclosure provide for the use of labeled nucleotides comprising dye moieties according to the present disclosure in polynucleotide sequencing-by-synthesis reactions. Sequencing-by-synthesis generally involves sequentially adding one or more nucleotides or oligonucleotides in a 5' to 3' direction to a growing polynucleotide chain using a polymerase or ligase to form an extended polynucleotide chain complementary to the template / target nucleic acid to be sequenced. The identity of the base present in one or more of the added nucleotides can be determined in a detection or "imaging" step. The identity of the added base can be determined after each nucleotide incorporation step. The sequence of the template can then be inferred using conventional Watson-Crick base pairing rules. Using nucleotides labeled with the dyes described herein to determine the identity of a single base can be useful, for example, in scoring single nucleotide polymorphisms, although such single base extension reactions are within the scope of the present disclosure.

[0168] In one embodiment of the present disclosure, the sequence of a template polynucleotide is determined by detecting the incorporation of one or more nucleotides into a nascent strand complementary to the template / target polynucleotide to be sequenced through detection of a fluorescent label attached to the incorporated nucleotide. The sequencing of the template polynucleotide can be primed with a suitable primer (or a primer prepared as a hairpin structure containing the primer as part of the hairpin), and the nascent strand is extended stepwise by adding nucleotides to the 3' end of the primer in a polymer catalyzed reaction.

[0169] In certain embodiments, each of the different nucleotide triphosphates (A, T, G, and C) may be labeled with a unique fluorophore and contain a blocking group at the 3' position to prevent uncontrolled polymerization. Alternatively, one of the four nucleotides may be unlabeled (dark). The polymerase enzyme will incorporate the nucleotide into the nascent strand complementary to the template / target polynucleotide, but the blocking group will prevent further incorporation of the nucleotide. Any unincorporated nucleotides can be washed away, and the fluorescent signal from each incorporated nucleotide can be optically "read" by suitable means, such as a charge-coupled device using a light source excitation and suitable emission filters. The 3'-blocking group and the fluorochrome compound can then be removed (deprotected) (simultaneously or sequentially) to subject the nascent strand to incorporation of further nucleotides. Typically, the identity of the incorporated nucleotide is determined after each incorporation step, although this is not strictly required. Similarly, US Pat. No. 5,302,509 (hereby incorporated by reference in its entirety) discloses a method for sequencing polynucleotides immobilized on a solid support.

[0170] The method, as exemplified above, incorporates fluorescently labeled 3' blocked nucleotides A, G, C, and T into a growing strand complementary to an immobilized polynucleotide in the presence of a DNA polymerase. The polymerase incorporates the base complementary to the target polynucleotide, but further addition is prevented by the 3'-blocking group. The label of the incorporated nucleotide can then be determined, after which the blocking group can be removed by chemical cleavage to allow further polymerization. The nucleic acid template to be sequenced in a sequencing by synthesis reaction can be any polynucleotide for which sequencing is desired. Nucleic acid templates for sequencing reactions typically contain a double-stranded region with a free 3' hydroxy group that serves as a primer or starting point for the addition of additional nucleotides in the sequencing reaction. The region of the template to be sequenced overhangs this free 3' hydroxy group on the complementary strand. The overhang region of the template to be sequenced may be single-stranded, but may also be double-stranded, except that there is a "nick" on the strand complementary to the template strand to be sequenced to provide a free 3'OH group to initiate the sequencing reaction. In such embodiments, sequencing may proceed by strand displacement. In certain embodiments, a primer with a free 3' hydroxy group may be added as a separate component (e.g., a short oligonucleotide) that hybridizes to the single-stranded region of the template to be sequenced. Alternatively, the primer and template strand to be sequenced may each form part of a partially self-complementary nucleic acid strand that can form an intramolecular duplex, such as a hairpin loop structure. Hairpin polynucleotides and methods by which they can be attached to solid supports are disclosed in PCT Publication Nos. 0157248 and 2005 / 047301, each of which is incorporated herein by reference. Nucleotides can be added sequentially to the growing primer to synthesize a polynucleotide strand in the 5' to 3' direction. The nature of the added base may be determined, particularly but not necessarily, after each nucleotide addition, and such determination provides sequence information about the nucleic acid template.Thus, a nucleotide is incorporated into a nucleic acid strand (or polynucleotide) by linking the nucleotide to a free 3' hydroxy group of the nucleic acid strand through formation of a phosphodiester linkage with the 5' phosphate group of the nucleotide.

[0171] The nucleic acid template to be sequenced may be DNA or RNA, or even a hybrid molecule consisting of deoxynucleotides and ribonucleotides. The nucleic acid template may contain naturally occurring and / or non-naturally occurring nucleotides and naturally occurring or non-naturally occurring backbone linkages, as long as this does not prevent copying of the template in the sequencing reaction.

[0172] In certain embodiments, the nucleic acid template to be sequenced can be attached to a solid support via any suitable linking method known in the art, for example, covalent bonding. In certain embodiments, the template polynucleotide can be directly attached to a solid support (e.g., a silica-based support). However, in other embodiments of the present disclosure, the surface of the solid support can be modified in some way to allow either direct covalent bonding of the template polynucleotide, or immobilization of the template polynucleotide through a hydrogel or polyelectrolyte layer (which itself is attached to the solid support in a manner other than covalent bonding).

[0173] Arrays in which polynucleotides are directly attached to a support (e.g., a silica-based support) are disclosed, for example, in WO 00 / 06770 (hereby incorporated by reference), in which polynucleotides are immobilized on a glass support by reaction between pendant epoxide groups on the glass and internal amino groups on the polynucleotide.In addition, polynucleotides can be attached to a solid support by reaction of a sulfur-based nucleophile with the solid support, as described, for example, in WO 2005 / 047301 (hereby incorporated by reference). Still further examples of solid-supported template polynucleotides include those in which the template polynucleotide is attached to a hydrogel supported on a silica-based or other solid support, as described, for example, in WO 00 / 31148, WO 01 / 01143, WO 02 / 12566, WO 03 / 014392, U.S. Pat. No. 6,465,178, and WO 00 / 53812, each of which is incorporated herein by reference.

[0174] Specific surfaces that template polynucleotides can be immobilized on include polyacrylamide hydrogels.Polyacrylamide hydrogels are described in the above references and in WO 2005 / 065814, which are incorporated herein by reference.Specific hydrogels that can be used include those described in WO 2005 / 065814 and US 2014 / 0079923.In one embodiment, the hydrogel is PAZAM (poly(N-(5-azidoacetamidylpentyl)acrylamide-co-acrylamide)).

[0175] The DNA template molecule can be attached to a bead or microparticle, for example, as described in U.S. Patent No. 6,172,218 (hereby incorporated by reference). Attachment to beads or microparticles can be useful for sequencing applications. Bead libraries can be prepared, with each bead containing a different DNA sequence. Exemplary libraries and methods for their creation are described in Nature, 437, 376-380 (2005), Science, 309, 5741, 1728-1732 (2005), each of which is hereby incorporated by reference. Sequencing of such arrays of beads using the nucleotides described herein is within the scope of this disclosure.

[0176] The template to be sequenced may form part of an "array" on a solid support, where the array may take any convenient form.Thus, the method of the present disclosure is applicable to all types of high-density arrays, including single molecule arrays, clustered arrays, and bead arrays.The nucleotides labeled with the dye compounds of the present disclosure may be used to sequence the template on essentially any type of array, including, but not limited to, those formed by immobilizing nucleic acid molecules on a solid support.

[0177] However, the nucleotides labeled with the dye compounds of the present disclosure are particularly advantageous in the context of clustered array sequencing. In clustered arrays, distinct regions (often called sites or features) on the array contain multiple polynucleotide template molecules. Generally, multiple polynucleotide molecules are not individually resolvable by optical means, but are instead detected as an ensemble. Depending on how the array is formed, each site on the array may contain multiple copies of one individual polynucleotide molecule (e.g., the site is homogeneous for a particular single-stranded or double-stranded nucleic acid species), or even multiple copies of a small number of different polynucleotide molecules (e.g., multiple copies of two different nucleic acid species). Clustered arrays of nucleic acid molecules can be generated using techniques commonly known in the art. By way of example, WO 98 / 44151 and WO 00 / 18957, each of which is incorporated herein, describe methods for amplifying nucleic acids, in which both the template and the amplification products remain immobilized on a solid support to form arrays consisting of clusters or "colonies" of immobilized nucleic acid molecules. The nucleic acid molecules present on the clustered arrays prepared according to these methods are suitable templates for sequencing using nucleotides labeled with the dye compounds of the present disclosure.

[0178] Nucleotides labeled with the dye compounds of the present application are also useful for sequencing templates on single molecule arrays. As used herein, the term "single molecule array" or "SMA" refers to a population of polynucleotide molecules distributed (or arrayed) on a solid support, with the separation distance of any individual polynucleotide from every other polynucleotide in the population being such that each individual polynucleotide molecule can be individually distinguished. Thus, the target nucleic acid molecules immobilized on the surface of a solid support can be distinguished by optical means in some embodiments. This means that one or more distinct signals, each representing one polynucleotide, are generated within an area that the particular imaging device used can distinguish.

[0179] Single molecule detection can be achieved where the spacing between adjacent polynucleotide molecules on the array is at least 100 nm, more particularly at least 250 nm, even more particularly at least 300 nm, even more particularly at least 350 nm, such that each molecule is individually resolvable and detectable as a single molecule fluorescent spot, the fluorescence from which also exhibits single-step photobleaching.

[0180] The terms "individually resolved" and "individual resolution" are used herein to specify that when visualized, it is possible to distinguish one molecule on an array from its neighboring molecules. The separation between individual molecules on an array is determined in part by the particular technique used to resolve the individual molecules. The general characteristics of single molecule arrays will be understood by reference to WO 00 / 06770 and WO 01 / 57248, each of which is incorporated herein by reference. One use of the labeled nucleotides of the present disclosure is a sequencing-by-synthesis reaction, although the utility of each nucleotide is not limited to such a method. Indeed, the labeled nucleotides described herein can be advantageously used in any sequencing method that requires detection of a fluorescent label attached to a nucleotide incorporated into a polynucleotide.

[0181] In particular, nucleotides labeled with the dye compounds of the present disclosure can be used in automated fluorescent sequencing protocols, in particular fluorescent dye terminator cycle sequencing, which is based on the chain termination sequencing method of Sanger and coworkers. Such methods generally incorporate fluorescently labeled dideoxynucleotides in primer extension sequencing reactions using enzymes and cycle sequencing. The so-called Sanger sequencing method and related protocols (Sanger-type) utilize randomized chain termini with labeled dideoxynucleotides.

[0182] Thus, the present disclosure also encompasses nucleotides labeled with dye compounds that are dideoxynucleotides that do not contain hydroxy groups at both the 3' and 2' positions, such modified deoxynucleotides being suitable for use in Sanger-type sequencing methods and the like.

[0183] It will be appreciated that nucleotides labeled with dye compounds of the present disclosure incorporating a 3' blocking group can also be useful in the Sanger method and related protocols, since the same effect achieved by using dideoxynucleotides can be obtained by using nucleotides with 3' hydroxy blocking groups (both of which prevent the incorporation of subsequent nucleotides). It will be understood that when nucleotides according to the present disclosure having a 3' blocking group are used in Sanger-type sequencing methods, the dye compound or detectable label attached to the nucleotide does not need to be connected via a cleavable linker, since in each instance where a labeled nucleotide of the present disclosure is incorporated, the nucleotide does not need to be subsequently incorporated and therefore the label does not need to be removed from the nucleotide.

[0184] Alternatively, the sequencing methods described herein can also be performed using unlabeled nucleotides and affinity reagents containing the fluorescent dyes described herein. For example, one, two, three, or each of the four different types of nucleotides (e.g., dATP, dCTP, dGTP, and dTTP or dUTP) in the incorporation mixture of step (a) can be unlabeled. Each of the four types of nucleotides (e.g., dNTPs) has a 3' hydroxy blocking group to ensure that only a single base can be added by the polymerase to the 3' end of the primer polynucleotide. After incorporation of the unlabeled nucleotide in step (b), the remaining unincorporated nucleotides are washed away. An affinity reagent is then introduced that specifically recognizes and binds to the incorporated dNTP to provide a labeled extension product containing the incorporated dNTP. The use of unlabeled nucleotides and affinity reagents in sequencing by synthesis is disclosed in WO 2018 / 129214 and WO 2020 / 097607. The modified sequencing method of the present disclosure using unlabeled nucleotides comprises the following steps: (a') contacting the primer polynucleotide / target polynucleotide complex with one or more unlabeled nucleotides (e.g., dATP, dCTP, dGTP, and dTTP or dUTP), where the primer polynucleotide is complementary to at least a portion of the target polynucleotide; (b') incorporating nucleotides into the primer polynucleotide to generate an extended primer polynucleotide; (c') contacting the extended primer polynucleotide with a series of affinity reagents, where one affinity reagent specifically binds to the incorporated unlabeled nucleotide to provide a labeled extended primer polynucleotide / target polynucleotide complex; (d') performing one or more fluorescence measurements of the labeled extended primer polynucleotide / target polynucleotide complex to determine the identity of the incorporated nucleotide.

[0185] In some embodiments of the modified sequencing methods described herein, each of the unlabeled nucleotides in the incorporation mixture contains a 3' blocking group. In further embodiments, the 3' hydroxy blocking group of the incorporated nucleotide is removed prior to the next incorporation cycle. In still further embodiments, the method further comprises removing the affinity reagent from the incorporated nucleotide. In still further embodiments, the 3' hydroxy blocking group and the affinity reagent are removed in the same reaction. In some embodiments, the series of affinity reagents may comprise a first affinity reagent that specifically binds to a first type of nucleotide, a second affinity reagent that specifically binds to a second type of nucleotide, and a third affinity reagent that specifically binds to a third type of nucleotide. In some further embodiments, each of the first, second, and third affinity reagents comprises one or more detectable labels that are spectrally distinguishable. In some embodiments, the affinity reagent may comprise a protein tag, an antibody (including, but not limited to, a binding fragment of an antibody, a single chain antibody, a bispecific antibody, etc.), an aptamer, a knottin, an affimer, or any other known agent that binds to the incorporated nucleotide with suitable specificity and affinity. In one embodiment, at least one affinity reagent is an antibody or a protein tag. In another embodiment, at least one of the first type, second type, and third type affinity reagents is an antibody or a protein tag that includes one or more detectable labels (e.g., multiple copies of the same detectable label), and the detectable label is or includes a bis-boron dye moiety as described herein. EXAMPLES

[0186] Additional embodiments are disclosed in further detail in the following examples, which are not intended to limit the scope of the claims.

[0187] Example 1. Synthesis of dyes containing fused bis-boron heterocycles

[0188] [ka]

[0189] 3,5-Dimethylpyrrole-2-carboxaldehyde (369 mg, 3.00 mmol) and 6-hydrazinonicotinic acid (459 mg, 3.00 mmol) in EtOH (20 mL) were treated with AcOH (100 μL) and heated to reflux for 5 h. The resulting precipitate was filtered under vacuum and washed with EtOH to give the corresponding hydrazone product (compound a) as a yellow solid (595 mg, 77%). 1 H NMR(400MHz,DMSO-d6)δ10.91(s,1H), 10.73(s,1H), 8.59(d,J=2.2Hz,1H), 8.07-7 .90(m,2H), 7.28(d,J=8.9Hz,1H), 5.66(d,J=2.5Hz,1H), 2.19(s,3H), 2.06(s,3H).

[0190] The bis-boron-containing fused pyrido- and pyrazino-heterocycle derivatives of formula (I) were prepared according to the general procedures described herein.

[0191] General Procedure A

[0192] [ka]

[0193] The relevant substituted {2-[(1H-pyrrol-2-yl)methylene]hydrazinyl}pyridine or -pyrazine (1.0 equiv.) in toluene was treated with TEA (18.0 equiv.). The reaction mixture was refluxed for 10 min, after which BF3·OEt2 (20.0 equiv.) was added dropwise. The reaction mixture was stirred at reflux for 5 h. The reaction solvent was removed under vacuum. The crude was dissolved in DCM and the organic layer was washed with H2O and then dried over anhydrous Na2SO4. The crude product was purified by flash chromatography.

[0194] [ka]

[0195] Hydrazone compound a (129 mg, 0.5 mmol) in toluene (10 mL) was treated with Et3N (1.25 mL) and stirred at room temperature for 10 min. Then, BF3OEt (1.5 mL) was added dropwise and the reaction mixture was stirred under reflux for 18 h. The mixture was cooled to room temperature, concentrated in vacuo and then purified by preparative reverse phase HPLC to give compound I-1 (70 μmol, 14%). Mass spectrometry: [M - ]=353

[0196] [ka]

[0197] Compound b was prepared from (Z)-2-chloro-6-(2-((4-ethyl-3,5-dimethyl-1 H-pyrrol-2-yl)methylene)hydrazinyl)pyrazine a according to general procedure A. The crude compound was purified by flash chromatography to give compound a as a bright yellow solid (yield: 61%). MS[M+H] + =374.

[0198] [ka]

[0199] Compound I-8 was prepared from compound c according to general procedure A. The crude compound was purified by flash chromatography to give I-8 as a bright yellow solid (yield: 67%). MS[M+H] + =356, [MH] - =354.

[0200] [ka]

[0201] Compound e was prepared from compound d according to general procedure A. The structure and composition were confirmed by NMR and LCMS.

[0202] [ka]

[0203] Compound I-7 was prepared as a bright yellow solid from 5-((2-(6-chloropyridin-2-yl)hydrazinylidene)methyl)-2,4-dimethyl-1H-pyrrole-3-carboxylic acid (compound f) (1.0 equiv.) according to general procedure A. The crude was purified by flash chromatography to give the final product as a bright yellow solid (yield: 59%). MS[MH] - =387.

[0204] Some new functional derivatives of the bis-boron-containing fused pyrido- and pyrazino-heterocycles of formula (I) can also be prepared by modification of the substituents, for example by replacement of the chlorine atom at the 2-position of the azine ring, for example according to general procedure B.

[0205] General Procedure B

[0206] [ka]

[0207] A mixture of the appropriate chloro-substituted compound (1 eq.), primary, secondary amine or amino acid (1.1 eq.) and TEA (2 eq.) in DMSO was stirred for 5 h at 95° C. The reaction mixture was then diluted with MeCN and 0.1 M TEAB and purified by reverse phase preparative HPLC.

[0208] [ka]

[0209] Compound I-3 was prepared by reacting compound e with sulfoalanine according to general procedure B. The reaction mixture was heated at 95° C. for 5 h to give the final product as a bright yellow solid (yield: 79%). MS[MH] - =504.

[0210] [ka]

[0211] Compound I-6 was prepared by reacting compound b with sulfoalanine according to general procedure B. The reaction mixture was heated at 95° C. for 16 h to give the final product as a bright yellow solid (yield: 5%). MS[MH] - =505.

[0212] [ka]

[0213] Compound I-9 was prepared by reacting compound e with azetidine-3-carboxylic acid according to general procedure B. The reaction mixture was heated at 95° C. for 5 h to give the final product as a bright yellow solid (yield: 99%). MS[MH] - =436.

[0214] [ka]

[0215] Compound I-13 was prepared by reacting compound I-7 with 2-oxa-6-azaspiro[3.3]heptane according to general procedure B. The product was isolated as a bright yellow solid (yield: 69%). MS[MH] - =450.

[0216] General Procedure C

[0217] [ka]

[0218] To a solution of the appropriate bis-difluoroboron-containing fused heterocycle (1 eq.) in DCM, BCl3 (4.5 eq.) was added dropwise. The reaction mixture was stirred for 30 min, then TEA (12.0 eq.) was added, followed by an aliphatic, aromatic mono- or dicarbonic acid, such as acetic acid (8.0 eq.) or malonic acid (4.0 eq.) or their derivatives. The reaction mixture was stirred for 16 h. The crude was filtered through Celite, and the Celite was washed with further DCM. The solvent was removed under vacuum, and the resulting residue was dissolved in MeCN containing 0.1 M TEAB and purified by reverse phase preparative HPLC.

[0219] [ka]

[0220] Compound I-4 was prepared from I-3 using acetic acid according to general procedure C. The reaction mixture was stirred at room temperature for 16 hours to give I-3 (yield: 5%). MS[MH] - =664.

[0221] [ka]

[0222] Compound I-10 was prepared from I-9 according to general procedure C. The reaction mixture was stirred at room temperature for 16 h to give I-10 (yield: 3%). MS[MH] - =596.

[0223] General Procedure D

[0224] [ka]

[0225] To a solution of the appropriate bis-difluoroboron-containing fused heterocycle (1 equiv.) in THF was added the Grignard reagent RMgBr (20.0 equiv.) dropwise at -78°C. The reaction mixture was stirred for 16 h. The solvent was removed under vacuum. The residue was dissolved in DCM and the organic layer was washed with H2O / NH4Cl, then dried over anhydrous Na2SO4 and purified by reverse phase preparative HPLC.

[0226] [ka]

[0227] Compound I-14 was prepared by reacting I-9 with PhMgBr according to general procedure D (yield: 12%). MS[MH] - =668.

[0228] [ka]

[0229] Compound I-15 was prepared by reacting I-3 with PhMgBr according to general procedure D (yield: 3%). MS[MH] - =736.

[0230] The fluorescence spectra of some exemplary dyes disclosed herein are summarized in Table 1 below.

[0231] [Table 1]

[0232] Example 2. Synthesis of ffN labeled with dyes containing bis-boron fused heterocycles The bis-boron-containing fused heterocyclic compounds described herein can be used for nucleotide labeling by coupling reactions with appropriate functionalized nucleotide derivatives containing an amino moiety.

[0233] General Procedure E The described bis-boron-containing fused heterocycles can be used for nucleotide labeling by coupling reactions with their appropriately functionalized nucleotides containing an amino moiety. The dye of formula (I) was dissolved in anhydrous N,N'-dimethylacetamide (DMA). N,N-diisopropylethylamine (DIPEA) was added, followed by TNTU. The reaction was stirred at room temperature under nitrogen for 30 min. The activated bis-boron dye solution was added to the 3'-blocked 2'-deoxynucleoside triphosphate-linker in triethylammonium bicarbonate (TEAB) solution and the reaction was stirred at room temperature for 18 h. The crude product was first purified by ion exchange chromatography on DEAE-Sephadex A25. The fractions containing the functionalized nucleotide were pooled and the solvent was evaporated to dryness under reduced pressure. The crude material was further purified by preparative scale RP-HPLC using a YMC-Pack-Pro C18 column. The final compound was characterized by LC-MS, analytical RP-HPLC and UV-visible spectroscopy.

[0234] [ka]

[0235] ffC-sPA-I-1 was prepared from I-1 based on the general procedure for ffN coupling (yield: 14%). MS[M - ]=1257.

[0236] [ka]

[0237] ffC-sPA-I-3 was prepared from I-3 based on the general procedure for ffN coupling (yield: 6%). MS[M - ]=1408.

[0238] [ka]

[0239] ffA-sPA-I-4 was prepared from I-4 based on the general procedure for ffN coupling (yield: 24%) MS [M-2H] 2- =795.

[0240] [ka]

[0241] ffA-sPA-I-6 was prepared from I-6 based on the general procedure for ffN coupling (yield: 8%). MS [M-2H] 2- =717.

[0242] [ka]

[0243] ffA-sPA-I-9 was prepared from I-9 based on the general procedure for ffN coupling (yield: 17%). MS [M-2H] 2- =681.

[0244] [ka]

[0245] ffA-sPA-I-13 was prepared from I-13 based on the general procedure for ffN coupling. The reaction mixture was heated at 40° C. for 48 h, and the final equivalents of A-SpA and DIPEA were (2.0 equiv.) and (20.0 equiv.), respectively, due to the slow coupling reaction. MS[MH] - =1378, [M+H] + =1380.

[0246] The fluorescence spectra of exemplary ffNs disclosed herein are summarized in Table 2 below.

[0247] [Table 2]

[0248] Example 3. ffN Spectral Characteristics Comparison In this example, the spectral properties of fully functionalized A nucleotides (ffA) conjugated with bis-boron dye I-4 (A-sPA-I-4) were characterized. Figure 1 shows the emission spectra of commercially available fully functionalized C nucleotides (ffC) labeled with A-spA-I-4 and reference dye A (C-sPA-reference dye A) in Universal Scan Mix (USM, 1 M Tris pH 7.5, 0.05% TWEEN®, 20 mM sodium ascorbate, 10 mM ethyl gallate). Spectra were acquired on an Agilent Cary 100 UV-Vis spectrophotometer and a Cary Eclipse fluorescence spectrophotometer using quartz or plastic cuvettes. A-sPA-I-4 was observed to have a shorter Stokes shift compared to reference dye A.

[0249] [ka]

[0250] Example 4. Stability of bis-boron dyes The stability of compounds I-1 and I-3 was evaluated and compared to commercially available ffC labeled with reference dye A by incubating the compounds in a loading buffer containing 50 mM ethanolamine at 37° C. in the dark for 2 days. The fluorescence intensity of the solutions was measured on an Agilent Cary 100 UV-Vis spectrophotometer and a Cary Eclipse fluorescence spectrophotometer using quartz cuvettes. In addition, aliquots of the solutions were taken and analyzed by analytical HPLC. Figure 2 shows that the fluorescence intensity of I-1 and I-3 decreased very slowly over time compared to C-sPA-reference dye A, indicating that the bis-boron dyes I-1 and I-3 were more stable compared to reference dye A under the same conditions.

[0251] Example 5. Sequencing experiments on the Illumina MiSeq™ platform The ffA labeled with the bis-boron dye I-4 was tested on an Illumina MiSeq™ instrument set to take a first image with blue excitation light (about 450 nm) and a second image with green excitation light (about 520 nm). The integration mix used in the experiment included the following five ffNs: A-sPA-I-4, ffA labeled with the known polymethine green dye NR550S0 (A-sPA-NR550S0), ffC labeled with a blue coumarin dye (C-sPA-reference dye B), ffT labeled with the green dye NR550S0 (T-sPA-NR550S0), and unlabeled ffG (dark G) in 50 mM ethanolamine buffer, pH 9.6, 50 mM NaCl, 1 mM EDTA, 0.2% CHAPS, 4 mM MgSO4, and DNA polymerase. FIG. 3 shows the sequencing matrix percent phasing of the ffN set containing ffA-spA-I-4 compared to the commercially available Reference 1 and Reference 2 ffN sets. The Reference 1 ffN set includes the following ffNs: Dark G, T-LN3-AF550POPOS0, C-sPA-Reference Dye A, C-LN3-SO7181, A-BL-Reference Dye A, A-BL-NR550S0. The Reference 2 ffN set includes the following ffNs: Dark G, T-LN3-AF550POPOS0, C-sPA-Reference Dye B, C-LN3-SO7181, A-sPA-BL-Reference Dye B, A-sPA-BL-NR550S0. The structure of C-sPA-Reference Dye B is as follows:

[0252] [ka]

[0253] The percent phasing of the ffN set containing bis-boron dye-labeled ffA was observed to be less than 0.1% after 26 cycles, however, increasing the light dose also increased the phasing value.

[0254] Figures 4A and 4B are scatter plots obtained for the incorporation mix containing ffA-spA-I-3 at cycle 26. Figures 4C and 4D are scatter plots obtained for the incorporation mix containing ffA-spA-I-4 at cycle 26. It was observed that five times (5x) the light dose caused photobleaching of the cloud of ffA labeled with I-3 (see Figure 4B, upper right quadrant). However, when the fluoro group was replaced with -OAc, the photostability of ffA labeled with I-4 was greatly improved as shown in Figure 4D, upper right quadrant.

Claims

1. A compound of formula (I) 【Chemical 1】 or a salt or mesomeric form thereof, wherein R 1 、 R 2 、 R 3 and R 4 each independently is H, unsubstituted or substituted C 1 -C 6 alkyl, C 1 -C 6 alkoxy, C 2 -C 6 alkenyl, C 2 -C 6 alkynyl, C 1 -C 6 haloalkyl, C 1 -C 6 haloalkoxy, C 1 -C 6 hydroxyalkyl, (C 1 -C 6 alkoxy)(C 1 -C 6 alkyl), unsubstituted or substituted amino, halo, cyano, carboxyl, hydroxy, nitro, sulfonyl, sulfino, sulfo, sulfonate, S-sulfonamide, N-sulfonamide, unsubstituted or substituted C 3 -C 10 carbocyclic, unsubstituted or substituted C 6 -C 10 aryl, unsubstituted or substituted 5- to 10-membered heteroaryl, or unsubstituted or substituted 3- to 10-membered heterocyclic, R a 、R b 、R c and R d each independently is halo, cyano, C 1 -C 6 alkyl, C 1 -C 6 haloalkyl, C 1 -C 6 alkoxy, C 1 -C 6 haloalkoxy, C 6 -C 10 aryl, C 6 -C 10 aryloxy, or -O-C(=O)R 5 wherein, Alternatively, R a and R b both are -O-C(=O)R 5 In the case where two Rs 5 together with the atoms to which they are attached form an unsubstituted or substituted 6- to 10-membered heterocyclyl, and both R c and R d both are -O-C(=O)R 5 In the case where two Rs 5 together with the atoms to which they are attached form an unsubstituted or substituted 6- to 10-membered heterocyclyl, R 5 is unsubstituted or substituted C 1 -C 6 alkyl, and Ring A is a 6- to 10-membered heteroaryl optionally substituted with one or more Rs 6 and is optionally substituted with one or more Rs Each R 6 is independently unsubstituted or substituted C 1 -C 6 alkyl, C 1 -C 6 alkoxy, C 2 -C 6 alkenyl, C 2 -C 6 alkynyl, C 1 -C 6 haloalkyl, C 1 -C 6 haloalkoxy, C 1 -C 6 hydroxyalkyl, (C 1 -C 6 alkoxy)(C 1 -C 6 alkyl), -NR 7 R 8 、halo, cyano, carboxyl, hydroxy, nitro, sulfonyl, sulfino, sulfo, sulfonate, S-sulfonamide, N-sulfonamide, unsubstituted or substituted C 3 -C 10 carbocyclic, unsubstituted or substituted C 6 -C 10 aryl, unsubstituted or substituted 5- to 10-membered heteroaryl, or unsubstituted or substituted 3- to 10-membered heterocyclyl, and R 7 and R 8 each independently is H, unsubstituted or substituted C 1 -C 6 alkyl, or R 7 and R 8 together with the nitrogen atom to which they are attached form an unsubstituted or substituted 3- to 10-membered heterocyclyl, However, R 1 、R 2 、R 3 、R 4 and at least one of ring A contains a carboxyl group, a compound, a salt thereof or a mesomeric form.

2. has a structure of formula (Ia) or (Ib) [Chemical 2] or a salt or mesomeric form thereof, wherein m is 0, 1, 2, or 3, the compound according to claim 1.

3. has a structure of formula (Ic), (Id) or (Ie) 【Chemical Formula 3】 or a salt or mesomeric form thereof, the compound according to claim 2.

4. Each R 6 is independently halo, cyano, carboxyl, unsubstituted or substituted C 1 -C 6 alkyl, unsubstituted phenyl, phenyl substituted with carboxyl, unsubstituted 5-membered heteroaryl, 5-membered heteroaryl substituted with carboxyl, or -NR 7 R 8 The compound according to claim 1, wherein is

5. R 6 is -NR 7 R 8 where R 7 is H, and R 8 is C substituted with one or more substituents selected from the group consisting of carboxyl, sulfo and sulfonate 1 -C 6 alkyl, or R 7 and R 8 together with the nitrogen atom to which they are attached form a 3 - to 10 - membered heterocyclyl optionally substituted with carboxyl, the compound according to claim 4.

6. R 6 is 【Chemical Formula 4】 wherein each of the ring structures is optionally substituted with carboxyl, the compound according to claim 5.

7. R 1 , R 2 and R 3 each of which is an independent H or unsubstituted C1-C6 alkyl, the compound according to claim 1.

8. R 1 and R 3 each is methyl, and R 2 is ethyl, the compound according to claim 1.

9. R 1 , R 2 and R 3 of which two are H or unsubstituted C 1 -C 6 -alkyl, and R 1 , R 2 and R 3 of which one is halo, carboxyl or C-substituted with carboxyl 1 -C 6 -alkyl, the compound according to claim 1.

10. R 4 is H, unsubstituted C 1 -C 6 alkyl, C1-C6 alkyl substituted with carboxyl, or phenyl substituted with carboxyl, the compound according to claim 1.

11. R a and R b each independently is fluoro, cyano, methyl, trifluoromethyl, methoxy, or -O-acyl (-OC(=O)CH 3 )), the compound according to claim 1.

12. R a and R b both are -OC(=O)R 5 and two Rs 5 together with the atoms to which they are attached form a structure 【Chemical Formula 5】 forms a 6-membered heterocyclyl having, the compound according to claim 1.

13. R c and R d each independently is fluoro, cyano, methyl, trifluoromethyl, methoxy, or -O-acyl (-OC(=O)CH 3 ) of the compound according to claim 1.

14. R c and R d both are -OC(=O)R 5 and the two Rs 5 together with the atoms to which they are attached form a structure ​ forms a 6-membered heterocyclyl having, the compound according to claim 1.

15. 【Chemical Formula 7】 【Chemical Formula 8】 【Chemical Formula 9】 or selected from the group consisting of salts or mesomeric forms thereof, the compound according to claim 1.

16. A nucleotide labeled with the compound of formula (I) according to claim 1.

17. The labeled nucleotide according to claim 16, wherein the compound of formula (I) is attached to the nucleotide via a carboxyl group of the compound of formula (I).

18. The labeled nucleotide according to claim 16, comprising a 3'-hydroxy blocking group covalently attached to the ribose sugar or deoxyribose sugar of the nucleotide.

19. An oligonucleotide or polynucleotide incorporated with the labeled nucleotide according to claim 16.

20. The oligonucleotide or polynucleotide according to claim 19, wherein the oligonucleotide or polynucleotide is at least partially complementary to a target polynucleotide immobilized on the surface of a solid support and hybridizes thereto.

21. The oligonucleotide or polynucleotide according to claim 20, wherein the solid support comprises an array of a plurality of target polynucleotides immobilized thereon.

22. A kit comprising the first type of labeled nucleotide according to claim 16.

23. The kit according to claim 22, wherein the kit contains four types of nucleotides, the first type of nucleotide is the labeled nucleotide according to any one of claims 16 to 18, the second type of nucleotide carries a second label, the third type of nucleotide carries a third label, and the fourth type of nucleotide is unlabeled (dark).

24. The kit according to claim 22, wherein the kit contains four types of nucleotides, the first type of nucleotide is the labeled nucleotide according to any one of claims 16 to 18, the second type of nucleotide carries a second label, the third type of nucleotide contains a mixture of third type nucleotides carrying two labels, and the fourth type of nucleotide is unlabeled (dark).

25. The kit according to claim 22, further comprising a DNA polymerase and one or more buffer compositions.

26. A method for determining the sequences of a plurality of target polynucleotides, comprising: (a) contacting a solid support with a solution containing a sequencing primer under hybridization conditions, wherein the solid support contains a plurality of different target polynucleotides immobilized thereon, and the sequencing primer is complementary to at least a portion of the target polynucleotide; (b) contacting the solid support with an aqueous solution containing a DNA polymerase and one or more of four different types of nucleotides under conditions suitable for DNA polymerase-mediated primer extension, wherein one of the nucleotides is the nucleotide according to claim 18 having a 3'-hydroxy blocking group covalently bonded to the deoxyribose sugar of the nucleotide, and each type of nucleotide has a 3'-hydroxy blocking group; (c) incorporating one type of nucleotide into the sequencing primer to generate an extended copy polynucleotide; (d) performing one or more fluorescence measurements of the extended copy polynucleotide to determine the identity of the incorporated nucleotide.