Water-soluble cyanine derivatives and their use in nucleic acid sequencing methods
By developing water-soluble cyanine derivatives as fluorescent dyes, the problem of insufficient performance in nucleic acid sequencing applications in existing technologies has been solved, and sequencing efficiency and accuracy have been improved.
Patent Information
- Application Number
- CN202380093132.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-07
- Filing Date
- 2023-12-07
- Publication Date
- 2025-09-12
AI Technical Summary
Existing cyanine dyes have insufficient performance in nucleic acid sequencing applications, and more suitable fluorescent dyes need to be developed to improve sequencing efficiency and accuracy.
Provided are a water-soluble cyanine derivative and its ionic derivatives, isomers or salts, which are used as fluorescent dyes in nucleic acid sequencing methods, and whose fluorescence performance is improved through specific chemical structure and functional group design.
The efficiency and accuracy of nucleic acid sequencing are improved, and the application effect of fluorescent dyes in sequencing methods is enhanced.
Smart Images

Figure CN120641508A_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to and the benefit of U.S. Provisional Application No. 63 / 430,993, filed December 7, 2022, which is incorporated herein by reference in its entirety for all purposes. Background Art
[0003] Cyanine dyes are particularly commonly used fluorophores and are widely used in many biological applications, including sequencing applications. Therefore, there is a need to develop cyanine derivatives that can be used in sequencing applications. The present disclosure addresses this need. Summary of the Invention
[0004] In some aspects, the present disclosure provides a compound of formula (I), (II) or (III):
[0005]
[0006] Ionic derivatives thereof, isomers thereof or salts thereof.
[0007] In some aspects, the present disclosure provides a sequencing method disclosed herein using a compound of the present disclosure (eg, as a fluorescent dye).
[0008] In some aspects, the present disclosure provides a compound of the present disclosure for use in the sequencing methods disclosed herein (eg, as a fluorescent dye).
[0009] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those generally understood by those of ordinary skill in the art to which this disclosure belongs. In this specification, unless the context clearly stipulates otherwise, the singular also includes the plural. Although methods and materials similar to or equivalent to the methods and materials described herein can be used to practice or test this disclosure, suitable methods and materials are described below. All publications, patent applications, patents and other references mentioned herein are incorporated by reference. The references cited herein are not considered to be prior art of the claimed invention. In the event of a conflict, this specification (including definitions) will be used as the standard. In addition, materials, methods and examples are only illustrative and are not intended to be restrictive. In the event of a conflict between the chemical structure and the name of the compound disclosed herein, the chemical structure will be used as the standard.
[0010] Other features and advantages of the disclosure will become apparent from the following detailed description and from the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] This patent or application file contains at least one drawing drawn in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the U.S. Patent and Trademark Office upon request and payment of the necessary fee.
[0012] Figure 1 is a schematic diagram of an exemplary low binding support comprising a glass substrate and alternating hydrophilic coatings covalently or non-covalently adhered to the glass, and which further comprises chemically reactive functional groups that serve as attachment sites for oligonucleotide primers (e.g., capture oligonucleotides). In an alternative embodiment, the support can be made of any material such as glass, plastic, or a polymeric material.
[0013] Figure 2 are schematic diagrams of various exemplary configurations of multivalent molecules. Left (Class I): Schematic diagrams of multivalent molecules having a "starburst" or "helter-skelter" configuration. Center (Class II): Schematic diagrams of multivalent molecules having a dendritic configuration. Right (Class III): Schematic diagrams of various multivalent molecules formed by reacting streptavidin with 4-arm or 8-arm PEG-NHS with biotin and dNTPs. The nucleotide unit is designated as 'N', biotin is designated as 'B', and streptavidin is designated as 'SA'.
[0014] Figure 3 is a schematic diagram of an exemplary multivalent molecule comprising a universal core attached to multiple nucleotide arms.
[0015] Figure 4 is a schematic diagram of an exemplary multivalent molecule comprising a dendritic core attached to multiple nucleotide arms.
[0016] Figure 5 A schematic diagram of an exemplary multivalent molecule comprising a core attached to multiple nucleotide arms, wherein the nucleotide arms comprise biotin, a spacer, a linker, and nucleotide units is shown.
[0017] Figure 6 is a schematic diagram of an exemplary nucleotide arm comprising a core attachment portion, a spacer, a linker, and a nucleotide unit.
[0018] Figure 7 Shown are the chemical structures of exemplary spacers (top), as well as the chemical structures of various exemplary linkers, including an 11-atom linker, a 16-atom linker, a 23-atom linker, and an N3 linker (bottom).
[0019] Figure 8 The chemical structures of various exemplary linkers (including linkers 1 to 9) are shown.
[0020] Figure 9A The chemical structures of various exemplary linkers joined / attached to the nucleotide units are shown.
[0021] Figure 9B The chemical structures of various exemplary linkers joined / attached to the nucleotide units are shown.
[0022] Figure 9C The chemical structures of various exemplary linkers joined / attached to the nucleotide units are shown.
[0023] Figure 9D The chemical structures of various exemplary linkers joined / attached to the nucleotide units are shown.
[0024] Figure 10 The chemical structure of an exemplary biotinylated nucleotide arm is shown. In this example, the nucleotide unit is linked to the linker via a propargylamine attachment at the 5-position of the pyrimidine base or the 7-position of the purine base.
[0025] Figure 11 is a schematic diagram of a guanine tetrad (eg, a G-tetrad).
[0026] Figure 12 is a schematic diagram of an exemplary intramolecular G-quadruplex structure. DETAILED DESCRIPTION
[0027] The present disclosure relates to compounds of the formula disclosed herein, their ionic derivatives, their isomers, and salts thereof. Without wishing to be bound by theory, these compounds can be used as dyes (e.g., fluorescent dyes) and thus can be used in sequencing methods. The present disclosure also relates to conjugates of the dyes, methods of using the dyes and their conjugates.
[0028] Compounds of the present disclosure
[0029] In some aspects, the present disclosure provides a compound of formula (I), (II) or (III):
[0030]
[0031]
[0032] An ionic derivative thereof, an isomer thereof, or a salt thereof, wherein:
[0033] n is 0 or 1;
[0034] m is 0 or 1;
[0035] p is 0 or 1;
[0036] R X and R Z are independently H, halogen, C1-12 Alkyl, C 1-12 Alkenyl, C 1-12 Alkynyl, C6-C 10 aryl or 5- to 10-membered heteroaryl; or R X and R Z Together with the atoms to which they are attached, they form C6-C 10 Arylene or C5-C 10 cycloalkylene;
[0037] R Y It is H, halogen, C 1-12 Alkyl, C 1-12 Alkenyl, C 1-12 Alkynyl, C6-C 10 aryl or 5- to 10-membered heteroaryl;
[0038] T A Is -S-, -O- or -C(R TA )2-;
[0039] Each R TA are independently H, halogen, C 1-12 Alkyl, C 1-12 Alkenyl, C 1-12 Alkynyl, C6-C 10 Aryl or 5 to 10 membered heteroaryl; wherein the C 1-12 Alkyl, C 1-12 Alkynyl, C6-C 10 Aryl or 5- to 10-membered heteroaryl is optionally substituted with one or more -S(=O)2OH or -C(=O)OH;
[0040] T B Is -S-, -O- or -C(R TB )2-;
[0041] Each R TB are independently H, halogen, C 1-12 Alkyl, C 1-12 Alkenyl, C 1-12 Alkynyl, C6-C 10 Aryl or 5 to 10 membered heteroaryl; wherein the C 1-12 Alkyl, C 1-12 Alkynyl, C6-C 10 Aryl or 5- to 10-membered heteroaryl is optionally substituted with one or more -S(=O)2OH or -C(=O)OH;
[0042] R NA It is C 1-12 Alkyl, C 1-12 Alkenyl or C 1-12 Alkynyl, wherein the C 1-12 Alkyl, C1-12 Alkenyl or C 1-12 Alkynyl is optionally substituted with one or more R NA1 replace;
[0043] Each R NA1 are independently -S(=O)2OH, -C(=O)OH, -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 Alkynyl), wherein the -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 alkynyl) is optionally substituted by one or more R NA2 replace;
[0044] Each R NA2 are independently -S(=O)2OH, -C(=O)OH, -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 Alkynyl), wherein the -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 alkynyl) is optionally substituted by one or more R NA3 replace;
[0045] Each R NA3 are independently -S(=O)2OH, -C(=O)OH, -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 Alkynyl), wherein the -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 alkynyl) is optionally substituted with one or more -S(=O)2OH or -C(=O)OH;
[0046] R NB It is C 1-12 Alkyl, C 1-12 Alkenyl or C 1-12 Alkynyl, wherein the C 1-12 Alkyl, C 1-12 Alkenyl or C1-12 Alkynyl is optionally substituted with one or more R NB1 replace;
[0047] Each R NB1 are independently -S(=O)2OH, -C(=O)OH, -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 Alkynyl), wherein the -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 alkynyl) is optionally substituted by one or more R NB2 replace;
[0048] Each R NB2 are independently -S(=O)2OH, -C(=O)OH, -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 Alkynyl), wherein the -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 alkynyl) is optionally substituted by one or more R NB3 replace;
[0049] Each R NB3 are independently -S(=O)2OH, -C(=O)OH, -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 Alkynyl), wherein the -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 alkynyl) is optionally substituted with one or more -S(=O)2OH or -C(=O)OH;
[0050] R 1A 、R 2A 、R 3A 、R 4A 、R 5A 、R 6A 、R 7A and R 8Aare independently H, halogen, -S(=O)2OH, -C(=O)OH, C 1-12 Alkyl, C 1-12 Alkenyl, C 1-12 Alkynyl, -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 Alkynyl), wherein the C 1-12 Alkyl, C 1-12 Alkenyl, C 1-12 Alkynyl, -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 alkynyl) is optionally substituted by one or more R AR replace;
[0051] Each R AR are independently -S(=O)2OH, -C(=O)OH, -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 Alkynyl), wherein the -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 alkynyl) is optionally substituted by one or more R AR1 replace;
[0052] Each R AR1 are independently -S(=O)2OH, -C(=O)OH, -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 Alkynyl), wherein the -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 alkynyl) is optionally substituted with one or more -S(=O)2OH or -C(=O)OH;
[0053] R 1B 、R 2B 、R 3B 、R 4B 、R 5B 、R 6B 、R 7Band R 8B are independently H, halogen, -S(=O)2OH, -C(=O)OH, C 1-12 Alkyl, C 1-12 Alkenyl, C 1-12 Alkynyl, -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 Alkynyl), wherein the C 1-12 Alkyl, C 1-12 Alkenyl, C 1-12 Alkynyl, -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 alkynyl) is optionally substituted by one or more R BR replace;
[0054] Each R AR are independently -S(=O)2OH, -C(=O)OH, -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 Alkynyl), wherein the -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 alkynyl) is optionally substituted by one or more R BR1 replace; and
[0055] Each R BR1 are independently -S(=O)2OH, -C(=O)OH, -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 Alkynyl), wherein the -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 alkynyl) is optionally substituted with one or more -S(=O)2OH or -C(=O)OH.
[0056] In some embodiments, Formula (I), (II) or (III) or an ionic derivative thereof, an isomer thereof or a salt thereof:
[0057] An ionic derivative thereof, an isomer thereof, or a salt thereof, wherein:
[0058] n is 0 or 1;
[0059] m is 0 or 1;
[0060] p is 0 or 1;
[0061] R X and R Z are independently H, halogen, C 1-12 Alkyl, C 1-12 Alkenyl, C 1-12 Alkynyl, C6-C 10 aryl or 5- to 10-membered heteroaryl; or R X and R Z Together with the atoms to which they are attached, they form C6-C 10 Arylene or C5-C 10 cycloalkylene;
[0062] R Y It is H, halogen, C 1-12 Alkyl, C 1-12 Alkenyl, C 1-12 Alkynyl, C6-C 10 aryl or 5- to 10-membered heteroaryl;
[0063] T A Is -S- or -C(R TA )2-;
[0064] Each R TA are independently H, halogen, C 1-12 Alkyl, C 1-12 Alkenyl, C 1-12 Alkynyl, C6-C 10 Aryl or 5 to 10 membered heteroaryl; wherein the C 1-12 Alkyl, C 1-12 Alkynyl, C6-C 10 Aryl or 5- to 10-membered heteroaryl is optionally substituted with one or more -S(=O)2OH or -C(=O)OH;
[0065] T B Is -S- or -C(R TB )2-;
[0066] Each R TB are independently H, halogen, C 1-12 Alkyl, C 1-12 Alkenyl, C 1-12 Alkynyl, C6-C 10 Aryl or 5 to 10 membered heteroaryl; wherein the C 1-12 Alkyl, C 1-12 Alkynyl, C6-C 10Aryl or 5- to 10-membered heteroaryl is optionally substituted with one or more -S(=O)2OH or -C(=O)OH;
[0067] R NA It is C 1-12 Alkyl, C 1-12 Alkenyl or C 1-12 Alkynyl, wherein the C 1-12 Alkyl, C 1-12 Alkenyl or C 1-12 Alkynyl is optionally substituted with one or more R NA1 replace;
[0068] Each R NA1 are independently -S(=O)2OH, -C(=O)OH, -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 Alkynyl), wherein the -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 alkynyl) is optionally substituted by one or more R NA2 replace;
[0069] Each R NA2 are independently -S(=O)2OH, -C(=O)OH, -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 Alkynyl), wherein the -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 alkynyl) is optionally substituted by one or more R NA3 replace;
[0070] Each R NA3 are independently -S(=O)2OH, -C(=O)OH, -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 Alkynyl), wherein the -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12alkynyl) is optionally substituted with one or more -S(=O)2OH or -C(=O)OH;
[0071] R NB It is C 1-12 Alkyl, C 1-12 Alkenyl or C 1-12 Alkynyl, wherein the C 1-12 Alkyl, C 1-12 Alkenyl or C 1-12 Alkynyl is optionally substituted with one or more R NB1 replace;
[0072] Each R NB1 are independently -S(=O)2OH, -C(=O)OH, -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 Alkynyl), wherein the -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 alkynyl) is optionally substituted by one or more R NB2 replace;
[0073] Each R NB2 are independently -S(=O)2OH, -C(=O)OH, -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 Alkynyl), wherein the -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 alkynyl) is optionally substituted by one or more R NB3 replace;
[0074] Each R NB3 are independently -S(=O)2OH, -C(=O)OH, -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 Alkynyl), wherein the -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12alkynyl) is optionally substituted with one or more -S(=O)2OH or -C(=O)OH;
[0075] R 1A 、R 2A 、R 3A 、R 4A 、R 5A 、R 6A 、R 7A and R 8A are independently H, halogen, -S(=O)2OH, -C(=O)OH, C 1-12 Alkyl, C 1-12 Alkenyl, C 1-12 Alkynyl, -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 Alkynyl), wherein the C 1-12 Alkyl, C 1-12 Alkenyl, C 1-12 Alkynyl, -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 alkynyl) is optionally substituted by one or more R AR replace;
[0076] Each R AR are independently -S(=O)2OH, -C(=O)OH, -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 Alkynyl), wherein the -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 alkynyl) is optionally substituted by one or more R AR1 replace;
[0077] Each R AR1 are independently -S(=O)2OH, -C(=O)OH, -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 Alkynyl), wherein the -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12alkenyl) or -C(=O)-NH-(C 1-12 alkynyl) is optionally substituted with one or more -S(=O)2OH or -C(=O)OH;
[0078] R 1B 、R 2B 、R 3B 、R 4B 、R 5B 、R 6B 、R 7B and R 8B are independently H, halogen, -S(=O)2OH, -C(=O)OH, C 1-12 Alkyl, C 1-12 Alkenyl, C 1-12 Alkynyl, -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 Alkynyl), wherein the C 1-12 Alkyl, C 1-12 Alkenyl, C 1-12 Alkynyl, -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 alkynyl) is optionally substituted by one or more R BR replace;
[0079] Each R AR are independently -S(=O)2OH, -C(=O)OH, -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 Alkynyl), wherein the -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 alkynyl) is optionally substituted by one or more R BR1 replace; and
[0080] Each R BR1 are independently -S(=O)2OH, -C(=O)OH, -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 Alkynyl), wherein the -C(=O)-NH-(C 1-12alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 alkynyl) is optionally substituted with one or more -S(=O)2OH or -C(=O)OH.
[0081] In some embodiments, the compound has Formula (IA), (II-A), (III-A), or (IV-A):
[0082]
[0083] Ionic derivatives thereof, isomers thereof or salts thereof.
[0084] In some embodiments, the compound has Formula (IB), (II-B), (III-B), or (IV-b):
[0085]
[0086] Ionic derivatives thereof, isomers thereof or salts thereof.
[0087] In some embodiments, the compound is of any of the formulae described herein, an ionic derivative thereof, an isomer thereof, or a salt thereof, wherein:
[0088] n is 0 or 1;
[0089] m is 0 or 1;
[0090] Each R TA is independently C optionally substituted with one or more -S(=O)2OH or -C(=O)OH 1-12 alkyl;
[0091] Each R TB is independently C optionally substituted with one or more -S(=O)2OH or -C(=O)OH 1-12 alkyl;
[0092] R NA is optionally replaced by one or more R NA1 Substituted C 1-12 alkyl;
[0093] Each R NA1 are independently -S(=O)2OH, -C(=O)OH or -C(=O)-NH-(C 1-12 alkyl), wherein -C(=O)-NH-(C 1-12 alkyl) is optionally substituted with one or more -S(=O)2OH, -C(=O)OH;
[0094] R NB is optionally replaced by one or more RNB1 Substituted C 1-12 alkyl;
[0095] Each R NB1 are independently -S(=O)2OH, -C(=O)OH or -C(=O)-NH-(C 1-12 alkyl), wherein -C(=O)-NH-(C 1-12 alkyl) is optionally substituted with one or more -S(=O)2OH, -C(=O)OH;
[0096] R 1A 、R 3A 、R 5A 、R 6A and R 8A are each independently H, halogen, -S(=O)2OH or -C(=O)OH;
[0097] R 2A 、R 4A 、R 5A and R 7A are independently H, halogen, -S(=O)2OH, -C(=O)OH, C 1-12 Alkyl or -C(=O)-NH-(C 1-12 alkyl), wherein C 1-12 Alkyl or -C(=O)-NH-(C 1-12 alkyl) is optionally substituted with -S(=O)2OH or -C(=O)OH;
[0098] R 1B 、R 3B 、R 5B 、R 6B and R 8B are each independently H, halogen, -S(=O)2OH or -C(=O)OH; and
[0099] R 2B 、R 4B 、R 5B and R 7B are independently H, halogen, -S(=O)2OH, -C(=O)OH, C 1-12 Alkyl or -C(=O)-NH-(C 1-12 alkyl), wherein C 1-12 Alkyl or -C(=O)-NH-(C 1-12 alkyl) is optionally substituted with -S(=O)2OH or -C(=O)OH.
[0100] In some embodiments, the compound is of any of the formulae described herein, an ionic derivative thereof, an isomer thereof, or a salt thereof, wherein:
[0101] n is 0 or 1;
[0102] m is 0 or 1;
[0103] Each R TA is independently C optionally substituted with one or more -S(=O)2OH or -C(=O)OH 1-12 alkyl;
[0104] Each R TB is independently C optionally substituted with one or more -S(=O)2OH or -C(=O)OH 1-12 alkyl;
[0105] R NA is C optionally substituted with one or more -S(=O)2OH or -C(=O)OH 1-12 alkyl;
[0106] R NB is C optionally substituted with one or more -S(=O)2OH or -C(=O)OH 1-12 alkyl;
[0107] R 1A 、R 3A 、R 5A 、R 6A and R 8A are each independently H or halogen;
[0108] R 2A 、R 4A 、R 5A and R 7A are independently H, halogen, -S(=O)2OH, -C(=O)OH, C 1-12 Alkyl or -C(=O)-NH-(C 1-12 alkyl), wherein C 1-12 Alkyl or -C(=O)-NH-(C 1-12 alkyl) is optionally substituted with -S(=O)2OH or -C(=O)OH;
[0109] R 1B 、R 3B 、R 5B 、R 6B and R 8B are each independently H or halogen; and
[0110] R 2B 、R 4B 、R 5B and R 7B are independently H, halogen, -S(=O)2OH, -C(=O)OH, C 1-12Alkyl or -C(=O)-NH-(C 1-12 alkyl), wherein C 1-12 Alkyl or -C(=O)-NH-(C 1-12 alkyl) is optionally substituted with -S(=O)2OH or -C(=O)OH.
[0111] In some embodiments, the compound is of any of the formulae described herein, an ionic derivative thereof, an isomer thereof, or a salt thereof, wherein:
[0112] n is 0 or 1;
[0113] m is 0 or 1;
[0114] Each R TA Independently C 1-12 alkyl;
[0115] Each R TB Independently C 1-12 alkyl;
[0116] R NA is C optionally substituted with one or more -S(=O)2OH or -C(=O)OH 1-12 alkyl;
[0117] R NB is C optionally substituted with one or more -S(=O)2OH or -C(=O)OH 1-12 alkyl;
[0118] R 1A 、R 3A 、R 5A 、R 6A and R 8A are each independently H or halogen;
[0119] R 2A 、R 4A 、R 5A and R 7A Each independently is -S(=O)2OH, -C(=O)OH, C 1-12 Alkyl or -C(=O)-NH-(C 1-12 alkyl), wherein C 1-12 Alkyl or -C(=O)-NH-(C 1-12 alkyl) is optionally substituted with -S(=O)2OH or -C(=O)OH;
[0120] R 1B 、R 3B 、R 5B 、R 6B and R 8B are each independently H or halogen; and
[0121] R 2B 、R 4B 、R 5B and R 7B Each independently is -S(=O)2OH, -C(=O)OH, C 1-12 Alkyl or -C(=O)-NH-(C 1-12 alkyl), wherein C 1-12 Alkyl or -C(=O)-NH-(C 1-12 alkyl) is optionally substituted with -S(=O)2OH or -C(=O)OH.
[0122] In some embodiments, the compound is of any of the formulae described herein, an ionic derivative thereof, an isomer thereof, or a salt thereof, wherein:
[0123] n is 0 or 1;
[0124] m is 0 or 1;
[0125] Each R TA Independently C 1-12 alkyl;
[0126] Each R TB Independently C 1-12 alkyl;
[0127] R NA is C optionally substituted with one or more -S(=O)2OH or -C(=O)OH 1-12 alkyl;
[0128] R NB is C optionally substituted with one or more -S(=O)2OH or -C(=O)OH 1-12 alkyl;
[0129] R 1A 、R 3A 、R 5A 、R 6A and R 8A Each independently is H;
[0130] R 2A 、R 4A 、R 5A and R 7A Each is independently -S(=O)2OH or -C(=O)OH;
[0131] R 1B 、R 3B 、R 5B 、R 6B and R 8B are each independently H; and
[0132] R 2B 、R 4B 、R 5B and R 7B Each is independently -S(=O)2OH or -C(=O)OH.
[0133] In some embodiments, R TA 、R TB 、R NA 、R NB 、R 1A 、R 2A 、R 3A 、R 4A 、R 5A 、R 6A 、R 7A 、R 8A 、R 1B 、R 2B 、R 3B 、R 4B 、R 5B 、R 6B 、R 7B and R 8B At least one of them includes -SO3H.
[0134] In some embodiments, R TA 、R TB 、R NA 、R NB 、R 1A 、R 2A 、R 3A 、R 4A 、R 5A 、R 6A 、R 7A 、R 8A 、R 1B 、R 2B 、R 3B 、R 4B 、R 5B 、R 6B 、R 7B and R 8B At least two, three, four, five, six, seven or eight of the above-mentioned radicals include -SO3H.
[0135] In some embodiments, R TA 、R TB 、R NA 、R NB 、R 1A 、R 2A 、R 3A 、R 4A 、R 1B 、R 2B 、R 3Band R 4B At least one of them includes -SO3H.
[0136] In some embodiments, R TA 、R TB 、R NA 、R NB 、R 1A 、R 2A 、R 3A 、R 4A 、R 1B 、R 2B 、R 3B and R 4B At least two, three, four, five, six, seven or eight of the above-mentioned radicals include -SO3H.
[0137] In some embodiments, R TA 、R TB 、R NA 、R NB 、R 1A 、R 2A 、R 5A 、R 6A 、R 7A 、R 8A 、R 1B 、R 2B 、R 3B and R 4B At least one of them includes -SO3H.
[0138] In some embodiments, R TA 、R TB 、R NA 、R NB 、R 1A 、R 2A 、R 5A 、R 6A 、R 7A 、R 8A 、R 1B 、R 2B 、R 3B and R 4B At least two, three, four, five, six, seven or eight of the above-mentioned radicals include -SO3H.
[0139] In some embodiments, R TA 、R TB 、R NA 、R NB 、R 1A 、R 2A 、R 5A 、R 6A 、R 7A 、R 8A 、R 1B 、R2B 、R 5B 、R 6B 、R 7B and R 8B At least one of them includes -SO3H.
[0140] In some embodiments, R TA 、R TB 、R NA 、R NB 、R 1A 、R 2A 、R 3A 、R 4A 、R 5A 、R 6A 、R 7A 、R 8A 、R 1B 、R 2B 、R 3B 、R 4B 、R 5B 、R 6B 、R 7B and R 8B At least two, three, four, five, six, seven or eight of the above-mentioned radicals include -SO3H.
[0141] In some embodiments, R TA 、R TB 、R 1A 、R 2A 、R 3A 、R 4A 、R 5A 、R 6A 、R 7A 、R 8A 、R 1B 、R 2B 、R 3B 、R 4B 、R 5B 、R 6B 、R 7B and R 8B At least one of them includes -SO3H.
[0142] In some embodiments, R TA 、R TB 、R 1A 、R 2A 、R 3A 、R 4A 、R 5A 、R 6A 、R 7A 、R 8A 、R 1B 、R 2B 、R 3B 、R4B 、R 5B 、R 6B 、R 7B and R 8B At least two, three, four, five, six, seven or eight of the above-mentioned radicals include -SO3H.
[0143] In some embodiments, R TA 、R TB 、R 1A 、R 2A 、R 3A 、R 4A 、R 1B 、R 2B 、R 3B and R 4B At least one of them includes -SO3H.
[0144] In some embodiments, R TA 、R TB 、R 1A 、R 2A 、R 3A 、R 4A 、R 1B 、R 2B 、R 3B and R 4B At least two, three, four, five, six, seven or eight of the above-mentioned radicals include -SO3H.
[0145] In some embodiments, R TA 、R TB 、R 1A 、R 2A 、R 5A 、R 6A 、R 7A 、R 8A 、R 1B 、R 2B 、R 3B and R 4B At least one of them includes -SO3H.
[0146] In some embodiments, R TA 、R TB 、R 1A 、R 2A 、R 5A 、R 6A 、R 7A 、R 8A 、R 1B 、R 2B 、R 3B and R 4B At least two, three, four, five, six, seven or eight of the above-mentioned radicals include -SO3H.
[0147] In some embodiments, R TA 、R TB 、R 1A 、R 2A 、R 5A 、R 6A 、R 7A 、R 8A 、R 1B 、R 2B 、R 5B 、R 6B 、R 7B and R 8B At least one of them includes -SO3H.
[0148] In some embodiments, R TA 、R TB 、R 1A 、R 2A 、R 3A 、R 4A 、R 5A 、R 6A 、R 7A 、R 8A 、R 1B 、R 2B 、R 3B 、R 4B 、R 5B 、R 6B 、R 7B and R 8B At least two, three, four, five, six, seven or eight of the above-mentioned radicals include -SO3H.
[0149] In some embodiments, R 1A 、R 2A 、R 3A 、R 4A 、R 5A 、R 6A 、R 7A 、R 8A 、R 1B 、R 2B 、R 3B 、R 4B 、R 5B 、R 6B 、R 7B and R 8B At least one of them includes -SO3H.
[0150] In some embodiments, R 1A 、R 2A 、R 3A 、R 4A 、R 5A 、R 6A 、R 7A 、R 8A、R 1B 、R 2B 、R 3B 、R 4B 、R 5B 、R 6B 、R 7B and R 8B At least two, three, four, five, six, seven or eight of the above-mentioned radicals include -SO3H.
[0151] In some embodiments, R 1A 、R 2A 、R 3A 、R 4A 、R 1B 、R 2B 、R 3B and R 4B At least one of them includes -SO3H.
[0152] In some embodiments, R 1A 、R 2A 、R 3A 、R 4A 、R 1B 、R 2B 、R 3B and R 4B At least two, three, four, five, six, seven or eight of the above-mentioned radicals include -SO3H.
[0153] In some embodiments, R 1A 、R 2A 、R 5A 、R 6A 、R 7A 、R 8A 、R 1B 、R 2B 、R 3B and R 4B At least one of them includes -SO3H.
[0154] In some embodiments, R 1A 、R 2A 、R 5A 、R 6A 、R 7A 、R 8A 、R 1B 、R 2B 、R 3B and R 4B At least two, three, four, five, six, seven or eight of the above-mentioned radicals include -SO3H.
[0155] In some embodiments, R 1A 、R 2A 、R 5A 、R 6A 、R7A 、R 8A 、R 1B 、R 2B 、R 5B 、R 6B 、R 7B and R 8B At least one of them includes -SO3H.
[0156] In some embodiments, R 1A 、R 2A 、R 3A 、R 4A 、R 5A 、R 6A 、R 7A 、R 8A 、R 1B 、R 2B 、R 3B 、R 4B 、R 5B 、R 6B 、R 7B and R 8B At least two, three, four, five, six, seven or eight of the above-mentioned radicals include -SO3H.
[0157] In some embodiments, the compound is of any of the formulae described herein, an ionic derivative thereof, an isomer thereof, or a salt thereof, wherein:
[0158] (a)R 2A 、R 5A 、R 7A 、R TA 、R NB 、R NA 、R TB 、R 2B 、R 5B 、R 6B 、R 7B 、R 4A 、R 4B 、R 8A 、R 8B At least one of which includes -SO3H;
[0159] (b) When T A and R TB When it is CH3, then (i) R 2A 、R 4A 、R 4B 、R 2B 、R 5B and R 7Bat least three of which are -SO3H, (ii) n is 1, and (iii) the compound does not have Formula III (e.g., Formula III-A, Formula III-B, Formula III-C, Formula III-D, Formula III-E, or Formula III-F);
[0160] (c) When the compound has Formula I (eg, Formula IA, Formula IB, Formula IC, Formula ID, Formula IE, or Formula IF) and n is 1, then (i) R 2A 、R 4A 、R 4B and R 2B At least one of them is C(=O)NHCH2CH2SO3H, or (ii) R 2A 、R 2B 、R 4A 、R 4B 、R 5B 、R 5B 、R 7A and R 7B The three in it are -SO3H;
[0161] (d) when (i) the compound has formula III (e.g., formula III-A, formula III-B, formula III-C, formula III-D, formula III-E or formula III-F) (ii) R 7 Yes – (C 2-12 Alkylene)-SO3H, and R 13 Yes – (C 2-12 alkylene)-C(=O)OH, and (iii) n is 2, then (i) R 5A 、R 7A 、R 5B and R 7B At least one of them is C(=O)OH, or (ii) R TA and R TB one of which is CH3; and / or
[0162] (e) when (i) the compound has formula III (e.g., formula III-A, formula III-B, formula III-C, formula III-D, formula III-E or formula III-F) (ii) R 7 Yes – (C 2-12 Alkylene)-SO3H, and R 13 Yes – (C 2-12 alkylene)-C(=O)OH, and (iii) n is 1, then (i) R 5A 、R 7A 、R 5B and R 7B At least one of them is C(=O)OH, or (ii) R TA and RTB Both are –(C 1-12 Alkylene)-SO3H.
[0163] In some embodiments, the compound has Formula (IC), (II-C), or (III-C):
[0164]
[0165] Ionic derivatives thereof, isomers thereof or salts thereof.
[0166] In some embodiments, the compound has Formula (ID), (II-D), or (III-D):
[0167]
[0168]
[0169] Ionic derivatives thereof, isomers thereof or salts thereof.
[0170] In some embodiments, the compound has Formula (IE), (II-E), or (III-E):
[0171]
[0172] Ionic derivatives thereof, isomers thereof or salts thereof.
[0173] In some embodiments, the compound has Formula (IF), (II-F), or (III-F):
[0174]
[0175] Ionic derivatives thereof, isomers thereof or salts thereof.
[0176] In some embodiments, the compound is of any of the formulae described herein, an ionic derivative thereof, an isomer thereof, or a salt thereof, wherein:
[0177] R NA is C optionally substituted with one or more -S(=O)2OH or -C(=O)OH 1-12 alkyl;
[0178] R NB is C optionally substituted with one or more -S(=O)2OH or -C(=O)OH 1-12 alkyl;
[0179] R 2A 、R 4A 、R 5A and R 7A Each independently is -S(=O)2OH, -C(=O)OH, C1-12 Alkyl or -C(=O)-NH-(C 1-12 alkyl), wherein C 1-12 Alkyl or -C(=O)-NH-(C 1-12 alkyl) is optionally substituted with -S(=O)2OH or -C(=O)OH; and
[0180] R 2B 、R 4B 、R 5B and R 7B Each independently is -S(=O)2OH, -C(=O)OH, C 1-12 Alkyl or -C(=O)-NH-(C 1-12 alkyl), wherein C 1-12 Alkyl or -C(=O)-NH-(C 1-12 alkyl) is optionally substituted with -S(=O)2OH or -C(=O)OH.
[0181] In some embodiments, the compound is of any of the formulae described herein, an ionic derivative thereof, an isomer thereof, or a salt thereof, wherein:
[0182] R NA is C optionally substituted with one or more -S(=O)2OH or -C(=O)OH 1-12 alkyl;
[0183] R NB is C optionally substituted with one or more -S(=O)2OH or -C(=O)OH 1-12 alkyl;
[0184] R 2A 、R 4A 、R 5A and R 7A are each independently -S(=O)2OH or -C(=O)OH; and
[0185] R 2B 、R 4B 、R 5B and R 7B Each is independently -S(=O)2OH or -C(=O)OH.
[0186] In some embodiments, R TA 、R TB 、R 2A 、R 5A 、R 7A 、R 2B 、R 5B and R 7B At least one of them includes -SO3H.
[0187] In some embodiments, R TA 、R TB 、R 2A 、R 5A 、R 7A 、R 2B 、R 5B and R 7B At least two, three, four or five of the above-mentioned groups include -SO3H.
[0188] In some embodiments, R TA 、R TB 、R 2A 、R 4A 、R 2B and R 4B At least one of them includes -SO3H.
[0189] In some embodiments, R TA 、R TB 、R 2A 、R 4A 、R 2B and R 4B At least two, three, four or five of the above-mentioned groups include -SO3H.
[0190] In some embodiments, one or more R TA and / or R TB It is -CH2S(=O)2OH.
[0191] In some embodiments, one or more R TA and / or R TB It is –(CH2)3-C(=O)OH.
[0192] In some embodiments, R NA and / or R NB It is -CH3.
[0193] In some embodiments, R X is H. In some embodiments, R X is halogen (e.g., F, Cl, Br, or I). In some embodiments, R X It is C 1-12 In some embodiments, R X is C optionally substituted by -S(=O)2OH or -C(=O)OH 1-12 Alkyl (eg, methyl, ethyl, propyl, butyl, pentyl, or hexyl).
[0194] In some embodiments, R Z is H. In some embodiments, R Zis halogen (e.g., F, Cl, Br, or I). In some embodiments, R Z It is C 1-12 In some embodiments, R Z is C optionally substituted by -S(=O)2OH or -C(=O)OH 1-12 Alkyl (eg, methyl, ethyl, propyl, butyl, pentyl, or hexyl).
[0195] In some embodiments, R X and R Z Together with the atoms to which they are attached, they form C 6-10 In some embodiments, R X and R Z , together with the atoms to which they are attached, form C5-C 10 Cycloalkylene (e.g., cyclopentylene or cyclohexylene).
[0196] In some embodiments, R Y is H. In some embodiments, R Y is halogen (e.g., F, Cl, Br, or I). In some embodiments, R Y It is C 1-12 In some embodiments, R Y is C optionally substituted by -S(=O)2OH or -C(=O)OH 1-12 In some embodiments, R Y It is C 6-10 In some embodiments, R Y is optionally replaced by -S(=O)2OH, -C(=O)OH or -C(=O)OC 1-12 Alkyl substituted C 6-10 In some embodiments, R Y is a 5-10 membered heteroaryl group (e.g., pyrrolyl, thienyl, furanyl, thiazolyl, pyridinyl, pyrazinyl, or pyrimidinyl). In some embodiments, R Y is optionally replaced by -S(=O)2OH, -C(=O)OH or -C(=O)OC 1-12 In some embodiments, R Y It is C 2-9In some embodiments, R Y is optionally replaced by -S(=O)2OH, -C(=O)OH or -C(=O)OC 1-12 Alkyl substituted C 2-9 In some embodiments, R Y is amino (e.g., unsubstituted amino, monoalkylamino, dialkylamino, monoarylamino, or diarylamino). In some embodiments, R Y Yes-SC 1-6 Alkyl or -SC 6-10 In some embodiments, R Y Yes-SC 1-6 Alkyl or -SC 6-10 Aryl, where C 1-6 Alkyl or C 6-10 The aryl group is optionally replaced by -S(=O)2OH, -C(=O)OH or -C(=O)OC 1-12 Alkyl substitution.
[0197] In some embodiments, T A In some embodiments, T A In some embodiments, T A Yes-C(R TA )2-. In some embodiments, each R TA is independently H. In some embodiments, each R TA is independently halogen (e.g., F, Cl, Br, or I). In some embodiments, each R TA Independently C 1-12 In some embodiments, each R TA is independently C optionally substituted with -S(=O)2OH or -C(=O)OH 1-12 In some embodiments, each R TA is independently methyl, C5 alkyl substituted with -C(=O)OH, or C3 alkyl substituted with -S(=O)2OH. In some embodiments, one R TA is CH3, and an R TA is a C5 alkyl substituted with -C(=O)OH, or a C3 alkyl substituted with -S(=O)2OH. In some embodiments, one R TA is a C5 alkyl group substituted with -C(=O)OH, and one RTA is a C3 alkyl substituted with -S(=O)2OH. In some embodiments, each R TA Independently C 6-10 In some embodiments, each R TA are independently optionally replaced by -S(=O)2OH, -C(=O)OH or -C(=O)OC 1-12 Alkyl substituted C 6-10 In some embodiments, each R TA is independently a 5-10 membered heteroaryl (e.g., pyrrolyl, thienyl, furanyl, thiazolyl, pyridinyl, pyrazinyl, or pyrimidinyl). In some embodiments, each R TA are independently optionally replaced by -S(=O)2OH, -C(=O)OH or -C(=O)OC 1-12 Alkyl-substituted 5-10 membered heteroaryl (eg, pyrrolyl, thienyl, furanyl, thiazolyl, pyridinyl, pyrazinyl, or pyrimidinyl).
[0198] In some embodiments, T B In some embodiments, T B In some embodiments, T B Yes-C(R TB )2-. In some embodiments, each R TB is independently H. In some embodiments, each R TB is independently halogen (e.g., F, Cl, Br, or I). In some embodiments, each R TB Independently C 1-12 In some embodiments, each R TB is independently C optionally substituted with -S(=O)2OH or -C(=O)OH 1-12 In some embodiments, each R TB is independently methyl, C5 alkyl substituted with -C(=O)OH, or C3 alkyl substituted with -S(=O)2OH. In some embodiments, one R TB is CH3, and an R TB is a C5 alkyl substituted with -C(=O)OH, or a C3 alkyl substituted with -S(=O)2OH. In some embodiments, one R TB is a C5 alkyl group substituted with -C(=O)OH, and one R TB is a C3 alkyl substituted with -S(=O)2OH. In some embodiments, each R TB Independently C6-10 In some embodiments, each R TB are independently optionally replaced by -S(=O)2OH, -C(=O)OH or -C(=O)OC 1-12 Alkyl substituted C 6-10 In some embodiments, each R TB is independently a 5-10 membered heteroaryl (e.g., pyrrolyl, thienyl, furanyl, thiazolyl, pyridinyl, pyrazinyl, or pyrimidinyl). In some embodiments, each R TB are independently optionally replaced by -S(=O)2OH, -C(=O)OH or -C(=O)OC 1-12 Alkyl-substituted 5-10 membered heteroaryl (eg, pyrrolyl, thienyl, furanyl, thiazolyl, pyridinyl, pyrazinyl, or pyrimidinyl).
[0199] In some embodiments, R NA It is C 1-12 In some embodiments, R NA is C optionally substituted by -S(=O)2OH or -C(=O)OH 1-12 In some embodiments, R NA is a C5 alkyl substituted with -C(=O)OH or a C3 alkyl substituted with -S(=O)2OH. In some embodiments, R NA Yes-(C 1-12 Alkylene)-C(=O)-NH-(CH2CH2O) q -(C 1-12 alkyl), wherein C 1-12 The alkyl group is optionally substituted with -S(=O)2OH, -C(=O)OH, and q is an integer from 20 to 30. In some embodiments, R NA Yes-(C 1-12 Alkylene)-C(=O)-NH-(C 1-12 alkyl), wherein C 1-12 The alkyl group is optionally substituted with -S(=O)2OH, -C(=O)OH. In some embodiments, R NA It is C 6-10 In some embodiments, R NA is optionally replaced by -S(=O)2OH, -C(=O)OH or -C(=O)OC 1-12 Alkyl substituted C 6-10 In some embodiments, R NAis a 5-10 membered heteroaryl group (e.g., pyrrolyl, thienyl, furanyl, thiazolyl, pyridinyl, pyrazinyl, or pyrimidinyl). In some embodiments, R NA is optionally replaced by -S(=O)2OH, -C(=O)OH or -C(=O)OC 1-12 Alkyl-substituted 5-10 membered heteroaryl (eg, pyrrolyl, thienyl, furanyl, thiazolyl, pyridinyl, pyrazinyl, or pyrimidinyl).
[0200] In some embodiments, R NB It is C 1-12 In some embodiments, R NB is C optionally substituted by -S(=O)2OH or -C(=O)OH 1-12 In some embodiments, R NB is a C5 alkyl substituted with -C(=O)OH or a C3 alkyl substituted with -S(=O)2OH. In some embodiments, R NB Yes-(C 1-12 Alkylene)-C(=O)-NH-(CH2CH2O) q -(C 1-12 alkyl), wherein C 1-12 The alkyl group is optionally substituted with -S(=O)2OH, -C(=O)OH, and q is an integer from 20 to 30. In some embodiments, R NB Yes-(C 1-12 Alkylene)-C(=O)-NH-(C 1-12 alkyl), wherein C 1-12 The alkyl group is optionally substituted with -S(=O)2OH, -C(=O)OH. In some embodiments, R NB It is C 6-10 In some embodiments, R NB is optionally replaced by -S(=O)2OH, -C(=O)OH or -C(=O)OC 1-12 Alkyl substituted C 6-10 In some embodiments, R NB is a 5-10 membered heteroaryl group (e.g., pyrrolyl, thienyl, furanyl, thiazolyl, pyridinyl, pyrazinyl, or pyrimidinyl). In some embodiments, R NB is optionally replaced by -S(=O)2OH, -C(=O)OH or -C(=O)OC 1-12 Alkyl-substituted 5-10 membered heteroaryl (eg, pyrrolyl, thienyl, furanyl, thiazolyl, pyridinyl, pyrazinyl, or pyrimidinyl).
[0201] In some embodiments, R 1A is H. In some embodiments, R 1A is -C(=O)OH. In some embodiments, R 1A is -S(=O)2OH. In some embodiments, R 1A It is C 1-12 In some embodiments, R 1A is C optionally substituted by -S(=O)2OH or -C(=O)OH 1-12 In some embodiments, R 1A is -C(=O)-NH-(C 1-12 In some embodiments, R 1A is -C(=O)-NH-(C 1-12 alkyl), wherein C 1-12 The alkyl group is optionally substituted with -S(=O)2OH or -C(=O)OH. 1A It is -C(=O)-NH-C2alkyl-S(=O)2OH.
[0202] In some embodiments, R 2A is H. In some embodiments, R 2A is -C(=O)OH. In some embodiments, R 2A is -S(=O)2OH. In some embodiments, R 2A It is C 1-12 In some embodiments, R 2A is C optionally substituted by -S(=O)2OH or -C(=O)OH 1-12 In some embodiments, R 2A is -C(=O)-NH-(C 1-12 In some embodiments, R 2A is -C(=O)-NH-(C 1-12alkyl), wherein C 1-12 The alkyl group is optionally substituted with -S(=O)2OH or -C(=O)OH. 2A It is -C(=O)-NH-C2alkyl-S(=O)2OH.
[0203] In some embodiments, R 3A is H. In some embodiments, R 3A is -C(=O)OH. In some embodiments, R 3A is -S(=O)2OH. In some embodiments, R 3A It is C 1-12 In some embodiments, R 3A is C optionally substituted by -S(=O)2OH or -C(=O)OH 1-12 In some embodiments, R 3A is -C(=O)-NH-(C 1-12 In some embodiments, R 3A is -C(=O)-NH-(C 1-12 alkyl), wherein C 1-12 The alkyl group is optionally substituted with -S(=O)2OH or -C(=O)OH. 3A It is -C(=O)-NH-C2alkyl-S(=O)2OH.
[0204] In some embodiments, R 4A is H. In some embodiments, R 4A is -C(=O)OH. In some embodiments, R 4A is -S(=O)2OH. In some embodiments, R 4A It is C 1-12 In some embodiments, R 4A is C optionally substituted by -S(=O)2OH or -C(=O)OH 1-12 In some embodiments, R 4A is -C(=O)-NH-(C 1-12In some embodiments, R 4A is -C(=O)-NH-(C 1-12 alkyl), wherein C 1-12 The alkyl group is optionally substituted with -S(=O)2OH or -C(=O)OH. 4A It is -C(=O)-NH-C2alkyl-S(=O)2OH.
[0205] In some embodiments, R 5A is H. In some embodiments, R 5A is -C(=O)OH. In some embodiments, R 5A is -S(=O)2OH. In some embodiments, R 5A It is C 1-12 In some embodiments, R 5A is C optionally substituted by -S(=O)2OH or -C(=O)OH 1-12 In some embodiments, R 5A is -C(=O)-NH-(C 1-12 In some embodiments, R 5A is -C(=O)-NH-(C 1-12 alkyl), wherein C 1-12 The alkyl group is optionally substituted with -S(=O)2OH or -C(=O)OH. 5A It is -C(=O)-NH-C2alkyl-S(=O)2OH.
[0206] In some embodiments, R 6A is H. In some embodiments, R 6A is -C(=O)OH. In some embodiments, R 6A is -S(=O)2OH. In some embodiments, R 6A It is C 1-12 In some embodiments, R 6Ais C optionally substituted by -S(=O)2OH or -C(=O)OH 1-12 In some embodiments, R 6A is -C(=O)-NH-(C 1-12 In some embodiments, R 6A is -C(=O)-NH-(C 1-12 alkyl), wherein C 1-12 The alkyl group is optionally substituted with -S(=O)2OH or -C(=O)OH. 6A It is -C(=O)-NH-C2alkyl-S(=O)2OH.
[0207] In some embodiments, R 7A is H. In some embodiments, R 7A is -C(=O)OH. In some embodiments, R 7A is -S(=O)2OH. In some embodiments, R 7A It is C 1-12 In some embodiments, R 7A is C optionally substituted by -S(=O)2OH or -C(=O)OH 1-12 In some embodiments, R 7A is -C(=O)-NH-(C 1-12 In some embodiments, R 7A is -C(=O)-NH-(C 1-12 alkyl), wherein C 1-12 The alkyl group is optionally substituted with -S(=O)2OH or -C(=O)OH. 7A It is -C(=O)-NH-C2alkyl-S(=O)2OH.
[0208] In some embodiments, R 8A is H. In some embodiments, R 8A is -C(=O)OH. In some embodiments, R8A is -S(=O)2OH. In some embodiments, R 8A It is C 1-12 In some embodiments, R 8A is C optionally substituted by -S(=O)2OH or -C(=O)OH 1-12 In some embodiments, R 8A is -C(=O)-NH-(C 1-12 In some embodiments, R 8A is -C(=O)-NH-(C 1-12 alkyl), wherein C 1-12 The alkyl group is optionally substituted with -S(=O)2OH or -C(=O)OH. 8A It is -C(=O)-NH-C2alkyl-S(=O)2OH.
[0209] In some embodiments, R 1B is H. In some embodiments, R 1B is -C(=O)OH. In some embodiments, R 1B is -S(=O)2OH. In some embodiments, R 1B It is C 1-12 In some embodiments, R 1B is C optionally substituted by -S(=O)2OH or -C(=O)OH 1-12 In some embodiments, R 1B is -C(=O)-NH-(C 1-12 In some embodiments, R 1B is -C(=O)-NH-(C 1-12 alkyl), wherein C 1-12 The alkyl group is optionally substituted with -S(=O)2OH or -C(=O)OH. 1BIt is -C(=O)-NH-C2alkyl-S(=O)2OH.
[0210] In some embodiments, R 2B is H. In some embodiments, R 2B is -C(=O)OH. In some embodiments, R 2B is -S(=O)2OH. In some embodiments, R 2B It is C 1-12 In some embodiments, R 2B is C optionally substituted by -S(=O)2OH or -C(=O)OH 1-12 In some embodiments, R 2B is -C(=O)-NH-(C 1-12 In some embodiments, R 2B is -C(=O)-NH-(C 1-12 alkyl), wherein C 1-12 The alkyl group is optionally substituted with -S(=O)2OH or -C(=O)OH. 2B It is -C(=O)-NH-C2alkyl-S(=O)2OH.
[0211] In some embodiments, R 3B is H. In some embodiments, R 3B is -C(=O)OH. In some embodiments, R 3B is -S(=O)2OH. In some embodiments, R 3B It is C 1-12 In some embodiments, R 3B is C optionally substituted by -S(=O)2OH or -C(=O)OH 1-12 In some embodiments, R 3B is -C(=O)-NH-(C 1-12 In some embodiments, R3B is -C(=O)-NH-(C 1-12 alkyl), wherein C 1-12 The alkyl group is optionally substituted with -S(=O)2OH or -C(=O)OH. 3B It is -C(=O)-NH-C2alkyl-S(=O)2OH.
[0212] In some embodiments, R 4B is H. In some embodiments, R 4B is -C(=O)OH. In some embodiments, R 4B is -S(=O)2OH. In some embodiments, R 4B It is C 1-12 In some embodiments, R 4B is C optionally substituted by -S(=O)2OH or -C(=O)OH 1-12 In some embodiments, R 4B is -C(=O)-NH-(C 1-12 In some embodiments, R 4B is -C(=O)-NH-(C 1-12 alkyl), wherein C 1-12 The alkyl group is optionally substituted with -S(=O)2OH or -C(=O)OH. 4B It is -C(=O)-NH-C2alkyl-S(=O)2OH.
[0213] In some embodiments, R 5B is H. In some embodiments, R 5B is -C(=O)OH. In some embodiments, R 5B is -S(=O)2OH. In some embodiments, R 5B It is C 1-12 In some embodiments, R 5B is C optionally substituted by -S(=O)2OH or -C(=O)OH 1-12 In some embodiments, R 5B is -C(=O)-NH-(C 1-12In some embodiments, R 5B is -C(=O)-NH-(C 1-12 alkyl), wherein C 1-12 The alkyl group is optionally substituted with -S(=O)2OH or -C(=O)OH. 5B It is -C(=O)-NH-C2alkyl-S(=O)2OH.
[0214] In some embodiments, R 6B is H. In some embodiments, R 6B is -C(=O)OH. In some embodiments, R 6B is -S(=O)2OH. In some embodiments, R 6B It is C 1-12 In some embodiments, R 6B is C optionally substituted by -S(=O)2OH or -C(=O)OH 1-12 In some embodiments, R 6B is -C(=O)-NH-(C 1-12 In some embodiments, R 6B is -C(=O)-NH-(C 1-12 alkyl), wherein C 1-12 The alkyl group is optionally substituted with -S(=O)2OH or -C(=O)OH. 6B It is -C(=O)-NH-C2alkyl-S(=O)2OH.
[0215] In some embodiments, R 7B is H. In some embodiments, R 7B is -C(=O)OH. In some embodiments, R 7B is -S(=O)2OH. In some embodiments, R 7B It is C 1-12 In some embodiments, R 7Bis C optionally substituted by -S(=O)2OH or -C(=O)OH 1-12 In some embodiments, R 7B is -C(=O)-NH-(C 1-12 In some embodiments, R 7B is -C(=O)-NH-(C 1-12 alkyl), wherein C 1-12 The alkyl group is optionally substituted with -S(=O)2OH or -C(=O)OH. 7B It is -C(=O)-NH-C2alkyl-S(=O)2OH.
[0216] In some embodiments, R 8B is H. In some embodiments, R 8B is -C(=O)OH. In some embodiments, R 8B is -S(=O)2OH. In some embodiments, R 8B It is C 1-12 In some embodiments, R 8B is C optionally substituted by -S(=O)2OH or -C(=O)OH 1-12 In some embodiments, R 8B is -C(=O)-NH-(C 1-12 In some embodiments, R 8B is -C(=O)-NH-(C 1-12 alkyl), wherein C 1-12 The alkyl group is optionally substituted with -S(=O)2OH or -C(=O)OH. 8B It is -C(=O)-NH-C2alkyl-S(=O)2OH.
[0217] Exemplary embodiments of compounds
[0218] In some embodiments, the compound is selected from the compounds described in Table 1, their ionic derivatives, their isomers, and their salts.
[0219] In some embodiments, the compound is selected from the compounds described in Table 1.
[0220] Table 1
[0221]
[0222]
[0223]
[0224]
[0225]
[0226]
[0227]
[0228]
[0229]
[0230]
[0231]
[0232]
[0233] Table 2
[0234]
[0235]
[0236]
[0237]
[0238]
[0239]
[0240]
[0241]
[0242]
[0243]
[0244]
[0245]
[0246]
[0247]
[0248]
[0249]
[0250] For the avoidance of doubt, it should be understood that in this specification, when a group is defined by "as described herein," that group encompasses the first and broadest definition as well as each and all specific definitions for that group.
[0251] Suitable salts of the compounds of the present disclosure are, for example, sufficiently basic acid addition salts of the compounds of the present disclosure, such as acid addition salts with, for example, inorganic or organic acids (e.g., hydrochloric acid, hydrobromic acid, sulfuric acid, phosphoric acid, trifluoroacetic acid, formic acid, citric acid, or maleic acid). In addition, suitable salts of the compounds of the present disclosure that are sufficiently acidic are alkali metal salts (e.g., sodium or potassium salts), alkaline earth metal salts (e.g., calcium or magnesium salts), ammonium salts, or salts formed with organic bases that provide cations (e.g., salts formed with methylamine, dimethylamine, diethylamine, trimethylamine, piperidine, morpholine, or tris-(2-hydroxyethyl)amine).
[0252] It should be understood that the compounds of the present disclosure and any salts thereof include stereoisomers, mixtures of stereoisomers, polymorphs of all isomeric forms of the compounds.
[0253] As used herein, the term "isomerization" refers to compounds that have the same molecular formula but differ in the sequence of bonding of their atoms or the arrangement of their atoms in space. Isomers that differ in the arrangement of their atoms in space are termed "stereoisomers." Stereoisomers that are not mirror images of one another are termed "diastereomers," and stereoisomers that are non-superimposable mirror images of one another are termed "enantiomers" or sometimes optical isomers. A mixture containing equal amounts of individual enantiomeric forms of opposite chirality is termed a "racemic mixture."
[0254] As used herein, the term "chiral center" refers to a carbon atom bonded to four different substituents.
[0255] As used herein, the term "chiral isomer" means a compound having at least one chiral center. Compounds having more than one chiral center may exist as individual diastereomers or as a mixture of diastereomers (referred to as a "diastereomeric mixture"). When there is one chiral center, the stereoisomers may be characterized by the absolute configuration (R or S) of the chiral center. Absolute configuration refers to the spatial arrangement of the substituents attached to the chiral center. The substituents attached to the chiral center under consideration are ordered according to the sequence rule of Cahn, Ingold, and Prelog. (Cahn et al., Angew. Chem. Inter. Edit. 1966, 5, 385; errata 511; Cahn et al., Angew. Chem. 1966, 78, 413; Cahn and Ingold, J. Chem. Soc. 1951 (London), 612; Cahn et al., Experientia 1956, 12, 81; Cahn, J. Chem. Educ. 1964, 41, 116).
[0256] As used herein, the term "geometric isomers" refers to the diastereomers that exist due to hindered rotation around a double bond or a cycloalkyl linker (e.g., 1,3-cyclobutyl). The names of these configurations are distinguished by the prefixes cis and trans, or Z and E, which indicate that these groups are located on the same or opposite sides of the double bond in the molecule according to the Cahn-Ingold-Prelog rules.
[0257] It should be understood that the compounds of the present disclosure may be depicted as different chiral isomers or geometric isomers. It should also be understood that when a compound has chiral or geometric isomeric forms, all isomeric forms are intended to be included within the scope of the present disclosure, and the naming of the compound does not exclude any isomeric form, but it should be understood that not all isomers have the same level of activity.
[0258] It should be understood that the structures and other compounds discussed in this disclosure include all atropicisomers thereof. It should also be understood that not all atropicisomers have the same level of activity.
[0259] As used herein, the term "atropisomer" refers to a class of stereoisomers in which the atoms of the two isomers are arranged differently in space. Atropisomers exist due to restricted rotation of bulky groups around a central bond. Such atropisomers typically exist as a mixture; however, recent advances in chromatographic techniques have made it possible in some cases to separate a mixture of two atropisomers.
[0260] As used herein, the term "tautomer" refers to one of two or more structural isomers that exist in equilibrium and are easily converted from one isomeric form to another. This conversion results in the migration of hydrogen atoms, accompanied by the conversion of adjacent conjugated double bonds. Tautomers exist as a mixture of tautomeric sets in solution. In solutions where tautomerism may occur, chemical equilibrium of the tautomers is reached. The exact ratio of tautomers depends on several factors, including temperature, solvent, and pH. The concept of tautomers converting into each other through tautomerization is called tautomerism. Of the various possible types of tautomerism, two are commonly observed. In keto-enol tautomerism, electrons and hydrogen atoms move simultaneously. Cyclic tautomerism is caused by the reaction of an aldehyde group (-CHO) in a sugar chain molecule with a hydroxyl group (-OH) in the same molecule, causing it to take on a cyclic (ring-shaped) form, as shown in glucose.
[0261] It should be understood that the compounds of the present disclosure can be depicted as different tautomers. It should also be understood that when a compound has tautomeric forms, all tautomeric forms are intended to be included within the scope of the present disclosure, and the naming of the compound does not exclude any tautomeric form. It should be understood that the activity level of certain tautomers may be higher than that of other tautomers.
[0262] Compounds that have the same molecular formula but differ in the nature or order of the bonding of their atoms or the spatial arrangement of their atoms are called "isomers". Isomers with different atomic spatial arrangements are called "stereoisomers". Stereoisomers that are not mirror images of each other are called "diastereomers", and those that are mirror images that do not superimpose on each other are called "enantiomers". When a compound has an asymmetric center, for example, when it is bonded to four different groups, a pair of enantiomers is possible. Enantiomers can be characterized by the absolute configuration of their asymmetric center and described by the R- and S-sequencing rules of Cahn and Prelog, or by the way in which the molecule rotates the plane of polarized light and is designated as right-handed or left-handed (i.e., referred to as (+) or (-)-isomers, respectively). Chiral compounds can exist as individual enantiomers or as mixtures thereof. A mixture containing equal proportions of enantiomers is called a "racemic mixture".
[0263] The compounds of the present disclosure may have one or more asymmetric centers; therefore, such compounds may be produced as individual (R)- or (S)-stereoisomers or as mixtures thereof. Unless otherwise indicated, the description or naming of a particular compound in the specification and claims is intended to include both individual enantiomers or their racemic or other mixtures. Methods for determining stereochemistry and separating stereoisomers are well known in the art (see discussion in Chapter 4 of "Advanced Organic Chemistry", 4th edition J. March, John Wiley and Sons, New York, 2001), for example, by synthesizing from optically active starting materials or by splitting racemic forms. Some compounds of the present disclosure may have geometric isomerization centers (E- and Z-isomers). It should be understood that the present disclosure encompasses all optical isomers, diastereomers and geometric isomers and mixtures thereof with inflammasome inhibitory activity.
[0264] The present disclosure also encompasses compounds of the present disclosure as defined herein that include one or more isotopic substitutions.
[0265] It should be understood that compounds of any formula described herein include the compounds themselves as well as salts thereof and solvates thereof (if applicable). For example, salts can be formed between anions and positively charged groups (e.g., amino groups) on the substituted compounds disclosed herein. Suitable anions include chloride, bromide, iodide, sulfate, bisulfate, sulfamate, nitrate, phosphate, citrate, methanesulfonate, trifluoroacetate, glutamate, glucuronide, glutarate, malate, maleate, succinate, fumarate, tartrate, toluenesulfonate, salicylate, lactate, naphthenate, and acetate (e.g., trifluoroacetate).
[0266] It should be understood that the compounds of the present disclosure (e.g., salts of the compounds) can exist in hydrate or non-hydrate (anhydrous) forms, or exist as solvates with other solvent molecules. Non-limiting examples of hydrates include monohydrates, dihydrates, etc. Non-limiting examples of solvates include ethanol solvates, acetone solvates, etc.
[0267] As used herein, the term "solvate" refers to a solvent addition form containing either stoichiometric or non-stoichiometric amounts of a solvent. Some compounds tend to trap fixed molar ratios of solvent molecules in their crystalline solid state, thereby forming solvates. If the solvent is water, the solvate is a hydrate; and if the solvent is an alcohol, the solvate formed is an alcoholate. Hydrates are formed by the association of one or more water molecules with a molecule of a substance, where the water remains in its molecular form as HO.
[0268] As used herein, the term "analog" refers to compounds that are structurally similar to one another but have slightly different compositions (e.g., an atom is replaced by an atom of a different element, or a particular functional group is present, or a functional group is replaced by another functional group). Thus, an analog is a compound that is similar or equivalent to a reference compound in function and appearance, but not in structure or origin.
[0269] As used herein, the term "derivative" refers to compounds having a common core structure and substituted with various groups as described herein.
[0270] As used herein, the term "bioisostere" refers to a compound produced by exchanging one atom or group of atoms with another substantially similar atom or group of atoms. The purpose of bioisostere replacement is to generate new compounds with similar biological properties to the parent compound. Bioisostere replacement can be based on physicochemical or topological principles. Examples of carboxylic acid bioisosteres include, but are not limited to, acylsulfonamides, tetrazoles, sulfonates, and phosphonates. See, for example, Patani and LaVoie, Chem. Rev. 96, 3147-3176, 1996.
[0271] It will also be understood that certain compounds of the present disclosure may exist in solvate as well as non-solvate forms (such as, for example, hydrate forms). Suitable solvates are, for example, hydrates, such as hemihydrates, monohydrates, dihydrates, or trihydrates. It will be understood that the present disclosure encompasses all such solvate forms having inflammasome inhibitory activity.
[0272] It should also be understood that certain compounds of the present disclosure may exhibit polymorphic forms, and the present disclosure encompasses all such forms or mixtures thereof having inflammasome inhibitory activity. In general, crystalline materials can be analyzed using conventional techniques (such as X-ray powder diffraction analysis, differential scanning calorimetry, thermogravimetric analysis, diffuse reflectance infrared Fourier transform (Diffuse Reflectance Infrared Fourier Transform, DRIFT) spectroscopy, near infrared (NIR) spectroscopy, solution and / or solid-state nuclear magnetic resonance spectroscopy). The water content of such crystalline materials can be determined by Karl Fischer analysis (Karl Fischer analysis).
[0273] Compound of the present disclosure can exist in multiple different tautomeric forms and quoting of compound of the present disclosure includes all such forms.For avoiding doubt, compound can exist in one of several tautomeric forms, and only specifically describe or show when one of them, every other form is all included in disclosed formula.The example of tautomeric form includes keto form, enol form and enolate form, as in for example following tautomerism pair: ketone / enol (as shown below), imines / enamine, amides / imino alcohol, amidine / amidine, nitroso-group / oxime, thioketone / enethiol and nitro / acid nitro.
[0274]
[0275] The compounds of the present disclosure containing amine functional groups can also form N-oxides. The compounds disclosed herein containing amine functional groups mentioned herein also include N-oxides. In the case where the compound contains several amine functional groups, one or more nitrogen atoms can be oxidized to form N-oxides. The specific examples of N-oxides are N-oxides of nitrogen atoms of tertiary amines or nitrogen-containing heterocycles. N-oxides can be formed by treating the corresponding amine with an oxidant such as hydrogen peroxide or a peracid (e.g., peroxycarboxylic acid), see, for example, Jerry March's Advanced Organic Chemistry, 4th edition, Wiley Interscience, pages. More specifically, N-oxides can be prepared by the procedure of LW Deady (Syn. Comm. 1977, 7, 509-514), wherein the amine compound reacts with meta-chloroperbenzoic acid (mCPBA), for example, in an inert solvent such as dichloromethane.
[0276] Combination of compounds
[0277] Without wishing to be bound by theory, it is understood that the compounds of the present disclosure can be used as fluorescent dyes, for example, in sequencing methods. In some embodiments, combinations of compounds are used, for example, to generate differential signals in sequencing methods.
[0278] In some aspects, the present disclosure provides a combination comprising two or more of the compounds disclosed herein.
[0279] In some embodiments, the combination comprises two compounds disclosed herein.
[0280] In some embodiments, the two compounds are selected from the compounds described in Table 1, their ionic derivatives, their isomers, and their salts.
[0281] In some embodiments, the two compounds are selected from the compounds described in Table 2, their ionic derivatives, their isomers, and their salts.
[0282] In some embodiments, the two compounds are selected from the compounds described in Table 1 and Table 2, their ionic derivatives, their isomers, and their salts.
[0283] In some embodiments, the two compounds are selected from Compound Nos. 1 to 4, ionic derivatives thereof, isomers thereof, and salts thereof.
[0284] In some embodiments, the combination comprises three or more compounds disclosed herein.
[0285] In some embodiments, the combination comprises three compounds disclosed herein.
[0286] In some embodiments, the three compounds are selected from the compounds described in Table 1, their ionic derivatives, their isomers, and their salts.
[0287] In some embodiments, the three compounds are selected from Compound Nos. 1 to 4, ionic derivatives thereof, isomers thereof, and salts thereof.
[0288] In some embodiments, the combination comprises four or more compounds disclosed herein.
[0289] In some embodiments, the combination comprises four compounds disclosed herein.
[0290] In some embodiments, the four compounds are selected from the compounds described in Table 1, their ionic derivatives, their isomers, and their salts.
[0291] In some embodiments, the four compounds are selected from Compound Nos. 1 to 4, their ionic derivatives, their isomers, and their salts.
[0292] In some embodiments, the combination comprises:
[0293] Compound No. 1, its ionic derivative, its isomer or its salt;
[0294] Compound No. 2, its ionic derivative, its isomer or its salt;
[0295] Compound No. 3, its ionic derivative, its isomer or a salt thereof; and
[0296] Compound No. 4, its ionic derivative, its isomer or a salt thereof.
[0297] In some embodiments, the combination comprises:
[0298] Compound No. 1, its ionic derivative, its isomer or its salt;
[0299] Compound No. 2, its ionic derivative, its isomer or its salt;
[0300] Compound No. 3, its ionic derivative, its isomer or a salt thereof; and
[0301] Compound No. 95, its ionic derivative, its isomer, or a salt thereof.
[0302] Methods of using the compounds
[0303] It will be understood that the compounds of the present disclosure can be used in various sequencing methods, including the sequencing methods described herein.
[0304] In some aspects, the present disclosure provides a sequencing method disclosed herein using a compound of the present disclosure (eg, as a fluorescent dye).
[0305] In some aspects, the present disclosure provides a compound of the present disclosure for use in the sequencing methods disclosed herein (eg, as a fluorescent dye).
[0306] Nucleic acid template molecules immobilized to a support or coated support
[0307] The present disclosure provides methods for sequencing a plurality of nucleic acid template molecules immobilized to a support (or to a coating on a support).In some embodiments, the individual template molecules comprise single-stranded nucleic acid molecules.
[0308] In some embodiments, the support is passivated / coated with at least one polymer layer (e.g., Figure 1 ). In some embodiments, at least one of the polymer layers comprises a plurality of capture primers tethered to the polymer layer. In some embodiments, a separate capture primer is used to attach the template molecule to the polymer layer. In some embodiments, the 5' or 3' end of the separate template molecule is covalently attached to the capture primer. In some embodiments, the 5' or 3' region of the separate template molecule hybridizes with the capture primer. In some embodiments, at least one polymer layer may further comprise a plurality of pinning primers tethered to the polymer layer. In some embodiments, a separate pinning primer is used to hybridize with a portion of the template molecule, thereby pinning the portion of the template molecule to the polymer layer.
[0309] In some embodiments, the nucleic acid template molecules can be generated by a clonal amplification workflow. In some embodiments, the nucleic acid template molecules are not generated by a clonal amplification workflow.
[0310] In some embodiments, the template molecule comprises one copy of the sequence of interest. For example, a single copy template molecule can be generated by bridge amplification. In some embodiments, bridge amplification can be performed using any combination of dATP, dGTP, dCTP, dTTP, and / or dUTP. In some embodiments, the single copy template molecule comprises at least one uridine nucleotide or lacks a uridine nucleotide.
[0311] In some embodiments, the template molecule comprises a concatemer comprising two or more serial copies of a polynucleotide unit. In some embodiments, the individual polynucleotide units of the concatemer comprise a sequence of interest (e.g., an insertion region) and at least one universal adapter sequence comprising any one of the following or any combination thereof: a capture primer binding site sequence; a pinning primer binding site sequence; a forward sequencing primer binding site sequence; a reverse sequencing primer binding site sequence; an amplification primer binding site sequence; a first sample index sequence; a second sample index sequence; a first unique molecular tag sequence; a second unique molecular tag sequence; a first compaction oligonucleotide binding site; and / or a second compaction oligonucleotide binding site. In some embodiments, concatemers can be produced by rolling circle amplification (RCA) using circularized library molecules, amplification primers (e.g., fixed to a support or soluble), a strand displacement polymerase, and a plurality of nucleotides. For example, a rolling circle amplification reaction can be performed in a template-guided manner to produce a concatemer having a sequence complementary to the circularized library molecule. The rolling circle amplification reaction can be performed under isothermal amplification conditions. In some embodiments, rolling circle amplification can be performed with any combination of dATP, dGTP, dCTP, dTTP, and / or dUTP. In some embodiments, the concatemer template molecule comprises at least one uridine nucleotide or lacks a uridine nucleotide.
[0312] In some embodiments, each polynucleotide unit of the concatemer comprises a sequence of interest (e.g., an insertion region) and at least one universal adapter sequence comprising any one of the following or any combination thereof: a capture primer binding site sequence; a pinning primer binding site sequence; a forward sequencing primer binding site sequence; a reverse sequencing primer binding site sequence; an amplification primer binding site sequence; a first sample index sequence; a second sample index sequence; a first unique molecular tag sequence; a second unique molecular tag sequence; a first compaction oligonucleotide binding site; and / or a second compaction oligonucleotide binding site.
[0313] The concatemer can self-collapse to form a DNA nanoball. The shape and size of the DNA nanoball can be further compressed by including a pair of inverted repeats in the circular template molecule, or by performing a rolling circle amplification reaction in the presence of one or more compacting oligonucleotides. In some embodiments, the compacting oligonucleotide comprises at least four consecutive guanines. The rolling circle amplification reaction produces a concatemer that comprises repeated copies of a universal binding sequence for the compacting oligonucleotide. At least one compacting oligonucleotide can form a guanine tetrad (e.g., Figure 11) and hybridizes with the universal binding sequence of the compacted oligonucleotide, and the resulting concatemer can fold to form an intramolecular G-quadruplex structure. The concatemer can self-collapse to form a compacted nanosphere. The formation of guanine quadruplexes and G-quadruplexes in the nanosphere can increase the stability of the nanosphere, maintaining its compact size and shape, so that it can withstand the repeated flow of reagents used to perform any sequencing workflow described herein.
[0314] In some embodiments, the compacted oligonucleotide may include at least one region having consecutive guanines. For example, the compacted oligonucleotide may include at least one region having 2, 3, 4, 5, 6 or more consecutive guanines. In some embodiments, the compacted oligonucleotide comprises four consecutive guanines that can form a guanine tetrad structure (see Figure 12 The guanine tetrad structure can be stabilized by Hoogsteen hydrogen bonding. The guanine tetrad structure can be stabilized by a central cation including potassium, sodium, lithium, rubidium, or cesium.
[0315] In certain embodiments, rolling circle amplification (RCA) can be carried out with compaction oligonucleotides, to produce single-stranded concatemer molecules with multiple copies of polynucleotide units arranged in series, wherein each polynucleotide unit comprises at least one binding site of a sequence of interest and a compaction oligonucleotide. In certain embodiments, the compaction oligonucleotide comprises a 5' district, an optional internal district (insertion district) and a 3' district. The 5' and 3' districts of the compaction oligonucleotides can hybridize with the binding sites in the concatemer to draw the distal portion of the concatemer together, thereby compacting the concatemer to form a DNA nanoball. For example, the 5' district of the compaction oligonucleotides is designed to hybridize with the first portion of the concatemer molecule, and the 3' district of the compaction oligonucleotides is designed to hybridize with the second portion of the concatemer molecule. Compared to the concatemers generated in the absence of the compaction oligonucleotides, including the compaction oligonucleotides during RCA can promote the formation of DNA nanoballs with tighter size and shape. The compact and stable properties of the DNA nanoballs improve sequencing accuracy by increasing signal intensity, and the DNA nanoballs maintain their shape and size during multiple sequencing cycles.
[0316] Methods used for sequencing
[0317] The present disclosure provides a method for sequencing a plurality of nucleic acid template molecules. In some embodiments, the nucleic acid template molecule comprises a copy of the sequence of interest, or the nucleic acid template molecule comprises a concatemer. In some embodiments, the sequencing reaction employs a nucleotide reagent comprising any one or any combination of nucleotides and / or multivalent molecules. In some embodiments, the nucleotide reagent comprises a canonical nucleotide. In some embodiments, the nucleotide reagent comprises a nucleotide analog comprising a detectably labeled nucleotide. In some embodiments, the nucleotide reagent comprises a nucleotide carrying a removable or non-removable chain termination portion. In some embodiments, the nucleotide reagent comprises a multivalent molecule, each of which comprises a central core attached to a plurality of polymer arms, each polymer arm having a nucleotide unit at the arm end. In some embodiments, the sequencing reaction employs binding to unlabeled nucleotides without incorporation. In some embodiments, the sequencing reaction employs incorporation of unlabeled nucleotide analogs. In some embodiments, the sequencing reaction employs incorporation of detectably labeled nucleotides with a removable chain termination portion. In some embodiments, the sequencing reaction employs a two-stage sequencing reaction comprising binding to a detectably labeled multivalent molecule without incorporation and incorporation of a nucleotide analog. Method for sequencing using nucleotide analogs
[0318] The present disclosure provides methods for sequencing nucleic acid template molecules, comprising the steps of (a) contacting (i) a plurality of sequencing polymerases, (ii) a plurality of nucleic acid template molecules, and (iii) a plurality of nucleic acid sequencing primers, wherein the contacting is performed under conditions suitable for forming a plurality of complex sequencing polymerases, each complex comprising a sequencing polymerase bound to a nucleic acid duplex, wherein the nucleic acid duplex comprises a nucleic acid template molecule hybridized to a nucleic acid sequencing primer. In some embodiments, the sequencing polymerase comprises a recombinant mutant sequencing polymerase that can bind to and incorporate nucleotide analogs. In some embodiments, the sequencing primer comprises a 3' extendable end or a 3' non-extendable end.
[0319] In some embodiments, the method for sequencing a nucleic acid template molecule further comprises step (b): contacting a plurality of sequencing polymerases with a plurality of nucleotides under conditions suitable for binding at least one nucleotide to one of the sequencing polymerases (the sequencing polymerase being bound to the nucleic acid duplex), and the conditions being suitable for promoting polymerase-catalyzed nucleotide incorporation. In some embodiments, the sequencing polymerase is contacted with the plurality of nucleotides in the presence of at least one catalytic cation comprising magnesium and / or manganese. In some embodiments, the plurality of nucleotides comprises at least one nucleotide analog having a chain termination moiety at the sugar 2' or 3' position. In some embodiments, the chain termination moiety can be removed from the sugar 2' or 3' position to convert the chain termination moiety into an OH or H group. In some embodiments, the plurality of nucleotides comprises at least one nucleotide lacking a chain termination moiety. In some embodiments, at least one nucleotide is labeled with a detectable reporter moiety (e.g., a fluorescent dye, such as a compound of the present disclosure). In some embodiments, the plurality of nucleotides comprises a type of nucleotide selected from the group consisting of dATP, dGTP, dCTP, dTTP, or dUTP. In some embodiments, the plurality of nucleotides comprises a mixture of any two or more types of nucleotides including dATP, dGTP, dCTP, dTTP, and / or dUTP.
[0320] In some embodiments, the method for sequencing a nucleic acid template molecule further comprises step (c): incorporating at least one nucleotide into the 3' end of an extendable sequencing primer of at least one multiplex sequencing polymerase. In some embodiments, the nucleotide incorporation reaction of step (c) comprises a primer extension reaction.
[0321] In some embodiments, the method for sequencing a nucleic acid template molecule further comprises step (d): repeating steps (b) and (c) at least once.
[0322] In certain embodiments, in step (b), dye is attached to nucleotide base.In certain embodiments, dye is attached to nucleotide base by the joint that can be cut / removed from base.In certain embodiments, at least one nucleotide in the nucleotide in multiple nucleotides is not labeled with detectable reporter gene moiety.In certain embodiments, the specific detectable reporter gene moiety (for example, fluorescent dye) attached to nucleotide can correspond to nucleotide base (for example, dATP, dGTP, dCTP, dTTP or dUTP), to allow detection and identification of nucleotide base.In certain embodiments, nucleotide analogs include the dye that is attached to the nucleotide base by the joint that can be cut / removed from base, and nucleotide analogs further include the chain termination part that is attached to 2 ' or 3 ' sugar position by a joint, and the joint can use the conditions (for example, chemical cutting conditions) identical with cutting dye from base to cut / remove.
[0323] In some embodiments, the method further comprises: detecting the at least one incorporated nucleotide at step (c) and / or (d). In some embodiments, the method further comprises: identifying the at least one incorporated nucleotide at step (c) and / or (d). In some embodiments, the sequence of the template molecule can be determined by detecting and identifying the nucleotides bound to the sequencing polymerase, thereby determining the sequence of the template molecule. In some embodiments, the sequence of the nucleic acid template molecule can be determined by detecting and identifying the nucleotides incorporated into the 3' end of the primer, thereby determining the sequence of the template molecule.
[0324] Two-stage method for sequencing nucleic acids
[0325] The present disclosure provides a two-stage method for sequencing nucleic acid template molecules. In some embodiments, the first stage generally includes combining a multivalent molecule with a complex polymerase to form a multivalent-complex polymerase, and detecting the multivalent-complex polymerase.
[0326] In some embodiments, the first phase method comprises step (a): contacting (i) a first plurality of sequencing polymerases, (ii) a plurality of nucleic acid template molecules, and (iii) a plurality of nucleic acid sequencing primers, wherein the contacting is performed under conditions suitable for forming a first plurality of complex sequencing polymerases, each complex comprising the first sequencing polymerase bound to a nucleic acid duplex, wherein the nucleic acid duplex comprises a nucleic acid template molecule hybridized to a nucleic acid sequencing primer. In some embodiments, the sequencing primer comprises a 3' extendable end or a 3' non-extendable end.
[0327] In some embodiments, the method for sequencing a nucleic acid template molecule further comprises step (b): contacting the first plurality of complexed polymerases with a plurality of multivalent molecules to form a plurality of multivalent-complexed polymerases (e.g., binding complexes). In some embodiments, an individual multivalent molecule in the plurality of multivalent molecules comprises a core attached to a plurality of nucleotide arms, and each nucleotide arm is attached to a nucleotide unit (e.g., a nucleotide portion) (e.g., Figures 2 to 5 In some embodiments, the contacting of step (b) is performed under conditions suitable for binding the complementary nucleotide units of the multivalent molecule to at least two of the first plurality of complex polymerases, thereby forming a plurality of multivalent-complex polymerases. In some embodiments, the conditions are suitable for inhibiting polymerase-catalyzed incorporation of the complementary nucleotide units into primers of the plurality of multivalent-complex polymerases. In some embodiments, the contacting of step (b) is performed in the presence of at least one non-catalytic cation that inhibits polymerase-catalyzed nucleotide incorporation. In some embodiments, the at least one non-catalytic cation comprises strontium, barium, and / or calcium.
[0328] In some embodiments, in the method of step (b), at least one of the multivalent molecules in the plurality of multivalent molecules is labeled with a detectable reporter moiety. In some embodiments, the detectable reporter moiety comprises a dye.
[0329] In some embodiments, in the method of step (b), the individual nucleotide arms of the multivalent molecule comprise: (i) a core attachment moiety, (ii) a spacer comprising a PEG moiety, (iii) a linker, and (iv) a nucleotide unit, wherein the core is attached to the plurality of nucleotide arms, wherein the spacer is attached to the linker, and wherein the linker is attached to the nucleotide unit. In some embodiments, the labeled multivalent molecule comprises a dye attached to the core, spacer, linker, and / or nucleotide unit of the multivalent molecule.
[0330] In some embodiments, in the method of step (b), the plurality of multivalent molecules comprises at least one having a plurality of nucleotide arms (e.g., Figures 2 to 5 ), each nucleotide arm having a nucleotide analog (e.g., a nucleotide analog unit) attached thereto, wherein the nucleotide analog comprises a chain terminating moiety located at the 2' and / or 3' position of the sugar. In some embodiments, the plurality of multivalent molecules comprises at least one multivalent molecule comprising a plurality of nucleotide arms, each nucleotide arm having a nucleotide unit attached thereto that lacks a chain terminating moiety.
[0331] In some embodiments, the method for sequencing further comprises step (c): detecting the plurality of multivalent-complexed polymerases. In some embodiments, the detection comprises detecting a multivalent molecule bound to a complexed polymerase in the first plurality of complexed polymerases, wherein complementary nucleotide units of the multivalent molecule bind to the primer but incorporation of the complementary nucleotide units is inhibited. In some embodiments, the multivalent molecule is labeled with a detectable reporter moiety to allow detection.
[0332] In some embodiments, the method for sequencing further comprises step (d): identifying the nucleobase of the complementary nucleotide unit bound to the first plurality of complexed polymerases, thereby determining the sequence of the nucleic acid template molecule. In some embodiments, the multivalent molecule is labeled with a detectable reporter moiety that corresponds to a specific nucleotide unit attached to the nucleotide arm to allow identification of the complementary nucleotide unit (e.g., the nucleotide base adenine, guanine, cytosine, thymine, or uracil) bound to the first plurality of complexed polymerases.
[0333] In some embodiments, the second stage of the two-stage sequencing method generally comprises nucleotide incorporation. In some embodiments, the method for sequencing further comprises step (e): dissociating the plurality of multivalent-complexed polymerases and removing the first plurality of sequencing polymerases and their associated multivalent molecules, and retaining the plurality of nucleic acid duplexes.
[0334] In some embodiments, the method for sequencing further comprises step (f): contacting the plurality of retained nucleic acid duplexes from step (e) with a second plurality of sequencing polymerases, wherein the contacting is performed under conditions suitable for binding the second plurality of sequencing polymerases to the plurality of retained nucleic acid duplexes, thereby forming a second plurality of complexed polymerases, each complex comprising the second sequencing polymerase bound to the nucleic acid duplexes. In some embodiments, the second sequencing polymerase comprises a recombinant mutant sequencing polymerase.
[0335] In some embodiments, the plurality of first sequencing polymerases of step (a) have an amino acid sequence that is 100% identical to the amino acid sequence of the plurality of second sequencing polymerases of step (f). In some embodiments, the plurality of first sequencing polymerases of step (a) have an amino acid sequence that is different from the amino acid sequence of the plurality of second sequencing polymerases of step (f).
[0336] In certain embodiments, the method for sequencing further comprises step (g): the polymerase of the second plurality of compounds is contacted with a plurality of nucleotides, wherein the contact is carried out under conditions suitable for combining the complementary nucleotides from a plurality of nucleotides with at least two of the composite polymerases of the second composite polymerase, thereby forming a plurality of nucleotide-complex polymerases. In certain embodiments, the contact of step (g) is carried out under conditions suitable for promoting the incorporation of the complementary nucleotides of the combined catalysis of the polymerase into the primer of the composite polymerase of the nucleotide-complex. In certain embodiments, the incorporation of nucleotides into the 3' end of the primer in step (g) includes a primer extension reaction. In certain embodiments, the contact of step (g) is carried out in the presence of at least one catalytic cation that promotes the incorporation of nucleotides. In certain embodiments, the at least one catalytic cation includes magnesium and / or manganese. In certain embodiments, a plurality of nucleotides include natural nucleotides (e.g., non-analog nucleotides) or nucleotide analogs. In certain embodiments, a plurality of nucleotides include removable or non-removable 2' and / or 3' chain termination parts. In certain embodiments, a plurality of nucleotides include a plurality of nucleotides labeled with a detectable reporter gene portion. The detectable reporter gene portion includes a dye. In certain embodiments, the dye is attached to the nucleotide base. In certain embodiments, the dye is attached to a nucleotide base with a joint that can be cut / removed from the base or can not be removed from the base. In certain embodiments, at least one nucleotide in the nucleotides in a plurality of nucleotides is not labeled with a detectable reporter moiety. In certain embodiments, a plurality of nucleotides are unlabeled nucleotides. In certain embodiments, a specific detectable reporter moiety (e.g., a fluorescent dye) attached to a nucleotide can correspond to a nucleotide base (e.g., dATP, dGTP, dCTP, dTTP, or dUTP) to allow detection and identification of nucleotide bases.
[0337] In some embodiments, when the plurality of nucleotides in step (g) comprises labeled nucleotides, the method for sequencing further comprises the step (h) of detecting complementary nucleotides incorporated into the primer of the nucleotide-complexed polymerase. In some embodiments, the plurality of nucleotides are labeled with a detectable reporter moiety to allow detection. In some embodiments, when the plurality of nucleotides in step (g) comprises unlabeled nucleotides, the detection of step (h) is omitted.
[0338] In some embodiments, when the plurality of nucleotides in step (g) includes labeled nucleotides, the method for sequencing further comprises step (i): identifying the base of the complementary nucleotide incorporated into the primer of the nucleotide-complexed polymerase. In some embodiments, the identification of the incorporated complementary nucleotide in step (i) can be used to confirm the identity of the complementary nucleotide of the multivalent molecule bound to the first plurality of complexed polymerases in step (d). In some embodiments, the identification in step (i) can be used to determine the sequence of the nucleic acid template molecule. In some embodiments, when the plurality of nucleotides in step (g) includes unlabeled nucleotides, the identification in step (i) is omitted.
[0339] In some embodiments, when the plurality of nucleotides in step (g) includes 2' and / or 3' chain-terminating nucleotides, the method for sequencing further comprises the step (j) of removing the chain-terminating moiety from the incorporated nucleotides.
[0340] In some embodiments, the method for sequencing further comprises step (k): repeating steps (a) to (j) at least once. In some embodiments, the sequence of the nucleic acid template molecule can be determined by detecting and identifying the multivalent molecule that binds to the sequencing polymerase but is not incorporated into the 3' end of the primer at steps (c) and (d). In some embodiments, the sequence of the nucleic acid template molecule can be determined (or confirmed) by detecting and identifying the nucleotide incorporated into the 3' end of the primer at steps (h) and (i).
[0341] Avidity complex formation
[0342] In some embodiments, in any method for sequencing a template molecule, a first plurality of complexed polymerases are bound to a plurality of multivalent molecules to form at least one affinity complex, the method comprising the following steps: (1) binding a first sequencing primer, a first sequencing polymerase, and a first multivalent molecule to a first portion of a concatemer template molecule to form a first binding complex, wherein the first nucleotide unit of the first multivalent molecule binds to the first sequencing polymerase; and (2) binding a second sequencing primer, a second sequencing polymerase, and the first multivalent molecule to a second portion of the same concatemer template molecule to form a second binding complex, wherein the second nucleotide unit of the first multivalent molecule binds to the second sequencing polymerase, wherein the first binding complex and the second binding complex of the same multivalent molecule form an affinity complex. The concatemer template molecule comprises a tandem repeat sequence of a sequence of interest and at least one universal site for binding a sequencing primer. The first sequencing primer and the second sequencing primer can bind to the sequencing primer binding site along the concatemer template molecule. Exemplary multivalent molecules are shown in Figures 2 to 5 middle.
[0343] Formation of affinity complexes and detection and identification
[0344] In some embodiments, in any method for sequencing a template molecule, wherein the method comprises binding a first plurality of complexed polymerases to a plurality of multivalent molecules to form at least one affinity complex, the method comprises the steps of: (1) contacting a plurality of sequencing polymerases and a plurality of sequencing primers to different portions of a concatemer template molecule to form at least a first complexed polymerase and a second complexed polymerase on the same concatemer template molecule; and (2) contacting the plurality of multivalent molecules with at least the first complexed polymerase and the second complexed polymerase on the same concatemer template molecule under conditions suitable for binding a single multivalent molecule from the plurality of multivalent molecules to the first complexed polymerase and the second complexed polymerase, wherein at least a first nucleotide unit of the single multivalent molecule binds to the first complexed polymerase, the first complexed polymerase comprising a first sequencing primer that hybridizes to a first portion of the concatemer template molecule, thereby forming a first bound complex (e.g., The method of claim 1 , wherein the contacting is performed under conditions suitable to inhibit polymerase-catalyzed incorporation of the bound first and second nucleotide units into the first and second binding complexes, and wherein the first and second binding complexes bound to the same multivalent molecule form an affinity complex; and (3) detecting the first and second binding complexes on the same concatemer template molecule, and (4) identifying the first nucleotide unit in the first binding complex, thereby determining the sequence of the first portion of the concatemer template molecule, and identifying the second nucleotide unit in the second binding complex, thereby determining the sequence of the second portion of the concatemer template molecule.
[0345] The concatemer template molecule contains tandem repeats of the sequence of interest and at least one universal site for binding a sequencing primer. Multiple nucleic acid primers can bind to the sequencing primer binding site along the concatemer template molecule. An exemplary multivalent molecule is shown in Figures 2 to 5 middle.
[0346] Combined sequencing
[0347] The present disclosure provides a method for sequencing a nucleic acid template molecule, the method comprising a sequencing-by-binding (SBB) procedure employing unlabeled chain-terminating nucleotides. In some embodiments, a sequencing-by-binding (SBB) method comprises the steps of (a) sequentially contacting a primed template nucleic acid (e.g., a template molecule hybridized to a sequencing primer) with at least two separate mixtures under conditions that stabilize a ternary complex, wherein each of the at least two separate mixtures comprises a polymerase and nucleotides, such that the sequential contacting results in contacting the primed template nucleic acid with nucleotides of base types homologous to a first, second, and third base types in the template under conditions that stabilize the ternary complex; (b) examining the at least two separate mixtures to determine whether a ternary complex is formed; and (c) identifying the next correct nucleotide of the primed template nucleic acid molecule, wherein if a ternary complex is detected in step (b), the next correct nucleotide is identified as a homolog of the first base type, the second base type, or the third base type, and wherein if a ternary complex is not present in step (b), the next correct nucleotide is inferred to be a nucleotide homolog of the fourth base type; (d) adding the next correct nucleotide to a primer of the primed template nucleic acid after step (b), thereby generating an extended primer; and (e) repeating steps (a) to (d) at least once for the primed template nucleic acid comprising the extended primer. Exemplary sequencing-by-binding methods are described in US Pat. Nos. 10,246,744 and 10,731,141 (the contents of both patents are hereby incorporated by reference in their entireties).
[0348] In some embodiments, in step (a) of all sequencing methods described herein, the plurality of nucleic acid template molecules are separated by about 10 2 to 10 15 Pieces / mm 2 The density is fixed to the support.
[0349] In some embodiments, in step (a) of all sequencing methods described herein, a plurality of nucleic acid template molecules are immobilized to a support, either at predetermined positions on the support, or at random positions on the support.
[0350] In some embodiments, in step (a) of all sequencing methods described herein, the support is passivated / coated with at least one polymer layer. In some embodiments, at least one of the polymer layers comprises a plurality of capture primers tethered to the polymer layer. In some embodiments, a separate capture primer is used to attach the template molecule to the polymer layer. In some embodiments, a plurality of nucleic acid template molecules are affixed to at least one polymer layer, either at predetermined locations on the polymer layer or at random locations on the polymer layer.
[0351] In some embodiments, in step (a) of all sequencing methods described herein, the plurality of nucleic acid template molecules comprises a plurality of single copy template molecules, which can be generated by bridge amplification. In some embodiments, the single copy template molecules comprise the same target sequence of interest or different target sequences of interest.
[0352] In some embodiments, in step (a) of all sequencing methods described herein, the plurality of nucleic acid template molecules comprises a plurality of concatemer molecules, wherein individual concatemer molecules comprise two or more tandem copies of a polynucleotide unit. In some embodiments, each polynucleotide unit of the concatemer comprises a sequence of interest (e.g., an insertion region) and at least one universal adapter sequence comprising any one of the following or any combination thereof: a capture primer binding site sequence; a pinning primer binding site sequence; a forward sequencing primer binding site sequence; a reverse sequencing primer binding site sequence; an amplification primer binding site sequence; a first sample index sequence; a second sample index sequence; a first unique molecular tag sequence; a second unique molecular tag sequence; a first compaction oligonucleotide binding site; and / or a second compaction oligonucleotide binding site. In some embodiments, the concatemer template molecules comprise the same target sequence of interest or different target sequences of interest.
[0353] Method for sequencing using nucleotides labeled with phosphate chains
[0354] The present disclosure provides a method for sequencing using an immobilized sequencing polymerase bound to an unimmobilized template molecule, wherein the sequencing reaction is performed using nucleotides labeled with phosphate chains. In some embodiments, the sequencing method includes step (a): providing a support having a plurality of sequencing polymerases immobilized thereon. In some embodiments, the sequencing polymerase includes a processive DNA polymerase. In some embodiments, the sequencing polymerase includes a wild-type or mutant DNA polymerase, including, for example, Phi29 DNA polymerase. In some embodiments, the support includes a plurality of separate compartments, and the sequencing polymerase is fixed to the bottom of the compartment. In some embodiments, the separate compartments include a silica bottom through which light can penetrate. In some embodiments, the separate compartments include a silica bottom configured with a nanophotonic confinement structure including holes in a metal coating (e.g., an aluminum coating). In some embodiments, the pore size of the hole in the metal coating is, for example, about 70 nm. In some embodiments, the height of the nanophotonic confinement structure is about 100 nm. In some embodiments, the nanophotonic confinement structure includes a zero-mode waveguide (ZMW). In some embodiments, the nanophotonic confinement structure contains a liquid.
[0355] In some embodiments, the sequencing method further comprises step (b): contacting a plurality of immobilized sequencing polymerases with a plurality of single-stranded circular nucleic acid template molecules and a plurality of oligonucleotide sequencing primers under conditions suitable for binding of a single immobilized sequencing polymerase to the single-stranded circular template molecule and hybridization of a single sequencing primer to the single-stranded circular template molecule, thereby generating a plurality of polymerase / template / primer complexes. In some embodiments, the single sequencing primer hybridizes to a universal sequencing primer binding site on the single-stranded circular template molecule.
[0356] In some embodiments, the sequencing method further includes step (c): multiple polymerase / template / primer complexes are contacted with multiple phosphate chain-labeled nucleotides, each phosphate chain-labeled nucleotide includes an aromatic base, a pentose (e.g., ribose or deoxyribose) and a phosphate chain including 3 to 20 phosphate groups, wherein the terminal phosphate group is connected to a detectable reporter gene portion (e.g., a fluorescent dye). The first phosphate group, the second phosphate group, and the third phosphate group can be referred to as α, β, and γ phosphate groups. In some embodiments, the specific detectable reporter gene portion attached to the terminal phosphate group corresponds to a nucleotide base (e.g., dATP, dGTP, dCTP, dTTP, or dUTP) to allow detection and identification of core bases. In some embodiments, under conditions suitable for polymerase-catalyzed nucleotide incorporation, multiple polymerase / template / primer complexes are contacted with multiple phosphate chain-labeled nucleotides. In some embodiments, sequencing polymerase can bind to the nucleotides labeled with complementary phosphate chains and incorporate complementary nucleotides relative to the nucleotides in the template molecule. In some embodiments, the polymerase-catalyzed nucleotide incorporation reaction cleaves between the alpha phosphate group and the beta phosphate group, thereby releasing the polyphosphate chain attached to the dye.
[0357] In some embodiments, the sequencing method further comprises step (d): detecting a fluorescent signal emitted by a phosphate-labeled nucleotide that is bound by a sequencing polymerase and incorporated into the end of a sequencing primer. In some embodiments, step (d) further comprises identifying the phosphate-labeled nucleotide that is bound by a sequencing polymerase and incorporated into the end of a sequencing primer.
[0358] In some embodiments, the sequencing method further comprises step (d): repeating steps (c) to (d) at least once. In some embodiments, the sequencing method using phosphate-labeled nucleotides can be performed according to the methods described in U.S. Patent Nos. 7,170,050; 7,302,146; and / or 7,405,281.
[0359] Sequencing polymerase
[0360] The present disclosure provides methods for sequencing nucleic acid template molecules, wherein a sequencing polymerase is capable of incorporating complementary nucleotides relative to nucleotides in a concatemer template molecule. The present disclosure provides methods for sequencing nucleic acid template molecules, wherein a sequencing polymerase is capable of incorporating complementary nucleotide units of a multivalent molecule relative to nucleotides in a nucleic acid template molecule. In some embodiments, the plurality of sequencing polymerases comprises a recombinant mutant polymerase.
[0361] Examples of suitable polymerases for sequencing with nucleotides and / or multivalent molecules include, but are not limited to, Klenow DNA polymerase; Thermus aquaticus DNA polymerase I (Taq polymerase); KlenTaq polymerase; Candidatus altiarchaeales archaea; Yellowstone subterranean archaea; Hadesarchaea; Euryarchaea; Thermoplasmata; Thermococcus polymerases, such as Thermococcus litoralis; bacteriophage T7 DNA polymerase; human α, δ, and ε DNA polymerases; bacteriophage polymerases, such as T4, RB69, and phi29 phage DNA polymerases; Pyrococcus furiosus DNA polymerase (Pfu polymerase); Bacillus subtilis DNA polymerase III; Escherichia coli DNA polymerase III α and ε; 9-degree N polymerase; reverse transcriptase, such as HIV type M or O reverse transcriptase; avian myeloblastosis virus reverse transcriptase; Moloney murine leukemia virus (MMLV) reverse transcriptase; or telomerase. Additional non-limiting examples of DNA polymerases include those from various Archaea genera (such as Aeropyrum, Archaeglobus, Desulfurococcus, Pyrococcus, Pyrococcus, Pyrolobus, Thermodiploides, Thermomyces, Stetella, Sulfolobus, Thermococcus, and Volcanicola, etc., or variants thereof), including such polymerases as are known in the art, such as 9 Degrees N, VENT, DEEP VENT, THERMINATOR, Pfu, KOD, Pfx, Tgo, and RB69 polymerases.
[0362] Nucleotides
[0363] The present disclosure provides a method for sequencing nucleic acid template molecules, wherein any sequencing method described herein adopts at least one nucleotide. Nucleotide includes a base, a sugar and at least one phosphate group. In certain embodiments, at least one nucleotide in a plurality of nucleotides includes an aromatic base, a pentose (e.g., ribose or deoxyribose) and one or more phosphate groups (e.g., 1 to 10 phosphate groups). A plurality of nucleotides may include at least one type of nucleotide selected from the group consisting of the following: dATP, dGTP, dCTP, dTTP and dUTP. A plurality of nucleotides may include a mixture of any combination of two or more types of nucleotides selected from the group consisting of the following: dATP, dGTP, dCTP, dTTP and / or dUTP. In certain embodiments, at least one nucleotide in a plurality of nucleotides is not a nucleotide analog. In certain embodiments, at least one nucleotide in a plurality of nucleotides includes a nucleotide analog.
[0364] In certain embodiments, in any method for sequencing nucleic acid template molecules described herein, at least one nucleotide in a plurality of nucleotides comprises a chain of one, two or three phosphorus atoms, wherein the chain is typically attached to the 5' carbon of the sugar moiety via an ester bond or a phosphoramide bond. In certain embodiments, at least one nucleotide in a plurality of nucleotides is an analog with a phosphorus chain, wherein the phosphorus atom is linked together with an O, S, NH, methylene or ethylene group in the middle. In certain embodiments, the phosphorus atom in the chain comprises a substituted side group (including O, S or BH3). In certain embodiments, the chain comprises a phosphate group substituted with an analog, including phosphoramide, phosphorothioate, phosphorodithioate and O-methylphosphoramidite groups.
[0365] In some embodiments, in any method described herein for sequencing a nucleic acid template molecule, at least one nucleotide in a plurality of nucleotides includes a terminator nucleotide analog having a chain termination portion (e.g., a blocking portion) at a sugar 2' position, at a sugar 3' position, or at a sugar 2' and 3' position. In some embodiments, the chain termination portion can inhibit the polymerase-catalyzed incorporation of subsequent nucleotide units or free nucleotides in the nascent chain during the primer extension reaction. In some embodiments, the chain termination portion is attached to the 3' sugar hydroxyl position, wherein the sugar includes a ribose or deoxyribose moiety. In some embodiments, the chain termination portion can be removed / cut from the 3' sugar hydroxyl position to produce a nucleotide with a 3'OH sugar group, which can be extended with subsequent nucleotides in a polymerase-catalyzed nucleotide incorporation reaction. In some embodiments, the chain termination portion includes an alkyl group, an alkenyl group, an alkynyl group, an allyl group, an aryl group, a benzyl group, an azide group, an amine group, an amide group, a keto group, an isocyanate group, a phosphate group, a thio group, a disulfide group, a carbonate group, a urea group, or a silyl group. In some embodiments, the chain-terminating moiety can be cleaved / removed from the nucleotide, for example, by reacting the chain-terminating moiety with a chemical agent, a change in pH, light, or heat. In some embodiments, the chain-terminating moiety alkyl, alkenyl, alkynyl, and allyl groups can be cleaved with tetrakis(triphenylphosphine)palladium(0) (Pd(PPh3)4), with piperidine, or with 2,3-dichloro-5,6-dicyano-1,4-benzoquinone (DDQ). In some embodiments, the chain-terminating moiety aryl and benzyl groups can be cleaved with H2Pd / C. In some embodiments, the chain-terminating moiety amine, amide, ketone, isocyanate, phosphate, thio, disulfide can be cleaved with phosphine or with a thiol group, including β-mercaptoethanol or dithiothreitol (DTT). In some embodiments, the chain-terminating moiety carbonate can be cleaved with potassium carbonate in MeOH (K2CO3), with triethylamine in pyridine, or with Zn in acetic acid (AcOH). In some embodiments, the chain terminating moieties urea and silyl can be cleaved with tetrabutylammonium fluoride, pyridine-HF, with ammonium fluoride, or with triethylamine trihydrofluoride.
[0366] In some embodiments, in any of the methods described herein for sequencing a nucleic acid template molecule, at least one nucleotide in the plurality of nucleotides comprises a terminator nucleotide analog having a chain-terminating moiety (e.g., a blocking moiety) at the 2' position of the sugar, at the 3' position of the sugar, or at both the 2' and 3' positions of the sugar. In some embodiments, the chain-terminating moiety comprises an azide, an azido, or an azidomethyl group. In some embodiments, the chain-terminating moiety comprises a 3'-O-azido or a 3'-O-azidomethyl group. In some embodiments, the chain-terminating moiety azide, an azido, and an azidomethyl group can be cleaved / removed using a phosphine compound. In some embodiments, the phosphine compound comprises a derivatized trialkylphosphine moiety or a derivatized triarylphosphine moiety. In some embodiments, the phosphine compound comprises tris(2-carboxyethyl)phosphine (TCEP), bissulfotriphenylphosphine (BS-TPP), or tris(hydroxypropyl)phosphine (THPP). In some embodiments, the cleaving agent comprises 4-dimethylaminopyridine (4-DMAP).
[0367] In some embodiments, in any of the methods described herein for sequencing a nucleic acid template molecule, the nucleotide comprises a chain terminating moiety selected from the group consisting of a 3'-deoxynucleotide, a 2',3'-dideoxynucleotide, a 3'-methyl, a 3'-azido, a 3'-azidomethyl, a 3'-O-azidoalkyl, a 3'-O-ethynyl, a 3'-O-aminoalkyl, a 3'-O-fluoroalkyl, a 3'-fluoromethyl, a 3'-difluoromethyl, a 3'-trifluoromethyl, a 3'-sulfonyl, a 3'-malonyl, a 3'-amino, a 3'-O-amino, a 3'-thiol, a 3'-aminomethyl, a 3'-ethyl, a 3'butyl, a 3'-tert-butyl, a 3'-fluorenylmethoxycarbonyl, a 3'-tert-butoxycarbonyl, a 3'-O-alkylhydroxyamino group, a 3'-phosphorothioate, and a 3-O-benzyl group, or derivatives thereof.
[0368] In certain embodiments, in any method for sequencing nucleic acid template molecules described herein, a plurality of nucleotides include a plurality of nucleotides labeled with a detectable reporter gene portion. The detectable reporter gene portion includes a dye. In certain embodiments, the dye is attached to a nucleotide base. In certain embodiments, the dye is attached to a nucleotide base by a joint that can be cut / removed from a base. In certain embodiments, at least one nucleotide in the nucleotides in a plurality of nucleotides is not labeled with a detectable reporter gene portion. In certain embodiments, the specific detectable reporter gene portion (e.g., a fluorescent dye) attached to a nucleotide can correspond to a nucleotide base (e.g., dATP, dGTP, dCTP, dTTP or dUTP) to allow detection and identification of a nucleotide base.
[0369] In some embodiments, in any of the methods described herein for sequencing nucleic acid template molecules, the cleavable linker on the nucleotide base comprises a cleavable moiety comprising an alkyl group, an alkenyl group, an alkynyl group, an allyl group, an aryl group, a benzyl group, an azide group, an amine group, an amide group, a keto group, an isocyanate group, a phosphate group, a thio group, a disulfide group, a carbonate group, a urea group, or a silyl group. In some embodiments, the cleavable linker on the base can be cleaved / removed from the base by reacting the cleavable moiety with a chemical agent, a pH change, light, or heat. In some embodiments, the cleavable moiety alkyl, alkenyl, alkynyl, and allyl groups can be cleaved with tetrakis(triphenylphosphine)palladium(0) (Pd(PPh3)4), with piperidine, or with 2,3-dichloro-5,6-dicyano-1,4-benzoquinone (DDQ). In some embodiments, the cleavable moiety aryl and benzyl groups can be cleaved with H2Pd / C. In some embodiments, the cleavable moiety amine, amide, ketone, isocyanate, phosphate, thio, disulfide can be cleaved with phosphine or with a thiol group, the thiol group including β-mercaptoethanol or dithiothreitol (DTT). In some embodiments, the cleavable moiety carbonate can be cleaved with potassium carbonate in MeOH (K2CO3), with triethylamine in pyridine, or with Zn in acetic acid (AcOH). In some embodiments, the cleavable moiety urea and silyl can be cleaved with tetrabutylammonium fluoride, pyridine-HF, with ammonium fluoride, or with triethylamine trihydrofluoride.
[0370] In some embodiments, in any of the methods described herein for sequencing a nucleic acid template molecule, the cleavable linker on the nucleotide base comprises a cleavable moiety comprising an azide, an azido, or an azidomethyl group. In some embodiments, the cleavable moieties azide, an azido, and an azidomethyl group can be cleaved / removed using a phosphine compound. In some embodiments, the phosphine compound comprises a derivatized trialkylphosphine moiety or a derivatized triarylphosphine moiety. In some embodiments, the phosphine compound comprises tris(2-carboxyethyl)phosphine (TCEP), bissulfotriphenylphosphine (BS-TPP), or tris(hydroxypropyl)phosphine (THPP). In some embodiments, the cleavage agent comprises 4-dimethylaminopyridine (4-DMAP).
[0371] In some embodiments, in any of the methods described herein for sequencing nucleic acid template molecules, the chain-terminating portion (e.g., at the sugar 2' and / or sugar 3' position) and the cleavable linker on the nucleotide base have the same or different cleavable portions. In some embodiments, the chain-terminating portion (e.g., at the sugar 2' and / or sugar 3' position) and the detectable reporter portion linked to the base can be chemically cleaved / removed with the same chemical agent. In some embodiments, the chain-terminating portion (e.g., at the sugar 2' and / or sugar 3' position) and the detectable reporter portion linked to the base can be chemically cleaved / removed with different chemical agents.
[0372] Multivalent molecules
[0373] The present disclosure provides methods for sequencing nucleic acid template molecules, wherein any of the sequencing methods described herein employ at least one monovalent molecule. In some embodiments, the multivalent molecule comprises a plurality of nucleotide arms attached to a core and having any configuration, including a starburst, spiral ladder, or bottle brush configuration (e.g., Figure 2 ). The multivalent molecule comprises: (1) a core; and (2) a plurality of nucleotide arms comprising (i) a core attachment portion, (ii) a spacer comprising a PEG portion, (iii) a linker, and (iv) nucleotide units, wherein the core is attached to the plurality of nucleotide arms, wherein the spacer is attached to the linker, and wherein the linker is attached to the nucleotide unit. In some embodiments, the nucleotide unit comprises a base, a sugar, and at least one phosphate group, and the linker is attached to the nucleotide unit via the base. In some embodiments, the linker comprises an aliphatic chain or an oligoethylene glycol chain, wherein both linker chains have 2 to 6 subunits. In some embodiments, the linker further comprises an aromatic portion. Exemplary nucleotide arms are shown in Figure 6 Exemplary multivalent molecules are shown in Figures 2 to 5 An exemplary spacer is shown in Figure 7 (top) and exemplary linkers are shown in Figure 7 (bottom) and Figure 8 Exemplary nucleotides attached to the linker are shown in 9A to 9D An exemplary biotinylated nucleotide arm is shown in Figure 10 middle.
[0374] In some embodiments, the multivalent molecule comprises a core attached to a plurality of nucleotide arms, and wherein the plurality of nucleotide arms have the same type of nucleotide units selected from the group consisting of: dATP, dGTP, dCTP, dTTP, and dUTP.
[0375] In certain embodiments, the multivalent molecule comprises a core attached to a plurality of nucleotide arms, wherein each arm comprises a nucleotide unit. The nucleotide unit comprises an aromatic base, a five-carbon sugar (e.g., ribose or deoxyribose) and one or more phosphate groups (e.g., 1 to 10 phosphate groups). A plurality of multivalent molecules may comprise a type of multivalent molecule with a type of nucleotide unit selected from the group consisting of the following: dATP, dGTP, dCTP, dTTP and dUTP. A plurality of multivalent molecules may comprise a mixture of any combination of two or more types of multivalent molecules, wherein the independent multivalent molecule in the mixture comprises a nucleotide unit selected from the group consisting of the following: dATP, dGTP, dCTP, dTTP and / or dUTP.
[0376] In certain embodiments, the nucleotide unit comprises a chain of one, two or three phosphorus atoms, wherein the chain is typically attached to the 5' carbon of the sugar moiety via an ester bond or a phosphoramidite bond. In certain embodiments, at least one nucleotide unit is a nucleotide analog with a phosphorus chain, wherein the phosphorus atom is linked together with an O, S, NH, methylene or ethylene group in the middle. In certain embodiments, the phosphorus atom in the chain comprises a substituted side group (including O, S or BH3). In certain embodiments, the chain comprises a phosphate group substituted with an analog, including phosphoramide, phosphorothioate, phosphorodithioate and O-methylphosphoramidite groups.
[0377] In some embodiments, the multivalent molecule comprises a core attached to a plurality of nucleotide arms, and wherein each nucleotide arm comprises a nucleotide unit, which is a nucleotide analog having a chain termination portion (e.g., a blocking portion) at a sugar 2' position, at a sugar 3' position, or at a sugar 2' and 3' position. In some embodiments, the nucleotide unit is included in a chain termination portion (e.g., a blocking portion) at a sugar 2' position, at a sugar 3' position, or at a sugar 2' and 3' position. In some embodiments, the chain termination portion can inhibit the polymerase-catalyzed incorporation of subsequent nucleotide units or free nucleotides in the nascent chain during a primer extension reaction. In some embodiments, the chain termination portion is attached to a 3' sugar hydroxyl position, wherein the sugar comprises a ribose or deoxyribose moiety. In some embodiments, the chain termination portion can be removed / cut from the 3' sugar hydroxyl position to produce a nucleotide with a 3'OH sugar group, which can be extended with subsequent nucleotides in a polymerase-catalyzed nucleotide incorporation reaction. In some embodiments, the chain-terminating moiety comprises an alkyl group, an alkenyl group, an alkynyl group, an allyl group, an aryl group, a benzyl group, an azide group, an amine group, an amide group, a keto group, an isocyanate group, a phosphate group, a thio group, a disulfide group, a carbonate group, a urea group, or a silyl group. In some embodiments, the chain-terminating moiety can be cleaved / removed from the nucleotide unit, for example, by reacting the chain-terminating moiety with a chemical agent, a pH change, light, or heat. In some embodiments, the chain-terminating moiety alkyl, alkenyl, alkynyl, and allyl groups can be cleaved with tetrakis(triphenylphosphine)palladium(0) (Pd(PPh3)4), with piperidine, or with 2,3-dichloro-5,6-dicyano-1,4-benzoquinone (DDQ). In some embodiments, the chain-terminating moiety aryl and benzyl groups can be cleaved with H2Pd / C. In some embodiments, the chain-terminating moieties amine, amide, ketone, isocyanate, phosphate, thio, disulfide can be cleaved with phosphine or with a thiol group, including β-mercaptoethanol or dithiothreitol (DTT). In some embodiments, the chain-terminating moiety carbonate can be cleaved with potassium carbonate in MeOH (K2CO3), with triethylamine in pyridine, or with Zn in acetic acid (AcOH). In some embodiments, the chain-terminating moieties urea and silyl can be cleaved with tetrabutylammonium fluoride, pyridine-HF, with ammonium fluoride, or with triethylamine trihydrofluoride.
[0378] In some embodiments, the nucleotide unit comprises a chain-terminating moiety (e.g., a blocking moiety) at the 2' position of the sugar, at the 3' position of the sugar, or at both the 2' and 3' positions of the sugar. In some embodiments, the chain-terminating moiety comprises an azide, an azido, or an azidomethyl group. In some embodiments, the chain-terminating moiety comprises a 3'-O-azido or a 3'-O-azidomethyl group. In some embodiments, the chain-terminating moieties azide, azido, and azidomethyl can be cleaved / removed using a phosphine compound. In some embodiments, the phosphine compound comprises a derivatized trialkylphosphine moiety or a derivatized triarylphosphine moiety. In some embodiments, the phosphine compound comprises tris(2-carboxyethyl)phosphine (TCEP), bissulfotriphenylphosphine (BS-TPP), or tris(hydroxypropyl)phosphine (THPP). In some embodiments, the cleaving agent comprises 4-dimethylaminopyridine (4-DMAP).
[0379] In some embodiments, the nucleotide unit comprises a chain terminating moiety selected from the group consisting of a 3'-deoxynucleotide, a 2',3'-dideoxynucleotide, a 3'-methyl, a 3'-azido, a 3'-azidomethyl, a 3'-O-azidoalkyl, a 3'-O-ethynyl, a 3'-O-aminoalkyl, a 3'-O-fluoroalkyl, a 3'-fluoromethyl, a 3'-difluoromethyl, a 3'-trifluoromethyl, a 3'-sulfonyl, a 3'-malonyl, a 3'-amino, a 3'-O-amino, a 3'-thiol, a 3'-aminomethyl, a 3'-ethyl, a 3'butyl, a 3'-tert-butyl, a 3'-fluorenylmethoxycarbonyl, a 3'-tert-butoxycarbonyl, a 3'-O-alkylhydroxyamino group, a 3'-phosphorothioate, and a 3-O-benzyl group, or derivatives thereof.
[0380] In some embodiments, the multivalent molecule comprises a core attached to a plurality of nucleotide arms, wherein the nucleotide arms comprise a spacer, a joint, and a nucleotide unit, and wherein the core, joint, and / or nucleotide unit are labeled with a detectable reporter moiety. In some embodiments, the detectable reporter moiety comprises a dye. In some embodiments, the specific detectable reporter moiety (e.g., a fluorescent dye) attached to the multivalent molecule can correspond to a base (e.g., dATP, dGTP, dCTP, dTTP, or dUTP) of a nucleotide unit to allow detection and identification of nucleotide bases.
[0381] In some embodiments, at least one nucleotide arm of the multivalent molecule has a nucleotide unit attached to a detectable reporter gene portion. In some embodiments, the detectable reporter gene portion is attached to a nucleotide base. In some embodiments, the detectable reporter gene portion comprises a dye. In some embodiments, the specific detectable reporter gene portion (e.g., a fluorescent dye) attached to the multivalent molecule can correspond to a base (e.g., dATP, dGTP, dCTP, dTTP, or dUTP) of the nucleotide unit to allow detection and identification of the nucleotide base.
[0382] In some embodiments, the core of the multivalent molecule comprises an avidin-like or streptavidin-like portion, and the core attachment portion comprises biotin. In some embodiments, the core comprises a streptavidin-type or avidin-type portion, which includes avidin protein, and any derivatives, analogs, and other non-natural forms of avidin that can be bound to at least one biotin portion. Other forms of the avidin portion include natural and recombinant avidin and streptavidin and derived molecules, such as non-glycosylated avidin and truncated streptavidin. For example, the avidin portion includes a deglycosylated form of avidin, bacterial streptavidin produced by Streptomyces (e.g., Streptomyces avidinii), and derivative forms, such as N-acyl avidin, such as N-acetyl, N-phthaloyl, and N-succinyl avidin, and commercially available products EXTRAVIDIN, CAPTAVIDIN, NEUTRAVIDIN, and NEUTRALIITE AVIDIN.
[0383] In some embodiments, any of the methods described herein for sequencing a nucleic acid molecule may comprise forming a binding complex, wherein the binding complex comprises (i) a polymerase, a nucleic acid concatemer molecule that forms a duplex with a primer, and nucleotides, or the binding complex comprises (ii) a polymerase, a nucleic acid concatemer molecule that forms a duplex with a primer, and a nucleotide unit of a multivalent molecule. In some embodiments, the residence time of the binding complex is greater than about 0.1 seconds, 0.2 seconds, 0.3 seconds, 0.4 seconds, 0.5 seconds, 0.6 seconds, 0.7 seconds, 0.8 seconds, 0.9 seconds, or 1 second. The residence time of the binding complex is greater than about 0.1 to 0.25 seconds, or about 0.25 to 0.5 seconds, or about 0.5 to 0.75 seconds, or about 0.75 to 1 second, or about 1 to 2 seconds, or about 2 to 3 seconds, or about 3 to 4 seconds, or about 4 to 5 seconds, and / or wherein the method is or can be performed at a temperature of at or above 15°C, at or above 20°C, at or above 25°C, at or above 35°C, at or above 37°C, at or above 42°C, at or above 55°C, at or above 60°C, or at or above 72°C, or at or above 80°C, or within the range defined by any of the foregoing. The binding complex (e.g., a ternary complex) remains stable until subjected to conditions that result in dissociation of the interaction between any of the polymerase, template molecule, primer, and / or nucleotide units or nucleotides. For example, dissociation conditions include contacting the binding complex with any one or any combination of detergent, EDTA, and / or water. In some embodiments, the present disclosure provides the methods, wherein the binding complex is deposited onto, attached to, or hybridized to a surface that exhibits a contrast-to-noise ratio in the detecting step of greater than 20. In some embodiments, the present disclosure provides the methods, wherein the contacting is performed under conditions that stabilize the binding complex when the nucleotide or nucleotide unit is complementary to the next base of the template nucleic acid, and destabilize the binding complex when the nucleotide or nucleotide unit is not complementary to the next base of the template nucleic acid.
[0384] Coated support
[0385] The present disclosure provides methods for sequencing nucleic acid template molecules, wherein the template molecules are fixed to a support. In some embodiments, at least one surface of the support can be modified with a compound capable of attaching a polymer coating to the support. For example, the support can be modified with a silane compound. In some embodiments, the silane compound can bind to the polymer coating. In some embodiments, at least one surface of the support is passivated with at least one polymer coating (e.g., Figure 1 In some embodiments, the support is passivated with 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more polymer coatings. In some embodiments, the coating forms a continuous layer on the support, wherein the coating does not form a predetermined pattern.
[0386] In some embodiments, the surface coating can be patterned so that the chemically modified layer is confined to one or more discrete areas of the support. For example, the coating can be patterned using photolithographic techniques to produce an ordered array or random pattern of chemically modified areas on the support. Alternatively or in combination, the coating can be patterned using, for example, contact printing and / or inkjet printing techniques. In some embodiments, the coating is distributed on the support in a predetermined pattern, for example, the predetermined pattern includes dots or other predetermined patterns arranged in rows and / or columns. In some embodiments, the coating with the predetermined pattern includes at least one gap area lacking a polymer coating. In some embodiments, the passivation layer forms a porous or semi-porous layer.
[0387] In some embodiments, at least one of the polymer coatings comprises a hydrophilic polymer layer. In some embodiments, at least one polymer coating comprises a polymer molecule having a molecular weight of at least 1000 daltons. The hydrophilic polymer coating may comprise polyethylene glycol (PEG). The hydrophilic polymer layer may comprise unbranched PEG. The hydrophilic polymer layer may comprise a branched PEG having at least 4 branches, for example, the branched PEG comprises 4 to 16 branches. In some embodiments, the hydrophilic polymer layer comprises crosslinking or lacks crosslinking. In some embodiments, the hydrophilic polymer layer comprises crosslinking to form a hydrogel.
[0388] In some embodiments, the hydrophilic polymer layer comprises a monolayer comprising an unbranched polymer, which can form a brush monolayer. In some embodiments, the brush monolayer can form an extended brush monolayer. In some embodiments, the brush monolayer comprises a plurality of unbranched polymers, wherein one end of a given unbranched polymer is attached to a support, and the other end of the same given unbranched polymer is attached to an oligonucleotide primer (e.g., a capture primer or a pinning primer). In some embodiments, the density of the plurality of oligonucleotide primers attached to the brush monolayer is about 10 2 to 10 15 per um 2 .
[0389] In some embodiments, the coating has a degree of hydrophilicity, measurable as a water contact angle, wherein the water contact angle does not exceed 45 degrees.
[0390] In some embodiments, any layer of the polymer coating comprises a plurality of oligonucleotide primers covalently tethered to the polymer layer. In some embodiments, the plurality of oligonucleotide primers are distributed at multiple depths throughout any polymer layer. In some embodiments, the density of the plurality of oligonucleotide primers in any polymer layer is about 10 2 to 10 15 per um 2. In certain embodiments, the independent oligonucleotide primer comprises a nucleic acid molecule, and these nucleic acid molecules include DNA, RNA, DNA / RNA chimera or their analogs. In certain embodiments, the length of multiple oligonucleotide primers is about 10 to 100 nucleotides. In certain embodiments, the independent oligonucleotide primer in multiple oligonucleotide primers includes a 3' extendable end or a 3' non-extendable end. In certain embodiments, the 3' non-extendable end includes a 3' chain termination portion. In certain embodiments, the 5' or 3' end or internal region of the independent oligonucleotide primer is attached to the polymer layer. In certain embodiments, the 5' end of multiple oligonucleotide primers is attached to the polymer layer. In certain embodiments, multiple oligonucleotide primers are randomly distributed in at least one of the polymer layers and embedded therein. In certain embodiments, multiple oligonucleotide primers are distributed in or on at least one polymer layer in a random manner or in a predetermined pattern. In certain embodiments, multiple oligonucleotide primers are distributed in or on at least one polymer layer in a non-random predetermined pattern, for example, the predetermined pattern includes strips or points or other predetermined patterns arranged in rows and / or columns.
[0391] In some embodiments, the support comprises a first layer comprising a first monolayer comprising hydrophilic polymer molecules tethered to the support. In some embodiments, at least some of the polymer molecules in the first layer are covalently tethered to oligonucleotide primers. In some embodiments, the tethered oligonucleotide primers in the first monolayer are arranged in a random manner or in a predetermined pattern. In some embodiments, the polymer molecules in the first layer are not tethered to oligonucleotide primers.
[0392] In some embodiments, the support further comprises a second layer comprising a second monolayer comprising hydrophilic polymer molecules tethered to the first monolayer. In some embodiments, at least some of the polymer molecules in the second layer are covalently tethered to oligonucleotide primers. In some embodiments, the tethered oligonucleotide primers in the second monolayer are arranged in a random manner or in a predetermined pattern. In some embodiments, the polymer molecules in the second layer are not tethered to oligonucleotide primers.
[0393] In some embodiments, the support further comprises a third layer comprising a third monolayer comprising hydrophilic polymer molecules tethered to the second monolayer. In some embodiments, at least some of the polymer molecules in the third layer are covalently tethered to oligonucleotide primers. In some embodiments, the tethered oligonucleotide primers in the third monolayer are arranged in a random manner or in a predetermined pattern. In some embodiments, the polymer molecules in the third layer are not tethered to oligonucleotide primers.
[0394] In some embodiments, the support comprises a functionalized polymer coating covalently bonded to at least a portion of the support via a chemical group on the support, a primer grafted to the functionalized polymer coating, and a water-soluble protective coating on the primer and the functionalized polymer coating. In some embodiments, the functionalized polymer coating comprises poly(N-(5-azidoacetamidopentyl)acrylamide-co-acrylamide) (PAZAM).
[0395] In certain embodiments, at least one of the polymer layers comprises an oligonucleotide primer, including a capture primer, a pinning primer, or a mixture of a capture primer and a pinning primer. In certain embodiments, a plurality of oligonucleotide primers comprise a mixture of one type of capture primer (e.g., with the capture primer sequence of the same batch) or 2 to 50 different types of capture primers (e.g., with 2 to 50 different batches of capture primer sequences). In certain embodiments, a plurality of oligonucleotide primers comprise a mixture of one type of pinning primer (e.g., with the pinning primer sequence of the same batch) or 2 to 50 different types of pinning primers (e.g., with 2 to 50 different batches of pinning primer sequences).
[0396] In some embodiments, separate capture primers (e.g., tethered to and / or embedded in a polymer layer) can be used in an on-support amplification reaction, wherein the separate capture primers hybridize to capture primer binding sites in circularized library molecules and rolling circle amplification can be performed to produce concatemer template molecules that are tethered to and / or embedded in a polymer layer.
[0397] In some embodiments, a separate capture primer (e.g., tethered to and / or embedded in a polymer layer) can be used in an in-solution amplification process, wherein the separate capture primer can hybridize to the capture primer binding site in the nascent concatemer molecule and rolling circle amplification can proceed on the polymer layer to produce a concatemer template molecule that is tethered to and / or embedded in the polymer layer.
[0398] In some embodiments, the density of capture primers in the polymer layer can be adjusted (e.g., increased or decreased) to achieve a desired density of immobilized concatemer template molecules on the support. Generally, a polymer layer with a high density of capture primers will produce a dense packing with a density of about 10 5 to 10 15 per mm 2 The concatemer template molecules are immobilized to the support at a density that is not achievable using supports fabricated to include nanoscale features for attaching the template molecules.
[0399] In some embodiments, a single pinned primer (e.g., which is tethered to or embedded in a polymer layer) can hybridize to a pinned primer binding site in a concatemer molecule to produce a concatemer template molecule that is tethered or embedded (e.g., pinned) in the polymer layer.
[0400] Nucleotide composition
[0401] The present disclosure provides the compositions comprising a plurality of nucleotides, wherein at least one nucleotide in a plurality of nucleotides is with any dye-labeled as herein described.In certain embodiments, a plurality of nucleotides comprise a mixture of dissimilar nucleotides with core base adenine, guanine, cytosine, thymine and / or uracil, wherein all dissimilar nucleotides are dye-labeled or wherein a type of nucleotide is not labeled.In certain embodiments, dyestuff is joined to the core base of nucleotide.In certain embodiments, dyestuff is joined to a phosphate group in the phosphate group in the phosphate chain.For example, dyestuff is joined to terminal phosphate group.In certain embodiments, dyestuff is joined to core base or phosphate group by joint.In certain embodiments, joint can be cut with chemicals, enzyme, heat or light.
[0402] Multivalent compositions
[0403] The present disclosure provides a composition comprising a plurality of multivalent molecules, wherein at least one of the plurality of multivalent molecules is labeled with a dye. In some embodiments, an individual multivalent molecule in the plurality of multivalent molecules comprises a core attached to a plurality of nucleotide arms, and each nucleotide arm is attached to a nucleotide unit (e.g., a nucleotide portion) (e.g., Figures 2 to 5 ). In some embodiments, the multivalent molecule comprises: (1) a core; and (2) a plurality of nucleotide arms comprising (i) a core attachment portion, (ii) a spacer comprising a PEG portion, (iii) a linker, and (iv) a nucleotide unit, wherein the core is attached to the plurality of nucleotide arms, wherein the spacer is attached to the linker, and wherein the linker is attached to the nucleotide unit. In some embodiments, the nucleotide unit comprises a base, a sugar, and at least one phosphate group, and the linker is attached to the nucleotide unit via the base. In some embodiments, the linker comprises an aliphatic chain or an oligoethylene glycol chain, wherein both linker chains have 2 to 6 subunits. In some embodiments, the linker further comprises an aromatic portion.
[0404] In some embodiments, the multivalent molecule comprises a core attached to a plurality of nucleotide arms, and wherein the plurality of nucleotide arms have the same type of nucleotide units selected from the group consisting of: dATP, dGTP, dCTP, dTTP, and dUTP.
[0405] In certain embodiments, the multivalent molecule comprises a core attached to a plurality of nucleotide arms, wherein each arm comprises a nucleotide unit. The nucleotide unit comprises an aromatic base, a five-carbon sugar (e.g., ribose or deoxyribose) and one or more phosphate groups (e.g., 1 to 10 phosphate groups). A plurality of multivalent molecules may comprise a type of multivalent molecule with a type of nucleotide unit selected from the group consisting of the following: dATP, dGTP, dCTP, dTTP and dUTP. A plurality of multivalent molecules may comprise a mixture of any combination of two or more types of multivalent molecules, wherein the independent multivalent molecule in the mixture comprises a nucleotide unit selected from the group consisting of the following: dATP, dGTP, dCTP, dTTP and / or dUTP.
[0406] In certain embodiments, the nucleotide unit comprises a chain of one, two or three phosphorus atoms, wherein the chain is typically attached to the 5' carbon of the sugar moiety via an ester bond or a phosphoramidite bond. In certain embodiments, at least one nucleotide unit is a nucleotide analog with a phosphorus chain, wherein the phosphorus atom is linked together with an O, S, NH, methylene or ethylene group in the middle. In certain embodiments, the phosphorus atom in the chain comprises a substituted side group (including O, S or BH3). In certain embodiments, the chain comprises a phosphate group substituted with an analog, including phosphoramide, phosphorothioate, phosphorodithioate and O-methylphosphoramidite groups.
[0407] In some embodiments, the plurality of multivalent molecules comprises a mixture of different types of multivalent molecules (e.g., any combination of dATP, dGTP, dCTP, dTTP, and / or dUTP), wherein all of the different types of multivalent molecules are labeled with a dye. In some embodiments, at least one type of multivalent molecule in the mixture is not labeled.
[0408] In certain embodiments, the multivalent molecule comprises a core attached to a plurality of nucleotide arms, wherein the nucleotide arms include a spacer, a joint, and a nucleotide unit. In certain embodiments, the core, joint, and / or nucleotide unit are labeled with a detectable reporter moiety. In certain embodiments, the detectable reporter moiety includes any fluorescent dye described herein. In certain embodiments, the specific dye attached to the multivalent molecule may correspond to a base (e.g., dATP, dGTP, dCTP, dTTP, or dUTP) of a nucleotide unit to allow detection and identification of nucleotide bases.
[0409] In some embodiments, at least one nucleotide arm of the multivalent molecule has a nucleotide unit attached to a detectable reporter gene portion. In some embodiments, the dye is attached to a phosphate group in a core base or a phosphate group. In some embodiments, the specific dye attached to the multivalent molecule can correspond to the base (e.g., dATP, dGTP, dCTP, dTTP, or dUTP) of the nucleotide unit to allow detection and identification of the nucleotide base.
[0410] In certain embodiments, the core of the multivalent molecule includes an avidin-like or streptavidin-like portion, and the core attachment portion includes biotin. In certain embodiments, the core is labeled with one or more dyes. In certain embodiments, the core includes a streptavidin-type or avidin-type portion, which includes avidin protein, and any derivatives, analogs, and other non-natural forms of the avidin that can be bound to at least one biotin portion. Other forms of the avidin portion include natural and recombinant avidin and streptavidin and derived molecules, such as non-glycosylated avidin and truncated streptavidin. For example, the avidin portion includes a deglycosylated form of avidin, bacterial streptavidin produced by Streptomyces (e.g., Streptomyces avidinii), and derivative forms, such as N-acyl avidin, such as N-acetyl, N-phthaloyl, and N-succinyl avidin, and commercially available products EXTRAVIDIN, CAPTAVIDIN, NEUTRAVIDIN, and NEUTRALIITE AVIDIN. In some embodiments, one or more lysine residues in the core can be labeled with a dye.
[0411] In certain embodiments, compositions comprises a plurality of multivalent molecules and a plurality of polymerases. In certain embodiments, independent multivalent molecule is not combined with independent polymerase. In certain embodiments, at least one multivalent molecule is combined with at least one polymerase. In certain embodiments, compositions further comprises a plurality of nucleic acid template molecules and / or a plurality of nucleic acid primer molecules. In certain embodiments, multivalent molecule, polymerase, template molecule and primer molecule may or may not be combined together.
[0412] In some embodiments, the composition comprises an affinity complex labeled with one or more dyes, wherein the affinity complex comprises (i) a first sequencing primer, a first sequencing polymerase, and a first multivalent molecule that binds to a first portion of a concatemer template molecule to form a first binding complex, wherein the first nucleotide unit of the first multivalent molecule binds to the first sequencing polymerase; and (ii) a second sequencing primer, a second sequencing polymerase, and the first multivalent molecule are bound to a second portion of the same concatemer template molecule to form a second binding complex, wherein the second nucleotide unit of the first multivalent molecule binds to the second sequencing polymerase, wherein the first binding complex and the second binding complex of the same multivalent molecule form an affinity complex. In some embodiments, the first multivalent molecule and / or the second multivalent molecule can be dye-labeled. In some embodiments, the concatemer template molecule comprises a tandem repeat sequence of a sequence of interest and at least one universal site for binding a sequencing primer. The first sequencing primer and the second sequencing primer can bind to different sequencing primer binding sites along the concatemer template molecule.
[0413] Compositions and methods for making and using multivalent molecules (also known as polymer-nucleotide conjugates) are described in U.S. 16 / 579,794, filed September 23, 2019, the contents of which are hereby expressly incorporated by reference in their entirety.
[0414] Additional suitable sequencing methods
[0415] Additional sequencing methods and / or devices suitable for use with the compounds disclosed herein are disclosed in the following international application numbers: PCT / EP2022 / 063647, filed May 19, 2022; PCT / US2022 / 022184, filed March 28, 2022; PCT / US2022 / 169972, filed February 3, 2022; PCT / EP2021 / 087044, filed December 21, 2021; PCT / EP2021 / 086349, filed December 16, 2021; PCT / US2021 / 018631, filed February 18, 2021; and PCT / US2021 / 022184, filed March 28, 2022. PCT / US2022 / 012306, PCT / US2021 / 013465 filed January 14, 2021, PCT / US2019 / 027292 filed April 12, 2019, PCT / US2017 / 049496 filed August 30, 2017, U.S. Patent No. 11,427,855 issued August 30, 2022, U.S. Patent No. 10,233,490 issued March 19, 2019, and U.S. patent application Nos. 16 / 783,301 filed February 6, 2022 and 14 / 784,605 filed April 17, 2014, which are hereby incorporated by reference.
[0416] definition
[0417] The headings provided herein are not limitations of the various aspects of the disclosure, which can be understood by reference to the specification as a whole.
[0418] As used herein, the term "ionic derivative" refers to the ionic form of the referenced structure. The ionic form can be a cation, anion, or zwitterion. In some embodiments, the ionic form is a zwitterion (i.e., a structure containing equal amounts of positively and negatively charged functional groups). For example, when a structure is described herein as an anion, the disclosure is intended to encompass the corresponding zwitterion of that structure (e.g., where one -S(=O)2OH forms -S(=O)2O). - ).
[0419] As used herein, "alkyl," "C1, C2, C3, C4, C5, or C6 alkyl," or "C1-C6 alkyl" is intended to include C1, C2, C3, C4, C5, or C6 straight-chain (linear) saturated aliphatic hydrocarbon groups and C3, C4, C5, or C6 branched-chain saturated aliphatic hydrocarbon groups. For example, C1-C6 alkyl is intended to include C1, C2, C3, C4, C5, and C6 alkyl groups. Examples of alkyl groups include moieties having from one to six carbon atoms, such as, but not limited to, methyl, ethyl, n-propyl, iso-propyl, n-butyl, sec-butyl, tert-butyl, n-pentyl, iso-pentyl, or n-hexyl. In some embodiments, a straight-chain or branched alkyl group has six or fewer carbon atoms (e.g., C1-C6 for a straight chain and C3-C6 for a branched chain), and in another embodiment, a straight-chain or branched alkyl group has four or fewer carbon atoms.
[0420] As used herein, the term "optionally substituted alkyl" refers to an unsubstituted alkyl group or an alkyl group having the specified substituents replacing one or more hydrogen atoms on one or more carbon atoms of the hydrocarbon backbone. Such substituents may include, for example, alkyl, alkenyl, alkynyl, halogen, hydroxy, alkylcarbonyloxy, arylcarbonyloxy, alkoxycarbonyloxy, aryloxycarbonyloxy, carboxylate, alkylcarbonyl, arylcarbonyl, alkoxycarbonyl, aminocarbonyl, alkylaminocarbonyl, dialkylaminocarbonyl, alkylthiocarbonyl, alkoxy, phosphate, phosphonate, phosphinate, amino (including alkylamino, dialkylamino, arylamino, diarylamino and alkylarylamino), acylamino (including alkylcarbonylamino, arylcarbonylamino, carbamoyl and ureido), amidino, imino, sulfhydryl, alkylthio, arylthio, thiocarboxylate, sulfate, alkylsulfinyl, sulfonate, sulfamoyl, sulfonamido, nitro, trifluoromethyl, cyano, azido, heterocyclyl, alkylaryl, or an aromatic or heteroaromatic moiety.
[0421] As used herein, the term "alkenyl" includes unsaturated aliphatic groups that are similar in length to the above-mentioned alkyl and may be substituted, but contain at least one double bond. For example, the term "alkenyl" includes straight-chain alkenyls (for example, vinyl, propenyl, butenyl, pentenyl, hexenyl, heptenyl, octenyl, nonenyl, decenyl) and branched alkenyl groups. In certain embodiments, the main chain of a straight or branched alkenyl group has six or fewer carbon atoms (for example, a straight chain is C2-C6, and a branched chain is C3-C6). The term "C2-C6" includes alkenyl groups containing two to six carbon atoms. The term "C3-C6" includes alkenyl groups containing three to six carbon atoms.
[0422] As used herein, the term "optionally substituted alkenyl" refers to unsubstituted alkenyl or alkenyl having specified substituents replacing one or more hydrogen atoms on one or more hydrocarbon backbone carbon atoms. Such substituents may include, for example, alkyl, alkenyl, alkynyl, halogen, hydroxy, alkylcarbonyloxy, arylcarbonyloxy, alkoxycarbonyloxy, aryloxycarbonyloxy, carboxylate, alkylcarbonyl, arylcarbonyl, alkoxycarbonyl, aminocarbonyl, alkylaminocarbonyl, dialkylaminocarbonyl, alkylthiocarbonyl, alkoxy, phosphate, phosphonate, phosphinate, amino (including alkylamino, dialkylamino, arylamino, diarylamino and alkylarylamino), acylamino (including alkylcarbonylamino, arylcarbonylamino, carbamoyl and ureido), amidino, imino, sulfhydryl, alkylthio, arylthio, thiocarboxylate, sulfate, alkylsulfinyl, sulfonate, sulfamoyl, sulfonamido, nitro, trifluoromethyl, cyano, heterocyclyl, alkylaryl, or an aromatic or heteroaromatic moiety.
[0423] As used herein, the term "alkynyl" includes unsaturated aliphatic groups similar in length to the above-mentioned alkyl and possibly substituted, but containing at least one triple bond. For example, "alkynyl" includes straight chain alkynyl groups (e.g., ethynyl, propynyl, butynyl, pentynyl, hexynyl, heptynyl, octynyl, nonynyl, decynyl) and branched chain alkynyl groups. In certain embodiments, the backbone of a straight or branched chain alkynyl group has six or fewer carbon atoms (e.g., straight chain is C2-C6, and branched chain is C3-C6). The term "C2-C6" includes alkynyl groups containing two to six carbon atoms. The term "C3-C6" includes alkynyl groups containing three to six carbon atoms. As used herein, "C2-C6 alkenylene linker" or "C2-C6 alkynylene linker" is intended to include C2, C3, C4, C5 or C6 chain (straight or branched) divalent unsaturated aliphatic hydrocarbon groups. For example, a C2-C6 alkenylene linker is intended to include C2, C3, C4, C5, and C6 alkenylene linker groups.
[0424] As used herein, the term "optionally substituted alkynyl" refers to unsubstituted alkynyl groups or alkynyl groups having specified substituents replacing one or more hydrogen atoms on one or more hydrocarbon backbone carbon atoms. Such substituents may include, for example, alkyl, alkenyl, alkynyl, halogen, hydroxy, alkylcarbonyloxy, arylcarbonyloxy, alkoxycarbonyloxy, aryloxycarbonyloxy, carboxylate, alkylcarbonyl, arylcarbonyl, alkoxycarbonyl, aminocarbonyl, alkylaminocarbonyl, dialkylaminocarbonyl, alkylthiocarbonyl, alkoxy, phosphate, phosphonate, phosphinate, amino (including alkylamino, dialkylamino, arylamino, diarylamino and alkylarylamino), acylamino (including alkylcarbonylamino, arylcarbonylamino, carbamoyl and ureido), amidino, imino, sulfhydryl, alkylthio, arylthio, thiocarboxylate, alkylsulfinyl, sulfonate, sulfamoyl, sulfonamido, nitro, trifluoromethyl, cyano, azido, heterocyclyl, alkylaryl, or an aromatic or heteroaromatic moiety.
[0425] Other optionally substituted parts (such as optionally substituted cycloalkyl, heterocycloalkyl, aryl or heteroaryl) include unsubstituted parts and parts with one or more specified substituents.For example, the substituted heterocycloalkyl includes the heterocycloalkyl substituted by one or more alkyl groups, such as 2,2,6,6-tetramethyl-piperidinyl and 2,2,6,6-tetramethyl-1,2,3,6-tetrahydropyridinyl.
[0426] As used herein, the term "cycloalkyl" refers to a group having 3 to 30 carbon atoms (e.g., C3-C 12 、C3-C 10 The saturated or partially unsaturated hydrocarbon monocyclic or polycyclic (e.g., fused ring, bridged ring, or spirocyclic) system of C-C (C or C-C). Examples of cycloalkyl groups include, but are not limited to, cyclopropyl, cyclobutyl, cyclopentyl, cyclohexyl, cycloheptyl, cyclooctyl, cyclopentenyl, cyclohexenyl, cycloheptenyl, 1,2,3,4-tetrahydronaphthyl, and adamantyl. In the case of polycyclic cycloalkyl groups, only one ring in the cycloalkyl group needs to be non-aromatic.
[0427] As used herein, the term "cycloalkylene" refers to a divalent moiety whose corresponding monovalent moiety is a cycloalkyl. It should be understood that the cycloalkylene can be saturated or partially unsaturated.
[0428] As used herein, the term "heterocycloalkyl" refers to a saturated or partially unsaturated 3-8 membered monocyclic ring, a 7-12 membered bicyclic ring (fused, bridged or spiro) or an 11-14 membered tricyclic ring system (fused, bridged or spiro) having one or more heteroatoms (such as O, N, S, P or Se), e.g., 1 or 1 to 2 or 1 to 3 or 1 to 4 or 1 to 5 or 1 to 6 heteroatoms, or e.g., 1, 2, 3, 4, 5 or 6 heteroatoms independently selected from the group consisting of nitrogen, oxygen and sulfur, unless otherwise specified. Examples of heterocycloalkyl groups include, but are not limited to, piperidinyl, piperazinyl, pyrrolidinyl, dioxanyl, tetrahydrofuranyl, isoindolyl, indolyl, imidazolidinyl, pyrazolidinyl, oxazolidinyl, isoxazolidinyl, triazolidinyl, oxiranyl, azetidinyl, oxetanyl, thietanyl, 1,2,3,6-tetrahydropyridinyl, tetrahydropyranyl, dihydropyranyl, pyranyl, morpholinyl, tetrahydrothiopyranyl, 1,4-diazepanyl, 1,4-oxazepanyl, 2-oxa-5-azabicyclo [2.2.1]heptyl, 2,5-diazabicyclo[2.2.1]heptyl, 2-oxa-6-azaspiro[3.3]heptyl, 2,6-diazaspiro[3.3]heptyl, 1,4-dioxa-8-azaspiro[4.5]decyl, 1,4-dioxaspiro[4.5]decyl, 1-oxaspiro[4.5]decyl, 1-azaspiro[4.5]decyl, 3'H-spiro[cyclohexane-1,1'-isobenzofuran]-yl, 7'H-spiro[cyclohexane-1,5'-furo[3,4-b]pyridine] -yl, 3'H-spiro[cyclohexane-1,1'-furo[3,4-c]pyridinyl]-yl, 3-azabicyclo[3.1.0]hexyl, 3-azabicyclo[3.1.0]hexan-3-yl, 1,4,5,6-tetrahydropyrrolo[3,4-c]pyrazolyl, 3,4,5,6,7,8-hexahydropyrido[4,3-d]pyrimidinyl, 4,5,6,7-tetrahydro-1H-pyrazolo[3,4-c]pyridinyl, 5,6,7,8-tetrahydroimidazo[1,2-a]pyridinyl, 6,7, 8,9-tetrahydro-5H-imidazo[1,2-a]azepinyl, 5,6,7,8-tetrahydropyrido[4,3-d]pyrimidinyl, 2-azaspiro[3.3]heptyl, 2-methyl-2-azaspiro[3.3]heptyl, 2-azaspiro[3.5]nonyl, 2-methyl-2-azaspiro[3.5]nonyl, 2-azaspiro[4.5]decyl, 2-methyl-2-azaspiro[4.5]decyl, 2-oxa-azaspiro[3.4]octyl, 2-oxa-azaspiro[3.4]octan-6-yl, etc. In the case of polycyclic heterocycloalkyl, only one ring of the heterocycloalkyl needs to be non-aromatic (for example, 4,5,6,7-tetrahydrobenzo[c]isoxazolyl).
[0429] As used herein, the term "aryl" includes groups having aromatic properties, including "conjugated" or polycyclic ring systems having one or more aromatic rings, and the ring structure does not contain any heteroatoms. The term aryl includes both monovalent and divalent species. Examples of aryl groups include, but are not limited to, phenyl, biphenyl, naphthyl, and the like. Conveniently, aryl is phenyl.
[0430] As used herein, the term "arylene" refers to a divalent moiety whose corresponding monovalent moiety is aryl.
[0431] As used herein, the term "heteroaryl" is intended to include stable 5-, 6-, or 7-membered monocyclic or 7-, 8-, 9-, 10-, 11-, or 12-membered bicyclic aromatic heterocycles consisting of carbon atoms and one or more heteroatoms, such as 1 or 1 to 2 or 1 to 3 or 1 to 4 or 1 to 5 or 1 to 6 heteroatoms, or such as 1, 2, 3, 4, 5, or 6 heteroatoms independently selected from the group consisting of nitrogen, oxygen, and sulfur. The nitrogen atom may be substituted or unsubstituted (i.e., N or NR, where R is H or other substituents as defined). The nitrogen and sulfur heteroatoms may be optionally oxidized (i.e., N→O and S(O)). p , where p = 1 or 2). It should be noted that the total number of S and O atoms in the aromatic heterocycle does not exceed 1. Examples of heteroaryl groups include pyrrole, furan, thiophene, thiazole, isothiazole, imidazole, triazole, tetrazole, pyrazole, oxazole, isoxazole, pyridine, pyrazine, pyridazine, pyrimidine, and the like. Heteroaryl groups can also be fused or bridged with non-aromatic alicyclic or heterocyclic rings to form polycyclic ring systems (e.g., 4,5,6,7-tetrahydrobenzo[c]isoxazolyl).
[0432] Furthermore, the terms "aryl" and "heteroaryl" include polycyclic aryl and heteroaryl groups, e.g., tricyclic, bicyclic, for example, naphthalene, benzoxazole, benzodioxazole, benzothiazole, benzimidazole, benzothiophene, quinoline, isoquinoline, naphthyridine, indole, benzofuran, purine, deazapurine, indolizine.
[0433] The cycloalkyl, heterocycloalkyl, aryl or heteroaryl ring may be substituted at one or more ring positions (e.g., a ring-forming carbon or heteroatom such as N) with such substituents as described above, for example, alkyl, alkenyl, alkynyl, halogen, hydroxy, alkoxy, alkylcarbonyloxy, arylcarbonyloxy, alkoxycarbonyloxy, aryloxycarbonyloxy, carboxylate, alkylcarbonyl, alkylaminocarbonyl, aralkylaminocarbonyl, alkenylaminocarbonyl, alkylcarbonyl, arylcarbonyl, aralkylcarbonyl, alkenylcarbonyl, alkoxycarbonyl, aminocarbonyl, , alkylthiocarbonyl, phosphate, phosphonate, phosphinate, amino (including alkylamino, dialkylamino, arylamino, diarylamino and alkylarylamino), acylamino (including alkylcarbonylamino, arylcarbonylamino, carbamoyl and ureido), amidino, imino, sulfhydryl, alkylthio, arylthio, thiocarboxylate, sulfate, alkylsulfinyl, sulfonate, sulfamoyl, sulfonamide, nitro, trifluoromethyl, cyano, azido, heterocyclic, alkylaryl or aromatic or heteroaromatic moiety. Aryl and heteroaryl groups can also be fused or bridged with non-aromatic alicyclic or heterocyclic rings to form polycyclic ring systems (e.g., tetralin, methylenedioxyphenyl, such as benzo[d][1,3]dioxol-5-yl).
[0434] As used herein, the term "substituted" means that any one or more hydrogen atoms on a designated atom are replaced by a group selected from a designated group, provided that the normal valence of the designated atom is not exceeded and that the substitution produces a stable compound. When the substituent is an oxo or keto (i.e., =O), 2 hydrogen atoms on the atom are replaced. Keto substituents are not present on aromatic moieties. As used herein, a ring double bond is a double bond (e.g., C=C, C=N, or N=N) formed between two adjacent ring atoms. "Stable compound" and "stable structure" mean that the compound is robust enough to be isolated from a reaction mixture to a useful degree of purity and formulated into an effective therapeutic agent.
[0435] When a substituent's bond crosses a bond connecting two atoms in a ring, the substituent may be bonded to any atom in the ring. When a substituent is listed without indicating the atom via which the substituent is bonded to the rest of the compound of a given formula, the substituent may be bonded via any atom in the formula. Combinations of substituents and / or variables are permissible only if such combinations result in stable compounds.
[0436] When any variable (e.g., R) occurs more than one time in any constituent or formula for a compound, its definition at each occurrence is independent of its definition at every other occurrence. Thus, for example, if a group is shown to be substituted with from 0 to 2 R moieties, then that group may be optionally substituted with up to two R moieties, and each occurrence of R is selected independently of the definition of R. Furthermore, combinations of substituents and / or variables are permissible only if such combinations result in stable compounds.
[0437] As used herein, the term "hydroxy" or "hydroxyl" includes a group having -OH or -O - group.
[0438] As used herein, the term "halo" or "halogen" refers to fluoro, chloro, bromo, and iodo.
[0439] The term "haloalkyl" or "haloalkoxy" refers to an alkyl or alkoxy group substituted with one or more halogen atoms.
[0440] As used herein, the term "optionally substituted haloalkyl" refers to unsubstituted haloalkyl or haloalkyl having specified substituents replacing one or more hydrogen atoms on one or more hydrocarbon backbone carbon atoms. Such substituents may include, for example, alkyl, alkenyl, alkynyl, halogen, hydroxy, alkylcarbonyloxy, arylcarbonyloxy, alkoxycarbonyloxy, aryloxycarbonyloxy, carboxylate, alkylcarbonyl, arylcarbonyl, alkoxycarbonyl, aminocarbonyl, alkylaminocarbonyl, dialkylaminocarbonyl, alkylthiocarbonyl, alkoxy, phosphate, phosphonate, phosphinate, amino (including alkylamino, dialkylamino, arylamino, diarylamino and alkylarylamino), acylamino (including alkylcarbonylamino, arylcarbonylamino, carbamoyl and ureido), amidino, imino, sulfhydryl, alkylthio, arylthio, thiocarboxylate, sulfate, alkylsulfinyl, sulfonate, sulfamoyl, sulfonamido, nitro, trifluoromethyl, cyano, azido, heterocyclyl, alkylaryl, or an aromatic or heteroaromatic moiety.
[0441] As used herein, the term "alkoxy" or "alkoxyl" includes substituted and unsubstituted alkyl, alkenyl, and alkynyl groups covalently linked to an oxygen atom. Examples of alkoxy groups or alkoxy radicals include, but are not limited to, methoxy, ethoxy, isopropoxy, propoxy, butoxy, and pentoxy groups. Examples of substituted alkoxy groups include halogenated alkoxy groups. The alkoxy group may be substituted with groups such as alkenyl, alkynyl, halogen, hydroxy, alkylcarbonyloxy, arylcarbonyloxy, alkoxycarbonyloxy, aryloxycarbonyloxy, carboxylate, alkylcarbonyl, arylcarbonyl, alkoxycarbonyl, aminocarbonyl, alkylaminocarbonyl, dialkylaminocarbonyl, alkylthiocarbonyl, alkoxy, phosphate, phosphonate, phosphinate, amino (including alkylamino, dialkylamino, arylamino, diarylamino and alkylarylamino), acylamino (including alkylcarbonylamino, arylcarbonylamino, carbamoyl and ureido), amidino, imino, sulfhydryl, alkylthio, arylthio, thiocarboxylate, sulfate, alkylsulfinyl, sulfonate, sulfamoyl, sulfonamido, nitro, trifluoromethyl, cyano, azido, heterocyclyl, alkylaryl, or an aromatic or heteroaromatic moiety. Examples of halogen-substituted alkoxy groups include, but are not limited to, fluoromethoxy, difluoromethoxy, trifluoromethoxy, chloromethoxy, dichloromethoxy, and trichloromethoxy.
[0442] It should be understood that throughout this specification, when a composition is described as having, including, or comprising a particular component, it is intended that the composition also consists essentially of or consists of said component. Similarly, when a method or process is described as having, including, or comprising particular process steps, such processes also consist essentially of or consist of said process steps. Furthermore, it should be understood that the order of the steps or the order in which certain actions are performed is not important so long as the present invention remains viable. Furthermore, two or more steps or actions may be performed simultaneously.
[0443] It should be understood that the synthetic methods disclosed herein can tolerate a wide variety of functional groups and therefore can use a variety of substituted starting materials. These methods generally provide the desired final compound at or near the end of the overall method, although in some cases it may be necessary to further convert the compound into a pharmaceutically acceptable salt thereof.
[0444] It will be understood that the compounds of the present disclosure can be prepared in a variety of ways using commercially available starting materials, compounds known in the literature, or from readily prepared intermediates by employing standard synthetic methods and procedures known to those skilled in the art or methods and procedures that will be apparent to those skilled in the art in light of the teachings herein. Standard synthetic methods and procedures for preparing organic molecules and functional group transformations and manipulations can be obtained from the relevant scientific literature or from standard textbooks in the field. Although not limited to any one or several sources, classic texts incorporated by reference herein, such as Smith, MB, March, J., March's Advanced Organic Chemistry: Reactions, Mechanisms, and Structure, 5th edition, John Wiley & Sons: New York, 2001; Greene, TW, Wuts, PGM, Protective Groups in Organic Synthesis, 3rd edition, John Wiley & Sons: New York, 1999; R. Larock, Comprehensive Organic Transformations, VCH Publishers (1989); L. Fieser and M. Fieser, Fieser and Fieser's Reagents for Organic Synthesis, John Wiley and Sons (1994); and L. Paquette, ed., Encyclopedia of Reagents for Organic Synthesis, John Wiley and Sons (1995), are useful and recognized reference textbooks on organic synthesis known to those skilled in the art.
[0445] Those of ordinary skill in the art will note that in the reaction sequence and synthesis scheme described herein, the order of some steps may change, such as the introduction and removal of protecting groups. Those of ordinary skill in the art will recognize that some groups may need to be protected by using protecting groups to protect them from the influence of reaction conditions. Protecting groups can also be used to distinguish similar functional groups in molecules. The list of protecting groups and how to introduce and remove these groups can be found in Greene, TW, Wuts, PGM, Protective Groups in Organic Synthesis, 3rd edition, John Wiley & Sons: New York, 1999.
[0446] Unless otherwise defined, the technology and scientific terms used herein have the meanings commonly understood by those of ordinary skill in the art. Generally speaking, the terms related to the technology of molecular biology as described herein, nucleic acid chemistry, protein chemistry, genetics, microbiology, transgenic cell production and hybridization are those terms well known in the art and commonly used. Technology and procedures as described herein are usually performed according to conventional methods well known in the art and as described in the various general and more specific references cited and discussed throughout this specification. For example, referring to Sambrook et al., Molecular Cloning:A Laboratory Manual (3rd edition, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY2000). Also referring to Ausubel et al., Current Protocols in Molecular Biology, Greene Publishing Associates (1992). The nomenclature used in conjunction with laboratory procedures as described herein and technology is those nomenclatures well known in the art and commonly used.
[0447] Unless the context otherwise requires, singular terms shall include pluralities and plural terms shall include the singular. The singular forms "a," "an," and "the," as well as any word used in the singular, include plural referents unless expressly and unequivocally limited to one referent.
[0448] It will be understood that use of alternative terms such as "or" is taken to mean one or both of the alternatives, or any combination thereof.
[0449] As used herein, the term "and / or" should be taken to mean a specific disclosure of each of the specified features or components with or without the other. For example, the term "and / or" as used in phrases such as "A and / or B" herein is intended to include: "A and B"; "A or B"; "A" (A alone); and "B" (B alone). Similarly, the term "and / or" as used in phrases such as "A, B, and / or C" herein is intended to cover each of the following aspects: "A, B, and C"; "A, B, or C"; "A or C"; "A or B"; "B or C"; "A and B"; "B and C"; "A and C"; "A" (A alone); "B" (B alone); and "C" (C alone).
[0450] As used herein and in the appended claims, the terms "comprises," "including," "having," and "containing," and grammatical variations thereof, as used herein, are intended to be non-limiting, such that one or more items in a list are not exclusive of other items that can be substituted or added to the listed items. It should be understood that whenever aspects are described herein with the language "comprising," other similar aspects described with "consisting of" and / or "consisting essentially of" are also provided.
[0451] As used herein, the terms "about" and "approximately" refer to values or compositions within an acceptable error range for a particular value or composition as determined by one of ordinary skill in the art, which will depend in part on how the value or composition is measured or determined, i.e., the limitations of the measurement system. For example, according to practice in the art, "about" or "approximately" may mean within one or more standard deviations. Alternatively, "about" or "approximately" may mean a range of up to 10% (i.e., ±10%) or greater, depending on the limitations of the measurement system. For example, about 5 mg may include any number between 4.5 mg and 5.5 mg. In addition, specifically with respect to biological systems or processes, the term may mean up to an order of magnitude or up to 5 times the value. When a specific value or composition is provided in the present disclosure, unless otherwise stated, the meaning of "about" or "approximately" should be assumed to be within an acceptable error range for that specific value or composition. In addition, where a range and / or subrange of values is provided, the range and / or subrange may include the endpoints of the range and / or subrange.
[0452] The term "biological sample" refers to a section of any one of a single cell, a plurality of cells, a tissue, an organ, an organism or these biological samples. A biological sample can be extracted from an organism (e.g., a biopsy) or obtained from a cell culture grown in a liquid or in a culture dish. A biological sample includes fresh, frozen, freshly frozen or archived samples (e.g., formalin-fixed paraffin-embedded; FFPE). A biological sample can be embedded in wax, resin, epoxy resin or agar. For example, a biological sample can be fixed in any one of the following or any combination of two or more of the following: acetone, ethanol, methanol, formaldehyde, paraformaldehyde-Triton or glutaraldehyde. A biological sample can be sliced or unsliced. A biological sample can be stained, decolorized or unstained.
[0453] Nucleic acids of interest can be extracted from biological samples using any of a variety of techniques known to those skilled in the art. For example, a typical DNA extraction procedure involves: (i) collecting a cell sample or tissue sample from which DNA is to be extracted, (ii) disrupting the cell membrane (i.e., cell lysis) to release DNA and other cytoplasmic components, (iii) treating the lysed sample with a concentrated salt solution to precipitate proteins, lipids, and RNA, followed by centrifugation to separate the precipitated proteins, lipids, and RNA, and (iv) purifying the DNA from the supernatant to remove detergents, proteins, salts, or other reagents used during cell membrane lysis. A variety of suitable commercial nucleic acid extraction and purification kits are consistent with the disclosure herein. Examples include, but are not limited to, the QIAamp kit (for isolating genomic DNA from human samples) and the DNAeasy kit (for isolating genomic DNA from animal or plant samples) from Qiagen (Germantown, Maryland), or the DNAeasy kit (for isolating genomic DNA from animal or plant samples) from Promega (Madison, Wisconsin). and ReliaPrep TM Series of kits.
[0454] As used herein, the terms "nucleic acid," "polynucleotide," and "oligonucleotide," as well as other related terms, are used interchangeably and refer to polymers of nucleotides and are not limited to any specific length. Nucleic acids include recombinant and chemically synthesized forms. Nucleic acids can be isolated. Nucleic acids include DNA molecules (e.g., cDNA or genomic DNA), RNA molecules (e.g., mRNA), analogs of DNA or RNA generated using nucleotide analogs (e.g., peptide nucleic acids (PNA) and non-naturally occurring nucleotide analogs), and chimeric forms containing DNA and RNA. Nucleic acids can be single-stranded or double-stranded. Nucleic acids include polymers of nucleotides, wherein the nucleotides include natural or non-natural bases and / or sugars. Nucleic acids include naturally occurring internucleoside bonds, such as phosphodiester bonds. Nucleic acids may lack phosphate groups. Nucleic acids include non-natural internucleoside bonds, including phosphorothioate, sulfur-containing phosphate, or peptide nucleic acid (PNA) bonds. In some embodiments, nucleic acids include a mixture of one type of polynucleotide or two or more different types of polynucleotides.
[0455] The terms "universal sequence," "universal adapter sequence," and related terms refer to a sequence in a nucleic acid molecule that is shared between two or more polynucleotide molecules. For example, an adapter having the same universal sequence can be joined to multiple polynucleotides such that the population of co-joined molecules carries the same universal adapter sequence. Examples of universal adapter sequences include amplification primer sequences, sequencing primer sequences, or capture primer sequences (e.g., soluble or support-immobilized capture primers).
[0456] As used herein, the terms "operably linked" and "operably joined" or related terms refer to the juxtaposition of components. The juxtaposed components can be covalently linked together. For example, two nucleic acid components can be enzymatically linked together, wherein the bond joining the two components together comprises a phosphodiester bond. A first nucleic acid component and a second nucleic acid component can be linked together, wherein the first nucleic acid component can confer a function on the second nucleic acid component. For example, the bond between a primer binding sequence and a sequence of interest forms a nucleic acid library molecule having a portion that can bind to a primer. In another example, a transgene (e.g., a nucleic acid encoding a polypeptide or a nucleic acid sequence of interest) can be linked to a vector, wherein the bond allows the transgene sequence contained in the vector to be expressed or function. In some embodiments, the transgene is operably linked to a host cell regulatory sequence (e.g., a promoter sequence) that affects transgene expression. In some embodiments, the vector comprises at least one host cell regulatory sequence comprising a promoter sequence, an enhancer, a transcription and / or translation initiation sequence, a transcription and / or translation termination sequence, a polypeptide secretion signal sequence, and the like. In some embodiments, the host cell regulatory sequence controls the level, timing, and / or location of expression of the transgene.
[0457] The terms "linked," "connected," "attached," "appended," and variations thereof include any type of fusion, binding, attachment, or association between any combination of compounds or molecules that is sufficiently stable to withstand use in a particular procedure. The procedure may include, but is not limited to, nucleotide binding; nucleotide incorporation; deblocking (e.g., removal of a chain terminating moiety); washing; removal; flow; detection; imaging and / or identification. Such bonds may include, for example, covalent bonding, ionic bonding, hydrogen bonding, dipole-dipole bonding, hydrophilic bonding, hydrophobic bonding, or affinity bonding, bonds or associations involving van der Waals forces, mechanical bonding, and the like. In some embodiments, such bonds occur within a molecule, such as connecting the ends of a single-stranded or double-stranded linear nucleic acid molecule together to form a circular molecule. In some embodiments, such bonds may occur between a combination of different molecules or between a molecule and a non-molecule, including, but not limited to, a bond between a nucleic acid molecule and a solid surface; a bond between a protein and a detectable reporter gene portion; a bond between a nucleotide and a detectable reporter gene portion; and the like. Some examples of bonds can be found in, e.g., Hermanson, G., “Bioconjugate Techniques”, 2nd ed. (2008); Aslam, M., Dent, A., “Bioconjugation: Protein Coupling Techniques for the Biomedical Sciences”, London: Macmillan (1998); Aslam, M., Dent, A., “Bioconjugation: Protein Coupling Techniques for the Biomedical Sciences”, London: Macmillan (1998).
[0458] The term "adapter" and related terms refer to oligonucleotides that can be operably linked (appended) to a target polynucleotide, wherein the adaptor pair conferring function to a co-linked adaptor-target molecule. The adaptor comprises DNA, RNA, chimeric DNA / RNA or its analogs. The adaptor may include at least one ribonucleoside residue. The adaptor may be single-stranded, double-stranded or have a single-stranded and / or double-stranded portion. The adaptor may be configured into a linear form, a stem-loop form, a hairpin form or a Y-shaped form. The adaptor may be of any length, including 4 to 100 nucleotides or longer. The adaptor may have a blunt end, an overhang end or a combination thereof. The overhang end may include a 5' overhang end and a 3' overhang end. The 5' end of a single-stranded adaptor or a chain of a double-stranded adaptor may have a 5' phosphate group or lack a 5' phosphate group. The adaptor may include a 5' tail (e.g., a tail adaptor) that is not hybridized with the target polynucleotide, or the adaptor may be tailless. The adapter may comprise a sequence that is complementary to at least a portion of a primer, such as an amplification primer, a sequencing primer, or a capture primer (e.g., a soluble or immobilized capture primer). The adapter may comprise a random sequence or a degenerate sequence. The adapter may comprise at least one inosine residue. The adapter may comprise at least one phosphorothioate, phosphorothiol, and / or phosphoramidite bond. The adapter may comprise a barcode sequence that can be used to distinguish polynucleotides (e.g., insert sequences) from different sample sources in a multiplex assay. The adapter may comprise a unique identification sequence (e.g., a unique molecular index, UMI; or a unique molecular tag) that can be used to uniquely identify the nucleic acid molecule to which the adapter is appended. In some embodiments, the unique identification sequence can be used to increase error correction and accuracy, reduce the rate of false positive variant identification, and / or increase the sensitivity of variant detection. The adapter may comprise at least one restriction enzyme recognition sequence, including any one or any combination of two or more selected from the group consisting of: type I, type II, type III, type IV, type Hs, or type IIB.
[0459] The terms "nucleic acid template," "template polynucleotide," "nucleic acid target," "target polynucleotide," "template strand," and other variations refer to a nucleic acid strand that serves as a base nucleic acid molecule for any of the analytical methods described herein (e.g., primer extension, amplification, and / or sequencing). The template nucleic acid can be single-stranded or double-stranded, or the template nucleic acid can have single-stranded or double-stranded portions. The template nucleic acid can be obtained from naturally occurring sources, recombinant forms, or chemically synthesized to include any type of nucleic acid analog. The template nucleic acid can be linear, circular, or in other forms. The template nucleic acid can include an insertion region having an insertion sequence also referred to as a sequence of interest. The template nucleic acid can also include at least one adapter sequence. The template nucleic acid can be a concatemer having two or tandem copies of a sequence of interest and at least one adapter sequence. The insertion region can be isolated in any form, including chromosomal, genomic, organelle (e.g., mitochondrial, chloroplast, or ribosomal) recombinant molecules, cloned, amplified cDNA, RNA (such as pre-mRNA or mRNA), oligonucleotides, whole genomic DNA obtained from fresh frozen paraffin embedded tissue, needle biopsy, circulating tumor cells, cell-free circulating DNA, or any type of nucleic acid library. The insertion region can be isolated from any source, including from organisms such as prokaryotes, eukaryotes (e.g., humans, plants, and animals), fungi, viruses, cells, tissues, normal or diseased cells or tissues, body fluids including blood, urine, serum, lymph, tumors, saliva, anal and vaginal secretions, amniotic fluid samples, sweat, semen, environmental samples, culture samples, or synthetic nucleic acid molecules prepared using recombinant molecular biology or chemical synthesis methods. Insertion regions can be isolated from any organ, including the head, neck, brain, breast, ovary, cervix, colon, rectum, endometrium, gallbladder, intestine, bladder, prostate, testis, liver, lung, kidney, esophagus, pancreas, thyroid, pituitary gland, thymus, skin, heart, larynx, or other organs. Template nucleic acids can be subjected to nucleic acid analysis (including sequencing and composition analysis).
[0460] As used herein, the term "polymerase" and its variants include enzymes comprising a domain that binds nucleotides (or nucleosides), wherein the polymerase can form a complex with a template nucleic acid and a complementary nucleotide. The polymerase may have one or more activities, including but not limited to: base analog detection activity, DNA polymerization activity, reverse transcriptase activity, DNA binding, strand displacement activity, and nucleotide binding and recognition. The polymerase can be any enzyme that can catalyze the polymerization of nucleotides (including analogs thereof) into nucleic acid chains. Typically, but not necessarily, such nucleotide polymerization can occur in a template-dependent manner. Typically, the polymerase includes one or more active sites at which nucleotide binding and / or nucleotide polymerization catalysis can occur. In some embodiments, the polymerase includes other enzymatic activities, such as 3' to 5' exonuclease activity or 5' to 3' exonuclease activity. In some embodiments, the polymerase has strand displacement activity. Polymerases can include, but are not limited to, naturally occurring polymerases and any subunits and truncations thereof, mutant polymerases, variant polymerases, recombinant, fused or otherwise engineered polymerases, chemically modified polymerases, synthetic molecules or assemblies, and any analogs, derivatives or fragments thereof (e.g., catalytically active fragments) that retain the ability to catalyze nucleotide polymerization. Polymerases include catalytically inactive polymerases, catalytically active polymerases, reverse transcriptases and other enzymes comprising a nucleotide binding domain. In certain embodiments, polymerases can be isolated from cells or produced using recombinant DNA technology or chemical synthesis methods. In certain embodiments, polymerases can be expressed in prokaryotes, eukaryotes, viruses or phage organisms. In certain embodiments, polymerases can be post-translationally modified proteins or fragments thereof. Polymerases can be derived from prokaryotes, eukaryotes, viruses or phages. Polymerases include DNA-guided DNA polymerases and RNA-guided DNA polymerases.
[0461] The term "strand displacement" refers to the ability of a polymerase to locally separate a double-stranded nucleic acid chain and synthesize a new chain in a template-based manner. The strand displacement polymerase displaces the complementary strand from the template strand and catalyzes the synthesis of a new chain. The strand displacement polymerase includes mesophilic and thermophilic polymerases. The strand displacement polymerase includes wild-type enzymes and variants (including exonuclease-minus mutants, mutant versions, chimeric enzymes, and truncated enzymes). Examples of strand displacement polymerases include phi29 DNA polymerase, Bst DNA polymerase large fragment, Bsu DNA polymerase large fragment (exo-), Bca DNA polymerase (exo-), the Klenow fragment of Escherichia coli DNA polymerase, T5 polymerase, M-MuLV reverse transcriptase, HIV virus reverse transcriptase, Deep Vent DNA polymerase, and KOD DNA polymerase. The phi29 DNA polymerase can be a wild-type phi29 DNA polymerase (eg, MagniPhi from Expedeon), or a variant EquiPhi29 DNA polymerase (eg, from Thermo Fisher Scientific), or a chimeric QualiPhi DNA polymerase (eg, from 4basebio).
[0462] As used herein, the term "DNA primer-polymerase" and related terms refer to an enzyme having the activity of a DNA polymerase and an RNA primer. DNA primer-polymerase can utilize deoxyribonucleotide triphosphates to synthesize DNA primers on a single-stranded DNA template in a template sequence-dependent manner, and can extend the primer chain via nucleotide polymerization (e.g., primer extension) in the presence of catalytic divalent cations (e.g., magnesium and / or manganese). DNA primer-polymerase includes enzymes that are members of DnaG-like primers (e.g., bacteria) and AEP-like primers (archaea and eukaryotes). Exemplary DNA primer-polymerase is Tth PrimPol from Thermus thermophilus HB27.
[0463] As used herein, the term "fidelity" refers to the accuracy of DNA polymerization by a template-dependent DNA polymerase. The fidelity of a DNA polymerase is typically measured by the error rate (the frequency of incorporation of inaccurate nucleotides (i.e., nucleotides that are not complementary to the template nucleotides). The accuracy or fidelity of DNA polymerization is maintained by the polymerase activity and 3'-5' exonuclease activity of the DNA polymerase.
[0464] As used herein, the term "binding complex" refers to a complex formed by binding together a nucleic acid duplex, a polymerase, and free nucleotides or nucleotide units of a multivalent molecule, wherein the nucleic acid duplex includes a nucleic acid template molecule hybridized to a nucleic acid primer. In the binding complex, the free nucleotides or nucleotide units may or may not bind to the 3' end of the nucleic acid primer at a position opposite to the complementary nucleotide in the nucleic acid template molecule. A "ternary complex" is an example of a binding complex formed by binding together a nucleic acid duplex, a polymerase, and free nucleotides or nucleotide units of a multivalent molecule, wherein the free nucleotides or nucleotide units bind to the 3' end of the nucleic acid primer (as part of the nucleic acid duplex) at a position opposite to the complementary nucleotide in the nucleic acid template molecule.
[0465] The term "residence time" and related terms refer to the length of time that a binding complex remains stable without any component dissociation, wherein the components of the binding complex include nucleic acid templates and nucleic acid primers, polymerases, nucleotide units or free (e.g., unconjugated) nucleotides of multivalent molecules. Nucleotide units or free nucleotides can be complementary or non-complementary to the nucleotide residues in the template molecule. Nucleotide units or free nucleotides can be combined with the 3' end of the nucleic acid primer at a position relative to the complementary nucleotide residues in the nucleic acid template molecule. The stability of the binding complex and the intensity of the binding interaction are indicated by the retention time. The retention time can be measured by observing the start and / or duration of the binding complex (e.g., by observing the signal of the labeled component from the binding complex). For example, labeled nucleotides or labeled reagents comprising one or more nucleotides can be present in the binding complex, thereby allowing the signal from the label to be detected during the residence time of the binding complex. An exemplary label is a fluorescent marker. Binding complex (e.g., a ternary complex) remains stable before being subjected to the conditions of the interactional dissociation that causes polymerase, template molecule, primer and / or nucleotide units or nucleotides. For example, dissociation conditions include contacting the bound complex with any one or any combination of detergent, EDTA, and / or water.
[0466] The term "primer" and related terms used herein refer to oligonucleotides that can hybridize with DNA and / or RNA polynucleotide templates to form duplex molecules. Primers include natural nucleotides and / or nucleotide analogs. Primers can be recombinant nucleic acid molecules. Primers can have any length, but generally range from 4 to 50 nucleotides. Typical primers include 5' ends and 3' ends. The 3' end of the primer can include a 3'OH portion that serves as a nucleotide polymerization initiation site in the primer extension reaction catalyzed by polymerase. Alternatively, the 3' end of the primer can lack a 3'OH portion or can include an end 3' blocking group that inhibits the nucleotide polymerization in the reaction catalyzed by polymerase. Any one nucleotide along the length of the primer or more than one nucleotide can be labeled with a detectable reporter gene portion. Primers can be in solution (e.g., soluble primers) or can be fixed to a carrier (e.g., capture primers).
[0467] When used to refer to nucleic acid molecules, the term "hybridize" or "hybridizing" or "hybridization" or other related terms refer to hydrogen bonding between two different nucleic acids to form a duplex nucleic acid. Hybridization also includes hydrogen bonding between two different regions of a single nucleic acid molecule to form a self-hybridizing molecule with a duplex region. Hybridization can include Watson-Crick or Hoogstein binding to form a duplex double-stranded nucleic acid or a double-stranded region within a nucleic acid molecule. The two different regions of a double-stranded nucleic acid or a single nucleic acid can be fully complementary or partially complementary. The complementary nucleic acid chains do not need to hybridize to each other across their entire length. Complementary base pairing can be standard AT or CG base pairing or can be other forms of base pairing interactions. The duplex nucleic acid can contain mismatched base-paired nucleotides.
[0468] When used in reference to nucleic acids, the terms "extend," "extending," "extension," and other variants refer to the incorporation of one or more nucleotides into a nucleic acid molecule. Nucleotide incorporation includes the polymerization of one or more nucleotides into the terminal 3' OH end of a nucleic acid chain (e.g., a nucleic acid primer), resulting in an extension of the nucleic acid chain (e.g., an extended primer). Nucleotide incorporation can be performed using natural nucleotides and / or nucleotide analogs. Typically, but not necessarily, nucleotide incorporation occurs in a template-dependent manner. Any suitable method for extending nucleic acid molecules can be used, including primer extension catalyzed by a DNA polymerase or an RNA polymerase.
[0469] In some embodiments, any of the amplification primer sequence, sequencing primer sequence, capture primer sequence (capture oligonucleotide), target capture sequence, circularization anchor sequence, sample barcode sequence, spatial barcode sequence, or anchor region sequence can be about 3 to 50 nucleotides in length, about 5 to 40 nucleotides in length, or about 5 to 25 nucleotides in length.
[0470] The term "nucleotide" and related terms refer to molecules comprising an aromatic base, a five-carbon sugar (e.g., ribose or deoxyribose) and at least one phosphate group. Standard or non-standard nucleotides are consistent with the use of the term. In certain embodiments, phosphates include monophosphates, diphosphates, or triphosphates or corresponding phosphate analogs. The term "nucleoside" refers to a molecule comprising an aromatic base and a sugar. Nucleotides and nucleosides can be unlabeled or labeled with a detectable reporter moiety.
[0471] Nucleotides (and nucleosides) typically contain heterocyclic bases, including substituted or unsubstituted nitrogen-containing parent heteroaromatic rings, which are commonly found in nucleic acids, including naturally occurring, substituted, modified or engineered variants or analogs thereof. The base of a nucleotide (or nucleoside) is capable of forming Watson-Crick and / or Hoostein hydrogen bonds with an appropriate complementary base. Exemplary bases include, but are not limited to, purines and pyrimidines, such as: 2-aminopurine, 2,6-diaminopurine, adenine (A), ethyleneadenine, N 6 -Δ 2 -Isopentenyl adenine (6iA), N 6 -Δ 2 -Isopentenyl-2-methylthioadenine (2ms6iA), N 6 -methyladenine, guanine (G), isoguanine, N 2 -dimethylguanine (dmG), 7-methylguanine (7mG), 2-thiopyrimidine, 6-thioguanine (6sG), hypoxanthine and O 6 -methylguanine; 7-deaza-purines, such as 7-deazaadenine (7-deazaadenine-A) and 7-deazaguanine (7-deaza-G); pyrimidines, such as cytosine (C), 5-propynylcytosine, isocytosine, thymine (T), 4-thiothymine (4sT), 5,6-dihydrothymine, O 4-methylthymine, uracil (U), 4-thiouracil (4sU) and 5,6-dihydrouracil (dihydrouracil; D); indoles such as nitroindole and 4-methylindole; pyrroles such as nitropyrrole; muscimol; inosine; hydroxymethylcytosine; 5-methylcytosine; base (Y); and methylated, glycosylated and acylated base moieties; etc. Additional exemplary bases can be found in Fasman, 1989, in "Practical Handbook of Biochemistry and Molecular Biology", pages 385-394, CRC Press, Boca Raton, Fla.
[0472] Nucleotides (and nucleosides) generally contain a sugar moiety, such as a carbocyclic moiety (Ferraro and Gotor 2000 Chem. Rev. 100:4319-48), an acyclic moiety (Martinez et al., 1999 Nucleic Acids Research 27:1271-1274; Martinez et al., 1997 Bioorganic & Medicinal Chemistry Letters Vol. 7:3013-3016), and other sugar moieties (Joeng et al., 1993 J. Med. Chem. 36:2627-2638; Kim et al., 1993 J. Med. Chem. 36:30-7; Eschenmosser 1999 Science 284:2118-2124; and U.S. Pat. No. 5,558,991). Sugar moieties include: ribosyl; 2'-deoxyribosyl; 3'-deoxyribosyl; 2',3'-dideoxyribosyl; 2',3'-didehydrodideoxyribosyl; 2'-alkoxyribosyl; 2'-azidoribosyl; 2'-aminoribosyl; 2'-fluororibosyl; 2'-thiolribosyl; 2'-alkylthioribosyl; 3'-alkoxyribosyl; 3'-azidoribosyl; 3'-aminoribosyl; 3'-fluororibosyl; 3'-thiolribosyl; 3'-alkylthioribosyl carbocyclic; acyclic or other modified sugars.
[0473] In certain embodiments, the nucleotide comprises a chain of one, two or three phosphorus atoms, wherein the chain is typically attached to the 5' carbon of the sugar moiety via an ester or phosphoramide bond. In certain embodiments, the nucleotide is an analog with a phosphorus chain, wherein the phosphorus atom is linked together with an O, S, NH, methylene or ethylene group in the middle. In certain embodiments, the phosphorus atom in the chain comprises a substituted side group (including O, S or BH3). In certain embodiments, the chain comprises a phosphate group substituted with an analog, including phosphoramide, phosphorothioate, phosphorodithioate and O-methylphosphoramidite groups.
[0474] The term "reporter moiety," "reporter moieties," or related terms refers to a compound that produces or causes a detectable signal. A reporter moiety is sometimes referred to as a "label." Any suitable reporter moiety can be used, including luminescence, photoluminescence, electroluminescence, bioluminescence, chemiluminescence, fluorescence, phosphorescence, chromophores, radioisotopes, electrochemistry, mass spectrometry, Raman, haptens, affinity tags, atoms, or enzymes. The reporter moiety produces a detectable signal caused by a chemical or physical change (e.g., heat, light, electricity, pH, salt concentration, enzymatic activity, or a neighboring event). A neighboring event includes two reporter moieties approaching each other, or associating with each other, or binding to each other. It is well known to those skilled in the art that the reporter moiety is selected so that each reporter moiety absorbs excitation radiation and / or fluoresces at a wavelength that is distinguishable from other reporter moieties, to allow monitoring of the presence of different reporter moieties in the same reaction or in different reactions. Two or more different reporter moieties can be selected that have spectrally distinct emission curves or have minimally overlapping spectral emission curves. The reporter moiety can be linked (eg, operably linked) to a nucleotide, nucleoside, nucleic acid, enzyme (eg, polymerase or reverse transcriptase), or carrier (eg, surface).
[0475] Reporter moieties (or labels) include fluorescent labels or fluorophores. Exemplary fluorescent moieties that can be used as fluorescent labels or fluorophores include, but are not limited to, fluorescein and fluorescein derivatives such as carboxyfluorescein, tetrachlorofluorescein, hexachlorofluorescein, carboxynaphthol fluorescein, fluorescein isothiocyanate, NHS-fluorescein, iodoacetamido fluorescein, maleimide fluorescein, SAMSA-fluorescein, thiosemicarbazide fluorescein, hydrazine carbonate methylthioacetamido fluorescein, rhodamine and rhodamine derivatives such as TRITC, TMR, lissamine rhodamine, Texas Red, Rhodamine B, Rhodamine 6G, Rhodamine 10, NHS-rhodamine, TMR-iodoacetamide, lissamine rhodamine B sulfonyl chloride, lissamine rhodamine B sulfonyl hydrazide, Texas Red sulfonyl chloride, Texas Red hydrazide, coumarin and coumarin derivatives such as AMCA, AMCA-NHS, AMCA-sulfo-NHS, AMCA-HPDP, DCIA, AMCE-hydrazide, BODIPY and derivatives such as BODIPY FL C3-SE, BODIPY 530 / 550C3, BODIPY 530 / 550C3-SE, BODIPY 530 / 550C3 hydrazide, BODIPY 493 / 503C3 hydrazide, BODIPY FL C3 hydrazide, BODIPY FLIA, BODIPY 530 / 551IA, Br-BODIPY 493 / 503, Cascade Blue and derivatives such as Cascade Blue acetyl triazoide, Cascade Blue cadaverine, Cascade Blue ethylenediamine, Cascade Blue hydrazide, Lucifer Yellow and derivatives such as Lucifer Yellow iodoacetamide, Lucifer Yellow CH, cyanines and derivatives such as indolium cyanine dyes, benzindolium cyanine dyes, pyridinium cyanine dyes, thiazolium cyanine dyes, quinolinium cyanine dyes, imidazolium cyanine dyes, Cy 3, Cy5, lanthanide chelates and derivatives such as BCPDA, TBP, TMT, BHHCT, BCOT, europium chelates, terbium chelates, Alexa Fluor dyes, DyLight dyes, Atto dyes, LightCycler Red dyes, CAL Flour dyes, JOE and its derivatives, Oregon Green dyes, WellRED dyes, IRD dyes, phycoerythrin and phycobilin dyes, malachite green, diphenylethylene, DEG dyes, NR dyes, near infrared dyes and other dyes known in the art, such as Haugland, Molecular Probes Handbook, (Eugene, Oreg.) 6th edition; Lakowicz,Principles of Fluorescence Spectroscopy, 2nd ed., Plenum Press New York (1999), or those described in Hermanson, Bioconjugate Techniques, 2nd ed., or derivatives thereof, or any combination thereof. Cyanine dyes may exist in sulfonated or non-sulfonated form and consist of two indolenine, benzindolium, pyridinium, thiazolium and / or quinolinium groups separated by a polymethine bridge between the two nitrogen atoms. Commercially available cyanine fluorophores include, for example, Cy3 (which may comprise 1-[6-(2,5-dioxopyrrolidin-1-yloxy)-6-oxohexyl]-2-(3-{1-[6-(2,5-dioxopyrrolidin-1-yloxy)-6-oxohexyl]-3,3-dimethyl-1,3-dihydro-2H-indol-2-ylidene}prop-1-en-1-yl)-3,3-dimethyl-3H-indolium or 1-[6-(2,5-dioxopyrrolidin-1-yloxy)-6-oxohexyl]-2 -(3-{1-[6-(2,5-dioxopyrrolidin-1-yloxy)-6-oxohexyl]-3,3-dimethyl-5-sulfo-1,3-dihydro-2H-indol-2-ylidene}prop-1-en-1-yl)-3,3-dimethyl-3H-indolium-5-sulfonate), Cy5 (which may include 1-(6-((2,5-dioxopyrrolidin-1-yl)oxy)-6-oxohexyl)-2-((1E,3E)-5-((E)-1-(6-((2,5-dioxopyrrolidin-1-yl)oxy)-6-oxohexyl)-2-((1E,3E)-5-((E)-1-(6-((2,5-dioxopyrrolidin-1 -yl)oxy)-6-oxohexyl)-3,3-dimethyl-5-indol-2-ylidene)penta-1,3-dien-1-yl)-3,3-dimethyl-3H-indol-1-ium or 1-(6-((2,5-dioxopyrrolidin-1-yl)oxy)-6-oxohexyl)-2-((1E,3E)-5-((E)-1-(6-((2,5-dioxopyrrolidin-1-yl)oxy)-6-oxohexyl)-3,3-dimethyl-5-sulfoindol-2-ylidene)penta-1,3-dien-1-yl)-3,3-dimethyl-3H-indol-1-ium 1-yl)-3,3-dimethyl-3H-indol-1-ium-5-sulfonate) and Cy7 (which may comprise 1-(5-carboxypentyl)-2-[(1E,3E,5E,7Z)-7-(1-ethyl-1,3-dihydro-2H-indol-2-ylidene)hepta-1,3,5-trien-1-yl]-3H-indolium or 1-(5-carboxypentyl)-2-[(1E,3E,5E,7Z)-7-(1-ethyl-5-sulfo-1,3-dihydro-2H-indol-2-ylidene)hepta-1,3,5-trien-1-yl]-3H-indolium5-trien-1-yl]-3H-indolium-5-sulfonate), where "Cy" stands for 'cyanine' and the first digit identifies the number of carbon atoms between the two indolenine groups. Cy2 is an oxazole derivative rather than an indolenine, and the benzo-derived Cy3.5, Cy5.5, and Cy7.5 are exceptions to this rule.
[0476] In some embodiments, the reporter moiety can be a FRET pair, so that multiple classifications can be performed in a single excitation and imaging step. As used herein, FRET can include excitation exchange (Forster) transfer or electron exchange (Dexter) transfer.
[0477] As used herein, the term "support" refers to a substrate designed for the deposition of biological molecules or biological samples for measurement and / or analysis. Examples of biological molecules to be deposited on a support include nucleic acids (e.g., DNA, RNA), polypeptides, carbohydrates, lipids, single cells or multiple cells. Examples of biological samples include, but are not limited to, saliva, sputum, mucus, blood, plasma, serum, urine, feces, sweat, tears, and fluids from tissues or organs.
[0478] In some embodiments, the support is solid, semi-solid, or a combination thereof. In some embodiments, the support is porous, semi-porous, non-porous, or any combination of porosity. In some embodiments, the support can be substantially planar, concave, convex, or any combination thereof. In some embodiments, the support can be cylindrical, for example, comprising a capillary or the inner surface of a capillary.
[0479] In some embodiments, the surface of the support can be substantially smooth. In some embodiments, the support can have a regular or irregular texture comprising bumps, etchings, holes, a three-dimensional scaffold, or any combination thereof.
[0480] In some embodiments, the support comprises beads having any shape, including spheres, hemispheres, cylinders, barrels, rings, disks, rods, cones, triangles, cubes, polygons, tubes, or wires.
[0481] The support can be made of any material, including but not limited to: glass, fused silica, silicon, polymers (e.g., polystyrene (PS), macroporous polystyrene (MPPS), polymethyl methacrylate (PMMA), polycarbonate (PC), polypropylene (PP), polyethylene (PE), high-density polyethylene (HDPE), cyclic olefin polymer (COP), cyclic olefin copolymer (COC), polyethylene terephthalate (PET)), or any combination thereof. Various compositions of both glass and plastic substrates are contemplated.
[0482] The support can have multiple (for example, two or more) nucleic acid templates fixed thereon.The multiple fixed nucleic acid templates have the same sequence or have different sequences.In certain embodiments, the independent nucleic acid template molecules in the multiple nucleic acid templates are fixed to different sites on the support.In certain embodiments, two or more independent nucleic acid template molecules in the multiple nucleic acid templates are fixed to sites on the support.
[0483] The term "array" refers to a support comprising a plurality of sites located at predetermined positions on the support to form an array of sites. The sites may be discrete and separated by gaps. In some embodiments, the predetermined sites on the support may be arranged in one dimension into rows or columns, or in two dimensions into rows and columns. In some embodiments, a plurality of predetermined sites are arranged in an organized manner on the support. In some embodiments, the plurality of predetermined sites are arranged in any organized pattern, including a straight line, a hexagonal pattern, a grid pattern, a pattern with reflection symmetry, or a pattern with rotational symmetry, etc. The spacing between different pairs of sites may be the same or may be different. In some embodiments, the support comprises at least 10 2 sites, at least 10 3 sites, at least 10 4 sites, at least 10 5 sites, at least 10 6 sites, at least 10 7 sites, at least 10 8 sites, at least 10 9 sites, at least 10 10 sites, at least 10 11 sites, at least 10 12 sites, at least 10 13 sites, at least 10 14 sites, at least 10 15 In some embodiments, the plurality of predetermined sites (e.g., 10 2 to 10 15 In some embodiments, the nucleic acid template is immobilized at a plurality of predetermined sites, for example, at 10 or more sites, to form a nucleic acid template array. In some embodiments, the nucleic acid template is immobilized at a plurality of predetermined sites by hybridization with an immobilized surface capture primer, or the nucleic acid template is covalently attached to a surface capture primer. In some embodiments, the nucleic acid template is immobilized at a plurality of predetermined sites, for example, at 10 or more sites. 2 to 10 15In certain embodiments, the fixed nucleic acid template is amplified in a clonal manner to produce fixed nucleic acid polymerase colonies at multiple predetermined sites. In certain embodiments, the independent fixed nucleic acid polymerase colonies comprise single-stranded or double-stranded concatemers.
[0484] In some embodiments, a support comprising a plurality of sites located at random positions on the support is referred to herein as a support having randomly positioned sites thereon. The positions of the randomly positioned sites on the support are not predetermined. The plurality of randomly positioned sites are arranged in a disordered and / or unpredictable manner on the support. In some embodiments, the support comprises at least 10 2 sites, at least 10 3 sites, at least 10 4 sites, at least 10 5 sites, at least 10 6 sites, at least 10 7 sites, at least 10 8 sites, at least 10 9 sites, at least 10 10 sites, at least 10 11 sites, at least 10 12 sites, at least 10 13 sites, at least 10 14 sites, at least 10 15 sites or more, wherein the sites are randomly positioned on the support. In some embodiments, the plurality of randomly positioned sites (e.g., 10 2 to 10 15 In some embodiments, the nucleic acid template is immobilized at a plurality of randomly located sites by hybridization with an immobilized surface capture primer, or the nucleic acid template is covalently attached to a surface capture primer. In some embodiments, the nucleic acid template is immobilized at a plurality of randomly located sites, for example, at 10 2 to 10 15 In certain embodiments, the nucleic acid template of the present invention can be fixed at a plurality of sites or more sites.In ...
[0485] When used to refer to a low binding surface coating, one or more layers of the multilayer surface coating may comprise a branched polymer or may be linear. Examples of suitable branched polymers include, but are not limited to, branched PEG, branched poly(vinyl alcohol) (branched PVA), branched poly(vinyl pyridine), branched poly(vinyl pyrrolidone) (branched PVP), branched), poly(acrylic acid) (branched PAA), branched polyacrylamide, branched poly(N-isopropylacrylamide) (branched PNIPAM), branched poly(methyl methacrylate) (branched PMA), branched poly(2-hydroxyethyl methacrylate) (branched PHEMA), branched poly(oligo(ethylene glycol) methyl ether methacrylate) (branched POEGMA), branched polyglutamic acid (branched PGA), branched polylysine, branched polyglucosides, and dextran.
[0486] In some embodiments, the branched polymer used to create one or more layers of any of the multi-layer surfaces disclosed herein may comprise: at least 4 branches, at least 5 branches, at least 6 branches, at least 7 branches, at least 8 branches, at least 9 branches, at least 10 branches, at least 12 branches, at least 14 branches, at least 16 branches, at least 18 branches, at least 20 branches, at least 22 branches, at least 24 branches, at least 26 branches, at least 28 branches, at least 30 branches, at least 32 branches, at least 34 branches, at least 36 branches, at least 38 branches, or at least 40 branches.
[0487] The linear, branched, or multi-branched polymers used to form one or more layers of any of the multi-layer surfaces disclosed herein can have a molecular weight of at least 500, at least 1,000, at least 2,000, at least 3,000, at least 4,000, at least 5,000, at least 10,000, at least 15,000, at least 20,000, at least 25,000, at least 30,000, at least 35,000, at least 40,000, at least 45,000, or at least 50,000 Daltons.
[0488] In some embodiments, for example, where at least one layer of the multilayer surface comprises a branched polymer, the number of covalent bonds between the branched polymer molecules of the layer being deposited and the molecules of the previous layer can range from about one covalent bond per molecule to about 32 covalent bonds per molecule. In some embodiments, the number of covalent bonds between the branched polymer molecules of the new layer and the molecules of the previous layer can be at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 12, at least 14, at least 16, at least 18, at least 20, at least 22, at least 24, at least 26, at least 28, at least 30, or at least 32 covalent bonds per molecule.
[0489] Any reactive functional groups remaining after coupling a material layer to a surface can optionally be blocked by coupling a small, inert molecule using a high-yield coupling chemistry. For example, where an amine coupling chemistry is used to attach a new material layer to a previous one, any residual amine groups can then be acetylated or inactivated by coupling with a small amino acid such as glycine.
[0490] The number of layers of low non-specific binding material (e.g., hydrophilic polymer material) deposited on the surface can range from 1 to about 10. In some embodiments, the number of layers is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10. In some embodiments, the number of layers can be at most 10, at most 9, at most 8, at most 7, at most 6, at most 5, at most 4, at most 3, at most 2, or at most 1. Any of the lower and upper limits described in this paragraph can be combined to form a range included in the present disclosure, for example, in some embodiments, the number of layers can be in the range of about 2 to about 4. In some embodiments, all layers can comprise the same material. In some embodiments, each layer can comprise different materials. In some embodiments, multiple layers can comprise multiple materials. In some embodiments, at least one layer can comprise a branched polymer. In some embodiments, all layers can comprise a branched polymer.
[0491] In some cases, one or more layers of low non-specific binding material can be deposited on the substrate surface and / or conjugated to the substrate surface using a polar protic solvent, a polar or polar aprotic solvent, a non-polar solvent, or any combination thereof. In some embodiments, the solvent used for layer deposition and / or conjugation can comprise an alcohol (e.g., methanol, ethanol, propanol, etc.), another organic solvent (e.g., acetonitrile, dimethyl sulfoxide (DMSO), dimethylformamide (DMF), etc.), water, an aqueous buffer solution (e.g., phosphate buffer, phosphate buffered saline, 3-(N-morpholino)propanesulfonic acid (MOPS), etc.), or any combination thereof. In some embodiments, the organic component of the solvent mixture used may comprise at least 1%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, or 99% of the total amount, with the balance being water or an aqueous buffer solution. In some embodiments, the aqueous component of the solvent mixture used may comprise at least 1%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, or 99% of the total amount, with the balance being organic solvent. The pH of the solvent mixture used may be less than 6, about 6, 6.5, 7, 7.5, 8, 8.5, 9, or greater than pH 9.
[0492] The term "branched polymer" and related terms refer to polymers with multiple functional groups that contribute to the conjugation of bioactive molecules (such as nucleotides), and the functional group can be on the side chain of the polymer or directly attached to the central core or central backbone of the polymer. Branched polymers can have a linear backbone, in which one or more functional groups are separated from the backbone to be conjugated. Branched polymers can also be polymers with one or more side chains, in which the side chains have sites suitable for conjugation. The example of functional group includes but is not limited to: hydroxyl, ester, amine, carbonate, acetal, aldehyde, aldehyde hydrate, alkenyl, acrylate, methacrylate, acrylamide, active sulfone, hydrazide, thiol, alkanoic acid, acyl halide, isocyanate, isothiocyanate, maleimide, vinyl sulfone, dithiopyridine, vinylpyridine, iodoacetamide, epoxide, glyoxal, diketone, mesylate, tosylate and trifluoroethylsulfonate (tresylate).
[0493] The term "immobilized" and related terms, when used to refer to immobilized nucleic acids, refers to nucleic acid molecules that are attached to a support, or to a coating on a support, or embedded within a matrix formed by a coating on a support, by covalent bonds or non-covalent interactions, wherein the nucleic acid molecules comprise surface capture primers, nucleic acid template molecules, and extension products of the capture primers. The extension products of the capture primers include nucleic acid concatemers that can form a nucleic acid polymerase community.
[0494] In certain embodiments, one or more nucleic acid templates are fixed on a support, for example, fixed on a site on a support. In certain embodiments, the one or more nucleic acid templates are for being amplified in a clonal manner. In some embodiments, the one or more nucleic acid templates are amplified (for example, in a solution) in a clonal manner from a support, and are subsequently deposited on a support and are fixed on a support. In certain embodiments, the clonal amplification reaction of the one or more nucleic acid templates is carried out on a support, resulting in fixing on a support. In some embodiments, nucleic acid amplification reaction is used to amplify the one or more nucleic acid templates (for example, in a solution or on a support) in a clonal manner, the nucleic acid amplification reaction includes any one of the following or any combination thereof: polymerase chain reaction (PCR), multiple displacement amplification (MDA), transcription-mediated amplification (TMA), amplification based on nucleic acid sequence (NASBA), chain displacement amplification (SDA), real-time SDA, bridge amplification, isothermal bridge amplification, rolling circle amplification (RCA), ring to ring amplification, helicase-dependent amplification, recombinase-dependent amplification and / or single-stranded binding (SSB) protein-dependent amplification.
[0495] The term "surface primer," "surface capture primer" and related terms refer to single-stranded oligonucleotides that are fixed to a support and comprise a sequence that can hybridize with at least a portion of a nucleic acid template molecule. Surface primers can be used for fixing template molecules to a support via hybridization. Surface primers can be fixed to a support in a manner that resists primer removal during flow, washing, suction, and changes in temperature, pH, salt, chemical, and / or enzyme conditions. Typically, but not necessarily, the 5' end of the surface primer can be fixed to a support. Alternatively, an inner portion or 3' end of the surface primer can be fixed to a support.
[0496] The surface primer comprises DNA, RNA or their analog.The surface primer can comprise the combination of DNA and RNA.The sequence of the surface primer can be fully complementary or partially complementary to at least a portion of the nucleic acid template molecule (for example, linear or circular template molecule) along its length.The support can comprise a plurality of fixed surface primers with the same sequence or with two or more different sequences.The surface primer can be any length, for example 4 to 50 nucleotides, or 50 to 100 nucleotides, or 100 to 150 nucleotides or longer length.
[0497] The surface primer may include a terminal 3' nucleotide with a sugar 3' OH portion that can be extended for nucleotide polymerization (for example, polymerase-catalyzed polymerization). The surface primer may include a terminal 3' nucleotide with a portion of the extension that blocks polymerase-catalyzed extension. The surface primer may include a terminal 3' nucleotide with a 3' sugar position that is connected to a chain termination portion that inhibits nucleotide polymerization. Blocking agents may be used to remove (for example, to block) the 3' chain termination portion, so that the 3' end is converted into an extendable 3' OH end. The example of the chain termination portion includes an alkyl group, an alkenyl group, an alkynyl group, an allyl group, an aryl group, a benzyl group, an azide group, an amine group, an amide group, a keto group, an isocyanate group, a phosphate group, a thio group, a disulfide group, a carbonate group, a urea group or a silyl group. The azido-type chain termination portion includes an azido, an azido and an azidomethyl group. Examples of deblocking agents include phosphine compounds such as tris(2-carboxyethyl)phosphine (TCEP) and bissulfotriphenylphosphine (BS-TPP) for azide, azido and azidomethyl chain terminating groups. Examples of deblocking agents include tetrakis(triphenylphosphine)palladium(0) (Pd(PPh3)4) with piperidine, or with 2,3-dichloro-5,6-dicyano-1,4-benzoquinone (DDQ) for alkyl, alkenyl, alkynyl and allyl chain terminating groups. Examples of deblocking agents include Pd / C for aryl and benzyl chain terminating groups. Examples of deblocking agents include phosphine, β-mercaptoethanol or dithiothreitol (DTT) for amine, amide, ketone, isocyanate, phosphate, thio and disulfide chain terminating groups. Examples of deblocking agents include potassium carbonate (K2CO3) in MeOH, triethylamine in pyridine and Zn in acetic acid (AcOH) for carbonate chain terminating groups. Examples of deblocking agents include tetrabutylammonium fluoride, pyridine-HF, ammonium fluoride, and triethylamine trihydrofluoride for the chain-terminating groups urea and silyl.
[0498] In some embodiments, the plurality of fixed surface capture primers on the support are fluidically connected to each other to allow a solution of reagents (e.g., linear or circular nucleic acid template molecules, soluble primers, enzymes, nucleotides, divalent cations, buffers, reagents, etc.) to flow onto the support so that the plurality of fixed surface capture primers on the support can react with the reagents substantially simultaneously in a large-scale parallel manner. In some embodiments, the fluidic connection of the plurality of fixed surface capture primers can be used to perform nucleic acid amplification reactions (e.g., RCA, MDA, PCR, and bridge amplification) on the plurality of fixed surface capture primers substantially simultaneously.
[0499] In some embodiments, a plurality of immobilized single-stranded nucleic acid concatemer template molecules on a support are in fluid communication with each other to allow a solution of reagents (e.g., soluble primers, enzymes, nucleotides, divalent cations, buffers, reagents, etc.) to flow onto the support, such that the plurality of immobilized concatemer template molecules on the support can react with the reagents substantially simultaneously in a massively parallel manner. In some embodiments, the fluid communication of the plurality of immobilized single-stranded nucleic acid concatemer template molecules can be used to perform nucleotide binding assays and / or nucleotide polymerization reactions (e.g., primer extension or sequencing) on the plurality of immobilized single-stranded nucleic acid concatemer template molecules substantially simultaneously, and optionally for detection and imaging to perform massively parallel sequencing.
[0500] The terms "amplify," "amplifying," "amplification," and other related terms when used in reference to nucleic acids include the production of multiple copies of an original polynucleotide template molecule, wherein the copies comprise a sequence that is complementary to the template sequence, or the copies comprise a sequence that is identical to the template sequence. In some embodiments, the copies comprise a sequence that is substantially identical to the template sequence or substantially identical to a sequence that is complementary to the template sequence.
[0501] The present disclosure provides various pH buffers. The full names of the pH buffers are listed herein.
[0502] The term "Tris" refers to the pH buffer tris(hydroxymethyl)-aminomethane. The term "Tris-HCl" refers to the pH buffer tris(hydroxymethyl)-aminomethane hydrochloride. The term "Tricine" refers to the pH buffer N-[tris(hydroxymethyl)methyl]glycine. The term "Bicine" refers to the pH buffer N,N-bis(2-hydroxyethyl)glycine. The term "Bis-Tris propane" refers to the pH buffer 1,3-bis(tris(hydroxymethyl)methylamino]propane. The term "HEPES" refers to the pH buffer 4-(2-hydroxyethyl)-1-piperazineethanesulfonic acid. The term "MES" refers to the pH buffer 2-(N-morpholino)ethanesulfonic acid. The term "MOPS" refers to the pH buffer 3-(N-morpholino)propanesulfonic acid. The term "MOPSO" refers to the pH buffer 3-(N-morpholino)-2-hydroxypropanesulfonic acid. The term "BES" refers to the pH buffer N,N-bis(2-hydroxyethyl)-2-aminoethanesulfonic acid. The term "TES" refers to the pH buffer 2-[(2-hydroxy-1,1-bis(hydroxymethyl)ethyl)amino]ethanesulfonic acid. The term "CAPS" refers to the pH buffer 3-(cyclohexylamino)-1-propanesulfonic acid. The term "TAPS" refers to the pH buffer N-[tris(hydroxymethyl)methyl]-3-aminopropanesulfonic acid. The term "TAPSO" refers to the pH buffer N-[tris(hydroxymethyl)methyl]-3-amino-2-hydroxypropanesulfonic acid. The term "ACES" refers to the pH buffer N-(2-acetamido)-2-aminoethanesulfonic acid. The term "PIPES" refers to the pH buffer piperazine-1,4-bis(2-ethanesulfonic acid).
[0503] All publications and patent documents cited herein are incorporated by reference as if each such publication or document was specifically and individually indicated to be incorporated by reference. Citation of publications and patent documents is not intended to be an admission of any relevant prior art and does not constitute an admission of their contents or date. The invention has now been described in the form of a written description, and those skilled in the art will recognize that the invention can be practiced in a variety of embodiments, and the foregoing description and the following examples are for illustrative purposes only and are not intended to limit the scope of the appended claims.
[0504] Examples
[0505] Unless otherwise stated, the values reported here are approximate and may be subject to normal instrumental and experimental errors.
[0506] Example 1. Synthesis of exemplary compounds
[0507] Exemplary compounds of the present disclosure, such as compounds 85 to 88, can be synthesized according to the following methods (eg, in Schemes 1A to 1D).
[0508]
[0509]
[0510] Example 2. Properties of Exemplary Compounds
[0511] The compounds disclosed herein can be used as dye molecules. Table 3 below reports the absorption and emission data for representative compounds of the present disclosure. Laser Condition 1 refers to a laser at 638 ± 7 nm and an emission range of 719.5 ± 32.5 nm. Laser Condition 2 refers to a laser at 638 ± 7 nm and an emission range of 674 ± 17 nm. Laser Condition 3 refers to a laser at 520 ± 7 nm and an emission range of 595.5 ± 24.5 nm. Laser Condition 4 refers to a laser at 520 ± 7 nm and an emission range of 555 ± 20 nm.
[0512] Table 3.
[0513]
[0514]
[0515]
[0516]
[0517] Equivalent
[0518] The details of one or more embodiments of the present disclosure are set forth in the above-appended description. Although any methods and materials similar or equivalent to those described herein can be used for the practice or testing of the present disclosure, preferred methods and materials are now described. Other features, objects, and advantages of the present disclosure will become apparent from this specification and the claims. In the specification and the appended claims, unless the context clearly provides otherwise, the singular form includes the plural indicator. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those of ordinary skill in the art to which the present invention belongs. All technical and scientific terms used herein have the same meaning as those of ordinary skill in the art to which the present disclosure belongs. All patents and publications cited in this specification are incorporated by reference.
[0519] The foregoing description is presented for purposes of illustration only and is not intended to limit the disclosure to the precise form disclosed, except as defined by the appended claims.
Claims
1. A compound of formula (I), (II) or (III): An ionic derivative thereof, an isomer thereof, or a salt thereof, wherein: An ionic derivative thereof, an isomer thereof, or a salt thereof, wherein: n is 0 or 1; m is 0 or 1; p is 0 or 1; R X and R Z are independently H, halogen, C 1-12 Alkyl, C 1-12 Alkenyl, C 1-12 Alkynyl, C6-C 10 aryl or 5- to 10-membered heteroaryl; or R X and R Z Together with the atoms to which they are attached, they form C6-C 10 Arylene or C5-C 10 cycloalkylene; R Y It is H, halogen, C 1-12 Alkyl, C 1-12 Alkenyl, C 1-12 Alkynyl, C6-C 10 aryl or 5- to 10-membered heteroaryl; T A Is -S-, -O- or -C(R TA )2-; Each R TA are independently H, halogen, C 1-12 Alkyl, C 1-12 Alkenyl, C 1-12 Alkynyl, C6-C 10 aryl or 5- to 10-membered heteroaryl; wherein said C 1-12 Alkyl, C 1-12 Alkynyl, C6-C 10 Aryl or 5- to 10-membered heteroaryl is optionally substituted with one or more -S(=O)2OH or -C(=O)OH; T B Is -S-, -O- or -C(R TB )2-; Each R TB are independently H, halogen, C 1-12 Alkyl, C 1-12 Alkenyl, C 1-12 Alkynyl, C6-C 10 aryl or 5- to 10-membered heteroaryl; wherein said C 1-12 Alkyl, C 1-12 Alkynyl, C6-C 10 Aryl or 5- to 10-membered heteroaryl is optionally substituted with one or more -S(=O)2OH or -C(=O)OH; R NA It is C 1-12 Alkyl, C 1-12 Alkenyl or C 1-12 Alkynyl, wherein the C 1-12 Alkyl, C 1-12 Alkenyl or C 1-12 Alkynyl is optionally substituted with one or more R NA1 replace; Each R NA1 are independently -S(=O)2OH, -C(=O)OH, -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 Alkynyl), wherein the -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 Alkynyl) is optionally substituted by one or more R NA2 replace; Each R NA2 are independently -S(=O)2OH, -C(=O)OH, -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 Alkynyl), wherein the -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 alkynyl) is optionally substituted by one or more R NA3 replace; Each R NA3 are independently -S(=O)2OH, -C(=O)OH, -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 Alkynyl), wherein the -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 alkynyl) is optionally substituted with one or more -S(=O)2OH or -C(=O)OH; R NB It is C 1-12 Alkyl, C 1-12 Alkenyl or C 1-12 Alkynyl, wherein the C 1-12 Alkyl, C 1-12 Alkenyl or C 1-12 Alkynyl is optionally substituted with one or more R NB1 replace; Each R NB1 are independently -S(=O)2OH, -C(=O)OH, -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 Alkynyl), wherein the -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 Alkynyl) is optionally substituted by one or more R NB2 replace; Each R NB2 are independently -S(=O)2OH, -C(=O)OH, -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 Alkynyl), wherein the -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 alkynyl) is optionally substituted by one or more R NB3 replace; Each R NB3 are independently -S(=O)2OH, -C(=O)OH, -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 Alkynyl), wherein the -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 alkynyl) is optionally substituted with one or more -S(=O)2OH or -C(=O)OH; R 1A 、R 2A 、R 3A 、R 4A 、R 5A 、R 6A 、R 7A and R 8A are independently H, halogen, -S(=O)2OH, -C(=O)OH, C 1-12 Alkyl, C 1-12 Alkenyl, C 1-12 Alkynyl, -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 Alkynyl), wherein the C 1-12 Alkyl, C 1-12 Alkenyl, C 1-12 Alkynyl, -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 alkynyl) is optionally substituted by one or more R AR replace; Each R AR are independently -S(=O)2OH, -C(=O)OH, -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 Alkynyl), wherein the -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 alkynyl) is optionally substituted by one or more R AR1 replace; Each R AR1 are independently -S(=O)2OH, -C(=O)OH, -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 Alkynyl), wherein the -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 alkynyl) is optionally substituted with one or more -S(=O)2OH or -C(=O)OH; R 1B 、R 2B 、R 3B 、R 4B 、R 5B 、R 6B 、R 7B and R 8B are independently H, halogen, -S(=O)2OH, -C(=O)OH, C 1-12 Alkyl, C 1-12 Alkenyl, C 1-12 Alkynyl, -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 Alkynyl), wherein the C 1-12 Alkyl, C 1-12 Alkenyl, C 1-12 Alkynyl, -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 alkynyl) is optionally substituted by one or more R BR replace; Each R AR are independently -S(=O)2OH, -C(=O)OH, -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 Alkynyl), wherein the -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 alkynyl) is optionally substituted by one or more R BR1 replace; and Each R BR1 are independently -S(=O)2OH, -C(=O)OH, -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 Alkynyl), wherein the -C(=O)-NH-(C 1-12 alkyl), -C(=O)-NH-(C 1-12 alkenyl) or -C(=O)-NH-(C 1-12 alkynyl) is optionally substituted with one or more -S(=O)2OH or -C(=O)OH.
2. The compound of claim 1 , which has formula (IA), (II-A), (III-A) or (IV-A): Ionic derivatives thereof, isomers thereof or salts thereof.
3. A compound according to any one of the preceding claims, having formula (IB), (II-B), (III-B) or (IV-B): Ionic derivatives thereof, isomers thereof or salts thereof.
4. A compound according to any one of the preceding claims, wherein R 2A 、R 5A 、R 7A 、R TA 、R NB 、R NA 、R TB 、R 2B 、R 5B 、R 6B 、R 7B 、R 4A 、R 4B 、R 8A 、R 8B At least one of them includes -SO3H.
5. A compound according to any one of the preceding claims, wherein when T A and R TB When it is CH3, then (i) R 2A 、R 4A 、R 4B 、R 2B 、R 5B and R 7B At least three of are -SO3H, (ii) n is 1, and (iii) the compound does not have Formula III.
6. A compound according to any one of the preceding claims, wherein when the compound has formula I and n is 1, then (i) R 2A 、R 4A 、R 4B and R 2B At least one of them is C(=O)NHCH2CH2SO3H, or (ii) R 2A 、R 2B 、R 4A 、R 4B 、R 5B 、R 5B 、R 7A and R 7B The three in it are -SO3H.
7. A compound according to any one of the preceding claims, wherein when (i) the compound has formula III, (ii) R 7 Yes – (C 2-12 Alkylene)-SO3H, and R 13 Yes – (C 2-12 alkylene)-C(=O)OH, and (iii) n is 2, then (i) R 5A 、R 7A 、R 5B and R 7B At least one of them is C(=O)OH, or (ii) R TA and R TB One of them is CH3.
8. A compound according to any one of the preceding claims, wherein when (i) the compound has formula III, (ii) R 7 Yes – (C 2-12 Alkylene)-SO3H, and R 13 Yes – (C 2-12 alkylene)-C(=O)OH, and (iii) n is 1, then (i) R 5A 、R 7A 、R 5B and R 7B At least one of them is C(=O)OH, or (ii) R TA and R TB Both are –(C 1-12 Alkylene)-SO3H.
9. A compound according to any one of the preceding claims, wherein each R TA is independently C optionally substituted with one or more -S(=O)2OH or -C(=O)OH 1-12 alkyl.
10. A compound according to any one of the preceding claims, wherein one or more R TA It is -CH2S(=O)2OH.
11. A compound according to any one of the preceding claims, wherein one or more R TA It is –(CH2)3-C(=O)OH.
12. A compound according to any one of the preceding claims, wherein each R TA Independently C 1-12 alkyl.
13. A compound according to any one of the preceding claims, wherein each R TB is independently C optionally substituted with one or more -S(=O)2OH or -C(=O)OH 1-12 alkyl.
14. A compound according to any one of the preceding claims, wherein one or more R TB It is -CH2S(=O)2OH.
15. A compound according to any one of the preceding claims, wherein one or more R TB It is –(CH2)3-C(=O)OH.
16. A compound according to any one of the preceding claims, wherein each R TB Independently C 1-12 alkyl.
17. A compound according to any one of the preceding claims, wherein R NA is optionally replaced by one or more R NA1 Substituted C 1-12 Alkyl, where each R NA1 are independently -S(=O)2OH, -C(=O)OH or -C(=O)-NH-(C 1-12 alkyl), wherein the -C(=O)-NH-(C 1-12 alkyl) is optionally substituted with one or more -S(=O)2OH, -C(=O)OH.
18. A compound according to any one of the preceding claims, wherein R NA is C optionally substituted with one or more -S(=O)2OH or -C(=O)OH 1-12 alkyl.
19. A compound according to any one of the preceding claims, wherein R NA It is -CH3.
20. A compound according to any one of the preceding claims, wherein R NB is optionally replaced by one or more R NB1 Substituted C 1-12 Alkyl, where each R NB1 are independently -S(=O)2OH, -C(=O)OH or -C(=O)-NH-(C 1-12 alkyl), wherein the -C(=O)-NH-(C 1-12 alkyl) is optionally substituted with one or more -S(=O)2OH, -C(=O)OH.
21. A compound according to any one of the preceding claims, wherein R NB is C optionally substituted with one or more -S(=O)2OH or -C(=O)OH 1-12 alkyl.
22. A compound according to any one of the preceding claims, wherein R NB It is -CH3.
23. A compound according to any one of the preceding claims, wherein R 1A 、R 3A 、R 5A 、R 6A and R 8A Each is independently H, halogen, -S(=O)2OH or -C(=O)OH.
24. A compound according to any one of the preceding claims, wherein R 1A 、R 3A 、R 5A 、R 6A and R 8A are each independently H or halogen.
25. A compound according to any one of the preceding claims, wherein R 1A 、R 3A 、R 5A 、R 6A and R 8A Each independently is H.
26. A compound according to any one of the preceding claims, wherein R 2A 、R 4A 、R 5A and R 7A are independently H, halogen, -S(=O)2OH, -C(=O)OH, C 1-12 Alkyl or -C(=O)-NH-(C 1-12 alkyl), wherein the C 1-12 Alkyl or -C(=O)-NH-(C 1-12 alkyl) is optionally substituted with -S(=O)2OH or -C(=O)OH.
27. A compound according to any one of the preceding claims, wherein R 2A 、R 4A 、R 5A and R 7A Each independently is -S(=O)2OH, -C(=O)OH, C 1-12 Alkyl or -C(=O)-NH-(C 1-12 alkyl), wherein the C 1-12 Alkyl or -C(=O)-NH-(C 1-12 alkyl) is optionally substituted with -S(=O)2OH or -C(=O)OH.
28. A compound according to any one of the preceding claims, wherein R 2A 、R 4A 、R 5A and R 7A Each is independently -S(=O)2OH or -C(=O)OH.
29. A compound according to any one of the preceding claims, wherein R 1B 、R 3B 、R 5B 、R 6B and R 8B Each is independently H, halogen, -S(=O)2OH or -C(=O)OH.
30. A compound according to any one of the preceding claims, wherein R 1B 、R 3B 、R 5B 、R 6B and R 8B are each independently H or halogen.
31. A compound according to any one of the preceding claims, wherein R 1B 、R 3B 、R 5B 、R 6B and R 8B Each independently is H.
32. A compound according to any one of the preceding claims, wherein R 2B 、R 4B 、R 5B and R 7B are independently H, halogen, -S(=O)2OH, -C(=O)OH, C 1-12 Alkyl or -C(=O)-NH-(C 1-12 alkyl), wherein the C 1-12 Alkyl or -C(=O)-NH-(C 1-12 alkyl) is optionally substituted with -S(=O)2OH or -C(=O)OH.
33. A compound according to any one of the preceding claims, wherein R 2B 、R 4B 、R 5B and R 7B Each independently is -S(=O)2OH, -C(=O)OH, C 1-12 Alkyl or -C(=O)-NH-(C 1-12 alkyl), wherein the C 1-12 Alkyl or -C(=O)-NH-(C 1-12 alkyl) is optionally substituted with -S(=O)2OH or -C(=O)OH.
34. A compound according to any one of the preceding claims, wherein R 2B 、R 4B 、R 5B and R 7B Each is independently -S(=O)2OH or -C(=O)OH.
35. A compound having the structure of any one of the compounds shown in Table 1, or an ionic derivative thereof, an isomer thereof, or a salt thereof.
36. The compound of claim 35, wherein the compound is selected from Compound 1, 2, 3, or 4.
37. The compound of claim 35, wherein the compound is selected from Compound 1, 2, 3, or 95.
38. A compound according to any one of claims 1 to 37 for use in a sequencing method.
39. A method of sequencing a nucleic acid comprising using a compound according to any one of claims 1 to 37.
40. The method of claim 39, wherein the sequencing comprises high-throughput sequencing.
41. A sequencing method comprising: (a) contacting (i) a plurality of polymerases, (ii) a plurality of nucleic acid template molecules, and (iii) a plurality of nucleic acid sequencing primers under conditions suitable for forming a plurality of complexes comprising the polymerases bound to nucleic acid duplexes, wherein the nucleic acid duplexes comprise nucleic acid template molecules hybridized to the primers; (b) contacting the plurality of complexes with a plurality of nucleotides under conditions suitable for binding of at least one nucleotide to one of the polymerases that bind to the nucleic acid duplex; as well as (c) incorporating at least one nucleotide into the 3' end of the extendable primer of at least one of the complexes, wherein at least one nucleotide of the plurality of nucleotides is labeled with a compound according to any one of claims 1 to 37.
42. The method of claim 41, wherein step (c) comprises a primer extension reaction.
43. The method of claim 41 or 42, comprising (d) repeating steps (b) and (c) at least once.
44. The method of any one of claims 41 to 43, wherein the compound is attached to the nucleotide base via a linker.
45. The method of claim 44, wherein the linker is a cleavable or removable linker.
46. The method of any one of claims 41 to 45, wherein the compound corresponds to a nucleotide base identity to allow detection and identification of the nucleotide base.
47. The method of any one of claims 41 to 46, comprising detecting at least one incorporated nucleotide at steps (c) and / or (d).
48. A sequencing method comprising: (a) subjecting (i) the first polymerase to a nucleic acid duplex under conditions suitable for forming a plurality of complexes comprising the first polymerase bound to the nucleic acid duplex; contacting a plurality of sequencing polymerases, (ii) a first plurality of nucleic acid template molecules, and (iii) a plurality of nucleic acid sequencing primers, wherein the nucleic acid duplex comprises nucleic acid template molecules hybridized to the nucleic acid sequencing primers; (b) contacting the plurality of complexes with a plurality of multivalent molecules to form a plurality of complexes, each complex comprising one or more polymerases, a nucleic acid template, and a sequencing primer, wherein the nucleic acid template and the sequencing primer are associated in a duplex form, wherein the individual multivalent molecules comprise a core attached to a plurality of nucleotide arms, and each nucleotide arm is attached to a nucleotide unit; and wherein an individual multivalent molecule of the plurality of multivalent molecules comprises a compound according to any one of claims 1 to 37; and (c) Detection of multiple multivalent-complexed polymerases.
49. The method of claim 48, wherein the contacting of step (b) occurs under conditions suitable to inhibit polymerase-catalyzed incorporation of the nucleotide units of the multivalent molecule into the 3' end of the sequencing primer of the complex.
50. The method of claim 48 or 49, wherein the individual nucleotide arms of the multivalent molecule comprise (i) a core attachment portion, (ii) a spacer, (iii) a linker, and (iv) a nucleotide unit, wherein the core is attached to the plurality of nucleotide arms, wherein the spacer is attached to the linker, and wherein the linker is attached to the nucleotide unit.
51. The method of any one of claims 48 to 50, wherein the multivalent molecule comprises the compound attached to the core, spacer, linker and / or nucleotide units.
52. The method of any one of claims 48 to 51, wherein the compound attached to a separate multivalent molecule corresponds to the identity of the nucleotide unit, thereby allowing identification of the complementary nucleotide in the nucleic acid template.
53. The method of any one of claims 48 to 52, comprising: (d) identifying the nucleobases of the nucleotide units bound to the first plurality of complexed polymerases, thereby determining the sequence of the nucleic acid template molecule.
54. A method according to any one of claims 48 to 53, comprising: (e) dissociating the plurality of multivalent-complexed polymerases and removing the first plurality of sequencing polymerases and their associated multivalent molecules, and retaining a plurality of nucleic acid duplexes.
55. The method of any one of claims 48 to 54, comprising: (f) contacting the plurality of retained nucleic acid duplexes of step (e) with a second plurality of sequencing polymerases under conditions suitable for binding of the second plurality of sequencing polymerases to the plurality of retained nucleic acid duplexes, thereby forming a second plurality of complexed polymerases comprising the second sequencing polymerase bound to the nucleic acid duplexes; (g) contacting the second plurality of complexed polymerases with a plurality of nucleotides, thereby incorporating complementary nucleotides into the sequencing primer of the nucleotide complexed polymerase in a primer extension reaction.
56. The method of claim 55, wherein a nucleotide in the plurality of nucleotides comprises the compound, and wherein the method comprises: (h) detecting the nucleotide incorporated into the sequencing primer; as well as (i) identifying the identity of the nucleobase of the nucleotide.
57. The method of any one of claims 48 to 56, wherein the nucleotide comprises a chain terminating moiety, and wherein the method comprises: (j) removing the chain terminating moiety from the incorporated nucleotide.
58. The method of claim 57, comprising repeating steps (a) to (j) at least once.
Citation Information
Patent Citations
Methods for assembling and reading nucleic acid sequences from mixed populations
US10233490B2
Method and system for sequencing nucleic acids
US10246744B2
Engineered polymerases for improved sequencing
US10731141B2
Multivalent binding composition for nucleic acid analysis
US10768173B1
Compositions and methods for pairwise sequencing
US11427855B1