Preparation and application of ribose coupled aliphatic chain nucleotide monomer
Through the click chemical reaction of the nucleotide monomer coupled to the fat chain at the 2' position of ribose and the GalNAc derivative, the problem of low internalization of oligonucleotide cells is solved, and efficient nucleic acid drug delivery is achieved.
Patent Information
- Application Number
- CN202410714637.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-04
- Publication Date
- 2025-08-01
AI Technical Summary
In the prior art, oligonucleotides containing 2’ azide or 2’ acetylene modification have low degree of internalization in the cell, resulting in a reduced delivery efficiency and the inability to achieve efficient targeted delivery of nucleic acid drugs.
A nucleotide monomer with a ribose 2' position coupled to a fat chain is provided, which stabilizes the conjugation of GalNAc derivatives or other ligand molecules through click chemical reactions, enhancing the cellular internalization ability of the oligonucleotide.
The cell internalization efficiency of oligonucleotide drugs is improved and efficient delivery of nucleic acid drugs is achieved.
Smart Images

Figure CN120398985A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of biomedicine, and in particular to the preparation and application of a ribose 2'-position coupled fatty chain nucleotide monomer. Background Art
[0002] Small nucleic acids, as functional molecules for gene regulation, have been widely used in the treatment of various diseases. Modulating the expression of disease-related genes through small nucleic acids to prevent, inhibit, or treat diseases has become a promising gene therapy approach. However, the precise and efficient delivery of small nucleic acids to disease-related tissues remains a key challenge in this field. The development of novel drug conjugates with high in vivo delivery efficiency is an urgent need in this field.
[0003] Oligonucleotide conjugates are a solution to the problem of small nucleic acid drug delivery. By conjugating molecules such as cholesterol, polycationic groups, polypeptide chains or fatty chains to oligonucleotides, the level of oligonucleotide recognition and internalization by cells can be significantly improved, thereby improving delivery efficiency. Since nucleotides modified at the 3' end are difficult to synthesize, and the modification and conjugation of nucleotide bases can easily interfere with base pairing, most oligonucleotide conjugates are modified molecules connected via the 5' end. As research deepens, in order to maintain the oligonucleotide in a 3'-endo configuration that is beneficial for delivery, thereby increasing the stability of the oligonucleotide and achieving the effect of enhancing drug delivery, the existing technology has proposed a strategy of modifying the oligonucleotide at the 2' end.
[0004] Click chemistry is a conjugation reaction with mild reaction conditions and high yields. Azide and alkyne click chemistry, in particular, is widely used in the pharmaceutical and biopharmaceutical fields due to its good biocompatibility and biorthogonality, and the formation of stable triazole bonds after the reaction. Oligonucleotides modified with 2'-azide or 2'-acetylene can be used to stably and efficiently synthesize various oligonucleotide conjugates and homologous conjugates, such as GalNAc derivatives or other ligand molecules through click chemistry. However, while oligonucleotides containing 2'-azide or 2'-acetylene modifications have strong ligand molecule binding stability, these additionally coupled compound chains containing azide or acetylene molecules can have unpredictable effects on the cellular internalization process of the oligonucleotide. In the case of low cellular internalization, the delivery efficiency of the oligonucleotide can be reduced, causing oligonucleotides containing 2'-azide or 2'-acetylene modifications to lose their original technical advantages. Therefore, at this stage, there is an urgent need to find a 2' azide or alkyne group-coupled oligonucleotide molecule that can achieve efficient cellular internalization, providing a basis for the targeted and efficient delivery of nucleic acid drugs. Summary of the Invention
[0005] In order to overcome the problem of cellular internalization of nucleotides modified with 2'-azide or alkyne groups in the prior art, the present invention provides a nucleotide monomer with a fatty chain conjugated at the ribose 2'-position. The ribose 2'-position of the nucleotide monomer contains a long-chain carbon chain conjugated through a triazole, and the functional group at the end of the long-chain carbon chain can be stably conjugated to a GalNAc derivative or other ligand molecules through click chemistry. The oligonucleotide chain containing this modified nucleotide monomer can not only stably conjugate other ligand molecules, but also achieve efficient cellular internalization, solving the problem of low cellular internalization efficiency in the delivery process of oligonucleotide drugs. At the same time, the present invention also provides a preparation method and use of the ribose 2'-position conjugated fatty chain nucleotide monomer and the oligonucleotide chain containing the above monomer.
[0006] On the one hand, the present invention provides a modified nucleotide monomer having the structure shown in formula (I),
[0007]
[0008] wherein R 1 and R 2 are independently selected from one or a combination of hydrogen, a sugar protecting group, a reactive group, a solid support or a compound containing a solid support, a phosphate ester, or an oligonucleotide, B is a base or a base analog, and R 3 is an alkynyl group, or a ligand containing a triazole group, or a fluorescent group containing a triazole group, and L 1 and / or L 2 is a linking structure.
[0009] In some embodiments, the nucleotide monomer further includes the structure shown in formula (I-C1):
[0010]
[0011] In some embodiments, B is selected from one or more of the following: guanine, adenine, cytosine, uracil, guanine analogs, adenine analogs, cytosine analogs, and uracil analogs.
[0012] In some embodiments, B is selected from one or more of the following groups: selected from xanthine, allylaminouridine, allylaminothymidine, hypoxanthine, dioxoadenine, dioxocytosine, dioxoguanine, dioxouracil, 6-chloropurine riboside, N6-methyladenine, methylpseudouracil, 2-thiocytosine, 2-thiouracil, 5-methyluracil, 4-thiothymidine, 4-thiouracil, 5,6-dihydro-5-methyluracil, 5,6-dihydrouracil, 5-[(3-indolyl)propionamide-N-allyl]uracil, 5-aminoallylcytosine, 5-aminoallyluracil, 5-bromouracil, 5-bromocytosine, 5-carboxycytosine, 5-carboxycytosine, 5-carboxyuracil, 5-fluorouracil, 5-formylcytosine, 5-formyluracil, 5-hydroxycytosine, 5-hydroxymethylcytosine, 5-hydroxymethyluracil, 5-hydroxyuracil, 5-iodocytosine, 5-iodouracil, 5-methoxycytosine, 5-methoxyuracil, 5-methylcytosine, 5-methyluracil, 5-propynylaminocytosine, 5-propynylaminuracil, 5-propynylcytosine, 5-propynyluracil, 6-azacytosine, 6-azauracil, 6-chloropurine, 6-thioguanine, 7-deazaadenine, 7-deazaguanine, 7-deaza-7-propylaminoadenine, 7-deaza-7-propynylaminoadenine, 8-azaguanine, 8-azidoadenine, 8-chloroadenine, 8-oxoadenine, 8-oxoguanine, biotin-16-7-deaza-7-propynylaminoguanine, biotin-16-aminoallylcytosine, biotin-16-aminoallyluracil, 3-5-propynylaminocytosine, 3-6-propynylaminuracil, cyano 3-aminoallylcytosine, cyano 3-aminoallyluracil, cyano 5-6-propynylaminocytosine, cyano 5-6-propynylaminuracil, cyano 5-aminoallylcytosine, cyano 5-aminoallyluracil, cyano 7-aminoallyluracil, Dabcyl-5-3-aminoallyluracil, desthiobiotin-16-aminoallyluracil, desthiobiotin-6-aminoallylcytosine, isoguanine, N1-ethylpseudouracil, N1-methoxymethylpseudouracil, N1-methyladenine, N1-methylpseudouracil, N1-propylpseudouracil, N2-methylguanine, N4-biotin-OBEA-cytosine, N4-methylcytosine, N6-methyladenine, 06-methylguanine, pseudoisocytosine, pseudouracil, thiophenocytosine, thiophenoguanine, thiophenouracil, xanthine, 3-deazaadenine, 2,6-diaminoadenine, 2,6 - biguanine, 5 - formyluracil, 5 - ethynyluracil, N6 - isopentenyladenine (i6A), 2 - methylthio - N6 - isopentenyladenine (ms2i6A), 2 - methylthio - N6 - methyladenine (ms2m6A), N6 - (cis - hydroxyisopentenyl)adenine (io6A), 2 - methylthio - N6 - (cis - hydroxyisopentenyl)adenine (ms2io6A), N6 - glycinylcarbamoyladenine (g6A), N6 - threonylcarbamoyladenine (t6A), 2 - methylthio - N6 - threonylcarbamoyladenine (ms2t6A), N6 - methyl - N6 - threonylcarbamoyladenine (m6t6A), N6 - hydroxyvalerylcarbamoyladenine (hn6A), 2 - methylthio - N6 - hydroxyvalerylcarbamoyladenine (ms2hn6A), N6,N6 - dimethyladenine (m62A) and N6 - acetyladenine (ac6A).
[0013] In some embodiments, the base or its analogue further comprises a base protecting group, and the base protecting group is selected from one of the following groups: fluorenylmethyloxycarbonyl (Fmoc), tert - butyloxycarbonyl (BOC), benzyloxycarbonyl (Cbz), optionally substituted acyl, trifluoroacetyl (TFA), benzyl, trityl (Tr), 4,4’ - dimethoxytrityl (DMTr) and tosyl (Ts).
[0014] In some embodiments, the L 1 is selected from one of the following structures: optionally substituted -(CH2) n -, -((CH2) m O) n - and -((CH2) m S) n -, where n is an integer from 0 to 20, preferably 1 - 10, more preferably 1 - 5; m is an integer from 0 to 20, preferably 0 - 10, more preferably 0 - 5. The groups of the optionally substituted substituents are selected from the following group: halogen, cyano, hydroxy, nitro, amino, alkylamino, cycloalkylamino, heterocyclic group, aminocarbonyl, sulfonyl, aminosulfonyl, carbonylamino, sulfonylamino, methyl, ethyl, aryl, methoxy, ethoxy, trifluoromethyl, trifluoroethyl, trifluoromethoxy, trifluoroethoxy, polyoxyethylene, polyoxypropylene, C1 - C3 alkyl polyoxyethylene, or C1 - C3 alkyl polyoxypropylene.
[0015] In some embodiments, the L 2 is selected from one of the following structures: optionally substituted -(CH2) n -, -((CH2) m O) n -, and -((CH2)m S) n -, wherein m is an integer from 0 to 20, preferably from 1 to 15; n is an integer from 0 to 20, preferably from 0 to 10, more preferably from 0 to 5. The group of any substituted substituent is selected from the following group: halogen, cyano, hydroxyl, nitro, amino, alkylamino, cycloalkylamino, heterocyclic group, aminocarbonyl, sulfonyl, aminosulfonyl, carbonylamino, sulfonylamino, methyl, ethyl, aryl, methoxy, ethoxy, trifluoromethyl, trifluoroethyl, trifluoromethoxy, trifluoroethoxy, polyoxyethylene, polyoxypropylene, C1-C3 alkyl polyoxyethylene, or C1-C3 alkyl polyoxypropylene.
[0016] In some embodiments, the phosphate ester comprises one of the following groups:
[0017] -P(=R Y2 )(R Y3 R Y1 )R Y1 、-P(=R Y2 )(R Y3 R Y1 )2. Wherein each R Y1 is independently selected from hydrogen, oxygen, hydroxyl or C1-C6 alkyl optionally substituted by one or more halogens or cyano, R Y2 is selected from O or S, and R Y3 is selected from O or S.
[0018] In some embodiments, the sugar protecting group comprises one of the following groups: acetyl (Ac), benzoyl (Bz), benzyl (Bn), β-methoxyethoxymethyl ether (MEM), dimethoxytriphenylmethyl (DMT), methoxymethyl ether (MOM), methoxytriphenylmethyl (MMT), p-methoxybenzyl ether (PMB), p-methoxyphenyl ether (PMP), pivaloyl (Piv), tetrahydrofuran (THF), triphenylmethyl (Tr), triphenylmethylsilane, triphenylmethylsilyl ether, trimethylsilyl (TMS), triisopropylsilyloxymethyl (TOM), triisopropylsilyloxymethyl, 4,4'-dimethoxytriphenylmethyl (DMTr), tert-butyldimethylsilyl (TBDMS).
[0019] In some embodiments, the reactive group is a phosphoramidite group.
[0020] In some embodiments, the phosphoramidite group is a 2-cyanoethyl N,N-diisopropylphosphoramidite group.
[0021] In some embodiments, the R 3 is an alkynyl group.
[0022] In some embodiments, the R 3is a ligand containing a triazole group.
[0023] In some embodiments, the R 3 is as shown in Chemical Formula (II):
[0024]
[0025] wherein, L 3 is a linking structure and L is a ligand.
[0026] In some embodiments, the L 3 is selected from one of the following structures: optionally substituted -(CH2)n-, -(O(CH2)m)n-, -(S(CH2)m)n-. Wherein, m is an integer from 0 to 20, preferably 1 to 15; n is an integer from 0 to 20, preferably 0 to 10, more preferably 0 to 5. The groups of the optionally substituted substituents are selected from the following group: halogen, cyano, hydroxyl, nitro, amino, alkylamino, cycloalkylamino, heterocyclic group, aminocarbonyl, sulfonyl, aminosulfonyl, carbonylamino, sulfonylamino, methyl, ethyl, aryl, methoxy, ethoxy, trifluoromethyl, trifluoroethyl, trifluoromethoxy, trifluoroethoxy, polyoxyethylene, polyoxypropylene, C1-C3 alkyl polyoxyethylene, or C1-C3 alkyl polyoxypropylene.
[0027] In some embodiments, the ligand is selected from one of the following groups: galactose, galactosamine, N-acetylgalactosamine, mannose, glucose, glucosamine, N-acetylglucosamine, fucose or lactose, N-acetylgalactosamine with all hydroxyl groups fully protected by acyl group, galactose with all hydroxyl groups fully protected by acyl group, galactosamine with all hydroxyl groups fully protected by acyl group, N-formyl-galactosamine with all hydroxyl groups fully protected by acyl group, N-propionyl-galactosamine with all hydroxyl groups fully protected by acyl group, N-n-butyryl-galactosamine with all hydroxyl groups fully protected by acyl group or N-isobutyryl-galactosamine with all hydroxyl groups fully protected by acyl group, wherein the acyl group is acetyl or benzoyl.
[0028] In some embodiments, the ligand L is as shown in Chemical Formula (III):
[0029]
[0030] In some embodiments, the R 3 has the structure as shown in Chemical Formula (III-L):
[0031]
[0032] In some embodiments, when R 3 is an alkynyl group, the nucleotide monomer includes the structures of formulas (I-1), (I-2), (I-3), (I-4) and (I-5):
[0033]
[0034] In some embodiments, when R 3 has the structure shown in Chemical Formula (III-L), the nucleotide monomer has the structure of the following formula:
[0035]
[0036] wherein R 1 and R 2 are phosphate esters, or phosphate esters linked to an oligonucleotide, and Ligand represents a ligand or a fluorophore. The fluorophore is labeled with 6-carboxyfluorescein FAM.
[0037] On the other hand, the present invention provides a modified oligonucleotide having the structure shown below:
[0038]
[0039] wherein R 2 is H, a sugar protecting group, a reactive group, a solid support or a compound containing a solid support, a phosphate ester, or a phosphate ester linked to an oligonucleotide. B is a base or a base analogue, and L 1 and / or L 2 is a linking structure, oligo1 represents an oligonucleotide, and R 3 is an alkynyl group, or a ligand containing a triazole group, or a fluorophore containing a triazole group.
[0040] In some embodiments, the oligonucleotide has the structure shown below:
[0041]
[0042] wherein B is a base or a base analogue, and L 1 and / or L 2 is a linking structure, oligo1 and oligo2 represent oligonucleotide fragments, and R 3 is an alkynyl group, or a ligand containing a triazole group, or a fluorophore containing a triazole group.
[0043] In some embodiments, the oligonucleotide has the structure shown below:
[0044]
[0045] wherein oligo1 and oligo2 represent oligonucleotide fragments.
[0046] In some embodiments, the sequence of the oligonucleotide fragment oligo1 is as shown in SEQ ID NO.: 1, and the sequence of the oligonucleotide fragment oligo2 is as shown in SEQ ID NO.: 2.
[0047] The present invention also provides methods for preparing and uses of the above nucleotide monomers.
[0048] Those skilled in the art can easily perceive other aspects and advantages of the present invention from the following detailed description. Only exemplary embodiments of the present invention are shown and described in the following detailed description. As those skilled in the art will recognize, the content of the present invention enables those skilled in the art to make changes to the disclosed specific embodiments without departing from the spirit and scope of the invention involved in the present invention. Accordingly, the descriptions in the drawings and the specification of the present invention are merely exemplary and not restrictive. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] The specific features of the invention involved in the present invention are shown in the appended claims. The features and advantages involved in the present invention can be better understood by referring to the exemplary embodiments and the drawings described in detail below. A brief description of the drawings is as follows:
[0050] Figure 1 Shows the 1H NMR spectrum of Compound 7 in the present invention.
[0051] Figure 2 Shows the 1H NMR spectrum of Compound 8 in the present invention.
[0052] Figure 3 Shows the 1H NMR spectrum of Compound 10 in the present invention.
[0053] Figure 4 Shows the 1H NMR spectrum of Compound 11 in the present invention.
[0054] Figure 5 Shows the 1H NMR spectrum of Compound 12 in the present invention.
[0055] Figure 6 Shows the HPLC purification spectrum of T-C14 oligonucleotide DNA in the present invention.
[0056] Figure 7 Shows the mass spectrum of T-C14 oligonucleotide DNA in the present invention.
[0057] Figure 8 Shows the purification schematic diagram of DNA-1 conjugated with FAM molecule in the invention.
[0058] Figure 9Shown are fluorescence micrographs of different concentrations of DNA-1 after being taken up by cells in the present invention. Among them, the FAM group represents the distribution of DNA-1 labeled with 6-carboxyfluorescein FAM, the SiR-Heochst group represents the distribution of cells stained with SiR-Heochst, and the merged group represents the chimerism of the first two figures. Detailed implementation mode
[0059] The following specific embodiments illustrate the implementation mode of the invention of the present application. Those skilled in this technology can easily understand other advantages and effects of the invention of the present application from the content disclosed in this specification.
[0060] Term definition
[0061] As used herein, the term "sugar protecting group" refers to a group that is attached to a nucleic acid molecule, an oligonucleotide chain, or a nucleotide monomer and has a protective effect on the pentose sugar structure and the functional groups on the structure of the nucleic acid molecule, oligonucleotide chain, or nucleotide monomer. The protective effect means that the sugar protecting group can avoid side reactions unrelated to nucleic acid chain elongation of any functional group on the pentose sugar structure during nucleic acid chain elongation. For example, the protected functional groups on the pentose sugar can be hydroxyl groups at the 2', 3', and / or 5' positions, etc. The sugar protecting groups include, but are not limited to, one of the following groups: acetyl (Ac), benzoyl (Bz), benzyl (Bn), β-methoxyethoxymethyl ether (MEM), dimethoxytrityl (DMT), methoxymethyl ether (MOM), methoxytrityl (MMT), p-methoxybenzyl ether (PMB), p-methoxyphenyl ether (PMP), pivaloyl (Piv), tetrahydrofuran (THF), trityl (Tr), tritylsilane, tritylsilyl ether, trimethylsilyl (TMS), triisopropylsilyloxymethyl (TOM), triisopropylsilyloxymethyl, 4,4'-dimethoxytrityl (DMTr), or tert-butyldimethylsilyl (TBDMS). The sugar protecting groups used for different types and positions of protected functional groups vary. For example, the hydroxyl group at the 5' position is commonly protected with 4,4'-dimethoxytrityl (DMTr), and tert-butyldimethylsilyl (TBDMS) is the most commonly used protecting group for the hydroxyl group at the 2' position.
[0062] As used herein, the term "solid support" refers to a solid-phase support for immobilizing synthesized or unsynthesized nucleic acid chains or monomeric nucleotides during the synthesis of the desired nucleic acid chain. The extension of a nucleic acid chain using a solid support is referred to as solid-phase synthesis of the nucleic acid chain. During solid-phase synthesis, it is usually necessary to covalently link and immobilize the initial monomeric nucleotide, i.e., the first nucleotide on the nucleic acid chain, to the solid support before extending the nucleic acid chain to ensure the correct synthesis of the nucleic acid chain product. The solid support can be directly covalently linked to the initial monomeric nucleotide or indirectly covalently linked to the monomeric nucleotide through the attachment of any intermediate compound. Generally speaking, the solid support can be linked to the hydroxyl functional group of the initial nucleotide, such as the 3'-hydroxyl functional group. The solid support includes, but is not limited to, controlled pore glass beads (CPG) and polystyrene (PS). The pore size of the solid support is proportional to the loading capacity of the synthesized nucleic acid chain. The larger the pore size of the solid support, the higher the loading capacity of the nucleic acid chain that can be synthesized.
[0063] As used herein, the term "reaction group" refers to a reaction activation group covalently linked to a nucleic acid chain. Reaction activation means that the attachment of the reaction group to the nucleotide chain enables a chemical reaction to proceed with a higher reaction efficiency compared to when the reaction group is not attached. For example, the linkage between nucleotide molecules proceeds with a higher reaction efficiency. The chemical reaction can be a chemical condensation reaction. In some embodiments, the chemical reaction also requires the addition of an activator for catalysis. Generally speaking, during the extension of a nucleic acid chain, the nucleotide linked with the reaction group can undergo a nucleophilic reaction and condensation with the 5'-position hydroxyl group of other nucleotide monomers through the reaction group to achieve the coupling of two nucleotide molecules. After the coupling reaction between the reaction groups of nucleotide molecules, any molecular structure can be removed or any molecular rearrangement can occur. The "reaction group" as used herein also includes a group whose molecular structure has changed compared to the original reaction group after the coupling ends. The reaction group can be any nucleophilic reaction group, including but not limited to ammonium phosphite groups and any group containing ammonium phosphite, such as 2-cyanoethyl N,N-diisopropyl phosphoramidite groups.
[0064] As used herein, the term "phosphate ester" refers to an ester derivative compound of phosphoric acid, generally referring to a derivative compound formed by the condensation of phosphoric acid, a salt containing phosphoric acid, and a compound containing a phosphate group with other hydroxyl-containing compounds through an esterification reaction. The phosphate ester described in the present invention can be a orthophosphate ester or a phosphite ester. The phosphate ester can contain one, two or three ester chemical bonds. For example, the phosphate ester can be a compound containing two ester chemical bonds formed by the esterification reaction of any molecule containing a phosphate group with the hydroxyl groups of two other compounds. In some embodiments, the phosphate ester is formed by the simultaneous or sequential esterification reaction of the hydroxyl groups of two nucleotide monomers with any molecule containing a phosphate group, and this reaction couples the two nucleotide molecules through a phosphodiester bond. One or more oxygen atoms of the phosphate ester can be replaced by one or more sulfur atoms. For example, any oxygen atom on the phosphodiester bond can be replaced by a sulfur atom.
[0065] As used herein, the term "oligonucleotide" refers to a nucleic acid chain synthesized by linking natural and / or modified nucleotides through natural and / or non-natural chemical bonds. The oligonucleotide is synthesized by the tandem connection of multiple monomeric nucleotides and generally has a length of 200 monomeric nucleotides or less. The length of the oligonucleotide can be 10 - 50 monomeric nucleotides, 50 - 100 monomeric nucleotides, 100 - 150 monomeric nucleotides, or 150 - 200 monomeric nucleotides. The oligonucleotide can be a single-stranded or double-stranded oligonucleotide, and can be a sense or antisense oligonucleotide.
[0066] As used herein, the term "base or base analogue" includes all known purine heterocycles and purine heterocycle analogues and variants, pyrimidine heterocycles and pyrimidine heterocycle analogues and variants. The base or base analogue may be naturally occurring or synthetic, such as a molecule obtained by modifying a natural base or base analogue. The base or base analogue is selected from one or more of the following groups: guanine, adenine, cytosine, uracil, guanine analogue, adenine analogue, cytosine analogue, uracil analogue, xanthine, allylaminouridine, allylaminothymidine, hypoxanthine, dioxoadenine, dioxocytosine, dioxoguanine, dioxouracil, 6-chloropurine riboside, N6-methyladenine, methylpseudouracil, 2-thiocytosine, 2-thiouracil, 5-methyluracil, 4-thiothymidine, 4-thiouracil, 5,6-Dihydro-5-methyluracil, 5,6-dihydrouracil, 5-[(3-indolyl)propionamide-N-allyl]uracil, 5-aminoallylcytosine, 5-aminoallyluracil, 5-bromouracil, 5-bromocytosine, 5-carboxycytosine, 5-carboxycytosine, 5-carboxyuracil, 5-fluorouracil, 5-formylcytosine, 5-formyluracil, 5-hydroxycytosine, 5-hydroxymethylcytosine, 5-hydroxymethyluracil, 5-hydroxyuracil, 5-iodocytosine, 5-iodouracil, 5-methoxycytosine, 5-methoxyuracil, 5-methylcytosine, 5-methyluracil, 5-propynylaminocytosine, 5-propynylaminuracil, 5-propynylcytosine, 5-propynyluracil, 6-azacytosine, 6-azauracil, 6-chloropurine, 6-thioguanine, 7-deazaadenine, 7-deazaguanine, 7-deaza-7-propylaminoadenine, 7-deaza-7-propynylaminoadenine, 8-azaaadenine, 8-azidoadenine, 8-chloroadenine, 8-oxoadenine, 8-oxoguanine, biotin-16-7-deaza-7-propynylaminoguanine, biotin-16-aminoallylcytosine, biotin-16-aminoallyluracil, 3-5-propynylaminocytosine, 3-6-propynylaminuracil, cyano 3-aminoallylcytosine, cyano 3-aminoallyluracil, cyano 5-6-propynylaminocytosine, cyano 5-6-propynylaminuracil, cyano 5-aminoallylcytosine, cyano 5-aminoallyluracil, cyano 7-aminoallyluracil, Dabcyl-5-3-aminoallyluracil, desthiobiotin-16-aminoallyluracil, desthiobiotin-6-aminoallylcytosine, isoguanine, N1-ethylpseudouracil, N1-methoxymethylpseudouracil, N1-methyladenine, N1-methylpseudouracil, N1-propylpseudouracil, N2-methylguanine, N4-biotin-OBEA-cytosine, N4-methylcytosine, N6-methyladenine, 06-methylguanine, pseudoisocytosine, pseudouracil, thiophenocytosine, thiophenoguanine, thiophenouracil, xanthine, 3-deazaadenine, 2,6-diaminoadenine, 2,6 - large aminoguanine, 5 - formyluracil, 5 - ethynyluracil, N6 - isopentenyladenine (i6A), 2 - methylthio - N6 - isopentenyladenine (ms2i6A), 2 - methylthio - N6 - methyladenine (ms2m6A), N6 - (cis - hydroxyisopentenyl)adenine (io6A), 2 - methylthio - N6 - (cis - hydroxyisopentenyl)adenine (ms2io6A), N6 - glycinylcarbamoyladenine (g6A), N6 - threonylcarbamoyladenine (t6A), 2 - methylthio - N6 - threonylcarbamoyladenine (ms2t6A), N6 - methyl - N6 - threonylcarbamoyladenine (m6t6A), N6 - hydroxyvalerylcarbamoyladenine (hn6A), 2 - methylthio - N6 - hydroxyvalerylcarbamoyladenine (ms2hn6A), N6,N6 - dimethyladenine (m62A) and N6 - acetyladenine (ac6A). The base or base analog may further include a base - protecting group or may not contain a base - protecting group. The base - protecting group refers to a group that is attached to a nucleic acid molecule, an oligonucleotide chain or a nucleotide monomer and has a protective effect on the base on the nucleic acid molecule, oligonucleotide chain or nucleotide monomer. The base - protecting group can avoid unnecessary additional reactions of any functional group on the base during the extension of the nucleic acid chain, such as the amino group on the base. The base - protecting group includes, but is not limited to, one of the following groups: fluorenylmethoxycarbonyl (Fmoc), tert - butyloxycarbonyl (BOC), benzyloxycarbonyl (Cbz), optionally substituted acyl, trifluoroacetyl (TFA), benzyl, trityl (Tr), 4,4’ - dimethoxytrityl (DMTr) and tosyl (Ts). The base - protecting groups used for different types of bases or base analogs are different. For example, the amino protection of guanine and its variants usually uses acetyl (Ac), phenoxyacetyl (Pac), 4 - isopropylphenoxyacetyl (iPrPac), etc.
[0067] As used herein, the term "triazole group" refers to a compound group containing a triazole structure, wherein the triazole structure refers to a five-membered double-bond heterocycle containing three nitrogen atoms and two carbon atoms. There are three nitrogen atoms in the triazole structure. According to the relative positions of the nitrogen atoms, the triazole structure can be divided into 1,2,3-triazole and 1,2,4-triazole, that is, the structure in which three nitrogen atoms are adjacent to each other and the structure in which two nitrogen atoms are adjacent to each other. One, two or three substituents may be present on the triazole structure, and the substituents may be attached to the nitrogen atoms or to the carbon atoms. For example, the 1,2,3-triazole structure may have a substituent attached to the nitrogen atom at the 1-position and the carbon atom at the 4-position, and a hydrogen atom attached to the carbon atom at the 5-position; or 1,2,4-triazole may have a substituent attached to the nitrogen atom at the 1-position and the carbon atom at the 3-position, and a hydrogen atom attached to the carbon atom at the 5-position.
[0068] As used herein, the term "ligand" refers to a chemical group or molecule that, after coupling to the desired molecule, imparts an additional biological function to the desired molecule. For example, the ligand can bind to a cell surface receptor, such as a mammalian cell surface receptor. After the ligand is linked to the desired molecule, the desired molecule can have the function of binding to the mammalian cell surface receptor. The ligand can be a carbohydrate or a carbohydrate derivative, such as a monosaccharide, disaccharide, trisaccharide or polysaccharide, or a modified monosaccharide, disaccharide, trisaccharide or polysaccharide. The ligand can be independently selected from the group consisting of the following sugars: glucose and its derivatives, mannan and its derivatives, galactose and its derivatives, xylose and its derivatives, ribose and its derivatives, fucose and its derivatives, lactose and maltose and their derivatives, arabinose and its derivatives, fructose and its derivatives or sialic acid. The ligand can be selected from one of the following groups: galactose, galactosamine, N-acetylgalactosamine, mannose, glucose, glucosamine, N-acetylglucosamine, fucose or lactose, N-acetylgalactosamine with all hydroxyl groups fully protected by acyl groups, galactose with all hydroxyl groups fully protected by acyl groups, galactosamine with all hydroxyl groups fully protected by acyl groups, N-formyl-galactosamine with all hydroxyl groups fully protected by acyl groups, N-propionyl-galactosamine with all hydroxyl groups fully protected by acyl groups, N-n-butyryl-galactosamine with all hydroxyl groups fully protected by acyl groups or N-isobutyryl-galactosamine with all hydroxyl groups fully protected by acyl groups, wherein the acyl group can be an acetyl group or a benzoyl group.
[0069] As used herein, the term "fluorescent group" refers to a chemical group containing a fluorescent label, where the fluorescent label refers to a moiety that, upon excitation by light at a specific wavelength, such as light at a wavelength of 520 nm, can absorb the light of that wavelength and then re-emit visible light of a color. The fluorescent group is typically covalently bound to other molecules and serves as an indicator for detecting the presence of other molecules in an environment. Common reactive groups include, but are not limited to, amine-reactive isothiocyanate derivatives such as FITC and TRITC (derivatives of fluorescein and rhodamine), amine-reactive succinimidyl esters such as NHS-fluorescein, and thiol-reactive maleimide-activated fluorescein (fluor) such as fluorescein-5-maleimide. In some embodiments, the fluorescent group is covalently linked to a nucleotide monomer or an oligonucleotide chain for detecting the presence of the nucleotide monomer or oligonucleotide chain in an intracellular environment. The fluorescent group can be indirectly linked to the nucleotide monomer or oligonucleotide chain through other nucleotide-modifying molecules or directly linked to the pentose sugar, base, and / or phosphate group of the nucleotide molecule. In some embodiments, the fluorescent group can be a 6-carboxyfluorescein FAM label.
[0070] As used herein, the term "optionally substituted" means that the group and / or atom being "optionally substituted" has one or more additional substituents or no additional substituents. The position and number of the additional substituents on the substituted group are also limited by the valence of each substituted group or atom. For example, a carbon atom can have at most four additional substituents, and a nitrogen atom can have at most three substituents. The additional substituents can include, but are not limited to, one or more of the following groups: halogen, cyano, hydroxy, nitro, amino, alkylamino, cycloalkylamino, heterocyclic group, aminocarbonyl, sulfonyl, aminosulfonyl, carbonylamino, sulfonylamino, methyl, ethyl, aryl, methoxy, ethoxy, trifluoromethyl, trifluoroethyl, trifluoromethoxy, trifluoroethoxy, polyoxyethylene, polyoxypropylene, C1-C3 alkyl polyoxyethylene, and C1-C3 alkyl polyoxypropylene. The position and number of the additional substituents on the substituted group are also limited by the steric hindrance of each additional substituent. Multiple sterically bulky additional substituents, such as tert-butyl or isopropyl, cannot be present simultaneously on the same substituted group or atom. Detailed Description of the Invention
[0072] 1. Modified Nucleotide Monomers
[0073] The present invention provides a nucleotide monomer with a fatty chain coupled to the 2'-position of ribose, and the fatty chain coupled to the 2'-position of ribose can contain an azide group or an alkyne group. The nucleotide monomer with a fatty chain coupled to the 2'-position of ribose provided by the present invention can stably conjugate other ligand molecules at the end of the fatty chain.
[0074] On the one hand, the present invention provides a modified nucleotide monomer having a structure represented by formula (I), wherein the 2'-position, 3'-position, and 4'-position of the nucleotide monomer each have the same or different substituents.
[0075]
[0076] Wherein R 1 and R 2 independently can each be selected from one or a combination of hydrogen, a sugar protecting group, a reactive group, a solid support or a compound containing a solid support, a phosphate ester, or an oligonucleotide, B can be a base or a base analogue, and R 3 can be an alkynyl group, or a ligand containing a triazole group, or a fluorescent group containing a triazole group, and L 1 and / or L 2 can be a linking structure. In some embodiments, the fluorescent group is 6-carboxyfluorescein FAM.
[0077] The nucleotide monomer can further include a structure represented by formula (I-C1):
[0078]
[0079]
[0080] 1.1 B group
[0081] The ribose structure of the nucleotide monomer provided by the present invention can include a B group, and B can be a base or a base analogue.
[0082] The base can be guanine, adenine, cytosine, thymine, and / or uracil.
[0083] The base analogue can be a guanine analogue, an adenine analogue, a cytosine analogue, a thymine analogue, and / or a uracil analogue.
[0084] The base analogue can be a modified base. The modified base can be adding and / or removing a certain chemical group to / from the original base structure. For example, base modification can include, but is not limited to, methylation, acetylation, phosphorylation, and sulfation. The modified base can also be a molecular rearrangement of the original base structure.
[0085] The modified base can be a modified base synthesized intracellularly. For example, mRNA synthesized intracellularly can carry specific modified bases, thereby rendering the mRNA in a non-translatable state. The modified base can be a modified base introduced artificially. For example, a nucleotide modified base can be present on an exogenously introduced nucleic acid drug, thereby enabling the nucleic acid drug to have better stability and targeting properties.
[0086] The modified base can impart functions or biological activities different from those before modification to the nucleotide monomer and / or oligonucleotide chain containing the modified base. The modified base on the nucleotide chain can regulate gene expression by altering the structure and function of DNA or RNA. For example, methylation of DNA nucleobases can inhibit gene expression, while the 5'-end cap structure and 3'-end tail structure of RNA can affect mRNA splicing and stability. The modified base of nucleotides on mRNA can also affect mRNA stability, subcellular localization, and the translation efficiency of the protein encoded by mRNA.
[0087] The base analogs can be selected from one or more of the following groups: xanthine, allylaminopyrimidine, allylaminothymidine, hypoxanthine, dioxoadenine, dioxocytosine, dioxoguanine, dioxouracil, 6-chloropurine riboside, N6-methyladenine, methylpseudouracil, 2-thiocytosine, 2-thiouracil, 5-methyluracil, 4-thiothymidine, 4-thiouracil, 5,6-dihydro-5-methyluracil, 5,6-dihydrouracil, 5-[(3-indolyl)propionamide-N-allyl]uracil, 5-aminoallylcytosine, 5-aminoallyluracil, 5-bromouracil, 5-bromocytosine, 5-carboxycytosine, 5-carboxycytosine, 5-carboxyuracil, 5-fluorouracil, 5-formylcytosine, 5-formyluracil, 5-hydroxycytosine, 5-hydroxymethylcytosine, 5-hydroxymethyluracil, 5-hydroxyuracil, 5-iodocytosine, 5-iodouracil, 5-methoxycytosine, 5-methoxyuracil, 5-methylcytosine, 5-methyluracil, 5-propynylaminocytosine, 5-propynylaminuracil, 5-propynylcytosine, 5-propynyluracil, 6-azacytosine, 6-azauracil, 6-chloropurine, 6-thioguanine, 7-deazaadenine, 7-deazaguanine, 7-deaza-7-propylaminoadenine, 7-deaza-7-propynylaminoadenine, 8-azaguanine, 8-azidoadenine, 8-chloroadenine, 8-oxoadenine, 8-oxoguanine, biotin-16-7-deaza-7-propynylaminoguanine, biotin-16-aminoallylcytosine, biotin-16-aminoallyluracil, 3-5-propynylaminocytosine, 3-6-propynylaminuracil, cyano 3-aminoallylcytosine, cyano 3-aminoallyluracil, cyano 5-6-propynylaminocytosine, cyano 5-6-propynylaminuracil, cyano 5-aminoallylcytosine, cyano 5-aminoallyluracil, cyano 7-aminoallyluracil, Dabcyl-5-3-aminoallyluracil, desthiobiotin-16-aminoallyluracil, desthiobiotin-6-aminoallylcytosine, isoguanine, N1-ethylpseudouracil, N1-methoxymethylpseudouracil, N1-methyladenine, N1-methylpseudouracil, N1-propylpseudouracil, N2-methylguanine, N4-biotin-OBEA-cytosine, N4-methylcytosine, N6-methyladenine, 06-methylguanine, pseudoisocytosine, pseudouracil, thiophenocytosine, thiophenoguanine, thiophenouracil, xanthine, 3-deazaadenine, 2,6-diaminoadenine, 2,6 - large aminoguanine, 5 - formyluracil, 5 - ethynyluracil, N6 - isopentenyladenine (i6A), 2 - methylthio - N6 - isopentenyladenine (ms2i6A), 2 - methylthio - N6 - methyladenine (ms2m6A), N6 - (cis - hydroxyisopentenyl)adenine (io6A), 2 - methylthio - N6 - (cis - hydroxyisopentenyl)adenine (ms2io6A), N6 - glycinylcarbamoyladenine (g6A), N6 - threonylcarbamoyladenine (t6A), 2 - methylthio - N6 - threonylcarbamoyladenine (ms2t6A), N6 - methyl - N6 - threonylcarbamoyladenine (m6t6A), N6 - hydroxyvalerylcarbamoyladenine (hn6A), 2 - methylthio - N6 - hydroxyvalerylcarbamoyladenine (ms2hn6A), N6,N6 - dimethyladenine (m62A) and N6 - acetyladenine (ac6A).
[0088] The base or base analog further may include a base - protecting group. A base - protecting group refers to a group that is attached to a nucleotide monomer and has a protective effect on the base of the nucleotide monomer. The base - protecting group can avoid unnecessary additional reactions of any functional group on the base of the nucleotide monomer during the nucleic acid chain elongation process, such as the amino group on the base. The base - protecting group can be selected from one of the following groups: fluorenylmethyloxycarbonyl (Fmoc), tert - butyloxycarbonyl (BOC), benzyloxycarbonyl (Cbz), optionally substituted acyl, trifluoroacetyl (TFA), benzyl, triphenylmethyl (Tr), 4,4’ - dimethoxytriphenylmethyl (DMTr) and tosyl (Ts).
[0089] 1.2L 1 and L 2
[0090] An L may exist on the 2’ - coupled aliphatic chain of the nucleotide monomer provided by the present invention. 1 and / or L 2 of the linking structure. The role of L 1 can be to link the ribose structure of the nucleotide and the modified triazole structure. The role of L 2 can be to link the modified triazole structure with R 3 connection.
[0091] L 1 The structure of L 1The structure can be an aliphatic chain containing 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 carbon atoms. Further, L 1 The structure can be an aliphatic chain containing 1, 2, 3, 4, or 5 carbon atoms. L 1 The structure can be a straight-chain aliphatic chain containing 0 - 20 carbon atoms, for example, it can be a straight-chain aliphatic chain containing 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 carbon atoms. To endow the modified nucleotide monomer with better cellular internalization ability and stability, L 1 The structure can be a straight-chain aliphatic chain containing 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 carbon atoms. Further, L 1 The structure can be a straight-chain aliphatic chain containing 1, 2, 3, 4, or 5 carbon atoms. L 1 The structure can be a straight-chain ether structure containing 0 - 20 carbon atoms, for example, it can be a straight-chain ether structure containing 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 carbon atoms. To endow the modified nucleotide monomer with better cellular internalization ability and stability, L 1 The structure can be a straight-chain ether structure containing 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 carbon atoms. Further, L 1 The structure can be a straight-chain ether structure containing 1, 2, 3, 4, or 5 carbon atoms.
[0092] L 2 The structure can be an aliphatic chain containing 0 - 20 carbon atoms, for example, it can be an aliphatic chain containing 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 carbon atoms. To endow the modified nucleotide monomer with better cellular internalization ability and stability, L 2 The structure can be an aliphatic chain containing 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 carbon atoms. Further, L 2 The structure can be an aliphatic chain containing 1, 2, 3, 4, or 5 carbon atoms. L 2 The structure can be a straight-chain aliphatic chain containing 0 - 20 carbon atoms, for example, it can be a straight-chain aliphatic chain containing 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 carbon atoms. To endow the modified nucleotide monomer with better cellular internalization ability and stability, L 2 The structure can be a straight-chain aliphatic chain containing 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 carbon atoms. Further, L 2The structure can be a straight-chain aliphatic chain containing 1, 2, 3, 4, or 5 carbon atoms. L 2 The structure can be a straight-chain ether structure containing 0 - 20 carbon atoms. For example, it can be a straight-chain ether structure containing 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 carbon atoms. To endow the modified nucleotide monomer with better cellular internalization ability and stability, L 2 The structure can be a straight-chain ether structure containing 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 carbon atoms. Further, L 2 The structure can be a straight-chain ether structure containing 1, 2, 3, 4, or 5 carbon atoms.
[0093] The said L 1 can be selected from one of the following structures: arbitrarily substituted -(CH2) n -, -((CH2) m O) n -, and -((CH2) m S) n -. Wherein, n can be an integer from 0 - 20, preferably 1 - 10, and more preferably can be 1 - 5; m can be an integer from 0 - 20, preferably 0 - 10, and more preferably can be 0 - 5. The said L 2 can be selected from one of the following structures: arbitrarily substituted -(CH2) n -, -((CH2) m O) n -, and -((CH2) m S) n -. Wherein m can be an integer from 0 - 20, preferably 1 - 15; n can be an integer from 0 - 20, preferably 0 - 10, and more preferably can be 0 - 5.
[0094] The said L 1 can be arbitrarily substituted -(CH2) n -, where n can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20. To endow the modified nucleotide monomer with better cellular internalization ability and stability, where n can be 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10. Further, where n can be 1, 2, 3, 4, or 5.
[0095] The said L 1 can be -(CH2)-, -(CH2)2-, -(CH2)3-, -(CH2)4-, -(CH2)5-, -(CH2)6-, -(CH2)7-, -(CH2)8-, -(CH2)9-, or -(CH2)10 -. For example, the L 1 can be -(CH2)-.
[0096] The L 1 can be any substituted -((CH2) m O) n -, where n can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20, and where m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20. In order for the modified nucleotide monomer to have better cell internalization ability and stability, where n can be 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10, and where m can be 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10. Further, where n can be 1, 2, 3, 4 or 5, and where m can be 1, 2, 3, 4 or 5.
[0097] The L 1 can be any substituted -((CH2) m S) n -, where n can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20, and where m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20. In order for the modified nucleotide monomer to have better cell internalization ability and stability, where n can be 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10, and where m can be 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10. Further, where n can be 1, 2, 3, 4 or 5, and where m can be 1, 2, 3, 4 or 5.
[0098] The L 1structure, where m and n can have the following quantitative relationships: when n is 1, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20; when n is 2, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20; when n is 3, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20; when n is 4, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20; when n is 5, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20; when n is 6, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20; when n is 7, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20; when n is 8, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20; when n is 9, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20; when n is 10, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20; when n is 11, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20; when n is 12, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20; when n is 13, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20; when n is 14, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20; when n is 15, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20; when n is 16, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20;When n is 17, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20; when n is 18, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20; when n is 19, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20; when n is 20, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20.
[0099] The said L 2 can be any substituted -(CH2) n -, where n can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20. In order to make the modified nucleotide monomer have better cell internalization ability and stability, where n can be 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10. Further, where n can be 1, 2, 3, 4 or 5.
[0100] The said L 2 can be -(CH2)-, -(CH2)2-, -(CH2)3-, -(CH2)4-, -(CH2)5-, -(CH2)6-, -(CH2)7-, -(CH2)8-, -(CH2)9-, -(CH2) 10 -, -(CH2) 11 -, -(CH2) 12 -, -(CH2) 13 -, -(CH2) 14 -, -(CH2) 15 -, -(CH2) 16 -, -(CH2) 17 -, -(CH2) 18 -, -(CH2) 19 -, or -(CH2) 20 -. For example, the said L 2 can be -(CH2) 12 -.
[0101] The said L 2 can be any substituted -((CH2) m O) n-, where n can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20, and where m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20. To enable the modified nucleotide monomer to have better cell internalization ability and stability, where n can be 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10, and where m can be 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10. Further, where n can be 1, 2, 3, 4 or 5, and where m can be 1, 2, 3, 4 or 5.
[0102] The L 2 can be any substituted -((CH2) m S) n -, where n can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20, and where m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20. To enable the modified nucleotide monomer to have better cell internalization ability and stability, where n can be 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10, and where m can be 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10. Further, where n can be 1, 2, 3, 4 or 5, and where m can be 1, 2, 3, 4 or 5.
[0103] The L 2The structure, where m and n can have the following quantitative relationships: when n is 1, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20; when n is 2, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20; when n is 3, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20; when n is 4, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20; when n is 5, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20; when n is 6, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20; when n is 7, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20; when n is 8, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20; when n is 9, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20; when n is 10, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20; when n is 11, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20; when n is 12, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20; when n is 13, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20; when n is 14, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20; when n is 15, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20; when n is 16, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20;When n is 17, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20; when n is 18, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20; when n is 19, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20; when n is 20, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20.
[0104] The groups of any substituted substituents can be selected from the following group: halogen, cyano, hydroxyl, nitro, amino, alkylamino, cycloalkylamino, heterocyclic group, aminocarbonyl, sulfonyl, aminosulfonyl, carbonylamino, sulfonylamino, methyl, ethyl, aryl, methoxy, ethoxy, trifluoromethyl, trifluoroethyl, trifluoromethoxy, trifluoroethoxy, polyoxyethylene, polyoxypropylene, C1-C3 alkyl polyoxyethylene, or C1-C3 alkyl polyoxypropylene.
[0105] The position and number of the said any substituted substituent groups on the substituted group are also restricted by the valence states of each substituted group or atom. For example, a carbon atom can have at most four substituent groups. The position and number of the said any substituted substituent groups on the substituted group are also restricted by the steric hindrance of each substituent group. Multiple sterically hindered substituent groups such as tert-butyl or isopropyl cannot exist simultaneously on the same substituted group or atom.
[0106] 1.3 Phosphate
[0107] The nucleotide monomer provided by the present invention can have a phosphate ester at the 5' or 3' position of the ribose structure. The phosphate ester can be an ester derivative compound of phosphoric acid, and generally can refer to a derivative compound formed by the condensation of phosphoric acid, salts containing phosphoric acid, and compounds containing a phosphate group with other hydroxyl-containing compounds after an esterification reaction.
[0108] The said phosphate ester can be a modified phosphate ester. The modified phosphate ester can enhance the activity of the nucleotide containing the modified phosphate ester and increase more biological functions. The modified phosphate ester can be one obtained by adding and / or removing a certain chemical group to / from the original phosphate ester structure, or by changing one or more atoms in the original phosphate ester structure. Common phosphate ester modifications can include phosphorothioate (PS), dithiophosphate (PS2), methylphosphonate (MP), methoxypropylphosphonate (MOP), and peptide nucleic acid (PNA), etc.
[0109] The modified phosphate ester can replace the oxygen atom in the phosphate ester bond with a sulfur atom, thereby replacing the phosphate ester bond with a phosphorothioate bond. For example, the non-bridging oxygen atom in the phosphodiester bond can be replaced with a sulfur atom, thereby replacing the phosphodiester bond with a phosphorothioate diester bond. Introducing the modified phosphate ester into the oligonucleotide chain can stabilize the structure of the oligonucleotide chain and maintain high specificity and high affinity of base pairing.
[0110] The phosphate ester can include one of the following groups: -P(=R Y2 )(R Y3 R Y1 )R Y1 or -P(=R Y2 )(R Y3 R Y1 )2. Wherein each R Y1 independently can be selected from hydrogen, oxygen, hydroxyl, or C1-C6 alkyl optionally substituted by one or more halogens or cyano groups, R Y2 can be selected from O or S, and R Y3 can be selected from O or S. For example, the phosphate ester can include one of the following groups:
[0111] -P(=O)(OH)2, -P(=S)(OH)2, -P(=O)(SH)2, -P(=S)(OH)(SH), -P(=O)(OH)(SH), or -P(=S)(SH)2.
[0112] 1.4 Sugar protecting groups
[0113] The nucleotide monomer provided by the present invention can have a sugar protecting group at the 5' or 3' position of the ribose structure.
[0114] The sugar protecting group can refer to a group that is attached to a nucleic acid molecule and has a protective effect on the pentose sugar structure and the functional groups on the structure of the nucleic acid molecule, such as a hydroxyl protecting group. The protecting group can make the chemical functional group insensitive to specific reaction conditions, and can also be attached to and removed from the functional group in the nucleic acid molecule without affecting the rest of the nucleic acid molecule. For example, the sugar protecting group can avoid side reactions unrelated to nucleic acid chain extension of any functional group on the pentose sugar structure, such as a hydroxyl group, during the nucleic acid chain extension process. For example, the protected functional groups on the pentose sugar can be hydroxyl groups at the 2', 3', and / or 5' positions, etc. The protecting group can maintain stable connection under alkaline conditions and can be removed under acidic conditions.
[0115] The sugar protecting group may include one of the following groups: acetyl (Ac), benzoyl (Bz), benzyl (Bn), β - methoxyethoxymethyl ether (MEM), dimethoxytrityl (DMT), methoxymethyl ether (MOM), methoxytrityl (MMT), p - methoxybenzyl ether (PMB), p - methoxyphenyl ether (PMP), pivaloyl (Piv), tetrahydrofuran (THF), trityl (Tr), tritylsilane, tritylsilyl ether, trimethylsilyl (TMS), triisopropylsilyloxymethyl (TOM), triisopropylsilyloxymethyl, 4,4'-dimethoxytrityl (DMTr), tert - butyldimethylsilyl (TBDMS).
[0116] The nucleotide monomer provided by the present invention may have a sugar protecting group at the 5' position of the ribose structure. The sugar protecting group may be 4,4'-dimethoxytrityl (DMTr). The sugar protecting group may be acetyl (Ac). The sugar protecting group may be benzoyl (Bz). The sugar protecting group may be benzyl (Bn). The sugar protecting group may be β - methoxyethoxymethyl ether (MEM). The sugar protecting group may be dimethoxytrityl (DMT). The sugar protecting group may be methoxymethyl ether (MOM). The sugar protecting group may be methoxytrityl (MMT). The sugar protecting group may be p - methoxybenzyl ether (PMB). The sugar protecting group may be p - methoxyphenyl ether (PMP). The sugar protecting group may be pivaloyl (Piv). The sugar protecting group may be tetrahydrofuran (THF). The sugar protecting group may be trityl (Tr). The sugar protecting group may be tritylsilane. The sugar protecting group may be tritylsilyl ether. The sugar protecting group may be trimethylsilyl (TMS). The sugar protecting group may be triisopropylsilyloxymethyl (TOM). The sugar protecting group may be triisopropylsilyloxymethyl. The sugar protecting group may be tert - butyldimethylsilyl (TBDMS).
[0117] The nucleotide monomers provided by the present invention may have a sugar protecting group at the 3'-position of the ribose structure. The sugar protecting group may be 4,4'-dimethoxytrityl (DMTr). The sugar protecting group may be acetyl (Ac). The sugar protecting group may be benzoyl (Bz). The sugar protecting group may be benzyl (Bn). The sugar protecting group may be β-methoxyethoxymethyl ether (MEM). The sugar protecting group may be dimethoxytrityl (DMT). The sugar protecting group may be methoxymethyl ether (MOM). The sugar protecting group may be methoxytrityl (MMT). The sugar protecting group may be p-methoxybenzyl ether (PMB). The sugar protecting group may be p-methoxyphenyl ether (PMP). The sugar protecting group may be pivaloyl (Piv). The sugar protecting group may be tetrahydrofuran (THF). The sugar protecting group may be trityl (Tr). The sugar protecting group may be tritylsilane. The sugar protecting group may be tritylsilyl ether. The sugar protecting group may be trimethylsilyl (TMS). The sugar protecting group may be triisopropylsilyloxymethyl (TOM). The sugar protecting group may be triisopropylsilyloxymethyl. The sugar protecting group may be tert-butyldimethylsilyl (TBDMS).
[0118] 1.5 Reaction group
[0119] The nucleotide monomers provided by the present invention may have a reaction group at the 5'- or 3'-position of the ribose structure.
[0120] The reaction group may be a phosphoramidite group. Solid-phase phosphoramidite synthesis is a common method for synthesizing oligonucleotide chains in the art. The phosphoramidite reaction group in the present invention also refers to nucleoside phosphoramidite, that is, the phosphoramidite coupled with a nucleotide. The phosphoramidite functional group can undergo a coupling reaction with a hydroxyl group on a solid-phase support, such as a resin, and be oxidized to form a solid-phase support connected by a phosphodiester bond. During the phosphoramidite solid-phase synthesis process, the hydroxyl group on the nucleotide is deprotected and then can undergo a coupling reaction with the phosphoramidite group on another nucleotide under coupling reaction conditions.
[0121] The nucleotide monomers provided by the present invention may have a reaction group at the 5'-position of the ribose structure, and the reaction group may be a 2-cyanoethyl N,N-diisopropylphosphoramidite group. The nucleotide monomers provided by the present invention may have a reaction group at the 3'-position of the ribose structure, and the reaction group may be a 2-cyanoethyl N,N-diisopropylphosphoramidite group.
[0122] 1.6R 3
[0123] The end of the fatty chain coupled at the 2'-position of the nucleotide monomers provided by the present invention may be R 3 .R3 The function can be a linking structure for connecting other molecules, such as containing functional groups that can undergo coupling reactions.
[0124] The said R 3 can be a structure containing an alkynyl group.
[0125] The said R 3 can be a structure containing a triazole group.
[0126] The said R 3 can be a fluorescent group containing a triazole group. The fluorescent group can be labeled with 6-carboxyfluorescein FAM.
[0127] The said R 3 can be a ligand containing a triazole group.
[0128] The said R 3 As shown in chemical formula (II):
[0129]
[0130] Wherein, L 3 can be a linking structure or a bond. L can be a ligand or a fluorescent group.
[0131] 1.6.1L 3
[0132] L 3 The function of L can link the triazole structure of R 3 with the ligand molecule.
[0133] The said L 3 can be selected from one of the following structures: optionally substituted -(CH2) n -, -(O(CH2) m ) n -, -(S(CH2) m ) ) n -. Wherein, n can be an integer from 0 to 20, preferably 1 to 15; m can be an integer from 0 to 20, preferably 0 to 10, more preferably can be 0 to 5. The optionally substituted substituent groups can be selected from the following group: halogen, cyano, hydroxyl, nitro, amino, alkylamino, cycloalkylamino, heterocyclic group, aminocarbonyl, sulfonyl, aminosulfonyl, carbonylamino, sulfonylamino, methyl, ethyl, aryl, methoxy, ethoxy, trifluoromethyl, trifluoroethyl, trifluoromethoxy, trifluoroethoxy, polyoxyethylene, polyoxypropylene, C1-C3 alkyl polyoxyethylene, or C1-C3 alkyl polyoxypropylene.
[0134] L 3The structure can be an aliphatic chain containing 0 - 20 carbon atoms. For example, it can be an aliphatic chain containing 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 carbon atoms. To enable the modified nucleotide monomer to have better cell internalization ability and stability, L 3 The structure can be an aliphatic chain containing 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 carbon atoms. Further, L 3 The structure can be an aliphatic chain containing 1, 2, 3, 4, or 5 carbon atoms. L 3 The structure can be a straight-chain aliphatic chain containing 0 - 20 carbon atoms. For example, it can be a straight-chain aliphatic chain containing 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 carbon atoms. To enable the modified nucleotide monomer to have better cell internalization ability and stability, L 3 The structure can be a straight-chain aliphatic chain containing 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 carbon atoms. Further, L 3 The structure can be a straight-chain aliphatic chain containing 1, 2, 3, 4, or 5 carbon atoms. L 3 The structure can be a straight-chain ether structure containing 0 - 20 carbon atoms. For example, it can be a straight-chain ether structure containing 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 carbon atoms. To enable the modified nucleotide monomer to have better cell internalization ability and stability, L 3 The structure can be a straight-chain ether structure containing 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 carbon atoms. Further, L 3 The structure can be a straight-chain ether structure containing 1, 2, 3, 4, or 5 carbon atoms.
[0135] The said L 3 can be any substituted -(CH2) n -, where n can be 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 1, 14, 15, 16, 17, 18, 19, or 20. To enable the modified nucleotide monomer to have better cell internalization ability and stability, where n can be 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10. Further, where n can be 0, 1, 2, 3, 4, or 5. For example, n can be 0.
[0136] The said L 3 can be any substituted -(O(CH2) m ) n-, where n can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20, and where m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20. To enable the modified nucleotide monomer to have better cell internalization ability and stability, where n can be 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10, and where m can be 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10. Further, where n can be 1, 2, 3, 4 or 5, and where m can be 1, 2, 3, 4 or 5.
[0137] The L 3 can be any substituted -(S(CH2) m ) n -, where n can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20, and where m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20. To enable the modified nucleotide monomer to have better cell internalization ability and stability, where n can be 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10, and where m can be 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10. Further, where n can be 1, 2, 3, 4 or 5, and where m can be 1, 2, 3, 4 or 5.
[0138] The L 3structure, where m and n can have the following quantitative relationships: when n is 1, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20; when n is 2, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20; when n is 3, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20; when n is 4, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20; when n is 5, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20; when n is 6, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20; when n is 7, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20; when n is 8, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20; when n is 9, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20; when n is 10, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20; when n is 11, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20; when n is 12, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20; when n is 13, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20; when n is 14, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20; when n is 15, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20; when n is 16, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20;When n is 17, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20; when n is 18, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20; when n is 19, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20; when n is 20, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20.
[0139] 1.6.2L
[0140] R 3 Furthermore, the L structure can be connected. The L structure can be a ligand or a fluorophore.
[0141] The ligand can be a ligand capable of binding to a cell surface receptor, such as a ligand capable of binding to a mammalian cell surface receptor, a ligand capable of binding to a liver surface receptor, or a ligand capable of binding to the asialoglycoprotein receptor (ASGPR) on the liver surface.
[0142] The ligand can be any ligand having an affinity for the asialoglycoprotein receptor (ASGPR) on the surface of mammalian hepatocytes. The ligand can be a carbohydrate or a carbohydrate derivative, which can be a monosaccharide, disaccharide, trisaccharide or polysaccharide. It can be a modified monosaccharide, disaccharide, trisaccharide or polysaccharide. The ligand can independently be selected from the group consisting of the following sugars: glucose and its derivatives, mannan and its derivatives, galactose and its derivatives, xylose and its derivatives, ribose and its derivatives, fucose and its derivatives, lactose and maltose and their derivatives, arabinose and its derivatives, fructose and its derivatives, or sialic acid.
[0143] The ligand can be selected from one of the following groups: galactose, galactosamine, N-acetylgalactosamine, mannose, glucose, glucosamine, N-acetylglucosamine, fucose or lactose, N-acetylgalactosamine with all hydroxyl groups fully protected by acyl groups, galactose with all hydroxyl groups fully protected by acyl groups, galactosamine with all hydroxyl groups fully protected by acyl groups, N-formyl-galactosamine with all hydroxyl groups fully protected by acyl groups, N-propionyl-galactosamine with all hydroxyl groups fully protected by acyl groups, N-n-butyryl-galactosamine with all hydroxyl groups fully protected by acyl groups, or N-isobutyryl-galactosamine with all hydroxyl groups fully protected by acyl groups, where the acyl group can be an acetyl group or a benzoyl group.
[0144] The ligand L structure can be as shown in chemical formula (III):
[0145]
[0146] Wherein the wavy line represents that the ligand structure can be connected to any atom on the triazole group, nucleotide monomer molecule or oligonucleotide chain, or can be connected to the triazole group, nucleotide monomer molecule or oligonucleotide chain through any intermediate linking structure.
[0147] Said R 3 can have the structure shown in Chemical Formula (III-L):
[0148]
[0149] Wherein the wavy line represents that the R 3 structure can be connected to any atom on the triazole group, nucleotide monomer molecule or oligonucleotide chain, or can be connected to the triazole group, nucleotide monomer molecule or oligonucleotide chain through any intermediate linking structure.
[0150] The fluorescent group refers to a chemical group containing a fluorescent label, and the fluorescent label refers to a substance that can absorb light of a specific wavelength, such as light with a wavelength of 520 nm, and then re-emit visible light of a color under the excitation of light of that wavelength. The fluorescent group is usually covalently bound to other molecules and is used as an indicator for detecting the presence of other molecules in the environment. Common reactive groups include, but are not limited to, amine-reactive isothiocyanate derivatives such as FITC and TRITC (derivatives of fluorescein and rhodamine), amine-reactive succinimidyl esters such as NHS-fluorescein, and thiol-reactive maleimide-activated fluorescein (fluor) such as fluorescein-5-maleimide. In some embodiments, the fluorescent group is covalently linked to a nucleotide monomer or oligonucleotide chain for detecting the presence of the nucleotide monomer or oligonucleotide chain in the intracellular environment. The fluorescent group can be indirectly linked to the nucleotide monomer or oligonucleotide chain through other nucleotide-modifying molecules, or can be directly linked to the pentose sugar, base, and / or phosphate group of the nucleotide molecule.
[0151] The fluorescent group can be labeled with 6-carboxyfluorescein FAM.
[0152] When R 3 is an alkynyl group, the nucleotide monomer can include structures such as Formula (I-1), (I-2), (I-3), (I-4) and (I-5):
[0153]
[0154]
[0155] When R 3 has the structure shown in Chemical Formula (II), the nucleotide monomer can have the structure of the following formula:
[0156]
[0157] wherein R 1 and R 2 independently may also be selected from one or a combination of hydrogen, a sugar protecting group, a reactive group, a solid support or a compound containing a solid support, a phosphate ester, a phosphate ester linked to an oligonucleotide, or a therapeutic agent.
[0158] The phosphate ester may be a modified phosphate ester, and the modified phosphate ester may enhance the activity of nucleotides containing the modified phosphate ester and increase more biological functions. The modified phosphate ester may be adding and / or removing a certain chemical group on the structure of the original phosphate ester, or changing one or more atoms on the structure of the original phosphate ester. Common phosphate ester modifications may include phosphorothioate (PS), dithiophosphate (PS2), methylphosphonate (MP), methoxypropylphosphonate (MOP), and peptide nucleic acid (PNA), etc. The modified phosphate ester may replace an oxygen atom in the phosphate ester bond with a sulfur atom, thereby replacing the phosphate ester bond with a phosphorothioate bond. For example, a non-bridging oxygen atom in a phosphodiester bond may be replaced with a sulfur atom, thereby replacing the phosphodiester bond with a phosphorothioate diester bond.
[0159] The phosphate ester may include one of the following groups: -P(=R Y2 )(R Y3 R Y1 )R Y1 or -P(=R Y2 )(R Y3 R Y1 )2. Wherein, each R Y1 independently may be selected from hydrogen, oxygen, a hydroxyl group, or a C1-C6 alkyl group optionally substituted by one or more halogens or cyano groups, R Y2 may be selected from O or S, and R Y3 may be selected from O or S. For example, the phosphate ester may include one of the following groups: -P(=O)(OH)2, -P(=S)(OH)2, -P(=O)(SH)2, -P(=S)(OH)(SH), -P(=O)(OH)(SH), or -P(=S)(SH)2.
[0160] Wherein the reactive group may be a phosphoramidite reactive group. The phosphoramidite functional group can undergo a coupling reaction with a hydroxyl group on a solid phase carrier, such as a resin, and is oxidized to form a solid phase carrier linked by a phosphodiester bond. During the phosphoramidite solid phase synthesis, the hydroxyl group on the nucleotide is deprotected, and then can undergo a coupling reaction with the phosphoramidite group on another nucleotide under coupling reaction conditions. The reactive group may be a 2-cyanoethyl N,N-diisopropylphosphoramidite group.
[0161] A sugar protecting group may refer to a group that is attached to a nucleic acid molecule and has a protective effect on the pentose sugar structure and the functional groups on the structure of the nucleic acid molecule, such as a hydroxyl protecting group. The sugar protecting group may include one of the following groups: acetyl (Ac), benzoyl (Bz), benzyl (Bn), β-methoxyethoxymethyl ether (MEM), dimethoxytrityl (DMT), methoxymethyl ether (MOM), methoxytrityl (MMT), p-methoxybenzyl ether (PMB), p-methoxyphenyl ether (PMP), pivaloyl (Piv), tetrahydrofuran (THF), trityl (Tr), tritylsilane, tritylsilyl ether, trimethylsilyl (TMS), triisopropylsilyloxymethyl (TOM), triisopropylsilyloxymethyl, 4,4'-dimethoxytrityl (DMTr), tert-butyldimethylsilyl (TBDMS).
[0162] The solid support refers to a solid-phase fixing agent used to fix the synthesized or unsynthesized nucleic acid chain or monomer nucleotide when the required nucleic acid chain is in the synthesis process. The solid support includes but is not limited to controlled pore glass beads (CPG) and polystyrene (PS).
[0163] Among them, the R 1 may be hydrogen. The R 1 may be a hydroxyl group. The R 1 may be 4,4'-dimethoxytrityl (DMTr). The R 1 may be acetyl (Ac). The R1 may be benzoyl (Bz). The R 1 may be benzyl (Bn). The R 1 may be β-methoxyethoxymethyl ether (MEM). The R 1 may be dimethoxytrityl (DMT). The R 1 may be methoxymethyl ether (MOM). The R 1 may be methoxytrityl (MMT). The R 1 may be p-methoxybenzyl ether (PMB). The R 1 may be p-methoxyphenyl ether (PMP). The R 1 may be pivaloyl (Piv). The R 1 may be tetrahydrofuran (THF). The R 1 may be trityl (Tr). The R 1 may be tritylsilane. The R 1 may be tritylsilyl ether. The R 1 may be trimethylsilyl (TMS). The R 1can be triisopropylsilyloxymethyl (TOM). The R 1 can be triisopropylsilyloxymethyl. The R 1 can be tert-butyldimethylsilyl (TBDMS).
[0164] Among them, the R 2 can be a phosphoramidite reaction group, and the R 2 can be 2-cyanoethyl N,N-diisopropylphosphoramidite group. The R 1 and R 2 can be 4,4'-dimethoxytriphenylmethyl (DMTr) and 2-cyanoethyl N,N-diisopropylphosphoramidite group respectively. Among them, R 1 can be 4,4'-dimethoxytriphenylmethyl (DMTr), and R 2 can be 2-cyanoethyl N,N-diisopropylphosphoramidite group.
[0165] Among them, the R 1 can have a structure of the following formula:
[0166]
[0167] where the wavy line indicates that the R 1 structure can be connected to other structures through the carbon atom on the benzyl group.
[0168] Among them, the R 2 can have a structure of the following formula:
[0169]
[0170] where the wavy line indicates that the R 2 structure can be connected to other structures through the phosphorus atom on the 2-cyanoethyl N,N-diisopropylphosphoramidite group.
[0171] Among them, the B group can be a base or a base analog. The base can refer to guanine, adenine, cytosine, thymine, and / or uracil. Among them, the B group can be guanine, the B group can be adenine, the B group can be cytosine, the B group can be thymine, and the B group can be uracil.
[0172] The base analog can be a guanine analog, an adenine analog, a cytosine analog, a thymine analog, and / or a uracil analog.
[0173] The base or base analogue may further include a base protecting group. A base protecting group refers to a group that is attached to a nucleotide monomer and has a protective effect on the base on the nucleotide monomer. The base protecting group can avoid unnecessary additional reactions of any functional group on the base of the nucleotide monomer during the extension of the nucleic acid chain, such as the amino group on the base. The base protecting group can be selected from one of the following groups: fluorenylmethyloxycarbonyl (Fmoc), tert-butoxycarbonyl (BOC), benzyloxycarbonyl (Cbz), optionally substituted acyl group, trifluoroacetyl group (TFA), benzyl group, triphenylmethyl group (Tr), 4,4'-dimethoxytriphenylmethyl group (DMTr), and tosyl group (Ts).
[0174] Wherein, B may have a structure as follows:
[0175]
[0176] Ligand represents a ligand or a fluorescent group. The fluorescent group refers to a chemical group containing a fluorescent label. The fluorescent label refers to a substance that can absorb light of a specific wavelength, such as light with a wavelength of 520 nm, and then re-emit visible light of a color under the excitation of light of that wavelength. The fluorescent group is usually covalently bound to other molecules and is used as an indicator to detect the presence of other molecules in the environment. Common reactive groups include, but are not limited to, amine-reactive isothiocyanate derivatives, such as FITC and TRITC (derivatives of fluorescein and rhodamine), amine-reactive succinimidyl esters, such as NHS-fluorescein, and thiol-reactive maleimide-activated fluorescein (fluor), such as fluorescein-5-maleimide. In some embodiments, the fluorescent group is used to detect the presence of the attached molecule in the intracellular environment. The fluorescent group can be labeled with 6-carboxyfluorescein FAM.
[0177] 2. Modified oligonucleotide
[0178] On the other hand, the present invention also provides a modified oligonucleotide that can be efficiently internalized by cells. The oligonucleotide contains the above-mentioned modified nucleotide monomers and has high stability when coupling other ligand molecules (such as Galnac molecules) through the modified nucleotide monomers. At the same time, the modified oligonucleotide can also achieve better cell internalization effect and higher in vivo delivery efficiency.
[0179] The oligonucleotide has a structure as shown below:
[0180]
[0181] Wherein R 2can be H, R1, a reactive group, a solid support or a compound containing a solid support, a phosphate ester, or a phosphate ester linked to an oligonucleotide. B can be a base or a base analogue, L 1 and / or L 2 can be a linking structure, oligo1 represents an oligonucleotide, R 3 can be an alkynyl group, or a ligand containing a triazole group, or a fluorophore containing a triazole group.
[0182] The oligonucleotide can have the structure shown below:
[0183]
[0184] wherein, B can be a base or a base analogue, L 1 and / or L 2 can be a linking structure, oligo1 and oligo2 represent oligonucleotide fragments, R 3 can be an alkynyl group, or a ligand containing a triazole group, or a fluorophore containing a triazole group.
[0185] The oligonucleotide can have the structure shown below:
[0186]
[0187] wherein oligo1 and oligo2 represent oligonucleotide fragments.
[0188] 2.1 B group
[0189] The modified oligonucleotide provided by the present invention may comprise a B group, and B can be a base or a base analogue.
[0190] The base may refer to guanine, adenine, cytosine, thymine and / or uracil.
[0191] The base analogue may be a guanine analogue, an adenine analogue, a cytosine analogue, a thymine analogue and / or a uracil analogue.
[0192] The base analogue may be a modified base. The modified base may be adding and / or removing a certain chemical group on the structure of the original base. For example, the base modification of nucleotides may include, but is not limited to, methylation, acetylation, phosphorylation and sulfation. The modified base may be a molecular rearrangement of the original base structure.
[0193] The modified base can be a modified base synthesized within the cell itself. For example, the mRNA synthesized within the cell itself can carry specific modified bases, thereby rendering the mRNA in a non-translated state. The modified base can also be a modified base introduced artificially. For example, a nucleic acid drug introduced from outside can have modified bases in its nucleotides, thereby endowing the nucleic acid drug with better stability and targeting ability.
[0194] The modified base can confer functions or biological activities different from those before modification to the nucleotide monomer and / or oligonucleotide chain containing the modified base. The modified base of the nucleotide can regulate gene expression by altering the structure and function of DNA or RNA. For example, DNA methylation can inhibit gene expression, while the 5'-end cap structure and 3'-end tail structure of RNA can affect the splicing and stability of mRNA. The modified base of the nucleotide on mRNA can affect the stability, subcellular localization of mRNA, and the translation efficiency of the protein encoded by mRNA.
[0195] The base analogs can be selected from one or more of the following groups: xanthine, allylaminopyrimidine, allylaminothymidine, hypoxanthine, dioxoadenine, dioxocytosine, dioxoguanine, dioxouracil, 6-chloropurine riboside, N6-methyladenine, methylpseudouracil, 2-thiocytosine, 2-thiouracil, 5-methyluracil, 4-thiothymidine, 4-thiouracil, 5,6-dihydro-5-methyluracil, 5,6-dihydro-uracil, 5-[(3-indolyl)propionamide-N-allyl]uracil, 5-aminoallylcytosine, 5-aminoallyluracil, 5-bromouracil, 5-bromocytosine, 5-carboxycytosine, 5-carboxycytosine, 5-carboxyuracil, 5-fluorouracil, 5-formylcytosine, 5-formyluracil, 5-hydroxycytosine, 5-hydroxymethylcytosine, 5-hydroxymethyluracil, 5-hydroxyuracil, 5-iodocytosine, 5-iodouracil, 5-methoxycytosine, 5-methoxyuracil, 5-methylcytosine, 5-methyluracil, 5-propynylaminocytosine, 5-propynylaminuracil, 5-propynylcytosine, 5-propynyluracil, 6-azacytosine, 6-azauracil, 6-chloropurine, 6-thioguanine, 7-deazaadenine, 7-deazaguanine, 7-deaza-7-propylaminoadenine, 7-deaza-7-propynylaminoadenine, 8-azaguanine, 8-azidoadenine, 8-chloroadenine, 8-oxoadenine, 8-oxoguanine, biotin-16-7-deaza-7-propynylaminoguanine, biotin-16-aminoallylcytosine, biotin-16-aminoallyluracil, 3-5-propynylaminocytosine, 3-6-propynylaminuracil, cyano 3-aminoallylcytosine, cyano 3-aminoallyluracil, cyano 5-6-propynylaminocytosine, cyano 5-6-propynylaminuracil, cyano 5-aminoallylcytosine, cyano 5-aminoallyluracil, cyano 7-aminoallyluracil, Dabcyl-5-3-aminoallyluracil, desthiobiotin-16-aminoallyluracil, desthiobiotin-6-aminoallylcytosine, isoguanine, N1-ethylpseudouracil, N1-methoxymethylpseudouracil, N1-methyladenine, N1-methylpseudouracil, N1-propylpseudouracil, N2-methylguanine, N4-biotin-OBEA-cytosine, N4-methylcytosine, N6-methyladenine, 06-methylguanine, pseudoisocytosine, pseudouracil, thiophencytosine, thiophenguanine, thiopheneuracil, xanthine, 3-deazaadenine, 2,6-diaminoadenine, 2,6-macroaminoguanine, 5-formamidouracil, 5-ethynyluracil, N6-isopentenyladenine (i6A), 2-methylthio-N6-isopentenyladenine (ms2i6A), 2-methylthio-N6-methyladenine (ms2m6A), N6-(cis-hydroxyisopentenyl)adenine (io6A), 2-methylthio-N6-(cis-hydroxyisopentenyl)adenine (ms2io6A), N6-glycylaminoformyladenine (g6A), N6-threonylaminoformyladenine (t6A), 2-methylthio-N6-threonylaminoformyladenine (ms2t6A), N6-methyl-N6-threonylaminoformyladenine (m6t6A), N6-hydroxyvalylaminoformyladenine (hn6A), 2-methylthio-N6-hydroxyvalylaminoformyladenine (ms2hn6A), N6,N6-dimethyladenine (m62A) and N6-acetyladenine (ac6A).
[0196] The base or base analog may further include a base protecting group. A base protecting group refers to a group attached to a nucleotide monomer that has a protective effect on the base on the nucleotide monomer. The base protecting group can prevent any functional group on the nucleotide monomer base from undergoing unnecessary additional reactions during the nucleic acid chain extension process, such as the amino group on the base. The base protecting group can be selected from one of the following groups: fluorenylmethyloxycarbonyl (Fmoc), tert-butyloxycarbonyl (BOC), benzyloxycarbonyl (Cbz), optionally substituted acyl, trifluoroacetyl (TFA), benzyl, trityl (Tr), 4,4'-dimethoxytrityl (DMTr) and toluenesulfonyl (Ts).
[0197] The modified oligonucleotide can interact with the functional target sequence, thereby affecting the normal function of the target sequence molecule, such as causing mRNA fragmentation or translation repression or exon skipping to trigger mRNA alternative splicing, etc. The modified oligonucleotide can be completely complementary to the bases of the target sequence, or can be complementary to 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or more of the bases of the target sequence.
[0198] 2.2L 1 and L 2
[0199] The modified oligonucleotide provided by the present invention may have L 1 and / or L 2 The connection structure. 1 The role of L can be to connect the ribose structure of the nucleotide and the modified triazole structure.2 may serve to link the modified triazole structure to R 3 and.
[0200] L 1 may have a structure of an aliphatic chain containing 0 - 20 carbon atoms, such as an aliphatic chain containing 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 carbon atoms. To endow the modified nucleotide monomer with better cell internalization ability and stability, L 1 may have a structure of an aliphatic chain containing 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 carbon atoms. Further, L 1 may have a structure of an aliphatic chain containing 1, 2, 3, 4 or 5 carbon atoms. L 1 may have a structure of a straight-chain aliphatic chain containing 0 - 20 carbon atoms, such as a straight-chain aliphatic chain containing 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 carbon atoms. To endow the modified nucleotide monomer with better cell internalization ability and stability, L 1 may have a structure of a straight-chain aliphatic chain containing 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 carbon atoms. Further, L 1 may have a structure of a straight-chain aliphatic chain containing 1, 2, 3, 4 or 5 carbon atoms. L 1 may have a structure of a straight-chain ether structure containing 0 - 20 carbon atoms, such as a straight-chain ether structure containing 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 carbon atoms. To endow the modified nucleotide monomer with better cell internalization ability and stability, L 1 may have a structure of a straight-chain ether structure containing 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 carbon atoms. Further, L 1 may have a structure of a straight-chain ether structure containing 1, 2, 3, 4 or 5 carbon atoms.
[0201] L 2 may have a structure of an aliphatic chain containing 0 - 20 carbon atoms, such as an aliphatic chain containing 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 carbon atoms. To endow the modified nucleotide monomer with better cell internalization ability and stability, L 2 may have a structure of an aliphatic chain containing 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 carbon atoms. Further, L 2The structure of L can be an aliphatic chain containing 1, 2, 3, 4 or 5 carbon atoms. 2 The structure of L can be a linear fatty chain containing 0-20 carbon atoms, for example, a linear fatty chain containing 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 carbon atoms. In order to make the modified nucleotide monomer have better cell internalization ability and stability, L 2 The structure of L can be a straight aliphatic chain containing 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 carbon atoms, further L 2 The structure of L can be a straight aliphatic chain containing 1, 2, 3, 4 or 5 carbon atoms. 2 The structure of L can be a linear ether structure containing 0-20 carbon atoms, for example, a linear ether structure containing 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 carbon atoms. In order to make the modified nucleotide monomer have better cell internalization ability and stability, L 2 The structure of L can be a linear ether structure containing 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 carbon atoms, further L 2 The structure of can be a straight chain ether structure containing 1, 2, 3, 4 or 5 carbon atoms.
[0202] The L 1 Can be selected from one of the following structures: optionally substituted -(CH2) n -、-((CH2) m O) n -and-((CH2) m S) n -. Wherein, n can be an integer of 0-20, preferably 1-10, more preferably 1-5; m can be an integer of 0-20, preferably 0-10, more preferably 0-5. 2 Can be selected from one of the following structures: optionally substituted -(CH2) n -、-((CH2) m O) n -and-((CH2) m S) n -, wherein m can be an integer from 0 to 20, preferably from 1 to 15; n can be an integer from 0 to 20, preferably from 0 to 10, more preferably from 0 to 5.
[0203] The L 1 -(CH2) may be optionally substituted n-, where n can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20. To enable the modified nucleotide monomer to have better cell internalization ability and stability, n can be 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10. Further, n can be 1, 2, 3, 4 or 5.
[0204] The L 1 can be -(CH2)-, -(CH2)2-, -(CH2)3-, -(CH2)4-, -(CH2)5-, -(CH2)6-, -(CH2)7-, -(CH2)8-, -(CH2)9- or -(CH2) 10 -. For example, the L 1 can be -(CH2)-.
[0205] The L 1 can be any substituted -((CH2) m O) n -, where n can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20, and m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20. To enable the modified nucleotide monomer to have better cell internalization ability and stability, n can be 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10, and m can be 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10. Further, n can be 1, 2, 3, 4 or 5, and m can be 1, 2, 3, 4 or 5.
[0206] The L 1 can be any substituted -((CH2) m S) n -, where n can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20, and m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20. To enable the modified nucleotide monomer to have better cell internalization ability and stability, n can be 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10, and m can be 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10. Further, n can be 1, 2, 3, 4 or 5, and m can be 1, 2, 3, 4 or 5.
[0207] The L 1structure, where m and n can have the following quantitative relationships: when n is 1, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20; when n is 2, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20; when n is 3, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20; when n is 4, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20; when n is 5, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20; when n is 6, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20; when n is 7, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20; when n is 8, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20; when n is 9, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20; when n is 10, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20; when n is 11, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20; when n is 12, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20; when n is 13, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20; when n is 14, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20; when n is 15, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20; when n is 16, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, ..........................................................11, 12, 13, 14, 15, 16, 17, 18, 19, or 20; It should be noted that there seems to be an ellipsis in the original text after "10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20" which is not fully reproduced in the translation for clarity. If the full text is required for a more accurate translation, please provide the complete and correct original text.When n is 17, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20; when n is 18, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20; when n is 19, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20; when n is 20, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20.
[0208] The said L 2 can be any substituted -(CH2) n -, where n can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20. In order to make the modified nucleotide monomer have better cell internalization ability and stability, where n can be 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10. Further, where n can be 1, 2, 3, 4 or 5.
[0209] The said L 2 can be -(CH2)-, -(CH2)2-, -(CH2)3-, -(CH2)4-, -(CH2)5-, -(CH2)6-, -(CH2)7-, -(CH2)8-, -(CH2)9-, -(CH2) 10 -, -(CH2) 11 -, -(CH2) 12 -, -(CH2) 13 -, -(CH2) 14 -, -(CH2) 15 -, -(CH2) 16 -, -(CH2) 17 -, -(CH2) 18 -, -(CH2) 19 - or -(CH2) 20 -. For example, the said L 2 can be -(CH2) 12 -.
[0210] The said L 2 can be any substituted -((CH2) m O) n-, where n can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20, and where m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20. To make the modified nucleotide monomer have better cell internalization ability and stability, where n can be 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10, and where m can be 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10. Further, where n can be 1, 2, 3, 4 or 5, and where m can be 1, 2, 3, 4 or 5.
[0211] The L 2 can be any substituted -((CH2) m S) n -, where n can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20, and where m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20. To make the modified nucleotide monomer have better cell internalization ability and stability, where n can be 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10, and where m can be 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10. Further, where n can be 1, 2, 3, 4 or 5, and where m can be 1, 2, 3, 4 or 5.
[0212] The L 2structure, where m and n can have the following quantitative relationships: when n is 1, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20; when n is 2, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20; when n is 3, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20; when n is 4, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20; when n is 5, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20; when n is 6, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20; when n is 7, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20; when n is 8, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20; when n is 9, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20; when n is 10, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20; when n is 11, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20; when n is 12, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20; when n is 13, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20; when n is 14, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20; when n is 15, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20; when n is 16, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20;When n is 17, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20; when n is 18, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20; when n is 19, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20; when n is 20, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20.
[0213] The group of any substituted substituent can be selected from the following group: halogen, cyano, hydroxyl, nitro, amino, alkylamino, cycloalkylamino, heterocyclic group, aminocarbonyl, sulfonyl, aminosulfonyl, carbonylamino, sulfonylamino, methyl, ethyl, aryl, methoxy, ethoxy, trifluoromethyl, trifluoroethyl, trifluoromethoxy, trifluoroethoxy, polyoxyethylene, polyoxypropylene, C1-C3 alkyl polyoxyethylene, or C1-C3 alkyl polyoxypropylene.
[0214] The position and number of the any substituted substituent group on the substituted group are also restricted by the valence state of each substituted group or atom. For example, at most four substituent groups can exist on a carbon atom. The position and number of the any substituted substituent group on the substituted group are also restricted by the steric hindrance of each substituent group. Multiple substituent groups with large steric hindrance, such as tert-butyl or isopropyl, cannot exist simultaneously on the same substituted group or atom.
[0215] In some embodiments, L 1 is -CH2-.
[0216] In some embodiments, L 2 is -(CH2) 12 -.
[0217] 2.3 Phosphate
[0218] There is a phosphate on the modified oligonucleotide provided by the present invention. The phosphate can be an ester derivative compound of phosphoric acid, and generally can refer to a derivative compound formed by condensation after an esterification reaction between phosphoric acid, a salt containing phosphoric acid, a compound containing a phosphate group and other compounds containing a hydroxyl group.
[0219] The phosphate ester may be a modified phosphate ester, and the modified phosphate ester can enhance the activity of nucleotides containing the modified phosphate ester and increase more biological functions. The modified phosphate ester may be adding and / or removing a certain chemical group on the structure of the original phosphate ester, or changing one or more atoms on the structure of the original phosphate ester. Common phosphate ester modifications may include phosphorothioate (PS), dithiophosphate (PS2), methylphosphonate (MP), methoxypropylphosphonate (MOP), and peptide nucleic acid (PNA), etc.
[0220] The modified phosphate ester may replace the oxygen atom in the phosphate ester bond with a sulfur atom, thereby replacing the phosphate ester bond with a phosphorothioate bond. For example, the non-bridging oxygen atom in the phosphodiester bond can be replaced with a sulfur atom, thereby replacing the phosphodiester bond with a phosphorothioate diester bond. Introducing the modified phosphate ester on the oligonucleotide chain can stabilize the structure of the oligonucleotide chain and maintain high specificity and high affinity of base pairing.
[0221] The phosphate ester may include one of the following groups: -P(=R Y2 )(R Y3 R Y1 )R Y1 、-P(=R Y2 )(R Y3 R Y1 )2. Wherein each R Y1 independently may be selected from hydrogen, oxygen, hydroxyl, or a C1-C6 alkyl group optionally substituted by one or more halogens or cyano groups, R Y2 may be selected from O or S, and R Y3 may be selected from O or S. For example, the phosphate ester may include one of the following groups: -P(=O)(OH)2, -P(=S)(OH)2, -P(=O)(SH)2, -P(=S)(OH)(SH), -P(=O)(OH)(SH), or -P(=S)(SH)2.
[0222] 2.4R1
[0223] The modified oligonucleotide provided by the present invention may have R1 at the 5' or 3' position of the ribose structure of the monomer nucleotide.
[0224] R1 can refer to a group that is attached to a nucleic acid molecule and protects the pentose sugar structure and the functional groups thereon of the nucleic acid molecule, such as a hydroxyl protecting group. The protecting group can make the chemical functional group insensitive to specific reaction conditions and can also be attached to and removed from the functional group in the nucleic acid molecule without affecting the rest of the nucleic acid molecule. For example, R1 can avoid any functional group on the pentose sugar structure, such as a hydroxyl group, from undergoing side reactions unrelated to nucleic acid chain elongation during the nucleic acid chain elongation process. For example, the protected functional groups on the pentose sugar can be hydroxyl groups at the 2', 3' and / or 5' positions, etc. The protecting group can maintain stable connection under alkaline conditions and can be removed under acidic conditions.
[0225] The R1 described above can include one of the following groups: acetyl (Ac), benzoyl (Bz), benzyl (Bn), β-methoxyethoxymethyl ether (MEM), dimethoxytrityl (DMT), methoxymethyl ether (MOM), methoxytrityl (MMT), p-methoxybenzyl ether (PMB), p-methoxyphenyl ether (PMP), pivaloyl (Piv), tetrahydrofuran (THF), trityl (Tr), tritylsilane, tritylsilyl ether, trimethylsilyl (TMS), triisopropylsilyloxymethyl (TOM), triisopropylsilyloxymethyl, 4,4'-dimethoxytrityl (DMTr), tert-butyldimethylsilyl (TBDMS).
[0226] The nucleotide monomer provided by the present invention can have R1 at the 5' of the ribose structure. The R1 can be 4,4'-dimethoxytrityl (DMTr). The R1 can be acetyl (Ac). The R1 can be benzoyl (Bz). The R1 can be benzyl (Bn). The R1 can be β-methoxyethoxymethyl ether (MEM). The R1 can be dimethoxytrityl (DMT). The R1 can be methoxymethyl ether (MOM). The R1 can be methoxytrityl (MMT). The R1 can be p-methoxybenzyl ether (PMB). The R1 can be p-methoxyphenyl ether (PMP). The R1 can be pivaloyl (Piv). The R1 can be tetrahydrofuran (THF). The R1 can be trityl (Tr). The R1 can be tritylsilane. The R1 can be tritylsilyl ether. The R1 can be trimethylsilyl (TMS). The R1 can be triisopropylsilyloxymethyl (TOM). The R1 can be triisopropylsilyloxymethyl. The R1 can be tert-butyldimethylsilyl (TBDMS).
[0227] The nucleotide monomer provided by the present invention may have R1 at the 3'-position of the ribose structure. The R1 may be 4,4'-dimethoxytriphenylmethyl (DMTr). The R1 may be acetyl (Ac). The R1 may be benzoyl (Bz). The R1 may be benzyl (Bn). The R1 may be β-methoxyethoxymethyl ether (MEM). The R1 may be dimethoxytriphenylmethyl (DMT). The R1 may be methoxymethyl ether (MOM). The R1 may be methoxytriphenylmethyl (MMT). The R1 may be p-methoxybenzyl ether (PMB). The R1 may be p-methoxyphenyl ether (PMP). The R1 may be pivaloyl (Piv). The R1 may be tetrahydrofuran (THF). The R1 may be triphenylmethyl (Tr). The R1 may be triphenylmethylsilane. The R1 may be triphenylmethylsilyl ether. The R1 may be trimethylsilyl (TMS). The R1 may be triisopropylsilyloxymethyl (TOM). The R1 may be triisopropylsilyloxymethyl. The R1 may be tert-butyldimethylsilyl (TBDMS).
[0228] 2.5 Reaction group
[0229] The modified oligonucleotide provided by the present invention may have a reaction group at the 5'- or 3'-position of the ribose structure of the monomer nucleotide.
[0230] The reaction group may be a phosphoramidite group. Solid-phase phosphoramidite synthesis is a common method for synthesizing oligonucleotide chains in the art. The phosphoramidite reaction group in the present invention also refers to nucleoside phosphoramidite, that is, the phosphoramidite coupled with a nucleotide. The phosphoramidite functional group can undergo a coupling reaction with a hydroxyl group on a solid-phase support, such as a resin, and is oxidized to form a solid-phase support connected by a phosphodiester bond. During the phosphoramidite solid-phase synthesis process, the hydroxyl group on the nucleotide is deprotected and then can undergo a coupling reaction with the phosphoramidite group on another nucleotide under the coupling reaction conditions.
[0231] The nucleotide monomer on the modified oligonucleotide provided by the present invention may have a reaction group at the 5'-position of the ribose structure, and the reaction group may be 2-cyanoethyl N,N-diisopropylphosphoramidite group. The nucleotide monomer provided by the present invention may have a reaction group at the 3'-position of the ribose structure, and the reaction group may be 2-cyanoethyl N,N-diisopropylphosphoramidite group.
[0232] 2.6 R 3
[0233] The modified oligonucleotide provided by the present invention may have R at the end of its monomer nucleotide 3 . R 3The function can be a linking structure for connecting other molecules, such as containing functional groups that can undergo coupling reactions.
[0234] The R 3 can be a structure containing an alkynyl group.
[0235] The R 3 can be a structure containing a triazole group.
[0236] The R 3 can be a fluorescent group containing a triazole group. The fluorescent group can be labeled with 6-carboxyfluorescein FAM.
[0237] The R 3 can be a ligand containing a triazole group.
[0238] The R 3 As shown in chemical formula (II):
[0239]
[0240] Wherein, L 3 can be a linking structure, L can be a ligand, or can be a fluorescent group.
[0241] 2.6.1L 3
[0242] L 3 The function of L 3 can connect the triazole structure of R
[0243] The L 3 can be selected from one of the following structures: optionally substituted -(CH2) n -, -(O(CH2) m ) n -, -(S(CH2) m ) n -. Wherein, n can be an integer from 0 to 20, preferably 1 to 15; m can be an integer from 0 to 20, preferably 0 to 10, more preferably can be 0 to 5. The groups of the optionally substituted substituents can be selected from the following group: halogen, cyano, hydroxyl, nitro, amino, alkylamino, cycloalkylamino, heterocyclic group, aminocarbonyl, sulfonyl, aminosulfonyl, carbonylamino, sulfonylamino, methyl, ethyl, aryl, methoxy, ethoxy, trifluoromethyl, trifluoroethyl, trifluoromethoxy, trifluoroethoxy, polyoxyethylene, polyoxypropylene, C1-C3 alkyl polyoxyethylene, or C1-C3 alkyl polyoxypropylene.
[0244] L 3The structure can be a fatty chain containing 0 - 20 carbon atoms. For example, it can be a fatty chain containing 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 carbon atoms. To enable the modified nucleotide monomer to have better cell internalization ability and stability, L 3 The structure can be a fatty chain containing 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 carbon atoms. Further, L 3 The structure can be a fatty chain containing 1, 2, 3, 4 or 5 carbon atoms. L 3 The structure can be a straight-chain fatty chain containing 0 - 20 carbon atoms. For example, it can be a straight-chain fatty chain containing 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 carbon atoms. To enable the modified nucleotide monomer to have better cell internalization ability and stability, L 3 The structure can be a straight-chain fatty chain containing 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 carbon atoms. Further, L 3 The structure can be a straight-chain fatty chain containing 1, 2, 3, 4 or 5 carbon atoms. L 3 The structure can be a straight-chain ether structure containing 0 - 20 carbon atoms. For example, it can be a straight-chain ether structure containing 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 carbon atoms. To enable the modified nucleotide monomer to have better cell internalization ability and stability, L 3 The structure can be a straight-chain ether structure containing 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 carbon atoms. Further, L 3 The structure can be a straight-chain ether structure containing 1, 2, 3, 4 or 5 carbon atoms.
[0245] The said L 3 can be any substituted -(CH2) n -, where n can be 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20. To enable the modified nucleotide monomer to have better cell internalization ability and stability, where n can be 0, 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10. Further, where n can be 0, 1, 2, 3, 4 or 5. For example, n can be 0.
[0246] The said L 3 can be any substituted -(O(CH2) m ) n-, where n can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20, and where m can be 1, 2, 3, 4, 5, 6, 7,
[0247] 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20. To endow the modified nucleotide monomer with better cellular internalization ability and stability, n can be 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10, and m can be 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10. Further, n can be 1, 2, 3, 4 or 5, and m can be 1, 2, 3, 4 or 5.
[0248] The L 3 can be any substituted -(S(CH2) m ) n -, where n can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20, and where m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20. To endow the modified nucleotide monomer with better cellular internalization ability and stability, n can be 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10, and m can be 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10. Further, n can be 1, 2, 3, 4 or 5, and m can be 1, 2, 3, 4 or 5.
[0249] The L 3The structure, where m and n can have the following quantitative relationships: when n is 1, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20; when n is 2, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20; when n is 3, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20; when n is 4, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20; when n is 5, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20; when n is 6, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20; when n is 7, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20; when n is 8, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20; when n is 9, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20; when n is 10, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20; when n is 11, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20; when n is 12, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20; when n is 13, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20; when n is 14, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20; when n is 15, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20; when n is 16, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20;When n is 17, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20; when n is 18, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20; when n is 19, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20; when n is 20, m can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20.
[0250] 2.6.2L
[0251] R 3 Furthermore, the L structure can be connected. The L structure can be a ligand or a fluorescent group.
[0252] The ligand can be a ligand capable of binding to a cell surface receptor, such as a ligand capable of binding to a mammalian cell surface receptor, a ligand capable of binding to a liver surface receptor, or a ligand capable of binding to the asialoglycoprotein receptor (ASGPR) on the liver surface.
[0253] The ligand can be any ligand having an affinity for the asialoglycoprotein receptor (ASGPR) on the surface of mammalian hepatocytes. The ligand can be a carbohydrate or a carbohydrate derivative, which can be a monosaccharide, disaccharide, trisaccharide or polysaccharide. It can be a modified monosaccharide, disaccharide, trisaccharide or polysaccharide. The ligand can be independently selected from the group consisting of the following sugars: glucose and its derivatives, mannan and its derivatives, galactose and its derivatives, xylose and its derivatives, ribose and its derivatives, fucose and its derivatives, lactose and maltose and their derivatives, arabinose and its derivatives, fructose and its derivatives, or sialic acid.
[0254] The ligand can be selected from one of the following groups: galactose, galactosamine, N-acetylgalactosamine, mannose, glucose, glucosamine, N-acetylglucosamine, fucose or lactose, N-acetylgalactosamine with all hydroxyl groups fully protected by acyl groups, galactose with all hydroxyl groups fully protected by acyl groups, galactosamine with all hydroxyl groups fully protected by acyl groups, N-formyl-galactosamine with all hydroxyl groups fully protected by acyl groups, N-propionyl-galactosamine with all hydroxyl groups fully protected by acyl groups, N-n-butyryl-galactosamine with all hydroxyl groups fully protected by acyl groups, or N-isobutyryl-galactosamine with all hydroxyl groups fully protected by acyl groups, where the acyl group can be an acetyl group or a benzoyl group.
[0255] The ligand L structure can be as shown in Chemical Formula (III):
[0256]
[0257] Wherein the wavy line represents that the ligand structure can be connected to any atom on the triazole group, nucleotide monomer molecule or oligonucleotide chain, or can be connected to the triazole group, nucleotide monomer molecule or oligonucleotide chain through any intermediate linking structure.
[0258] Said R 3 can have the structure shown in Chemical Formula (III-L):
[0259]
[0260] Wherein the wavy line represents that the R 3 structure can be connected to any atom on the triazole group, nucleotide monomer molecule or oligonucleotide chain, or can be connected to the triazole group, nucleotide monomer molecule or oligonucleotide chain through any intermediate linking structure.
[0261] The fluorescent group refers to a chemical group containing a fluorescent label, and the fluorescent label refers to a substance that can absorb light of a specific wavelength, such as light with a wavelength of 520 nm, and then re-emit visible light of a color under the excitation of light of that wavelength. The fluorescent group is usually covalently bound to other molecules and is used as an indicator to detect the presence of other molecules in the environment. Common reactive groups include, but are not limited to, amine-reactive isothiocyanate derivatives such as FITC and TRITC (derivatives of fluorescein and rhodamine), amine-reactive succinimidyl esters such as NHS-fluorescein, and thiol-reactive maleimide-activated fluorescein (fluor) such as fluorescein-5-maleimide. In some embodiments, the fluorescent group is covalently linked to a nucleotide monomer or oligonucleotide chain for detecting the presence of the nucleotide monomer or oligonucleotide chain in the intracellular environment. The fluorescent group can be indirectly linked to the nucleotide monomer or oligonucleotide chain through other nucleotide-modifying molecules, or can be directly linked to the pentose sugar, base, and / or phosphate group of the nucleotide molecule.
[0262] The fluorescent group can be labeled with 6-carboxyfluorescein FAM.
[0263] When R 3 is an alkynyl group, the nucleotide monomer can include structures such as Formula (I-1), (I-2), (I-3), (I-4) and (I-5):
[0264]
[0265]
[0266] The sequence of the oligonucleotide fragment oligo1 can be as shown in SEQ ID NO.: 1, and the sequence of the oligonucleotide fragment oligo2 can be as shown in SEQ ID NO.: 2. The sequence of the oligonucleotide fragment oligo1 can be as shown in SEQ ID NO.: 2, and the sequence of the oligonucleotide fragment oligo2 can be as shown in SEQ ID NO.: 1.
[0267] The oligonucleotide fragment oligo1 and / or oligo2 can contain modifications with fluorescent groups. For example, the 1st position of oligo1 can contain a 6-carboxyfluorescein FAM label.
[0268] 3. Preparation method
[0269] The modified nucleotide monomers provided by the present invention can be prepared by any feasible synthesis route in the art. For example, the modified nucleotide monomers provided by the present invention can couple a fatty chain containing an azide group or an alkyne group at the 2'-position of ribose. The modified nucleotide monomers can be prepared by the following method:
[0270] The modified oligonucleotides provided by the present invention can be prepared by any feasible synthesis route in the art. For example, the modified oligonucleotides provided by the present invention can be prepared by the following method: Under the conditions of phosphoramidite solid-phase synthesis, the nucleoside monomers are sequentially linked in the 3'-to-5' direction according to the nucleotide types and sequences corresponding to the used oligonucleotides. The connection of each nucleoside monomer includes four steps of deprotection, coupling, capping, and oxidation or sulfurization.
[0271] The method can also, in the presence of a coupling reagent, contact the modified nucleotide monomers and / or unmodified nucleotide monomers with the modified nucleotide and / or unmodified nucleotide sequence linked to a solid-phase carrier, so that the free nucleotide monomers and the nucleotides on the solid-phase carrier are linked to the nucleotide sequence through a coupling reaction.
[0272] The method can also include the steps of removing protecting groups and cleaving from the solid-phase carrier, separation and purification steps, and an optional annealing step.
[0273] The separation and purification methods used in the present invention can be methods of conventional operations in the art, generally including cleaving the synthesized nucleotide sequence from the solid-phase carrier, removing the protecting groups on the bases, phosphate groups, and ligands, and purification and desalting.
[0274] Among them, the synthesized nucleotide sequence is cleaved from the solid support, and the protecting groups on the bases, phosphate groups, and ligands are removed. The nucleotide sequence attached to the solid support can be contacted with concentrated ammonia water; during the deprotection process, the corresponding nucleoside with a free 2'-hydroxyl group is obtained. The concentrated ammonia water may refer to 25-30% ammonia water.
[0275] When there is at least one 2'-TBDMS protection on the synthesized nucleotide sequence, the method further includes contacting the nucleotide sequence from which the solid support has been removed with triethylamine trihydrofluoride to remove the 2'-TBDMS protection, and obtaining the corresponding nucleoside with a free 2'-hydroxyl group.
[0276] The methods of purification and desalting used in the present invention are methods of conventional operations in the art. For example, a preparative ion chromatography purification column can be used to complete the purification of nucleic acids by gradient elution with NaBr or NaCl. After the product is collected and combined, a reverse-phase chromatography purification column can be used for desalting.
[0277] During the synthesis of the modified oligonucleotide, the purity and molecular weight of the nucleic acid sequence can be detected at any time, so as to better control the synthesis quality. The detection methods are well known to those skilled in the art. For example, the nucleic acid purity can be detected by ion exchange chromatography, and the molecular weight can be determined by liquid chromatography-mass spectrometry.
[0278] After obtaining the modified oligonucleotide of the present disclosure, methods such as liquid chromatography-mass spectrometry can also be used to characterize the synthesized modified oligonucleotide by molecular weight detection and the like, to determine that the synthesized modified oligonucleotide is the target-designed modified oligonucleotide, and the sequence of the synthesized oligonucleotide is consistent with the sequence of the oligonucleotide to be synthesized, for example, consistent with the sequence listed in any one of SEQ ID NO.: 1-3.
[0279] Without being bound by any theory, the following examples are only for explaining the modified nucleotide molecules, preparation methods, uses, etc. of the present invention, and are not used to limit the scope of the invention of this application.
[0280] The present application also discloses the following embodiments:
[0281] 1. A compound, including the structure shown in formula (I):
[0282]
[0283] Wherein R 1 and R 2 are independently selected from one or a combination of hydrogen, sugar protecting groups, reactive groups, solid supports or compounds containing solid supports, phosphate esters, or oligonucleotides;
[0284] B is a base or a base analog;
[0285] R 3 is an alkynyl group, or a ligand containing a triazole group, or a fluorophore containing a triazole group;
[0286] L 1 and / or L 2 is a linking structure.
[0287] 2. The compound according to Embodiment 1, further comprising a structure represented by formula (I-C1):
[0288]
[0289] 3. The compound according to any one of Embodiments 1-2, wherein B is selected from one or more of the following: guanine, adenine, cytosine, uracil, guanine analogs, adenine analogs, cytosine analogs, and uracil analogs.
[0290] 4. The compound according to Embodiment 3, wherein B is selected from one or more of the following group: xanthine, allylaminopyrimidine, allylaminothymidine, hypoxanthine, dioxoadenine, dioxocytosine, dioxoguanine, dioxouracil, 6-chloropurine riboside, N6-methyladenine, methylpseudouracil, 2-thiocytosine, 2-thiouracil, 5-methyluracil, 4-thiothymidine, 4-thiouracil, 5,6-dihydro-5-methyluracil, 5,6-dihydrouracil, 5-[(3-indolyl)propionamide-N-allyl]uracil, 5-aminoallylcytosine, 5-aminoallyluracil, 5-bromouracil, 5-bromocytosine, 5-carboxycytosine, 5-carboxycytosine, 5-carboxyuracil, 5-fluorouracil, 5-formylcytosine, 5-formyluracil, 5-hydroxycytosine, 5-hydroxymethylcytosine, 5-hydroxymethyluracil, 5-hydroxyuracil, 5-iodocytosine, 5-iodouracil, 5-methoxycytosine, 5-methoxyuracil, 5-methylcytosine, 5-methyluracil, 5-propynylaminocytosine, 5-propynylaminuracil, 5-propynylcytosine, 5-propynyluracil, 6-azacytosine, 6-azauracil, 6-chloropurine, 6-thioguanine, 7-deazaadenine, 7-deazaguanine, 7-deaza-7-propylaminoadenine, 7-deaza-7-propynylaminoadenine, 8-azaguanine, 8-azidoadenine, 8-chloroadenine, 8-oxoadenine, 8-oxoguanine, biotin-16-7-deaza-7-propynylaminoguanine, biotin-16-aminoallylcytosine, biotin-16-aminoallyluracil, 3-5-propynylaminocytosine, 3-6-propynylaminuracil, cyano 3-aminoallylcytosine, cyano 3-aminoallyluracil, cyano 5-6-propynylaminocytosine, cyano 5-6-propynylaminuracil, cyano 5-aminoallylcytosine, cyano 5-aminoallyluracil, cyano 7-aminoallyluracil, Dabcyl-5-3-aminoallyluracil, desthiobiotin-16-aminoallyluracil, desthiobiotin-6-aminoallylcytosine, isoguanine, N 1 -ethylpseudouracil, N 1 -methoxymethylpseudouracil, N1-methyladenine, N 1 -methylpseudouracil, N 1 -propylpseudouracil, N2-methylguanine, N 4 -biotin-OBEA-cytosine, N4-methylcytosine, N 6-Methyladenine, 06-methylguanine, pseudoisocytosine, pseudouracil, thiophenocytosine, thiophenoguanine, thiophenouracil, xanthine, 3-deazaadenine, 2,6-diaminoadenine, 2,6-diaminoguanine, 5-formyluracil, 5-ethynyluracil, N 6 -Isopentenyladenine (i6A), 2-methylthio-N 6 -Isopentenyladenine (ms2i6A), 2-methylthio-N 6 -Methyladenine (ms2m6A), N 6 -(Cis-hydroxyisopentenyl)adenine (io6A), 2-methylthio-N 6 -(Cis-hydroxyisopentenyl)adenine (ms2io6A), N 6 -Glycylcarbamoyladenine (g6A), N 6 -Threonylcarbamoyladenine (t6A), 2-methylthio-N 6 -Threonylcarbamoyladenine (ms2t6A), N 6 -Methyl-N 6 -Threonylcarbamoyladenine (m6t6A), N 6 -Hydroxyvalylcarbamoyladenine (hn6A), 2-methylthio-N 6 -Hydroxyvalylcarbamoyladenine (ms2hn6A), N 6 , N 6 , N-Dimethyladenine (m62A) and N 6 -Acetyladenine (ac6A).
[0291] 5. The compound according to any one of Embodiments 3-4, wherein the base or its analogue further comprises a base protecting group, and the base protecting group is selected from one of the following groups: fluorenylmethyloxycarbonyl (Fmoc), tert-butoxycarbonyl (BOC), benzyloxycarbonyl (Cbz), optionally substituted acyl, trifluoroacetyl (TFA), benzyl, triphenylmethyl (Tr), 4,4'-dimethoxytriphenylmethyl (DMTr), and tosyl (Ts).
[0292] 6. The compound according to any one of Embodiments 1-4, wherein L 1 is selected from one of the following structures:
[0293] Optionally substituted -(CH2) n -, -((CH2) m O) n -, and -((CH2) m S) n -; wherein,
[0294] n is an integer from 0 to 20, preferably from 1 to 10, more preferably from 1 to 5; m is an integer from 0 to 20, preferably from 0 to 10, more preferably from 0 to 5; the groups of any substituted substituents are selected from the following groups: halogen, cyano, hydroxy, nitro, amino, alkylamino, cycloalkylamino, heterocyclic group, aminocarbonyl, sulfonyl, aminosulfonyl, carbonylamino, sulfonylamino, methyl, ethyl, aryl, methoxy, ethoxy, trifluoromethyl, trifluoroethyl, trifluoromethoxy, trifluoroethoxy, polyoxyethylene, polyoxypropylene, C1-C3 alkyl polyoxyethylene, or C1-C3 alkyl polyoxypropylene.
[0295] 7. The compound according to any one of Embodiments 1-5, wherein L 2 is selected from one of the following structures:
[0296] Optionally substituted -(CH2) n -, -((CH2) m O) n - and -((CH2) m S) n -; wherein
[0297] n is an integer from 0 to 20, preferably from 1 to 15; m is an integer from 0 to 20, preferably from 0 to 10, more preferably from 0 to 5;
[0298] The groups of any substituted substituents are selected from the following group: halogen, cyano, hydroxy, nitro, amino, alkylamino, cycloalkylamino, heterocyclic group, aminocarbonyl, sulfonyl, aminosulfonyl, carbonylamino, sulfonylamino, methyl, ethyl, aryl, methoxy, ethoxy, trifluoromethyl, trifluoroethyl, trifluoromethoxy, trifluoroethoxy, polyoxyethylene, polyoxypropylene, C1-C3 alkyl polyoxyethylene, or C1-C3 alkyl polyoxypropylene.
[0299] 8. The compound according to any one of Embodiments 1-6, wherein the phosphate ester includes one of the following groups:
[0300] -P(=R Y2 )(R Y3 R Y1 )R Y1 、-P(=R Y2 )(R Y3 R Y1 )2; wherein,
[0301] Each R Y1 is independently selected from hydrogen, oxygen, hydroxy or C1-C6 alkyl optionally substituted by one or more halogens or cyano;
[0302] R Y2 is selected from O or S;
[0303] RY3 Selected from O or S.
[0304] 9. The compound according to any one of Embodiments 1-8, wherein the sugar protecting group comprises one of the following groups: acetyl (Ac), benzoyl (Bz), benzyl (Bn), β-methoxyethoxymethyl ether (MEM), dimethoxytrityl (DMT), methoxymethyl ether (MOM), methoxytrityl (MMT), p-methoxybenzyl ether (PMB), p-methoxyphenyl ether (PMP), pivaloyl (Piv), tetrahydrofuran (THF), trityl (Tr), tritylsilane, tritylsilyl ether, trimethylsilyl (TMS), triisopropylsilyloxymethyl (TOM), triisopropylsilyloxymethyl, 4,4'-dimethoxytrityl (DMTr), or tert-butyldimethylsilyl (TBDMS).
[0305] 10. The compound according to any one of Embodiments 1-9, wherein the reactive group is a phosphoramidite group.
[0306] 11. The compound according to Embodiment 10, wherein the phosphoramidite group is a 2-cyanoethyl N,N-diisopropylphosphoramidite group.
[0307] 12. The compound according to any one of Embodiments 1-10, wherein R 3 is an alkynyl group.
[0308] 13. The compound according to any one of Embodiments 1-10, wherein R 3 is a ligand containing a triazole group.
[0309] 14. The compound according to Embodiment 13, wherein R 3 is as shown in Chemical Formula (II)
[0310]
[0311] wherein, L 3 is a linking structure, and L is a ligand.
[0312] 15. The compound according to any one of Embodiments 14, wherein L 3 is selected from one of the following structures:
[0313] Optionally substituted -(CH2) n -, -(O(CH2) m ) n -, -(S(CH2) m ) n -; wherein
[0314] n is an integer from 0 to 20, preferably from 1 to 15; m is an integer from 0 to 20, preferably from 0 to 10, more preferably from 0 to 5;
[0315] The optionally substituted substituents are selected from the group consisting of: halogen, cyano, hydroxy, nitro, amino, alkylamino, cycloalkylamino, heterocyclic group, aminocarbonyl, sulfonyl, aminosulfonyl, carbonylamino, sulfonylamino, methyl, ethyl, aryl, methoxy, ethoxy, trifluoromethyl, trifluoroethyl, trifluoromethoxy, trifluoroethoxy, polyoxyethylene, polyoxypropylene, C1-C3 alkyl polyoxyethylene, or C1-C3 alkyl polyoxypropylene.
[0316] 16. The compound according to any one of embodiments 14-15, wherein the ligand L is selected from one of the following groups: galactose, galactosamine, N-acetylgalactosamine, mannose, glucose, glucosamine, N-acetylglucosamine, fucose or lactose, N-acetylgalactosamine with all hydroxyl groups fully protected by acyl, galactose with all hydroxyl groups fully protected by acyl, galactosamine with all hydroxyl groups fully protected by acyl, N-formyl-galactosamine with all hydroxyl groups fully protected by acyl, N-propionyl-galactosamine with all hydroxyl groups fully protected by acyl, N-n-butyryl-galactosamine with all hydroxyl groups fully protected by acyl or N-isobutyryl-galactosamine with all hydroxyl groups fully protected by acyl, wherein the acyl group is acetyl or benzoyl.
[0317] 17. The compound according to any one of embodiments 14-16, wherein the ligand L is as shown in chemical formula (III):
[0318]
[0319] 18. The compound according to any one of embodiments 14-17, wherein R 3 is as shown in chemical formula (III-L)
[0320] 19. The compound according to any one of embodiments 1-12, having a structure of the following formula
[0321]
[0322]
[0323] 20. The compound according to any one of embodiments 1-18, having a structure of the following formula:
[0324]
[0325] wherein R 1 and R 2 are phosphate esters, or phosphate esters linked with oligonucleotides, and Ligand represents a ligand or a fluorescent group.
[0326] 21. A conjugate comprising an oligonucleotide moiety having the structure shown below:
[0327]
[0328] wherein R 2 is H, a sugar protecting group, a reactive group, a solid support or a compound containing a solid support, a phosphate ester, or a phosphate ester linked to an oligonucleotide;
[0329] B is a base or a base analogue;
[0330] L 1 and / or L 2 is a linking structure, oligo1 represents an oligonucleotide,
[0331] R 3 is an alkynyl group, or a ligand containing a triazole group, or a fluorophore containing a triazole group.
[0332] 22. The conjugate according to embodiment 21, having the structure shown below
[0333]
[0334] B is a base or a base analogue;
[0335] L 1 and / or L 2 is a linking structure, wherein oligo1 and oligo2 represent oligonucleotides,
[0336] R 3 is an alkynyl group, or a ligand containing a triazole group, or a fluorophore containing a triazole group.
[0337] 23. The conjugate according to any one of embodiments 21-22, having the structure shown below:
[0338]
[0339] wherein oligo1 and oligo2 represent oligonucleotides.
[0340] 24. The conjugate according to any one of embodiments 22-23, wherein the sequence of oligonucleotide oligo1 is as shown in SEQ ID NO.: 1, and the sequence of oligonucleotide oligo2 is as shown in SEQ ID NO.: 2.
[0341] The conjugate according to embodiment 24, wherein the 21st position of oligonucleotide oligo1 is labeled with 6-carboxyfluorescein FAM.
[0342] Example
[0343] Example 1: Preparation of ACS8270-ZJ991009 (T-C14) monomer
[0344]
[0345] First step:
[0346] Add 40 g (0.15 mol) of Compound 1 (purchased from Shanghai Haohong Biopharmaceutical Technology Co., Ltd., CAS: 1463-10-1) to a 1 L single-necked flask, dissolve it with 0.5 L (Compound 1 is 0.3 mol / L) of pyridine, add 63.5 g (0.2 mol) of TiPDSCl2 reagent at 20 °C, stir at room temperature for 3 hours, and TLC shows that the reaction is completed. Rotavaporize pyridine, dissolve the obtained crude product of Compound 2 with 200 ml of ethyl acetate, wash it once with 300 ml of 1 mol / L HCl, separate the organic phase, then wash it once with 300 ml of 10% Na2CO3 solution, and then wash it with 10% NaCl solution. Rotavaporize the organic phase, and column chromatography (eluent ratio: ethyl acetate / n-heptane = 1:5) gives 67 g of white solid of Compound 2 (purity 87%).
[0347]
[0348] Second step:
[0349] Add 66 g (0.13 mol) of Compound 2, 132 ml of dimethyl sulfoxide, 203 ml of acetic acid, and 132 ml of acetic anhydride to a 1 L single-necked flask in sequence. Stir at 50 °C for 15 hours, and TLC shows that the reaction is completed. Pour the reaction solution into 1.5 L of water to quench. Extract with 1000 ml of ethyl acetate, adjust the pH of the organic phase to neutral with 10% Na2CO3 solution, separate the organic phase and rotavaporize it, and column chromatography (eluent ratio: ethyl acetate / n-heptane = 1:10) gives 52 g of white solid of Compound 3 (purity 71%).
[0350]
[0351] Third step:
[0352] Add 35 g (0.062 mol) of Compound 3 and 500 ml of dichloromethane to a 1 L single-necked flask in sequence. Dropwise add 9.3 g (0.069 mol) of sulfonyl chloride at room temperature. Stir at 25 °C for 2 hours, and TLC shows that the reaction is completed. Pour the reaction solution into methanol to quench, and rotavaporize the reaction solution to obtain 38 g (111%), which is directly used for the next step.
[0353]
[0354] To a 0.5 L single-necked flask, 34 g (0.062 mol) of compound 4, 240 ml of dimethylformamide, and 20 g (0.31 mol) of solid sodium azide were added successively. After stirring at 25 °C for 24 hours, the reaction was terminated by TLC (quenched with methanol). The reaction solution was poured into 0.8 L of water for quenching, extracted with 500 ml of ethyl acetate, the organic phase was separated and dried over sodium sulfate, and the organic phase was rotary evaporated to obtain compound 5 (purity 123%).
[0355]
[0356] Step 5:
[0357] To a 2 L three-necked flask, 60 g (0.61 mol) of trimethylsilylacetylene and 0.3 L of dimethylformamide (trimethylsilylacetylene at 0.5 mol / L) were added. At -70 °C, 0.38 L (0.61 mol) of n-butyllithium solution was added dropwise. After the addition was complete, the temperature was raised to 0 °C and stirred for 1 hour, then 110 g (0.61 mol) of hexamethylphosphoramide was added, and then a 100 ml tetrahydrofuran solution of compound 6 (50 g) was added dropwise. After the addition was complete, the ice bath was removed, and the reaction was carried out at room temperature for 4 hours. The reaction was terminated by TLC. The reaction solution was poured into 1.5 L of aqueous ammonium chloride solution for quenching, extracted with 1000 ml of ethyl acetate, the organic phase was separated, dried, and rotary evaporated, and column chromatography was carried out with n-heptane to obtain 40 g of compound 7 (purity 73%) ( Figure 1 ).
[0358] 1 1H NMR (400 MHz, DMSO): δ 2.02 - 2.09 (m, 4H), 1.31 - 1.42 (m, 5H), 1.16 - 1.3 (m, 4H), 1.1 - 1.18 (m, 13H), 0.0 (m, 18H) ppm.
[0359]
[0360] Step 6:
[0361] To a 1 L single-necked flask, 90 g (0.248 mol) of compound 7 was added, dissolved in 0.5 L of dimethylformamide (compound 7 at 0.5 mol / L), then 28 g (0.124 mol) of tetrapropylammonium fluoride was added, and the reaction was carried out at room temperature for 2 hours. The reaction was terminated by TLC. The dimethylformamide was rotary evaporated, the reaction solution was poured into 0.5 L of water for quenching, extracted with 200 ml of ethyl acetate, the organic phase was separated, dried, and rotary evaporated, and column chromatography was carried out with n-heptane to obtain 55 g of white solid compound 8 (purity 102%) ( Figure 2 ).
[0362] 1H NMR (400MHz, DMSO): δ2.23(m,4H),1.9(t,2H),1.5(m,4H),1.3(m,4H),1.26(m,13H)ppm.
[0363]
[0364] Step 7:
[0365] To a 100 mL single-necked flask, 2 g (0.00442 mol) of compound 5, 70 mg (0.000442 mol) of copper sulfate, 174 mg (0.00088 mol) of ascorbic acid sodium salt, and 0.33 g (0.00442 mol) of compound 8 were added in sequence. The mixture was dissolved in 10 mL (0.5 mol / L of compound 5) of tetrahydrofuran and reacted at 60°C for 4 hours. After the reaction was completed by TLC, the organic phase was separated, dried, and filtered through a column (developing solvent: ethyl acetate / n-heptane) to obtain 4 g of compound 9 as a white solid (purity 105%).
[0366]
[0367] Step 8:
[0368] 25 g (0.032 mol) of compound 9 was added to a 0.5 L single-necked bottle, dissolved in 110 ml (0.3 mol / L of compound 9) of dimethylformamide, and then 7.4 g (0.032 mol) of tetrapropylammonium fluoride was added. The reaction was allowed to react at room temperature for 2 hours. After the TLC reaction was completed, the reaction solution was extracted with 150 ml of water and 150 ml of ethyl acetate. The organic phase was separated, dried and spin-dried, and washed with 10% ethyl acetate / n-heptane solution to obtain 17 g of white solid compound 10 (purity 92%) ( Figure 3 ).
[0369] 1 H NMR (400MHz, DMSO): δ11.3(s,1H),7.8(s,1H),7.5(s,1H),5.7-5.8(t,1H),5.69 -5.7(t,1H),5.66(s,1H),5.16(s,1H),5.14(s,1H),4.13-4.14(m,2H),3.85(d,1 H),3.74(m,1H),3.72(m,1H),3.14(m,1H),2.75(s,1H),2.6(m,2H),2.16(m,2H) ,1.75(m,3H),1.55(m,3H),1.45(m,2H),1.35(m,6H),1.1(m,1H),0.9(m,2H)ppm.
[0370]
[0371] Step 9:
[0372] 16g (0.03mol) of compound 10 was added to a 250mL single-necked bottle, dissolved in 150ml (0.2mol / L of compound 10) of dichloromethane, and then 15g (0.12mol) of N,N-diisopropylethylamine was added. 20g (0.06mol) of 4,4'-dimethoxytriphenylmethane was added in two batches under an ice bath and reacted at room temperature for 2 hours. After the TLC reaction was completed, the reaction solution was extracted with 150ml of water and 150ml of dichloromethane, the organic phase was separated, dried and spin-dried, and passed through a column (developer ratio: methanol / dichloromethane = 1:20) to obtain 24g of solid compound 11 (purity 96%) ( Figure 4 ).
[0373] 1 H NMR (400MHz, DMSO): δ11.3(s,1H),7.9(s,1H),7.25-7.36(m,2H),7.03(m,3H),6.9(d,2H),5.8-5.66(m,2H),5.44(s,1H),4.28 (m,2H),3.99(m,2H),3.74(d,3H),3.2(m,2H),2.75(d,1H),2.6(m,1H),2.16(m,1H),1.55(m,1H),1.45(m,1H),1.35(m,2H)ppm.
[0374]
[0375] Step 10:
[0376] 14g (0.017mol) of compound 11 was added to a 250mL single-necked bottle, dissolved in 150ml of dichloromethane, and then 4.3g (0.033mol) of N,N-diisopropylethylamine was added. PAMCl (7.7g, 0.033mmol) was added under ice bath, nitrogen protection was removed, and the reaction was carried out at room temperature for 2.5 hours. Most of the dichloromethane was dried and passed through a column (developing solvent ratio: methanol / dichloromethane = 1:50) to obtain 15.7g of solid compound 12 (purity 89%) ( Figure 5 ).
[0377] 11H NMR (400 MHz, DMSO): δ 11.3 (s, 1H), 7.2 - 7.38 (m, 10H), 6.9 (m, 4H), 5.7 - 5.8 (m, 3H), 4.5 (m, 2H), 4.1 - 4.2 (d, 1H), 3.7 (m, 7H), 3.3 (m, 4H), 3.2 (m, 2H), 2.7 (d, 1H), 2.49 - 2.51 (m, 3H), 2.12 (m, 2H), 1.5 (m, 2H), 1.4 - 1.1 (m, 2H), 0.9 (m, 3H) ppm.
[0378]
[0379] Example 2: Chemical Synthesis of T-C14 Oligonucleotide DNA
[0380] DNA-1 sequence: 5’-CACCTTAAAAATTTTTTCGATCTGGCCCATTTGGGACAAGTTCC-3’; where the T at position 30 is T-C14.
[0381] Oligo1: 5’-CACCTTAAAAATTTTTTCGATCTGGCCCA-3’.
[0382] Oligo2: 5’-TTGGGACAAGTTCC-3’.
[0383] The synthesis of the DNA-1 sequence was carried out using the classical solid-phase synthesis method of oligonucleotides. Starting from the solid-phase support as the initial cycle, nucleoside monomers were sequentially connected in the 3’-5’ direction according to the nucleotide arrangement order. Each connection of a nucleoside involved four steps: deprotection, coupling, capping, and oxidation or thiolation. Finally, a DNA-1 molecule with a solid-phase support was obtained. The conditions for each step of the reaction are as follows:
[0384] (1) Nucleoside monomer: Dissolved in an acetonitrile solution with a concentration of 0.1 mol / L.
[0385] (2) Deprotection: Add a 3% dichloroacetic acid-dichloromethane solution.
[0386] (3) Coupling reaction: Add a 0.3 mol / L ETT acetonitrile solution.
[0387] (4) Oxidation reaction: Add a 0.05 mol / L iodine solution in tetrahydrofuran / pyridine / water (70 / 20 / 10, v / v / v).
[0388] (5) Thiolation reaction: Add a 0.2 mol / L pyridine solution of hydrogen xanthate.
[0389] (6) Capping reaction: Add 20% acetic anhydride-acetonitrile and pyridine / N-methylimidazole / acetonitrile (10 / 14 / 76, v / v / v) solution.
[0390] Add the synthesized DNA-1 with solid support into a 2 ml centrifuge tube, add 25 - 28% ammonia water, react at 55 °C for 16 hours, filter and then wash 3 times with 1 mL of 50% ethanol aqueous solution to remove the solid support. After the filtrate is concentrated and dried, the crude product of single-stranded DNA-1 is obtained and waiting for purification. Dissolve the crude product of single-stranded DNA-1 with 1 ml of RNase-free water, and purify it by ion-pair reverse-phase chromatography or ion-exchange chromatography. Detect the collected samples, and the qualified samples are combined for desalting to obtain the pure product of single-stranded DNA-1 with amino modification. The chromatographic identification of the purified product of DNA-1 synthesis is as Figure 6 shown. The chromatographic results show that the maximum peak area is detected at a retention time of 12.961 seconds. The mass spectrometry (MS) identification of the synthesized product of DNA-1 is as Figure 7 shown. The mass spectrometry results show that the maximum peak intensity is detected at a molecular weight of 14265 Da.
[0391] Example 3 Cell-free uptake assay of DNA-1
[0392] 3.1 Coupling of T-C14 and FAM
[0393] Add 6.25 uL of DMSO solution of fluorescein azide (100 mM, 625 nmol) and 10 uL of freshly prepared copper bromide solution (solute: 25 mM copper bromide: 250 nmol ligand = 1:1, solvent: water:DMSO:tert-butanol = 4:3:1) to 25 uL of DNA-1 aqueous solution (0.5 mM, 12.5 nmol). After reacting at 15 °C for 1 hour, add 200 uL of water for dilution, and then take samples for gel electrophoresis detection.
[0394] After the click chemical reaction, add EDTA·2Na with the same molar amount as copper ions to complex copper ions, then precipitate overnight with ethanol. After centrifuging to remove the supernatant, redissolve the precipitate, perform PAGE gel electrophoresis, and then cut the gel for purification. Crush the cut gel strip, add 6 mL of Elution Buffer (100 mM Tris + 100 mM EDTA + 500 mM NaCl; pH = 8.00; filtered through a 0.22 um filter membrane), centrifuge at 7500 rpm for 2 minutes and then take out the supernatant, and then filter with a 0.22 um or 0.45 um filter membrane and filter with an ultrafiltration tube with a molecular weight cut-off of 3K (4 mL - 7000 rpm - 20 minutes / time), and wash with water three times to complete the purification ( Figure 8 ). It can be seen from the gel electrophoresis results that the purity of the purified FAM-coupled DNA-1 molecule is significantly improved.
[0395] 3.2 Free cellular uptake of DNA-1
[0396] Hep-G2 cells were selected and cultured in DMEM cell culture medium supplemented with 10% fetal bovine serum in a 5% CO2 incubator at 37 °C. The cells were seeded into 12-well plates for culture, with 2×10 5 cells plated per well, and 1 mL of DMEM medium containing 10% fetal bovine serum was added. After 12 hours of culture, DNA-1 was added, and the final concentrations of DNA-1 were 20 nM and 50 nM, respectively. After continued culture for 12 hours, the cells were washed 3 times with fresh DMEM, and 200 nM of SiR-Hoechst was added and the cells were cultured for another 4 hours. Then, the fluorescence distribution and intensity of FAM and SiR-Hoechst in the cells were observed and photographed under a fluorescence microscope. The FAM group represents the distribution of DNA-1 labeled with 6-carboxyfluorescein FAM, the SiR-Heochst group represents the distribution of Hep-G2 cells stained with SiR-Heochst, and the merged group represents the chimerism of the first two images.
[0397] By means of click chemistry, the C14-terminal alkyne was reacted with azide-FAM, and the FAM group could be efficiently coupled to the C14 terminus. The results showed ( Figure 8 ) that in the 44-nt DNA-1 sequence, the introduction of a T-C14 monomer could effectively promote the entry of DNA-1 into cells and showed a dose-dependent effect, demonstrating good cellular internalization and delivery effects.
Claims
1. A compound comprising a structure represented by formula (I): wherein R 1 and R 2 are independently selected from hydrogen, a sugar protecting group, a reactive group, a solid support or a compound containing a solid support, a phosphate ester, or an oligonucleotide, or a combination thereof; B is a base or a base analogue; R 3 is an alkynyl group, or a ligand containing a triazole group, or a fluorescent group containing a triazole group; L 1 and / or L 2 is a connecting structure.
2. The compound according to claim 1, further comprising a structure represented by formula (I-C1):
3. The compound according to any one of claims 1-2, wherein the reactive group is a phosphoramidite group.
4. The compound according to claim 3, wherein the phosphoramidite group is a 2-cyanoethyl N,N-diisopropyl phosphoramidite group.
5. The compound according to any one of claims 1-10, wherein R 3 is a ligand containing a triazole group.
6. The compound according to claim 5, wherein R 3 as shown in Chemical Formula (II) Among them, L 3 is a connecting structure, and L is a ligand.
7. The compound according to any one of claims 1-6 has a structure of the following formula 8. The compound according to any one of claims 1-7 has a structure of the following formula: wherein R 1 and R 2 are phosphate esters, or phosphate esters linked with oligonucleotides, and Ligand represents a ligand or a fluorescent group.
9. A conjugate comprising an oligonucleotide moiety having a structure as shown below: wherein R 2 is H, a sugar protecting group, a reactive group, a solid support or a compound containing a solid support, a phosphate ester, or a phosphate ester linked to an oligonucleotide; B is a base or a base analogue; L 1 and / or L 2 is a connecting structure, where oligo1 represents an oligonucleotide, R 3 is an alkynyl group, or a ligand containing a triazole group, or a fluorophore containing a triazole group.
10. The conjugate according to claim 9 has a structure as shown below B is a base or a base analogue; L 1 and / or L 2 is a linking structure, where oligo1 and oligo2 represent oligonucleotides, and R 3 is an alkynyl group, or a ligand containing a triazole group, or a fluorophore containing a triazole group.