Small molecule regulatory selective splicing
Nucleic acid molecules with splice regulator binding sites and minigenes enable precise regulation of gene expression, addressing the limitations of current methods by ensuring conditional and tissue-specific protein expression, thereby enhancing therapeutic efficacy.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- RGENTA THERAPEUTICS INC
- Filing Date
- 2024-04-13
- Publication Date
- 2026-05-01
AI Technical Summary
Current methods for regulating gene expression, particularly through transcription and splicing, lack specificity and efficiency in controlling the timing and location of gene expression, which is crucial for therapeutic applications.
The use of nucleic acid molecules with specific splice regulator binding sites, such as DGAGTDDGHV or DGAGTDDNHV, and minigenes containing exons and introns, to control the splicing process and regulate transgene expression, combined with splice regulators that bind to these sites.
This approach allows for precise regulation of gene expression, enhancing the therapeutic potential by ensuring conditional, temporal, or tissue-specific expression of proteins, improving efficacy and safety of therapeutic proteins.
Smart Images

Figure 2026514051000001_ABST
Abstract
Description
[Technical Field]
[0001] (Cross-reference to related applications) This application claims the benefit and priority of U.S. Provisional Patent Application No. 63 / 496,199, filed on 14 April 2023, the disclosure of which is incorporated herein by reference in its entirety for all purposes.
[0002] This disclosure provides nucleic acid molecules for use in regulating the transcription of linked transgenes via the exon inclusion mechanism. [Background technology]
[0003] The regulation of gene expression involves a wide range of mechanisms used by cells to increase or decrease the production of specific gene products. Sophisticated programs of gene expression are widely observed in biology, for example, to induce developmental pathways, respond to environmental stimuli, or adapt to new food sources. In the pathology of disease or disorder, the regulation of various gene pathways influences the onset and development of many diseases or disorders. Several steps of gene expression, from transcription initiation to RNA processing and post-translational modification of proteins, can be regulated. In gene regulatory networks, often one gene regulator controls a second gene regulator, and a third gene regulator controls a third gene regulator, and the pattern may continue to form a subsequent cascade of gene regulators controlling other gene regulators.
[0004] By regulating transcription, we control the timing of transcription and the amount of RNA produced. The transcription of genes by RNA polymerase and subsequent splicing of premRNA can be regulated by several endogenous mechanisms. Protein translation can also be regulated. For example, one mechanism involves the regulation of mRNA translation at the initiation level, resulting from the recruitment of small ribosomal subunits, which can be regulated by mRNA secondary structure, antisense RNA binding, protein, or ligand (e.g., small molecule) binding. In some embodiments, a class of transcripts (e.g., riboswitches) acts as ribozymes (RNAs with autocatalytic activity), self-regulating their expression and acting as either a negative or positive feedback loop.
[0005] Transcriptional, splicing, and translational regulators may be used in synthetic biology to influence the increase or decrease of the level of a protein of interest. While not limited by theory, an increase in polypeptide levels may provide a therapeutic effect by providing polypeptides whose expression is reduced or absent in the tissue of interest, or a decrease in polypeptide levels may provide a therapeutic benefit by providing a reduction of polypeptides whose expression is increased or abnormal in the tissue of interest. While not limited by theory, controlling the timing or location of gene expression by, for example, the application or removal of splice regulators may improve the efficacy and / or safety of such therapeutic proteins by ensuring that expression is regulated conditionally (e.g., temporally or tissue-specifically). Non-limiting examples of genes that may be expressed include proteins, RNA, microRNAs (miRNAs), or small hairpin RNAs (shRNAs). Therefore, this disclosure, in part, provides nucleic acid molecules useful for turning on transgene expression using splice regulators (e.g., small molecules). This disclosure also provides vectors and pharmaceutical compositions containing such nucleic acid molecules and splice regulators (e.g., small molecules), intended for use in methods of regulating gene expression. [Overview of the project] [Means for solving the problem]
[0006] In this disclosure, in certain embodiments, a nucleic acid molecule is provided comprising a minigene located adjacent to the 5' side of a transgene, wherein the minigene comprises (a) a first exon adjacent to the 5' side of a first intron, (b) a start codon, and (c) a splice regulator binding site, the splice regulator binding site comprising the nucleic acid sequence DGAGTDDGHV (SEQ ID NO: 82) or DGAGTDDNHV (SEQ ID NO: 83), where D is A, G, or T, R is A, G, N is A, C, G, or T, H is A, C, or T, and V is A, G, or C.
[0007] In some embodiments, the splice regulatory factor binding site includes the nucleic acid sequence DGAGTRRGHV (SEQ ID NO: 1) or DGAGTRRNHV (SEQ ID NO: 2), where D is A, G, or T, R is A or G, N is A, C, G, or T, H is A, C, or T, and V is A, G, or C.
[0008] In some embodiments, the splice regulatory factor binding site includes the nucleic acid sequence DGAGTTTGHV (SEQ ID NO: 84), where D is A, G, or T, H is A, C, or T, and V is A, G, or C.
[0009] In some embodiments, the nucleic acid includes a splice regulator binding site containing one of sequence numbers 75 to 81.
[0010] In some embodiments, the nucleic acid comprises a first exon containing a start codon, a first intron, and a transgene, in the order of 5' to 3'.
[0011] In some embodiments, the nucleic acid further comprises a stop codon.
[0012] In some embodiments, the nucleic acid further comprises a second exon.
[0013] In some embodiments, the nucleic acid comprises a first exon containing a start codon, a second exon, and a transgene, in the order of 5' to 3'.
[0014] In some embodiments, the nucleic acid comprises, in order from 5' to 3', a first exon containing a start codon, a first intron, a second exon, and a transgene.
[0015] In some embodiments, the nucleic acid comprises, in the order of 5' to 3', a first exon, a second exon containing a start codon, and a transgene.
[0016] In some embodiments, the nucleic acid comprises, in order from 5' to 3', a first exon, a second exon containing a start codon, a first intron, and a transgene.
[0017] In some embodiments, the nucleic acid comprises, in 5' to 3' order, a first exon, a first intron, a second exon containing a start codon, a second intron, and a transgene.
[0018] In some embodiments, the nucleic acid contains at least about 80% sequence identity with respect to one of sequence numbers 21 to 73.
[0019] In some embodiments, the nucleic acid contains at least about 90% sequence identity with respect to one of sequence numbers 21 to 73.
[0020] In some embodiments, the nucleic acid includes one of sequence numbers 21 to 73.
[0021] In some embodiments, the nucleic acid further comprises a third exon. In some embodiments, the nucleic acid further comprises a third intron.
[0022] In some embodiments, the nucleic acid comprises, in order from 5' to 3', a first exon, a second exon containing a stop codon, a third exon containing a start codon, and a transgene.
[0023] In some embodiments, the nucleic acid comprises, in 5' to 3' order, a first exon, a first intron, a second exon containing a stop codon, a second intron, a third exon containing a start codon, a third intron, and a transgene.
[0024] In some embodiments, the nucleic acid comprises, in 5' to 3' order, a first exon, a first intron, a second exon containing a stop codon, a third exon containing a start codon, a second intron, and a transgene.
[0025] In some embodiments, the nucleic acid comprises, in order from 5' to 3', a first exon, a second exon containing a start codon, a third exon containing a stop codon, and a transgene.
[0026] In some embodiments, the nucleic acid comprises, in 5' to 3' order, a first exon, a first intron, a second exon containing a start codon, a second intron, a third exon containing a stop codon, a third intron, and a transgene.
[0027] In some embodiments, the nucleic acid comprises, in 5' to 3' order, a first exon, a first intron, a second exon containing a start codon, a third exon containing a stop codon, a second intron, and a transgene.
[0028] In this disclosure, in certain embodiments, a nucleic acid molecule is provided comprising a minigene located adjacent to the 5' side of a transgene, wherein the minigene comprises: (a) a first exon located at the 5' side of a first intron; (b) a start codon comprising a first portion and a second portion, wherein the first and second portions of the start codon are not in the same exon; and (c) a splice regulator binding site, which in some embodiments comprises the nucleic acid sequence DGAGTDDGHV (SEQ ID NO: 82) or DGAGTDDNHV (SEQ ID NO: 83), where D is A, G, or T, R is A, G, N is A, C, G, or T, H is A, C, or T, and V is A, G, or C.
[0029] In some embodiments, the splice regulatory factor binding site includes the nucleic acid sequence DGAGTRRGHV (SEQ ID NO: 1) or DGAGTRRNHV (SEQ ID NO: 2), where D is A, G, or T, R is A or G, N is A, C, G, or T, H is A, C, or T, and V is A, G, or C.
[0030] In some embodiments, the splice regulatory factor binding site includes the nucleic acid sequence DGAGTTTGHV (SEQ ID NO: 84), where D is A, G, or T, H is A, C, or T, and V is A, G, or C.
[0031] In some embodiments, the first portion of the start codon comprises one or two nucleotides of the start codon, and the second portion of the start codon comprises one or two nucleotides of the start codon.
[0032] In some embodiments, the first portion of the start codon is located in the first exon, and the second portion of the start codon is located in the transgene.
[0033] In some embodiments, the nucleic acid includes a first exon containing a first portion of the start codon, a first intron, and a transgene containing a second portion of the start codon, in the order of 5' to 3'.
[0034] In some embodiments, the nucleic acid further comprises a second exon. In some embodiments, the nucleic acid further comprises a second intron.
[0035] In some embodiments, the nucleic acid comprises, in order from 5' to 3', a first exon containing a first portion of the start codon, a second exon containing a second portion of the start codon, and a transgene.
[0036] In some embodiments, the nucleic acid comprises, in 5' to 3' order, a first exon containing a first portion of the start codon, a first intron, a second exon containing a second portion of the start codon, a second intron, and a transgene.
[0037] In some embodiments, the nucleic acid includes a transgene comprising a first exon in the order of 5' to 3', a second exon containing the first portion of the start codon, and a second portion of the start codon.
[0038] In some embodiments, the nucleic acid comprises a transgene including, in 5' to 3' order, a first exon, a second exon containing a first portion of the start codon, a first intron, and a second portion of the start codon.
[0039] In some embodiments, the nucleic acid comprises a first exon containing the first portion of the start codon, a second exon, and a transgene containing the second portion of the start codon, in the order of 5' to 3'.
[0040] In some embodiments, the nucleic acid comprises a transgene containing a first exon with a first portion of the start codon, a second exon, a first intron, and a second portion of the start codon, in the order of 5' to 3'.
[0041] In some embodiments, the nucleic acid comprises, in order from 5' to 3', a first exon containing a first portion of the start codon, a second exon containing a second portion of the start codon, and a transgene.
[0042] In some embodiments, the nucleic acid comprises, in order from 5' to 3', a first exon containing a first portion of the start codon, a first intron, a second exon containing a second portion of the start codon, and a transgene.
[0043] In some embodiments, the nucleic acid further comprises a stop codon.
[0044] In some embodiments, the nucleic acid comprises a first exon containing a first portion of the start codon, a second exon containing a stop codon, and a transgene containing a second portion of the start codon, in the order of 5' to 3'.
[0045] In some embodiments, the nucleic acid comprises a first exon containing a first portion of the start codon, a first intron, a second exon containing a stop codon, and a transgene containing a second portion of the start codon, in the order of 5' to 3'.
[0046] In some embodiments, the second exon includes a splice regulator binding site.
[0047] In some embodiments, the nucleic acid further comprises a third exon.
[0048] In some embodiments, the nucleic acid further includes a third intron.
[0049] In some embodiments, the nucleic acid comprises, in order from 5' to 3', a first exon containing the first portion of the start codon, a second exon containing the stop codon, a third exon containing the second portion of the start codon, and a transgene.
[0050] In some embodiments, the nucleic acid comprises, in 5' to 3' order, a first exon containing a first portion of the start codon, a first intron, a second exon containing a stop codon, a third exon containing a second portion of the start codon, a second intron, and a transgene.
[0051] In some embodiments, the nucleic acid comprises a transgene containing, in order from 5' to 3', a first exon, a second exon containing the first portion of the start codon, a third exon containing the stop codon, and a second portion of the start codon.
[0052] In some embodiments, the nucleic acid comprises a transgene including, in 5' to 3' order, a first exon, a first intron, a second exon containing a first portion of the start codon, a third exon containing a stop codon, a second intron, and a second portion of the start codon.
[0053] In some embodiments, the nucleic acid includes a transgene comprising, in the order from 5' to 3', a first exon, a second exon containing a stop codon, a third exon containing the first portion of the start codon, and a second portion of the start codon.
[0054] In some embodiments, the nucleic acid comprises a transgene including, in 5' to 3' order, a first exon, a first intron, a second exon containing a stop codon, a third exon containing a first portion of a start codon, a second intron, and a second portion of a start codon.
[0055] In some embodiments, the nucleic acid comprises a transgene including, in 5' to 3' order, a first exon, a second exon, a third exon containing the first portion of the start codon, and a second portion of the start codon.
[0056] In some embodiments, the nucleic acid comprises a transgene including, in 5' to 3' order, a first exon, a first intron, a second exon, a third exon including the first portion of the start codon, a second intron, and the second portion of the start codon.
[0057] In some embodiments, the nucleic acid comprises a transgene comprising a first exon in the order of 5' to 3', a second exon containing a stop codon, a third exon containing the first portion of the start codon, and a second portion of the start codon.
[0058] In some embodiments, the nucleic acid comprises a transgene including, in 5' to 3' order, a first exon, a first intron, a second exon containing a stop codon, a second intron, a third exon containing the first portion of a start codon, a third intron, and a second portion of a start codon.
[0059] In some embodiments, the nucleic acid comprises a transgene containing, in order from 5' to 3', a first exon, a second exon containing the first portion of the start codon, a third exon containing the stop codon, and a second portion of the start codon.
[0060] In some embodiments, the nucleic acid comprises a transgene including, in 5' to 3' order, a first exon, a first intron, a second exon containing a first portion of the start codon, a second intron, a third exon containing a stop codon, a third intron, and a second portion of the start codon.
[0061] In some embodiments, the first intron includes a splice regulator binding site. In some embodiments, the splice regulator binding site is located at the junction between the second exon and the second intron. In some embodiments, the splice regulator binding site is located within the second intron.
[0062] In certain embodiments of this disclosure, a nucleic acid molecule is provided, comprising, in order from 5' to 3', (a) a first exon containing a first portion of a start codon at the 3' end of a first exon, (b) a second exon containing a second portion of a start codon at the 5' end of a second exon, and (c) a third exon. In some embodiments, the third exon is a transgene.
[0063] In this disclosure, in certain embodiments, a nucleic acid molecule is provided that includes a transgene comprising a first exon, a second exon containing a stop codon, a third exon containing a first portion of a start codon, and a second portion of a start codon, in the order of 5' to 3'.
[0064] In this disclosure, in certain embodiments, a nucleic acid molecule is provided that includes a transgene comprising a first exon, a second exon containing a first portion of a start codon, a third exon containing a stop codon, and a second portion of a start codon, in the order of 5' to 3'.
[0065] In some embodiments, the nucleic acid molecule further comprises a first exon, a second exon, a third exon, and a first intron, a second intron, and a third intron interposed between the transgene.
[0066] In some embodiments, the splice regulatory factor binding site does not include the nucleic acid sequence AAGAGT (SEQ ID NO: 3), ATGAGT (SEQ ID NO: 4), TAGAGT (SEQ ID NO: 5), TTGAGT (SEQ ID NO: 6), GAGAGT (SEQ ID NO: 7), GTGAGT (SEQ ID NO: 8), AAGAGT (SEQ ID NO: 9), ATGAGT (SEQ ID NO: 10), ACGAGT (SEQ ID NO: 11), AGGAGT (SEQ ID NO: 12), AGAGGTAGAG (SEQ ID NO: 13), TGAGGTTGAG (SEQ ID NO: 14), GGAGGTGGAG (SEQ ID NO: 15), TAG (SEQ ID NO: 16), CAG (SEQ ID NO: 17), TAG (SEQ ID NO: 18), or NAGAGTNNNN (SEQ ID NO: 19), (wherein N is A, C, G, or T).
[0067] In some embodiments, the splice regulator binding site includes a splice regulator binding site that includes any one of sequence numbers 75 to 81.
[0068] In this disclosure, in certain embodiments, a composition is provided comprising a nucleic acid molecule of any embodiment described herein and a splice modifier, wherein the splice modifier is a compound of formula (I), [ka] or comprising a pharmaceutically acceptable salt, solvate, or prodrug thereof, wherein W is -S- or -HC=CH-, and R1 is H, halogen, hydroxyl, cyano, C1-C6 alkyl, C2-C6 alkenyl, C2-C6 alkynyl, C1-C6 haloalkyl, C1-C6 alkoxyl, C1-C6 haloalkoxyl, -(CH2) 0-2 -C3-C8 cycloalkyl, NH2, NH(C1-C6 alkyl), N(C1-C6 alkyl)2, or -(CH2) 0-2- is a heterocyclyl, where the heterocyclyl is a 4- to 7-membered ring containing 1, 2, or 3 heteroatoms independently selected from N, O, and S, and the alkyl, alkenyl, alkynyl, alkoxyl, cycloalkyl, or heterocyclyl may be optionally substituted with one or more C3-C8 cycloalkyl, aryl, or 4- to 7-membered heterocyclyls containing 1, 2, or 3 heteroatoms independently selected from N, O, and S, and R2 is an aryl, 5- to 7-membered cycloalkyl, 5, 6, or 9-membered heterocyclyl containing 1, 2, or 3 heteroatoms independently selected from N, O, and S, and the aryl, cycloalkyl, heterocyclyl, or heteroaryl may be optionally substituted with one or more R4, and each R3 may independently be a halogen, C1-C6 alkyl, C 2-C6 alkenyl, C2-C6 alkynyl, C1-C6 alkoxyl, C3-C8 cycloalkyl, NH2, NH(C1-C6 alkyl), or N(C1-C6 alkyl)2, wherein alkyl, alkenyl, alkynyl, alkoxyl, and cycloalkyl may be optionally substituted with one or more hydroxyl or NH2, and each R4 independently is halogen, hydroxyl, cyano, nitro, C1-C6 alkyl, C2-C6 alkenyl, C2-C6 Alkynyl, C1-C6 haloalkyl, C1-C6 alkoxyl, C1-C6 haloalkoxyl, C3-C8 cycloalkyl, NH2, NH(C1-C6 alkyl), N(C1-C6 alkyl)2, or C(O)NH2, wherein alkyl, alkenyl, alkynyl, alkoxyl, and cycloalkyl are 4-7 member heterocyclyl, NH2, NH(C1-C6 alkyl), containing one or more heteroatoms independently selected from hydroxyl, N, O, and S.Alternatively, R5 may be optionally substituted with N(C1-C6 alkyl)2, where R5 is H, C1-C6 alkyl, C2-C6 alkenyl, C2-C6 alkynyl, C1-C6 alkoxyl, C3-C8 cycloalkyl, -CH2C3-C8 cycloalkyl, heterocyclyl, -CH2 heterocisyl, -CH2CH2 heterocisyl, -CH2-(5-6 membered heteroaryl), where the heterocyclyl is a 4-7 membered ring and contains 1, 2, or 3 heteroatoms independently selected from N, O, and S, where R5 is one or more halo R6 is H, halogen, C1-C6 alkyl, C1-C6 haloalkyl, C1-C6 heteroalkyl, C1-C6 alkoxyl, C3-C8 cycloalkyl, spiro C3-C8 cycloalkyl, spiro 4-7 member heterocyclyl, 5-6 member heteroaryl, oxo, cyano, or hydroxyl, and R6 is H, halogen, C1-C6 alkyl, or C1-C6 haloalkyl, R7 is H, halogen, C1-C6 alkyl, or C1-C6 haloalkyl, and n is 0, 1, 2, 3, 4, or 5.
[0069] In a particular embodiment of this disclosure, a composition comprising (i) a nucleic acid molecule comprising a minigene located adjacent to the 5' side of a transgene, wherein the minigene comprises a) an exon located at the 5' side of an intron, b) a start codon having a first portion and a second portion, wherein the first and second portions of the start codon are not in the same exon, and c) a splice regulator binding site, and (ii) a substance bound to the splice regulator binding site, formula (I): [ka] Splice regulators including structures that conform to (I), The present invention provides a composition comprising either a pharmaceutically acceptable salt, solvate, or prodrug thereof, wherein W is -S- or -HC=CH-, and R1 is H, halogen, hydroxyl, cyano, C1-C6 alkyl, C2-C6 alkenyl, C2-C6 alkynyl, C1-C6 haloalkyl, C1-C6 alkoxyl, C1-C6 haloalkoxyl, -(CH2) 0-2 -C3-C8 cycloalkyl, NH2, NH(C1-C6 alkyl), N(C1-C6 alkyl)2, or -(CH2) 0-2- is a heterocyclyl, where the heterocyclyl is a 4- to 7-membered ring containing 1, 2, or 3 heteroatoms independently selected from N, O, and S, and the alkyl, alkenyl, alkynyl, alkoxyl, cycloalkyl, or heterocyclyl may be optionally substituted with one or more C3-C8 cycloalkyl, aryl, or 4- to 7-membered heterocyclyls containing 1, 2, or 3 heteroatoms independently selected from N, O, and S, and R2 is an aryl, 5- to 7-membered cycloalkyl, 5, 6, or 9-membered heterocyclyl containing 1, 2, or 3 heteroatoms independently selected from N, O, and S, or a 5, 6, or 9-membered heteroaryl containing 2, 3, or 4 heteroatoms independently selected from N, O, and S, and the aryl, cycloalkyl, heterocyclyl, or heteroaryl may be optionally substituted with one or more R4, and each R3 is independently a halogen, C1-C6 alkyl, C2 -C6 alkenyl, C2-C6 alkynyl, C1-C6 alkoxyl, C3-C8 cycloalkyl, NH2, NH(C1-C6 alkyl), or N(C1-C6 alkyl)2, wherein alkyl, alkenyl, alkynyl, alkoxyl, and cycloalkyl may be optionally substituted with one or more hydroxyl or NH2, and each R4 independently is halogen, hydroxyl, cyano, nitro, C1-C6 alkyl, C2-C6 alkenyl, C2-C6 a Lukinyl, C1-C6 haloalkyl, C1-C6 alkoxyl, C1-C6 haloalkoxyl, C3-C8 cycloalkyl, NH2, NH(C1-C6 alkyl), N(C1-C6 alkyl)2, or C(O)NH2, wherein alkyl, alkenyl, alkynyl, alkoxyl, and cycloalkyl are 4- to 7-membered heterocyclyls, NH2, NH(C1-C6 alkyl), containing one or more heteroatoms independently selected from hydroxyl, N, O, and S.Alternatively, R5 may be optionally substituted with N(C1-C6 alkyl)2, where R5 is H, C1-C6 alkyl, C2-C6 alkenyl, C2-C6 alkynyl, C1-C6 alkoxyl, C3-C8 cycloalkyl, -CH2C3-C8 cycloalkyl, heterocyclyl, -CH2 heterocisyl, -CH2CH2 heterocisyl, -CH2-(5-6 membered heteroaryl), where the heterocyclyl is a 4-7 membered ring and contains 1, 2, or 3 heteroatoms independently selected from N, O, and S, where R5 is one or more halo R6 is H, halogen, C1-C6 alkyl, C1-C6 haloalkyl, C1-C6 heteroalkyl, C1-C6 alkoxyl, C3-C8 cycloalkyl, spiro C3-C8 cycloalkyl, spiro 4-7 member heterocyclyl, 5-6 member heteroaryl, oxo, cyano, or hydroxyl, and R6 is H, halogen, C1-C6 alkyl, or C1-C6 haloalkyl, R7 is H, halogen, C1-C6 alkyl, or C1-C6 haloalkyl, and n is 0, 1, 2, 3, 4, or 5.
[0070] In this disclosure, in certain embodiments, a composition comprising: (i) a nucleic acid molecule comprising a minigene located adjacent to the 5' side of a transgene, wherein the minigene comprises (a) a first exon adjacent to the 5' side of a first intron, (b) a start codon, and (c) a splice regulator binding site; and (ii) a splice regulator comprising a structure that binds to the splice regulator binding site and conforms to formula (I): [ka] The present invention provides a composition comprising either a pharmaceutically acceptable salt, solvate, or prodrug thereof, wherein W is -S- or -HC=CH-, and R1 is H, halogen, hydroxyl, cyano, C1-C6 alkyl, C2-C6 alkenyl, C2-C6 alkynyl, C1-C6 haloalkyl, C1-C6 alkoxyl, C1-C6 haloalkoxyl, -(CH2) 0-2-C3-C8 cycloalkyl, NH2, NH(C1-C6 alkyl), N(C1-C6 alkyl)2, or -(CH2) 0-2-A heterocyclyl, wherein the heterocyclyl is a 4- to 7-membered ring containing 1, 2, or 3 heteroatoms independently selected from N, O, and S, and the alkyl, alkenyl, alkynyl, alkoxyl, cycloalkyl, or heterocyclyl may be optionally substituted with one or more C3-C8 cycloalkyl, aryl, or 4- to 7-membered heterocyclyls containing 1, 2, or 3 heteroatoms independently selected from N, O, and S, and R2 is aryl, 5- to 7-membered cycloalkyl, 1, 2, or or a 5, 6, or 9-membered heterocyclyl containing 3 heteroatoms, or a 5, 6, or 9-membered heteroaryl containing 2, 3, or 4 heteroatoms independently selected from N, O, and S, wherein the aryl, cycloalkyl, heterocyclyl, or heteroaryl may be optionally substituted with one or more R4s, each R3 independently being a halogen, C1-C6 alkyl, C2-C6 alkenyl, C2-C6 alkynyl, C1-C6 alkoxyl, C3-C8 cycloalkyl, NH2, NH(C1-C6 alkyl), or N(C1-C6 The alkyl, alkenyl, alkynyl, alkoxyl, and cycloalkyl groups may be optionally substituted with one or more hydroxyls or NH2 groups, and each R4 is independently a halogen, hydroxyl, cyano, nitro, C1-C6 alkyl, C2-C6 alkenyl, C2-C6 alkynyl, C1-C6 haloalkyl, C1-C6 alkoxyl, C1-C6 haloalkoxyl, C3-C8 cycloalkyl, NH2, NH(C1-C6 alkyl), N(C1-C6 alkyl)2, or C(O)NH2, and the alkyl, alkenyl Nyl, alkynyl, alkoxyl, and cycloalkyl may be optionally substituted with a 4- to 7-membered heterocyclyl, NH2, NH(C1-C6 alkyl), or N(C1-C6 alkyl)2 containing one or more heteroatoms independently selected from hydroxyl, N, O, and S, and R5 may be H, C1-C6 alkyl, C2-C6 alkenyl, C2-C6 alkynyl, C1-C6 alkoxyl, C3-C8 cycloalkyl, -CH2C3-C8 cycloalkyl, heterocyclyl, -CH2 heterocisyl, -CH2CH2 heterocisyl,-CH2-(a 5- to 6-membered heteroaryl), wherein the heterocyclyl is a 4- to 7-membered ring and contains 1, 2, or 3 heteroatoms independently selected from N, O, and S; wherein R5 is optionally substituted with one or more halogens, C1-C6 alkyl, C1-C6 haloalkyl, C1-C6 heteroalkyl, C1-C6 alkoxyl, C3-C8 cycloalkyl, spiro C3-C8 cycloalkyl, spiro 4- to 7-membered heterocyclyl, 5- to 6-membered heteroaryl, oxo, cyano, or hydroxyl; R6 is H, halogen, C1-C6 alkyl, or C1-C6 haloalkyl; R7 is H, halogen, C1-C6 alkyl, or C1-C6 haloalkyl; and n is 0, 1, 2, 3, 4, or 5.
[0071] In certain embodiments of the present disclosure, there is provided a composition comprising: (i) a nucleic acid molecule comprising a minigene located adjacent to the 5'-side of a transgene, the minigene comprising: (a) a first exon adjacent to the 5'-side of a first intron, (b) a start codon, and (c) a splice regulatory factor binding site; and (ii) a splice regulatory factor that binds to the splice regulatory factor binding site and comprises a structure according to Formula (II):
Chemical formula
化
Chemical formula
[0072] In this disclosure, in certain embodiments, a composition comprising (i) a nucleic acid molecule comprising a minigene located adjacent to the 5' side of a transgene, wherein the minigene comprises (a) a first exon adjacent to the 5' side of a first intron, (b) a start codon, and (c) a splice regulator binding site, and (ii) a splice regulator comprising a structure that binds to the splice regulator binding site and conforms to formula (III): [ka] The present invention provides a composition comprising either a pharmaceutically acceptable salt, solvate, or prodrug thereof, wherein A is a saturated or partially unsaturated monocyclic or bicyclic 4- to 9-membered heterocycloalkyl or NR1R2, the heterocycloalkyl comprising one or two nitrogen ring atoms optionally substituted with 1, 2, 3, or 4 R6s, R1 being a heterocycloalkyl comprising one nitrogen ring atom which may optionally be substituted with 1, 2, 3, or 4 R6s, and R2 being hydrogen, C 1-7 Alkyl, or C 3-8 It is a cycloalkyl group, and R3 is H, halo, C 1-7 Alkyl, OR5, N(R5)2, C 3-8 A cycloalkyl or heterocycloalkyl, where R4 is an aryl or bicyclic 9-membered heteroaryl containing 2, 3, or 4 heteroatoms independently selected from N, O, and S, where R4 is optionally substituted with 1, 2, or 3 R7 atoms, and each R5 is independently C 1-7 Alkyl, C 3-8 It is a cycloalkyl or heterocycloalkyl group, where each R6 is independently C 1-7 Alkyl, amino, amino-C 1-7 Alkyl, C 3-8 Cycloalkyl, heterocycloalkyl, or C 1-7It is either an alkoxy-heterocycloalkyl group, or both R6s are C 1-7 It forms an alkylene, and each R7 independently forms a halo, cyano, or C 1-7 Alkyl, C 1-7 Haloalkyl, C 1-7 Alkoxy, C 1-7 Haloalkoxy, or C 3-8 It is a cycloalkyl, C 1-7 The alkyl group may be optionally substituted with an OH group.
[0073] In some embodiments, the splice regulator that binds to the splice regulator binding site is selected from the group consisting of compounds 1A-192A and 100B-135B.
[0074] In some embodiments, the splice regulators that bind to the splice regulator binding site are selected from the group including 3A, 6A, 8A, 10A, 15A, 24A, 86A, 100B, 111B, 117B, 121B, 135B, and 192A.
[0075] In some embodiments, the splice regulators that bind to the splice regulator binding site are selected from the group including 116B, 100B, 1A, 22A, 24A, 2A, and 34A.
[0076] In some embodiments, the splice regulator binding site comprises the nucleic acid sequence DGAGTDDGHV (SEQ ID NO: 82) or DGAGTDDNHV (SEQ ID NO: 83), where D is A, G, or T, R is A or G, N is A, C, G, or T, H is A, C, or T, and V is A, G, or C.
[0077] In some embodiments, the splice regulator binds to a splice regulator binding site containing the nucleic acid sequence DGAGTRRGHV (SEQ ID NO: 1) or DGAGTRRNHV (SEQ ID NO: 2), where D is A, G, or T, R is A or G, N is A, C, G, or T, H is A, C, or T, and V is A, G, or C.
[0078] In some embodiments, the splice regulator binding site comprises the nucleic acid sequence DGAGTTTGHV (SEQ ID NO: 84), where D is A, G, or T, H is A, C, or T, and V is A, G, or C.
[0079] In some embodiments, the splice regulatory factor binding site does not include the nucleic acid sequence AAGAGT (SEQ ID NO: 3), ATGAGT (SEQ ID NO: 4), TAGAGT (SEQ ID NO: 5), TTGAGT (SEQ ID NO: 6), GAGAGT (SEQ ID NO: 7), GTGAGT (SEQ ID NO: 8), AAGAGT (SEQ ID NO: 9), ATGAGT (SEQ ID NO: 10), ACGAGT (SEQ ID NO: 11), AGGAGT (SEQ ID NO: 12), AGAGGTAGAG (SEQ ID NO: 13), TGAGGTTGAG (SEQ ID NO: 14), GGAGGTGGAG (SEQ ID NO: 15), TAG (SEQ ID NO: 16), CAG (SEQ ID NO: 17), TAG (SEQ ID NO: 18), or NAGAGTNNNN (SEQ ID NO: 19), (wherein N is A, C, G, or T).
[0080] In some embodiments, the splice regulator binds to a splice regulator binding site that includes one of the nucleic acid sequences from SEQ ID NOs. 75 to 81.
[0081] In some embodiments, the transgene encodes the target protein. In some embodiments, the transgene encodes the target miRNA. In some embodiments, the transgene encodes the target shRNA. In some embodiments, the transgene encodes the target functional or regulatory RNA.
[0082] In some embodiments, the nucleic acid molecule further includes a promoter. In some embodiments, the promoter is the GFAP promoter, nestin promoter, S100B promoter, Nefh promoter, dystrophin promoter, H1 promoter, 7SK promoter, apolipoprotein E-human-alpha-1-antitrypsin promoter, CK8 promoter, mU1a, EF-1α promoter, TBG promoter, PKG promoter, CAG, SV40 early promoter, mouse mammary tumor virus LTR promoter, Ad MLP, HSV promoter, CMV promoter such as CMV-IE, RSV promoter, U6 promoter or its variant, hSyn promoter, hexaribonucleotide-binding protein-3 (NeuN) promoter, CaMKII promoter, Tα-1 promoter, neuron-specific enolase (NSE) promoter, PDGFβ promoter, VGLUT promoter, SST promoter, NPY promoter, VIP promoter, PV promoter, GAD65 or GAD67 promoter, DRD1 and DRD2 promoters, MAP1B, C1ql2 promoter, POMC promoter, PROX1 promoter, or any suitable promoter.
[0083] In some embodiments, the molecule comprises polyA, optionally in this case polyA being SV40 polyA, HGH polyA, BGH polyA, betaglobin polyA, alphaglobin polyA, ovalbumin polyA, kappa light chain polyA, synthetic polyA, or any suitable polyA.
[0084] In certain embodiments, a vector comprising any nucleic acid molecule or composition of any embodiment described herein is provided herein, and optionally the vector is a plasmid, DNA vector, RNA vector, virion, or viral vector.
[0085] In some embodiments, the vector is a viral vector. In some embodiments, the viral vector is adeno-associated virus (AAV), lentivirus, adenovirus, simian virus 40, vaccinia virus, measles virus, herpesvirus, or poxvirus. In some embodiments, the viral vector is AAV. In some embodiments, the AAV comprises a capsid protein derived from an AAV serotype selected from the group consisting of AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAVrh10, and AAVrh74. In some embodiments, the AAV is pseudotype AAV.
[0086] In certain embodiments of this disclosure, a pharmaceutical composition is provided comprising a nucleic acid molecule of any embodiment described herein, or a composition of any embodiment described herein, or a vector of any embodiment described herein, and a pharmaceutically acceptable carrier, diluent, or excipient.
[0087] In certain embodiments of this disclosure, a method is provided for modulating the expression of a protein, RNA, or other biomolecule in a subject requiring such modification, the method comprising administering a therapeutically effective amount of a nucleic acid molecule or composition of any embodiment described herein, or a vector or pharmaceutical composition of any embodiment described herein, to the subject.
[0088] In some embodiments, the method further includes administering a therapeutically effective dose of a splice modifier to the target.
[0089] In some embodiments, the protein is expressed in the presence of a splice regulator.
[0090] In some embodiments, the method involves administering a therapeutically effective dose of a splice modifier to induce the inclusion of one of two or more exons, one or more exons, a first exon, or a second exon. In some embodiments, one of the two or more exons, or one of the one or more exons, is a second exon.
[0091] In some embodiments, the splice regulator binds to a segment of the RNA-binding protein and / or splice regulator binding site of any nucleic acid of the embodiments described herein. [Brief explanation of the drawing]
[0092] [Figure 1] Figure 1A is a schematic diagram of a splicing event that may occur in an exemplary nucleic acid molecule of this disclosure. The nucleic acid molecule may contain a start codon (e.g., ATG or AUG) having at least one nucleotide located within an exon, which is adjacent to 5' and 3' splice sites located within adjacent introns, respectively. Such a start codon may be excised during splicing, thereby preventing the downstream transgene from being translated due to the absence of the start codon ("off"). Figure 1B shows the same exemplary nucleic acid molecule, except that a splice regulator (e.g., a small molecule) binds to an exon containing a splice regulator binding site. Binding of the splice regulator to the splice regulator binding site initiates an alternative splicing event in the exemplary nucleic acid molecule, thereby incorporating the start codon into the 5' end of the transgene and translating the transgene ("on"). [Figure 2]Figure 2 is a schematic diagram of an exemplary nucleic acid molecule (e.g., a DNA molecule) of the present disclosure, which includes a minigene linked to a transgene, the minigene comprising a first exon, a first intron, a second exon, a start codon, a second intron, and a polyadenylation signal (poly-A). In the original nucleic acid molecule, the three codons of the start codon are located in the second exon. In the exemplary nucleic acid molecule of the present disclosure (e.g., "start"), at least one nucleotide of the start codon is located in the exon, and at least one nucleotide of the start codon is located in the transgene. In some embodiments, one nucleotide of the start codon is located in the second exon, and two nucleotides of the start codon are located in the transgene. Alternatively, for example, in some embodiments, two nucleotides of the start codon are located in the second exon, and one nucleotide of the start codon is located in the transgene. [Figure 3] Figure 3 is a schematic diagram of an adeno-associated virus (AAV) encoding a nucleic acid molecule (e.g., a DNA molecule) containing a minigene linked to a transgene, the minigene containing a first exon, a first intron, i) at least one nucleotide of a start codon, and ii) a splice regulator binding site adjacent to an inverted terminal repeat (ITR), a second intron, and a second exon containing poly(A). As shown in Figure 1B, in the presence of a splice regulator (e.g., a small molecule) that binds to the splice regulator binding site located in the second exon, the splicing event occurs such that the start codon (e.g., a start codon having at least one nucleotide of a start codon located in the second exon and at least one nucleotide of a start codon located in the transgene) is incorporated into the 5' end of the transgene and the transgene is translated. [Figure 4-1]Figures 4A–4E are schematic diagrams of several exemplary mechanisms by which nucleic acid molecules of this disclosure may perform exon inclusion by a splicing event. All panels provide nucleic acid molecules containing a two-site start codon (e.g., ATG or AUG) having at least one nucleotide located in an exon adjacent to the 5' and 3' splice sites located in adjacent introns (e.g., the second exon in the schematic diagram). In all panels, the third illustrated exon may be a transgene, which is not translated in the absence of the reconstituted two-site start codon. Figure 4A shows a nucleic acid molecule in which the second exon is excised during splicing unless a splice regulator (e.g., a small molecule) binds to and stabilizes the attachment of the spliceosome to the exon-intron junction of the second exon and the second intron (left), so that an alternative splicing event is initiated, the start codon is rearranged at the 5' end of the third illustrated exon (e.g., a transgene), and the third exon is translated. Figure 4B shows a nucleic acid molecule in which a second exon is excised by a spliceosome in the presence of a splice repressor, which suppresses splicing at the exon-intron junction of the second exon and second intron (left). The splice regulatory factor (e.g., a small molecule) binds to a segment of the nucleic acid molecule, blocking or destabilizing the binding of the repressor, thereby initiating an alternative splicing event, in which the start codon is reassembled at the 5' end of the third illustrated exon (e.g., a transgene), and the third exon is translated (right). Figure 4C depicts a nucleic acid molecule in which the second exon is excised by the spliceosome in the absence of a splice enhancer, supporting splicing at the exon-intron junction (left) of the second exon and second intron, where a splice regulator (e.g., a small molecule) binds to a segment of the nucleic acid molecule, stabilizing enhancer binding, thereby initiating an alternative splicing event, where the start codon is reassembled at the 5' end of the third illustrated exon (e.g., a transgene) and the third exon is translated (right).Figure 4D shows a nucleic acid molecule (left) in which the second exon is excised during splicing unless a splice regulator (e.g., a small molecule) binds to and stabilizes a splice regulator binding site (e.g., a riboswitch, e.g., an aptamer). Such an aptamer stabilizes the spliceosome binding to the exon-intron junction of the second exon and second intron so that an alternative splicing event is initiated, the start codon is rearranged at the 5' end of the third illustrated exon (e.g., a transgene), and the third exon is translated (right). Figure 4E shows a nucleic acid molecule (left) in which the second exon is excised during splicing due to the presence of a splice repressor that suppresses splicing at the exon-intron junction of the second exon and second intron unless a splice regulator (e.g., a small molecule) binds to and stabilizes a splice regulator binding site (e.g., a riboswitch, e.g., an aptamer). These aptamers block or destabilize the spliceosome's binding to the splicing repressor so that an alternative splicing event is initiated, the start codon is rearranged at the 5' end of the third illustrated exon (e.g., the transgene), and the third exon is translated (right). [Figure 4-2] Same as above. [Figure 5] Figure 5 is a schematic diagram of an exemplary nucleic acid molecule (e.g., a DNA molecule) of the present disclosure, which includes a minigene linked to a transgene, the minigene comprising a first exon, a first intron, a second exon, a start codon, a second intron, and a polyadenylation signal (poly-A). In the original nucleic acid molecule, the three codons of the start codon are located in the second exon. In the exemplary nucleic acid molecule of the present disclosure, at least one nucleotide of the start codon is located in the first exon, and at least one nucleotide of the start codon is located in the second exon. [Figure 6-1]Figures 6A–6K are schematic diagrams of several exemplary mechanisms by which a nucleic acid molecule of the present disclosure (e.g., comprising one or more exons and / or one or more introns) may undergo exon inclusion by a splicing event. Figure 6A is an exemplary nucleic acid molecule of the present disclosure in which at least one nucleotide of the start codon is located in a first exon and at least one nucleotide of the start codon is located in a transgene. Figure 6B is an exemplary nucleic acid molecule of the present disclosure in which at least one nucleotide of the start codon is located in a first exon and at least one nucleotide of the start codon is located in a second exon. Figure 6C is an exemplary nucleic acid molecule of the present disclosure in which at least one nucleotide of the start codon is located in a first exon and at least one nucleotide of the start codon is located in a second exon and the stop codon is located in a second exon upstream of the at least one nucleotide of the start codon located in the second exon. Figure 6D is an exemplary nucleic acid molecule of the present disclosure, in which at least one nucleotide of the start codon is located in a second exon, the stop codon is located in a second exon downstream of at least one nucleotide of the start codon in the second exon, and at least one nucleotide of the start codon is located in the transgene. Figure 6E is an exemplary nucleic acid molecule of the present disclosure, in which the stop codon is located in a second exon upstream of at least one nucleotide of the start codon located in the second exon, and at least one nucleotide of the start codon is located in the transgene. Figure 6F is an exemplary nucleic acid molecule of the present disclosure, in which at least one nucleotide of the start codon is located in a second exon, and at least one nucleotide of the start codon is located in the transgene. Figure 6G is an exemplary nucleic acid molecule of the present disclosure, in which at least one nucleotide of the start codon is located at the 3' end of a second exon, and at least one nucleotide of the start codon is located in the transgene. Figure 6H shows an exemplary nucleic acid molecule of the present disclosure, in which at least one nucleotide of the start codon is located in the first exon and at least one nucleotide of the start codon is located in the transgene.Figure 6I is an exemplary nucleic acid molecule of the present disclosure, in which at least one nucleotide of the start codon is located in the first exon, at least one nucleotide of the start codon is located in the transgene, and the stop codon is located in the transgene upstream of at least one nucleotide of the start codon located within the transgene. Figure 6J is an exemplary nucleic acid molecule of the present disclosure, in which the stop codon is located in the second exon, at least one nucleotide of the start codon is located in the third exon, and at least one nucleotide of the start codon is located in the transgene. Figure 6K is an exemplary nucleic acid molecule of the present disclosure, in which at least one nucleotide of the start codon is located in the second exon upstream of the stop codon in the third exon, and at least one nucleotide of the start codon is located in the transgene. [Figure 6-2] Same as above. [Figure 7] Figures 7A–7H are schematic diagrams of several exemplary mechanisms by which nucleic acid molecules of the present disclosure may perform exon inclusion by a splicing event. Figure 7A is an exemplary nucleic acid molecule of the present disclosure in which the start codon is located in a single intron and in a first exon upstream of the transgene. Figure 7B is an exemplary nucleic acid molecule of the present disclosure in which the start codon is located in a first exon upstream of the intron and the stop codon is located in the transgene. Figure 7C is an exemplary nucleic acid molecule of the present disclosure in which the start codon is located at the 3' end of a second exon upstream of the intron and the transgene. Figure 7D is an exemplary nucleic acid molecule of the present disclosure in which the stop codon is located in a second exon and the start codon is located in a third exon upstream of the transgene. Figure 7E is an exemplary nucleic acid molecule of the present disclosure in which the start codon is located in a second exon and the stop codon is located in a third exon upstream of the transgene. Figure 7F shows an exemplary nucleic acid molecule of the present disclosure, where the stop codon is located in the second exon 5' of the start codon in the same exon upstream of the transgene. Figure 7G shows an exemplary nucleic acid molecule of the present disclosure, where the start codon is located in the second exon 5' of the stop codon in the same exon upstream of the transgene. Figure 7H shows an exemplary nucleic acid molecule of the present disclosure, where the start codon is located in the second exon 5' of the upstream of the transgene. [Figure 8A] Figure 8A shows the scheme of the splicing assay used to optimize the switch sequence. [Figure 8B] Figure 8B is a graph of luciferase signaling from HEK-293T cells transiently transfected with an expression plasmid containing either RS1-10, which controls the firefly luciferase gene, or a variant switch identified from a switch sequence screening containing a 28-nucleotide intron deletion. Transfected cells were treated with 85 nM 24A for 24 hours. Each bar represents the mean ± standard deviation. [Figure 9A] Figure 9A is a graph of screened variants for the induction factor of the luciferase signal of RS1-1 relative to the vehicle, and for the induction of the luminescence signal using either the DMSO vehicle or a specified dose of 1A. Each bar represents the mean of >3 wells normalized relative to the vehicle control. [Figure 9B] Figure 9B is a graph of the luciferase signal induction ratios relative to the vehicle for RS1-1 variants screened for induction of luminescence signal using either a DMSO vehicle or a specified dose of 1A. Each bar represents the mean of >3 wells normalized to the vehicle control. [Figure 9C] Figure 9C is a bar graph showing the induction factor of the luciferase signal of RS1-10 against the vehicle using 100 nM or 1,000 nM indicated compounds or DMSO. [Figure 9D] Figure 9D is a graph of the induction ratio of the luciferase signal of RS1-10 to the vehicle in response to a specified concentration of 22A. [Figure 9E] Figure 9E is a graph of the induction ratio of the luciferase signal to the vehicle of RS1-10 in response to a specified concentration of 24A. [Figure 9F] Figure 9F is a graph of the induction ratio of the luciferase signal to the RS1-10 vehicle in response to a specified concentration of 34A. [Figure 9G]Figure 9G is a graph showing the proportion of luciferase signaling in Fln-In 293 cell lines containing a single copy insertion of the RS1-10 sequence fused to the luciferase gene in response to treatment at a specific concentration of 24A, compared to luciferase signaling in Fln-In 293 cell lines constitutively expressing luciferase. [Figure 10A] Figure 10A is a graph of the induction factor of luciferase signaling to the vehicle of RS1-10 in response to a specified concentration of 24A in HEK-293T cells. [Figure 10B] Figure 10B is a graph of the induction ratio of luciferase signaling of RS1-10 to the vehicle in response to a specified concentration of 24A in NIH-3T3 cells. [Figure 10C] Figure 10C is a graph of the induction multipliers of luciferase signaling to the vehicle in response to 1A in SH-SY5Y cells having either a CBA promoter or a human synapsin promoter. [Figure 10D] Figure 10D is a graph of the induction ratio of luciferase signaling to the vehicle in response to 1A in HepG2 cells with a CBA promoter. [Figure 11A] Figure 11A shows a scheme of RS1-10 injected into C57BL / 6 mice and the resulting vector containing the AAV-PHP.eB capsid. [Figure 11B] Figure 11B is an image of a C57BL / 6 mouse 6 hours after administration of vehicle or 24A. [Figure 11C] Figure 11C shows an image of a C57BL / 6 mouse 7 days after administration of vehicle or 24A. [Figure 12] Figure 12 is a graph showing the quantification of radiance in the head region of mice 6 hours after administration of each treatment. [Figure 13A]Figure 13A is a graph of the luciferase signal induction ratios for screened variants for induction of luminescence signals using either the RS2-1 vehicle and either the DMSO vehicle or 100B. Each bar represents the mean ± standard deviation of the 3 wells normalized relative to the vehicle control. [Figure 13B] Figure 13B is a graph of the luciferase signal induction ratios for additional variants screened for the induction of luminescence signals using either the RS2-1 vehicle or either the DMSO vehicle or 100B. Each bar represents the mean ± standard deviation of the 3 wells normalized relative to the vehicle control. [Figure 13C] Figure 13C is a graph of the induction factor of the luciferase signal of RS2-3 to the vehicle using 100 nM or 1,000 nM indicated compounds or DMSO. [Figure 14A] Figure 14A is a graph of the luciferase signal induction ratios for variants screened for induction of luminescence signals using either the RS3-1 vehicle or either the DMSO vehicle or 116B. Each bar represents the mean ± standard deviation of the 3 wells normalized to the vehicle control. [Figure 14B] Figure 14B is a graph of the luciferase signal induction ratio relative to the vehicle for RS3-1 variants screened for luminescence signal induction using either a DMSO vehicle or 116B. Each bar represents the mean ± standard deviation of the 3 wells normalized relative to the vehicle control. [Figure 14C] Figure 14C is a graph of the luciferase signal induction ratios relative to the vehicle for RS3-1 variants screened for luminescence signal induction using either a DMSO vehicle or a specified concentration of 116B. Each bar represents the mean ± standard deviation of the 3 wells normalized relative to the vehicle control. [Figure 14D] Figure 14D is a graph of the induction factor of the luciferase signal of RS3-13 to the vehicle using 100 nM or 1,000 nM indicated compounds or DMSO. [Figure 15A] Figure 15A is a graph of the luciferase signal induction ratios for screened variants for induction of luminescence signals using either the RS4-1 vehicle or either the DMSO vehicle or 24A. Each bar represents the mean ± standard deviation of the 3 wells normalized to the vehicle control. [Figure 15B] Figure 15B is a graph of the luciferase signal induction ratios for screened variants for induction of luminescence signals using either the RS4-1 vehicle and either the DMSO vehicle or 116B. Each bar represents the mean ± standard deviation of the 3 wells normalized relative to the vehicle control. [Modes for carrying out the invention]
[0093] Nucleic acid molecules of the present disclosure This disclosure provides a highly efficient means of regulating transgene expression, based at least in part on the surprising discovery that splicing events regulated by exogenous splice regulators can influence the inclusion of a two-site start codon (e.g., a start codon containing at least one nucleotide in an exon and at least one nucleotide in the transgene). This disclosure is also based at least in part on surprising discoveries of nucleic acid molecular design that improve upon the idea of regulating transgene expression using splicing. For example, by splitting the start codon between exons (e.g., molecular design shown in Figure 2, see Start), engagement of the translational mechanism on the transcript can be avoided, as long as the rearrangement of the start codon does not generate an on state when a splice regulator is present. This design is expected to reduce background expression and prevent the production of cleaved proteins. In embodiments of this disclosure, the nucleic acid molecules herein generally function by an exon inclusion mechanism. In some embodiments, the splice regulator binds to an RNA-binding protein (RBP) that intrinsically promotes the splicing event and stimulates its function. In some embodiments, the splice regulators of the Disclosure facilitate the binding of RBP to nucleic acid molecules, thereby enabling splicing to occur at the junction between a second exon and a second intron, thereby including the exon. In some embodiments, the splice regulators bind to a riboswitch so that splicing occurs at the junction between a second exon and a second intron, and the exon is included. In both embodiments, and in all embodiments of the Disclosure, the nucleic acid molecules of the Disclosure function by an exon inclusion mechanism. Five exemplary mechanisms of action by which the nucleic acid molecules of the Disclosure perform exon inclusion by a splicing event are shown in Figures 4A–4E. In some embodiments, the compositions and methods described herein are used to regulate the expression of proteins, functional RNAs, regulatory RNAs, microRNAs (miRNAs), and small hairpin RNAs (shRNAs) (e.g., proteins, miRNAs, or shRNAs encoded by transgenes) in subjects requiring such regulation.
[0094] In all embodiments, the nucleic acid molecules of this disclosure function by an exon inclusion mechanism. In some embodiments, a splice regulator binds to and stimulates the function of RBP, which intrinsically promotes the splicing event. In such embodiments, the splice regulator of this disclosure facilitates the binding of RBP to the spliceosome, thereby enabling splicing to occur at the junction between a second exon and a second intron, resulting in exon inclusion. In another embodiment of this disclosure, the splice regulator binds to a riboswitch so that splicing occurs at the junction between a second exon and a second intron, and exon inclusion occurs. In both embodiments, and in all embodiments of this disclosure, the nucleic acid molecules of this specification function by an exon inclusion mechanism.
[0095] In some embodiments, the second exon of a nucleic acid molecule is excised during splicing (Figure 4A, left) unless a price regulator (e.g., a small molecule) binds to and stabilizes the second exon and second intron of the spliceosome at the exon-intron junction, initiating an alternative splicing event, where the start codon is reassembled at the 5' end of a third illustrated exon (e.g., a transgene), and the third exon is translated (Figure 4A, right). In some embodiments, the second exon of a nucleic acid molecule is excised by the spliceosome due to the presence of a splicing repressor that inhibits splicing at the exon-intron junction of the second exon and second intron (as shown in Figure 4B, left). Unless a splice regulator (e.g., a small molecule) binds to a segment of the nucleic acid molecule and blocks or destabilizes the binding of the repressor, an alternative splicing event is initiated, the start codon is reassembled at the 5' end of a third illustrated exon (e.g., a transgene), and the third exon is translated (Figure 4B, right). In some embodiments, the second exon of a nucleic acid molecule is excised by the spliceosome due to the absence of a splice enhancer supporting splicing at the exon-intron junction of the second exon and second intron (as shown in Figure 4C, left). Unless a splice regulator (e.g., a small molecule) binds to a segment of the nucleic acid molecule and stabilizes the binding of the enhancer, an alternative splicing event is initiated, the start codon is reassembled at the 5' end of a third illustrated exon (e.g., a transgene), and the third exon is translated (Figure 4C, right). In some embodiments, the second exon of a nucleic acid molecule is excised during splicing unless a splice regulator (e.g., a small molecule) binds to and stabilizes a splice regulator binding site (e.g., a riboswitch, e.g., an aptamer) (Figure 4D, left). These aptamers stabilize the spliceosome's attachment to the exon-intron junction of the second exon and second intron so that an alternative splicing event is initiated, the start codon is rearranged at the 5' end of the third illustrated exon (e.g., the transgene), and the third exon is translated (Figure 4D, right) (Figure 4D).In some embodiments, the second exon of a nucleic acid molecule is excised during splicing due to the presence of a splice repressor that inhibits splicing at the exon-intron junction of the second exon and second intron unless a splice regulator (e.g., a small molecule) binds to and stabilizes the splice regulator binding site (e.g., a riboswitch, e.g., an aptamer) (Figure 4E, left). These aptamers block or destabilize the binding of the spliceosome to the splicing repressor so that an alternative splicing event is initiated, the start codon is rearranged at the 5' end of a third illustrated exon (e.g., a transgene), and the third exon is translated (Figure 4E, right).
[0096] In a particular embodiment of this disclosure, the nucleic acid molecule of the disclosure is a nucleic acid molecule comprising a minigene located adjacent to the 5' side of a transgene, wherein the minigene comprises (a) a first exon adjacent to the 5' side of a first intron, (b) a start codon, and (c) a splice regulator binding site, the splice regulator binding site comprising the nucleic acid sequence DGAGTDDGHV (SEQ ID NO: 82) or DGAGTDDNHV (SEQ ID NO: 83), where D is A, G, or T, R is A, G, N is A, C, G, or T, H is A, C, or T, and V is A, G, or C. In some embodiments, the splice regulator binding site comprises the nucleic acid sequence DGAGTTTGHV (SEQ ID NO: 84), where D is A, G, or T, H is A, C, or T, and V is A, G, or C. Furthermore, in certain embodiments of this disclosure, the nucleic acid molecule of this disclosure is a nucleic acid molecule comprising a minigene located adjacent to the 5' side of a transgene, wherein the minigene comprises (a) a first exon located at the 5' side of a first intron, (b) a start codon comprising a first portion and a second portion, wherein the first and second portions of the start codon are not in the same exon, and (c) a splice regulator binding site, wherein in some embodiments, the splice regulator binding site comprises the nucleic acid sequence DGAGTRRGHV(SEQ ID NO: 1) or DGAGTRRNHV(SEQ ID NO: 2), where D is adenine (A), guanine (G), or thymidine (T), N is A, C, G, or T, R is A or G, H is A, C, or T, and V is A, G, or C. See Figures 6A-6K.
[0097] In some embodiments, the first portion of the start codon contains one or two nucleotides of the start codon. In some embodiments, the second portion of the start codon has one or two nucleotides of the start codon. In some embodiments, the first portion of the start codon is located in the first exon, and the second portion of the start codon is located in the transgene.
[0098] In some embodiments, the nucleic acid includes a first exon containing the first portion of the start codon, a first intron, and a transgene containing the second portion of the start codon, in 5' to 3' order. See Figure 6A.
[0099] In some embodiments, the nucleic acid molecule further comprises a second exon. In some embodiments, the nucleic acid molecule further comprises a second intron.
[0100] In some embodiments, the nucleic acid includes a first exon containing the first portion of the start codon, a second exon containing the second portion of the start codon, and the transgene in 5' to 3' order. See Figure 6B.
[0101] In some embodiments, the nucleic acid includes, in 5' to 3' order, a first exon containing the first portion of the start codon, a first intron, a second exon containing the second portion of the start codon, a second intron, and a transgene. See Figure 6B.
[0102] In some embodiments, the nucleic acid includes a transgene comprising a first exon in the order of 5' to 3', a second exon containing the first portion of the start codon, and a second portion of the start codon. See Figure 6G.
[0103] In some embodiments, the nucleic acid includes a transgene comprising, in 5' to 3' order, a first exon, a second exon containing the first portion of the start codon, a first intron, and a second portion of the start codon. See Figure 6G.
[0104] In some embodiments, the nucleic acid includes a first exon containing the first portion of the start codon, a second exon, and a transgene containing the second portion of the start codon, in the order of 5' to 3'. See Figure 6H.
[0105] In some embodiments, the nucleic acid includes a first exon containing the first portion of the start codon, a second exon, a first intron, and a transgene containing the second portion of the start codon, in the order of 5' to 3'. See Figure 6H.
[0106] In some embodiments, the nucleic acid includes a first exon containing the first portion of the start codon, a second exon containing the second portion of the start codon, and the transgene in 5' to 3' order. See Figure 6B.
[0107] In some embodiments, the nucleic acid comprises, in 5' to 3' order, a first exon containing the first portion of the start codon, a first intron, a second exon containing the second portion of the start codon, and a transgene. See Figure 6B.
[0108] In some embodiments, the nucleic acid molecule further includes a stop codon.
[0109] In some embodiments, the nucleic acid includes a first exon containing the first portion of the start codon, a second exon containing the stop codon, and a transgene containing the second portion of the start codon, in the order of 5' to 3'. See Figure 6I.
[0110] In some embodiments, the nucleic acid includes, in 5' to 3' order, a first exon containing the first portion of the start codon, a first intron, a second exon containing the stop codon, and a transgene containing the second portion of the start codon. See Figure 6I.
[0111] In some embodiments, the second exon includes a splice regulator binding site.
[0112] In some embodiments, the nucleic acid molecule further comprises a third exon. In some embodiments, the nucleic acid molecule further comprises a third intron.
[0113] In some embodiments, the nucleic acid includes, in 5' to 3' order, a first exon containing the first portion of the start codon, a second exon containing the stop codon, a third exon containing the second portion of the start codon, and the transgene. See Figure 6C.
[0114] In some embodiments, the nucleic acid comprises, in 5' to 3' order, a first exon containing the first portion of the start codon, a first intron, a second exon containing the stop codon, a third exon containing the second portion of the start codon, a second intron, and a transgene. See Figure 6C.
[0115] In some embodiments, the nucleic acid includes a transgene containing, in the order of 5' to 3', a first exon, a second exon containing the first portion of the start codon, a third exon containing the stop codon, and a second portion of the start codon. See Figure 6D.
[0116] In some embodiments, the nucleic acid includes a transgene comprising, in 5' to 3' order, a first exon, a first intron, a second exon containing the first portion of the start codon, a third exon containing the stop codon, a second intron, and a second portion of the start codon. See Figure 6D.
[0117] In some embodiments, the nucleic acid includes a transgene containing, in the order of 5' to 3', a first exon, a second exon containing a stop codon, a third exon containing the first portion of the start codon, and a second portion of the start codon. See Figure 6E.
[0118] In some embodiments, the nucleic acid includes a transgene comprising, in 5' to 3' order, a first exon, a first intron, a second exon containing a stop codon, a third exon containing the first portion of a start codon, a second intron, and a second portion of a start codon. See Figure 6D.
[0119] In some embodiments, the nucleic acid includes a transgene comprising, in 5' to 3' order, a first exon, a second exon, a third exon containing the first portion of the start codon, and a second portion of the start codon. See Figure 6F.
[0120] In some embodiments, the nucleic acid includes a transgene comprising, in 5' to 3' order, a first exon, a first intron, a second exon, a third exon containing the first portion of the start codon, a second intron, and the second portion of the start codon. See Figure 6F.
[0121] In some embodiments, the nucleic acid includes a transgene comprising a first exon in the order of 5' to 3', a second exon containing a stop codon, a third exon containing the first portion of the start codon, and a transgene containing the second portion of the start codon. See Figure 6J.
[0122] In some embodiments, the nucleic acid includes a transgene comprising, in 5' to 3' order, a first exon, a first intron, a second exon containing a stop codon, a second intron, a third exon containing the first portion of a start codon, a third intron, and a second portion of a start codon. See Figure 6J.
[0123] In some embodiments, the nucleic acid includes a transgene containing, in 5' to 3' order, a first exon, a second exon containing the first portion of the start codon, a third exon containing the stop codon, and a second portion of the start codon. See Figure 6K.
[0124] In some embodiments, the nucleic acid includes a transgene comprising, in 5' to 3' order, a first exon, a first intron, a second exon containing the first portion of the start codon, a second intron, a third exon containing the stop codon, a third intron, and a second portion of the start codon. See Figure 6K.
[0125] In some embodiments, the first intron includes a splice regulator binding site.
[0126] In some embodiments, the splice regulator binding site is located at the junction between the second exon and the second intron.
[0127] In some embodiments, the splice regulator binding site is located within a second intron.
[0128] In this disclosure, in certain embodiments, a nucleic acid molecule is described comprising a minigene located adjacent to the 5' side of a transgene, wherein the minigene comprises (a) a first exon adjacent to the 5' side of a first intron, (b) a start codon, and (c) a splice regulator binding site, the splice regulator binding site comprising the nucleic acid sequence DGAGTDDGHV (SEQ ID NO: 82) or DGAGTDDNHV (SEQ ID NO: 83), where D is A, G, or T, R is A, G, N is A, C, G, or T, H is A, C, or T, and V is A, G, or C. In some embodiments, the splice regulator binding site comprises the nucleic acid sequence DGAGTTTGHV, where D is A, G, or T, H is A, C, or T, and V is A, G, or C. In this disclosure, in certain embodiments, a nucleic acid molecule comprising a minigene located adjacent to the 5' end of a transgene is further described, wherein the minigene comprises (a) a first exon adjacent to the 5' end of a first intron, (b) a start codon, and (c) a splice regulator binding site, the splice regulator binding site comprising DGAGTRRGHV (SEQ ID NO: 1) or DGAGTRRNHV (SEQ ID NO: 2), where D is A, G, or T, R is A or G, N is A, C, G, or T, H is A, C, or T, and V is A, G, or C. See Figures 7A-7G. In some embodiments, the splice regulator binding site comprises the nucleic acid sequence DGAGTTTGHV (SEQ ID NO: 84), where D is A, G, or T, H is A, C, or T, and V is A, G, or C. In some embodiments, the splice regulator binding site includes TAGAGTAAGACA (SEQ ID NO: 75). In some embodiments, the splice regulator binding site includes ATGAGTATGACA (SEQ ID NO: 76). In some embodiments, the splice regulator binding site includes ATGAGTAAGCAG (SEQ ID NO: 77). In some embodiments, the splice regulator binding site includes ATGAGTATGT (SEQ ID NO: 78).In some embodiments, the splice regulator binding site includes ATGAGTTTGT (SEQ ID NO: 79). In some embodiments, the splice regulator binding site includes ATGAGTAAGT (SEQ ID NO: 80). In some embodiments, the splice regulator binding site includes ATGAGTTAGT (SEQ ID NO: 81).
[0129] In some embodiments, the nucleic acid includes a first exon containing the start codon, a first intron, and a transgene, in the order of 5' to 3'. See Figure 7A.
[0130] In some embodiments, the nucleic acid molecule further includes a stop codon.
[0131] In some embodiments, the nucleic acid molecule further comprises a second exon.
[0132] In some embodiments, the nucleic acid includes a first exon containing a start codon, a second exon containing a stop codon, and a transgene, in the order of 5' to 3'. See Figure 7B.
[0133] In some embodiments, the nucleic acid includes, in 5' to 3' order, a first intron, a first exon containing a start codon, a second exon containing a stop codon, and a transgene. See Figure 7B.
[0134] In some embodiments, the nucleic acid includes, in the order of 5' to 3', a first exon, a second exon containing the start codon, and a transgene. See Figure 7C.
[0135] In some embodiments, the nucleic acid includes, in 5' to 3' order, a first exon, a second exon containing the start codon, a first intron, and a transgene. See Figure 7C.
[0136] In some embodiments, the nucleic acid molecule further comprises a third exon.
[0137] In some embodiments, the nucleic acid molecule further includes a third intron.
[0138] In some embodiments, the nucleic acid includes, in the order of 5' to 3', a first exon, a second exon containing a stop codon, a third exon containing a start codon, and a transgene. See Figure 7D.
[0139] In some embodiments, the nucleic acid includes, in 5' to 3' order, a first exon, a first intron, a second exon containing a stop codon, a second intron, a third exon containing a start codon, a third intron, and a transgene. See Figure 7D.
[0140] In some embodiments, the nucleic acid includes, in 5' to 3' order, a first exon, a first intron, a second exon containing a stop codon, a third exon containing a start codon, a second intron, and a transgene. See Figure 7F.
[0141] In some embodiments, the nucleic acid includes, in 5' to 3' order, a first exon, a second exon containing a start codon, a third exon containing a stop codon, and a transgene. See Figures 7E and 7G.
[0142] In some embodiments, the nucleic acid includes, in 5' to 3' order, a first exon, a first intron, a second exon containing a start codon, a second intron, a third exon containing a stop codon, a third intron, and a transgene. See Figure 7E.
[0143] In some embodiments, the nucleic acid includes, in 5' to 3' order, a first exon, a first intron, a second exon containing a start codon, a third exon containing a stop codon, a second intron, and a transgene. See Figure 7G.
[0144] In some embodiments, the nucleic acid includes, in 5' to 3' order, a first exon, a first intron, a second exon containing a start codon, a second intron, and a transgene. See Figure 7H.
[0145] In some embodiments, the nucleic acid includes a sequence that is at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to any one of SEQ ID NOs: 21-73. In some embodiments, the nucleic acid includes a sequence that is at least about 70% identical to any one of SEQ ID NOs: 21-73. In some embodiments, the nucleic acid includes a sequence that is at least about 75% identical to any one of SEQ ID NOs: 21-73. In some embodiments, the nucleic acid includes a sequence that is at least about 80% identical to any one of SEQ ID NOs: 21-73. In some embodiments, the nucleic acid includes a sequence that is at least about 85% identical to any one of SEQ ID NOs: 21-73. In some embodiments, the nucleic acid includes a sequence having at least about 90% identity with any one of sequence numbers 21 to 73. In some embodiments, the nucleic acid includes a sequence having at least about 95% identity with any one of sequence numbers 21 to 73. In some embodiments, the nucleic acid includes a sequence having at least about 96% identity with any one of sequence numbers 21 to 73. In some embodiments, the nucleic acid includes a sequence having at least about 97% identity with any one of sequence numbers 21 to 73. In some embodiments, the nucleic acid includes a sequence having at least about 98% identity with any one of sequence numbers 21 to 73. In some embodiments, the nucleic acid includes a sequence having at least about 99% identity with any one of sequence numbers 21 to 73. In some embodiments, the nucleic acid includes a sequence having 100% identity with any one of sequence numbers 21 to 73. In some embodiments, sequence numbers 21 to 73 include a first exon, a first intron, a second exon containing a start codon, and a second intron. In some embodiments, the designs of sequence numbers 21-73 are reflected in Figure 7H.
[0146] In certain embodiments of this disclosure, a nucleic acid molecule is described comprising, in the order from 5' to 3', (a) a first exon containing a first portion of the start codon at the 3' end of the first exon, (b) a second exon containing a second portion of the start codon at the 5' end of the second exon, and (c) a third exon. See Figure 6B.
[0147] In some embodiments, the third exon is a transgene.
[0148] In this disclosure, in certain embodiments, a nucleic acid molecule is described that includes a transgene comprising a first exon, a second exon containing a stop codon, a third exon containing a first portion of a start codon, and a second portion of a start codon, in the order of 5' to 3'. See Figure 6J.
[0149] In this disclosure, in certain embodiments, a nucleic acid molecule comprising a transgene is described, comprising a first exon, a second exon containing a first portion of the start codon, a third exon containing a stop codon, and a second portion of the start codon, in the order of 5' to 3'. See Figure 6K.
[0150] In some embodiments, the nucleic acid molecule further comprises a first exon, a second exon, a third exon, and a first intron, a second intron, and a third intron interposed between the transgene.
[0151] In some embodiments, the splice regulatory factor binding site does not contain the nucleic acid sequence AAGAGT (SEQ ID NO: 3), ATGAGT (SEQ ID NO: 4), TAGAGT (SEQ ID NO: 5), TTGAGT (SEQ ID NO: 6), GAGAGT (SEQ ID NO: 7), GTGAGT (SEQ ID NO: 8), AAGAGT (SEQ ID NO: 9), ATGAGT (SEQ ID NO: 10), ACGAGT (SEQ ID NO: 11), AGGAGT (SEQ ID NO: 12), AGAGGTAGAG (SEQ ID NO: 13), TGAGGTTGAG (SEQ ID NO: 14), GGAGGTGGAG (SEQ ID NO: 15), TAG (SEQ ID NO: 16), CAG (SEQ ID NO: 17), TAG (SEQ ID NO: 18), or NAGAGTNNNN (SEQ ID NO: 19), (wherein N is A, C, G, or T).
[0152] In certain embodiments, nucleic acid molecules comprising a minigene ligated to a transgene are described herein, where the minigene is designed to regulate the transcription of the transgene via an exon inclusion mechanism. In some embodiments, the minigene comprises (a) two or more introns (e.g., about 3, about 4, or about 5) and (b) two or more exons (e.g., about 3, about 4, or about 5). In some embodiments, the minigene comprises (a) two introns and (b) two exons. In some embodiments, one or two nucleotides at the 5' end of a start codon are located in an exon positioned on the 3' side of the first intron of the minigene. In some embodiments, one or two nucleotides at the 3' end of a start codon are located within the transgene. In some embodiments, the nucleic acid molecules of this disclosure do not contain a two-site stop codon.
[0153] In some embodiments, the minigene includes (a) a first intron and a second intron, and (b) a first exon and a second exon.
[0154] In certain embodiments, nucleic acid molecules comprising minigenes designed to regulate the transcription of a transgene via an exon exclusion mechanism are described herein. The minigene comprises a first exon, a first intron, one or more (e.g., about 2, about 3, about 4, or about 5) exons located at the 3' end of the first intron, a start codon, and a second intron. In some embodiments, one or two nucleotides at the 5' end of the start codon are located in any of the one or more exons (e.g., multiple exons located at the 3' end of the first intron). In some embodiments, one or two nucleotides at the 3' end of the start codon are located within the transgene. In some embodiments, the nucleic acid molecules of the present disclosure function by an exon inclusion mechanism. In some embodiments, the nucleic acid molecules of the present disclosure do not contain a two-site stop codon.
[0155] In some embodiments, one or more exons located on the 3' side of the first exon include a second exon.
[0156] In some embodiments, one or two nucleotides at the 5' end of the start codon are located in a second exon.
[0157] In this disclosure, in certain embodiments, nucleic acid molecules comprising a minigene designed to regulate the transcription of a transgene via an exon inclusion mechanism are described. The minigene comprises a first exon, a first intron, a second exon, a start codon, and a second intron, and is ligated to a transgene in which one or two nucleotides at the 5' end of the start codon are located in the second exon and one or two nucleotides at the 3' end of the start codon are located in the transgene.
[0158] In some embodiments, the two nucleotides at the 5' end of the start codon are located in a second exon, and the one nucleotide at the 3' end of the start codon is located in the transgene. In some embodiments, the nucleic acid molecule of the disclosure functions by an exon inclusion mechanism. In some embodiments, the nucleic acid molecule of the disclosure does not contain a two-site stop codon.
[0159] In some embodiments, the start codon is discontinuous (e.g., split). Such splitting can occur across exons, introns, and / or transgenes. For example, in some embodiments, at least one (e.g., one or two) nucleotides of the start codon are located in a second exon. In some embodiments, at least one (e.g., one or two) nucleotides of the start codon are located within the transgene. In some embodiments, at least one (e.g., one or two) nucleotides of the start codon are located in a second exon, and at least one (e.g., one or two) nucleotides of the start codon are located within the transgene.
[0160] In some embodiments, the nucleic acid molecule includes at least one splicing element (e.g., a 5' splice site and / or a 3' splice site). In some embodiments, the splicing element is adjacent to the start codon. In some embodiments, the splicing element is on one or both sides of the start codon.
[0161] None of the nucleic acid molecules in this disclosure contain two-site stop codons.
[0162] Start codon The nucleic acid molecules described herein encode a translation initiation sequence, for example, a start codon (START). In some embodiments, the translation initiation sequence includes a Kossack or Shine-Dalgamo sequence. In some embodiments, the translation initiation sequence includes a Kossack sequence. Further examples of translation initiation sequences are found in paragraphs of International Patent Publication WO2019 / 118919.
[0163] This is described in
[0165] .
[0163] In some embodiments, the nucleic acid molecule includes an exon and / or an expression sequence (e.g., a transgene) with a start codon located adjacent to or within it.
[0164] In some embodiments, the start codon is discontinuous (e.g., split). Such splitting can occur across exons, introns, and / or transgenes. For example, in some embodiments, at least one (e.g., one or two) nucleotides of the start codon are located in a second exon. In some embodiments, at least one (e.g., one or two) nucleotides of the start codon are located within the transgene. In some embodiments, at least one (e.g., one or two) nucleotides of the start codon are located in a second exon and at least one (e.g., one or two) nucleotides of the start codon are located in the transgene. In some embodiments, at least one nucleotide of the start codon is located in a second exon and at least one nucleotide of the start codon is located in the transgene. In some embodiments, two nucleotides of the start codon are located in a second exon and one nucleotide of the start codon is located in the transgene. In some embodiments, one nucleotide of the start codon is located in the second exon, and the two nucleotides of the start codon are located in the transgene.
[0165] In some embodiments, the start codon is a non-coding start codon.
[0166] Any suitable start codon can be used.
[0167] In some embodiments, the start codon is a 3-nucleotide codon.
[0168] Splice regulator binding site The nucleic acid molecules described herein may include a splice regulator binding site. For example, in some embodiments, such a splice regulator binding site includes the nucleic acid sequence DGAGUNNGBD(SEQ ID NO: 1) or DGAGUNNNBD(SEQ ID NO: 2), where D is adenine (A), guanine (G), or uracil (U), where N is A, cytosine (C), G, or U, and where B is A, C, or U. For example, in some embodiments, the splice regulator binding site includes the nucleic acid sequence DGAGUNNGBD(SEQ ID NO: 1). In some embodiments, the splice regulator binding site includes the nucleic acid sequence DGAGUNNNBD(SEQ ID NO: 2).
[0169] In some embodiments, the splice regulator binding site does not include the nucleic acid sequence AAGAGU (SEQ ID NO: 3), AUGAGU (SEQ ID NO: 4), UAGAGU (SEQ ID NO: 5), UUGAGU (SEQ ID NO: 6), GAGAGU (SEQ ID NO: 7), GUGAGU (SEQ ID NO: 8), ACGAGU (SEQ ID NO: 9), AUGAGU (SEQ ID NO: 10), ACGAGU (SEQ ID NO: 11), AGGAGU (SEQ ID NO: 12), AGAGGTAGAG (SEQ ID NO: 13), UGAGGTUGAG (SEQ ID NO: 14), GGAGGTGGAG (SEQ ID NO: 15), TAG (SEQ ID NO: 16), CAG (SEQ ID NO: 17), UAG (SEQ ID NO: 18), or NAGAGTNNNN (SEQ ID NO: 19), (wherein N is A, C, G, or T). For example, in some embodiments, the splice regulator binding site does not include the nucleic acid sequence AAGAGU (SEQ ID NO: 3). In some embodiments, the splice regulator binding site does not include the nucleic acid sequence AUGAGU (SEQ ID NO: 4). In some embodiments, the splice regulatory factor binding site does not include the nucleic acid sequence UAGAGU (SEQ ID NO: 5). In some embodiments, the splice regulatory factor binding site does not include the nucleic acid sequence UUGAGU (SEQ ID NO: 6). In some embodiments, the splice regulatory factor binding site does not include the nucleic acid sequence GAGAGU (SEQ ID NO: 7). In some embodiments, the splice regulatory factor binding site does not include the nucleic acid sequence GUGAGU (SEQ ID NO: 8). In some embodiments, the splice regulatory factor binding site does not include the nucleic acid sequence AAGAGU (SEQ ID NO: 9). In some embodiments, the splice regulatory factor binding site does not include the nucleic acid sequence AUGAGU (SEQ ID NO: 10). In some embodiments, the splice regulatory factor binding site does not include the nucleic acid sequence ACGAGU (SEQ ID NO: 11). In some embodiments, the splice regulatory factor binding site does not include the nucleic acid sequence AGGAGU (SEQ ID NO: 12). In some embodiments, the splice regulatory factor binding site does not include the nucleic acid sequence AGAGGTAGAG (SEQ ID NO: 13). In some embodiments, the splice regulatory factor binding site does not include the nucleic acid sequence UGAGGTUGAG (SEQ ID NO: 14). In some embodiments, the splice regulator binding site does not include the nucleic acid sequence GGAGGTGGAG (SEQ ID NO: 15).In some embodiments, the splice regulator binding site does not include the nucleic acid sequence TAG (SEQ ID NO: 16). In some embodiments, the splice regulator binding site does not include the nucleic acid sequence CAG (SEQ ID NO: 17). In some embodiments, the splice regulator binding site does not include the nucleic acid sequence UAG (SEQ ID NO: 18). In some embodiments, the splice regulator binding site does not include the nucleic acid sequence NAGAGTNNNN (SEQ ID NO: 19).
[0170] In some embodiments, the splice regulator binding site is located in the second exon of the nucleic acid molecule. Furthermore, the splice regulator may bind to the splice regulator binding site, and as a result, the binding may lead to a splicing event and / or transcription of the linked transgene.
[0171] The splice regulator binding site may consist of a set of four or more nucleotides (e.g., five, six, seven, or eight).
[0172] In some embodiments, the splice regulator binding site includes a sequence recognized by an RNA-binding protein (RBP). In some embodiments, the RBP is tissue-specific. In some embodiments, the RBP is exclusively expressed in the brain. In accordance with this principle, in some embodiments, the nucleic acid molecules of this disclosure operate by the principle of exon inclusion by using a splice regulator that targets an RBP having tissue-specific expression.
[0173] In some embodiments, RBP is a splicing enhancer (e.g., serine and arginine-rich (SR) protein) (for a review, see, e.g., Jeong S. Mol Cells. 2017 Jan;40(1):1-9). In some embodiments, RBP is a splicing repressor (e.g., heterogeneous nuclear ribonucleoprotein (hnRNP)), see, for example, Han et al. Biochem J. 2010 Sep 15;430(3):379-9, 2.
[0174] In some embodiments, the splice regulator binding site is located at the junction between the second exon and the second intron.
[0175] In some embodiments, splice regulator binding sites are designed by screening more than 100 (e.g., more than 1,000, more than 10,000, more than 100,000, or more than 1,000,000) candidate splice regulator binding sites for their ability to functionally respond to specific splice regulators. In some embodiments, the screening evaluates the inclusion of both basal regulator stimulation and splice regulator stimulation of two-site start codons, as well as the resulting increase or decrease in transgene expression for each. In some embodiments, the screening utilizes, for example, DNA constructs containing minigenes derived from naturally occurring mammalian genes, or synthetic intron-containing constructs having standard 5' and 3' splice site sequences. In some embodiments, each minigene is linked to a reporter gene, for example, firefly luciferase, and assembled via DNA synthesis and molecular cloning techniques known in the art. A library of minigene constructs having point mutations, insertions, and / or deletions of splicing regulator binding sites can be generated and introduced into mammalian cells using electroporation, chemical transfection, virus-mediated integration, or virus-mediated episomal delivery. In some embodiments, evaluation of candidate splicing regulator binding sites and locations in the minigenes is achieved by including a reporter gene in the DNA structure, such as a luminescent enzyme (e.g., firefly luciferase, nanolucinase, renyral luciferase, and Gaussian luciferase), a fluorescent protein (e.g., green fluorescent protein, blue fluorescent protein, and red fluorescent protein), or a colorimetric enzyme (e.g., beta-lactamase or secreted placental alkaline phosphatase), which generates a quantifiable signal proportional to the splicing of an exon (e.g., the second exon) and is readable by, for example, flow cytometry, microscopy, or a multimode microplate reader using a photomultiplier tube. Alternatively, for example, candidate splice regulator-dependent biological activity can be evaluated by sequencing mRNA transcripts produced by the DNA construct using RNA-Seq or quantitative reverse transcription PCR.In some embodiments, following such exemplary screening, candidate splice regulators are applied to cells containing DNA constructs for up to 6, 12, 24, or 72 hours, and then evaluated for start codon-containing activity. In some embodiments, a series of candidate splice regulator binding site design, synthesis, and assay evaluation are performed to optimize the splice regulator binding site of the transgene regulatory minigene.
[0176] In some embodiments, the splice regulator binding site is designed by screening various substances and / or chemical derivatives of substances for their ability to regulate the expression of a reporter gene linked to a minigene construct in mammalian cells. In some embodiments, the minigene template is derived from a naturally occurring mammalian gene or a synthetic intron-containing construct having standard 5' and 3' splice site sequences. In some embodiments, the minigene template is linked to a reporter gene, such as firefly luciferase, to enable the detection of changes in gene expression resulting from a chemical that acts as a splice regulator to mediate alternative splicing of the minigene.
[0177] In some embodiments, the splice regulator binding site includes an intron or exon splice silencer or enhancer element. For example, in some embodiments, the splice regulator binding site includes an intron splice silencer. In some embodiments, the splice regulator binding site includes an exon splice silencer. In some embodiments, the splice regulator binding site includes an enhancer element.
[0178] In some embodiments, the splice regulator binding site is an aptamer. In some embodiments, the aptamer is part of a riboswitch.
[0179] Riboswitch Riboswitches are regulatory segments of nucleic acid molecules that influence the expression of genetic elements. Therefore, nucleic acid molecules containing riboswitches directly participate in regulating their activity in response to the concentration of the effector molecule by forming an alternative structure in response to effector binding.
[0180] More specifically, a riboswitch consists of (i) a splice regulator binding site domain (e.g., an aptamer) that binds to a defined ligand with high affinity, and (ii) an expression platform that brings about an effect on gene expression through splice regulator binding (e.g., binding of a splice regulator to an aptamer). Regulation occurs either at the transcriptional level (e.g., by the formation of a terminator or anti-terminator structure) or at the translational level (e.g., by the presentation or sequestration of a ribosome binding site). These elements can be manipulated by combining different splice regulator binding sites (e.g., aptamers) and expression platforms through modular compositions.
[0181] Riboswitches are known to respond to RNA derivatives such as coenzymes (13 riboswitch classes), nucleotide derivatives (7 riboswitch classes), signaling molecules (5 riboswitch classes), as well as ions (5 riboswitch classes), amino acids (3 riboswitch classes), and other metabolites (5 riboswitch classes). To date, 38 different classes of riboswitches have been discovered with their homologous ligands (see, for example, McCown, Phillip J., et al. Rna 23.7. 2017. 995-1011). In particular, the more than 100,000 representative riboswitches can be classified into 38 validated riboswitch classes. Each riboswitch class is named according to its ligand. As shown in Table 1 below, these classes include thiamine pyrophosphate (TPP), adenosylcobalamin (AdoCbl), or coenzyme B 12S-adenosylmethionine (SAM), cyclic-di-GMP (C-di-GMP), glycine, flavin mononucleotide (FMN), divalent manganese (Mn 2+ ), lysine, cyclic-di-AMP (C-di-AMP), fluoride, prequeuosine 1 (PreQ1), guanine, 5-aminoimidazole-4-carboxamidorivonucleoside-5'-triphosphate (ZTP), glucosamine-6-phosphate (GlcN6P), tetrahydrofolate (THF), glutamine, molybdenum cofactor (Moco), divalent magnesium (Mg 2+ Examples include S-adenosylhomocysteine (SAH), guanidine, azaa aromatic, tungsten cofactor (Wco), aquacobalamin (AqCbl), divalent nickel and divalent cobalt (NiCo), cyclic AMP-GMP (c-AMP-GMP), 2'-deoxyguanosine (2'-dG), and FMN riboswitch variants (FMN-Var). In some embodiments, multiple structural classes have been identified for the same ligand (for example, as with the S-adenosylmethionine (SAM) class). Any riboswitch of any of the described classes, or any other classes currently known or to be discovered later, may be included in the nucleic acid molecules described herein. Such riboswitches may function as splice regulator binding sites described herein. [Table 1]
[0182] In some embodiments, the riboswitch is a riboswitch in the TPP riboswitch class. In some embodiments, the riboswitch is AdoCbl or coenzyme B 12It is a riboswitch in the class of riboswitches. In some embodiments, the riboswitch is a riboswitch in the class of SAM riboswitches. In some embodiments, the riboswitch is a riboswitch in the class of C-di-GMP riboswitches. In some embodiments, the riboswitch is a riboswitch in the class of glycine riboswitches. In some embodiments, the riboswitch is a riboswitch in the class of FMN riboswitches. In some embodiments, the riboswitch is a riboswitch in the class of Mn 2+ This is a riboswitch in the class of riboswitches. In some embodiments, the riboswitch is a riboswitch in the class of lysine riboswitches. In some embodiments, the riboswitch is a riboswitch in the class of C-di-AMP riboswitches. In some embodiments, the riboswitch is a riboswitch in the class of fluoride riboswitches. In some embodiments, the riboswitch is a riboswitch in the class of PreQ1 riboswitches. In some embodiments, the riboswitch is a riboswitch in the class of guanine riboswitches. In some embodiments, the riboswitch is a riboswitch in the class of ZTP riboswitches. In some embodiments, the riboswitch is a riboswitch in the class of GlcN6P riboswitches. In some embodiments, the riboswitch is a riboswitch in the class of THF riboswitches. In some embodiments, the riboswitch is a riboswitch in the class of glutamine riboswitches. In some embodiments, the riboswitch is a riboswitch in the class of Moco riboswitches. In some embodiments, the riboswitch is Mg 2+This is a riboswitch in the class of riboswitches. In some embodiments, the riboswitch is a riboswitch in the class of Sah riboswitches. In some embodiments, the riboswitch is a riboswitch in the class of guanidine riboswitches. In some embodiments, the riboswitch is a riboswitch in the class of aza aromatic riboswitches. In some embodiments, the riboswitch is a riboswitch in the class of Wco riboswitches. In some embodiments, the riboswitch is a riboswitch in the class of AqCbl riboswitches. In some embodiments, the riboswitch is a riboswitch in the class of NiCo riboswitches. In some embodiments, the riboswitch is a riboswitch in the class of c-AMP-GMP riboswitches. In some embodiments, the riboswitch is a riboswitch in the class of 2'-dG riboswitches. In some embodiments, the riboswitch is a riboswitch in the class of FMN-Var riboswitches.
[0183] In some embodiments, the riboswitch responds to RNA-based coenzymes thiamine pyrophosphate (TPP), B12, SAM, and FMN.
[0184] In some embodiments, the riboswitch is c-di-GMP (c-di-AMP and c-AMP-GMP), 5-aminoimidazole-4-carboxamidorivoside 5'-monophosphate (ZTP), glucosamine-6-phosphate (GlcN6P), aza aromatic ligand, guanidine, lysine, glutamine, divalent cation (e.g., Mg 2+ Ni 2+ , and Co 2+It responds to a monoanion and / or fluoride. In some embodiments, the riboswitch responds to c-di-GMP (c-di-AMP and c-AMP-GMP). In some embodiments, the riboswitch responds to ZMP or its triphosphorylated form ZTP. In some embodiments, the riboswitch responds to GlcN6P. In some embodiments, the riboswitch responds to azaa aromatic ligands. In some embodiments, the riboswitch responds to guanidine. In some embodiments, the riboswitch responds to lysine. In some embodiments, the riboswitch responds to glutamine. In some embodiments, the riboswitch responds to a divalent cation (e.g., Mg 2+ Ni 2+ , and Co 2+ ) responds to a monoanion. In some embodiments, the riboswitch responds to a monoanion. In some embodiments, the riboswitch responds to a fluoride.
[0185] The riboswitch is located in an exon of the nucleic acid molecule of this disclosure.
[0186] As described in this disclosure, in some embodiments, the riboswitch is a riboswitch of any suitable riboswitch class that is currently known (e.g., by the riboswitch classes described in this disclosure, but not limited to) or will be discovered later. In some embodiments, the splice regulator binding site described in this disclosure is an aptamer that is part of the riboswitch. Any suitable aptamer is used.
[0187] Transgene In some embodiments, the nucleic acid molecules of this disclosure include a minigene construct linked to a transgene. In some embodiments, the transgene encodes a protein of interest, RNA (e.g., a functional or regulatory RNA of interest), miRNA, or shRNA.
[0188] In some embodiments, the target protein, RNA, miRNA, or shRNA is expressed in the presence of a splice regulator (see the Splice Regulators section). In some embodiments, the target protein, RNA, miRNA, or shRNA is expressed only in the presence of a splice regulator.
[0189] In some embodiments, the transgene encodes a protein. In some embodiments, the protein is expressed only in the presence of a splice regulator.
[0190] In some embodiments, the start codon is located in the second exon of the minigene.
[0191] In some embodiments, the two nucleotides at the 5' end of the bisite start codon are located in the second exon, and the single nucleotide at the 3' end of the start codon is located in the transgene.
[0192] In some embodiments, one nucleotide at the 5' end of the bisite start codon is located in the second exon, and the two nucleotides at the 3' end of the start codon are located in the transgene.
[0193] control array The nucleic acid molecules disclosed herein may be desirable to be expressed at sufficiently high levels to derive therapeutic benefits. Therefore, polynucleotide expression may be mediated by promoter sequences capable of driving robust expression of the disclosed nucleic acid molecules. According to the methods and compositions disclosed herein, the promoter is, in some embodiments, a heterologous promoter. Useful heterologous control sequences generally include sequences derived from sequences encoding mammalian or viral genes. For the purposes of this disclosure, both heterologous promoters and other regulatory elements such as tissue-specific and inducible promoters, enhancers, splicing enhancers, and splicing silencers may be particularly used.
[0194] In some embodiments, the promoter is either entirely derived from a native gene or composed of different elements derived from different native promoters. In some embodiments, the promoter includes a synthetic polynucleotide sequence. Different promoters direct gene expression in different tissues or cell types, or at different developmental stages, or in response to different environmental conditions, or in the presence or absence of drugs or transcription cofactors. Ubiquitous, cell type-specific, tissue-specific, developmental stage-specific, and conditional promoters, such as drug-responsive promoters (e.g., tetracycline-responsive promoters), are well known in the art.
[0195] Examples of promoters useful for the expression of disclosed nucleic acid molecular agents in mammalian cells include, for example, GFAP, nestin, S100B, Nefh, dystrophin H1 promoter, 7SK promoter, apolipoprotein E-human-alpha-1-antitrypsin promoter, CK8 promoter, mouse U1 promoter (mU1a), elongation factor 1α (EF-1α) promoter, thyroxine-binding globulin (TBG) promoter, phosphoglycerate kinase (PKG) promoter, ubiquitous promoters such as CAG ((CMV) cytomegalovirus enhancer, chicken beta-actin promoter (CBA), and rabbit beta-globin intron), SV40 early promoter, and mouse mammary tumor virus LTR promoter, ubiquitous promoters such as adenovirus major late promoter (Ad MLP), herpes simplex virus (HSV) promoter, CMV promoter such as CMV immediate early promoter region (CMV-IE), Roussarcoma virus (RSV) promoter, and U6 promoter or variants thereof. A cell-type specific promoter may be used to drive cell-type specific expression of the nucleic acid molecules disclosed herein.In some embodiments, neuron-specific expression of nucleic acid molecules is conferred using neuron-specific promoters, such as the human synapsin 1 (hSyn) promoter, hexaribonucleotide-binding protein-3 (NeuN) promoter, Ca2+ / calmodulin-dependent protein kinase II (CaMKII) promoter, tubulin alpha I (Tα-1) promoter, neuron-specific enolase (NSE) promoter, platelet-derived growth factor beta chain (PDGFβ) promoter, vesicular glutamate transporter (VGLUT) promoter, somatostatin (SST) promoter, neuropeptide Y (NPY) promoter, vasoactive intestinal peptide (VIP) promoter, parvalbumin (PV) promoter, glutamate decarboxylase (GAD65 or GAD67) promoter, dopamine-1 receptor (DRD1) and dopamine-2 receptor (DRD2) promoters, microtubule-associated protein 1B (MAP1B), and complement component 1. Examples include the q subcomponent-like 2 (C1ql2) promoter, the pro-opiomelanocortin (POMC) promoter, and the prosperohomeobox protein 1 (PROX1) promoter. In some embodiments, the promoter is derived either entirely or from a functional modifier of any of the exemplified promoters.
[0196] Synthetic promoters, hybrid promoters, and the like may also be used in conjunction with the methods and compositions disclosed herein. In addition, sequences derived from non-viral genes, such as the mouse metallothionein gene, may also be found to be used in this disclosure. Such promoter sequences are commercially available, for example, from Stratagene (San Diego, CA).
[0197] In some embodiments, the promoter is a chimeric promoter or a hybrid promoter. For example, in some embodiments, the promoter is a chimeric promoter. In some embodiments, the promoter is a hybrid promoter.
[0198] In some embodiments, the promoter may or may not contain an intragene intron. For example, in some embodiments, the promoter contains an intragene intron. In some embodiments, the promoter does not contain an intragene intron.
[0199] Other DNA sequence elements that may be included in the polynucleotides used in the compositions and methods described herein are enhancer sequences. Enhancers represent another class of regulators that induce conformational changes in the polynucleotide containing the gene of interest so that the DNA adopts a preferred three-dimensional orientation for binding of transcription factors and RNA polymerase at the transcription start site. In some embodiments, the polynucleotides for use in the compositions and methods described herein include mammalian enhancer sequences. In some embodiments, the enhancers are derived from genes encoding mammalian globin, elastase, albumin, α-fetoprotein, and insulin. Enhancers for use in the compositions and methods described herein also include enhancers derived from the genetic material of viruses that can infect eukaryotic cells. Examples include the SV40 enhancer (bp100-270) on the late side of the replication origin, the cytomegalovirus early promoter enhancer, the polyoma enhancer on the late side of the replication origin, and the adenovirus enhancer. Additional enhancer sequences that induce activation of eukaryotic gene transcription are disclosed in Yaniv et al., Nature 297:17 (1982). In some embodiments, the enhancer is spliced, for example, at the 5' or 3' position of the gene into a vector containing a polynucleotide encoding the nucleic acid molecule of the Disclosure. In some embodiments, the enhancer is positioned 5' to the promoter, which is then 5' relative to the polynucleotide encoding the nucleic acid molecule of the Disclosure.
[0200] Additional modulo elements of this disclosure include polyadenylated sequences (polyA sequences). Exemplary polyA of this disclosure include, but are not limited to, SV40 polyA, human growth hormone (HGH) polyA, bovine growth hormone (BGH) polyA, betaglobin polyA, alphaglobin polyA, ovalbumin polyA, kappa light chain polyA, or synthetic polyA. For example, in some embodiments, polyA is SV40 polyA. In some embodiments, polyA is human growth hormone (HGH) polyA. In some embodiments, polyA is bovine growth hormone (BGH) polyA. In some embodiments, polyA is betaglobin polyA. In some embodiments, polyA is alphaglobin polyA. In some embodiments, polyA is ovalbumin polyA. In some embodiments, polyA is kappa light chain polyA. In some embodiments, polyA is synthetic polyA.
[0201] In some embodiments, the adjustment element includes an intron or exon splice silencer or enhancer element. For example, in some embodiments, the adjustment element includes an intron splice silencer. In some embodiments, the adjustment element includes an exon splice silencer. In some embodiments, the adjustment element includes an enhancer element.
[0202] In some embodiments, the nucleic acid molecule further includes a regulatory site that can be recognized by the spliceosome complex.
[0203] Splice regulators The disclosure also provides a composition comprising a nucleic acid molecule according to any one of the embodiments described above and a splice regulator, wherein the splice regulator comprises the structure of formula (I) or a pharmaceutically acceptable salt, solvate, or prodrug thereof.
[0204] In certain embodiments of this disclosure, we describe a nucleic acid molecule comprising (i) a minigene located immediately 5' to the transgene, the minigene comprising a) an exon located 5' to the intron, b) a start codon having a first portion and a second portion, wherein the first and second portions of the start codon are not in the same exon, and c) a splice regulator binding site; and (ii) a splice regulator that binds to the splice regulator binding site and comprises a structure according to formula (I), a pharmaceutically acceptable salt thereof, a solvate thereof, or a prodrug thereof. See Figures 6A-6K.
[0205] In this disclosure, in certain embodiments, a composition is described comprising (i) a nucleic acid molecule comprising a minigene located immediately 5' to the transgene, comprising (a) a first exon located immediately 5' to the first intron, (b) a start codon, and (c) a splice regulator binding site, and (ii) a splice regulator comprising a structure according to formula (I), or a pharmaceutically acceptable salt, solvate, or prodrug thereof, which binds to the splice regulator binding site. See Figures 7A–7G.
[0206] In some embodiments, the splice regulatory factor binding site includes the nucleic acid sequence DGAGTDDGHV (SEQ ID NO: 82) or DGAGTDDNHV (SEQ ID NO: 83), where D is A, G, or T, R is A or G, N is A, C, G, or T, H is A, C, or T, and V is A, G, or C. In some embodiments, the splice regulatory factor binding site includes the nucleic acid sequence DGAGTTTGHV (SEQ ID NO: 84), where D is A, G, or T, H is A, C, or T, and V is A, G, or C. In some embodiments, the splice regulator comprises the nucleic acid sequence DGAGTRRGHV (SEQ ID NO: 1) or DGAGTRRNHV (SEQ ID NO: 2), where D is A, G, or T; R is A or G; N is A, C, G, or T; H is A, C, or T; and V is A, G, or C.
[0207] In some embodiments, the splice regulatory factor binding site does not include the nucleic acid sequences AAGAGT (SEQ ID NO: 3), ATGAGT (SEQ ID NO: 4), TAGAGT (SEQ ID NO: 5), TTGAGT (SEQ ID NO: 6), GAGAGT (SEQ ID NO: 7), GTGAGT (SEQ ID NO: 8), AAGAGT (SEQ ID NO: 9), ATGAGT (SEQ ID NO: 10), ACGAGT (SEQ ID NO: 11), AGGAGT (SEQ ID NO: 12), AGAGGTAGAG (SEQ ID NO: 13), TGAGGTTGAG (SEQ ID NO: 14), GGAGGTGGAG (SEQ ID NO: 15), TAG (SEQ ID NO: 16), CAG (SEQ ID NO: 17), TAG (SEQ ID NO: 18), or NAGAGTNNNN (SEQ ID NO: 19) (where N is A, C, G, or T).
[0208] The splice regulatory factor of the present disclosure has the structure of formula (I):
Chemical formula
[0209] In some embodiments, this disclosure provides, in particular, the structure of formula (Ic), [ka] , or comprising a pharmaceutically acceptable salt, solvate, or prodrug thereof, in the formula, R1 is H, halogen, hydroxyl, cyano, C1-C6 alkyl, C2-C6 alkenyl, C2-C6 alkynyl, C1-C6 haloalkyl, C1-C6 alkoxyl, C1-C6 haloalkoxyl, -(CH2) 0-2 -C3-C8 cycloalkyl, NH2, NH(C1-C6 alkyl), N(C1-C6 alkyl)2, or -(CH2) 0-2-A heterocyclyl, wherein the heterocyclyl comprises one, two, or three heteroatoms independently selected from N, O, and S, and is a 4- to 7-membered ring, and the alkyl, alkenyl, alkynyl, alkoxyl, cycloalkyl, or heterocyclyl may be optionally substituted with one or more C3-C8 cycloalkyl, aryl, or 4- to 7-membered heterocyclyls comprising one, two, or three heteroatoms independently selected from N, O, and S. R2 is a 5, 6, or 9-membered heterocyclyl containing 1, 2, or 3 heteroatoms independently selected from aryl, 5- to 7-membered cycloalkyl, N, O, and S, or a 5, 6, or 9-membered heteroaryl containing 2, 3, or 4 heteroatoms independently selected from N, O, and S, wherein the aryl, cycloalkyl, heterocyclyl, or heteroaryl may be optionally substituted with one or more R4s. Each R3 is independently a halogen, C1-C6 alkyl, C2-C6 alkenyl, C2-C6 alkynyl, C1-C6 alkoxyl, C3-C8 cycloalkyl, NH2, NH(C1-C6 alkyl), or N(C1-C6 alkyl)2, and the alkyl, alkenyl, alkynyl, alkoxyl, and cycloalkyl may be optionally substituted with one or more hydroxyls or NH2. Each R4 is independently a halogen, hydroxyl, cyano, nitro, C1-C6 alkyl, C2-C6 alkenyl, C2-C6 alkynyl, C1-C6 haloalkyl, C1-C6 alkoxyl, C1-C6 haloalkoxyl, C3-C8 cycloalkyl, NH2, NH(C1-C6 alkyl), N(C1-C6 alkyl)2, or C(O)NH2, and the alkyl, alkenyl, alkynyl, alkoxyl, and cycloalkyl may be optionally substituted with a 4- to 7-membered heterocyclyl, NH2, NH(C1-C6 alkyl), or N(C1-C6 alkyl)2 containing one or more heteroatoms independently selected from hydroxyl, N, O, and S. R5 is H, C1-C6 alkyl, C2-C6 alkenyl, C2-C6 alkynyl, C1-C6 alkoxyl, C3-C8 cycloalkyl, -CH2C3-C8 cycloalkyl, heterocyclyl, -CH2 heterocisyl, -CH2CH2 heterocisyl, -CH2-(5-6 member heteroaryl), where the heterocyclyl is a 4-7 member heterocyclyl containing 1, 2, or 3 heteroatoms independently selected from N, O, and S, and in the formula, R5 may be optionally substituted with one or more halogens 'C1-C6 alkyl, C1-C6 haloalkyl, C1-C6 heteroalkyl, C1-C6 alkoxyl, C3-C8 cycloalkyl, spiroC3-C8 cycloalkyl, spiro4-7 member heterocyclyl, 5-6 member heteroaryl, oxo, cyano, or hydroxyl'. R6 is H, halogen, C1-C6 alkyl, or C1-C6 haloalkyl. R7 is H, halogen, C1-C6 alkyl, or C1-C6 haloalkyl. n is 0, 1, 2, 3, 4, or 5.
[0210] In some embodiments, this disclosure provides, in particular, the structure of formula (Id), [ka] or comprising a pharmaceutically acceptable salt, solvate, or prodrug thereof, in the formula, W is -S- or -HC=CH-, R1 is a 4- to 7-membered heterocycline containing H, halogen, hydroxyl, cyano, C1-C6 alkyl, C2-C6 alkenyl, C2-C6 alkynyl, C1-C6 haloalkyl, C1-C6 alkoxyl, C1-C6 haloalkoxyl, C3-C8 cycloalkyl, NH2, NH(C1-C6 alkyl), N(C1-C6 alkyl)2, or 1, 2, or 3 heteroatoms independently selected from N, O, and S, and the alkyl, alkenyl, alkynyl, alkoxyl, cycloalkyl, or heterocycline may be optionally substituted with one or more C3-C8 cycloalkyl, aryl, or 4- to 7-membered heterocyclines containing 1, 2, or 3 heteroatoms independently selected from N, O, and S. R2 is a 5, 6, or 9-membered heterocyclyl containing 1, 2, or 3 heteroatoms independently selected from aryl, 5- to 7-membered cycloalkyl, N, O, and S, or a 5, 6, or 9-membered heteroaryl containing 2, 3, or 4 heteroatoms independently selected from N, O, and S, wherein the aryl, cycloalkyl, heterocyclyl, or heteroaryl may be optionally substituted with one or more R4s. Each R3 is independently a halogen, C1-C6 alkyl, C2-C6 alkenyl, C2-C6 alkynyl, C1-C6 alkoxyl, C3-C8 cycloalkyl, NH2, NH(C1-C6 alkyl), or N(C1-C6 alkyl)2, and the alkyl, alkenyl, alkynyl, alkoxyl, and cycloalkyl may be optionally substituted with one or more hydroxyls or NH2. Each R4 is independently a halogen, hydroxyl, cyano, nitro, C1-C6 alkyl, C2-C6 alkenyl, C2-C6 alkynyl, C1-C6 haloalkyl, C1-C6 alkoxyl, C1-C6 haloalkoxyl, C3-C8 cycloalkyl, NH2, NH(C1-C6 alkyl), N(C1-C6 alkyl)2, or C(O)NH2, and the alkyl, alkenyl, alkynyl, alkoxyl, and cycloalkyl may be optionally substituted with a 4- to 7-membered heterocyclyl, NH2, NH(C1-C6 alkyl), or N(C1-C6 alkyl)2 containing one or more heteroatoms independently selected from hydroxyl, N, O, and S. R5 is H, C1-C6 alkyl, C2-C6 alkenyl, C2-C6 alkynyl, C1-C6 alkoxyl, C3-C8 cycloalkyl, -CH2C3-C8 cycloalkyl, heterocyclyl, -CH2 heterocisyl, -CH2CH2 heterocisyl, -CH2-(5-6 membered heteroaryl), where the heterocyclyl is a 4-7 membered ring and contains 1, 2, or 3 heteroatoms independently selected from N, O, and S, wherein R5 may be optionally substituted with one or more halogens, C1-C6 alkyl, C1-C6 haloalkyl, C1-C6 heteroalkyl, C1-C6 alkoxyl, C3-C8 cycloalkyl, spiroC3-C8 cycloalkyl, spiro4-7 membered heterocyclyl, 5-6 membered heteroaryl, oxo, cyano, or hydroxyl. n is 0, 1, 2, 3, 4, or 5.
[0211] In some embodiments, W is -S-.
[0212] In some embodiments, W is -HC=CH-.
[0213] In some embodiments, R1 is H, halogen, hydroxyl, cyano, C1-C6 alkyl, C2-C6 alkenyl, C2-C6 alkynyl, C1-C6 haloalkyl, C1-C6 alkoxyl, C1-C6 haloalkoxyl, C3-C8 cycloalkyl, NH2, NH(C1-C6 alkyl), N(C1-C6 alkyl)2, or a 4- to 7-membered heterocyclyl containing 1, 2, or 3 heteroatoms independently selected from N, O, and S.
[0214] In some embodiments, R1 is H.
[0215] In some embodiments, when R1 is halogen, hydroxyl, cyano, C1-C6 alkyl, C2-C6 alkenyl, C2-C6 alkynyl, C1-C6 haloalkyl, C1-C6 alkoxyl, C1-C6 haloalkoxyl, C3-C8 cycloalkyl, NH2, NH(C1-C6 alkyl), N(C1-C6 alkyl)2, or a 4- to 7-membered heterocyclyl containing 1, 2, or 3 heteroatoms independently selected from N, O, and S, the alkyl, alkenyl, alkynyl, alkoxyl, cycloalkyl, heterocyclyl may be optionally substituted with one or more C3-C8 cycloalkyl, aryl, or a 4- to 7-membered heterocyclyl containing 1, 2, or 3 heteroatoms independently selected from N, O, and S.
[0216] In some embodiments, R1 is halogen, hydroxyl, cyano, C1-C6 alkyl, C2-C6 alkenyl, C2-C6 alkynyl, C1-C6 haloalkyl, C1-C6 alkoxyl, C1-C6 haloalkoxyl, C3-C8 cycloalkyl, NH2, NH(C1-C6 alkyl), N(C1-C6 alkyl)2, or a 4- to 7-membered heterocyclyl containing 1, 2, or 3 heteroatoms independently selected from N, O, and S heteroaryl.
[0217] In some embodiments, R1 is halogen, hydroxyl, or cyano.
[0218] In some embodiments, R1 is a halogen. In some embodiments, R1 is F, Cl, Br, or I. In some embodiments, R1 is F, Cl, or Br. In some embodiments, R1 is F or Cl. In some embodiments, R1 is F. In some embodiments, R1 is Cl. In some embodiments, R1 is Br. In some embodiments, R1 is I.
[0219] In some embodiments, R1 is a hydroxyl group.
[0220] In some embodiments, R1 is cyanoacrylate.
[0221] In some embodiments, R1 is a C1-C6 alkyl, C2-C6 alkenyl, C2-C6 alkynyl, C1-C6 haloalkyl, C1-C6 alkoxyl, C1-C6 haloalkoxyl, C3-C8 cycloalkyl, NH2, NH(C1-C6 alkyl), N(C1-C6 alkyl)2, or a 4- to 7-membered heterocycline containing one, two, or three heteroatoms independently selected from N, O, and S, wherein the alkyl, alkenyl, alkynyl, alkoxyl, cycloalkyl, or heterocycline may be optionally substituted with one or more C3-C8 cycloalkyl, aryl, or 4- to 7-membered heterocyclines containing one, two, or three heteroatoms independently selected from N, O, and S.
[0222] In some embodiments, R1 is a 4- to 7-membered heterocycline containing one, two, or three heteroatoms independently selected from C1-C6 alkyl, C2-C6 alkenyl, C2-C6 alkynyl, C1-C6 haloalkyl, C1-C6 alkoxyl, C1-C6 haloalkoxyl, C3-C8 cycloalkyl, NH2, NH(C1-C6 alkyl), N(C1-C6 alkyl)2, or N, O, and S.
[0223] In some embodiments, R1 is a C1-C6 alkyl, a C2-C6 alkenyl, or a C2-C6 alkynyl.
[0224] In some embodiments, R1 is a C1-C6 alkyl group.
[0225] In some embodiments, R1 is methyl. In some embodiments, R1 is ethyl. In some embodiments, R1 is propyl. In some embodiments, R1 is butyl. In some embodiments, R1 is isopropyl. In some embodiments, R1 is isobutyl. In some embodiments, R1 is sec-butyl. In some embodiments, R1 is tert-butyl.
[0226] In some embodiments, R1 is a C1-C6 alkyl group that may be optionally substituted with one or more C3-C8 cycloalkyl, aryl, or 4- to 7-membered heterocyclines containing 1, 2, or 3 heteroatoms independently selected from N, O, and S. In some embodiments, R1 is a C1-C6 alkyl group that is substituted with one or more C3-C8 cycloalkyl, aryl, or 4- to 7-membered heterocyclines containing 1, 2, or 3 heteroatoms independently selected from N, O, and S. In some embodiments, R1 is a C1-C6 alkyl group that is substituted with one C3-C8 cycloalkyl, aryl, or 4- to 7-membered heterocyclines containing 1, 2, or 3 heteroatoms independently selected from N, O, and S. In some embodiments, R1 is a C1-C6 alkyl group that is substituted with two C3-C8 cycloalkyl, aryl, or 4- to 7-membered heterocyclines containing 1, 2, or 3 heteroatoms independently selected from N, O, and S. In some embodiments, R1 is a C1-C6 alkyl group substituted with three C3-C8 cycloalkyl, aryl, or 4- to 7-membered heterocycline containing one, two, or three heteroatoms independently selected from N, O, and S.
[0227] In some embodiments, R1 is a C2-C6 alkenyl. In some embodiments, R1 is a C2 alkenyl. In some embodiments, R1 is a C3 alkenyl. In some embodiments, R1 is a C4 alkenyl. In some embodiments, R1 is a C5 alkenyl. In some embodiments, R1 is a C6 alkenyl.
[0228] In some embodiments, R1 is a C2-C6 alkenyl in which R1 may be optionally substituted with one or more C3-C8 cycloalkyl, aryl, or 4- to 7-membered heterocyclines containing 1, 2, or 3 heteroatoms independently selected from N, O, and S. In some embodiments, R1 is a C2-C6 alkenyl in which R1 is substituted with one or more C3-C8 cycloalkyl, aryl, or 4- to 7-membered heterocyclines containing 1, 2, or 3 heteroatoms independently selected from N, O, and S. In some embodiments, R1 is a C2-C6 alkenyl in which R1 is substituted with one C3-C8 cycloalkyl, aryl, or 4- to 7-membered heterocyclines containing 1, 2, or 3 heteroatoms independently selected from N, O, and S. In some embodiments, R1 is a C2-C6 alkenyl substituted with a 4- to 7-membered heterocycline containing two C3-C8 cycloalkyl, aryl, or 1, 2, or 3 heteroatoms independently selected from N, O, and S. In some embodiments, R1 is a C2-C6 alkenyl substituted with a 4- to 7-membered heterocycline containing three C3-C8 cycloalkyl, aryl, or 1, 2, or 3 heteroatoms independently selected from N, O, and S.
[0229] In some embodiments, R1 is a C2-C6 alkynyl. In some embodiments, R1 is a C2 alkynyl. In some embodiments, R1 is a C3 alkynyl. In some embodiments, R1 is a C4 alkynyl. In some embodiments, R1 is a C5 alkynyl. In some embodiments, R1 is a C6 alkynyl.
[0230] In some embodiments, R1 is a C2-C6 alkynyl which may be optionally substituted with one or more C3-C8 cycloalkyl, aryl, or 4- to 7-membered heterocyclines containing 1, 2, or 3 heteroatoms independently selected from N, O, and S. In some embodiments, R1 is a C2-C6 alkynyl which is substituted with one or more C3-C8 cycloalkyl, aryl, or 4- to 7-membered heterocyclines containing 1, 2, or 3 heteroatoms independently selected from N, O, and S. In some embodiments, R1 is a C2-C6 alkynyl which is substituted with one C3-C8 cycloalkyl, aryl, or 4- to 7-membered heterocyclines containing 1, 2, or 3 heteroatoms independently selected from N, O, and S. In some embodiments, R1 is a C2-C6 alkynyl substituted with two C3-C8 cycloalkyl, aryl, or 4- to 7-membered heterocyclines containing 1, 2, or 3 heteroatoms independently selected from N, O, and S. In some embodiments, R1 is a C2-C6 alkynyl substituted with three C3-C8 cycloalkyl, aryl, or 4- to 7-membered heterocyclines containing 1, 2, or 3 heteroatoms independently selected from N, O, and S.
[0231] In some embodiments, R1 is a C1-C6 haloalkyl, C1-C6 alkoxyl, C1-C6 haloalkoxyl, C3-C8 cycloalkyl, NH2, NH(C1-C6 alkyl), N(C1-C6 alkyl)2, or a 4- to 7-membered heterocycline containing 1, 2, or 3 heteroatoms independently selected from N, O, and S, and the alkoxyl, cycloalkyl, and heterocycline may be optionally substituted with one or more C3-C8 cycloalkyl, aryl, or 4- to 7-membered heterocyclines containing 1, 2, or 3 heteroatoms independently selected from N, O, and S.
[0232] In some embodiments, R1 is a 4- to 7-membered heterocycline containing one, two, or three heteroatoms independently selected from C1-C6 haloalkyl, C1-C6 alkoxyl, C1-C6 haloalkoxyl, C3-C8 cycloalkyl, NH2, NH(C1-C6 alkyl), N(C1-C6 alkyl)2, or N, O, and S.
[0233] In some embodiments, R1 is a C1-C6 haloalkyl, a C1-C6 alkoxyl, or a C1-C6 haloalkoxyl, and the alkoxyl may be optionally substituted with one or more C3-C8 cycloalkyl, aryl, or 4-7 membered heterocyclines containing one, two, or three heteroatoms independently selected from N, O, and S.
[0234] In some embodiments, R1 is a C1-C6 haloalkyl, a C1-C6 alkoxyl, or a C1-C6 haloalkoxyl.
[0235] In some embodiments, R1 is a C1-C6 haloalkyl. In some embodiments, R1 is a halomethyl. In some embodiments, R1 is a haloethyl. In some embodiments, R1 is a halopropyl. In some embodiments, R1 is a halobutyl. In some embodiments, R1 is a halopentyl. In some embodiments, R1 is a halohexyl.
[0236] In some embodiments, R1 is a C1-C6 alkoxyl. In some embodiments, R1 is a methoxyl. In some embodiments, R1 is an ethoxyl. In some embodiments, R1 is a propoxyl. In some embodiments, R1 is a butioxyl. In some embodiments, R1 is a pentioxyl. In some embodiments, R1 is a hexoxyl.
[0237] In some embodiments, R1 is a C1-C6 haloalkoxyl. In some embodiments, R1 is a halomethoxyl. In some embodiments, R1 is a haloethoxyl. In some embodiments, R1 is a halopropoxyl. In some embodiments, R1 is a halobutoxyl. In some embodiments, R1 is a halopentoxyl. In some embodiments, R1 is a halohexoxyl.
[0238] In some embodiments, R1 is a C3-C8 cycloalkyl or a 4- to 7-membered heterocycline containing one, two, or three heteroatoms independently selected from N, O, and S, and the cycloalkyl and heterocycline may be optionally substituted with one or more C3-C8 cycloalkyl, aryl, or 4- to 7-membered heterocyclines containing one, two, or three heteroatoms independently selected from N, O, and S.
[0239] In some embodiments, R1 is a C3-C8 cycloalkyl group, or a 4- to 7-membered heterocycline containing one, two, or three heteroatoms independently selected from N, O, and S.
[0240] In some embodiments, R1 is a C3-C8 cycloalkyl group. In some embodiments, R1 is a cyclopropyl group. In some embodiments, R1 is a cyclobutyl group. In some embodiments, R1 is a cyclopentyl group. In some embodiments, R1 is a cyclohexyl group.
[0241] In some embodiments, R1 is a crosslinked C3-C8 cycloalkyl. In some embodiments, R1 is a condensed C3-C8 cycloalkyl. In some embodiments, R1 is a spiro-C3-C8 cycloalkyl.
[0242] In some embodiments, R1 is a C3-C8 cycloalkyl that may be optionally substituted with one or more C3-C8 cycloalkyl, aryl, or 4- to 7-membered heterocyclyl groups containing one, two, or three heteroatoms independently selected from N, O, and S. In some embodiments, R1 is a C3-C8 cycloalkyl that may be optionally substituted with one or more C3-C8 cycloalkyl or aryl groups. In some embodiments, R1 is a C3-C8 cycloalkyl that is substituted with one or more C3-C8 cycloalkyl or aryl groups. In some embodiments, R1 is a C3-C8 cycloalkyl that is substituted with one or more C3-C8 cycloalkyl groups. In some embodiments, R1 is a C3-C8 cycloalkyl that is substituted with one or more aryl groups.
[0243] In some embodiments, R1 is a 4- to 7-membered heterocycline containing 1, 2, or 3 heteroatoms independently selected from N, O, and S.
[0244] In some embodiments, R1 is a four-membered heterocycline containing one, two, or three heteroatoms independently selected from N, O, and S.
[0245] In some embodiments, R1 is a five-membered heterocycline containing one, two, or three heteroatoms independently selected from N, O, and S.
[0246] In some embodiments, R1 is a six-membered heterocycline containing one, two, or three heteroatoms independently selected from N, O, and S.
[0247] In some embodiments, R1 is a seven-membered heterocycline containing one, two, or three heteroatoms independently selected from N, O, and S.
[0248] In some embodiments, R1 is a 4- to 7-membered heterocycline, aryl, or 4- to 7-membered heterocycline containing 1, 2, or 3 heteroatoms independently selected from N, O, and S. In some embodiments, R1 is a 4- to 7-membered heterocycline containing 1, 2, or 3 heteroatoms independently selected from N, O, and S, and may be optionally substituted with one or more C3-C8 cycloalkyl or aryl atoms. In some embodiments, R1 is a 4- to 7-membered heterocycline containing 1, 2, or 3 heteroatoms independently selected from N, O, and S, substituted with one or more C3-C8 cycloalkyl or aryl atoms. In some embodiments, R1 is a 4- to 7-membered heterocycline containing 1, 2, or 3 heteroatoms independently selected from N, O, and S, substituted with one or more C3-C8 cycloalkyl atoms. In some embodiments, R1 is a 4- to 7-membered heterocycline comprising one, two, or three heteroatoms independently selected from N, O, and S, which are substituted with one or more aryl atoms.
[0249] In some embodiments, R1 is NH2, NH(C1-C6 alkyl), or N(C1-C6 alkyl)2.
[0250] In some embodiments, R1 is NH2.
[0251] In some embodiments, R1 is NH(C1-C6 alkyl) or N(C1-C6 alkyl)2.
[0252] In some embodiments, R1 is NH(C1-C6 alkyl). In some embodiments, R1 is NH(methyl). In some embodiments, R1 is NH(ethyl). In some embodiments, R1 is NH(propyl). In some embodiments, R1 is NH(butyl). In some embodiments, R1 is NH(pentyl). In some embodiments, R1 is NH(hexyl).
[0253] In some embodiments, R1 is N(C1-C6 alkyl)2. In some embodiments, R1 is N(methyl)2. In some embodiments, R1 is N(ethyl)2. In some embodiments, R1 is N(propyl)2. In some embodiments, R1 is N(butyl)2. In some embodiments, R1 is N(pentyl)2. In some embodiments, R1 is N(hexyl)2.
[0254] In some embodiments, R1 is H or F.
[0255] In some embodiments, R2 is an aryl.
[0256] In some embodiments, R2 is an aryl that may be optionally replaced by one or more R4s. In some embodiments, R2 is an aryl that is replaced by one or more R4s. In some embodiments, R2 is an aryl that is replaced by one R4. In some embodiments, R2 is an aryl that is replaced by two R4s. In some embodiments, R2 is an aryl that is replaced by three R4s.
[0257] In some embodiments, R2 is a 5- to 7-membered cycloalkyl group.
[0258] In some embodiments, R2 is a 5- to 7-membered saturated cycloalkyl group. In some embodiments, R2 is a 5- to 7-membered partially saturated cycloalkyl group.
[0259] In some embodiments, R2 is a 5- to 7-membered cycloalkyl group optionally substituted with one or more R4 groups. In some embodiments, R2 is a 5- to 7-membered cycloalkyl group substituted with one or more R4 groups. In some embodiments, R2 is a 5- to 7-membered cycloalkyl group substituted with one R4 group. In some embodiments, R2 is a 5- to 7-membered cycloalkyl group substituted with two R4 groups. In some embodiments, R2 is a 5- to 7-membered cycloalkyl group substituted with three R4 groups.
[0260] In some embodiments, R2 is a 5, 6, or 9-membered heterocycline containing 1, 2, or 3 heteroatoms independently selected from N, O, and S.
[0261] In some embodiments, R2 is a 5, 6, or 9-membered heterocycline containing 1, 2, or 3 heteroatoms independently selected from N and O.
[0262] In some embodiments, R2 is a 5, 6, or 9-membered heterocycline containing one or two heteroatoms independently selected from N, O, and S.
[0263] In some embodiments, R2 is a 5, 6, or 9-membered heterocycline containing one or two heteroatoms independently selected from N and O.
[0264] In some embodiments, R2 is a 5, 6, or 9-membered heterocycline containing 1, 2, or 3 heteroatoms independently selected from N, O, and S, and R2 may be optionally substituted with one or more R4s.
[0265] In some embodiments, R2 is a 5, 6, or 9-membered heteroaryl compound containing 2, 3, or 4 heteroatoms independently selected from N, O, and S.
[0266] In some embodiments, R2 is a 5, 6, or 9-membered heteroaryl compound containing 2, 3, or 4 heteroatoms independently selected from N and O.
[0267] In some embodiments, R2 is a 5, 6, or 9-membered heteroaryl comprising 2, 3, or 4 heteroatoms independently selected from N, O, and S, and the heteroaryl may be optionally substituted with one or more R4s.
[0268] In some embodiments, R2 is a 5, 6, or 9-membered heteroaryl comprising 2, 3, or 4 heteroatoms independently selected from N and O, and the heteroaryl may be optionally substituted with one or more R4s.
[0269] In some embodiments, R2 is a 5, 6, or 9-membered heteroaryl containing 2, 3, or 4 heteroatoms independently selected from N, O, and S, and the heteroaryl is substituted with one or more R4s.
[0270] In some embodiments, R2 is a 5, 6, or 9-membered heteroaryl containing 2, 3, or 4 heteroatoms independently selected from N and O, and the heteroaryl is substituted with one or more R4s.
[0271] In some embodiments, R2 is a 5, 6, or 9-membered heteroaryl containing 2, 3, or 4 heteroatoms independently selected from N, O, and S, and the heteroaryl is substituted with one R4.
[0272] In some embodiments, R2 is a 5, 6, or 9-membered heteroaryl containing 2, 3, or 4 heteroatoms independently selected from N and O, and the heteroaryl is substituted with one R4.
[0273] In some embodiments, R2 is a 5, 6, or 9-membered heteroaryl containing 2, 3, or 4 heteroatoms independently selected from N, O, and S, and the heteroaryl is substituted with 2 R4s.
[0274] In some embodiments, R2 is a 5, 6, or 9-membered heteroaryl containing 2, 3, or 4 heteroatoms independently selected from N and O, and the heteroaryl is substituted with 2 R4s.
[0275] In some embodiments, R2 is a 5, 6, or 9-membered heteroaryl containing 2, 3, or 4 heteroatoms independently selected from N, O, and S, and the heteroaryl is substituted with 3 R4 atoms.
[0276] In some embodiments, R2 is a 5, 6, or 9-membered heteroaryl containing 2, 3, or 4 heteroatoms independently selected from N and O, and the heteroaryl is substituted with 3 R4 atoms.
[0277] In some embodiments, R2 is a nine-membered heteroaryl comprising two, three, or four heteroatoms independently selected from N, O, and S, and the heteroaryl may be optionally substituted with one or more R4s.
[0278] In some embodiments, R2 is a nine-membered heteroaryl comprising two, three, or four heteroatoms independently selected from N and O, and the heteroaryl may be optionally substituted with one or more R4s.
[0279] In some embodiments, R2 is a nine-membered heteroaryl comprising two, three, or four heteroatoms independently selected from N, O, and S, and the heteroaryl is substituted with one or more R4s.
[0280] In some embodiments, R2 is a nine-membered heteroaryl comprising two, three, or four heteroatoms independently selected from N and O, and the heteroaryl is substituted with one or more R4s.
[0281] In some embodiments, R2 is a nine-membered heteroaryl compound containing two, three, or four heteroatoms independently selected from N, O, and S, and the heteroaryl compound is substituted with one R4.
[0282] In some embodiments, R2 is a nine-membered heteroaryl compound containing two, three, or four heteroatoms independently selected from N and O, and the heteroaryl compound is substituted with one R4.
[0283] In some embodiments, R2 is a nine-membered heteroaryl compound containing two, three, or four heteroatoms independently selected from N, O, and S, and the heteroaryl compound is substituted with two R4 atoms.
[0284] In some embodiments, R2 is a nine-membered heteroaryl compound containing two, three, or four heteroatoms independently selected from N and O, and the heteroaryl compound is substituted with two R4 atoms.
[0285] In some embodiments, R2 is a nine-membered heteroaryl compound containing two, three, or four heteroatoms independently selected from N, O, and S, and the heteroaryl compound is substituted with three R4 atoms.
[0286] In some embodiments, R2 is a nine-membered heteroaryl compound containing two, three, or four heteroatoms independently selected from N and O, and the heteroaryl compound is substituted with three R4 atoms.
[0287] In some embodiments, R2 is a nine-membered heteroaryl comprising two heteroatoms independently selected from N, O, and S, and the heteroaryl may be optionally substituted with one or more R4s.
[0288] In some embodiments, R2 is a nine-membered heteroaryl compound containing two heteroatoms independently selected from N and O, and the heteroaryl compound may be optionally substituted with one or more R4s.
[0289] In some embodiments, R2 is a nine-membered heteroaryl compound containing two heteroatoms independently selected from N, O, and S, and the heteroaryl compound is substituted with one or more R4s.
[0290] In some embodiments, R2 is a nine-membered heteroaryl compound containing two heteroatoms independently selected from N and O, and the heteroaryl compound is substituted with one or more R4s.
[0291] In some embodiments, R2 is a nine-membered heteroaryl compound containing two heteroatoms independently selected from N, O, and S, and the heteroaryl compound is substituted with one R4.
[0292] In some embodiments, R2 is a nine-membered heteroaryl compound containing two heteroatoms independently selected from N and O, and the heteroaryl compound is substituted with one R4.
[0293] In some embodiments, R2 is a nine-membered heteroaryl compound containing two heteroatoms independently selected from N, O, and S, and the heteroaryl compound is substituted with two R4s.
[0294] In some embodiments, R2 is a nine-membered heteroaryl compound containing two heteroatoms independently selected from N and O, and the heteroaryl compound is substituted with two R4s.
[0295] In some embodiments, R2 is a nine-membered heteroaryl compound containing two heteroatoms independently selected from N, O, and S, and the heteroaryl compound is substituted with three R4 atoms.
[0296] In some embodiments, R2 is a nine-membered heteroaryl compound containing two heteroatoms independently selected from N and O, and the heteroaryl compound is substituted with three R4 atoms.
[0297] In some embodiments, R2 is a nine-membered heteroaryl comprising three heteroatoms independently selected from N, O, and S, and the heteroaryl may be optionally substituted with one or more R4s.
[0298] In some embodiments, R2 is a nine-membered heteroaryl compound comprising three heteroatoms independently selected from N and O, and the heteroaryl compound may be optionally substituted with one or more R4s.
[0299] In some embodiments, R2 is a nine-membered heteroaryl compound containing three heteroatoms independently selected from N, O, and S, and the heteroaryl compound is substituted with one or more R4s.
[0300] In some embodiments, R2 is a nine-membered heteroaryl compound containing three heteroatoms independently selected from N and O, and the heteroaryl compound is substituted with one or more R4s.
[0301] In some embodiments, R2 is a nine-membered heteroaryl compound containing three heteroatoms independently selected from N, O, and S, and the heteroaryl compound is substituted with one R4.
[0302] In some embodiments, R2 is a nine-membered heteroaryl compound containing three heteroatoms independently selected from N and O, and the heteroaryl compound is substituted with one R4.
[0303] In some embodiments, R2 is a nine-membered heteroaryl compound containing three heteroatoms independently selected from N, O, and S, and the heteroaryl compound is substituted with two R4 atoms.
[0304] In some embodiments, R2 is a nine-membered heteroaryl compound containing three heteroatoms independently selected from N and O, and the heteroaryl compound is substituted with two R4 atoms.
[0305] In some embodiments, R2 is a nine-membered heteroaryl compound containing three heteroatoms independently selected from N, O, and S, and the heteroaryl compound is substituted with three R4 atoms.
[0306] In some embodiments, R2 is a nine-membered heteroaryl compound containing three heteroatoms independently selected from N and O, and the heteroaryl compound is substituted with three R4 atoms.
[0307] In some embodiments, R2 is a nine-membered heteroaryl comprising four heteroatoms independently selected from N, O, and S, and the heteroaryl may be optionally substituted with one or more R4s.
[0308] In some embodiments, R2 is a nine-membered heteroaryl comprising four heteroatoms independently selected from N and O, and the heteroaryl may be optionally substituted with one or more R4s.
[0309] In some embodiments, R2 is a nine-membered heteroaryl compound containing four heteroatoms independently selected from N, O, and S, and the heteroaryl compound is substituted with one or more R4s.
[0310] In some embodiments, R2 is a nine-membered heteroaryl compound containing four heteroatoms independently selected from N and O, and the heteroaryl compound is substituted with one or more R4s.
[0311] In some embodiments, R2 is a nine-membered heteroaryl compound containing four heteroatoms independently selected from N, O, and S, and the heteroaryl compound is substituted with one R4.
[0312] In some embodiments, R2 is a nine-membered heteroaryl compound containing four heteroatoms independently selected from N and O, and the heteroaryl compound is substituted with one R4.
[0313] In some embodiments, R2 is a nine-membered heteroaryl compound containing four heteroatoms independently selected from N, O, and S, and the heteroaryl compound is substituted with two R4 atoms.
[0314] In some embodiments, R2 is a nine-membered heteroaryl compound containing four heteroatoms independently selected from N and O, and the heteroaryl compound is substituted with two R4 atoms.
[0315] In some embodiments, R2 is a nine-membered heteroaryl compound containing four heteroatoms independently selected from N, O, and S, and the heteroaryl compound is substituted with three R4 atoms.
[0316] In some embodiments, R2 is a nine-membered heteroaryl compound containing four heteroatoms independently selected from N and O, and the heteroaryl compound is substituted with three R4 atoms.
[0317] In some embodiments, R2 is a biring 9-membered heteroaryl.
[0318] In some embodiments, R2 is [ka] And, In the equation, each R8 is independently R4.
[0319] In some embodiments, R2 is [ka] That is the case.
[0320] In some embodiments, R3 is independently a halogen, a C1-C6 alkyl, a C2-C6 alkenyl, a C2-C6 alkynyl, a C1-C6 alkoxyl, a C3-C8 cycloalkyl, NH2, NH(C1-C6 alkyl), or N(C1-C6 alkyl)2.
[0321] In some embodiments, at least one R3 is a halogen. In some embodiments, at least one R3 is F, Cl, Br, or I. In some embodiments, at least one R3 is F, Cl, or Br. In some embodiments, at least one R3 is F or Cl. In some embodiments, at least one R3 is F. In some embodiments, at least one R3 is Cl. In some embodiments, at least one R3 is Br. In some embodiments, at least one R3 is I.
[0322] In some embodiments, R3 is independently a C1-C6 alkyl, a C2-C6 alkenyl, a C2-C6 alkynyl, a C1-C6 alkoxyl, or a C3-C8 cycloalkyl.
[0323] In some embodiments, R3 is independently a C1-C6 alkyl, a C2-C6 alkenyl, a C2-C6 alkynyl, a C1-C6 alkoxyl, or a C3-C8 cycloalkyl, and the alkyl, alkenyl, alkynyl, alkoxyl, and cycloalkyl may be optionally substituted with one or more hydroxyls or NH2 groups.
[0324] In some embodiments, R3 is independently a C1-C6 alkyl, a C2-C6 alkenyl, or a C2-C6 alkynyl.
[0325] In some embodiments, R3 is independently a C1-C6 alkyl, a C2-C6 alkenyl, or a C2-C6 alkynyl, and the alkyl, alkenyl, or alkynyl may be optionally substituted with one or more hydroxyls or NH2 groups.
[0326] In some embodiments, at least one R3 is a C1-C6 alkyl group. In some embodiments, at least one R3 is methyl. In some embodiments, at least one R3 is ethyl. In some embodiments, at least one R3 is propyl. In some embodiments, at least one R3 is butyl. In some embodiments, R3 is isopropyl. In some embodiments, at least one R3 is isobutyl. In some embodiments, at least one R3 is sec-butyl. In some embodiments, at least one R3 is tert-butyl.
[0327] In some embodiments, each R3 is a C1-C6 alkyl group optionally substituted with one or more hydroxyls or NH2 groups.
[0328] In some embodiments, R3 is a C2-C6 alkenyl. In some embodiments, R3 is a C2 alkenyl. In some embodiments, R3 is a C3 alkenyl. In some embodiments, R3 is a C4 alkenyl. In some embodiments, R3 is a C5 alkenyl. In some embodiments, R3 is a C6 alkenyl.
[0329] In some embodiments, R3 is a C2-C6 alkenyl optionally substituted with one or more hydroxyls or NH2 groups.
[0330] In some embodiments, R3 is C2-C6 alkynyl. In some embodiments, R3 is C2 alkynyl. In some embodiments, R3 is C3 alkynyl. In some embodiments, R3 is C4 alkynyl. In some embodiments, R3 is C5 alkynyl. In some embodiments, R3 is C6 alkynyl.
[0331] In some embodiments, R3 is a C2-C6 alkynyl which is optionally substituted with one or more hydroxyls or NH2 groups.
[0332] In some embodiments, each R3 is independently a C1-C6 alkoxyl or a C3-C8 cycloalkyl.
[0333] In some embodiments, R3 is independently a C1-C6 alkoxyl or a C3-C8 cycloalkyl, and the alkoxyl and cycloalkyl may be optionally substituted with one or more hydroxyls or NH2 groups.
[0334] In some embodiments, each R3 is independently a C1-C6 alkoxyl optionally substituted with one or more hydroxyls or NH2s.
[0335] In some embodiments, at least one R3 is a C1-C6 alkoxyl. In some embodiments, at least one R3 is a methoxyl. In some embodiments, at least one R3 is an ethoxyl. In some embodiments, at least one R3 is a propoxyl. In some embodiments, at least one R3 is a butoxyl. In some embodiments, at least one R3 is a pentoxyl. In some embodiments, at least one R3 is a hexoxyl.
[0336] In some embodiments, each R3 is independently a C3-C8 cycloalkyl group optionally substituted with one or more hydroxyls or NH2 groups.
[0337] In some embodiments, at least one R3 is independently C3-C8 cycloalkyl. In some embodiments, at least one R3 is independently cyclopropyl. In some embodiments, at least one R3 is independently cyclobutyl. In some embodiments, at least one R3 is independently cyclopentyl. In some embodiments, at least one R3 is independently cyclohexyl. In some embodiments, at least one R3 is independently cycloheptyl. In some embodiments, at least one R3 is independently cyclooctyl.
[0338] In some embodiments, each R3 is independently NH2, NH(C1-C6 alkyl), or N(C1-C6 alkyl)2.
[0339] In some embodiments, at least one R3 is NH2.
[0340] In some embodiments, each R3 is NH(C1-C6 alkyl) or N(C1-C6 alkyl)2.
[0341] In some embodiments, at least one R3 is NH(C1-C6 alkyl).
[0342] In some embodiments, at least one R3 is NH (methyl). In some embodiments, at least one R3 is NH (ethyl). In some embodiments, at least one R3 is NH (propyl). In some embodiments, at least one R3 is NH (butyl). In some embodiments, at least one R3 is NH (pentyl). In some embodiments, at least one R3 is NH (hexyl).
[0343] In some embodiments, at least one R3 is N(C1-C6 alkyl)2. In some embodiments, at least one R3 is N(methyl)2. In some embodiments, at least one R3 is N(ethyl)2. In some embodiments, at least one R3 is N(propyl)2. In some embodiments, at least one R3 is N(butyl)2. In some embodiments, at least one R3 is N(pentyl)2. In some embodiments, at least one R3 is N(hexyl)2.
[0344] In some embodiments, R3 is F or methyl.
[0345] In some embodiments, each R4 is independently a halogen, hydroxyl, cyano, nitro, C1-C6 alkyl, C2-C6 alkenyl, C2-C6 alkynyl, C1-C6 haloalkyl, C1-C6 alkoxyl, C1-C6 haloalkoxyl, C3-C8 cycloalkyl, NH2, NH(C1-C6 alkyl), N(C1-C6 alkyl)2, or C(O)NH2.
[0346] In some embodiments, each R4 is independently a halogen, hydroxyl, cyano, or nitro.
[0347] In some embodiments, at least one R4 is a halogen. In some embodiments, at least one R4 is F, Cl, Br, or I. In some embodiments, at least one R4 is F, Cl, or Br. In some embodiments, at least one R4 is F or Cl. In some embodiments, at least one R4 is F. In some embodiments, at least one R4 is Cl. In some embodiments, at least one R4 is Br. In some embodiments, at least one R4 is I.
[0348] In some embodiments, each R4 is independently hydroxyl, cyano, or nitro.
[0349] In some embodiments, at least one R4 is independently a hydroxyl group.
[0350] In some embodiments, at least one R4 is independently cyanoacrylate.
[0351] In some embodiments, at least one R4 is independently nitro.
[0352] In some embodiments, each R4 is independently a C1-C6 alkyl, C2-C6 alkenyl, C2-C6 alkynyl, C1-C6 haloalkyl, C1-C6 alkoxyl, C1-C6 haloalkoxyl, C3-C8 cycloalkyl, NH2, NH(C1-C6 alkyl), N(C1-C6 alkyl)2, or C(O)NH2, and the alkyl, alkenyl, alkynyl, alkoxyl, and cycloalkyl may be optionally substituted with one or more hydroxyls, 4-7 membered heterocyclines, containing one, two, or three heteroatoms independently selected from N, O, and S, NH2, NH(C1-C6 alkyl), or N(C6 alkyl).
[0353] In some embodiments, each R4 is independently a C1-C6 alkyl, C2-C6 alkenyl, C2-C6 alkynyl, C1-C6 haloalkyl, C1-C6 alkoxyl, C1-C6 haloalkoxyl, C3-C8 cycloalkyl, NH2, NH(C1-C6 alkyl), N(C1-C6 alkyl)2, or C(O)NH2.
[0354] In some embodiments, each R4 is independently a C1-C6 alkyl, C2-C6 alkenyl, or C2-C6 alkynyl, and the alkyl, alkenyl, or alkynyl may be optionally substituted with one or more hydroxyl, 4-7 membered heterocyclines containing one, two, or three heteroatoms independently selected from N, O, and S, NH2, NH(C1-C6 alkyl), or N(C1-C6 alkyl)2.
[0355] In some embodiments, each R4 is independently a C1-C6 alkyl, a C2-C6 alkenyl, or a C2-C6 alkynyl.
[0356] In some embodiments, R4 is a C1-C6 alkyl group. In some embodiments, R4 is methyl. In some embodiments, R4 is ethyl. In some embodiments, R4 is propyl. In some embodiments, R4 is butyl. In some embodiments, R4 is isopropyl. In some embodiments, R4 is isobutyl. In some embodiments, R4 is sec-butyl. In some embodiments, R4 is tert-butyl.
[0357] In some embodiments, R4 is a C1-C6 alkyl group that may be optionally substituted with one or more hydroxyl, 4- to 7-membered heterocyclines containing one, two, or three heteroatoms independently selected from N, O, and S, NH2, NH(C1-C6 alkyl), or N(C1-C6 alkyl)2.
[0358] In some embodiments, R4 is a C2-C6 alkenyl. In some embodiments, R4 is a C2 alkenyl. In some embodiments, R4 is a C3 alkenyl. In some embodiments, R4 is a C4 alkenyl. In some embodiments, R4 is a C5 alkenyl. In some embodiments, R4 is a C6 alkenyl.
[0359] In some embodiments, R4 comprises N, O, and one, two, or three heteroatoms independently selected from S, NH2, NH(C1-C6 alkyl), or N(C1-C6 alkyl)2. 、1 It is a C2-C6 alkenyl which may be optionally substituted with one or more hydroxyls or 4- to 7-membered heterocyclines.
[0360] In some embodiments, R4 is a C2-C6 alkynyl. In some embodiments, R4 is a C2 alkynyl. In some embodiments, R4 is a C3 alkynyl. In some embodiments, R4 is a C4 alkynyl. In some embodiments, R4 is a C5 alkynyl. In some embodiments, R4 is a C6 alkynyl.
[0361] In some embodiments, R4 comprises N, O, and one, two, or three heteroatoms independently selected from S, NH2, NH(C1-C6 alkyl), or N(C1-C6 alkyl)2. 、1 It is a C2-C6 alkynyl molecule that may be optionally substituted with one or more hydroxyls or 4- to 7-membered heterocyclines.
[0362] In some embodiments, each R4 is independently a C1-C6 alkoxyl or a C3-C8 cycloalkyl.
[0363] In some embodiments, each R4 is independently a C1-C6 alkoxyl or a C3-C8 cycloalkyl, and the alkoxyl and cycloalkyl may be optionally substituted with one or more hydroxyls, 4- to 7-membered heterocyclines, NH2, NH(C1-C6 alkyl), or N(C1-C6 alkyl)2.
[0364] In some embodiments, each R4 is independently a C1-C6 alkoxyl. In some embodiments, each R4 is independently a methoxyl. In some embodiments, each R4 is independently an ethoxyl. In some embodiments, each R4 is independently a propoxyl. In some embodiments, each R4 is independently a butioxyl. In some embodiments, each R4 is independently a pentioxyl. In some embodiments, each R4 is independently a hexoxyl.
[0365] In some embodiments, each R4 is a C1-C6 alkoxyl that may be independently and optionally substituted with one or more hydroxyls, 4- to 7-membered heterocyclines, NH2, NH(C1-C6 alkyl), or N(C1-C6 alkyl)2.
[0366] In some embodiments, each R4 is independently a C3-C8 cycloalkyl. In some embodiments, each R4 is independently a cyclopropyl. In some embodiments, each R4 is independently a cyclobutyl. In some embodiments, each R4 is independently a cyclopentyl. In some embodiments, each R4 is independently a cyclohexyl. In some embodiments, each R4 is independently a cycloheptyl. In some embodiments, each R4 is independently a cyclooctyl.
[0367] In some embodiments, each R4 is independently a C3-C8 cycloalkyl which may be optionally substituted with one or more hydroxyls, 4- to 7-membered heterocyclines, NH2, NH(C1-C6 alkyl), or N(C1-C6 alkyl)2.
[0368] In some embodiments, each R4 is independently a C1-C6 haloalkyl or C1-C6 haloalkoxyl.
[0369] In some embodiments, each R4 is independently a C1-C6 haloalkyl. In some embodiments, each R4 is independently a halomethyl. In some embodiments, each R4 is independently a haloethyl. In some embodiments, each R4 is independently a halopropyl. In some embodiments, each R4 is independently a halobutyl. In some embodiments, each R4 is independently a halopentyl. In some embodiments, each R4 is independently a halohexyl.
[0370] In some embodiments, each R4 is independently CF3, CHF2, or CH2F.
[0371] In some embodiments, each R4 is independently a C1-C6 haloalkoxyl. In some embodiments, each R4 is independently a halomethoxyl. In some embodiments, each R4 is independently a haloethoxyl. In some embodiments, each R4 is independently a halopropoxyl. In some embodiments, each R4 is independently a halobutoxyl. In some embodiments, each R4 is independently a halopentoxyl. In some embodiments, each R4 is independently a halohexoxyl.
[0372] In some embodiments, each R4 is independently NH2, NH(C1-C6 alkyl), N(C1-C6 alkyl)2, or C(O)NH2.
[0373] In some embodiments, each R4 is independently NH2, NH(C1-C6 alkyl), or N(C1-C6 alkyl)2.
[0374] In some embodiments, each R4 is independently NH2.
[0375] In some embodiments, each R4 is independently C(O)NH2.
[0376] In some embodiments, each R4 is independently NH(C1-C6 alkyl) or N(C1-C6 alkyl)2.
[0377] In some embodiments, each R4 is independently NH(C1-C6 alkyl). In some embodiments, each R4 is independently H(methyl). In some embodiments, each R4 is independently NH(ethyl). In some embodiments, each R4 is independently NH(propyl). In some embodiments, each R4 is independently NH(butyl). In some embodiments, each R4 is independently NH(pentyl). In some embodiments, each R4 is independently NH(hexyl).
[0378] In some embodiments, each R4 is independently N(C1-C6 alkyl)2. In some embodiments, each R4 is independently N(methyl)2. In some embodiments, each R4 is independently N(ethyl)2. In some embodiments, each R4 is independently N(propyl)2. In some embodiments, each R4 is independently N(butyl)2. In some embodiments, each R4 is independently N(pentyl)2. In some embodiments, each R4 is independently N(hexyl)2.
[0379] In some embodiments, each R4 is independently methyl, ethyl, F, or CF3.
[0380] In some embodiments, each R4 is independently methyl or ethyl.
[0381] In some embodiments, each R4 is independently F or CF3.
[0382] In some embodiments, R5 is H.
[0383] In some embodiments, R5 is a C1-C6 alkyl, C1-C6 haloalkyl, C2-C6 alkenyl, C2-C6 alkynyl, C1-C6 alkoxyl, C3-C8 cycloalkyl, or heterocyclyl, where the heterocyclyl is a 4- to 7-membered ring containing 1, 2, or 3 heteroatoms independently selected from N, O, and S, and the cycloalkyl and heterocyclyl may be optionally substituted with one or more halogens 'C1-C6 alkyl' or hydroxyl'.
[0384] In some embodiments, R5 is a C1-C6 alkyl, a C2-C6 alkenyl, or a C2-C6 alkynyl.
[0385] In some embodiments, R5 is a C1-C6 alkyl group. In some embodiments, R5 is methyl. In some embodiments, R5 is ethyl. In some embodiments, R5 is propyl. In some embodiments, R5 is butyl. In some embodiments, R5 is isopropyl. In some embodiments, R5 is isobutyl. In some embodiments, R5 is sec-butyl. In some embodiments, R5 is tert-butyl.
[0386] In some embodiments, R5 is a C2-C6 alkenyl. In some embodiments, R5 is a C2-C6 alkenyl. In some embodiments, R5 is a C2 alkenyl. In some embodiments, R5 is a C3 alkenyl. In some embodiments, R5 is a C4 alkenyl. In some embodiments, R5 is a C5 alkenyl. In some embodiments, R5 is a C6 alkenyl.
[0387] In some embodiments, R5 is C2-C6 alkynyl. In some embodiments, R5 is C2-C6 alkynyl. In some embodiments, R5 is C2 alkynyl. In some embodiments, R5 is C3 alkynyl. In some embodiments, R5 is C4 alkynyl. In some embodiments, R5 is C5 alkynyl. In some embodiments, R5 is C6 alkynyl.
[0388] In some embodiments, R5 is a C1-C6 alkoxyl or a C3-C8 cycloalkyl.
[0389] In some embodiments, R5 is independently a C1-C6 alkoxyl. In some embodiments, R5 is independently a methoxyl. In some embodiments, R5 is independently an ethoxyl. In some embodiments, R5 is independently a propoxyl. In some embodiments, R5 is independently a butoxyl. In some embodiments, R5 is independently a pentoxyl. In some embodiments, R5 is independently a hexoxyl.
[0390] In some embodiments, R5 is a C3-C8 cycloalkyl group. In some embodiments, R5 is a cyclopropyl group. In some embodiments, R5 is a cyclobutyl group. In some embodiments, R5 is a cyclopentyl group. In some embodiments, R5 is a cyclohexyl group. In some embodiments, R5 is a heptyl group. In some embodiments, R5 is a cyclooctyl group.
[0391] In some embodiments, R5 is a heterocyclyl, which is a 4- to 7-membered ring containing one, two, or three heteroatoms independently selected from N, O, and S, and the cycloalkyl and heterocyclyl may be optionally substituted with one or more halogens, C1-C6 alkyl, or hydroxyl. In some embodiments, R5 is a heterocyclyl, which is a 4- to 6-membered ring containing one or two heteroatoms independently selected from N and O, and the cycloalkyl and heterocyclyl may be optionally substituted with one or more halogens, C1-C6 alkyl, or hydroxyl. In some embodiments, R5 is tetrahydropyranyl or tetrahydrofuranyl.
[0392] In some embodiments, R5 is H, C 1~6The molecule is an alkyl or heterocyclyl, where the heterocyclyl is a 4- to 7-membered ring containing one or two heteroatoms independently selected from N and O.
[0393] In some embodiments, R5 is H. In some embodiments, R6 is H. In some embodiments, both R5 and R6 are H.
[0394] In some embodiments of the compound of formula I, n is 0, 1, 2, 3, 4, or 5. In some embodiments of the compound of formula I, n is 0, 1, 2, 3, or 4. In some embodiments of the compound of formula I, n is 0, 1, 2, or 3. In some embodiments of the compound of formula I, n is 0, 1, or 2. In some embodiments of the compound of formula I, n is 0 or 1. In some embodiments of the compound of formula I, n is 0. In some embodiments of the compound of formula I, n is 1. In some embodiments of the compound of formula I, n is 2. In some embodiments of the compound of formula I, n is 3. In some embodiments of the compound of formula I, n is 4. In some embodiments of the compound of formula I, n is 5.
[0395] The splice modifier of this disclosure comprises the structure of formula (II), [ka] or comprising a pharmaceutically acceptable salt, solvate, or prodrug thereof, in the formula, A is a saturated or partially unsaturated monocyclic or bicyclic 4- to 9-membered heterocycloalkyl or NR1R2, where the heterocycloalkyl contains one or two nitrogen ring atoms and is optionally substituted with 1, 2, 3, or 4 R6 atoms. R1 is a heterocycloalkyl group containing one nitrogen ring atom which may be optionally substituted with 1, 2, 3, or 4 R6 atoms. R2 is hydrogen, C 1-7 Alkyl, or C 3-8 It is a cycloalkyl, R3 stands for H, Halo, C 1-7Alkyl, OR5, N(R5)2, C 3-8 It is a cycloalkyl or heterocycloalkyl, R4 is an aryl or bicyclic 9-membered heteroaryl containing 2, 3, or 4 heteroatoms independently selected from N, O, and S, where R4 is optionally substituted with 1, 2, or 3 R7 atoms. Each R5 is independent, C 1-7 Alkyl, C 3-8 It is a cycloalkyl or heterocycloalkyl, Each R6 is a halogen, hydroxyl, cyano, -COOH, -C(O)-C1-C6 alkyl, -C(O)O-C1-C6 alkyl, C1-C7 alkyl, C1-C8 heteroalkyl, C 1-7 Alkoxy-heterocycloalkyl, C2-C6 alkenyl, C2-C6 alkynyl, C1-C6 alkoxy, -(CH2) 0-2 A group independently selected from the group consisting of -C3-C8 cycloalkyl, 4-7 membered monocyclic heterocycloalkyl, NH2, NH(C1-C6 alkyl), N(C1-C6 alkyl)2, -NHC(O)-C1-C6 alkyl, -N(C1-C6 alkyl)-C(O)-C1-C6 alkyl, -C(O)-NH2, -C(O)-NH(C1-C6 alkyl), and -C(O)-N(C1-C6 alkyl)2, wherein alkyl, alkenyl, alkynyl, and alkoxy may be optionally substituted with one or more halogens, hydroxyl, or NH2, and cycloalkyl and heterocycloalkyl may be optionally substituted with one or more halogens, hydroxyl, C1-C6 alkyl, C1-C6 heteroalkyl, C1-C6 alkoxy, or NH2, or Two Rs on the same carbon 6 Can they be combined as keto (=O), or Both R6s are C 1-7 Forms alkylene, Each R7 independently controls Halo, Cyano, and C 1-7 Alkyl, C 1-7 Haloalkyl, C 1-7 Alkoxy, C 1-7 Haloalkoxy, or C 3-8It is a cycloalkyl, C 1-7 Alkyl groups may be optionally substituted with OH groups. R 16 H, Halo, C 1-7 Alkyl, OR5, N(R5)2, C 3-8 It is a cycloalkyl or heterocycloalkyl, R 17 H, Halo, C 1-7 Alkyl, OR5, N(R5)2, C 3-8 A composition that is cycloalkyl or heterocycloalkyl.
[0396] The splice modifier of this disclosure comprises the structure of formula (III), [ka] or comprising a pharmaceutically acceptable salt, solvate, or prodrug thereof, in the formula, A is a saturated or partially unsaturated monocyclic or bicyclic 4- to 9-membered heterocycloalkyl or NR1R2, where the heterocycloalkyl contains one or two nitrogen ring atoms and is optionally substituted with 1, 2, 3, or 4 R6 atoms. R1 is a heterocycloalkyl group containing one nitrogen ring atom which may be optionally substituted with 1, 2, 3, or 4 R6 atoms. R2 is hydrogen, C 1-7 Alkyl, or C 3-8 It is a cycloalkyl, R3 stands for H, Halo, C 1-7 Alkyl, OR5, N(R5)2, C 3-8 It is a cycloalkyl or heterocycloalkyl, R4 is an aryl or bicyclic 9-membered heteroaryl containing 2, 3, or 4 heteroatoms independently selected from N, O, and S, where R4 is optionally substituted with 1, 2, or 3 R7 atoms. Each R5 is independent, C 1-7 Alkyl, C 3-8 It is a cycloalkyl or heterocycloalkyl, Each R6 independently, C 1-7Alkyl, amino, amino-C 1-7 Alkyl, C 3-8 Cycloalkyl, heterocycloalkyl, or C 1-7 Is alkoxy-heterocycloalkyl, or two R6s together form C 1-7 An alkylene, Each R7 is independently halo, cyano, C 1-7 Alkyl, C 1-7 Haloalkyl, C 1-7 Alkoxy, C 1-7 Haloalkoxy, or C 3-8 Cycloalkyl, and C 1-7 The alkyl may optionally be substituted with OH.
[0397] In some embodiments, the compound is selected from prodrugs of the compounds described in Table 3 and pharmaceutically acceptable salts thereof.
Table 3-1
Table 3-2
Table 3-3
Table 3-4
Table 3-5
Table 3-6
Table 3-7
Table 3-8
Table 3-9
Table 3-10
[0398] In some embodiments, the compound is selected from the group consisting of 1A to 192A and 100B to 135B. In some embodiments, the compound is selected from the group including 3A, 6A, 8A, 10A, 15A, 24A, 86A, 100B, 111B, 117B, 121B, 135B, and 192A. In some embodiments, the compound is compound 116B. In some embodiments, the compound is compound 100B. In some embodiments, the compound is compound 1A. In some embodiments, the compound is compound 24A. In some embodiments, the compound is compound 2A. In some embodiments, the compound is compound 22A. In some embodiments, the compound is compound 34A.
[0399] To avoid any doubt, it should be understood that where a group is defined as “as described herein,” that group encompasses the first and broadest definition, as well as each and all of the specific definitions relating to that group.
[0400] The various functional groups and substituents constituting the compound of formula (I) are typically selected so that the molecular weight of the compound does not exceed 1000 daltons. More commonly, the molecular weight of the compound is less than 900, for example, less than 800, or less than 750, or less than 700, or less than 650 daltons. More conveniently, the molecular weight is less than 600, for example, 550 daltons or less.
[0401] Any compound of any of the formulas disclosed herein and any pharmaceutically acceptable salt thereof will be understood to include stereoisomers of all isomers of the compound, or mixtures of stereoisomers.
[0402] It should be understood that any compound of any formula described herein includes, where applicable, the compound itself, as well as salts thereof and solvates thereof.
[0403] The in vivo effect of any one of the compounds of the formulas disclosed herein may be partially exerted by one or more metabolites formed in the body of a human or animal after administration of any one of the compounds of the formulas disclosed herein. As mentioned above, the in vivo effect of any one of the compounds of the formulas disclosed herein may also be exerted by the metabolism of a precursor compound (prodrug).
[0404] It is appropriate to exclude any individual compound that does not possess the biological activity as defined herein.
[0405] In some embodiments, the target protein, RNA, miRNA, or shRNA (for example, a protein, RNA, miRNA, or shRNA encoded by one of the genes described above) is expressed after co-administration of (i) the encoding nucleic acid molecule or vector and (ii) a splice regulator described herein.
[0406] The splice modifier is administered to the subject, for example, in a therapeutically effective dose. In some embodiments, administering a therapeutically effective dose of the splice modifier to the subject results in the inclusion of one of two or more exons, one or more exons, a first exon, or a second exon. In some embodiments, one of two or more exons, or one or more exons, is a second exon.
[0407] In some embodiments, the nucleic acid molecules of the Disclosure include a regulatory site that can be recognized by a spliceosome complex, and in some embodiments, stabilization of the interaction between the spliceosome complex and the regulatory site that can be recognized by the spliceosome complex occurs upon target administration of a therapeutically effective amount of the splice regulator.
[0408] In some embodiments, splice modifiers are tested for their ability to modify splicing so that a net increase in the inclusion of two-site stop codons and transgene expression is produced. Screens may utilize DNA constructs containing minigenes, which may be assembled via DNA synthesis and molecular cloning techniques known in the art and inserted into mammalian cells using electroporation, chemical transfection, virus-mediated integration, or virus-mediated episomal delivery. In some embodiments, candidate splicing regulators are evaluated by fusing a luminescent enzyme (e.g., firefly luciferase, nanolucinol luciferase, renyral luciferase, or Gaussian luciferase), a fluorescent protein (e.g., green fluorescent protein, blue fluorescent protein, or red fluorescent protein), or a colorimetric enzyme (e.g., beta-lactamase, or secreted embryonic alkaline phosphatase) to a DNA construct such that the construct generates a quantifiable signal that is proportional to the splicing of exons (e.g., the second exon) and is readable, for example, by flow cytometry, microscopy, or a multimode microplate reader using a photomultiplier tube. Alternatively, candidate splicing regulator-dependent biological activity is evaluated by sequencing mRNA transcripts produced by the DNA construct using RNASeq, or by quantitative reverse transcription PCR. In some embodiments, following such exemplary screening, candidate splice regulators are applied to cells containing DNA constructs for up to 6, 12, 24, or 72 hours, and then evaluated for start codon-containing activity.
[0409] In some embodiments, splice regulators are tested by screening various substances and / or chemical derivatives of substances for their ability to regulate the expression of a reporter gene linked to a minigene construct in mammalian cells. The minigene template may be derived from a naturally occurring mammalian gene or from a synthetic intron-containing construct having standard 5' and 3' splice site sequences. In some embodiments, the minigene template is linked to a reporter gene, such as firefly luciferase, to enable the detection of changes in gene expression resulting from chemicals that can act as splice regulators to mediate alternative splicing of the minigene.
[0410] In some embodiments, splice regulators are tested for their ability to bind to a segment of RNA corresponding to a segment of a transcribed minigene. In some embodiments, the RNA segment screened is a regulator or splice regulator binding site that controls the splicing of an exon (e.g., the second exon) in a minigene. In some embodiments, the RNA segment is generated, for example, by in vitro transcription, by chemoRNA synthesis, or by transient or stable expression in cells. In some embodiments, evaluation of splice regulators that bind to regulatory RNA segments is assessed by biophysical methods such as affinity-selective mass spectrometry, affinity chromatography, or similar techniques relating to the detectable association of splice regulators with RNA or RNA-protein complexes, where differences in liquid-phase migration or movement of the splice regulator are assessed. In one embodiment, the biophysical method includes assessing the liquid-phase migration or movement of the splice regulator. In some embodiments, evaluation of splice regulators that bind to regulatory RNA segments is also assessed by functional methods. In one embodiment, the functional method includes assessing the minigene reporter expression affected by the splice regulator. In one embodiment, the splice modifier is modified in a detectable manner.
[0411] In some embodiments, the splice regulator binding site is screened for various nucleic acid modifications (e.g., single and / or multiple point mutations, insertions, and / or deletions) to the minigene template. The minigene template may be derived from a spontaneously occurring mammalian gene or a synthetic intron-containing construct having standard 5' and 3' splice site sequences. In some embodiments, the minigene template is linked to a reporter gene, such as firefly luciferase, to enable detection of changes in gene expression resulting from alternative splicing of the minigene by a splice regulator. In some embodiments, single and / or multiple point mutations, insertions, and / or deletions are introduced into the minigene template and assayed in mammalian cells to generate a combined library that selects one or more minigene sequences that optimally respond to a splice regulator and control the expression of the reporter gene.
[0412] In some embodiments, the splice regulator binds to a segment of the RNA-binding protein and / or splice regulator binding site of any one nucleic acid of the embodiment.
[0413] In some embodiments, the splice regulator binds to a segment of the RNA-binding protein and / or splice regulator binding site of any one nucleic acid of the embodiment.
[0414] In some embodiments, the splice regulator binds to RBP. In some embodiments, the expression of such RBP is tissue-specific. In some embodiments, RBP is exclusively expressed in the brain or any other specific tissue of the subject. In accordance with this principle, in some embodiments, the nucleic acid molecule of the Disclosure operates by the principle of exon inclusion by using a splice regulator that targets RBP having tissue-specific expression. For example, the nucleic acid molecule is ligated to a tissue-specific promoter that matches the expression profile of the RBP of the subject.
[0415] In some embodiments, the RBP is a splicing enhancer (e.g., a serine and arginine-rich (SR) protein). In some embodiments, the RBP is a splicing repressor (e.g., a heterogeneous nuclear ribonucleoprotein (hnRNP)).
[0416] Synthesis method As an example only, a scheme for preparing the splicing modifiers described herein is provided.
[0417] In some embodiments, a scheme for preparing splicing modifiers is described in Scheme 1. [ka]
[0418] In some embodiments, a scheme for preparing splicing modifiers is described in Scheme 2. [ka]
[0419] The compounds of this disclosure can be prepared by any suitable technique known in the art. Specific processes for the preparation of these compounds are further described in the attached examples.
[0420] In some embodiments, the synthesis of any of the compounds disclosed herein, compounds 1A-192A and 100B-135B, is described in PCT / US2022 / 080352 (WO2023092149) or PCT / US2022 / 079748 (WO2023086959), the contents of which are incorporated herein by reference.
[0421] It will be understood that in the description of the synthesis method described herein and any reference synthesis method used to prepare the starting materials, all proposed reaction conditions, including the choice of solvent, reaction atmosphere, reaction temperature, duration of the experiment, and work-up procedure, can be selected by those skilled in the art.
[0422] Those skilled in organic synthesis will understand that the functionalities present in various parts of a molecule must be compatible with the reagents and reaction conditions used.
[0423] It will be understood that during the synthesis of the compounds of this disclosure in the processes defined herein, or during the synthesis of certain starting materials, it may be desirable to protect certain substituents to prevent undesirable reactions. Those skilled in the art will understand when such protection is necessary and how such protecting groups may be placed in place and subsequently removed. For examples of protecting groups, see one of the many general texts on the subject, e.g., 'Protective Groups in Organic Synthesis' by Theodora Green (publisher: John Wiley & Sons). Protecting groups may be removed by any simple method known to those skilled in the art, where described in the literature or appropriate for the removal of the protecting group in question, and such method is selected to result in the removal of the protecting group with minimal interference to other groups in the molecule. Thus, when reactants contain groups such as amino, carboxy, or hydroxy, it may be desirable to protect these groups in some of the reactions referred to herein.
[0424] As an example, suitable protecting groups for amino groups or alkylamino groups include, for example, acyl groups, such as alkanoyl groups like acetyl; alkoxycarbonyl groups, such as methoxycarbonyl, ethoxycarbonyl, or t-butoxycarbonyl groups; arylmethoxycarbonyl groups, such as benzyloxycarbonyl; or aroyl groups, such as benzoyl. The deprotection conditions for the above protecting groups inevitably change depending on the choice of protecting group. Therefore, for example, acyl groups such as alkanoyl or alkoxycarbonyl groups or aroyl groups can be removed by hydrolysis using a suitable base such as lithium or an alkali metal hydroxide such as sodium hydroxide. Alternatively, acyl groups such as tert-butoxycarbonyl groups may be removed by treatment with a suitable acid such as hydrochloric acid, sulfuric acid, phosphoric acid, or trifluoroacetic acid, and arylmethoxycarbonyl groups such as benzyloxycarbonyl groups may be removed by hydrogenation on a catalyst such as palladium carbon, or by treatment with a Lewis acid such as boron tris(trifluoroacetic acid). Suitable alternative protecting groups for primary amino groups include, for example, alkylamines such as dimethylaminopropylamine, or phthaloyl groups that can be removed by treatment with hydrazine.
[0425] Suitable protecting groups for hydroxyl groups include, for example, acyl groups, such as alkanoyl groups like acetyl, alloyl groups like benzoyl, or arylmethyl groups like benzyl. The deprotection conditions for the above protecting groups inevitably change depending on the choice of protecting group. Therefore, for example, acyl groups such as alkanoyl or alloyl groups can be removed by hydrolysis using a suitable base such as an alkali metal hydroxide such as lithium, sodium hydroxide, or ammonia. Alternatively, arylmethyl groups such as benzyl groups may be removed by hydrogenation on a catalyst such as palladium-carbon.
[0426] Suitable protecting groups for the carboxyl group include, for example, esterifying groups such as methyl or ethyl groups which can be removed by hydrolysis with a base such as sodium hydroxide, or tert-butyl groups which can be removed by treatment with an acid such as an organic acid such as trifluoroacetic acid, or a benzyl group which can be removed by hydrogenation on a catalyst such as palladium carbon.
[0427] Once the structure of formula (I) is synthesized by any one of the processes defined in this disclosure, the process may then further include the additional steps of (i) removing any protecting groups present, (ii) converting the structure of formula (I) to another structure of formula (I) to form a pharmaceutically acceptable salt, hydrate, or solvate thereof, and / or (iv) forming a prodrug thereof.
[0428] In some embodiments, the compound obtained as a result of formula (I) is isolated and purified using techniques well known in the art.
[0429] Conveniently, the reaction of the compounds takes place in the presence of a suitable solvent, which is preferably inert under the respective reaction conditions. Examples of suitable solvents include hydrocarbons such as hexane, petroleum ether, benzene, toluene, or xylene; chlorinated hydrocarbons such as trichloroethylene, 1,2-dichloroethane, tetrachloromethane, chloroform, or dichloromethane; alcohols such as methanol, ethanol, isopropanol, n-propanol, n-butanol, or tert-butanol; ethers such as diethyl ether, diisopropyl ether, tetrahydrofuran (THF), 2-methyltetrahydrofuran, cyclopentyl methyl ether (CPME), methyl tert-butyl ether (MTBE), or dioxane; and ethylene glycoside. Examples of solvents include, but are not limited to, glycol ethers such as monomethyl or monoethyl ether or ethylene glycol dimethyl ether (diglym), ketones such as acetone, methyl isobutyl ketone (MIBK), or butanone, amides such as acetamide, dimethylacetamide, dimethylformamide (DMF), or N-methylpyrrolidinone (NMP), nitriles such as acetonitrile, sulfoxides such as dimethyl sulfoxide (DMSO), nitro compounds such as nitromethane or nitrobenzene, esters such as ethyl acetate or methyl acetate, or mixtures of solvents or mixtures with water.
[0430] The reaction temperature is appropriately between approximately -100°C and 300°C, depending on the reaction process and the conditions under which it is used.
[0431] The reaction time generally ranges from one minute to several days, depending on the reactivity of each compound and the respective reaction conditions. A suitable reaction time can be easily determined by methods known in the art, such as reaction monitoring. Based on the reaction temperature mentioned above, a suitable reaction time is generally in the range of 10 minutes to 48 hours.
[0432] Furthermore, in some embodiments, additional compounds of the Disclosure can be readily prepared by utilizing the procedures described herein, in conjunction with ordinary skill in the art. In some embodiments, those skilled in the art will readily understand that known variations of the conditions and processes of the following preparation procedures are used to prepare these compounds.
[0433] As will be understood by those skilled in the art of organic synthesis, the compounds of this disclosure are readily accessible by a variety of synthetic routes, some of which are illustrated in the accompanying examples. Those skilled in the art will readily recognize what kinds of reagents and reaction conditions to use to obtain the compounds of this disclosure, and how they are applied and adapted. Furthermore, in some embodiments, some of the compounds of this disclosure are readily synthesized by reacting them with other compounds of this disclosure under preferred conditions, for example, by converting a particular functional group present in the compounds of this disclosure or in a preferred precursor molecule to another by applying standard synthetic methods, such as oxidation, addition, or substitution, which are well known to those skilled in the art. Similarly, those skilled in the art may apply synthetic protecting (or protecting) groups as needed or useful, and preferred protecting groups and methods for introducing and removing them are well known to those skilled in the art of chemical synthesis, and are described in more detail, for example, PGM Wuts, TW Greene, “Greene's Protective Groups in Organic Synthesis”, 4th edition (2006) (John Wiley & Sons).
[0434] A common route for preparing the compounds of this application is described in this disclosure.
[0435] nucleic acid vectors In some embodiments, effective intracellular concentrations of the nucleic acid molecules disclosed herein are also achieved through the stable expression of a vector encoding the nucleic acid molecule, such as a nucleic acid molecule containing a minigene linked to the transgene, as described herein (e.g., by integration into the nucleus or mitochondrial genome of a mammalian cell), where the minigene comprises a first exon, a first intron, a second exon, a start codon, and a second intron. To introduce such nucleic acid molecules into mammalian cells, the nucleic acid molecules can be incorporated into a vector.
[0436] Vectors can be introduced into cells by a variety of methods, including transformation, transfection, direct uptake, projectile shock, and encapsulation of the vector in liposomes. Examples of suitable methods for transfecting or transforming cells include calcium phosphate precipitation, electroporation, microinjection, infection, lipofection, and direct uptake. These methods are described in detail, for example, Green et al., Molecular Cloning: A Laboratory Manual, Fourth Edition (Cold Spring Harbor University Press, New York (2014)) and Ausubel et al., Current Protocols in Molecular Biology (John Wiley & Sons, New York (2015)).
[0437] In some embodiments, the nucleic acid molecules disclosed herein are introduced into mammalian cells by targeting cell membrane phospholipids with a vector containing a polynucleotide encoding such nucleic acid molecules. For example, in some embodiments, the vector targets phospholipids on the extracellular surface of the cell membrane by ligating the vector molecule to a VSV-G protein, which is a viral protein having affinity for all cell membrane phospholipids. In some embodiments, the constructs are prepared using conventional and routine methods of the present technology. In addition to achieving high-speed transcription and translation, stable expression of exogenous polynucleotides in mammalian cells can be achieved by incorporating gene-containing polynucleotides into the nuclear genome of mammalian cells. Various vectors have been developed for the delivery and integration of polynucleotides encoding exogenous proteins or RNA products (e.g., miRNA or shRNA) into the nuclear DNA of mammalian cells. Examples of expression vectors are disclosed, for example, in WO1994 / 011026. Expression vectors for use in the compositions and methods described herein contain a polynucleotide sequence encoding a nucleic acid molecule, as well as additional sequence elements used, for example, for the expression of these nucleic acid molecules and / or the integration of these polynucleotide sequences into the genome of mammalian cells. In some embodiments, the specific vectors used include plasmids containing regulatory sequences such as promoter and enhancer regions that induce gene transcription. Other useful vectors contain polynucleotide sequences that enhance the translation rate of these genes or improve the stability or nuclear transport of mRNA resulting from gene transcription. These sequence elements include, for example, 5' and 3' UTR regions, internal ribosome entry sites (IRESs), and poly(A) to guide the efficient transcription of genes supported on the expression vector. Expression vectors suitable for use in the compositions and methods described herein may also contain polynucleotides encoding markers for selecting cells containing such vectors. Examples of suitable markers are genes encoding resistance to antibiotics such as ampicillin, chloramphenicol, kanamycin, and norseotricin.
[0438] Viral vector Viral genomes provide a rich vector source that can be used for the efficient delivery of exogenous polynucleotides to mammalian cells. Viral genomes are particularly useful vectors for gene delivery because the polynucleotides contained within such genomes are typically incorporated into the nuclear genome of mammalian cells by generalized or specific transduction. These processes occur as part of the natural viral replication cycle and do not require added proteins or reagents to induce gene integration. Examples of viral vectors include parvoviruses (e.g., adeno-associated virus (AAV)), retroviruses (e.g., retroviridae virus vectors), adenoviruses (e.g., Ad5, Ad26, Ad34, Ad35, and Ad48), coronaviruses, orthomyxoviruses (e.g., influenza virus), rhabdoviruses (e.g., rabies and vesicular stomatitis virus), paramyxoviruses (e.g., measles and senai virus), positive-strand RNA viruses (such as picornaviruses and alphaviruses) and double-stranded DNA viruses including adenoviruses, herpesviruses (e.g., herpes simplex virus types 1 and 2, Epstein-Barr virus, cytomegalovirus), and poxviruses (e.g., vaccinia, modified vaccinia ankara (MVA), poultry and capillary). Other viruses include, for example, Norwalk virus, togavirus, flavivirus, reovirus, papovavirus, hepadnavirus, human papillomavirus, human foam virus, and hepatitis viruses. Examples of retroviruses include avian leukemia sarcoma, avian type C virus, mammalian type C, B, and D viruses, oncotrovirus, HTLV-BLV group, lentivirus, alpha retrovirus, gamma retrovirus, and spumavirus (Coffin, JM, Retroviridae: The viruses and their replication, Virology, Third Edition (Lippincott-Raven, Philadelphia, (1996))).Other examples include mouse leukemia virus, mouse sarcoma virus, mouse mammary tumor virus, bovine leukemia virus, feline leukemia virus, feline sarcoma virus, avian leukemia virus, human T-cell leukemia virus, sheep endogenous virus, Gibbon ape leukemia virus, nucleic acid molecule Park monkey virus, simian immunodeficiency virus, simian sarcoma virus, Simian virus 40 (SV40), Rous sarcoma virus, and lentivirus. Other examples of vectors are described, for example, in McVey et al., (U.S. Patent No. 5,801,030).
[0439] Retrovirus vectors The delivery vectors used in the methods and compositions described herein may be retroviral vectors. Retroviruses may be selected as gene delivery vectors because of their ability to integrate their genes into the host genome, their ability to transfer large amounts of exogenous genetic material, their ability to infect a wide range of species and cell types, and their ability to be packaged into specific cell lines. Furthermore, retroviral vectors can infect a wide variety of cell types.
[0440] One type of retroviral vector that may be used in the methods and compositions described herein is a lentiviral vector. Lentiviral vectors (LVs), a subset of retroviruses, efficiently transduce a wide range of dividing and non-dividing cell types, resulting in stable, long-term expression of polynucleotides. For an overview of optimization strategies for LV packaging and transduction, see Delenda, The Journal of Gene Medicine 6: S125 (2004). The use of lentiviral-based gene transfer techniques relies on the in vitro production of recombinant lentiviral particles carrying a highly deletable viral genome containing the polynucleotide of interest. In particular, the recombinant lentivirus is restored via trans co-expression in a cell line tolerant of a transfer vector into which the sequence to be expressed is inserted, comprising (1) a vector expressing a Gag-Pol precursor together with a packaging construct, i.e., Rev (which is alternatively expressed in trans), (2) a vector expressing an envelope receptor, which is generally heterogeneous in nature, and (3) viral cDNA that has all open reading frames deleted but maintains the sequences necessary for replication, capsid formation, and expression. The LV used in the methods and compositions described herein may consist of one or more of the following: a 5'-long-term repeat (LTR), an HIV signaling sequence, an HIV Psi signaling 5'-splice site (SD), a delta-GAG element, a Rev Responsive element (RRE), a 3'-splice site (SA), an elongation factor (EF) 1-alpha promoter, and a 3'-self-inactivating LTR (SIN-LTR). The lentiviral vector optionally includes a central polypurine tube (cPPT) and a Woodchuck hepatitis virus post-transcriptional regulatory element (WPRE), as described in US 6,136,597. The lentiviral vector may further include a pHR' skeleton, which may include, for example, the configuration described herein.Lentigen LV, as described in Lu et al., Journal of Gene Medicine 6:963 (2004), may be used to express DNA molecules and / or transduce cells.
[0441] The LVs used in the methods and compositions described herein may include a 5'-long-term repeat (LTR), an HIV signal sequence, an HIV Psi signal 5'-splice site (SD), a delta-GAG element, a Rev response element (RRE), a 3'-splice site (SA), an elongation factor (EF) 1-alpha promoter, and a 3'-self-inactivating LTR (SIN-LTR). Optionally, one or more of these regions are replaced with another region performing a similar function. Enhancer elements can be used to increase the expression of modified DNA molecules or to increase lentiviral integration efficiency. The LVs used in the methods and compositions described herein may also include a nef sequence.
[0442] The LV used in the methods and compositions described herein may be configured to include a cPPT sequence that enhances vector integration. The cPPT acts as a secondary source of (+) strand DNA synthesis, introducing a partial strand duplication into the middle of its native HIV genome. Introducing the cPPT sequence into the import vector backbone resulted in a strong increase in the total amount of genome integrated into the DNA of the target cell through nuclear transport.
[0443] The LV used in the methods and compositions described herein may be a configuration comprising a Woodchuck post-transcriptional regulator (WPRE). WPRE acts at the transcriptional level by promoting nuclear transport of transcripts and / or by increasing the efficiency of polyadenylation of nascent transcripts, and thus increases the total amount of mRNA in cells. Addition of WPRE to the LV results in a substantial improvement in the level of polynucleotide expression from several different promoters, both in vitro and in vivo. The LV used in the methods and compositions described herein may be a configuration comprising both a cPPT sequence and a WPRE sequence.
[0444] The vector may also be configured to include an IRES sequence that enables the expression of multiple polypeptides from a single promoter. In addition to the IRES sequence, other elements that enable the expression of multiple polynucleotides are useful. The vectors used in the methods and compositions described herein may be configured to include multiple promoters that enable the expression of multiple polynucleotides.
[0445] Other elements that enable the expression of multiple polynucleotides to be identified in the future are useful and may be utilized in vectors suitable for use in the compositions and methods described herein.
[0446] The vectors used in the methods and compositions described herein may be clinical-grade vectors. Therefore, retroviral vectors may be used in conjunction with the disclosed methods and compositions.
[0447] To construct a retroviral vector, the nucleic acid encoding the target gene is inserted into the viral genome in place of a specific viral sequence to produce a replication-deficient virus. To produce virions, a packaging cell line is constructed that contains the gag, pol, and / or env genes but does not contain the LTR and / or packaging components. When a recombinant plasmid containing cDNA is introduced into this cell line along with the retroviral LTR sequence and packaging sequence (e.g., by calcium phosphate precipitation), the packaging sequence allows the RNA transcript of the recombinant plasmid to be packaged into viral particles, which are then secreted into the culture medium. The medium containing the recombinant retrovirus is then collected, optionally concentrated, and used for gene transfer.
[0448] AAV Vector In some embodiments, the nucleic acid molecules described herein are incorporated into recombinant AAV (rAAV) vectors to facilitate their introduction into cells, such as target cells. rAAV vectors useful in conjunction with the compositions and methods described herein include a recombinant nucleic acid construct containing (1) a nucleic acid molecule and (2) a nucleic acid that promotes the promotion and expression of a heterologous gene. The viral nucleic acid may be configured to include the AAV sequence required in cis for the replication and packaging of DNA into virions (e.g., functional ITR). Such rAAV vectors may also contain a marker or reporter gene.
[0449] Useful rAAV vectors include those in which one or more naturally occurring AAV genes are deleted in whole or in part, but which retain functionally adjacent ITR sequences. The AAV ITR may be from any serotype suitable for a particular application (e.g., derived from serotype 2 or 5). Methods for using rAAV vectors are described, for example, in Tal et al., J. Biomed. Sci. 7:279-291 (2000), and Monahan and Samulski, Gene Delivery 7:24-30 (2000). These disclosures are incorporated herein by reference as relating to AAV vectors for gene delivery.
[0450] In some embodiments, the nucleic acids and vectors described herein are incorporated into an rAAV virion to facilitate the introduction of the nucleic acid or vector into a cell. The AAV capsid protein consists of the external non-nucleic acid portion of the virion and is encoded by the AAVcap gene. The cap gene encodes three viral coat proteins, VP1, VP2, and VP3, necessary for virion assembly. Construction of the rAAV virion is described, for example, in U.S. Patents 5,173,414, 5,139,941, 5,863,541, 5,869,305, 6,057,152, and 6,376,237, as well as in Rabinowitz et al., J. Virol. 76:791-801 (2002) and Bowles et al., J. Virol. 77:423-432 (2003).
[0451] rAAV virions useful in conjunction with the compositions and methods described herein include rAAV virions derived from various AAV serotypes, including AAV 1, 2, 3, 4, 5, 6, 7, 8, and 9. The construction and use of AAV vectors and AAV proteins of different serotypes are described, for example, in Chao et al., Mol. Ther. 2:619-623 (2000), Davidson et al., Proc. Natl. Acad. Sci. USA 97:3428-3432 (2000), Xiao et al., J. Virol. 72:2224-2232 (1998), Halbert et al., J. Virol. 74:1524-1532 (2000), Halbert et al., J. Virol. 75:6615-6624 (2001), and Auricchio et al., Hum. Molec. Genet. 10:3075-3081 (2001).
[0452] Also useful in conjunction with the compositions and methods described herein are pseudotyped rAAV vectors. Pseudotyped vectors include AAV vectors of a given serotype that are pseudotyped with a capsid gene derived from a serotype other than a given serotype (e.g., AAV1, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, or AAV9). For example, a typical pseudotyped vector is an AAV2 vector encoding a therapeutic protein that is pseudotyped with a capsid gene derived from AAV serotype 8 or AAV serotype 9. Techniques involving the construction and use of pseudotype rAAV virions are publicly known in the art and are described, for example, in Duan et al., J. Virol. 75:7662-7671 (2001), Halbert et al., J. Virol. 74:1524-1532 (2000), Zolotukhin et al., Methods, 28:158-167 (2002), and Auricchio et al., Hum. Molec. Genet., 10:3075-3081 (2001).
[0453] Using AAV virions with mutations within the virion capsid allows for more effective infection of specific cell types than using non-mutant capsid virions. For example, a suitable AAV variant may be a configuration with ligand insertion mutations to enhance AAV targeting to specific cell types. The construction and characterization of AAV capsid variants, including insertion variants, alanine screening variants, and epitope tag variants, are described in Wu et al., J. Virol. 74:8635-45 (2000).
[0454] Other rAAV virions that may be used in the methods of this disclosure include their capsid hybrids produced by molecular breeding of the virus and by exon shuffling. See, for example, Soong et al., Nat. Genet., 25:436-439 (2000) and Kolman and Stemmer, Nat. Biotechnol. 19:423-428 (2001). Methods for delivering exogenous nucleic acids to target cells
[0455] In some embodiments, techniques used to introduce nucleic acid molecules into mammalian cells are well known in the art. For example, electroporation is used to permeate mammalian cells (e.g., human target cells) by applying an electrostatic potential to the target cells. Mammalian cells, such as human cells, thus exposed to an external electric field are then predisposed to taking up exogenous nucleic acids. Electroporation of mammalian cells is described in detail, for example, Chu et al., Nucleic Acids Research 15:1311 (1987). A similar technique is Nucleofection. TM This method utilizes an applied electric field to stimulate the uptake of exogenous polynucleotides into the nucleus of eukaryotic cells.
[0456] Nucleofection is useful for implementing this technology. TM The protocols are described in detail, for example, Distler et al., Experimental Dermatology 14:315 (2005), and US 2010 / 0317114.
[0457] An additional technique useful for transfection of target cells is squeeze-poration. This technique induces rapid mechanical deformation of cells to stimulate the uptake of exogenous DNA through membrane pores formed in response to applied stress. This technique is advantageous in that a vector is not required for the delivery of nucleic acids to cells such as human target cells. Squeeze-poration is described in detail, for example, Sharei et al., JoVE 81:e50980 (2013).
[0458] Lipofection represents another technique useful for transfection of target cells. This method involves loading nucleic acids into liposomes, which often present cationic functional groups, such as quaternary amines or protonated amines, toward the outside of the liposome. This facilitates electrostatic interactions between the liposome and the cell due to the anionic nature of the cell membrane, which ultimately leads to the uptake of exogenous nucleic acids, for example, by direct fusion of the liposome with the cell membrane or endocytosis of the complex. Lipofection is described in detail, for example, U.S. Patent No. 7,442,386.
[0459] A similar technique that utilizes ionic interactions with the cell membrane to induce the uptake of exogenous nucleic acids involves contacting cells with cationic polymer-nucleic acid complexes. Exemplary cationic molecules that associate with polynucleotides to confer a favorable positive charge for interaction with the cell membrane include activated dendrimers (e.g., described in Dennig, Topics in Current Chemistry 228:227 (2003)), polyethyleneimines, and diethylaminoethyl (DEAE)-dextran, whose use as transfection agents is described in detail, for example, Gulick et al., Current Protocols in Molecular Biology 40:1:9.2:9.2.1 (1997). Magnetic beads are another tool that can be used to transfect target cells in a gentle and efficient manner, as this methodology utilizes an applied magnetic field to direct nucleic acid uptake. This technique is described in detail, for example, US 2010 / 0227406.
[0460] Another useful tool for inducing the uptake of exogenous nucleic acids by target cells is laser-fection, also known as optical transfection. This technique involves exposing cells to electromagnetic radiation of specific wavelengths to gently permeate them, allowing polynucleotides to penetrate the cell membrane. The bioactivity of this technique has been found to be similar to, and in some cases superior to, electroporation.
[0461] Impale infection is another technique that can be used to deliver genetic material to target cells. This relies on the use of nanomaterials such as carbon nanofibers, carbon nanotubes, and nanowires.
[0462] Needle-like nanostructures are synthesized perpendicular to the surface of a substrate. DNA containing genes intended for intracellular delivery is bound to the surface of the nanostructures. A chip with an array of these needles is then pressed onto cells or tissue. Cells implicated by the nanostructures can express the delivered genes. An example of this technique is described in Shalek et al., PNAS 107:1870 (2010).
[0463] Magnetic infection can also be used to deliver nucleic acids to target cells. The principle of magnetic infection is to associate nucleic acids with cationic magnetic nanoparticles. The magnetic nanoparticles are made from iron oxide, which is completely biodegradable, and coated with specific cationic proprietary molecules that vary depending on the application. Their association with gene vectors (DNA, viral vectors, etc.) is achieved by salt-induced colloidal aggregation and electrostatic interactions. The magnetic particles are then enriched on the target cells under the influence of an external magnetic field generated by a magnet. This technique is described in detail in Scherer et al., Gene Therapy 9:102 (2002).
[0464] Another useful tool for inducing the uptake of exogenous nucleic acids by target cells is sonoporation, a technique that involves the use of sound (typically the frequency of nucleic acid molecules) to modify the permeability of the cell plasma membrane, allowing polynucleotides to penetrate the cell membrane. This technique is described in detail, for example, Rhodes et al., Methods in Cell Biology 82:309 (2007).
[0465] Vesicles represent another potential vehicle that may be used to modify the genome of target cells according to the methods described herein. For example, microvesicles induced by co-overexpression of glycoprotein VSV-G, such as genome-modifying proteins like nucleases, can be used to efficiently deliver proteins into cells, where they can then catalyze site-specific cleavage of endogenous polynucleotide sequences to prepare the cell's genome for covalent incorporation of target polynucleotides, such as genes or regulatory sequences. The use of such vesicles, also called gesicles, for genetic modification of eukaryotic cells is described in detail, for example, Quinn et al., Genetic Modification of Target Cells by Direct Delivery of Active Protein [abstract] In: Methylation changes in early embryonic genes in cancer [abstract], in: Proceedings of the 18th Annual Meeting of the American Society of Gene and Cell Therapy; 2015 May 13, Abstract No. 122.
[0466] Detection of nucleic acids and proteins The splicing events mediated by one or more of the compounds described above can be characterized using a variety of assays known to those skilled in the art to determine whether the transgene has been transcribed. For example, in some embodiments, the transgene is characterized by assays including, but not limited to, those described herein to determine whether its nucleic acid is expressed. In some embodiments, nucleic acid levels are determined using Northern blotting, Southern blotting, nuclease protection assays (NPAs), insight hybridization (ISH), reverse transcription polymerase chain reaction (RT-PCR), or RNA sequencing (RNA-Seq). In some embodiments, RNA sequencing is performed using Sanger sequencing, Illumina sequencing, ion Torrent sequencing, 454 sequencing, SOLiD sequencing, or nanopore sequencing.
[0467] In some embodiments, gene expression levels are determined using a microarray-based platform.
[0468] In some embodiments, amplification-based assays such as PCR or qPCR are also used to measure the expression levels of one or more markers (e.g., genes).
[0469] In some embodiments, the splicing event mediated by one or more of the above-described compounds is characterized using a variety of assays known to those skilled in the art to determine whether the transgene has been translated. In some embodiments, the transgene is characterized by assays including, but not limited to, Western blotting, immunoblotting, enzyme-linked immunosorbent assay (ELISA), radioimmunoassay (RIA), immunoprecipitation, immunofluorescence, surface plasmon resonance, chemiluminescence, fluorescence polarization, phosphorescence, immunohistochemistry, MALDI-TOF (matrix-associated laser desorption / ionization time of light) mass spectrometry, liquid chromatography (LC) mass spectrometry, microcytometry, microscopy, fluorescence-activated cell coating (FAC), mass spectrometry (MS), flow cytometry, and the following:
[0470] Pharmaceutical composition and route of administration Any one of the compositions described herein, such as a nucleic acid molecule, a nucleic acid vector encoding it, a composition, or a splice modulator, may be formulated into a pharmaceutical composition for administration to a mammalian (e.g., human) subject in a biocompatible form suitable for in vivo administration.
[0471] Pharmaceutical compositions may be manufactured by commonly known methods, for example, by mixing, dissolving, granulation, raising, emulsification, encapsulation, or lyophilization processes. Pharmaceutical compositions may also be formulated using one or more pharmaceutically acceptable carriers, which include excipients and / or adjuvants that facilitate the treatment of the active compound into a pharmaceutically usable preparation, the appropriate formulation depending on the chosen route of administration.
[0472] The active compound can be prepared on a pharmaceutically acceptable carrier that protects the compound from rapid elimination from the body, such as controlled-release formulations including implants and microencapsulation delivery systems. Biodegradable and biocompatible polymers such as ethylene vinyl acetate, polyanhydride, polyglycolic acid, collagen, polyorthoesters, and polylactic acid can be used.
[0473] It should be understood that the pharmaceutical compositions of this disclosure are formulated to conform to their intended route of administration. Examples of routes of administration include systemic administration, parenteral administration, e.g., intravenous administration, intradermal administration, subcutaneous administration, oral administration (e.g., ingestion), inhalation administration, transdermal administration (topical administration), and transmucosal administration. Solutions or suspensions used for parenteral administration, intradermal or subcutaneous application may contain components such as sterile diluents such as water for injection, physiological saline, fixative oil, polyethylene glycol, glycerin, propylene glycol or other synthetic solvents, antimicrobial agents such as benzyl alcohol or methylparaben, antioxidants such as ascorbic acid or sodium sulfite, chelating agents such as ethylenediaminetetraacetic acid, buffers such as acetates, citrates or phosphates, and agents for adjusting tonicity such as sodium chloride or dextrose. pH can be adjusted with an acid or base such as hydrochloric acid or sodium hydroxide. Parenteral preparations may be sealed in ampoules, disposable syringes, or multi-dose vials made of glass or plastic.
[0474] It should be understood that pharmaceutical compositions may be contained in a container, pack, or dispenser along with instructions for administration.
[0475] It should be understood that for any compound, the therapeutically effective dose can first be estimated in a cell culture assay, e.g., a neoplastic cell assay, or in an animal model, usually a rat, mouse, rabbit, dog, or pig. Using an animal model, appropriate concentration ranges and routes of administration may be determined. Such information can then be used to determine useful doses and routes for administration in humans. Therapeutic / prophylactic efficacy and toxicity can be determined in cell cultures or experimental animals, e.g., ED. 50 (A therapeutically effective dose for 50% of the population) and LD 50 This can be determined by standard medical procedures at a dose that is lethal to 50% of the population. The dose-to-toxicity ratio is the therapeutic index, LD50. 50 / ED50 It can be expressed as a ratio. Pharmaceutical compositions exhibiting a large therapeutic indicator are preferred. The dose may vary within this range depending on the dosage form used, the sensitivity of the target, and the route of administration.
[0476] Dosage and administration are adjusted to provide a sufficient level of activator or to maintain the desired effect. Factors to consider include the severity of the disease state, the subject's overall health status, age, weight, and sex, diet, timing and frequency of administration, drug combinations, response sensitivity, and tolerance / response to therapy.
[0477] Accordingly, this disclosure relates to pharmaceutical compositions containing nucleic acid molecules disclosed herein. In some embodiments, this disclosure relates to compositions comprising a vector comprising the nucleic acid molecules of this disclosure. In certain embodiments, this disclosure provides pharmaceutical compositions comprising a vector (e.g., a lentiviral vector or AAV vector) comprising the nucleic acid molecules of this disclosure ligated to a promoter, as disclosed herein. Such compositions may be administered co-administered with a pharmaceutically acceptable splice regulator. The pharmaceutical composition may subsequently comprise an AAV vector comprising (a) a viral capsid and (b) an artificial polynucleotide comprising an expression cassette adjacent to the AAV ITR, wherein the expression cassette comprises a polynucleotide encoding a two-site start codon that regulates the expression of the transgene when ligated by a splice regulator.
[0478] Treatment method Diseases associated with altered RNA transcription levels are often treated by focusing on abnormal protein expression. However, if processes involved in abnormal RNA-level changes, such as components of the splicing process, or related transcription factors, or related stability factors, can be targeted by small molecule treatment, it may be possible to reverse the undesirable effects of abnormal levels of RNA transcript or related protein expression. Therefore, there is a need for methods to regulate the amount of RNA transcript encoded by specific genes as a way to prevent or treat diseases associated with abnormal expression of RNA transcript or related proteins.
[0479] Abnormal splicing of mRNA, such as premRNA, can result in defective proteins and cause disease or impairment in a subject. In some embodiments, the compositions and methods described herein reduce this abnormal splicing of mRNA, such as premRNA, and treat disease or impairment caused by this abnormal splicing.
[0480] Any of the compositions described herein may be used in therapeutic methods, for example, in methods that modulate the expression level of a protein in a subject requiring such treatment. Such treatments may, for example, obtain desired pharmacological and / or physiological effects. The effects may be therapeutic in that they partially or completely cure a disease and / or adverse effects resulting from the disease.
[0481] definition Herein, the structure and other details of this disclosure are described in more detail. Specific terms used in this specification, examples, and appended claims are collected herein. These definitions should be taken in consideration of the remainder of this disclosure and understood by those skilled in the art. Unless otherwise defined, all technical and scientific terms used in this disclosure have the same meaning as commonly understood by those skilled in the art.
[0482] As described in this disclosure, the term “administer” or its grammatical derivatives refers to the delivery of a nucleic acid molecule or a vector containing a nucleic acid molecule to an object requiring it. Any preferred method of administration can be selected by those skilled in the art in consideration of this disclosure.
[0483] Where used in this disclosure, the term “aptamer” means a term commonly used and understood in the art. This disclosure includes, but is not limited to, an aptamer being a fragment (or domain) of nucleic acid that selectively binds to a ligand or molecule. Introduction of a ligand into a ligand-specific aptamer causes a conformational change within the aptamer, affecting the nucleic acid adjacent to the aptamer. In some embodiments, such conformational changes may contribute to a splicing event.
[0484] Where used in this disclosure, the term "binding" refers to a term commonly used and understood in the art. Without limitation, this disclosure includes, but is not limited to, splice regulator binding sites (e.g., nucleic acid molecules or proteins) and splice regulators, molecular interactions (e.g., hydrogen bonding, hydrophobic interactions, electrostatic interactions, van der Waals interactions), splicing (e.g., alternative splicing by spliceosomes at 5' or 3' splice sites), or, more generally, splicing domains or RNA-binding proteins (RBPs) associated with their regulatory sequences on premRNA. Furthermore, splice regulator binding sites may include configurations with additional modifications.
[0485] As used in this disclosure, the term "two-site" refers to a discontinuous (e.g., a codon whose nucleotides are split) start codon that can be reconstructed during a splicing event of a nucleic acid molecule.
[0486] As used in this disclosure, the term codon refers to a specific amino acid, or any set of three nucleotide bases in a given messenger RNA molecule, or any set of DNA coding strands that designates a translation start signal or stop signal. The term codon also refers to a nucleotide triplet in a DNA strand. Such nucleotide bases may be contiguous or discontinuous (for example, a two-site codon).
[0487] Throughout this specification and the claims, the words “comprise,” or variations such as “comprises,” or “comprising,” will be understood to mean that they include the words or groups of words described, but not that they exclude any other words or groups of words.
[0488] Where used in this disclosure, and where used in relation to the compositions described herein, such as nucleic acid molecules or vectors encoding such nucleic acid molecules, the term “effective amount” means an amount sufficient to produce a beneficial or desired result (e.g., expression of the protein, RNA, miRNA, or shRNA of the subject) when administered to a subject, such as a mammal (e.g., a human). For example, an effective amount of one or more compositions described herein (e.g., nucleic acid molecules and splice regulators) may achieve the expression of the protein, RNA, miRNA, or shRNA of the subject compared to the expression of such protein, RNA, miRNA, or shRNA without administration of the composition of the subject. In some embodiments, the protein, RNA, miRNA, or shRNA is expressed in the presence of splice regulators. “Effective amount,” “therapeutic effective amount,” etc., of compositions such as nucleic acid molecules or vectors encoding such nucleic acid molecules also include an amount that produces a beneficial or desired result in a subject compared to a control.
[0489] As used in this disclosure, the term exon refers to a region within the coding or non-coding region of a gene whose nucleotide sequence determines the nucleotide sequence of the corresponding mRNA and / or the amino acid sequence of the corresponding protein. The term exon also refers to the corresponding region of RNA transcribed from a gene. As described above, a gene may contain several exons separated by intervening introns. Exons may be transcribed into premRNA and incorporated into the mRNA in accordance with alternative splicing of the gene.
[0490] As used in this disclosure, the term intron refers to a region within the non-coding region of a gene whose nucleotide sequence is not incorporated into the mRNA of the corresponding gene. The term intron also refers to a corresponding region of RNA transcribed from a gene. In some embodiments, a gene may contain, for example, several introns, each forming an intervening sequence between two exons. Introns are transcribed into premRNA but are removed during processing and are not included in mature mRNA.
[0491] As used in this disclosure, the terms “excise,” “excised,” “excision,” and “excising” refer to the removal of one or more exons and / or one or more introns from a nucleic acid molecule that occurs during splicing (e.g., splicing out).
[0492] As used in this disclosure, the term “exon-inclusion” refers to a ligand-responsive riboswitch that facilitates binding to the 5' or 3' splice site of a spliceosome, or a mechanism that facilitates the binding of an RNA-binding protein (RBP) to its regulatory target site.
[0493] The terms “linked” refer to terms commonly used and understood in the art. Without limitation, this disclosure includes, but is not limited to, “linked” referring to juxtapositions in which the components described as such are related in a way that enables them to function in their intended manner. In embodiments, a promoter is linked to a coding sequence if the promoter affects its transcription or expression.
[0494] As used in this disclosure, the term “minigene” means an isolated nucleic acid sequence encoding a recombinant protein, RNA, miRNA, or shRNA, wherein one or more elements of the corresponding gene encoding a native protein, miRNA, or shRNA have been removed, and the protein, RNA, miRNA, or shRNA encoded by the minigene retains a specific segment of the corresponding native or synthetic protein, RNA, miRNA, or shRNA.
[0495] As used in this disclosure, the terms “nucleic acid molecule” and “nucleic acid” refer to polymers of any length composed of monomeric nucleotides. Nucleic acids include, but are not limited to, ribonucleic acid (RNA), deoxyribonucleic acid (DNA), single-stranded nucleic acids, double-stranded nucleic acids, small interfering ribonucleic acid (siRNA), short hairpin RNA (shRNA), and microRNA (miRNA).
[0496] As used in this disclosure, nucleotide means a nucleoside having a phosphate group covalently bonded to the sugar portion of the nucleoside.
[0497] The term “pharmaceutically acceptable” means safe for administration to mammals, such as humans. In some embodiments, a pharmaceutically acceptable composition is approved by a federal or state regulatory authority or is listed in the United States Pharmacopeia or any other generally recognized pharmacopoeia for use in animals (e.g., humans). As used in this disclosure, “pharmaceutically acceptable” means a compound, anion, cation, material, composition, carrier, and / or dosage form suitable for use in contact with human and animal tissues without excessive toxicity, irritation, allergic reactions, or other problems or complications, in proportion to a reasonable benefit / risk ratio, within the bounds of sound medical judgment.
[0498] Where used in this disclosure, the terms “pharmaceutically acceptable carrier” or “pharmaceutically acceptable excipient” interchangeably refer to any and all solvents, dispersions, coatings, isotonic agents, and absorption retarders, etc., that are compatible with pharmaceutically active substances. The use of such media and agents for pharmaceutically active substances is well known in the art. The composition may also contain, together with one or more pharmaceutically acceptable excipients, other active compounds that provide supplemental, additional, or enhanced therapeutic function.
[0499] It should be understood that this disclosure also provides pharmaceutical compositions comprising any of the compounds described herein in combination with at least one pharmaceutically acceptable excipient or carrier.
[0500] As used in this disclosure, the term "pharmaceutical composition" refers to a formulation containing the compound of this disclosure in a form suitable for administration to a subject. In one embodiment, the pharmaceutical composition is a bulk or unit dosage form. A unit dosage form is any of various forms, including, for example, a capsule, an IV bag, a tablet, a single pump on an aerosol inhaler, or a vial. The amount of the active ingredient (e.g., a formulation of the disclosed compound or a salt, hydrate, solvate, or isomer) in a unit dose of the composition is an effective dose and varies according to the specific treatment involved. Those skilled in the art will understand that it may be necessary to make routine adjustments to the dosage depending on the age and condition of the subject. The dosage also depends on the route of administration. Various routes are intended, including oral, pulmonary, rectal, parenteral, transdermal, subcutaneous, intravenous, intramuscular, intraperitoneal, inhalation, buccal, sublingual, intrapleural, subarachnoid, and intranasal. Dosage forms for topical or transdermal administration of the compounds of this disclosure include powders, sprays, ointments, pastes, creams, lotions, gels, solutions, patches, and inhalants. In one embodiment, the active compound is mixed under sterile conditions with a pharmaceutically acceptable carrier and any necessary preservatives, buffers, or propellants.
[0501] As used in this disclosure, the term "pharmaceutically acceptable excipient" means an excipient that is generally safe, non-toxic, and not biologically undesirable, and is useful in the preparation of pharmaceutical compositions that are acceptable for veterinary and human pharmaceutical use. As used herein and in the claims, "pharmaceutically acceptable excipient" includes both one and more such excipients.
[0502] As used in this disclosure, the term therapeutic effective dose refers to the amount of a drug used to treat, improve, or prevent a specific disease or condition, or to exhibit a detectable therapeutic or inhibitory effect. The effect can be detected by any assay method known in the art. The exact effective dose for a subject depends on the subject's weight, size, and health, the nature and severity of the condition, and the therapeutic agent or combination of therapeutic agents selected for administration. The therapeutic effective dose for a given situation can be determined by routine experimentation, which is within the skill and judgment of the clinician.
[0503] In this disclosure, the terms “polyadenylation signal” (PolyA) or “polyadenylation site” refer to a nucleic acid sequence sufficient to command the addition of one or more polyadenosine ribonucleic acids to an RNA molecule expressed in a cell.
[0504] As used in this disclosure, the term “promoter” refers to a recognition site on DNA to which an RNA polymerase is bound. The polymerase drives the transcription of the transgene. Exemplary promoters suitable for use in the compositions and methods described herein are described herein. Furthermore, the term “promoter” may refer to a synthetic promoter, such as a regulatory DNA sequence, that does not occur naturally in a biological system. Synthetic promoters include portions of natural promoters that do not occur naturally and can be combined with polynucleotide sequences that can be optimized to express recombinant DNA.
[0505] The terms “recipient,” “individual,” “subject,” “host,” and “patient” are used interchangeably herein and, in some embodiments, refer to any mammalian subject, particularly humans, for whom diagnosis, treatment, or therapy is desired. A mammal for therapeutic purposes refers to any animal classified as a mammal, such as humans, farm animals and livestock, as well as laboratory animals, zoo animals, sports animals, or pet animals, such as dogs, horses, cattle, cattle, sheep, goats, pigs, mice, rats, rabbits, guinea pigs, and monkeys. In some embodiments, the mammal is human. None of these terms require the supervision of a medical professional.
[0506] Splicing refers to the process in which intron sequences are removed from nascent pre-messenger RNA (pre-mRNA), and exons are joined together to form mRNA. Splice sites are located at the junction between exons and introns and are defined by different consensus sequences at the 5' and 3' ends of the intron (i.e., the splice donor site and splice acceptor site, respectively). Alternative pre-mRNA splicing, or alternative splicing, is a widespread process that occurs in many genes containing multiple exons. This is accomplished by many multicomponent structures called spliceosomes, which are aggregates of nuclear ribonucleoproteins (snRNPs) and a wide variety of coproteins. By recognizing various cis-regulatory sequences, spliceosomes define exon / intron boundaries, remove intron sequences, and incorporate splice exons into the final translatable message (e.g., mRNA). In alternative splicing, certain exons may or may not be included to ultimately alter the coding message, thereby altering the resulting expressed protein.
[0507] As used in this disclosure, a splicing domain refers to a nucleic acid sequence having a motif that is recognized by a spliceosome and mediates splicing (e.g., by alternative splicing). A splicing domain includes a splice site. The splice site may also include other regulatory elements. For example, in some embodiments, a splicing domain includes a splicing enhancer (e.g., an exon splicing enhancer or an intron splicing enhancer). In some embodiments, a splicing domain includes a branching point (e.g., a strongly conserved branching point), a branching point sequence, or a polypyrimidine tube. In some embodiments, a splicing domain includes a splice acceptor or a splice donor.
[0508] As used in this disclosure, the term splice donor site refers to a splice site located at the 5' end of an intron or the 3' end of an exon. The term splice donor site is used interchangeably with the term 5' splice site. As used in this disclosure, the term splice receptor site refers to a splice site located at the 3' end of an intron or the 5' end of an exon. The term splice acceptor site is used interchangeably with the term 3' splice site.
[0509] As used in this disclosure, the terms “transduction” and “transduction” refer to a method of introducing a viral vector construct or a portion thereof into a cell and subsequently expressing the transgene encoded by the vector construct or a portion thereof in the cell.
[0510] As used in this disclosure, “transfection” means any of the various techniques commonly used to introduce exogenous DNA into prokaryotic or eukaryotic host cells, such as electroporation, lipofection, calcium phosphate precipitation, diethylaminoethyl (DEAE)-dextran transfection, and NUCLEOFECTION. TM Squeeze electroporation, sonoporation, optical transfection, MAGNETOFECTION TMThis refers to conditions such as impalement.
[0511] Where used in this disclosure, the term “transgene” means a recombinant nucleic acid (e.g., DNA or cDNA) that codes for a gene product (e.g., a protein, RNA, miRNA, or shRNA of interest as described herein). The gene product may be RNA (e.g., miRNA or shRNA), a peptide, or a protein. In addition to the coding region of the gene product, the transgene may be a configuration that includes, or is linked to, one or more elements for promoting or enhancing expression, such as promoters, enhancers, destabilization domains, response elements, reporter elements, insulating elements, polyadenylation signals, and other functional elements. Embodiments of this disclosure may utilize any known and suitable promoters, enhancers, destabilization domains, response elements, reporter elements, insulating elements, polyadenylation signals, and / or other functional elements.
[0512] As used in this disclosure, the term vector includes nucleic acid vectors, such as plasmids, RNA vectors, or DNA vectors such as other suitable replicons (e.g., viral vectors). Various vectors have been developed for delivering exogenous polynucleotides or polynucleotides encoding proteins to prokaryotic or eukaryotic cells. Examples of such expression vectors are disclosed, for example, in WO1994 / 011026. Expression vectors suitable for use in the compositions and methods described herein contain polynucleotide sequences as well as additional sequence elements used for the expression of heterologous nucleic acid materials (e.g., nucleic acid molecules) in mammalian cells. Specific vectors that may be used for the expression of nucleic acid molecules described herein include plasmids containing regulatory sequences such as promoter and enhancer regions that induce gene transcription. Other useful vectors for the expression of nucleic acid molecular agents disclosed herein contain polynucleotide sequences that enhance the translation rate of these polynucleotides or improve the stability or nuclear transport of RNA resulting from gene transcription. Examples of these sequence elements include, for example, 5' and 3' untranslated regions, IRESs, 2A ribosome skipping peptides, and polyA sequences to induce efficient transcription of genes supported on the expression vector. Expression vectors suitable for use in the compositions and methods described herein may also contain polynucleotides encoding markers for selecting cells containing such vectors. Examples of suitable markers are genes encoding resistance to antibiotics such as ampicillin, chloramphenicol, kanamycin, noortheotricin, or zeosin.
[0513] Where used in this disclosure, the term "compounds of this disclosure" generally refers to the compounds disclosed herein.
[0514] When used in this disclosure, the terms "alkyl," "C1, C2, C3, C4, C5, or C6 alkyl," or "C1-C6 alkyl (C)" are used. 1~"C1-C6 alkyl)" is intended to encompass linear saturated aliphatic hydrocarbon groups of C1, C2, C3, C4, C5, or C6, and branched saturated aliphatic hydrocarbon groups of C3, C4, C5, or C6. For example, C1-C6 alkyl is intended to include C1, C2, C3, C4, C5, and C6 alkyl groups. Examples of alkyls include, but are not limited to, methyl, ethyl, n-propyl, i-propyl, n-butyl, s-butyl, t-butyl, n-pentyl, i-pentyl, or n-hexyl, and include moieties having 1 to 6 carbon atoms. In some embodiments, linear or branched alkyls have 6 or fewer carbon atoms (e.g., C1-C6 for linear, C3-C6 for branched), and in other embodiments, linear or branched alkyls have 4 or fewer carbon atoms.
[0515] As used in this disclosure, “alkenyl” is intended to encompass linear or branched hydrocarbon groups (“C2-C6 alkenyl”) having 2 to 6 carbon atoms, one or more carbon-carbon double bonds, and no triple bonds. One or more carbon-carbon double bonds may be internal (e.g., 2-butenyl) or terminal (e.g., 1-butenyl). 2- Examples of C6 alkenyl groups include ethenyl (C2), 1-propenyl (C3), 2-propenyl (C3), 1-butenyl (C4), 2-butenyl (C4), and butadienyl (C4).
[0516] As used in this disclosure, “alkynyl” is intended to encompass a linear or branched hydrocarbon group having 2 to 6 carbon atoms, one or more carbon-carbon triple bonds, and optionally one or more double bonds (C2-C6 alkynyl). One or more carbon-carbon triple bonds may be internal (e.g., 2-butynyl) or terminal (e.g., 1-butynyl). 2- Examples of C4 alkynyl groups include, but are not limited to, ethynyl (C2), 1-propynyl (C3), 2-propynyl (C3), 1-butynyl (C4), and 2-butynyl (C4).
[0517] As used in this disclosure, the term “optionally substituted alkyl” refers to an unsubstituted alkyl or alkyl having specified substituents that substitute one or more hydrogen atoms on one or more carbons of a hydrocarbon skeleton. These substituents include, for example, alkyl, alkenyl, alkynyl, halogen, hydroxyl, alkylcarbonyloxy, arylcarbonyloxy, alkoxycarbonyloxy, aryloxycarbonyloxy, carboxylate, alkylcarbonyl, arylcarbonyl, alkoxycarbonyl, aminocarbonyl, alkylaminocarbonyl, dialkylaminocarbonyl, alkylthiocarbonyl, alkoxyl, phosphoric acid, phosphonato, phosphinato, amino (including alkylamino, dialkylamino, arylamino, diarylamino, and alkylarylamino), acylamino (including alkylcarbonylamino, arylcarbonylamino, carbamoyl, and ureido), amidino, imino, sulfhydryl, alkylthio, arylthio, thiocarboxylate, sulfate, alkylsulfinyl, sulfonate, sulfamoyl, sulfonamide, nitro, trifluoromethyl, cyano, azide, heterocycloalkyl, alkylaryl, or aromatic or heteroaromatic moieties.
[0518] Other optionally substituted moieties (such as optionally substituted cycloalkyl, heterocycloalkyl, aryl, or heteroaryl moieties) include both unsubstituted moieties and moieties having one or more of the specified substituents. For example, substituted heterocycloalkyls include those substituted with one or more alkyl groups, such as 2,2,6,6-tetramethyl-piperidinyl and 2,2,6,6-tetramethyl-1,2,3,6-tetrahydropyridinyl.
[0519] As used in this disclosure, the terms “alkoxy” or “alkoxyl” refer to a group-OR where R is alkyl. Certain alkoxy groups are methoxy, ethoxy, n-propoxy, isopropoxy, n-butoxy, tert-butoxy, sec-butoxy, n-pentoxy, n-hexoxy, and 1,2-dimethylbutoxy. Certain alkoxy groups are lower alkoxys, i.e., having 1 to 6 carbon atoms.
[0520] When used in this disclosure, "heteroalkyl," "C1, C2, C3, C4, C 5、 "C6 heteroalkyl" or "C1-C6 heteroalkyl" refers to a linear saturated aliphatic hydrocarbon group of C1, C2, C3, C4, C5, or C6, in which at least one of the carbon atoms is substituted with N, O, or S, and C3, C4, C 5、 Alternatively, it is intended to mean a branched saturated aliphatic hydrocarbon group of C6. The heteroatom may be bonded to any desired hydrogen to satisfy the valence of the heteroatom (for example, CH2 may be substituted with O or NH, and CH may be substituted with N). Examples of such substituents include -O-CH(CH3)2, -CH2-N(CH3)-CH2CH2OCH3, and -S-CH2CH2-O-CH2CH3.
[0521] As used in this disclosure, the term "cycloalkyl" means a group of 3 to 30 carbon atoms (e.g., C3-C3-C3). 12 , C3-C 10 This refers to monocyclic or polycyclic (e.g., condensed, cross-linked, or spirocyclic) saturated or partially unsaturated hydrocarbon systems having C3-C8 rings. Examples of cycloalkyls include, but are not limited to, cyclopropyl, cyclobutyl, cyclopentyl, cyclohexyl, cycloheptyl, cyclooctyl, cyclopentenyl, cyclohexenyl, cycloheptenyl, 1,2,3,4-tetrahydronaphthalenyl, and adamantyl. In the case of polycyclic cycloalkyls, only one of the rings in the cycloalkyl must be non-aromatic.
[0522] As used in this disclosure, the terms "heterocycloalkyl" or "heterocyclyl" mean, unless otherwise specified, a saturated or partially unsaturated 3- to 8-membered monocyclic, 7- to 12-membered bicyclic (condensed, bridging, or spiro-ring), or 11- to 14-membered tricyclic ring system (condensed, bridging, or spiro-ring), having, for example, one, 1-2, 1-3, 1-4, 1-5, or 1-6 heteroatoms, or one or more heteroatoms (O, N, S, P, or Se) independently selected from the group consisting of 1, 2, 3, 4, 5, or 6 heteroatoms, nitrogen, oxygen, and sulfur.Examples of heterocycloalkyl groups include, but are not limited to, piperidinyl, piperazinyl, pyrrolidinyl, dioxanyl, tetrahydrofuranyl, isoindolinyl, indolinyl, imidazolidinyl, pyrazolidinyl, oxazolidinyl, isoxazolidinyl, triazolidinyl, oxyranyl, azetidinyl, oxetanyl, thietanyl, 1,2,3,6-tetrahydropyridinyl, tetrahydropyranyl, dihydropyranyl, pyranyl, morpholinyl, tetrahydrothiopyranyl, 1,4 -Diazepanyl, 1,4-Oxazepanyl, 2-Oxa-5-azabicyclo[2.2.1]heptanyl, 2,5-Diazabicyclo[2.2.1]heptanyl, 2-Oxa-6-azaspiro[3.3]heptanyl, 2,6-Diazaspiro[3.3]heptanyl, 1,4-Dioxa-8-azaspiro[4.5]decanyl, 1,4-Dioxaspiro[4.5]decanyl, 1-Oxaspiro[4.5]decanyl, 1-Azaspiro[4.5]decanyl, 3'H-Spiro[cyclohexane-1,1'-isobenzoph [lan]-yl, 7'H-spiro[cyclohexane-1,5'-flo[3,4-b]pyridine]-yl, 3'H-spiro[cyclohexane-1,1'-flo[3,4-c]pyridine]-yl, 3-azabicyclo[3.1.0]hexanyl, 3-azabicyclo[3.1.0]hexane-3-yl, 1,4,5,6-tetrahydropyrrolo[3,4-c]pyrazolyl, 3,4,5,6,7,8-hexahydropyrido[4,3-d]pyrimidinyl, 4,5,6,7-tetrahydro-1H-pyrazo Examples include lo[3,4-c]pyridinyl, 5,6,7,8-tetrahydropyrido[4,3-d]pyrimidinyl, 2-azaspiro[3.3]heptanyl, 2-methyl-2-azaspiro[3.3]heptanyl, 2-azaspiro[3.5]nonanyl, 2-methyl-2-azaspiro[3.5]nonanyl, 2-azaspiro[4.5]decanyl, 2-methyl-2-azaspiro[4.5]decanyl, 2-oxazaspiro[3.4]octanyl, and 2-oxazaspiro[3.4]octan-6-yl. In the case of polycyclic heterocycloalkyls, only one of the rings in the heterocycloalkyl must be non-aromatic (e.g., 4,5,6,7-tetrahydrobenzo[c]isoxazolyl).
[0523] As used in this disclosure, the term "aryl" refers to a monocyclic or polycyclic (e.g., bicyclic or tricyclic) 4n+2 aromatic ring system (e.g., having 6, 10, or 14 π electrons shared in a cyclic array) having 6 to 14 ring carbon atoms and zero heteroatoms provided to the aromatic ring system. Examples of aryl groups include, but are not limited to, phenyl and naphthyl. Conveniently, aryl is phenyl.
[0524] As used in this disclosure, the term heteroaryl is intended to include a carbon atom and one or more heteroatoms independently selected from the group consisting of nitrogen, oxygen, and sulfur, e.g., 1, 1-2, or 1-3, or 1-4, or 1-5, or 1-6 heteroatoms, or a stable 5, 6, or 7-membered monocyclic or 7, 8, 9, 10, 11, or 12-membered bicyclic aromatic heterocycle consisting of, for example, 1, 2, 3, 4, 5, or 6 heteroatoms. The nitrogen atom may be substituted or unsubstituted (i.e., N or NR, where R is H or another defined substituent). The nitrogen and sulfur heteroatoms may optionally be oxidized (i.e., N → O and S(O)). p (wherein p=1 or 2). Note that the total number of S and O atoms in the aromatic heterocycle is 1 or less. Examples of heteroaryl groups include pyrrole, furan, thiophene, thiazole, isothiazole, imidazole, triazole, tetrazole, pyrazole, oxazole, isoxazole, pyridine, pyrazine, pyridazine, and pyrimidine. Heteroaryl groups may also be fused or bridged with non-aromatic alicyclic or heterocyclic rings to form polycyclic systems (e.g., 4,5,6,7-tetrahydrobenzo[c]isoxazolyl).
[0525] Furthermore, the terms "aryl" and "heteroaryl" encompass polycyclic aryl and heteroaryl groups, such as tricyclic and bicyclic groups, including naphthalene, benzoxazole, benzodiaxazole, benzothiazole, benzimidazole, benzothiophene, quinoline, isoquinoline, naphtholidine, indole, benzofuran, purine, benzofuran, deazapurine, and indoridine.
[0526] Cycloalkyl, heterocycloalkyl, aryl, or heteroaryl rings include, for example, alkyl, alkenyl, alkynyl, halogen, hydroxyl, alkoxy, alkylcarbonyloxy, arylcarbonyloxy, alkoxycarbonyloxy, aryloxycarbonyloxy, carboxylate salt, alkylcarbonyl, alkylaminocarbonyl, aralkylaminocarbonyl, alkenylaminocarbonyl, alkylcarbonyl, arylcarbonyl, aralkylcarbonyl, alkenylcarbonyl, alkoxycarbonyl, aminocarbonyl, alkylthiocarbonyl, phosphate, phosphonato, phosphinato, amino(containing alkylamino) The aryl and heteroaryl groups may be substituted at one or more ring positions (e.g., ring-forming carbons or heteroatoms such as N) having substituents as described above, such as dialkylaminos, arylaminos, diarylaminos and alkylarylaminos, acylaminos (including alkylcarbonylaminos, arylcarbonylaminos, carbamoyls and ureidos), amidinos, iminos, sulfhydryls, alkylthios, arylthios, thiocarboxylates, sulfates, alkylsulfinyls, sulfonates, sulfamoyls, sulfonamides, nitros, trifluoromethyls, cyanos, azides, heterocycloalkyls, alkylaryls, or aromatic or heteroaromatic moieties. The aryl and heteroaryl groups may also be condensed or crosslinked with non-aromatic alicyclic or heterocyclic rings to form polycyclic systems (e.g., methylenedioxyphenyls such as tetralin and benzo[d][1,3]dioxol-5-yl).
[0527] As used in this disclosure, the term “substitution” means that any one or more hydrogen atoms on a given atom are replaced by a selection from a given group, provided that the substitution does not exceed the normal valence of the given atom and that the substitution results in a stable compound. When a substituent is oxo or keto (i.e., =O), two hydrogen atoms on the atom are substituted. Keto substituents are not present in aromatic moieties. A ring double bond, as used in this disclosure, is a double bond formed between two adjacent ring atoms (e.g., C=C, C=N, or N=N). “Stable compound” and “stable structure” mean a compound that survives isolation from a reaction mixture to a useful purity and is robust enough to be formulated into an effective therapeutic agent.
[0528] If a bond to a substituent is shown to intersect with a bond connecting two atoms in the ring, such substituents may be bonded to any atom in the ring. If substituents are enumerated without indicating the atom to which such substituents are bonded to the rest of the compound in a given formula, such substituents may be bonded to any atom in such formula. Combinations of substituents and / or variables are permitted, but only if such combinations result in a stable compound.
[0529] If any variable (e.g., R) occurs more than once in any component or formula of a compound, its definition in each occurrence is independent of its definition in all other occurrences. Therefore, for example, if a group is shown to be substituted with 0 to 2 R moieties, the group may, at its discretion, be substituted with up to 2 R moieties, and in each occurrence, R is selected independently of the definition of R. Furthermore, combinations of substituents and / or variables are permitted, but only if such combinations result in a stable compound.
[0530] Where used in this disclosure, the terms "hydroxy" or "hydroxyl" mean -OH or -O - It contains a group having a group.
[0531] As used in this disclosure, the term cyano refers to the -CN group.
[0532] As used in this disclosure, the term nitro refers to the radical-NO2.
[0533] As used in this disclosure, the term oxo refers to =O.
[0534] As used in this disclosure, the terms "halo" or "halogen" refer to fluoro, chloro, bromo, and iodine.
[0535] As used in this disclosure, the term haloalkyl refers to a branched or unbranched alkyl group substituted with one or more halogens. For example, C 1-6 A haloalkyl is an alkyl group consisting of one to seven carbon atoms, in which at least one hydrogen atom is substituted with a halogen. Examples of haloalkyls include, but are not limited to, CFH2, CF2H, CF3, CH2CF3, CF2CF3, C(F)(CH3)2, CH2CH2Br, CH(I)CH2F, and CH2Cl.
[0536] As used in this disclosure, the term "haloalkoxy" refers to an alkoxy structure substituted with one or more halo groups or combinations thereof. For example, the terms "fluoroalkyl" and "fluoroalkoxy" include a haloalkyl group and a haloalkoxy group, respectively, where the halo is fluorine.
[0537] As used in this disclosure, the term “optionally substituted haloalkyl” refers to an unsubstituted haloalkyl having specified substituents that substitute one or more hydrogen atoms on one or more hydrocarbon backbone carbon atoms. These substituents may include, for example, alkyl, alkenyl, alkynyl, halogen, hydroxyl, alkylcarbonyloxy, arylcarbonyloxy, alkoxycarbonyloxy, aryloxycarbonyloxy, carboxylate, alkylcarbonyl, arylcarbonyl, alkoxycarbonyl, aminocarbonyl, alkylaminocarbonyl, dialkylaminocarbonyl, alkylthiocarbonyl, alkoxyl, phosphoric acid, phosphonato, phosphinato, amino (including alkylamino, dialkylamino, arylamino, diarylamino, and alkylarylamino), acylamino (including alkylcarbonylamino, arylcarbonylamino, carbamoyl, and ureido), amidino, imino, sulfhydryl, alkylthio, arylthio, thiocarboxylate, sulfate, alkylsulfinyl, sulfonate, sulfamoyl, sulfonamide, nitro, trifluoromethyl, cyano, azide, heterocycloalkyl, alkylaryl, or aromatic or heteroaromatic moieties.
[0538] When used in this disclosure, expression means one or more of A, B, or C, one or more of A, B, or C, one or more of A, B, and C, A, B, and C selected from the group consisting of A, B, and C, and all similar terms are used interchangeably, and all mean selected from the group consisting of A, B, and / or C, i.e., one or more As, one or more B, one or more C, or any combination thereof.
[0539] It should be understood that this disclosure provides methods for the synthesis of any of the compounds of the formulas described herein. This disclosure also provides detailed methods for the synthesis of various disclosed compounds of this disclosure by the schemes shown below, as well as by the schemes shown in the examples.
[0540] Throughout the description in which a composition is described as having, containing, or containing certain components, it should be understood that the composition is also intended to be essentially composed of or consisting of the listed components. Similarly, if a method or process is described as having, containing, or containing certain process steps, the process is also intended to be essentially composed of or consisting of the listed processing steps. Furthermore, naturally, the order of steps or sequences for performing a particular action is not important, as long as this disclosure remains operational. Moreover, two or more steps or actions can be performed simultaneously.
[0541] It should be understood that the synthesis process of this disclosure can accommodate a wide variety of functional groups, and therefore a variety of substituted starting materials can be used. The process generally provides the desired final compound at or near the end of the entire process, but in certain examples, it may be desirable to further convert the compound to its pharmaceutically acceptable salt.
[0542] It should be understood that the compounds of this disclosure can be prepared in a variety of ways using commercially available starting materials, compounds known in the literature, or readily prepared intermediates by employing standard synthetic methods and procedures that are known to those skilled in the art or would be apparent to those skilled in the art in light of the teachings herein. Standard synthetic methods and procedures for the preparation of organic molecules and the transformation and manipulation of functional groups can be obtained from relevant scientific literature or standard textbooks in the field. Not limited to any one or more sources, but including, Smith, MB, March, J., March's Advanced Organic Chemistry: Reactions, Mechanisms, and Structure, 5 th edition, John Wiley & Sons: New York, 2001, Greene, TW, Wuts, PGM, Protective Groups in Organic Synthesis, 3 rdJohn Wiley & Sons, New York, 1999; R. Larock, Comprehensive Organic Transformations, VCH Publishers (1989); L. Fieser and M. Fieser, Fieser and Fieser's Reagents for Organic Synthesis, John Wiley and Sons (1994); and L. Paquette, ed., Encyclopedia of Reagents for Organic Synthesis, John Wiley and Sons (1995) are useful and recognized as classic literature on organic synthesis known to those skilled in the art.
[0543] Those skilled in the art will note that the order of certain steps, such as the introduction and removal of protecting groups, may be altered during the reaction sequences and synthetic schemes described herein. Those skilled in the art will also recognize that certain groups may require protection from reaction conditions through the use of protecting groups. Protecting groups may also be used to distinguish similar functional groups in a molecule. A list of protecting groups, as well as methods for introducing and removing these groups, can be found in Greene, TW, Wuts, PGM, Protective Groups in Organic Synthesis, 3. rd See edition, John Wiley & Sons: New York, 1999.
[0544] With respect to the compounds of this disclosure that can further form salts, it should be understood that all of these forms are also intended to be within the scope of the claimed disclosure.
[0545] As used in this disclosure, the term “pharmaceutically acceptable salt” refers to a derivative of a compound of this disclosure in which the parent compound is modified by forming an acid or a base salt thereof. Examples of pharmaceutically acceptable salts include, but are not limited to, mineral or organic salts of basic residues such as amines, and alkali or organic salts of acidic residues such as carboxylic acids. Pharmaceutically acceptable salts include, for example, conventional non-toxic or quaternary ammonium salts of the parent compound formed from non-toxic inorganic or organic acids. For example, such conventional non-toxic salts include, but are not limited to, 2-acetoxybenzoic acid, 2-hydroxyethanesulfonic acid, acetic acid, ascorbic acid, benzenesulfonic acid, benzoic acid, bicarbonate, carbonic acid, citrate, edetic acid, ethanesulfonic acid, 1,2-ethanesulfonic acid, fumaric acid, glucoheptonic acid, gluconic acid, glutamic acid, glycolic acid, glycolylarsanilic acid, hexylresorcinolic acid, hydrabamic acid, hydrobromic acid, hydrochloric acid, hydroiodic acid, hydroxymaleic acid, hydroxynaphthoic acid, Examples include isethionic acid, lactic acid, lactobionic acid, lauryl sulfonic acid, maleic acid, malic acid, mandelic acid, methanesulfonic acid, napsylic acid nitrate, oxalate, pamoic acid, pantothenic acid, phenylacetic acid, phosphoric acid, polygalacturonic acid, propionic acid, salicylic acid, stearic acid, acetic acid, succinic acid, sulfate, sulfate, sulfate, tannic acid, tartaric acid, toluenesulfonic acid, and inorganic and organic acids selected from commonly occurring amine acids, such as glycine, alanine, phenylalanine, and arginine.
[0546] In some embodiments, pharmaceutically acceptable salts include sodium salts, potassium salts, calcium salts, magnesium salts, diethylamine salts, choline salts, meglumine salts, benzathine salts, trometamic acid salts, ammonia salts, arginine salts, or lysine salts.
[0547] Other examples of pharmaceutically acceptable salts include hexanoic acid, cyclopentanepropionic acid, pyruvate, malonic acid, 3-(4-hydroxybenzoyl)benzoic acid, cinnamic acid, 4-chlorobenzenesulfonic acid, 2-naphthalenesulfonic acid, 4-toluenesulfonic acid, camphorsulfonic acid, 4-methylbicyclo-[2.2.2]-octa-2-ene-1-carboxylic acid, 3-phenylpropionic acid, trimethylacetic acid, tertiary butylacetic acid, and muconic acid. The disclosure also includes salts formed by either replacing the acidic proton present in the parent compound with a metal ion, such as an alkali metal ion, an alkaline earth ion, or an aluminum ion, or by coordinating with an organic base such as ethanolamine, diethanolamine, triethanolamine, tromethamine, or N-methylglucamine. In salt form, it is understood that the ratio of the compound to the salt cation or anion can be 1:1, or any other ratio, such as 3:1, 2:1, 1:2, or 1:3.
[0548] All references to pharmaceutically acceptable salts should be understood to include the solvent-added form (solvate) of the same salt as defined herein.
[0549] The compound or a pharmaceutically acceptable salt thereof may be administered orally, nasally, percutaneously, pulmonaryly, by inhalation, buccally, sublingually, intraperitoneally, subcutaneously, intramuscularly, intravenously, rectally, intrapleurally, intrathecally, and parenterally. In one embodiment, the compound is administered orally. Those skilled in the art will recognize the advantages of specific routes of administration.
[0550] Salts can be formed, for example, between an anion and a positively charged group (e.g., amino) on a substituted compound disclosed herein. Suitable anions include chlorides, bromides, iodides, sulfates, bissulfates, sulfates, nitrates, phosphates, citrates, methanesulfons, trifluoroacetates, glutamates, glucurons, glutarates, malates, maleates, succinates, fumarates, tartrates, tosylates, salicylates, lactates, naphthalenesulfons, and acetates (e.g., trifluoroacetates).
[0551] As used in this disclosure, the term “pharmaceutically acceptable anion” refers to an anion suitable for forming a pharmaceutically acceptable salt. Similarly, salts may also be formed between a cation on a substituted compound disclosed herein and a negatively charged group (e.g., a carboxylate salt). Suitable cations include sodium, potassium, magnesium, calcium ions, and ammonium cations such as tetramethylammonium or diethylamine ions. The substituted compounds disclosed herein also include salts thereof containing a quaternary nitrogen atom.
[0552] It should be understood that the compounds of this disclosure, for example, salts of the compounds, may exist in either a hydrated or unhydrated (anhydrous) form, or as solvates with other solvent molecules. Non-limiting examples of hydrates include monohydrates and dihydrates. Non-limiting examples of solvates include ethanol solvate and acetone solvate.
[0553] As used in this disclosure, the term solvate means a solvent-added form containing either a stoichiometric or non-stoichiometric amount of solvent. Some compounds tend to capture a fixed molar ratio of solvent molecules in their crystalline solid state and thus form solvates. When the solvent is water, the solvate formed is a hydrate; when the solvent is an alcohol, the solvate formed is an alcoholic acid salt. Hydrates are formed by a combination of one or more water molecules and one molecule of a substance in which water retains its molecular state as H2O.
[0554] As used in this disclosure, the term "analog" refers to a compound that is structurally similar to another compound but has a slightly different composition (such as the substitution of one atom with an atom of a different element, or substitution due to the presence of a particular functional group, or substitution of one functional group with another functional group). Thus, an analog is a compound that is similar or equivalent in function and appearance but is not of the same structure or origin as the reference compound.
[0555] As used in this disclosure, the term derivative refers to a compound having a common core structure and being substituted with one of the various groups described herein.
[0556] As used in this disclosure, the term bioisoster refers to a compound resulting from the exchange of one atom or group of atoms with another broadly similar atom or group of atoms. The purpose of bioisosteric substitution is to create a novel compound having similar biological properties to the parent compound. Bioisosteric substitution may be based on physicochemical or morphological factors. Examples of carboxylic acid bioisosters include, but are not limited to, acylsulfonamides, tetrazoles, sulfonates, and phosphonates. See, for example, Patani and LaVoie, Chem. Rev. 96, 3147-3176, 1996.
[0557] It will also be understood that any particular compound of any of the formulas disclosed herein may exist in both solvated and non-solvated forms, such as hydrated forms. Preferred pharmaceutically acceptable solvates are hydrates, such as hemihydrate, monohydrate, dihydrate, or trihydrate.
[0558] Any one of the compounds of the formulas disclosed herein may exist in many different tautomerized forms, and a reference to a compound in formula (I) includes all such forms. To avoid any doubt, if a compound may exist as one of several tautomerized forms, but only one is specifically described or illustrated, then all the others are nevertheless encompassed by formula (I).
[0559] Examples of tautomers include, for example, the keto, enol, and enolate forms as shown in the following tautomer pairs: keto / enol (illustrated below), imine / enamine, amide / iminoalcohol, amidine / amidine, nitroso / oxime, thioketone / enthiol, and nitro / acidic nitro. [ka]
[0560] Any compound of any of the formulas disclosed herein that contains an amine function may also form an N-oxide. The references in this disclosure to the structure of formula (I) containing an amine function also include N-oxides. If a compound contains several amine functions, one or more nitrogen atoms may be oxidized to form an N-oxide. Specific examples of N-oxides are N-oxides of tertiary amines or nitrogen atoms of nitrogen-containing heterocycles. N-oxides may be formed by treatment of the corresponding amine with an oxidizing agent such as hydrogen peroxide or a peracid (e.g., a peroxycarboxylic acid). See, for example, Advanced Organic Chemistry, by Jerry March, 4th Edition, Wiley Interscience, pages. More specifically, N-oxides may be prepared by a procedure described in LWDady (Syn. Comm. 1977, 7, 509-514) in which an amine compound reacts with meta-chloroperoxybenzoic acid (mCPBA) in an inert solvent, for example, dichloromethane.
[0561] Any one of the compounds of the formulas disclosed herein may be administered in the form of a prodrug that is broken down in the body of a human or animal to release the compound disclosed herein. The prodrug may be used to alter the physical and / or pharmacokinetic properties of the compound disclosed herein. A prodrug may be formed when the compound disclosed herein contains a suitable group or substituent to which a characterizing group can be attached. Examples of prodrugs include derivatives of any one of the ester or amide groups of the formulas disclosed herein that contain an alkyl or acyl substituent that can be cleaved in vivo.
[0562] As used in this disclosure, the term isomer means compounds having the same molecular formula but differing in the arrangement of their atoms or the spatial arrangement of their atoms. Isomers that differ in the spatial arrangement of their atoms are called stereoisomers. Stereoiomers that are not mirror images of each other are called diastereoisomers, and stereoisomers that are non-superimposable mirror images of each other are called enantiomers or optical isomers. A mixture containing equal amounts of individual enantiomer forms of opposing chirality is called a racemic mixture.
[0563] As used in this disclosure, the term chiral center refers to a carbon atom bonded to four non-identical substituents.
[0564] As used in this disclosure, the term chiral isomer means a compound having at least one chiral center. Compounds having two or more chiral centers may exist either as individual diastereomers or as a mixture of diastereomers called a diastereomer mixture. Where there is one chiral center, the stereoisomer may be characterized by the absolute configuration (R or S) of that chiral center. Absolute configuration refers to the spatial arrangement of substituents attached to the chiral center. The substituents attached to the chiral center in question are ranked according to the Sequence Rule of Cahn, Ingold, and Prelog. (Cahn et al., Angew. Chem. Inter. Edit. 1966, 5, 385, errata 511, Cahn et al., Angew. Chem. 1966, 78, 413, Cahn and Ingold, J. Chem. Soc. 1951 (London), 612; Cahn et al., Experientia 1956, 12, 81; Cahn, J. Chem. Educ. 1964, 41, 116).
[0565] As used in this disclosure, the term “geometric isomer” means diastereomers resulting from the presence of a double bond or a cycloalkyl linker (e.g., 1,3-cyclobutyl) that prevents rotation. These configurations are distinguished in their names by the prefixes cis and trans, or Z and E, which indicate that the group is on the same side or opposite side of the double bond in the molecule according to the Kahn-Ingold-Prelogue rule.
[0566] It should be understood that the compounds of this disclosure may be represented as different chiral or geometric isomers. Where a compound has chiral or geometric isomeric forms, all isomeric forms are intended to be included within the scope of this disclosure, and it should be understood that the naming of the compounds does not exclude any isomeric form, and that not all isomers may have the same level of activity.
[0567] It should be understood that the structures and other compounds discussed in this disclosure include all of their atropic isomers. Naturally, not all atropic isomers may possess the same level of activity.
[0568] As used in this disclosure, the term “atropic isomer” refers to a type of stereoisomer in which the atoms of two isomers are arranged differently in space. Anisotropic isomers are characterized by limited rotation caused by the obstruction of rotation of large groups around a central bond. While such atropic isomers typically exist as a mixture, recent advances in chromatographic techniques have made it possible to separate a mixture of two atropic isomers in some cases.
[0569] As used in this disclosure, the term tautomer is one of two or more structural isomers that exist in equilibrium and are readily convertible from one isomeric form to another. This conversion results in a formal transfer of hydrogen atoms, with the switching of adjacent conjugated double bonds. Tautomers exist as a mixture of tautomer sets in solution. In solutions where tautomerization is possible, a chemical equilibrium of tautomers is reached. The exact ratio of tautomers depends on several factors, including temperature, solvent, and pH. The concept of tautomers that can be interconverted by tautomerization is called tautomerism. Of the various types of tautomerism possible, two are commonly observed. Keto-enol tautomerism involves a simultaneous shift of electrons and hydrogen atoms. Ring chain tautomerism arises as a result of an aldehyde group (-CHO) in a sugar chain molecule reacting with one of the hydroxyl groups (-OH) in the same molecule, giving it the cyclic (ring-shaped) form represented by glucose.
[0570] It should be understood that the compounds of this disclosure may be represented as different tautomers. Where a compound has tautomers, all tautomers are intended to be included within the scope of this disclosure, and the naming of the compounds does not exclude any tautomers. It should be understood that certain tautomers may have higher levels of activity than others.
[0571] Compounds that have the same molecular formula but differ in the nature or arrangement of their atomic bonds, or in the arrangement of their atoms in space, are called isomers. Isomers that differ in the arrangement of their atoms in space are called stereoisomers. Stereoiomers that are not mirror images of each other are called diastereomers, and stereoisomers that are non-superimposable mirror images of each other are called enantiomers. If a compound has an asymmetric center, for example, if it is bonded to four different groups, a pair of enantiomers is possible. Enantiomers can be characterized by the absolute arrangement of their asymmetric center, by the Kahn and Prelog R and S sequence rules, or by the way the molecule rotates in the plane of polarized light and is designated as revolutational or revotrotatory (i.e., as (+) or (-) isomers, respectively). Chiral compounds can exist as individual enantiomers or as mixtures thereof. A mixture containing equal proportions of enantiomers is called a racemic mixture.
[0572] The compounds of this disclosure may have one or more asymmetric centers, and therefore such compounds may be produced as individual (R) or (S) stereoisomers, or as mixtures thereof. Unless otherwise indicated, the descriptions or nomenclature of specific compounds in this specification and the claims are intended to include individual enantiomers and mixtures thereof, racemates, or both. Methods for determining stereochemistry and separating stereoisomers are well known in the art, for example, by synthesis from optically active starting materials or decomposition of racemic forms (see Chapter 4 of “Advanced Organic Chemistry”, 4th edition J. March, John Wiley and Sons, New York, 2001). Some of the compounds of this disclosure may have geometric isomer centers (E and Z isomers).
[0573] Accordingly, this disclosure includes any one of the formulas defined above when made available by organic synthesis and when made available in the body of a human or animal by cleavage of its prodrug. Accordingly, this disclosure includes any one of the formulas disclosed herein produced by organic synthesis means, and also such compounds produced in the body of a human or animal by metabolism of precursor compounds, i.e., any one of the formulas disclosed herein may be a compound produced synthetically or a compound produced metabolically.
[0574] A suitable pharmaceutically acceptable prodrug of any one of the compounds of the formulas disclosed herein is based on reasonable medical judgment that it is suitable for administration to a subject without undesirable pharmacological activity and without excessive toxicity.
[0575] A suitable pharmaceutically acceptable prodrug of any one compound of the formulas disclosed herein having a hydroxyl group is, for example, an in vivo cleavable ester or ether thereof. An in vivo cleavable ester or ether of any one compound of the formulas disclosed herein containing a hydroxyl group is, for example, a pharmaceutically acceptable ester or ether that is cleaved in a subject to produce a parent hydroxyl compound. Suitable pharmaceutically acceptable ester-forming groups for the hydroxyl group include inorganic esters such as phosphate esters (including phosphoramide cyclic esters). More suitable pharmaceutically acceptable ester-forming groups for the hydroxyl group include C1-C groups such as acetyl, benzoyl, phenylacetyl, and substituted benzoyl and phenylacetyl groups. 10Examples of ring substituents on phenylacetyl and benzoyl groups include alkanoyl groups, C1-C2 alkoxycarbonyl groups such as ethoxycarbonyl, N,N-(C1-C6 alkyl)2-carbamoyl, 2-dialkylaminoacetyl, and 2-carboxyacetyl groups. Examples of ring substituents on phenylacetyl and benzoyl groups include aminomethyl, N-alkylaminomethyl, N,N-dialkylaminomethyl, morpholinomethyl, piperazine-1-ylmethyl, and 4-(C1-C4 alkyl)piperazine-1-ylmethyl. Suitable pharmaceutically acceptable ether-forming groups for the hydroxyl group include acetoxymethyl and α-acyloxyalkyl groups such as pivaloyloxymethyl.
[0576] Examples of suitable pharmaceutically acceptable prodrugs of any one of the formulas disclosed herein having a carboxyl group include, for example, amides formed with amines such as ammonia, such as amides that can be cleaved in vivo, such as methylamine, such as C 1-4 Examples include alkylamines, (C1-C4 alkyl)2 amines such as dimethylamine, N-ethyl-N-methylamine or diethylamine, C1-C4 alkoxy-C2-C4 alkylamines such as 2-methoxyethylamine, phenyl-C1-C4 alkylamines such as benzylamine, and amino acids such as glycine or its esters.
[0577] A suitable pharmaceutically acceptable prodrug of any one of the compounds of the formulas disclosed herein having an amino group is, for example, an in vivo cleavable amide derivative thereof. Suitable pharmaceutically acceptable amides from an amino group include, for example, acetyl, benzoyl, phenylacetyl, and C1-C such as substituted benzoyl and phenylacetyl groups. 10 Examples of amides formed by alkanoyl groups include aminomethyl, N-alkylaminomethyl, N,N-dialkylaminomethyl, morpholinomethyl, piperazine-1-ylmethyl, and 4-(C1-C4 alkyl)piperazine-1-ylmethyl.
[0578] The administration regimen utilizing this compound is selected according to various factors, including the type, species, age, weight, sex, and medical condition of the subject, the severity of the condition being treated, the route of administration, the subject's renal and hepatic function, and the specific compound or salt used. A physician or veterinarian with the usual skills can easily determine and prescribe the effective dose of the drug necessary to prevent, counteract, or halt the progression of the condition. A physician or veterinarian with the usual skills can easily determine and prescribe the effective dose of the drug necessary to counteract or halt the progression of the condition.
[0579] The formulation and administration techniques of the compounds disclosed in this disclosure are described in Remington: the Science and Practice of Pharmacy, 19 th This can be found in edition, Mack Publishing Co., Easton, PA (1995). In one embodiment, the compounds described herein and their pharmaceutically acceptable salts are used in a pharmaceutical preparation in combination with a pharmaceutically acceptable carrier or diluent. Suitable pharmaceutically acceptable carriers include inert solid fillers or diluents and sterile aqueous or organic solutions. The compounds are present in such pharmaceutical composition in an amount sufficient to provide a desired dose within the range described herein.
[0580] All proportions and ratios used in this disclosure are by weight unless otherwise indicated. Other components and advantages of this disclosure are evident from different embodiments. The provided embodiments illustrate different components and methods useful for carrying out this disclosure. The embodiments do not limit the claimed disclosure. Based on this disclosure, those skilled in the art can identify and adopt other components and methods useful for carrying out this disclosure.
[0581] In the synthetic schemes described herein, compounds may be stretched in one particular configuration for simplification. Such a particular configuration should not be construed as limiting this disclosure to one or another isomer, tautomer, positional isomer, or stereoisomer, nor does it exclude mixtures of isomers, tautomers, positional isomers, or stereoisomers; however, it will be understood that a given isomer, tautomer, positional isomer, or stereoisomer may have a higher level of activity than another isomer, tautomer, positional isomer, or stereoisomer.
[0582] The references to publications and patent documents are not intended to acknowledge any of them as relevant prior art, nor do they constitute any endorsement of their content or date. While this disclosure is described herein by written description, those skilled in the art will recognize that this disclosure may be implemented in various embodiments, and that the foregoing descriptions and examples herein are for illustrative purposes only and not to limit the following claims. [Examples]
[0583] This disclosure is further illustrated by the following embodiments. The embodiments are provided for illustrative purposes only and should not be construed as limiting the scope or content of this disclosure in any way.
[0584] Example 1: Selection of a switch sequence containing a splice regulatory factor binding site This example demonstrates the identification of minigene sequences using a library of variant switch sequences having unique barcodes and splice regulators.
[0585] Splice regulator binding sites were designed by screening over 20,000 switch sequence variants, each labeled with a unique nucleotide barcode. Sequence variants included nucleotide changes such as insertions, deletions, and mutagenesis in various patterns. The oligonucleotide library was then cloned into expression plasmids to obtain a plasmid pool. The plasmid pool was transfected into cells such as HEK-293T using nuclear transduction. After transfection, cells were treated with specific concentrations of splice regulator compounds. After incubation with the compounds for 24–72 hours, cells were harvested, total RNA was purified, and cDNA was synthesized from the purified RNA. Figure 8A shows the scheme of this assay. Target RNA-seq libraries were generated by amplifying switch transcription sequences from the cDNA pool. For each switch transcript, identified paired terminal RNA sequences were determined: variant identification barcode, and whether or not the target exon was spliced into a given transcript. For each barcoded variant, the splicing-in percentage (spliced reads / total reads) was calculated for each vehicle or drug state used in the screening. If the splicing-in percentage for a vehicle state was smaller than the percentage for the wild-type control, and the multiplicative change in the splicing-in percentage after compound treatment exceeded a desired cutoff, such as the 95th percentile, the variant was selected for further analysis.
[0586] A plasmid pool transfected with HEK-293T cells was treated with 85 nM splice regulator 24A for 24 hours. RS1-10 is the start minigene sequence, and the variant switch sequence was identified by switch sequence screening, which contains a 28-nucleotide intron deletion compared to the RS1-10 sequence RS1-X10. RS1-10, which controls the firefly luciferase gene, or RS1-X1, which controls the firefly luciferase gene, were cloned into individual expression plasmids containing the CMV promoter. HEK 293T cells were transfected with either the RS1-10 or RS1-X1 expression plasmid and then treated with 85 nM for 24 hours while incubated in a tissue culture incubator (37°C, 5% CO2). Luminescence was measured as an endpoint analysis by adding a luciferase substrate, such as Steady-Luc Firefly luciferase substrate, to each well of an assay plate. The liquid contents of each well were transferred to a 96-well plate, mixed again on an orbital shaker for 2 minutes, and incubated at room temperature for 5 minutes. Finally, the luminescence signal from each well was measured with a plate reader (Figure 8B). RS1-X1 produced higher luminescence (RLU) than RS1-10.
[0587] Example 2: Selection of splice modifiers This embodiment describes an assay for identifying candidate splice regulators.
[0588] Splice regulators are generated by screening more than 100 (e.g., more than 1,000, more than 10,000, more than 100,000, or more than 1,000,000) candidate splice regulators for their ability to modify the splicing of the minigene from Example 1 to increase the inclusion of two-site start codons and enable transgene expression. The screen utilizes a DNA construct containing a minigene linked to a reporter gene, which is assembled via DNA synthesis and molecular cloning techniques known in the art and inserted into mammalian cells using electroporation, chemical transfection, virus-mediated integration, or virus-mediated episomal introduction. The reporter gene is a luminescent enzyme (e.g., firefly luciferase, nanoluciferase, renyral luciferase, or gausyl luciferase), a fluorescent protein (e.g., green fluorescent protein, blue fluorescent protein, or red fluorescent protein), or a colorimetric enzyme (e.g., beta-lactamase or secretory embryonic alkaline phosphatase). The constructs generate a quantifiable signal proportional to the splicing event leading to the inclusion of a start codon, readable by flow cytometry, microscopy, or a multimode microplate reader using a photomultiplier tube. Alternatively, candidate splice regulator-dependent biological activity is evaluated by sequencing the alternatively spliced mRNA transcript produced by the DNA construct using RNA-Seq, or by quantitative reverse transcription PCR (RT-qPCR) in singleton or multiplexed formats. After screening, candidate splice regulators are applied to cells containing the DNA construct for up to 6, 12, 24, or 72 hours, and then evaluated for start codon inclusion activity.
[0589] Example 3: Nucleic acid molecule containing two-site start codons for selective expression of a transgene This embodiment demonstrates the ability of the nucleic acid molecules described herein to selectively control transgene expression in vitro.
[0590] In this embodiment, the nucleic acid molecules of the present disclosure (e.g., Figures 2, 6A–6L, and 7A–7H) incorporated into the vector (Figure 3) are designed as described in Example 1 and produced using methods known in the art. The designed nucleic acid molecules function by an exon inclusion mechanism, five exemplary mechanisms are shown in Figures 4A–4E. Cultured cells (e.g., HEK293 cells) are transfected with the described vector, which contains a two-site start codon (e.g., at least one nucleotide of the start codon is located in a second exon, and at least one nucleotide of the start codon is located in the transgene of the respective nucleic acid molecule) and a transgene, under the control of a splice regulator-dependent switch. After culturing at 37°C and 5% CO2 for 24 hours, the cells are treated with or without a splice regulator (e.g., a small molecule). In the presence of a splice regulator, the bisite start codon is incorporated into the 5' end of the transgene mRNA transcript to produce a full start codon ("on"; Figure 1B), thereby enabling transgene translation. Alternatively, in the absence of a splice regulator (e.g., a small molecule), the start codon is omitted by splicing (Figure 1A), and the transgene is not expressed. After 24 hours of continuous culture, cells are fixed in 4% paraformaldehyde at room temperature for 15 minutes, blocked, and permeabilized in 10% FBS in PBS containing 0.2% surfactant for 45 minutes. The samples are then incubated with a primary antibody against the reporter gene and a secondary antibody conjugated with Alexa 488 in PBS containing 10% PBS and 0.1% surfactant for 2 hours at room temperature or overnight at 4°C. The cells are then washed with PBS and subsequently analyzed using a fluorescence microscope to detect reporter gene expression. The reporter gene is expressed when a splice regulator (e.g., a small molecule) is administered in combination with the vector described above. In the absence of splice regulators, reporter gene expression is not observed.
[0591] Example 4: Regulation of proteins in subjects requiring such regulation by administration of a vector containing a nucleic acid molecule with two-site start codons. Using conventional molecular biology techniques known in the art, nucleic acid molecules containing a dual-site start codon (e.g., at least one nucleotide of the start codon is located in the second exon, and at least one nucleotide of the start codon is located in the transgene of each nucleic acid molecule), and transgenes encoding a target protein, RNA, miRNA, or shRNA (e.g., a therapeutic protein or RNA product) are generated. One or more nucleic acid molecules are then incorporated into a vector, such as a viral vector, and administered to a subject in need, e.g., a subject suffering from a disease associated with a deficiency of the target protein, RNA, miRNA, or shRNA. The subject is administered a viral vector encoding a target protein, RNA, miRNA, or shRNA under the control of a splice regulator-dependent switch that facilitates the incorporation of a dual-site start codon (e.g., at least one nucleotide of the start codon is located in the second exon, and at least one nucleotide of the start codon is located in the transgene) by an alternative splicing event in the presence of a splice regulator. An AAV vector, e.g., a pseudotype AAV2 / 8 or AAV2 / 9 vector, is generated by incorporating the nucleic acid molecule described herein between the 5' terminal repeat and the 3' terminal inverted sequence of the vector, and a two-site start codon (e.g., at least one nucleotide of the start codon is located in the second exon, and at least one nucleotide of the start codon is located in the transgene) is placed under the control of a splice regulator-dependent switch as described in Examples 3 and 4 above. The AAV vector is administered to the subject by various routes, particularly intravenous, intramuscular, or subcutaneous.
[0592] After administering the vector to the target, those skilled in the art can monitor the expression of the transgene by various methods.
[0593] Example 5: Screening of RS1-1 and its derivatives using splice regulatory molecules This embodiment demonstrates the identification of RS1-1 variant constructs that induce gene expression in response to the presence of splice regulators.
[0594] Luciferase Induction Assay HEK-293T cells were cultured in DMEM + 10% FBS, collected after trypsin treatment, and dispensed into 96-well plates (cell culture treatment). 1.0 × 10⁶ cells were placed in 100 μL of culture medium in each well of the plate. 5 Nine cells were seeded. After incubation for 1 hour in a tissue culture incubator (5% CO2, 37°C), plasmid DNA was conjugated with a transfection reagent such as lipofectamine 3000 and added to each well on the plate. 100 ng of plasmid DNA conjugated with 0.15 μL of lipofectamine 3000 was injected into each well. The plasmid used in the experiment encoded the CMV promoter upstream of the RS1-1 variant, which controls the expression of the firefly luciferase gene and subsequently the SV40 polyadenylation signal. After incubation for 10 minutes at room temperature, either DMSO vehicle or a splice regulator compound was added to the wells of the plate at the specified dose. The plate was then incubated in a tissue culture incubator for 18 hours. Endpoint analysis was performed by adding a luciferase substrate, such as Steady-Luc Firefly luciferase substrate, to each well of the assay plate. Each liquid was transferred to a 96-well plate, mixed again on an orbital shaker for 2 minutes, and incubated at room temperature for 5 minutes. Finally, the luminescence signal from each well was measured using a plate reader.
[0595] RS1-1 and its variants were screened for induction of luminescence signals using either a DMSO vehicle or 1A (Figure 9A), and this assay was performed. RS1-1, RS1-2, RS1-3, RS1-4, RS1-7, and RS1-9 show an increase in luciferase induction multiplier with increasing concentration of splice regulators.
[0596] The RS1-1 variant was screened for induction of luminescence signaling using either a DMSO vehicle or 1A (Figure 9B), and this assay was performed. The constructs show an increase in the luciferase induction factor with increasing concentrations of splice regulators.
[0597] Exemplary RS1-10s are screened with splice regulator molecules. HEK-293T cells were cultured in DMEM + 10% FBS, harvested after trypsin treatment, and then aliquoted into 96-well plates (cell culture treatment). 1.0 × 10⁶ cells were placed in 100 μL of culture medium in each well of the plate. 5Nine cells were seeded. After incubation for 1 hour in a tissue culture incubator (5% CO2, 37°C), plasmid DNA was complexed with a transfection reagent such as lipofectamine 3000 and added to each well on the plate. Each well contained 0.15 μL of lipofectamine and an equivalent of 100 ng of plasmid DNA complexed with lipofectamine. Triple wells were used for each condition tested in the experiment. The plasmids used in the experiment encoded the CMV promoter upstream of the RS1-10 variant, which controls the expression of the firefly luciferase gene and subsequently the SV40 polyadenylation signal. After incubation for 10 minutes at room temperature, either a DMSO vehicle or a specified dose of a particular compound was added to the wells of the plate. The plate was then incubated in a tissue culture incubator for 18 hours. Endpoint analysis was performed by adding 100 μL of Steady-Luc Firefly luciferase substrate to each well of the assay plate. After mixing, each liquid content was transferred to a 96-well plate, mixed again on an orbital shaker for 2 minutes, and incubated at room temperature for 5 minutes. Finally, the luminescence signal from each well was measured using a PerkinElmer Envision plate reader. The results of this assay show that RS1-10 responded to most splice regulator compounds used in a dose-responsive manner (Figure 9C). The bars represent the mean ± standard deviation of the 3 wells normalized to the vehicle control. The mean fold inductance of luciferase (relative to the vehicle control) for the various compounds tested is summarized in Table 4.
[0598] Dose-response profiling assay Single copy insertion of the switch construct and subsequent cell line generation were performed in Flp-In-293 cells according to the manufacturer's protocol. The RS1-10 sequence fused to the Firefly luciferase gene was cloned into a Flp-In compatible pcDNA5 / FRT expression vector using the CMV promoter and BGH polyadenylation signal. The resulting Flp-In cell line carried a single copy expression cassette (RSwitch controlled) integrated into the genome. RS1-10 Flp-In cell lines were cultured in DMEM + 10% FBS, collected after trypsin treatment, and then aliquoted into 96-well plates (cell culture treatment). 1.0 × 10⁶ cells were added to each well of the plate in 100 μL of culture medium. 5 Cells were seeded. Then, either a DMSO vehicle or a specific dose of a splice regulator compound was added to the plate wells. Triple wells were used for each condition tested in the experiment. The plates were then incubated in a tissue culture incubator (5% CO2, 37°C) for 18 hours. Endpoint analysis was performed by adding a luciferase substrate, such as 100 μL of Steady-Luc Firefly luciferase substrate, to each well of the assay plate. Each liquid content was transferred to a 96-well plate, mixed again on an orbital shaker for 2 minutes, and incubated at room temperature for 5 minutes. Finally, the luminescence signal from each well was measured with a plate reader. Each plotted data point on the graph in the figure represents the mean ± standard deviation of the 3 wells normalized to the vehicle control. Nonlinear regression curve fittings of the plotted data points were generated using graphing software.
[0599] The dose-response of RS1-10 for specific concentrations of 22A (Figure 9D), 24A (Figure 9E), or 34A (Figure 9F) was measured. RS1-10 indicates an increase in the induction factor of the luciferase signal with increasing concentration of 22A (Figure 9D), 24A (Figure 9E), or 34A (Figure 9F).
[0600] Comparison of RS1-10 Flp-In cell lines and constitutive Flp-In cell lines Single copy insertion of the RS1-10 sequence and subsequent cell line generation were performed in Flp-In-293 cells according to the manufacturer's protocol. The RS1-10 sequence fused to the Firefly luciferase gene was cloned into a Flp-In compatible pcDNA5 / FRT expression vector using the CMV promoter and BGH polyadenylation signal. The resulting Flp-In cell line carried a single copy expression cassette of RS1-10 integrated into the genome. Similarly, a cell line constitutively expressing firefly luciferase was generated in Flp-In-293 cells from a single copy genome insertion. For the constitutive cell line, only the firefly luciferase gene was cloned into a pcDNA5 / FRT expression vector, which was used to generate the single copy genome insertion. Both cell lines were cultured in DMEM + 10% FBS, collected after trypsin treatment, and then aliquoted into 96-well plates (cell culture treatment). In each well of the plate, add 1.0 × 10⁶ of culture medium in 100 μL. 5 The cells were seeded. Subsequently, either the specified dose of DMSO vehicle or 24A was added to the wells of a plate containing RS1-10 Flp-In 293 cells. Triple wells were used for each condition tested in the experiment. The plates were then incubated in a tissue culture incubator (5% CO2, 37°C) for 18 hours. Endpoint analysis was performed by adding a luciferase substrate, such as Steady-Luc Firefly luciferase substrate, to each well of the assay plate. Each liquid content was transferred to a 96-well plate, mixed again on an orbital shaker for 2 minutes, and incubated at room temperature for 5 minutes. The results are shown in Figure 9G. The bars on the graph represent the mean ± standard deviation of the 3 wells normalized to 100% for constitutively expressed firefly luciferase cell lines. RS1-10 Flp-In-293 cells showed at least 60% luciferase signaling at 7 nM in the presence of 24A, and 200 nM of 24A resulted in luciferase expression similar to that of the Flp-In 293 cell line, which constitutively expresses firefly luciferase.
[0601] Example 6: RS1-10 performance in various cell lines This example demonstrates the response of RS1-10 to the presence of splice regulators in various cell lines.
[0602] Human HEK-293T cells or NIH-3T3 mouse fibroblasts were cultured in DMEM + 10% FBS, collected after trypsin treatment, and then dispensed into 96-well plates (cell culture treatment). 1.0 × 10⁶ cells were placed in 100 μL of culture medium in each well of the plate. 5Nine cells were seeded. After incubation for 1 hour in a tissue culture incubator (5% CO2, 37°C), plasmid DNA was complexed with a transfection reagent such as lipofectamine 3000 and added to each well on the plate. Each well received 100 ng of plasmid DNA equivalent complexed with 0.15 μL of lipofectamine 3000. Triple wells were used for each condition tested in the experiment. The plasmids used in the HEK-293T cell and NIH-3T3 mouse fibroblast experiments encoded the CMV promoter upstream of the RS1-10 construct, which controls the expression of the firefly luciferase gene and subsequently the SV40 polyadenylation signal. After incubation for 10 minutes at room temperature, either DMSO vehicle or the specified dose of 24A was added to the wells of the plate. The plate was then incubated in a tissue culture incubator for 18 hours. Endpoint analysis was performed by adding a luciferase substrate, such as Steady-Luc Firefly luciferase substrate, to each well of the assay plate. Each liquid content was transferred to a 96-well plate, mixed again on an orbital shaker for 2 minutes, and incubated at room temperature for 5 minutes. Finally, the luminescence signal from each well was measured with a plate reader. The results are shown in Figures 10A and 10B. RS-10 shows the increase in luciferase induction ratio with increasing concentrations of 24A in HEK-293T cells (Figure 10A) and NIH-373 cells (Figure 10B). The bars represent the mean ± standard deviation of the 3 wells normalized to the vehicle control. RS1-10 shows the increase in luciferase induction ratio with increasing concentrations of 24A.
[0603] Human SH-SY5Y neuroblastoma cell lines or human liver HepG2 cell lines were cultured in EMEM + 10% FBS, collected after trypsin treatment, and then aliquoted into 96-well plates (cell culture treatment). 1.0 × 10¹⁴ cells were placed in 100 μL of culture medium in each well of the plate. 5Nine cells were seeded. After incubation for 1 hour in a tissue culture incubator (5% CO2, 37°C), plasmid DNA was complexed with a transfection reagent such as lipofectamine 3000 and added to each well on a plate. Each well received 100 ng of plasmid DNA equivalent complexed with 0.15 μL of lipofectamine 3000 (Opti-MEM I was used as a diluent). Triple wells were used for each condition tested in the experiment. The plasmid used in the human SH-SY5Y neuroblastoma cell line encoded either the CBA promoter upstream of the RS1-10 construct, which controls the expression of the firefly luciferase gene and subsequently the SV40 polyadenylation signal, or the cloned human synapsin promoter. The plasmid used in HepG2 cells encoded the CBA promoter upstream of the RS1-10 construct, which controls the expression of the firefly luciferase gene and subsequently the SV40 polyadenylation signal.
[0604] After incubation at room temperature for 10 minutes, either DMSO vehicle or 1A was added to the plate wells. The plates were then incubated in a tissue culture incubator for 18 hours. Endpoint analysis was performed by adding a luciferase substrate, such as Steady-Luc Firefly luciferase substrate, to each well of the assay plate. Each liquid content was transferred to a 96-well plate, mixed again on an orbital shaker for 2 minutes, and incubated at room temperature for 5 minutes. Finally, the luminescence signal from each well was measured with a plate reader. RS-10 shows the luciferase induction factor in SH-SY5Y cells in the presence of 1A with either a CBA promoter or a synapsin promoter (Figure 10C). The bars represent the mean ± standard deviation of 3 wells normalized to the vehicle control. RS-10 shows the luciferase induction factor in HepG2 cells with a CBA promoter in the presence of 1A (Figure 10D). The bars represent the mean ± standard deviation of 3 wells normalized to the vehicle control.
[0605] Example 7: RS1-10 performance in mice This embodiment demonstrates the design of an AAV cassette including RS1-10, where the exemplary switch RS1-10 rapidly expresses luciferase after treatment with a splice regulator, and the expression is non-constitutive.
[0606] In vivo mouse study to demonstrate proof of concept for viral (AAV) gene therapy using RS1-10 C57BL / 6 mice received tail vein injection of the 3.00E+11 vector (shown in Figure 11A), a genome / mouse AAV vector consisting of a genome containing a cloned human synapsin promoter that drives the expression of the firefly luciferase gene. Luciferase gene expression was gated by the RS1-10 switch. The AAV vector utilized a brain-permeable AAV-PHP.eB capsid (Figure 11A). Approximately three weeks after AAV delivery, mice were administered either the compound vehicle or 10 mg / kg 24A as a single oral gastric tube feeding. Six hours after oral administration, mice were intraperitoneally administered D-luciferin (150 mg / kg), anesthetized, and imaged using a bioluminescence imaging (BLI) system to analyze RS1-10 activa...
Claims
1. A nucleic acid molecule containing a minigene located adjacent to the 5' side of the introduced gene, wherein the minigene is (a) The first exon adjacent to the 5' side of the first intron, (b) Start codon and (c) A nucleic acid molecule comprising a splice regulatory factor binding site, wherein the splice regulatory factor binding site includes the nucleic acid sequence DGAGTDDGHV (SEQ ID NO: 82) or DGAGTDDNHV (SEQ ID NO: 83), wherein D is A, G, or T, R is A or G, N is A, C, G, or T, H is A, C, or T, and V is A, G, or C.
2. The nucleic acid molecule according to claim 1, wherein the splice regulatory factor binding site comprises the nucleic acid sequence DGAGTRRGHV (SEQ ID NO: 1) or DGAGTRRRNHV (SEQ ID NO: 2), wherein D is A, G, or T, R is A or G, N is A, C, G, or T, H is A, C, or T, and V is A, G, or C.
3. The nucleic acid molecule according to claim 1, wherein the splice regulatory factor binding site comprises the nucleic acid sequence DGAGTTTGHV (SEQ ID NO: 84), where D is A, G, or T, H is A, C, or T, and V is A, G, or C.
4. The nucleic acid molecule according to any one of claims 1 to 4, wherein the nucleic acid includes the splice regulatory factor binding site which includes any one of sequence numbers 75 to 81.
5. The nucleic acid molecule according to any one of claims 1 to 4, wherein the nucleic acid comprises, in order from 5' to 3', the first exon containing the start codon, the first intron, and the transgene.
6. A nucleic acid molecule according to any one of claims 1 to 5, further comprising a stop codon.
7. The nucleic acid molecule according to any one of claims 1 to 5, further comprising a second exon.
8. The nucleic acid molecule according to claim 7, wherein the nucleic acid comprises, in the order from 5' to 3', the first exon containing the start codon, the second exon, and the introduced gene.
9. The nucleic acid molecule according to claim 7, wherein the nucleic acid comprises, in order from 5' to 3', the first exon containing the start codon, the first intron, the second intron, and the transgene.
10. The nucleic acid molecule according to claim 7, wherein the nucleic acid comprises, in the order from 5' to 3', the first exon, the second exon containing the start codon, and the transgene.
11. The nucleic acid molecule according to claim 7, wherein the nucleic acid comprises, in order from 5' to 3', the first exon, the second exon containing the start codon, the first intron, and the transgene.
12. The nucleic acid molecule according to claim 7, wherein the nucleic acid comprises, in order from 5' to 3', the first exon, the first intron, the second exon including the start codon, the second intron, and the transgene.
13. The nucleic acid molecule according to claim 8, wherein the nucleic acid has at least about 80% sequence identity with any one of sequence numbers 21 to 73.
14. The nucleic acid molecule according to claim 8, wherein the nucleic acid has at least about 90% sequence identity with any one of sequence numbers 21 to 73.
15. The nucleic acid molecule according to claim 8, wherein the nucleic acid comprises one of sequence numbers 21 to 73.
16. A nucleic acid molecule according to any one of claims 6 to 15, further comprising a third exon.
17. A nucleic acid molecule according to any one of claims 6 to 16, further comprising a third intron.
18. The nucleic acid molecule according to claim 16, wherein the nucleic acid comprises, in order from 5' to 3', the first exon, the second exon containing the stop codon, the third exon containing the start codon, and the introduced gene.
19. The nucleic acid molecule according to claim 17, wherein the nucleic acid comprises, in order from 5' to 3', the first exon, the first intron, the second exon including the stop codon, the second intron, the third exon including the start codon, the third intron, and the transgene.
20. The nucleic acid molecule according to claim 16, wherein the nucleic acid comprises, in order from 5' to 3', the first exon, the first intron, the second exon containing the stop codon, the third exon containing the start codon, the second intron, and the transgene.
21. The nucleic acid molecule according to claim 16, wherein the nucleic acid comprises, in order from 5' to 3', the first exon, the second exon containing the start codon, the third exon containing the stop codon, and the introduced gene.
22. The nucleic acid molecule according to claim 16, wherein the nucleic acid comprises, in order from 5' to 3', the first exon, the first intron, the second exon including the start codon, the second intron, the third exon including the stop codon, the third intron, and the transgene.
23. The nucleic acid molecule according to claim 16, wherein the nucleic acid comprises, in order from 5' to 3', the first exon, the first intron, the second exon including the start codon, the third exon including the stop codon, the second intron, and the transgene.
24. A nucleic acid molecule containing a minigene located adjacent to the 5' side of an introduced gene, wherein the minigene is (a) The first exon located on the 5' side of the first intron, (b) A start codon comprising a first part and a second part, wherein the first part and the second part of the start codon are not in the same exon, (c) A nucleic acid molecule comprising a splice regulatory factor binding site, which in some embodiments comprises the nucleic acid sequence DGAGTDDGHV (SEQ ID NO: 82) or DGAGTDDNHV (SEQ ID NO: 83), wherein D is A, G, or T, R is A or G, N is A, C, G, or T, H is A, C, or T, and V is A, G, or C.
25. The nucleic acid molecule according to claim 24, wherein the splice regulatory factor binding site comprises the nucleic acid sequence DGAGTRRGHV (SEQ ID NO: 1) or DGAGTRRRNHV (SEQ ID NO: 2), wherein D is A, G, or T, R is A or G, N is A, C, G, or T, H is A, C, or T, and V is A, G, or C.
26. The nucleic acid molecule according to claim 24, wherein the splice regulatory factor binding site comprises the nucleic acid sequence DGAGTTTGHV (SEQ ID NO: 84), where D is A, G, or T, H is A, C, or T, and V is A, G, or C.
27. The nucleic acid molecule according to any one of claims 24 to 26, wherein the first portion of the start codon comprises one or two nucleotides of the start codon, and the second portion of the start codon comprises one or two nucleotides of the start codon.
28. The nucleic acid molecule according to any one of claims 24 to 27, wherein the first portion of the start codon is located in the first exon, and the second portion of the start codon is located in the transgene.
29. The nucleic acid molecule according to claim 28, wherein the nucleic acid comprises, in order from 5' to 3', a first exon including the first portion of the start codon, a first intron, and the transgene including the second portion of the start codon.
30. A nucleic acid molecule according to any one of claims 24 to 29, further comprising a second exon.
31. A nucleic acid molecule according to any one of claims 24 to 30, further comprising a second intron.
32. The nucleic acid molecule according to claim 31, wherein the nucleic acid comprises, in order from 5' to 3', a first exon containing the first portion of the start codon, a second exon containing the second portion of the start codon, and the introduced gene.
33. The nucleic acid molecule according to claim 31, wherein the nucleic acid comprises, in order from 5' to 3', a first exon containing the first portion of the start codon, a first intron, a second exon containing the second portion of the start codon, a second intron, and the transgene.
34. The nucleic acid molecule according to claim 31, wherein the nucleic acid comprises, in the order from 5' to 3', the first exon, the second exon including the first portion of the start codon, and the transgene including the second portion of the start codon.
35. The nucleic acid molecule according to claim 31, wherein the nucleic acid comprises, in order from 5' to 3', the first exon, the second exon including the first portion of the start codon, the first intron, and the transgene including the second portion of the start codon.
36. The nucleic acid molecule according to claim 31, wherein the nucleic acid comprises, in the order from 5' to 3', a first exon including the first portion of the start codon, a second exon, and the transgene including the second portion of the start codon.
37. The nucleic acid molecule according to claim 31, wherein the nucleic acid comprises, in order from 5' to 3', a first exon including the first portion of the start codon, a second exon, a first intron, and the transgene including the second portion of the start codon.
38. The nucleic acid molecule according to claim 31, wherein the nucleic acid comprises, in the order from 5' to 3', a first exon containing the first portion of the start codon, a second exon containing the second portion of the start codon, and the introduced gene.
39. The nucleic acid molecule according to claim 31, wherein the nucleic acid comprises, in order from 5' to 3', a first exon containing the first portion of the start codon, a first intron, a second exon containing the second portion of the start codon, and the introduced gene.
40. A nucleic acid molecule according to any one of claims 24 to 39, further comprising a stop codon.
41. The nucleic acid molecule according to claim 40, wherein the nucleic acid comprises, in the order of 5' to 3', a first exon containing the first portion of the start codon, a second exon containing the stop codon, and the transgene containing the second portion of the start codon.
42. The nucleic acid molecule according to claim 40, wherein the nucleic acid comprises, in order from 5' to 3', a first exon containing the first portion of the start codon, a first intron, a second exon containing the stop codon, and the transgene containing the second portion of the start codon.
43. The nucleic acid molecule according to any one of claims 30 to 42, wherein the second exon includes the splice regulator binding site.
44. A nucleic acid molecule according to any one of claims 30 to 43, further comprising a third exon.
45. A nucleic acid molecule according to any one of claims 30 to 44, further comprising a third intron.
46. The nucleic acid molecule according to claim 45, wherein the nucleic acid comprises a first exon containing the first portion of the start codon in the order of 5' to 3', a second exon containing the stop codon, a third exon containing the second portion of the start codon, and the introduced gene.
47. The nucleic acid molecule according to claim 45, wherein the nucleic acid comprises, in order from 5' to 3', a first exon containing the first portion of the start codon, a first intron, a second exon containing the stop codon, a third exon containing the second portion of the start codon, a second intron, and the transgene.
48. The nucleic acid molecule according to claim 45, wherein the nucleic acid comprises the first exon, the second exon including the first portion of the start codon, the third exon including the stop codon, and the transgene including the second portion of the start codon, in the order of 5' to 3'.
49. The nucleic acid molecule according to claim 45, wherein the nucleic acid comprises, in order from 5' to 3', the first exon, the first intron, the second exon including the first portion of the start codon, the third exon including the stop codon, the second intron, and the transgene including the second portion of the start codon.
50. The nucleic acid molecule according to claim 45, wherein the nucleic acid comprises the first exon in the order of 5' to 3', the second exon comprising the stop codon, the third exon comprising the first portion of the start codon, and the introduced gene comprising the second portion of the start codon.
51. The nucleic acid molecule according to claim 45, wherein the nucleic acid comprises, in order from 5' to 3', the first exon, the first intron, the second exon including the stop codon, the third exon including the first portion of the start codon, the second intron, and the transgene including the second portion of the start codon.
52. The nucleic acid molecule according to claim 45, wherein the nucleic acid comprises, in the order from 5' to 3', the first exon, the second exon, the third exon including the first portion of the start codon, and the transgene including the second portion of the start codon.
53. The nucleic acid molecule according to claim 45, wherein the nucleic acid comprises, in order from 5' to 3', the first exon, the first intron, the second exon, the third exon including the first portion of the start codon, the second intron, and the transgene including the second portion of the start codon.
54. The nucleic acid molecule according to claim 45, wherein the nucleic acid comprises, in order from 5' to 3', the first exon, the second exon containing the stop codon, the third exon containing the first portion of the start codon, and the transgene containing the second portion of the start codon.
55. The nucleic acid molecule according to claim 45, wherein the nucleic acid comprises, in order from 5' to 3', the first exon, the first intron, the second exon including the stop codon, the second intron, the third exon including the first portion of the start codon, the third intron, and the transgene including the second portion of the start codon.
56. The nucleic acid molecule according to claim 45, wherein the nucleic acid comprises, in the order of 5' to 3', the first exon, the second exon comprising the first portion of the start codon, the third exon comprising the stop codon, and the transgene comprising the second portion of the start codon.
57. The nucleic acid molecule according to claim 45, wherein the nucleic acid comprises, in order from 5' to 3', the first exon, the first intron, the second exon including the first portion of the start codon, the second intron, the third exon including the stop codon, the third intron, and the transgene including the second portion of the start codon.
58. The nucleic acid molecule according to any one of claims 24 to 58, wherein the first intron includes the splice regulator binding site.
59. The nucleic acid molecule according to any one of claims 24 to 58, wherein the splice regulator binding site is located at the junction between the second exon and the second intron.
60. The nucleic acid molecule according to any one of claims 24 to 58, wherein the splice regulatory factor binding site is located within the second intron.
61. A nucleic acid molecule, in the order from 5' to 3', (a) A first exon containing the first portion of the start codon at the 3' end of the first exon, (b) A second exon having the second portion of the start codon at its 5' end, (c) A nucleic acid molecule containing a third exon.
62. The nucleic acid molecule according to claim 61, wherein the third exon is a transgene.
63. A nucleic acid molecule containing a transgene, in the order of 5' to 3', the first exon, the second exon containing the stop codon, the third exon containing the first portion of the start codon, and the second portion of the start codon.
64. A nucleic acid molecule containing a first exon, a second exon containing the first portion of the start codon, a third exon containing the stop codon, and a transgene containing the second portion of the start codon, in the order of 5' to 3'.
65. The nucleic acid molecule according to any one of claims 63 to 64, further comprising a first intron, a second intron, and a third intron interposed between the first exon, the second exon, the third exon, and the introduced gene, respectively.
66. The nucleic acid molecule according to any one of claims 1 to 65, wherein the splice regulatory factor binding site does not contain the nucleic acid sequence AAGAGT (SEQ ID NO: 3), ATGAGT (SEQ ID NO: 4), TAGAGT (SEQ ID NO: 5), TTGAGT (SEQ ID NO: 6), GAGAGT (SEQ ID NO: 7), GTGAGT (SEQ ID NO: 8), AAGAGT (SEQ ID NO: 9), ATGAGT (SEQ ID NO: 10), ACGAGT (SEQ ID NO: 11), AGGAGT (SEQ ID NO: 12), AGAGGTAGAG (SEQ ID NO: 13), TGAGGGTTTGAG (SEQ ID NO: 14), GGAGGTGGAG (SEQ ID NO: 15), TAG (SEQ ID NO: 16), CAG (SEQ ID NO: 17), TAG (SEQ ID NO: 18), or NAGAGTNNNNN (SEQ ID NO: 19) (wherein N is A, C, G, or T).
67. The nucleic acid molecule according to any one of claims 1 to 66, wherein the splice regulatory factor binding site includes the splice regulatory factor binding site which includes any one of sequence numbers 75 to 81.
68. A composition comprising a nucleic acid molecule according to any one of claims 1 to 67 and a splice regulator, wherein the splice regulator is a compound of formula (I), 【Chemistry 27】 or comprising a pharmaceutically acceptable salt, solvate, or prodrug thereof, in the formula, W is -S- or -HC=CH-, R 1 is H, halogen, hydroxyl, cyano, C 1 -C 6 alkyl, C 2 -C 6 alkenyl, C 2 -C 6 alkynyl, C 1 -C 6 haloalkyl, C 1 -C 6 alkoxyl, C 1 -C 6 haloalkoxyl, -(CH 2 ) 0-2 -C 3 -C 8 cycloalkyl, NH 2 , NH(C 1 -C 6 alkyl), N(C 1 -C 6 alkyl) 2 , or -(CH 2 ) 0-2 -heterocyclyl (wherein the heterocyclyl is a 4- to 7-membered ring or contains 1, 2, or 3 heteroatoms independently selected from N, O, and S, and the alkyl, alkenyl, alkynyl, alkoxyl, cycloalkyl, heterocyclyl may be optionally substituted with one or more C 3 -C 8 cycloalkyl, and the aryl or 4- to 7-membered heterocyclyl contains 1, 2, or 3 heteroatoms independently selected from N, O, and S), R 2 However, the aryl, cycloalkyl, heterocyclil, or heteroaryl is a 5, 6, or 9-membered heterocyclil containing 1, 2, or 3 heteroatoms independently selected from aryl, 5-7 membered cycloalkyl, N, O, and S, or a 5, 6, or 9-membered heteroaryl containing 2, 3, or 4 heteroatoms independently selected from N, O, and S, wherein the aryl, cycloalkyl, heterocyclil, or heteroaryl contains one or more R 4 It may be optionally replaced by, Each R 3 These are independent of halogen and C 1 -C 6 Alkyl, C 2 -C 6 Alkenil, C 2 -C 6 Alkinyl, C 1 -C 6 Alkoxyl, C 3 -C 8 Cycloalkyl, NH 2 NH(C 1 -C 6 Alkyl), or N(C) 1 -C 6 Alkyl) 2 The alkyl, alkenyl, alkynyl, alkoxyl, and cycloalkyl groups consist of one or more hydroxyls or NH groups. 2 It may be optionally replaced by, Each R 4 These independently comprise halogen, hydroxyl, cyano, nitro, and C. 1 -C 6 Alkyl, C 2 -C 6 Alkenil, C 2 -C 6 Alkinyl, C 1 -C 6 Haloalkyl, C 1 -C 6 Alkoxyl, C 1 -C 6 Haloalkoxyl, C 3 -C 8 Cycloalkyl, NH 2 NH(C 1 -C 6 Alkyl), N (C 1 -C 6 Alkyl) 2 , or C(O)NH 2 The alkyl, alkenyl, alkynyl, alkoxyl, and cycloalkyl elements are N, O, and S, NH 2 NH(C 1 -C 6 Alkyl), or N(C) 1 -C 6 Alkyl) 2 It may be optionally substituted with one or more hydroxyl, 4-membered to 7-membered heterocyclines containing one, two, or three heteroatoms independently selected from the above, R 5 is H, C 1 -C 6 alkyl, C 2 -C 6 alkenyl, C 2 -C 6 alkynyl, C 1 -C 6 alkoxyl, C 3 -C 8 cycloalkyl, -CH 2 C 3 -C 8 cycloalkyl, heterocyclyl, -CH 2 heterosilyl, -CH 2 CH 2 heterosilyl, -CH 2 -(5- to 6-membered heteroaryl), where the heterocyclyl is a 4- to 7-membered ring and contains 1, 2, or 3 heteroatoms independently selected from N, O, and S, and wherein R 5 is one or more halogens, C 1 -C 6 alkyl, C 1 -C 6 haloalkyl, C 1 -C 6 heteroalkyl, C 1 -C 6 alkoxyl, C 3 -C 8 cycloalkyl, spiro C 3 -C 8 cycloalkyl, spiro 4- to 7-membered heterocyclyl, 5- to 6-membered heteroaryl, oxo, cyano, or hydroxyl, and may be optionally substituted R 6 H, halogen, C 1 -C 6 Alkyl, or C 1 -C 6 It is a haloalkyl, R 7 However, H, halogen, C 1 -C 6 Alkyl, or C 1 -C 6 It is a haloalkyl, A composition in which n is 0, 1, 2, 3, 4, or 5.
69. A composition, (i) A nucleic acid molecule containing a minigene located adjacent to the 5' side of the transgene, wherein the minigene is a) The exon located at the 5' end of the intron, b) A start codon having a first part and a second part, wherein the first part and the second part of the start codon are not in the same exon, c) A nucleic acid molecule comprising a splice regulatory factor binding site, (ii) Binds to the splice regulatory factor binding site, formula (I): 【Chemistry 28】 Splice regulators including structures that follow (I), or comprising a pharmaceutically acceptable salt, solvate, or prodrug thereof, in the formula, W is -S- or -HC=CH-, R 1 However, H, halogen, hydroxyl, cyano, C 1 -C 6 Alkyl, C 2 -C 6 Alkenil, C 2 -C 6 Alkinyl, C 1 -C 6 Haloalkyl, C 1 -C 6 Alkoxyl, C 1 -C 6 Haloalkoxyl, -(CH 2 ) 0-2 -C 3 -C 8 Cycloalkyl, NH 2 NH(C 1 -C 6 Alkyl), N (C 1 -C 6 Alkyl) 2 , or - (CH 2 ) 0-2 - A heterocyclyl, wherein the heterocyclyl is a 4- to 7-membered ring and contains 1, 2, or 3 heteroatoms independently selected from N, O, and S, and the alkyl, alkenyl, alkynyl, alkoxyl, cycloalkyl, heterocyclyl contains one or more C 3 -C 8 It may be optionally substituted with a cycloalkyl, aryl, or 4- to 7-membered heterocycline containing one, two, or three heteroatoms independently selected from N, O, and S. R 2 This is a 5, 6, or 9-membered heterocyclyl comprising one, two, or three heteroatoms independently selected from aryl, 5- to 7-membered cycloalkyl, N, O, and S, or a 5, 6, or 9-membered heteroaryl comprising two, three, or four heteroatoms independently selected from N, O, and S, wherein the aryl, cycloalkyl, heterocyclyl, or heteroaryl comprises one or more R 4 It may be optionally replaced by, Each R 3 These are independently halogen, C 1 -C 6 Alkyl, C 2 -C 6 Alkenil, C 2 -C 6 Alkinyl, C 1 -C 6 Alkoxyl, C 3 -C 8 Cycloalkyl, NH 2 NH(C 1 -C 6 Alkyl), or N(C) 1 -C 6 Alkyl) 2 The alkyl, alkenyl, alkynyl, alkoxyl, and cycloalkyl groups consist of one or more hydroxyls or NH groups. 2 It may be optionally replaced by, Each R 4 These are independently halogen, hydroxyl, cyano, nitro, and C 1 -C 6 Alkyl, C 2 -C 6 Alkenil, C 2 -C 6 Alkinyl, C 1 -C 6 Haloalkyl, C 1 -C 6 Alkoxyl, C 1 -C 6 Haloalkoxyl, C 3 -C 8 Cycloalkyl, NH 2 NH(C 1 -C 6 Alkyl), N (C 1 -C 6 Alkyl) 2 , or C(O)NH 2 The alkyl, alkenyl, alkynyl, alkoxyl, and cycloalkyl are 4-membered to 7-membered heterocyclines containing one, two, or three heteroatoms independently selected from one or more hydroxy, N, O, and S. 2 NH(C 1 -C 6 Alkyl), or N(C) 1 -C 6 Alkyl) 2 It may be optionally replaced by, R 5 However, H, C 1 -C 6 Alkyl, C 2 -C 6 Alkenil, C 2 -C 6 Alkinyl, C 1 -C 6 Alkoxyl, C 3 -C 8 Cycloalkyl, -CH 2 C 3 -C 8 Cycloalkyl, heterocyclyl, -CH 2 Heterocisyl, -CH 2 CH 2 Heterocisyl, -CH 2 - (5-6 membered heteroaryl), wherein the heterocyclyl is a 4-7 membered ring and contains 1, 2, or 3 heteroatoms independently selected from N, O, and S, where R 5 is one or more halogens, C 1 -C 6 Alkyl, C 1 -C 6 Haloalkyl, C 1 -C 6 Heteroalkyl, C 1 -C 6 Alkoxyl, C 3 -C 8 Cycloalkyl, Spiro C 3 -C 8 They may be optionally substituted with cycloalkyl, spiro-4-7 member heterocyclyl, 5-6 member heteroaryl, oxo, cyano, or hydroxyl. R 6 H, halogen, C 1 -C 6 Alkyl, or C 1 -C 6 It is a haloalkyl, R 7 However, H, halogen, C 1 -C 6 Alkyl, or C 1 -C 6 It is a haloalkyl, A composition in which n is 0, 1, 2, 3, 4, or 5.
70. A composition, (i) A nucleic acid molecule containing a minigene located adjacent to the 5' side of the transgene, wherein the minigene is (a) The first exon adjacent to the 5' side of the first intron, (b) Start codon and (c) A nucleic acid molecule containing a splice regulatory factor binding site, (ii) A splice regulator that binds to the splice regulator binding site and includes a structure according to formula (I): 【Chemistry 29】 or comprising a pharmaceutically acceptable salt, solvate, or prodrug thereof, in the formula, W is -S- or -HC=CH-, R 1 However, H, halogen, hydroxyl, cyano, C 1 -C 6 Alkyl, C 2 -C 6 Alkenil, C 2 -C 6 Alkinyl, C 1 -C 6 Haloalkyl, C 1 -C 6 Alkoxyl, C 1 -C 6 Haloalkoxyl, -(CH 2 ) 0-2 -C 3 -C 8 Cycloalkyl, NH 2 NH(C 1 -C 6 Alkyl), N (C 1 -C 6 Alkyl) 2 , or - (CH 2 ) 0-2 - A heterocyclyl, the heterocyclyl being a 4- to 7-membered ring containing 1, 2, or 3 heteroatoms independently selected from N, O, and S, wherein the alkyl, alkenyl, alkynyl, alkoxyl, cycloalkyl, or heterocyclyl is composed of one or more C 3 -C 8 It may be optionally substituted with a cycloalkyl, aryl, or 4- to 7-membered heterocycline containing 1, 2, or 3 heteroatoms independently selected from N, O, and S. R 2 This is a 5, 6, or 9-membered heterocyclyl comprising one, two, or three heteroatoms independently selected from aryl, 5- to 7-membered cycloalkyl, N, O, and S, or a 5, 6, or 9-membered heteroaryl comprising two, three, or four heteroatoms independently selected from N, O, and S, wherein the aryl, cycloalkyl, heterocyclyl, or heteroaryl comprises one or more R 4 It may be optionally replaced by, Each R 3 These are independently halogen, C 1 -C 6 Alkyl, C 2 -C 6 Alkenil, C 2 -C 6 Alkinyl, C 1 -C 6 Alkoxyl, C 3 -C 8 Cycloalkyl, NH 2 NH(C 1 -C 6 Alkyl), or N(C) 1 -C 6 Alkyl) 2 The alkyl, alkenyl, alkynyl, alkoxyl, and cycloalkyl groups consist of one or more hydroxyls or NH groups. 2 It may be optionally replaced by, Each R 4 These are independently halogen, hydroxyl, cyano, nitro, and C 1 -C 6 Alkyl, C 2 -C 6 Alkenil, C 2 -C 6 Alkinyl, C 1 -C 6 Haloalkyl, C 1 -C 6 Alkoxyl, C 1 -C 6 Haloalkoxyl, C 3 -C 8 Cycloalkyl, NH 2 NH(C 1 -C 6 Alkyl), N (C 1 -C 6 Alkyl) 2 , or C(O)NH 2 The alkyl, alkenyl, alkynyl, alkoxyl, and cycloalkyl are 4-membered to 7-membered heterocyclines, NH3, and 1, 2, or 3 heteroatoms independently selected from one or more hydroxy, N, O, and S. 2 NH(C 1 -C 6 Alkyl), or N(C) 1 -C 6 Alkyl) 2 It is also acceptable for it to be replaced by an optional substitution. R 5 However, H, C 1 -C 6 Alkyl, C 2 -C 6 Alkenil, C 2 -C 6 Alkinyl, C 1 -C 6 Alkoxyl, C 3 -C 8 Cycloalkyl, -CH 2 C 3 -C 8 Cycloalkyl, heterocyclyl, -CH 2 Heterocisyl, -CH 2 CH 2 Heterocisyl, -CH 2 - (5-6 membered heteroaryl), wherein the heterocyclyl is a 4-7 membered ring and contains 1, 2, or 3 heteroatoms independently selected from N, O, and S, where R 5 is one or more halogens, C 1 -C 6 Alkyl, C 1 -C 6 Haloalkyl, C 1 -C 6 Heteroalkyl, C 1 -C 6 Alkoxyl, C 3 -C 8 Cycloalkyl, Spiro C 3 -C 8 They may be optionally substituted with cycloalkyl, spiro-4-7 member heterocyclyl, 5-6 member heteroaryl, oxo, cyano, or hydroxyl. R 6 H, halogen, C 1 -C 6 Alkyl, or C 1 -C 6 It is a haloalkyl, R 7 However, H, halogen, C 1 -C 6 Alkyl, or C 1 -C 6 It is a haloalkyl, A composition in which n is 0, 1, 2, 3, 4, or 5.
71. A composition, (i) A nucleic acid molecule containing a minigene located adjacent to the 5' side of the transgene, wherein the minigene is (a) The first exon adjacent to the 5' side of the first intron, (b) Start codon and (c) A nucleic acid molecule containing a splice regulatory factor binding site, (ii) A splice regulator that binds to the splice regulator binding site and includes a structure according to formula (II): 【Transformation 30】 or comprising a pharmaceutically acceptable salt, solvate, or prodrug thereof, in the formula, A is a saturated or partially unsaturated monocyclic or bicyclic 4- to 9-membered heterocycloalkyl or NR 1 R 2 A heterocycloalkyl group contains one or two nitrogen ring atoms and 1, 2, 3, or 4 R atoms. 6 It is arbitrarily replaced with, R 1 However, 1, 2, 3, or 4 R 6 It is a heterocycloalkyl group containing one nitrogen ring atom which may be optionally substituted, R 2 However, hydrogen, C 1-7 Alkyl, or C 3-8 It is a cycloalkyl, R 3 However, H, Haro, C 1-7 Alkyl, OR 5 , N(R 5 ) 2 , C 3-8 It is a cycloalkyl or heterocycloalkyl, R 4 However, it is an aryl or bicyclic nine-membered heteroaryl containing two, three, or four heteroatoms independently selected from N, O, and S, where R 4 is 1, 2, or 3 R 7 It is optionally replaced with, Each R 5 C 1-7 Alkyl, C 3-8 It is a cycloalkyl or heterocycloalkyl, Each R 6 However, halogen, hydroxy, cyano, -COOH, -C(O)-C 1 -C 6 Alkyl, -C(O)O-C 1 -C 6 Alkyl, C 1 -C 7 Alkyl, C 1 -C 8 Heteroalkyl, C 1-7 Alkoxy-heterocycloalkyl, C 2 -C 6 Alkenil, C 2 -C 6 Alkinyl, C 1 -C 6 Alkoxy, -(CH 2 ) 0-2 -C 3 -C 8 Cycloalkyl, 4-7 membered monocyclic heterocycloalkyl, NH 2 NH(C 1 -C 6 Alkyl), N (C 1 -C 6 Alkyl) 2 , -NHC(O)-C 1 -C 6 Alkyl, -N(C) 1 -C 6 Alkyl)-C(O)-C 1 -C 6 Alkyl, -C(O)-NH 2 , -C(O)-NH(C 1 -C 6 Alkyl), and -C(O)-N(C 1 -C 6 Alkyl) 2 Independently selected from the group consisting of, the alkyl, alkenyl, alkynyl, and alkoxy are one or more halogens, hydroxyl, or NH 2 The cycloalkyl and heterocycloalkyl groups may be optionally substituted with one or more halogens, hydroxyls, and C 1 -C 6 Alkyl, C 1 -C 6 Heteroalkyl, C 1 -C 6 Alkoxy, or NH 2 It may be optionally replaced by, Two Rs on the same carbon 6 Can they be combined as keto (=O), or Two R's 6 Both are C 1-7 Forms alkylene, Each R 7 They became independent, Halo, Cyano, C 1-7 Alkyl, C 1-7 Haloalkyl, C 1-7 Alkoxy, C 1-7 Haloalkoxy, or C 3-8 It is a cycloalkyl, and the C 1-7 Alkyl groups may be optionally substituted with OH groups. R 16 However, H, Haro, C 1-7 Alkyl, OR 5 , N(R 5 ) 2 , C 3-8 It is a cycloalkyl or heterocycloalkyl, R 17 However, H, Haro, C 1-7 Alkyl, OR 5 , N(R 5 ) 2 , C 3-8 A composition that is cycloalkyl or heterocycloalkyl.
72. A composition, (i) A nucleic acid molecule containing a minigene located adjacent to the 5' side of the transgene, wherein the minigene is (a) The first exon adjacent to the 5' side of the first intron, (b) Start codon and (c) A nucleic acid molecule containing a splice regulatory factor binding site, (ii) A splice regulator that binds to the splice regulator binding site and includes a structure according to formula (III): 【Chemistry 31】 or comprising a pharmaceutically acceptable salt, solvate, or prodrug thereof, in the formula, A is a saturated or partially unsaturated monocyclic or bicyclic 4- to 9-membered heterocycloalkyl or NR 1 R 2 A heterocycloalkyl group contains one or two nitrogen ring atoms and 1, 2, 3, or 4 R atoms. 6 It is arbitrarily replaced with, R 1 However, 1, 2, 3, or 4 R 6 It is a heterocycloalkyl group containing one nitrogen ring atom which may be optionally substituted, R 2 However, hydrogen, C 1-7 Alkyl, or C 3-8 It is a cycloalkyl, R 3 However, H, Haro, C 1-7 Alkyl, OR 5 , N(R 5 ) 2 , C 3-8 It is a cycloalkyl or heterocycloalkyl, R 4 However, it is an aryl or bicyclic nine-membered heteroaryl containing two, three, or four heteroatoms independently selected from N, O, and S, where R 4 is 1, 2, or 3 R 7 It is optionally replaced with, Each R 5 C 1-7 Alkyl, C 3-8 It is a cycloalkyl or heterocycloalkyl, Each R 6 However, independently, C 1-7 Alkyl, amino, amino-C 1-7 Alkyl, C 3-8 Cycloalkyl, heterocycloalkyl, or C 1-7 It is either an alkoxy-heterocycloalkyl group or two R groups. 6 Both are C 1-7 Forms alkylene, Each R 7 They became independent, Halo, Cyano, C 1-7 Alkyl, C 1-7 Haloalkyl, C 1-7 Alkoxy, C 1-7 Haloalkoxy, or C 3-8 It is a cycloalkyl, and the C 1-7 A composition in which the alkyl group may be optionally substituted with an OH group.
73. The composition according to any one of claims 69 to 72, wherein the splice regulator that binds to the splice regulator binding site is selected from the group consisting of compounds 1A to 192A and 100B to 135B.
74. The composition according to any one of claims 69 to 73, wherein the splice regulator that binds to the splice regulator binding site is selected from the group comprising 3A, 6A, 8A, 10A, 15A, 24A, 86A, 100B, 111B, 117B, 121B, 135B, and 192A.
75. The composition according to any one of claims 69 to 73, wherein the splice regulator that binds to the splice regulator binding site is selected from the group comprising 116B, 100B, 1A, 22A, 24A, 2A, and 34A.
76. The composition according to any one of claims 69 to 73, wherein the splice regulatory factor binding site comprises the nucleic acid sequence DGAGTDDGHV (SEQ ID NO: 82) or DGAGTDDNHV (SEQ ID NO: 83), wherein D is A, G, or T, R is A or G, N is A, C, G, or T, H is A, C, or T, and V is A, G, or C.
77. The composition according to claim 76, wherein the splice regulator binds to the splice regulator binding site comprising the nucleic acid sequence DGAGTRRGHV (SEQ ID NO: 1) or DGAGTRRRNHV (SEQ ID NO: 2), wherein D is A, G, or T, R is A or G, N is A, C, G, or T, H is A, C, or T, and V is A, G, or C.
78. The composition according to claim 76, wherein the splice regulatory factor binding site comprises the nucleic acid sequence DGAGTTTGHV (SEQ ID NO: 84), where D is A, G, or T, H is A, C, or T, and V is A, G, or C.
79. A composition according to any one of claims 76, wherein the splice regulatory factor binding site does not contain the nucleic acid sequence AAGAGT (SEQ ID NO: 3), ATGAGT (SEQ ID NO: 4), TAGAGT (SEQ ID NO: 5), TTGAGT (SEQ ID NO: 6), GAGAGT (SEQ ID NO: 7), GTGAGT (SEQ ID NO: 8), AAGAGT (SEQ ID NO: 9), ATGAGT (SEQ ID NO: 10), ACGAGT (SEQ ID NO: 11), AGGAGT (SEQ ID NO: 12), AGAGGTAGAG (SEQ ID NO: 13), TGAGTTTGAG (SEQ ID NO: 14), GGAGGTGGAG (SEQ ID NO: 15), TAG (SEQ ID NO: 16), CAG (SEQ ID NO: 17), TAG (SEQ ID NO: 18), or NAGAGTNNNNN (SEQ ID NO: 19) (wherein N is A, C, G, or T).
80. The composition according to claim 79, wherein the splice regulator binds to the splice regulator binding site which includes any one nucleic acid sequence of sequence numbers 75 to 81.
81. The nucleic acid molecule according to any one of claims 1 to 67, or the composition according to any one of claims 68 to 80, wherein the introduced gene encodes the target protein.
82. The nucleic acid molecule according to any one of claims 1 to 67, or the composition according to any one of claims 68 to 80, wherein the introduced gene encodes the target miRNA.
83. The nucleic acid molecule according to any one of claims 1 to 67, or the composition according to any one of claims 68 to 80, wherein the introduced gene encodes the target shRNA.
84. A nucleic acid molecule according to any one of claims 1 to 67, or a composition according to any one of claims 68 to 80, wherein the introduced gene encodes a functional RNA or regulatory RNA of the target.
85. The nucleic acid molecule according to any one of claims 1 to 67, or the composition according to any one of claims 68 to 80, further comprising a promoter.
86. A nucleic acid molecule according to any one of claims 1 to 67, or a composition according to any one of claims 68 to 85, wherein the promoter is a GFAP promoter, a nestin promoter, an S100B promoter, a Nefh promoter, a dystrophin promoter, an H1 promoter, a 7SK promoter, an apolipoprotein E-human-alpha-1-antitrypsin promoter, a CK8 promoter, an mU1a promoter, an EF-1α promoter, a TBG promoter, a PKG promoter, a CAG promoter, an SV40 early promoter, a mouse mammary gland tumor virus LTR promoter, or an Ad A nucleic acid molecule or composition that is MLP, HSV promoter, CMV promoter such as CMV-IE, RSV promoter, U6 promoter or its variant, hSyn promoter, hexaribonucleotide-binding protein-3 (NeuN) promoter, CaMKII promoter, Tα-1 promoter, neuron-specific enolase (NSE) promoter, PDGFβ promoter, VGLUT promoter, SST promoter, NPY promoter, VIP promoter, PV promoter, GAD65 or GAD67 promoter, DRD1 and DRD2 promoters, MAP1B, C1ql2 promoter, POMC promoter, PROX1 promoter, or any suitable promoter.
87. The nucleic acid molecule according to any one of claims 1 to 67 or the composition according to any one of claims 68 to 80, wherein the molecule comprises polyA, and optionally the polyA is SV40 polyA, HGH polyA, BGH polyA, betaglobin polyA, alphaglobin polyA, ovalbumin polyA, kappa light chain polyA, synthetic polyA, or any suitable polyA.
88. A vector comprising a nucleic acid molecule according to any one of claims 1 to 67 or a composition according to any one of claims 68 to 80, wherein the vector is optionally a plasmid, a DNA vector, an RNA vector, a virion, or a viral vector.
89. The vector according to claim 88, wherein the vector is a viral vector.
90. The vector according to claim 89, wherein the viral vector is adeno-associated virus (AAV), lentivirus, adenovirus, monkin virus 40, vaccinia virus, measles virus, herpesvirus, or poxvirus.
91. The vector according to claim 90, wherein the viral vector is AAV.
92. The vector according to claim 91, wherein the AAV comprises a capsid protein derived from an AAV serotype selected from the group consisting of AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAVrh10, and AAVrh74.
93. The vector according to any one of claims 91 to 92, wherein the AAV is a pseudotype AAV.
94. A pharmaceutical composition comprising a nucleic acid molecule according to any one of claims 1 to 67, or a composition according to any one of claims 68 to 80, or a vector according to any one of claims 88 to 93, and a pharmaceutically acceptable carrier, diluent, or excipient.
95. A method for regulating the expression of a protein, RNA, or other biomolecule in a subject requiring such regulation, comprising administering to the subject a therapeutically effective amount of a nucleic acid molecule according to any one of claims 1 to 67, a composition according to any one of claims 68 to 80, a vector according to any one of claims 88 to 93, or a pharmaceutical composition according to claim 94.
96. The method according to claim 95, further comprising administering a therapeutically effective amount of a splice modifier to the subject.
97. The method according to claim 95, wherein the protein is expressed in the presence of the splice regulator.
98. The method according to any one of claims 96 to 97, wherein administering a therapeutically effective amount of the splice modifier to the subject causes inclusion of one of the two or more exons, one or more exons, the first exon, or the second exon.
99. The method according to claim 98, wherein one of the two or more exons, or one of the one or more exons, is the second exon.
100. The method according to any one of claims 96 to 99, wherein the splice regulator binds to the RNA-binding protein of the nucleic acid described in any one of claims 1 to 67 and / or to the segment of the splice regulator binding site.